# Data guide

Every figure in this guide describes a dataset in `data/` or a structure in the
portal index. Where a number comes from a document rather than from the data, the
document number is given inline.


How to read the NYC 9/11 portal, and how to read the datasets built from
it. Written for someone who has not been in the archive. Current to
10 Sept 2026.

If you read only one section, read **§6, Ways to get this wrong**.

---

## 1. What the portal is

The city published records on 8 Sept 2026 under a settlement with 9/11
Health Watch. It is a **rolling release**: this is the first tranche, more
arrives monthly for twelve months, and the parties agree search terms as
they go. Stated omissions are personally identifiable information and
security evaluations of Lower Manhattan buildings.

Consequence for anyone writing from it: **absence is not evidence of
withholding.** Anything you cannot find has at least four possible causes
— not yet released, released but not found by your search, deliberately
omitted under the stated policy, or never created. Date-stamp every claim
about what is missing.

## 2. How it is organised

~48,900 index entries. Each document appears roughly twice, so halve for
document counts. Four sources:

| Source | Entries | What it is |
|---|---|---|
| DEP Hard Copies (68 Boxes) | 42,800 | The substance. Everything below. |
| WTC 7 | 6,050 | DCAS/FDNY/DDC leases and contracts for 7 WTC |
| DORIS Giuliani | 42 | 21 folders of clippings and code reprints |
| Other (DDC, DOB) | ~40 | Incl. the DOB building code task force report |

**No Health Department, Education, Emergency Management, Mayor's Office,
Law or Sanitation collection exists in this release.** Every DOH document
in hand survives only because DEP filed a copy. That is a structural bias
in the record and belongs in any method note.

Documents are PDFs named by the Bates ID of their first page
(`NYC-WTC_000148192`). Pages run consecutively; the next document starts
at last page + 1. Mid-document IDs 404. Direct link:

```
https://sept11documents.cityofnewyork.us/apps/content/September11_MD/NYC-WTC_<9 digits>.pdf
```

## 3. The five record types

The 68 DEP boxes divide cleanly, and the type determines what a document
can and cannot tell you.

| Type | Boxes | ~Docs | Answers | Does not answer |
|---|---|---|---|---|
| Laboratory batches | 29 | 8,000 | that an analysis ran, when, with what QC | what the result was, or where the sample came from |
| Per-building files | 20 | 7,300 | what happened at an address | anything city-wide |
| Subject files | 13 | 3,700 | policy, reports, correspondence | systematic coverage |
| Email | 4 | 1,700 | who decided what, when | verified numbers |
| Administrative | 2 | 716 | nothing (timesheets, routing slips) | — |

**Laboratory batches** (boxes 04, 12–18, 20–25, 27–30, 57–60, 62–68) are
named by instrument or EPA method plus dates: `GCMS X 11/20/01 A
12/20/01`. A sample-log legend inside 148192 p.421 gives the codes as
**X = Collected, A = Analyzed, C = Cleared, E = Exceeded, N = Not
Evaluated**, so that folder means collected 20 Nov, analysed 20 Dec — a
30-day gap between taking a sample and reading it. The legend comes from
the pilot cleaning log rather than from the lab boxes themselves, so treat
the decode as strong but not certain. GCMS = organics;
ICP/ICPMS = metals; 508 = pesticides and PCBs; 524 = volatiles; 525 =
semivolatiles. Inside are handwritten extraction bench sheets, instrument
quantitation reports, and daily performance checks. Samples are
identified by lab number only.

**Per-building files** (boxes 01, 02, 33–40, 45–52, 54, 55) are one folder
per address with block, lot and BIN. Contents are largely contractor
close-out packages. See the warning in §6 about what those actually
measure.

**Subject files** include Box 10 (the air-monitoring core, and the
smallest box in the release), Box 05 (EPA HVAC indoor cleanup, 1,580
entries), Box 07 (hotlines), Box 08 (schools, Nadler), Box 31 and 35
(named buildings).

## 4. Fetching and searching

There is no robots.txt. Etiquette is self-imposed: truthful user-agent,
serial requests, ~2.5 s pause, backoff on 429/503, stop on 403. See
`crawl.py`. Large files need ranged resume (`curl -C -`); the server
supports it and a single request will drop on the 360 MB file.

The Mindbreeze API takes unauthenticated POSTs at `/api/v2/search`. Three
things will waste your time if you do not know them:

1. **Properties return empty unless you request a format.** Without
   `"formats": ["VALUE"]` every property's `data` array is `{}`. Without it the API looks broken.
2. **It caps at 100 results and does not paginate.** `paging_state`,
   `offset` and larger `count` values all return the same first 100. The
   only way to get more is to make the query itself narrower.
3. **`count: 0` breaks `estimated_count`.** Use `count: 1`.

Field-scoped queries work and are the main lever:
`box_name:"DEP Box 10"`, `folder_name:"GCMS X 11/20/01 A 12/20/01"`,
`agency:"Fire Department"`, `source:"WTC 7"`.

To enumerate something larger than 100, partition the query. For the lab
boxes I iterated one query per calendar day — `folder_name:"11/20/01"` —
which recovered 630 distinct folders across 487 days. See
`harvest_dates.py`.

## 5. The datasets

Every file below is in `data/`. Each row carries the document and, where
the source is paginated, the page.

### readings_unified.csv — start here
Assembled from the other files in this folder; every row keeps the document and page it came from.
1,766 measurements in one shape: `date`, `place`, `substance`, `value`,
`unit`, `method`, `bench`, `bench_name`, `bclass`, `ratio`, `censored`,
`kind`, `doc`, `page`. Everything else in this list is a subset of it.

`bench` is the limit the reading was judged against and `bclass` is what
kind of limit that is — workplace, public-health guideline, clearance
threshold, ambient standard, background comparison. **Never compare ratios
across classes.** A workplace limit sits close to where harm begins; a
public-health guideline carries safety margins a hundred to a thousand
times wider.

`censored` marks a non-detection. Those rows carry the detection floor in
`value`, not a measurement. Treat them as "below this", never as a number.

### schools_readings.csv — the best data here
1,079 readings, four downtown schools, 4 Oct 2001 – 30 Jun 2002, from
NYC-WTC_000149346 pp. 8–19.

Complete, not a sample: 154 days × 4 schools = 616 readings to 6 Mar 2002,
and DEP's own slide (142716) states 616 school samples as of that date.
Two cells did not parse and are listed by `parse_schools.py` on every run.

Verified two ways: the header and the single exceedance row were read from
the scan by eye, and the series was compared cell by cell against two
independently filed copies (161895, 161897). 324 of 335 agreed in one and
107 of 111 in the other, about 96.6 per cent. Every disagreement is the
same error in the same direction — a faint "less than" sign read as a
measurement — and all sit at the detection floor.

### control_schools_full.csv
183 readings from five outer-borough schools, 14 Sept – 10 Oct 2001, from
the daily tables in 148192. All below 0.01 f/cc, maximum 0.008. **The
school-by-school attribution is unverified**: columns were assigned by
position on the page, so treat the series as an aggregate.

### dep_exceedances.csv
The nine. Every DEP sample over 70 structures per square millimetre to
6 Mar 2002, by date, site number and corner, from 148192 p. 351. Complete
as the city's own accounting of its own network.

### dep_reported_counts.csv
From NYC-WTC_000148192 page 360, hand-encoded from the scan and checked against the image.
21 rows: the city's own tally for 15–23 Sept 2001, by area — samples taken,
samples over the threshold, maximum, and the corner where it was found.
Hand-encoded from the scan and image-verified. It counts results
**received**, not samples taken, so it lags the field by whatever the
laboratory turnaround was.

### indoor_readings.csv
15 indoor samples, 14–22 Sept 2001, with floor, from the daily tables.
Five over the threshold. This is what the city had in hand about indoor air
in the first fortnight, not a description of indoor air.

### metals_air.csv
14 rows from the metals table in 148192 pp. 332–333: lead at ten locations
plus copper, iron and calcium at the highest. Three lead readings exceed
the 1.5 µg/m³ national standard, which is written as a 90-day average.
Column alignment was verified against the scan by eye after a first pass
put the locations one column out.

### voc_exceedances.csv
From the US EPA draft exposure evaluation, NYC-WTC_000148952, which is marked draft.
18 solvent readings over a health benchmark, from the federal draft
evaluation. **Grab samples and 24-hour samples are both here and are not
comparable**: a few minutes of air at the pile against a continuous day's
sampling a few streets away. In one such pair the two readings are about ten
thousand times apart. Neither is wrong; they measure different things.

### lab_batches.csv
624 laboratory folders read from folder names alone — no document was
opened. Columns include `assay`, `collected`, `analysed`, `lag_days`.

Process metadata, not results. Only 370 record both dates, and the missing
ones are not random: every ICP and ICP/MS folder carries a single date, so
any turnaround figure describes the section running gas chromatography, not
the whole laboratory. The `X`/`A` codes are read as Collected and Analyzed
on the strength of a legend at 148192 p. 421; that decode is strong but not
certain.

### building_index.csv and stations.csv
From folder labels harvested across the building-file boxes; no document was opened to build it.
824 Lower Manhattan addresses with block, lot, BIN and their document
numbers, covering 4,564 folder entries of roughly 7,300 building documents. Nineteen
documents are filed at two addresses each, so those entries are 4,545 distinct
documents. The rest are
in folders whose labels could not be matched to an address.

45 distinct ambient monitoring locations with sample counts, merged from 68
name variants in the source. Thirty-four carry coordinates from Geoclient, the New York City
Department of City Planning's address geocoder (v2 API), covering 1,263 of 1,481 samples; the remaining eleven, mostly
Battery Park City parks and waterfront, could not be placed.

### statements_and_decisions.csv
Compiled from documents in the release and, where the status column says so, from contemporaneous press reports.
44 dated statements and decisions with source, URL and status. `status`
distinguishes in hand, in hand (quoted), in hand (web), and not in hand.
Nine still rest on press reports rather than a document.

### What is not here

The perimeter sampling tables in the daily monitoring reports. An
extraction of 594 rows exists but is not published: only 91 carry a
numeric value, and the highest of those is on the electron-microscope
scale in a column labelled for fibre counting, so some rows are
mislabelled by method. The claim audit sets this out.

## 6. Ways to get this wrong

**Personal samples are not ambient samples.** The building close-out
packages are largely *personal air monitoring* — samples taken in the
breathing zone of abatement workers during the work. They run high by
design. Mixing them into an ambient dataset would badly overstate what
residents and passers-by breathed. The 114 Liberty close-out
(NYC-WTC_000101404) is 267 pages and roughly 250 of them are worker
exposure sheets. In `readings_unified.csv` these carry
`kind = Worker, personal`; 141 of the 1,766 rows are theirs, and 42 of
those have no date on the sheet.

**PCM and TEM are different measurements.** PCM counts all fibres by
optical microscope and reports fibres/cc. TEM identifies asbestos
specifically by electron microscope and reports structures per square
millimetre. A PCM number is not an asbestos number.

**Method changes look like air changes.** The school series switches from
f/cc to s/sqmm on 19 Oct 2001. Nothing in the air changed that day. Any
chart running one line through 18–19 Oct draws a cliff that did not
happen. Break the series, use separate panels, or plot each reading as a ratio to
its own benchmark.

**Most of the data is censored.** 1,168 of the 1,766 rows are below what
the laboratory could quantify. In the school series it is nine in ten. A
censored row carries the detection limit in `value`, not a measurement;
treat it as "below this", never as zero and never as the limit itself. A
chart of raw values will mostly show laboratory sensitivity.

**Detection floors move.** The limit is set sample by sample, so the floor is
not one value: `<13.33 s/sqmm` runs to 9 January, `<14.3` to 12 June and
`<13.3` after, with six one-day values elsewhere. On six days the four schools
do not share a limit, which is why the floor tracks the analysis and not the
air.

**The two exceedance counts are not comparable.** DEP: 9 of 4,564 (as of
6 Mar 2002; 3,540 ground-zero network + 616 schools + 408 outer boroughs).
EPA NCEA draft: 22 of 8,870 TEM samples in lower Manhattan. Different
agencies, databases, date ranges and networks. Neither refutes the other.

**Two different 4,564s.** The city's sample total above and the count of
building-file documents are the same number by coincidence. Nothing connects
them.

**Nine exceedances are not nine places.** The nine came from five corners.
Murray and Church accounts for four of them, on four days inside a week.
Counting them as nine sites understates how concentrated they were.

**Why a building has a file.** In `building_index.csv`, a file exists because
someone complained, a landlord called, or a contractor filed paperwork. It does not mean the building was contaminated, and no file does
not mean it was clean.

**Coordinates are approximate.** Building positions come from tax lots and
building identifiers, station positions from the city's geocoder. 721 of
824 addresses and 34 of 45 stations are placed; the rest are absent from
the maps but present in the data. A station sits at its corner, not at the
instrument.

**Station names were merged.** 68 name variants in the source collapse to
45 places. The merge is on resolved position where coordinates exist and on
normalised street pairs otherwise, so sample counts in `stations.csv` are
sums across spellings.

**148952 is a draft** and its own header forbids quoting. Paraphrase.

**Coverage cannot be expressed as one fraction** for the ground-zero
network: the numerator is unclean, the date ranges do not align, and it is
unclear whether DEP's category matches this one. The school subset is the
exception — 616 of 616, same document family, same cutoff.

**Lab batches cannot be tied to places** without the chain-of-custody log
sheets in Box 42 ("WTC- Air Data", e.g. 161813), which are handwritten
carbon forms that OCR cannot read.

## 7. Selection bias, and what to do about it

Most of what was found early came from phrase searches — "barrage of
lawsuits", "not yet suitable for reoccupancy". Those queries only return documents where someone worried or hedged, and
routine documents go unfound because nothing distinctive is written in
them. A corpus assembled that way can show that something was said on a
date. It cannot show what was typical.

The working distinction:

- **Complete series** — a defined frame, everything in it read, boring
  parts included. Can support conclusions. `schools_readings.csv` and
  `lab_batches.csv` qualify.
- **Found documents** — surfaced by search. Can establish that a specific
  thing was said on a specific date. Cannot establish what officials
  generally knew.

Practical rules: define the frame before looking; log searches that
returned nothing; write down in advance what a finding *against* your
question would look like. If nothing would count as disconfirming, it is
a position, not a question.

## 8. Verification protocol

Applied to displayed values, not everything.

- **Tier 1, cross-copy.** The same reading appears in multiple fax copies;
  agreement across independent OCR reads is evidence. Falls out of the
  dedupe for free — *if* the dedupe groups by sample rather than by value.
- **Tier 2, arithmetic.** Re-derive concentration from fibres, fields,
  filter area and volume. Requires those columns to have been captured.
- **Tier 3, image.** Crop the cell and read it. Reserved for
  disagreements, exceedances, and anything the page displays prominently.

Publish the agreement rate. Show held-back values as a count, not silently.

## 9. Scripts

- `crawl.py` — polite fetcher, manifest, daily cap
- `mbsearch.py` — phrase search
- `enumerate_box.py` — list a box or folder
- `survey_boxes.py` — structural survey of all 68 boxes
- `harvest_dates.py` — date-partitioned folder harvest
- `parse_schools.py` — coordinate-based table extraction

**On parsing tables:** flat `pdftotext` output lies about landscape scans.
It interleaves columns and silently drops the `<` from "less than" cells.
Extract words with coordinates (PyMuPDF `page.get_text("words")`), bucket
by y for rows, slice by x for columns. That recovered 1,079 of 1,080 cells
from the school table where flat text produced plausible nonsense. Use the
same approach on any remaining table of this shape.
