Every figure in this guide describes a dataset in data/ or a structure in the portal index. Where a number comes from a document rather than from the data, the document number is given inline.
How to read the NYC 9/11 portal, and how to read the datasets built from it. Written for someone who has not been in the archive. Current to 10 Sept 2026.
If you read only one section, read §6, Ways to get this wrong.
—-
The city published records on 8 Sept 2026 under a settlement with 9/11 Health Watch. It is a rolling release: this is the first tranche, more arrives monthly for twelve months, and the parties agree search terms as they go. Stated omissions are personally identifiable information and security evaluations of Lower Manhattan buildings.
Consequence for anyone writing from it: absence is not evidence of withholding. Anything you cannot find has at least four possible causes — not yet released, released but not found by your search, deliberately omitted under the stated policy, or never created. Date-stamp every claim about what is missing.
~48,900 index entries. Each document appears roughly twice, so halve for document counts. Four sources:
| Source | Entries | What it is |
|---|---|---|
| DEP Hard Copies (68 Boxes) | 42,800 | The substance. Everything below. |
| WTC 7 | 6,050 | DCAS/FDNY/DDC leases and contracts for 7 WTC |
| DORIS Giuliani | 42 | 21 folders of clippings and code reprints |
| Other (DDC, DOB) | ~40 | Incl. the DOB building code task force report |
No Health Department, Education, Emergency Management, Mayor's Office, Law or Sanitation collection exists in this release. Every DOH document in hand survives only because DEP filed a copy. That is a structural bias in the record and belongs in any method note.
Documents are PDFs named by the Bates ID of their first page (NYC-WTC_000148192). Pages run consecutively; the next document starts at last page + 1. Mid-document IDs 404. Direct link:
https://sept11documents.cityofnewyork.us/apps/content/September11_MD/NYC-WTC_<9 digits>.pdf
The 68 DEP boxes divide cleanly, and the type determines what a document can and cannot tell you.
| Type | Boxes | ~Docs | Answers | Does not answer |
|---|---|---|---|---|
| Laboratory batches | 29 | 8,000 | that an analysis ran, when, with what QC | what the result was, or where the sample came from |
| Per-building files | 20 | 7,300 | what happened at an address | anything city-wide |
| Subject files | 13 | 3,700 | policy, reports, correspondence | systematic coverage |
| 4 | 1,700 | who decided what, when | verified numbers | |
| Administrative | 2 | 716 | nothing (timesheets, routing slips) | — |
Laboratory batches (boxes 04, 12–18, 20–25, 27–30, 57–60, 62–68) are named by instrument or EPA method plus dates: GCMS X 11/20/01 A 12/20/01. A sample-log legend inside 148192 p.421 gives the codes as X = Collected, A = Analyzed, C = Cleared, E = Exceeded, N = Not Evaluated, so that folder means collected 20 Nov, analysed 20 Dec — a 30-day gap between taking a sample and reading it. The legend comes from the pilot cleaning log rather than from the lab boxes themselves, so treat the decode as strong but not certain. GCMS = organics; ICP/ICPMS = metals; 508 = pesticides and PCBs; 524 = volatiles; 525 = semivolatiles. Inside are handwritten extraction bench sheets, instrument quantitation reports, and daily performance checks. Samples are identified by lab number only.
Per-building files (boxes 01, 02, 33–40, 45–52, 54, 55) are one folder per address with block, lot and BIN. Contents are largely contractor close-out packages. See the warning in §6 about what those actually measure.
Subject files include Box 10 (the air-monitoring core, and the smallest box in the release), Box 05 (EPA HVAC indoor cleanup, 1,580 entries), Box 07 (hotlines), Box 08 (schools, Nadler), Box 31 and 35 (named buildings).
There is no robots.txt. Etiquette is self-imposed: truthful user-agent, serial requests, ~2.5 s pause, backoff on 429/503, stop on 403. See crawl.py. Large files need ranged resume (curl -C -); the server supports it and a single request will drop on the 360 MB file.
The Mindbreeze API takes unauthenticated POSTs at /api/v2/search. Three things will waste your time if you do not know them:
1. Properties return empty unless you request a format. Without "formats": ["VALUE"] every property's data array is {}. Without it the API looks broken. 2. It caps at 100 results and does not paginate. paging_state, offset and larger count values all return the same first 100. The only way to get more is to make the query itself narrower. 3. count: 0 breaks estimated_count. Use count: 1.
Field-scoped queries work and are the main lever: box_name:"DEP Box 10", folder_name:"GCMS X 11/20/01 A 12/20/01", agency:"Fire Department", source:"WTC 7".
To enumerate something larger than 100, partition the query. For the lab boxes I iterated one query per calendar day — folder_name:"11/20/01" — which recovered 630 distinct folders across 487 days. See harvest_dates.py.
Every file below is in data/. Each row carries the document and, where the source is paginated, the page.
Assembled from the other files in this folder; every row keeps the document and page it came from. 1,766 measurements in one shape: date, place, substance, value, unit, method, bench, bench_name, bclass, ratio, censored, kind, doc, page. Everything else in this list is a subset of it.
bench is the limit the reading was judged against and bclass is what kind of limit that is — workplace, public-health guideline, clearance threshold, ambient standard, background comparison. Never compare ratios across classes. A workplace limit sits close to where harm begins; a public-health guideline carries safety margins a hundred to a thousand times wider.
censored marks a non-detection. Those rows carry the detection floor in value, not a measurement. Treat them as "below this", never as a number.
1,079 readings, four downtown schools, 4 Oct 2001 – 30 Jun 2002, from NYC-WTC_000149346 pp. 8–19.
Complete, not a sample: 154 days × 4 schools = 616 readings to 6 Mar 2002, and DEP's own slide (142716) states 616 school samples as of that date. Two cells did not parse and are listed by parse_schools.py on every run.
Verified two ways: the header and the single exceedance row were read from the scan by eye, and the series was compared cell by cell against two independently filed copies (161895, 161897). 324 of 335 agreed in one and 107 of 111 in the other, about 96.6 per cent. Every disagreement is the same error in the same direction — a faint "less than" sign read as a measurement — and all sit at the detection floor.
183 readings from five outer-borough schools, 14 Sept – 10 Oct 2001, from the daily tables in 148192. All below 0.01 f/cc, maximum 0.008. The school-by-school attribution is unverified: columns were assigned by position on the page, so treat the series as an aggregate.
The nine. Every DEP sample over 70 structures per square millimetre to 6 Mar 2002, by date, site number and corner, from 148192 p. 351. Complete as the city's own accounting of its own network.
From NYC-WTC_000148192 page 360, hand-encoded from the scan and checked against the image. 21 rows: the city's own tally for 15–23 Sept 2001, by area — samples taken, samples over the threshold, maximum, and the corner where it was found. Hand-encoded from the scan and image-verified. It counts results received, not samples taken, so it lags the field by whatever the laboratory turnaround was.
15 indoor samples, 14–22 Sept 2001, with floor, from the daily tables. Five over the threshold. This is what the city had in hand about indoor air in the first fortnight, not a description of indoor air.
14 rows from the metals table in 148192 pp. 332–333: lead at ten locations plus copper, iron and calcium at the highest. Three lead readings exceed the 1.5 µg/m³ national standard, which is written as a 90-day average. Column alignment was verified against the scan by eye after a first pass put the locations one column out.
From the US EPA draft exposure evaluation, NYC-WTC_000148952, which is marked draft. 18 solvent readings over a health benchmark, from the federal draft evaluation. Grab samples and 24-hour samples are both here and are not comparable: a few minutes of air at the pile against a continuous day's sampling a few streets away. In one such pair the two readings are about ten thousand times apart. Neither is wrong; they measure different things.
624 laboratory folders read from folder names alone — no document was opened. Columns include assay, collected, analysed, lag_days.
Process metadata, not results. Only 370 record both dates, and the missing ones are not random: every ICP and ICP/MS folder carries a single date, so any turnaround figure describes the section running gas chromatography, not the whole laboratory. The X/A codes are read as Collected and Analyzed on the strength of a legend at 148192 p. 421; that decode is strong but not certain.
From folder labels harvested across the building-file boxes; no document was opened to build it. 824 Lower Manhattan addresses with block, lot, BIN and their document numbers, covering 4,564 folder entries of roughly 7,300 building documents. Nineteen documents are filed at two addresses each, so those entries are 4,545 distinct documents. The rest are in folders whose labels could not be matched to an address.
45 distinct ambient monitoring locations with sample counts, merged from 68 name variants in the source. Thirty-four carry coordinates from Geoclient, the New York City Department of City Planning's address geocoder (v2 API), covering 1,263 of 1,481 samples; the remaining eleven, mostly Battery Park City parks and waterfront, could not be placed.
Compiled from documents in the release and, where the status column says so, from contemporaneous press reports. 44 dated statements and decisions with source, URL and status. status distinguishes in hand, in hand (quoted), in hand (web), and not in hand. Nine still rest on press reports rather than a document.
The perimeter sampling tables in the daily monitoring reports. An extraction of 594 rows exists but is not published: only 91 carry a numeric value, and the highest of those is on the electron-microscope scale in a column labelled for fibre counting, so some rows are mislabelled by method. The claim audit sets this out.
Personal samples are not ambient samples. The building close-out packages are largely personal air monitoring — samples taken in the breathing zone of abatement workers during the work. They run high by design. Mixing them into an ambient dataset would badly overstate what residents and passers-by breathed. The 114 Liberty close-out (NYC-WTC_000101404) is 267 pages and roughly 250 of them are worker exposure sheets. In readings_unified.csv these carry kind = Worker, personal; 141 of the 1,766 rows are theirs, and 42 of those have no date on the sheet.
PCM and TEM are different measurements. PCM counts all fibres by optical microscope and reports fibres/cc. TEM identifies asbestos specifically by electron microscope and reports structures per square millimetre. A PCM number is not an asbestos number.
Method changes look like air changes. The school series switches from f/cc to s/sqmm on 19 Oct 2001. Nothing in the air changed that day. Any chart running one line through 18–19 Oct draws a cliff that did not happen. Break the series, use separate panels, or plot each reading as a ratio to its own benchmark.
Most of the data is censored. 1,168 of the 1,766 rows are below what the laboratory could quantify. In the school series it is nine in ten. A censored row carries the detection limit in value, not a measurement; treat it as "below this", never as zero and never as the limit itself. A chart of raw values will mostly show laboratory sensitivity.
Detection floors move. The limit is set sample by sample, so the floor is not one value: <13.33 s/sqmm runs to 9 January, <14.3 to 12 June and <13.3 after, with six one-day values elsewhere. On six days the four schools do not share a limit, which is why the floor tracks the analysis and not the air.
The two exceedance counts are not comparable. DEP: 9 of 4,564 (as of 6 Mar 2002; 3,540 ground-zero network + 616 schools + 408 outer boroughs). EPA NCEA draft: 22 of 8,870 TEM samples in lower Manhattan. Different agencies, databases, date ranges and networks. Neither refutes the other.
Two different 4,564s. The city's sample total above and the count of building-file documents are the same number by coincidence. Nothing connects them.
Nine exceedances are not nine places. The nine came from five corners. Murray and Church accounts for four of them, on four days inside a week. Counting them as nine sites understates how concentrated they were.
Why a building has a file. In building_index.csv, a file exists because someone complained, a landlord called, or a contractor filed paperwork. It does not mean the building was contaminated, and no file does not mean it was clean.
Coordinates are approximate. Building positions come from tax lots and building identifiers, station positions from the city's geocoder. 721 of 824 addresses and 34 of 45 stations are placed; the rest are absent from the maps but present in the data. A station sits at its corner, not at the instrument.
Station names were merged. 68 name variants in the source collapse to 45 places. The merge is on resolved position where coordinates exist and on normalised street pairs otherwise, so sample counts in stations.csv are sums across spellings.
148952 is a draft and its own header forbids quoting. Paraphrase.
Coverage cannot be expressed as one fraction for the ground-zero network: the numerator is unclean, the date ranges do not align, and it is unclear whether DEP's category matches this one. The school subset is the exception — 616 of 616, same document family, same cutoff.
Lab batches cannot be tied to places without the chain-of-custody log sheets in Box 42 ("WTC- Air Data", e.g. 161813), which are handwritten carbon forms that OCR cannot read.
Most of what was found early came from phrase searches — "barrage of lawsuits", "not yet suitable for reoccupancy". Those queries only return documents where someone worried or hedged, and routine documents go unfound because nothing distinctive is written in them. A corpus assembled that way can show that something was said on a date. It cannot show what was typical.
The working distinction:
parts included. Can support conclusions. schools_readings.csv and lab_batches.csv qualify.
thing was said on a specific date. Cannot establish what officials generally knew.
Practical rules: define the frame before looking; log searches that returned nothing; write down in advance what a finding against your question would look like. If nothing would count as disconfirming, it is a position, not a question.
Applied to displayed values, not everything.
agreement across independent OCR reads is evidence. Falls out of the dedupe for free — if the dedupe groups by sample rather than by value.
filter area and volume. Requires those columns to have been captured.
disagreements, exceedances, and anything the page displays prominently.
Publish the agreement rate. Show held-back values as a count, not silently.
crawl.py — polite fetcher, manifest, daily capmbsearch.py — phrase searchenumerate_box.py — list a box or foldersurvey_boxes.py — structural survey of all 68 boxesharvest_dates.py — date-partitioned folder harvestparse_schools.py — coordinate-based table extractionOn parsing tables: flat pdftotext output lies about landscape scans. It interleaves columns and silently drops the < from "less than" cells. Extract words with coordinates (PyMuPDF page.get_text("words")), bucket by y for rows, slice by x for columns. That recovered 1,079 of 1,080 cells from the school table where flat text produced plausible nonsense. Use the same approach on any remaining table of this shape.
Back to the start, or read the claim audit. Plain text version