Inside the archive

What the 68 boxes hold, how the readings were extracted and checked, and what on these pages is new against what has been public for twenty years.

Every reading, against its own limit

1,766 measurements, each shown as a multiple of the limit it was judged by, grouped by the kind of limit. Filter by where the sample was taken.

Highest reading in each group. Full data in readings_unified.csv, every row carrying its document and page.

What is in the boxes

The release is not what its page count suggests. Most of it is neither policy nor correspondence but laboratory paperwork and per-building case files. Two boxes hold staff attendance sheets.

laboratory, 29 boxes building files, 20 boxes subject files, 13 boxes email, 4 boxes administrative, 2 boxes
hollow dot — documents fetched and searched, 20 boxes solid dot — a figure on this page comes from it, 7 boxes

Boxes 1 to 68 in order. Hover, tap or focus a square for its number, what it holds, and how many documents came out of it.

Each box was classified based on a sample of its folder labels, about a third of the index. This approach was validated against Box 10, which was enumerated in full: content_census.csv. Classification only needs the dominant type. A box with mixed contents could still be labelled wrongly. The labels can also mislead: a 177-page folder marked “NO INFO” holds the federal inspector general’s report. Labels are quoted, never paraphrased. Method in the data guide. Six of the 118 documents have no box recorded in the portal, so they are not counted here.
What each kind of record can answer
RecordAnswersCannot answer
Laboratorythat an analysis ran, when, with what quality controlthe result, or where the sample came from — samples carry lab numbers only
Building fileswhat happened at one addressanything citywide; mostly worker exposure, not street air
Subject filespolicy, reports, advisoriessystematic coverage of anything
Emailwho decided what, and whenverified numbers

No Health Department collection is in this release. Every health advisory quoted anywhere here survives only because the environmental department filed a copy.

What is new here, and what is not

The White House Council on Environmental Quality shaping EPA's press releases, the whistleblower's memoranda, the congressional criticism — all public, most since 2003, and all reported again in the coverage of this release. Referred to here as context because not everyone knows it, and labelled wherever it appears.

Not seen published elsewhere: the four-school series as data rather than a scanned table; turnaround times for the laboratory section that handled solvents, derived from folder labels; a map of what the 68 boxes contain; and an address index to the building files. If any of them exists already, tell me and I will link to it.

The box map is above. The four-school series and the turnaround figures are on the front page; the address index is the building lookup.

How the numbers were read

These are faxes of typed tables, scanned. Ordinary text extraction reads across columns, merges unrelated samples, and drops the "less than" sign that separates a measurement from a non-detection. Everything here was extracted by position instead — each word with its coordinates, grouped geometrically. On the school table that recovered 1,079 of 1,080 cells; the one it does not return is listed rather than guessed at.

The missing cell is P.S. 234 on 24 June 2002, which reads “Overloaded”; parse_schools.py lists it on every run. Data: schools_readings.csv.

How they were checked

None of this would catch an error in the source document itself.

Errors found so far

Three so far. A count of days without readings went in at 57 when it was 42, and measured my gaps rather than the city's. A table was read confidently in the wrong orientation. A parser ate a decimal point and turned 0.018 into 18 — in the single most-quoted reading on the site. All are listed, dated, on the corrections page.