What the 68 boxes hold, how the readings were extracted and checked, and what on these pages is new against what has been public for twenty years.
1,766 measurements, each shown as a multiple of the limit it was judged by, grouped by the kind of limit. Filter by where the sample was taken.
The release is not what its page count suggests. Most of it is neither policy nor correspondence but laboratory paperwork and per-building case files. Two boxes hold staff attendance sheets.
Boxes 1 to 68 in order. Hover, tap or focus a square for its number, what it holds, and how many documents came out of it.
| Record | Answers | Cannot answer |
|---|---|---|
| Laboratory | that an analysis ran, when, with what quality control | the result, or where the sample came from — samples carry lab numbers only |
| Building files | what happened at one address | anything citywide; mostly worker exposure, not street air |
| Subject files | policy, reports, advisories | systematic coverage of anything |
| who decided what, and when | verified numbers |
No Health Department collection is in this release. Every health advisory quoted anywhere here survives only because the environmental department filed a copy.
The White House Council on Environmental Quality shaping EPA's press releases, the whistleblower's memoranda, the congressional criticism — all public, most since 2003, and all reported again in the coverage of this release. Referred to here as context because not everyone knows it, and labelled wherever it appears.
Not seen published elsewhere: the four-school series as data rather than a scanned table; turnaround times for the laboratory section that handled solvents, derived from folder labels; a map of what the 68 boxes contain; and an address index to the building files. If any of them exists already, tell me and I will link to it.
These are faxes of typed tables, scanned. Ordinary text extraction reads across columns, merges unrelated samples, and drops the "less than" sign that separates a measurement from a non-detection. Everything here was extracted by position instead — each word with its coordinates, grouped geometrically. On the school table that recovered 1,079 of 1,080 cells; the one it does not return is listed rather than guessed at.
parse_schools.py lists it on every run. Data: schools_readings.csv.None of this would catch an error in the source document itself.
Three so far. A count of days without readings went in at 57 when it was 42, and measured my gaps rather than the city's. A table was read confidently in the wrong orientation. A parser ate a decimal point and turned 0.018 into 18 — in the single most-quoted reading on the site. All are listed, dated, on the corrections page.