How the data gets here
Every source is downloaded from the agency that publishes it — the bulk file, the open-data API or the enforcement portal, in that order of preference. There is no intermediary vendor, no purchased database and no scraped aggregator anywhere in the chain. Where an agency offers a full extract (OSHA and DOL through the Labor Department’s enforcement portal, EPA through ECHO, MSHA through its open-government archive, the CFPB through its complaint export), we take the whole file rather than a query result, so that our copy can be reproduced by anyone with the same download.
Each file is then normalized into one SQLite schema: dates parsed into a common format, dollar columns coerced to numbers, state codes standardized, and a per-source table kept intact so that no record ever loses its provenance. Aggregation happens on top of those tables and never in a spreadsheet. The inventory you see above is generated by the same export step that writes the site’s data files, which is why the totals on this page, in the footer and on the homepage are arithmetically the same number rather than three hand-maintained approximations.
Records are never deleted. When an agency reissues a file, the affected table is rebuilt from the new extract and the record count moves — up or down — with it. The release history at the bottom of this page shows every such movement since the first published inventory in March 2026, including the ones that were corrections of our own mistakes.
For eleven of these datasets we go one step further and publish the aggregated tables themselves as citable research with a permanent DOI, so that a figure quoted from Settlement Insight can be traced to a frozen, archived version of the data it came from. Those are listed below.