A plain-language look at exactly what this database holds — and what it doesn't.
This archive holds 1,201,245 documents that the U.S. Department of Justice released as part of the Epstein Files, organized into the 12 numbered batches ("datasets") the DOJ published them in.
About 98% have machine-readable text; the rest are scans or handwriting that optical character recognition couldn't read. Roughly 87% carry a detected date — explore them on the homepage timeline.
Tommy Carstensen converted the scanned DOJ PDFs into text using optical character recognition (Tesseract OCR). We then ran automated name-detection over that text to pull out people, organizations, and places.
That yielded 1,140,556 named entities, 8,804,176 links between documents and names, and 8,690,938 co-mention pairs. The names are machine-guessed and imperfect — OCR errors and ambiguous names introduce noise, so we merge clear variants of the same entity (alternate spellings, OCR fragments, and email-header bleed) into a single record — the many forms of Jeffrey Epstein, for example, now resolve to one. Browse entities.
After publishing, the DOJ has quietly removed 3,145 documents and silently modified 12,393 more. We track both so the record is preserved.
Removal waves
Of the removed files, 2,129 have an archived copy linked on their document page. Deletion and modification tracking comes from Tommy Carstensen's archive.
This is a point-in-time snapshot. We hold 1,201,245 documents — about 87% of the ~1,380,987 files the DOJ has released. The rest are mostly the most recent releases (datasets 11–12), captured after our snapshot.
EFTA ID numbers run higher than the document count because 1,350,893 of them are individual pages inside larger documents, not separate files. Beyond that, 35,915 multi-page documents are missing from our copy entirely. We don't host the original PDFs — we link to justice.gov and to archived copies — and we don't have files the DOJ never released.
Source: the DOJ's justice.gov/epstein release, OCR'd and structured by Carstensen's archive; the search index, name-detection, and DOJ-status tracking are ours. Extraction and name-detection are automated and imperfect — see the About page.
Spotted an error or have something to add? Corrections & additions →