Research Datasets

Every curated Profoundd archive, published as a downloadable dataset — one JSONL file per collection, regenerated automatically as the archives grow. Free for research under CC-BY 4.0 (attribute "Profoundd, profoundd.com").

Poll manifest.json to detect updates cheaply — it lists every collection's SHA-256 checksum, document count, and last-updated timestamp. Transcripts of video/audio are verbatim ASR captions, not certified transcripts; documents keep their original source URLs for verification.

Dataset Documents Size Updated Download
FBI Vault 8,895 438.7 MB 2026-09-02 JSONL
Southern Poverty Law Center (Wayback) 7,830 6.4 MB 2026-09-02 JSONL
Pfizer Documents (PHMPT/FDA) 2,373 225.2 MB 2026-09-02 JSONL
CDC ACIP — Vaccine Advisory Committee 449 9.5 MB 2026-09-02 JSONL
NIH Pandemic-Era Grants 9,873 78.8 MB 2026-09-02 JSONL
Survival, Water, Medical Field Manuals 1,435 98.5 MB 2026-09-02 JSONL
Rudolf Steiner Archive 342 65.4 MB 2026-09-02 JSONL
Hesperian Health Guides (Where There Is No Doctor) 19 3.2 MB 2026-09-02 JSONL
Robots — Open-Source 3D-Printable Designs 27 22.5 KB 2026-09-02 JSONL
Smart Cities & Surveillance — WEF, Contracts, FOIA 151 164.2 KB 2026-09-02 JSONL
Fauci Files — DNI Gabbard Release + Rand Paul Diaries (2026) 50 6.8 MB 2026-09-02 JSONL
Mail Tribune (Medford, OR — Wayback) 144,941 761.7 MB 2026-09-02 JSONL
Ashland Daily Tidings (Ashland, OR — Wayback) 32,377 190.7 MB 2026-09-02 JSONL
Charlie Kirk / Tyler Robinson Case — Court Transcripts & Filings 270 21.4 MB 2026-09-02 JSONL
Medical Talks — Integrative & Longevity Medicine 771 13.1 MB 2026-09-02 JSONL
Robert Malone — Collected Works 233 4.2 MB 2026-09-02 JSONL
Peter McCullough — Collected Works 169 725.5 KB 2026-09-02 JSONL
Pierre Kory — Collected Works 317 6.1 MB 2026-09-02 JSONL
Oregon LUBA — Land Use Board of Appeals Oral Arguments 36 1.0 MB 2026-09-02 JSONL
Iran "Lego" Propaganda — Under Analysis 12 26.0 KB 2026-09-02 JSONL

How to use

Each file is newline-delimited JSON — one document per line with title, content (full text), source_url, doc_date, doc_kind, and tags. Load with pandas.read_json(url, lines=True), stream with jq, or ingest into any search/analysis tool. Collections marked commentary/news/critique in doc_kind are labeled third-party material indexed for comparison — not the primary record.