Research Datasets
Every curated Profoundd archive, published as a downloadable dataset — one JSONL file per collection, regenerated automatically as the archives grow. Free for research under CC-BY 4.0 (attribute "Profoundd, profoundd.com").
Poll manifest.json
to detect updates cheaply — it lists every collection's SHA-256 checksum,
document count, and last-updated timestamp. Transcripts of video/audio are
verbatim ASR captions, not certified transcripts; documents keep their
original source URLs for verification.
How to use
Each file is newline-delimited JSON — one document per line with
title, content (full text), source_url,
doc_date, doc_kind, and tags.
Load with pandas.read_json(url, lines=True), stream with
jq, or ingest into any search/analysis tool. Collections
marked commentary/news/critique in
doc_kind are labeled third-party material indexed for
comparison — not the primary record.