For researchers
Bulk access
WARC files, CDX indexes and derivative datasets, for work that cannot be done through a replay interface.
Derivative datasets
| File | Type | Size | SHA-256 | |
|---|---|---|---|---|
| link-graph-2026-q2.parquet | graph | 1.9 TB | a640406ee42cb441… | HTTPSTorrentsig |
| text-extract-2026-q2.parquet | text | 8.4 TB | ebb6fa4fd50baa4d… | HTTPSTorrentsig |
| cdx-index-full.zst | index | 412 GB | 6b678b35de084c29… | HTTPSTorrentsig |
| domain-status-history.parquet | metadata | 88 GB | 415bd12a33b7092b… | HTTPSTorrentsig |
| collection-news-2026.warc.gz | warc | 6.1 PB | 3180897024aaf932… | HTTPSTorrentsig |
| collection-gov-2026.warc.gz | warc | 2.9 PB | 7de06c3bee2791c5… | HTTPSTorrentsig |
Derivative datasets are CC BY 4.0. Raw WARC access is granted per project — email the research team with a short description and we will arrange transfer, including shipping physical media for the very large requests.
Research access
Write to the research team with the collection, the date range and what you intend to do. Academic use is normally approved within two weeks; we ask for a citation and a copy of the resulting publication.
Yes. Above about 100 TB it is faster and cheaper than the network for most institutions. We use encrypted disks and a signed manifest.
On the derivative text dataset, yes, under CC BY. On raw WARC it depends on the collection and the underlying rights, and the request form asks you to be specific.
Yes — CDX query, replay and capture-status endpoints, documented and rate limited generously for registered research accounts.