Chronofoldest. 2011

Web preservation

The page you cited will change. This one will not.

Forty-one billion captures since 2011, each with a permanent identifier, a cryptographic digest and a replay that renders the page as it was — scripts, stylesheets and all.

 41 B captures since 2011 permanent citations
41.2B
captures
34PB
WARC stored
1.9TB
crawled per day
204M
replays per month

What makes an archive citable

A permanent identifier

Every capture gets an identifier that includes the URL, the timestamp and the content digest. Cite it and the reader sees exactly what you saw.

Integrity you can check

WARC records are hashed and the hashes are published in a monthly signed digest. Anyone can verify a replay was not altered after the fact.

Replay that works

Client-side rewriting so archived JavaScript runs against archived resources, not against the live web.

Full text

Searchable across the whole archive, not just titles and URLs, with time-sliced results.

Collections

Curated crawls around events, institutions and at-risk sites, each with a documented scope and cadence.

Bulk access

WARC and CDX exports for researchers, and a derivative dataset programme for text and link graphs.

Datasets

Link rot

Half of what a paper cites will be gone in a decade

We measured it on our own corpus: of URLs cited in journal articles published in 2015, 48% no longer resolve to the cited content, and 21% do not resolve at all.

  • 48% of 2015 citations no longer show the cited content
  • 21% return an error or a parked domain
  • Median lifespan of a cited page: 7.4 years
  • Government and news domains fare worst, not best

Common questions

For crawling, yes, with a narrow exception for on-demand captures requested by a person who can see the page. We do not retroactively hide old captures because a robots.txt changed years later — that practice made archives unreliable for research.