Web preservation
The page you cited will change. This one will not.
Forty-one billion captures since 2011, each with a permanent identifier, a cryptographic digest and a replay that renders the page as it was — scripts, stylesheets and all.
What makes an archive citable
A permanent identifier
Every capture gets an identifier that includes the URL, the timestamp and the content digest. Cite it and the reader sees exactly what you saw.
Integrity you can check
WARC records are hashed and the hashes are published in a monthly signed digest. Anyone can verify a replay was not altered after the fact.
Replay that works
Client-side rewriting so archived JavaScript runs against archived resources, not against the live web.
Full text
Searchable across the whole archive, not just titles and URLs, with time-sliced results.
Collections
Curated crawls around events, institutions and at-risk sites, each with a documented scope and cadence.
Bulk access
WARC and CDX exports for researchers, and a derivative dataset programme for text and link graphs.
Datasets
Link rot
Half of what a paper cites will be gone in a decade
We measured it on our own corpus: of URLs cited in journal articles published in 2015, 48% no longer resolve to the cited content, and 21% do not resolve at all.
- 48% of 2015 citations no longer show the cited content
- 21% return an error or a parked domain
- Median lifespan of a cited page: 7.4 years
- Government and news domains fare worst, not best