A URL is a promise that somebody keeps paying for. The default state of a link is broken, and the only question is how long it takes.
A hundred pages in each grid. Press the button and the ones that are gone go dark. Both figures are verbatim from Pew Research Center, May 2024.
Pages that existed in 2013
0 of 100 gone by October 2023, ten years
Pages captured in 2021
0 of 100 gone two years later
These are two different cohorts, not two points on one line. Ten years took 38. Two years took about a fifth. That is not the same measurement run twice, and this plate will not draw a curve between them, because the shape in the middle is not in the data.
The places that cite the most carefully rot the hardest, because a citation is a link to something you do not control.
Read the bottom two carefully. That is reference rot, which means the link is dead or it still resolves and no longer says what was cited. The second kind is worse, because nothing looks broken. A footnote pointing at a page that quietly changed is a citation to something that never existed.
The scholarly record measures the same way. One in five articles across a million web references suffers reference rot, and seven in ten among articles that cite web resources at all.
In March 2013 Google published a post called A second spring of cleaning. It is the post that announced Google Reader was closing, which is the moment a lot of people point at when they talk about the platform turn.
1googleblog.blogspot.com/2013/03/a-second-spring-of-cleaning.html
301 Moved Permanently
2blog.google/topics/inside-google/a-second-spring-of-cleaning/
The announcement that a service was being shut down has itself been moved to a different host. It still resolves, so nothing counts it as rot, and every link anyone wrote to it between 2013 and the move now depends on a redirect somebody is choosing to maintain. Observed 2 August 2026, and you can run it yourself.
An archived page is not a screenshot. The Archive rewrites every URL inside it so the page's own assets resolve back into the archive, and it does that with two letters in the timestamp.
/web/20091026132241cs_/http://...
stylesheets
/web/20091026132241js_/http://...
scripts
/web/20091026132241im_/http://...
images
/web/20091026132241id_/http://...
the original bytes, nothing rewritten
Everything this guide has quoted was pulled with the last one. id_ returns the original unmodified bytes with no toolbar and no rewriting, which is the only way to read a 1996 page as a 1996 browser received it.
And one thing worth knowing about that machine. Read the script tags at the top of any archived page and the load order is bundle-playback.js, then wombat.js, then ruffle.js. The Wayback Machine ships a Flash emulator so that Flash content inside archived pages still plays. The archive is not only storing the old web, it is running an interpreter for a format that was switched off.
Hrm.
Wayback Machine has not archived that URL.
The real failure screen. Everything not captured before it went is simply gone, and this is the sentence that tells you.
How to spot rot before it bites.
/2013/03/a-second-spring-of-cleaning.htmlA link into a path that encodes a date
/2013/03/ in a URL means the publisher was organizing by time, and time-organized paths are the first thing a migration rewrites.
blogspot.comtumblr.commedium.comA hostname that is a product rather than an institution
blogspot.com, tumblr.com, medium.com. Products get sunset; institutions get budgets.
301 →blog.google/...A redirect chain with no canonical at the end
If a link 301s and the destination does not link back to the old address, the old address is now unverifiable, not preserved.
[accessed 2013]A citation with no capture date
A footnote that says only "accessed" and gives no archived copy is a promise about a page nobody can check.
200 OKand the page no longer says itReference rot without link rot
The URL resolves, the page loads, and it no longer says the thing. Nothing anywhere reports this as an error.
example.com/a/b/c.html404A deep link into a site with no index
If the surrounding site cannot be browsed, a dead deep link cannot be recovered by hand.
Nothing on this page is a technology that failed. Rot is what happens by default when a name has to be paid for, forever, by someone who eventually stops. Domains lapse, companies reorganize, a directory gets renamed, and the link that pointed there was never a copy of anything. It was an address.
What we have instead is one non-profit that decided to keep a copy, and every plate in this guide is built on material that only exists because of it. That is a remarkable thing to have and a precarious way to hold the record of a medium. The counterweight is unglamorous: save the bytes, not the address, and cite the capture.
What breaks: a link is a promise that a stranger will keep paying rent on a name. Most of them stop.