YAHEW!
A FIELD GUIDE TO THE OLD WEB

Sources: The GeoCities deletion V · Extinction

← Back to the The GeoCities deletion plate

Source. The address system is real and verbatim, scraped by scripts/make-geocities-lots.mjs from two captured OOcities indexes: 38 neighborhood names, 47 street names and 288 numbered lots. The shipped typo pairs are kept, dungeom beside dungeon and labrinth beside labyrinth, because somebody typed that index by hand. The 2009 Yahoo! palette and the closure notice, quoted word for word including its three cross-sell links, are from the Wayback capture of geocities.yahoo.com on 8 November 2009.

The pages are real pages, and they are held here. Each address on the map was resolved against the Internet Archive's CDX index to a capture that actually returned 200, and that capture was then mirrored onto this origin with its images: the markup is the bytes the archive holds, the title is the page's own, and the images are the originals. An earlier version generated a room per address out of stock furniture, which was rightly called slop: 335 rooms with identical text is not a recreation of anything. There are no generated rooms now. An address the archive has nothing for is simply not on the map, because the map is what survives.

Why it is mirrored rather than framed from the archive. The first working version pointed an iframe at web.archive.org. That was slow, rate limited, dependent on somebody else's uptime, and it told a third party who was reading. It is also the wrong argument for this plate in particular: a page about who was holding your work should not need to ask another server for permission to show it. The archive is the source, not the dependency. Scripts are stripped from the mirrored pages; everything else is as captured.

A number this plate does not print, and one it nearly did. "38 million pages" traces to Wikipedia, which the Internet Archive quoted into its own collection description, which people now cite as though the Archive had counted. Archive Team, who did the crawling, cite 23 million from Yahoo's own Site Explorer, guess around 10 TB, and say they could be wildly off in any direction. Yahoo never published a figure and refused when asked. Separately, the first run of the capture resolver treated rate-limit errors as historical absence and reported that only five of 335 addresses survived; every one of those requests was returning HTTP 429. That number was about the fetcher, not about history, and it is recorded here because a plate about uncounted things has no business inventing a count.

Deliberately absent. Any figure for how many pages existed or what share was saved. The archive size is given as a range, about 640 GB compressed and roughly 900 GB unpacked, because Jason Scott's own posts drift between 640, 641, 643 and 652 GB as the set was patched. The 26 October 2009 date is Yahoo's announced one; Archive Team recorded the service going dark around 12:30pm Pacific on the 27th, and that is a single witness, so it is phrased as their record. The sixty second clock and the five save limit are a mechanic, not a measurement.

← Back to the The GeoCities deletion plate · The directory