frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Where did the old web go? We followed 657,607 links to find out

https://0.mk/blog/link-rot
26•tdx•1h ago

Comments

tdx•1h ago
I found an old database backup of 0.mk on a disk I had kept.

0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.

The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.

Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.

I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.

There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.

A few things I did not expect:

- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.

Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.

I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.

Happy to answer questions about the crawl, the old data, or the rebuild.

hyperionultra•8m ago
How did you managed to obtain that domain? Usually single digit or letter domains are “reserved”.
exitnode•31m ago
Wow, that is a great domain!
shevy-java•20m ago
Webpages dying is probably one of the biggest design flaws of the original web.

I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.

efskap•13m ago
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.

We hot-linked to all those image hosts because we couldn't imagine them disappearing.

Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.

ChadNauseam•8m ago
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
rcxdude•4m ago
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
tokai•10m ago
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.

Gemini 3.7 Flash

https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-fl...
247•thisisauserid•1h ago•164 comments

Accelerating GPT-5.6 Sol Ultrafast

https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
62•pr337h4m•40m ago•9 comments

Choose Boring Technology (2015)

https://mcfunley.com/choose-boring-technology
64•tosh•1h ago•21 comments

Mistral OCR 4.1

https://docs.mistral.ai/models/ocr-4-1
98•spelk•1h ago•26 comments

Donkey.bas is 45 Years Old – 131 line of Glory

https://donkeybas.com/
42•jkrauska•1h ago•16 comments

Spaghettifying DRAM

https://github.com/xoreaxeaxeax/skitter-creek-bath-salts
319•matt_d•4h ago•98 comments

Tocharian Online

https://lrc.la.utexas.edu/eieol/tokol/0
26•Bluestein•1h ago•1 comments

Where did the old web go? We followed 657,607 links to find out

https://0.mk/blog/link-rot
26•tdx•1h ago•9 comments

AI At Home Part 1: A Box Of Scraps

https://jdagostino.github.io/ai-pt1-box-o-scraps/index.html
30•timmmmmmay•2h ago•12 comments

Come for ENIAC, Stay for UNIVAC and Skeduflo

https://uniqueatpenn.wordpress.com/2026/08/05/come-for-eniac-stay-for-univac-and-skeduflo/
45•cainxinth•2d ago•11 comments

DeepSeek Harness developer preview

https://deepseek.com/harness/en/
448•bjin•5h ago•203 comments

How art invented humanity

https://aeon.co/essays/humans-did-not-invent-art-it-was-the-other-way-around
51•prismatic•20h ago•8 comments

Kubernetes on Oxide: How customer needs shaped our integrations

https://oxide.computer/blog/kubernetes-on-oxide
97•stevehipwell•4h ago•37 comments

Gloomberb

https://gloom.sh/
299•rbanffy•4h ago•155 comments

Ordinary abundance

https://ordinaryabundance.com/
106•yen223•5h ago•36 comments

Choosing an AI model: one prompt, 11 models, different results

https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
129•toddmorey•5h ago•57 comments

Codex in ChatGPT desktop app for Linux is now in preview

https://community.openai.com/t/codex-in-chatgpt-desktop-app-for-linux-is-now-in-preview/1390027
406•allanrbo•13h ago•279 comments

GoAccess – Open-source real-time log analyzer and interactive viewer

https://goaccess.io/
19•gregsadetsky•2h ago•1 comments

JDK 27 G1/Parallel/Serial GC Changes

https://tschatzl.github.io/2026/08/10/jdk27-g1-serial-parallel-gc-changes.html
13•0x54MUR41•1h ago•1 comments

ATG (YC F25) Is Hiring Member of Technical Staff (Data Platform)

https://atg.science/careers
1•dkobran•6h ago

I built a 500k-domain search engine for makers in a weekend for $10

https://alexmorleyfinch.github.io/marlin/history/v1/article/the_birth.html
94•dreamforever•5h ago•53 comments

Graduate student proves a quantum uncertainty principle for fractals

https://www.quantamagazine.org/graduate-student-proves-the-fractal-uncertainty-principle-20260812/
46•bookofjoe•4h ago•6 comments

Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5

https://github.com/fellowgeek/mcp-memory
47•pcbmaker20•4h ago•28 comments

We eliminated 1,400 CVEs in NanoClaw's container images

https://www.echo.ai/blog/echo-xnanoclaw-under-the-hood
59•omrimaya•4h ago•39 comments

Launch HN: Bullet (YC S26) – A Faster Coding Agent

https://www.codewithbullet.com
24•adi1•10h ago•26 comments

Show HN: OJCP – an open protocol for agent-consumable job data

https://ojcp.dev/
16•fraywing•1d ago•3 comments

Better Gaussian Splatting in Julia

https://pxl-th.github.io/blog/better-gs-julia/
97•pxl-th•4d ago•13 comments

I requested a copy of my data from McDonald’s loyalty program

https://www.wired.com/story/mcdonalds-built-a-515-page-dossier-on-me-it-says-ill-never-leave/
158•thehoff•4h ago•192 comments

The Indo-European Family Tree

https://djbinder.com/language-tree/
40•benbreen•20h ago•30 comments

Deutsche Bank becomes first foreign yuan clearing bank in Europe

https://tradersunion.com/news/central-banks/show/2973571-deutsche-bank-becomes/
343•Markoff•6h ago•373 comments