出典:Hacker News原文を見る ↗
原文の著作権は出典元に帰属します。当サイトでは収録、翻訳、体裁調整のみを行います。
解説と影響
旧网页都去哪了?一项追踪 657,607 个链接的研究揭示了互联网的"腐烂"速度
研究发现了什么?
该研究追踪了超过 65 万个链接,试图回答一个看似简单却难以精确回答的问题:旧网页都去哪了?所谓"链接腐烂",指的是曾经有效的 URL 随着时间推移逐渐失效——页面被删除、域名过期、网站改版后未做重定向、内容迁移至新地址却未保留旧链接等。研究通过大规模爬取和验证这些链接的当前状态,量化了不同年份、不同来源的链接失效率。原文未提供具体的失效率百分比数据,但 HN 评论区 129 条讨论表明,许多开发者对自身项目文档、技术博客和 API 参考链接的快速失效有切身体会。
为什么链接腐烂是个工程问题
链接腐烂不仅是互联网考古学的趣味话题,更是实实在在的工程挑战。对于依赖外部文档链接的软件项目、引用在线参考文献的学术论文、以及需要长期维护的技术博客而言,链接失效意味着知识链条的断裂。开发者社区长期以来的应对方案包括:使用 Internet Archive 的 Wayback Machine 存档关键页面、在文档中同时提供多个镜像来源、采用 DOI 等持久标识符系统。然而,这些方案各有局限——Wayback Machine 的存档覆盖并不完整,DOI 系统主要面向学术出版,普通网页缺乏统一的持久化机制。
更广泛的语境:互联网平台的结构性变迁
这项研究发布的时机值得注意。同期的 Hacker News 上,Bluesky 协议服务的发布(93 分)代表了去中心化协议对传统集中式平台架构的挑战;而 Apple 与 Epic 关于外部链接收费的纠纷以及法官要求 Google 简化竞品应用商店安装的报道,则揭示了平台封闭生态对内容可达性的影响。链接腐烂的根源之一,正是内容被锁定在商业平台的围墙花园内——当平台调整策略、关闭服务或修改 URL 结构时,大量外部链接随之失效。从这个角度看,开放协议和可移植的内容标识符,或许是缓解链接腐烂的结构性方案。
参考資料
出典原文
← Back to blog July 25, 2026 · Updated August 11, 2026
Where did the old web go? We followed 657,607 links to find out.
An old 0.mk database backup held 657,958 links created between 2009 and 2014, along with their click counts.** We restored 657,607 of those records as pre-2015 links and followed every destination in August 2026. Of 655,178 safe, crawlable link records, 76.7% no longer returned a loading page.
Most 0.mk users were in Macedonia, so this is not a census of the entire web. It is a large surviving record of what one online community shared during that period, including local news, personal blogs, photo hosts, forums, and the major platforms of the time.
When 0.mk started in 2009, it was a passion project built by a team of three. We worked on it when we could, usually for a few hours a week around our regular jobs. Seventeen years later, one of us found an old database backup on a disk and decided to bring it back.
Here is what those six years of link creation look like, with the long silence after them:
2009: 3,668 links 2009 4k 2010: 20,283 links 2010 20k 2011: 103,053 links 2011 103k 2012: 23,148 links 2012 23k 2013: 224,931 links 2013 225k 2014: 282,524 links 2014 283k 2026: 404 links 2026 404 Raw link records, not users. The 2011 spike includes one 83,398-link batch; 97.8% of 2013 records and 99.9% of 2014 records are not attached to a recovered account. The green sliver is the 2026 relaunch.
The survival test
The crawl covers all 657,607 restored link records dated through December 2014. We excluded 2,429 records whose targets were malformed, internal, credentialed, or policy-blocked, leaving 655,178 crawlable historical links:
Could not connect: 51.24% HTTP error: 25.44% Loaded: 23.32% 51.24 % could not connect ( DNS, timeout, TLS ) 25.44 % http error ( 4xx / 5xx ) 23.32 %** loaded ( 2xx / 3xx response ) Even that 23.3% overstates how much survived. A login wall, a parked domain full of ads, or a "this content is no longer available" notice all count as loading. A working page does not mean the original content is still there.
Why 657,607 links but 494,781 URLs? Multiple short links sometimes point to the exact same destination. There are 162,826 such repeat records. Counting each destination once leaves 494,781 distinct URLs, of which 492,620 were crawlable. Only 21.3% of those loaded. The percentage barely moves when repeated destinations are removed: 78.7% still did not load.
At the unique-URL level, 55.0% failed at the network layer after retrying uncertain results from a second network, and 23.7% returned an HTTP error. The most common HTTP result was 404, across 76,403 distinct URLs. Another 29,663 returned 403 or 429; those pages did not load for the crawler, but may be blocking automated requests rather than missing. A 403 or 429 can mean the site blocked our crawler, so "did not load" is more honest than saying every one of those pages is gone.
The same pattern appears at the domain level. Of 133,605 crawlable hostnames, only 34,827 had even one URL load. The other 98,778 had none.
2009 URLs 64.58 % Hosts 61.53 % 2010 URLs 60.39 % Hosts 60.36 % 2011 URLs 92.53 % Hosts 61.74 % 2012 URLs 59.43 % Hosts 62.46 % 2013 URLs 75.06 % Hosts 75.24 % 2014 URLs 78.16 % Hosts 75.45 % Share with no loading page in the complete crawl. URL-level results count every distinct path; host-level results count each hostname once. The 2011 split explains the strange annual totals. One account created 83,398 distinct links to pelaphptutorials.com. At URL level, 92.5% of 2011 destinations did not load. Count that host once and the figure is 61.7%, almost identical to 2010 and 2012.
The annual totals do not show a collapse in ordinary usage during 2012. Remove that one batch and 2011 falls from 103,053 records to 19,655; 2012 had 23,148. Almost all records from 2013 and 2014 are anonymous in the recovered data, and three quarters of their hostnames have no loading URL. Raw link volume is not a user-growth curve.
Many of the recognizable survivors are giants: YouTube, Wikipedia, and Google properties. Personal blogs, forums, local news sites, and photo hosts appear throughout the unavailable set. The centralized web has generally held up better than the small web.
A walk through the graveyard
The database reads like a museum of the 2010s internet. Some residents, with the number of links pointing at them:
Facebook photo CDN (fbcdn.net), none loaded 835 links Google Code, now redirects many URLs to its archive 803 links PureVolume, the old service is gone but its domain responds 796 links Rapidshare, no URL loaded 139 links Megaupload, no URL loaded 71 links Picasa Web Albums, no URL loaded 69 links People shared Facebook photos as direct CDN links; none of the 789 distinct fbcdn.net URLs behind those 835 records loaded. Yet PureVolume now returns pages for 633 of 653 distinct URLs, and Google Code loads or redirects 628 of 754. The original services are gone, but their domains respond. An HTTP response is not the same as preserved content.
The Macedonian layer
0.mk was the first Macedonian URL shortener, so the data is also a record of a national web that partly no longer exists. The links point to A1 Television (shut down 2011), and to the newspapers Utrinski Vesnik, Dnevnik, and Vest, all of which stopped publishing in 2017. Hundreds of links to local news that can no longer be read anywhere except, sometimes, the Internet Archive. The short links outlived the newsrooms.
The gems
Seventeen years of other people's bookmarks contain some treasures:
The first link ever shortened** (July 14, 2009, 3:52 AM) was not a manifesto or a launch post. It was a CSS stylesheet on someone's WordPress blog. Two clicks, ever. Empires begin humbly.
On day two, someone shortened localhost.** 0.mk/localhost pointed at http://127.0.0.1/ . It got two clicks, each of which sent the visitor to their own machine. The shortest URL for the loneliest destination.
The shortest link points at the longest domain.** 0.mk/1 has recorded 10,415 clicks while pointing to thelongestlistofthelongeststuffatthelongestdomainnameatlonglast.com , a 2000s curiosity that is itself now gone. Four characters pointing at sixty-three, for 17 years.
The longest URL we ever shortened is 38,753 characters**, a 2012 CodePen link whose query string literally repeats TRYINGTHEMAXIMUM_URL. Someone was testing us. We passed, and we still have their test.
4,478 of our links point at other URL shorteners**: bit.ly, TinyURL, goo.gl. A short link to a short link, twice the fragility. Google shut goo.gl down in 2025, so every one of those is now a chain with a missing middle: our half still works, and points at a service that no longer resolves. Link rot squared.
And the immortal one:** 0.mk/7 , created July 15, 2009, points at google.com. 95,999 clicks and counting, and it still points there today, seventeen years and one resurrection later.
Why 0.mk came back
By 2014, 0.mk's revenue did not cover hosting or the work required to keep it running. Spam was constant. Filtering it meant more engineering, and abuse reports needed someone to review them. The original team closed the service.
Seventeen years later, AI has changed that equation. It now handles much of the development, spam detection, abuse review, support, and monitoring that the old project could not afford. That made bringing 0.mk back realistic. The recovered links now run from edge machines in more than 300 cities. More about how that works is here.
How we checked
We tested every restored link dated before January 1, 2015. The crawler followed up to five redirects and retried connection failures from a second network. We counted HTTP 2xx and 3xx responses as loading, reported HTTP errors separately, and never contacted local or unsafe addresses.