The Internet Is Losing Its Memory, and AI Is Speeding It Up
Open a bookmark you saved ten years ago. Odds are decent you get a 404. But there’s a newer failure mode, and it’s stranger: the page is alive and well, and nobody goes there anyway. The AI read it for you and gave you the gist. The internet has found a new way to forget.
Link rot is an old disease
The benchmark study here is still Pew Research Center’s 2024 analysis. Of the web pages that existed in 2013, 38% were unreachable by 2023. Narrow it to pages created in that single year and it gets worse: roughly a quarter of them didn’t survive the decade.
Government sites aren’t immune. 21% of US federal web pages contained at least one broken link. News sites hit 23%. And 11% of the references on English Wikipedia now point at pages that no longer exist. The citation survives; the evidence evaporated.
This isn’t a legacy-web problem. It’s happening right now. That link you dropped into a doc this morning has, statistically, about a one-in-three chance of being dead in ten years.
AI summaries cut off the path back to the source
Now add a new variable. Search something and an AI answer sits at the top, pre-digested.
Pew tracked how 900 US adults actually searched in March 2025. When an AI summary appeared, users clicked a standard search result 8% of the time. Without a summary, 15%. Roughly half. The click-through rate on the citation links inside the AI summary was worse: 1%.
Flip that number around. When AI summarizes, one person in a hundred checks the original. For the other 99, that source might as well not exist.
Publishers report the mirror image. Many have seen search referral traffic fall by close to half since late 2024, and some informational sites took steeper hits than that. No traffic means no ad revenue. No revenue means the site closes. And when the site closes, its pages go with it.
Crawlers in, humans out
Cloudflare published a stat that captures the imbalance neatly: the ratio between how often AI crawlers scrape a site and how often a human actually arrives via that AI varies wildly by company, and it is lopsided in every case. Thousands of scrapes on one side, one visitor on the other.
The old bargain was simple. A search engine crawled your site and sent you readers. You gave up crawling, you got traffic, and the web ran on that equilibrium.
Now only one side keeps taking. So sites throw up login walls, put content behind paywalls, and slam the door with robots.txt. Perfectly rational defense. It also has a side effect: the door that blocks AI crawlers blocks archival bots too. The measure meant to protect content from AI shuts the last passage that would carry it into the future.
The archive is a thin seawall
You might reasonably say: the Internet Archive has this covered. The Wayback Machine holds more than 900 billion web page snapshots. It is one of the largest libraries humanity has ever assembled.
It is also alarmingly fragile. The Internet Archive is run by a single nonprofit. In October 2024, a breach exposed 31 million account records and knocked the service offline for days. It lost its lawsuit against publishers. The memory of the entire web is staked on one organization’s balance sheet and legal luck.
There are technical limits too. Archives preserve static HTML well. What they don’t preserve well: anything behind a login, JavaScript-rendered views, algorithmic feeds that show every user something different. A large share of the last fifteen years of online conversation happened in exactly those places.
The scary part is the quiet disappearance
A 404 is honest, at least. It tells you something is gone.
The dangerous case is the page that lives on while nobody can reach it. Pushed down in search, absorbed into an AI summary, no longer linked by anyone — technically present, functionally vanished. It shows up in no statistic. It exists only in a server log.
Then it compounds. AI learns from the web, its output goes back onto the web, and the next generation of models trains on that. Summaries of summaries fill the space the originals left. What gets shaved off first is exactly what mattered: minority positions, specifics, context, edge cases. What remains is a memory sanded down to the average.
What you can actually do
I don’t have a grand fix. A few individual habits do help.
For anything that matters, don’t just link it — push a snapshot to the Wayback Machine. It takes thirty seconds. When you cite something, include both the original URL and the archive URL, and somebody in 2036 will be grateful. For writing you genuinely care about, save a local copy, PDF or plain text. The cloud disappears too.
And the more plausible an AI answer sounds, the more it’s worth clicking through to the source. That click is a vote to keep the original alive.
We were told the internet was the place where everything lives forever. It turned out to be closer to the opposite, and AI has now put a foot on the accelerator. Look back at the 2026 web from twenty years out and what will be left? Probably just whatever somebody thought to save.
Comments
Loading comments...