Aaron Swartz Faced 35 Years for Downloading Papers. Meta Torrented 80 Million Books.
In January 2013, a 26-year-old programmer took his own life. His crime: connecting a laptop to MIT’s network and downloading academic papers in bulk. Federal prosecutors were seeking up to 35 years. Twelve years later, Meta pulled millions of copyrighted books off pirate torrent sites to train its AI models, and a federal judge largely ruled in its favor. Same category of act — unauthorized mass acquisition. Wildly different endings.
What Aaron Swartz Actually Did
Swartz co-authored the RSS 1.0 spec at 14. He was an early Reddit developer. He helped build the technical architecture of Creative Commons. He was also, unambiguously, an activist who believed publicly funded research belonged to the public.
In late 2010, he plugged a laptop into MIT’s campus network and ran a script that pulled roughly 4.8 million articles from JSTOR, the academic journal database. Here’s the part that gets skipped: MIT granted JSTOR access to anyone on campus. The access itself was authorized. What tripped the wire was the speed and the volume.
And Swartz never distributed a single paper. JSTOR — the supposed victim — dropped its civil claim and publicly said it considered the matter closed. Federal prosecutors pressed ahead anyway.
Thirteen Counts, Built on a Law from 1986
The weapon was the Computer Fraud and Abuse Act, written in 1986 — before the commercial web existed. Congress passed it in the shadow of the movie WarGames. It criminalizes “unauthorized access” without ever defining, with any precision, what authorization means. For decades, that ambiguity let prosecutors treat terms-of-service violations as felonies.
Prosecutors opened with four counts. They later expanded to 13, stacking a maximum exposure of 35 years and a $1 million fine. The escalation followed Swartz’s refusal to take a plea deal — a pattern critics called what it looked like: leverage. He was found dead in his Brooklyn apartment on January 11, 2013, three months before trial.
The reform effort that followed, nicknamed Aaron’s Law, never passed. What eventually narrowed the CFAA was the Supreme Court’s 2021 decision in Van Buren v. United States, which rejected the expansive reading prosecutors had relied on. That was eight years after Swartz died.
What Meta Actually Did
Court filings unsealed through 2025 laid out the mechanics. To train Llama, Meta sourced data from LibGen and Anna’s Archive — pirate shadow libraries, not gray-area scrapes. Estimates from the litigation run into the millions of books and, by some counts, north of 80 million documents.
The method matters more than the volume. BitTorrent is not a download protocol; it is a swarm protocol. While you pull chunks, you serve chunks to everyone else in the swarm. Downloading and uploading happen simultaneously. Internal communications surfaced in discovery show employees explicitly worried about torrenting from corporate IP addresses, and evidence suggesting efforts to minimize the seeding footprint.
Run that through copyright doctrine and it is, on its face, the heavier conduct. Swartz rapidly downloaded material he was authorized to access and shared none of it. Meta knowingly pulled from pirate sources and, by the nature of the protocol, redistributed as it went.
So Why Did the Court Split the Other Way
In June 2025, Judge Vince Chhabria of the Northern District of California ruled for Meta in Kadrey v. Meta. Read the actual opinion, though, and the win is narrower than the headline.
The court accepted that training a model on copyrighted text is transformative use — closer to a person reading novels and absorbing style than to reprinting them. But Chhabria went out of his way to write that the plaintiffs lost because they failed to develop a “market dilution” theory, not because Meta’s conduct was clean. He effectively left a roadmap for the next set of lawyers.
Judge William Alsup drew a sharper line in the parallel Anthropic case. Scanning lawfully purchased books to train a model: fair use. Downloading pirated copies to build a permanent internal library: a separate act of infringement, full stop. He split acquisition from training. Apply that framework to Swartz and yes, the download itself is the exposed step. It still doesn’t get you to 35 years in federal prison.
That’s the real fault line. Swartz was charged under criminal law. Big Tech gets sued under civil copyright law. Criminal charges are a prosecutor’s discretionary choice. One road ends in a cell; the other ends in a settlement line item. Someone decides which road you’re on.
The Asymmetry Has a Simpler Name
The tech community keeps resurfacing this comparison for a reason, and it isn’t nostalgia. The rules aren’t different. The capacity to absorb the rules is.
Meta has an in-house legal department and a litigation budget larger than most companies’ revenue. Losing does not threaten its existence. When news broke that Anthropic had agreed to a $1.5 billion settlement with authors, the industry reaction was not “ruinous.” It was closer to “cost of doing business” — a big number that still prices out as a line in an operating budget. For Swartz, the trial itself was the punishment: bankruptcy, a felony record, a life dismantled before any verdict.
There’s a second asymmetry, and it’s uglier. Swartz was trying to return taxpayer-funded research to the public. There was no revenue model. Meta is building a product line worth hundreds of billions. The law was gentler with the second one. Infringement that generates enterprise value gets called innovation. Infringement that generates none gets called a crime.
What’s Still Open
Treating this purely as a 12-year-old tragedy misses the live problem. Right now, researchers scraping public data, data journalists pulling records, and small startups building datasets are all operating in the same exposed zone. The CFAA is narrower post-Van Buren, but terms-of-service liability and prosecutorial discretion haven’t gone anywhere. The entities that got a safe harbor are the ones with “AI training” in the justification and a legal department to argue it.
Worth noting: this is based on published opinions and unsealed filings, and a substantial share of these cases are still in active litigation. Kadrey is not the last word.
When the same act gets two different names, someone chooses the name. Whoever gets charged next will find out which one they were assigned.
Deepen your perspective
Comments
Loading comments...