Google 6 min read

Google Tried to Copyright Its Own Search Results. A Judge Said No.

Name the company that has scraped more of the internet than anyone else. It’s not close. Google indexed the web, built a trillion-dollar business on top of it, and spent two decades arguing that reading public pages is not copyright infringement.

Then someone started scraping Google’s search results. Suddenly the argument flipped. Google reached for the DMCA — and a court told it no.

In the middle of the AI training data wars, that ruling matters far beyond a handful of SEO tool vendors.

What Actually Happened

Rank trackers, SEO platforms, and price comparison services all scrape Google’s search engine results pages. They run a query, record which sites rank where, and do it again tomorrow. Multiply by millions of keywords. That daily snapshot is the raw material for an entire industry — Semrush, Ahrefs, and every agency that bills clients for “we moved you from position 8 to position 3.”

Google hates this. Its terms of service ban automated collection, and its engineering teams have fought scrapers with bot detection, CAPTCHAs, and rate limits for years.

But a ToS violation is just a breach of contract. Contract suits are expensive, slow, and the damages are fuzzy — what’s the dollar value of someone querying your public webpage too many times? So Google went looking for a bigger stick. It found copyright law.

The Real Weapon Was Section 1201

The DMCA is more than the takedown notices that pull videos off YouTube. That’s Section 512, the polite part. Section 1201 is the sharp end: the anti-circumvention provision.

The logic is simple. If a copyrighted work is behind a technological lock, breaking that lock is illegal on its own — regardless of what you do once you’re inside. Congress wrote it in 1998 to stop DVD ripping and game console modchips. It has been stretched ever since, most famously to sue farmers for repairing their own John Deere tractors.

Google’s argument was that the same logic applies to the open web. Bot detection, CAPTCHAs, and rate limiting are access controls. Routing around them to collect search results is picking a lock.

If that had landed, scraping would have stopped being a contract problem and become copyright infringement — with statutory damages up to $150,000 per work willfully infringed. That’s not a different lawsuit. That’s a different weapon class.

Where the Court Drew the Line

The court wasn’t buying it, on two independent grounds.

First: is a search results page even Google’s creative work? Look at what’s actually on it. The links are other people’s URLs. The snippets are sentences lifted from other people’s pages. Under Feist v. Rural Telephone — the 1991 Supreme Court case that killed copyright in phone books — facts aren’t copyrightable, and neither is sweat-of-the-brow compilation. You can argue that selection and arrangement deserve protection, but Google’s arrangement is generated by an algorithm. It’s hard to claim authorship over an ordering no human chose.

Second: is bot detection an access control? Google search results are public. No password, no paywall, no encryption. Anyone with a browser sees them. The question is whether selectively turning away certain visitors at an open door counts as a lock in the statutory sense. The court was skeptical.

This tracks where US courts have been heading for years. In hiQ Labs v. LinkedIn, the Ninth Circuit held that scraping publicly available profiles probably doesn’t violate the Computer Fraud and Abuse Act, because you can’t gain “unauthorized access” to a page the whole world can already load. Different statute — CFAA then, copyright now — same underlying instinct: if you publish it openly, you don’t get to relitigate that choice through whichever law is most convenient.

Why This Is Really an AI Story

Now look at the timing.

The biggest fight on the web right now is over AI training data. The New York Times is suing OpenAI. Authors, publishers, artists, Reddit, Getty Images — the docket is crowded and getting more so.

The AI defense runs roughly like this: reading the public web is lawful, training extracts patterns rather than copying expression, and the output is transformative. It is, almost word for word, the argument Google made for twenty years to justify its index.

Then Google’s own data became the target and the company argued the opposite. This page is ours. Our technical barriers are access controls. Getting past them is a federal offense.

The company that scraped the entire web to build a search empire tried to make scraping its search results a copyright crime.

And here’s the part that should make you sit up: if that argument had worked, it wouldn’t have stayed contained. Any website could slap on a bot-detection script and convert every AI crawler into a Section 1201 defendant. News sites. Forums. Stack Overflow. Your personal blog. Statutory damages, no need to prove the model memorized anything, no fair use fight over outputs — just “you got past our lock.”

That is precisely the weapon publishers currently suing AI companies would love to have. Google nearly forged it for them, then dodged the outcome most dangerous to its own business. Whether Google’s lawyers fully appreciated that irony is a question for someone with access to the billing records.

What Doesn’t Change

None of this makes scraping a free-for-all.

Terms of service still bind you. Breach of contract claims survive intact. And the technical arms race continues regardless of what any judge says — Cloudflare now blocks AI crawlers by default across its network and has stood up a marketplace for paid crawling deals. The door is being closed at the infrastructure layer, not the courthouse. That’s faster, cheaper, and doesn’t require winning an argument about 1998 legislative intent.

There’s a stranger shift underneath all of this, too. The search results page is disappearing. AI Overviews now sit above the ten blue links, and on a growing share of queries the user never scrolls past them. Rank tracking assumes there’s a ranking to track. Winning the legal right to scrape a page is a hollow victory if the page stops mattering.

The Takeaway

The principle here is old and unglamorous: publish something publicly, and you accept that others will read it. Google, the Times, this blog — same rule, no exceptions carved for size.

But nobody writing that rule imagined readers that consume a million pages an hour, never click an ad, and turn what they read into a product that competes with the source. Whether the law should treat human reading and machine reading identically is the actual question, and courts are answering it one narrow ruling at a time.

I’m genuinely undecided. Should a writer be able to stop an AI from training on their work? Or is that the deal you sign the moment you hit publish?

Google DMCA web scraping AI copyright

Comments

    Loading comments...