Is AI stealing from websites?

Updated 2026-07-15720 searches/moRanked #410 of 519· AI explained
Short answer

Courts are drawing a line: training itself may be legal, but how you get the data isn't. In Bartz v. Anthropic (June 2025), Judge William Alsup ruled training on lawfully bought books was fair use — while using millions of pirated copies was not. Anthropic settled for $1.5 billion. The New York Times v. OpenAI, filed December 2023, is still unresolved in 2026.

Why — the first-principles explanation

"Stealing" is the wrong frame legally, and that's not a technicality — it's why the cases turn out the way they do. Theft means depriving someone of a thing. Copying doesn't: the Times still has its articles. So the actual legal question is copyright infringement, which asks whether an unlicensed copy was made and whether it qualifies as fair use.

Fair use is not a yes/no switch; it's a four-factor balance, and the factor doing the heavy lifting is whether the use is transformative — does it serve a genuinely different purpose than the original? AI labs argue that reading millions of articles to learn statistical patterns of language is nothing like reading an article for the news, so it's transformative. That argument has real force, and in June 2025 it partly won.

Here's the ruling everyone misreports. In Bartz v. Anthropic, Judge William Alsup granted summary judgment that training on digital copies was fair use. But the same ruling found Anthropic had used millions of pirated library copies, and that piracy could not be fair use. The case headed to trial on damages for the pirated copies, and Anthropic settled in September 2025 for $1.5 billion — roughly $3,000 per book, described as the largest copyright resolution in US history. Read those together and the emerging principle is sharp: the learning is probably fine; the acquisition is what costs you a billion dollars.

The second front is output, not input, and it's where NYT v. Microsoft and OpenAI lives. Filed December 27, 2023, it alleges not just unauthorized training but that ChatGPT and Copilot reproduced near-verbatim Times reporting and fabricated attributions. Judge Sidney Stein denied most motions to dismiss in March 2025, letting core copyright claims proceed. The fight has since turned ugly over evidence: on July 9, 2026, the news plaintiffs asked the court to sanction OpenAI, alleging it told the court it couldn't search its models for their content while having already done so internally — including assembling a database of about 78 million de-identified ChatGPT conversations to assess its own infringement exposure, and building detection tooling. Those are allegations, not findings. But they show the live question is no longer "is training theft?" It's "how much does the model regurgitate, and who has to prove it?" Meanwhile the economic complaint publishers actually feel — that AI answers satisfy the query so readers never click — isn't copyright at all. There is no law entitling you to traffic.

An example that makes it click

Two art students copy from the same museum.

The first buys a ticket, sits with a Vermeer for three hundred hours, studies how light falls on cloth, then paints entirely original work that's informed by Vermeer. Nobody thinks she stole anything. That's training on lawfully obtained material — and it's roughly what Judge Alsup blessed.

The second breaks in at night to study. Even if his resulting paintings are just as original, he committed a crime getting in. The originality of his output doesn't undo the break-in. That's the pirated-books problem — and it's the $1.5 billion.

Now a third: she buys a ticket, studies, and then paints a canvas so close to the Vermeer that a buyer would take hers instead. That's not learning anymore, that's substitution — and that's the near-verbatim regurgitation claim at the heart of the Times case, still being fought.

Key facts

Infographic: Is AI stealing from websites — short answer and key facts
Visual summary — Is AI stealing from websites?
▶ The 60-second explainer (script)

Is AI stealing from websites? Courts are actually answering this now, and the answer is sharper than either side wants. First, drop the word stealing. Theft means depriving someone of a thing. Copying doesn't — the Times still has its articles. The real question is copyright infringement, and whether it's fair use. Now the ruling everyone gets backwards. Bartz versus Anthropic, June 2025. Judge William Alsup ruled that training on digital copies was fair use. Training. Was. Fair use. But the same ruling found Anthropic had used millions of pirated library copies — and piracy could not be fair use. Anthropic settled in September 2025 for one point five billion dollars. About three thousand dollars a book. Largest copyright resolution in US history. Put those together and the principle is brutal and clear: the learning is probably fine. The acquisition is what costs you a billion dollars. Two art students copy from the same museum. The first buys a ticket, studies a Vermeer for three hundred hours, paints original work informed by it. Nobody thinks she stole. The second breaks in at night. Even if his paintings are just as original, he still committed a crime getting in. The originality doesn't undo the break-in. That's the pirated books. Then there's a third student — buys a ticket, studies, then paints something so close a buyer takes hers instead of the Vermeer. That's not learning. That's substitution. And that's the New York Times case: filed December 2023, alleging ChatGPT reproduced near-verbatim reporting. Most motions to dismiss were denied in March 2025. It's still unresolved, and it's gotten ugly — in July 2026 the papers asked the court to sanction OpenAI, alleging it claimed it couldn't search its models while it had already done so internally. Those are allegations, not findings. One last thing. The complaint publishers actually feel — AI answers the question so nobody clicks — isn't copyright at all. Nobody has a legal right to your traffic.

What authoritative sources say

Bartz v. Anthropic — Wikipediaorg — Judge William Alsup granted summary judgment on June 23, 2025 that Anthropic's use of digital copies for training was fair use, but found that its use of millions of pirated library copies could not be fair use; Anthropic settled in September 2025 for $1.5 billion (~$3,000 per book plus interest), the largest copyright resolution in US history. source ↗
The New York Times v. Microsoft and OpenAI — Wikipediaorg — The New York Times v. Microsoft and OpenAI was filed December 27, 2023 in the Southern District of New York alleging unauthorized training on millions of articles, near-verbatim reproduction, and false attribution; Judge Sidney H. Stein denied most motions to dismiss in March 2025, and OpenAI asserts a transformative fair use defense. source ↗
TechCrunch — New York Times says OpenAI hid evidence in ChatGPT copyright trialmedia — On July 9, 2026, the NYT-led group asked the court to sanction OpenAI, alleging it misrepresented its ability to search training data; a deposition of OpenAI engineer Vinnie Monaco allegedly revealed internal searches of the training corpus and a database of ~78 million de-identified ChatGPT conversations used to assess infringement. source ↗

People also ask

Did a court actually rule that AI training is legal?

Partly. In Bartz v. Anthropic, Judge Alsup ruled training on lawfully obtained copies was fair use. But he ruled that using pirated copies was not — and that piece cost Anthropic $1.5 billion.

So is it illegal to scrape my website for AI training?

Unsettled. The rulings so far turn on how material was acquired and whether output substitutes for the original. Publicly accessible isn't the same as licensed, and the NYT case is still being fought in 2026.

Can I stop AI companies from using my site?

You can block known AI crawlers via robots.txt and your host or CDN, and many publishers now do. Compliance is voluntary, and robots.txt doesn't retroactively remove anything already trained on.

Is losing traffic to AI answers illegal?

No. It's real economic harm but not a copyright injury — there's no legal entitlement to referral traffic. That's why publishers litigate on reproduction and licensing instead of on lost clicks.

What's the difference between the Anthropic and OpenAI cases?

Anthropic's was about inputs — pirated books. The Times case is largely about outputs — whether ChatGPT reproduces near-verbatim articles and fabricates attributions. Different questions, different answers.

Related questions