Short

USA Today sues OpenAI for more than 250 million dollars

USA Today and its local papers sued OpenAI on October 8 for more than 250 million dollars, saying hundreds of thousands of articles were copied to train ChatGPT. Here is what the filing claims, and what is still unproven.

On Thursday, October 8, 2026, USA Today Co. sued OpenAI in federal court in New York. The publisher, together with the local newspapers it owns, says OpenAI copied hundreds of thousands of articles to train its models and then used that reporting inside ChatGPT. The filing asks for damages of more than 250 million dollars. The Verge reported the suit the same afternoon, citing an earlier Reuters dispatch, and OpenAI had not commented when The Verge asked.

This is not a small blog picking a fight with a lab. USA Today is one of the largest newspaper companies in the United States, the company formerly known as Gannett. If its account of the training data holds up, a very large share of American local news sat inside the corpora that taught the GPT models how to write. That is the claim. It is not yet a finding. What follows is what the public reporting and the complaint, as described by those reports, actually say.

Illustration: a pencil sketch of a folded newspaper on a desk beside a laptop showing a chat window, with a red legal seal on the paper

The three things you need to know

  • The claim. USA Today Co. and the local papers it owns say OpenAI copied hundreds of thousands of articles without permission, including material behind paywalls, and used it to train GPT models.
  • The money. The complaint seeks more than 250 million dollars. Under the statute it cites, willful copyright infringement can mean up to 150,000 dollars per work, and stripping copyright notices can add up to 25,000 dollars per violation.
  • The pattern. The New York Times sued OpenAI and Microsoft in December 2023. The Verge lists further actions from The Intercept, Ziff Davis, CBC/Radio-Canada, Encyclopaedia Britannica, Merriam-Webster, The Seattle Times, and a coalition of nearly 400 local newspapers. OpenAI did not immediately respond to this filing. In earlier cases it has argued that training on publicly available text is fair use.

Who filed, and where

The Verge's account is the clean public version. USA Today Co., along with several local newspapers it owns, filed on Thursday. The publisher asks for damages of more than 250 million dollars and says OpenAI's unauthorized use of its content "has done real and continuing" harm to its outlets. The complaint, quoted by The Verge, says: "OpenAI's commercial success rests on large-scale copyright infringement." It alleges that OpenAI never asked for permission and instead "took" USA Today's content to "build products worth hundreds of billions of dollars."

A longer writeup of the docket by Unite.AI adds detail that The Verge does not. It says the case is styled USA Today Co., Inc. v. OpenAI Foundation, filed in the U.S. District Court for the Southern District of New York by Steven Lieberman of Rothwell, Figg, Ernst & Manbeck. The complaint was entered on October 8, 2026, with an exhibit of copyright registrations and a second exhibit of GPT-5.6 output examples. Same day filings included a civil cover sheet, a copyright form, a notice of appearance, and a request for a summons. A statement of relatedness filed the same day asks that the action be treated as related to the consolidated OpenAI copyright cases already pending in that court. Another report gives the docket number as 1:26-cv-08892. I have not read the complaint itself, so the procedural details below are what those reports say the filing contains.

The plaintiffs, as Unite.AI describes them, hold copyrights in content from 19 publications, all owned by USA Today Co. The list is long because that is the point of the case: this is not one flagship site. It includes USA TODAY, The Tennessean, the Indy Star, The Bergen Record, The Enquirer, the Asbury Park Press, the Democrat and Chronicle, The Knoxville News-Sentinel, the Naples Daily News, The Oklahoman, the Milwaukee Journal Sentinel, The Columbus Dispatch, The Arizona Republic, The Courier-Journal, The Des Moines Register, the Detroit Free Press, The Detroit News, The Palm Beach Post, and the Star News. The Verge names The Tennessean, Indy Star, The Columbus Dispatch and The Oklahoman as examples. A Forbes writeup syndicated by Yahoo describes USA Today Co. and 13 affiliated entities covering those 19 publications. The counts differ by a few entities depending on which story you read. The publications do not.

The complaint, again according to Unite.AI, describes USA Today Co. as a diversified media company that traces its history to 1906, publishes USA TODAY and hundreds of daily publications, says its papers have won dozens of Pulitzer Prizes, and employs hundreds of journalists across thirteen states. That corporate description matters because damages and standing turn on who actually owns the copyrights, not on the brand name on the masthead.

What the complaint says was copied

The core allegation is familiar from the New York Times case and from almost every publisher suit since late 2023. OpenAI's models, the complaint says, were trained on copyrighted material scraped from the internet without authorization, regardless of paywalls or other access limits. It also says OpenAI used programs designed to strip copyright management information, the bylines, titles and notices that identify who owns a work. Removing that information can be its own violation, separate from the copying, which is why the damages section has two numbers rather than one.

Where this filing tries to be specific is in the training sets. Unite.AI's account of the complaint says content from the plaintiffs' publications makes up more than 160,000 entries in WebText, the corpus OpenAI built to train GPT-2. Of those, about 83,266 entries came from usatoday.com and 12,994 from freep.com, the Detroit Free Press. It also says the publications' domains account for more than 122 million tokens in C4, the filtered English subset of a 2019 Common Crawl snapshot. The complaint then points at OpenAI's own published GPT-3 training mix, which weighted Common Crawl at 60 percent and WebText2, an expanded WebText, at 22 percent. If those counts are right, USA Today content was not a rounding error in at least two of the datasets the field has talked about for years.

A separate report, from CryptoBriefing, gives a slightly different cut of the same idea: the plaintiffs say they found over 160,000 articles from their domains inside large training datasets including WebText and C4, with approximately 83,266 from usatoday.com. I am treating the WebText and C4 figures as allegations in the complaint, reported by two outlets, not as numbers I have verified against the datasets.

Illustration: a pencil sketch of a tall stack of newspapers dissolving into scraps that flow toward an abstract model, with a red ribbon

The complaint, as described, also reaches for a paper trail inside OpenAI. It cites a written submission OpenAI made to a British House of Lords inquiry in December 2023, in which the company said that because copyright covers virtually every sort of human expression, limiting training data to public domain works would not produce AI systems that meet current needs. That letter has been quoted in publisher cases for almost three years. It is not a confession of infringement. It is OpenAI saying, in public, that a public domain only dataset would not be enough for the products it wanted to build. Plaintiffs read that as proof the company knew it needed copyrighted text. OpenAI has read the same sentence as a description of how language works.

Unite.AI says the complaint quotes internal communications as well. Co-founder Greg Brockman allegedly told colleagues the models were particularly good at predicting the text of news articles. A 2020 presentation by then research leader Dario Amodei allegedly listed news generation among GPT-3's skills. OpenAI's VP of Research is quoted as saying, in substance, that the company trains its networks to memorize the training data and that this is their objective. June 2022 internal messages allegedly show employees acknowledging that GPT-4 would have memorized a large amount of data and would be highly effective at regurgitating it. I am relaying those lines as the complaint's quotations, via a secondary report. They are not documents I have seen, and internal slides from 2020 are not the same thing as a 2026 product.

There is also a Microsoft thread in the complaint, according to the same report. Over a three year period, Microsoft allegedly provided OpenAI with a copy of the Bing Index, a compilation of billions of webpages, under an initiative codenamed Project Taxi, and the index included the plaintiffs' content. Microsoft allegedly also ran a crawler called Project Mango on OpenAI's behalf, which OpenAI paid for. And the complaint allegedly says OpenAI's output filters did not suppress content from any entity that had not sued, an approach a Microsoft executive supposedly described internally as an accidental cover up. If that last point is in the filing, it is aimed at the idea that OpenAI only started blocking a publisher's text after the publisher sued, which plaintiffs will argue shows the company could have filtered the material and chose not to.

Outputs, not only training

Training is half the case. The other half is what ChatGPT does with the result. The filing, Unite.AI says, reproduces examples in which GPT-5.6, asked to find and summarize a specific article by title, produced extensive multi section summaries that paraphrased the originals and followed their structure. The examples named in the report come from the Indianapolis Star, the Detroit Free Press, The Knoxville News-Sentinel, The Palm Beach Post, The Tennessean, The Enquirer, The Des Moines Register, The Courier-Journal, the Naples Daily News, the Asbury Park Press, the Milwaukee Journal Sentinel, The Columbus Dispatch, The Oklahoman, and the Star News, with full versions attached as an exhibit.

The legal theory here is substitution. A model that can reconstruct the point of a reported story, section by section, gives the reader a reason not to click, and not to subscribe. The complaint allegedly quotes OpenAI's Head of ChatGPT acknowledging that once ChatGPT gives an answer there is "no good reason to click" a link to the underlying source, and an engineer saying that no matter how prominently links are displayed, users will not click. Plaintiffs say OpenAI post trained its models to summarize copyrighted articles instead of returning the articles themselves, producing substitutes that serve the same purpose as the originals.

That theory is the live fight in the publisher cases, more than the old question of whether a training run made a copy. Courts have not settled it. A summary can be fair use, criticism, or news reporting in its own right. It can also be a replacement. The exhibit of GPT-5.6 outputs is there so a jury can look at the before and after and decide which one this is. The models named in the complaint, according to Unite.AI, run from GPT-1 through GPT-6.1 and GPT-OSS, including Instant, Thinking, mini, nano and Pro variants. The plaintiffs demand a jury trial.

How the damages number is built

"More than 250 million dollars" is a demand, not a verdict, and it is not a simple invoice. US copyright law lets a plaintiff who registered in time choose statutory damages instead of proving actual lost sales. For willful infringement the ceiling is 150,000 dollars per work. A separate provision, aimed at the removal of copyright management information, allows up to 25,000 dollars per violation. Multiply either number by "hundreds of thousands" of articles and you get figures far above 250 million. The 250 million is the number the plaintiffs chose to put in the demand. A court can award much less, including nothing, if it finds fair use, if registrations are missing, or if the works are not treated as separate.

The plaintiffs also want an injunction, a court order to stop the alleged copying. That is often the part publishers care about more than the check, because an order can force a company to filter outputs, drop a dataset, or negotiate a license. CryptoBriefing notes the request for injunctive relief explicitly.

The defendants, as Unite.AI lists them from the complaint, are not a single logo. They are OpenAI Foundation, OpenAI GP, LLC, OAI International, Inc., OpenAI OpCo, LLC, OpenAI Global, LLC, OAI Corporation, and OpenAI Group PBC, each described as a San Francisco entity that was directly involved in or profited from the alleged infringement. The complaint recounts the October 28, 2025 recapitalization, under which the nonprofit became the OpenAI Foundation, holding equity the company described as valued at about 130 billion dollars, and the for profit became OpenAI Group PBC, a public benefit corporation. Citing OpenAI's own disclosures, the complaint says ChatGPT had more than 900 million weekly active users and over 50 million paying subscribers as of March 2026. Those user figures are older than the 1.2 billion weekly figure OpenAI has used more recently for the GPT-6 rollout. Both are company numbers. The point in the complaint is scale: a product this large, built, they say, on unlicensed news.

Illustration: a pencil sketch of courthouse columns and steps, with a single red folder on the stairs

The longer line of cases

The Verge is right to frame Thursday's filing as the latest, not the first. The New York Times sued OpenAI and Microsoft in December 2023 over training, outputs, trademarks and false attribution. Since then the docket in the Southern District of New York has become the main room where American news copyright and generative AI get argued. The Verge's list of other plaintiffs is worth reading slowly, because it is not one industry. The Intercept is an investigative nonprofit. Ziff Davis owns CNET and a pile of tech and health sites. CBC/Radio-Canada is a public broadcaster. Encyclopaedia Britannica and Merriam-Webster are reference publishers, which matters because reference text is exactly what a chatbot is good at replacing. The Seattle Times is a regional daily. And a coalition of nearly 400 local newspapers has its own action. USA Today sits closer to that last group than to the Times. Its value, in this lawsuit, is the breadth of local reporting: city halls, high school sports, statehouses, obituaries, the stories that do not get rewritten by a national desk.

Licensing has moved in parallel with the lawsuits. Some large publishers have signed deals with OpenAI or with other labs. Others have refused and sued. A reader can hold both facts at once. A license with one outlet does not answer a claim from another, and a deal signed in 2024 does not automatically cover a model released in 2026. USA Today's complaint, as reported, does not say the company was offered a license and turned it down. It says OpenAI never asked.

Fair use, and what OpenAI has said before

OpenAI did not immediately respond to The Verge's request for comment on Thursday. I am not going to invent a statement. What the company has said, repeatedly, in the earlier publisher cases and in public posts, is that training on publicly available text is fair use, that models learn patterns rather than storing a library, and that output filters exist to stop verbatim copies. The House of Lords letter is the version of that argument that plaintiffs like to quote back: if you want capable systems, you cannot train only on text whose copyright has expired.

Fair use in the United States is a four factor test, and news is an awkward input for it. News articles are factual, and facts are not copyrightable, which helps a defendant. They are also expressive, highly commercial when reused inside a paid product, and easy to substitute for, which helps a plaintiff. Memorization research has shown that models can sometimes emit training text verbatim. Plaintiffs treat that as a copy. Defendants treat rare regurgitation as a bug to be filtered, not as the design. Thursday's complaint, if the internal quotes are accurate, tries to close that gap by saying employees described memorization as the objective, not the accident.

None of that is decided by a blog post or by this article. It is decided by judges, and maybe by a jury in this case, on a record that includes the exhibit of GPT-5.6 summaries. Until then, "OpenAI stole USA Today's archive" and "this is ordinary fair use" are both slogans. The filing is a set of allegations with a damages number attached.

Why it matters

Three practical consequences sit on top of the legal one.

First, local news is the part of journalism that is already the most fragile, and it is also the part a general chatbot is most likely to answer from memory. If you ask what the county board did last week, you are often asking for a story that only one paper reported. A system that can paraphrase that story without sending the reader back is a direct hit on the subscription and the ad impression. National outlets have other revenue. A paper in Knoxville or Des Moines often does not.

Second, the case arrives in the same week OpenAI put GPT-6, with interactive answers, in front of free ChatGPT users. Summaries that used to be a paragraph of text can now be a tidy card, a list, a chart. That does not change the copyright question, but it does change how complete the substitute feels. The more finished the answer looks, the weaker the reason to click. USA Today's lawyers are clearly writing with that product in mind. The exhibit is GPT-5.6, and the model list in the complaint runs through GPT-6.1.

Third, statutory damages are a negotiating weapon. A demand north of 250 million dollars, with a per work ceiling of 150,000 dollars behind it, is how you get a licensing conversation even if you never expect a jury to award the headline number. Other publishers will read Thursday's filing as a template: count your URLs in WebText and C4, attach output examples, file in the Southern District, and ask to be related to the consolidated cases. If that template works, the next year of AI copyright news will look a lot like this one.

There is a counterpoint worth keeping. A world where every newsroom must be paid before a model can learn from the open web is a world where only the labs that can write the biggest checks get to build models, and where small open projects are simply illegal. That is a real policy cost, and it is why fair use exists. The honest version of this story is not "AI versus journalism" as a team sport. It is a fight over whether a statistical model of the web is more like reading, which nobody licenses, or more like a database of articles, which you do.

Dany's take

I am glad this one was filed by the local papers and not only by a national brand. The public argument about AI and news has been stuck on the New York Times for three years, and most of what people actually read is not the Times. If USA Today's numbers are even roughly right, tens of thousands of its own stories were in WebText before most of us had heard the word ChatGPT. That does not make the legal answer obvious. It does make the "we only used the open web, so nobody was harmed" line harder to say with a straight face, because the open web is where those papers published, and the harm they describe is readers who never come back.

I also do not want the remedy to be a ban on training. The useful outcome is a license, a filter that stops near copies, and a link that is worth clicking. OpenAI already pays some publishers. It can pay more, or it can win on fair use and keep the status quo. What it cannot do, after a filing this specific, is pretend the question is still abstract. The exhibit is sitting in a New York court with summaries of stories from Indianapolis, Detroit, Knoxville and Palm Beach. Either those summaries are fair, or they are the product. I would like to see the exhibit. Until it is public in full, I am not going to pretend a secondary writeup is the whole case.

If you publish anything online, this is the week to decide what you think. Not in general. On the facts in front of you: a newspaper company, 19 publications, a demand above 250 million dollars, a training set count someone will now have to prove, and a product that answers the question so you do not have to open the paper. Should publishers get paid when a model trains on their work, or is that just how reading works now? Tell me on Reddit, X or YouTube.

Source

Primary report: The Verge, "USA Today becomes the latest publisher to sue OpenAI". Reuters report, as linked by other coverage: Reuters via TradingView. Complaint detail as reported by Unite.AI and CryptoBriefing. I have not read the complaint PDF myself. Allegations in the filing are not findings.

Source: theverge.com

Newsletter

The AI news that matters, in your inbox.