A coalition of newspapers led by The New York Times has asked a Manhattan federal court to punish OpenAI, arguing the company spent two years falsely claiming it could not search its own systems for stolen journalism while, all along, it was quietly running exactly those searches. The motion for sanctions, filed July 9, 2026, transforms a technical discovery squabble into an accusation of deliberate deception aimed at the heart of one of the most closely watched copyright cases of the AI era.
The filing landed in the consolidated litigation before the U.S. District Court for the Southern District of New York, where The Times, The New York Daily News, the Chicago Tribune and other outlets accuse OpenAI and Microsoft of illegally training generative models on their articles and reproducing that reporting through ChatGPT. Discovery has been overseen by Magistrate Judge Ona T. Wang, whose rulings have repeatedly gone against OpenAI and were affirmed in January by District Judge Sidney H. Stein.
At the center of the dispute is a question that sounds mundane but carries enormous stakes: can OpenAI actually search its training corpus and its vast archive of ChatGPT conversation logs for the publishers' copyrighted work? For roughly two years, OpenAI told the court it could not, or that doing so would be prohibitively burdensome and would compromise the privacy of users whose logs would need to be retrieved, processed and de-identified. The publishers wanted that data to prove two things: that their journalism sat inside OpenAI's training set, and that ChatGPT regurgitated it to users.
According to the plaintiffs, that story collapsed in an April 2026 court-ordered deposition of OpenAI data privacy engineer Vinnie Monaco. Monaco allegedly acknowledged that OpenAI had already conducted internal searches of its training corpus for copyrighted journalism, and that beginning before The Times even sued, the company had assembled a database of roughly 78 million de-identified ChatGPT conversations to study internally how much it was infringing others' work. Shortly after the suit was filed, the plaintiffs say, OpenAI deployed a "Bloom" filter as part of an internal toolset dubbed "Project Giraffe" that detected and logged instances of models regurgitating text.
"If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it," said Ian B. Crosby, lead counsel for the plaintiffs.
The publishers frame the alleged concealment as part of a broader pattern of discovery misconduct. They had originally sought a sample of 120 million chat logs; OpenAI negotiated that down to 20 million. When the company finally handed over the sample last December, the plaintiffs say it arrived so heavily redacted that the court itself described it as "unusable," and that OpenAI had substituted millions of logs in the requested set. The motion also revives an earlier flashpoint: the plaintiffs' claim that OpenAI deleted billions of ChatGPT outputs after the suit was filed, in violation of the court's evidence-preservation order.
The remedies the newspapers are demanding are severe. They ask the court to bar OpenAI from relying on the 20-million-log sample as evidence, on the grounds that it is unreliable; to find as an established fact that the logs would have shown substantial and systematic reproduction of the plaintiffs' journalism; to prohibit OpenAI from arguing the produced logs fail to demonstrate such regurgitation; to award attorneys' fees for the effort spent chasing the evidence; and to instruct any eventual jury about OpenAI's alleged destruction of evidence.
OpenAI rejects the accusations and casts the motion as a distraction by a weakening plaintiff. "As the Times' case weakens and they've been forced to drop claims against us, they're persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations," said OpenAI spokesperson Drew Pusateri. "We'll continue defending our users' privacy and the long-established principles of fair use."
Why It Matters
Sanctions motions are the sharp end of civil litigation, and this one targets the evidentiary foundation of the entire case. If Judge Wang or Judge Stein credits the publishers' account, adverse-inference instructions and evidentiary bars could effectively hand the plaintiffs findings they would otherwise have had to prove at trial. The dispute also crystallizes a structural problem in AI copyright litigation: the most probative evidence lives inside opaque model weights and proprietary logging systems that only the defendant can meaningfully search. Courts are being asked to police preservation and candor in a domain where verifying a company's claims about its own technical capabilities is extraordinarily difficult. A finding of spoliation or misrepresentation against OpenAI would reverberate across the dozens of pending suits brought by authors, artists, music labels and other publishers, signaling that "we can't search that" will no longer be accepted at face value.
What to Watch
OpenAI's formal opposition brief and the plaintiffs' reply will flesh out the factual record, likely including deposition excerpts and internal documents about "Project Giraffe." Watch for whether Judge Wang orders an evidentiary hearing or additional depositions before ruling, and whether any sanctions decision is appealed to Judge Stein. The outcome will shape both the leverage each side carries into settlement talks and the discovery playbook every AI defendant faces next.
"If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it."— Ian B. Crosby, Lead counsel for the news plaintiffs