Key Takeaways
The New York Times, the New York Daily News, and a group of fellow publishers asked a Manhattan federal judge Thursday to sanction OpenAI for allegedly withholding and destroying evidence in their copyright lawsuit over ChatGPT.
An April deposition revealed OpenAI had built a database of about 78 million de-identified ChatGPT conversations and a regurgitation-detection toolset called Project Giraffe while telling the court it could not search its own systems, the publishers say.
The motion asks the court to throw out OpenAI's 20 million chat log sample, accept as fact that the logs would show substantial regurgitation, and order OpenAI to pay legal fees.
The two-year copyright fight between news publishers and OpenAI escalated Thursday. The New York Times, the New York Daily News, and a group of fellow publishers asked a Manhattan federal judge to sanction the company, alleging it withheld and destroyed evidence at the center of their case over how ChatGPT was trained on their journalism.
What the April Deposition Revealed
OpenAI has argued for two years that searching its own training data and chat logs was technically burdensome and a privacy risk. In an April court-ordered deposition, OpenAI data privacy engineer Vinnie Monaco revealed the company had already run internal searches of its training corpus for copyrighted journalism, according to the motion reported by TechCrunch. The publishers say OpenAI also built a database of about 78 million de-identified ChatGPT conversations before the Times ever sued, and stood up a detection toolset called Project Giraffe that logged regurgitation in outputs shortly after the case began.
Deleted Logs and an Unusable Sample
The outlets originally sought 120 million chat logs and negotiated down to a sample of 20 million, which arrived last December so heavily redacted the court called it unusable. The motion says OpenAI deleted billions of ChatGPT outputs after the suit was filed, in violation of a preservation order, and substituted millions of logs in the sample. Ian B. Crosby, lead counsel for the publishers, said in a statement that if OpenAI believed the copying was legal, "it wouldn't have hid the truth about having done it."
What the Publishers Want the Court to Do
The motion asks the court to bar OpenAI from using the 20 million log sample, to accept as established fact that the logs would have shown substantial regurgitation of the publishers' work, and to make OpenAI pay the legal fees spent chasing the evidence. OpenAI denies the allegations. Spokesperson Drew Pusateri said the publishers are trying to invade the privacy of people unconnected to the case as their claims weaken, and that the company will keep defending fair use.
The dispute lands on questions this newsroom has tracked for months, from cryptographic receipts for AI training data to AI-fabricated citations spreading through published research. A judge with sanction power now gets to decide what AI accountability looks like in practice. Worth watching.
People Also Ask
What is the New York Times lawsuit against OpenAI about?
The Times sued OpenAI in December 2023, alleging the company violated copyright law by training its models on Times journalism and reproducing that content in ChatGPT outputs. The case is now in its discovery and pretrial phase in Manhattan federal court.
What sanctions are the publishers asking for?
They want the court to exclude OpenAI's 20 million chat log sample as unreliable, accept as established fact that ChatGPT logs would show substantial regurgitation of their content, bar OpenAI from arguing otherwise, and award legal fees.
Did OpenAI delete ChatGPT evidence?
The publishers allege OpenAI deleted billions of ChatGPT outputs after the lawsuit was filed, in violation of a court preservation order, and swapped millions of logs in the requested sample. OpenAI denies the allegations and calls them blatantly false.
What is Project Giraffe?
According to the sanctions motion, Project Giraffe is an internal OpenAI toolset that included a Bloom filter to detect and record when ChatGPT outputs regurgitated copyrighted material. The publishers say it was built shortly after the lawsuit was filed and never disclosed.
