OpenAI faces sanctions for allegedly hiding 80M ChatGPT logs from NYT lawsuit
OpenAI hid evidence of copyright infringement for two years, NYT alleges.
OpenAI is facing calls for serious sanctions after news organizations—led by The New York Times—accused the AI company of repeatedly lying to conceal evidence that ChatGPT could regurgitate paywalled articles. In a Thursday sanctions motion, plaintiffs claimed OpenAI misled the court for two years about the cost and feasibility of searching ChatGPT logs. The alleged deception came to light during a re-deposition of OpenAI privacy engineer Vincent Monaco, who inadvertently revealed that OpenAI had already conducted searches on large de-identified log samples before the lawsuit began.
According to the filing, OpenAI had two large samples—spanning 10 million and 78 million logs—which had already been de-identified and could have been made available to plaintiffs early in discovery. Instead, OpenAI allegedly forced plaintiffs to spend eight months searching in a heavily restricted sandbox. The news organizations argue that OpenAI withheld this evidence to protect its fair use defense, while OpenAI counters that the motion is a desperate attempt to invade user privacy as the Times' case weakens. The court has deemed the provided log sample unusable, escalating the dispute over whether ChatGPT's use of copyrighted content is transformative fair use or infringement.
- NYT alleges OpenAI misled court for two years about ability to search ChatGPT logs for copyrighted content
- OpenAI reportedly hid two de-identified log samples (10M and 78M) that had already been searched for NYT content
- Court has labeled OpenAI's provided log sample as 'unusable'; sanctions motion filed Thursday
Why It Matters
This case could set precedent on AI copyright liability and fair use for training data.