Newly unredacted filings in The New York Times’ three-year-old copyright lawsuit against OpenAI and Microsoft reveal internal admissions that executives privately viewed the companies’ AI training practices as theft and an existential threat to publishers. The material, drawn primarily from the Times’ legal brief, alleges the companies scraped paywalled content undetected, built massive training datasets, and deliberately stripped copyright notices before content reached their models.
Microsoft’s director of Applied Science, Brent Hecht, reportedly described declining click-through traffic to Times content, down as much as 93% following Copilot’s launch, as a “doom loop” threatening the underlying content supply chain in internal documents, at one point calling the scraping “the largest theft of labor in human history.” Microsoft CEO Satya Nadella testified that paywalled content should be licensed before use in training, and said he would have required OpenAI to retrain its models had he known paywalled material was involved.
OpenAI’s head of ChatGPT, Nick Turley, reportedly described the product as increasingly substitutive for publisher content, while President Greg Brockman acknowledged awareness of a method to bypass the Times’ paywall. The filings claim OpenAI’s training datasets contain more than 91,000 copies of Times, Daily News and Center for Investigative Reporting material. Neither company responded to requests for comment.