The New York Times' copyright lawsuit against OpenAI and Microsoft has recently revealed more unedited materials. New documents show that the two companies have discussed multiple times internally the dependency of AI on news content for training, which could directly impact publishers' traffic, revenue, and employment base.
Internal documents mention "theft"
According to the litigation documents, a Microsoft executive once privately described the act of capturing AI as "theft." In another internal memo from January 2023, Microsoft's Director of Applied Sciences, Brent Hecht, referred to it as "an unprecedented and astonishing theft," and wrote that it was "the largest scale of labor theft in human history."
The document also shows that similar risks were discussed within OpenAI. Nick Turley, who is in charge of ChatGPT, stated in internal communications that chatbot products pose a "survival threat" to publishers, and that this substitutability will become even stronger as the model capabilities improve.
News website traffic under pressure
Litigation documents claim that Microsoft's own data shows that the “answer engine” of Copilot leads users to click less on the original news websites. Compared to traditional Bing searches, the click-through rate to the New York Times domain name decreased by a maximum of 93%.
In an internal presentation in January 2024, Hecht referred to this trend as a "doom cycle." The document stated that if the commercial foundation of the content providers is weakened, both the model itself and the entire network ecosystem would be affected negatively.
Microsoft CEO Satya Nadella also stated in his testimony this year that for any content protected by paywalls, anyone who wishes to use it for training or enhancement purposes should obtain authorization. He also mentioned that if he had known at the time that OpenAI was using content behind paywalls for training purposes, Microsoft could have required OpenAI to retrain the models.
Crawling methods and the scale of data exposed
The newly disclosed materials also describe how two companies obtained news content. The lawsuit claims that OpenAI and Microsoft established training datasets through large-scale scraping, with some of the content coming from the Bing index, as well as a large amount of web page data from Common Crawl.
- Mid-term training data contains over 91,692 copies of news articles
- nytimes.com contains over 2 million documents in a set of data.
- Project Mango contains at least 160,903 copies of independent works.
The document also alleges that OpenAI employees discussed how to bypass The New York Times' paywall without being detected, and removed copyright notices before training to prevent the model from providing users with relevant copyright information.
Additional information:At present, some of the underlying evidence is still under seal. The newly disclosed information this time mainly comes from court documents submitted by The New York Times; neither OpenAI nor Microsoft has responded to requests for comments from TechCrunch.











