NEWS
NEWS

How millions of newspaper articles were used by OpenAI and Microsoft to train their AI and end the web: "Perhaps the biggest theft in history"

Updated

The lawsuit documents from The New York Times prove how techniques were devised to bypass paywalls without paying for licenses and the strategy to become the new single window of information on the web

Techniques were devised to bypass paywalls.
Techniques were devised to bypass paywalls.JOSETXU PIÑEIRO

"It's an astonishing theft, of unprecedented proportions," perhaps "the biggest theft of work in human history." With these words, the Director of Applied Sciences at Microsoft referred to the practices used by OpenAI to feed the ChatGPT models, as reported in one of the documents presented by the defense of The New York Times in the lawsuit they maintain with 11 publications against the two tech companies for using their content without consent to train AI, and whose initial documents have been published by a New York court.

Over nearly 90 pages, the media's lawyers compile internal documents, quotes, and interviews with company executives openly acknowledging that they are illegally extracting protected content from the media and the damage they anticipate their business model will cause to the media industry.

The only divergent mention in this regard was made by Microsoft's CEO, Satya Nadella, who is quoted in a sworn statement saying that licensed content should be paid for if used. "If I had known that OpenAI had scraped and trained on information behind paywalls, we would have invoked our right to ask them to retrain the models," the executive states. However, the communications between the teams of both companies are very different.

The level of web cannibalization by the tech companies even worried researchers at Microsoft. In a document, they claimed that a "dumbing loop" was being created that would damage both the AI world and the entire web, referring to the fact that there would be less original content and more content created by chatbots, resulting in AI models having fewer sources to train with.

At OpenAI, there were no qualms about obtaining as much information as possible. The lawyers' document from The New York Times refers to a chain of messages in which a researcher reaches out to Greg Brockman, the company's president, offering him "a tool to bypass The New York Times paywall." "Oh, great," Brockman replies. Other developers of the models feeding ChatGPT agree that there was no policy in place to prevent downloading protected content from the internet, nor to subsequently remove it from the training datasets. In fact, the platform even had several third-party-developed applications explicitly designed to bypass paywalls and available to any user.

These practices were also not detected at Microsoft, which, for example, had access to everything used to train GPT-3 before uploading the models to its cloud. Additionally, the Seattle-based company developed its own projects like Mango, aimed at helping OpenAI obtain more data, and also never sought permission from news editors to gather information directly from their web pages.

OpenAI did not start implementing controls on the pages being downloaded until September 2023, almost a year after launching ChatGPT. These mechanisms were further reinforced in January 2025, in addition to endorsing the creation of a joint protocol to prevent mass downloads for content creators, which involved including the robot.txt trace in their codes, a text file indicating to bots downloading content that this page had prohibited it. However, The New York Times's lawyers claim that the text generated by ChatGPT related to their publication shows how OpenAI researchers created a way to bypass these warnings and work with that data regardless.

The analyses conducted during the investigation show that ChatGPT was trained with millions of articles from The New York Times and other media outlets like the Chicago Tribune, which had also refused to provide their content for free. To achieve this, the tech company even bought all copies of the New York newspaper from 1987 to 2007 under the pretext of using them for research purposes.

In one document, a ChatGPT product leader acknowledged that media outlets were facing an "existential threat" from a product that allowed for a "broad" replacement and would increasingly substitute them. "No matter how prominently we display the links, people won't click," noted an engineer involved in ChatGPT development.

The data collected by the parties supports this theory. According to Microsoft's information, a user accessing Times information through Copilot had between 87% and 91% less chance of clicking on a link compared to a user using traditional search, a situation that would have worsened for websites with the launch of AI-generated previews.