Top staff at OpenAI and Microsoft knew that scraping news websites to train their artificial intelligence models may eventually damage or even destroy those publications. They did it anyway, according to a legal filing unsealed Thursday in a high-stakes copyright infringement case.
"Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time," Brent Hecht, Microsoft's director of applied science, wrote in an internal document soon after the New York Times sued, according to the filing. "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"
An OpenAI executive referred to an "existential threat" to publishers, according to the filing. Meanwhile, a Microsoft researcher told colleagues that compensating content creators is "in the best interest of my employer, of my country, and of many other groups I belong to."
OpenAI employees discussed opportunities to skirt news sites' paywalls, and leaders at both companies said their products replaced the need for users to check out original sources, according to the filing. Many publishers historically relied on web traffic to generate ad revenue and gain new subscribers.
The Times sued Microsoft and OpenAI in December 2023, alleging that ChatGPT and Microsoft's Copilot trained on millions of pieces of Times content and drew on that material to serve up answers to users' queries. The Chicago Tribune, New York Daily News and other large newspapers owned by Alden Global Capital sued both companies in April 2024. Those and other cases were later consolidated.
The tech companies have argued that their actions constituted "fair use," which allows for some copyright material to be used without explicit permission, and that their use of the material was ultimately transformative.
Hecht's comments "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views," a Microsoft representative said. OpenAI didn't immediately respond to a request for comment.
The newly unsealed filing cites internal company documents, as well as testimony from Microsoft CEO Satya Nadella and comments from OpenAI President Greg Brockman and ChatGPT leader Nick Turley, among others. The details were included in a brief filed jointly by the news organizations to support their request that a federal judge rule in their favor ahead of a trial.
One major topic detailed in the filing was executives' thoughts on sidestepping the paywalls many publishers have in place.
After an OpenAI researcher informed Brockman about a hack to get around the New York Times paywall when scraping content, Brockman responded, "ah nice," according to the filing.
Meanwhile, Nadella testified that "anything that is paywalled should be licensed by anyone who wants to use it," for purposes including AI training. He said he would have had Microsoft force OpenAI to retrain its models, had he known they were scraping and training on paywalled information.
Nadella also testified that conversing with chatbots "substituted" for websites, "versus needing to go to the underlying source," according to the filing.
Users can now regularly get answers to their questions about major world events, stock market moves and recipes via AI-generated search results and chatbots, limiting the need to visit news sites directly.
The Microsoft representative said the company lays out in court filings why its use of the information is transformative and consistent with copyright law and "why Copilot is not a substitute for publishers' journalism."
The representative said Nadella had spoken about broad changes in how people find information, and "those observations should not be confused with conclusions about copyright questions before the Court."
Turley, the head of ChatGPT, used the word "substitutive," according to the filing, and said the AI products "will get more and more substitutive as they get better." He also wrote that publishers faced an "existential threat" from AI products.
OpenAI internal documents characterized ChatGPT as a "modern newsstand," according to the filing. And Anthropic CEO Dario Amodei, then a top OpenAI researcher, said in a presentation that "news generation" was a top skill of an earlier ChatGPT model, offering up a sample query of "What's the NYT saying today?"
Anthropic didn't immediately comment.
The Justice Department recently filed a statement of interest in the suit, arguing that training on news content is within the bounds of fair use under U.S. copyright law, and narrowing such carve-outs could hurt competition.
As AI advancements upend the media industry, many publishers have employed a dual strategy: striking lucrative content-licensing deals with select tech companies, while suing others for copyright infringement.
Wall Street Journal parent News Corp has content deals with OpenAI and Meta. Two News Corp subsidiaries have sued another AI company, Perplexity.
Comments