Unsealed court documents show Microsoft and OpenAI personnel privately discussing whether AI scraping and chatbot substitution could undermine the news organizations whose reporting their systems depend on. (Source: Image by RR)

Internal Microsoft Document Warned of Potential Generative AI ‘Doom Loop’

Newly unsealed court documents in a copyright battle involving OpenAI and Microsoft reveal internal discussions about the potential impact of generative AI on news organizations whose content helps train and inform AI systems. According to a summary-judgment motion filed by news publishers led by The New York Times, Microsoft Director of Applied Science Brent Hecht described large-scale scraping of news for AI training as an extraordinary appropriation of human labor and questioned whether the practice could reasonably be considered fair use. Internal OpenAI messages cited by the publishers similarly acknowledged that AI products capable of answering questions directly could become substitutes for the news organizations producing the underlying information.

The documents also reveal concern inside the companies about a potential economic “doom loop.” One Microsoft document warned that generative AI could undermine the financial foundation of the very publishers supplying information that AI systems depend on, damaging both the web and eventually the models themselves. The publishers, as noted in an article at arstechnica.com, cited Microsoft data showing click-through declines ranging from 83 to 93 percent for some plaintiffs and 51 to 94 percent for others, while OpenAI employees reportedly discussed how users often have little reason to visit original sources once a chatbot has already supplied the requested information.

News organizations argue the internal communications strengthen their claim that OpenAI and Microsoft understood their products could substitute for publishers while using copyrighted journalism to build and operate those products. The filing also alleges OpenAI personnel discussed ways of circumventing The New York Times’ paywall and that the companies reproduced substantial portions of articles in response to certain prompts. The publishers are asking the court for summary judgment on a narrower group of articles where they say outputs demonstrate extensive verbatim overlap, while reserving other copyright claims for trial.

Microsoft rejects the publishers’ interpretation of the evidence. A spokesperson told Ars Technica that Hecht’s comments reflected one employee’s perspective rather than Microsoft’s position or a legal analysis, and the company maintains that its AI products represent transformative fair use rather than substitutes for news websites. The broader case could nevertheless influence how courts define the relationship between AI developers and the copyrighted material used to build and ground their systems, while raising a larger economic question: whether AI can sustainably depend on high-quality human-created information if the products built from that information reduce the revenue supporting its creation.

read more at arstechnica.com