For months, Microsoft and OpenAI have publicly described their use of the open web as a fair, transformative and largely benign act of training. Newly unsealed court documents in The New York Times' copyright lawsuit against the two companies suggest their own employees saw something considerably darker.
The filings, reported this week by The Verge, Ars Technica, Decrypt and Fast Company, capture internal debate over how large language models are built: by ingesting vast quantities of human-produced text, much of it journalism, without payment or permission. In those documents, Microsoft's own Director of Applied Science, Brent Hecht, warned that the practice was initiating a "doom loop" for the web, characterized the scraping as "the largest theft of labor in human history," and concluded that it made "a complete mockery of the idea of fair use."
"These comments ref…" — Microsoft spokesperson Alex Haurek's response to The Verge, which reported that the company has sought to distance itself from Hecht's assertions.
What the documents actually say
The material emerged through the discovery process in the Times' suit, filed in December 2023, which accuses OpenAI and Microsoft of building commercial products on the back of millions of copyrighted news articles. While the litigation itself is about money and licensing, the unsealed papers are notable for what they reveal about internal awareness — the degree to which the people building the technology understood the collateral damage it could cause.
The "theft of labor" phrasing is striking because of who used it. Hecht is not a critic or a litigant; he is a senior applied-science executive inside one of the two companies named in the suit. His warning that scraping would damage the web — the very corpus the industry depends on — is effectively a self-diagnosis delivered in private.
The 'doom loop,' explained
The fear Hecht articulated has a straightforward logic. AI companies scrape publishers' content to train models. Those same models then answer users' questions directly, so readers no longer click through to the original reporting. Traffic and advertising revenue fall. Newsrooms shrink or close. Fewer humans produce the high-quality text that made the models useful in the first place — forcing the next generation of systems to train on a web increasingly filled with synthetic, AI-generated material.
That cycle is the doom loop: a self-inflicted erosion of the ecosystem that feeds the industry. It is a concern that academics and publishing executives have raised for years, but seeing it spelled out by a Microsoft executive gives the argument a different weight in court.
An 'existential threat' to publishers
The documents reportedly go further, capturing OpenAI's own leadership acknowledging the stakes for the news business. According to the unsealed filings, the executive who oversees ChatGPT warned that publishers faced an "existential threat" — an internal admission that the product being built was positioned to disrupt the economics of the industry whose work it consumed.
Multiple outlets covering the filings framed the story around that tension: companies marketing themselves as partners to journalism while internal documents describe the relationship in terms of theft and threat. Decrypt focused on the raw language of the "largest theft of labor" line; Fast Company led with the doom-loop warning; Ars Technica and The Verge emphasized the fair-use mockery quote, which lands directly on one of the core legal defenses the defendants are expected to press.
The fair-use problem
That last point may prove the most legally consequential. OpenAI and Microsoft have broadly argued that training models on publicly available text is a transformative use protected under copyright law. If the companies' own personnel privately concluded that the practice was a "mockery" of fair use, plaintiffs will argue that the defense is a litigation posture rather than a genuine belief.
Microsoft has pushed back on that reading. Spokesperson Alex Haurek told The Verge that the company disputed the framing of Hecht's remarks, part of an effort to distance the corporation from an executive's characterization. The company has not, however, disputed that the documents are authentic — and the filings remain part of the record in an active case.
Why it matters beyond the courtroom
The disclosures arrive at a delicate moment. Regulators in the United States and Europe are weighing how copyright law should apply to generative AI; publishers are negotiating licensing deals from a position of weakness; and news organizations are experimenting with AI tools even as they sue over how their archives were used to build them.
- The legal front: The Times' case is one of several high-stakes copyright suits that could define what AI companies owe rights holders.
- The business front: Internal warnings of publisher collapse complicate claims that AI is a net positive for the media ecosystem.
- The technical front: A degraded, AI-saturated web threatens the quality of future training data — a risk the industry is now on record acknowledging.
None of this proves liability. Internal musings are not admissions, and companies routinely employ people whose job is to stress-test the ethics of their employer's strategy. But the documents do something subtler and, for the industry, more awkward: they show that the architects of the AI boom foresaw the damage early, wrote it down, and shipped anyway.



