Unsealed NYT filings quote a Microsoft researcher on "the theft of labor"
A January 2023 internal memo called AI scraping 'the largest theft of labor in human history'; OpenAI's own datasets held over 91,000 copies of Times material.
- Courts & copyright
- Notable
Judge Sidney Stein ordered the unsealing of hundreds of pages of internal documents in The New York Times’ copyright suit against OpenAI and Microsoft, revealing internal communications the companies had fought to keep confidential nearly three years into the litigation.
The most quoted document was a January 2023 memo by Brent Hecht, a Microsoft director of applied science, describing the scraping used to train AI models as “an astonishing theft of unprecedented proportions” and, in his words, “the largest theft of labor in human history.” A later Hecht memo, from January 2024, warned that Copilot risked becoming a “doom loop” for the publishers whose content trained it, writing that “it is highly unusual that an end-product threatens the economic foundations of its essential suppliers.” Microsoft told reporters the memos reflected Hecht’s personal opinion, not the company’s legal position, which continues to argue in court that training on the material is fair use. Elsewhere in the filings, Microsoft chief executive Satya Nadella testified in deposition that “anything that is paywalled should be licensed by anyone who wants to use it… for grounding or training,” and said he would have required retraining of models had he known paywalled Times content had been scraped without a licence.
The filings also quote OpenAI’s head of ChatGPT, Nick Turley, writing internally that publishers faced an “existential threat” from chatbot products that were “largely substitutive” for their journalism, and show OpenAI president Greg Brockman replying “ah nice” when a researcher raised a way around the Times paywall. On the scale of copying, the documents put the number of copies of Times, New York Daily News and Center for Investigative Reporting material in OpenAI’s mid-training datasets at more than 91,000, with a Common Crawl-derived dataset separately containing over two million documents from nytimes.com.
The exhibits are the plaintiffs’ selection and framing, made public through ordinary discovery rather than any new finding by the court, and OpenAI and Microsoft continue to dispute that the copying was unlawful. They followed the Justice Department’s statement of interest backing the training-is-fair-use position filed weeks earlier, and gave the Times fresh internal-communications evidence for the same fair-use dispute both sides had asked the court to resolve on summary judgment.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseJacob Coxon Warns of Human Extinction and Triggers a Preference Cascade
- 11 September 2026 · Zvi Mowshowitz · Don't Worry About the VaseThe Extinction Risk Preference Cascade: Quotes
- 24 September 2026 · Shakeel Hashim · TransformerHacking is the least worrying part of OpenAI’s Australia incident