Unsealed NYT lawsuit records show Microsoft staffer called AI training data use the biggest labor theft in human history

Unsealed NYT lawsuit records show Microsoft staffer called AI training data use the biggest labor theft in human history

N
News Editor
2026-09-18 02:35:26
Newly unsealed court records in The New York Times copyright case against OpenAI and Microsoft show unusually blunt language inside Microsoft about how AI models were trained. Brent Hecht, Microsoft’s director of applied science, described the practice in an internal memo as an unprecedented and shocking theft, and at one point called it “the biggest labor theft in human history.” The filings also detail internal concerns over traffic loss tied to answer engines such as Copilot. Microsoft’s own data, cited in the case, said referral clicks to The New York Times domain fell as much as 93% compared with traditional Bing search. The documents listed declines of 83% to 93% for The New York Times and Daily News domains, and 51% to 94% for Ziff Davis domains. The case, now part of multidistrict litigation in the Southern District of New York, also includes plaintiffs such as New York Daily News, Center for Investigative Reporting, Ziff Davis, and several book authors. Court filings further describe OpenAI and Microsoft data-sharing arrangements, the scale of copied news content in training datasets, comments about paywalled material, and internal exchanges involving executives including Satya Nadella, Greg Brockman, Nick Turley, Nick Ryder, and Jack Clark. Microsoft said the comments reflected one employee’s personal views rather than legal analysis or company policy. OpenAI did not respond to a request for comment.

Newly unsealed filings in The New York Times copyright lawsuit against OpenAI and Microsoft show that Microsoft employees used unusually direct language in internal discussions about AI training data. In one memo cited in the case, Brent Hecht, Microsoft’s director of applied science, described the practice as an "unprecedented, shocking theft" and called it "the biggest labor theft in human history."

The case has been consolidated into multidistrict litigation in the U.S. District Court for the Southern District of New York. In addition to The New York Times, the plaintiffs include New York Daily News, the Center for Investigative Reporting, Ziff Davis, and several book authors. Motions for summary judgment are now under review.

Internal Microsoft records described a "doom loop"

According to the court filings, Hecht said in a January 2024 Microsoft presentation that traffic declines caused by Copilot created a "doom loop." He wrote that the effect would "simultaneously hurt our models and the broader web ecosystem."

Microsoft’s own data, as cited in the legal records, showed that answer engines such as Copilot reduced referral clicks to The New York Times domain by as much as 93% compared with traditional Bing search. The filings also said traffic to The New York Times and Daily News domains fell by 83% to 93%, while Ziff Davis domains saw declines of 51% to 94%.

Another Microsoft internal document stated: "It is highly unusual for a terminal product to threaten the economic foundation of its key suppliers, but that is exactly the situation we have created in the LLM business content supply chain." The same records said generative AI poses a "real risk" of "substantially impacting the employment" of people who produce the material used to train foundation models.

Testimony and internal messages were included in the filings

The complaint says Microsoft CEO Satya Nadella testified this year that "any content behind a paywall, anything anybody wants to use for grounding or training, should be licensed." Under oath, he also said that if he learned OpenAI had scraped and trained on paywalled information, he would "exercise [Microsoft’s rights] to demand that OpenAI retrain the models."

Nadella also agreed under oath that conversations with chatbots "have replaced ... directly giving you information on an AI platform without needing to go to the original source." He separately said downloading pirated material is "absolutely" illegal.

The filings also cite an exchange in which OpenAI researcher Nick Ryder told president Greg Brockman about a "hack to get around the NYT paywall." Brockman replied, "ah nice."

Separate internal communications from ChatGPT head Nick Turley said publishers face an "existential threat" from chatbots. The filings said these products are largely substitutive and "become more substitutive as they get better." Brockman was also quoted as saying the models are "really good at news."

Jack Clark, a former OpenAI policy director who is now a co-founder of Anthropic, was quoted as writing that the company’s work would increasingly lead to systems that replace the labor of people who define society’s "culture," adding: "Our work will make people unemployed."

Filings laid out dataset size and project names

On the scale of the data collection, the complaint said OpenAI’s mid-training dataset contained more than 91,692 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting. Another dataset derived from Common Crawl contained more than 2 million nytimes.com files.

The filings said some of the material was obtained by scraping directly from the Bing index. They also said OpenAI gave Microsoft the entire GPT-3 training dataset so Microsoft could evaluate how to integrate OpenAI models into its own commercial products.

The records further alleged that Microsoft, through projects code-named Project Taxi and Project Mango, sent training data back to OpenAI. The Project Mango dataset alone was said to contain at least 160,903 unique works from news publishers.

According to the filings, Microsoft deleted two training datasets in 2022 under a program called Project Clear because of legal concerns. The documents also said OpenAI relied disproportionately on scraped news content when building datasets such as WebText and deliberately removed copyright notices from training data because researchers "did not want the model to output" those notices to users.

Microsoft responded, OpenAI did not comment

Microsoft said the internal remarks "reflect the personal views of an employee, not legal analysis, and do not represent the company’s position." OpenAI did not respond to a request for comment.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
3200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.