AI Newsway

A Microsoft Memo Called AI Training 'the Largest Theft of Labor in Human History'

Unsealed material in the New York Times case puts internal Microsoft and OpenAI candor about paywalls, substitution and publisher harm on the record

|5 min read0
AI Summary
Unsealed portions of The New York Times' brief against OpenAI and Microsoft quote a January 2023 Microsoft memo calling AI training on scraped work the largest theft of labor in human history. The filing also alleges Copilot cut click-throughs to nytimes.com by up to 93%, that OpenAI datasets held more than 91,692 copies of the plaintiffs' works, and that staff discussed bypassing paywalls. The underlying exhibits remain sealed.
The New York Times copyright case against OpenAI and Microsoft has entered its third year, with newly unsealed filings quoting internal company documents.
The New York Times copyright case against OpenAI and Microsoft has entered its third year, with newly unsealed filings quoting internal company documents.

Newly unsealed material in the copyright suit The New York Times brought against OpenAI and Microsoft nearly three years ago contains something the companies' public filings never have: their own people describing the practice in the plainest possible terms. In a January 2023 internal memo, Microsoft's director of Applied Science, Brent Hecht, called AI training on scraped work "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The quotes surfaced Thursday and were first reported by TechCrunch.

One caveat frames everything below. Most of the newly visible language comes from the plaintiffs' own brief rather than the underlying exhibits, which remain sealed, so the excerpts arrive stripped of the context in which they were written. The allegations have not been tested at trial.

Key takeaways

  • A January 2023 Microsoft memo described AI training on scraped content as the largest theft of labor in human history, according to the unsealed brief.
  • Microsoft's own data allegedly showed Copilot cutting click-throughs to nytimes.com by as much as 93% versus traditional Bing search, which an internal presentation called a "doom loop."
  • The brief alleges OpenAI's mid-training datasets held more than 91,692 copies of works from the three news plaintiffs, and a separate joint dataset at least 160,903 unique works.

Why the internal language matters legally

Courts have generally been receptive to the argument that training a model on copyrighted text is fair use, and the Trump administration filed a brief earlier this month supporting OpenAI's unlicensed use of copyrighted material. Fair use, however, turns partly on whether the new product substitutes for the original in the market — and that is precisely where the unsealed statements cut against the defense.

OpenAI's head of ChatGPT, Nick Turley, is quoted describing an existential threat to publishers from products that are largely substitutive and, in his framing, will only become more so as they improve. President Greg Brockman is quoted calling the models excellent at news. Microsoft CEO Satya Nadella agreed under oath that talking to a chatbot supplies the information on the AI platform instead of sending a reader to the underlying source. None of that reads as transformation; it reads as replacement.

The traffic numbers attached to the filing sharpen the point. Microsoft's internal measurements allegedly showed that its Copilot answer engine reduced click-through rates for the Times' domain by as much as 93% compared with ordinary Bing search. A January 2024 presentation by Hecht described the dynamic as a doom loop that would damage both the models and the wider web at once.

It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its "content supply chain."

The scale and the shortcuts

The brief also puts numbers on the copying. It alleges that OpenAI's mid-training datasets alone hold more than 91,692 copies of works from the Times, the New York Daily News and the Center for Investigative Reporting, and that a Common Crawl-derived dataset contained upward of two million documents from nytimes.com. Data moved in both directions between the partners: OpenAI is said to have handed Microsoft the entire GPT-3 training set, while Microsoft supplied data through efforts named Project Taxi and Project Mango. The Mango dataset allegedly assembled at least 160,903 unique works from the news plaintiffs.

Some of the most damaging passages concern intent rather than volume. The filing describes OpenAI staff discussing a method to get past the Times' paywall without detection — researcher Nick Ryder raising it and Brockman responding approvingly — and alleges that copyright notices were deliberately stripped from training data so models would not reproduce them. Nadella testified that paywalled material should be licensed by anyone using it for training or grounding, and said that had he known OpenAI scraped behind paywalls, he would have invoked Microsoft's right to demand the models be retrained.

What happens next

Counsel for the New York Daily News said the evidence shows for the first time that both companies knew the practice was wrong. Neither Microsoft nor OpenAI responded to requests for comment. Whether the quotes survive contact with their sealed context is the question that will decide how much they are worth, and the exhibits themselves are still under seal.

The disclosure lands as the industry negotiates the same problem in public through consent standards rather than litigation — Microsoft among the companies that recently signed on to Cloudflare's crawler rules. What the filing adds is evidence that the economics of web scraping for large language models were understood internally, and described in unflattering terms, years before the public argument began.

FAQ

What exactly was unsealed?

Unredacted portions of a legal brief filed by The New York Times in its copyright case against OpenAI and Microsoft, which quotes internal memos, presentations and depositions. The underlying exhibits those quotes come from remain sealed, so the excerpts appear without their original context.

Does this settle the fair use question?

No. Judges have largely accepted that model training can qualify as fair use, and the case has not gone to trial. The unsealed statements matter because fair use weighs whether a product substitutes for the original work, and several quoted executives describe exactly that kind of substitution.

Have Microsoft and OpenAI responded?

Neither company returned requests for comment on the unsealed material. The allegations in the brief are the plaintiffs' characterization of the evidence and have not been adjudicated.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Judge Voids Pentagon Supply Chain Risk Label on Anthropic as Unlawful Retaliation
Tech & Business

Judge Voids Pentagon Supply Chain Risk Label on Anthropic as Unlawful Retaliation

A 59-page order found the Pentagon retaliated against Anthropic for criticizing the government, violating the First and Fifth Amendments.

Seung Jung21 days ago
OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months
Tech & Business

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeño posts its first InferenceX benchmark results.

Seung Jung23 days ago
New Mexico's Top Court Fines a Lawyer $5,000 Over ChatGPT's Fake Witnesses
Tech & Business

New Mexico's Top Court Fines a Lawyer $5,000 Over ChatGPT's Fake Witnesses

New Mexico's Supreme Court fined attorney Stephen Aarons $5,000 and held him in contempt after ChatGPT invented witnesses and police testimony in an appeal.

Seung Jung6 days ago
Washington's $1 ChatGPT Deal Expires. Its Replacement Bills by the Token.
Tech & Business

Washington's $1 ChatGPT Deal Expires. Its Replacement Bills by the Token.

OpenAI's new 27-month GSA deal waives a $15 per-seat license and halves token rates, but replaces a $1-a-year flat fee with metered billing from October 1.

Seung Jung5 days ago
Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.
Tech & Business

Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.

Antitrust enforcers have opened their first real examination of the deal structure that has replaced the acquisition in AI: Nvidia has received a formal Justice...

Seung Jung5 days ago
Altman Rules Out an OpenAI IPO in 2026 Even With a Confidential S-1 on File
Tech & Business

Altman Rules Out an OpenAI IPO in 2026 Even With a Confidential S-1 on File

OpenAI will not complete an initial public offering this year, chief executive Sam Altman said in an interview with Fortune editor in chief Alyson Shontell, ans...

Seung Jung5 days ago