A Microsoft top executive called OpenAI’s data collection the biggest theft of labor in human history
Declassified court documents made public as part of a lawsuit brought by The New York Times reveal that executives at Microsoft and OpenAI privately described training large language models on news content in far harsher terms than their public statements; the newspaper’s 2023 lawsuit against OpenAI and Microsoft produced the correspondence now available to the public.
Among the most pointed lines, Microsoft’s Director of Applied Science, Dr. Brent Hecht, wrote that OpenAI’s methods represented “the largest theft of labor in human history,” and warned they could create a “doom loop.” OpenAI’s top executive, Nick Terly, described the practices as an “existential threat to publishers,” and elsewhere called AI products “largely substitutive” with respect to journalism.
The partially released documents include not only these quoted reactions but also technical descriptions of how training data were assembled. According to the reporting, OpenAI and partners circumvented paid subscriptions on publisher websites, compiled training datasets drawn from millions of documents, and removed copyright notices from those materials. In one exchange, an employee alerted OpenAI president Greg Brockman to a new technique for bypassing paywalls; Brockman’s apparent response recorded in the files was “ah, great.”
Other internal communications in 2023 captured concern about the broader impact on creators and the news ecosystem. An internal Microsoft document warned that “millions of people worldwide will soon feel that large models, 'sucking up' all their work, are committing a staggering theft of unprecedented scale.” Hecht also described large models as “a product that destroys its own supply chain.” An OpenAI engineer cautioned colleagues that “no matter how prominently we show links, users won't click on them,” a point Terly echoed in his remarks about substitution of journalism.
Legal experts note that several related lawsuits have already been decided in favor of AI companies, but judges in those cases emphasized the absence of a settled legal framework for generative models. The New York Times’ suit is widely viewed as a potentially pivotal test of whether training neural networks on copyrighted news content can qualify as “fair use,” the doctrine that can permit use of protected material without permission in certain contexts such as parody or commentary.
Microsoft has sought to distance the company from the tone of some employees’ messages. The company pointed to its filings, saying: “Microsoft's position is set out in its court documents, explaining why these transformative uses comply with copyright law and why Copilot does not replace publishers' journalism.”
Only selected fragments of the correspondence have been made public so far, and many of the released passages lack surrounding context, the reporting notes. The documents nonetheless provide a rare look into internal assessments of both the technical practices used to gather training data and the potential consequences for publishers and other content creators.