A major legal battle over artificial intelligence, copyright and the future of journalism has taken a new turn after previously sealed court documents revealed internal concerns at Microsoft and OpenAI about how AI systems were being trained using news content.
The New York Times alleges that OpenAI used millions of news articles to develop its artificial intelligence models without obtaining permission from publishers. According to newly unsealed filings, more than 10 million articles were involved in the material examined by the plaintiffs, with a substantial portion reportedly originating from The New York Times.
Among the most striking details is an internal statement attributed to Brent Hecht, Microsoft’s director of applied science. According to the court filings, Hecht described the scale of the alleged copying as “an astonishing theft of unprecedented proportions” and potentially the “largest theft of labor in human history.”
The documents also indicate that some people inside the companies were concerned about a much broader problem: what happens to journalism when AI systems can provide information directly to users without requiring them to visit the original publisher.
OpenAI’s head of ChatGPT, Nick Turley, reportedly warned internally that publishers faced an “existential threat” from AI products and suggested that these products could become increasingly substitutive as the technology improved. Other internal discussions reportedly questioned whether AI-generated answers could reduce traffic to news websites.
The issue is particularly important because publishers depend heavily on readers visiting their websites. If people increasingly receive summaries, explanations and answers from AI systems instead of reading original reporting, publishers could potentially lose advertising revenue, subscriptions, engagement and opportunities to build relationships with their audiences.
The court filings also raise questions about the scale and methods used to collect training material. Reporting based on the unsealed documents says the datasets examined in the litigation contained large numbers of copies of works produced by news organizations. The filings further allege that OpenAI and Microsoft exchanged or used datasets containing material from publishers.
Microsoft has pushed back against the interpretation of the internal statements. A company spokesperson said Hecht’s comments represented his individual perspective rather than Microsoft’s official position. Microsoft continues to argue in court that its use of copyrighted material is transformative and protected by copyright law.
OpenAI and Microsoft have also argued that training AI models on copyrighted material can qualify as fair use because the systems transform the underlying information rather than simply reproducing the original articles. The publishers strongly dispute that argument, maintaining that the technology can compete directly with the original sources and undermine the economic value of journalism.
The dispute has grown beyond a disagreement between one newspaper and two technology companies. Other publishers have joined related litigation, making the case part of a much larger fight over who should benefit when artificial intelligence learns from human-created work.
The legal question could have consequences far beyond newsrooms. Writers, photographers, authors, musicians, researchers and other creators are watching closely because the outcome could help shape the rules governing how copyrighted work can be used to train increasingly powerful AI systems.
The US Department of Justice has also entered the debate. In September, the department filed a brief supporting Microsoft and OpenAI’s position in the case, arguing that the development of AI has implications for scientific progress, economic growth and national security.
The New York Times and other publishers are seeking a ruling in their favour without a full trial. The case is being considered by US District Judge Sidney H. Stein, who is reviewing the parties’ arguments on summary judgment.
At the heart of the dispute is a question that technology is forcing the legal system to answer: when an AI system learns from millions of pieces of human-created work, where should the line be drawn between technological innovation and the rights of the people who created that work?
The answer could influence how AI companies build their models, how publishers protect their content and how creators are compensated in the years ahead.
There is also a deeply human side to this story. Behind every article is a reporter who spent hours researching, interviewing, checking facts and putting information into words. Behind every photograph, investigation or book is someone’s time, experience and creative effort. AI may change how the world finds information, but the people who create that information still have to live from their work. The debate now is not simply about machines learning from data. It is about how technology and human creativity can exist together without making the people who create the original work invisible.
As this case moves forward, the decision could help define one of the most important boundaries of the AI era: how innovation can move forward while still recognizing the value of human-created knowledge.

