The Day Generative AI Creates a "Doom Loop" That Reduces Click-Through Rates by Up to 93% and Erases News

The Day Generative AI Creates a "Doom Loop" That Reduces Click-Through Rates by Up to 93% and Erases News

When you ask a generative AI a question, it provides a well-structured answer in seconds. Behind this, there is a vast amount of work accumulated by journalists conducting interviews, editors verifying information, photographers, researchers, writers, and programmers.

If AI is trained by collecting this work without permission, and the finished AI takes away access from the original creators, is this "learning" or "theft"?

In a copyright lawsuit filed by media outlets like The New York Times against OpenAI and Microsoft, parts of the plaintiff's motion for summary judgment, which had been redacted, were made public on September 17, 2026. The document contained harsh words not from external critics but from people within the AI development industry.

Brent Hecht, Director of Applied Science at Microsoft, reportedly described the act of large-scale models absorbing the achievements of people worldwide as unprecedented theft, potentially "the largest theft of labor in human history," in an internal document. Nick Turley, who leads ChatGPT at OpenAI, also recognized that products like chatbots pose an "existential threat" to publishers and are replacing existing services.

These are strong words. However, there is an important point to note first. The documents released this time primarily consist of legal claims constructed by the media side, citing evidence. Some of the source evidence remains sealed, making it impossible to fully verify the context of the statements from the outside. Therefore, these should be read as the plaintiff's claims and internal statements cited therein, not as established facts recognized by the court.

Nevertheless, the significance of this disclosure lies in the possibility that there were completely different concerns inside and outside AI companies.


The issue is not just about "copying"

OpenAI and Microsoft have argued that using copyrighted works for AI learning is not about selling the original articles but about creating new functions, which constitutes "transformative use" under U.S. copyright law's fair use doctrine.

However, in determining fair use, not only the purpose of use and the nature and amount of the copyrighted work used are considered, but also whether it encroaches on the market for the original work. Here, internal statements carry weight.

According to the plaintiff's documents, Turley noted that AI products have already become substitutes for news, and the better their performance, the stronger their substitutability. Microsoft's CEO Satya Nadella reportedly acknowledged in sworn testimony that chatbots have replaced websites in the sense that users obtain information directly on the AI screen without visiting the original sites.

Microsoft explained to Reuters that Nadella's statement was about the broader principle of changing how information is searched and consumed, not a conclusion on the copyright issues in the lawsuit. They also stated that Hecht's expression was a personal view of an employee, not a legal analysis or the company's view.

In other words, harsh internal expressions do not immediately constitute an admission of illegality. However, for the plaintiffs, it serves as evidence that the defendant companies themselves foresaw "substitution" and "damage to the market." The focus of the trial is shifting from the mere fact that new technology was created through learning to how much the finished product has replaced the original reporting.


Up to 93% decrease—Shift from "search" to "answer"

The most concrete figures relate to changes in click-through rates.

The plaintiff's documents cite Microsoft's data, claiming that in Copilot's answer-type searches, click-through rates to New York Times-related sites decreased by 87-93%, Daily News-related sites by 83-91%, and Ziff Davis-related sites by 51-94% compared to traditional Bing searches.

Traditional search engines served as gateways directing users to information sources. Search results displayed headlines and brief descriptions, and those wanting more details would click the links. Sites would earn advertising revenue or subscription contracts from these visits.

In contrast, answer-type AI reads multiple sources, reorganizes them into the desired form for the user, and presents conclusions on the spot. The convenience centers on "not having to follow links." What is an advantage for AI services becomes a loss of revenue opportunities for information sources.

Microsoft's internal documents reportedly perceived this structure as a "doom loop." AI uses articles. Users stop visiting the original sites. Publishers' revenues decrease, reducing their ability to invest in reporting and editing. The amount of high-quality new information decreases, and eventually, even the reliable information AI can reference diminishes.

This is not just a battle between AI and the media industry. The "knowledge supply chain" that supports AI's performance is weakening its own upstream. Today's AI appears smart because of the wealth of information humans have published until yesterday. There is no guarantee that the same quality of information will be produced tomorrow.


Millions of articles and two joint projects

The plaintiffs also presented new figures regarding the scale of copying.

According to the disclosed documents, OpenAI's intermediate learning dataset contained over 91,692 works from the plaintiffs, and a dataset derived from Common Crawl contained over 2.06 million documents from nytimes.com. The initiative where Microsoft provided Bing index data to OpenAI was called "Project Taxi," and the joint effort by both companies to collect web documents was called "Project Mango." It is claimed that the learning data derived from Mango included at least 160,903 works from the plaintiffs.

Furthermore, there is a record cited where OpenAI researchers communicated a method to bypass The New York Times' paywall, to which co-founder Greg Brockman reportedly responded positively. There are also claims that copyright notices were systematically removed from the learning data.

Nadella reportedly testified that a license would be necessary to use information within a paywall for learning or as a basis for responses, and if he had known OpenAI used paid articles for learning, he would have demanded retraining using contractual rights.

There are multiple issues intertwined here. Learning from publicly available web texts, collecting against terms of use or robots.txt, bypassing paywalls, repurposing datasets with licensing conditions for commercial learning, and reproducing article expressions during responses are not the same legal or ethical issues.

The core of this case cannot be captured by the binary choice of "all AI learning is theft" or "learning is the same as human reading and thus all permissible." It is necessary to examine step by step what data was obtained, how it was used, and what impact it had on the original market.


Outrage erupted on social media and caution against oversimplification

The report spread rapidly on social media. On X, Jason Kint, a representative of a digital media industry group, pointed out that the unredacted documents were "eye-opening," suggesting that company insiders practically wrote the headlines for the plaintiffs. Ed Newton-Rex, who opposes unauthorized AI learning, criticized that internal statements directly conflict with AI companies' defenses.

 

In the technology community on Reddit, there was notable anger with comments like "Will they remove stolen data from Copilot?" "AI is made from everyone's work, so profits should be returned to society," and "There's too much difference in how large corporations' unauthorized use and individual theft are treated." There were also voices questioning the double standards of AI companies protecting their own products' intellectual property while claiming broad fair use for learning data.

On the other hand, there was backlash against the expression "the largest in human history," arguing that it is exaggerated and should not be compared to historical forced labor like slavery. Additionally, some opinions noted that the current copyright system itself hinders the sharing of knowledge, and equating learning with reproduction could stifle open-source and research activities.

These are representative reactions from some communities, not public opinion polls. Nonetheless, the axes of the debate are visible. People are not just angry about the possibility of articles being copied. They are upset about the structure where enormous value is created using countless human labor as raw material, with the profits and decision-making power concentrated in a few companies.


This is not just a distant issue for Japan

The same problem is approaching Japanese publishers, newspapers, creators, and web operators.

High-quality Japanese learning data is less abundant than in English-speaking regions. Therefore, carefully researched and edited Japanese articles and expert explanations are highly valuable. If AI search becomes widespread and site traffic significantly decreases, not only major companies but also regional media, specialized media, and individually operated sites could be hit first.

On the other hand, generative AI provides powerful production tools to small businesses and individuals. It offers benefits in translation, summarization, research assistance, programming, video and music production, and more, which should not be denied. The important thing is not whether to stop or advance AI, but how to distribute compensation and choice between those who create knowledge and the companies providing AI.

Possible mechanisms include clear licenses for learning and search use, distribution according to usage volume, designs that encourage actual visits to sources, transparency of learning data, effective means of refusal, and maintaining copyright notices and source information. If individual article contracts are impractical, centralized management like music copyrights or industry-wide standard contracts are worth considering.

However, if the licensing system favors only large companies, small media and individual creators will be left behind again. If open knowledge and non-profit research are overly restricted, it will harm innovation and freedom of expression. What is needed is a system design that carefully distinguishes between unauthorized mass acquisition and public research, commercial services that replace original works, and citation and criticism.


The question is not about the future of AI but about a future where humans can continue to create

The conclusion of this lawsuit has not yet been reached. It is unknown how the court will judge fair use or how much weight will be given to internal statements. The U.S. government has submitted an opinion supporting OpenAI, emphasizing the transformative nature of AI learning, and there are strong policy arguments focusing on industrial competitiveness and market entry.

However, the reason the phrase "the largest theft of labor in human history" attracted attention is clear. It brought the debate over generative AI back from abstract terms like data and copyright to the time and labor of the humans who created that data.

AI does not create text from nothing. Journalists go to the field, researchers conduct experiments, writers revise drafts, developers write code, and countless people share knowledge on the web. If the tools created using that accumulation also take away the revenue that sustains the original work, it cannot be said to be sustainable based on convenience alone.

The real question is not how smart AI can become. It is whether we can leave behind a society where humans can continue to create new knowledge, journalism, and culture even after AI has spread.



Source URL

※The descriptions in the disclosed documents are the plaintiff's claims, and some of the cited evidence remains undisclosed. The illegality of the defendants has not been confirmed by judicial judgment.