TechNewsReel
Live

AI Training Hits Legal Wall as Berne Convention Defines Scraping as 'Reproduction'

A new legal analysis warns that fragmented national copyright exceptions are creating a governance crisis for generative AI developers.

TechNewsReel Newsroom · August 17, 2026

The training of generative AI models is colliding with foundational international copyright law, creating a precarious legal environment for developers operating across borders. A new analysis suggests that the act of scraping and storing copyrighted data for AI training constitutes a reproduction of work, triggering protections under one of the world's oldest intellectual property treaties.

In Roundtable #36, titled "Scraping By: Generative AI and the Limits of Cross-Border Governance," the Columbia Undergraduate Law Review examines the intersection of AI training and the Berne Convention. The piece asserts that when an AI model stores a copyrighted work as part of its training process, a reproduction has occurred under Article 9 of the Berne Convention. This classification places the technical process of data ingestion directly within the scope of international copyright regulation.

The Framework of Global Copyright

The Berne Convention serves as the primary international agreement on copyright, signed by more than 180 countries. At its core, the treaty establishes the "reproduction right," which grants authors the exclusive authority to permit or deny the copying of their original works. However, the treaty is not a monolithic set of rules; it allows member nations to implement their own specific exceptions under Article 9(2).

This flexibility has led to a fragmented global landscape. For example, the United States relies on the broad doctrine of "fair use" to determine if scraping is permissible, while the European Union has implemented specific text and data mining (TDM) exceptions. Because these exceptions vary by jurisdiction, the same act of scraping a dataset can be viewed as a legal innovation in one country and a copyright infringement in another.

The Governance Crisis

This discrepancy creates a significant "cross-border governance" crisis. Generative AI companies typically scrape data from a global internet, meaning their training sets inevitably include works protected by the laws of multiple different nations. When international treaties like the Berne Convention provide the baseline definition of reproduction but leave the exceptions to individual states, they cannot provide a uniform answer for a technology that ignores national borders.

For AI developers, this results in a state of profound legal uncertainty. A model trained on a global scale may be compliant with the laws of the country where the company is headquartered but remain vulnerable to litigation in every other jurisdiction where the scraped content originated. For content creators, the lack of a unified standard makes it nearly impossible to enforce their reproduction rights against foreign AI entities.

The Path Forward

As the industry matures, the tension between global data needs and national copyright sovereignty is expected to intensify. The legal community is now watching to see whether international bodies will attempt to harmonize AI-specific exceptions or if the burden of compliance will fall entirely on developers to filter training data by jurisdiction. Until a unified framework emerges, the boundary between permissible training and illegal reproduction remains a moving target.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.