AI’s Piracy Problem: You Can’t Build an Empire on Stolen Goods
This is part one of a four-part blog series about AI companies use of pirated copies of copyrighted works taken from illicit sources to train their AI models.
As generative artificial intelligence races from novelty to infrastructure, an uncomfortable truth sits beneath many of today’s most powerful AI models: they are built on pirated copyright-protected works taken from so-called shadow libraries. This practice is not a technical inevitability nor a legal gray area—it is a deliberate choice that treats copyright law, creative labor, and ethical norms as obstacles to be overrun rather than principles to be respected. When AI companies quietly ingest entire libraries of books, articles, music, and art scraped from illegal sources, they are not just training algorithms; they are normalizing large-scale theft, supporting illicit and harmful foreign piracy networks, and shifting the costs of innovation onto creators who never consented to participate.
Throughout U.S. history, there have been companies that have tried to build businesses on illegal goods and get away with it. In each instance, the company has failed in its attempt. But today’s AI companies believe that they are exempt from this historical precedent. They believe that they should be able to build their AI models from stolen goods in the form of pirated copyrighted works that they knowingly copied from so-called shadow libraries.
Over 130 cases have been brought in federal court against AI companies for copyright infringement and, in many of those cases, courts will decide whether AI companies are legally allowed to knowingly use pirated works to build their AI models. It will be some time before those decisions start coming down, but when they do, we will eventually learn whether history will repeat itself and AI companies will be liable for their transgressions (as was the case with Anthropic, which settled last summer) or whether AI companies get the special treatment they claim they are entitled to.
One of the biggest differences between historical attempts to build illicit businesses based on illegal goods and what AI companies are doing today is that these prior “businesses” knew what they were doing was illegal and were not brazen enough to attempt to persuade the public and the courts that their actions were reasonable, ethical, and legal. These illicit enterprises knew they were engaged in illegal acts and did their best to hide it. While today’s Big AI companies also try to conceal their illicit acts, they hedge their bets by using their billions to try to convince policymakers, judges, and others that their use of illegal goods to support the foundation of their AI models should be excused.
AI Companies That Scrape From Pirate Websites Are Today’s Digital Equivalent of the 1970s Chop Shop
Let’s look at some prior examples of firms that tried to build a business on illegal goods and get away with it. Back in the 70s, policymakers first confronted the issue of auto “chop shop” fencing operations. Car thieves would steal a car and then sell it and/or its parts to a chop shop, which would dismantle it and sell the parts for profit. There was no doubt that what these chop shops were doing was wrong, so legislators and law enforcement sprang into action to pass and enforce laws aimed at stopping these fencing operations. What AI companies are doing isn’t much different—they acquire massive repositories of stolen works and profit from the “parts” by using the copyrighted works contained therein for training.
Now consider a used‑parts dealer who keeps their business afloat by knowingly buying engines, airbags, and transmissions from these chop shops. A used‑parts dealer that keeps its prices low by buying engines and airbags from a chop shop is not a neutral middleman, as it is knowingly exploiting a criminal supply chain. The law has never excused that dealer simply because he didn’t steal the cars (or the parts) himself because knowingly relying on a criminal supply chain is illegal and unethical.
Regardless of whether big AI companies are more akin to the chop shop or the used-parts dealer, AI companies that train their models on pirated copyright-protected works scraped from pirate shadow libraries are engaging in the digital equivalent. They are not passively consuming publicly available data; they are deliberately seeking out and ingesting stolen copyrighted works from known illicit sources to fuel their products. Rebranding their transgressions as “algorithms” does not erase the fact that the underlying supply chain is illegal.
There are numerous other examples: Silk Road thrived on a steady supply of illegal drugs and forged documents. Backpage generated revenue from ads for illegal sex trafficking. Pioneer Import Corporation traded in conflict diamonds to be used as engagement rings. Megaupload profited from pirated movies and music it knew came from repeat infringers. Each of these examples demonstrates that a company cannot legitimize its business by transforming illegal inputs into a lawful-looking product or service, whether its stolen car parts used to repair vehicles, conflict diamonds to make engagement rings, or pirated copyrighted works used to make AI models. No amount of alleged “transformation” can cleanse the illegality of the supply chain.
Law enforcement shut down these operations and prosecuted their operators because knowingly relying on illegal source materials has never been tolerated as a legitimate business strategy. AI companies that train models on works scraped from pirate shadow libraries are attempting the same gambit in a digital guise, and the law, along with basic ethical norms, should treat it no differently.
Use of Illicitly Obtained Copyrighted Works is the Norm for Big AI
Unfortunately, the use of illicitly obtained inputs by AI companies is not uncommon. It is typical for most big AI companies to obtain copies of the copyrighted works they use to train their AI systems from so-called pirate shadow libraries, bypassing licensing markets and undermining the legal frameworks that govern creative development and distribution. AI companies’ use of pirated works copied from so-called shadow libraries is neither accidental nor harmless. These big AI companies purposefully cut corners by knowingly building their AI products by copying massive amounts of infringing material from pirate shadow libraries that were “constructed” in clear violation of the law. When infringement becomes a business strategy rather than an exception, policymakers and the courts must confront whether existing copyright law, and (as discussed in more detail in part II of the blog) the fair use defense in particular, is being distorted and misapplied.
This isn’t innovation operating at the edge of legality; it’s large-scale infringement laundered through technical jargon and venture capital. By ingesting stolen books, articles, images, and music without consent or compensation, big AI companies have made a calculated decision to weigh growth and market dominance over creators’ rights, the rule of law, and basic ethical restraint.
Should we reward such activities? Should we vest AI companies with special privileges that no other company in the history of our great country has been granted? Should we brush off the immense harm being caused to America’s creative community? Should we allow big AI companies to cut corners instead of using some of their billions to source legitimate goods?
We all know the uncomfortable truth: AI companies can’t build an empire on stolen goods.
To stay up to date with the latest news in artificial intelligence (AI) and copyright, sign up for our AI Copyright Alert. You can also visit our AI and Copyright hub for additional resources on federal court cases, current licensing, and more. Additionally, if you aren’t already a member of the Copyright Alliance, you can join today by completing our Creator Membership form! Members gain access to monthly newsletters, educational webinars, and so much more — all for free!
