Fair Use Is Not a Shield AI Companies Can Use to Shelter Their Use of Pirate Copies: Analysis of Fair Use Factor 1
The following blog is part two of a four-part series on the use of pirated copies of copyrighted works by AI companies and whether such use should be considered to be fair use.
In the first part of this blog series, we discussed the long U.S. history of companies being held liable for trying to build a business by using goods acquired from illicit supply chains and why AI companies should be treated no differently than their predecessors. In this blog, we consider the primary justification AI companies have used to attempt to absolve themselves of liability—copyright law’s fair use defense and, more specifically, whether the fair use defense gives AI companies a free pass to knowingly use pirated copies of copyrighted works from illicit sources (where “pirated copies” includes both illegal copies and illegally acquired copies).
Fair use is an affirmative defense often raised by a user of a copyrighted work who is accused of copyright infringement. When a use qualifies as a fair use under Section 107 of the Copyright Act, the user may use the copyrighted work without the copyright owner’s permission and without compensating the copyright owner for the use. There are no bright-line rules for determining fair use, since it is determined on a case-by-case basis. But Section 107 of the Copyright Act lists four factors that must be considered when deciding whether a use constitutes a fair use. These factors are:
- The purpose and character of the use, including whether such use is of a commercial nature or is for non-profit educational purposes;
- The nature of the copyrighted work;
- The amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
- The effect of the use upon the potential market for or value of the copyrighted work.
Each of the four factors must be considered and, while no one factor is dispositive, typically factor four is considered the most important (particularly in the context of a fair use analysis around the use of pirated works). So, this blog considers each of the four factors to determine whether the ingestion and use of pirated works from illicit sources by AI companies should qualify as a fair use.
Factor One: The Purpose and Character of the Use
In evaluating whether the first fair use factor weighs in favor of or against AI companies when they ingest pirated copyright-protected works to train their AI models, the court will consider many subfactors. Two of the most important subfactors are whether the use is commercial and whether the use is transformative. Neither of these subfactors is likely to be impacted by the fact that the sources of the inputs are pirated copies of copyrighted works.
But there are other subfactors within the first fair use factor that may be impacted by the use of pirated copies. One of those subfactors is whether the purpose justified the use and the other is whether the user acted in bad faith.
a. Justification
As the Supreme Court explained in its most recent fair use decision, Andy Warhol Found. for the Visual Arts, Inc. v. Goldsmith, when there is a commercial use that is similar in purpose as that of the original, “a particularly compelling justification is needed.” On the other hand, as the Tenth Circuit recently recognized in the Whyte Monkee Prods., LLC v. Netflix, Inc. case, when the purpose of a secondary use is completely distinct from the original work’s purpose, justification plays less of a role because it would be more inherently transformative and thus “advances the goals of copyright.”
Most cases of generative AI use will be commercial uses that are similar in purpose to the uses of the works being ingested.[1] Thus, the AI developer will need a compelling justification to use copyrighted works to train its AI model. Where those copyrighted works are not just copyrighted works but are pirated versions obtained unlawfully or from illicit sites, it’s hard to imagine a court finding a compelling justification for such use, especially when non-pirated copies of the works can be acquired through legitimate sources and means.
At least one AI developer has attempted to justify its copying by arguing that it requires the ingestion of massive amounts of content that cannot reasonably be obtained through legal means. That claim is patently false. There are many successful AI developers that do not indiscriminately scrape the internet, or pirate sites to train their models. Simply because some AI developers want to use everything to train their systems or find it is easier and cheaper to do so, doesn’t mean that it is necessary. The desire by AI companies to purposely scrape these pirate sites and services and avoid obtaining copyrighted works through legitimate means cannot be a “compelling justification” that would weigh in favor of fair use.
Legitimate versions of copyrighted works are available from copyright owners, book distributors, public and private libraries, and others. This availability undercuts any arguments that using pirated works for AI training is somehow necessary and justified. The fact that the copyrighted works that AI companies wish to use for training are largely available through legitimate sources contrasts starkly to past fair use cases where works were not licensable or otherwise available. In sum, sourcing mass amounts of pirated copyrighted works is not necessary for the development of generative AI models. It is simply the cheapest, most convenient, and quickest way for some AI companies to obtain the vast quantity of works they want. Thus, this lack of compelling justification should weigh heavily against a finding of fair use.
b. Bad Faith
Another subfactor that might impact the first factor fair use analysis when the works being used are pirated copies is whether and to what extent the user has acted in bad faith.
To be clear, there is no doubt that AI companies’ use of pirated works because it’s less expensive and easier than legally obtaining the works is, at the very least, an act of bad faith. These AI companies know that the copy of the work is illicit and illegal and that the sources of these pirated copies are not legitimate. Yet, they attempt to camouflage the multitude of wrongs and harms arising from participating in and perpetuating piratical activities and networks by arguing that their piratical uses are justified by the goal of AI training. This type of commercial deception is a form of bad faith that is worthy of some weight in a fair use analysis. In fact, as discussed later, it may be something that goes beyond mere bad faith. The real question that needs to be addressed here is whether and to what extent their actions, which constitute bad faith and likely something more nefarious, are considered in a fair use analysis.
When it comes to the role of bad faith in fair use analyses, court decisions in this area are all over the map. When it does get considered by courts, it tends to fall within the scope of the first fair use factor. The two most important Supreme Court cases to consider bad faith are Harper & Row v. Nation Enterprises and Oracle v Google. In Harper & Row, the type of bad faith contemplated by the Court was akin to egregious intentional copying for commercial gain and was something that should be considered in a fair use analysis. Egregious copying for commercial gain certainly sounds a lot like what AI companies are doing when they use pirated copies from illicit sites.
The Court in Oracle, was a little more amorphous, saying:
“As for bad faith, our decision in Campbell expressed some skepticism about whether bad faith has any role in a fair use analysis. 510 U. S., at 585, n. 18. We find this skepticism justifiable, as “[c]opyright is not a privilege reserved for the well-behaved.” Leval 1126. We have no occasion here to say whether good faith is as a general matter a helpful inquiry. We simply note that given the strength of the other factors pointing toward fair use and the jury finding in Google’s favor on hotly contested evidence, that factbound consideration is not determinative in this context.
It would seem from that comment that the Court is saying that bad faith is not dispositive of fair use, and that, depending on the facts of the case and how the other fair use factors are weighed, bad faith might play a role in a fair use analysis, but it did not in the Oracle case.
There are numerous district and appellate court decisions in which the defendants’ bad faith was taken into account during the fair use analysis. There are a few conclusions that can be drawn from these decisions, as well as those handed down by the Supreme Court, as to the application of bad faith in a fair use analysis. It is clear from these decisions, that an act of bad faith will not be dispositive of the fair use analysis. It is also clear that failure to seek permission or a license, standing alone, will likely not qualify as an act of bad faith, but that affirmative misconduct in obtaining or using the copyrighted work would likely qualify as an act of bad faith. This last point is particularly important because many bad faith cases involve acts of illegally obtaining a legal copy of a copyrighted work. But it is also important to understand that AI companies are not obtaining legal copies, they are obtaining illegal, pirated copies and perpetuating illegal, pirate networks and operations. That makes their acts much worse than an act of mere bad faith (more on that shortly).
In the Andy Warhol Foundation v. Goldsmith decision, in evaluating whether a use was a transformative use, the Supreme Court said that whether a secondary use has a further purpose or different character is a matter of degree, and that degree of difference must be balanced against the commercial nature of the use. It would appear that bad faith is also a matter of degree. At one end of the bad faith spectrum, are bad faith acts that are more akin to a moral wrong and that don’t cause any harms outside of an ethical/moral violation. That type of bad faith typically should not have much, if any, impact on a fair use analysis. At the other end of the bad faith spectrum, there are egregious acts done for commercial gain, which typically should have an impact on the fair use analysis while also not being dispositive.
There can be little doubt what side of the bad faith spectrum the AI companies fall into. In fact, a strong argument can be made that their actions are so egregious that they exceed mere bad faith (and, as discussed in more detail below, that is something the judges in two district court AI cases, Bartz and Kadrey, seem to agree with). When AI companies’ use of pirated copies obtained from illicit sites for commercial gain seems like an undertaking that is closer to a criminal act, then a court should not only consider that illicit activity in its fair use analysis, but it should have a significant impact on the first factor analysis.
To be found liable for criminal copyright infringement, which is codified in Section 506 of the Copyright Act, one element that must be proven is that the alleged infringer acted willfully. This is important in this context because (as noted in more detail in part I of this blog series), many AI companies have knowingly sourced their training material (i.e., the pirated copyrighted works) from criminal enterprises. Indeed, to date, there have only been two AI cases to reach the summary judgement stage and in both those cases the defendant AI companies (Meta and Anthropic) admitted to downloading vast amounts of pirated material from known illicit websites and services.
One of those illicit sources is Z-Library. Z-Library was one of the world’s largest online repositories of pirated literary works, hosting over 13 million pirated books and 84 million pirated articles. In November 2022, the FBI seized over 240 domains associated with Z-Library and charged two Russian nationals, Anton Napolsky and Valeriia Ermakova, with criminal copyright infringement, wire fraud, and money laundering.[2]
AI companies scrape numerous illicit sites that mirror Z-Library’s illicit activities. These include:
- Library Genesis (LibGen): One of the largest and most-used shadow libraries on the web with no download limits. Contains over 33 terabytes of books, scientific papers, and comics.
- Anna’s Archive: A shadow library that aggregates from Z-Library, Sci-Hub, and LibGen, boasting over 100 million files and functions as a search engine across multiple shadow library sources.
- Sci-Hub: Provides access to over 85.4 million scientific articles and research papers.
- PDF Drive: Hosts around 75-80 million pirated eBooks in PDF format with no download limits.
- Bibliotik: An invite-only private torrent tracker and shadow library that has been cited in major copyright infringement class-action lawsuits against AI companies for supplying unauthorized books used in training datasets. (Books3 originated from Bibliotik)
Many of these piracy operations and websites have been sued by copyright owners but often reside outside the jurisdiction of the U.S. court system. Anna’s Archive was sued by Spotify and UMG and a group of major book publishers for offering gob smacking amounts of pirated sound recordings and musical works for AI training. Those book publishers recently won a collective $19.5 million default judgment against Anna’s Archive. Sci-Hub was sued by Elsevier and the American Chemical Society, which both also won default judgments for million in damages, though the money was never collected as the operators remained outside U.S. legal jurisdiction. Library.nu (formerly Gigapedia) was shut down following a lawsuit from seventeen publishers. In sum, it has been alleged in dozens of AI infringement cases, and proven in many, that AI companies have sourced training material from scraping or downloading directly from the illicit sites listed above.
While state of mind of the user may rarely play a decisive role in the typical fair use analysis, what AI companies are doing is far from typical. The fact that AI companies’ intentional use of these illicit sites does not presumptively disqualify them from claiming fair use is antithetical to the foundations and goals of our copyright system and the rights it guarantees. When AI companies knowingly use massive amounts of pirated works obtained from illicit sources and criminal enterprises it’s not just an act of mere bad faith; it’s an act that borders on criminal infringement that should weigh heavily against fair use. Judge Alsup recognized this in Bartz v. Anthropic, finding that Anthropic’s sourcing of pirated works did not qualify as fair use and that “bad faith is not the basis for this decision.”
There are many key differences between past cases that have involved bad faith and the actions of AI companies in the current AI infringement cases, such as:
- Past bad faith cases involved the use of one or, at most, a few copies. In contrast, AI companies are copying millions of copyrighted works;
- Past bad faith cases involved the use of works from one or a few creators. In contrast, AI companies are copying works from millions of creators;
- Past bad faith cases involved the use of legitimate copies (that, in some instances, were obtained through questionable means). In contrast, in the AI infringement cases, the copies used are known pirated copies;
- Past bad faith cases often involved the defendant obtaining the copyrighted work through a legitimate third party. In contrast, AI companies are knowingly sourcing the copyrighted works from known criminal enterprises or otherwise obtaining lawful copies of works illicitly.
- Past bad faith cases involved copies that were often not otherwise available and the only way the user could obtain the copy was through conduct deemed to be in bad faith. In contrast, the copyrighted works scraped by the AI companies from illicit sources can typically be obtained through other means, such as from copyright owners, book distributors, private and public libraries, and so on.
AI companies’ knowing and intentional use of pirated copies obtained unlawfully or from illicit sources to cut corners and get a leg up on their AI competitors, while causing massive harm to creators, is no more a mere act of bad faith than (as noted in part 1 of this blog series) stealing car to sell the parts to a chop shop (i.e. they acquire massive repositories of stolen works and profit from the “parts” by using the copyrighted works contained therein for training). What these acts have in common is that most people agree they are wrong by their very nature. What AI companies are doing is not only in bad faith, but it is also inherently wrong and borderline criminal activity. To simply shrug it off as merely an act of bad faith makes little sense.
Judge Alsup acknowledged this in Bartz v Anthropic. In denying Anthropic’s fair use defense for scraping pirated works from illicit sites, he specifically states that “bad faith is not the basis for this decision” and finds that “the person who copies the textbook from a pirate site has infringed already, full stop” and “can[not] be excused as fair use merely because some will eventually be used to train LLMs.” He goes on to say that “piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.” (emphasis added) and “[h]ere, piracy was the point: To build a central library that one could have paid for, just as Anthropic later did, but without paying for it.”
In Kadrey v. Meta, which was decided in the same district court two days later,Judge Chhabria used somewhat similar reasoning in his analysis but, because he did not have the necessary evidence before him, reached a different outcome. After correctly noting that “[t]he law is in flux about whether bad faith is relevant to fair use,” he seems to find that Meta’s acts may qualify as something beyond any bad faith consideration, stating that:
“…downloading copyrighted material from shadow libraries would be relevant if it benefitted those who created the libraries and thus supported and perpetuated their unauthorized copying and distribution of copyrighted works. In the vast majority of cases, this sort of peer-to-peer file-sharing will constitute copyright infringement. … So if Meta’s act of downloading propped up these libraries or perpetuated their unlawful activities—for instance, if they got ad revenue from Meta’s visits to their websites—then that could affect the “character” of Meta’s use. But the plaintiffs have not submitted any evidence about this.”
The evidence that Chhabria was missing from this case can be found directly from one of Anna’s Archive’s mirror sites (PiLiMi (Piracy Library Mirror))—one of the sources that Anthropic used—thanks to the research of Ed Newton-Rex detailed in his LinkedIn post. In late 2025, Newton-Rex “was curious to know how much Anna’s Archive was charging AI developers for access to their massive library of pirated works for training.” So, he emailed them with his interest in buying access. They replied back to say that they are charging $200,000 for high-speed access to the full, pirated collection, which includes more than 60 million books. Newton Rex goes on to note that training on pirated works is rife in the AI industry and has been embraced by some of the biggest AI companies. “When they claim what they’re doing is fair [use], it’s worth remembering that not only is it theft – it supports further theft by funding pirates.”
Here is the actual response Ed received:
How’s that for evidence! Pretty good, but not perfect. Better evidence would be to get a list of all the AI companies that paid Anna’s Archive and/or the other illicit sites. But because Anna’s Archive, Z-Library, and most other illicit sites are located outside the United States it will be very difficult, if not impossible, to get this kind of information. Thus, likely, the only way to get that evidence is directly from the AI companies through discovery taking place in the more than 130 pending lawsuits.
With so many lawsuits against AI companies it’s likely that this information will start to get out eventually. When that happens, we may begin to see other courts follow the lead of Judges Alsup and Chhabria and find that AI companies’ knowing use of pirated works from known illicit sources is not bad faith, but rather borderline criminal activity that is beyond bad faith considerations. When this illicit use is considered by the courts along with the lack of compelling justification and the commercial nature of the use, it’s likely the first fair use factor will weigh heavily against a finding of fair use regardless of whether the court finds the use to be transformative or not.
The analysis of the first fair use factor was “fairly” long, so we split the fair use analysis into two blog posts. Stay tuned for the next blog post, in which we examine factors 2-4, and more.
[1] For example, compare the purpose of a music model and the purpose of the ingested music. The purpose of a human creating music is for the end user to listen to and enjoy the music for all of the myriad reasons that humans seek out recorded music. Now consider what the purpose of ingesting the copyrighted works is to the AI developer. The AI developer’s purpose is to train the AI to generate music that the end user can listen to and enjoy. Thus, the purposes of the ingested works and the AI-generated outputs are the same.
[2] The criminal case remains pending resolution because the defendants avoided extradition after escaping house arrest.
