Fair Use Is Not a Shield AI Companies Can Use to Shelter Their Use of Pirate Copies: Analysis of Fair Use Factors 2-4

The following blog post is part three of a four-part series on the use of pirated copies of copyrighted works by AI companies and whether such use should be considered to be fair use. Previously, we discussed the history of companies being held liable for trying to make a business from using goods acquired from illicit supply chains and why AI companies should be treated no different than their predecessors (part I) and why fair use is not a shield for AI companies to use to shelter their use of pirate copies (part II). In part three of this blog series, we continue to analyze the remaining fair use factors.

Factor Two: The Nature of the Copyrighted Work

The second fair use factor involves a consideration of the nature of the underlying work, specifically whether it is more creative or more factual. The publication status of the copyrighted work is also considered as part of this factor. When the underlying copyrighted work is unpublished, the use of it is less likely to be a fair use.

What is considered to be published or unpublished under the copyright law is very different than the typical meaning attributed to these terms. Under the copyright law, Section 101 of the Copyright Act defines the act of “publication” as “the distribution of copies or phonorecords of a work to the public by sale or other transfer of ownership, or by rental, lease, or lending. The offering to distribute copies or phonorecords to a group of persons for purposes of further distribution, public performance, or public display, constitutes publication.” And clarifies that “a public performance or display of a work does not of itself constitute publication.” Countless creators post their copyright-protected songs, written material, videos, photographs, and other works of visual art to the internet with no intention to sell or transfer ownership, and without “purposes of further distribution, public performance, or public display.”

Therefore, it’s unlikely that the mere posting of the work online would constitute a publication of that work. As Section 1008.3(B) of the Copyright Office Compendium explains, “the Office does not consider a work to be published if it is merely displayed or performed online, unless the author or copyright owner clearly authorized the reproduction or distribution of that work, or clearly offered to distribute the work to a group of intermediaries for purposes of further distribution, public performance, or public display.” (emphasis added) Section 1008.3(C) of the Compendium is even more clear, stating that “[a] critical element of publication is that the distribution of copies or phonorecords to the public must be authorized by the copyright owner.” and adding that “[w]hile a fair use may be lawful, it is not considered an authorized reproduction or distribution that publishesthe copyright owner’s work. Similarly, an infringing reproduction or distribution does not constitute publication, even if the unauthorized copies or phonorecords are dispersed among large number of people.” (emphasis added)

Based on these statements, many of the pirated works on these illicit sites that AI companies are scraping could be unpublished works for purposes of a fair use analysis. While Section 107 makes clear that “[t]he fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the [fair use] factors,” the unpublished nature of the works being copied would undoubtedly weigh heavily against a finding of fair use.

Of course, the pirated copies of mainstream books, movies and sound recordings used by the AI companies have very likely been previously published. Even in those cases where the pirated works were “published,” it is not unreasonable to think that—similar to consideration of the unpublished and published nature of the work—courts may consider whether the work is an authorized or unauthorized copy of the work when evaluating the nature of the work. If they do so, then that should also weigh heavily against a finding of fair use.

While it does not appear that any court has yet to expressly take into account the unauthorized nature of a work under the second fair use factor, that may simply be due to the fact that no defendant has had the temerity to pursue a fair use defense when they are knowingly using pirated copies obtained from a criminal enterprise for their commercial purposes. There are instances where a defendant has pilfered a copy or two to try to make a fair use of a work, but that is very different from the generative AI training fact pattern—where millions of pirated copies are being knowingly downloaded by AI companies from criminal enterprises to train their AI systems. Thus, just like when the work used is a pirated copy obtained from a criminal enterprise, that should weigh against fair use.

Thus, just like when the copyrighted work is unpublished, the use is less likely to be a fair use when the work used is a pirated copy obtained from a criminal enterprise, that should also weigh against fair use.

Factor Three: The Amount Used in Relation to the Copyrighted Work as a Whole

The third fair use factor considers the amount of the copyrighted work that was used compared to the copyrighted work as a whole. This factor also considers the qualitative amount of the copyrighted work used. AI companies are copying entire works, so this factor would likely weigh against fair use unless the AI companies can establish that taking the whole pirated work was necessary in light of the purpose of the use.

In addition to the amount of any particular work that is being copied, another factor that could be considered here—especially in the class action cases—is the vast amounts of pirated works that are being scraped and copied in their entirety. In many instances, AI companies are not just copying a handful of works; they are of copying millions of works.

Ultimately, because the third fair use factor considers the amount taken and number of works used and not the nature of those works—being pirated or legal—factor three is unlikely to be impacted by the fact that AI companies are sourcing copyrighted works from shadow libraries.

Factor Four: The Effect of the Use on the Market and Value of the Copyrighted Work

The fourth factor, which is widely recognized as the most important of the fair use factors, requires the court to consider whether the defendant’s use harms the actual or potential markets or the value of the work. The court must also consider whether there may be market harm if the defendant’s use were to become widespread.

The existence of an AI licensing market weighs against a finding that the unauthorized use is excused by the fair use defense. The fact that licenses are available for the use of copyrighted material for AI ingestion, and the fact that countless licensing deals have been and continue to be struck between copyright owners and AI developers, demonstrates that there is an actual and potential licensing market for copyrighted works to be used for AI training.

As with the other three fair use factors, the question here is whether the fourth factor analysis differs at all when the copyrighted works used are pirate copies obtained from illicit sources rather than legitimate copies. The answer is a resounding “yes.”

As the Supreme Court recently explained in Warhol v. Goldsmith, “the problem of substitution [is]-copyright’s bête noire.” Obtaining a pirated copy instead of buying or licensing a copy is an exact substitute—it is the exact activity that fair use is not intended to allow. If the AI company was merely using a legitimate copy or two, they might argue that the copyright owner was getting paid something initially. While there is harm to the copyright owner’s market for the work(s) in that situation, that harm is less than when illegal copies are being used and the copyright owner is being completely cut out of the picture. Ratifying AI companies’ decision to source copyrighted works from illicit, pirate, and often criminal sources under the guise of fair use would destroy the AI training market and may also eviscerate the entire market for certain works.

In addition, it has also been alleged in many ongoing infringement cases that AI companies used BitTorrent technology to simultaneously upload pirated copies of copyrighted works at the same time they downloaded them. This type of redistribution and downstream dissemination of stolen works would undoubtedly compound the harms of piracy detailed above, further eviscerating the market for the original works and diminishing their value.

Other Fair Use Considerations

While federal courts must apply the four statutory fair-use factors outlined above, they sometimes consider additional factors that help them determine whether a use should qualify as a fair use. Additional considerations the courts often take into account in a fair use analysis are: (1) whether the alleged fair use is consistent with the purposes and goals of the copyright law; (2) whether the use is of public interest. These are very important considerations in the context of the AI companies’ ingestion and use of pirated works from illicit sources websites and services.

a. Goals of Copyright

There can be no doubt that allowing AI companies to ingest pirated works from criminal enterprises is wholly contrary to the purposes and goals of copyright. Piracy is the very antithesis of copyright. Nothing is more antithetical to the purpose and goals of copyright than piracy. These acts by their very nature are wholly contrary to the goals and purposes of copyright and must weigh very strongly against a finding of fair use.

As if knowingly scraping pirate sites to source their AI models with copyrighted works was not bad enough, AI companies are not only using works from these sites but, in some instances, they are also financially supporting them. In many cases, AI companies are keeping these criminal enterprises in business and allowing them to expand their operations by paying hundreds of thousands of dollars to get improved access to the pirated works. These payments will only make the criminal enterprises stronger and more accessible so that their use becomes even more widespread, causing exponentially more damage. Thus, when considering the use of pirated works from shadow libraries by AI companies, courts need to consider not only the harm to the AI market but also the harm caused by AI companies that financially support the operation of pirate sites.

b. Public Interest

Another consideration courts often take into account is the public interest. New generative AI technologies have the potential to advance the public interest and free speech. But is that potential a goal we should seek at any cost? There must be some line that an AI company steps over where the public benefits do not outweigh the costs. There can be little doubt that furthering and supporting piracy crosses that line.

If the courts allow pirate sites to prosper and AI models to dilute the market for copyrighted works under the false flag of fair use, the incentive to create will be dramatically reduced, resulting in fewer creators and fewer creative works for the public to enjoy. That most certainly is not in the public’s interest.

Furthermore, if there is no economic incentive for people to create and the number, type, and diversity of creations dwindles as a result, AI companies will have fewer creative works to ingest. With fewer works they will only be capable of regurgitating synthetic slop, which will eventually lead to model collapse. That is not beneficial to AI companies or the public.

Conclusion

It’s long been said that big tech’s motto is “move fast and break things.” AI companies are trying to bend fair use until it breaks. To claim that building their commercial models from millions of pirated copies qualifies as “fair use” boggles the mind. What AI companies are doing is borderline criminal activity in itself. That type of activity cannot possibly be permitted under the fair use doctrine. If it is, then the fair use doctrine is broken beyond repair.

Anthropic settled the class action brought against it for $1.5 billion after the judge in the case vociferously condemned Anthropic’s use of pirated works from illicit sites as “irredeemably infringing.” Anthropic likely did the same fair use analysis explained above and realized they were fighting a losing battle because the four fair use factors, along with additional considerations, likely weigh heavily against a finding of fair use. How long will it be before other courts and AI companies realize the same?


To stay up to date with the latest news in artificial intelligence (AI) and copyright, sign up for our AI Copyright Alert. You can also visit our AI and Copyright hub for additional resources on federal court cases, current licensing, and more. Additionally, if you aren’t already a member of the Copyright Alliance, you can join today by completing our Creator Membership form! Members gain access to monthly newsletters, educational webinars, and so much more — all for free!

get blog updates