Litigation vs. Licensing for AI Training

Scrabble tiles spelling 'LITIGATION vs LICENSING' on a game board.

Image: Author

There is an ongoing struggle between the tech world of AI training and the cultural world of content creation. It has led to lots of litigation but also an increasing number of licensing agreements, the obvious market solution. Litigation has helped convince AI companies to share some of the wealth by pursuing licensing. Yet the AI world continues to try to find ways to avoid the basic step of seeking permission from rightsholders for using their valuable content to create their products.

Anyone who has seen the striking graphic “Who is Suing Whom in AI”, created by the design website Information is Beautiful, will be struck by the enormity and breadth of the issue which is so cleverly displayed, with the big AI developers such as Perplexity, Anthropic, Meta, Google, Open AI, Midjourney, Cohere and others at the centre with the creators (every content entity from Conde Nast, Getty Images, Universal Music Group, CNN, Disney and Thomson Reuters to Elsevier, Dow Jones, New York Times and others) ranged around the periphery, a stunning visual encompassing more than 100 lawsuits in the United States. That graphic was up-to-date as of June 26 of this year. Since then, at least one more major lawsuit has been filed, by a group of textbook authors against Meta. The graphic does not include the first such case in Canada where a group of media organizations (Canadian Press, Torstar, The Globe and Mail, Postmedia and CBC/Radio-Canada) is suing OpenAI, or the Getty Images case in the UK, or indeed any cases outside the US. From this graphic, it would seem that to resolve the issue of how copyrighted content is going to be used in AI development and training, litigation is the inevitable route. But is it?

As far as I am aware, Information is Beautiful has not created a similar graphic to display the range of licensing deals that have taken place, many of them between some of the same actors that appear on the litigation chart. If they did it would be similar, but encompassing even more licensing agreements than lawsuits. Licensing deals are being struck so frequently it is just as hard to keep up with them as it is to track all the litigation underway. The University of Glasgow’s CREATe Centre says it has documented 274 licensing deals and has a chart that tracks 109 of them. Whatever the number, it is a lot and it is growing. That is not to say that the AI industry has finally accepted the need to pay for the content they are using to create their products, just as they pay for software engineers or data processing capacity. This is where the link between litigation and licensing becomes interesting.

In a perfect world, AI developers would obtain their inputs through the market on the basis of permission, which would encompass both compensation (in most cases) plus transparency or accountability, i.e. documenting what content was used. But we don’t live in a perfect world, which is why we have the rule of law and courts to enforce those laws. In some cases, AI platforms did begin negotiations with rightsholders but when it was not possible to reach an agreement, the AI industry switched tactics and took the content anyway, arguing it was legal to do so for a variety of reasons. This is precisely the scenario that led to the New York Times suing OpenAI. These cases are even more egregious because there was initially a tacit acknowledgement by the user that the content had value. Then, when the price or conditions did not suit the potential licencee, suddenly it was okay to take the content anyway under the guise of fair use. Various arguments have been deployed ranging from the claim that no copying actually occurs, to the dubious assertion that what is copied is data not content, to the invocation of the US “transformation” doctrine.

On the issue of copying, a study by the Atlantic (AI’s Memorization Crisis: Large language models don’t “learn”—they copy. And that could change everything for the tech industry) convincingly demonstrated the uncomfortable truth that LLMs can reproduce long excerpts from books they have been trained on. The inputs are not just ones and zeros, they are content— someone else’s content that was taken without permission. Whether the use was fair according to US fair use interpretations is still an open question. US courts and other countries are trying to come to grips with this issue. In countries such as Canada or Australia, where there is no statutory copyright exception for Text and Data Mining (TDM) that would permit permissionless AI training on content, the AI industry has been floating various workaround proposals. The “incentives” would include (in Australia) establishing a government-managed fund to compensate rightsholders according to some sort of formula, plus investments in AI data centres. What is missing from proposals such as this is the concept of permission from those who actually own the content, or even discussion of the proposal with them. As Prof. Rod Sims, former Chair of Australian Competition and Consumer Commission, put it in a recent opinion piece in Canada’s National Post, “what other sector refuses to negotiate with suppliers and instead goes to government to bypass such a step?”

Let me use a food industry analogy to make the point even more clearly. When you run a restaurant you have labour costs, rent, taxes, etc. and the cost of ingredients to consider. You don’t get to raid the farmer’s field to obtain your inputs for free, just because you are able to root out crops without the farmer being able to stop you or even know it is happening. Setting up a fund to “compensate” farmers for their stolen crops, on terms set by the government rather than the market, doesn’t even begin to make this right. Legalization of this theft would remove any possibility of litigation or legal protection, for the farmer—or for content owners. Litigation, while protracted, costly and potentially leading to uncertain outcomes, is nonetheless the stick that is needed to facilitate licensing.

The obvious route for the AI industry to take is to license the content they want to use. That may not seem as “efficient” as just taking it for free but with the threat of litigation hanging over the proceedings, licensing suddenly becomes the more efficient alternative. It is also win/win for both AI developers and the content industries. And, it is simply the “right thing to do”.

© Hugh Stephens, 2026. All Rights Reserved

I am pleased to note that this blog was recognized by Feedspot as being among the “40 Best Copyright Blogs to Follow in 2026”. In fact, we hit the middle of the pack at No. 20. I am honoured to be included in such distinguished company.  

Feedspot is an RSS Reader that lets readers subscribe to blogs, news sites, and any website they wish to follow.

Using Copyrighted Content to Train AI: Can Licensing Bridge the Gap?

Image: Shutterstock

The struggle between authors (writers, artists, musicians) and AI developers over the unauthorized and uncompensated use of copyrighted works to train AI applications continues, both in the courts (here is a summary of the current state of play in the US where most of the litigation is taking place) and in the political arena, such as the UK government’s latest initiative to put its thumb on the scale in favour of the AI industry, now slowed down by opposition within Parliament. The creative industries in Britain are still nervous, however, as demonstrated by the coordinated “Make it Fair” campaign organized by leading UK newspapers on February 25. While the courts may provide some guidance, it is unlikely to be dispositive and is almost certainly to be somewhat contradictory and lengthy, given the appeal process that will play out. With new applications being rolled out every day, AI appears to be unstoppable. Let’s accept that this is the case. If so, what then will be the rules governing the use of AI training content, particularly content that is protected by copyright, such as books, journalistic output, paintings, musical compositions etc.?

It is already apparent that at least some of the AI output trained on these materials will compete in the marketplace with the original works. If that is the case, then surely some of that additional value should be shared with those who helped create the content initially. The way this will most likely be done is through licensing in the form of payment and permission for use of the copyrighted creative output that enabled the training to take place. Licensing would also help resolve another potential issue, the possibility that the final product produced by the AI algorithm infringes on the copyright of the works on which it was trained. This is unlikely to happen in the case of written works but is certainly potentially possible with graphic or musical works.

While many have called for licensing as a solution, there are many challenges to be overcome to make it work effectively. Yet some licensing is already taking place between AI developers and owners of well delineated data sets. As an example, various newspaper and magazine publishers have already reached licensing agreements with AI providers. OpenAI has signed licensing deals with the Wall Street Journal, Times of London, the Financial Times, Time, Le Monde, Axel Springer and others. This is in marked contrast to OpenAI’s relationship with the New York Times, which has led to one of the most prominent lawsuits in the field, with the Times suing OpenAI for copyright infringement. The reason for this lawsuit, of course, is because licensing negotiations between the two entities broke down. Some photo and image licensing companies have concluded AI deals (Shutterstock is the most prominent example) while others, such as Getty Images have not. (Getty is suing StabilityAI in the UK). Eventually most of the institutional or corporate holders of valuable content in one form or another will likely reach, or attempt to reach, licensing deals with the major AI developers. But that still leaves out an awful lot of copyright-protected content.

The conundrum is how to deal with the millions of individual creators who produce content in different formats, and tie them into a workable licensing regime. The first challenge is how to even figure out who is producing content that is likely to be used by AI developers. The second is to calculate how much that use is worth. Then there is the challenge of how to administer a collective licensing scheme in a way that is both practical and affordable and where the small amount of royalties for individual works are not swamped by the administrative costs of collection and disbursement. Finally, there is the question of how to resolve the issue of competing licensing organizations in order to provide more or less one-stop-shopping for the AI industry.

It is worth noting that one-stop-shopping currently does not exist in any area of collective licensing. Different collectives represent creators in different fields so music, publishing, art, broadcasting and visual arts licensing are all represented by different organizations, in some cases with more than one collective in a given field. The Copyright Board of Canada lists 36 copyright collectives on its website. I haven’t seen a definitive list for the US but this university website lists about the same number.

Whereas users of music only have to deal with a handful of CMOs (collective management organizations), and users of text based content (online or offline) need only to acquire a reprographic license from the major licensing collectives for published works, such as the Copyright Clearance Center in the US or Access Copyright in Canada, AI developers access the full gamut of content. It will be challenging to make access easy for the AI development industry, a point developed by Dr. Pamela Samuelson of the University of California, Berkeley, well-known copyright scholar (and skeptic, let it be added). In her recent paper in the UCLA Law Review (“Fair Use Defenses in Disruptive Technology Cases”), Samuelson focuses primarily on the question of fair use—as suggested by the title—but also examines the issue of a collective licensing regime for generative AI development. She manages to raise just about every objection conceivable (see pp.80-86 of the document for more details);

generative AI uses all forms of content therefore the licence would have to be very broad
-an issue would arise as to whether content for training was used just once, or on repeat occasions
-it would be very difficult and costly to administer given that there could be literally billions of creators involved

-creators would get very little revenue; the bulk would go to the administering agencies, the CMOs
-it would be difficult to determine value and to set a price on each transaction
-what about orphan works?
-differing national regimes might create confusion; alternatively some countries might not require a licence payment, giving them an unfair advantage
-it would be unfair to startups since the incumbents have already scooped volumes of content without payment.

She notes that creators may lose out, but since AI will affect the livelihoods of so many others, this is not exclusively a copyright problem. Tough luck creators.

Clearly Dr. Samuelson is not in favour of a collective licensing regime for content appropriated by AI developers, yet despite her firehose of cold water, there are a number of promising developments in this area. For example, the Copyright Clearance Center (CCC) in the US recently announced it would provide AI re-use rights within its Annual Copyright Licenses, making the CCC’s licence “the first-ever collective licensing solution for the internal use of copyrighted materials in AI systems.” Note the caveat. While covering re-use of content for AI applications, the CCC makes it clear that;

The license enables participating rightsholders to fulfill the needs of companies that require an efficient way to legally acquire the rights to use copyrighted materials within AI systems for internal use.”

Not training. The Copyright Agency in Australia has done something very similar.

Starting from February 2025, Copyright Agency will extend its Annual Business Licence to cover staff of licensed businesses who include third party material in prompts for AI tools (and) copy and share outputs from AI tools with colleagues”.

However, it does not apply to AI training and does not allow capture of the content outside the business, such as by an externally provided AI tool.

Likewise, the Copyright Licensing Agency (CLA) in the UK issues a Text and Data Mining (TDM) Licence. The CLA’s website explains that TDM “is the process of transforming unstructured content into a structured format to analyse, extract and identify meaningful information and insights. By using TDM, organisations can harness the power of vast volumes of information and data, capturing and revealing key concepts, trends, and hidden relationships.” Sounds quite a bit like training generative AI, but it’s not.

CLA’s TDM licence extension includes rights covering use of published content for TDM purposes. This does not cover the use of content in training or prompting Generative AI models.

Canada’s equivalent CMO, Access Copyright, is actively examining the issue, as it notes in its new strategic plan for 2025-2028;

Like collective rights management organizations around the world, we will actively explore how we might enhance our corporate licence offerings to include uses related to AI, providing Canadian rights holders who wish to participate in the emerging market for AI licensing to do so, either in Canada or by virtue of reciprocal agreements with sister organizations.”

It is clear that these Reproduction Rights Organizations (CMOs by another name) are cautiously feeling their way forward to find the appropriate role for collective licensing. Meanwhile the private sector has not been sitting idly by. Forbes reports that so many content aggregation startups have been established that they have formed a Data Providers Alliance. Recently launched “Created by Humans” is another commercial entrant that is pitching itself to authors.

Take control of your work’s AI Rights and get compensated for its use by AI companies.”

As these new enterprises enter the market, it threatens to become quite crowded. Just as there are more and more AI companies, including new entrants like DeepSeek, a proliferation of new sector-specific content-aggregators will make licensing more challenging. If the CMOs wait too long, they will face entrenched competition. Not all these new aggregators will survive. In the end, AI developers will not subscribe to multiple content licensors; they will go with the ones that provide the broadest coverage. It will be a Darwinian selection process.

While this is happening, other countries are experimenting with the concept of extended collective licensing for AI content. This allows CMOs to grant licenses on behalf of both their members and non-members alike. An extended collective licence is not a compulsory licence but it could lead to such a system being established. Spain was first out of the gate, but has since pulled back after the proposed Royal Decree attracted the criticism from many rights holders that it would proscribe their options. Yet it is one of many solutions being tested.

The recently released US Copyright Office report “Identifying the Economic Implications of Artificial Intelligence for Copyright Policy” includes an extensive discussion of licensing possibilities, including examining the pros and cons of a new statutory blanket licence. This would need to include a provision excluding rightsholders (such as entities that have already reached licensing agreements with AI developers) who have the ability and wish to issue voluntary licences that generate greater remuneration than a statutory payout would earn. This raises thorny opt-in/opt-out issues. Compromises will be required, but the challenges are not insurmountable.

The trick is to devise a system that will capture as much content as possible while allowing some flexibility to rightsholders, allocating payments in way that is fair and efficient (the USCO paper suggests that revenues associated with a work could serve as a rough proxy for its relative value), at the same time minimizing administrative costs so that expenses do not exceed potential revenues for rightsholders holding limited content inventory. Can it be done?

Despite the many obstacles identified by Dr. Samuelson and others, I am convinced that in the end collective licensing for content used in AI development and applications will become as accepted as the collective licensing regimes for use of various forms of copyrighted content today. The way forward won’t be straightforward; there will be zigs and zags. The courts and legislatures will play a role, as will authors, publishers, and the AI developers themselves. But in the end we will get there. Licensing, including some form of collective licensing, is the inevitable bridge that will bring AI developers and copyright holders together.

© Hugh Stephens, 2025. All Rights Reserved.