AI Training and Copyright: It’s a Global Issue—and There is a Global Solution

Infographic titled 'AI Training & Copyright: How Jurisdictions Approach AI Training and Copyrighted Works' detailing various jurisdictions including the USA, European Union, Japan, South Korea, China, United Kingdom, Canada, and Australia. It summarizes their legal frameworks, approaches to AI training with copyrighted works, and key points regarding copyright considerations.

Image: Used with permission (Maryna Hryhorieva)

This striking graphic, created by IP lawyer Maryna Hryhorieva, is a great summation of the current state of play regarding the still very fluid issue of AI training and copyright. It compares how eight different regimes, the US, EU, Japan, South Korea, China, the UK, Canada and Australia, are addressing this thorny issue, and highlights the different frameworks used to determine what is legal and what’s not. Each is different with its own peculiarities. Navigating through them is not easy. What may apply in one jurisdiction will not apply in another. A use that is acceptable in Country A may result in an unfavourable court ruling in Country B. An AI company can decide to base its content use practices on the most AI-friendly regime in which it operates, staying onside in that jurisdiction, but then can still be found to be infringing elsewhere. Cross-border jeopardy is almost certain where an AI company wants to scoop up local cultural content, pitting it against determined rightsholders who will, among other things, invoke cultural sovereignty.

The essence of the issue goes back to the way in which international copyright law is structured. Each country sets its own rules, regulations and interpretations within a broad commonly agreed framework. That framework, based on the Berne Convention, results in common basic standards centred on some fundamental principles, incarnated in the famous “three step test”.[i] The result is a generally consistent pattern of protection for rightsholders but with some significant differences by jurisdiction, like duration of copyright protection for example. The same inconsistencies apply to data mining (or what I would call unauthorized content reproduction) for AI training. So, what is a poor little ol’ trillion-dollar AI developer to do? Guess what? There is a universally applicable solution. Can you guess what it is? I reveal all in the final paragraph.

First, let’s briefly summarize the chart.

The US approach is based squarely on litigation, letting the courts decide the parameters of fair use regarding the unauthorized taking of protected content for AI training. This case-by-case approach will provide case-by-case answers, depending on the circumstances, the judge in question and even the circuit within the US federal justice system in which the case is heard. A legal approach inevitably leads to risk and uncertainty. That is not to say that in the end, the US will not resort to legislation. Thee US Justice Department recently submitted a “Statement of Interest” in the OpenAI v New York Times case, stating that a ruling in favour of the Times and other news industry plaintiffs would “threaten US national security, give a competitive advantage to foreign adversaries”, and thwart “creative and scientific progress while hindering American prosperity”. The Statement of Interest tries to put its thumb on the scale of justice, weighing in in support of the OpenAI argument that its use of NYT content was “transformational”, and thus constituted fair use. The Court is under no obligation to accept this interpretation; its duty is to interpret the law as it stands. But if the DOJ statement represents the Trump Administration’s interpretation of the law, and that interpretation is not upheld by the courts, that potentially opens to door to legislative change, assuming that Congress can be convinced to act in the name of big tech. Either way, there is continued uncertainty.

The EU approach is based on a narrow text and data mining exception embedded in the Copyright Directive. Resort to the exception is limited in two ways; if rightsholders cannot opt-out, the exception is limited to scientific organizations and cultural heritage institutions for research purposes only. There is also a general exception that allows for commercial use but in this case, rightsholders can reserve their rights by opting out.

Japan is considered by many to have a very AI-friendly regime in terms of allowing copying for training purposes, but it is not as wide as many have claimed, as I pointed out in this blog posting a couple of years ago. South Korea is struggling with the issue, being whipsawed between AI developers promising the moon and its vibrant cultural industry. China likewise is still trying to determine which way to go. It does not have a text and data mining (TDM) in its law, nor does it apply a fair use doctrine. However, it has announced plans to deal with AI training rules through its copyright development plan covering the period 2026-30.

The UK has had several false starts. A TDM exception limited to non-commercial purposes only has been on the books for several years, but the Labour government tried to expand this by advocating for a regime that would allow AI developers to freely use copyrighted content for algorithm training unless rightsholders opted out. After a massive outcry and pressure from British cultural industries and prominent performers, the government retreated. Now Britain has gone back to the drawing board.

In the case of Canada, everything is up in the air. Canada has a fair dealing regime with no TDM exception. It has held several public consultations and is developing an AI strategy but has been very unclear about where the government comes down on the issue of freely using copyrighted material for AI training. Meanwhile several Canadian media outlets are suing OpenAI for copyright infringement. OpenAI lost its argument that Canadian courts had no jurisdiction in this case. Now the Canadian music collective SOCAN is suing AI music company Suno. While AI companies are arguing for a wide TDM exception in Canadian law, rightsholders are seeking to retain the current law and use it to protect their rights.

Australia on the other hand has spoken out strongly in defence of its cultural sector and has declared that it will not legislate a TDM exception. Australian Prime Minister Albanese has taken that a step further declaring that, “An artist’s creative endeavour is their work and their property. No company should use Australian books, music, art or news to build or train AI without the artist’s control. That includes the artist’s control of the price and value of their work.” What will happen next in Australia is unclear although some tech platforms continue to try to push various compulsory licensing regimes.

As you can see, the rules are diffuse and in flux. Frankly, it is all over the map. One thing seems to be clear. The push by the tech industry for widespread TDM exceptions seems to be stalling. The litigation route so favoured in the US is yielding uncertain results. Legal actions in fair dealing countries like Canada, the UK and Australia may bring varying results with complex rules governing applicability of judgements. All this leads to greater uncertainty, and costs. However, as mentioned above, there is one universal remedy, one that transcends all the jurisdictional issues and conflicting interpretations. It is called voluntary licensing.

Voluntary licensing between rightsholders and AI developers eliminates the jurisdictional problems, assuming the agreements are structured to provide wide coverage and indemnity. With a voluntary licence, it doesn’t matter if there is or isn’t a text and data mining exception, or if the use is for commercial purposes, or if the use is sufficiently transformational. It doesn’t matter whether the AI developer is truly “enjoying” (accessing the essence of the content), as in the case of Japan. Some jurisdictions have toyed with the idea of a compulsory licence, as in India, but not only has this been opposed by both rightsholders and the AI industry, it is jurisdictionally bound. Voluntary licences deal with this issue and, if properly structured, remove the hazard of conflicting jurisdictional rulings.

While it is not always easy to reach such deals, as the OpenAI v New York Times case illustrates, it is surely better and less costly than the alternative of endless litigation. This is why it is becoming the preferred solution. More and more licensing agreements are being negotiated across the full range of copyrighted content, including audio-visual, music and publishing industries for AI training and use. Voluntary licensing is a global solution–but it requires AI companies to recognize the value of the content they want for training, and to accept the right of creators to control the use of their content. If it requires a willing buyer, it also requires a willing seller and while not all rightsholders will be willing to sign a licensing agreement, the holdouts will be the exceptions if fair deals are offered. It is a far better solution for both sides, AI developers and rightsholders, than endless litigation or constant lobbying for (or pushing back against) legislative change. It’s a global solution to a global problem.

© Hugh Stephens, 2026. All Rights Reserved.


[i] Any exception must satisfy three requirements: (1) Limited to certain special cases:. i.e. the exception must be narrowly defined and clearly circumscribed, rather than acting as a broad or general exemption; (2) No conflict with normal exploitation: i.e.the use must not interfere with the ways the copyright owner routinely makes money or manages the market for their work; (3) No unreasonable prejudice to legitimate interests of rightsholders, i.e. the use must not cause unfair or excessive economic or legal harm to the author or right-holder

We need more Canada in the Training Data, but through Licensing not Loopholes

Canada Has a Choice When it Comes to AI Training Content

Scrabble tiles arranged to display the words 'LOOPHOLES' and 'LICENSING' on a game board.

Michael Geist, Canada Research Chair in Internet and E-Commerce Law at the University of Ottawa has argued, in an appearance before the Heritage Committee of the House of Commons, that “we need more Canada in the training data”. He is absolutely right, but just not in the way he proposes. Dr. Geist is what I would call a well-known skeptic when it comes to the intrinsic value of copyright, a copyright “minimalist” if you will (probably an understatement).

With respect to the unauthorized and uncompensated use of copyrighted content for AI training, he states that “in the context of AI, the application of copyright isn’t clear cut. The outputs of AI systems rarely rise (to) the level of actual infringement given that the expression may be similar or inspired by another source, but it is not a direct copy of the original.” Whether the outputs mirror the inputs is not the sole issue. In some cases, such as when music and images have provided the inputs, they do. This is an infringement of the reproduction right, and likely also an infringement of the distribution right and the right to produce a derivative copy (under US law). In Canada the right to create another work from an original work comes from the right of adaptation. However, even without a mirrored output, full reproduction still takes place at the input stage, creating an infringement unless the copies meet a fair dealing purpose and fulfill fair dealing criteria, even if the copies are later deleted. As Keith Kupferschmid, CEO of the Washington DC based Copyright Alliance has pointed out in a recent blog post discussing the copyright principles that apply in AI training cases,

“Some people mistakenly believe that in order to establish an infringement during the input stage, the copyright owner needs to establish substantial similarity between the ingested copyrighted work and AI-generated output and if no substantial similarity exists there is no infringement in this stage. That is incorrect.”

Even without mirrored outputs, full non-transitory copies of copyrighted works are being made at the ingestion stage of AI training. That is an infringement, just as making a photocopy of a complete work, such as a book, would be an infringement unless covered by an explicit exception such as preservation purposes by a library or archive. 

Dr. Geist’s second line of argument is that if Canada makes it more difficult or costly to develop large language models, AI development will shift outside the country. This is a tried-and-true but tired pretext frequently employed by those seeking to justify the appropriation of copyright protected content in the name of “innovation”, as I pointed out in an earlier blog post. (CanLII v CasewayAI: Defendant Trots Out AI Industry’s Misinformation and Scare Tactics -But Don’t Panic, Canada). This is a race to the bottom, throwing the content industry under the bus on the pretext that everyone is doing it, even though that is untrue. One provision that has been selectively incorporated into the laws of some jurisdictions, like the UK and the EU, is an exception for “text and data mining” (TDM). Dr. Geist states this is why Canada also needs to introduce a similar statutory exception to promote AI.

However, not everyone is engaged in this race to the bottom. In fact, there are increasing doubts that establishing a statutory TDM exception for AI training is the best way to go. Australia has just firmly rejected the creation of a TDM exception in its copyright law even though it is also grappling with the same issue of how to incentivize AI training and research in that country. The UK’s current TDM exception is limited to non-commercial research purposes and in the face of strong opposition from its creative sector, Britain has put proposals to expand TDM on hold. Even the EU’s TDM law, which has two aspects, one limiting the data mining to non-commercial scientific research conducted by scientific research organizations or cultural heritage institutions while the other is a general purpose TDM that is open to commercial organizations, has guardrails. These include an opt-out provision whereby rightsholders can block ingestion of their content through technical measures, contract provisions or other means, in which case the TDM exception does not apply.

While opting-out by rightsholders is one way to limit the damage of unrestricted text and data mining, this is controversial because it places the onus on the rightsholder to take action whereas normally a party wanting to use someone else’s property would have to obtain permission in advance. Opting out is not a preferred solution for the creative community. It doesn’t work well in practice as rightsholders often lack the technical means or awareness to apply their opt-out rights. Because of this, the European Parliament’s Committee on Legal Affairs has just published a study examining how generative artificial intelligence interacts with European Union copyright law. The study recommends moving from opt-out to opt-in for rightsholders.

Thus, far from TDM being or becoming the norm, it is being rejected or constrained in a number of countries where the AI industry has been pushing it as the ultimate solution. The Canadian creative community, like the creative sector in Australia,  has spoken out strongly against introducing a TDM exception into Canadian law. Indeed, there is no need to do so as licensing solutions allowing AI training and text and data mining are becoming more and more common, including in Canada. For example, the Writers Union of Canada is studying a proposed agreement between select nonfiction authors, HarperCollins, and Microsoft to license full texts for the purpose of training artificial intelligence. Licensing agreements have taken off big-time in the US and elsewhere as the AI industry begins to understand this is the safest way to protect their investments. Canadian creators risk being left by the roadside if Canada brings in a TDM exception that would allow AI developers to steam ahead, appropriating content without payment or permission and ignoring licensing requirements by hiding behind a TDM exception.  The surest way to kill a nascent and growing licensing market is to give the AI sector a TDM loophole to exploit, removing any incentive to reach licensing agreements with rightsholders.  The solution is licensing, not loopholes.

Dr. Geist stated in his testimony to the Heritage Committee that AI developers would take the view that if they had to pay for (i.e. to license) content from Canadian creators, they would simply exclude it. The record of licensing deals being reached elsewhere suggests this is completely off base. Instead, the record shows that when AI developers want reliable, curated content to make their product better than the competition, they are ready to pay for it. But they will never pay for it if they are given a blank cheque through a legislated loophole. He also claims the position of the creative community is “Don’t use my stuff”. Again, the record of licensing deals to date and in the pipeline disproves this characterization in spades. Rather than blocking use of their content, creators are saying, “If you want to use my content, let’s talk”. Finally, Dr. Geist managed to completely mischaracterize the position of the creative community with regard to licensing. He said in his testimony that creators are advocating for a change to copyright law to mandate payments for AI training use. On the contrary, the creative community is simply asking that existing copyright law not be gutted. There is no need to create a mandatory payment requirement; existing copyright law is fit for purpose in dealing with how those wishing to use copyrighted content for purposes that fall outside fair dealing can do so. Negotiate a licence.

If any proof is needed of how the creation of a loophole will kill a licensing market is, all one needs to do is look at the sorry state of educational publishing in Canada. The industry has been decimated, and many authors have lost their livelihood because of the ill-conceived educational exception that was introduced into Canada’s Copyright Act in 2012. With that loophole in place, educational institutions across the country, with the notable exception of Quebec, began to tear up the reproduction licenses they had held from Access Copyright, the copyright collective representing authors. The educational exemption as part of fair dealing criteria could still be fixed, but the educational sector, facing severe financial pressures, has a powerful lobby working against it. The financial pressures are real, but taking a free ride on educational publishers and authors is wrong.

What happened with educational publishing is a cautionary tale for Canada. It should not make the same mistake twice. The way to promote a strong AI industry, alongside vibrant content industries, is licensing, not loopholes. Building a robust AI/TDM licensing market is the way to get more Canada into the training data, not giving the AI industry a blank cheque to help itself to the proprietorial content of others. With voluntary licensing everyone benefits. AI developers get secure access to quality content; the creative sector is rewarded for its efforts and becomes a partner in developing responsible AI. It’s a shame that the Canada Research Chair at the University of Ottawa doesn’t understand this.

© Hugh Stephens, 2025. All Rights Reserved.