Copyright, Cultural Issues and Canada’s General Election, 2025

Image: Shutterstock (AI generated)

As we complete the first few days in what is the shortest election campaign in Canadian history, the minimum 37 days required by law, where do the copyright and cultural industries stand with respect to electoral platforms and public consciousness? Given the overwhelming focus on dealing with economic and even potential political disruption coming from south of the border, along with traditional bread and butter issues like the cost of living, especially food and housing, one could be tempted to say that cultural and copyright issues are largely invisible. Party platforms have not yet been released (and are probably still being worked on) and by the time they are made public, the election will be well underway. So while there still may be a couple of small references to copyright issues in party platforms (as occurred in the 2021 election, none of which led to any substantive legislation), they will simply be part of a laundry list of possible actions in many disparate areas. However, that has not stopped the cultural sector from outlining its policy proposals, which have been laid out articulately by the Coalition for the Diversity of Cultural Expressions (CDCE), an umbrella group that represents more than 350,000 creators and artists, and more than 3,000 cultural enterprises. Despite the fact that copyright issues are not at or even near the top of the agenda, there is a strong undercurrent of Canadian nationalism in this election that will inevitably have an influence on policies in the cultural sector.

In 2021 the governing Trudeau Liberals included a promise to “protect Canadian artists, creators and copyright holders by making changes to the Copyright Act including amending the Act to allow resale rights for artists”. They were re-elected but did nothing. The Conservatives for their part undertook “recognize and correct the adverse economic impact for creators and publishers from the uncompensated use of their works…”. They weren’t elected so the commitment was meaningless. This time proposed changes to copyright legislation are unlikely to move the needle for any party although the issue of the unauthorized use of copyrighted content to train AI still needs to be resolved, since AI will become a front-burner issue for any party elected. The CDCE’s paper addresses this issue, among others, in its 9 recommendations. Broken down into 4 buckets, the CDCE’s proposals address (1) International Trade and Cultural Sovereignty (2) Broadcasting and CBC/Radio Canada (3) Copyright and (4) Artificial Intelligence and Culture.

The CDCE proposal under “International Trade” is to insist that the cultural exemption clause be retained if the CUSMA/USMCA is renegotiated, and that cultural activities, goods and services be excluded from all future agreements. The cultural exemption clause, (Article 32.6 of the CUSMA) is based on a similar exemption in NAFTA and the original US-Canada bilateral trade agreement of 1989 but is more of a political fig-leaf than a real protection since if the provision is invoked, the US can retaliate with equivalent effect in any trade sector. However, it provided comfort to the cultural sector at a time when free trade with the US was seen to make Canada vulnerable culturally. Thirty plus years of bilateral, and now trilateral, trade proved that fear to be unfounded—until now—and the cultural exemption has never been used. During the period from 1989 to the present, even through the ups and downs of Trump 1.0, the fundamentals of the initial bilateral Free Trade Agreement, then NAFTA, and now the CUSMA/USMCA were basically respected by all parties. Under Trump 2.0 this has all been called into question. If the Trump Administration is going to disavow the basic elements of the CUSMA, having a cultural exemption clause becomes less than meaningless.

On April 2, the US will unveil its “reciprocal tariff” regime. It has arrogated to itself the right to include, in addition to tariffs imposed by other countries, self identified non-tariff measures in its calculations. Among these may be various cultural support measures imposed by Canada on foreign entities operating in Canada requiring them to make financial contributions to Canadian content. If that happens, the US will be violating yet again the provisions of the CUSMA/USMCA as it has already done with regard to the imposition of tariffs on some products on the specious grounds of fentanyl trafficking from Canada to the US, (less than 20kg in all of 2024). However, given the surge in Canadian nationalism as a result of the tariff threats but more particularly the verbal diarrhea coming daily from President Trump about Canada becoming the 51st state, it is unlikely that any Canadian government would throw Canada’s cultural identity under the bus for the sake of preserving tariff-free access to the US market for some commodities. Thus, seeing Canada sacrifice cultural support measures that may annoy some US businesses operating in Canada (like online streaming content providers) in return for a degree of tariff relief is an unlikely outcome in the present circumstances.

This surge of nationalism relates to the second of the CDCE’s “demands”, protecting the CBC and the Canadian broadcasting environment. Ever since Pierre Poilievre became leader of the opposition Conservative Party, one of the Party’s mantras has been “defund the CBC”. There is no question that the CBC business model is in need of reform, particularly its English language entertainment television service which captures a very small market share, but CBC radio, CBC news broadcasts and CBC’s French language service, Radio-Canada, remain highly relevant, as this CBC explainer attempts to show. Given the need to protect national identity in the face of the Trumpian onslaught, and the recent rediscovery that perhaps Canada is not so “broken” after all, if ever there was a need for this national institution, it is now.

The third basket of issues raised in the CDCE position paper relates to copyright concerns, which get very little traction among the general electorate but are important to the creative and cultural community. Once again, the CDCE reminds parties of the lack of an Artists Resale Right in Canada (noting previous promises to establish this measure), as well as some other longstanding issues like fair remuneration for writers and publishers for the use of their works in the education sector and extending the private copying regime to electronic devices. This would impose a small levy (about $3) paid by manufacturers and embedded in the cost of a smartphone to compensate for unregulated widespread copying of music on these devices, with the funds flowing back to music creators.

The final bucket deals with Artificial Intelligence (AI) and copyrighted content. At the present time there are some 40 lawsuits in the US pitting rightsholders against AI developers, and even a couple of cases in Canada. Canada has been slow off the mark in addressing this issue; at the moment there is no Text and Data Mining exception in Canadian copyright law and both rightsholders and AI developers are not clear on the ground rules. The CDCE is asking that a legislative framework be adopted that includes the key principles of (1) Authorization (by the rightsholder) (2) Remuneration (payment for use of copyrighted content) and (3) Transparency (the establishment of disclosure rules as to what training data is used in AI systems and ensuring that all AI-generated content is clearly identified). These are reasonable asks but there is no guarantee they will be respected.

In the US, AI developers are pushing the Trump Administration to give them a pass on respecting author’s copyright, notwithstanding the cases before the courts, using the argument that the US will lose the AI race to China if US developers cannot help themselves freely to the content of others. OpenAI (which is being sued by the New York Times) and Google argued in submissions to the US government that giving them unfettered access to data, including content owned by others, is essential for national security. Described by blogger David Newhoff as “tech bro bombast”, OpenAI’s attempt to wrap itself in the national security blanket is a cynical ploy to get around the inconvenient fact that it and other AI developers are hijacking the creative work of authors, artists, and musicians without permission or compensation while creating outputs that in a number of cases can compete with or even displace the original works that contributed to their training. A similar situation is developing in the UK where the creative community is pushing back against the original copyright carte blanche that the UK government seemed inclined to give to the tech community, in the name of AI competitiveness. Canadian governments are not beyond succumbing to the siren calls of the AI community and it is timely to establish some guiding principles, of which Authorization, Remuneration and Transparency are a good place to start.

However, while AI and copyright are not going to become election issues, national identity, which is closely intertwined with cultural sovereignty, surely is. Indirectly, copyright will be important as it is one of the foundation stones of cultural sovereignty, an issue that would have played second fiddle to economic issues like food inflation, carbon pricing, cost of housing, fuel and utility costs etc until Donald Trump started spouting his annexationist nonsense.

Frankly, had Trump really wanted to absorb Canada (eventually) he should have brought Canada inside the US economic tent and made the country even more reliant on the US market, by providing it with an exception to his attempts to take on the world trading system. Instead, he has woken Canadians from a restful, dependent slumber brought on by three decades of relatively uncontroversial free trade and economic integration and made them realize that they have no one to depend on but themselves. In doing so, he has revitalized a sense of nationalism that will play out in this election. Who can best defend Canadian interests has become the litmus test for Canadian voters, leading to a remarkable resurgence for the Liberal Party under new leader Mark Carney after the political corpse of Justin Trudeau was removed from the electoral scene. This may or may not change during the course of this short campaign. One thing is certain; while copyright issues per se will not get much profile, cultural identity issues will certainly be in the spotlight. This is a shift in emphasis that in the long run is likely to benefit the creative sector.

© Hugh Stephens, 2025. All Rights Reserved.

Using Copyrighted Content to Train AI: Can Licensing Bridge the Gap?

Image: Shutterstock

The struggle between authors (writers, artists, musicians) and AI developers over the unauthorized and uncompensated use of copyrighted works to train AI applications continues, both in the courts (here is a summary of the current state of play in the US where most of the litigation is taking place) and in the political arena, such as the UK government’s latest initiative to put its thumb on the scale in favour of the AI industry, now slowed down by opposition within Parliament. The creative industries in Britain are still nervous, however, as demonstrated by the coordinated “Make it Fair” campaign organized by leading UK newspapers on February 25. While the courts may provide some guidance, it is unlikely to be dispositive and is almost certainly to be somewhat contradictory and lengthy, given the appeal process that will play out. With new applications being rolled out every day, AI appears to be unstoppable. Let’s accept that this is the case. If so, what then will be the rules governing the use of AI training content, particularly content that is protected by copyright, such as books, journalistic output, paintings, musical compositions etc.?

It is already apparent that at least some of the AI output trained on these materials will compete in the marketplace with the original works. If that is the case, then surely some of that additional value should be shared with those who helped create the content initially. The way this will most likely be done is through licensing in the form of payment and permission for use of the copyrighted creative output that enabled the training to take place. Licensing would also help resolve another potential issue, the possibility that the final product produced by the AI algorithm infringes on the copyright of the works on which it was trained. This is unlikely to happen in the case of written works but is certainly potentially possible with graphic or musical works.

While many have called for licensing as a solution, there are many challenges to be overcome to make it work effectively. Yet some licensing is already taking place between AI developers and owners of well delineated data sets. As an example, various newspaper and magazine publishers have already reached licensing agreements with AI providers. OpenAI has signed licensing deals with the Wall Street Journal, Times of London, the Financial Times, Time, Le Monde, Axel Springer and others. This is in marked contrast to OpenAI’s relationship with the New York Times, which has led to one of the most prominent lawsuits in the field, with the Times suing OpenAI for copyright infringement. The reason for this lawsuit, of course, is because licensing negotiations between the two entities broke down. Some photo and image licensing companies have concluded AI deals (Shutterstock is the most prominent example) while others, such as Getty Images have not. (Getty is suing StabilityAI in the UK). Eventually most of the institutional or corporate holders of valuable content in one form or another will likely reach, or attempt to reach, licensing deals with the major AI developers. But that still leaves out an awful lot of copyright-protected content.

The conundrum is how to deal with the millions of individual creators who produce content in different formats, and tie them into a workable licensing regime. The first challenge is how to even figure out who is producing content that is likely to be used by AI developers. The second is to calculate how much that use is worth. Then there is the challenge of how to administer a collective licensing scheme in a way that is both practical and affordable and where the small amount of royalties for individual works are not swamped by the administrative costs of collection and disbursement. Finally, there is the question of how to resolve the issue of competing licensing organizations in order to provide more or less one-stop-shopping for the AI industry.

It is worth noting that one-stop-shopping currently does not exist in any area of collective licensing. Different collectives represent creators in different fields so music, publishing, art, broadcasting and visual arts licensing are all represented by different organizations, in some cases with more than one collective in a given field. The Copyright Board of Canada lists 36 copyright collectives on its website. I haven’t seen a definitive list for the US but this university website lists about the same number.

Whereas users of music only have to deal with a handful of CMOs (collective management organizations), and users of text based content (online or offline) need only to acquire a reprographic license from the major licensing collectives for published works, such as the Copyright Clearance Center in the US or Access Copyright in Canada, AI developers access the full gamut of content. It will be challenging to make access easy for the AI development industry, a point developed by Dr. Pamela Samuelson of the University of California, Berkeley, well-known copyright scholar (and skeptic, let it be added). In her recent paper in the UCLA Law Review (“Fair Use Defenses in Disruptive Technology Cases”), Samuelson focuses primarily on the question of fair use—as suggested by the title—but also examines the issue of a collective licensing regime for generative AI development. She manages to raise just about every objection conceivable (see pp.80-86 of the document for more details);

–generative AI uses all forms of content therefore the licence would have to be very broad
-an issue would arise as to whether content for training was used just once, or on repeat occasions
-it would be very difficult and costly to administer given that there could be literally billions of creators involved

-creators would get very little revenue; the bulk would go to the administering agencies, the CMOs
-it would be difficult to determine value and to set a price on each transaction
-what about orphan works?
-differing national regimes might create confusion; alternatively some countries might not require a licence payment, giving them an unfair advantage
-it would be unfair to startups since the incumbents have already scooped volumes of content without payment.

She notes that creators may lose out, but since AI will affect the livelihoods of so many others, this is not exclusively a copyright problem. Tough luck creators.

Clearly Dr. Samuelson is not in favour of a collective licensing regime for content appropriated by AI developers, yet despite her firehose of cold water, there are a number of promising developments in this area. For example, the Copyright Clearance Center (CCC) in the US recently announced it would provide AI re-use rights within its Annual Copyright Licenses, making the CCC’s licence “the first-ever collective licensing solution for the internal use of copyrighted materials in AI systems.” Note the caveat. While covering re-use of content for AI applications, the CCC makes it clear that;

“The license enables participating rightsholders to fulfill the needs of companies that require an efficient way to legally acquire the rights to use copyrighted materials within AI systems for internal use.”

Not training. The Copyright Agency in Australia has done something very similar.

“Starting from February 2025, Copyright Agency will extend its Annual Business Licence to cover staff of licensed businesses who include third party material in prompts for AI tools (and) copy and share outputs from AI tools with colleagues”.

However, it does not apply to AI training and does not allow capture of the content outside the business, such as by an externally provided AI tool.

Likewise, the Copyright Licensing Agency (CLA) in the UK issues a Text and Data Mining (TDM) Licence. The CLA’s website explains that TDM “is the process of transforming unstructured content into a structured format to analyse, extract and identify meaningful information and insights. By using TDM, organisations can harness the power of vast volumes of information and data, capturing and revealing key concepts, trends, and hidden relationships.” Sounds quite a bit like training generative AI, but it’s not.

“CLA’s TDM licence extension includes rights covering use of published content for TDM purposes. This does not cover the use of content in training or prompting Generative AI models.”

Canada’s equivalent CMO, Access Copyright, is actively examining the issue, as it notes in its new strategic plan for 2025-2028;

“Like collective rights management organizations around the world, we will actively explore how we might enhance our corporate licence offerings to include uses related to AI, providing Canadian rights holders who wish to participate in the emerging market for AI licensing to do so, either in Canada or by virtue of reciprocal agreements with sister organizations.”

It is clear that these Reproduction Rights Organizations (CMOs by another name) are cautiously feeling their way forward to find the appropriate role for collective licensing. Meanwhile the private sector has not been sitting idly by. Forbes reports that so many content aggregation startups have been established that they have formed a Data Providers Alliance. Recently launched “Created by Humans” is another commercial entrant that is pitching itself to authors.

“Take control of your work’s AI Rights and get compensated for its use by AI companies.”

As these new enterprises enter the market, it threatens to become quite crowded. Just as there are more and more AI companies, including new entrants like DeepSeek, a proliferation of new sector-specific content-aggregators will make licensing more challenging. If the CMOs wait too long, they will face entrenched competition. Not all these new aggregators will survive. In the end, AI developers will not subscribe to multiple content licensors; they will go with the ones that provide the broadest coverage. It will be a Darwinian selection process.

While this is happening, other countries are experimenting with the concept of extended collective licensing for AI content. This allows CMOs to grant licenses on behalf of both their members and non-members alike. An extended collective licence is not a compulsory licence but it could lead to such a system being established. Spain was first out of the gate, but has since pulled back after the proposed Royal Decree attracted the criticism from many rights holders that it would proscribe their options. Yet it is one of many solutions being tested.

The recently released US Copyright Office report “Identifying the Economic Implications of Artificial Intelligence for Copyright Policy” includes an extensive discussion of licensing possibilities, including examining the pros and cons of a new statutory blanket licence. This would need to include a provision excluding rightsholders (such as entities that have already reached licensing agreements with AI developers) who have the ability and wish to issue voluntary licences that generate greater remuneration than a statutory payout would earn. This raises thorny opt-in/opt-out issues. Compromises will be required, but the challenges are not insurmountable.

The trick is to devise a system that will capture as much content as possible while allowing some flexibility to rightsholders, allocating payments in way that is fair and efficient (the USCO paper suggests that revenues associated with a work could serve as a rough proxy for its relative value), at the same time minimizing administrative costs so that expenses do not exceed potential revenues for rightsholders holding limited content inventory. Can it be done?

Despite the many obstacles identified by Dr. Samuelson and others, I am convinced that in the end collective licensing for content used in AI development and applications will become as accepted as the collective licensing regimes for use of various forms of copyrighted content today. The way forward won’t be straightforward; there will be zigs and zags. The courts and legislatures will play a role, as will authors, publishers, and the AI developers themselves. But in the end we will get there. Licensing, including some form of collective licensing, is the inevitable bridge that will bring AI developers and copyright holders together.

© Hugh Stephens, 2025. All Rights Reserved.

The Height of Hypocrisy! OpenAI Accuses DeepSeek of Stealing its Content


Image: Shutterstock (edited)

Am I the only one, or did anyone else have just a touch of schadenfreude when they read the story in the New York Times that OpenAI is claiming the Chinese start-up DeepSeek may have “improperly harvested” its data. What irony! DeepSeek caught everyone’s attention earlier this week when it announced a new AI application that appears to outperform or at least match OpenAI’s ChatGPT. Not only that, it is also open source and completely free to download and use. More important, its alleged development costs were but a fraction of the development cost of US models, reported to be in the hundreds of millions whereas DeepSeek claims that it produced its results with an investment of as little as $6 million. (This clearly does not include the value of earlier R&D, but the question is whether or not DeepSeek covered these costs).

We saw the shock this caused on the NASDAQ especially with respect to chip-designer Nvidia’s share price, with over $600 billion in value wiped off its valuation in one day. As often happens, there was a rebound the following day as saner heads digested the news and found a silver lining in the fact that AI development costs could be greatly reduced yet spending would continue. Of course, the spectre of “unfair” Chinese competition was raised, while others wondered how DeepSeek did it in the face of US high-tech embargos on the sale of advanced Nvidia chips to China. “They must have cheated” was the mantra.

It appears that part of DeepSeek’s success is based on what is called “distillation” in the AI industry. As explained in this tech article, distillation is a technique that “focuses on creating efficient models by transferring knowledge from large, complex models to smaller, deployable ones”. The earlier models do the heavy-lifting with respect to research and as they produce results, those results are incorporated into newer training models that take advantage of the earlier work. To my untrained mind, this sounds like building on knowledge created by others, as happens all the time or, to look at it negatively, by free riding on the investment of others. The question is, what knowledge is protectable and proprietary? This dichotomy is at the heart of the debate over copyright. You can’t copyright an idea, but the specific expression of an idea is protectable. Likewise, the functionality of software code cannot be copyrighted although a specific software program is considered a “literary work” and is protected.

There is also the issue of open source. Release of code as open source enables further advancements, pushing the boundaries of knowledge. This is a common feature of the digital revolution and one reason for rapid advancements in Silicon Valley. However, not all content is fully open source. In the case of OpenAI it would seem it considers its content to be proprietary to the extent that it can control the use to which it is put. The accusation is that DeepSeek took and distilled OpenAI’s results to create a competing application without permission. In effect, DeepSeek used ChatGPT to improve its own model.

OpenAI’s position that it can dictate the uses to which ChatGPT can be put is, in my view, contradictory, hypocritical and in the end morally if not legally indefensible. OpenAI has no problem enabling and encouraging people to use ChatGPT to “improve on” or create works in any field, from AI written novels to AI created art or music, resulting in works that directly compete with authors, artists and musicians. Remember that OpenAI has used their original copyrighted works without permission to build the AI machine that now threatens their livelihood and ability to create. Yet when that same AI application, ChatGPT, is used to improve on or create a new and better AI platform, this is declared to be infringement.

While distillation is common across the AI field, OpenAI claims its terms of service prohibit any use of data generated by its systems to build technologies that compete in the same market. This caveat would be similar to that which is applied to copyrighted content made publicly available on websites, with a disclaimer that it is copyright protected and potential users should contact the rightsholder. Did that stop OpenAI from helping itself without permission to this protected content to train its AI algorithm? Absolutely not. In fact, while it justified its activities by saying that all it was doing was taking “publicly available” content, not even paywalls and terms of service were allowed to get in their way. This was clearly demonstrated in the case brought against it by the New York Times. (When Giants Wrestle, the Earth Moves (NYT v OpenAI/Microsoft).

It seems that from OpenAI’s perspective, use of other people’s content without permission is okay, but when it’s their content, not so much. OpenAI is partially owned by Microsoft which is itself engaged in rolling out its own AI application, Copilot, trained in part through the unwitting contribution of hundreds of millions of users of Microsoft software, like MS-Word, as I wrote about last month. (Writers! Do You Know your Drafts on MS Word are being Scooped by Microsoft to Build its AI Algorithm? But You Can Stop This From Happening (Read On).

Given all that has transpired, and the struggle that authors and rightsholders are facing to protect and get paid for the use of their works in AI training, it is hard to have much, if any, sympathy for OpenAI. I certainly don’t. Poetic justice.

© Hugh Stephens, 2025. All Rights Reserved.

Writers! Do You Know your Drafts on MS Word are being Scooped by Microsoft to Build its AI Algorithm? But You Can Stop This From Happening (Read On).

Image: Shutterstock

Although I post my blog content on WordPress, I usually use MS Word to draft my content initially. I am used to it, and it is easy to use. Little did I know that, according to the blogsite and forum nixCraft, Microsoft recently (September Privacy update) switched on a feature that allows them to ingest everything you write on Word to help develop their AI Algorithm, called Copilot. The setting is turned on by default in the Privacy settings and must be unchecked manually. Did Microsoft tell you this? Well, kinda, sorta. Microsoft says, “we don’t use your customer data to train Copilot or its AI features unless you provide consent to do so”. Did you provide consent? You no doubt did, unknowingly, when Microsoft updated its Terms of Use, which it does on a regular basis. If you continued to use Office 365, you granted consent.

In the last few days, a pop up has appeared when I am doing something on Word through Office 365.

“Thank you for using Office! We’ve made some updates to the privacy settings to give you more control”.

If you believe that, I have a bridge to sell you.

If you go to Privacy Settings there is a summary blurb on how the Terms of Use were updated on September 30. If you look hard enough you will find this reference;

“We added a section on AI services to set out certain restrictions, use of Your Content and requirements associated with the use of the AI services.”

You really should read the full Terms but as Microsoft notes, this will take an hour of your time (ESTIMATED READING TIME: 55 Minutes; 14268 words).

Having waded through it, this I believe is the relevant wording;

b. To the extent necessary to provide the Services to you and others, to protect you and the Services, and to improve Microsoft products and services, you grant to Microsoft a worldwide and royalty-free intellectual property license to use Your Content, for example, to make copies of, retain, transmit, reformat, display, and distribute via communication tools Your Content on the Services.

There is nothing here about opting out. You have to go to Privacy settings and do some digging to get to that. By masking these changes to make them appear that your privacy has been strengthened (whereas in fact it is just a content grab), Microsoft has stood things on its head by putting the onus on you, the user, to exercise your privacy rights. If you don’t want your creative work used to help train its AI algorithm, which in the end might compete directly or indirectly with your work, you need to opt out (unless you want to stop using MS Word altogether). Microsoft, however, is not suggesting this as a preferred option, or even letting it be widely known that it exists as an option. In fact, when you go into Settings to opt out you are presented with this little gem; “The Trust Center contains security and privacy settings. These settings help keep your computer safe. We recommend that you do not change these settings”.

But that is exactly what you must do if you want to keep your creative content out of the hands of Microsoft’s AI developers. Here is how to do it, based on instructions from nixCraft.

On a Windows computer, when in a Word file, go to File in the top left-hand corner. There is a drop-down menu. You want to go to Options. On my computer the Options choice does not show up unless you hit the arrow at the bottom of the page for “More”. When you get to Options, go to Trust Center (left side menu), then Trust Centre Settings. Next up is Privacy Options which leads you to Privacy Settings. There is a drop down menu, including Connected Experiences. There is a heading labelled “Experiences that analyze your content“. This box is checked for you. You want to uncheck it. To save the setting you will have to log out of Word and then log back in. (Update: I have just discovered a quicker way to do this. Go File-Options-General (top of list)-Privacy Settings-Connected Experiences-Experiences that analyze your content-Uncheck).

Eliminating this option will come at a price, according to all the “Learn More” button provided by MS, but it is your choice. For my part, I will forgo the bells and whistles for privacy.

The opt out process is not simple and not intuitive, but worth doing, even if only as a matter of principle. Office365, unlike Google Search or Bing, is not free. We pay to use it through an annual subscription. Even the tired old argument that you are providing your data as a sort of payment for “free” use of a platform’s service does not apply in this case. Microsoft needs more data to feed its AI machine and yours will do just fine, thank you very much. Don’t let them get away with it.

© Hugh Stephens, 2025. All Rights Reserved.

Looking Back at 2024: It’s All About AI and Copyright (And a Few Other Things)

Image: Shutterstock

A retrospective on the year now coming to a close is what one expects this time of year, so I will try not to disappoint. However, when I look back at the copyright developments I wrote about in 2024, the dominant issues that jump out are AI, AI and AI. You can’t read or think about copyright without Artificial Intelligence, or to be more correct, Generative Artificial Intelligence (GAI), occupying most of the space despite many other issues on the copyright agenda. The mantra of “AI, AI and AI”, as in “Location, Location and Location” is apt because there are at least three important copyright dimensions related to AI; training of AI models; copyright protection for outputs generated by AI; and infringement of copyright by works created with or by AI. Of the three, the use of copyrighted content for AI training is the most salient.

Last year in my year-ender, I also discussed AI and the numerous lawsuits that were emerging as rightsholders pushed back on having their content vacuumed up by AI developers to train their algorithms. Those lawsuits have only multiplied. At last count, there are more that 30 cases in the US, ranging from big media vs big AI (New York Times v OpenAI/Microsoft) to class action suits brought by artists and authors, as well as litigation in the UK, EU, and now in Canada (see here and here). That is just on the input side.

In terms of output, i.e. whether works produced by an AI can be copyrighted, there are a couple of interesting cases in the US where applications for copyright registration have been refused by the US Copyright Office (USCO) because of a lack of human creativity. A couple of months ago, I discussed two such high profile cases, one brought by Stephen Thaler, and the other by Jason Allen. To date the USCO is not budging, although it is undertaking an extensive study of the issue. Part 1 of its study, on digital replicas, was published in July of this year. The next section on copyrightability is expected to be published in January with the issues of ingestion for training and licensing in Q1 2025.

While the USCO has to date denied applications for copyright registration of AI-generated works, the Canadian copyright office (CIPO-Canadian Intellectual Property Office) has been caught up in a problem of its own making. This is because Canadian copyright registration is granted automatically, so long as tombstone data and the prescribed fee is provided. The work for which registration is sought is not examined. As a result, copyright certificates have been issued to works created by AI, notwithstanding the general presumption that copyright protection is only accorded to human created work (although this is not explicitly stated in the Act). In July a legal challenge was launched against copyright registrant Ankit Sahni, who successfully registered a work with CIPO claiming an AI as co-author. The case was brought by the Canadian Internet Policy and Public Interest Clinic (CIPPIC) at the University of Ottawa, as I wrote about here. (Canadian Copyright Registration and AI-Created Works: It’s Time to Close the Loophole).

While the courts in the US, UK, Canada and elsewhere are grappling with various issues related to AI and copyright, governments are studying the issue.

In Australia, the Select Committee on Adopting Artificial Intelligence issued its final report in November. While the report was wide-ranging, three of its recommendations related to copyright;

• engagement with the creative Industry to address unauthorized use of their works by AI developers and tech companies,

• transparency in Training Data by requiring AI developers to disclose the use of copyrighted works in training datasets and ensure proper licensing and payment for these works, and

• remuneration for AI Outputs, with an appropriate mechanism to be determined through further consultation

These are important principles, but how they will be implemented in practice remains to be determined.

In Canada, a consultation on AI and copyright was launched late in 2023 with submissions to be received by January 15, 2024. The Canadian cultural community put forth three key demands;

• No weakening of copyright protection for works currently protected (i.e. no exception for text and data mining to use copyrighted works without authorization to train AI systems)

• Copyright must continue to protect only works created by humans (AI generated works should not qualify)

• AI developers should be required to be transparent and disclose what works have been ingested as part of the training process (transparency and disclosure).

Submissions to the consultation were published in mid-year but since then there has been no apparent action. Given the current political crisis facing the Trudeau government, none is expected in the near term although the issue will inevitably have to be addressed after the general election in 2025.

While the EU has already established some parameters dealing with use of copyrighted materials for AI training, the new UK Labour government is taking another run at the issue after various proposals in Britain to find a modus vivendi between the AI and content industries under the Tories went nowhere. The current UK discussion paper on Copyright and Artificial Intelligence, which seems excessively tilted in favour of the AI industry, has aroused plenty of controversy. While it says some of the right things, such as proclaiming that one of the objectives of the consultation is to “support…right holders’ control of their content and ability to be remunerated for its use” the thrust of the paper is to find ways to encourage the AI industry to undertake more research in the UK by establishing a more permissive regime with respect to use of copyrighted content. It is based on three self-declared principles; (notice how these things always seem to come in threes?);

• Control: Right holders should have control over, and be able to license and seek remuneration for, the use of their content by AI models

• Access: AI developers should be able to access and use large volumes of online content to train their models easily, lawfully and without infringing copyright, and

• Transparency: The copyright framework should be clear and make sense to its users, with greater transparency about works used to train AI models, and their outputs.

These three objectives then lead to what is clearly the preferred solution;

“A data mining exception which allows right holders to reserve their rights, underpinned by supporting measures on transparency”

Fine in principle, but the devil is always in the detail and the details in this case revolve around transparency (how detailed, what form, what about content already taken?) and, in particular, reservation of rights, aka “opting out”. This is easy to proclaim in principle but difficult to do in practice. British creators are up in arms, led by artists such as Paul McCartney, and supported by the creative industries in the US. The British composer Ed Newton-Rex has penned a brilliant satire explaining how AI development in the UK will work if current proposal is enacted. The problem with an opt-out solution is essentially twofold; it doesn’t deal with content already absorbed by AI developers and it would be cumbersome if not impossible for many rightsholders to use.

Other governments have addressed the issue in different ways. Singapore has taken a very loose approach toward copyright protection, putting its thumb firmly on the scale in favour of AI developers. It is currently considering additional proposals that would strip even more protection from rights-holders, who are pushing back strongly. Japan had been widely and incorrectly reported to have been on the same path, resulting in a welcome clarification this year from the Agency for Cultural Affairs regarding the limits of Japan’s text and data mining (TDM) exception.

While AI dominated the copyright agenda in 2024, there were other issues relating to copyright and copyright industries that I wrote about. The ongoing question of payment for news content by large digital platforms continued to play out in different ways. In Canada, the struggle between the government and US tech giants Google and META was finally “resolved” (after a fashion) at the end of last year. Google agreed to “voluntarily” pay $100 million annually into a fund for Canadian journalism in return for being exempted from the Online News Act (ONA) while META called the government’s bluff by blocking Canadian news providers from its platform thus, in theory, avoiding being subject to the ONA. However, META has a very subjective interpretation as to what is Canadian news content, allowing some news providers to post to it, while many users have found workarounds, as documented by McGill’s Media Ecosystem Observatory. While the CRTC investigated, the issue is still unresolved.

Meanwhile in Australia, it seems that META intends to go down the same road of blocking news, announcing it will not renew the content deals it initially signed with Australian media in response to Australia’s News Media Bargaining Code, the model upon which Canada’s legislation was based. Unlike in Canada, the Australian government is planning a robust response. (More on this in a future blog post). Finally, on the same topic, California (which was threatening to introduce its own version of legislation to require digital platforms to compensate news content providers) emerged with an outcome very similar to that reached in Canada, with Google offering up some funding (although proportionally less than in Canada) while META appears to have walked away.

Controlled Digital Lending (CDL) was another copyright issue finally settled in 2024 (in the US). The Internet Archive, after losing a lawsuit brought against it by a consortium of publishers who argued that the digital copying of their works constituted copyright infringement, notwithstanding the Archive’s theory that they were simply lending a digital version of a legally obtained physical work held by them (or someone else associated with them), lost its appeal. In December, the deadline for further appeals expired, thus effectively ending this saga. Whether Canadian university libraries, some of whom are avid devotees of CDL, will take note remains to be seen.

The issue of circumventing a TPM (“Technological Protection Measure”), commonly referred to as a “digital lock” and often represented by a password allowing access to content behind a paywall, was also front and centre this year in Canada. In the case of Blacklock’s Reporter v Attorney General for Canada, the Federal Court found that an employee of Parks Canada, who shared a single subscription to Blacklock’s with a number of other employees by providing them with the password did not infringe Blacklock’s copyright since the employee did not circumvent (in the meaning of the law) the TPM and the purpose of the sharing was for “research“, which is a specified fair dealing purpose. Blacklock’s is a digital research service that sells access to its content and protects its content with a paywall, as is common for many online content providers, like magazines and newspapers.

Despite the hoo-ha of anti-copyright commentators asserting the Court had found that “digital lock rules do not trump fair dealing“, it was equally clear the Court had ruled that fair dealing does not trump digital locks (TPMs). The Court did not undermine the protection afforded to businesses to protect their content through use of TPMs. Rather, it determined that sharing a licitly obtained password did not constitute circumvention as outlined in the Act, as I explained here. (Fair Dealing, Passwords and Technological Protection Measures (TPMs) in Canada: Federal Court Confirms Fair Dealing Does Not Trump TPMs (Digital Lock Rules). Although the Court did not legitimize circumvention of a TPM for fair dealing purposes, contrary to claims stating the opposite, its acceptance of password sharing is an outcome that legal experts have disagreed with, (as do I for what it is worth). The law is very clear that fair dealing cannot be used as a pretext or a defence against violation of the anti-circumvention provisions of the Copyright Act. The decision now under appeal by Blacklock’s.

Finally, the last copyright point of note for 2024 is that this year marked the bicentenary of the introduction of the first copyright legislation in Canada, in the Assembly of Lower Canada, in 1824. It also marked the centenary of the entry in force of the first truly Canadian Copyright Act on January 1, 1924. This two hundred years of domestic copyright history is worth celebrating. The first legislation was introduced “for the Encouragement of Learning” so that more local school texts would be written and printed. Given the current standoff between the secondary and post-secondary educational establishment and Canadian authors and their copyright collective over license payments for use of copyrighted works in teaching, one wonders whether we have really learned anything about the role copyright plays in our society. (Copyright and Education in Canada: Have We Learned Nothing in the Past Two Centuries? (From the “Encouragement of Learning” to the “Great Education Free Ride”).

Leaving that question with you to ponder, gentle Reader, is probably a good way to end this look back over the past 12 months. Stay tuned for more commentary on copyright developments in 2025.

© Hugh Stephens, 2024. All Rights Reserved.

CanLII v CasewayAI: Defendant Trots Out AI Industry’s Misinformation and Scare Tactics (But Don’t Panic, Canada)

Image: Pixabay

Last month I highlighted the first AI/Copyright case in Canada to reach the courts, CanLII v CasewayAI. CanLII, (the Canadian Legal Information Institute), a non-profit established in 2001 by the Federation of Law Societies of Canada, sued Caseway AI, a self-described AI-driven legal research service, for copyright infringement and for violating CanLII’s Terms of Use through a massive downloading of 3.5 million files which Caseway allegedly used to populate its AI based services. Now the principal of CasewayAI, Alistair Vigier, through an article (Don’t Scare AI Companies Away, Canada – They’re Building the Future) published in Techcouver, has responded publicly by trotting out many of the tired and specious arguments put forward by the AI industry to justify the unauthorized “taking” of copyrighted content to use in or to train generative AI models. Let’s have a closer look at these arguments.

Vigier opens by referencing another AI/Copyright case in Canada where a consortium of Canadian media companies is suing OpenAI for copyright infringement. He claims this is all based on a misunderstanding of how AI training works, stating that “AI systems like OpenAI rely on publicly available data to learn and improve. This does not equate to stealing content.” Whether data is “publicly available” or not is irrelevant when it comes to determining whether copyright infringement (aka stealing content) is concerned. Books in libraries are publicly available, or so is a book that you purchase in a bookstore, or content on the internet that is not behind a paywall. (It is worth noting that the Canadian media companies also claim that OpenAI circumvented their paywalls to access their content when copying it). But in none of these cases is copying permitted unless the copying falls within a fair dealing exception, which is very precise in its definition. Labelling copied material as “publicly available” is a red herring.

Vigier’s next argument is to equate the ingestion of content by various AI development models with a human being reading a book. We know that humans enhance their knowledge through reading and are thus able, presumably, to better reason based on the content they have absorbed. Vigier says, “This is how AI works. The AI “reads” as much as it can, gets really “smart,” and then explains what it knows when you ask it a question. Like a human learns from reading the news, so does an AI.”

Really? A human does not make a copy, not even a temporary copy, of the content although some elements of the content are no doubt retained in the human brain. But AI operates differently. It makes a copy of the content. This should be beyond dispute although the AI industry continues to muddy the waters by claiming that when content is “ingested” it is converted to numeric data and is thus not actually copied. This is a fallacious argument. Just because the form changes, this does not mean there is no reproduction. When you make a digital copy of a book, there is still reproduction even though the digital form is different from the original hard copy version. When a work is converted to data, the content is still represented in the dataset.

Vigier dubiously states, with regard to OpenAI, “OpenAI’s models do not reproduce articles verbatim; they process vast datasets to identify patterns, enabling insights and efficiency.” Apart from the fact that the New York Times in its separate lawsuit in the US has been able to demonstrate that by typing in leads of articles, it can prompt OpenAI to reproduce verbatim the rest of the article (OpenAI claimed that the Times “tricked” the algorithm), copying is copying even if the result of the copying is somewhat different from the original. The Copyright Act is crystal clear on this point. Section 3 (1) of the Act states that, “For the purposes of this Act, copyright, in relation to a work, means the sole right to produce or reproduce the work or any substantial part thereof in any material form whatever…“. If copyright protected content is reproduced in its entirety without permission for a commercial purpose (eg for AI training), that is infringement, unless the use qualifies as a fair dealing under Canadian law or fair use in the US.

The issue of whether ingestion of content to train an AI application results in copying (reproduction) has been carefully studied and documented. One of the most thorough examples is a recent SSRN (Social Science Research Network) paper, entitled, “The Heart of the Matter: Copyright, AI Training, and LLMs” with noted scholar Daniel Gervais (a Canadian by the way) of Vanderbilt University as lead author. The article goes into a detailed discussion on how copying of content occurs during AI scraping to build a Large Language Model (LLM), including the stages of tokenization, embedding, leading to reward modelling and reinforcement learning. The section of the article explaining how copying occurs (pp. 1-6) is dense, technical text but the conclusion is clear, “LLMs make copies of the documents on which they are trained, and this copying takes various forms, and as a result, with appropriate prompting, applications that use the LLMs are able to reproduce original works.” A shorter (and earlier) version explaining how the LLM copyright process works can be found in this article (“Heart of the Matter: Demystifying Copying in the Training of LLMs“), produced by the Copyright Clearance Center in the US. It is also worth noting that these explanations refer only to ingestion of text. AI models that train on images and music are even more likely to produce exact or close-to-exact reproductions of some of the works they have been built and trained on.

So much for the misinformation in Vigier’s article. Now to the scare tactics. He says that the recent Canadian media lawsuit against OpenAI sends a negative message to innovators that Canada may not be open to AI development.

“If Canada wishes to remain relevant in this (AI) sector, it must balance protecting intellectual property and promoting technological progress.”

The fact that there are currently more than 30 lawsuits in the US, including the seminal New York Times v OpenAI case, does not seem to have slowed down the AI companies in the US. In the UK, legislation has been introduced that would, according to British media reports, “ensure that operators of web crawlers (internet bots that copy content to train GAI, generative AI) and GAI firms themselves comply with existing UK copyright law. These amendments would provide creators with crucial transparency regarding how their content is copied and used, ensuring tech firms are held to account in cases of copyright infringement.” There is lots of AI innovation ongoing in Britain.

The Australian Senate Select Committee Report on Adopting AI has recommended, among other findings, that there be mandatory transparency requirements and compensation mechanisms for rightsholders. The EU is already way out in front on this issue. Its new AI Act stipulates that providers of AI generative models will be required to provide a detailed summary of content used for training in a way that allows rightsholders to exercise and enforce their rights under EU law. Even India now has its own version of the US and Canadian media cases against OpenAI. (OpenAI’s defence in part is based on the argument that no copying took place in India because no OpenAI servers are located there!)

If that is what the “competition” is doing, who does Vigier cite as being the jurisdictions most likely to attract innovators away from Canada? Why, it is those AI powerhouses of Switzerland, Dubai—and the Bahamas!

The argument that if legislators and the courts don’t give AI innovators a free pass on helping themselves to copyrighted content for AI training purposes, this will either slow down innovation or chase it elsewhere is a common fearmongering strategy of the AI industry. This is a race-to-the-bottom mentality whereby content industries are thrown under the AI bus. Vigier, having been the subject of his own lawsuit, argues that instead of resorting to litigation, the Canadian media companies should have sought a licensing solution. But the fact that no licensing agreement was reached with OpenAI is undoubtedly the reason for the lawsuit in the first place. That is certainly the reason behind the NYT v OpenAI lawsuit in the US; licensing negotiations broke down. If someone has taken your content without authorization, and then offers you pennies on the dollar in comparison to what that content is actually worth, then the stage for a lawsuit is set.

In explaining CasewayAI’s position in the litigation brought by CanLII, Vigier says that Caseway approached CanLII with an offer to collaborate but was rebuffed. As a result they developed other extensive web crawling technology that pulled the needed material from elsewhere. (Where exactly the material was downloaded from is the crux of the matter). Regardless, this makes it sound as if it was CanLII’s fault for refusing to share their content. Surely a rightsholder has the right to determine the terms on which their content is to be shared with others, if at all.

The fact that Caseway went to CanLII in the first place suggests that CanLII had developed the content that Caseway wanted. Caseway claims the material it accessed was on the public record, such as court documents and decisions. CanLII, on the other hand, claims that it had reviewed, indexed, analyzed, curated and otherwise enhanced the content in question, thus adding a wrapping of copyright protection to what otherwise would be public documents. Who is right, and whether the material was scraped from CanLII’s website without authorization, will be determined by the BC Supreme Court.

If the material taken by CasewayAI was not copyright protected, they are in the clear, at least with respect to copyright infringement. That is quite different, however, from arguing that no copying takes place during AI training or that if rightsholders use the courts to protect their rights, Canada will be a laggard when it comes to AI development. Robust AI development needs to go hand in hand with robust copyright protection for creators, with an appropriate sharing of the spoils of the new wealth generated from the creative work of authors, artists, musicians and other rightsholders. To say, as Vigier does in his concluding paragraph that;

“Canada has a choice to make. Will we embrace AI as the transformative force it is, or will we let fear and litigation stifle innovation? The lawsuits against Caseway and OpenAI message tech companies: you’re not welcome here. If this continues, Canada won’t just lose its AI startups; it will lose the future of job creation.”

What sheer self-interested nonsense!. This is fearmongering of the worst kind, based on an inaccurate and misinformed knowledge of how AI is developed and trained, that moreover impugns the legitimate right of a rightsholder to seek the protection of the law to protect their creativity and investment in content. Vigier might be correct when he says that licensing of content is a win/win for both parties. I agree with that. But licensing negotiations are about money and conditions of use and require willing parties on both sides. When licensing discussions break down, or when one party decides to do an end run on licensing because they have been rebuffed, then the way to gain clarity is through the courts whose job it is to interpret what the legislation means.

Canada still needs to come to grips with the question of how copyrighted content will interface with AI development. As I noted earlier, both sides in the debate made their cases in the public consultation launched a year ago, but since then there has been no movement in Ottawa. The law could be strengthened to ensure adequate protection of rightsholder interests in an age of AI, resulting in facilitating licensing solutions. In the meantime, misinformation and scare tactics need to be called out for what they are.

Adequate protection for rightsholders does not mean the end of AI innovation or investment in Canada. There is no need for panic. We can walk and chew gum at the same time.

© Hugh Stephens, 2024. All Rights Reserved.

AI-Scraping Copyright Litigation Comes to Canada (CANLII v Caseway AI)

Image: Shutterstock (with AI assist)

It was inevitable. After all the lawsuits in the US (and some in the UK) pitting various copyright holders against AI development companies alleging the AI platforms were infringing copyright by reproducing and ingesting copyrighted materials without authorization to train their algorithms to produce outputs based on the ingested content–outputs that in some cases compete directly with the original work—AI scraping litigation has finally come to Canada. As reported by the CBC, CanLII (the Canadian Legal Information Institute), a non-profit established in 2001 by the Federation of Law Societies of Canada “to provide efficient and open online access to judicial decisions and legislative documents” is suing Caseway AI, a self described AI-driven legal research service, for copyright infringement and for violating CanLII’s Terms of Use through a massive downloading of 3.5 million files.

In its civil claim brought before the Supreme Court of British Columbia, CanLII alleges that the defendants, doing business as Caseway AI, violated its Terms of Use that prohibit bulk or systematic download of CanLII material and that in doing so, the defendants also engaged in copyright infringement by reproducing, publishing and creating a derivative work based on the copied works for the defendants own commercial purposes. There is no question that Caseway is providing legal material for commercial gain. Caseway’s services start at $49.99 a month , or $499.99 a year, and offer an AI driven service that “leverages advanced AI to find relevant case law in less than a minute… Designed with a user-friendly chatbot interface powered by proprietary technology, Caseway (is) a robust tool tailored specifically for the legal profession.” Caseway’s Terms of Service have all sorts of disclaimers, however.

In his defence, Caseway’s Canadian principal (and defendant) Alistair Vigier is reported to have said that “court documents are public record, not owned by any organization, including CanLII. Numerous other websites also make these decisions available.” It is true that court documents and decisions are public documents not subject to copyright protection. However, CanLII claims that its database contains more than just the court’s decisions. It says in its claim that it spends significant time to “review, analyze, curate, aggregate, catalogue, annotate, index and otherwise enhance the data” prior to publication. It is this creative effort that turns public documents into a copyright protected document (or so the argument goes). To use another copyright analogy, you cannot copyright a recipe (a “list of ingredients”) but we all know that cookbooks containing recipes are always copyrighted. This is because of the display and illustrations of the recipes, the layout, commentary and other editorial touches. Julia Child’s sole amandine recipe is not just any old recipe for fried sole. Is CanLII’s compilation of “judicial ingredients” protectable? We will have to wait to find out.

CanLII’s case is reminiscent of a similar case in the US, Thomson Reuters v Ross Intelligence. Thomson Reuters operates a subscription-based legal research service called Westlaw. One of Westlaw’s employees allegedly copied Westlaw content to enable Ross Intelligence to build a machine learning platform that competed with Westlaw. Part of Ross’ defence was that the judicial decisions themselves are public domain documents, so there could be no infringement. Westlaw maintained that its case head notes, summaries that described the cases, were copyrightable material. Ross also brought forward a fair use defence arguing transformation, i.e. they had produced something new and different that did not compete directly with Westlaw’s product. Here is a good summary of the case. The court determined that Ross had copied the headnotes but the copyrightability of Westlaw’s numbering system and headnotes needed to go to a jury to determine. While Ross’ anti-trust case against Westlaw has been dismissed, the copyright case is still pending.

Another case that has been cited as a possible precedent is the famous 2004 CCH Canadian Ltd v Law Society of Upper Canada case in which the Supreme Court of Canada ruled that copies of CCH materials made by the Law Society library for its members did not infringe CCH’s copyright because the library was exercising the fair dealing research exception on behalf of the individuals requesting the copies. I personally don’t see the relevance of this case (but I am not a lawyer) since the Great Library’s users were copying only relevant parts of certain documents, for a specified fair dealing purpose. In the CanLII case, Caseway has apparently inhaled the full collection of documents and is doing so for a commercial purpose, with the resultant product (although not identical to the original) competing with it. Moreover, since there is no text and data mining exception in Canadian law, the “transformation” defences available to US-based AI companies (i.e transforming the original materials to produce something different) are not applicable in Canada. This will be an interesting one for the lawyers.

What the case demonstrates is a crying need for some legislative guidance on the question of AI scraping of copyrighted materials in Canada. It may be that CanLII’s collection cannot be protected by copyright, which would provide Caseway a defence without settling the fundamental issue of whether it is a violation of the Copyright Act to do what Caseway did, assuming the material they used was protectable by copyright. A consultation exercise was launched by the government of Canada (through the Ministry of Innovation, Science and Economic Development, ISED) last October, closing in January with submissions posted in June. Since then, there has been silence on the part of the government. With Parliament at a standstill, and the current government hanging on to power by its fingernails, don’t expect clarity any time soon.

© Hugh Stephens, 2024. All Rights Reserved

Canadian Copyright Registration and AI-Created Works: It’s Time to Close the Loophole

Image: Shutterstock

In July, the Canadian Internet Policy and Public Interest Clinic (CIPPIC) at the University of Ottawa filed an application in the Federal Court to expunge or amend a Canadian copyright registration that claimed an AI program, the RAGHAV AI Painting App, as co-author of a registered work. While the other co-author, an Indian IP lawyer by the name of Ankit Sahni is named as the respondent, the real defendant ought to be the Canadian Intellectual Property Office (CIPO), the organ within the Department of Industry (ISED) responsible for managing copyright registration. It is CIPO’s “rubber stamp, content-blind, absence of judgement” automated system of registration that has led to this situation, putting Canada in a significantly different place from that of the United States or many other countries when it comes to granting copyright protection to works produced by AI algorithms with no or little human intervention.

Last month I wrote a couple of blog posts on the issue of whether content produced with or by generative AI could or should qualify for copyright protection. I looked at the ongoing uphill struggle that two “creators”, Stephen Thaler and Jason Allen, have experienced with the US Copyright Office (USCO) in their attempts to get the USCO to register their works. Thaler claims his submitted work (“A Recent Entrance to Paradise”) was created exclusively by his AI algorithm (the “Creativity Machine”) and, accordingly, it should be recognized as the “author”. However, as the human behind the machine, having invested in creating it, the benefits of the registration should fall to him. He argues that the algorithm carried out the work at his behest, much like a work for hire. Allen, by contrast, claims that although his award-winning work (“Théâtre D’Opéra Spatial) was produced with AI assists, he was the creator through control and manipulation of the prompt process. In neither case has the USCO budged from its position that the works do not qualify for copyright protection on the basis they were not human-created. The same goes for the courts to which the USCO’s rejection has been appealed.

That is the current situation in the US; in Canada it is quite different. Works produced exclusively with AI have been accorded copyright registration, more than once. I have even done it myself! (See “Canadian Copyright Registration for my 100 Percent AI-Generated Work”).

Because Canadian copyright registration is automated and done through a website, an applicant must provide the author’s address, contact details and date of death, if deceased. To work around that, I clearly specified that the work was created entirely by two AI programs (DALL E-2 and CHAT-GPT) with virtually no exercise of “skill and judgement” on my part, this supposedly being the threshold in Canada for creative content that can be afforded copyright protection.

This is how my Canadian copyright certificate No. 1201819, issued April 11, 2023 (I should have tried to register it on April Fools Day), reads in terms of describing the registered work;

“SUNSET SERENITY, BEING AN IMAGE AND POEM ABOUT SUNSET AT AN ONTARIO LAKE CREATED ENTIRELY BY AI PROGRAMS DALL-E2 AND CHATGPT (POEM) ON THE BASIS OF PROMPTS DEMONSTRATING MINIMAL SKILL AND JUDGEMENT ON THE PART OF THE HUMAN AUTHOR CLAIMING COPYRIGHT”.

While this little exercise in inanity was fun, (and was done to expose the failings of the current system), I was not the first to register an AI created work in Canada. That honour, as far as I can tell, belongs to Sahni, the named respondent in the CIPPIC case who, in December 2021, managed to register the artistic work Suryast, listing the AI-powered RAGHAV painting app as co-author. That was a neat way of getting around the requirement to provide an address, contact details etc. Sahni could provide his contact details yet still claim the AI algorithm was an author, even if a co-author. Clever. What Sahni’s motivation was I cannot say, but apparently he has been active in registering the work in as many jurisdictions as he can, maybe to boost the marketability of RAHGHAV. CIPPIC claims he is seeking registration to force various countries to address the AI authorship issue. Canada must have been one of the easiest registrations he received. Now he is being called to account. The application brought by CIPPIC seeks a declaration either that there is no copyright in Sahni’s image, Suryast, or, alternatively, if there is copyright in Suryast, that the Respondent (Sahni) is it sole author. It also seeks an order to expunge the copyright certificate in question or to rectify it by deleting the painting app as a co-author.

The fundamental problem of course is not Sahni or his AI app, (although like me, he may have been mischievous) but rather the way in which copyright registration is offered and maintained in Canada. It was not always this way. Once upon a time, to register a work in Canada you were required to not only pay a registration fee, (which is still the case today) but submit three copies of the work, one for the Copyright Branch (which was part of the Department of Agriculture), one for the Canadian Parliamentary Library and one for the British Museum. Because of these depository requirements, today we have a record of many early copyrighted works in Canada, such as the famous early 20th Century Inuit photographs of Canada’s first professional female photographer, Geraldine Moodie, about whom I wrote a few years ago (“Geraldine Moodie and her Pioneering Photographs: A Piece of Canada’s Copyright History”).

When the first international copyright convention, the Berne Convention of 1886, was established among a limited number of countries, there was a push by authors to abolish the registration requirement because it was burdensome to have to register in all Berne countries. Initially, registration in the home country was supposed to provide protection in all member states of the Convention, but this proved difficult to put into practice. Consequently, in 1908 at the Berlin revision of the Convention, the following provision (which is today part of Article 5(2) was adopted, “The enjoyment and the exercise of these rights shall not be subject to any formality”. Canada was a member of Berne because Britain had acceded, but was nonetheless a reluctant conscript (even though then PM Sir John A. Macdonald had acquiesced to Canada’s inclusion). In 1921 Canada finally passed its own Copyright Act (coming into force in 1924, a century ago this year), and subsequently joined Berne in its own right in 1928. I suspect that registration as a requirement, along with depository and examination conditions, was dropped at that time. That is probably when the current (but non-automated) voluntary registration process was established.

Certainly such a system was in place in the early 1950s when broadcaster Gil Seabrook of Vernon, BC registered an “untitled and unpublished artistic work” entitled “Ogopogo”. The registration of that undocumented work became the source of the urban myth that the City of Vernon owned the intellectual property rights to the mythical lake monster Ogopogo (Seabrook had donated his copyright to the City in an attempt to upstage Vernon’s rival town to the south, Kelowna, that claimed it was the “home of Ogopogo”). As a result, in 2022 Vernon Council went to great lengths to “return” the rights to Ogopogo to the local First Nation as an act of “reconciliation”. Of course, they never had the rights to Ogopogo in the first place. If you want more information, you can read all about it here. (“Copyrighting the Ogopogo: The © Story Behind the News Story”).

Despite the abolition of a registration requirement by Berne Convention countries, Canada is not the only country that maintains one. In the US, which only joined Berne in 1989, both registration and renewal were required for a work to enjoy copyright protection. When the US joined Berne, it maintained the registration requirement for US citizens who wished to take legal action to enforce their copyright. This is allowed under Berne. As such, the US has maintained a robust registration system where a legal deposit of the work is required, registrations are examined and can be challenged or refused.

We know that is not the case in Canada, but Canada is not the only country to have a voluntary registration system. In a recent study by WIPO (World Intellectual Property Organization), some 95 countries were identified as having either a voluntary registration system, a recordation system (for transfer of copyrights) or a legal deposit requirement. What is notable, however, is that of all these countries, only three (Canada, Japan and Madagascar) do not require a deposit of the work seeking registration. Canada does review applications but only to ensure they meet all the formality requirements (name and address of the owner of the copyright; a declaration that the applicant is the author, owner of the copyright or an assignee; the category of the work; its title; name of the author and, if dead, the date of the author’s death, if known. For a published work, the date and place of first publication must be provided and, perhaps most important, payment of the prescribed fee). Nothing else. In fact, if Mr. Mickey Mouse, address Disneyland Way, filed a copyright application for a work and paid the required fee of $63, I am sure a Canadian copyright certificate would be issued. It used to come in the mail, printed on nice quality paper but, alas, in the interests of efficiency, it is now only available in PDF format on CIPO’s website. Print it yourself.

That is the current situation, but why has CIPPIC gone to the Federal Court to dispute the wording of Sahni’s copyright certificate, No. 1188619? While the Registrar of Copyrights can accept requests for correction of a copyright certificate (either because of an error in filing or because the Office itself made a mistake), it cannot by itself amend or remove a registered work from the Register. Instead, the Registrar needs the Federal Court to effect such action. Section 57 of the Copyright Act states, with respect to Rectification of Register by the Court;

(4) The Federal Court may, on application of the Registrar of Copyrights or of any interested person, order the rectification of the Register of Copyrights by
(a) the making of any entry wrongly omitted to be made in the Register,
(b) the expunging of any entry wrongly made in or remaining on the Register, or
(c) the correction of any error or defect in the Register

However , while CIPPIC is seeking expungement of this particular copyright registration, it is the system it is really going after. This is clear from its memorial to the Court;

(23) “In automating its copyright registration process, CIPO is derogating from its obligations to administer copyright in a fair and balanced manner under the Copyright Act.”
(24) “The consequence of this system is that content that does not merit copyright can…easily obtain the benefits of registration.”
(25) “Copyright registrants obtain certain benefits under the Act – such as litigation presumptions – and users and defendants are correspondingly burdened. Once a “work” is registered, the Copyright Act…shifts certain presumptions such as subsistence and ownership….In this very case, as a result of CIPO’s oversight failures, the burden rests on CIPPIC to prove the image Suryast lacks originality and that an AI program cannot be an author.”

Moreover, CIPPIC notes that it brought this case to the attention of CIPO but it refused to correct the Copyright Register, instead encouraging CIPPIC to seek resolution in court. Assuming it is granted standing, CIPPIC may well prevail and have the Suryast registration amended or expunged. But will that really achieve its goals? If its goals are to get CIPO to stop “derogating from its obligations”, then simply cancelling or amending this one registration won’t do it. What is the solution?

One option would be to eliminate the voluntary registration requirement altogether, but is this the right course of action? The WIPO document referenced earlier points out some of the advantages of a voluntary registration system. It can ensure that information about authorship and copyright, including date of registration, become publicly available. This benefits not only authors and rightsholders, who can use the registration as a rebuttable presumption of copyright in court, as in Canada, but also provides information to the public to verify ownership claims and trace title. A voluntary system does not, however, provide a definitive list of what works are under copyright and which are not. Another factor is that an automated voluntary system, such as the one operated by CIPO, is not burdensome for registrants and presents no meaningful obstacle. The problem is that its barriers to registration are so low that it is easy to trick the system. Is a Canadian copyright certificate worth the paper it is printed on if there is no verification?

A second option is to improve the registration process to make it meaningful, but this will require resources. Current fees are low (but the US system which is much more robust has a similar fee structure). Nonetheless, to institute a USCO type system would require substantial additional resources that are unlikely to be forthcoming in the present fiscal environment. One would have to ask whether the extra cost could be justified. It’s a conundrum. Meanwhile, the government has circulated a paper on the issue of Copyright and AI and the Canadian cultural community has weighed in with its views. Prominent among these is the position that copyright protection should be accorded only to human-created works. (This is not currently specified in the Copyright Act).

CIPPIC’s court action puts the spotlight on the current copyright dilemma. The current system seems to be not fit-for-purpose, but an economically viable alternative is not immediately apparent. At the very least, Canada should amend the Copyright Act to prevent AI-created works from obtaining copyright registration.

© Hugh Stephens, 2024. All Rights Reserved.

If AI Tramples Copyright During its Training and Development, Should AI’s Output Benefit from Copyright Protection? Part Two: Jason Allen

Image: Théâtre D’Opéra Spatial, Jason B. Allen (not protected by copyright)

Last week I wrote about Stephen Thaler’s quixotic and determined approach to obtain copyright registration in the US for his AI generated artwork, “A Recent Entrance to Paradise”, created (he claims) exclusively by his AI “machine”, the so-called Creativity Machine. So far, despite repeated efforts, he has drilled a dry hole. An alternative approach to claiming copyright for an AI-generated work is by asserting that the AI used to produce it was simply a technological assist. The essence of the work was produced through human creativity, using AI only as a tool, and therefore the work should be eligible for copyright protection, or so the argument goes. Unlike the example of Kristina Kashtanova, discussed in last week’s blog, under this theory the entire work is protectable because it was human created, with AI playing only a facilitating or assistive role. This line of attack has most recently been pushed by the creator of the work “Théâtre D’Opéra Spatial”, Jason B. Allen.

Allen made headlines a couple of years back (September 2022) when he entered Théâtre into the Colorado State Fair’s annual art competition in the category of “digital art/digitally manipulated photography.” He labelled the piece as having been created by him, “via Midjourney”, the popular generative AI art algorithm that had recently been released. He won first prize, incurring the opprobrium of many artists who accused him of crashing a contest for human creators. Writing just a month later, in October of 2022, I posted my own AI produced artwork, based on the style of Monet (whose works are in the public domain) in this blog post, (AI and Computer-Generated Art: Its Impact on Artists and Copyright). My effort was substantially less artistic than Allen’s but was an original work of sorts, created with the help of AI. I used the program DALL-E2, which is similar to Midjourney. Both were freely available and a tool that any rank amateur “artist” (like me) could use.

Generative AI as a source of art burst into the public’s consciousness in 2022 because of the public release of these programs, but AI generated art has been around for a few years before that although used exclusively by art specialists. The New York Times reports on a sale at Christie’s Auction House four years earlier, in October 2018 (not exactly eons ago, but generations in internet/AI time). A portrait with blurred and distorted lines produced by an AI algorithm sold in New York for $432,500 (with fees). Christie’s, in inimitable auctioneering style, billed it as “the first portrait generated by an algorithm to come up for auction”, according to the Times. Now AI generated art is a dime a dozen. In fact, often it is hard to tell what is AI generated and what is not.

If the question of copyright protection for AI generated content is a big issue, an even bigger one is currently being played out currently in the courts; content owners, ranging from the New York Times to Getty Images to music labels to authors, are suing various AI development companies for unlicensed use of their content to train AI programs. The Copyright Alliance has a good summary of the various lawsuits in play here. The ultimate outcome is undecided, but if the courts find that the wholesale unlicensed and unauthorized ingestion of copyrighted content to train AI algorithms is not fair use (in the US) or does not fall within specified text and data mining exceptions in other jurisdictions, then the table will be set for serious negotiation between rights holders and AI developers. Some of the parameters of this negotiation are already pretty obvious;

  • a transparent inventory of what copyrighted works were accessed for training;
  • the ability of rights holders to be able to opt out or opt in;
  • various options for licensing content for training purposes.

If these conditions governing inputs were met, rights holders might be somewhat more sympathetic to arguments for copyrighting the output of generative AI programs. As it is, AI developers, and the users of AI programs, want to have their cake and eat it too. Jason B. Allen of “Théâtre D’Opéra Spatial” is Exhibit No. 1.

Allen is currently appealing in court the US Copyright Office’s ruling that his work cannot be registered under copyright. He claims that because of all the publicity about his work, and the USCO’s subsequent decision to deny copyright registration on the grounds that it was an AI generated rather than human creation, the work has lost value and impacted his ability to charge industry-standard licensing fees. Moreover, he claims that without copyright protection, he has no ability to stop others from using his work without authorization. (Like me, posting Théâtre on this blog post). Apparently, people are selling copies of the work on Etsy. One has some sympathy for his position, as it is one faced by many artists whose copyrighted works are also being ripped off on the internet.

Allen argues he had substantial creative input into the production of the work, using no fewer than 624 prompts to create the work to his mental specifications. Of course, we have no idea of what those specifications were. What Midjourney produced may have been an accurate reflection of what Allen had in his mind and intentionally created, or it may have taken him on a journey where he eventually settled on the output offered. One thing is likely, if not certain. Were he to enter those exact same prompts into Midjourney today, the outcome would not be identical to the current “Théâtre D’Opéra Spatial”. This raises the question of who, exactly, is guiding the creative process, the artist making the prompts, or the algorithm responding to the prompts.

Because of the way AI works, there is a large degree of randomness in the results, requiring more and more precise prompts to narrow the range of possibilities and guide the algorithm to the desired destination. But it is almost impossible to recreate precisely the route to an outcome. This suggests to me that, in the end, it is the algorithm that is in control, not the human issuing the prompts. (Although not everyone agrees with this thesis). To my mind, this is what distinguishes AI-generated art from photography, where the photographer, while using a mechanical assist, is nevertheless in full control at all times and has the ability to adjust for extraneous inputs such as light, shadow etc. rather than being controlled by them.

Allen’s appeal of the USCO’s rejection of his copyright claim takes place against a backdrop where Midjourney, the AI program he used, is itself being sued by a group of artists for appropriating their work without permission in order to train Midjourney’s art generation algorithms. Does anyone see any irony in this? However, the fact that Midjourney takes the works of others without permission for training and development purposes is not really Allen’s fault. He and other users of the program could perhaps be considered victims almost as much as the artists whose works have been appropriated. Nevertheless, if Allen is ever successful in getting Théâtre registered, this will be not just to his benefit, but also to the benefit of Midjourney and all the other AI developers who are in a similar position. On the other hand, if the output of their programs cannot be copyright protected, it diminishes the value of the AI product. So, perversely, the more Allen pursues registration for Théâtre, the more he undercuts those who make a living from producing art.

I see one possible scenario that could help resolve the issue of AI outputs being unprotectable. If the key elements of respect for copyrighted work (transparent inventory, opt in/out, licensing) were to be adopted by AI developers, then perhaps rights-holders would be more amenable to accepting at least some degree of copyright protection for AI created or assisted outputs. But right now, the AI industry wants it both ways; total freedom to appropriate copyrighted works for training and development purposes while claiming the same copyright protection they have just trampled for AI generated outputs. Jason Allen and other digital artists who use AI to produce art or other works are caught in the middle.

At the moment there is no clear solution. The most likely outcome–after all the legal dust has settled—is probably going to be some ability to copyright works produced with AI, dependent on the extent and degree of human intervention in a given work (which could possibly be carefully tailored prompts), balanced by a commitment by the AI industry to recognize the property rights of those holding copyright over the content it is using to create the AI program in the first place. This will necessarily involve an appropriate sharing of the added value being produced by AI in the form of licensing fees. This will take a few more years, a few more lawsuits, a few court decisions, and a few government interventions in the form of legislation–but I can see no other way forward. In the interim, neither Stephen Thaler nor Jason Allen are likely to get what they want.

© Hugh Stephens 2024. All Rights Reserved.

If AI Tramples Copyright During its Training and Development, Should AI’s Output Benefit from Copyright Protection? Part One: Stephen Thaler

” A Recent Entrance to Paradise”, Stephen Thaler (not protected by copyright)

One of the ongoing debates about works made with generative AI is whether they qualify for copyright protection. Should they? Let’s consider the essence of copyright. What is its raison d’être? According to the classical European definition, it is to respect the property rights of the author (droit d’auteur), sometimes described in the simple terms of the Eighth Commandment (“Thou shalt not steal”). According to the more utilitarian Anglo-Saxon rationale for copyright, it is to benefit society by rewarding and incentivizing authors, thus stimulating further production of works for the greater good. In either case, it gives one pause to wonder how works created by AI fit into either school of thought. Is there an inherent property right in content produced by an algorithm? How does copyright protection incentivize an algorithm to produce more “useful arts”?

In my view, the only way one can square this circle is by attributing human creation, or at least a degree of human creation, to AI generated works. But this opens Pandora’s box; if there is to be human attribution, to whom in the chain of creation does the credit fall? How much human creativity is required? It also raises the spectre of hypocrisy, as the AI industry hijacks the creative work of others, without recognition or recompense, yet has the gall to claim that AI outputs are unique and worthy of protection. I have written about these issues before (here, here, and here), but a couple of recent cases in the US have brought these fundamental issues bubbling back to the surface.

When it comes to trying to prove that AI and copyright protection go together, there are a couple of different approaches. One, most notably espoused by an AI technologist in the US named Stephen Thaler, is to claim that a given work was produced exclusively by AI but should nonetheless be protected. In Thaler’s case, the work for which he is seeking copyright registration was created by a particular AI “machine” (or algorithm), specifically the one he “invented”, the so-called “Creativity Machine”. Thaler claims his machine should hold the copyright but behind the machine, of course, stands Thaler. This is not dissimilar to existing British copyright law where, under Section 178 of the Copyright, Designs and Patents Act, 1988, works “generated by computer in circumstances such that there is no human author of the work” are nevertheless accorded copyright protection for fifty years from date of creation, with the copyright being held by “the person by whom the arrangements necessary for the creation of the work are undertaken”, even if there was no creative act undertaken by that person.

As I noted in an earlier blog post (The Humanity of Copyright), Thaler began his (so far) unsuccessful pursuit of US copyright registration for his professed 100% AI generated art work, “A Recent Entrance to Paradise”, back in 2018. Despite several reversals, both in the application process at the USCO, its Review Board and in the District of Columbia courts, Thaler persists in his quixotic journey. He has unsuccessfully argued various precedents for non-human copyright ownership, such as the “work for hire” doctrine and corporate copyright, although it is worth noting that humans stand behind both. He is now apparently pursuing the common law theory of “fruit of the tree” in his attempt to get the USCO to register his work. To my non-legal mind, this is the ultimate stretch.

What Thaler could do is to claim that at least part of the work is the result of his personal creative efforts. That was the USCO outcome for the graphic novel Zarya of the Dawn produced by writer Kristina Kashtanova (who identifies as “they”). The novel contained both generative artwork and human story and design elements. After initially registering Kashtanova’s work, the Copyright Office cancelled the registration after they (Kashtanova, that is) claimed it was AI produced. The Office subsequently reconsidered and granted copyright protection to the parts of the work Kashtanova had created, namely the text, and selection and arrangement of the work’s written and visual elements. That, however, is a step too far for Thaler who continues to push for recognition by the Copyright Office that a work produced exclusively by his “Creativity Machine” can be protected by copyright. That seems very unlikely to happen.

While Thaler doggedly pursues copyright registration for “A Recent Entry to Paradise” (featured on this blog post—after all, it is not copyright protected), others who have created art using AI are following a different track. One of these is Jason B. Allen, whose award-winning digital art creation “Théâtre D’Opéra Spatial” has also been denied copyright registration. Allen’s approach is the opposite to that taken by Thaler. In contrast to Thaler’s insistence that the work is a creation of AI (his AI “machine”), Allen insists that he is the source of the creative inspiration behind the work, notwithstanding that it was created by an AI algorithm, Midjourney. Allen’s pursuit of copyright registration for “Théâtre” will be the subject of my blog post next week.

© Hugh Stephens, 2024. All Rights Reserved.