In early June Canada issued its national AI strategy paper, “AI for All”. As I noted in a blog post at the time, while the strategy covered many elements of AI in its 50 pages outlining policy objectives and planned actions, it managed to avoid using the word “copyright” even once. Australia has just come out with its own updated AI policy statement “AI in Australia’s interest”, which builds on its own “National AI Plan”, released last December. But whereas the Carney government in its AI strategy managed to completely avoid putting copyright into the AI equation, Prime Minister Albanese, after discussing the importance of developing AI for Australia, had this to say;
“But let me make this crystal clear: not everything produced in Australia is up for grabs.
Not at all.
Australian writers, musicians, artists and journalists must retain ownership and control of their work.
Our laws will spell that out, plain as day.
An artist’s creative endeavour is their work and their property.
No company should use Australian books, music, art or news to build or train AI without the artist’s control.
That includes the artist’s control of the price and value of their work.
Anything less, is theft.”
Blunt, clear and refreshing. If Australia can protect its cultural community while promoting policies for sensible AI adoption and development, then so can Canada.
Both Canada and Australia currently have no Text and Data Mining (TDM) exception in their copyright law. This legal loophole would allow AI developers to appropriate content without permission for training purposes. In both countries there have been calls from the tech community to introduce a TDM exception, a carte blanche that would allow AI companies to ingest copyrighted content without authorization, payment or even acknowledgement. In its December “National AI Plan”, which is much more analogous to Canada’s “AI for All” than Albanese’s recent short AI policy statement–in that it outlined a range of detailed policy proposals for AI adoption in Australia– the Australian government nonetheless managed to grasp the copyright nettle unambiguously.
Among the issues highlighted under “AI Risks and Harms” was the following:
“Reviewing application of copyright law in AI contexts: The Attorney-General’s Department is engaging with stakeholders through the Copyright and AI Reference Group to consult on possible updates to Australia’s copyright laws as they relate to AI. The government has provided certainty to Australian creators andmedia workers by ruling out a text and data mining exception in Australian copyright law”(emphasis added)
Just as the Australian government has sensibly ruled out a TDM option. Canada needs to do the same, as called for Canadian cultural umbrella groups, such as the Coalition for Diversity of Cultural Expression (CDCE).
So far Canada has danced around the issue. Heritage and Identity Minister Marc Miller has said that “the current copyright law does and should protect those that have created material, and people need to be compensated properly”, but he is just one minister among several. Evan Solomon, Minister of Artificial Intelligence and Digital Innovation, and Minister of Industry Melanie Joly, both have a big piece of this file. One can expect that both can be counted on to be more sympathetic to tech bros than cultural mavens. What is needed is a prime ministerial pronouncement clarifying that Canada’s creative community–artists, writers, publishers, musicians, filmmakers, photographers, journalists and more– is not going to be thrown under the bus on the pretence of keeping Canada competitive in the global AI game.
In the wake of Australia’s announcement that a TDM exception was off the table, the tech industry tried a new approach by suggesting the creation of a centralized fund that would be used to compensate rightsholders for the permissionless use of their works in AI training. Specifically, AI company Anthropic reportedly tied a proposed $15 billion USD ($21.6 billion AUD) investment in data centres in Australia to creation of the creatives fund in order to allow to access Australian content without licensing or negotiation with rightsholders. Australia’s creative community quickly mobilized. Their concerns were heard. Along with setting clear guardrails ruling out the unauthorized use of copyrighted creative works, Albanese has created a new Office of AI within the Prime Minister’s Office, recognizing the need for policy coordination given the breadth of AI’s policy impact. This is something that Canada might consider. It has Evan Solomon, Minister of Artificial Intelligence and Digital Innovation, but there seem to be very few cultural community voices within Solomon’s hearing range.
Australia has the same goal as Canada of getting its fair share of the AI pie while managing AI adoption and its impact on society. But there is one big difference. In so doing, the Australian government has made it clear it will pursue its AI goals while simultaneously respecting and protecting its culture and its creators. Canada’s cultural and creative community deserves no less consideration.
There is an ongoing struggle between the tech world of AI training and the cultural world of content creation. It has led to lots of litigation but also an increasing number of licensing agreements, the obvious market solution. Litigation has helped convince AI companies to share some of the wealth by pursuing licensing. Yet the AI world continues to try to find ways to avoid the basic step of seeking permission from rightsholders for using their valuable content to create their products.
Anyone who has seen the striking graphic “Who is Suing Whom in AI”, created by the design website Information is Beautiful, will be struck by the enormity and breadth of the issue which is so cleverly displayed, with the big AI developers such as Perplexity, Anthropic, Meta, Google, Open AI, Midjourney, Cohere and others at the centre with the creators (every content entity from Conde Nast, Getty Images, Universal Music Group, CNN, Disney and Thomson Reuters to Elsevier, Dow Jones, New York Times and others) ranged around the periphery, a stunning visual encompassing more than 100 lawsuits in the United States. That graphic was up-to-date as of June 26 of this year. Since then, at least one more major lawsuit has been filed, by a group of textbook authors against Meta. The graphic does not include the first such case in Canada where a group of media organizations (Canadian Press, Torstar, The Globe and Mail, Postmedia and CBC/Radio-Canada) is suing OpenAI, or the Getty Images case in the UK, or indeed any cases outside the US. From this graphic, it would seem that to resolve the issue of how copyrighted content is going to be used in AI development and training, litigation is the inevitable route. But is it?
As far as I am aware, Information is Beautiful has not created a similar graphic to display the range of licensing deals that have taken place, many of them between some of the same actors that appear on the litigation chart. If they did it would be similar, but encompassing even more licensing agreements than lawsuits. Licensing deals are being struck so frequently it is just as hard to keep up with them as it is to track all the litigation underway. The University of Glasgow’s CREATe Centre says it has documented 274 licensing deals and has a chart that tracks 109 of them. Whatever the number, it is a lot and it is growing. That is not to say that the AI industry has finally accepted the need to pay for the content they are using to create their products, just as they pay for software engineers or data processing capacity. This is where the link between litigation and licensing becomes interesting.
In a perfect world, AI developers would obtain their inputs through the market on the basis of permission, which would encompass both compensation (in most cases) plus transparency or accountability, i.e. documenting what content was used. But we don’t live in a perfect world, which is why we have the rule of law and courts to enforce those laws. In some cases, AI platforms did begin negotiations with rightsholders but when it was not possible to reach an agreement, the AI industry switched tactics and took the content anyway, arguing it was legal to do so for a variety of reasons. This is precisely the scenario that led to the New York Times suing OpenAI. These cases are even more egregious because there was initially a tacit acknowledgement by the user that the content had value. Then, when the price or conditions did not suit the potential licencee, suddenly it was okay to take the content anyway under the guise of fair use. Various arguments have been deployed ranging from the claim that no copying actually occurs, to the dubious assertion that what is copied is data not content, to the invocation of the US “transformation” doctrine.
On the issue of copying, a study by the Atlantic (AI’s Memorization Crisis: Large language models don’t “learn”—they copy. And that could change everything for the tech industry) convincingly demonstrated the uncomfortable truth that LLMs can reproduce long excerpts from books they have been trained on. The inputs are not just ones and zeros, they are content— someone else’s content that was taken without permission. Whether the use was fair according to US fair use interpretations is still an open question. US courts and other countries are trying to come to grips with this issue. In countries such as Canada or Australia, where there is no statutory copyright exception for Text and Data Mining (TDM) that would permit permissionless AI training on content, the AI industry has been floating various workaround proposals. The “incentives” would include (in Australia) establishing a government-managed fund to compensate rightsholders according to some sort of formula, plus investments in AI data centres. What is missing from proposals such as this is the concept of permission from those who actually own the content, or even discussion of the proposal with them. As Prof. Rod Sims, former Chair of Australian Competition and Consumer Commission, put it in a recent opinion piece in Canada’s National Post, “what other sector refuses to negotiate with suppliers and instead goes to government to bypass such a step?”
Let me use a food industry analogy to make the point even more clearly. When you run a restaurant you have labour costs, rent, taxes, etc. and the cost of ingredients to consider. You don’t get to raid the farmer’s field to obtain your inputs for free, just because you are able to root out crops without the farmer being able to stop you or even know it is happening. Setting up a fund to “compensate” farmers for their stolen crops, on terms set by the government rather than the market, doesn’t even begin to make this right. Legalization of this theft would remove any possibility of litigation or legal protection, for the farmer—or for content owners. Litigation, while protracted, costly and potentially leading to uncertain outcomes, is nonetheless the stick that is needed to facilitate licensing.
The obvious route for the AI industry to take is to license the content they want to use. That may not seem as “efficient” as just taking it for free but with the threat of litigation hanging over the proceedings, licensing suddenly becomes the more efficient alternative. It is also win/win for both AI developers and the content industries. And, it is simply the “right thing to do”.
I am pleased to note that this blog was recognized by Feedspot as being among the “40 Best Copyright Blogs to Follow in 2026”. In fact, we hit the middle of the pack at No. 20. I am honoured to be included in such distinguished company.
Feedspot is an RSS Reader that lets readers subscribe to blogs, news sites, and any website they wish to follow.
If it does, you would never know it from reading Canada’s new AI strategy just released by the Minister of Artificial Intelligence and Digital Innovation, Evan Solomon. It is a magistral document, addressing key elements of AI under six pillars: (with my shorthand summary in brackets)
Protecting Canadians and safeguarding democracy (addressing trust, safety and privacy concerns)
Empowering Canadians (promotingAI literacy and economic opportunity)
Powering AI adoption for shared prosperity (accelerating adoption, especially for SMEs)
Building a sovereign AI foundation (building domestic compute, cloud and connectivity infrastructure)
Scaling Canadian champions (more government funding for domestic AI development)
Building trusted economic and governance partnerships and global alliances (leading the creation of a multinational middle power alliance to curb the power of hegemons and hyperscalers)
The latter objective will no doubt go down really well with the Trump Administration!
Those six headings cover just about all aspects of AI, from its creation to its use to its impact on the economy, on society and on individuals. But in all 50 pages of the document, as far as I can ascertain, you won’t find the word “copyright”, although “protecting intellectual property” is certainly featured. The intellectual property rights that are mentioned have nothing to do with the rights of those whose content was used without authorization to create AI but rather relate to protecting the intellectual output of AI developers in Canada. John Degen, CEO of the The Writers’ Union of Canada (TWUC) was the first to call this out. Given the make-up of the task force that produced the report, this is not surprising. While it was made up of the great and the good from the AI world, with academics, financiers, CEOs, cybersecurity experts, innovators, educators and so on as part of the roster, there was not a single representative from the cultural community.
There are many elements of AI this document tries to address, all of them important to a country like Canada, although there are limits to what can be done by a middle power given that the lead on development has been seized by a handful of large companies, mostly in the US. The US government itself is caught in the dilemma of wanting the US to lead AI development yet not becoming overwhelmed by it to the point that a few major corporations are calling all the shots.
As for content issues, including what must surely include some copyrighted content, they are addressed only indirectly in the Canadian strategy. The three principal issues relating to content are; (1) privacy and access to data; (2) Canadian identity and culture; and (3) AI misuse, such as creation of deepfakes and misinformation.
On privacy and data, the document notes that AI is only as powerful as the data it can access (how true!). It reminds us that governments in Canada hold vast amounts of data that should be treated as a strategic national asset and mobilized to fuel innovation and productivity (i.e. provided for AI research). Thankfully, there is a tip of the hat to the need for “strong privacy protections” but there is no mention of the unauthorized scraping of databases and protected content by AI developers, both domestic and international. Privacy is important but so is ownership of content, and the right to grant permission to use it. Unfortunately, this latter point is not mentioned.
Protecting and promoting Canadian identity and culture is also mentioned as an important goal. It is obvious that if AI developers are blocked or hindered from ingesting Canadian content, then there will be less of Canada reflected in AI outputs. That argument was put forward recently by Michael Geist in a blog post criticizing recommendations issued by the Parliamentary Standing Committee on Heritage that had called for protection of the property rights and interests of artists through the Copyright Act on the basis of authorization, remuneration and transparency. This would lead to “AI without Canada”, according to Prof. Geist. This could be true if AI developers did not need or want curated Canadian content, but they do. The solution, as I pointed out, is not to give away everything in the shop window by creating a broad AI training exception in Canadian copyright law–which would amount to legalized theft, but instead to facilitate licensing solutions by resisting the smash-and-grab. Applying the existing legislation will incentivize the AI industry to strike deals with rightsholders. In other words, they will pay a negotiated amount for the products on display. That’s the best way to get more Canadian content into AI.
On the identity issue, the government’s summary document has this to say:
“Canadian AI must support, reflect, and project Canadian culture, which includes our customs, our history, and our heritage. Canadian voices, languages, communities, and knowledge must also be represented in how AI systems are designed, built, and used. Given our diverse and multicultural society, our approach to AI must acknowledge and support this rich diversity, including strengthening the French language by capturing and projecting its idioms, expressions, and cultural contexts.”
The best way to do this is to ensure that quality content in both official languages is made available to AI developers. As I have stated above, the fairest and most efficacious way to do this is through content licensing. Broad copyright exceptions will not facilitate licensing discussions. In fact, they do just the opposite by encouraging avoidance of dealing with rightsholders.
Regarding misinformation and deepfakes, this is a huge concern, and not just in Canada. Various legislative solutions have been proposed such as the bipartisan NO FAKES Act, currently working its way through the US Congress (opposed, as usual, by the internet libertarian organization, the Electronic Frontier Foundation). Other countries, such as Denmark, are addressing the issue through amendments to copyright law, giving individuals the reproduction rights to their image and voice. The UK has an anti-deepfake law on the books, introduced earlier this year, but Canada is still struggling to get its Online Harms legislation, after a couple of false starts, finalized and across the line. Re-introduction of that legislation is expected imminently, and will likely include social media restrictions on children, a highly controversial issue.
Privacy in relation to access to data, cultural identity, and misinformation including deepfakes are all content issues that Canada’s AI strategy will need to address. And so is copyright, although not mentioned in the strategy. Putting the best possible gloss on things, perhaps it is just as well there was not some throwaway line in the strategy pointing to the need to provide wider access to copyrighted content to ensure that Canada remains competitive on AI. That is the argument often employed by those who want freer access to “OPC” (Other Peoples’ Content). The argument is that “Everyone else is doing it (i.e. giving it away–which is factually untrue), so we have to as well in order to stay competitive”. Maybe silence was better than saying the wrong thing in this document.
In the absence of any reference to copyright issues, the last word must rest with Heritage and Identity Minister Marc Miller who spoke recently to the press after the National Summit on Artificial Intelligence and Culture in Banff, AB. The Minister is quoted as saying that Canadian copyright law is already clear that artists’ work needs to be respected, and that…”the current copyright law does and should protect those that have created material, and people need to be compensated properly.”
While that is encouraging, it would have been nice to have had this reaffirmed in the AI strategy document.
It was as predictable as wasps at a picnic. Within days of the Canadian Parliament’s Heritage Committee releasing its report on “The Impact of Artificial Intelligence on the Creative Industries”, with its lead recommendation being (my highlights)…
That the Government of Canada protect the property rights and interests of artists through the principles of the Copyright Act, in accordance with the ART principle—authorization, remuneration and transparency:
a) The Government of Canada must take the necessary steps and ensure that the scope of the Copyright Act applies to AI-generated content in order to guarantee copyright protection.
b) The Government of Canada must mandate greater transparency from AI developers regarding copyrighted works used to train their models, including disclosure of training data sources, to enable proper authorization and licensing.
c) The Government of Canada must establish a clear opt-in consent requirement for the use of copyrighted works in the training of artificial intelligence systems, ensuring that creators’ works may not be used for text and data mining or model development without their prior authorization.
…prolific tech and copyright commentator Michael Geist of the University of Ottawa was attacking its conclusions, issuing warnings that unless the tech industry is allowed (without authorization or compensation from rightsholders) to help itself to copyrighted content for the purpose of AI training, we will have “AI without Canada”. In other words, unless the tech industry is allowed to plunder Canadian content in the same way that it has been doing to date in the US (although this is meeting legal challenges and is quickly changing as licensing solutions take hold), there will be less Canadian content in the training data. This, apparently, will leave Canada as an “outlier” compared to peer jurisdictions. The AI developers will turn their back on Canada and rush off elsewhere. (This is a standard threat deployed by the AI industry to play off one country against another). He cites the EU, Japan, Singapore and Israel, as well as the US in support of this interpretation. Not mentioned as “peer jurisdictions” are the UK and Australia but then that would not have served the purpose of his narrative. Australia has recently declared it will not be legislating a Text and Data Mining (TDM) exception to its copyright laws to legalize unauthorized ingestion of copyrighted works for AI training, while the UK has just hit the pause button on a series of ill thought-out and badly received proposals to allow AI developers to freely use copyrighted content to train their AI algorithms unless rightsholders specifically opt out.
Singapore and Israel are among a small minority of countries that, under US pressure, have adopted US-style fair use laws that potentially allow for a weakening of copyright protection through a hodge-podge of court rulings. While many cite Japan as a jurisdiction that has given carte blanche to tech interests and AI developers, the facts are quite different as I pointed out in this blog post a couple of years ago. Japan has a strong cultural industry that it wants to nourish and protect and has defined its TDM exception very narrowly and carefully. The EU, has two provisions in its Copyright Directive related to AI training (Article 3 which permits TDM carried out only for non-commercial scientific research purposes, and Article 4, which permits TDM for any purpose, including commercial, as long as rightsholders have not opted-out, subject to strict transparency provisions by AI companies). Both impose constraints on AI developers, although there are differing views on opt-out.
Opting-out may sound like a compromise that both rightsholders and the AI industry could support but Britain’s example demonstrates otherwise. In its now aborted public consultation, the UK government put forward several options including its “preferred” option of opt-out. Fully 97 percent of respondents, from both the tech and creative communities, trashed this option. For creators, opting out not only stands copyright on its head (it is a property right, so why should holders of that right be required to notify someone who wants to infringe on that right that they may not do so, i.e. it’s like passing a law allowing anyone to picnic on my front lawn unless I post a “No Trespassing” sign), but it is technically difficult to do, especially for individuals and small-scale rightsholders. The robots.txt protocol is not binding and is in many cases not very effective. The tech industry doesn’t like opt-out because it imposes constraints on their untrammelled ability to access anyone’s copyright-protected content, anywhere, anytime. Instead the Committee recommends “a clear opt-in consent requirement” for the use of copyrighted works in the training of artificial intelligence systems.
Now it’s my turn to quibble. IMHO, there should be no explicit need for a rightsholder to “opt in”. I think that Canada’s copyright laws, properly interpreted, already provide sufficient protection to prevent unauthorized use. A rightsholder can “opt in” to AI training or any other unauthorized use not subject to fair dealing by granting a license to use their content. If that is an “opt-in” requirement then I am in favour. If yet another opt-in step is required, this would seem to be unnecessary. Licensing is a growing phenomenon. AI developers want reliable, curated content to develop their applications. As long as they are prevented from simply helping themselves, there is incentive for them to reach licensing deals with content owners. However, giving the tech industry a pass by allowing themselves to take for free whatever they want in the name of developing AI applications (for their commercial advantage) removes the needed incentive to negotiate with rightsholders. As to whether unauthorized use for AI training constitutes fair dealing, as Dr. Geist claims (“most TDM for AI training purposes would likely qualify as fair dealing under existing law”), this is doubtful to say the least. It is hard to imagine which fair dealing purpose currently applicable in Canadian law (research, private study, education, parody or satire, criticism or review, news reporting) would apply particularly when there are fair dealing limits to the amount of a work that can be used for such purposes, and specific factors that must be applied as to the effect of the dealing on the work.
The Committee’s lead recommendation is not the only complaint that Dr. Geist has about the Committee’s report. He feels it is unbalanced because the majority of its witnesses represented the cultural industries. It’s true that its lead recommendation is very much in line with the mainstream views of the Canadian cultural community. It was, after all, the Report of the Standing Committee on Canadian Heritage. This reminds me of the conflicting reports on copyright issued a few years ago by the Heritage Committee and its counterpart the INDU Committee. The 2019 Heritage Committee report, titled Shifting Paradigms, was attacked at the time by Dr. Geist as “the most one-sided Canadian copyright report issued in the past 15 years”. He claimed that there was “no attempt to engage with a broad range of stakeholders”, even though he himself appeared along with a number of others who shared his perspective on copyright. Shortly after issuing its own report, the INDU committee then issued a tone-deaf “We’re in charge” press release reminding the world that it had “sole responsibility” for administering the Copyright Act. (This is not strictly accurate). Dr. Geist’s main complaint, whether with “Shifting Paradigms” in 2019 or the current Heritage Committee report seems to be that the Committee members, in their wisdom, did not take his expert advice.
What is the function of Parliamentary Committees? It is to hear evidence, draw conclusions and make recommendations. He complains that while there were different points of view, including notably his, on how to tackle the issue under study, the Committee’s conclusions did not reflect these views. Was it because, numerically, there were more pro-copyright witnesses from the creative community that those from the Geist camp? That is theoretically possible if it were just a mathematical exercise of adding up comments in a pro and con column. But that is not the case. While the Report made a conscientious effort to capture the full range of comments, including those of Dr. Geist, in the end the members (from three political parties) made a judgement and reached consensus conclusions. (Although the Conservative Party members provided their own addendum that added to but did not refute the Committee’s conclusions). Presumably the members of the Committee were more convinced by the force of the arguments presented by some witnesses than others. Given the range and similarity of concerns presented by disparate members of the creative community it is not surprising where they came out in terms of conclusions.
Dr. Geist is entitled to disagree with these conclusions and recommendations. To be fair, his blog commentary echoes the position he presented to the Committee, except for his complaints about process. As I said at the outset, his attack on the Committee’s report is entirely predictable, like wasps at a picnic. And those wasps can be so annoying, distracting from the main event with the occasional bite and annoying buzzing, but as any determined picnic-goer knows, it’s important to not let them become the centre of attention. The Heritage Committee’s report was carefully considered and drafted by an all-party group after hearing from a wide range of experts. It provides important recommendations that the government would be well advised to take into account as it develops a legal framework in which both the AI and creative industries can co-exist and flourish.
I don’t know whether to feel offended or flattered, but I’ve been scraped–by AI. And I can prove it. The first intimation I had of this signal occurrence was a notice from WordPress asking me to approve a comment on my recent blog post, “Copyright, AI and the Legal Profession: Who Blinked?”. I logged in to find it wasn’t a comment but rather a link to this website.
A quick click took me to the article “CanLII settling with Caseway signals shift in legal-tech power dynamics”, dated April 20, the same day I had posted my blog. It was under a byline “London News”, which initially I naively assumed referred to London, Ont, (shows how parochial I can be) but quickly realized that this was some kind of online journal for commuters heading toward Picadilly Circus. London News appears to be written by a bot called Noah News Service, managed by the company HBM Advisory, based in London (England). There was no direct reference or link to my blog post in the article, but when I read it, it seemed eerily similar. The words were all different but the thread (with one exception that I will come to later) was the same. When I searched further, I found a footnote indicating the London News story was “inspired by” my blog post. What does this mean in reality?
My original post is protected by copyright, but anyone (even a bot I suppose) can take “inspiration” from a copyrighted work and produce something new. However, the “inspiration” I provided the bot is substantially different, in my view, from the sort of inspiration I would get from reading, say, an Agatha Christie mystery and then deciding to write my own mystery novel. In the case of my blog post, the bot did not really take “inspiration” from the content to create a new original work but rather engaged in rewriting the story using AI analysis of its key points to recreate what I had said using different words. That’s not true inspiration; it’s paraphrasing. Moreover, I’ll wager that an unauthorized copy of my work was made in order to feed the content to the bot to undertake its rewrite. While facts cannot be copyrighted (only someone’s expression of the facts), this rewrite was not based on the facts of the case. It was based on my blog post. Although the bot has not hijacked my precise words (i.e. my expression) it has nevertheless replicated the structure of my work, its flow and its arguments. It’s sailing very close to the wind, but probably still legal. This is not dissimilar to the challenge faced by news organizations who find their expensively created content being scraped and repackaged by online platforms such as Google, META, and others. According to the National Post, in a recent survey commissioned by News Media Canada, more than seven in 10 Canadians (of those surveyed) think the federal government should prevent artificial intelligence companies from taking and repackaging news content without permission or compensation.
But back to the London News article. Scrolling down to the end, I found an analysis of my blog post, produced by Noah. The post was rated according to various categories. It earned a “Freshness Check” score of 8/10 (i.e. the story was relevant), a “Quotes” check of 7/10; a “Source Reliability” score of just 6/10, a “Plausibility” rating of 8/10 but, sadly, an Overall Assessment for credibility of “Fail”, based on a “Medium” degree of confidence in this assessment. OMG, where did I fail to make the bot happy? How did I not meet its standards?
The Source reliability score would have been higher, according to the bot, if it had been published by an “established news organisation”, rather than on a personal blog;
“While the author, Hugh Stephens, has expertise in international copyright issues, (thanks, bot) the blog’s content is not subject to the same scrutiny as mainstream media.”
Well, I can live with that. The whole point of a personal blog is to offer a different perspective from Fox News, the BBC or the Globe and Mail.
The bot’s analysis continued:
“The article references reputable sources, but the lack of direct links to these sources raises concerns about transparency and verifiability.”
In other words, stuff your blog post with direct links to “mainstream media” and you might improve your report card. I could do that, but it might not be appreciated by my readers. The need for more direct links is repeated in the Quotes section (Score: 7/10) as well.
As for my failing grade, the bot’s summary says;
“The article provides a speculative analysis of the CanLII-Caseway AI settlement, referencing reputable sources but lacking direct links for independent verification. Its opinion-based nature and the author’s personal blog platform contribute to concerns about reliability and independence. Given these factors, the content does not meet the standards for publication under our editorial indemnity.”
But they published it anyway, as they do all kinds of content scraped from the web. I am not sure what the editorial indemnity policy is, but I suppose it is some sort of guaranteed reliability indicator, designed to separate the loony conspiracy theories (alternate facts?) from “real news”.
I wondered who would pass the bot’s scrutiny. Of the ten AI related stories posted on the front page of London News on the day I selected, 5 passed, 4 failed, and one was Conditional. The sources were all specialized but non-mainstream tech publications, or informed blogs, but certainly not conspiracy-theory outlets. Yet about half failed to gain Noah’s approval. I started to feel a bit better. Perhaps I’m not such an outlier.
I wonder if could write a blog post that would get an “A” from the bot. First, I would have to catch its attention, which I guess I could do by making sure there were lots of references to “AI” in the text, and then I would have to suppress my instinct to offer views on the topic. I would also have to stuff in lots of links to mainstream sources, like the Guardian and its ilk. But what is the fun in that? And what is the point? If people want to read “just the facts”, they can turn over the screening of content (and thinking) to their mainstream media subscriptions. However, I will say that the idea of assessing the reliability of a story on any topic, whether it’s on AI or the war in the Middle East, is not a bad thing. In the case of HBM, the assessment is used as a teaser to convince users (individuals, but more likely businesses) to sign up for more comprehensive, paid analysis. Part of the problem is that the assessment is done by an AI bot, and we know that AI is far from perfect.
HBM claims it uses AI and statistical modelling blended with human expertise and oversight to do its assessments. There is a thin but cursory layer of human involvement; fact-checking, source verification, style refinement etc. I think this is borne out by one missing key paragraph from HBM’s rewrite of my blog post. I had taken aim at Deloitte as an example of a large multinational company, that should know better, having been caught red-handed using unattributed AI that produced inaccurate, “hallucinated” results in a consulting report it prepared for the Newfoundland government. (“Deloitte’s AI Nightmare: Top Global Firm Caught Using AI-Fabricated Sources to Support its Policy Recommendations”). While HBM’s rewrite included almost all the key points in my post, there was zero reference to Deloitte. I am sure that “human expertise” decided that there was no point in gratuitously antagonizing an actual or potential client. Can I prove it? No, I guess its just another conspiracy theory.
I wonder if this blog post will be picked up and analyzed by Noah and if so, whether I would get a “Pass” this time. After all, it is “Fresh” and I have used lots of quotes from Noah. Having referred to the London News, I should get a 10/10 for Source Reliability (although I am not mainstream media, but neither is Noah). As for Plausibility what could be more plausible than an AI bot ripping off an author’s work through an unauthorized rewrite? Would all that land me a “Pass” from Noah? I will probably never know.
I wonder what really happened? Maybe we’ll never know. On March 23 it was announced that Caseway AI and CanLII (The Canadian Legal Information Institute) had reached a settlement in the copyright infringement case brought by CanLII against Caseway in 2024. As the saying goes, “Somebody knows something”, but they aren’t saying. The settlement is confidential and both sides are very tight-lipped, although Caseway is willing to riff a bit on social media. The CanLII announcement that each party will move forward independently, and that both consider the matter fully and finally resolved with no further comment, is particularly buttoned-up leading one (the “one” being me) to suspect it was maybe CanLII that blinked, not Caseway. But I could be wrong. There is no announcement that Caseway will be licensing CanLII content, or any hints that money has changed hands. Maybe Caseway agreed to stop what they were doing even though they denied doing it.
The facts of the case are as follows. According to its website, CanLII is “a non-profit organization founded in 2001 by the Federation of Law Societies of Canada on behalf of its 14-member law societies. Its mandate is to provide efficient and open online access to judicial decisions and legislative documents.”
Not only that but,
“CanLII supports members of the legal profession in the performance of their duties while providing the public with permanent open access to laws and legal decisions from all Canadian jurisdictions.”
Caseway AI says it is a company that is applying AI techniques to the legal profession “to make legal knowledge accessible, affordable, and usable for everyone.”
This being the case, you might think that CanLII would be delighted when an AI company like Caseway came along to use CanLII’s “free” resources to develop an AI-based legal platform, which would arguably improve access to legal information on the part of the public, plus simplify the research function for legal firms. You would be wrong. Part of the problem, no doubt, was that the AI company, Caseway, charges for its services while not being part of the profession.(i.e. take but not give).
Caseway’s sales pitch also might not endear it to the legal profession;
“We believe the justice system should not feel closed off to those without deep pockets or institutional power…By combining trusted legal sources with modern technology, Caseway levels the playing field—empowering solo lawyers, small firms, businesses, and individuals navigating legal challenges on their own…”
Oh oh. The self representation bogey. Maybe the real reason for CanLII’s suit was that Caseway AI and others like it were setting themselves up as a direct threat to the legal profession. Apart from the threat of more self representation, AI is a two-edged sword for many lawyers. Yes, it simplifies a number of routine duties and research functions, but at the end of the day it could also result in a lot fewer lawyers. The threat is no different than the threat posed to accountants, radiologists, stock market analysts and soothsayers, but needs to be taken seriously.
The nub of the CanLII case was that while it provides public, non-copyrightable judicial decisions, these public documents are compiled in a proprietary database. CanLII argued that it spends considerable time, effort and money to “review, analyze, curate, aggregate, catalogue, annotate, index and otherwise enhance the data” prior to publication and that this creative effort converts public information into copyright protected content. CanLII might be right, based on the US case of Thomson Reuters v Ross where a US court found that Ross Intelligence, an AI research firm, had infringed on the copyrighted legal materials, indexing system and case headnotes (summaries of judicial cases) of Westlaw, a legal research platform owned by Thomson Reuters. Notably Ross had tried to license the Westlaw content, but Thomson Reuters had refused, viewing Ross as a competitor to Westlaw. Ross then helped itself to the material. In both the CanLII and Reuters/Ross cases, the foundational content, (judicial decisions) were in the public domain, but the issue revolved around the secondary, interpretive materials and processes. In presenting its defence, Caseway did not argue that it was entitled to use CanLII’s content under fair dealing or because it was in the public domain. Instead, it argued that it didn’t access CanLII’s content at all. It got its content from other public sources. CanLII had to prove the contrary.
When I asked Google’s AI mode “How strong was the CanLII case against Caseway” I got a summary of various Canadian LawyerMagazine articles which discussed the pros and cons of the case, and an unsubstantiated assertion that “Caseway agreed to respect CanLII’s terms of service and cease any unauthorized automated data extraction.” Whether that is true or not I cannot say, but it is clear that both CanLII and Caseway will continue on their respective paths. Indeed, Caseway has just burnished its image a bit by cutting a deal with UBC (University of British Columbia, in Vancouver) to research ways to improve the accuracy of AI legal research tools. This is an ongoing problem for legal researchers and more than one lawyer has been sanctioned by the courts for presenting supposed legal precedents that were in fact non-existent, having been hallucinated by AI.
Apart from the AI hallucination problem, which does not limit itself to the legal profession (Deloitte Consulting being a prominent example of a major company being caught with its hand in the AI error-ridden cookie jar, without disclosure to the client), there is also the question of whether an AI platform should be allowed to provide legal advice. It is not licensed to do so and as a regulated profession, lawyers are jealous of their prerogatives. The profession is regulated for good reasons; to ensure competence and integrity to protect clients and the public. There are strict regulations against unlicensed practitioners providing legal advice, with severe penalties. In March of this year, ChatGPT’s parent company, OpenAI, was sued for engaging in the unauthorized practice of law, in this case by providing legal advice through a consumer‑facing chatbot. The seriousness of unlicensed persons or entities providing legal advice explains the many warnings posted on websites and blogs when discussing legal issues. “The foregoing does not constitute legal advice”. The case is pending.
Back to Caseway AI. Do I think that if you have a legal problem, you can solve it with a $49.99 a month subscription to Caseway instead of engaging a lawyer? Well, if you are determined to self-represent, it might be better to try it out rather than heading to the library to borrow a copy of the Highways Act, or Criminal Code, or searching for legal precedents that might be relevant to your case. On the other hand, remember the old saying, often attributed to Abraham Lincoln, that “A man who is his own lawyer has a fool for a client”.
In the scramble to jump on the global AI bandwagon, Korea has floated a proposal that would supposedly remove “legal uncertainty” for AI developers who use copyrighted content to train AI platforms. Unfortunately, the proposed “solution” threatens to throw Korea’s globally renowned creative sector under the bus. Nor does it remove the uncertainty.
As part of President Lee Jae-Myung’s National Artificial Intelligence Strategy, its Presidential Council has put forward a 98 point “Action Plan”, a blueprint for implementation. There are many aspects to an AI strategy, but a key element is to ensure legal clarity with respect to the use of content for AI training, especially copyrighted content. The Action Plan purports to do this. Its Point 32 proposes the introduction of “explicit exceptions under the Copyright Act to allow copyrighted works to be used without legal uncertainty (emphasis added) in the processes of collecting and analyzing data available on the web”. In other words, introduction of a copyright exception for text and data mining (TDM), subject to certain conditions such as some form of remuneration, transparency, and opt-out features for rightsholders.
If the goal is to remove “legal uncertainty” regarding the use of copyrighted works, this proposal falls short of the mark. No exception can provide 100% certainty given the Berne Convention requirement that any exception meet the so-called “three step test”, meaning that an exception is permitted only in certain special cases, provided that it does not conflict with normal exploitation of the work and does not unreasonably prejudice the author’s legitimate interests. While there will always be a degree of uncertainty regarding exceptions, the good news is there is a ready alternative. The surest way to ensure legal certainty is to encourage licensing of content from rightsholders. The problem with the introduction of a TDM exception—or even the discussion of a possible TDM loophole–is that it diminishes the likelihood of reaching licensing solutions by reducing the pressure on AI developers to open their wallets to reach licensing deals.
It is even more bizarre that the tech industry is pressing for a TDM exception given that Korea is one of the few countries, alongside the United States, that has adopted a fair use provision in its Copyright law. This was done in 2011 as part of the implementation of the Korea-US Free Trade Agreement after heavy lobbying by the tech sector. Fair use allows courts to make case-by-case judgements as to whether a given use meets fair use criteria, thus potentially allowing reproduction of copyrighted material without advance permission from the rightsholder. If free use of copyrighted material for AI training can be shown to be “fair”, why is a TDM exception needed? Even in countries where fair use does not apply (which is most of the world), there is no convincing case or consensus on the need for a TDM exception; there is even less reason for one in a state that has already adopted fair use.
Point 32 of the Action Plan takes note of the existence of fair use in Korea, commenting that the Ministry of Culture, Sports and Tourism is preparing fair use guidelines. These are to provide interpretive guidance on the exemption provisions of the Copyright Act to enable companies to utilize copyrighted data “with greater confidence”. Despite this, the Action Plan claims these guidelines alone are unlikely to fully eliminate uncertainty and judicial risks. Voluntary licensing, however, would eliminate both.
It is well established that AI developers need vast amounts of data to improve the performance of their AI platforms. To date they have largely employed a “take first, ask later” policy. This has led to numerous lawsuits pitting rightsholders against the tech industry, mostly but not exclusively in the US. AI developers in the US have argued that what they are doing amounts to “fair use” because the final AI product is used for a different purpose from the original and thus does not compete with it. That is highly debatable, especially with image and music-based AI works. To date, the results from the US courts have been mixed.
The legality of the tech industry’s unauthorized use of copyrighted content is an issue that a number of countries, in Asia and around the world, are looking at. Various solutions have been proposed to eliminate the uncertainties that arise from leaving the decision to the courts. Among these are TDM exceptions which have been introduced, albeit with strict limitations, in the UK, the EU and Japan. In the UK for example, use is limited to non-commercial purposes. In the EU, it must be accompanied by transparency requirements and opt-out provisions for authors. In Japan, if the unauthorized user derives commercial benefit from the content, the safe harbour does not apply. Australia has explicitly ruled out introducing a TDM exception in order to protect its creative sector, while many others (eg. India, Canada) have no TDM provision in their copyright law. As noted above, the clearest way to remove any uncertainty about the legality of using copyrighted works is to incentivize and recognize voluntary licensing as the solution. This ensures that rightsholders receive appropriate compensation for the work they have put into creating content, while guaranteeing legal certainty for licensees.
The Korean strategy paper argues that AI companies are required to obtain individual consent from each copyright holder “leading to significant costs and time burdens in securing high-quality training data”. But large amounts of high-quality content can be accessed through voluntary licensing agreements with major content creation companies such as studios, publishers, broadcasters, music labels and so on. As for individual authors and artists, one possibility is to look at the model currently used for licensing print and music content through Collective Rights Management Organizations as a supplement to voluntary licenses signed with major rightsholders.
In addition to being instructed to prepare the necessary amendment to the Copyright Law for presentation to the National Assembly by Q2 of this year, the Culture, Sports and Tourism Ministry, in cooperation with the Ministry of Science and Technology, is to “promulgate standard contract templates for the licensing and transfer of copyrights for AI training”. This is the kind of heavy-handed market intervention that is guaranteed to stifle voluntary licensing. Not only that, it amounts to expropriating the rights of Korean creators to manage their works.
The Action Plan gives a nod to the importance of compensating rightsholders and claims it wants to establish a system that respects the rights of creators. However, given the size and importance of Korea’s cultural industries, from film to K-Pop to literature, it is surprising there isn’t greater recognition of what an important strategic and economic asset this sector represents for Korea. Although the Strategy acknowledges that content industries should be able to share the benefits of growth in the AI industry, the proposed solution is unbalanced and biased toward clearing any so-called “obstacles” to unimpeded use of content. As a result, just days after the extremely brief (20 day) consultation period on the Strategy had closed, in mid-January sixteen creator and rightsholder groups issued a strong statement condemning the Action Plan, labelling it “an attempt to fundamentally undermine copyright as a private property right”.
While paying lip service to creator’s rights, the Plan does not address how creators can enforce these rights (other than through the creation of opt-out protocols, which stands the normal copyright procedure of seeking permission prior to usage on its head). The Strategy seems to lead to what has been described by many as a “use now, pay later” system, with little information on how payment would be calculated or implemented. On the other hand, prior, voluntary licensing of content for AI training is a solution that would respect the rights of Korea’s creators while providing the welcome revenue sharing and income stream for which the Strategy advocates. Strong content industries benefit AI development in Korea by encouraging continued creation of the valuable Korean language content so necessary to refine and improve AI models. Conversely, providing the tech industry with an escape hatch to avoid licensing by instituting a TDM exception is the surest way to kill a licensing market for AI content. It will only continue the legal uncertainty that the Presidential Council seems to feel is hindering AI development in Korea.
The one-sided formulation of the Strategy to date has provoked an inevitable negative reaction from Korea’s cultural industries. This is not surprising since the strategy of the tech industry, in Korea and elsewhere, is to avoid dealing with ministries directly responsible for culture and copyright and instead lobby industry, technology and science ministries to bring pressure for changes to copyright law. This adversarial stance is unfortunate as the content and tech industries need and can help each other. The Strategy needs to be amended so rather than throwing the cultural and copyright industries under the bus in the name of facilitating AI development, Korea provides the framework for a mutually beneficial and legally certain relationship. This is best done by upholding longstanding copyright principles and encouraging the growth of a voluntary licensing market for content used in AI training.
It seems everyday new applications and new threats emerge from the AI world. This applies in particular to creators who see growing AI challenges to their livelihoods; graphic art and album covers spat out by AI generators; voice actors replaced by AI clones; authors struggling to make their works known in a sea of AI-generated slop; now AI artists are even making the Billboard charts. At the same time, AI has many other functions and produces a host of products that have little to do with artistic creation. In particular, it can be used as a crutch to assist and enable research in a wide range of fields. Today it is routinely used by everyone from school kids to law firms to health care researchers. And that is where the risks of mainlining AI are the most evident, because of the propensity of AI platforms to fabricate plausible sounding misinformation.
In a blog post earlier this year (AI’s Habit of Information Fabrication (“Hallucination”): Where’s the Human Factor?) I discussed some examples of law firms caught submitting non-existent case precedents in court as a result of sloppy legal research using AI. Judges have very limited tolerance for this practice, which wastes valuable court time, and they are increasingly imposing significant penalties—that is, if the fabricated information is actually spotted. The problem is not going away. This website maintained by Paris-based legal scholar Damien Charlotin has compiled a database of more than 550 legal cases in 25 countries where generative AI has produced hallucinated content. These are typically fake citations, but also include other types of AI-generated arguments. The US wins the lottery at 373 cases, but Canada is second with 39. Even Papua-New Guinea has one case.
As you can well imagine, AI hallucinated results in health care could be fatal. As Dr. Peter Bonis, Chief Medical Officer at Wolters Kluwer Health points out, hallucination in the health care field has led to various consequences such as recommending surgery when it was not needed, advising that a specific drug could be safely stopped abruptly when this was known to be dangerous, avoiding recommending vaccinations based on known allergies even though it was safe to do so, proposing wrong starting treatments for patients with rheumatoid arthritis and so on. You get the picture. You don’t want your family doc using AI search for the remedy for whatever ails you. The fact that the models present incorrect information with such confidence, and that potentially dangerous incorrect information is embedded with a lot of correct information, makes proper use of AI outputs particularly challenging.
How is it that AI platforms consistently produce unreliable results? This MIT Sloan article identifies three elements;
Training data sources (the uneven quality of inputs, including pirated, biased and otherwise unreliable content)
Limitations of generative models (generative AI models are designed to predict the next word or sequence based on observed patterns and to generateplausible content, not to verify its accuracy)
Inherent Challenges in AI Design (The technology isn’t designed to differentiate between what’s true and what’s not true)
This is all pretty concerning if people are going to surrender personal judgement to AI and use it to cut corners without verification. One way to address part of the problem is to ensure the training data used is reliable and of high quality. That is where licensing of accurate, curated data and content as training inputs is important and that is why a licensing market is developing as AI companies seek out better quality data to distinguish their product from that of their competitors. This can be very helpful where the AI platform is limited to discrete areas of knowledge, such as in the medical field for example, where usage can be limited to professionals who are prepared to pay for a bespoke AI product and who are qualified to interpret the results properly. AI for the general public is another matter, and this is where most of the problems arise. Unfortunately, while improving the quality of training data helps reduce hallucinations, it does not completely eliminate them. As the New York Times has reported,
“Because the internet is filled with untruthful information, the technology learns to repeat the same untruths. And sometimes the chatbots make things up. They produce new text, combining billions of patterns in unexpected ways. This means even if they learned solely from text that is accurate, they may still generate something that is not.”
User beware. Nonetheless, better inputs lead to better outputs. As AI developers work to take their products to the next level by refining their training processes and making outputs more predictable and trustworthy, they will need access to curated, proprietorial content and closer collaboration with content owners. Dr. Bonis noted that for specialized areas like health care, AI companies will get better quality feedstock while creators of the content will receive funding allowing them to continue research. A virtuous circle.
Users bear a big responsibility to ensure AI is employed effectively. The mindless, unjudgemental use of AI to reach conclusions in areas where the user has little knowledge can be dangerous. By all means use AI as a tool to sort and categorize, but don’t rely on it to produce the answers on which substantive decisions will be based. Any sensible user of AI has a pretty good idea of the answer to the question before it is even asked. It is also a good idea to refine the question, so you narrow the range of possibilities.
Some proprietary AI models offer RAG (Retrieval Augmented Generation) where the AI will retrieve relevant information from trusted sources to supplement its preliminary analysis. This can increase reliability. However RAG, where the AI goes after specific inputs to bolster its results, can also expose AI developers to charges of copyright infringement, as is currently the case with Canadian AI company, Cohere, which is being sued by a number of newspaper publishers, including the Toronto Star, for copyright infringement. As Canadian lawyer Barry Sookman has pointed out in a recent blog, use of RAG can create risk for the AI platform. In the case of Cohere, when its RAG feature was switched on, it reproduced large amounts of almost verbatim text pulled directly from the litigating news sources. But if the RAG function was switched off, it produced fabricated information (hallucinations) yet still identified this false information as coming from an identified reliable news source, leading to charges of trademark dilution. The value of the brand was diminished by the attribution of false information to it. This trademark dilution issue is also part of the New York Times case against OpenAI.
At the end of the day, it is a case of user beware as a recent case in Newfoundland demonstrates well. The Government of Newfoundland commissioned an in-depth study on the future of education in the province. The 410 page report containing over 110 recommendations, authored by two university professors, was released with great fanfare at the end of August. No doubt a great deal of careful research had gone into producing the study over the 18-month production period. But then cracks started appearing in the edifice. It was chock full of made-up citations. The more people started checking, the more they found. The Department of Education and Early Childhood Development tried to whitewash the issue by saying it was aware of a “small number of potential errors in citations” in the report. But even one fabricated citation is one too many! If you search for the report online now you get the classic “404 Not Found” message. A lot of work has potentially gone down the drain, and possibly the credibility of two academics has been destroyed by careless use of AI. This is a cautionary tale that I have no doubt will be repeated.
In fact, it was repeated just a few days later. It seems Newfoundland is particularly prone to victimization by hallucinating AI platforms. After the education report debacle, new reports have surfaced that a $1.5 million study on the health care system conducted by none other than Deloitte also contains fabricated information included made up references. The opposition party is demanding the government insist on a refund.
In our rush to embrace AI, many seem to have forgotten the value of human creativity and judgement. Coming back to the creative industries and AI, some of those whose livelihoods may be threatened by this new phenomenon are bravely trying to find a silver lining. Some voice actors are generating an additional revenue stream by licensing their voice clips for AI training, and many graphic artists use AI as an assist. Are they putting themselves out of work in the long run or are they simply adapting? The jury is still out, but the generally low quality of AI produced art, music and literature, as well as the ongoing problem of hallucination, suggests that there will always be a need for real human input. Anyone planning on substituting AI for “real work” had better think again.
Last month I highlighted the first AI/Copyright case in Canada to reach the courts, CanLII v CasewayAI. CanLII, (the Canadian Legal Information Institute), a non-profit established in 2001 by the Federation of Law Societies of Canada, sued Caseway AI, a self-described AI-driven legal research service, for copyright infringement and for violating CanLII’s Terms of Use through a massive downloading of 3.5 million files which Caseway allegedly used to populate its AI based services. Now the principal of CasewayAI, Alistair Vigier, through an article (Don’t Scare AI Companies Away, Canada – They’re Building the Future) published in Techcouver, has responded publicly by trotting out many of the tired and specious arguments put forward by the AI industry to justify the unauthorized “taking” of copyrighted content to use in or to train generative AI models. Let’s have a closer look at these arguments.
Vigier opens by referencing another AI/Copyright case in Canada where a consortium of Canadian media companies is suing OpenAI for copyright infringement. He claims this is all based on a misunderstanding of how AI training works, stating that “AI systems like OpenAI rely on publicly available data to learn and improve. This does not equate to stealing content.” Whether data is “publicly available” or not is irrelevant when it comes to determining whether copyright infringement (aka stealing content) is concerned. Books in libraries are publicly available, or so is a book that you purchase in a bookstore, or content on the internet that is not behind a paywall. (It is worth noting that the Canadian media companies also claim that OpenAI circumvented their paywalls to access their content when copying it). But in none of these cases is copying permitted unless the copying falls within a fair dealing exception, which is very precise in its definition. Labelling copied material as “publicly available” is a red herring.
Vigier’s next argument is to equate the ingestion of content by various AI development models with a human being reading a book. We know that humans enhance their knowledge through reading and are thus able, presumably, to better reason based on the content they have absorbed. Vigier says, “This is how AI works. The AI “reads” as much as it can, gets really “smart,” and then explains what it knows when you ask it a question. Like a human learns from reading the news, so does an AI.”
Really? A human does not make a copy, not even a temporary copy, of the content although some elements of the content are no doubt retained in the human brain. But AI operates differently. It makes a copy of the content. This should be beyond dispute although the AI industry continues to muddy the waters by claiming that when content is “ingested” it is converted to numeric data and is thus not actually copied. This is a fallacious argument. Just because the form changes, this does not mean there is no reproduction. When you make a digital copy of a book, there is still reproduction even though the digital form is different from the original hard copy version. When a work is converted to data, the content is still represented in the dataset.
Vigier dubiously states, with regard to OpenAI, “OpenAI’s models do not reproduce articles verbatim; they process vast datasets to identify patterns, enabling insights and efficiency.” Apart from the fact that the New York Times in its separate lawsuit in the US has been able to demonstrate that by typing in leads of articles, it can prompt OpenAI to reproduce verbatim the rest of the article (OpenAI claimed that the Times “tricked” the algorithm), copying is copying even if the result of the copying is somewhat different from the original. The Copyright Act is crystal clear on this point. Section 3 (1) of the Act states that, “For the purposes of this Act, copyright, in relation to a work, means the sole right to produce or reproduce the work or any substantial part thereof in any material form whatever…“. If copyright protected content is reproduced in its entirety without permission for a commercial purpose (eg for AI training), that is infringement, unless the use qualifies as a fair dealing under Canadian law or fair use in the US.
The issue of whether ingestion of content to train an AI application results in copying (reproduction) has been carefully studied and documented. One of the most thorough examples is a recent SSRN (Social Science Research Network) paper, entitled, “The Heart of the Matter: Copyright, AI Training, and LLMs” with noted scholar Daniel Gervais (a Canadian by the way) of Vanderbilt University as lead author. The article goes into a detailed discussion on how copying of content occurs during AI scraping to build a Large Language Model (LLM), including the stages of tokenization, embedding, leading to reward modelling and reinforcement learning. The section of the article explaining how copying occurs (pp. 1-6) is dense, technical text but the conclusion is clear, “LLMs make copies of the documents on which they are trained, and this copying takes various forms, and as a result, with appropriate prompting, applications that use the LLMs are able to reproduce original works.” A shorter (and earlier) version explaining how the LLM copyright process works can be found in this article (“Heart of the Matter: Demystifying Copying in the Training of LLMs“), produced by the Copyright Clearance Center in the US. It is also worth noting that these explanations refer only to ingestion of text. AI models that train on images and music are even more likely to produce exact or close-to-exact reproductions of some of the works they have been built and trained on.
So much for the misinformation in Vigier’s article. Now to the scare tactics. He says that the recent Canadian media lawsuit against OpenAI sends a negative message to innovators that Canada may not be open to AI development.
“If Canada wishes to remain relevant in this (AI) sector, it must balance protecting intellectual property and promoting technological progress.”
The fact that there are currently more than 30 lawsuits in the US, including the seminal New York Times v OpenAI case, does not seem to have slowed down the AI companies in the US. In the UK, legislation has been introduced that would, according to British media reports, “ensure that operators of web crawlers (internet bots that copy content to train GAI, generative AI) and GAI firms themselves comply with existing UK copyright law. These amendments would provide creators with crucial transparency regarding how their content is copied and used, ensuring tech firms are held to account in cases of copyright infringement.” There is lots of AI innovation ongoing in Britain.
The Australian Senate Select Committee Report on Adopting AI has recommended, among other findings, that there be mandatory transparency requirements and compensation mechanisms for rightsholders. The EU is already way out in front on this issue. Its new AI Act stipulates that providers of AI generative models will be required to provide a detailed summary of content used for training in a way that allows rightsholders to exercise and enforce their rights under EU law. Even India now has its own version of the US and Canadian media cases against OpenAI. (OpenAI’s defence in part is based on the argument that no copying took place in India because no OpenAI servers are located there!)
If that is what the “competition” is doing, who does Vigier cite as being the jurisdictions most likely to attract innovators away from Canada? Why, it is those AI powerhouses of Switzerland, Dubai—and the Bahamas!
The argument that if legislators and the courts don’t give AI innovators a free pass on helping themselves to copyrighted content for AI training purposes, this will either slow down innovation or chase it elsewhere is a common fearmongering strategy of the AI industry. This is a race-to-the-bottom mentality whereby content industries are thrown under the AI bus. Vigier, having been the subject of his own lawsuit, argues that instead of resorting to litigation, the Canadian media companies should have sought a licensing solution. But the fact that no licensing agreement was reached with OpenAI is undoubtedly the reason for the lawsuit in the first place. That is certainly the reason behind the NYT v OpenAI lawsuit in the US; licensing negotiations broke down. If someone has taken your content without authorization, and then offers you pennies on the dollar in comparison to what that content is actually worth, then the stage for a lawsuit is set.
In explaining CasewayAI’s position in the litigation brought by CanLII, Vigier says that Caseway approached CanLII with an offer to collaborate but was rebuffed. As a result they developed other extensive web crawling technology that pulled the needed material from elsewhere. (Where exactly the material was downloaded from is the crux of the matter). Regardless, this makes it sound as if it was CanLII’s fault for refusing to share their content. Surely a rightsholder has the right to determine the terms on which their content is to be shared with others, if at all.
The fact that Caseway went to CanLII in the first place suggests that CanLII had developed the content that Caseway wanted. Caseway claims the material it accessed was on the public record, such as court documents and decisions. CanLII, on the other hand, claims that it had reviewed, indexed, analyzed, curated and otherwise enhanced the content in question, thus adding a wrapping of copyright protection to what otherwise would be public documents. Who is right, and whether the material was scraped from CanLII’s website without authorization, will be determined by the BC Supreme Court.
If the material taken by CasewayAI was not copyright protected, they are in the clear, at least with respect to copyright infringement. That is quite different, however, from arguing that no copying takes place during AI training or that if rightsholders use the courts to protect their rights, Canada will be a laggard when it comes to AI development. Robust AI development needs to go hand in hand with robust copyright protection for creators, with an appropriate sharing of the spoils of the new wealth generated from the creative work of authors, artists, musicians and other rightsholders. To say, as Vigier does in his concluding paragraph that;
“Canada has a choice to make. Will we embrace AI as the transformative force it is, or will we let fear and litigation stifle innovation? The lawsuits against Caseway and OpenAI message tech companies: you’re not welcome here. If this continues, Canada won’t just lose its AI startups; it will lose the future of job creation.”
What sheer self-interested nonsense!. This is fearmongering of the worst kind, based on an inaccurate and misinformed knowledge of how AI is developed and trained, that moreover impugns the legitimate right of a rightsholder to seek the protection of the law to protect their creativity and investment in content. Vigier might be correct when he says that licensing of content is a win/win for both parties. I agree with that. But licensing negotiations are about money and conditions of use and require willing parties on both sides. When licensing discussions break down, or when one party decides to do an end run on licensing because they have been rebuffed, then the way to gain clarity is through the courts whose job it is to interpret what the legislation means.
Canada still needs to come to grips with the question of how copyrighted content will interface with AI development. As I noted earlier, both sides in the debate made their cases in the public consultation launched a year ago, but since then there has been no movement in Ottawa. The law could be strengthened to ensure adequate protection of rightsholder interests in an age of AI, resulting in facilitating licensing solutions. In the meantime, misinformation and scare tactics need to be called out for what they are.
Adequate protection for rightsholders does not mean the end of AI innovation or investment in Canada. There is no need for panic. We can walk and chew gum at the same time.
In July, the Canadian Internet Policy and Public Interest Clinic (CIPPIC) at the University of Ottawa filed an application in the Federal Court to expunge or amend a Canadian copyright registration that claimed an AI program, the RAGHAV AI Painting App, as co-author of a registered work. While the other co-author, an Indian IP lawyer by the name of Ankit Sahni is named as the respondent, the real defendant ought to be the Canadian Intellectual Property Office (CIPO), the organ within the Department of Industry (ISED) responsible for managing copyright registration. It is CIPO’s “rubber stamp, content-blind, absence of judgement” automated system of registration that has led to this situation, putting Canada in a significantly different place from that of the United States or many other countries when it comes to granting copyright protection to works produced by AI algorithms with no or little human intervention.
Last month I wrote a couple of blog posts on the issue of whether content produced with or by generative AI could or should qualify for copyright protection. I looked at the ongoing uphill struggle that two “creators”, Stephen Thaler and Jason Allen, have experienced with the US Copyright Office (USCO) in their attempts to get the USCO to register their works. Thaler claims his submitted work (“A Recent Entrance to Paradise”) was created exclusively by his AI algorithm (the “Creativity Machine”) and, accordingly, it should be recognized as the “author”. However, as the human behind the machine, having invested in creating it, the benefits of the registration should fall to him. He argues that the algorithm carried out the work at his behest, much like a work for hire. Allen, by contrast, claims that although his award-winning work (“Théâtre D’Opéra Spatial) was produced with AI assists, he was the creator through control and manipulation of the prompt process. In neither case has the USCO budged from its position that the works do not qualify for copyright protection on the basis they were not human-created. The same goes for the courts to which the USCO’s rejection has been appealed.
That is the current situation in the US; in Canada it is quite different. Works produced exclusively with AI have been accorded copyright registration, more than once. I have even done it myself! (See “Canadian Copyright Registration for my 100 Percent AI-Generated Work”).
Because Canadian copyright registration is automated and done through a website, an applicant must provide the author’s address, contact details and date of death, if deceased. To work around that, I clearly specified that the work was created entirely by two AI programs (DALL E-2 and CHAT-GPT) with virtually no exercise of “skill and judgement” on my part, this supposedly being the threshold in Canada for creative content that can be afforded copyright protection.
This is how my Canadian copyright certificate No. 1201819, issued April 11, 2023 (I should have tried to register it on April Fools Day), reads in terms of describing the registered work;
“SUNSET SERENITY, BEING AN IMAGE AND POEM ABOUT SUNSET AT AN ONTARIO LAKE CREATED ENTIRELY BY AI PROGRAMS DALL-E2 AND CHATGPT (POEM) ON THE BASIS OF PROMPTS DEMONSTRATING MINIMAL SKILL AND JUDGEMENT ON THE PART OF THE HUMAN AUTHOR CLAIMING COPYRIGHT”.
While this little exercise in inanity was fun, (and was done to expose the failings of the current system), I was not the first to register an AI created work in Canada. That honour, as far as I can tell, belongs to Sahni, the named respondent in the CIPPIC case who, in December 2021, managed to register the artistic work Suryast, listing the AI-powered RAGHAV painting app as co-author. That was a neat way of getting around the requirement to provide an address, contact details etc. Sahni could provide his contact details yet still claim the AI algorithm was an author, even if a co-author. Clever. What Sahni’s motivation was I cannot say, but apparently he has been active in registering the work in as many jurisdictions as he can, maybe to boost the marketability of RAHGHAV. CIPPIC claims he is seeking registration to force various countries to address the AI authorship issue. Canada must have been one of the easiest registrations he received. Now he is being called to account. The application brought by CIPPIC seeks a declaration either that there is no copyright in Sahni’s image, Suryast, or, alternatively, if there is copyright in Suryast, that the Respondent (Sahni) is it sole author. It also seeks an order to expunge the copyright certificate in question or to rectify it by deleting the painting app as a co-author.
The fundamental problem of course is not Sahni or his AI app, (although like me, he may have been mischievous) but rather the way in which copyright registration is offered and maintained in Canada. It was not always this way. Once upon a time, to register a work in Canada you were required to not only pay a registration fee, (which is still the case today) but submit three copies of the work, one for the Copyright Branch (which was part of the Department of Agriculture), one for the Canadian Parliamentary Library and one for the British Museum. Because of these depository requirements, today we have a record of many early copyrighted works in Canada, such as the famous early 20th Century Inuit photographs of Canada’s first professional female photographer, Geraldine Moodie, about whom I wrote a few years ago (“Geraldine Moodie and her Pioneering Photographs: A Piece of Canada’s Copyright History”).
When the first international copyright convention, the Berne Convention of 1886, was established among a limited number of countries, there was a push by authors to abolish the registration requirement because it was burdensome to have to register in all Berne countries. Initially, registration in the home country was supposed to provide protection in all member states of the Convention, but this proved difficult to put into practice. Consequently, in 1908 at the Berlin revision of the Convention, the following provision (which is today part of Article 5(2) was adopted, “The enjoyment and the exercise of these rights shall not be subject to any formality”. Canada was a member of Berne because Britain had acceded, but was nonetheless a reluctant conscript (even though then PM Sir John A. Macdonald had acquiesced to Canada’s inclusion). In 1921 Canada finally passed its own Copyright Act (coming into force in 1924, a century ago this year), and subsequently joined Berne in its own right in 1928. I suspect that registration as a requirement, along with depository and examination conditions, was dropped at that time. That is probably when the current (but non-automated) voluntary registration process was established.
Despite the abolition of a registration requirement by Berne Convention countries, Canada is not the only country that maintains one. In the US, which only joined Berne in 1989, both registration and renewal were required for a work to enjoy copyright protection. When the US joined Berne, it maintained the registration requirement for US citizens who wished to take legal action to enforce their copyright. This is allowed under Berne. As such, the US has maintained a robust registration system where a legal deposit of the work is required, registrations are examined and can be challenged or refused.
We know that is not the case in Canada, but Canada is not the only country to have a voluntary registration system. In a recent study by WIPO (World Intellectual Property Organization), some 95 countries were identified as having either a voluntary registration system, a recordation system (for transfer of copyrights) or a legal deposit requirement. What is notable, however, is that of all these countries, only three (Canada, Japan and Madagascar) do not require a deposit of the work seeking registration. Canada does review applications but only to ensure they meet all the formality requirements (name and address of the owner of the copyright; a declaration that the applicant is the author, owner of the copyright or an assignee; the category of the work; its title; name of the author and, if dead, the date of the author’s death, if known. For a published work, the date and place of first publication must be provided and, perhaps most important, payment of the prescribed fee). Nothing else. In fact, if Mr. Mickey Mouse, address Disneyland Way, filed a copyright application for a work and paid the required fee of $63, I am sure a Canadian copyright certificate would be issued. It used to come in the mail, printed on nice quality paper but, alas, in the interests of efficiency, it is now only available in PDF format on CIPO’s website. Print it yourself.
That is the current situation, but why has CIPPIC gone to the Federal Court to dispute the wording of Sahni’s copyright certificate, No. 1188619? While the Registrar of Copyrights can accept requests for correction of a copyright certificate (either because of an error in filing or because the Office itself made a mistake), it cannot by itself amend or remove a registered work from the Register. Instead, the Registrar needs the Federal Court to effect such action. Section 57 of the Copyright Act states, with respect to Rectification of Register by the Court;
(4) The Federal Court may, on application of the Registrar of Copyrights or of any interested person, order the rectification of the Register of Copyrights by (a) the making of any entry wrongly omitted to be made in the Register, (b) the expunging of any entry wrongly made in or remaining on the Register, or (c) the correction of any error or defect in the Register
However , while CIPPIC is seeking expungement of this particular copyright registration, it is the system it is really going after. This is clear from its memorial to the Court;
(23) “In automating its copyright registration process, CIPO is derogating from its obligations to administer copyright in a fair and balanced manner under the Copyright Act.” (24) “The consequence of this system is that content that does not merit copyright can…easily obtain the benefits of registration.” (25) “Copyright registrants obtain certain benefits under the Act – such as litigation presumptions – and users and defendants are correspondingly burdened. Once a “work” is registered, the Copyright Act…shifts certain presumptions such as subsistence and ownership….In this very case, as a result of CIPO’s oversight failures, the burden rests on CIPPIC to prove the image Suryast lacks originality and that an AI program cannot be an author.”
Moreover, CIPPIC notes that it brought this case to the attention of CIPO but it refused to correct the Copyright Register, instead encouraging CIPPIC to seek resolution in court. Assuming it is granted standing, CIPPIC may well prevail and have the Suryast registration amended or expunged. But will that really achieve its goals? If its goals are to get CIPO to stop “derogating from its obligations”, then simply cancelling or amending this one registration won’t do it. What is the solution?
One option would be to eliminate the voluntary registration requirement altogether, but is this the right course of action? The WIPO document referenced earlier points out some of the advantages of a voluntary registration system. It can ensure that information about authorship and copyright, including date of registration, become publicly available. This benefits not only authors and rightsholders, who can use the registration as a rebuttable presumption of copyright in court, as in Canada, but also provides information to the public to verify ownership claims and trace title. A voluntary system does not, however, provide a definitive list of what works are under copyright and which are not. Another factor is that an automated voluntary system, such as the one operated by CIPO, is not burdensome for registrants and presents no meaningful obstacle. The problem is that its barriers to registration are so low that it is easy to trick the system. Is a Canadian copyright certificate worth the paper it is printed on if there is no verification?
A second option is to improve the registration process to make it meaningful, but this will require resources. Current fees are low (but the US system which is much more robust has a similar fee structure). Nonetheless, to institute a USCO type system would require substantial additional resources that are unlikely to be forthcoming in the present fiscal environment. One would have to ask whether the extra cost could be justified. It’s a conundrum. Meanwhile, the government has circulated a paper on the issue of Copyright and AI and the Canadian cultural community has weighed in with its views. Prominent among these is the position that copyright protection should be accorded only to human-created works. (This is not currently specified in the Copyright Act).
CIPPIC’s court action puts the spotlight on the current copyright dilemma. The current system seems to be not fit-for-purpose, but an economically viable alternative is not immediately apparent. At the very least, Canada should amend the Copyright Act to prevent AI-created works from obtaining copyright registration.