AI Training and Copyright Licensing: The Changing Tide

Sunset over the ocean with golden reflections on the wet sand and gentle waves lapping at the shore.

Image: Author (Tofino, BC)

Australian Prime Minister Anthony Albanese caught global attention in mid-July with his speech at the University of Sydney, “AI in Australia’s Interests”. That address, which outlined the Australian government’s approach to AI, included important statements regarding the work of Australia’s creative community, which Albanese declared is “not up for grabs”. Having previously ruled out the creation of a copyright loophole for AI training, known in the trade as a “Text and Data Mining” (TDM) exception, Albanese went on to say that;

Australian writers, musicians, artists and journalists must retain ownership and control of their work…No company should use Australian books, music, art or news to build or train AI without the artist’s control…of the price and value of their work. Anything less is theft.”

Bravo! Hopefully this means what it says, that rightsholders will retain control and receive compensation on their terms if their works are used for AI training, even though there is still strong pressure from segments of the AI world for Australia to loosen its terms of copyright protection. Anthropic is dangling a $20 billion carrot in the form of potential investment in AI data centres in Australia but only if there is a copyright carve-out for AI training. One Anthropic proposal was for the creation of a $350 million fund to compensate rightsholders, paid for by the AI industry, but it was unclear how this would operate or if rightsholders would be required to opt-out if they didn’t want their content used. For creators to be able to control their work within the accepted framework of copyright law, they must have the ability to accept or reject voluntary licensing and not be required to opt-out of a compulsory scheme where what they receive is decided by bureaucrats on the basis of limited contributions to a common fund. The commitments that Albanese made in his speech indicate that voluntary licensing is the preferred solution, but the issue remains under study (hopefully this will be done with full transparency, following suggestions earlier this year of secret dealings on this critical matter of public interest). Nonetheless, the bright line that Albanese has drawn with regard to creatives retaining control of their work is welcome and will hopefully encourage other governments who are developing AI policy to do the same.

As noted, the AI industry never gives up. Having been clearly told that creators should retain control over how their work is used, and be paid for that use, some companies are floating new objections based on the supposed “long tail” argument. “Long tail” is the term being used to describe the many small rightsholders who create and therefore (in theory but maybe not in practice) control use of their content, as opposed to larger rightsholders such as major publishers of books, journals and newspapers, music labels, film studios and so on with whom it is easier to engage in licensing discussions. The argument is that, even if the AI industry wanted to license the content that it uses in training (by no means a given with every company), it can’t possibly deal with the myriad of small rightsholders. Elsewhere, some AI developers have dealt with this problem by ignoring it, simply helping themselves to whatever content they wanted, without licence. In some cases, they even used pirate libraries as sources, almost daring rightsholders to bring legal challenges. Anthropic is a high-profile culprit that got caught doing this.

One idea currently being promoted by Anthropic is to create a statutory licence that would be restricted just to Australian rightsholders, with funds directed only to them. It is hard to see how this is a viable solution. Not only would it violate international norms that Australia has committed to by discriminating against foreign rightsholders in the Australian market, if it is applied universally, it would lead to a stream of royalties flowing out of Australia while doing little to help local creators. Elsewhere, various non-statutory collective licensing schemes have been proposed to deal with the issue of small rightsholders, although to date none have emerged that provide true one-stop shopping for AI developers. But a compulsory licence cannot be the solution. These large companies with stratospheric valuations (especially those getting ready to launch an IPO!) surely have the wherewithal to find ways to license the content they use and will continue to use as they develop and refine their AI models.

For the past several years the tech community has followed the “better to ask for forgiveness after rather than permission before” approach when appropriating copyright protected content to develop AI platforms. The inevitable result has been a spate of lawsuits, most of which are still ongoing, with results varying from case to case and jurisdiction to jurisdiction. A second front opened by some AI platforms has been to attempt to weaken copyright protection through legislation, focussing particularly on the introduction of a broad TDM exception. As a result of this push, many countries began studying whether to bring in a TDM carve-out, or if they had a narrow TDM exception (e.g. for non-commercial research purposes) whether to widen it. From the perspective of rightsholders, it began to look like a race to the bottom. Now, the tide seems to be changing. Albanese’s speech is a good example of the shift, but Australia is not alone in taking a more considered and balanced approach.

Hong Kong is a case in point. Hong Kong has always had one of the stronger intellectual property regimes in the region, a source of competitive advantage. But under the pressure of blandishments from the tech community, it issued a consultation paper in 2024 that, among other things, proposed a TDM exception–including for commercial purposes. When the paper went out for public comment, the predictable suspects urged adoption of a TDM provision while rightsholder groups argued that existing law and voluntary licensing was suitably flexible to handle issues arising from new technologies. Now the Hong Kong Government has concluded there is no pressing need to pursue a TDM exception through legislation. The priority will be to issue best-practice guidelines grounded in the existing legal framework, providing reference for stakeholders on copyright protection and infringement liability relating to AI-generated works. Voluntary licensing arrangements between copyright owners and AI developers are seen as the answer.

I noted in a blog post last year that the TDM issue was under review in a number of Asian jurisdictions, including India, Malaysia and Korea. In India, a government-commissioned working group recommended a compulsory licence scheme that displeased just about everyone, from the tech gurus to India’s cultural creators, although to the working group’s credit, it also rejected a TDM exception for India. The lack of momentum for this idea means that current copyright legislation continues to apply, albeit with Indian courts starting to weigh in, (but unfortunately in ways that could potentially undermine India’s fair dealing balance). Malaysia continues to attract significant investment from the hi-tech sector despite not yet adopting a TDM exception, in stark contrast to neighbouring Singapore. In Thailand, which has one of the richest cultural traditions in Asia through its audiovisual and music sectors, the tech sector has been agitating for introduction of a TDM, as elsewhere. This would put at risk a content sector that is a significant contributor to national GDP as well as driver of the all-important tourist industry.

Korea is another country where the AI industry is bringing pressure to bear. Korea’s unique and thriving culture is one of its national assets, along with a domestic innovation and hi-tech industry that is second to none. To allow offshore tech companies to plunder Korea’s rich cultural tradition would be a shortsighted and shameful sellout. But neither Korea, nor Thailand, nor Malaysia, India or Australia are opposed to the development of AI, and all want their share of the benefits of this emerging if not already emergent industry. All recognize that AI development requires vast amounts of content, including the kind of high-quality content produced by their respective creative sectors.

There is a proven path for the AI industry to access this rich content. It is called voluntary licensing, and it already exists within the established four corners of copyright law. Moreover, it is increasingly becoming the solution as AI developers finally come to realize (a) that it is not worth risking the viability of the company on an unpredictable lawsuit that could result in crippling damages (or as in the case of India, years of court trials) and (b) equally important, their global push to get a legislated copyright loophole through TDM provisions in national law is going nowhere. In fact, the TDM tide is receding, as the examples of Australia and Hong Kong clearly show.

As the TDM tide goes out, the voluntary licensing tide flows in, floating the boats of both the creative and hi-tech sectors. Australia hoisted the first signal, but other jurisdictions in Asia seem ready to follow.

© Hugh Stephens, 2026. All Rights Reserved.

AI Training and Copyright: Australia Gets it Right—Now it’s Canada’s Turn

Flags of Australia and Canada displayed side by side, showcasing their national colors and symbols.

Image: Shutterstock

In early June Canada issued its national AI strategy paper, “AI for All”. As I noted in a blog post at the time,  while the strategy covered many elements of AI in its 50 pages outlining policy objectives and planned actions, it managed to avoid using the word “copyright” even once. Australia has just come out with its own updated AI policy statement “AI in Australia’s interest”, which builds on its own “National AI Plan”, released last December. But whereas the Carney government in its AI strategy managed to completely avoid putting copyright into the AI equation, Prime Minister Albanese, after discussing the importance of developing AI for Australia, had this to say;

“But let me make this crystal clear: not everything produced in Australia is up for grabs.

Not at all.

Australian writers, musicians, artists and journalists must retain ownership and control of their work.

Our laws will spell that out, plain as day.

An artist’s creative endeavour is their work and their property.

No company should use Australian books, music, art or news to build or train AI without the artist’s control.

That includes the artist’s control of the price and value of their work.

Anything less, is theft.”

Blunt, clear and refreshing. If Australia can protect its cultural community while promoting policies for sensible AI adoption and development, then so can Canada.

Both Canada and Australia currently have no Text and Data Mining (TDM) exception in their copyright law. This legal loophole would allow AI developers to appropriate content without permission for training purposes. In both countries there have been calls from the tech community to introduce a TDM exception, a carte blanche that would allow AI companies to ingest copyrighted content without authorization, payment or even acknowledgement. In its December “National AI Plan”, which is much more analogous to Canada’s “AI for All” than Albanese’s recent short AI policy statement–in that it outlined a range of detailed policy proposals for AI adoption in Australia– the Australian government nonetheless managed to grasp the copyright nettle unambiguously.

Among the issues highlighted under “AI Risks and Harms” was the following:

Reviewing application of copyright law in AI contexts: The Attorney-General’s Department is engaging with stakeholders through the Copyright and AI Reference Group to consult on possible updates to Australia’s copyright laws as they relate to AI. The government has provided certainty to Australian creators and media workers by ruling out a text and data mining exception in Australian copyright law” (emphasis added)

Just as the Australian government has sensibly ruled out a TDM option. Canada needs to do the same, as called for Canadian cultural umbrella groups, such as the Coalition for Diversity of Cultural Expression (CDCE).

So far Canada has danced around the issue. Heritage and Identity Minister Marc Miller has said that “the current copyright law does and should protect those that have created material, and people need to be compensated properly”, but he is just one minister among several. Evan Solomon, Minister of Artificial Intelligence and Digital Innovation, and Minister of Industry Melanie Joly, both have a big piece of this file. One can expect that both can be counted on to be more sympathetic to tech bros than cultural mavens. What is needed is a prime ministerial pronouncement clarifying that Canada’s creative community–artists, writers, publishers, musicians, filmmakers, photographers, journalists and more– is not going to be thrown under the bus on the pretence of keeping Canada competitive in the global AI game.

In the wake of Australia’s announcement that a TDM exception was off the table, the tech industry tried a new approach by suggesting the creation of a centralized fund that would be used to compensate rightsholders for the permissionless use of their works in AI training. Specifically, AI company Anthropic reportedly tied a proposed $15 billion USD ($21.6 billion AUD) investment in data centres in Australia to creation of the creatives fund in order to allow to access Australian content without licensing or negotiation with rightsholders. Australia’s creative community quickly mobilized. Their concerns were heard. Along with setting clear guardrails ruling out the unauthorized use of copyrighted creative works, Albanese has created a new Office of AI within the Prime Minister’s Office, recognizing the need for policy coordination given the breadth of AI’s policy impact. This is something that Canada might consider. It has Evan Solomon, Minister of Artificial Intelligence and Digital Innovation, but there seem to be very few cultural community voices within Solomon’s hearing range.

Australia has the same goal as Canada of getting its fair share of the AI pie while managing AI adoption and its impact on society. But there is one big difference. In so doing, the Australian government has made it clear it will pursue its AI goals while simultaneously respecting and protecting its culture and its creators. Canada’s cultural and creative community deserves no less consideration.

© Hugh Stephens, 2026. All Rights Reserved.

Litigation vs. Licensing for AI Training

Scrabble tiles spelling 'LITIGATION vs LICENSING' on a game board.

Image: Author

There is an ongoing struggle between the tech world of AI training and the cultural world of content creation. It has led to lots of litigation but also an increasing number of licensing agreements, the obvious market solution. Litigation has helped convince AI companies to share some of the wealth by pursuing licensing. Yet the AI world continues to try to find ways to avoid the basic step of seeking permission from rightsholders for using their valuable content to create their products.

Anyone who has seen the striking graphic “Who is Suing Whom in AI”, created by the design website Information is Beautiful, will be struck by the enormity and breadth of the issue which is so cleverly displayed, with the big AI developers such as Perplexity, Anthropic, Meta, Google, Open AI, Midjourney, Cohere and others at the centre with the creators (every content entity from Conde Nast, Getty Images, Universal Music Group, CNN, Disney and Thomson Reuters to Elsevier, Dow Jones, New York Times and others) ranged around the periphery, a stunning visual encompassing more than 100 lawsuits in the United States. That graphic was up-to-date as of June 26 of this year. Since then, at least one more major lawsuit has been filed, by a group of textbook authors against Meta. The graphic does not include the first such case in Canada where a group of media organizations (Canadian Press, Torstar, The Globe and Mail, Postmedia and CBC/Radio-Canada) is suing OpenAI, or the Getty Images case in the UK, or indeed any cases outside the US. From this graphic, it would seem that to resolve the issue of how copyrighted content is going to be used in AI development and training, litigation is the inevitable route. But is it?

As far as I am aware, Information is Beautiful has not created a similar graphic to display the range of licensing deals that have taken place, many of them between some of the same actors that appear on the litigation chart. If they did it would be similar, but encompassing even more licensing agreements than lawsuits. Licensing deals are being struck so frequently it is just as hard to keep up with them as it is to track all the litigation underway. The University of Glasgow’s CREATe Centre says it has documented 274 licensing deals and has a chart that tracks 109 of them. Whatever the number, it is a lot and it is growing. That is not to say that the AI industry has finally accepted the need to pay for the content they are using to create their products, just as they pay for software engineers or data processing capacity. This is where the link between litigation and licensing becomes interesting.

In a perfect world, AI developers would obtain their inputs through the market on the basis of permission, which would encompass both compensation (in most cases) plus transparency or accountability, i.e. documenting what content was used. But we don’t live in a perfect world, which is why we have the rule of law and courts to enforce those laws. In some cases, AI platforms did begin negotiations with rightsholders but when it was not possible to reach an agreement, the AI industry switched tactics and took the content anyway, arguing it was legal to do so for a variety of reasons. This is precisely the scenario that led to the New York Times suing OpenAI. These cases are even more egregious because there was initially a tacit acknowledgement by the user that the content had value. Then, when the price or conditions did not suit the potential licencee, suddenly it was okay to take the content anyway under the guise of fair use. Various arguments have been deployed ranging from the claim that no copying actually occurs, to the dubious assertion that what is copied is data not content, to the invocation of the US “transformation” doctrine.

On the issue of copying, a study by the Atlantic (AI’s Memorization Crisis: Large language models don’t “learn”—they copy. And that could change everything for the tech industry) convincingly demonstrated the uncomfortable truth that LLMs can reproduce long excerpts from books they have been trained on. The inputs are not just ones and zeros, they are content— someone else’s content that was taken without permission. Whether the use was fair according to US fair use interpretations is still an open question. US courts and other countries are trying to come to grips with this issue. In countries such as Canada or Australia, where there is no statutory copyright exception for Text and Data Mining (TDM) that would permit permissionless AI training on content, the AI industry has been floating various workaround proposals. The “incentives” would include (in Australia) establishing a government-managed fund to compensate rightsholders according to some sort of formula, plus investments in AI data centres. What is missing from proposals such as this is the concept of permission from those who actually own the content, or even discussion of the proposal with them. As Prof. Rod Sims, former Chair of Australian Competition and Consumer Commission, put it in a recent opinion piece in Canada’s National Post, “what other sector refuses to negotiate with suppliers and instead goes to government to bypass such a step?”

Let me use a food industry analogy to make the point even more clearly. When you run a restaurant you have labour costs, rent, taxes, etc. and the cost of ingredients to consider. You don’t get to raid the farmer’s field to obtain your inputs for free, just because you are able to root out crops without the farmer being able to stop you or even know it is happening. Setting up a fund to “compensate” farmers for their stolen crops, on terms set by the government rather than the market, doesn’t even begin to make this right. Legalization of this theft would remove any possibility of litigation or legal protection, for the farmer—or for content owners. Litigation, while protracted, costly and potentially leading to uncertain outcomes, is nonetheless the stick that is needed to facilitate licensing.

The obvious route for the AI industry to take is to license the content they want to use. That may not seem as “efficient” as just taking it for free but with the threat of litigation hanging over the proceedings, licensing suddenly becomes the more efficient alternative. It is also win/win for both AI developers and the content industries. And, it is simply the “right thing to do”.

© Hugh Stephens, 2026. All Rights Reserved

I am pleased to note that this blog was recognized by Feedspot as being among the “40 Best Copyright Blogs to Follow in 2026”. In fact, we hit the middle of the pack at No. 20. I am honoured to be included in such distinguished company.  

Feedspot is an RSS Reader that lets readers subscribe to blogs, news sites, and any website they wish to follow.

Copyright Developments in New Zealand: Going in the Right Direction

Flag of New Zealand featuring a blue field with the Union Jack in the canton and four red stars with white borders representing the Southern Cross constellation.

Image: Wikimedia (Public domain)

New Zealand is proposing to introduce a number of optional updates to its Copyright Act when it enacts required changes to bring legislation into compliance with two treaties it has signed. This is good news for creators. Still to be addressed, however, is the thorny issue of AI training on copyrighted content.

New Zealand needs to make some required legislative changes to its Copyright ordinance as part of implementing two treaties it has signed, the UK-New Zealand Free Trade Agreement (FTA) and New Zealand’s FTA with the European Union. In both cases New Zealand has agreed to extend its term of copyright protection from life of the author plus 50 years to “life plus 70”, as well as preventing the circumvention of TPMs (technical protection measures, aka “digital locks”) except in specified narrow situations. These provisions must be enacted by May of 2028. They will bring New Zealand’s copyright law into alignment with most of its major trading partners. However, while there is a legal requirement to address the above two issues, the Ministry of Business, Innovation and Employment (MBIE) has proposed that a number of other copyright issues also be addressed as part of the process of updating the Act. These include;

  • supporting not-for-profit gallery, library, archive and museum (GLAM) organisations to preserve and provide access to collections, including by allowing use of orphan works, making digital copies for preservation and access, and applying research and private study copying rules across all GLAM organisations, with safeguards for copyright owners
  • introducing a new fair dealing exception for parody and satire, applying across a wide range of works while maintaining authors’ moral rights
  • providing courts with a framework to order internet service providers to block access to overseas websites primarily engaged in copyright infringement, with appropriate safeguards and flexibility
  • removing an outdated peer-to-peer file-sharing enforcement regime that is no longer used, reducing compliance costs for internet service providers
  • enabling copyright licensing organisations to take collective action on behalf of copyright owners to prevent infringement
  • clarifying that the first distribution right is only exhausted where the copyright owner has consented to the overseas sale of copies, supporting control over parallel imports of infringing copies
  • changing the default rule for commissioned works so that creators are the first copyright owners unless agreed otherwise
  • extending resale royalty rights for visual artists by 20 years to align with the longer copyright term.

It is encouraging to see New Zealand take this opportunity to review and update its copyright framework while it implements the needed changes to meet its trade agreement commitments. Canada was also required to extend its copyright term as a result of the new NAFTA agreement with the United States, and it did so, at the last minute. However, it did the minimum required and passed on the opportunity to address wider issues, of which many have been identified by Parliamentary committees, while more are coming forward as a result of developments in AI.

The proposed changes in New Zealand should be welcomed by the copyright and copyright-using community. They will provide legal protection for the sort of digital replication that the GLAM sector needs to preserve older and orphan works, although more information on what how the research and private copy rules will be implemented is needed. Widening fair dealing to include satire and parody has been done in a number of jurisdictions, and this will bring New Zealand in line with other Commonwealth countries like Australia, Canada and the UK that have such exceptions (“parody, caricature, and pastiche” in the wording of the UK legislation). In the application of the defence, New Zealand courts should follow the Australian lead, where courts have kept a tight rein on this defence. Parody is a tricky exception to invoke, as a recent UK case well illustrates. The moral rights of the author are also a factor to consider.

For the first time, site-blocking (that is, requiring ISPs to block pirate offshore websites, after legal review) will have a firm foundation in New Zealand law. Australia has had such legislation on the books for more than a decade, and the UK for longer than that. Both the UK and EU treaties required New Zealand to allow the courts to issue injunctions “against an intermediary whose services are used by a third party to infringe intellectual property rights.” Canada has dealt with this issue through the courts exercising their inherent jurisdiction without the enactment of specific site-blocking legislation, with initial challenges from some ISPs being dismissed on appeal. The process has now become routine. It seems the New Zealand government intends to ensure clarity by amending copyright legislation to “provide courts with a framework to order internet service providers to block access to overseas websites”. IP scholars in New Zealand, such as Prof. Graeme Austin, have been calling for the government to take the lead. It seems they have been heard.

The empowering of collective management organizations (CMOs) to take legal action against infringers on behalf of their members is also an important step. Under present provisions, CMOs cannot bring actions because they do not hold the rights to individual works. This requires multiple authors either to take individual actions or join in a joint action. Given the cost of such an exercise, this is not feasible (large publishers who have licensed rights from authors may be in a position to do this, but authors themselves are hamstrung). Giving their collective management organization the right to represent them is a positive move. This is a move that Canada could well replicate to enable CMOs like Access Copyright to represent authors.

Changing the default rule for commissioned works will, for example, give photographers greater control over their work. Clients can contract for the right to display copies of the work but the copyright in the original work will belong to the creator. The same is true for artistic works unless there is a specific agreement that the work is created under an employment contract. Canada enacted this provision in 2012 when it passed the Copyright Modernization Act. Extending the resale royalty rights for authors to match the longer copyright term keeps these two provisions in alignment. New Zealand, like Australia and the UK, and EU member states, has enacted an Artists’ Resale Right (ARR), which allows a small portion of the proceeds of a resale of artwork through a professional dealer to be paid to the original artist (or their estate). Canada has been promising for several years to enact an ARR but has not yet done so.

The one big issue this round of copyright amendments will not address is use of copyrighted content for AI training. That is a rapidly evolving issue in many countries and is a moving target. The solution, as suggested in this article by Prof. Austin, is to foster market solutions, that is facilitating the licensing of content to AI developers. The way not to do this is to provide a wide exemption for AI training, as many in the tech world are advocating, but to ensure that rightsholders have the right to protect their content and to grant access to it on terms that they agree to. This is already happening in a number of areas such as licensing agreements between major publishers, news enterprises, and the AI industry, but individual authors are still being left out of the discussions.

 Australia has just ruled out creating a fair dealing exemption for AI training (known as the TDM or Text and Data Mining exemption). Even the notoriously anti-copyright Productivity Commission supports this position. Such an exemption would remove any incentive for AI developers to negotiate with rightsholders for use of content. Hopefully New Zealand will follow suit in this regard. While we will have to wait for further developments when it comes to dealing with AI issues, the current set of proposals will be very useful in renewing and updating the copyright framework in New Zealand.

© Hugh Stephens, 2026. All Rights Reserved.

The AI Copyright Crisis Contains an Opportunity for which Publishers have Waited Centuries

A promotional graphic for Citations LLC, featuring the tagline 'Rights-aware AI access infrastructure' and three services: REVEAL™ (Semantic extraction engine), CITATIONS GATEWAY™ (Access & transaction engine), and CITATIONS CORE™ (Settlement & analytics engine). The design has a dark blue background with gold text.

We read daily about new lawsuits brought by rightsholders against AI developers, strategy papers floated by governments seeking to solve the riddle of reconciling copyright and AI, and declarations issued by authors proclaiming the end of human creativity. The creative community seems to have coalesced around the principles of transparency, permission and remuneration but the tools to effect those key elements remain elusive. The AI community would generally prefer not to pay or ask permission but is gradually accepting the need to license content. Yet there is still a technical gap in terms of knowing what content has been used, when and how. Without that knowledge, the principles of permission and remuneration are left treading water. The blog post below by Jim Bryant, Co-Founder and CEO, Citations LLC, offers potential solutions to this challenge, and I offer it to you as a possible pathway forward. I have no financial interest in Citations LLC, nor did they pay me to post this information. It is presented as a contribution to the search for a world where copyright and AI can co-exist for mutual benefit. (Hugh Stephens)

A problem or an opportunity?

Imagine a student in Montreal asks an AI assistant a question about traditional Chinese medicine, in French. The AI answers fluently — in French — drawing on the Encyclopedia of China, a monumental work with over 125 million characters that has never been translated into any language.  Now imagine the same student switches to English and asks a follow-up question. The AI answers again, equally fluently, in English. The student is satisfied. The publisher gets nothing. No notification, no attribution, no compensation. They don’t even know it happened.

This scenario is entirely plausible with current AI technology. And while it represents a genuine copyright problem — real-time AI translation of a protected work, without license, in a jurisdiction whose law was not written to contemplate it — it also represents something else: an extraordinary, unrealized opportunity.

For the first time in the history of publishing, the technology exists to know, at the moment it happens, that someone in Montreal, Mumbai, or Mexico City is asking a question that your content just answered. The question is whether publishers will help build the systems to capture that signal — or whether they will leave it entirely to the AI companies, who are already building without them.

Publishers have always been flying blind.

Think about what publishers have never been able to know — and what AI companies, for the first time, can. An AI system that has trained on your works without permission is, in effect, drawing on your content every time it answers a relevant question. Which of your titles is it using right now, and where? Which readers are getting answers derived from your content without ever being directed back to the original? Which backlist titles are generating AI responses in markets where you have no distribution and no visibility? Which gaps in your catalogue are readers repeatedly trying to fill — and how would you know, if the only signal is buried inside a system you have no access to? The argument for independent monitoring infrastructure is not only about compensation. It is about visibility. Publishers are currently funding AI responses with their content and receiving nothing in return — not money, not data, not even the knowledge that it is happening.

For centuries, publishers sent their works into the world and largely lost sight of them. Sales data arrived months or years later, filtered through agents, booksellers, distributors, and described what sold — not what readers wanted but couldn’t find. The feedback loop from reader demand to editorial decision has always been slow, indirect, and incomplete.

A properly instrumented knowledge access infrastructure changes all of that. Real-time query data across AI systems is, in effect, a continuous signal of what readers want — more granular, more current, and more honest than any market research tool the industry has ever had. That data is a byproduct of the same system that creates the copyright exposure publishers are currently fighting in court.

The moment of demand is the moment to act.

Here is the specific opportunity that AI creates, and that no prior technology has made possible: when an AI system surfaces content in response to a query, it creates a demonstrated moment of demand. A reader who just received an AI-generated answer drawn from a specific book is, at that moment, maximally interested in that book. That is the moment to offer them the chance to borrow it from a library, purchase it from a retailer, or access an authorized digital edition.

Rather than substituting for the book, the AI interaction becomes the discovery mechanism that leads to it. Publishers have spent decades trying to close the distance between the moment a reader becomes interested in a title and the moment they act on that interest. AI closes that distance to zero — but only if the infrastructure exists to capture it. Without that infrastructure, the moment passes, the reader moves on, and the publisher never knew the opportunity existed.

Libraries are being bypassed — and publishers are losing their best customers. Libraries are among the largest single customers some publishers have. A major academic or reference publisher may depend on library subscriptions for a substantial share of its revenue. AI is disrupting that relationship in ways that have received too little attention. When a patron who would previously have borrowed a book — or prompted their library to acquire it — instead receives an AI-generated answer derived from that same book, the library never makes the purchase, the publisher never sees the revenue, and neither institution knows the transaction occurred. The AI company captures the value; the library loses a use case; the publisher loses a sale. The institution most structurally committed to legal, compensated access to knowledge is being systematically bypassed by systems that obtained that knowledge without payment.

The same logic applies to real-time trend identification. Aggregate query patterns across an AI knowledge system are a leading indicator of what readers want — not what they bought last quarter, but what they are looking for right now. Which subjects are rising? Which titles are being asked about in markets where they have no distribution? Which authors are generating interest that isn’t yet reflected in sales? This intelligence, continuously available, would transform publishing from a reactive industry into a responsive one.

The translation question is the hardest — and the most important.

The Encyclopedia of China example is worth dwelling on, because it illustrates both the opportunity and the complexity in their sharpest form. That encyclopedia has never been translated — into French, English, or any other language. The economics of translation have made it prohibitive: 125 million characters, uncertain commercial return, no obvious path to a global audience. As a result, it has been accessible only to readers of Chinese. That constraint has nothing to do with the quality or the importance of the content.

AI removes that constraint entirely. In this hypothetical, a reader anywhere in the world could ask the encyclopedia a question in their own language and receive an answer. This is, genuinely, one of the most remarkable things that AI makes possible: the dissolution of language as a barrier to knowledge, overnight, at no marginal cost.

But it raises a set of copyright questions that existing law is not equipped to answer. A real-time AI translation is, in the most precise legal sense, the creation of a derivative work — at the point of query, in a foreign jurisdiction, without a license, without attribution, and without compensation to the original publisher. It is not covered by any existing text-and-data-mining exception, because it is not mining — it is real-time derivation. It is not covered by fair use or fair dealing analysis that was designed for static reproduction, not dynamic on-the-fly translation.

And yet the underlying interest of the publisher is not to prevent this from happening — it is to be compensated when it does, and to have some say in how their content is represented. A framework that would allow the publisher of the Encyclopedia of China to authorize AI-mediated translation under defined conditions, receive a per-query payment, and have the source attributed, would serve everyone’s interests. The absence of such a framework means the publisher gets nothing, the AI company gets everything, and the reader gets an answer of uncertain provenance.

The ten copyright challenges — briefly.

It is worth cataloguing the specific challenges, because they are often discussed in isolation when they actually share a common cause. The publishing industry currently faces at least ten major copyright issues arising from AI:

AI training — whether training on copyrighted works requires permission and compensation, currently being litigated in multiple jurisdictions.

Transparency — AI developers do not disclose what content their models were trained on, making it impossible for rights holders to assess exposure or negotiate terms.

Reproduction — models can and do reproduce passages that closely approximate protected expression, as documented in peer-reviewed computer science research.

Market impact — AI summaries and Q&A responses can substitute for the original work, displacing revenues that would otherwise flow to the publisher.

Derivative works — the degree of transformation required to render AI output non-infringing remains genuinely unsettled, particularly for outputs that blend multiple protected sources.

Attribution — AI outputs routinely fail to identify the works they draw on, undermining both the moral rights of authors and the practical basis for any royalty mechanism.

Compensation — no industry standard governs AI licensing fees; per-query, per-token, and blanket models are all being proposed, with no settled framework.

Retrieval — retrieval-augmented generation systems access copyrighted content at inference time, raising rights questions distinct from and additional to those arising from training.

International law — training data crosses borders; copyright law does not; EU, US, UK, Japanese, and Canadian frameworks diverge in ways that create genuine compliance complexity.

Auditability — without verifiable records of what was accessed, when, and in what context, no licensing agreement is enforceable and no royalty calculation is credible.

These are not ten separate legal problems. They are ten symptoms of one missing piece of infrastructure: a neutral, independent system for monitoring how AI systems access and use copyrighted content, reporting on that usage in real time, and enabling settlement between AI platforms and rights holders on the basis of verified data rather than estimates.

Why the infrastructure must be independent.

This point deserves emphasis, because there is a tempting shortcut that would not actually work. Publishers cannot rely on AI developers to build and operate the systems that monitor AI’s use of their content. The conflict of interest is structural: the party whose compliance is being measured cannot be the party doing the measuring.

What is required is a neutral layer — operated independently of both AI developers and publishers — that records access events, aggregates usage data, reports to rights holders, and enables automated settlement. Think of it as the knowledge economy’s equivalent of a financial clearinghouse: not owned by any single participant, trusted by all of them, and essential to the functioning of the market.

This is not a novel concept — it is exactly the model that makes collective rights management organizations function in the music industry and payment card networks function in financial services. Every industry that has needed to account for consumption at scale and distribute revenues to multiple rights holders has eventually built a neutral clearinghouse.

The window is open — but not indefinitely.

Canada’s AI strategy, recently released, makes almost no mention of copyright or the rights of content creators — a significant omission that Hugh has written about on this blog. The European Parliament’s work on AI and copyright has moved further, but still focuses primarily on training rather than on the access and retrieval layer where the most tractable opportunities lie.

The practices governing how AI systems access knowledge are being established right now, largely by default. The companies building AI systems are not waiting for a legal or regulatory framework; they are building, and the norms are hardening around what they build. Publishers who are not at the table when that infrastructure is designed will find themselves subject to whatever framework others have built for them.

Copyright law exists to balance access and incentive — to ensure that knowledge can circulate while the conditions that make knowledge production sustainable are preserved. AI does not change that objective. It changes the technical conditions under which the balance has to be achieved. The good news is that those technical conditions, for the first time, make real-time monitoring, attribution, and settlement not just possible but straightforward.

The question is not whether AI will access books. It will. The question is whether publishers will be watching when it does — and whether they will have built the systems to act on what they see.

That system already exists. It logs the moment, attributes the source, and settles the account — not as a future framework, but as infrastructure operating today. It’s called Citations, and it’s already watching.

* * *

About the author

Jim Bryant is the co-founder and CEO of Citations LLC, which has built the independent infrastructure for rights-aware AI access to authoritative content — enabling real-time monitoring, attribution, and settlement between AI platforms and publishers. See how it works at: citationslogic.ai.  Jim previously founded ProCD, one of the first CD-ROM reference publishing companies; managed Information Please, which became one of the most visited reference destinations of the early internet; and founded Trajectory, which developed and deployed natural language processing algorithms to read and extract structured metadata from over one million books in English and Chinese.

(c) Citations LLC, 2026

Canada’s National AI Strategy “AI for All”: Does Copyright Exist in the AI World?

A futuristic robotic figure with glowing blue accents, portrayed in a tech-inspired environment. A 'no copyright' symbol is visible in the corner.

Image: Shutterstock.com (adapted, clumsily)

If it does, you would never know it from reading Canada’s new AI strategy just released by the Minister of Artificial Intelligence and Digital Innovation, Evan Solomon. It is a magistral document, addressing key elements of AI under six pillars: (with my shorthand summary in brackets)

  • Protecting Canadians and safeguarding democracy (addressing trust, safety and privacy concerns)
  • Empowering Canadians (promoting AI literacy and economic opportunity)
  • Powering AI adoption for shared prosperity (accelerating adoption, especially for SMEs)
  • Building a sovereign AI foundation (building domestic compute, cloud and connectivity infrastructure)
  • Scaling Canadian champions (more government funding for domestic AI development)
  • Building trusted economic and governance partnerships and global alliances (leading the creation of a multinational middle power alliance to curb the power of hegemons and hyperscalers)

The latter objective will no doubt go down really well with the Trump Administration!

Those six headings cover just about all aspects of AI, from its creation to its use to its impact on the economy, on society and on individuals. But in all 50 pages of the document, as far as I can ascertain, you won’t find the word “copyright”, although “protecting intellectual property” is certainly featured. The intellectual property rights that are mentioned have nothing to do with the rights of those whose content was used without authorization to create AI but rather relate to protecting the intellectual output of AI developers in Canada. John Degen, CEO of the The Writers’ Union of Canada (TWUC) was the first to call this out. Given the make-up of the task force that produced the report, this is not surprising. While it was made up of the great and the good from the AI world, with academics, financiers, CEOs, cybersecurity experts, innovators, educators and so on as part of the roster, there was not a single representative from the cultural community.

There are many elements of AI this document tries to address, all of them important to a country like Canada, although there are limits to what can be done by a middle power given that the lead on development has been seized by a handful of large companies, mostly in the US. The US government itself is caught in the dilemma of wanting the US to lead AI development yet not becoming overwhelmed by it to the point that a few major corporations are calling all the shots.

As for content issues, including what must surely include some copyrighted content, they are addressed only indirectly in the Canadian strategy. The three principal issues relating to content are; (1) privacy and access to data; (2) Canadian identity and culture; and (3) AI misuse, such as creation of deepfakes and misinformation.

On privacy and data, the document notes that AI is only as powerful as the data it can access (how true!). It reminds us that governments in Canada hold vast amounts of data that should be treated as a strategic national asset and mobilized to fuel innovation and productivity (i.e. provided for AI research). Thankfully, there is a tip of the hat to the need for “strong privacy protections” but there is no mention of the unauthorized scraping of databases and protected content by AI developers, both domestic and international. Privacy is important but so is ownership of content, and the right to grant permission to use it. Unfortunately, this latter point is not mentioned.

Protecting and promoting Canadian identity and culture is also mentioned as an important goal. It is obvious that if AI developers are blocked or hindered from ingesting Canadian content, then there will be less of Canada reflected in AI outputs. That argument was put forward recently by Michael Geist in a blog post criticizing recommendations issued by the Parliamentary Standing Committee on Heritage that had called for protection of the property rights and interests of artists through the Copyright Act on the basis of authorization, remuneration and transparency. This would lead to “AI without Canada”, according to Prof. Geist. This could be true if AI developers did not need or want curated Canadian content, but they do. The solution, as I pointed out, is not to give away everything in the shop window by creating a broad AI training exception in Canadian copyright law–which would amount to legalized theft, but instead to facilitate licensing solutions by resisting the smash-and-grab. Applying the existing legislation will incentivize the AI industry to strike deals with rightsholders. In other words, they will pay a negotiated amount for the products on display. That’s the best way to get more Canadian content into AI.

On the identity issue, the government’s summary document has this to say:

“Canadian AI must support, reflect, and project Canadian culture, which includes our customs, our history, and our heritage. Canadian voices, languages, communities, and knowledge must also be represented in how AI systems are designed, built, and used. Given our diverse and multicultural society, our approach to AI must acknowledge and support this rich diversity, including strengthening the French language by capturing and projecting its idioms, expressions, and cultural contexts.”

The best way to do this is to ensure that quality content in both official languages is made available to AI developers. As I have stated above, the fairest and most efficacious way to do this is through content licensing. Broad copyright exceptions will not facilitate licensing discussions. In fact, they do just the opposite by encouraging avoidance of dealing with rightsholders.

Regarding misinformation and deepfakes, this is a huge concern, and not just in Canada. Various legislative solutions have been proposed such as the bipartisan NO FAKES Act, currently working its way through the US Congress (opposed, as usual, by the internet libertarian organization, the Electronic Frontier Foundation). Other countries, such as Denmark, are addressing the issue through amendments to copyright law, giving individuals the reproduction rights to their image and voice. The UK has an anti-deepfake law on the books, introduced earlier this year, but Canada is still struggling to get its Online Harms legislation, after a couple of false starts, finalized and across the line. Re-introduction of that legislation is expected imminently, and will likely include social media restrictions on children, a highly controversial issue.

Privacy in relation to access to data, cultural identity, and misinformation including deepfakes are all content issues that Canada’s AI strategy will need to address. And so is copyright, although not mentioned in the strategy. Putting the best possible gloss on things, perhaps it is just as well there was not some throwaway line in the strategy pointing to the need to provide wider access to copyrighted content to ensure that Canada remains competitive on AI. That is the argument often employed by those who want freer access to “OPC” (Other Peoples’ Content). The argument is that “Everyone else is doing it (i.e. giving it away–which is factually untrue), so we have to as well in order to stay competitive”. Maybe silence was better than saying the wrong thing in this document.

In the absence of any reference to copyright issues, the last word must rest with Heritage and Identity Minister Marc Miller who spoke recently to the press after the National Summit on Artificial Intelligence and Culture in Banff, AB. The Minister is quoted as saying that Canadian copyright law is already clear that artists’ work needs to be respected, and that…”the current copyright law does and should protect those that have created material, and people need to be compensated properly.”

While that is encouraging, it would have been nice to have had this reaffirmed in the AI strategy document.

© Hugh Stephens, 2026. All Rights Reserved.

Like Wasps at a Picnic: (Distracting from the Canadian Heritage Committee Report on AI and Creative Industries)

Close-up of a wasp drinking from a metallic surface with blurred green background.

Image: Pixabay.com

It was as predictable as wasps at a picnic. Within days of the Canadian Parliament’s Heritage Committee releasing its report on “The Impact of Artificial Intelligence on the Creative Industries”, with its lead recommendation being (my highlights)…

That the Government of Canada protect the property rights and interests of artists through the principles of the Copyright Act, in accordance with the ART principle—authorization, remuneration and transparency:

a) The Government of Canada must take the necessary steps and ensure that the scope of the Copyright Act applies to AI-generated content in order to guarantee copyright protection.

b) The Government of Canada must mandate greater transparency from AI developers regarding copyrighted works used to train their models, including disclosure of training data sources, to enable proper authorization and licensing.

c) The Government of Canada must establish a clear opt-in consent requirement for the use of copyrighted works in the training of artificial intelligence systems, ensuring that creators’ works may not be used for text and data mining or model development without their prior authorization.

…prolific tech and copyright commentator Michael Geist of the University of Ottawa was attacking its conclusions, issuing warnings that unless the tech industry is allowed (without authorization or compensation from rightsholders) to help itself to copyrighted content for the purpose of AI training, we will have “AI without Canada”. In other words, unless the tech industry is allowed to plunder Canadian content in the same way that it has been doing to date in the US (although this is meeting legal challenges and is quickly changing as licensing solutions take hold), there will be less Canadian content in the training data. This, apparently, will leave Canada as an “outlier” compared to peer jurisdictions. The AI developers will turn their back on Canada and rush off elsewhere. (This is a standard threat deployed by the AI industry to play off one country against another). He cites the EU, Japan, Singapore and Israel, as well as the US in support of this interpretation. Not mentioned as “peer jurisdictions” are the UK and Australia but then that would not have served the purpose of his narrative. Australia has recently declared it will not be legislating a Text and Data Mining (TDM) exception to its copyright laws to legalize unauthorized ingestion of copyrighted works for AI training, while the UK has just hit the pause button on a series of ill thought-out and badly received proposals to allow AI developers to freely use copyrighted content to train their AI algorithms unless rightsholders specifically opt out.

Singapore and Israel are among a small minority of countries that, under US pressure, have adopted US-style fair use laws that potentially allow for a weakening of copyright protection through a hodge-podge of court rulings. While many cite Japan as a jurisdiction that has given carte blanche to tech interests and AI developers, the facts are quite different as I pointed out in this blog post a couple of years ago. Japan has a strong cultural industry that it wants to nourish and protect and has defined its TDM exception very narrowly and carefully. The EU, has two provisions in its Copyright Directive related to AI training (Article 3 which permits TDM carried out only for non-commercial scientific research purposes, and Article 4, which permits TDM for any purpose, including commercial, as long as rightsholders have not opted-out, subject to strict transparency provisions by AI companies). Both impose constraints on AI developers, although there are differing views on opt-out.

Opting-out may sound like a compromise that both rightsholders and the AI industry could support but Britain’s example demonstrates otherwise. In its now aborted public consultation, the UK government put forward several options including its “preferred” option of opt-out. Fully 97 percent of respondents, from both the tech and creative communities, trashed this option. For creators, opting out not only stands copyright on its head (it is a property right, so why should holders of that right be required to notify someone who wants to infringe on that right that they may not do so, i.e. it’s like passing a law allowing anyone to picnic on my front lawn unless I post a “No Trespassing” sign), but it is technically difficult to do, especially for individuals and small-scale rightsholders. The robots.txt protocol is not binding and is in many cases not very effective. The tech industry doesn’t like opt-out because it imposes constraints on their untrammelled ability to access anyone’s copyright-protected content, anywhere, anytime. Instead the Committee recommends “a clear opt-in consent requirement” for the use of copyrighted works in the training of artificial intelligence systems.

Now it’s my turn to quibble. IMHO, there should be no explicit need for a rightsholder to “opt in”. I think that Canada’s copyright laws, properly interpreted, already provide sufficient protection to prevent unauthorized use. A rightsholder can “opt in” to AI training or any other unauthorized use not subject to fair dealing by granting a license to use their content. If that is an “opt-in” requirement then I am in favour. If yet another opt-in step is required, this would seem to be unnecessary. Licensing is a growing phenomenon. AI developers want reliable, curated content to develop their applications. As long as they are prevented from simply helping themselves, there is incentive for them to reach licensing deals with content owners. However, giving the tech industry a pass by allowing themselves to take for free whatever they want in the name of developing AI applications (for their commercial advantage) removes the needed incentive to negotiate with rightsholders. As to whether unauthorized use for AI training constitutes fair dealing, as Dr. Geist claims (“most TDM for AI training purposes would likely qualify as fair dealing under existing law”), this is doubtful to say the least. It is hard to imagine which fair dealing purpose currently applicable in Canadian law (research, private study, education, parody or satire, criticism or review, news reporting) would apply particularly when there are fair dealing limits to the amount of a work that can be used for such purposes, and specific factors that must be applied as to the effect of the dealing on the work.

The Committee’s lead recommendation is not the only complaint that Dr. Geist has about the Committee’s report. He feels it is unbalanced because the majority of its witnesses represented the cultural industries. It’s true that its lead recommendation is very much in line with the mainstream views of the Canadian cultural community.  It was, after all, the Report of the Standing Committee on Canadian Heritage. This reminds me of the conflicting reports on copyright issued a few years ago by the Heritage Committee and its counterpart the INDU Committee. The 2019 Heritage Committee report, titled Shifting Paradigms, was attacked at the time by Dr. Geist as “the most one-sided Canadian copyright report issued in the past 15 years”. He claimed that there was “no attempt to engage with a broad range of stakeholders”, even though he himself appeared along with a number of others who shared his perspective on copyright. Shortly after issuing its own report, the INDU committee then issued a tone-deaf “We’re in charge” press release reminding the world that it had “sole responsibility” for administering the Copyright Act. (This is not strictly accurate). Dr. Geist’s main complaint, whether with “Shifting Paradigms” in 2019 or the current Heritage Committee report seems to be that the Committee members, in their wisdom, did not take his expert advice.

What is the function of Parliamentary Committees? It is to hear evidence, draw conclusions and make recommendations. He complains that while there were different points of view, including notably his, on how to tackle the issue under study, the Committee’s conclusions did not reflect these views. Was it because, numerically, there were more pro-copyright witnesses from the creative community that those from the Geist camp? That is theoretically possible if it were just a mathematical exercise of adding up comments in a pro and con column. But that is not the case. While the Report made a conscientious effort to capture the full range of comments, including those of Dr. Geist, in the end the members (from three political parties) made a judgement and reached consensus conclusions. (Although the Conservative Party members provided their own addendum that added to but did not refute the Committee’s conclusions). Presumably the members of the Committee were more convinced by the force of the arguments presented by some witnesses than others. Given the range and similarity of concerns presented by disparate members of the creative community it is not surprising where they came out in terms of conclusions.

Dr. Geist is entitled to disagree with these conclusions and recommendations. To be fair, his blog commentary echoes the position he presented to the Committee, except for his complaints about process. As I said at the outset, his attack on the Committee’s report is entirely predictable, like wasps at a picnic. And those wasps can be so annoying, distracting from the main event with the occasional bite and annoying buzzing, but as any determined picnic-goer knows, it’s important to not let them become the centre of attention. The Heritage Committee’s report was carefully considered and drafted by an all-party group after hearing from a wide range of experts. It provides important recommendations that the government would be well advised to take into account as it develops a legal framework in which both the AI and creative industries can co-exist and flourish.

© Hugh Stephens, 2026. All Rights Reserved

An AI Bot Rewrote my Blog Post—And then Gave Me a Failing Grade for Credibility!

A humanoid robot sitting at a desk, using a typewriter while looking at a sheet of paper in a cozy, modern interior with soft lighting.
Image: Shutterstock.com

I don’t know whether to feel offended or flattered, but I’ve been scraped–by AI. And I can prove it. The first intimation I had of this signal occurrence was a notice from WordPress asking me to approve a comment on my recent blog post, “Copyright, AI and the Legal Profession: Who Blinked?”. I logged in to find it wasn’t a comment but rather a link to this website.

A quick click took me to the article “CanLII settling with Caseway signals shift in legal-tech power dynamics”, dated April 20, the same day I had posted my blog. It was under a byline “London News”, which initially I naively assumed referred to London, Ont, (shows how parochial I can be) but quickly realized that this was some kind of online journal for commuters heading toward Picadilly Circus. London News appears to be written by a bot called Noah News Service, managed by the company HBM Advisory, based in London (England). There was no direct reference or link to my blog post in the article, but when I read it, it seemed eerily similar. The words were all different but the thread (with one exception that I will come to later) was the same. When I searched further, I found a footnote indicating the London News story was “inspired by” my blog post. What does this mean in reality?

My original post is protected by copyright, but anyone (even a bot I suppose) can take “inspiration” from a copyrighted work and produce something new. However, the “inspiration” I provided the bot is substantially different, in my view, from the sort of inspiration I would get from reading, say, an Agatha Christie mystery and then deciding to write my own mystery novel. In the case of my blog post, the bot did not really take “inspiration” from the content to create a new original work but rather engaged in rewriting the story using AI analysis of its key points to recreate what I had said using different words. That’s not true inspiration; it’s paraphrasing. Moreover, I’ll wager that an unauthorized copy of my work was made in order to feed the content to the bot to undertake its rewrite. While facts cannot be copyrighted (only someone’s expression of the facts), this rewrite was not based on the facts of the case. It was based on my blog post. Although the bot has not hijacked my precise words (i.e. my expression) it has nevertheless replicated the structure of my work, its flow and its arguments. It’s sailing very close to the wind, but probably still legal. This is not dissimilar to the challenge faced by news organizations who find their expensively created content being scraped and repackaged by online platforms such as Google, META, and others. According to the National Post, in a recent survey commissioned by News Media Canada, more than seven in 10 Canadians (of those surveyed) think the federal government should prevent artificial intelligence companies from taking and repackaging news content without permission or compensation.

But back to the London News article. Scrolling down to the end, I found an analysis of my blog post, produced by Noah. The post was rated according to various categories. It earned a “Freshness Check” score of 8/10 (i.e. the story was relevant), a “Quotes” check of 7/10; a “Source Reliability” score of just 6/10, a “Plausibility” rating of 8/10 but, sadly, an Overall Assessment for credibility of “Fail”, based on a “Medium” degree of confidence in this assessment. OMG, where did I fail to make the bot happy? How did I not meet its standards?

The Source reliability score would have been higher, according to the bot, if it had been published by an “established news organisation”, rather than on a personal blog;

While the author, Hugh Stephens, has expertise in international copyright issues, (thanks, bot) the blog’s content is not subject to the same scrutiny as mainstream media.”

Well, I can live with that. The whole point of a personal blog is to offer a different perspective from Fox News, the BBC or the Globe and Mail.

The bot’s analysis continued:

“The article references reputable sources, but the lack of direct links to these sources raises concerns about transparency and verifiability.”

In other words, stuff your blog post with direct links to “mainstream media” and you might improve your report card. I could do that, but it might not be appreciated by my readers. The need for more direct links is repeated in the Quotes section (Score: 7/10) as well.

As for my failing grade, the bot’s summary says;

“The article provides a speculative analysis of the CanLII-Caseway AI settlement, referencing reputable sources but lacking direct links for independent verification. Its opinion-based nature and the author’s personal blog platform contribute to concerns about reliability and independence. Given these factors, the content does not meet the standards for publication under our editorial indemnity.”

But they published it anyway, as they do all kinds of content scraped from the web. I am not sure what the editorial indemnity policy is, but I suppose it is some sort of guaranteed reliability indicator, designed to separate the loony conspiracy theories (alternate facts?) from “real news”.

I wondered who would pass the bot’s scrutiny. Of the ten AI related stories posted on the front page of London News on the day I selected, 5 passed, 4 failed, and one was Conditional. The sources were all specialized but non-mainstream tech publications, or informed blogs, but certainly not conspiracy-theory outlets. Yet about half failed to gain Noah’s approval. I started to feel a bit better. Perhaps I’m not such an outlier.

I wonder if could write a blog post that would get an “A” from the bot. First, I would have to catch its attention, which I guess I could do by making sure there were lots of references to “AI” in the text, and then I would have to suppress my instinct to offer views on the topic. I would also have to stuff in lots of links to mainstream sources, like the Guardian and its ilk. But what is the fun in that? And what is the point? If people want to read “just the facts”, they can turn over the screening of content (and thinking) to their mainstream media subscriptions. However, I will say that the idea of assessing the reliability of a story on any topic, whether it’s on AI or the war in the Middle East, is not a bad thing. In the case of HBM, the assessment is used as a teaser to convince users (individuals, but more likely businesses) to sign up for more comprehensive, paid analysis. Part of the problem is that the assessment is done by an AI bot, and we know that AI is far from perfect.

HBM claims it uses AI and statistical modelling blended with human expertise and oversight to do its assessments. There is a thin but cursory layer of human involvement; fact-checking, source verification, style refinement etc. I think this is borne out by one missing key paragraph from HBM’s rewrite of my blog post. I had taken aim at Deloitte as an example of a large multinational company, that should know better, having been caught red-handed using unattributed AI that produced inaccurate, “hallucinated” results in a consulting report it prepared for the Newfoundland government. (“Deloitte’s AI Nightmare: Top Global Firm Caught Using AI-Fabricated Sources to Support its Policy Recommendations”). While HBM’s rewrite included almost all the key points in my post, there was zero reference to Deloitte. I am sure that “human expertise” decided that there was no point in gratuitously antagonizing an actual or potential client. Can I prove it? No, I guess its just another conspiracy theory.

I wonder if this blog post will be picked up and analyzed by Noah and if so, whether I would get a “Pass” this time. After all, it is “Fresh” and I have used lots of quotes from Noah. Having referred to the London News, I should get a 10/10 for Source Reliability (although I am not mainstream media, but neither is Noah). As for Plausibility what could be more plausible than an AI bot ripping off an author’s work through an unauthorized rewrite?  Would all that land me a “Pass” from Noah? I will probably never know.

© Hugh Stephens, 2026. All Rights Reserved.

Update: Noah picked up and summarized (using much more direct language this time) the blog post above and then (drumroll) gave me a “Pass”.

Korea’s AI Action Plan: Declaring War on Creators?

A young woman in a sparkling silver outfit poses next to a large robotic figure, adorned with South Korean symbols and colors, in a vibrant city background.

Image: Shutterstock

In the scramble to jump on the global AI bandwagon, Korea has floated a proposal that would supposedly remove “legal uncertainty” for AI developers who use copyrighted content to train AI platforms. Unfortunately, the proposed “solution” threatens to throw Korea’s globally renowned creative sector under the bus. Nor does it remove the uncertainty.

As part of President Lee Jae-Myung’s National Artificial Intelligence Strategy, its Presidential Council has put forward a 98 point “Action Plan”, a blueprint for implementation. There are many aspects to an AI strategy, but a key element is to ensure legal clarity with respect to the use of content for AI training, especially copyrighted content. The Action Plan purports to do this. Its Point 32 proposes the introduction of “explicit exceptions under the Copyright Act to allow copyrighted works to be used without legal uncertainty (emphasis added) in the processes of collecting and analyzing data available on the web”. In other words, introduction of a copyright exception for text and data mining (TDM), subject to certain conditions such as some form of remuneration, transparency, and opt-out features for rightsholders.

If the goal is to remove “legal uncertainty” regarding the use of copyrighted works, this proposal falls short of the mark. No exception can provide 100% certainty given the Berne Convention requirement that any exception meet the so-called “three step test”, meaning that an exception is permitted only in certain special cases, provided that it does not conflict with normal exploitation of the work and does not unreasonably prejudice the author’s legitimate interests. While there will always be a degree of uncertainty regarding exceptions, the good news is there is a ready alternative. The surest way to ensure legal certainty is to encourage licensing of content from rightsholders. The problem with the introduction of a TDM exception—or even the discussion of a possible TDM loophole–is that it diminishes the likelihood of reaching licensing solutions by reducing the pressure on AI developers to open their wallets to reach licensing deals.

It is even more bizarre that the tech industry is pressing for a TDM exception given that Korea is one of the few countries, alongside the United States, that has adopted a fair use provision in its Copyright law. This was done in 2011 as part of the implementation of the Korea-US Free Trade Agreement after heavy lobbying by the tech sector. Fair use allows courts to make case-by-case judgements as to whether a given use meets fair use criteria, thus potentially allowing reproduction of copyrighted material without advance permission from the rightsholder. If free use of copyrighted material for AI training can be shown to be “fair”, why is a TDM exception needed? Even in countries where fair use does not apply (which is most of the world), there is no convincing case or consensus on the need for a TDM exception; there is even less reason for one in a state that has already adopted fair use.

Point 32 of the Action Plan takes note of the existence of fair use in Korea, commenting that the Ministry of Culture, Sports and Tourism is preparing fair use guidelines. These are to provide interpretive guidance on the exemption provisions of the Copyright Act to enable companies to utilize copyrighted data “with greater confidence”. Despite this, the Action Plan claims these guidelines alone are unlikely to fully eliminate uncertainty and judicial risks. Voluntary licensing, however, would eliminate both.

It is well established that AI developers need vast amounts of data to improve the performance of their AI platforms. To date they have largely employed a “take first, ask later” policy. This has led to numerous lawsuits pitting rightsholders against the tech industry, mostly but not exclusively in the US. AI developers in the US have argued that what they are doing amounts to “fair use” because the final AI product is used for a different purpose from the original and thus does not compete with it. That is highly debatable, especially with image and music-based AI works. To date, the results from the US courts have been mixed.

The legality of the tech industry’s unauthorized use of copyrighted content is an issue that a number of countries, in Asia and around the world, are looking at. Various solutions have been proposed to eliminate the uncertainties that arise from leaving the decision to the courts. Among these are TDM exceptions which have been introduced, albeit with strict limitations, in the UK, the EU and Japan. In the UK for example, use is limited to non-commercial purposes. In the EU, it must be accompanied by transparency requirements and opt-out provisions for authors. In Japan, if the unauthorized user derives commercial benefit from the content, the safe harbour does not apply. Australia has explicitly ruled out introducing a TDM exception in order to protect its creative sector, while many others (eg. India, Canada) have no TDM provision in their copyright law. As noted above, the clearest way to remove any uncertainty about the legality of using copyrighted works is to incentivize and recognize voluntary licensing as the solution. This ensures that rightsholders receive appropriate compensation for the work they have put into creating content, while guaranteeing legal certainty for licensees.

The Korean strategy paper argues that AI companies are required to obtain individual consent from each copyright holder “leading to significant costs and time burdens in securing high-quality training data”. But large amounts of high-quality content can be accessed through voluntary licensing agreements with major content creation companies such as studios, publishers, broadcasters, music labels and so on. As for individual authors and artists, one possibility is to look at the model currently used for licensing print and music content through Collective Rights Management Organizations as a supplement to voluntary licenses signed with major rightsholders.

In addition to being instructed to prepare the necessary amendment to the Copyright Law for presentation to the National Assembly by Q2 of this year, the Culture, Sports and Tourism Ministry, in cooperation with the Ministry of Science and Technology, is to “promulgate standard contract templates for the licensing and transfer of copyrights for AI training”. This is the kind of heavy-handed market intervention that is guaranteed to stifle voluntary licensing. Not only that, it amounts to expropriating the rights of Korean creators to manage their works.

The Action Plan gives a nod to the importance of compensating rightsholders and claims it wants to establish a system that respects the rights of creators. However, given the size and importance of Korea’s cultural industries, from film to K-Pop to literature, it is surprising there isn’t greater recognition of what an important strategic and economic asset this sector represents for Korea. Although the Strategy acknowledges that content industries should be able to share the benefits of growth in the AI industry, the proposed solution is unbalanced and biased toward clearing any so-called “obstacles” to unimpeded use of content. As a result,  just days after the extremely brief (20 day) consultation period on the Strategy had closed, in mid-January sixteen creator and rightsholder groups issued a strong statement condemning the Action Plan, labelling it “an attempt to fundamentally undermine copyright as a private property right”.

While paying lip service to creator’s rights, the Plan does not address how creators can enforce these rights (other than through the creation of opt-out protocols, which stands the normal copyright procedure of seeking permission prior to usage on its head). The Strategy seems to lead to what has been described by many as a “use now, pay later” system, with little information on how payment would be calculated or implemented. On the other hand, prior, voluntary licensing of content for AI training is a solution that would respect the rights of Korea’s creators while providing the welcome revenue sharing and income stream for which the Strategy advocates. Strong content industries benefit AI development in Korea by encouraging continued creation of the valuable Korean language content so necessary to refine and improve AI models. Conversely, providing the tech industry with an escape hatch to avoid licensing by instituting a TDM exception is the surest way to kill a licensing market for AI content. It will only continue the legal uncertainty that the Presidential Council seems to feel is hindering AI development in Korea.

The one-sided formulation of the Strategy to date has provoked an inevitable negative reaction from Korea’s cultural industries. This is not surprising since the strategy of the tech industry, in Korea and elsewhere, is to avoid dealing with ministries directly responsible for culture and copyright and instead lobby industry, technology and science ministries to bring pressure for changes to copyright law. This adversarial stance is unfortunate as the content and tech industries need and can help each other. The Strategy needs to be amended so rather than throwing the cultural and copyright industries under the bus in the name of facilitating AI development, Korea provides the framework for a mutually beneficial and legally certain relationship. This is best done by upholding longstanding copyright principles and encouraging the growth of a voluntary licensing market for content used in AI training.

© Hugh Stephens, 2026. All Rights Reserved

No Surprise:  Ontario Court Asserts Jurisdiction in Canadian Media Lawsuit Against OpenAI

A judge sitting at a bench in a courtroom, wearing a black robe with a red collar, Canadian flags in the background.

Image: Shutterstock

The Ontario Superior Court has ruled it has jurisdiction to hear the case against ChatGPT owner OpenAI brought by a consortium of Canadian media companies led by the Toronto Star. The media enterprises, who include the Globe and Mail, PostMedia, CBC/Radio Canada, Canadian Press and Metroland Media Group, are suing the US company for copyright infringement, circumvention of technological protection measures (TPMs), breach of contract, and unjust enrichment as a result of OpenAI’s scraping of their websites to obtain content to train its AI algorithm. The allegations also cover OpenAI’s use of Retrieval Augmented Generation (RAG) to produce contemporary search results from paywall-protected content that augment ChatGPT’s AI-generated responses. When the suit was brought in November 2024, OpenAI had challenged the Ontario court’s jurisdiction on the basis, among others, that it had no physical presence in Canada. As pointed out by this legal blog, a court may presumptively assume jurisdiction over a dispute where one of five factors is present:

  • The defendant is domiciled or resident in the province.
  • The defendant carries on business in the province.
  • The tort was committed in the province.
  • A contract connected with the dispute was made in the province.
  • Property related to the asserted claims is located in the province.

The court found that OpenAI carries on business in Ontario notwithstanding its lack of a physical presence and was a party to contracts in Ontario as a result of tacitly accepting the terms of service regarding access to the media companies websites when it scraped them.

OpenAI wanted the venue of the litigation changed to the United States to take advantage of developments in US law regarding unauthorized reproduction of copyright protected content for use as AI training inputs. To date, while many cases are still ongoing, US courts have tended to support a fair use argument by AI developers allowing them to access copyrighted content without permission on the basis that the end use is “transformational”, resulting in a new product that does not compete with the original work. In Canada, the fair use doctrine does not apply and exceptions to copyright protection are either explicitly laid out in the law (e.g. for law enforcement or archival preservation purposes) or are governed by the fair dealing provisions of the Copyright Act. These require that an unauthorized use fall into one of eight categories (research, private study, education, parody, satire, criticism, review and news reporting) that is in turn subject to various court-interpreted criteria such as amount of the work copied, the purpose of the copying, market impact etc. AI developers have been lobbying for the introduction of a text and data mining (TDM) exception into Canadian copyright law, but so far this has been successfully resisted by Canada’s creative community. All this to say that it is more difficult for AI companies to avoid liability for unauthorized use of copyright protected material in Canada than in the US, thus the importance of whether the Ontario court has jurisdiction.

Back in September, on the basis of previous Canadian court rulings where courts ranging from provincial courts to the Supreme Court of Canada asserted jurisdiction over large digital US companies operating virtually in Canada, such as Google (who challenged Canadian legal authority over them on the basis of lack of a physical presence), I predicted (guessed would be a more accurate term) that the Ontario court would be loath to surrender jurisdiction simply because the company was headquartered in the US. The earlier cases were for defamation rather than copyright infringement, and my “prediction” was based more on a hunch than legal analysis, but I am satisfied that I called it right. OpenAI has no compunction about selling services and collecting revenues in Canada and presumably (I hope) pays taxes here, although it is not subject to the Digital Services Tax (DST) that the Carney government threw overboard in a vain attempt to placate Donald Trump. Recall that Trump had threatened to terminate trade talks if Canada proceeded to implement the long-planned DST, so Canada blinked. Trade talks resumed until Trump found another excuse to end the talks, in this case the anti-tariff ads on US television placed and paid for by the Ontario government to which he took offence. But there is no doubt that OpenAI does business here; it just doesn’t want to be subject to Canadian law and Canadian courts. It can’t have it both ways.

While this is a victory for Canadian sovereignty, just because the Ontario Superior Court has confirmed its jurisdiction, this doesn’t mean that once the substantive proceedings begin copyright infringement will be found. Lawyer Barry Sookman, in an analytical  blog post on this topic, has noted that in determining whether the alleged copyright infringements occurred in Canada, “the court relied heavily on the Supreme Court decision in SOCAN for the proposition that the territorial jurisdiction of the CCA (Canadian Copyright Act) extended to where Canada is the country of transmission or reception.” However, “SOCAN applied the real and substantial connection test to the communication to the public right” whereas the alleged copying involved the right of reproduction.

Sookman continues;

“…that test does not apply to the reproduction right. (The Federal Court has) held that the only relevant factor is the location in which copies of a work are fixed into some material form. The locations where source copies reside or acts of copying onto servers located outside of Canada, are not infringements” (according to the cases cited).

Inside baseball information but important when it comes to determining copyright infringement. On the other hand, it seems to me that the infringement involved not just, potentially, the reproduction right (the copying) but also the communication right, because OpenAI, through Microsoft, provided RAG content to users in Canada and elsewhere purloined from behind the paywalls of the media companies. So, we will have to see. Lots of fodder for IP lawyers.

In the meantime, deep-pocketed OpenAI will appeal the jurisdictional ruling—and will likely lose again. The appeal will buy time for it to negotiate licensing deals with the complainants. This is increasingly the model in the US as AI developers, including OpenAI, are reaching licensing agreements with content owners, particularly media organizations. To date, OpenAI has signed licensing deals with the Associated Press, the Atlantic, Financial Times, News Corp, Vox Media, Business Insider, People, and Better Homes & Gardens, among others, while being sued (in addition to Toronto Star et al), by the New York Times and a collection of daily newspapers consisting of the New York Daily News, the Chicago Tribune, the Orlando Sentinel, the Sun Sentinel of Florida, San Jose Mercury News, The Denver Post, the Orange County Register and the St. Paul Pioneer Press. Even META, that arch-opponent of paying for media content–which it claims adds no value to its users– has struck a media deal with news publishers, including USA Today, People, CNN, Fox News, The Daily Caller, Washington Examiner and Le Monde. (One wonders if this will cause it to rethink its position of thumbing its nose at Canada’s Online News Act, where it “complied” with the legislation by blocking all Canadian news links).

In another content area, OpenAI and Disney have just agreed on a three-year output deal, allowing it to use Disney characters (subject to certain limitations) in its AI creations. (Meanwhile Disney is suing Google for using its characters in Google’s AI offering). Open AI is currently facing 20 lawsuits, including the Toronto Star case, and needs to resolve these legal challenges before its expected public offering next year or 2027. The spectre of impending lawsuits will inevitably lower the IPO price.

Most if not all of these lawsuits are going to end in settlements via voluntary licensing agreements, but that will only happen if OpenAI thinks the alternative (losing a major lawsuit) is a worse outcome. If it can wriggle out from the Toronto Star case by invoking some specious argument related to jurisdiction, it will. If it can’t it, will eventually open its chequebook and provide the Canadian media outlets some compensation for the valuable curated content it has hijacked. Canadian courts need to stay the course to help ensure that this happens.

© Hugh Stephens, 2025. All Rights Reserved.