When the End Does Not Justify the Means, Anthropic’s $1.5 Billion Lesson

“Fair Use” Does Not Justify Piracy

A hand-written note on a white paper that reads 'END ≠ MEANS'.

Image: Author

The stunning announcement on September 5 that AI company Anthropic had agreed to a USD$1.5 billion out-of-court settlement to settle a class-action lawsuit brought by a group of authors was ground breaking in terms of its size, and goes to disprove the old adage that “the end justifies the means”. It is still not clear if the “end” (i.e. using copyrighted content without authorization to train AI algorithms) is legal, although preliminary indications are that at least in the US this may be the case. However, even if what Anthropic and other AI companies have been doing is ultimately determined to be fair use under US law—which is by no means certain—downloading and storing pirated content is clearly not legal, even if it is to be used for a fair use purpose. In other words, the piracy stands alone and must be judged as such, separate from whatever ultimate use to which the pirated content may be put.

Ironically, in the end, Anthropic did not even use much of the pirated content it had collected for training its platform, Claude. It seems to have had second thoughts about using content from online pirate libraries such as LibGen (Library Genesis) and PiLiMi (Pirate Library Mirror) and instead went out and purchased single physical copies of many works, disassembling and then digitizing them page by page for its Central Library, after which it destroyed the hard copies. Why go to all this trouble? Why not just access a legal online library? That’s because when you access a digital work, you don’t actually purchase it. You purchase a licence to use it, and that licence comes with conditions, such as likely prohibiting use for AI training. Anthropic would have been exposing itself to additional legal risk by violating the terms of the licence, so instead of negotiating a training licence, they took the easy way out by downloading content from pirate sites LibGen and PiLiMi. Later, having second thoughts, they purchased physical copies of the works they wanted to ingest and then scanned them. But it was too late. The piracy had already occurred.

When the decision in the Bartz v Anthropic case was released this summer, I commented that the findings were a mixed bag for AI developers. A very expensive mixed bag, it turns out. In the Anthropic case, there were clearly some interim “wins” for the AI industry. Anthropic’s unauthorized use of the works of the plaintiffs (authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson, who filed a class action suit) was ruled by the judge (William Alsup) to be “exceedingly tranformative” thus tipping the scales to qualify as a fair use. In addition, he ruled that Anthropic’s unauthorized digitization of the purchased books to also be fair and not infringing. However, it was the downloading and storing of the pirated works that got Anthropic into hot water. Even though the intended use of the pirated works was to train Claude, a so-called transformative fair use, this did not excuse the piracy. While Alsup did not specifically rule that use of pirated materials invalidates a fair use determination (i.e. he ruled that the piracy and the AI training were separate acts), his ruling exposes a weak flank for the AI companies. For example, the US Copyright Office has stated that the knowing use of pirated or illegally accessed works as training data weighs against a fair-use defence. In short, the end does not justify the means.

The piracy finding was significant because Judge Alsup decreed that this element of the case would be sent to a jury to determine the extent of damages. (In Canada and the UK, judges rather than juries normally play this role). Given that under US law statutory damages start at $750 for each work infringed but can go up to $150,000 per work for willful infringement, Anthropic could have been on the hook for tens of billions of dollars in damages for the almost 500,000 works at issue. (Over 7 million works were inventoried by the pirate websites and downloaded by Anthropic but the limitations on who qualifies for the class action reduced the number of actionable works to just 7 percent of the total). As deep as its pockets are (Anthropic is backed by Amazon), if a jury awarded damages toward the higher end of the scale, the company could have been bankrupted.

Thus, Anthropic had lots of incentive to settle (including keeping the fair use findings unchallenged). As it stands, the $1.5 billion payout, while large in total, amounts only to about $3000 per infringed work, not the minimum but not really financially significant for the plaintiffs. This amount will probably have to be split between authors and publishers, with some of the funds covering costs, so no authors are going to be buying a new house on the proceeds. The real beneficiaries will be the law firms that represented them. The messy process of deciding who gets what that has led Judge Alsup to suspend the proposed settlement in its current form and require greater clarity as to how the payouts will be managed. The number of works eligible for payment is limited by the fact that to qualify they have to meet three criteria;

1) they were downloaded by Anthropic from LibGen or PiLiMi in August 2022

2) they have an ISBN or ASIN (Amazon Standard Identification Number) and, importantly,

 3) they were registered with the US Copyright Office (USCO) within five years of publication, and prior to either June 2021 or July 2022, (depending on the library at issue).

Any other works do not qualify. Registration with the USCO is not a requirement for copyright protection but in a peculiarity of US law, without registration a copyright holder cannot bring legal action in the US.

While the settlement has been welcomed in copyright circles, and could set a standard for settlement in other pending cases where pirated material has been downloaded for AI training by companies such as META and OpenAI, it doesn’t settle the overriding question of whether the unauthorized use of non-pirated materials for AI training is legal. With the settlement, the Anthropic case is closed, including with respect to the fair use findings. There will be no appeal, another benefit for Anthropic. However, there are still a number of other cases working their way through the US courts, so the question of whether unauthorized use of copyrighted content for AI training constitutes fair use is far from settled.

The Anthropic settlement, especially its size, has caught people’s attention. It may result in AI developers deciding it is better to resort to licensing solutions to access content rather than risking the uncertain results of litigation. On the other hand, payments like this could be one-offs, a speed bump for deep pocketed AI companies who will continue to trample on the rights of creators if they can get away with it. In the Anthropic case, while the company must destroy its pirated database, it is not required to “unlearn” the pirated content that it ingested. Moreover, even if this case leads to more payments to authors, which would be welcome, there are still many copyright-related conundra to be resolved. It should not be necessary to have to constantly resort to litigation to assert creator’s rights given that, as the Anthropic case shows, only a very limited number of rightsholders benefit from specific cases. Broad licensing solutions are required. This would also help address the problem of AI platforms producing outputs that bear close resemblance to, or compete with, the content on which they have been trained.

While Bartz v Anthropic is a decision that applies only to the US, and only to this one very specific circumstance, it will be studied closely elsewhere in countries that do not follow the unpredictable US process of determining fair use, for example in fair dealing countries like the UK, Canada, Australia, New Zealand and elsewhere, and in EU countries. In Canada, the unauthorized use of copyrighted works for training commercial AI models is a live issue. With the possible exception of research, unauthorized use such as that undertaken by Anthropic is unlikely to fall into any of the fair dealing categories (in Canada, they are education, research, private study, criticism, review, news reporting, parody and satire) nor is there a Text and Data Mining (TDM) exception in Canadian law. As Canada and other countries come to grips with the copyright/AI training dilemma, the principle of how content is accessed will surely be an important principle. Just as fair use (if indeed AI training is determined to be fair use) does not justify piracy in the US, licit access is required in Canada to exercise fair dealing user rights, including where TPM’s (technological protection measures, aka digital locks) are in place to protect that content.

Judge Alsup’s decision upholds the important principle that the end (if legal) does not justify the means (if illegal). This is a key takeaway from the Anthropic case, imperfect as the outcomes of that case were. Meanwhile the legal process of determining how and on what terms AI developers should have access to copyrighted content to train their algorithms continues.

© Hugh Stephens, 2025. All Rights Reserved.

Sci-Hub Blocked in India: Has the Last Domino Fallen for this Notorious Academic Pirate Site?

Sci-Hub Undermines both Paywall and Open Access Models

A stylized black bird holding a red key in its beak, against a starry background.

Image: Logopedia (CC-BY-SA Licence)

As reported by TorrentFreak, Sci-Hub, the notorious pirate site for scientific and academic journals, has been blocked in India by court order after a 5 year court process. Obstinacy and failure to appear or offer a defence on the part of Sci-Hub’s operator, Kazakhstan-based Alexandra Elbakyan, appear to have been factors in finally deciding the case. Although Sci-Hub has been blocked or banned in a number of countries, India was a holdout. Sci-Hub had been effective in mobilizing the “exploited Global South/knowledge should be free” argument to delay proceedings. Sci-Hub has accused the academic publishers, in this case Elsevier, Wiley and the American Chemical Society, of monopolizing knowledge and sealing it off behind paywalls that block access. In India, where the cost of western IP is a political issue (often played out in the patent domain in the area of pharmaceuticals, where India’s widespread production of generic drugs is controversial among western pharmaceutical companies), that argument has political legs. But in the end, it did not stop the court from putting an end to the delays and reaching a clear decision. Perhaps the last major domino has fallen.

Sci-Hub dates back to 2011 when it was founded by Elbakyan explicitly to do an end run on publishers of academic journals. Its motto, highlighted on its site, is “breaking academic paywalls since 2011”. What are Elbakyan’s motivations? They appears to be altruistic, i.e. making knowledge “free” as proclaimed on Sci-Hub’s website, although it is easy to be altruistic with someone else’s property. There doesn’t seem to be much of a business model behind Sci-Hub, with donations being the prime source of funding, a pipeline made more difficult when PayPal and Twitter agreed to block the platform. Justifying Sci-Hub’s piracy by arguing that it frees up academic and scientific knowledge is fed by academics who are unhappy with the academic publishing model. For example, this India-based academic argues that, as a right, researchers should have “complete, paywall-free access to every paper published everywhere.” In other words, all free, all the time. Publishers would argue that widespread access is already provided through university libraries to those who need it, although not every institution has access to all the key journals. However, it is always possible to contact the author directly and ask for a copy. But that is a hassle. It is far easier to go to Sci-Hub and download pirated articles.

The issue of access to academic and scientific journals is complex. Academic and scientific publication is essential for several reasons. A key one is to advance and share knowledge through accurate, peer-reviewed, properly edited, credible publications. Part of that process involves career advancement and development in academe; scholars and academics are often evaluated by their institutions based on the quality and sometimes quantity of adjudicated publications they author or co-author. Research without sharing the knowledge gained is essentially pointless. The issue is how best and most credibly to disseminate that knowledge.

There are lots of dodgy so-called academic journals that will publish anything, accurate or peer reviewed or not, simply on a pay-to-play basis. Frankly, no self-respecting academic would publish in such journals; it would undermine their reputation and credibility. Therefore, publication needs to be in a recognized and respected journal where there is a high academic barrier to entry. Those journals are generally published by a few major publishing houses, such as Elsevier, Wiley, Springer and Sage. Elsevier publishes almost 3000 journals, concentrating on science and health, Springer about the same but including social sciences and humanities, Wiley 1500 and Sage about 1000, with a focus on social sciences. Other institutions such as the Royal Society, American Chemical Society and various university presses also publish peer reviewed journals.

The normal model of a publisher paying an author a royalty for the right to publish a work is stood on its head in the academic world. In a sense, the publication is doing the author a favour by agreeing to publish, assuming the work meets editorial standards. At least, that is the way the publishers see it. The peer review process is unpaid work undertaken by other academics as part of their research commitments (although many academics are unhappy with this unpaid labour) while the publisher absorbs the costs of editing, publishing, archiving and distributing. (Sometime editors are senior academics who are not separately compensated for these services). In the digital world distribution and archiving is far less costly than in the pre-digital days of printing and distributing physical copies of journals, mostly to university libraries. Today, authors are not expected to pay for publication but also normally receive no compensation for granting the right to publish. The publisher recoups its costs, (and turns a profit) by charging for access to the journals. Most universities purchase bulk access for their students, but of course they cannot subscribe to every single journal. One time access is normally available through payment of anything from $20 or $30 to much more to get the paywall key.

This business model has been modified over time with the growth of the Open Access model and other means to provide content without going through a paywall. Under Open Access (OA) the article is published in the journal but not put behind the publisher’s paywall, i.e. it is disseminated without charge. Another is the pre-print process whereby authors can post a pre-proof copy of their paper on a preprint server. Preprints are copies of papers that can be posted prior to peer review and are freely available, typically with Creative Commons or similar licenses. Pre-prints were originally opposed by the publishers, but the arrival of open access mandates led to a sudden shift where publishers now risk having their journals abandoned by authors if they refuse to accept articles with posted preprints. Many new preprint servers, some “altruistically” funded by publishers, others funded by Foundations, have sprung up in response.

As for post-proof Open Access copies published in the journals, there are still costs to be covered so if the user is not going to pay for access, who covers the costs? Why, the author of course! (The cost would possibly be covered from whatever grant the author was using to research the topic, but often there is insufficient funding to cover these costs. Sometimes there is assistance provided by the author’s institution). Cost can vary significantly but are typically in the low thousands of dollars although it can cost over $12,000 to publish an Open Access article in Nature. The point is there are legitimate costs to be covered and somebody needs to pay, unless the author has uploaded a pre-print version. But that is not how Sci-Hub works. Peer-reviewed, published works are hijacked and placed in Sci-Hub’s repository through various means such as illicit sharing of passwords, leaked credentials from students or faculty who have legitimate access through their institutional libraries, or apparently through more nefarious means such as phishing. Police in the UK reported that 42 UK universities had been “hacked” by Sci-Hub by tricking students into revealing their log in credentials.

For the academic and scientific publishers, Sci-Hub is a growing problem and so they have taken action, as in the recent case of India. In 2015 Sci-Hub was sued in the US by both Elsevier and the American Chemical Society (ACS). Elsevier won a $15 million judgement; ACS was awarded $4.8 million. Sci-Hub was not represented in court and despite the judgements, continues to operate despite losing its domain. Not surprisingly, the damages were not paid. In the UK, the publishers took a different approach, successfully obtaining site blocking orders, an approach also taken in a number of EU member states, including France, Germany and  Sweden. But India was the big test, given significant support in India for Sci-Hub and the sensitive “decolonialization of knowledge” argument. An “inconvenient truth” regarding this argument, however, is that, according to a 2022 study published in Nature (sorry, paywalled) the country with the second largest number of users of Sci-Hub, after China, is the United States, followed by France, Brazil, India, Indonesia and Germany. The widespread use of VPNs also hides where many users reside.

Why do students in developed countries use Sci-Hub? It is sometimes–perhaps often–quicker and easier than going through an institutional library, where they may have to be physically present or have updated and valid credentials. Piracy is often the course of least resistance for the user, although it comes with costs and risks, such as malware, whether it is pirating academic journals or streaming content. It is for this reason, among others, that many academic institutions warn their students against using Sci-Hub. Sci-Hub has been accused of obtaining and potentially misusing all sorts of other personal information, such as email addresses and social insurance numbers, obtained through using “borrowed” library access credentials.

It has also been accused of undermining the legal Open Access movement. Counter-intuitively, I think Sci-Hub’s role initially likely provided impetus to expand Open Access. As an academic friend put it, while policy makers were trying to figure out how to share the keys to the library, Sci-Hub had broken in through the window and was giving away the books in the back alley. Finding legal ways to minimize the negative impacts of paywalls helped energize Open Access, but now that it is well established, resorting to Sci-Hub for documents weakens the case for Open Access options. As this university library notes;

“The OA movement is a way to transform the research dissemination in a healthy and safe way for the long term without putting users and institutions at risk. If as researchers we are unsatisfied with the current limitations of academic publishing, then the solution is to push for change in how we disseminate our work that don’t necessitate responses like Sci-Hub.”

Now the “Indian domino” has fallen, with a blocking order issued by the Delhi High Court, will this stop Sci-Hub? Not likely, which means there will continue to be a need for regular pushback by the publishers, along with continued work on improving legal open access. Otherwise, the established and essential system of academic and scientific publishing will be undermined by Sci-Hub’s piracy. No-one benefits from that outcome.

While the motivation for academics to get paid for the content they produce is generally not as intense as with other authors (by “authors”, I include artists, photographers, musicians, etc)–because most academics and scientists already get compensated by their employing institutions for the research and writing that goes into an academic publication–they still need compensation. That compensation comes in the form of being able to publish in recognized, academically respected journals. As with any discussion of piracy, the reality is that when free-riding begins to overtake the legitimate dissemination of content, the content creation and distribution model is undermined. By taking various legal means to disable Sci-Hub, the publishers are of course protecting their business model–but they are also protecting the future production and distribution of quality, credible and verifiable knowledge, including content distributed through Open Access models.

© Hugh Stephens, 2025. All Rights Reserved.

“Just the facts, Ma’am”: Facts and Copyright

Historical scene depicting a gathering in a church, showcasing a group of people, many in period costumes, attentively listening as a military officer reads a document. The setting features wooden benches and large windows, suggesting an atmosphere of important decision-making.

Painting by C.W. Jeffreys, Public Domain via Wikimedia Commons

Resurrecting this 1950s-era Joe Friday (played by Jack Webb) quote from the TV series Dragnet may date me but it is a classic. The request seems so simple.  Just the facts, and nothing but the facts. Facts are integral to interpretation of copyright law because “the facts” cannot be copyright protected. As almost everyone knows–but it bears repeating–copyright does not protect ideas or facts, only original expressions of ideas (and expressions that may be based on facts). Thus, Van Gogh’s painting of a vase of flowers could be copyright protected, but anyone can paint their own version of flowers in a vase. Winston Churchill wrote several copyrighted volumes about his interpretation of what happened in the Second World War, but the circumstances and facts of that conflict are open to anyone to write about. I was thinking about history and facts this summer when my wife and I visited one of Canada’s National Historic Sites, Grand Pré in Nova Scotia, site of the “Acadian Expulsion”. (known in French as Le Grand Dérangement).

The “facts” are probably generally if imprecisely known to many, especially in Canada and the US. The reasons why it happened, also part of the narrative, are less definitive. Here are the essential facts. In 1755, the British governor of Nova Scotia, Charles Lawrence, ordered the removal of some 6,000 Acadians settled in the areas of the Fundy marshes, close to present day Wolfville, NS. It was a brutal event; the settlers’ houses were torched, livestock killed, families often separated. The Acadians were widely scattered, with many being settled in the New England colonies and later in Britain and France. They were not, contrary to popular belief, deported to Louisiana, but many of those who eventually ended up in France were recruited by the Kingdom of Spain to settle in Louisiana, at that time under Spanish control. Spain wanted reliable Catholic settlers for colonization purposes. This is the foundation of the Cajun (Cadian) people of Louisiana. Some members of the Acadian diaspora eventually found their way back to what are now Canada’s Maritime provinces. Their return was permitted after 1764 following the defeat of France in North America and the fall of Québec. They settled largely on the north shore of what is now New Brunswick since New England Planters from the Thirteen Colonies had taken up much of their original lands. This part of New Brunswick has become an Acadian stronghold, and a strong sense of Acadian nationality and pride has developed over the years. Today it is common to see Acadian flags flying, and the Acadian star adorning houses in areas where Acadians reside.

The resurgence of Acadian awareness and pride can, ironically, be traced to a 19th century American poet, Henry Wadsworth Longfellow, who in 1847 published the opus “Evangeline”. This was a poetic work of fiction based loosely on the Acadian Expulsion (or Upheaval, as it is sometimes called), centred around a deported Acadian heroine, Evangeline, who for years engaged in a fruitless search for her deported betrothed, Gabriel, only to find him in Pennsylvania on his deathbed many years later, afflicted by the plague. As a work of fiction, Longfellow’s work was of course copyrighted. The poem was his expression of what had happened to the Acadians. Given its date of publication, its copyright protection has long since lapsed, and it has been in the public domain for many years. But behind Longfellow’s epic poem, and a few other early works written about the Acadians, (offering various interpretations and perspectives), there are “the facts” explaining what actually happened, and why. But what did actually happen? What are the bare facts?

You would think it would be relatively straightforward to recount the factual story but recall this all happened 250 years ago. The most public display of “the facts” is at the Grand Pré exhibition hall, run by Parks Canada. And this is where it becomes somewhat difficult to get to the unvarnished truth, “just the facts”.

The site is a place of pilgrimage for Acadians, part of their national story. People come to visit the Evangeline Chapel built in 1924 when the site was first established as a memorial. But one must not forget the Mic’maw people who populated the area before the Acadians arrived, and after they left. They are still around, still active and very vocal. And then there are the New England settlers who were brought in to take over the lands of the Acadians and establish a “loyal”, non-Catholic presence. Many of the current residents of rural areas of Nova Scotia are direct descendants of what is known as the New England Plantation. In addition, there is the interpretation of the positions of the then British and French governments, and the role they played, and the responsibility they bore. The display panels at the Grand Pré site, in English and French, do a good job of trying to manage the interpretation of the facts in a way that meets contemporary needs, or at least will offend the least possible number of people!

While we were at Grand Pré there was an historical re-enactment. An actor dressed up as an Acadian told us about the shock of the expulsion, being separated from her children, and so on  but then in an aside made it very plain that the New England settlers who arrived a year or two after the Acadians had been forcibly removed were not to blame for the expulsion. In fact, when it came to explaining the reasons for the expulsion, there were several different options to choose from in the interactive displays. It was clear that the Acadians were expelled for refusing to take the oath of allegiance to Britain, which was the governing power in Nova Scotia since the area had been ceded to them by the French in 1713 in the Treaty of Utrecht. That treaty still left the French in possession of Québec and what today is called Cape Breton Island, where they established the fortress of Louisbourg. Those French fortresses were perceived to pose a threat to the English settlements further south, including Rhode Island, Massachusetts, and Connecticut. Between the French and the English settlements lay Acadia, governed by Britain but populated largely by French speaking settlers whose loyalty was, at best, dubious. Some had actively aided French expeditions sent south to penetrate the area, although others had remained neutral.  Was it unreasonable for the British to be concerned about what in later years would be called a “Fifth Column” in their midst? Was it unreasonable to pressure the Acadians to pledge loyalty to Britain, which some did but most resisted? I guess it depends on your point of view and the judgement of history.

But what about the Acadians? Why didn’t they accept their fate and realize they had been abandoned by France? Perhaps they had hopes that the outcome of Utrecht would be eventually reversed (since Britain and France seemed to go to war with each other every 10 years or so). Or was it because they were concerned their Catholic faith would be in jeopardy, given that Britain would not guarantee this? Another theory has it that the Acadians, who had a close but not completely satisfactory relationship with the powerful Mik’maw, were afraid that they would be subject to attack if the Mik’maw thought the Acadians were accommodating the British. There were no doubt many reasons for their refusal.

The Mik’maw also play a role. They were originally dispossessed by the Acadians, but had been converted to Catholicism, and thus there was a certain bond between the two groups. Apparently intermarriage was not infrequent. The Mik’maw fiercely resisted the British for a number of years to the point that Governor Edward Cornwallis offered a bounty for Mik’maw scalps. This has put Cornwallis, who founded Halifax and who is considered the father of Nova Scotia, in bad odour in today’s climate of reconciliation. (Edward was the uncle of Charles Cornwallis who surrendered British forces to George Washington in 1783). Cornwallis issued his proclamation after the Mik’maw had attacked and killed settlers. Today that would be called defending your land. Who is right? What are the “facts”?

You won’t find “just the facts” in Longfellow’s poem, or even definitively in the Parks Canada panels at Grand Pré describing the events of the day. As I noted, great effort has been made to present all views and, presumably, to let visitors decide for themselves as to what led to the Expulsion. This interpretation of the “facts” by Parks Canada historians, who have to answer to all segments of public opinion, could certainly be copyrighted as one expression of what happened. I am not aware that copyright has been asserted although it is possible that the interpretive panels and content are under Crown copyright. That would be appropriate as the “balanced and blended” interpretation of the reasons for what happened to the Acadians between 1755-1764, and why, is still only one version. Joe Friday didn’t realize what a complicated question he was asking when he uttered that famous expression, “Just the facts, ma’am”.

© Hugh Stephens 2025. All Rights Reserved.

Paywalls and News Publishing: There Should be No Ambiguities

An illustration of a laptop displaying a padlock icon, surrounded by images, documents, and a credit card, symbolizing online content security and paywalls.

Image: Shutterstock.com

In trying to search for the right analogy to explain the obvious,–i.e. if a news organization puts up a paywall, it is illegal to bypass or hack it to get at the content and, if you’ve paid for access, that doesn’t mean you can share it with all your friends–I have come up with the old-fashioned movie ticket as the comparator. When there is a show at the local cinema you want to see, there is normally only one legal way to watch it. Buy a ticket. And once you have used your ticket to watch the show, you don’t get to give it to someone else to see the next showing, and the next, and the next. Perhaps this is all too obvious, yet people seem to have great difficulty in getting their heads around the simple fact of what a news site paywall is, and why it is there. It’s really pretty simple. It is there to provide access to content (like a ticket) but also to limit access (for those who don’t have a ticket). It is the basic element of the online business model, for news access certainly but also for other forms of content, such as streaming entertainment.

I don’t understand why some people think they should have free access to content that others pay for. There is always someone who wants to beat the system and then thinks up some excuse to justify their actions. I freely admit that paywalls can be annoying, especially if you are surfing the web and come across a random article that you want to read in the Moose Jaw Monitor or the Peoria Progress. It’s almost always all or nothing. No free samples; just an annual subscription, although likely discounted for the first year. But if all you want is that one article there never seems to be a “pay by the item” option. It’s all or nothing. I have come to the conclusion that in most such cases I can either live without it or sometimes, if I search hard enough, I can find the same thing elsewhere, unpaywalled. What I don’t do is hack it.

A recent discussion of the ethics of paywalls examined the thorny question of whether it was ok to cheat—but just once in a while and under certain circumstances. The unconvincing conclusion: it depends. However, if you believe in the value of curated news–and most people do although they are remarkably resistant to paying for it (a recent Pew Research Center survey indicated that 83% of Americans had not paid for news in the past year; in Canada the numbers are comparable with 15% saying they were willing to pay, up from 11% a year earlier)—then it is only logical that the more free rides people take, the less responsible news coverage is going to be produced. You are eating your own seed grain. As someone put it, the garbage is free and the quality stuff has to be paid for.

The most recent example of success in combatting paywall-busters was the recent announcement that the News Media Alliance, the trade association in the US for major news publishers, had secured the removal of a website that existed to enable users to bypass paywalls, known as 12ft.io. It was self-described as a 12 ft. ladder to get over a 10 ft. wall. According to an article in the Verge, the site also allowed users to view webpages without ads, trackers, or pop-ups by disguising a user’s browser as a web crawler, giving them unfettered access to a webpage’s contents. Now it is out of business. The Alliance doesn’t say how it achieved this feat but does say that it will “continue to take similar actions against other purveyors of unauthorized paywall bypassing technologies.”

While certain elements of the public (academics? other journalists? researchers?) seem to think they can lay claim to justifications to bypass paywalls, an act that is illegal in both the US and Canada if a “technological protection measure” (TPM), aka a digital lock, is circumvented, the most egregious example of paywall-busting is the Government of Canada itself, the same government that is responsible for the Copyright Act. As I have noted a couple of times, (“The integrity of journalism paywalls is under threat. The Government of Canada should settle the Blacklock’s case”; “Does the Trudeau government really support Canadian media? Saying One Thing but Doing Another”), if the Government of Canada, which spends millions on media and communications, cannot be bothered to obtain a licence to access paywall-protected material, how can they expect ordinary citizens to respect paywalls.

 The issue is the $148 individual subscription taken out by an employee of Parks Canada back in 2013, a subscription that the Government of Canada through the Attorney-General argues should allow it to reproduce and pass around individual articles within a large government department without obtaining an institutional subscription. This all hinges on a complicated case, originally brought by Blacklock’s Reporter, an online investigative journalism enterprise, but then pursued by the Crown after Blacklock’s withdrew, in which the A-G argued that it was entitled to access the paywalled content on the basis of fair dealing. While it is illegal to circumvent a TPM/digital lock for the purpose of accessing TPM-protected content, the issue was whether a password constitutes a TPM. Logically, I think most people would assume that it does, but the judge ruled that evidence had not been presented to conclude that was the case.

One line of argument is that a password is not a digital lock; rather it is a digital key to a digital lock (TPM). Therefore if someone licitly obtains the key (the password), are they entitled to share it with others as long as the purpose of the sharing is for a fair dealing purpose, such as research or education? That is the nub of the issue. Some observers proclaimed that this case proved that fair dealing trumped or allowed the bypassing of a TPM. That was not the court’s conclusion, as I pointed out (here) but the ruling effectively gutted password protection for businesses. A recent internal memorandum produced for the Minister of Canadian Identity and Culture by his department noted that “The use of passwords to limit access to copyright protected content is a common business practice among online platforms including news sites, streaming services and video game digital distribution services,” …“Rights holders may be concerned that passwords and paywalls are no longer seen as effective technological protection measures.”

The ruling is now under appeal, with a decision expected later this fall. Blacklock’s is a small David pitted against the taxpayer-funded, deep-pocketed Goliath of the Government of Canada, but I understand they may be getting some financial help from other paywall-dependent businesses. I hope so. The right thing to do would be for the Government of Canada to settle with Blacklock’s but maybe this wouldn’t remove ambiguities about the role of paywalls. An appeal court ruling may be needed. Stay tuned.

In the meantime, inconvenient as it may be, and recognizing that it is impossible to subscribe to everything you could possibly want at any given time, respect the integrity and the work of the journalists and their employers who bring you curated news, commentary and valuable reportage. Pay for what you can–and play by the rules for the rest.

© Hugh Stephens, 2025. All Rights Reserved.  

Bespoke Bookstore Tourism? Why Not? People Visit Gardens, Museums and Battlefields

Interior view of an ornate library with high wooden archways, numerous bookshelves, and visitors milling about. Busts line the walls, highlighting the historical significance of the space.

The Old Library: Trinity College, Dublin (Photo: Author)

Last month I published a guest column (Literary Pilgrimages: How Iconic Independent Bookstores are Becoming Travel Destinations) volunteered by reader Jennifer Mark who discussed how independent bookstores (or “bookshops” if you are from the other side of the Pond) are adding to their offerings to attract customers and survive. The enticements include offering a range of refreshments from coffee and tea with scones, to wine and beer, as well as providing interesting venues that are destinations in themselves. The writer’s interest had been piqued by my recent blog post highlighting the “Harry Potter” bookstore in Porto, Portugal, Livraria Lello, which my wife and I visited earlier this year. Lello, aided by a somewhat tenuous connection to the Boy Wizard given that JK Rowling lived in Porto for a short time while writing the first of the Potter books, has become a tourist attraction in its own right. It is even able to charge admission, a welcome new revenue stream. Jennifer took this a step further and wrote about the culinary attractions of some of the most interesting bookstores in the US, UK and Australia, and referenced booktowns in Japan and Argentina. It made me realize that bookstore tourism is definitely a “thing”. If you Google it, a cornucopia of information will spill forth. Not surprisingly, (since there is not much that is “new” these days), this concept has been around for a while.

It seems the idea of arranging customized tours to select bookstores for booklovers was first commercialized by Pennsylvania-based writer Larry Portzline, who wrote a book about it back in 2004. (Bookstore Tourism: The Book Addict’s Guide To Planning & Promoting Bookstore Road Trips For Bibliophiles & Other Bookshop Junkies). According to Portzline’s website;

“I first began leading “bookstore road trips” to New York City in 2003. I’d load 50 people on a chartered bus in Harrisburg, PA, and we’d spend the day visiting the 20 or so indie bookshops in and around Greenwich Village…I promoted the concept with a website, a blog, podcasts, and even a how-to book….Bookstore Tourism eventually grew into a grassroots effort in various locations around the U.S. It was never huge, but there was a nice ripple of interest and support.… I think Bookstore Tourism could rise again…It offers nothing but benefits: to the organizations that sponsor them, to the bookselling and travel industries, to literacy and reading efforts…”

Well said, Larry.

It looks as if Bookstore Tourism is back, whether it means visits to famous bookstores around the world or trips to “booktowns”, where no one bookstore is the destination but rather the collection of browsing options available, usually in a smaller but quaint and well situated (i.e. not too far from major population centres) town. The “book town” concept appears to have been first conceived in the 1970’s in Hay-on-Wye, Wales. It all started with a bookseller named Richard Booth in 1962, who opened his first bookstore in Hay-on-Wye’s old fire station. He subsequently opened six bookstores in the town, one in a ramshackle Norman castle, others in old warehouses. This critical mass led to others to open their own bookshops, over 30 in all. After becoming known as the world’s first booktown, Hay-on-Wye’s role as the place for literary mavens to gather was further cemented by the launch of its well known annual Literary Festival, which began in the late 1980s and continues to this day. It has provided new economic life for a small Welsh village of just 1500 residents. The idea has caught on as a means of revitalizing small towns that may have previously had an industrial base (one industry towns where the industry is long gone) but are now looking for new ways to attract tourists and visitors. In fact, there is even an association of booktowns, which, interestingly, does not seem to include Hay-on-Wye although it attributes the idea to Booth.

While there are all kinds of innovative ideas out there to attract patrons to bookstores and booktowns, it is still a challenging environment for many in the industry. A recent blog reported that there are now about 20 bookstores in Hay-on-Wye; at one point there were as many as 38. There was the onslaught of online booksellers, led by Amazon, that was supposed to kill the industry. It didn’t, and the innovation and entrepreneurialism of independent booksellers helps explain why, but it is still a struggle with many ups and downs. In the UK, the Bookseller.com reports that the number of independent bookshops slightly declined in 2024, continuing the previous year’s downward trend but still well above the 2016 low. Despite the small decline, almost 50 new independent bookshops opened in 2024 in Britain. In the US, the story is slightly different. According to this AP article, the number of members of the  American Booksellers Association has doubled since 2016, to almost 3000 members, but as Jennifer Mark pointed out in her blog on Literary Pilgrimages, store owners have had to be creative to survive.

In Canada, it appears that indies are hanging on but face a tough market, including potential new challenges from threatened Canadian tariff retaliation against Donald Trump’s tariffs on Canadian goods. The idea of creating a booktown in Canada has also been tried. At one point, Sidney, BC, just north of Victoria, had a dozen book outlets, several of them owned by the same couple, Clive and Christine Tanner. Sidney was also marketed as a book town. Those days are past, and there are now just a couple, although Sidney continues to be a great place to visit for other reasons.

But what about bespoke bookstore tourism for well-heeled book lovers who have already been on safari, visited Antarctica and done all those Viking river cruises? After all, there are specialized garden tours, museum tours, art gallery tours, history and battlefield tours, culinary (foodie) tours, wine tours (of course), but apparently no bookstore tours.  I tried to find a travel agency that offers tailored tours for booklovers but couldn’t identify any. Larry Portzline’s idea has never really taken off. Now there is a travel niche to be exploited! For now, booklovers may have to settle for a virtual tour. A good one can be obtained from “1000 Libraries”, a publisher that is marketing a glossy coffee table took featuring “The Most Beautiful Book Places in the World”. (The book looks pretty good on the website but I have not actually seen one; this is not a paid announcement).

Maybe virtual bookstore tourism is not such a bad alternative. After all, much of the pleasure of a book is to be able to visit places and enjoy stories from the comfort of your armchair/deckchair/chaise longue or wherever you choose to read. However, buying that book in an interesting place and having a memorable experience doing so adds to the enjoyment. A book can be forever and if you have the added benefit of a sight, sound, taste or smell sensation when buying it, I can guarantee that it will mean even more to you.

© Hugh Stephens, 2025. All Rights Reserved.

Echovita is Still Going Strong: The Sleazy (but Apparently Legal) Business of Monetizing Obituaries Without Consent

A close-up of delicate white flowers in bloom, alongside softly glowing white candles, creating a tranquil and reflective atmosphere.

Image: Shutterstock

The phone call came out of the blue. It was from a distressed family member who had just suffered a close and tragic personal loss. That in itself was obviously difficult enough to deal with. What was really upsetting was the unauthorized dissemination of that heart-rending event through Echovita, (or Echovita Canada) including postings on Facebook which, to add insult to injury, included inaccurate information. I wrote about Echovita last fall (Distasteful Yes, But Not Copyright Infringement: Publishing “Basic Fact” Unauthorized Obituaries is Going Strong (And Often Getting it Wrong). The caller had seen this post and was reaching out for advice and help. Unfortunately I was not able to provide much of either because, it seems, having learned its lesson through its predecessor operation, Afterlife, which was fined $20 million for copyright infringement and promptly went out of business, Echovita (which is based in Quebec) manages to stay just within the law. Nonetheless, its business model preys on the bereaved and from a moral perspective is about as sleazy a business as one can imagine. Moreover, it frequently gets some of the facts wrong because it is probably scraping hundreds if not thousands of obituaries daily and uploading data which is not verified or checked by anything other than a bot. And then it puts the onus on the bereaved family to reach out to correct the error, which it undertakes to do–within 3 days! If you Google “Echovita” you will see that the internet is rife with complaints and negative comments about this company, including from individuals on Reddit and from legitimate funeral homes warning consumers not to do business with them. Here is but one example, drawn from Facebook;

“A website called ECHOVITA and other third-party websites incorrectly rewrite obituaries that are posted on funeral home websites and then urge people to make a donation. The family does NOT benefit from this, so please DO NOT donate money unless you are legitimately on a funeral home website. If you are searching the internet for an obituary, the name may appear on different websites. Always look for the obituary hosted by the funeral home that is coordinating the services”

In other words, avoid Echovita. However, if you have lost a loved one, it is almost impossible to avoid them since Echovita doesn’t ask anyone’s permission for what it is doing, scraping authorized obituaries, extracting the so-called “public information” from them, and then posting an obituary rewrite alongside ads for various memorial offerings (flowers, planting trees, memorial books etc,) as well as rolling ads for various services. Needless to say, funeral homes don’t like these “freeriders” because what they are doing upsets their clients and might even be cutting into their own revenue for follow up services. As a result the Bereavement Authority of Ontario, a funeral industry regulatory body, has called out Echovita in a “Consumer Alert” published earlier this year. While it gets lots of negative publicity and media “exposes”, like this one recently on CTV, this doesn’t deter Echovita’s owners because they are basically chasing quantity over quantity. If they harvest enough obituaries and flood the internet with them, there will be enough people who believe they are dealing with the genuine obituary and will use the Echovita platform to order flowers and other services thorough third party suppliers (like Blooms Today—which itself offers very questionable service if you believe online reviews) allowing Echovita to turn a comfortable profit. However, maybe if enough people boycotted them they might go away? One can always hope.

Given the ingenuity that people display when trying to extract revenue from the internet through Youtube, it is not surprising that a Youtube version of obituary freeriding also exists. As reported in Wired, there is a fairly recent phenomenon on Youtube of videos featuring men reading obituaries using information harvested from funeral home websites, sometimes using voiceovers of “funereal” images like candles, sometimes just reading deadpan. On occasion these low-quality videos promote direct sale of products but generally the object is to attract enough aggregate views to qualify for advertisement revenue sharing from Youtube. It’s morbid filler for the internet. What next?

Just as with Echovita’s business model, these videos avoid copyright infringement by extracting basic information (dates of birth and death, location of death etc) without using the full obituary or any visuals. As I explained in earlier blog posts, an obituary is often a creative work embodying original expression and is thus protected by copyright laws. It is the story of a person’s life, often written by a close relative or in some cases in advance by the deceased person themself. However, basic facts and ideas cannot be protected, only the expression of an idea based on those facts or ideas. The same is as true for obituaries as it is with news, and this is where the obituary harvesters enter the picture, picking up the unprotectable basic facts.

What can be done about it? Unfortunately, not a great deal except to contact Echovita to request that an unauthorized obituary be either taken down or corrected. But be careful in doing so because if, in the process of dealing with Echovita, you create an account you are agreeing to their Terms of Service which gives them rights to use your data. You can email them without creating an account but at least one Reddit listing stated that in order to approve the takedown request, you have to click on an Echovita link which then gives Echovita access to your data. Whether this is true or not, I can’t say because fortunately I have not been in the position of having to opt out of Echovita’s “services”.

Echovita encourages people to use their website to post authorized obituaries and also provides other information such as listings of funeral homes and a searchable database. This is all in an attempt to appear legitimate and drive traffic but their basic business model is based on unauthorized scraping of basic obituary information off other websites, including Legacy.com, another obituary-based business that I wrote about recently (Obituaries and Copyright: If You Publish an Obituary in the Globe and Mail (and many other papers), Be Prepared for Legacy.com and its Upsell Business Model). While Legacy.com also profits from sales of memorial items associated with obituaries of individuals, at least it is based on consent as it draws its content from obituary listings in newspapers where those placing and paying for the printed obits acknowledge and accept that the paid-for newspaper listing will be put up on the internet and given a wider reach through Legacy.com. Often, they pay extra for the Legacy.com listing. In other words, it is an “opt-in” service, unlike Echovita where people not wanting their services are required to “opt-out”.

The solution to this scourge must surely reside in privacy rather than copyright law, or perhaps in some kind of consumer protection legislation. I hope so. But for now it seems that entities like Echovita have been able to find a sweet spot that enables them to continue to take advantage of a very personal and private part of life. It should be possible to respectfully and lovingly bid adieu to departed friends and family with dignity without crass commercialization. The last thing a bereaved family needs is having to deal with unwanted and inaccurate memorializations of their loved one by online businesses that see someone’s life as just another opportunity to generate a quick buck. Very sad.

© Hugh Stephens, 2025.

Hold the Champagne: The Two AI Training/Copyright Decisions Released in the US Last Week Were a Mixed Bag for AI Developers

Illustration of a champagne bottle being popped, enclosed in a red circle with a slash indicating 'no champagne'.

Image: Shutterstock.com

Last week I wrote about the questionable ethics of META’s use of pirated content to train its AI model, Llama, pointing out the ethical issues involved with META’s admitted use of pirated online libraries, such as LibGen (Library Genesis), to feed content to Llama for training purposes. This is quite apart from whatever legal issues that may arise from the widespread practice of ingesting copyrighted content for AI training by making an unauthorized copy from any source (such as a legitimate library, through purchase of a single copy of a work, or from publicly available internet sources, for example) not to mention the additional element of taking that content from pirate sources. The day after that blog was posted the first of what will be a series of legal decisions in the US regarding cases brought by authors and copyright holders against AI companies was issued, followed by another a day later. Both cases were heard in the Northern District of California, in the same San Franciso court house, but handled by different judges.

I updated last week’s blog to make reference to the Bartz v Anthropic case (hereafter “Anthropic”), but given the importance of that decision, combined with a decision released in another California court room a day later (Kadrey et al v META), these cases merit further exploration–especially since they were widely trumpeted by AI advocates as opening the door to unauthorized use of copyrighted content for AI training on the basis of “fair use”.

Fair use is the complex legal doctrine used in the US to determine exceptions to copyright protection. US readers are well aware of the intricacies and idiosyncrasies of fair use but for those not overly familiar with how it works, here is a short summation I drew from a blog post on fair use vs fair dealing that I wrote a few years ago.

In the US context, fair use is an affirmative defence against copyright infringement and is determined by the courts on a case by case basis, judged against several fairness factors (purpose and character of the use, the nature of the work copied, the amount and substantiality of the amount of the work used, and the effect of the use on the value of the original work)… Fair use is not defined by law. Some examples are given in US law of areas where the use is likely to be fair (criticism, comment, news reporting, teaching, scholarship, research) but these are illustrative and not exhaustive. In short, it is the courts that decide. This in turn can lead to extensive litigation as to what is and is not fair use, and it is worth noting that different judicial circuits in the US have at times come up with conflicting interpretations.

Or, for that matter, two different judges in the same circuit delivering decisions just days apart on similar issues but with some significantly different outcomes, as we saw last week (although in these cases both found fair use by AI developers with regard to the copyrighted works at issue).

On the Anthropic case, US District Judge William Alsup ruled, on summary judgement, that the use of copyrighted works for AI training, even though done without authorization, is highly transformative and does not substitute for the original work (“The technology at issue was among the most transformative many of us will see in our lifetimes”). It thus qualifies, according to Alsup, as fair use because the transformative nature of the use overrides or swallows the three other fair use factors, including the important fourth factor (effect of the use on the value of the work). He notes there was no allegation that the output of Anthropic’s model, known as “Claude”, produced content infringing the works of the plaintiffs. However, Judge Alsup then went on to consider the legality of Anthropic’s actions to download more than 7 million works from pirate libraries (such as Books3, Library Genesis and the Pirate Library Mirror) to constitute its reference library, which it initially planned to use for AI training. He concluded this was a prima facie case of copyright infringement, whether Anthropic intended to use some or all of the pirated works to train Claude or not. (“Anthropic seems to believe that because some of the works it copied were sometimes used in training LLMs (Large Language Models), Anthropic was entitled to take for free all the works in the world and keep them forever with no further accounting “.) Damages, to be decided at trial, could be substantial. Alsop did not, however, rule explicitly on whether or not the use of pirated works for AI training purposes could be a fair use.

Because of the controversial nature of Alsup’s findings on transformation and fair use, there is no question that this case will be appealed. While there have been many criticisms of the fair use elements of Alsup’s ruling, a particularly clear and trenchant analysis was put forth by Kevin Madigan of the Copyright Alliance (Fair Use Decision Fumbles Training Analysis but Sends Clear Piracy Message).

The second case last week to reach the decision stage was Kadrey et al v META. In this case District Judge Vince Chhabria found that META’s use of the works of the plaintiffs, thirteen noted fiction writers, to train its AI model (“Llama”) was also fair use. Chhabria, like Alsup, found that META’s use was transformative on the first fairness factor dealing with the purpose and character of the use (“There is no serious question that Meta’s use of the plaintiffs’ books had a “further purpose” and “different character” than the books—that it was highly transformative.”) but unlike Alsup, Chhabria put much greater emphasis on market harm, (the fourth fairness factor dealing with the effect of use on the value of the work) suggesting that it could be determinative. Unfortunately for the plaintiffs, however, Chhabria considered their arguments with respect to market harm to be unconvincing. There was no evidence that Llama’s output reproduced their works in any substantial way or substituted for the specific works at play nor was there evidence, according to the judge, that the unauthorized copying deprived the authors of licensing opportunities.

Chhabria suggested that a far more cogent argument would have been that use (unauthorized reproduction) of copyrighted books to train a Large Language Model might harm the market for those works by enabling the rapid generation of countless similar works that compete with the originals, even if the works themselves are not infringing. In other words, causing indirect substitution for the works rather than direct substitution. This is the theory of “market dilution”, which was also put forward speculatively by the US Copyright Office in its recent Pre-Publication Report on AI and copyright. Since this wasn’t presented as an argument, Chhabria could not rule on it but in effect he is inviting future litigants to pursue this line of argument, noting that his decision on fair use relates only to the works of the thirteen authors who brought the case.

The clearest way to illustrate his line of reasoning is to quote directly,

“In cases involving uses like Meta’s, it seems like the plaintiffs will often win, at least where those cases have better-developed records on the market effects of the defendant’s use. No matter how transformative LLM training may be, it’s hard to imagine that it can be fair use to use copyrighted books to develop a tool to make billions or trillions of dollars while enabling the creation of a potentially endless stream of competing works that could significantly harm the market for those books”.

This editorializing, known in legal circles as obiter dicta, is not binding nor precedential, yet will undoubtedly have some influence given Chhabria’s stature. It is likely that one of these days Judge Chhabria will have the opportunity to put these theories into practice when ruling on a similar case, but one where the plaintiffs have made a better case for market harm. He has provided them a roadmap.

While these two cases have fired the first shots in what is going to be a lengthy war, they do not seem to be dispositive. There are enough caveats and nuances to be able to conclude that the AI developers are far from being out of the woods. Both “victories” have a sting in their tail, especially Judge Alsup’s finding on piracy. Neither copyright advocates nor AI developers should be breaking out the champagne just yet. But whichever way it turns out, there will be some sure winners; the lawyers for each side.

© Hugh Stephens, 2025.

Is it Ethical to Use Pirated Content for Commercial Purposes? META Thinks So

Two signs hanging on a string, one labeled 'ETHICAL' in green and the other labeled 'LEGAL' in red, against a purple background.

Image: Shutterstock.com

There is the question of what is ethical, and then there is the question of what is legal. Sometimes they are the same, often not. The legality of using copyrighted content without authorization for commercial purposes, such as in training AI models—as META and a number of other companies have done—is being decided in court. In META’s case, however, there is the further complaint (not denied by META) that many of the unauthorized copies it made were taken from pirated content. While this revelation may not change the fundamentals of the copyright infringement case against it, there is still the ethical question for META to answer. On this, it comes up short. Very short.

META, the parent company of Facebook, Instagram and WhatsApp, used vast amounts of copyrighted content, without permission or licensing, to train its AI model. It is not alone in doing so. This practice may or may not be legal. A number of cases are working their way through the courts, most of them in the US, with copyright owners from Getty Images to Disney and Universal, from the New York Times to the Authors Guild and on to music labels, all claiming that their content was unfairly and illegally copied to provide training fodder for training AI models, such as META’s model, Llama. META and other AI developers claim that their use was a “fair use” under US law. We’ll see. However, as part of its giant vacuuming of publicly available (but in many cases protected) content, META also ingested content from various pirate sites and databases, notably the notorious “shadow library”, LibGen (Library Genesis). LibGen originated in Russia and contains up to 80 million scientific and academic articles, as well as millions of novels and nonfiction books, most unauthorized, unlicensed copies. It has a been sued by major academic and textbook publishers. In 2017 Elsevier won a $15 million judgement against LibGen, and another pirate website, SciHub. Last year Elsevier was awarded a $30 million default judgement. However, both LibGen and SciHub remain available online.

The extent of the copyrighted content held by LibGen was revealed in an investigative report published recently by The Atlantic. You can search through the LibGen database as published by The Atlantic to find out what works are included, and whether your work has been pirated. Authors from Newfoundland to New York and lots of places in-between and elsewhere found their works included in the database when they did the search. The Authors Guild advises writers to fight back by sending a formal notice to META and other AI companies asserting their rights, as well as adding a “No AI Training” notice on the copyright page of works. This is in addition, as would be expected, to joining the Authors Guild to help them fight what is happening.

Consuming pirated content can result in costly penalties, as some unfortunate downloaders have found out to their regret. Using it for commercial purposes is even more egregious. It’s like running a pirate streaming service based on stolen content. META didn’t use pirated content in this way, but they used it commercially just the same, in their case for AI training. Were they aware of what they were doing. You bet they were.

The discovery process in the US suit of Kadrey et. al. v META revealed a series of email exchanges in which some META employees expressed concerns over the ethics of using pirated content. The concerns went up the chain and back, with “MZ” (guess who? No, not Moses Znaimer) giving approval to proceed. Following on these revelations in Kadrey v META in the US, two class action lawsuits have been filed in Canada, one in Quebec on behalf of a number of French language authors and one in British Columbia. The Quebec suit specifically flags the piracy issue. Among the listed complaints is the following:

“Rather than acting within the law and respecting the rights of class members, it (META) deliberately chose to train its LLMs (Large Language Models) from datasets containing illicit copies of works from all over the world, including those of class members.”

Damages sought are $20,000 per work.

It is clear that META wilfully torrented content from LibGen, knowing that many or most of the works on LibGen were infringing, pirated copies. They just didn’t care.

If it turns out that somehow, inexplicably, META’s unauthorized use of copyrighted content for AI training is ruled by the US courts to be fair use, would the fact that the source of some of the content was from a pirate source be relevant? I am not sure, but a judgement that has just been delivered in California in the “Anthropic” case suggests that even if unauthorized copying can be justified as fair use because it is considered “transformative”, that does not excuse piracy–which is still an infringement. In this case, Anthropic copied both purchased and pirated works to train its AI model, and kept the copies in its central library. It was sued by some authors and journalists in a class action suit alleging copyright infringement. The judge, in one of the first such cases to reach a decision point, concluded on summary judgement that Anthropic’s unauthorized reproduction of copyrighted works for AI training was fair use under the transformation doctrine but added, with respect to those works drawn from pirate sources such as LibGen and others,

“piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and
immediately discarded”
.

A trial will be held to determine the damages from the piracy.

Being one of the first AI training cases out of the gate, Anthropic will certainly be appealed, so this is not the last word. However, this ruling when added to the US Copyright Office’s views expressed in its Pre-Publication Report on Generative AI issued last month on May 9, a day before Register Shira Perlmutter was dismissed, that “the copying of expressive works from pirate sources in order to generate unrestricted content that competes in the marketplace, when licensing is reasonably available, is unlikely to qualify as fair use”, suggests that META could be in both ethical and legal trouble.

In other jurisdictions, such as Singapore, content used in AI training under a Text and Data Mining exception has to be legally accessed, although this is very thin legal protection because technology companies can legally purchase just one copy of a work to comply. In Canada you cannot break the law (i.e. circumvent a technological protection measure) to exercise a fair dealing right. But whether or not using a pirated source puts META offside the law in the US with respect to fair use, (and the Anthropic case suggests that it could at least with respect to the pirated works), think of the ethics and the image this presents to the public.

A company like META, capitalized at something like $2 trillion, cannot be bothered to even access content legitimately, let alone use it legitimately. Why? Because MZ said it was ok to proceed. Sadly, even though they are not the only ones to use pirated content to train their AI models, that tells me all I need to know about the values and ethics of this particular company.

© Hugh Stephens, 2025. All Rights Reserved.

This post has been updated to include reference to the decision in the Anthropic case, released after the initial publication of this blog post.

Copyright Litigation in China: Some Interesting AI-Related Decisions from Chinese Courts

A wooden gavel resting on a circular base in front of a red backdrop featuring the flag of China.

Image: Shutterstock

These days just about any information in North America related to China, especially regarding intellectual property (IP), is highly negative. The narrative is along the lines of “China is an adversary with deliberately lax IP laws who has stolen and continues to steal our IP, etc.”. This characterization of China is reinforced by our political leaders (When asked during the Leaders’ debate what was the greatest security threat to Canada, Prime Minister Carney replied with one word. “China”). Donald Trump continues to have an obsession with China, the latest manifestation of which is the recent announcement that the US will revoke the visa status of an undetermined number of Chinese students currently studying in the US. (Over a quarter of a million students from China are currently studying at American colleges and universities, many simply seeking an alternative to studying in the hyper-competitive environment at home). The “China as IP thief” narrative is supported by government publications such as the annual Special 301 Report produced by the Office of the US Trade Representative (USTR) which this year had ten full pages on China. One excerpt will suffice to give you the flavour of the report. “In 2024, the pace of reforms in China aimed at addressing intellectual property (IP) protection and enforcement remained slow…Concerns remain about longstanding issues, including technology transfer, trade secrets, counterfeiting, online piracy, copyright law, patent and related policies, bad faith trademarks, and geographical indications.” Well, that covers the waterfront. One wonders how Chinese brands, innovators and creators manage to survive in such an environment.

This is not to dismiss the darker side of China’s long IP history. Have there been cases of industrial espionage involving China? Yes, certainly. There have also reportedly been more than 1200 intellectual property theft lawsuits brought by US companies against Chinese entities in either the US or China over the past 25 years. There is no question that IP protection in China is not all it could and should be, or that some Chinese companies and other entities have been aggressive in seeking to acquire IP by less than transparent means. But that is not the whole story. While the number of IP infringement lawsuits against Chinese entities over the years sounds like a lot, this business website estimates that the number of IP litigation cases globally totals around 12,000 annually. There are several thousand patent litigation cases alone in the US each year. A lot of US companies sue other US companies in the patent, trademark and copyright field. And Chinese companies sue Chinese companies.

In the past, Chinese IP laws had loopholes, were often weakly enforced and were dealt with by courts that had scant knowledge and training in IP matters. That is rapidly changing as China not only climbs the innovation ladder, but has come to dominate it in some areas, such as EV’s and EV batteries, cashless payment systems, renewable energy and others. It is rapidly catching up in generative AI. While this has been happening, Chinese courts have been producing some interesting and increasingly sophisticated decisions when it comes to AI and copyright. China–like other countries–is grappling with several aspects of this issue. There is the question of finding the right balance between protecting creators and innovators while using domestic creative works to spur AI training, development and research. Another element is the extent to which AI assisted or created works qualify for copyright protection. There is currently no Text and Data Mining (TDM) exception in Chinese law to allow AI training on copyrighted content nor is there a definitive interpretation as to whether content produced by AI can be protected by copyright. However, several court decisions, which we examine below, have shed some light on this complex question.

Dreamwriter Case

In one of the earlier cases, which I wrote about back in 2020, (the Dreamwriter case), a Chinese court (in Shenzhen) ruled that an automated article written by an AI program (Dreamwriter), created by Tencent, which had been copied and published without permission by another Chinese company, Yinxun, was nevertheless subject to copyright protection because it met the originality test through the involvement of a creative group of editors. These people had performed a number of functions to direct the program, such as arranging the data input and format, selecting templates for the structure of the article, and training the algorithm model. The article was ruled to be a protectable work, and Yinxun was found to have infringed.

Li v Liu Case

The relatively loose interpretation regarding the degree of human engagement required to protect the output of an AI program in the Dreamwriter case has been supported by other Chinese courts. In the prominent Li v Liu case, the Beijing Internet Court ruled that Mr. Li, who had created the image of a young woman using the AI program Stable Diffusion, had provided “significant intellectual input and personalized expression” in creating the image through a series of prompts. As explained in detail by this article from Technollama, the prompts (along with a number of negative prompts) were sufficient for the court to decide that Li had met the standard of creative expression.

These were Li’s prompts;

“ultra-photorealistic: 1.3), extremely high quality highdetail RAW color photo, in locations, Japan idol, highly detailed symmetrical attractive face, angular symmetrical face, perfect skin, skin pores, dreamy black eyes, reddish-brown plaits hairs, uniform, long legs, thighhighs, soft focus, (film grain, vivid colors, Film emulation, kodak gold portra 100, 35mm, canon50 f1,2), Lens Flare, Golden Hour, HD, Cinematic, Beautiful Dynamic Lighting”

Liu, who had been sued by Li for using the AI generated image without authorization, was found liable for infringement and fined 500 CNY (about USD75).

At that time (late 2023), this decision was considered ground-breaking for image-based works given the position of the US Copyright Office (USCO). USCO had denied copyright registration to several generative-AI created image works owing to insufficient human creativity. (see If AI Tramples Copyright During its Training and Development, Should AI’s Output Benefit from Copyright Protection? Part One: Stephen Thaler and Part Two: Jason Allen). Since then (in January of this year) the USCO has taken a more nuanced position, permitting registration of an AI assisted work (an image called A Single Piece of American Cheese, created by graphic artist Kent Kiersey). Although Kiersey used InvokeAI to create the work, in the view of the US Copyright Office, sufficient human creativity was involved through the “selection, coordination, and arrangement of material generated by artificial intelligence”.

Plastic Chair Case

If China has been in the forefront of acknowledging that human control over AI tools used to generate content qualifies the works for copyright protection, a more recent case has reset the pendulum somewhat. As recounted in this blog by UK-based market research firm IAM, very recently a court in Jiangsu Province dismissed a copyright infringement claim brought by a designer against a company that manufactured, without a licence, children’s plastic chairs based on her AI-based designs. The designer, Feng Runjuan, had created three designs using the AI program Midjourney and posted them to social media, including the prompt she had used. Her prompt was “Children’s chair with jelly texture, shape of cute pink butterfly, glass texture, light background“. The company manufacturing the chairs approached Feng to license the designs but was unable to reach an agreement with her. They then went ahead anyway (without a licence) to produce chairs that bore some similarity to the original designs, using Feng’s original prompt with some tweaks. Feng sued. There was little doubt that the chair manufacturing company had used her prompts to produce the chair design, but the key question was whether the AI generated designs qualified as original works meriting copyright protection.

Feng was unable to reproduce the original images using her prompts owing to the randomness of the AI program. This suggested to the court that it was the AI program making the design decisions, not the person providing the prompts. As outlined in the IAM article referenced above, the court held that a user must provide a verifiable creative process that shows the:

  • adjustment, selection and embellishment of the original images by adding prompts and changing parameters; and
  • deliberate, individualised choices and substantial intellectual input over the visual expression elements, such as layout, proportion, perspective, arrangement, colour and lines.

It concluded that the original images did not qualify as original works and thus they could not be protected. Feng’s lawsuit failed.

So now we have a situation where one Chinese court has ruled that the prompts generated by Li in what I will call the “young girl image” case constituted sufficient intellectual input and personalized expression to qualify for copyright protection, even though the actual image was generated by an AI program, whereas another court has denied copyright protection for a work also produced with prompts, albeit simpler and far fewer. The difference seems to be the degree of human involvement in creating the prompts, although the fact that Ms. Feng in the plastic chair case was unable to reproduce the original images seems to have also weighed against her. As anyone who has ever used an AI program will know, identical prompts will produce different images owing to the way the program works. Does that disqualify the artist? I would hope not, but the degree of control is clearly a key factor, as both the rulings of Chinese courts and the recent USCO decision to register the work A Single Piece of American Cheese would seem to show. Both Chinese court decisions are defensible, demonstrating careful and reasoned consideration, and are helpful in establishing parameters for use in determining whether works are AI assisted or AI created.

Ultraman Case

Another area where Chinese courts have left their mark is on the topic of AI liability for copyright infringement. In what is known as the “Ultraman” case, a Chinese court (the Guangzhou Internet Court, upheld on appeal by the Intermediate Peoples’ Court in Hangzhou) delivered a ruling of contributory infringement against a company that provided AI generated text-picture services through its website. The complainant was the Chinese licensee of the Japanese company that owns the rights to the cartoon character Ultraman. When the defendant’s website (effectively a chat-bot capable of generating AI images at its users’ request) was asked to generate an Ultraman-related image, it generated a character that appeared to be substantially similar to the claimant’s licensed Ultraman. The court had to decide whether the defendant had infringed the plaintiff’s reproduction and derivative production rights and if so, what remedies were applicable.

In its ruling the court decided that even though the defendant did not directly infringe the licensee’s rights, its failure to exercise a reasonable duty of care to prevent infringements (for example, by cautioning users or providing adequate filtering or blocking mechanisms), rendered it liable for contributory infringement. It was ordered to compensate the claimant the amount of CNY 10,000, about USD1500 (considerably less than the damages sought of CNY300,000). Here we have another sophisticated and well reasoned decision, which appears to have been the first instance globally of recognizing the liability of an AI platform for contributory copyright infringement. It does not create any legal precedents but is a useful contribution to the emerging debate.

These cases well illustrate the growing sophistication and complexity of IP rulings in China and are reflective, in my view, of an economy that is rapidly moving up the innovation and creativity ladder. When it comes to IP protection in China, is the glass half empty or half full? I would argue the latter, even though this may not be the most popular interpretation these days. One thing that I am willing to predict with certainty is that we can expect more interesting and thoughtful IP legal decisions from the Chinese legal system in the months and years ahead.

© Hugh Stephens, 2025. All Rights Reserved.

AI’s Habit of Information Fabrication (“Hallucination”): Where’s the Human Factor?

An illustration of a cartoonish robot face on a computer screen with the text 'THE WORLD IS FLAT' above it.

Image: Shutterstock (with AI assist)

It is well known that when AI applications can’t respond to a query, instead of admitting they don’t know the answer, they often resort to “making stuff up”—a phenomenon commonly called “hallucination” but which should more accurately be called for what it is, total fabrication. This was one of the legal issues raised by the New York Times in its lawsuit against OpenAI, with the Times complaining, among other things, that false information attributed to the journal by OpenAI’s bot undermined the credibility of Times journalism and diminished its value, leading to trademark dilution. According to a recent article in the Times, the incidence of hallucination is growing, not shrinking, as AI models develop. One would have thought that as the models ingest more material, including huge swathes of copyrighted and curated material such as content from reputable journals like the Times (without permission in most instances), its accuracy would improve. That doesn’t seem to be the case. Given AI’s hit and miss record of accuracy, it should be evident that AI output cannot be trusted or, at the very least, can only be trusted if verified. Not only is AI built on the backs of human creativity (with a potentially disastrous impact on creators unless the proper balance is struck between AI training and development, and the rights of creators to authorize and benefit from the use of their work), but human oversight and judgement is required to make it a useful and reliable tool. AI on auto-pilot can be downright dangerous.

The most recent outrageous example of AI going astray is the publication by the Chicago Sun-Times and Philadelphia Inquirer, both reputable papers (or at least they used to be), of a summer reading list in which only five of fifteen books listed were real. The authors were real but most of the book titles and plots were just made up. Pure bullshit produced by AI. The publishers did a lot of backing and filling, pointing to a freelancer who had produced the insert on behalf of King Features, a unit of Hearst. Believe it or not, it was actually licensed content! That freelancer, reported to be one Marco Buscaglia, a Chicago “writer”, admitted that he had used AI to create the piece and had not checked it. “It’s 100% on me”, he is reported to have said. No kidding. Pathetic. Readers used to have an expectation that when a paper or magazine published a feature recommending something, like a summer reading list, the recommendation represented the intellectual output of someone who had done some research, exercised some judgement, and had presumably even read or at least heard about the books on the list. How could anyone recommend non-existent works? The readers trusted the newspaper, the paper trusted the licensor, the licensor trusted the freelancer, the so-called author. Nobody checked. Where was the human element? The list wasn’t worth the paper it was printed on.

The same problem of irresponsible dependence on unverified information produced by AI is a growing problem in the legal field. Prominent lawyer and blogger Barry Sookman has just published a cautionary tale about the consequences of using hallucinatory AI legal references. Counsel for an applicant in a divorce proceeding in Ontario cited several legal references using the CanLII database (for more information on CanLII see “AI-Scraping Copyright Litigation Comes to Canada (CANLII v Caseway AI”) that the presiding judge could not locate—because they did not exist. He suspected the factum had been prepared using Generative AI and threatened to cite the lawyer in question for contempt of court, noting that putting forward fake cases in court filings is an abuse of process, and a waste of the court’s time. The lawyer in question has now confirmed that AI was used by her law clerk, that the citations were unchecked, and has apologized, thus avoiding a contempt citation. Again, nobody checked (until the judge went to the references cited).

This is not even the first case in Canada where legal precedents fabricated by AI were presented to a court. Last year in a child custody case in the BC Supreme Court, the lawyer for the applicant was reprimanded by the presiding judge for presenting false cases as precedents. The fabricated information was discovered by the defence attorneys when they went to check the applicant’s lawyer’s arguments. As a result, the applicant’s lawyer was ordered to personally compensate the defence lawyers for the time they took to track down the truth. The perils of using AI to argue legal cases first came to prominence in the US in 2023 when a New York federal judge fined two lawyers $5000 each for submitting legal briefs written by ChatGPT, which included citations of non-existent court opinions and fake quotes.

Another area fraught with consequences for using unacknowledged AI generated references is academia. The issue extends well beyond undergraduate student essays being researched and written by AI to include graduate students, PhD candidates and professors taking shortcuts. This university library website, in its guide to students on use of AI generated content, notes that LLMs (Large Language Models used in AI) can hallucinate as much as 27% of the time and that factual errors are found in 46% of the output. The solution is pretty simple. When writing a research paper, don’t cite sources that you didn’t consult.

This brings up the question of “you don’t know what you don’t know”. If your critical faculties are so weak as to not be able to detect a fabricated response, you are in trouble. Of course, some hallucinations are easier to spot than others. Some of the checking is to simply verify that a fact stated in an AI response is accurate or that a cited reference actually exists (but then it should be read to determine relevance). In other cases, it may be more subtle, with the judgement and creativity of the human mind being brought into play to detect a hallucination. That requires experience, knowledge, context—all of which may be lacking in the position of a junior clerk or student intern assigned the task of compiling information. This is all the more reason why it is important for those using AI to check sources, and to exercise quality control. Part of the process is to ensure transparency. If AI is used as an assist, that should be disclosed.

At the end of the day, AI depends on human creativity and accurate information produced by humans. Without these inputs, it is nothing. This brings us to the fundamental issue of whether and how copyright protected content should be used in AI training to produce AI generated outputs.

The US Copyright Office has just released a highly anticipated study on the use of copyrighted content in generative AI training. Here is a good summary produced by Roanie Levy for the Copyright Clearance Center. The USCO report is clear in stating that the training process for AI implicates the right of reproduction. That is not in doubt. It then examines fair use arguments under the four factors used in the US. Notably, with respect to the purpose and character of the work used for training, USCO notes that the use of copyrighted content for AI training may not be transformative if the resulting model is used to generate expressive content or potentially reproduce copyrighted expression. It notes that the copying involved in AI training can threaten significant harm to the market for, or value of, copyrighted works especially where a model can produce substantially similar outputs that directly substitute for works used in the training data. This report is not binding on the courts but is a considered and well researched opinion by a key player.

It is interesting to note that the report was released quickly in a pre-publication version on May 9, just a day before the Register of Copyrights (the Head of the Office) Shira Perlmutter was dismissed by the Trump Administration and a day after the Librarian of Congress, Carla Hayden (to whom Perlmutter reports) was fired. Washington is rife with speculation on the causes for, and the legality of, the dismissals. We will no doubt hear more on this. With respect to fair use in general, the study concludes that “making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets…goes beyond established fair use boundaries”. The anti-copyright Electronic Frontier Foundation (EFF), of course, disagrees. (Which probably further validates the USCO’s conclusions).

The USCO study is about infringement, not hallucination or fabrication, yet both stem from the indiscriminate application and use of AI where the human factor is largely ignored and devalued. Human creativity and judgement is needed to set guardrails on both. Transparency as to what content has been used to train an AI model, along with licensing reputable and reliable content for training purposes, are important factors in helping AI to get its outputs right. Not taking an AI output as gospel but applying a degree of diligence, common sense, fact verification or experienced judgement are other important factors in deploying AI as it should be used, as an aide and assist to make human creativity and human directed output more efficient but not as a substitute for thinking or original research. Generative AI must be the servant, not the master. Human creativity and judgement are needed to ensure it stays that way.

© Hugh Stephens, 2025. All Rights Reserved.