India’s Proposed Compulsory Licensing Scheme for AI Training: It Could be the Worst of All Worlds.

Used with permission

In early December the Indian government through the Department for Promotion of Industry and Internal Trade (DPIIT), launched a public consultation on its proposal for a “One Nation, One Licence, One Payment” regime to govern AI training on copyrighted content. The comment period closes early next month. The stated intent is to establish a framework to allow AI developers to access proprietary content for AI algorithm training, without rightsholder authorization, while ostensibly taking into account the concerns and economic interests of those same rightsholders. The proposal from a semi-official committee established by DPIIT is based on a compulsory licence regime covering all content, past, present and future, with no opt outs for content owners. Given the unrepresentative nature of the committee, which failed to include rightsholders from many creative sectors, it is not surprising it has managed to come up with a solution that pleases absolutely no-one, not rightsholders nor the AI industry. Moreover, it is probably unworkable and could rightly be described as the “worst of all worlds”. How did we end up here?

It is widely accepted that to further refine and develop AI, more data is constantly required. Quality data is important. The better the data, the better the product. Many AI developers have, without authorization, already helped themselves to vast amounts of copyright protected material either through accessing pirate databases or by simply hoovering up publicly accessible (but nonetheless copyright protected) content on the internet. The result has been a series of lawsuits launched by rightsholders, primarily in the US but also in India, the UK and Canada. In the US the parameters of the fair use doctrine are being tested, with mixed results. There have been some settlements, notably the Bartz v Anthropic case in which AI developer Anthropic agreed to pay a group of authors $1.5 billion for having accessed and reproduced their works through pirate websites (and without authorization), as well as a large number of voluntary licensing agreements between corporate rightsholders such as media outlets, publishers and music labels, and AI companies. Clearly, legal leverage is needed to convince AI developers to negotiate with rightsholders (no one is volunteering, that’s for sure). In the US this is backed up by a statutory damages’ regime where the risk of being hit with large statutory damages per infraction is a strong incentive for the AI industry to reach licensing deals. Unfortunately, many countries do not have a statutory damages provision in their copyright law, and India is one of them.

Ironically, outside the US where the fair use case-by-case legal doctrine does not exist, AI developers are arguably in even greater legal jeopardy. As a result, they have pushed for a wide statutory text and data mining (TDM) exception to be introduced into national copyright laws. The AI industry is doing this in India, although there is strong opposition to introducing a TDM escape hatch. When a TDM exception applies, rightsholders are paid nothing. The DPIIT “solution” proposes to address the non-payment issue–but is going about it in precisely the wrong way. Appropriating the IP of all rightsholders by imposing a draconian compulsory license regime penalizes rather than rewards rightsholders and imposes a lowest-common-denominator value on all content, not distinguishing between premium and pedestrian content while taking away from rightsholders the inherent right to determine how and where their content is used.

The compulsory licence regime proposed by DPIIT would be administered by a new non-profit body, a Collective Management Organization (CMO) to be established by statute. This would create another layer of bureaucracy harking back to the days of the “licence raj”. A government appointed committee would set rates in conjunction with the new non-profit. Royalties would be applied retroactively as well as prospectively. Existing voluntary licensing agreements between content providers and AI companies would have to be terminated and future licence agreements prohibited. Rightsholders who do not want to licence their content could not opt-out. This is overreach at its worst.

The tech industry represented by NASSCOM, the National Association of Software and Service Companies, that also includes Google and Microsoft, vehemently opposes the One Nation, One Licence, One Payment proposal, arguing instead for a TDM exception. It also argues that rightsholders who wish to opt out should be able to do so, thus avoiding AI companies having to pay royalties for content it does not want or has already licensed. It also doesn’t want the precedent of compulsory payments, as India would be the first country globally to institute such a regime. NASSCOM claims that a compulsory licence regime would slow innovation, the card always played by the AI industry when the issue of protecting rightsholders is discussed. However, a TDM exception is not currently on offer in India. It is a “solution” that is facing increasing pushback in many countries. Australia has just rejected the concept, and other jurisdictions are examining it critically.  Where a TDM exception does exist, as in the UK and EU, its use is constrained and subject to several conditions, such as in Britain where it is limited to non-commercial research.

Not only is the tech industry strongly opposed to the compulsory licence proposal, so are rightsholders such as the broadcasting industry and Bollywood. Not only would a compulsory licence be extremely difficult to implement given the nature of CMOs which are generally highly inefficient and administratively cost heavy– not to mention that the Committee’s working paper proposes two layers of CMO (thus double handling and processing)– but the draconian and sweeping nature of the compulsory licence is of great concern. Rightsholders are given no ability to opt out or to refuse to have their content conscripted. In fact, the proposal includes a provision granting AI developers access to content as a “matter of right”. What happened to the fundamental right of an author to determine how, when and even if their content is to be reproduced? Compulsory licences are extraordinarily blunt instruments and do not work when sophisticated content like audio-visual and music products are involved. There are many elements to licensing, including contractual issues that go beyond price that are carefully negotiated and carry with them specific obligations and privileges. A compulsory licence is a “one size fits all” solution. It strips away the rights of content owners and is, in effect, a form of “compensated expropriation”. And the compensation is minimal.

Having pleased no-one with its proposal, it remains to be seen where DPIIT will go next once all comments are received and evaluated. A follow-up proposal is expected that will deal with outputs, just as the initial one did with inputs. India, like many countries, wants to participate in the global development of AI. Its rich local-language content is a strategic asset. Imposing a compulsory licence would stifle the creativity that drives this cultural comparative advantage. In the absence of a market failure there is no rationale for resorting to the sledgehammer of a compulsory licence for all content. The preferred solution for rightsholders, voluntary licensing, is also becoming a preferred solution for AI developers as the legal ground on which they are operating in accessing content without authorization is looking increasingly shaky. Voluntary licensing is growing globally as AI developers absorb this reality. If wide TDM exceptions are off the table, AI developers will have no recourse but to negotiate licence agreements. (A statutory damages regime would provide even greater incentive to do so). Bringing in a wide TDM exception or worse, introducing a blunt and highly bureaucratic compulsory licence regime, is a sure way to kill the growing voluntary licensing market.

DPIIT’s next steps will be crucial. If the goal is to create the conditions for a “made in India” AI industry while nurturing and protecting India’s valuable cultural assets there is no better solution than to create the conditions for the growth of a voluntary licensing market. Strong cultural industries and robust AI development go hand-in-hand. This means dispensing with forced, bureaucratic solutions like the proposed compulsory licensing regime while holding firm on rejecting a TDM loophole that would allow AI companies to plunder India’s cultural richness without any compensation to creators and rightsholders. Let’s hope India gets it right.

© Hugh Stephens, 2026.  All Rights Reserved

AI Training and Nurturing Cultural Industries in Asia: Finding the Right Balance

Text and Data Mining (TDM) Exceptions and Compulsory Licensing Solutions Carry Heavy Risks

A scale balancing two labeled blocks, one marked 'TDM' and the other marked '©', representing the debate between Text and Data Mining and copyright.

Text and data mining (TDM) is a hot topic in many countries. In jurisdictions where exceptions to copyright protection are embedded in legislation rather than determined by the courts on a case-by-case basis (as in the US), TDM has become a favoured vehicle of AI developers, although compulsory licensing has also been floated by some as a potential solution. AI developers see TDM as a loophole allowing access to copyright protected works for algorithm training without payment or permission. Compulsory licencing would establish a statutory regime requiring rights-holders to provide access to their content upon payment by users. While seemingly offering a middle ground, it is fraught with problems. Meanwhile, content industry stakeholders have been vocal on the need to protect their intellectual property, while in some cases resorting to legal action.

TDM has been on the front burner in the UK, Australia and Canada, and Asia is facing many of the same issues. From India to Malaysia to Japan, and from Korea to Hong Kong to Singapore, access to copyrighted content for AI training is front and centre although being played out in different ways. Some countries already have instituted limited TDM exceptions while others are reviewing options. In India, which has long used compulsory licences in the patent field, and which has provision for compulsory licences under certain narrowly specified circumstances in its Copyright Act, both TDM and wider compulsory licensing are being pushed by the AI industry. A common thread in all countries is the concern by rightsholders that their valuable proprietorial content is being or may be taken and reproduced to provide training inputs to a commercial process without authorization or compensation. These concerns are not misplaced.

Compulsory licensing is a “solution” (actually opposed by many in the AI industry who believe that all content should be “free”) that strips away the rights of content owners to determine how their valuable intellectual property will be used. In effect, it is a form of expropriation. While compulsory licences may set a price for use (which may or may not be seen as fair), they don’t address other issues that are normally included in licensing deals such as how the work is to be used, or any specific limitations related to the content. There is also the difficult issue of equitably distributing collected funds.

Voluntary licensing where rights-holders can opt-in is a fairer and more feasible solution, offering mutual benefit to both the content and AI industries. A growing voluntary licensing market exists for print, AV and music content—but AI developers have been slow to respond, a key reason being the mixed signals they are receiving from various governments. Rather than negotiate, the AI industry would rather push for a broad exemption legalizing the practice of helping themselves to protected content owned by others. The pretexts advanced are either a) they are not really copying (just turning content into data tokens is the argument) or, b) if they are, they should be allowed to continue doing so in the name of “innovation”. There is also the implicit threat that if laws and regulations are too protective of the creative sector, AI development funds will go elsewhere, to more compliant jurisdictions.

This argument does not hold water as many factors go into making investment decisions regarding facilities such as data centres, notably the availability and cost of talent, land, power, etc. It is worth noting that while Malaysia does not have a TDM exception in its copyright law (whereas Singapore does), investment is pouring into Johore Bahru–just across the causeway from Singapore–because of Malaysia’s relative competitive advantage in input costs. The AI industry’s “fear factor” threatens to start a race to the bottom as governments around the world don’t want to be left behind as the AI race heats up. While it is clear that AI will transform some industries and has the potential to increase productivity in many areas, it may lead to more job losses than gains whereas the cultural sector is both a key economic driver in all the Asian economies in question and an important pillar of national identity.

India

India is a good case in point. It is a well known cultural and technological powerhouse with a  creative economy that was estimated by WIPO to be valued at over $30 billion (USD) in 2023, with 20% growth in creative exports generating over $11 billion. Prime Minister Modi has called on the creative sector to further increase its share of GDP. Yet the TDM issue has raised its head in India, especially after OpenAI was sued by several Indian media entities for copyright infringement. In May Reuters reported that the Ministry of Commerce had set up an expert panel to examine the AI training issue. Both domestic and international content industries in India are concerned that creation of a TDM copyright exception or widening of compulsory licensing in India’s copyright law will undermine the incentive to create new content, and stall the development of a voluntary licensing market for AI training. Careless implementation of TDM or bringing in a misplaced compulsory licensing regime risks throwing out the baby with the bathwater.  

Malaysia

As in India, AI industry lobbyists in Malaysia have called for implementation of a TDM exception. The case of neighbouring Singapore is often cited, but Singapore is a particularly poor example to follow. Singapore’s overly-broad TDM exceptions, referred to locally as exceptions to facilitate “computational data analysis”, combined with severe limitations on use of contract law to control access to copyright protected works, have weakened Singapore’s creative sector and held back the development of licensing options. There is no need for introduction of a TDM exception in Malaysia. Kuala Lumpur can distinguish itself by offering an appropriate balance between AI development and fostering important cultural industries, encouraging the development of a mutually beneficial licensing market. Its other attributes have helped it to successfully attract significant high-tech investment without undermining its investment in content creation.

Japan

Japan, which has a TDM exception in its copyright law, is often held out by AI developers as a model for the kind of copyright law they would like to see replicated elsewhere, but the impression that anything goes in Japan with respect to use of copyrighted content is mistaken and based on misunderstandings. As I outlined in a blog post last year (Japan’s Text and Data Mining (TDM) Copyright Exception for AI Training: A Needed and Welcome Clarification from the Responsible Agency), Japan’s TDM exception does not apply if the user of the copyrighted data “enjoys” the content. As an example, this means that if a user derives benefit through using the copied material to create outputs based on the reproduced content, the TDM exception does not apply. As this website succinctly puts it, “Expressive intent invalidates the safe harbour.” As is the case elsewhere, the limits of the law are being tested in court. Yomiuri Shinbun, Japan’s largest paper, as well as Nikkei and Asahi Shinbun, are suing Perplexity AI for copyright infringement in Tokyo District Court. Meanwhile a market for licensing content is beginning to develop.

Korea

Korean content companies are also turning to the courts for redress against unrestricted copying by digital platforms. Korea’s three terrestrial broadcasters, KBS, MBC, and SBS filed suit in January against Korean tech giant Naver claiming the platform used their news content to train its AI application. The broadcasters had earlier put Naver on notice not to use their content without permission. Naver is, broadly speaking, the Korean version of Google. It has recently been reported that more lawsuits are pending against Naver, this time from the Korean Newspaper Association.

Korea does not have a TDM exception in its copyright law, but it has (at least in theory), adopted the US fair use doctrine as a result of the US-Korea Free Trade Agreement. However, although fair use was incorporated into Korean law in 2011, its has seldom been used and the Korean courts have been very reluctant to apply it, and where they have, the application has been very narrow, essentially limited to non-commercial use. To date there have been no fair use cases brought to the Supreme Court, and lower courts tend to rely on the specified exceptions that apply in Korean law. Because of this there have been attempts to introduce a TDM exception, and more are expected in the current National Assembly. Various versions have been proposed that are of concern to rightsholders, including broad interpretations that would not distinguish between commercial and non-commercial use. Korea, one of the cultural giants in Asia, needs to tread carefully if it wants to maintain this leading cultural export, while encouraging development of content licensing.

Hong Kong

Hong Kong does not have a TDM exception in its current copyright law but under pressure from the AI sector is considering the idea. The Intellectual Property Department launched a public consultation late last year, receiving input from stakeholders representing both sides of the argument and has come forth with recommendations to the legislature (Legco). It has proposed a TDM exception for both commercial and non-commercial use but with a number of limitations; 1) access to content must be lawful (i.e. no use of pirated content); 2) a public record must be kept of copyrighted works used in AI training (transparency requirement); 3) the TDM exception will not apply where licensing schemes (i.e. licences that have been issued by the Copyright Tribunal) exist; and 4) rightsholders can reserve their rights by opting out.

There are problems with this proposal, despite the limitations. Requiring rightsholders to opt-out stands the existing basis of copyright on its head, as it has in the EU (i.e. users normally need to obtain permission from rightsholders in advance) while the licensing provision provides limited relief.  While not as potentially destructive as some proposed TDM exceptions elsewhere, it is questionable if Hong Kong needs a TDM exception given that voluntary licensing alternatives are increasingly available. At present, the recommendations are with the Legco; given public skepticism about the proposal, legislation is not expected until 2026 at the earliest.

Conclusion

Lawmakers and regulators in Asia are grappling with a common problem; how to incentivize the development of responsible AI while continuing to encourage and promote all-important content industries. Cultural expression is particularly important in Asia as an expression of values, and throwing the cultural sector under the bus in the hopes of attracting some ephemeral hi-tech AI jobs is a false bargain. It’s like eating the seed grain from which the bounty of cultural creativity springs. Undermining the nurturing environment provided by sound copyright protection, whether through compulsory licensing or creation of TDM exceptions, is bad public policy.

Strong cultural industries enable the development of strong content licensing markets for AI development, enabling a virtuous circle of further creativity. A strong cultural sector and strong, sustainable digital industries, especially those powered by AI, go hand-in-hand. Asian regulators need to exercise prudence and weigh the consequences of rash action. The winners will be those that find the right balance between encouraging innovation and fostering creativity.

© Hugh Stephens, 2025. All Rights Reserved