Praharsh Gour

Abstract: In light of the ongoing dispute between ANI and OpenAI before the Delhi High Court, broader questions concerning the sustainability of digital journalism and small digital media outlets require greater attention. Distinguishing between two concerns, namely, the use of copyrighted works as training data and the potential displacement of publishers when AI-generated responses substitute for visits to original sources, this piece argues that while copyright may remain relevant to unauthorised training and outputs involving memorisation or substantial reproduction, it may offer limited answers to source substitution and readership displacement. The piece therefore considers alternative regulatory tools, including attribution, linking, collective bargaining, competition law, and transparency obligations, arguing for a differentiated regulatory approach to distinct harms.
I find myself in a somewhat awkward position to be writing for this Mini-Symposium on ANI v. OpenAI. When the good folks at the Law School Policy Review first reached out to me to contribute a piece on the much-talked-about Delhi High Court’s decision not to grant an interim injunction to the news agency in this case, I had a very different set of arguments in mind. Something around the interpretation of fair dealing in the context of copyright and input for training AI models, stemming from the High Court’s interim order. However, as the litigation progressed, my stance on what to write changed to something more fundamental. This largely stemmed from a conversation I had with Swaraj Barooah about the relevance of a small platform/ online publisher in the day and age of AI, where readers can just ask ChatGPT to respond to a query/ give a summary rather than engage with the analytical posts on a development.
In the whole hustle and bustle around the interim decision of the Delhi High Court, I thought that the question that deserves our attention is– what the future looks like for the small news media publishers and whether copyright is the right response to resolve the grievances that publishers are raising against generative AI (or Gen AI) models. At the outset, it is important to differentiate one thing here. The term “publishers” that I have used is a broader one which refers to entities that monetise content (either directly from the readers through subscriptions or through advertisements) by distributing it directly to their audiences; organisations like ANI operate substantially as a news-gathering and syndication organisation which licenses its feeds and content to subscribing news organisations in addition to distributing content through its own website and social media channels. This differentiation is important in light of the prima facie finding on market substitution by the Delhi High Court in the interim order.
This piece does not dismiss the copyright concerns around the works used to train these Gen AI models (regardless, see here and here for an interesting counter). Also, I agree with the fact that copyright may have a role to play in instances of memorization, regurgitation, or substantial reproduction while assessing the generated output. However, the piece deals with a parallel concern about these models’ output and the consequent readership displacement. Breaking this down further, there are two different problems here:-
- When works are collected, processed, and used as training material before generating an output, and
- When a Gen AI model produces an output where the user receives the information, they are looking for without necessarily visiting the website from which that information originated.
To explain this further, training a GenAI model, specifically a Large Language Model like GPT a corpus of data is required which may include data from publicly available sources (this is not to be confused with works available in the public domain which are free from copyright), licensed material. The material used for training may include books, articles, blogposts which are divided into tokens and converted into numerical representation which the model can process. (The working of LLM part of the ANI OpenAI order explains the process of training LLMs pretty nicely.) During pre-training, these models are trained to predict the next token based on the tokens that precede it, with the model’s parameters being iteratively adjusted based on the accuracy of its predictions. After this, the models are further fined tuned to reduce prediction errors and to follow instructions better. So, copyright law may get involved here, where material (many of which might be unauthorised and copyright protected works) are used to train a model.
Now when a user puts in a prompt for the LLM asking the model to do something, the model uses probability and statistics based on the above training to predict a sequence of likely tokens. The wording of the output could be something very different from the wording on which the model was trained on. So, let’s say when a model is trained on an article by an economist explaining the 2026 budget, and the prompt is to “explain the key points of the 2026 budget”, the model may use the input from the above article to generate/ predict an output which might not be verbatim of the above article but in essence states the same things. In this case, the output essentially substitutes the original source with its output.
The Delhi High Court’s interim order also demonstrates how different these questions are by considering the issue of training separate from the issue of ChatGPT’s output amounting to memorisation, regurgitation or substantial reproduction. Regardless, the Court addressed these issues through the lens of copyright law. My argument focuses on the second question and is slightly different. In this piece, I argue that despite all the temptation to apply copyright as the one quick fix for all the ailments, we need to broaden our horizons and look at other mechanisms that might be more viable to resolve the problem faced by the publishers.
Why is this not a Copyright Problem?
The reaction to look at copyright law as the relevant regime to offer solutions to this problem posed by Gen AI is only obvious, considering that these platforms use enormous amounts of works (both protected under copyright law and works which are in the public domain). The concern that is being flagged here is that copyrighted material is being used without adequately compensating the authors and owners of the work. To resolve this issue, one of the proposed responses has been to set up a licensing framework around access to copyrighted works for AI training. Part I of the DPIIT Working Paper on Generative AI and Copyright offers the same solution–a blanket licensing mechanism to address the use of copyrighted works as training data through a statutory licensing framework and collective royalty distribution. [Sidenote: It must be clarified here that the DPIIT paper talks about the licensing mechanism for using works only at the training/ input stage, and Part II of the Working Paper will deal with the copyright concerns at the output stage.]
But as scholars have pointed out, concerns at the output stage are a different beast altogether and cannot be tamed by copyright. Payment for training does not necessarily compensate the economic loss arising from subsequent source substitution and readership displacement. Copyright may enable the publishers to receive royalties for the use of their work in the training corpus of a Gen AI model, compensating them for the use of their copyright-protected works at the input stage. But what about the displacement of readership when a reader, instead of visiting the publisher’s website, visits a Gen AI platform or relies on a Gen AI-generated overview or answers (which is not a reproduction of the original article) to questions about the very subject on which that publisher spent money, time, and editorial resources reporting? If the platform gives the reader a satisfactory answer, the reader may never visit the publisher’s website, and the publisher loses not only a page view, but an advertising opportunity, a potential subscriber (incurring economic displacement), or simply the chance to establish a relationship with that reader. After all, copyright (and broadly IP laws) does not/ cannot/ should not guarantee a copyright owner a particular amount of readership, traffic, or revenue, since these are not the concerns that copyright was devised to address. Nor can it ordinarily confer a right to prevent competitors from communicating the same underlying facts (also see here).
This economic displacement can significantly affect journalism and future reporting, especially in areas where reporting is expensive, speciali sed, or locali sed. The broader digital environment already shows the significance of this ‘zero-click information’ consumption, and early research suggests documented reductions in traditional publisher referrals in some contexts (see, for instance, here, here). But why should Gen AI platforms be held accountable here? Without a proper regulatory structure, we are seeing a catch-22 situation where the business model of these platforms depends on information produced by an information ecosystem whose economic sustainability it may simultaneously undermine. (For more on this, interested readers can refer here, here, and here and generally see the Knowledge Futures set of posts for a more detailed discussion on these issues.) The relevant question thus is not whether Gen AI platforms have taken something from the publishers but whether the legal framework should intervene when a new intermediary can capture the value of information.
To sum it up, a platform may threaten the economic position of an entire industry without infringing the industry’s copyright, and thus, the actual question is not how we should extend copyright as a solution for this new problem. Instead, we should be looking for obligations that these AI platforms should be held to. Treating it as a copyright problem risks designing regulation around acquisition of information rather than sustainability of the information ecosystem that produces it.
Can Licensing at the Output Stage Fix the Problem?
Now let’s assume that AI models and publishers do enter into an agreement where the AI platform is supposed to provide explicit summaries, clear attributions, and links to the original article. On the face of it, this seems like a workable solution for source substitution and readership displacement. Right? Attributing to the source could preserve some of the reputational and informational value of the original publication. A reader who would receive a summary with a link to the source might be encouraged to visit it, thereby preserving the provenance. Such mechanisms that operate at the point of inference, such as requirements concerning explicit summaries, attribution, and links to the publishers’ sources, may address some of the concerns directly. And this is on top of license fees for using the publishers’ works for training the model. This shifts the discussion from getting a license at the input/ training stage to the point where source substitution and readership displacement occur. The above model could then be implemented in different forms, with publishers negotiating attribution standards or minimum linking requirements.
But then there is another difficulty– a licensing market does not necessarily produce equal bargaining power (indicatively see here and here). Big players in the market generally have a repository of valuable catalogues, extensive archives and (if push comes to shove) deep pockets to pay hefty legal fees for endless negotiations. Smaller players like regional publications, independent journalists and specialised outlets comparatively have much less leverage.
So, an obvious solution might be to look at collective bargaining mechanisms to improve the position of smaller publishers that cannot individually negotiate with large technology platforms. However, such aggregation would then raise questions about distribution, valuation, and representation such as: Who gets paid and how much? What would be the basis of this payment? Does payment correspond to the volume of material used, the popularity of a publisher, the amount of content retrieved by users, the number of citations, or the social value of the journalism? And who’ll check if these payments are fair? Should there be a FRAND-sort of mechanism (like in the case of Standard Essential Patents) to regulate these payments? And each response to these questions could have its own set of challenges that could unintentionally reinforce the existing concentration of the media market in the country. If the value of content is measured primarily through readership, citations, or negotiating leverage, the largest publishers may capture a disproportionate share of the resulting revenues and smaller publishers may receive little, even though their reporting may be particularly important to local or specialised information ecosystems.
These approaches may still require scrutiny, particularly under antitrust law and IT laws to address concerns around market power, coordination, and safe harbour provision.
Different Harms Require Different Regulatory Tools
This leads us to the idea that perhaps a fragmented approach to solve this problem might be a better one. Now, I understand that a temptation to search for one comprehensive solution might be feasible. However, as explained above, multiple problems require multiple interventions.
For instance, regarding training, we may need clarity on copyright concerns around unauthorised use and whether fair dealing provisions would be applicable at this stage; Regarding outputs, two things need to be looked at– if the response is memorized, regurgitated, or substantially reproduced, then in those cases a copyright inquiry will be useful; however, if the output is not verbatim to the source and the issue is of source substitution or readership displacement, then in that case attribution, linking, and other transparency mechanisms will be useful. For the issue of bargaining asymmetry, as discussed above, competition law might be the relevant mechanism to look at.
I understand that each of these suggestions may come with its own set of challenges; for instance, in the case of attribution, it is to be seen as a pathway to the original source, preserving the provenance, but it cannot be assumed that merely by attributing the lost economic value can be restored completely. The point that I am trying to make here is that each method tries to address a different set of harms. Trying to solve all this through a single mechanism risks ignoring the nuanced distinctions between these issues. This may eventually result in setting up a system which focuses on acquisition of information as the central problem instead of focusing on wider public access to the information and to the organisations which produce it. This does not mean that the acquisition of information bit is not a question which deserves attention, nor does it mean that with the age of AI all the publishers and media outlets will shut shop in the years to come. Rather, if the concern is copying and substantive reproduction of the source as the output, then copyright might be the appropriate place to look. However, if the concern is source substitution or readership displacement, then copyright is not the most important solution to the problem.
With an appeal pending before the Delhi High Court against the interim order in the ANI and OpenAI case, and the DPIIT working on coming up with the second part of the Working Paper, which is supposed to deal with output, India’s judicial and regulatory institutions are at a juncture where they have the opportunity to expand the scope of the problem (and consequently the solution) beyond copyright and look at the value which needs to be preserved. For individual literary work, copyright may seem like the best place to check for solutions; however, for preserving the economic sustainability of journalism and media, a diverse news ecosystem, and users’ ability to access information without an increasingly expensive chain of permissions, we may need to look beyond copyright.
Praharsh Gour is an IP lawyer currently working as an Editor and Researcher at SpicyIP. The views expressed, and any errors in this piece, are attributable to the author alone. Comments and counterarguments are welcome.
Categories: Law & Technology
