Podcast

Panel Discussion on ANI v. OpenAI: The Future of the Information Ecosystem

Sr. Adv. Diya Kapur, senior advocate at the Delhi High Court and the Supreme Court, moderates a panel of Prof. (Dr.) Arul George Scaria, court-appointed amicus curiae in the case, Prof. (Dr.) Dev Gangjee of the Oxford IP Research Centre, Adv. Madhav Khosla of Khaitan & Co, and Prof. (Dr.) Zakir T. Thomas, former Registrar of Copyright, Government of India, for a conversation on the ANI v. OpenAI judgment and its implications for India’s copyright and AI regulatory landscape.

The panel works through the four issues framed by the Delhi High Court regarding training, storage, fair dealing under Section 52, and jurisdiction, before they turn to memorization and how courts overseas have approached the question of whether a trained model can itself be a copy. The conversation moves into where the real locus of infringement lies, the judgment’s jurisdictional reach over training done outside India, and closes on the unsettled question of what counts as lawfully acquired training data.


LISTEN TO THE PANEL DISCUSSION


Shreya: Okay, so good evening everyone. My name is Shreya and this is Aastha and we’re editors at the Law School Policy Review or LSPR, which is a student-run platform dedicated to generating open access discussions on contemporary law and policy in India. And we are organizing today’s event in partnership with the Law and Technology Society or the L-Tech, which is India’s premier forum for law and technology engagement and has convened leading voices on these questions since 2007. Now over to you, Aastha.

Aastha Nayak: Thank you, Shreya. This panel discussion examines the ANI Media versus Open AI case and its implications for the future of our information ecosystem. This case raises significant questions at the intersection of copyright law, generative artificial intelligence, access to information, platform regulation, and the future of the digital information ecosystem. Our discussion today will examine the legal and policy questions emerging from the case including the implications of AI training on copyrighted works, fair dealing, the rights of news publishers, and the broader regulatory framework governing generative AI in India. Our panel today brings together leading scholars and practitioners. Now, just to give a brief overview about the speakers, first of all, we have Professor Dr. Arul George Scaria, who is an IP and competition law scholar at NLSIU, and served as a court-appointed amicus curiae in the ANI v Open AI case. His research focuses on copyright reform, open science, access to knowledge, and the intersection of IP and market competition. Hello, Professor. Our next speaker is Professor Dev Gangjee. He is a Professor of Intellectual Property Law and Director of the Oxford IP Research Centre at the University of Oxford. His research examines copyright law and the legal boundaries of generative AI training, and he regularly advises international organizations and governments on complex IP policy. Thank you for joining us, Professor. We are also delighted to have advocate Madhav Khosla, who is a partner at Khaitan and Co, specializing in AI training, data, and platform regulation disputes. His practice focuses on LLM pre-training copyright liabilities, data set licensing, and fair dealing defenses, and he has also authored Khaitan and Co’s Global AI Brief. We’re very happy that you’re here.

Madhav Khosla: Thank you for having me.

Aastha Nayak: Our next speaker is Dr. Zakir Thomas. He is the DPIIT IPR Chair Professor at NLSIU and former Registrar of Copyright, Government of India. His research evaluates copyright law and Section 52 of the Copyright Act, fair dealing provisions applied to generative AI development in India. Thank you for joining us, Professor. Now we have senior advocate Diya Kapoor practicing at the Delhi High Court as our moderator for this panel. And she also practices at the Supreme Court with expertise in digital rights, intermediary liability and IP disputes. She is also an alumna of NLSIU and the Harvard Law School. We’re really delighted to have you here.

Diya Kapur: Delighted to be here.

Aastha Nayak: Thank you all so much for joining us today. We would like to hand it over to Diya Kapur now. Thank you.

Diya Kapur: Thank you. Thank you all for putting together this incredible panel of some very learned people and I don’t know about the rest of you, but I am very excited to hear from them what they have to say. Open AI versus ANI judgment, I mean, I think everybody would have read, but just for the benefit of sort of starting out this conversation, I will take it over in a moment to Professor Scaria because he was in fact an amicus in the case. So what you’re not going to read in the judgment, he’s going to tell you. But effectively what this case decided, and raised issues on whether Open AI was in breach of the copyright of ANI. And while many issues were raised, effectively, this was at an interim stage and the only issue the court really went into was whether training of the data set was something which violated. The other two issues, which are actually the other big issues on whether once the AI stores the data, and the output that is produced by the AI, those issues have actually been left open for after trial. But I’m going to just hand it over to Professor Scaria to give you an overview of the ANI judgment.

Arul George Scaria: Thank you, Ms. Kapur. So if I may share my slides, just give me one second. Yeah. I hope the slides are visible, right? So yeah, thanks to the student team for putting up this discussion. It’s always a pleasure to be part of a panel like this. So what I would like to do in the first couple of minutes is to just briefly introduce the four issues that were discussed as part of the judgment and then maybe focus on the two core issues, particularly to discuss the question of whether this is a balanced outcome, because there has been a lot of discussions on whether this is like a one-sided judgment. So it is very important that maybe we spend a little bit of time to see how maybe the court might have attempted a balancing through this particular judgment and more importantly, whether I see some challenges with regards to this particular balance or how we address this issue of balance in the long term. So as most of you would be aware of, when the matter came before the court in November 2024, it had identified 4 core issues to be discussed. The first one was related to training. The second one was related to the outputs. The third one was relating to the applicability of exceptions. And the 4th one was relating to the jurisdiction issue. And I’m hoping that maybe the different issues that were discussed as part of the judgment will also be discussed by different panelists. As I mentioned earlier, I would like to focus on two of the core issues, i.e. issues one and three, which the court also considered as interlinked issues. Because to me, this is a crux of the, I would say, issue that we need to discuss because the issues on the output side are going to vary substantially based on the facts of each and every case that are going to come in the future. But on the training site, maybe we need to have more clarity. So issue one, as I mentioned earlier, was focusing on the training side. But when the matter came up for discussion, and also when you look at the judgment, you will notice that this issue cannot be discussed without also looking at the third issue, that is how or to what extent the exceptions are applicable with respect to this particular issue. So when it comes to the determination of issue one, you would have noticed that the judge is coming to a conclusion that the storage of a copyrighted work in any medium would at least technically amount to reproduction of the work. So one of the things which I also want to highlight here, which may not be so evident in the judgment, is that during the course of proceedings, a lot of discussion went into the question of whether there is an independent right of storage existing under our copyright law. Of course, people like me are aware of the view that maybe we should read it in the broader context of the right of reproduction. But a lot of energy and time was actually spent on deliberating whether there is an independent storage, right?

Diya Kapur: Dr. Scaria, can I just pause you here for a moment? I mean, just to sort of, you know, contextualize this. Under the Copyright Act, the Section 14, right, which is under Section 14, we’re sort of talking about whether there was an infringement of copyright or not. And the question before the court was under 14(1), whether you know storage in an electronic medium, the mere storage amounted to a copyright infringement. And I guess the question before the court was when Open AI is using all the data that it is taking to train its model, whether it is actually storing that data while training that model, and therefore whether there’s an infringement. And I think that OpenAI actually made an admission that they are in fact storing their data in order to train.

Arul George Scaria: Yes, and also one of the things which we need to very specifically observe in 14(a)(1)is that it doesn’t make a differentiation between temporary and permanent storage. So it is an undisputed matter that if you have to train, at least there should be a temporary storage of the material. So that’s a broader context in which the court had to address this particular issue. So that’s a very relevant point and thanks for highlighting that dimension here. But what is quite interesting here is that yes, once a judge is taking the position that storage amounts to reproduction of the work, that is, that is technically a potential violation of 14(a)(1). The judge is also then immediately saying that things don’t end there. We have to also look at section 52. To me, yes that precisely is what a judge needs to do when it encounters an issue like this, because you cannot treat them as separated things. You have to read them as a whole. So exceptions are an integral part of the analysis of infringement, because the moment someone is able to show that the concerned activity falls within the ambit of section 52, then the presumption is that there is no infringement at all. So that analysis had to happen together and that’s definitely happening in this particular case. And when it comes to issue 3, that is the fair dealing analysis. Again, you will notice that the court rightfully engages in that two-step analysis, wherein the first step focuses on the question of whether it is for one of the specific purposes mentioned under Section 52(1)(a). And at the next stage, the court is engaging in what is only known as a fairness test. And when it comes to that first step, the court is taking a liberal and purposely interpretation there. The court is looking at the term private use because if you look at Section 52(1)(a), you will notice that both private use and personal use are mentioned. And to the best of my knowledge, this is for the first time a judge is clarifying what could be the scope of private use. And in the context of this judgment, you will notice that since the training is happening in a closed space, and since the materials are available and used only for the purpose of training, the judge is of the or taking the view that the private use that purpose related requirement is met. You would also observe the, yeah.

Arul George Scaria: Yeah, the two important areas wherein, the judge is engaging in a very liberal and purposely interpretation are with respect to the terms private use and research. You will see these deliberations in the paragraphs which I have specifically mentioned here. It is also important to add here that the court also can be seen deliberating on two of the other very contentious issues that is, are commercial users allowed under Section 52(1)(a) and also the question of whether there should be lawful access to exerciser privileges that are provided under Section 52(1)(a). There also the court can be seen observing that just because something is a commercial use, it doesn’t automatically go outside the ambit of fair dealing exception. The judge also can be seen pointing out that wherever lawful access was intended in our statute, the court has specifically clarified it. Most importantly, when it came to the fairness analysis, the court can also be seen asking 3 questions. And in my opinion, these three questions are also the ones that contribute to the balancing. So when it comes to, say, for example, the first question, you will notice that the court observes that it was only used for the limited purpose of training, that the copyrighted materials were only used for the limited purpose of training. When it comes to the second, question, which again, I think is extremely important – the court is looking at the question of whether there is economic competition and whether this is negatively prejudicing the legitimate interest of the plaintiff. The court is reaching the conclusion that there is no such harm happening in this particular case. And on the third question- whether there is public interest promoted, you will also notice that the court is acknowledging the role of LLMs in promoting scientific knowledge and information. So these Three questions contributed substantially in terms of balancing the interest of both the sides. And that’s something which I am emphasizing because in a different set of factual circumstances -for example, if you are talking about a music service, it is quite possible that a court might reach a very different conclusion. For example, if the concerned platform acts as a competitor for the copyright holder’s work, it is quite possible that the finding with respect to fair dealing might reach a very different conclusion. So to summarize, what I am trying to convey here is that if you look at how the judge has approached the issues in this case, at least in my opinion, there is a strong attempt to strike a fair balance between the interests of both the creators as well as users of copyrighted works. But to me, the bigger challenge is retaining that fair balance. And there we might have to look beyond infringement related issues. One of the very specific things which I want to point out is that when we are taking a liberal approach with respect to the use of copyrighted materials for the purpose of training, maybe we should also complement it with a stricter approach with respect to how we are going to deal with AI generated outputs. I’m emphasizing on that because just two days back, the Indian Copyright Office issued the order in the DABUS case, wherein you will notice that, of course, in that case, the concerned applicant had mentioned DABUS as the order. For that reason, the court, the Copyright Office has rejected the application. But in all the other parts, it is indicating that maybe India wants to take a fairly liberal approach with respect to grant of copyright protection with respect to AI generated works. To me, that is a dangerous part because that is going to considerably upset the fair balance which we might have seen through the ANI v OpenAI case. The other three things also which I’ll briefly mention and maybe we’ll get more chance to discuss that as part of the deliberations today. Maybe we should also have stronger transparency and disclosure mandates with respect to AI generated outputs. We are already seeing some initiatives in this regard in other parts of the globe. And maybe a better clarification is required in section 14(a)(i) because 14(a)(i) is talking about storing, but maybe a clarification should be made there which says that storing for non-expressive uses. or storing for the purpose of AI training does not constitute infringement. The problem currently is that fair dealing analysis can be extremely subjective and there wouldn’t be sufficient ex-ante clarity for all the stakeholders. So if you make a clarification in Section 14(a)(i), maybe we’ll be able to achieve substantial extended clarity for all the stakeholders. So with that, maybe I will close my introductory remarks and I’ll be really happy to take up this during the rest of the discussion.

Diya Kapur: Thank you very much, sir. To just summarize one last point that was made, which is on striking that balance. And I think what you said was that in order to strike that balance, perhaps we need to strike that balance with the way we regulate the output. I’ll come to some of the other speakers to discuss the output in a moment. But before we get there, between the input and the output, there’s somewhere in between.

Diya Kapur: And I feel like for me that Section 14, the way the court has dealt with Section 14 by saying that there is copyright infringement because there’s storage. The question that I want to ask the panelists and specifically to Professor Gangjee, since you know he’s dealt with this in some other context – when Section 14 of the Copyright Act says that storage itself is a copyright infringement, and then it goes to 52 for a fair use for training, if an AI system is actually storing that data, even post-training, in order to actually answer the question and retrieve and perform its function, then I think we have a problem which is not going to be saved by Section 52 fair dealing. So the question for you, Professor Gangjee, is how do, and I’m going to come to you also, Madhav, for the same thing -on the technicality of how the LLMs function. Do they actually store data after training? What is that memorization, as they call it, that happens by the AI?

Dev Gangjee: Very happy to pick up on that. In fact, I have a few slides which will directly go to that question. So, Miss Kapur’s question ‘Distinguish between the training of the model, and then whether the trained model actually stores the information’. And that’s the unresolved question. So I want to spend the next few minutes talking about that question I was directly asked. But just to go back to Professor Scaria’s presentation, you need to make lots of copies when you’re training a model. You need to make lots of copies to create data sets, more specialized sub data sets, higher resolution versions of images, so you scale them up, you scale them down, you tokenize them and you break them up into smaller parts. All of these are creating copies or parts of copies. And one of the arguments we heard in ANI was, well, all of this copy making activity is not usual for copyright law because it’s non-expressive. It’s not people sitting down and reading books and listening to musical works and copying them for those purposes. It’s copying to learn. And copyright law should not be getting in the way of copying to learn. But the court disagreed. And the court went with many other courts around the world and said, any copying is copying, which is copyright infringement. That’s the end of the story. The pressure then shifts to whether a defense lets that kind of copying off the hook. In the EU, you have the text and data mining exception, which allows you to make copies for data mining purposes and do this within certain constraints. In the US, you have the flexible four-factor fair use test, where already courts in California have been saying, if you’re copying for this, yes, a copy is an infringing copy. but we’ll allow it under the defense of fair use. So that first question seems to have a consensual answer emerging. Any copy is likely to be infringing, and we need to turn to defenses to save the situation, as the Delhi High Court did. The harder question is whether the trained copy retains the data in some recognizable form, which we might also call storage, and that takes us to memorization. And how do we know it happens? Well, when you prod or you prompt the model, it regurgitates the copy, or something very close to a copy of the original data, or something that is substantially similar as a copy of the original data. And that’s the bit I want to dig into, which sets up… I suspect Madhav’s contribution which follows. So what is or isn’t memorization? Just being exposed to the data and the data set is not memorization. In fact, that first left-hand side of the slide, the first point there, is really just a repeat of the point we discussed. You’re making lots of copies to train the model, so having the copies in the data set is not memorization. The second point is a little more nuanced. Does training the model involve exposure to the works and therefore some form of copying? And again, the answer, at least from a technology perspective, seems to be quite clear. That’s not copying because you’re training the model to recognize general patterns, and I’ll give you some examples in a minute. You’re not training it to memorize the works like a parrot and reproduce them. You’re training the model to learn what dogs or cats look like and reproduce basic images of dogs or cats and not remember a particular dog or a particular cat. So the third point on the slide really is where memorization begins. Has the information specific to the work’s expression, not the facts, not the general ideas behind the work, But the copyright protectable expression become encoded in the parameters of the model, the weights, the parameters, the linkages of the model, such that it can be reconstructed by the trained model itself without access to the original or access to the internet where it can get access to the original, so in the closed off sealed away model. Somewhere in the guts of the model itself, does it remember enough to recreate the copy? That’s memorization. Now, this is a really important issue, because if the answer is yes, then it taints the model. The model is potentially carrying all of these copies of images with it. And anytime the model is downloaded, anytime the model crosses borders, anytime the model is given to the public to use, that’s potentially infringing. So the model itself is tainted if memorization is a thing. What is stored in the model? There’s no literal storage of files in any region. It’s not like some giant database with little pigeonholes which actually stores these images or these texts or these songs. The information is instead stored as a configuration of interrelated weights, a probabilistic network of distribution of probabilities, If you’re talking in the fantasy, children’s fantasy context, and your word is Harry, the next word statistically is likely to be Potter. Those are the kinds of probabilities we’re talking about, and that’s how the model stores information. But we also know that numbers can encode information, and we know from the history of computer science that file storage has been done through increasingly compressing images and using mathematical algorithms to compress images. So you’re not storing it pixel for pixel, but you’re storing it in some mathematical compression formula to save space. So A compressed version is still a copy in some senses. So is this, does the model really contain all of these copies in this indirect way? Is the legal question. A practical sort of reminder for us before that, models don’t aim to memorize. That’s not what they’re in the business of doing. Models aim for generalization. They’re trained to try and answer general queries for things that they haven’t even been trained on to respond to new situations. So the model might learn, for example, if you look at the left of the slide, Many stories begin with once upon a time. And so therefore the model will trot out once upon a time in a suggestion for a story, because that’s how many stories start. Or the model learns that Paris co-occurs with France so much that it’s probably the capital of France. Or many detective stories start with a crime, investigation, clues that help, red herrings that lead you away. and then the revelation at the end. But memorization of particular examples, so it takes the examples and generalizes from them, but sometimes it gets stuck on the particular examples themselves and the specific expressive information is then locked into the model. That can happen when it’s exposed to the same idiosyncratic example many, many, many times over. Or if it’s the only example of its kind and someone asks a query and that’s the only example a model has, then the probabilities will lead it back to recreating that as well. So memorization can happen, but models aren’t designed to reproduce by memorization. Here’s just a brief example. If you look at the two images on the left, that’s an example where the model’s got it wrong. It’s memorized. That’s a mistake. In the training set is the clearer image on the left of the lady, Ann Graham Lotz. And the model, when prompted, can generate the image on the right, which is a pretty close copy of it. But what the model is actually trying to be trained to do is take general ideas and develop them. So the general idea of Darth Vader having a bike accident in a place called Telford is what the model is supposed to be generating. and is not supposed to be generating the memorized image you can see there. How have different judgments around the world approached the idea of memorization? If you look at the UK High Court’s decision in Getty Images versus Stability, there’s no direct argument for memorization in that judgment, because when they’re trying to argue that the training of the model is infringing the first stage, which we’ve already spoken about, making lots of copies to train the model, the defendant’s stable diffusion’s answer is, but we train the model in the US, and the US is a different jurisdiction, and you can’t blame us for copyright infringement for what happens in the US. So the claimant is forced to then argue that in allowing the model to be downloaded in the UK, the model contains infringing copies inside it. in it and therefore downloading the model itself is downloading an infringing copy. And really this is an importing, infringing, copyright infringing materials provision. It was drafted back in the old days of people importing books and importing DVDs and importing CDs for those of you who like a bit of history in your lives. So that is really the problem with this. It was designed for physical things like books with chapters being imported, and now it’s being applied to a digital model and being asked really to approach memorization indirectly. So the claimant Getty argued that importing an infringing copy means if the making of the copy is infringing, and you trained it in the US with lots of access to unauthorized copies, then the model is tainted and the model is infringing. But stability argued, for it to be an infringing importation, it has to be an infringing copy. That means the model as a whole has to contain copies, and that’s simply not possible because the model holds mathematical statistical representations. So the judge agreed. with the defendant stability here and said, yes, in principle, memorization is possible. But if you actually look at the model, the model is a mathematical representation of generalized patterns it’s learned. And therefore, you cannot say that this generalized representation of the materials it’s been trained on is something we can fairly call a copy. There are some arguments to say it may not be a direct copy, it may be a derivative work, it may be some kind of distant copy, and that’s in other litigation. But the court here just rejects the argument, saying it’s too much of A stretch to say the train model is a copy of any of the individual works that went into it. And one of the reasons why the court may be going in this direction is None of the images that the claimant produced to say that the model has memorized things were very close copies at all. If you look at the synthetic output on the right and the original image on the left, you can see it’s in the same zone, but maybe in the same zone as ideas or scene software, or not a specific direct copy enough to say. We know the synthetic image has been taken from the original image. It could have been taken from other images of the football manager holding up the t-shirt and a general idea behind it. So memorization is indirectly rejected in Getty. It’s also rejected in the ANI versus OpenAI judgment for a simple technical reason. Models are trained in iterative stages, and at a certain endpoint, you have a particular version of the model. And when the claimant ANI said its news stories were used in the training data, Open AI could simply point to the version of the trained model and say, these news stories hadn’t come out by then as yet because our model stopped training in April 2024. And the news stories were all published in August and September, so they were simply not part of the training data set. So we don’t have to argue the hard question of how memorization happens and whether memorization is a copy. We simply know these weren’t in the training data, they weren’t memorized, end of story. And to the extent that there is similarity, as Professor Scaria has pointed out, that was because of retrieval augmented generation. The model went out onto the internet, found the stories, and brought bits back. But the bits it brought back were quite factual because these are new stories, and those factual bits don’t deserve protection. So it didn’t have to answer memorization, because the stories weren’t in the training data, they couldn’t have been memorized. But a case where memorization is established, and I’ll end with this, is 2 cases from Germany. GEMA is the claimant in both of them, which is the German musical collecting society, representing lyricists and authors, creators of music. Here it takes 6 musical works including Mambo Number Five, for those of you with sort of good musical memories, and it says that the defendant here is Suno, and as many of you know, Suno is a text prompt to musical output generative AI generator, and it says, “Look, we think your model stores our music”, because you can submit very easy prompts and get melodies and music back out of it. What are the very easy prompts we submit? We submit the title of the song, we submit the lyrics of the song, and we ask you through your generative AI to reproduce something. And what you reproduce is pretty much the melody, the rhythm, the musical aspects of the song itself. So you’ve learned our songs. How do they do this? They don’t get into the complex statistics and mathematics of memorization. They simply say, we know our works were in your training data, and we’re going to compare what the model outputs with our original works, the melody of the music in both cases. And we are clearly on strong evidential ground when we say there is a substantial similarity, if not an identity, between the melody of this song and the melody of our work. And this similarity is not explained by random chance, and it’s not explained by clever prompting, because we give very simple prompts. Here’s the title of a song, here’s the lyrics of the song. Now tell us what the music should sound like, and your system gives us back the music of the original song. That means it must be memorized. So really, they’ve got a simplified model here saying, if our copyrighted work is in the input, and if we can very easily prompt it to get it in the output without adversarial or tricky prompting, then it must be memorized. So they’re using the clever technique of a legal presumption to say that we have the input. We have a very easy way of getting it in the output. It must live in the model. Therefore, it must be a copy. It must be stored. And even the model itself satisfies infringement. Why does this matter? Because if you can prove it in this way, it means all models are potentially tainted unless models build in filters at the output stage. So even if it’s memorized and lives in the model, it’s going to be blocked from being shown to you, users and prompters. And unless that direction opens up, you could very easily argue that all models are tainted and you don’t need to prove the mathematics of storage within the model. You can just say it was in the input, it’s in the output, it must live in there somewhere, and you have infringement as a result of it. So that’s my sense of why this is a potentially powerful set of claims. Thank you.

Diya Kapur: Yeah, thank you so much, Dr. Gangjee. That was extremely helpful, at least for me, because I have been struggling with this one question ever since the Open AI judgment where they basically kicked that down to the end of the road to say, look, that’ll be a matter of trial. We’ll see ultimately whether it stores it or it doesn’t. We don’t know that yet. We don’t have enough prima facie evidence. But the moment, you know, the courts start to decide that there is evidence of storage in the model, then we’re stuck because we don’t have a S. 52 exception and, S. 14 straight away on storage brings you in. So unless we have some legislative change, which sort of, brings about to actually facilitate AI and says that, look, in our public interest exceptions, which is really what S. 52 was about, in a pre-AI world, they came up with a list of public interest exceptions, and you know, possibly makes a public interest exception in S. 52, we have a problem, right? And Madhav, to come to you on the technical aspects. I think what Dr. Gangjee pointed out to us is that we don’t really need to get into the technical aspect of how does an LLM actually functions because, when you do input output, you simply see that it must have memorized. But if you can add some color to anything on the technical functioning of the LLM that we do know that it memorizes or it doesn’t memorize, just as a matter of technicality, it would be helpful.

Madhav Khosla: Okay, great. Thank you, Ms. Kapur. And also I just wanted to, before we start, say, as my wife points out, that you have very narrowly saved us from being a manel for which we, I think, are all grateful. But just if we can take it back a step on what are the technical aspects, right? As Professor Gangjee was saying, The model is a set of statistical parameters. The parameters, the weights of the model reflect the patterns that the LLM learns from the whole of the training data. a statistic is by its very nature removed from the expression itself. And it’s You can never say, for example, that you know that if this is the gender ratio of a particular class at NLS that derives from any of the particular students in that class, it derives from a whole. So, for those very reasons, the parameters themselves, the statistical information that the LLM then uses to generate outcomes. that those are themselves infringing is an argument that has been made both ways. And I think I will really come back to something that Professor Gangjee said is that this is potentially infringed, that the model may then produce output that infringes. So basically, in essence, you’re saying because the output is infringing, allegedly, then in that case, the model itself becomes infringed. But, to my mind, that’s the same logic as saying that a human being who has memorized a work can always create a handwritten copy of something, can always type something out on a computer, and then you get a copy which is infringed. Again, I think this neatly ties back to what Professor Scaria was saying, is that if the control is on the output, because what you are concerned with in a case where you manage to prove memorization, is that there will then be infringing output. You’re not concerned with the rest of the LLM. You’re not saying, and in Indian law, where you’re really concerned with an injunction, saying that in this one situation, there is a work that has been infringed. That brings up, I think, questions of proportionality, but also questions of attribution. What is the actual infringement we’re talking about? Because if the model itself is a set of disassociated parameters that do not store the actual particular expression then by saying that the model memorized and is therefore infringing is not really accurate. You’re saying the model has memorized and they may therefore potentially infringe on the output.

Diya Kapur: I mean, just to pause you there for a moment, Madhav. So effectively, we had identified three issues, one is at the training stage, the other is the storage of the trained model, and then is the output. And I think what you’re saying is that that midpoint becomes irrelevant if the output that is being produced is not infringing. So effectively, people, as you know, plaintiffs or claimants, would really only be concerned with, is my data being used to train? And in the output, do you see anything infringing? They don’t really care about what’s happening in the middle. Is that what you’re saying? Because coming back to what, you know, Professor Gangjee was saying, which is that let’s then have filters at the output stage. And again, going back to, you know, what Professor Scaria was saying, which is that how we balance it is by really controlling at the output level. So, I mean, I think, Madhav, you’re making the same point to say that, look, that’s really what’s relevant. We don’t really care about the middle. So maybe you can tell us more about how we can then balance the output.

Madhav Khosla: The court is saying if two individuals had produced the same output, would we be treating this as infringement? If the example that we used, that sadly, you know, like I’m always really happy when something I’ve written makes up its way for judgment, but we’d call it a two-journalist test and said that if two journalists… are making, had written one in column A and one in column B, would you call this infringement? And I think the court really uses that same standard. Is there substantial similarity in the same way that if two human beings had produced this output? And if there isn’t, then it isn’t an infringement. And you would apply the same test to my mind. In the context of music, would you apply it in the context of video, the context of images? If those aren’t infringements to that standard, does it warrant protection? Which, if you look at the judgment in Anthropic in the California District Court’s view, protection of that middle layer isn’t really that against potential competition from works that are not sharing the same particular expression isn’t what copyright law is trying to protect. The Anthropic analysis of fair use looks at it and says on the economic harm limb. The court says, is this economic harm of the nature that copyright law is concerned? You are saying this model may tomorrow create works that do not actually infringe on your expression. But they compete with you nonetheless. I use my LLM, which has studied your work, to create 20 new detective novels. Now, somebody may say that the fact that there are now 20 instead of 1 or 25 detective novels instead of five means I now have a 4% market share, not a 20% market share. and that harms me economically. But the Anthropic court’s view was that that is not the sort of harm that copyright law is meant to protect you against. It’s not meant to protect you against competition. It’s meant to protect you against the use of your particular expression. So if training itself does not harm you, the harm really comes when the output is infringing. So that’s the same reason. My view is that when you’re looking at memorization: memorization is saying that I’m concerned that they may be potential infringement of my work. But the same, if the court is applying the test that memorization is proved by regurgitation, then what you’re really concerned with is output infringement. And it’s a sort of convenient bypass of the fact that how do I determine how much infringement has actually occurred? What is the scale of this? What harms are actually occurring from the output to say, well, the model is infringed? I think it sort of skips the hard question and takes an easy bypass to it. That’s my view.

Diya Kapur: Dr. Thomas, would you like to weigh in on this?

Dr Zakir Thomas: Thank you very much, NLS team, for organizing this and inviting us to this. Now, I’m a professor who teaches this, you know, case in a classroom. Any classroom you walk in today, students are very keen to know about copyright and AI. The first thing which I, you know, will encounter when I meet these bright students at, you know, at our law school or law schools across the country is when I put across to them the issues that have been framed by the court. I’ll quickly read it whether storage of the plaintiff’s data for training, its software- ChatGPT, would amount to infringement of plaintiff’s copyright. Whether use by the defendants of plaintiff’s copyrighted data in order to generate responses for its users would amount to infringement of plaintiff’s copyright. Whether the defendants use of plaintiffs copyrighted data qualifies for fair use in terms of Section 52. Now, the first thing which we teach kids in a copyright class is copyright protects work, patents protects invention, trademarks deal with mass. And here the issues are framed. Plaintiffs copyrighted data? How am I going to explain this to the kids? The smart kids will fry me up in the class. Let me not weigh in too much on this. I’ll leave a few questions over here, then provide any answers. The panelists before me are much more serious academicians. Professor Arun just pointed out the diverse case, where the the registrar of copyright almost has decided that an output is copyrightable and the owner of copyright is the one who caused that output to be generated. Now, I’m just taking, since this question was discussed, by Madhav also, the example which was given here. The ANI made the model to generate the response of Neeraj Chopra’s mother, which was translated and produced it. Now, apply the copyright registrar’s logic. Who caused the work to be generated? ANI. Who is actually the owner of the speech which was made? It is Neeraj’s mother. So will Neeraj’s mother now be able to sue ANI for copyright infringement? Because the natural logic that follows from what the copyright registrar’s decision is going to be. Now, another issue, which I’m sure my brilliant students are going to fry me up in the class, is about jurisdiction. Which is Berne Convention. Right, Article 5, lex loci protectionis. Where infringement occurs, the place where infringement occurs is where the action has to be taken. That’s why we apply the principle of national treatment. Now, there are two aspects here. One is, Dave mentioned and Madhav also mentioned about the output, which is coming, infringing output. Infringing output happens in India. Getty Images also was about the infringing output that happens in the UK and they give a decision there. I have no issue. I can explain to a student on that aspect why probably the court may exercise jurisdiction in such cases. Now, there is the other part of it, the storage part of it. It is undisputed that the storage: temporary, transient, the storage of the work for training has happened in servers located in the United States. Now, what are we doing? We are applying Section 52 of the Indian copyright law onto an event, an infringement that has happened in the jurisdiction of the United States. Now, tomorrow, if my bright students ask me, sir, our movies, music gets infringed in the United States. Well, can I move Delhi High Court? And just the other day, again, on jurisdiction itself, luckily, there was a decision in the case of Hindustan Unilever Limited, the Quick Living Private Limited, where the jurisdiction matter in the Delhi High Court has been referred to a Division Bench. But for example, like Bangladesh, and our state of Bengal, our neighboring state, they are the borders, the same language, books, both sides gets read – to some extent, you can pirate it. And tomorrow, if a Bangladeshi court gives a decision following the logic of Delhi High Court, then what? Now, let me conclude by asking one question which Madhav just pointed out, which is Indian courts deal with injunctions. Okay, now my kids are going to ask me, what will you inject? ChatGPT? On the same line, you inject Cloud, Gemini, Llama, whatever AI tools are – within the borders of this country, right? So where does that take us? These are some questions which I’m sure that we will deliberate in the classroom, but for now, I rest in this real pressure to be in a panel like this.

Madhav Khosla: Thank you. Actually, I’m sorry to jump in, but I actually would like to hear Ms. Kapur’s view, because I think of the questions that this case deals with, I think the court is very careful to be confined itself to the facts of this case, except to my mind, this question, the question of whether the Copyright Act applies to training occurring outside of India, seems a little bit more colored by how this would affect other infringement cases. If it’s online, does it really happen in India? Because the court’s chain of reasoning is, well, the act of copying starts from India and finishes there. And that, you know, we don’t really get into the technical aspects of how copying occurs on a server, but that’s equally applicable in intermediary liability cases, in online infringement, in all of these cases where you’re talking about shopping, and as someone who pays in a lot of these IP cases, that’d be very curious to hear your view.

Diya Kapur: Yeah, I think this actually a very, very interesting and important issue, because what the ANI judgment takes by way of jurisdiction is they say, well, you’ve taken the data from here. They find jurisdiction by saying, well, the ANI data was here, you’ve taken that data from here, right? And that data has therefore been copied. You’re right, they don’t get into the aspect of where does this reside. And I think this jurisdiction point becomes important because while we have held over here that look, using for training purposes is actually accepted by fair dealing, where the server resides actually becomes important from the storage point of view, right? The whole issue of memorization and that the AI model is infringing, will actually then depend on where that server is sitting, where is that storage happening. I think Professor Gangjee brought up that there is a server, I mean, there’s an exception in Europe, which allows you to store, right? And I think Dr. Thomas also said, you know, what are we talking about? How is this copyright infringement when you’re just storing data at the end of the day? So given that our 14 and the way ANI has interpreted it and said that, look, you know, mere storage is actually an infringement of copyright. If that server is not sitting in India, that middle layer is completely protected. And so we’ve got the first layer protected by fair dealing. The middle layer gets protected depending on where the server is sitting and those laws. And output, I mean, just to sort of raise the issue of output, is isn’t the output would be governed, I would imagine, by traditional copyright laws, right? I don’t see anything different in an AI output versus a non-AI output, right? You know, is it substantially similar? Is it an exact copy? Does it have the same tonality? Those same tests that would apply generally for any copyright infringement would apply to the AI output.

Madhav Khosla: Yep.

Diya Kapur: So we’re actually then not looking at, you know, needing any new regime really in India, as long as the servers are not sitting in India. But if the servers are sitting in India, then possibly we do need a legislative change. Am I right?

Dev Gangjee: If I could just, I was just wondering if I could jump in on that. And my sense of the worrying aspect of this judgment in terms of the long armed jurisdiction it sets up is the argument seems to be, it goes back to Madhav’s chain point and a point Arul made as well, where the original, the news stories were taken from India. And that’s the start of the chain. And then you can track them anywhere. And so long as the output eventually comes back to India, then the copying of those new stories to train the model, which should have been, according to the UK Getty Images logic, extra jurisdiction, and therefore out of the scope of infringement, come within the scope of infringement. So really, The making of copies is infringing is what the Delhi High Court holds, even if the making of copies happen over in the US, because the copies originated from India, and that’s the worrying bit.

Dr Zakir Thomas: They, they, I have a question on that, the, the, the see what is it that is happening there, like Indian, say, for example, and a server, let’s imagine that’s in India, they’re only accessing. On linking a website and accessing an information that’s available on the website, is that is that infringement? That’s what’s what what what what the your server is doing, right? They’re only accessing it, they’re not coming in like scrapping from India, that’s not how the internet works, right?

Diya Kapur: And I think that was Madhav’s point, which is that, you know, you’re accessing it anywhere. How does the Indian jurisdiction come just because it’s created here? The jurisdiction should come based on where it’s being accessed. And I think that’s actually a very interesting and valid point.

Dev Gangjee: I agree. I think, in fact, what is becoming more clear to me as I’m discussing this with all of you is I initially thought the starting point of the chain was the story accessed in India, i.e. on an Indian server. But I wonder if the judge’s conception is even more upstream, saying the story was created and recognized under Indian copyright law. And therefore, it’s got an Indian anchor, and that is the starting point for this chain, and therefore, no matter where it goes, we’re going to claim jurisdiction over the copying.

Madhav Khosla: I mean, I would hope it doesn’t go as far upstream as that. I would hope that you’re really talking about the fact that it’s hosted from an Indian server and therefore, when the copy is made, it is made in some sense from an Indian server. But I don’t think that logic, while

Diya Kapur: I think so.

Madhav Khosla: appealing doesn’t apply, I think, technically, but I think it is. One hopes it doesn’t go to the fact that an Indian author’s work is always protectable in India.

Dev Gangjee: I share the same concerns, yeah.

Diya Kapur: Before we open it up to questions from the audience, I have one question for all of you. On the test of lawfully acquired, because I know some jurisdictions have actually, you know, taken that, but we haven’t. So what are your thoughts on whether, you know, for training purposes, the data ought to be lawfully acquired? Should it be licensed if it’s behind a paywall, et cetera? Can you use pirated? data to train models, you know, the concept of lawfully acquired. Just quick thoughts on that and then we’ll open it up.

Dev Gangjee: Happy to briefly jump in and point people to a resource. The European Copyright Society a couple of weeks ago just published a fantastic opinion paper on what lawfully acquired should mean. And they made a very simple and clear point. They said lawfully acquired doesn’t have to mean acquired with permission of the copyright holder. They said it could be because you get it via a defense. They said it could be because it’s available on the internet and there’s a presumption of an implied license. to be able to use it. So there’s a whole spectrum of uses between absolutely free public domain works and works where you must have the permission of the copyright owner. And there are lots of different access points which could still fall within lawfully acquired that don’t necessarily need the permission of the copyright holder. So it’s a fantastic resource that draws a map for all of this.

Diya Kapur: Super, because this is something I think gay and ice skirts.

Arul George Scaria: Thank you there for bringing in that. And I just wanted to add to that, that many people have very different conceptions of what lawful access means in the context of a, if people are demanding that there should be permission or there should be a license, to constitute local access, then we can forget about AI development in India. I don’t think any firm in India is going to be able to take census for all the works that are required for training any large language model. So I’m happy that the court rejected that argument and said that wherever. the framers of the statute intended lawful access, they have clarified it in the statute. So I’m happy with what the Delhi Court has done in this case with respect to lawful access.

Madhav Khosla: Actually, if I can just jump in there, I think one of the interesting things is that the court also dealt with an argument that not just of lawful access, but that… One, the explanation to 52(1)(a) would mean that if somebody did not have permission when they made the copy, the 52(1)(a) defense would not be available to them. But that argument is, as the court correctly notes, circular that if you say that you must have permission in order to make a copy to avail of a fair dealing defense, then no one can ever have a fair deal. So I think these sorts of, it’s important to consider that that balance lies in leaving all of these options open.

Shauryaveer Chaudhry: Thing VM, I just got disconnected.

Arul George Scaria: You’re muted person, Zakir.

Madhav Khosla: Professor Thomas, I think you are on mute.

Dr Zakir Thomas: Sorry. See, like lawful access, even within the permitted uses in Section 52, it is like that becomes lawful access is what you know, probably European Union’s standards, and that sounds interesting, but you know, there again, I just wanted to point out, like, here we are. making 52 almost absolute. That’s what, you know, it’s not dependent on 51. That’s the kind of argument that has probably come up in the court, which is actually a development from our damasory photocopy case that, you know, even that is Section 52, then an absolute right. Or an exception? You know, that is another question that leaves open over here.

Diya Kapur: Wonderful. Can we now see if there are any questions? Aastha, I’ll give it over to you to see if there are any questions from the audience.

Aastha Nayak: Yes, thank you all so much. We’ve opened the chat.

Diya Kapur: Because we have a super panel and I know we’re right at the end of time. Maybe we have time for just one or two questions.

Aastha Nayak: Yes, we’ve opened the chat, so I invite anybody who has a question to type it in the chat, please.

Diya Kapur: I think we have time literally for one question. 30. So if there’s any one person that wants to just ask a question very quickly, Aastha, you can ask on their behalf.

Shauryaveer Chaudhry: Let me open the chat, so whoever has a question, I guess they put it there, else we can, we give it a minute. If nobody can type anything by that time, we can just control. Okay, we got a question. I’ll read it out for the benefit of the transcription data. Okay, we got multiple, but I lost the first one that came in the interest of fairness. Does the physical destruction unlicensed digital ingestion of copyrighted books by AI developers violate core copyright principles and creative ethics, and should physical literature be legally protected as non-renewable cultural heritage?

Dev Gangjee: I’m happy to jump in briefly. I think this takes us to a very important question that’s outside of copyright law. So books are destroyed in the process of digitally scanning them in high volumes. And that’s what the question is asking about. And in recent years, in recent months, secondhand booksellers have been saying lots of their stock is being bought up, which is great. but including some very valuable volumes where the physical copy is then lost forever. So I think this is a hard question for copyright law to answer. It’s more of a question for cultural heritage law and heritage institutions to answer, but I see the problem and it’s a real one.

Dr Zakir Thomas: But they’ve they’re the principle of doctrine of first sale exotion comes in right in physical books. So it brings it back within the copyright law and it takes it out of any infringing action within the copyright term.

Dev Gangjee: I think you’re absolutely right, Professor Thomas. I think this is an activity that is technically legal. You bought the book, it’s your property. You can do what you want with it because of exhaustion. But it feels immoral. And I think that’s the distinction. You’re destroying A valuable book.

Dr Zakir Thomas: Yes.

Diya Kapur: Wonderful. Thank you so much, panelists. This was an extremely engaging discussion, at least for me. I learned a lot. And my big take away at the end of it is this jurisdiction issue is going to re-agitate itself. And while one was very excited about the way OpenAI had kind of, you know, settled the training model, I think that, you know, if we get into the question of jurisdiction, we don’t know whether the court’s judgment holds at all or not. So we’ll be wide open all over again. I’m looking forward to seeing Madhav in court agitating this issue and see where it takes us.

Madhav Khosla: Thank you so much. I really appreciate the opportunity.

Dev Gangjee: Thank you for hosting us, and yeah.

Dr Zakir Thomas: Thank you.

Shauryaveer Chaudhry: I’d like to invite the co-convener of L-Tech Laavanya to kind of give the vote of thanks.

Arul George Scaria: Thanks.

Laavanya: Hello, everyone. First of all. I’ll introduce myself. I’m Laavanya and I’m the co-convener of the Law and Technology Society of NLS. On behalf of both L-Tech and Law School Policy Review, I’d like to close today’s session with a few words of thanks. First of all, of course, our deepest gratitude to all of our panelists, Professor Arul George Scaria. Professor Dev Gangjee, Advocate Madhav Khosla, and Dr. Zakir T Thomas for so generously sharing their time, expertise, and insight with us this morning. Each of you brought a very distinct vantage point to this conversation, and that range of expertise is exactly what made this discussion as rich as it was. We’d also like to extend our sincere thanks to Senior Advocate Diya Kapur for moderating today’s discussion with such skill and insight. To our colleagues at the Law School Policy Review, Shauryaveer, Shreya, Aastha, Arghav, and the entire editorial board, thank you so much for your partnership in putting this together and for your work in bringing this panel to light. To life, and finally, to all our friends at L-Tech and all our colleagues, Shivam, Mohit, everyone who joined us this evening, and everyone who joined us this evening, we would like to thank you for being here and for your attention and engagement throughout this discussion. You all asked a lot of… Great questions and thank you for this engagement. And for anyone who would like to revisit today’s conversation, LSPR will be uploading a transcript of this session and new posts. Given how quickly this area of law is developing, we hope this is only the first of many such conversations. Thank you all and very good evening to all of you.

Dr Zakir Thomas: Thank you.

Diya Kapur: Thank you. Thank you.

Dev Gangjee: Thank you. Bye.

Madhav Khosla: Thank you.

Categories: Podcast