Contents
- 01 I. Two questions, not one
- 02 II. The European concept of reproduction
- 03 III. What the Munich court held
- A. GEMA v. OpenAI
- B. GEMA v. Suno
- 04 IV. The affirmative case: what the output cannot have invented
- 05 V. The objections, and why they fail
- A. The technical objections
- B. The doctrinal objection: identifiability
- C. A note on the American position
- 06 VI. The limiting principle: where the line falls
A generative model is trained on protected works. In some cases, when prompted, it returns those works, or parts of them, in a form that is recognisably the original. The argument that follows from this is uncomfortable because it is simple. If the work can be obtained from the model, then the work must be present in the model, and a thing that holds a fixation of a protected work from which the work can be reproduced is, in the language of European copyright, a reproduction. The infringement, on this view, is not located only in the output that a user sees on the screen. It is located in the model itself, and the model, once its interface and its branding are set aside, is a set of numerical parameters: its weights.
Two judgments of a single German court have now accepted that reasoning, and in doing so have carried a question that until recently belonged to the technical literature and the academic journals into the operative part of an enforceable injunction. This article defends their conclusion for European Union law, and defends a particular version of it. The position taken here is that the correct test is causal: a model contains a reproduction of a work where the appearance of that work in the output is best explained by the model having encoded the work during training, rather than by anything the prompt supplied or by the model having learned the ordinary conventions of the genre. On this test the difficulty or ease of obtaining the output goes to the strength of the evidence and not to the existence of the reproduction. That position aligns with the two German judgments and departs, deliberately and for reasons rooted in the structure of Union copyright law, from the most careful analysis in the American literature. It is a position about the European Union specifically, and the views of the United States Copyright Office and the decision of the English High Court are used here to test the European analysis and to throw its choices into relief, not as authorities that govern it.
Section I sets out the two questions that the debate has tended to run together, and explains why this article is concerned with the second of them. Section II establishes the European concept of reproduction on which the analysis depends. Section III examines what the two Munich judgments held and on what reasoning. Section IV states the affirmative case, built on an argument about where the information in an output can come from. Section V confronts the strongest objections, technical and doctrinal, and answers them. Section VI addresses the limiting principle, because a test that proves too much proves nothing, and sets out with precision where the line falls.
I. Two questions, not one
Much of the confusion in this area comes from a failure to keep two distinct questions apart, each of which has had its turn at the centre of the litigation and the commentary.
The first question is about training. Does the ingestion of protected works into the training corpus, and their reproduction in the course of the training process, infringe the reproduction right, or is it authorised by an exception, in Europe the text and data mining exceptions and in the United States the fair use doctrine? This was the question that dominated the first wave of litigation. In Bartz v. Anthropic (No. 3:24-cv-05417, N.D. Cal., 23 June 2025) and Kadrey v. Meta Platforms (No. 3:23-cv-03417, N.D. Cal., 25 June 2025), two judges of the Northern District of California held the particular training uses before them to be fair use, although Bartz dealt separately with the retention of pirated library copies, and Kadrey turned in large part on the inadequacy of the plaintiffs’ evidence of harm to the market for their works.
The second question is about output. When the model returns text or music or an image that resembles a protected work, is that output an infringing reproduction, and if so, who is responsible for it? This question moved the focus from the input to the result, and much of the compliance machinery around generative AI has been built to address it, through filters and comparison tools designed to stop the model from returning recognisable training data at the point of generation.
Neither of these is the question this article asks. The training question concerns the preparatory and intermediate reproductions made in acquiring, processing and using the corpus, some of which are fleeting and some of which persist as stored datasets, preprocessed files, validation sets and backups. The output question concerns the reproductions made at the point of generation, which are individual and tied to particular prompts. Between the two sits the model itself, and the question that has drawn the least attention is whether the model, in its resting state, considered simply as a stored artefact, is itself a reproduction of the works it has memorised. That is the question that matters most, because its answer settles the legal character of the artefact that is downloaded, copied, distributed and deployed, and it is the question the two Munich judgments have now answered.
II. The European concept of reproduction
The reproduction right in Union law is the author’s exclusive right to authorise or prohibit reproduction of the work, and it is granted in terms of unusual breadth. Article 2 of Directive 2001/29/EC (the Information Society Directive, ‘InfoSoc’) requires Member States to grant the right to authorise or prohibit “direct or indirect, temporary or permanent reproduction by any means and in any form, in whole or in part.” Every part of that formula has been drafted to resist a narrow reading, and each of them bears on the present question.
The reproduction may be indirect, which means it need not be perceptible to a human being without a machine. A recording on a compact disc reproduces the recorded work even though no one can hear the work by looking at the disc, because the work can be perceived through a device. The reproduction may be temporary or permanent, so duration does not decide whether a reproduction has taken place. It may be by any means and in any form, so the technical format in which the work is fixed carries no legal weight: a work held in a compressed file, or as a set of numerical values, is reproduced as fully as one printed on paper. And it may be in whole or in part, so the reproduction of a fragment engages the right, provided the fragment is protected expression.
The Court of Justice of the European Union set out the governing approach to that last point in Infopaq International v. Danske Dagblades Forening (Case C-5/08, EU:C:2009:465). The Court held that there is no minimum length in the concept of reproduction in part, and that the question is instead whether what has been reproduced contains elements that are the expression of the author’s own intellectual creation. An extract of eleven words could infringe if those words carried the author’s intellectual creation, and would not infringe if they did not. What matters is the originality of the fragment, not its length.
Two features of this framework are decisive for what follows. The first is that Union law defines a reproduction by what it does rather than by what it looks like: a reproduction exists where there is a fixation of the work from which the work can be perceived, directly or indirectly, and it does not matter whether that fixation is legible, transparent or convenient to inspect. The Court has applied the right to copies held on digital storage media in Copydan Båndkopi v. Nokia Danmark A/S (Case C-463/12, EU:C:2015:144), and to fleeting fragments held in the memory of a decoder in Football Association Premier League v. QC Leisure(Joined Cases C-403/08 and C-429/08, EU:C:2011:631), in each case without suggesting that the reproduction had to be readable in the medium that held it.
The second feature is that the concept is meant to be neutral as between technologies, an objective the Directive states in its Recital 5 and which the Court has reinforced in the neighbouring field of the reproduction exceptions, holding in Austro-Mechana Gesellschaft zur Wahrnehmung mechanisch-musikalischer Urheberrechte Gesellschaft mbH v. Strato AG (Case C-433/20, EU:C:2022:217) that the Directive pursues a high level of protection for authors. A concept written to keep pace with new technologies is not one that can be escaped by pointing to the novelty of the technology in question.
III. What the Munich court held
The Regional Court of Munich I, the Landgericht München I, has now applied that framework to model weights in two judgments, both from its 42nd Civil Chamber, a chamber that specialises in copyright. Neither judgment is final. The first is under appeal to the Higher Regional Court, the Oberlandesgericht München. The second remains appealable and, at the time of writing, is not final. They are decisions of a court of first instance, and everything that follows is subject to that qualification. But they are reasoned, they are consistent with each other, and the second in particular engages the technical objections at a level of detail that no earlier European decision had attempted.
A. GEMA v. OpenAI
In GEMA v. OpenAI (LG München I, 42 O 14139/24, 11 November 2025), the German collecting society for musical rights showed that simple prompts requesting or otherwise eliciting the lyrics caused the GPT-4 and GPT-4o models to reproduce substantial parts of nine protected song lyrics substantially verbatim. The court held that the lyrics were reproducibly contained in the models, that the models had memorised them, and that this memorisation was itself a reproduction under Section 16 of the German Copyright Act (Urheberrechtsgesetz, ‘UrhG’), the provision that enacts Article 2 InfoSoc. It treated the technical form of the storage, distributed across the parameters of a neural network in a way that no expert could isolate, as beside the point, because the concept of reproduction reaches any fixation from which the work can be perceived with the help of a machine.
B. GEMA v. Suno
The second judgment, GEMA v. Suno (LG München I, 42 O 763/25, 31 July 2026), is the more important of the two, and not only because it is the more recent. It concerned compositions rather than lyrics: the melody, harmony and rhythm of six works, among them “Atemlos durch die Nacht” and “Big in Japan.” That distinction matters, because it removes the objection that had been available against the first judgment. Lyrics are text, and the verbatim reproduction of text is easy to demonstrate and easy to dismiss as a peculiarity of language models. Music is the harder case, where resemblance is a matter of degree and not of literal correspondence, and the court found reproduction there as well.
The way the outputs were obtained is central to the analysis, and the judgment sets it out with a precision that rewards close reading. For each work, the claimant entered into the generator the original lyrics of the song, a general indication of musical style, and the title. The prompts contained no melody, no harmony, no rhythm, no tempo, no key. Suno characterised the claimant’s testing as extensive and targeted. The court, however, focused on the 338 identical prompts concerning the six works in issue, of which 176 related to “Atemlos durch die Nacht,” 124 to “Big in Japan,” and single figures to the others.
The court held those prompts to be simple and open as to outcome, and its reasons deserve to be set out because they anticipate the objections examined below. Entering the same prompt again in a fresh run changes nothing: the input is identical, the parameters are identical, and the probability distribution from which the output is drawn is therefore identical, so repetition does not steer the model toward a particular result in the way that a suggestive or evaluative prompt would. And the lyrics supplied in the prompt do not fix the music that came back. The court spent several pages establishing that the words of a song set, at most, a frame of atmosphere and stress, and do not determine the sequence of notes, the intervals, the shape of the melody, the time signature, the tempo, the key or the harmonic progression, all of which remain open and interchangeable for the same words. The music was not in the prompt.
On the central question the court held that the compositions were reproducibly contained in the models, that the model versions that contained them were stored on the defendant’s servers in Germany, and that this storage was a reproduction under Section 16 UrhG committed within the jurisdiction, whatever the position as to the training, which had taken place in the United States. It grounded the reproduction analysis on the breadth of Article 2 InfoSoc, on the Infopaq line, and on the authorities holding that reproduction reaches modified and compressed fixations. It relied on Copydan, which concerned digital storage media used for copies, as support for the proposition that technically encoded works remain indirectly perceptible reproductions, and it cited, for the point that a fixation expressed only in probability values still counts, the German commentary discussed in Section V.
Three features of the reasoning should be noted, because they dispose in advance of arguments that might otherwise be raised against the position defended here. The court expressly declined to treat the model as a database or the output as an exact copy of a training item, and it refused the defendant’s application for an expert opinion on whether the works were stored in the model, holding that whether a distribution across weights and parameters amounts to a reproduction is a question of law and not one for expert evidence. It ruled out chance as the explanation for the outputs, on the ground that the complexity and length of the works, together with the simplicity of the prompts, excluded coincidence, and it recorded that the defendant had offered no alternative account of how the outputs could have arisen. And it disposed of the capacity objection, that the model is mathematically too small to store the training corpus, by observing that the claimant had never asserted that the whole corpus was stored, only that these six works were recognisably present, which the claimant was not required to justify by explaining why these six and not others.
IV. The affirmative case: what the output cannot have invented
The case for infringement in the weights does not rest on any analogy, and the analogies usually offered, the model as a library, as a lossy archive, as a music player, tend to cloud the argument rather than carry it. The case rests on a single observation about where the information in an output can come from.
An output produced by a generative model is assembled from three sources and no others. There is the prompt, which the user supplies. There is the randomness introduced at the sampling step, by which the model selects among the possible continuations it has computed. And there are the parameters, the weights, fixed at the end of training and identical for every user. If an output contains information that was not in the prompt and cannot be accounted for by the ordinary workings of the model on any input, that information came from the parameters, because there is nowhere else for it to have come from.
A necessary refinement belongs here, because without it the argument would prove too much. A generative model does not draw evenly from every arrangement of notes or words that is mathematically possible. Training has already shaped its probability distribution, so that it assigns much higher probabilities to the ordinary furniture of a genre: conventional melodic phrases, familiar chord progressions, the standard patterns that a great deal of music shares. Some resemblance between an output and a protected work can therefore come from the model having learned these unprotected conventions, and not from its having encoded any particular work. This is the real difficulty in the argument, and it has to be met head on rather than assumed away.
It is met by asking what best explains the output. Where an output reproduces a combination of protected musical features that is complex and distinctive enough that the ordinary conventions of the genre do not account for it, where that combination was not supplied by the prompt, and where the model was trained on the work in which the combination appears, the natural explanation is that the model encoded that work during training, and the burden shifts to the provider to offer some other plausible account. On the facts in Suno there was no other account on offer. The melody of a specific composition came back, the prompt did not contain it, generic musical convention did not produce it, the model had been trained on it, and the defendant put forward nothing to explain the correspondence. In those conditions the inference that the work is encoded in the parameters is not merely available; it is the only explanation left standing.
This is why the number of prompts, which has drawn so much attention, does not carry the weight placed on it. The attempt count is a measurement. Because each run of an identical prompt is an independent draw from an unchanged distribution, the number of attempts needed to obtain a given output shows how much probability the model assigns to that output, and nothing more. A work that emerges once in four attempts is more strongly memorised than one that emerges once in 176, and the first is better evidence than the second. But both are evidence of the same fact, that the work is encoded in the parameters, because in each case the competing explanations, the prompt and generic convention, have been excluded. Mark A. Lemley and A. Feder Cooper, in “Probabilistic ‘Copies’ in Generative AI Models” (forthcoming in 41 Berkeley Technology Law Journal, 2026), are right that a smaller number of prompts makes for a stronger demonstration, and an advocate seeking to prove memorisation should always prefer the most nearly deterministic extraction available. But the strength of the demonstration is one thing and the existence of the reproduction is another, and the two should not be run together.
The output, then, does two things at once. It is the evidence from which the reproduction in the model is inferred, the visible symptom of an encoding that is itself the reproduction and without which the output could not have occurred. And it may, at the same time, be a further reproduction, or an adaptation, in its own right: the Suno court treated the storage of the work in the model and the generation of the recognisable output as distinct acts, and characterised the generated music as an adaptation of the composition. The two are not rivals. The model-level reproduction is the subject of this article, and the output is both the proof of it and, on the Munich analysis, a separate infringement alongside it. If the work were not in the weights, no sequence of prompts, however long, could have produced it, because the information would have had no source, and it is in that sense that the model is the infringement and the output its symptom.
V. The objections, and why they fail
The position defended here has able opponents, and the objections divide into the technical and the doctrinal. The technical objections have for the most part already been addressed by the Suno court, and it is worth recording that they were addressed by a court that had before it detailed technical submissions, scientific studies and an extensive evidentiary record, rather than dismissed by one that had not. The doctrinal objection is more serious and needs a fuller answer.
A. The technical objections
The first technical objection is that the weights do not contain the training data in any real sense, because they are a lossy, distributed, statistical representation of the whole training set, from which no particular work can be located or extracted by inspection. The parameters, on this view, encode patterns and correlations, not works, and the appearance of a work in the output is an emergent behaviour of the system rather than the retrieval of stored content. This is the argument that succeeded in England, in Getty Images v. Stability AI ([2025] EWHC 2863 (Ch)), where the court found on the evidence that the Stable Diffusion weights did not store or contain copies of the training images and that the model was the product of patterns and features learned during training.
The objection proves less than it claims. It is true, and nothing in the argument here denies, that the parameters cannot be inspected to locate a work, and that the encoding is distributed and lossy. But the reproduction right has never required that a copy be locatable within the medium that carries it. No one can find the chorus of a song among the charge states of a flash memory chip, or the article within the compressed data of a zip file, and it has never been suggested that the file therefore fails to reproduce the work. The copy is identified by what can be obtained from the carrier, and the Court of Justice has applied the reproduction right to compressed and modified fixations on that footing. To require that the work be legible in the weights would be to strip copyright protection from every storage format that is not transparent to inspection, a result the technology-neutral concept of reproduction cannot support. The Suno court reached the same conclusion by a shorter route, refusing expert evidence on whether the works were stored precisely because the question is not a technical one about the architecture of the model but a legal one about the concept of reproduction.
The second technical objection is one of capacity: the model is too small, by a wide margin, to hold the works of the training corpus, so it cannot be storing them. This objection answers a claim that no one makes. The argument is not that the entire corpus is stored in the weights, which is indeed impossible, but that particular works, those that are heavily represented in the training data and so are memorised, are recoverably encoded. The Suno court disposed of the point in exactly these terms, holding that the claimant had never asserted that all training data was stored and needed only to show that the six works in issue were present. A general impossibility is not an answer to a specific demonstration.
The third technical objection concerns the way the outputs were obtained: they were produced by people entering prompts, sometimes many prompts, and the result is therefore a joint product of the model and the prompter rather than a readout of the model’s contents. This is the most substantial of the technical objections, and it is the one the argument in Section IV is built to meet. The objection has force only where the prompt might have supplied the information that appears in the output. Where it did not, the objection falls away, and the structure of the evidence in Suno is what makes the demonstration clean. The prompt supplied the lyrics, and the music came back. The one thing the prompter contributed is the one thing that was not in issue, and the composition, which was in issue, is the thing the prompter did not supply. Holding constant what the prompt contains and observing what the model adds is what separates the model’s contribution from the prompter’s, and here the separation is plain.
B. The doctrinal objection: identifiability
The strongest argument against the position defended here is not technical at all, and it is made by serious scholars of Union copyright law. Matthias Leistner and Lucie Antoine, in “TDM and AI Training in the European Union – From ‘LAION’ to Possible Ways Ahead?” (74 GRUR International 1027–1044 (2025), DOI: 10.1093/grurint/ikaf114), argue from the case law of the Court of Justice that the subject matter of protection must be identifiable with sufficient precision and objectivity, a requirement the Court established in Levola Hengelo v. Smilde Foods (Case C-310/17, EU:C:2018:899) when it held that the taste of a food could not be a protected work because it could not be pinned down with the necessary precision. Applied to model weights, the argument runs that a work supposedly reproduced in the parameters is not identifiable with the required precision and objectivity, because no expert can detect it there, and because its reappearance in an output is a matter of probability and depends on a prompt supplied by a third party. What exists in the model, on this view, is a disposition to produce the work under certain conditions rather than an embodiment of it, and a disposition is not a reproduction. The reproduction right, they conclude, is not engaged at the level of the model, and rightholders must be left to their claims against infringing outputs, where identification is straightforward.
The argument is coherent, and it has to be answered on its own terms, in four steps.
The first is that Levola is concerned with the identifiability of the work, not with the identifiability of its location inside a carrier. The requirement of precision and objectivity exists to define the subject matter of protection, so that the scope of the right is knowable. In every case under discussion the work is identified with complete precision: “Atemlos durch die Nacht” is a specific composition, and no one is in any doubt as to what it is or where it begins and ends. What cannot be identified is the position of the work within the parameters, which is a different thing. Copyright has never required that a copy be locatable within its medium, for the reasons already given about the flash chip and the compressed file. Levola keeps out of protection subject matter that is inherently imprecise; it does not keep out precise subject matter that happens to be held in an opaque medium.
The second step meets the distinction between disposition and embodiment on which the objection rests, and here the information argument does its heaviest work. A disposition to produce a specific melody, reliably, far more often than the ordinary conventions of the genre would explain, in response to prompts that contain no musical information, is possible only if the information that makes up that melody is present in the parameters. Information does not exist apart from a physical carrier, and it cannot come out of a process unless it went into it. If the melody comes out, and it was not in the prompt, and generic convention does not account for it, then it is in the weights, and the disposition is simply the embodiment seen through its behaviour rather than through its structure. Whether the parameters are called a propensity or a fixation is a choice of words, not a choice about the facts, and Article 2 InfoSoc, which reaches reproduction by any means and in any form, was written to make the vocabulary of the carrier irrelevant.
The third step concerns the objectivity that Levola requires, and finds it supplied by the very process the objection treats as the problem. An extraction procedure is a specification: a fixed prompt, fixed generation settings, a set number of attempts, a measured degree of correspondence between output and work. Anyone can run it, and because the model behaves in a fixed way up to its sampling step, as the Suno court explained, it returns the same probability distribution to everyone who runs it. That is not less objective than opening a file with a media player; it is the same operation with a probability attached, and the probability can be measured. What Levola excludes is subject matter that can be known only through variable and subjective sensory impression, the taste on an individual tongue. A repeatable extraction rate is the opposite of that.
The fourth step concerns the consequence that Leistner and Antoine treat as decisive, that a contrary holding would destroy legal certainty because a provider cannot know which of millions of works its model contains. The premise is true and the conclusion does not follow. Uncertainty about how many infringements one has committed has never meant that none was committed, and a publisher who digitises a library without keeping records is not excused by its own ignorance of what it copied. The European system deals with blameless uncertainty where such uncertainty belongs, in the assessment of fault for the purposes of damages and in the proportionality of remedies, and not by narrowing the definition of a reproduction so that the infringement is treated as never having happened. Their proposed alternative, that rightholders may sue on outputs, is not doctrinally neutral. It places model-level copying and distribution outside the reproduction right, and that allocation calls for an affirmative justification of its own rather than following from Levola. That justification has to be made as a matter of policy, where it meets a powerful consideration on the other side: an output-only regime is one that providers of closed models can escape through filtering and that providers of open models cannot.
C. A note on the American position
It is worth pausing on the United States, not because American law governs the European question, but because the convergence is instructive and its limits more instructive still. The United States Copyright Office, in the third part of its report on copyright and artificial intelligence (Copyright and Artificial Intelligence, Part 3: Generative AI Training, pre-publication version, May 2025), concluded that model weights may contain copies where they memorise protectable expression capable of being reproduced with the aid of a machine, reasoning by the same route as the Munich court that content need not be directly perceptible to constitute a copy. Lemley and Cooper, in “Probabilistic ‘Copies’ in Generative AI Models,” accept in a qualified form that a heavily memorised work is a stored copy, drawing the line at whether the work can be extracted with relatively little effort. The related work of A. Feder Cooper and James Grimmelmann, “The Files Are in the Computer: On Copyright, Memorization, and Generative AI” (100 Chicago-Kent Law Review 141 (2025)), develops the technical basis for treating extraction as evidence of memorisation rather than as its cause.
That qualification is where the European analysis parts company with the American, and the divergence is one of principle. The American concept in issue is fixation, a statutory requirement that a work be embodied in a copy stably enough to be perceived, and the effort-based threshold is an attempt to give that concept a workable boundary against the concern that any large model would otherwise contain copies of everything distinctive it had ingested. The European concept is reproduction under Article 2 InfoSoc, which contains no effort criterion and which the Court of Justice has consistently declined to narrow in order to solve problems that arise downstream, preferring to solve them downstream. Under Union law the concerns that drive the American threshold have other homes: proportionality lives in the law of remedies, fault lives in the assessment of damages, and the fear that everything becomes a copy is answered by the requirement of protected expression and by the ordinary rules of evidence, under which only works whose recoverable presence is shown are ever in issue. The single threshold that Union law requires is factual and causal, and the second, legal threshold that the American analysis builds into the definition of a copy is, for the European system, both unnecessary and out of place.
VI. The limiting principle: where the line falls
An argument that any work obtainable from a model is for that reason contained in it would prove far too much, and an honest defence of the position has to say where the line falls. A large model can, in principle, produce almost any short passage of text or music if enough attempts are made and the prompt is pushed toward the result, and much of what it produces will resemble existing works simply because it has learned the conventions those works share. If every resemblance counted as proof of storage, the concept of a reproduction would lose all limiting force.
The line defended here is causal, and it is drawn at explanation rather than at improbability. The question is not merely whether a particular output was more likely than it would have been under pure randomness, because a shaped distribution makes a great many conventional outputs likely without any work having been encoded. The question is whether the probability and the correspondence are best explained by the model having encoded the specific protected expression of a particular work, rather than by its having learned the unprotected patterns and conventions that works of the kind have in common. That is the floor, and it sits low, well below the threshold of easy extraction that the American analysis requires, but it is not the vanishing floor of mere improbability.
For the claimant to cross it, a combination of things has to be shown, and they can be stated without a list because they form a single chain of reasoning. The work must have been in the training data, so that encoding is possible in the first place. The output must reproduce a combination of protected expression complex and distinctive enough that the ordinary conventions of the form do not account for it. That combination must not have been supplied by the prompt. And the provider must have offered no plausible alternative explanation for the correspondence, whether generalised learning of convention or anything else. Where all of this holds, as it did in Suno, the output supports the inference that the work is encoded in the parameters, and the difficulty of the extraction bears only on how strong that inference is, never on whether the reproduction exists.
The English holding in Getty can be turned to account here, and the turn is worth making because it uses the sceptics’ own authority. Before the English court reached its finding that the weights contained no copies, it had to decide a prior question: whether model weights are an article at all, capable in principle of being an infringing copy. It held that they are, that an intangible set of parameters can be an article for the purposes of the Copyright, Designs and Patents Act 1988. That holding is English and not European, and the caution matters, because Getty is a decision under a national statute of a state that has left the Union and it does not bind the reading of Article 2 InfoSoc. But the point it settles is not peculiar to English law: the weights are a thing, a fixed and copyable artefact, and not a mere capacity. This is the answer to the objection, sometimes made, that treating the model as a reproduction would equally treat a skilled musician who can play a song on request as containing an infringing copy. The musician’s ability is not an artefact. It cannot be copied, distributed, downloaded onto a server, or seized to satisfy a judgment. The weights can be, and are, all of these things, and it is as an artefact, and not as a capacity, that the model falls within a right that has always attached to fixations in things.
The line, then, is drawn where the structure of Union copyright law indicates it should be drawn. A model contains a reproduction of a work where the work is protected expression, where the work appears in the output, and where the appearance of the work is best explained by the model having encoded it rather than by the prompt or by generic convention. The difficulty of the extraction is evidence of the strength of the memorisation and nothing more. The impossibility of locating the work in the parameters is a fact about the medium and not about the copy. And the artefact that results, the trained model considered as its weights, is a reproduction of the protected works it has memorised, one that may be copied, distributed and deployed only within the limits the reproduction right imposes. What follows from that for the companies now running such models on their own servers, and for the open-weight ecosystem in particular, is a separate question, and the subject of a separate analysis.