29 jul 2026

United States Copyright Office (2025) Copyright and Artificial Intelligence, Part 3: Generative AI Training. Pre-publication version. Washington, DC: United States Copyright Office.



The United States Copyright Office frames generative AI training as a sequence of legally differentiated acts rather than a single technological event. Data acquisition, curation, copying, training, memorisation, retrieval and output are separated so that copyright analysis can attach to specific operations and factual conditions. The report’s theoretical importance lies in resisting categorical answers: fair use cannot be presumed for all training, nor rejected in the abstract, because purpose, commerciality, source acquisition, amount used, market substitution and public benefit vary across systems. Licensing appears not merely as payment but as an infrastructural problem involving metadata, transaction costs, collective representation and sector-specific practices. The method is institutional jurisprudence: technical description is translated into an analytic framework for courts, legislators and markets. Its bridge to platform governance lies in showing that model development depends upon contested regimes of permission. Innovation and authorship are not external values to be balanced after the fact; they are produced through the rules that determine who may convert cultural expression into computational capacity.