The University of Oxford authorized digitized materials from the Bodleian Library to be incorporated into OpenAI's training sets, according to internal documents obtained by the Guardian. That purpose had not been made explicit in the public announcement of the partnership, presented in 2025 mainly as a project to digitize and expand access to historical material.
The records indicate that, by June 2025, about 125,000 images of historical dissertations had already been shared with OpenAI. The project mainly involves public-domain materials, including European and American dissertations from the 19th and 20th centuries.
Oxford had officially described the initiative as a pilot to test digitization at scale and make collections previously unavailable online searchable and accessible. The project's current page also states that approximately 429,000 catalog card files were digitized.
Oxford says use was not hidden
The university told the Guardian that the volume of material is “modest in scale”, limited to works out of copyright and made available to OpenAI on a non-exclusive basis. The Bodleian retains the rights to the digitizations and intends to publish the content openly on the internet.
Oxford also rejected the interpretation that the use for training had been hidden. According to the institution, digitization was the main objective of the partnership, but there was transparency about the contribution of the materials to artificial intelligence systems.
The original announcement, however, did not explicitly say that the digitized texts would be added to OpenAI's training sets. Internal documents also recorded employee concerns about possible reputational risks associated with the partnership.
The case reinforces the strategic value of academic and historical collections that are underrepresented on the internet and raises questions about transparency and governance when cultural institutions provide data for commercial AI models.



