Oxford University’s Bodleian Library supplied OpenAI with digital scans of historic theses and texts to train its AI models, a collaboration first disclosed publicly in early 2025 but recently detailed through internal documents.
- Oxford made the partnership with OpenAI public in March 2025.
- Over 125,000 scans of old theses were shared by June 2025.
- Oxford retains rights and plans to make scans publicly available online.
What happened
Oxford University’s Bodleian Library partnered with OpenAI to scan and digitize a vast collection of historic academic texts, primarily old PhD theses from European and American universities dating back to the 19th and 20th centuries. These materials, numbering approximately 125,000 scans by mid-2025, were used not only to preserve and increase access but also to train OpenAI’s artificial intelligence models.
This collaboration was publicly announced in March 2025 with a focus on the scanning and preservation benefits. However, internal documents obtained recently confirm that the digitized texts were incorporated into OpenAI’s training dataset. Oxford emphasized that the materials were out of copyright, nonexclusive, and that the library maintains ownership of the scans, intending to release them online for wider access.
Why it matters
The use of rare and historic academic texts for AI training underscores growing industry demand for high-quality, reliable source data amid concerns about AI-generated content diluting information quality online. As AI models rely heavily on training data, access to unique historical documents can enhance cultural and historical breadth, influencing how these models understand and generate content.
Nevertheless, the deal has raised ethical questions among some Bodleian staff about reputational risks for Oxford and the environmental impact of AI training. Transparency around data use is critical as universities and cultural institutions navigate how to support AI innovation while safeguarding intellectual property and public trust.
What to watch next
Oxford plans to publish the scanned materials online in the coming months, making this unique academic resource more broadly accessible. How these texts are ultimately integrated into AI systems and if further partnerships emerge with other institutions will be closely observed as the AI training landscape evolves.
Meanwhile, the approach taken by the Bodleian Library — preserving physical books while enabling digital training data — contrasts with more controversial practices in the industry, such as books being destroyed for scanning. Monitoring responses from academia, libraries, and the wider public will reveal evolving standards for ethical AI training data acquisition.