Large language models (LLMs) that power AI tools like ChatGPT until now have mainly been trained on vast datasets that include internet data, licensed books and articles, code repositories and human feedback.

But with much of the content available online now exhausted — and increasingly polluted with poor-quality, AI-generated writing — AI labs are turning to uncorrupted texts published pre-2022.

One company, ISBNdb, which boasts that it has the “world’s largest book database”, now offers bulk book buying for AI labs, “up to one million titles per order”, including of “older, rare and specialist volumes”.

“The world’s best AI training data is sitting on a shelf,” it states on its website.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.” {read}