Destroying Books and Developing AI: For What Good?

Criticism has not been lacking of a recent project to train AI models that resulted in the destruction of books. According to reports, tech companies have been approaching dealers of rare and antiquarian books and placing large orders – sometimes for over a hundred volumes – in order to train their AI models on the content. The use of high-speed scanning machines in the process involves spines being cut off pages being shredded or pulped, thus destroying the books. While one of the companies involved has claimed that its offer to acquire content was made to explore demand and was never actually pursued, the nature of the project has provoked disapproval.

Criticisms: The (De)Valuation of Human Endeavour

There are numerous reasons for which people object to the mass ingestion and destruction of books: breach of copyright, the threat of human obsolescence and the straightforward destruction of books – a practice with grim historical precedents. All have merit, albeit to varying degrees.

If the term of copyright has expired on books, then feeding their content into AI is not a breach of the rights of estates or authors, but for titles used in this way that are still copyright protected, there certainly are questions of ownership. Indeed, many publishers have licensed their content to tech companies for the purposes of training their AI models, though there have been suggestions that in reality, such agreements constituted a surrender on the part of publishers who were struggling to protect their content from being scraped by online bots anyway: if the content is going to be taken and used without an agreement, there is sense in reaching a formal licensing arrangement with a fee.

The argument regarding human obsolescence is more complicated. It is probable that human work in some fields will be replaced by AI and this may well become more likely if training AI models on human-produced content renders them better able to compete with human authors. However, it is far from clear that the development of AI will result in a ‘jobs apocalypse’. It is quite possible that while rendering some forms of work obsolete, AI will create new jobs or render existing roles more rewarding by removing a measure of repetitive work. Moreover, while the writing prepared by AI models is improving, most readers seem to agree that it is formulaic and inelegant. Perhaps the human (in literature) will not go the way of the horse (in agriculture) as a result of automation. Nevertheless, while perhaps over-stated, potential human obsolescence occasioned by a mere likeness of human effort, itself made possible by the use of genuine human-produced content, is a legitimate area of concern.

Opinions might vary but the destruction of rare or antiquarian books is surely distasteful. Books are cultural objects, the products of human endeavour and part of humanity’s collective heritage. Even if the content of those books is preserved in digital form and made accessible in its entirety to a wider readership – which is surely a service to culture and education – something of importance is irreplaceably lost when the original object is destroyed. Old or rare books have a value not simply because of their scarcity, but because they are the incarnation of human effort. Their embodiment of the intellectual concerns or literary (and manufacturing) endeavours of those who produced them goes beyond their monetary price: no pecuniary reward can replace them or compensate for their loss as products of human thought and creativity.

While these objections each have a different focus, they share an underlying concern with the devaluing or inadequate recognition of human effort and creativity. This fact invites an observation on the nature of AI and a question about the trade-offs involved in its development, as illustrated by the AI training project that provoked such visceral condemnation in some quarters.

Training AI and the Value of Human Content

Books can be scanned without destroying them, so a certain amount of the opprobrium occasioned by the project in question is connected to the needless destruction involved. Nevertheless, it is important to keep in mind the fact that the destruction, though thoughtless, lazy or wasteful, was not capricious: there was no desire to destroy purely for the sake of doing so. There was a purpose, of sorts: to train AI models. This can take place in various ways, sometimes through human interaction as ‘trainers’ correct the model’s ability to recognise images, for example. Alternatively, models can crawl the internet and scrape information from websites. It is now possible for more advanced AI models to train other models and there is talk of recursive AI development, whereby models improve their own code and capabilities without the need for human input.

However, the use of books reminds us that the training of models also relies upon human-produced content. As the website that originally reported the story states, the models have absorbed so much data and the internet is now so full of AI-generated ‘slop’ that fresh, human-made data is required if the models are to improve.

Human Exceptionalism and AI

This in itself perhaps tells us something of significance: no matter how sophisticated it has become, no matter how much data it is fed, AI remains substantially less than human. In order to advance further, it needs content created by human beings. What it is crucial to recognise, however, is that even were it to absorb all of the books in the world, it would still neither know nor understand anything. Creativity, understanding, critical thought and imagination are fundamentally human traits, and for all the advances in artificial intelligence, such capacities remain – and one might believe will forever remain – beyond the reach of machines. Put simply, in important respects human beings are by their very nature fundamentally different from AI. Human creativity has a unique value and this is recognised by AI developers – albeit perhaps for different reasons – even when they are destroying its fruits.

Addressing AI Trade-Offs

In light of this observation about the uniqueness of human thought and effort, how are the trade-offs involved in the development and implementation of AI to be addressed? In short, are the costs ‘worth it’? The answer will depend on what it is that AI is believed to be capable of and what it is that is being lost. One social media user commented that books that had survived wars and fires were being destroyed so that AI could write a better marketing email. If that is the case, the standpoint is obvious: valuable and rare artefacts are being lost to pursue an end of highly questionable value.

The trade-offs connected with artificial intelligence are unlikely always to be so clear-cut. In order to assess them, it is important to keep in mind the status of AI and the uniqueness of human beings. Artificial intelligence is emphatically not human and remains only an artifice or tool. As such, the question that must be asked is what AI is being developed and used for. Is it in the service of human beings – and if so, does it meet some genuine human need or realise some definite good?

Much is said at present about the capacity of artificial intelligence to transform the way we work, to increase productivity, to assist with medical research and improve healthcare. The destruction of books – some of which were rare and valuable – to increase incrementally the power of a large language model is of dubious worth. It is far from clear what good this achieves. Will faster or more accurate answers to users’ online queries compensate for what has been lost? Would the development of an ‘agentic’ AI bot that can book flights for busy executives be sufficient? Perhaps tech companies have reason to believe that the more advanced models that emerge will lead to medical breakthroughs. Or are they simply engaged in a development race based on ever greater levels of investment and spending without a clear object in view?

Moreover, the costs – real or potential – beyond the destruction of books must also be considered. Are these – the surrender of copyright or the loss of livelihoods – justified by concrete human goods realised by advancing AI?

Conclusion: Human-Centred AI

Put simply, the development of AI, as a device, is properly conducted in a manner that focuses on genuine goods and human flourishing. (For a fuller discussion, see Andrei Rogobete’s work on human-centred artificial intelligence). In the case of this project, much of the opprobrium online was fuelled by the destruction of books and putative breaches of copyright. However, even with proper licensing of copyright-protected material, even with the preservation of rare books or their return to suitable custodians or re-circulation in the market, the question that needs to be addressed is, ‘To what end?’ or more specifically, ‘For what human good?’ It is the failure adequately to value such human goods that lies at the heart of the major criticisms levelled at the project in question. This single instance illustrates a far broader principle: without some clearly defined human good in view, AI-development risks appearing speculative, indeterminate and of questionable value.

About the Author

Neil Jordan

Neil Jordan is Senior Editor at the Centre for Enterprise, Markets and Ethics. For more information about Neil please click here.