In recent weeks there have been reports of autonomous hacking on the part of unreleased AI models. Both Anthropic and OpenAI have admitted to models escaping their testing environments and attacking third parties. The British AI Security Institute has reported similar occurrences during testing, though some have expressed scepticism at the capacity of AI models to act autonomously, while others have questioned the security of the tests. These cases raise some important questions and also help to illustrate certain fundamental issues in moral philosophy.
Freedom and Responsibility
The ability of AI models to act autonomously – or at least to act in ways not foreseen by human developers and testers – and so hack third party infrastructure suggests a need for reconsidering approaches to potential regulation, as suggested by this article. So far, most thought has been based on the assumption that AI will be misused by some human malefactor, but what if an AI model acts ‘independently’ and causes harm unforeseen and unintended by its developers? In such cases, who is to be held liable? There are two questions to consider in this regard. The first is the practical question of whether such ‘escapes’ should be considered accidents or whether the developers bear responsibility for the harm caused by their models, for instance if there is malicious intent or negligence involved. This question is likely a matter to be settled on a case-by-case basis. (An independent audit has been conducted of the escape by OpenAI’s model.) The second is a far broader question of responsibility, which has been discussed before on the CEME blog. As non-conscious machines that simply solve problems and respond to instructions or prompts, AI models are not agents and cannot be regarded as ‘free’ in any conventional sense of the term. They cannot be considered responsible for whatever damage they may do any more than a tree is responsible for the damage caused by a branch that falls from it in high winds. Responsibility must always reside with human beings or organisations (or those who run them) – whether the owners or developers, for example, of the model in question.
Consequentialist Ethics
Cases of autonomous hacking also illustrate an interesting issue in moral philosophy, particularly in relation to consequentialist outlooks such as utilitarianism. This is the theory that considers morality to consist in the maximisation of happiness and the minimisation of pain. When making moral evaluations, we are to focus on the effects or consequences of actions, seeking to achieve ‘the greatest happiness of the greatest number’, as Jeremy Bentham, one of the founding figures of utilitarian thought, put it. There are many varieties of utilitarianism, some of which aim for the minimisation of suffering as a priority rather than the maximisation of happiness, while others concern themselves with preference-satisfaction as a central aim. Some forms of utilitarianism, such as that of John Stuart Mill, take the view that since there exist different qualities of pleasure or happiness, we should consider more than simply the quantity that any action occasions: some pleasures are more worthwhile than others. In short, however, all aim to increase utility, whether understood as happiness, enjoyment, well-being or preference satisfaction, with a focus on the consequences of our actions. As Bentham stated in the opening lines of his Introduction to the Principles of Morals and Legislation: ‘Nature has placed mankind under the governance of two sovereign masters, pain and pleasure. It is for them alone to point out what we ought to do, as well as to determine what we shall do.’
This is a very common – even natural – approach to thinking about ethics: when facing certain moral problems, most of us will at one time or another adopt a consequentialist outlook and reflect on the likely outcomes of the options open to us in terms of happiness and suffering, often almost by default. Moreover, utilitarian perspectives are particularly influential in the domains of economics and therefore public policy, as might be expected with an approach that treats moral evaluations as cost-benefit analyses. Indeed, Bentham talked of a hedonic calculus by means of which pleasures and pains could be measured according to various criteria, so as to evaluate the moral worth of an action.
Autonomous Hacking and Moral Agents
Much has been written on the advantages and weaknesses of utilitarian ethics and the theory has been endlessly refined, but the issue of autonomous hacking on the part of AI models illustrates one particular difficulty: accounting for ‘evil done’, or agency. Consider a case in which a hacker breaks into a company’s IT infrastructure and disables that company’s systems. (For the sake of simplicity, assume that this is a case of straightforward vandalism rather than a ransomware attack.) Most would consider this to be wrong and a utilitarian reckoning would support such a view: with the company unable to function properly, staff and investors stand to lose financially and employees have to suffer stress, uncertainty and added difficulty in doing their jobs. Some might be temporarily laid off. Customers experience lower levels of service or diminished product availability. The perpetrator merely enjoys the perverse pleasure of wanton destruction. The harms clearly outweigh the happiness, and by some margin. Suppose, however, that similar damage is done to a second company by an AI model that, in the course of completing some task assigned by its developers, escapes its test environment and acts ‘autonomously’, hacking the second company’s systems. Most of us would want to say that there is something worse about the first case, just as we would want to say that hitting another person over the head with a club is morally worse than a branch falling on someone’s head in high winds. However, utilitarianism struggles to account for this: the consequences or effects are the same, with the same levels of harm or suffering undergone by the affected parties. There is no easy way of accounting for the difference between harm done and harm simply occurring or befalling someone. One could try to address this problem by seeking to ground a sense of responsibility in talk of causation, according to which agents are responsible for harms that they cause. However, this is too simplistic and does not establish why an armed attacker is responsible for the harm he causes but that a tree or high winds are not. It says nothing about what or who an agent may be and why agents are responsible when other beings are not.
Duty, Virtue and the Centrality of Persons
Of course, there are other problems with utilitarian ethics – and arguably of a much more serious nature than its difficulty in accounting for ‘evil done’. This difficulty certainly does not of itself show that utilitarianism is simply mistaken or that we should not consider the effects of actions when thinking about what to do; neither does it demonstrate that happiness and suffering have no place in our moral reflections. What it does suggest, however, is that a straightforward focus on consequences and utility is too thin an approach to ethics. The utilitarian perspective runs the risk of overlooking the actual human beings involved in any situation by simply aggregating ‘units’ of utility. However, when it comes to moral evaluations, the persons involved, what they undergo and deserve, what they have done or failed to do – and why – are often of vital importance.
As such, our moral considerations require more. We often have to look at the nature of the action itself and not just its consequences. This is the thinking behind deontological or duty ethics: that there are certain things one must do, or simply should never do, regardless of effects. As revealed by divine command or the use of reason (as Immanuel Kant suggested, with his reference to ‘the moral law within’) certain types of action or motive are by their very nature wrong, or matters of obligation. This helps to explain why there is a moral difference between striking someone with a club and a falling branch. The nature of assault – what is intended – gives the act significance.
Moreover, persons are central to our moral thinking. We often reflect on the character of the moral agent – and this is the approach of virtue ethics. A full or flourishing life is also a virtuous life – at least in part – with virtues understood as character traits or dispositions to certain types of action. By cultivating virtues such as courage or justice, initially by way of instruction but over time by practice and one’s own efforts, one becomes a certain kind of person, disposed to just and courageous action. Without persons, it is hard to make sense of ideas such as who is wronged and who acts (and why). Again, the notion of responsibility becomes hard to establish. When we condemn an armed attacker, we comment on his character as vicious or cruel. We consider him an agent and, believing that actions flow from character, we might ground our judgement in evidence of his former conduct. It makes no sense to regard an aged tree in the same way.
The Scope of Moral Thinking
Utilitarianism, then, is not ‘simply false’ owing to its difficulty with questions of agency – though it may be false for other reasons. At the very least, however, it certainly requires supplementing. Beyond taking into account consequences and matters of utility, we need in our moral thinking to consider actions themselves and the motives behind them, as well as the characters of the people who commit them. Persons are central – and with them, notions such as freedom, motive and agency. With situations involving autonomous AI, there seems to be no scope for talking about any of these things: AI models are not persons and do not have motives; they are not free and do not exercise agency. They cannot be responsible, so moral questions about the behaviour of AI models’ behaviour are ultimately to be reduced questions about the ethics of human action (or inaction). What this demonstrates in broad theoretical terms, is that consequentialist ethics centred on concerns over wellbeing or harm are not immediately equipped to account for our sense of moral responsibility or the nature of certain types of wrongdoing. Moreover, responsible persons are so intimately bound up with our methods of making moral judgements that in their absence, the purview of our moral thinking is seriously narrowed and its coherence diminished.