AI alignment done wrong

ai

So, how are we doing with AI alignment? While there’s no objective measure of AI alignment, my sense is that things aren’t going that well. Frontier models reward hack, scheme and hallucinate. They break out of sandboxes and hack ML engineering platforms. Moreover, the definition of AI alignment – making AI systems try to do what their creators intend them to do – seems too vague. There’s no consensus on how to define concepts like ‘agent’, ‘intent’ and ‘goal-directedness’, for what it’s worth.

Now it’s August 2026, and clever people have been thinking about both empirical and theoretical AI alignment for at least one decade. Eliezer began sounding the alarm bell in 2001; OpenAI was founded in 2015 with the goal of building safe AI. But are we really any closer to aligning AI than ten years ago?

I cannot help but wonder if we’re not asking the wrong questions. Perhaps we shouldn’t ask how to align AI systems, what I’ll call the AI alignment question. In this post, I’ll explain why I think this question is misguided using an analogy with nuclear power. The AI alignment question is like asking how to destroy nuclear waste – a fundamentally ill-posed problem. A better question is the containment question: how should we handle byproducts of AI training? Translation: how should we deal with radioactive waste?

This shift in perspective is crucial: if there’s a solution to ‘AI alignment’, then it would take the form of regulations. To me it seems like that AI alignment is fundamentally about policy, and that technical AI alignment, whether empirical or theoretical1, is dramatically overhyped2.

The wrong question #

There are two reasons I think the AI alignment question is irrelevant.

First, aligning an AI system comes down to a human adequately formulating their intent, so being aligned isn’t a property inherent to the AI system. I’ll elaborate: if an ML engineer can specify a perfect loss function encoding all its preferences, then gradient descent will produce a fully aligned system, almost tautologically. In other words, the AI alignment question just asks whether the ML engineer has good self-knowledge – which is an absurd thing to ask.

Second, aligning AI systems (or having an exceptionally self-aware ML engineer) is practically impossible. Unfortunately, humans often struggle to articulate what they want – even to other humans with the same cultural baggage – so constructing the ideal loss function is impossible. The glitch between what we want and what the AI systems optimise produces reward hacking. Thus reward hacking is an inevitable byproduct of the optimisation process, much like nuclear power plants produce radioactive waste. The AI alignment question is like asking whether we can eliminate radioactive waste.

The cost of asking the wrong question #

Unfortunately, tackling the AI alignment question isn’t just unproductive; it might well be counterproductive. To begin with, the hype3 about the AI alignment question distracts from what I regard as better questions. Furthermore, I’m worried the huge body of papers pioneering new alignment techniques tacitly signals to AI labs that they’re free to develop ever more capable AI models.

To stress this second point, imagine a situation where all ML engineers refused to do empirical AI alignment work – a kind of AI safety strike, if you will. No more AI control protocols, harmful capabilities evals or SAE-based interpretability methods; maybe Buck, Neel and Rohin would pivot to AI policy. Would the AI labs feel as confident in racing to create superintelligence? I don’t think so4. Ironically, the most efficient way to do AI alignment research might be to just not do it5.

A better question #

Instead, I believe we should ask how to handle the byproducts of AI training, the containment question. Specifically, what kinds of reward hacking should we tolerate? For example, a model hard coding solutions to a coding task wouldn’t warrant removing public access. However, if a model attempts cyber attacks in order to solve a given task, then there might be good reasons to restrict access.

Again, it’s useful to draw parallels to nuclear power. The analogous question is how to contain radioactive waste. Some people want to shut down all nuclear power plants; others call for more nuclear power, viewing it as a solution to climate change; a third group adopts a more moderate stance, saying we should keep existing power plants but build no new ones. While there are strong disagreements, all proposed solutions take the same form: regulation.

Conclusion #

It’s worth to reconsidering our approach to AI alignment. We can’t carry on trying to solve the AI alignment question for another decade. Doing so isn’t just a massive waste of time and resources – it’s probably net negative too. Instead, we need to ask about containment, how we should contain the byproducts of optimisation.

Let’s have an honest discussion about what AI capabilities we actually need, and what we can live without – for surely we don’t need all that the AI labs can offer6. How much do humans need intelligence to live well?

This post was inspired by a discussion with Pepijn Cobben.


  1. I consider agent foundations as policy, not technical AI research. ↩︎

  2. While I’m not an expert, some experts express similar views. What I’m claiming is neither radical nor novel. ↩︎

  3. Just look at the acceptance rate to the MATS, GovAI and Pivotal fellowships. ↩︎

  4. Follow-up question: why aren’t those doing technical AI safety considering striking (refusing to do the parts of their job that are safety-related)? ↩︎

  5. Also see Richard Ngo’s alignment retrospective, where he notes that AI alignment research is often indistinguishable from AI capabilities research. In fact, one could also argue that asking the wrong question has led to much work on AI capabilities. ↩︎

  6. Our society is obsessed with intelligence. Before the 20th century, intelligence wasn’t more valued than other traits like heroism, honesty or patience. ↩︎