Recursive safe-improvement?
As AIs automate AI R&D more and more, shouldn’t we make AIs automate AI safety R&D too? Before giving it much thought, I would’ve said ‘yes, of course’. Others, like Marius Hobbahn, Joe Carlsmith and Geoffrey Irving, share this intuition, arguing in favour of some kind of research automation.
At Iliad, the question of whether to automate AI safety research (in a strong sense) became a hot potato. In this post, I wish to survey some arguments for and against research automation.
Varying degrees of automation are possible; we’ve already automated coding, which is a big part of AI safety research. By AI safety automation, I mean research where the only human input is setting a broad research agenda. Agents would be responsible for derivations, experiments, writing, formulating new questions and so on, iterating until a satisfactory answer is reached (whether positive or negative).
It’s worth stressing this kind of workflow would be radically different from how AI safety works today. The ability to replicate papers quickly would be useless; instead, the main skill required would be conceptual clarity. We probably wouldn’t need nearly as many researchers.
The primary reason for wanting to automate AI safety research is to speed up progress. If you can figure out a good technical setup for automation and you draft solid research agendas, running the automated pipeline would likely accelerate progress by several orders of magnitude. The exact success metric isn’t important here; the point is that research would be much, much quicker. Hopefully, progress in AI safety would match progress in AI capabilities, so that the latest safety techniques can always keep frontier models in check. If AI slowdown isn’t happening, safety workers need to keep up.
But in the second sentence of the previous paragraph, that’s a big, jarring “If”. Many objections to automating research simply question whether it’s possible finding an adequate technical setup and reasonable research agendas. While this is mostly an empirical question, I’ll offer two conceptual arguments.
Firstly, how do you align your research system? A system comprising five moderately misaligned models will probably be somewhat misaligned too (no, we don’t get to assume the constituent AI systems were aligned). For example, I’m rather concerned about reward hacking: it’s an inevitable biproduct of optimisation, so it can’t be trained away.
Secondly, specifying good, well-scoped research agendas is difficult. The distinction between AI safety research and general AI capabilities research is fuzzy1 – there’s no consensus within the AI safety community. As pointed out by Richard Ngo at Iliad, an attempt at building an automated researcher might well fournish AI labs with a tool for quickly improving AI capabilities. It could also worsen the race dynamics among AI labs2. Moreover, some aspects of AI safety, like character training, broach more philosophical questions. Do we need to vote about the research agendas? That’s hardly going to speed things up.
The idea of automating AI safety research has been around for some time. Now that it’s becoming possible3, it’s worth grasping what it really means. Whether or not we should automate AI safety research, is ultimately an empirical question: can we construct a robust research scaffold? The only way to find out is to try. But at what price?
This post was inspired by a discussion with Matteo Bulloni.
See Joe Carlsmith’s distinction between the AI capabilities feedback loop and AI safety feedback loop. ↩︎
See his retrospective on AI alignment. ↩︎
Companies like Resolution are investing heavily in research automation. ↩︎