The 80,000 Hours Podcast on Artificial Intelligence
Imagine you’re an orphaned eight-year-old whose parents left you a $1 trillion company, with no trusted adult to guide you. You have to hire a smart adult to run that company, guide your life the way a parent would, and administer your vast wealth. You have to hire them based on a work trial or interview that you design. You don’t get to see any resumes or do reference checks. And because you’re so rich, tonnes of people apply — for all sorts of reasons.
Ajeya Cotra argues this peculiar setup resembles the situation humanity finds itself in as we train very general and very capable AI models using current deep learning methods. Ajeya was a senior research analyst at Coefficient Giving at the time of this interview, and she now works at METR (Model Evaluation & Threat Research).
As she explains, this eight-year-old faces a challenging problem. In the candidate pool there are likely some truly nice people, who sincerely want to help and make decisions that are in your interest. But there are probably other characters too — like people who will pretend to care while you’re monitoring them, but intend to exploit the job to enrich themselves as soon as they think they can get away with it.
Like a child trying to judge adults, at some point humans will need to judge the trustworthiness and reliability of machine learning models that are as goal-oriented as people, and greatly outclass them in knowledge, experience, breadth, and speed. Tricky!
Can’t we rely on models' performance during training tasks to guide us? Ajeya worries this won’t work. The trouble is that three different sorts of models will all produce the same output during training, but could behave very differently once deployed in a setting that allows their true colours to come through. She describes three such motivational archetypes:
In principle, a machine learning training process based on reinforcement learning could spit out any of these three attitudes, because all three would perform roughly equally well on the tests we give them, and ‘performs well on tests’ is how these models are selected.
But while that’s true in principle, maybe it’s not something that could plausibly happen in the real world. After all, if we train an agent based on positive reinforcement for accomplishing X, shouldn’t the training process produce a model that just does X and doesn’t have complex thoughts and goals beyond that?
According to Ajeya, this is one thing we don’t know, and should be trying to test empirically as these models get more capable. For reasons she explains in the interview, the Sycophant or Schemer models may in fact be simpler and easier for the learning algorithm to creep towards than their Saint counterparts.
But there are also ways we could end up actively selecting for motivations that we don’t want.
For example, let’s say you train an agentic AI model to run a small business, selecting for behaviours that make money and measuring success by the balance in its bank account. During training, a highly capable model may experiment with the strategy of tricking its trainers into thinking it has made money legitimately when it hasn’t. Maybe instead it steals some money and covers that up. This isn’t a hypothetical worry: models often come up with creative — sometimes undesirable — approaches during training that their developers didn’t anticipate.
If such deception isn’t caught, a model like this may be rated as particularly successful, and the training process will reinforce its tendency to engage in deceptive behaviour. A model that could deceive without being caught would, in effect, have a competitive advantage.
What if deception is picked up, but just some of the time? Would the model then learn that honesty is the best policy? Perhaps. But it might learn a different lesson instead: that deception does pay, as long as it’s done selectively and carefully enough to avoid detection. Would that actually happen? We don’t yet know, but it’s possible.
In this conversation, Ajeya and host Rob Wiblin discuss the above, as well as:
Learn more and read the full transcript on the 80,000 Hours website.
This episode was originally released in May 2023, but we still think it’s one of the best episodes we have at explaining core risks from power-seeking AI.
Chapters:
Producer: Keiran Harris
Audio mastering: Ryan Kessler and Ben Cordell
Transcriptions: Katy Moore