#236 – Max Harms on why teaching AI right from wrong could get everyone killed

#236 – Max Harms on why teaching AI right from wrong could get everyone killed

Author: The 80,000 Hours team February 24, 2026 Duration: 2:40:22

Most people in AI are trying to give AIs ‘good’ values. Max Harms wants us to give them no values at all. According to Max, the only safe design is an AGI that defers entirely to its human operators, has no views about how the world ought to be, is willingly modifiable, and completely indifferent to being shut down — a strategy no AI company is working on at all.

In Max’s view any grander preferences about the world, even ones we agree with, will necessarily become distorted during a recursive self-improvement loop, and be the seeds that grow into a violent takeover attempt once that AI is powerful enough.

It’s a vision that springs from the worldview laid out in If Anyone Builds It, Everyone Dies, the recent book by Eliezer Yudkowsky and Nate Soares, two of Max’s colleagues at the Machine Intelligence Research Institute.

To Max, the book’s core thesis is common sense: if you build something vastly smarter than you, and its goals are misaligned with your own, then its actions will probably result in human extinction.

And Max thinks misalignment is the default outcome. Consider evolution: its “goal” for humans was to maximise reproduction and pass on our genes as much as possible. But as technology has advanced we’ve learned to access the reward signal it set up for us, pleasure — without any reproduction at all, by having sex while on birth control for instance.

We can understand intellectually that this is inconsistent with what evolution was trying to design and motivate us to do. We just don’t care.

Max thinks current ML training has the same structural problem: our development processes are seeding AI models with a similar mismatch between goals and behaviour. Across virtually every training run, models designed to align with various human goals are also being rewarded for persisting, acquiring resources, and not being shut down.

This leads to Max’s research agenda. The idea is to train AI to be “corrigible” and defer to human control as its sole objective — no harmlessness goals, no moral values, nothing else. In practice, models would get rewarded for behaviours like being willing to shut themselves down or surrender power.

According to Max, other approaches to corrigibility have tended to treat it as a constraint on other goals like “make the world good,” rather than a primary objective in its own right. But those goals gave AI reasons to resist shutdown and otherwise undermine corrigibility. If you strip out those competing objectives, alignment might follow naturally from AI that is broadly obedient to humans.

Max has laid out the theoretical framework for “Corrigibility as a Singular Target,” but notes that essentially no empirical work has followed — no benchmarks, no training runs, no papers testing the idea in practice. Max wants to change this — he’s calling for collaborators to get in touch at maxharms.com.


Links to learn more, video, and full transcript: https://80k.info/mh26

This episode was recorded on October 19, 2025.

Chapters:

  • Cold open (00:00:00)
  • Who’s Max Harms? (00:01:20)
  • If anyone builds it, will everyone die? The MIRI perspective on AGI risk (00:01:56)
  • Evolution failed to ‘align’ us, just as we'll fail to align AI (00:24:28)
  • We're training AIs to want to stay alive and value power for its own sake (00:42:56)
  • Objections: Is the 'squiggle/paperclip problem' really real? (00:52:24)
  • Can we get empirical evidence re: 'alignment by default'? (01:05:02)
  • Why do few AI researchers share Max's perspective? (01:10:17)
  • We're training AI to pursue goals relentlessly — and superintelligence will too (01:18:34)
  • The case for a radical slowdown (01:24:51)
  • Max's best hope: corrigibility as stepping stone to alignment (01:27:53)
  • Corrigibility is both uniquely valuable, and practical, to train (01:32:34)
  • What training could ever make models corrigible enough? (01:45:06)
  • Corrigibility is also terribly risky due to misuse risk (01:51:38)
  • A single researcher could make a corrigibility benchmark. Nobody has. (01:58:57)
  • Red Heart & why Max writes hard science fiction (02:12:20)
  • Should you homeschool? Depends how weird your kids are. (02:34:08)

Video and audio editing: Dominic Armstrong, Milo McGuire, Luke Monsour, and Simon Monsour
Music: CORBIT
Coordination, transcripts, and web: Katy Moore


The 80,000 Hours Podcast, from The 80,000 Hours team, digs into the complex and often overlooked questions surrounding how we can best use our careers to tackle the world's most pressing problems. While artificial intelligence is a recurring and critical theme, framing some of the most important conversations you won't hear elsewhere, the discussions range far wider into the intersections of technology, policy, philosophy, and global priorities. Hosts Rob Wiblin, Luisa Rodriguez, and Zershaaneh Qureshi guide in-depth interviews with researchers, policymakers, and practitioners, breaking down daunting ideas into actionable insights. You'll hear nuanced analyses of career paths, ethical dilemmas in emerging tech, and evidence-based strategies for creating a positive impact. This isn't about quick tips; it's about deep, substantive exploration of how specific choices and systemic changes can lead to a better future. The podcast lives in the Education and Technology categories because it fundamentally aims to equip listeners with the knowledge and perspective to navigate a rapidly changing world thoughtfully. Each episode is built on rigorous research, challenging assumptions while maintaining a conversational and accessible tone. Tune in for a consistently engaging and intellectually honest look at the forces shaping our century and the practical steps individuals can take within their own 80,000-hour working lives to make a meaningful difference.
Author: Language: English Episodes: 50

80,000 Hours Podcast
Podcast Episodes
Max Nadeau on why ambitious people should start AI safety nonprofits [not-audio_url] [/not-audio_url]

Duration: 1:03:45
There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a list of dozens of ideas…
Inside the first AI-coordinated cyberattack on a real company [not-audio_url] [/not-audio_url]

Duration: 22:03
In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenA…
#252 – Owain Evans on accidentally training AI models to be evil [not-audio_url] [/not-audio_url]

Duration: 2:15:28
Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested…
#250 – Toby Ord on where AGI timelines go wrong [not-audio_url] [/not-audio_url]

Duration: 2:46:03
Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make big mistakes when th…
What the hell happened with AGI timelines in 2026? – Rob Wiblin [not-audio_url] [/not-audio_url]

Duration: 49:28
Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, describing them as “alien tools” that are “rocking the profession.”He was far from alone in his whiplas…
#248 – Jasmine Sun on what the people building AI really believe [not-audio_url] [/not-audio_url]

Duration: 1:06:21
Many AI researchers believe mass job displacement is coming — and some even think there’s a chance their technology will kill everyone. But they’re building it anyway. Writer and journalist Jasmine Sun has been documenti…