Highlights: #214 – Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway

Highlights: #214 – Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway

Author: The 80,000 Hours team April 18, 2025 Duration: 41:26

Most AI safety conversations centre on alignment: ensuring AI systems share our values and goals. But despite progress, we’re unlikely to know we’ve solved the problem before the arrival of human-level and superhuman systems in as little as three years.

So some — including Buck Shlegeris, CEO of Redwood Research — are developing a backup plan to safely deploy models we fear are actively scheming to harm us: so-called “AI control.” While this may sound mad, given the reluctance of AI companies to delay deploying anything they train, not developing such techniques is probably even crazier.

These highlights are from episode #214 of The 80,000 Hours Podcast: Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway, and include:

  • What is AI control? (00:00:15)
  • One way to catch AIs that are up to no good (00:07:00)
  • What do we do once we catch a model trying to escape? (00:13:39)
  • Team Human vs Team AI (00:18:24)
  • If an AI escapes, is it likely to be able to beat humanity from there? (00:24:59)
  • Is alignment still useful? (00:32:10)
  • Could 10 safety-focused people in an AGI company do anything useful? (00:35:34)

These aren't necessarily the most important or even most entertaining parts of the interview — so if you enjoy this, we strongly recommend checking out the full episode!

And if you're finding these highlights episodes valuable, please let us know by emailing podcast@80000hours.org.

Highlights put together by Ben Cordell, Milo McGuire, and Dominic Armstrong


From the team behind 80,000 Hours comes 80k After Hours, a companion podcast that ventures beyond the main feed's structured career advice. Here, the researchers and producers gather for more informal, wide-ranging conversations that reflect their ongoing curiosities. You'll hear discussions that dig into the nuances of effective altruism, grapple with complex ideas in philosophy and policy, and explore unexpected topics in culture and science-all through the lens of how we might better understand and improve the world. The tone is conversational and exploratory, often feeling like you're listening in on a lively lunch debate among experts who don't take themselves too seriously. While the core mission of finding impactful ways to do good remains, this podcast allows for the tangents, personal reflections, and deeper dives that don't always fit a standard format. It’s a space for the team to think out loud, test arguments, and share the interesting, sometimes quirky, material they're engaging with in their own work. For listeners of the main show, it offers valuable background and new perspectives; for newcomers, it provides an accessible entry point into a community of ideas focused on practical problem-solving and self improvement. Tune in for a blend of thoughtful documentary-style analysis, societal deep dives, and candid chats that aim to be both intellectually substantive and genuinely enjoyable.
Author: Language: English Episodes: 100

80k After Hours
Podcast Episodes
Off the Clock #8: Leaving Las London with Matt Reardon [not-audio_url] [/not-audio_url]

Duration: 1:43:21
Watch this episode on YouTube! https://youtu.be/fJssGodnCQgConor and Arden sit down with Matt in his farewell episode to discuss the law, their team retreat, his lessons learned from 80k, and the fate of the show.
Off the Clock #7: Getting on the Crazy Train with Chi Nguyen [not-audio_url] [/not-audio_url]

Duration: 1:24:27
Watch this episode on YouTube! https://youtu.be/IRRwHCK279EMatt, Bella, and Huon sit down with Chi Nguyen to discuss cooperating with aliens, elections of future past, and Bad Billionaires pt. 2.Check out: Matt’s summer…