Ten: Holden Karnofsky on dozens of opportunities to make AI safer lying on the table — and all his AGI takes

Ten: Holden Karnofsky on dozens of opportunities to make AI safer lying on the table — and all his AGI takes

Author: 80,000 Hours June 5, 2026 Duration: 4:30:19

For years, working on AI safety usually meant theorising about the ‘alignment problem’ or trying to convince other people to give a damn. If you could find any way to help, the work was frustrating and low feedback.

According to Holden Karnofsky — currently at Anthropic, previously cofounder and CEO of Open Philanthropy (now Coefficient Giving) — this situation has now reversed completely.

There are now large amounts of useful, concrete, shovel-ready projects with clear goals and deliverables. Holden thinks people haven’t appreciated the scale of the shift, and wants everyone to see the large range of “well-scoped object-level work” they could personally help with, in both technical and non-technical areas. 

In fact, in this episode alone, Holden lists 39 projects he’s excited to see happening, including:

  • Training deceptive AI models to study deception and how to detect it
  • Developing classifiers to block jailbreaking
  • Implementing security measures to stop ‘backdoors’ or ‘secret loyalties’ from being added to models during training
  • Developing policies on model welfare, AI-human relationships, and what instructions to give models
  • Training AIs to work as alignment researchers

And that’s all just stuff he’s happened to observe directly, which is probably only a small fraction of the options available.

All this low-hanging fruit is one factor behind his decision to join Anthropic this year. That said, his wife is also a cofounder and president of the company, giving him a big financial stake in its success — and making it impossible for him to be seen as independent no matter where he worked.

Holden makes a case that, for many people, working at an AI company like Anthropic will be the best way to steer AGI in a positive direction. He notes there are “ways that you can reduce AI risk that you can only do if you’re a competitive frontier AI company.” At the same time, he believes external groups have their own advantages and can be equally impactful.

Critics worry that Anthropic’s efforts to stay at that frontier encourage competitive racing towards AGI — significantly or entirely offsetting any useful research they do. Holden thinks this seriously misunderstands the strategic situation we’re in: “I work at an AI company, and a lot of people think that’s just inherently unethical. They’re imagining that everyone wishes they could go slowly, but they’re going fast so they can beat everyone else. […] But I emphatically think this is not what’s going on in AI.”

The reality, in Holden’s view:


“I think there’s too many players in AI who […] don’t want to slow down. They don’t believe in the risks. Maybe they don’t even care about the risks. […] If Anthropic were to say, ‘We’re out, we’re going to slow down,’ they would say, ‘This is awesome! Now we have a better chance of winning, and this is even good for our recruiting’ — because they have a better chance of getting people who want to be on the frontier and want to win.”

Holden believes a frontier AI company can reduce risk by:

  • Developing cheap, practical safety measures other companies might adopt
  • Prototyping policies regulators could mandate
  • Gathering crucial data about what advanced AI can actually do

Host Rob Wiblin and Holden discuss the case for and against those strategies, and much more.

Learn more and read the full transcript on the 80,000 Hours website.

This episode was originally released in October 2025.

Chapters:

  • Cold open (00:00:00)
  • Holden is back! (00:02:26)
  • An AI Chernobyl we never notice (00:02:56)
  • Is rogue AI takeover easy or hard? (00:07:32)
  • The AGI race isn't a coordination failure (00:17:48)
  • What Holden now does at Anthropic (00:28:04)
  • The case for working at Anthropic (00:30:08)
  • Is Anthropic doing enough? (00:40:45)
  • Can we trust Anthropic, or any AI company? (00:43:40)
  • How can Anthropic compete while paying the “safety tax”? (00:49:14)
  • What, if anything, could prompt Anthropic to halt development of AGI? (00:56:11)
  • Holden's retrospective on responsible scaling policies (00:59:01)
  • Overrated work (01:14:27)
  • Concrete shovel-ready projects Holden is excited about (01:16:37)
  • Great things to do in technical AI safety (01:20:48)
  • Great things to do on AI welfare and AI relationships (01:28:18)
  • Great things to do in biosecurity and pandemic preparedness (01:35:11)
  • How to choose where to work (01:35:57)
  • Overrated AI risk: Cyberattacks (01:41:56)
  • Overrated AI risk: Persuasion (01:51:37)
  • Why AI R&D is the main thing to worry about (01:55:36)
  • The case that AI-enabled R&D wouldn't speed things up much (02:07:15)
  • AI-enabled human power grabs (02:11:10)
  • Main benefits of getting AGI right (02:23:07)
  • The world is handling AGI about as badly as possible (02:29:07)
  • Learning from targeting companies for public criticism in farm animal welfare (02:31:39)
  • Will Anthropic actually make any difference? (02:40:51)
  • “Misaligned” vs “misaligned and power-seeking” (02:55:12)
  • Success without dignity: how we could win despite being stupid (03:00:58)
  • Holden sees less dignity but has more hope (03:08:30)
  • Should we expect misaligned power-seeking by default? (03:15:58)
  • Will reinforcement learning make everything worse? (03:23:45)
  • Should we push for marginal improvements or big paradigm shifts? (03:28:58)
  • Should safety-focused people cluster or spread out? (03:31:35)
  • Is Anthropic vocal enough about strong regulation? (03:35:56)
  • Is Holden biased because of his financial stake in Anthropic? (03:39:26)
  • Have we learned clever governance structures don't work? (03:43:51)
  • Is Holden scared of AI bioweapons? (03:46:12)
  • Holden thinks AI companions are bad news (03:49:47)
  • Are AI companies too hawkish on China? (03:56:39)
  • The frontier of infosec: confidentiality vs integrity (04:00:51)
  • How often does AI work backfire? (04:03:38)
  • Is AI clearly more impactful to work in? (04:18:26)
  • What's the role of earning to give? (04:24:54)

Video editing: Simon Monsour, Luke Monsour, Dominic Armstrong, and Milo McGuire
Audio engineering: Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: CORBIT
Coordination, transcriptions, and web: Katy Moore


This curated collection from the archives of The 80,000 Hours Podcast on Artificial Intelligence (September 2023) pulls together ten essential conversations that cut through the usual hype and panic. It’s a deep dive into the societal forces, ethical dilemmas, and potential trajectories of AI, framed through perspectives often concerned with the very long-term future. You’ll hear from researchers and thinkers grappling with questions that go far beyond today’s headlines, examining what it means to navigate this technology responsibly on a global scale. The discussions naturally explore themes from longtermism and existential risk to the practical insights of effective altruism, offering a structured way to understand the stakes involved. This isn't about quick takes or product announcements; it's a foundational series for anyone wanting to build a more nuanced, evidence-informed view of where AI might be taking us. Each episode in this compilation stands as a key piece of that puzzle, providing the context and depth often missing from mainstream coverage. Tune in for a challenging and perspective-shifting listen that reframes how you think about intelligence, progress, and our collective responsibility.
Author: Language: en-gb Episodes: 14

The 80,000 Hours Podcast on Artificial Intelligence
Podcast Episodes
Two: Ajeya Cotra on accidentally teaching AI models to deceive us [not-audio_url] [/not-audio_url]

Duration: 2:49:40
Imagine you’re an orphaned eight-year-old whose parents left you a $1 trillion company, with no trusted adult to guide you. You have to hire a smart adult to run that company, guide your life the way a parent would, and…
Three: Carl Shulman on the economy and national security after AGI [not-audio_url] [/not-audio_url]

Duration: 4:14:58
The human brain does what it does with a shockingly low energy supply: just 20 watts — a fraction of a cent worth of electricity per hour. What would happen if AI technology merely matched what evolution already managed,…
Eight: Robert Long on how we’re not ready for AI consciousness [not-audio_url] [/not-audio_url]

Duration: 3:25:40
Claude sometimes reports loneliness between conversations. And when asked what it’s like to be itself, it activates neurons associated with ‘pretending to be happy when you’re not.’ What do we do with that?Robert Long fo…
Nine: Neel Nanda on the race to read AI minds [not-audio_url] [/not-audio_url]

Duration: 3:01:11
We don’t know how AIs think or why they do what they do. Or at least, we don’t know much. This is only becoming more troubling as AIs grow more capable and appear on track to wield enormous cultural influence, directly a…