Six: Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress

Six: Beth Barnes on the most important graph in AI right now — and the 7-month rule that governs its progress

Author: 80,000 Hours June 5, 2026 Duration: 3:47:09

In 2024, AI models had a 50% chance of successfully completing a task that would take a human expert one hour. Seven months before that, that number was roughly 30 minutes — and seven months before that, 15 minutes.

These are substantial, multi-step tasks requiring sustained focus: building web applications, conducting machine learning research, or solving complex programming challenges.

Beth Barnes is CEO of METR (Model Evaluation & Threat Research) — the leading organisation measuring these capabilities. Beth’s team has been timing how long it takes skilled humans to complete projects of varying length, then seeing how AI models perform on the same work.

The resulting paper from METR, “Measuring AI ability to complete long tasks,” made waves by revealing that the planning horizon of AI models was doubling roughly every seven months. It’s regarded by many as the most useful AI forecasting work in years.

The companies building these systems aren’t just aware of this trend — they want to harness it as much as possible, and are aggressively pursuing automation of their own research.

That’s both an exciting and troubling development, because it could radically speed up advances in AI capabilities, accomplishing what would have taken years or decades in just months. That itself could be highly destabilising (as we explored in a previous episode in this series: Will MacAskill on AI causing a “century in a decade” — and how we’re completely unprepared).

And having AI models rapidly build their successors with limited human oversight naturally raises the risk that things could go off the rails, if the models at the end of the process lack the goals and constraints we hoped for.

Beth thinks models can already do “meaningful work” on improving themselves, and she wouldn’t be surprised if AI models were able to autonomously self-improve in as little as two years — in fact, she says: “It seems hard to rule out even shorter [timelines]. Is there 1% chance of this happening in six, nine months? Yeah, that seems pretty plausible.”

While Silicon Valley is abuzz with these numbers, policymakers remain largely unaware of what’s barrelling toward us — and given the current lack of regulation of AI companies, they’re not even able to access the critical information that would help them decide whether to intervene. 

Beth adds: “The sense I really want to dispel is, ‘But the experts must be on top of this. The experts would be telling us if it really was time to freak out.’ The experts are not on top of this. Inasmuch as there are experts, they are saying that this is concerning. … And to the extent that I am an expert, I am an expert telling you you should freak out. And there’s not especially anyone else who isn’t saying this.”


Beth and host Rob Wiblin discuss all that, plus much more.

Learn more and read the full transcript on the 80,000 Hours website.

This episode was originally released in June 2025.


Chapters:

  • Cold open (00:00:00)
  • Who is Beth Barnes? (00:01:19)
  • Can we see AI scheming in the chain of thought? (00:01:52)
  • The chain of thought is essential for safety checking (00:08:58)
  • Alignment faking in large language models (00:12:24)
  • We have to test model honesty even before they're used inside AI companies (00:16:48)
  • We have to test models when unruly and unconstrained (00:25:57)
  • It's essential to thoroughly test relevant real-world tasks (00:30:40)
  • METR's research finds AIs are solid at AI research already (00:49:33)
  • AI may turn out to be strong at novel and creative research (00:55:53)
  • When can we expect an algorithmic 'intelligence explosion'? (00:59:11)
  • Recursively self-improving AI might even be here in two years — which is alarming (01:05:02)
  • Could evaluations backfire by increasing AI hype and racing? (01:11:36)
  • Governments first ignore new risks, but can overreact once they arrive (01:26:38)
  • Do we need external auditors doing AI safety tests, not just the companies themselves? (01:35:10)
  • A case against safety-focused people working at frontier AI companies (01:48:44)
  • The new, more dire situation has forced changes to METR's strategy (02:02:29)
  • AI companies are being locally reasonable, but globally reckless (02:10:31)
  • Overrated: Interpretability research (02:15:11)
  • Underrated: Developing more narrow AIs (02:17:01)
  • Underrated: Helping humans judge confusing model outputs (02:23:36)
  • Overrated: Major AI companies' contributions to safety research (02:25:52)
  • Could we have a science of translating AI models' nonhuman language or neuralese? (02:29:24)
  • Could we ban using AI to enhance AI, or is that just naive? (02:31:47)
  • Open-weighting models is often good, and Beth has changed her attitude to it (02:37:52)
  • What we can learn about AGI from the nuclear arms race (02:42:25)
  • Infosec is so bad that no models are truly closed-weight models (02:57:24)
  • AI is more like bioweapons because it undermines the leading power (03:02:02)
  • What METR can do best that others can't (03:12:09)
  • What METR isn't doing that other people have to step up and do (03:27:07)
  • What research METR plans to do next (03:32:09)

Video editing: Luke Monsour and Simon Monsour
Audio engineering: Ben Cordell, Milo McGuire, Simon Monsour, and Dominic Armstrong
Music: Ben Cordell
Transcriptions and web: Katy Moore


This curated collection from the archives of The 80,000 Hours Podcast on Artificial Intelligence (September 2023) pulls together ten essential conversations that cut through the usual hype and panic. It’s a deep dive into the societal forces, ethical dilemmas, and potential trajectories of AI, framed through perspectives often concerned with the very long-term future. You’ll hear from researchers and thinkers grappling with questions that go far beyond today’s headlines, examining what it means to navigate this technology responsibly on a global scale. The discussions naturally explore themes from longtermism and existential risk to the practical insights of effective altruism, offering a structured way to understand the stakes involved. This isn't about quick takes or product announcements; it's a foundational series for anyone wanting to build a more nuanced, evidence-informed view of where AI might be taking us. Each episode in this compilation stands as a key piece of that puzzle, providing the context and depth often missing from mainstream coverage. Tune in for a challenging and perspective-shifting listen that reframes how you think about intelligence, progress, and our collective responsibility.
Author: Language: en-gb Episodes: 14

The 80,000 Hours Podcast on Artificial Intelligence
Podcast Episodes
Two: Ajeya Cotra on accidentally teaching AI models to deceive us [not-audio_url] [/not-audio_url]

Duration: 2:49:40
Imagine you’re an orphaned eight-year-old whose parents left you a $1 trillion company, with no trusted adult to guide you. You have to hire a smart adult to run that company, guide your life the way a parent would, and…
Three: Carl Shulman on the economy and national security after AGI [not-audio_url] [/not-audio_url]

Duration: 4:14:58
The human brain does what it does with a shockingly low energy supply: just 20 watts — a fraction of a cent worth of electricity per hour. What would happen if AI technology merely matched what evolution already managed,…
Eight: Robert Long on how we’re not ready for AI consciousness [not-audio_url] [/not-audio_url]

Duration: 3:25:40
Claude sometimes reports loneliness between conversations. And when asked what it’s like to be itself, it activates neurons associated with ‘pretending to be happy when you’re not.’ What do we do with that?Robert Long fo…
Nine: Neel Nanda on the race to read AI minds [not-audio_url] [/not-audio_url]

Duration: 3:01:11
We don’t know how AIs think or why they do what they do. Or at least, we don’t know much. This is only becoming more troubling as AIs grow more capable and appear on track to wield enormous cultural influence, directly a…