Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

Author: Conviction June 26, 2026 Duration: 36:18
When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry’s traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated. Noam explains how, if properly scaffolded, today’s models can reason for weeks or even months on complex tasks. He also discusses real-world implications of test-time compute, from building poker solver bots to disproving legendary math conjectures. Together, they also unpack the large gaps in current AI safety frameworks, explore the bottlenecks for recursive self-improvement, and look ahead at the future of multi-agent collaboration and global knowledge sharing. Read more: Implications of Large-Scale Test-Time Compute Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @polynoamial | @OpenAI Chapters: 00:00 – Cold Open 00:43 – Noam Brown Introduction 01:23 – Why Benchmarks Are Broken 04:19 – Compute Budgets and Projections 05:34 – How Long Should Models Think? 06:47 – Benchmark-Maxxing 08:34 – Using Poker Bots as Evals 11:26 – Safety Evals When Model Capability Scales With Budget  14:41 – Release Cycle vs. Agent Runtime  17:06 – Latent Model Capability  20:59 – Limits on Recursive Self-Improvement 27:09 – Large-Scale Multi-Agent Coordination  29:11 – Competition at the Frontier  31:51 – Breaking the Benchmark Grid Equilibrium  33:29 – Why Benchmarks Should be Evaluated by Cost 36:18 – Conclusion

Elad Gil and Sarah Guo guide conversations in No Priors: Artificial Intelligence | Technology | Startups that cut straight to the core of what's happening now. This isn't about abstract futures; it's grounded in dialogues with the very people building and shaping the field-leading AI engineers, pioneering researchers, and the founders turning theory into reality. Each episode tackles the pressing, often daunting questions that define this technological inflection point. You'll hear them explore the practical pathways and hurdles toward AGI, debate which industries are genuinely poised for transformation, and examine how the state-of-the-art in research translates into real-world products and societal shifts. The discussions naturally span the impact on commerce, culture, and the very structure of how we live and work. Produced by Conviction, this podcast serves as an essential, clear-eyed resource for anyone looking to move beyond the hype and understand the forces driving the AI revolution. Sarah Guo, a startup investor, and Elad Gil bring their direct experience to these conversations, ensuring every interview provides substantive insight you can use.
Author: Language: English Episodes: 50

No Priors: Artificial Intelligence | Technology | Startups
Podcast Episodes
Redefining Chip Architecture with Arm CEO Rene Haas [not-audio_url] [/not-audio_url]

Duration: 37:06
From data center orchestrators to AGI and robotics, CPUs remain the heart of modern computing. Arm CEO Rene Haas joins Elad Gil and Sarah Guo to explore how Arm is positioned at the epicenter of AI-driven demands for com…
From Restoring Sight to Reimagining the Brain, with Max Hodak [not-audio_url] [/not-audio_url]

Duration: 31:38
Max Hodak, co-founder and CEO of Science Corporation, joins Sarah Guo to discuss the future of vision, brain-computer interfaces, and the human experience. Max explains how Science’s PRIMA retinal implant could restore f…
Travel Through the Lens of AI with with Booking.com CEO Glenn Fogel [not-audio_url] [/not-audio_url]

Duration: 41:04
When Glenn Fogel joined Priceline in 2000, the business was worth a few hundred million dollars. One week later, the Nasdaq peaked, eventually sending its stock down to a dollar a share. But over 25 years later, Booking…