THIS BILLION SECONDS: AI@70 EP 2 -- THE WATERSHED

THIS BILLION SECONDS: AI@70 EP 2 -- THE WATERSHED

Author: Ampel August 21, 2026 Duration: 10:00

THE NEXT BILLION SECONDS: AI@70

EP 2: THE WATERSHED

ChatGPT opened the door to artificial intelligence. Almost immediately, researchers built 'agents' using it - but three years of hard yards passed before agents were 'good enough' to be useful. In this episode we trace the path from 'bad' to 'good enough' agents - and what it unlocked as we 'crossed the watershed'.

No one really thought that on the last day of November in 2022, the world would change completely. Not even the folks closest to it. Like Sam Altman, CEO of OpenAI, who said, "We always knew we'd hit a tipping point," 

"But we didn't know what the moment would be."

His firm modestly promoted a 'research preview' of their new AI chatbot.ChatGPT.

G'day, I'm Mark Pesce and this billion seconds are already unfolding as the most significant of this century.

In this miniseries, celebrating the 70th anniversary of artificial intelligence, we're looking at where we’ve come from, how we got here - and where we seem to be going.

Because we’re travelling at the speed of thought.

ChatGPT promised artificial intelligence on tap for everyone. But it took another three years for that promise to be realised. 

This episode traces that story, from something kids used to do their homework for them - into a tool reshaping our economy.

That's on this episode of This Billion Seconds.

I reckon anyone who used a chatbot in 2023 or 2024 deserves a gold star.

Two gold stars if they used one to get work done.

ChatGPT was interesting.

But it wasn't very accurate.

The earliest chatbots, trained to be eager to please, would sometimes make up their answers. 

With no basis in fact. We call these 'hallucinations', but that's not what they are.

It's that we'd caught the AI out. Asked a question it didn't know how to answer. So, it did its best with whatever it could pull together.

It didn't mean to lie. It didn't even know it was lying. It was just generating a response. In a chatbot conversation that isn't a big deal. You could fact check - against another chatbot, against the web - even, if you can imagine it, ask another human being.

Those sorts of hallucinations could be caught - if you went to some effort.

But then again - why ask a chatbot if you want to make an effort?

Folks took chatbots at their word.

And some them paid a price.

Hallucinations are an annoyance for people.

But they're show-stoppers for autonomous agents.

Oh yes, the nearly forty year old dream of Apple CEO John Sculley of autonomous agents running around and doing all of our work for us - that dream flickered back into life alongside ChatGPT.

Researchers reckoned that with 'good enough' AI, they'd be able to build agents.

To read your email, Keep your calendar. Book a reservation. That sort of thing.

Just three months after ChatGPT landed, the first of those autonomous agents popped up.

A piece of software known as AutoGPT. AutoGPT turns an AI like ChatGPT into an agent.

It does that by providing the three things an agent needs. Memory - so that the agent can remember what it's supposed to be doing, and keep notes on its progress as it does it.

Tools - So that it can do things like read and write files, respond to emails, or add items to the calendar, And goal logic - this is the thing that turns an AI into a single-minded goal-oriented piece of software.

Basically, it's the same quality as the Terminator: the agent has one goal and it will stop at nothing to achieve it. AutoGPT gave ChatGPT memory, tools and goal logic - everything need to turn it into an autonomous agent.

And it should have been absolutely amazing. Except for one thing.

Here's how an agent works: you give it a goal, and it 'decomposes' that goal into a series of steps, then breaks those steps down into discrete actions.

The agent then methodically works its way through each of the actions.

At the end of every action, an agent 'reflects' - it checks the results of that action.  Did it work? Does the agent get to go on to the next action, or does it need to do this action again?

And this is where the problems arise. Because if ChatGPT happens to hallucinate in the midst of this process, the whole thing very quickly falls over.

Maybe an agent thinks an action worked, when it didn't. Or thinks it didn't, even though it did.

Now there's always a change that - for any given action - there will be a hallucination that will make the agent fall over.

You can deal with that if there are just a few actions. Chances are it will get through all of the actions before problems arise.

But for anything even modestly complex, there are many actions. Tens to hundreds. Maybe even thousands.

So the probability for just one hallucination - that's all it takes to make the agent fall over - grows higher and higher as the list of actions grows longer and longer.

Now here's the thing: people measured this.

A group known as Model Evaluation and Threat Research or METR, they've kept a running record of how long agents can run before they fall over.

And back in early 2023, they'd make it no more than about 4 minutes.

An agent can't do a lot in 4 minutes. So although we could make agents using ChatGPT, we couldn't make them work well.

For that, we'd need better AI. Funny thing about that. All of us can take some credit for making AI better.

You see, most everyone who's using ChatGPT or Claude or Gemini or any of the others, is making those models better with every question we put the them. Those questions and the answers to them get fed back in, to train the next generation of models.

It's why AI is so much better now than it was three years ago - and why I say anyone who used AI day-to-day in 2023 or 2024 deserves a gold star. Those were hard yards because the AI just wasn't very good. How do we know AI is better to day than it was two or three years ago? Ah, here's where we come back to those folks at METR.

They've been tracking what they call the 'task horizon' of agents. How long they can perform actions before they fall over. And they noticed that with every subsequent generation of AI, that task horizon got longer. Dependably. It soon became clear that the task horizon doubled an average in seven months. 4 minutes becomes eight minutes toward the end of 2023, then sixteen minutes in mid 2024 thirty-two minutes in early 2025, and then sixty-four minutes. Just over an hour In November 2025. Three years after ChatGPT launched. Now that moment in time - just under a year ago - saw the launch of three brand new models: Google Gemini 3, Anthropic Claude Opus 4.5, and OpenAI GPT-5.2.

Each of these models could be used to create agents with task horizons greater than an hour.

An hour is "long enough" that you can assign an agent a reasonably complex task, let it go off and do the work, with confidence that the work will be done correctly.

That's kind of a magic length of time. It takes agents out of the theoretical - where they'd been for nearly 40 years - and makes them very practical. That's the reason I call this moment "The Watershed". It's the moment when AI gets "good enough" to do real work. I'm far from the only person to notice this. Anyone using AI tools to write software noticed a big shift, as they stopped fighting with their tools. Because the tools had gotten 'smart enough' to handle the task.

This is the moment where we first hear the term 'vibe coding' - tell the agent what kind of software you want to create, and the agent will go off and build it for you. And now that we're on the other side of the watershed, it's all downhill. Gaining speed. Because the task horizon didn't stop growing at an hour. We're more than seven months beyond November 2025, and task horizons have doubled again. Two hours.

By early next year, four hours. And by the end of next year - eight hours. An entire work day. Tell the agent what to do - for the whole of the day - and let it do its thing. We've come a long, long way since November 2023. But we're through the hardest bits. You're using the worst AI you'll ever use.

And boy, will it get better from here. In our next episode, we'll look at what autonomous agents mean for business - and whether business is actually prepared to pay for them. That's on the next episode of This Billion seconds.

THIS BILLION SECONDS was written and recorded by Mark Pesce. Produced with assistance from Myrtle and Pine. If you like this show, please share it with a friend. And make sure to follow or subscribe to get all of the episodes in this series.This is Mark Pesce, thanking you for listening.

See omnystudio.com/listener for privacy information.


Ever feel like the world is shifting beneath your feet? That's the sensation The Next Billion Seconds with Mark Pesce captures and explores. Hosted by award-winning futurist and journalist Mark Pesce, this podcast digs into the profound technological and societal transformations happening right now-at a pace humanity has never before experienced. It’s not just speculation; it’s about understanding the forces reshaping how we work, connect, and even think about value and money, so we can navigate today with more clarity. Mark has a knack for untangling complex ideas, making the dizzying rate of change feel comprehensible and, more importantly, actionable. Each episode serves as a guide, helping you piece together what’s coming next from the signals already here. For anyone curious about where science, technology, and global news are steering us, this podcast from Ampel is an essential companion. Tune in to hear thoughtful analysis that connects the dots between emerging trends and your daily decisions, all delivered with an engaging, accessible style. The future isn't a distant abstraction on this show-it's the next billion seconds, and they're already unfolding.
Author: Language: en-au Episodes: 50

The Next Billion Seconds with Mark Pesce
Podcast Episodes
SXSW 2024 ZEEKR: HOW DOES A NEW BRAND SURVIVE? With Gustaf Gunér [not-audio_url] [/not-audio_url]

Duration: 29:25
A special episode of The Next Billion Cars with Drew Smith and Mark Pesce - car brand ZEEKR: HOW DOES A NEW BRAND SURVIVE? With Gustaf Gunér - head of brand, Zeekr Design. Sign up for 'The Practical Futurist' newsletter…
Ghosts in the Machine [not-audio_url] [/not-audio_url]

Duration: 11:04
Sign up for 'The Practical Futurist' newsletter here. In this series of THE NEXT BILLION SECONDS we'll explore the enormous changes that are taking place right now - almost everywhere we look. Because things are moving f…
AI UNEMPLOYED [not-audio_url] [/not-audio_url]

Duration: 12:22
On this episode, we stare into the abyss of one of our deepest fears - that AI has suddenly made us all obsolete. Or has it? Sign up for 'The Practical Futurist' newsletter here. In this series of THE NEXT BILLION SECOND…
CHIPS AND CHAINS [not-audio_url] [/not-audio_url]

Duration: 12:03
On this episode, we look at what's happened to one of the most important areas of the economy - the semiconductor industry. The very few firms who make the chips that power our civilisation - TSMC, Apple, AMD and Nvidia.…
THE FOUR-DAY WEEK IS ALREADY HERE [not-audio_url] [/not-audio_url]

Duration: 12:48
Sign up for 'The Practical Futurist' newsletter here. We've tacitly adopted the four-day week from two directions... The Future has arrived... The last two years have seen more change and the previous 20. Are we ready fo…
The Future Has Arrived [not-audio_url] [/not-audio_url]

Duration: 7:37
Sign up for 'The Practical Futurist' newsletter here The Future has arrived... The last two years have seen more change and the previous 20. Are we ready for that? In this series of THE NEXT BILLION SECONDS we'll explore…
VALE: Vernor Vinge - creator of a "Technological Singularity" [not-audio_url] [/not-audio_url]

Duration: 43:36
Science fiction legend Vernor Vinge inspired the title of this podcast - and his influence extends far beyond fiction. His novella "True Names" gave readers a first taste of the metaverse, and in a 1993 talk for NASA, Vi…
Cryptonomics - So long and thanks for all the grift! [not-audio_url] [/not-audio_url]

Duration: 27:19
Mark started working with cryptocurrencies back in 2014. Ten years at the coalface has convinced him that - despite incredible promise - cryptocurrencies are pwned by gamblers and grifters. After a decade advising financ…