109. Danijar Hafner - Gaming our way to AGI

109. Danijar Hafner - Gaming our way to AGI

Author: The TDS team January 12, 2022 Duration: 50:06

Until recently, AI systems have been narrow — they’ve only been able to perform the specific tasks that they were explicitly trained for. And while narrow systems are clearly useful, the holy grain of AI is to build more flexible, general systems.

But that can’t be done without good performance metrics that we can optimize for — or that we can at least use to measure generalization ability. Somehow, we need to figure out what number needs to go up in order to bring us closer to generally-capable agents. That’s the question we’ll be exploring on this episode of the podcast, with Danijar Hafner. Danijar is a PhD student in artificial intelligence at the University of Toronto with Jimmy Ba and Geoffrey Hinton and researcher at Google Brain and the Vector Institute.

Danijar has been studying the problem of performance measurement and benchmarking for RL agents with generalization abilities. As part of that work, he recently released Crafter, a tool that can procedurally generate complex environments that are a lot like Minecraft, featuring resources that need to be collected, tools that can be developed, and enemies who need to be avoided or defeated. In order to succeed in a Crafter environment, agents need to robustly plan, explore and test different strategies, which allow them to unlock certain in-game achievements.

Crafter is part of a growing set of strategies that researchers are exploring to figure out how we can benchmark and measure the performance of general-purpose AIs, and it also tells us something interesting about the state of AI: increasingly, our ability to define tasks that require the right kind of generalization abilities is becoming just as important as innovating on AI model architectures. Danijar joined me to talk about Crafter, reinforcement learning, and the big challenges facing AI researchers as they work towards general intelligence on this episode of the TDS podcast.

***

Intro music:

- Artist: Ron Gelinas

- Track Title: Daybreak Chill Blend (original mix)

- Link to Track: https://youtu.be/d8Y2sKIgFWc

***

Chapters:

  • 0:00 Intro
  • 2:25 Measuring generalization
  • 5:40 What is Crafter?
  • 11:10 Differences between Crafter and Minecraft
  • 20:10 Agent behavior
  • 25:30 Merging scaled models and reinforcement learning
  • 29:30 Data efficiency
  • 38:00 Hierarchical learning
  • 43:20 Human-level systems
  • 48:40 Cultural overlap
  • 49:50 Wrap-up

While the active production of Towards Data Science has concluded, its archive remains a vital resource. Created by The TDS team, this collection captures a specific moment in the rapid evolution of data science and artificial intelligence. Each conversation pulls you directly into the room with leading researchers and practitioners who were shaping the tools and theories of their time. The discussions are not abstract lectures; they are grounded explorations of real-world problems, ethical dilemmas, and technical challenges that defined the field's trajectory. You'll hear experts dissect the implications of their work, from algorithmic fairness to the practicalities of deploying models at scale. This podcast served as a forum for nuanced debate, where complex ideas were unpacked with clarity and depth. Listening now offers a unique historical perspective, a chance to understand the foundational conversations that continue to influence where technology is headed next. The archive of Towards Data Science stands as a substantive record of insight, preserving the voices and questions from the forefront of a digital revolution.
Author: Language: en-us Episodes: 50

Towards Data Science
Podcast Episodes
87. Evan Hubinger - The Inner Alignment Problem [not-audio_url] [/not-audio_url]

Duration: 1:09:32
How can you know that a super-intelligent AI is trying to do what you asked it to do? The answer, it turns out, is: not easily. And unfortunately, an increasing number of AI safety researchers are warning that this is a…
86. Andy Jones - AI Safety and the Scaling Hypothesis [not-audio_url] [/not-audio_url]

Duration: 1:25:44
When OpenAI announced the release of their GPT-3 API last year, the tech world was shocked. Here was a language model, trained only to perform a simple autocomplete task, which turned out to be capable of language transl…
85. Brian Christian - The Alignment Problem [not-audio_url] [/not-audio_url]

Duration: 1:06:19
In 2016, OpenAI published a blog describing the results of one of their AI safety experiments. In it, they describe how an AI that was trained to maximize its score in a boat racing game ended up discovering a strange ha…
83. Rosie Campbell - Should all AI research be published? [not-audio_url] [/not-audio_url]

Duration: 52:37
When OpenAI developed its GPT-2 language model in early 2019, they initially chose not to publish the algorithm, owing to concerns over its potential for malicious use, as well as the need for the AI industry to experime…
82. Jakob Foerster - The high cost of automated weapons [not-audio_url] [/not-audio_url]

Duration: 54:07
Automated weapons mean fewer casualties, faster reaction times, and more precise strikes. They’re a clear win for any country that deploys them. You can see the appeal. But they’re also a classic prisoner’s dilemma. Once…
81. Nicolas Miailhe - AI risk is a global problem [not-audio_url] [/not-audio_url]

Duration: 56:03
In December 1938, a frustrated nuclear physicist named Leo Szilard wrote a letter to the British Admiralty telling them that he had given up on his greatest invention — the nuclear chain reaction. "The idea of a nuclear…