96. Jan Leike - AI alignment at OpenAI

96. Jan Leike - AI alignment at OpenAI

Author: The TDS team September 29, 2021 Duration: 1:05:17

The more powerful our AIs become, the more we’ll have to ensure that they’re doing exactly what we want. If we don’t, we risk building AIs that use dangerously creative solutions that have side-effects that could be undesirable, or downright dangerous. Even a slight misalignment between the motives of a sufficiently advanced AI and human values could be hazardous.

That’s why leading AI labs like OpenAI are already investing significant resources into AI alignment research. Understanding that research is important if you want to understand where advanced AI systems might be headed, and what challenges we might encounter as AI capabilities continue to grow — and that’s what this episode of the podcast is all about. My guest today is Jan Leike, head of AI alignment at OpenAI, and an alumnus of DeepMind and the Future of Humanity Institute. As someone who works directly with some of the world’s largest AI systems (including OpenAI’s GPT-3) Jan has a unique and interesting perspective to offer both on the current challenges facing alignment researchers, and the most promising future directions the field might take.

--- 

Intro music:

➞ Artist: Ron Gelinas

➞ Track Title: Daybreak Chill Blend (original mix)

➞ Link to Track: https://youtu.be/d8Y2sKIgFWc

--- 

Chapters:  

0:00 Intro

1:35 Jan’s background

7:10 Timing of scalable solutions

16:30 Recursive reward modeling

24:30 Amplification of misalignment

31:00 Community focus

32:55 Wireheading

41:30 Arguments against the democratization of AIs

49:30 Differences between capabilities and alignment

51:15 Research to focus on

1:01:45 Formalizing an understanding of personal experience

1:04:04 OpenAI hiring

1:05:02 Wrap-up


While the active production of Towards Data Science has concluded, its archive remains a vital resource. Created by The TDS team, this collection captures a specific moment in the rapid evolution of data science and artificial intelligence. Each conversation pulls you directly into the room with leading researchers and practitioners who were shaping the tools and theories of their time. The discussions are not abstract lectures; they are grounded explorations of real-world problems, ethical dilemmas, and technical challenges that defined the field's trajectory. You'll hear experts dissect the implications of their work, from algorithmic fairness to the practicalities of deploying models at scale. This podcast served as a forum for nuanced debate, where complex ideas were unpacked with clarity and depth. Listening now offers a unique historical perspective, a chance to understand the foundational conversations that continue to influence where technology is headed next. The archive of Towards Data Science stands as a substantive record of insight, preserving the voices and questions from the forefront of a digital revolution.
Author: Language: en-us Episodes: 50

Towards Data Science
Podcast Episodes
119. Jaime Sevilla - Projecting AI progress from compute trends [not-audio_url] [/not-audio_url]

Duration: 48:34
There’s an idea in machine learning that most of the progress we see in AI doesn’t come from new algorithms of model architectures. instead, some argue, progress almost entirely comes from scaling up compute power, datas…
118. Angela Fan - Generating Wikipedia articles with AI [not-audio_url] [/not-audio_url]

Duration: 51:44
Generating well-referenced and accurate Wikipedia articles has always been an important problem: Wikipedia has essentially become the Internet's encyclopedia of record, and hundreds of millions of people use it do unders…
117. Beena Ammanath - Defining trustworthy AI [not-audio_url] [/not-audio_url]

Duration: 46:46
Trustworthy AI is one of today’s most popular buzzwords. But although everyone seems to agree that we want AI to be trustworthy, definitions of trustworthiness are often fuzzy or inadequate. Maybe that shouldn’t be surpr…
116. Katya Sedova - AI-powered disinformation, present and future [not-audio_url] [/not-audio_url]

Duration: 54:24
Until recently, very few people were paying attention to the potential malicious applications of AI. And that made some sense: in an era where AIs were narrow and had to be purpose-built for every application, you’d need…
115. Irina Rish - Out-of-distribution generalization [not-audio_url] [/not-audio_url]

Duration: 50:12
Imagine, for example, an AI that’s trained to identify cows in images. Ideally, we’d want it to learn to detect cows based on their shape and colour. But what if the cow pictures we put in the training dataset always sho…
114. Sam Bowman - Are we *under-hyping* AI? [not-audio_url] [/not-audio_url]

Duration: 47:48
Google the phrase “AI over-hyped”, and you’ll find literally dozens of articles from the likes of Forbes, Wired, and Scientific American, all arguing that “AI isn’t really as impressive at it seems from the outside,” and…
113. Yaron Singer - Catching edge cases in AI [not-audio_url] [/not-audio_url]

Duration: 35:20
It’s no secret that AI systems are being used in more and more high-stakes applications. As AI eats the world, it’s becoming critical to ensure that AI systems behave robustly — that they don’t get thrown off by unusual…