124. Alex Watson - Synthetic data could change everything

124. Alex Watson - Synthetic data could change everything

Author: The TDS team May 18, 2022 Duration: 51:47

There’s a website called thispersondoesnotexist.com. When you visit it, you’re confronted by a high-resolution, photorealistic AI-generated picture of a human face. As the website’s name suggests, there’s no human being on the face of the earth who looks quite like the person staring back at you on the page.

Each of those generated pictures are a piece of data that captures so much of the essence of what it means to look like a human being. And yet they do so without telling you anything whatsoever about any particular person. In that sense, it’s fully anonymous human face data.

That’s impressive enough, and it speaks to how far generative image models have come over the last decade. But what if we could do the same for any kind of data?

What if I could generate an anonymized set of medical records or financial transaction data that captures all of the latent relationships buried  in a private dataset, without the risk of leaking sensitive information about real people? That’s the mission of Alex Watson, the Chief Product Officer and co-founder of Gretel AI, where he works on unlocking value hidden in sensitive datasets in ways that preserve privacy.

What I realized talking to Alex was that synthetic data is about much more than ensuring privacy. As you’ll see over the course of the conversation, we may well be heading for a world where most data can benefit from augmentation via data synthesis — where synthetic data brings privacy value almost as a side-effect of enriching ground truth data with context imported from the wider world.

Alex joined me to talk about data privacy, data synthesis, and what could be the very strange future of the data lifecycle on this episode of the TDS podcast.

***

Intro music:

- Artist: Ron Gelinas

- Track Title: Daybreak Chill Blend (original mix)

- Link to Track: https://youtu.be/d8Y2sKIgFWc

***

Chapters:

  • 2:40 What is synthetic data?
  • 6:45 Large language models
  • 11:30 Preventing data leakage
  • 18:00 Generative versus downstream models
  • 24:10 De-biasing and fairness
  • 30:45 Using synthetic data
  • 35:00 People consuming the data
  • 41:00 Spotting correlations in the data
  • 47:45 Generalization of different ML algorithms
  • 51:15 Wrap-up

While the active production of Towards Data Science has concluded, its archive remains a vital resource. Created by The TDS team, this collection captures a specific moment in the rapid evolution of data science and artificial intelligence. Each conversation pulls you directly into the room with leading researchers and practitioners who were shaping the tools and theories of their time. The discussions are not abstract lectures; they are grounded explorations of real-world problems, ethical dilemmas, and technical challenges that defined the field's trajectory. You'll hear experts dissect the implications of their work, from algorithmic fairness to the practicalities of deploying models at scale. This podcast served as a forum for nuanced debate, where complex ideas were unpacked with clarity and depth. Listening now offers a unique historical perspective, a chance to understand the foundational conversations that continue to influence where technology is headed next. The archive of Towards Data Science stands as a substantive record of insight, preserving the voices and questions from the forefront of a digital revolution.
Author: Language: en-us Episodes: 50

Towards Data Science
Podcast Episodes
119. Jaime Sevilla - Projecting AI progress from compute trends [not-audio_url] [/not-audio_url]

Duration: 48:34
There’s an idea in machine learning that most of the progress we see in AI doesn’t come from new algorithms of model architectures. instead, some argue, progress almost entirely comes from scaling up compute power, datas…
118. Angela Fan - Generating Wikipedia articles with AI [not-audio_url] [/not-audio_url]

Duration: 51:44
Generating well-referenced and accurate Wikipedia articles has always been an important problem: Wikipedia has essentially become the Internet's encyclopedia of record, and hundreds of millions of people use it do unders…
117. Beena Ammanath - Defining trustworthy AI [not-audio_url] [/not-audio_url]

Duration: 46:46
Trustworthy AI is one of today’s most popular buzzwords. But although everyone seems to agree that we want AI to be trustworthy, definitions of trustworthiness are often fuzzy or inadequate. Maybe that shouldn’t be surpr…
116. Katya Sedova - AI-powered disinformation, present and future [not-audio_url] [/not-audio_url]

Duration: 54:24
Until recently, very few people were paying attention to the potential malicious applications of AI. And that made some sense: in an era where AIs were narrow and had to be purpose-built for every application, you’d need…
115. Irina Rish - Out-of-distribution generalization [not-audio_url] [/not-audio_url]

Duration: 50:12
Imagine, for example, an AI that’s trained to identify cows in images. Ideally, we’d want it to learn to detect cows based on their shape and colour. But what if the cow pictures we put in the training dataset always sho…
114. Sam Bowman - Are we *under-hyping* AI? [not-audio_url] [/not-audio_url]

Duration: 47:48
Google the phrase “AI over-hyped”, and you’ll find literally dozens of articles from the likes of Forbes, Wired, and Scientific American, all arguing that “AI isn’t really as impressive at it seems from the outside,” and…
113. Yaron Singer - Catching edge cases in AI [not-audio_url] [/not-audio_url]

Duration: 35:20
It’s no secret that AI systems are being used in more and more high-stakes applications. As AI eats the world, it’s becoming critical to ensure that AI systems behave robustly — that they don’t get thrown off by unusual…
110. Alex Turner - Will powerful AIs tend to seek power? [not-audio_url] [/not-audio_url]

Duration: 46:57
Today’s episode is somewhat special, because we’re going to be talking about what might be the first solid quantitative study of the power-seeking tendencies that we can expect advanced AI systems to have in the future.…