107. Kevin Hu - Data observability and why it matters

107. Kevin Hu - Data observability and why it matters

Author: The TDS team December 15, 2021 Duration: 49:56

Imagine for a minute that you’re running a profitable business, and that part of your sales strategy is to send the occasional mass email to people who’ve signed up to be on your mailing list. For a while, this approach leads to a reliable flow of new sales, but then one day, that abruptly stops. What happened?

You pour over logs, looking for an explanation, but it turns out that the problem wasn’t with your software; it was with your data. Maybe the new intern accidentally added a character to every email address in your dataset, or shuffled the names on your mailing list so that Christina got a message addressed to “John”, or vice-versa. Versions of this story happen surprisingly often, and when they happen, the cost can be significant: lost revenue, disappointed customers, or worse — an irreversible loss of trust.

Today, entire products are being built on top of datasets that aren’t monitored properly for critical failures — and an increasing number of those products are operating in high-stakes situations. That’s why data observability is so important: the ability to  track the origin, transformations and characteristics of mission-critical data to detect problems before they lead to downstream harm.

And it’s also why we’ll be talking to Kevin Hu, the co-founder and CEO of Metaplane, one of the world’s first data observability startups. Kevin has a deep understanding of data pipelines, and the problems that cap pop up if you they aren’t properly monitored. He joined me to talk about data observability, why it matters, and how it might be connected to responsible AI on this episode of the TDS podcast.

Intro music:

➞ Artist: Ron Gelinas

➞ Track Title: Daybreak Chill Blend (original mix)

➞ Link to Track: https://youtu.be/d8Y2sKIgFWc 0:00

Chapters: 

  • 0:00 Intro
  • 2:00 What is data observability?
  • 8:20 Difference between a dataset’s internal and external characteristics
  • 12:20 Why is data so difficult to log?
  • 17:15 Tracing back models
  • 22:00 Algorithmic analyzation of a date
  • 26:30 Data ops in five years
  • 33:20 Relation to cutting-edge AI work
  • 39:25 Software engineering and startup funding
  • 42:05 Problems on a smaller scale
  • 46:40 Future data ops problems to solve
  • 48:45 Wrap-up

While the active production of Towards Data Science has concluded, its archive remains a vital resource. Created by The TDS team, this collection captures a specific moment in the rapid evolution of data science and artificial intelligence. Each conversation pulls you directly into the room with leading researchers and practitioners who were shaping the tools and theories of their time. The discussions are not abstract lectures; they are grounded explorations of real-world problems, ethical dilemmas, and technical challenges that defined the field's trajectory. You'll hear experts dissect the implications of their work, from algorithmic fairness to the practicalities of deploying models at scale. This podcast served as a forum for nuanced debate, where complex ideas were unpacked with clarity and depth. Listening now offers a unique historical perspective, a chance to understand the foundational conversations that continue to influence where technology is headed next. The archive of Towards Data Science stands as a substantive record of insight, preserving the voices and questions from the forefront of a digital revolution.
Author: Language: en-us Episodes: 50

Towards Data Science
Podcast Episodes
119. Jaime Sevilla - Projecting AI progress from compute trends [not-audio_url] [/not-audio_url]

Duration: 48:34
There’s an idea in machine learning that most of the progress we see in AI doesn’t come from new algorithms of model architectures. instead, some argue, progress almost entirely comes from scaling up compute power, datas…
118. Angela Fan - Generating Wikipedia articles with AI [not-audio_url] [/not-audio_url]

Duration: 51:44
Generating well-referenced and accurate Wikipedia articles has always been an important problem: Wikipedia has essentially become the Internet's encyclopedia of record, and hundreds of millions of people use it do unders…
117. Beena Ammanath - Defining trustworthy AI [not-audio_url] [/not-audio_url]

Duration: 46:46
Trustworthy AI is one of today’s most popular buzzwords. But although everyone seems to agree that we want AI to be trustworthy, definitions of trustworthiness are often fuzzy or inadequate. Maybe that shouldn’t be surpr…
116. Katya Sedova - AI-powered disinformation, present and future [not-audio_url] [/not-audio_url]

Duration: 54:24
Until recently, very few people were paying attention to the potential malicious applications of AI. And that made some sense: in an era where AIs were narrow and had to be purpose-built for every application, you’d need…
115. Irina Rish - Out-of-distribution generalization [not-audio_url] [/not-audio_url]

Duration: 50:12
Imagine, for example, an AI that’s trained to identify cows in images. Ideally, we’d want it to learn to detect cows based on their shape and colour. But what if the cow pictures we put in the training dataset always sho…
114. Sam Bowman - Are we *under-hyping* AI? [not-audio_url] [/not-audio_url]

Duration: 47:48
Google the phrase “AI over-hyped”, and you’ll find literally dozens of articles from the likes of Forbes, Wired, and Scientific American, all arguing that “AI isn’t really as impressive at it seems from the outside,” and…
113. Yaron Singer - Catching edge cases in AI [not-audio_url] [/not-audio_url]

Duration: 35:20
It’s no secret that AI systems are being used in more and more high-stakes applications. As AI eats the world, it’s becoming critical to ensure that AI systems behave robustly — that they don’t get thrown off by unusual…