Are Evals Dead?

Are Evals Dead?

Author: Demetrios September 26, 2025 Duration: 25:24

AI Conversations Powered by Prosus Group 


Your AI agent isn’t failing because it’s dumb—it’s failing because you refuse to test it. Chiara Caratelli cuts through the hype to show why evaluations—not bigger models or fancier prompts—decide whether agents succeed in the real world. If you’re not stress-testing, simulating, and iterating on failures, you’re not building AI—you’re shipping experiments disguised as products.


Guest speaker: Chiara Caratelli - Data Scientist @ Prosus Group

Host: Demetrios Brinkmann - Founder of MLOps Community


~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

Join our Slack community [https://go.mlops.community/slack]

Follow us on X/Twitter [@mlopscommunity](https://x.com/mlopscommunity) or [LinkedIn](https://go.mlops.community/linkedin)]

Sign up for the next meetup: [https://go.mlops.community/register]

MLOps Swag/Merch: [https://shop.mlops.community/]


Hosted by Demetrios, MLOps.community is a space for honest, meandering talks about the real work of making artificial intelligence systems actually work. This isn't about hype or theoretical papers; it's about the messy, practical, and often surprising journey of taking models from a notebook into a live environment. You'll hear from engineers and practitioners who are in the trenches, discussing the tools, the frustrations, and the occasional breakthroughs that define the day-to-day. The conversations are deliberately relaxed, covering everything from traditional machine learning pipelines to the new world of large language models and even the intangible "vibes" of team culture and process. Each episode peels back a layer on what "production" really means, whether that involves deploying a predictive service, managing an agentic system, or maintaining reliability as everything scales. Tuning into this podcast feels like grabbing a coffee with colleagues who aren't afraid to dig into the technical nitty-gritty while keeping the tone conversational and accessible. It's for anyone who builds, manages, or is just curious about the operational backbone that allows AI to deliver value, offering a grounded perspective often missing from the broader conversation.
Author: Language: en-us Episodes: 100

MLOps.community
Podcast Episodes
A New Way of Building with AI [not-audio_url] [/not-audio_url]

Duration: 1:04:49
Thanks to MLflow for supporting this episode — the platform helping teams track, manage, and deploy ML and GenAI projects with ease. Try it free at mlflow.org.What if AI could build and maintain your software—like a co-w…
Inside Uber’s AI Revolution - Everything about how they use AI/ML [not-audio_url] [/not-audio_url]

Duration: 45:23
Kai Wang joins the MLOps Community podcast LIVE to share how Uber built and scaled its ML platform, Michelangelo. From mission-critical models to tools for both beginners and experts, he walks us through Uber’s AI playbo…
The Missing Data Stack for Physical AI [not-audio_url] [/not-audio_url]

Duration: 52:42
The Missing Data Stack for Physical AI // MLOps Podcast #328 with Nikolaus West, CEO of Rerun.Join the Community: https://go.mlops.community/YTJoinInGet the newsletter: https://go.mlops.community/YTNewsletter// AbstractN…
Greg Kamradt: Benchmarking Intelligence | ARC Prize [not-audio_url] [/not-audio_url]

Duration: 48:30
What makes a good AI benchmark? Greg Kamradt joins Demetrios to break it down—from human-easy, AI-hard puzzles to wild new games that test how fast models can truly learn. They talk about hidden datasets, compute tradeof…
The Creator of FastAPI’s Next Chapter // Sebastián Ramírez // #324 [not-audio_url] [/not-audio_url]

Duration: 1:09:37
The Creator of FastAPI’s Next Chapter // MLOps Podcast #324 with Sebastián Ramírez, Developer at FastAPI Labs.Join the Community: https://go.mlops.community/YTJoinInGet the newsletter: https://go.mlops.community/YTNewsle…
Everything Hard About Building AI Agents Today [not-audio_url] [/not-audio_url]

Duration: 47:02
Willem Pienaar and Shreya Shankar discuss the challenge of evaluating agents in production where "ground truth" is ambiguous and subjective user feedback isn't enough to improve performance.The discussion breaks down the…
Tricks to Fine Tuning // Prithviraj Ammanabrolu // #318 [not-audio_url] [/not-audio_url]

Duration: 54:01
Tricks to Fine Tuning // MLOps Podcast #318 with Prithviraj Ammanabrolu, Research Scientist at Databricks. Join the Community: https://go.mlops.community/YTJoinIn Get the newsletter: https://go.mlops.community/YTNewslett…