Product Metrics are LLM Evals // Raza Habib CEO of Humanloop // #320

Author: Demetrios June 3, 2025 Duration: 53:06

Technology

Raza Habib, the CEO of the LLM Eval platform Humanloop, talks to us about how to make your AI products more accurate and reliable by shortening the feedback loop of your evals. Quickly iterating on prompts and testing what works, along with some of his favorite Dario from Anthropic AI Quotes.

// Bio

Raza is the CEO and Co-founder at Humanloop. He has a PhD in Machine Learning from UCL, was the founding engineer of Monolith AI, and has built speech systems at Google. For the last 4 years, he has led Humanloop and supported leading technology companies such as Duolingo, Vanta, and Gusto to build products with large language models. Raza was featured in the Forbes 30 Under 30 technology list in 2022, and Sifted recently named him one of the most influential Gen AI founders in Europe.

// Related Links

Websites: https://humanloop.com

~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Raza on LinkedIn: /humanloop-raza

Timestamps:

[00:00] Cracking Open System Failures and How We Fix Them

[05:44] LLMs in the Wild — First Steps and Growing Pains

[08:28] Building the Backbone of Tracing and Observability

[13:02] Tuning the Dials for Peak Model Performance

[13:51] From Growing Pains to Glowing Gains in AI Systems

[17:26] Where Prompts Meet Psychology and Code

[22:40] Why Data Experts Deserve a Seat at the Table

[24:59] Humanloop and the Art of Configuration Taming

[28:23] What Actually Matters in Customer-Facing AI

[33:43] Starting Fresh with Private Models That Deliver

[34:58] How LLM Agents Are Changing the Way We Talk

[39:23] The Secret Lives of Prompts Inside Frameworks

[42:58] Streaming Showdowns — Creativity vs. Convenience

[46:26] Meet Our Auto-Tuning AI Prototype

[49:25] Building the Blueprint for Smarter AI

[51:24] Feedback Isn’t Optional — It’s Everything

MLOps.community

Hosted by Demetrios, MLOps.community is a space for honest, meandering talks about the real work of making artificial intelligence systems actually work. This isn't about hype or theoretical papers; it's about the messy, practical, and often surprising journey of taking models from a notebook into a live environment. You'll hear from engineers and practitioners who are in the trenches, discussing the tools, the frustrations, and the occasional breakthroughs that define the day-to-day. The conversations are deliberately relaxed, covering everything from traditional machine learning pipelines to the new world of large language models and even the intangible "vibes" of team culture and process. Each episode peels back a layer on what "production" really means, whether that involves deploying a predictive service, managing an agentic system, or maintaining reliability as everything scales. Tuning into this podcast feels like grabbing a coffee with colleagues who aren't afraid to dig into the technical nitty-gritty while keeping the tone conversational and accessible. It's for anyone who builds, manages, or is just curious about the operational backbone that allows AI to deliver value, offering a grounded perspective often missing from the broader conversation.

Author: Demetrios Language: en-us Episodes: 100

Official website RSS

Podcast Episodes

[not-audio_url]

[/not-audio_url]

Does AgenticRAG Really Work?

12.12.2025

Duration: 1:01:39

Satish Bhambri is a Sr Data Scientist at Walmart Labs, working on large-scale recommendation systems and conversational AI, including RAG-powered GroceryBot agents, vector-search personalization, and transformer-based ad…

[not-audio_url]

[/not-audio_url]

How Sierra AI Does Context Engineering

10.12.2025

Duration: 1:04:03

Zack Reneau-Wedeen is the Head of Product at Sierra, leading the development of enterprise-ready AI agents — from Agent Studio 2.0 to the Agent Data Platform — with a focus on richer workflows, persistent memory, and hig…

[not-audio_url]

[/not-audio_url]

Overcoming Challenges in AI Agent Deployment: The Sweet Spot for Governance and Security // Spencer Reagan // #349

05.12.2025

Duration: 54:17

Spencer Reagan leads R&D at Airia, working on secure AI-agent orchestration, data governance systems, and real-time signal fusion technologies for regulated and defense environments.Overcoming Challenges in AI Agent Depl…

[not-audio_url]

[/not-audio_url]

Hardening Agents for E-commerce Scale: From RL Alignment to Reliability // Panel 2

02.12.2025

Duration: 29:16

Thanks to Prosus Group for collaborating on the Agents in Production Virtual Conference 2025.Abstract //The discussion centers on highly technical yet practical themes, such as the use of advanced post-training technique…

[not-audio_url]

[/not-audio_url]

Building Cursor: A Fireside Chat with VP Solutions Ricky Doar

27.11.2025

Duration: 26:44

Ricky Doar is the VP of Solutions at Cursor, where he leads forward-deployed engineers. A seasoned product and technical leader with over a decade of experience in developer tools and data platforms, Ricky previously ser…

[not-audio_url]

[/not-audio_url]

Relational Foundation Models: Unlocking the Next Frontier of Enterprise AI // Jure Leskovec // #348

25.11.2025

Duration: 49:00

Dr. Jure Leskovec is the Chief Scientist at Kumo.AI and a Stanford professor, working on relational foundation models and graph-transformer systems that bring enterprise databases into the foundation-model era.Relational…

[not-audio_url]

[/not-audio_url]

Context Engineering, Context Rot, & Agentic Search with the CEO of Chroma, Jeff Huber

21.11.2025

Duration: 44:55

Jeff Huber is the CEO of Chroma, working on context engineering and building reliable retrieval infrastructure for AI systems. Context Engineering, Context Rot, & Agentic Search with the CEO of Chroma, Jeff Huber // MLO…

[not-audio_url]

[/not-audio_url]

Reliable Voice Agents

18.11.2025

Duration: 38:21

Brooke Hopkins is the CEO of Coval, a company making voice agents more reliable. Reliable Voice Agents // MLOps Podcast #347 with Brooke Hopkins, Founder of Coval.Join the Community: https://go.mlops.community/YTJoinInGe…

[not-audio_url]

[/not-audio_url]

The Future of AI Operations: Insights from PwC AI Managed Services

14.11.2025

Duration: 41:27

Rani Radhakrishnan is a Principal at PwC US, leading work on AI-managed services, autonomous agents, and data-driven transformation for enterprises.The Future of AI Operations: Insights from PwC AI Managed Services // ML…

[not-audio_url]

[/not-audio_url]

GPU Uptime with VAST Data CTO

11.11.2025

Duration: 1:33:45

Andy Pernsteiner is the Field CTO at VAST Data, working on large-scale AI infrastructure, serverless compute near data, and the rollout of VAST’s AI Operating System.The GPU Uptime Battle // MLOps Podcast #346 with Andy…