GPU Uptime with VAST Data CTO

Author: Demetrios November 11, 2025 Duration: 1:33:45

Technology

Andy Pernsteiner is the Field CTO at VAST Data, working on large-scale AI infrastructure, serverless compute near data, and the rollout of VAST’s AI Operating System.

The GPU Uptime Battle // MLOps Podcast #346 with Andy Pernsteiner, Field CTO of VAST Data.Huge thanks to VAST Data for supporting this episode!

Join the Community:

https://go.mlops.community/YTJoinIn

Get the newsletter:

https://go.mlops.community/YTNewsletter

// Abstract

Most AI projects don’t fail because of bad models; they fail because of bad data plumbing. Andy Pernsteiner joins the podcast to talk about what it actually takes to build production-grade AI systems that aren’t held together by brittle ETL scripts and data copies. He unpacks why unifying data - rather than moving it - is key to real-time, secure inference, and how event-driven, Kubernetes-native pipelines are reshaping the way developers build AI applications. It’s a conversation about cutting out the complexity, keeping data live, and building systems smart enough to keep up with your models.

// Bio

Andy is the Field Chief Technology Officer at VAST, helping customers build, deploy, and scale some of the world’s largest and most demanding computing environments.

Andy has spent the past 15 years focused on supporting and building large-scale, high-performance data platform solutions. From humble beginnings as an escalations engineer at pre-IPO Isilon, to leading a team of technical Ninjas at MapR, he’s consistently been in the frontlines solving some of the toughest challenges that customers face when implementing Big Data Analytics and next-generation AI solutions.

// Related Links

Website: www.vastdata.com

https://www.youtube.com/watch?v=HYIEgFyHaxk

https://www.youtube.com/watch?v=RyDHIMniLro

The Mom Test by Rob Fitzpatrick: https://www.momtestbook.com/

~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

Join our Slack community

[https://go.mlops.community/slack]

Follow us on X/Twitter [@mlopscommunity](https://x.com/mlopscommunity) or [LinkedIn](https://go.mlops.community/linkedin)]

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Andy on LinkedIn: /andypernsteiner

Timestamps:

[00:00] Prototype to production gap

[00:21] AI expectations vs reality

[03:00] Prototype vs production costs

[07:47] Technical debt awareness

[10:13] The Mom Test

[15:40] Chaos engineering

[22:25] Data messiness reflection

[26:50] Small data value

[30:53] Platform engineer mindset shift

[34:26] Gradient description comparison

[38:12] Empathy in MLOps

[45:48] Empathy in Engineering

[51:04] GPU clusters rolling updates

[1:03:14] Checkpointing strategy comparison

[1:09:44] Predictive vs Generative AI

[1:17:51] On Growth, Community, and New Directions

[1:24:21] UX of agents

[1:32:05] Wrap up

MLOps.community

Hosted by Demetrios, MLOps.community is a space for honest, meandering talks about the real work of making artificial intelligence systems actually work. This isn't about hype or theoretical papers; it's about the messy, practical, and often surprising journey of taking models from a notebook into a live environment. You'll hear from engineers and practitioners who are in the trenches, discussing the tools, the frustrations, and the occasional breakthroughs that define the day-to-day. The conversations are deliberately relaxed, covering everything from traditional machine learning pipelines to the new world of large language models and even the intangible "vibes" of team culture and process. Each episode peels back a layer on what "production" really means, whether that involves deploying a predictive service, managing an agentic system, or maintaining reliability as everything scales. Tuning into this podcast feels like grabbing a coffee with colleagues who aren't afraid to dig into the technical nitty-gritty while keeping the tone conversational and accessible. It's for anyone who builds, manages, or is just curious about the operational backbone that allows AI to deliver value, offering a grounded perspective often missing from the broader conversation.

Author: Demetrios Language: en-us Episodes: 100

Official website RSS

Podcast Episodes

[not-audio_url]

[/not-audio_url]

Does AgenticRAG Really Work?

12.12.2025

Duration: 1:01:39

Satish Bhambri is a Sr Data Scientist at Walmart Labs, working on large-scale recommendation systems and conversational AI, including RAG-powered GroceryBot agents, vector-search personalization, and transformer-based ad…

[not-audio_url]

[/not-audio_url]

How Sierra AI Does Context Engineering

10.12.2025

Duration: 1:04:03

Zack Reneau-Wedeen is the Head of Product at Sierra, leading the development of enterprise-ready AI agents — from Agent Studio 2.0 to the Agent Data Platform — with a focus on richer workflows, persistent memory, and hig…

[not-audio_url]

[/not-audio_url]

Overcoming Challenges in AI Agent Deployment: The Sweet Spot for Governance and Security // Spencer Reagan // #349

05.12.2025

Duration: 54:17

Spencer Reagan leads R&D at Airia, working on secure AI-agent orchestration, data governance systems, and real-time signal fusion technologies for regulated and defense environments.Overcoming Challenges in AI Agent Depl…

[not-audio_url]

[/not-audio_url]

Hardening Agents for E-commerce Scale: From RL Alignment to Reliability // Panel 2

02.12.2025

Duration: 29:16

Thanks to Prosus Group for collaborating on the Agents in Production Virtual Conference 2025.Abstract //The discussion centers on highly technical yet practical themes, such as the use of advanced post-training technique…

[not-audio_url]

[/not-audio_url]

Building Cursor: A Fireside Chat with VP Solutions Ricky Doar

27.11.2025

Duration: 26:44

Ricky Doar is the VP of Solutions at Cursor, where he leads forward-deployed engineers. A seasoned product and technical leader with over a decade of experience in developer tools and data platforms, Ricky previously ser…

[not-audio_url]

[/not-audio_url]

Relational Foundation Models: Unlocking the Next Frontier of Enterprise AI // Jure Leskovec // #348

25.11.2025

Duration: 49:00

Dr. Jure Leskovec is the Chief Scientist at Kumo.AI and a Stanford professor, working on relational foundation models and graph-transformer systems that bring enterprise databases into the foundation-model era.Relational…

[not-audio_url]

[/not-audio_url]

Context Engineering, Context Rot, & Agentic Search with the CEO of Chroma, Jeff Huber

21.11.2025

Duration: 44:55

Jeff Huber is the CEO of Chroma, working on context engineering and building reliable retrieval infrastructure for AI systems. Context Engineering, Context Rot, & Agentic Search with the CEO of Chroma, Jeff Huber // MLO…

[not-audio_url]

[/not-audio_url]

Reliable Voice Agents

18.11.2025

Duration: 38:21

Brooke Hopkins is the CEO of Coval, a company making voice agents more reliable. Reliable Voice Agents // MLOps Podcast #347 with Brooke Hopkins, Founder of Coval.Join the Community: https://go.mlops.community/YTJoinInGe…

[not-audio_url]

[/not-audio_url]

The Future of AI Operations: Insights from PwC AI Managed Services

14.11.2025

Duration: 41:27

Rani Radhakrishnan is a Principal at PwC US, leading work on AI-managed services, autonomous agents, and data-driven transformation for enterprises.The Future of AI Operations: Insights from PwC AI Managed Services // ML…

[not-audio_url]

[/not-audio_url]

The Evolution of AI in Cyber Security // Jeff Schwartzentruber // #344

04.11.2025

Duration: 35:14

Dr. Jeff Schwartzentruber is a Senior Machine Learning Scientist at eSentire, working on anomaly detection pipelines and the use of large language models to enhance cybersecurity operations.The Evolution of AI in Cyber S…