We Cut LLM Latency by 70% in Production

We Cut LLM Latency by 70% in Production

Author: Demetrios April 10, 2026 Duration: 1:05:20

Maher Hanafi is an engineering leader who went from zero AI experience to self-hosting LLMs at enterprise scale — managing GPU costs, optimizing inference with TensorRT LLM, and building an AI platform for HR tech. In this conversation, he breaks down exactly how his team cut latency by 70%, reduced GPU spend through counterintuitive scaling strategies, and navigated the messy reality of taking AI from proof-of-concept to production.


How We Cut LLM Latency 70% With TensorRT in Production // MLOps Podcast #369 with Maher Hanafi, SVP of Engineering at Betterworks


Key topics covered:

The AI Iceberg — Why the invisible work behind AI (performance, latency, throughput, cost, accuracy) is harder than building the features themselves

GPU Cost Optimization — How upgrading to more expensive GPUs actually saved money by reducing total runtime hours

TensorRT LLM Deep Dive — Rewiring neural networks to match GPU architecture for 50-70% latency reduction

Cold Start Solutions — Using AWS FSx, baking models into container images, and cutting minutes off spin-up times

KV Cache & In-Flight Batching — Why using one model per GPU with maximum KV cache beats cramming multiple models together

Scheduled & Dynamic Scaling — Pattern-based scaling for HR tech workloads (nights, weekends, end-of-quarter spikes)

Verticalized AI Platform — Building horizontal AI infrastructure that serves multiple HR product verticals

AI Engineering Lab — How junior vs. senior engineers adopted AI coding tools differently, and the cultural shift that followed

Agentic Coding in Practice — Navigating AI coding agent costs, quality control, and redefining the SDLC

Chinese Models & Compliance — Why enterprise customers block DeepSeek/Qwen and the geopolitics of model training data


This episode is for engineering leaders building AI in production, MLOps engineers optimizing GPU infrastructure, and anyone navigating the gap between AI demos and enterprise-scale deployment.


Links & Resources:

TensorRT LLM: https://github.com/NVIDIA/TensorRT-LLM

NVIDIA Run: ai Model Streamer (cold start optimization): https://developer.nvidia.com/blog/reducing-cold-start-latency-for-llm-inference-with-nvidia-runai-model-streamer/

vLLM vs TensorRT-LLM comparison: https://northflank.com/blog/vllm-vs-tensorrt-llm-and-how-to-run-them


Timestamps:

[00:00] Optimizing GPU Usage and Latency

[00:21] Learning AI as Leadership

[04:34] AI Cost Centers

[13:56] Throughput and Infrastructure Efficiency

[18:10] Scaling and Unit Economics

[24:14] Championing AI ROI

[36:11] Queue to Value Engine

[41:30] Failed Product Features

[46:12] Agentic Engineering Costs

[58:49] AI Self-Hosting in Engineering

[1:04:40] Wrap up


Hosted by Demetrios, MLOps.community is a space for honest, meandering talks about the real work of making artificial intelligence systems actually work. This isn't about hype or theoretical papers; it's about the messy, practical, and often surprising journey of taking models from a notebook into a live environment. You'll hear from engineers and practitioners who are in the trenches, discussing the tools, the frustrations, and the occasional breakthroughs that define the day-to-day. The conversations are deliberately relaxed, covering everything from traditional machine learning pipelines to the new world of large language models and even the intangible "vibes" of team culture and process. Each episode peels back a layer on what "production" really means, whether that involves deploying a predictive service, managing an agentic system, or maintaining reliability as everything scales. Tuning into this podcast feels like grabbing a coffee with colleagues who aren't afraid to dig into the technical nitty-gritty while keeping the tone conversational and accessible. It's for anyone who builds, manages, or is just curious about the operational backbone that allows AI to deliver value, offering a grounded perspective often missing from the broader conversation.
Author: Language: en-us Episodes: 50

Agentic Conversations (formally mlops.community)
Podcast Episodes
The Dark Side of MCP Servers [not-audio_url] [/not-audio_url]

Duration: 1:09:59
Sam Partee (CTO & co-founder of Arcade.dev) and Nate Barbettini (Founding Engineer at Arcade.dev) sit down at the MCP Dev Summit to unpack what nobody wants to admit about the Model Context Protocol: the security model i…
Sandboxing, Agent Harnesses, and Agent Teamwork [not-audio_url] [/not-audio_url]

Duration: 1:19:53
Shahram Anver is the Co-Founder and CEO of Cleric, the autonomous AI SRE that investigates and root-causes production issues like an experienced teammate — often in under two minutes. Before Cleric, Shahram led MLOps, De…
MCP Servers Are Becoming the UI for AI Agents [not-audio_url] [/not-audio_url]

Duration: 47:21
Naseem Al-Naji is the co-founder of MCPcat.io and the creator of Opal — a builder with deep roots in privacy-first developer tooling. In this conversation, he breaks down why MCP servers have become a black box in produc…
Agents & the $40M Bet on Multiplayer AI [not-audio_url] [/not-audio_url]

Duration: 1:20:46
Stanislas Polu is Co-Founder & CTO of Dust — the enterprise AI agent platform used by 51,000 workers at 3,000+ companies. Before Dust, he spent three years on OpenAI's research team under Ilya Sutskever, working on mathe…
From Single-Player to Multi-Player: Operating AI Agents at Scale [not-audio_url] [/not-audio_url]

Duration: 55:54
James Everingham is the CEO and Co-founder of Guild.ai — the AI agent control plane for production teams. With roots at Netscape, Instagram (Head of Engineering), and Meta (Head of Dev Infra, leading a 1,000-person org),…
The Control-vs-Magic Spectrum Building Agents [not-audio_url] [/not-audio_url]

Duration: 43:18
Thiago Cardoso is the Director of Data & AI at iFood and the architect behind iFood Pago's AI agent platform. This fintech system serves millions of restaurants across Brazil through WhatsApp and the iFood app. In this e…
Logs Are All You Need: Rethinking Observability with AI Agents [not-audio_url] [/not-audio_url]

Duration: 46:39
Sherwood Callaway is the founder of Sazabi (YC P26), the AI-native observability platform built for engineering teams who ship fast. He previously founded and exited a YC company — now he's back, betting that logs are al…
AI Is Fast. AI Projects Are Slow. Let's Fix That. [not-audio_url] [/not-audio_url]

Duration: 56:47
Joe Maionchi (Co-founder & COO) and Rod Christensen (Co-founder & Chief Architect) of RocketRide join the MLOps Community to walk through AIDE — the AI Integrated Development Environment. RocketRide is an open-source AI…
Architecting Modern AI Systems: Platforms, Agents, and Integration [not-audio_url] [/not-audio_url]

Duration: 56:59
BuzzHPC Roundtable episode: Architecting Modern AI Systems: Platforms, Agents, and Integration Join the Community: https://go.mlops.community/YTJoinInGet the newsletter: https://go.mlops.community/YTNewsletterMLOps GPU G…