AI and Proactive Reliability with Kolton Andrus

Author: Software Huddle April 8, 2026 Duration: 55:11

Technology

Today we're talking with Kolton Andrus, the Founder and CEO of Gremlin, about what happens to reliability when AI is writing most of the code. Kolton helped build the Chaos Engineering practice of both Amazon and Netflix before starting Gremlin. In our conversation we talk about scar tissue, the intuition engineers develop from being woken up at 3:00 AM to fix production outages and how AI doesn't have any of it. It generates code in an afternoon that maybe took a team previously weeks to build, but none of those painful lessons come along for the ride. We dig into why 10x more code might mean 10x more failures. The concept of reliability guardrails, think ethical guardrails, but for keeping your systems up. Why you still have to test in production no matter how good your staging environment is? How Gremlin is rethinking their product for the world where agents, not engineers, are essentially the primary users.And why we're entering a painful, narrow part of the hourglass before AI gets good enough to handle all of this on its own.

Software Huddle

Every week on Software Huddle, Alex DeBrie and Sean Falconer sit down with a different expert from across the tech landscape. The conversations are less about quick tips and more about substantive discussions, digging into the real challenges and decisions behind building software, launching products, and navigating the industry's constant shifts. You'll hear from practitioners who have been in the trenches, offering perspectives that blend deep technical knowledge with hard-won business and entrepreneurial experience. Alex brings his specialized expertise as the author of The DynamoDB Book and an AWS Data Hero, while Sean contributes a unique viewpoint shaped by over two decades as an engineer, founder, and marketing executive, recognized as a Snowflake Data Superhero. Together, they create a space where complex topics in software development and technology trends become accessible and genuinely engaging. This podcast is for anyone who wants to move beyond surface-level news and understand the "why" behind the tools and strategies shaping our digital world. Tune in for a thoughtful huddle that feels more like a candid conversation between colleagues than a formal interview.

Author: Software Huddle Language: en-us Episodes: 79

Official website RSS

Podcast Episodes

[not-audio_url]

[/not-audio_url]

Rewriting in Rust + Being a Learning Machine with AJ Stuyvenberg

06.05.2025

Duration: 1:21:36

Today's guest is AJ Stuyvenberg, a Staff Engineer at Datadog working on their Serverless observability project. He had a great article recently about how they rewrote their AWS Lambda extension in Rust. It's a really int…

[not-audio_url]

[/not-audio_url]

Software Reliability Agents with Amal Kiran

29.04.2025

Duration: 51:07

So if you're writing code or keeping systems running, you probably know the drill. Late night pages, chasing down weird bugs, dealing with alert storms. It's tough! It costs money when things break, and honestly, nobody…

[not-audio_url]

[/not-audio_url]

From ORM to Infra: Prisma Postgres with Søren Bramer Schmidt

22.04.2025

Duration: 1:02:22

Today we have Søren from Prisma on the show. Prisma has been the most popular ORM in the TypeScript world for a while, and now they’re moving more into hosted infrastructure. We spend a lot of time talking about their ne…

[not-audio_url]

[/not-audio_url]

Fast Inference with Hassan El Mghari

08.04.2025

Duration: 53:06

Today we have Hassan back on the show. Hassan was one of our first guests for Huddle when he was working at Vercel, but since then, he's joined Together AI, one of the hottest companies in the world. They just raised a m…

[not-audio_url]

[/not-audio_url]

Seattle Startups, AI’s Future & Big Acquisitions with Yujian Tang

14.03.2025

Duration: 1:02:54

Today on the show, we talked with Yujian Tang. He was on the show previously when he worked at Zilliz, when we talked about vector databases and RAG. He's since branched out on his own, building the tech startup scene in…

[not-audio_url]

[/not-audio_url]

Faster & Cheaper on PlanetScale Metal with Sam Lambert

12.03.2025

Duration: 1:19:43

Today, we have Sam Lambert back on the show! Sam is the CEO of PlanetScale, and if you follow him on X, you know he’s one of the sharpest voices in the database space—cutting through the hype with deep experience and a n…

[not-audio_url]

[/not-audio_url]

Redis but Faster With Roman Gershman

04.03.2025

Duration: 1:00:51

Redis is consistently one of the most beloved pieces of infrastructure for developers. And in the last few years, we've seen a number of new Redis-compatible projects that aim to improve on the core of Redis in some way.…

[not-audio_url]

[/not-audio_url]

Lessons from Building Tagged.com + AI-Driven Database Optimization with Johann Schleier-Smith

11.12.2024

Duration: 56:13

Today, we’re joined by Johann Schleier-Smith. Johann co-founded Tagged during the early days of social media, a time when building scalable systems for the web was uncharted territory. Back then, cloud computing didn’t e…

[not-audio_url]

[/not-audio_url]

Building + Evolving Sentry's Architecture and Funding Open Source with David Cramer

13.11.2024

Duration: 1:13:06

Today, we have David Cramer on the show. David is one of the co-founders of Sentry, an application monitoring tool that's one of the most widely-adopted tools for developers. Sentry does over 300,000 events per second on…

[not-audio_url]

[/not-audio_url]

Deep Dive into Inference Optimization for LLMs with Philip Kiely

06.11.2024

Duration: 1:04:05

Today we have Philip Kiely from Baseten on the show. Baseten is a Series B startup focused on providing infrastructure for AI workloads. We go deep on Inference Optimization. We cover choosing a model, discuss the hype a…