Lessons from Transcribing and Indexing 3.5 Million Podcasts with Arvid Kahl

Lessons from Transcribing and Indexing 3.5 Million Podcasts with Arvid Kahl

Author: Software Huddle July 8, 2025 Duration: 1:18:00
Big time guest today as Arvid Kahl joins us. Arvid is my favorite type of guest -- a deeply technical founder that can talk about both the technical and business challenges of a startup. Lots to enjoy from this episode. Arvid is known as the Bootstrapped Founder and has documented his path to selling Feedback Panda back in 2019. He's now building Podscan and sharing his journey as he goes. Podscan is a fascinating project. It's making the content of *every* podcast episode around the world fully searchable. He currently has 3.5 million episodes transcribed and adds another 30,000 - 50,000 episodes every day. This involves a ton of technical challenges, including how to get the best transcription results from the latest LLMs, whether you should use APIs from public providers or run your own LLMs, and how to efficiently provide full-text search across terabytes of transcription data. Arvid shares the lessons he's learned and the various strategies he's tried over the years. But there are also unique business challenges. For most technical businesses, your infrastructure costs grow in line with your customers. More customers == more data == more servers. With Podscan, Arvid has to index the entire podcast ecosystem regardless of his customers. This means a lot of upfront investment as he looks to grow his customer base. Arvid tells us how he's optimized his infrastructure to account for this unique challenge.

Every week on Software Huddle, Alex DeBrie and Sean Falconer sit down with a different expert from across the tech landscape. The conversations are less about quick tips and more about substantive discussions, digging into the real challenges and decisions behind building software, launching products, and navigating the industry's constant shifts. You'll hear from practitioners who have been in the trenches, offering perspectives that blend deep technical knowledge with hard-won business and entrepreneurial experience. Alex brings his specialized expertise as the author of The DynamoDB Book and an AWS Data Hero, while Sean contributes a unique viewpoint shaped by over two decades as an engineer, founder, and marketing executive, recognized as a Snowflake Data Superhero. Together, they create a space where complex topics in software development and technology trends become accessible and genuinely engaging. This podcast is for anyone who wants to move beyond surface-level news and understand the "why" behind the tools and strategies shaping our digital world. Tune in for a thoughtful huddle that feels more like a candid conversation between colleagues than a formal interview.
Author: Language: en-us Episodes: 79

Software Huddle
Podcast Episodes
Architecting for SaaS with Bill Tarr [not-audio_url] [/not-audio_url]

Duration: 1:09:53
This week on the show, we talk with Bill Tarr, Principal Solutions Architect at AWS SaaS Factory. He's a super thoughtful guy, expert in SaaS architecture and architectural patterns. We talk about tenancy, infrastructure…
High Performance Postgres with Andrew Atkinson [not-audio_url] [/not-audio_url]

Duration: 1:13:55
Database performance is likely the biggest factor in whether your application is slow or not, and yet too many developers don't take the time to properly understand how their database works. In today's episode, we have A…
Postgres for Search + Analytics with Philippe Noël [not-audio_url] [/not-audio_url]

Duration: 42:39
ParadeDB is Postgres for search and analytics. As Postgres continues to rise in popularity, the "Just Use Postgres'' movement is getting stronger and stronger. Yet there are still things that standard Postgres doesn't do…
AI Agents and Long Context Windows with Mark Huang [not-audio_url] [/not-audio_url]

Duration: 50:24
Today we have Mark Huang on the show. Mark has previously held roles in Data Science and ML at companies like Box and Splunk and is now the co-founder and chief architect of Gradient, an enterprise AI platform to build a…
Vector Databases with Bob van Luijt [not-audio_url] [/not-audio_url]

Duration: 47:19
Today we have Bob van Luijt, the CEO and founder of Weaviate on the show. Bob talks about building AI native applications and what that means, the role a vector database will play in the future of AI applications, and ho…
Akamai: From CDN to Full Cloud Provider with Talia Nassi [not-audio_url] [/not-audio_url]

Duration: 41:48
Today, we have Talia Nassi on the show. Talia’s been leading Developer Advocacy at Akamai. Akamai is in a really interesting space where they've been around for a long time, as a CDN provider, as a security provider, and…
Jamstack and Composable Web Architecture with Brian Rinaldi [not-audio_url] [/not-audio_url]

Duration: 53:59
Today we have Brian Rinaldi from LaunchDarkly on the show. This is the final episode of our in person coverage at the SHIFT Conference in Miami. And although Brian works at LaunchDarkly, we actually didn't talk at all ab…
Practical AI for LLMs with Emanuel Lacić [not-audio_url] [/not-audio_url]

Duration: 51:16
Today we have Emanuel Lacić on the show. He was in academia for a while. Now he’s been working at Infobip for the last couple of years, building some of this AI stuff and putting it into production. We picked his brain a…
Why Building an API for Email is Hard with Christine Spang [not-audio_url] [/not-audio_url]

Duration: 56:33
Today, on the show we have Christine Spang, Co-founder and CTO of Nylas. Christine was the keynote at the recent Shift Developer Conference in Miami, and we caught up with her there. Nylas is a unified API for email, cal…