Ship It Conversations: Mat Ryer of Grafana Labs on AI Observability, Agents, Evals, and Operating AI in Production

Ship It Conversations: Mat Ryer of Grafana Labs on AI Observability, Agents, Evals, and Operating AI in Production

Author: Teller's Tech - DevOps, SRE and Cloud Podcast July 20, 2026 Duration: 45:06

This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.

In this Ship It Conversations episode, I talk with Mat Ryer of Grafana Labs about AI observability, production agents, evals, telemetry cost, guardrails, and what changes once AI moves beyond demos and into systems teams actually depend on.

Mat is Senior Director of AI at Grafana Labs, where he focuses on how AI fits into observability and production systems.

We talk about Grafana Assistant, why AI observability is not just logs, latency, and HTTP 200s, and how teams can measure whether agents are actually helping. Mat gets into evals, LLM-as-judge patterns, traces as a way to think about conversations, user feedback, tool choice, model changes, and the cost of collecting new telemetry.

We also dig into UX and trust. If an AI assistant gives you a wall of text, you still have to decide whether to believe it. If it can show the graph, deep link into Grafana, apply filters, and expose the source data, that becomes a much more useful operating experience.

The big takeaway: start small, enhance workflows you already have, build feedback loops, and treat production AI like something you actually have to operate.

Highlights

• Why AI demos are easy, but production AI is harder

• Why agents need observability, evals, and guardrails

• How Grafana thinks about AI Assistant and AI observability

• Where LLM-as-judge patterns, traces, tool calls, and feedback fit

• Why telemetry cost problems may repeat with AI workloads

• Why UX matters when operators need to trust the answer

• Where AI can help SRE and platform teams today

Links

Mat Ryer on LinkedIn: https://www.linkedin.com/in/matryer/

Mat Ryer on GitHub: https://github.com/matryer

Grafana Labs: https://grafana.com

Grafana Assistant: https://grafana.com/products/cloud/ai-assistant/

Grafana AI Observability: https://grafana.com/docs/grafana-cloud/machine-learning/ai-observability/

Grafana Adaptive Telemetry: https://grafana.com/products/cloud/adaptive-telemetry/

Grafana MCP server: https://github.com/grafana/mcp-grafana

OpenTelemetry: https://opentelemetry.io

Prometheus: https://prometheus.io

Grafana Loki: https://grafana.com/docs/loki/latest/

More episodes and show notes: https://shipitweekly.fm

On Call Brief: https://oncallbrief.com


For anyone building or running modern systems, the sheer volume of news, tools, and incident reports can be overwhelming. Ship It Weekly cuts through that noise. This isn't a surface-level scan of headlines. Host Brian Teller digs into the latest significant outages, major software releases, and insightful post-mortems, focusing squarely on the practical implications for DevOps, SRE, and platform engineering work. Each episode of the podcast breaks down a couple of key stories, providing the crucial context often missing from tech news. You'll hear analysis that translates events into actionable insights, answering the "so what?" for your own infrastructure and processes. The show also includes a quick rundown of tools or updates actually worth your attention, saving you hours of browsing. The tone is direct and informed, favoring depth over breadth. It’s designed for engineers and technical leaders who need a concise, reliable filter for the week's most relevant developments. Listen to this podcast for a focused recap that prioritizes what actually matters, delivered without fluff. You get the news, plus the necessary interpretation to understand how it might affect your systems, your team, and your on-call rotation. It's a weekly briefing that respects your time while aiming to make you more effective.
Author: Language: English Episodes: 50

Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News
Podcast Episodes
GitHub Outage, PleaseFix Agentic Browser Vulnerability, AWS Certificate Manager Drops Email Validation, Cloudflare TypeScript CI Workflows, AI Observability Consolidation, and the Hidden Cost of “Simple” Platform Changes [not-audio_url] [/not-audio_url]

Duration: 17:39
This week on Ship It Weekly: GitHub suffers another widespread outage affecting the web interface, APIs, Actions, authentication, Copilot, and other critical developer workflows. Zenity Labs demonstrates PleaseFix attack…