Most teams assume better operational intelligence comes down to better tools. AWS's Michael Hausenblas — who works on Prometheus, Grafana and OpenTelemetry on AWS's open source observability team, and wrote Cloud Observability in Action — disagrees.
In this episode, he argues the real barrier is almost always organisational: teams that split "build" from "run" struggle no matter what they buy. He talks through why AWS rotates engineers through on-call instead of hiring dedicated operators, why blameless postmortems only work if leadership keeps repeating the message, how to think about hot versus cold data retention, and where AI actually helps — spotting patterns across postmortems, not automating the response.
Chapters
- 0:01:08 — Michael's path from Red Hat to leading AWS observability
- 0:02:42 — Why observability is just a subset of operational intelligence
- 0:06:14 — Operational data as an untapped product signal
You can't just buy a tool and everything will work. It's part of the solution, but you need to do much more than just the tool itself.
- 0:09:34 — The dev-vs-ops incentive problem, and AWS's on-call rotation fix
It's really about being very intentional about what you want to invest in, and then you will find the tools.
- 0:14:31 — Blameless postmortems and psychological safety
This is not about you making a mistake — this is about the fact that you were able to make that mistake, which means we as an organization have failed.
- 0:18:26 — How much you can share across teams before it breaks down
- 0:20:17 — What it really costs to store and act on operational data
- 0:24:28 — Where AI adds the most value: spotting patterns humans miss
Key takeaways
- Observability is a subset of a bigger idea Hausenblas calls "operational intelligence" — it's not just about collecting the right signals, it's what an organization does with incidents, learning, and process afterward.
- The biggest barrier most teams hit isn't a tooling or standards choice — it's organizational. Teams that separate "build" and "run" into different roles with opposite incentives struggle regardless of what they buy.
- AWS deliberately rotates its own engineers through on-call rather than hiring dedicated operators, specifically so the people writing the code also feel the pain of running it.
- Blameless postmortems only work if leadership actively reinforces that a mistake reveals an organizational gap, not an individual failure — and that message has to be repeated, not assumed.
- Not all observability data needs to be "hot": most of it only needs a short, expensive retention window for active debugging, with the rest moved to cheap cold storage for compliance or historical analysis.
- The most underrated use for AI in this space isn't automation — it's pattern-spotting across postmortems and tickets at a scale humans struggle to hold in their heads.