Most teams assume the path from monitoring to real operational intelligence runs through better tools. Michael Hausenblas, who has spent years building observability products at AWS, says the tools were never really the bottleneck.
Michael Hausenblas is a Solution Engineering Lead on AWS's open source observability team, where he works on Prometheus, Grafana and OpenTelemetry. Before Amazon he worked at Red Hat, Mesosphere (now D2iQ) and MapR, and spent close to a decade in academic applied research. He's also the author of the book "Cloud Observability in Action" and writes the o11y.news newsletter.
Ask most engineering teams what's standing between them and better operational intelligence, and the answer is usually a tool: a better dashboard, a new platform, a bigger budget for data retention. Hausenblas has heard that answer from customers for years, and his own experience points somewhere else entirely — toward the incentives that shape how people on a team actually behave.
Hausenblas draws a clear line between three overlapping ideas. Monitoring means knowing upfront what you want to ask — a threshold, an alert, a known failure mode. Observability goes further: it gives you the data to ask ad hoc questions when something breaks that you didn't anticipate. But both of those, in his framing, sit inside something larger still, which he and this podcast call operational intelligence — everything an organization does with that data once it has it, including the processes around incidents, learning, and continuous improvement.
That distinction matters because it's easy to stop at observability and assume the job is done. Hausenblas points to service level objectives as a concrete example of the gap: an SLO is not just another metric sitting next to CPU utilization, it's a measure of how your service is actually perceived from the outside — by the customer, not the infrastructure. A system can be technically healthy by every internal metric and still be failing the people using it.
The same logic applies to postmortems. AWS calls its version a Correction of Error, but the label matters less than the discipline behind it: taking the time after an incident to ask why, repeatedly, until the underlying pattern is visible — and treating that as a deliberate practice rather than paperwork nobody reads.
This is not about you making a mistake — this is about the fact that you were able to make that mistake, which means we as an organization have failed.
Most teams treat operational data as something you check when things break and otherwise ignore. Hausenblas thinks that undersells it badly. The same data that helps you fix an incident faster is often a source of genuine product insight — what he describes as looking for "dogs not barking": gaps in usage that hint at unmet demand, or patterns that point toward a feature nobody had explicitly asked for yet.
That reframing changes what observability data is for. Instead of treating it purely as a cost center for firefighting, teams that mine it for product signal are getting a second use out of infrastructure they're already paying for. Hausenblas is candid that this angle is chronically undervalued — most organizations treat their operational tooling as a baseline necessity rather than a lens onto where the product itself could go next.
You can't just buy a tool and everything will work. It's part of the solution, but you need to do much more than just the tool itself.
Ask Hausenblas what the biggest barrier to operational intelligence actually is, and the answer isn't technical. Talking to customers daily, he finds the obstacle is almost always organizational: a classic split where developers are incentivized to ship new features and operators are incentivized to avoid any change that might introduce risk. Push those two groups far enough apart and you get a "throw it over the wall" dynamic where nobody feels responsible for the whole picture.
AWS's own answer is structural rather than cultural in the abstract sense — it changes who's actually on the hook. Engineers rotate through on-call rather than handing that job permanently to a separate operations team. The weeks they're not on call, they're building; the weeks they are, they're living with whatever they shipped. That rotation builds the incentive to instrument well directly into the role, because the same person who under-instruments a service is the one who'll be paged for it later.
Hausenblas is careful not to oversell this as a universal fix, but he's equally clear that culture isn't something you either have or don't: it's something you can design, the way you'd plan a garden, by being deliberate about incentives rather than assuming good people alone will sort it out.
It's really about being very intentional about what you want to invest in, and then you will find the tools.
The AWS practice Hausenblas keeps returning to is what happens after something breaks. A blameless postmortem sounds simple in theory and is genuinely hard to sustain in practice, because the instinct to find who made the mistake is strong. His framing for why that instinct is counterproductive: if someone was able to make that mistake, the organization — not the individual — has failed, because the system should have made that mistake harder to make in the first place.
That message doesn't stick on its own. Hausenblas is explicit that it has to come from leadership, repeatedly, or the blameless framing collapses back into quiet finger-pointing the moment something expensive breaks. The payoff is also not immediate — it's what he calls a long game, where the value of a genuinely blameless culture shows up over months and years of people being willing to surface problems early rather than hide them.
Inside AWS, that safety shows up in how incidents are actually run: engineers focus purely on diagnosing and fixing the technical problem, while someone else handles customer communication — over-communicating deliberately, confirming understanding back and forth, because sloppy communication during a live incident compounds the confusion rather than reducing it.
Cost is a real constraint, and Hausenblas doesn't pretend otherwise — but he draws a sharper line than most teams do between what's actually expensive and what just feels that way. A tool's monthly price is easy to quantify. The opportunity cost of the time spent adopting it is much harder, and he thinks that's where most of the real cost analysis breaks down.
On data specifically, his advice is to be deliberate about what's "hot" and what isn't. Most debugging only needs a short online retention window — 30 to 100 days is common — kept in expensive, fast storage. Everything else, kept for compliance or long-term historical analysis, can move to a cheap cold tier like object storage, where you give up instant query speed but keep the data for a fraction of the cost. Being explicit about that policy, rather than defaulting everything to hot storage out of inertia, is itself part of doing operational intelligence well.
He also pushes back on blindly copying hyperscale practices. Benchmarking against AWS, Google or Microsoft doesn't tell a ten-person team much, because the constraints are entirely different. A better reference point is an organization of similar size and maturity — close enough in context that their trade-offs are actually comparable to yours.
Hausenblas is enthusiastic about AI's role in this space, but specific about where. The strongest opportunity, in his view, isn't automating responses — it's pattern recognition at a scale humans simply aren't built for. Manually reading through hundreds of postmortems or support tickets looking for a recurring root cause is exactly the kind of task where AI keeps surfacing insights he says genuinely surprised him, even before large language models made this a mainstream conversation — he points back to early open source machine learning tools he used over a decade ago as evidence this value isn't new, just newly accessible.
His caution is about scope, not potential: the industry is still working out where AI adds the most value, and applying it everywhere indiscriminately is a worse strategy than being deliberate about which specific, high-volume pattern-recognition problems it's actually suited to solve right now.
Michael Hausenblas is a Solution Engineering Lead on AWS's open source observability team, working on Prometheus, Grafana and OpenTelemetry. Before Amazon he worked at Red Hat, Mesosphere (now D2iQ) and MapR, and spent close to a decade in academic applied research. He's the author of Cloud Observability in Action (Manning) and writes the o11y.news newsletter.
Most of what Hausenblas describes — spotting patterns across incidents, being deliberate about what data actually gets watched — starts with being able to see it all in one place. If that's still scattered across your organization, you can start using SquaredUp for free.
Getting started with SquaredUp is free and easy.