
Ganesh Datta
HostCTO & Co-founder of Cortex

Steve Flanders
Senior Director of Engineering at Splunk
August 13, 2026
In This Episode
Steve Flanders is a Senior Director of Engineering at Splunk, now part of Cisco, where he oversees the Splunk Observability Cloud platform. He has spent nearly two decades in observability and was a founding member of OpenCensus and OpenTelemetry, helping build the service that became the OpenTelemetry Collector. He wrote the book on it, literally: Mastering OpenTelemetry and Observability. He's now also leading AI transformation for his business unit, which he digs into during the episode.
Steve joins Cortex CTO Ganesh Datta to trace where the real constraints in software delivery are moving now that AI writes more of the code. They get into why review and rework break first, why "garbage in, garbage out" now governs observability, and how to instrument for a world where an agent, not a human, is querying your telemetry. Steve closes with his advice for teams under pressure to just ship: fix one thing at a time.
You’ll learn
Code was never the bottleneck, it was the most expensive constraint. Senior engineering talent is the premium line item, and the bet is that AI commoditizes the junior work. But generating code was never the slow part of the SDLC. Getting it to production is, and that's where AI has barely reached.
Kill one constraint and the next one surfaces. Post-code, two things break first in the AI era: review time climbs and DORA's rework rate rises.
Do you even need a PRD in an AI world? The stages before code (product and engineering requirements, architecture review, security and compliance) are mostly manual, and speeding up only the coding just moves the slowdown upstream. Either infuse AI into those stages or redesign them.
Garbage in, garbage out applies to your telemetry too. Observability vendors are bolting on AI chat assistants, but if the underlying data isn't correlated and lacks the right metadata, you're asking AI to search over nonsense.
Define intent upfront or don't bother. Whether it's an engineering requirements doc or a telemetry spec, AI can't hit an outcome you never specified.
Start with a skateboard. You won't convert everything to OpenTelemetry overnight, and you don't have to.
Quotes
If you're ingesting telemetry data that isn't well correlated, doesn't have the right metadata, you're asking AI to search over nonsense.
Steve Flanders
Senior Director of Engineering at Splunk
If you don't define the intent correctly to the AI, there's no way you're going to achieve the outcome you asked it for in the first place.
Steve Flanders
Senior Director of Engineering at Splunk
This isn't a new problem. AI just makes the cracks that have always been there much more apparent.
Steve Flanders
Senior Director of Engineering at Splunk
Just because you have traces, metrics, and logs doesn't mean you have observability. What are you actually instrumenting?
Steve Flanders
Senior Director of Engineering at Splunk
You're not trying to build the car and only ship it when it's fully ready. If I can get you a skateboard quicker, I'll get you a skateboard so you can get going, then a bicycle next, and eventually I’ll get you a car.
Steve Flanders
Senior Director of Engineering at Splunk
Timestamps
(01:14)
Steve's background, two decades in observability, now leading AI transformation for his business unit.
(01:58)
The OpenTelemetry origin story: Omnition, OpenCensus, and the Collector.
(02:27)
"Code was never the bottleneck, it was the most expensive constraint"
(05:13)
Theory of constraints, and the two post-code shifts: review time on bigger PRs and DORA's rework rate.
(07:06)
Beyond engineering, PRDs, ERDs, security reviews, and whether you still need a PRD.
(09:02)
Observability without good guardrails: rising failure rates and customer-found defects.
(10:30)
Garbage in, garbage out, and why AI assistants fail on poorly correlated telemetry.
(11:41)
Why metrics and logs are useless if you can't stitch them together.
(13:28)
Designing observability upfront, as part of the feature instead of after.
(16:27)
What "good" looks like: standardizing on OpenTelemetry and semantic conventions.
(21:27)
The future of dashboards and alerts in an agentic, MCP-driven world.
(23:27)
Where to start when you've shipped code you don't fully understand.
(28:22)
Handling "just ship" pressure with iterative delivery, and the brownfield OTel Collector shortcut.
Other episodes
July 30, 2026Your Platform is your Business, Encoded onto your Infra, with Syntasso's Abby Bangser
Abby Bangser
Principal Engineer at Syntasso
July 16, 2026Trust is for People, Confidence is for Tools: Tealium's Dr. Martin Nettling on Reviewing AI-generated Work
Martin Nettling
Senior Director of Engineering
July 9, 2026Your Ops Review is Theater, and That's the Point: Aleks Rudzitis on Turning Reliability into a Shared Value
Aleks Rudzitis
Principal Engineer

