The Five Levels of the Enterprise AI Software Factory
Back to Blog
AIBest PracticeDRIVEEVOLVE

The Five Levels of the Enterprise AI Software Factory

The Five Levels of the Enterprise AI Software Factory
Anish Dhar

Anish Dhar

CEO & Co-founder

Ganesh Datta

Ganesh Datta

CTO & Co-founder

October 8, 2026

Scroll LinkedIn for ten minutes and you'd think every engineering organization already runs an autonomous SDLC, with agents writing, reviewing, and shipping code while humans watch. Inside large enterprises, the picture looks different. Coding agents there have to work around sensitive customer data, recurring compliance audits, a larger attack surface, downtime that costs millions, and a CFO who wants to know what last year's token spend bought.

Over the past two years, we've watched hundreds of enterprises roll out AI across their engineering organizations. Some rollouts went well. Others got chaotic fast, usually because leaders treated access to better models as evidence of a more mature organization. The five levels below come from those patterns, and we presented them together in our EVOLVE 2026 keynote.

Model capability does not equal organizational maturity

Many of the AI adoption measurement frameworks shared online run from autocomplete to the "dark factory," and they're mostly a proxy for how good the models and harnesses have become. They say very little about whether an organization is ready for an AI rollout, especially one with thousands of engineers that has to protect reliability and security while keeping costs in check.

For a large enterprise, breadth matters as much as depth. A platform team running background agents on every dependency upgrade is impressive, but it tells you little about the other 200 teams. That scale also changes the path to an autonomous software factory: an enterprise running thousands of services owned by hundreds of teams, under compliance requirements, has far more at stake than a startup when something breaks, so every step toward autonomy has to be staged and measured.

We saw a similar pattern with cloud migrations. Hyperscalers were ready years before most enterprises moved, and plenty of large organizations are still migrating more than a decade later. Cost, reliability, and security concerns slowed migrations down, but the hardest part was change management: getting thousands of people to work in a new way.

AI rollouts for enterprise organizations need the same change management: an executive sponsor, a phased rollout that earns trust team by team, and deliberate changes to the processes that strain first, like code review, incident response, and ownership.

The five levels of the enterprise AI software factory at a glance

We modeled the levels loosely on the scale used to rate self-driving cars, and built them specifically for enterprises with thousands of engineers or complex environments.

Level

Description

How to get there

Level 1: Evangelists

Influential senior engineers experimenting with large, flexible token budgets

Pick engineers with a track record of org-wide impact, give light guidance, and keep the focus on value

Level 2: Automations

Task-specific flows like code review, vulnerability and package upgrades, and incident support

Productionize the best level 1 experiments and show measurable ROI

Level 3: Human-driven, agent-assisted

Engineers working alongside coding agents like Claude Code, Codex, and Cursor

Roll out in phases gated on readiness evidence, with clear ownership and a recurring OpEx review

Level 4: Human-assisted, agent-driven

Mostly automated reviews, verification harnesses, and deploys, with humans stepping in at set points

Use OpEx signals to find teams and services ready for less oversight, and keep iterating on your harnesses

Level 5: Autonomous software delivery

Spec in, production code out, with humans designing the factory and the architecture

Out of reach for enterprises today. Prepare for it instead of building it.

Level 1: Evangelists

What it looks like: A handful of senior engineers with org-wide influence experiment with new models on real problems, then share what works through demos and lunch-and-learns.

What can go wrong: Experiments chase whatever shipped that week, and the excitement never turns into anything the rest of the organization can use.

How to get it right: Pick engineers with a track record of org-wide impact, give them generous token budgets and light guidance, and keep their attention on outcomes the business cares about.

Level 2: Automations

What it looks like: Narrow production workflows for tasks you understand well, like a code review agent, automated vulnerability and package upgrades, or an incident support agent.

What can go wrong: Automating a workflow you don't understand, or one whose value you can't measure, leaves leadership with nothing to justify the next round of investment.

How to get it right: Productionize the best level 1 experiments, keep the scope small enough to monitor for quality and cost, and report the ROI. Automations that take toil off engineers who haven't touched AI yet also make the next level an easier sell.

Level 3: Human-driven, agent-assisted

What it looks like: Engineers across the organization work alongside coding agents like Claude Code, Codex, or Cursor. Many enterprises are here today, and most will stay here for several years.

What can go wrong: Handing an agent to every engineer who asks multiplies code volume faster than CI, code review, incident response, and cost controls can absorb it. Ownership gaps that were tolerable at a few hundred services surface first. For every repository an agent helped create, someone needs to know who created it, whether they still work here, what depends on it, and whether it's safe to delete.

How to get it right: Roll agents out in phases, gated on evidence that a team is ready. Scorecards give you a quantitative read on each team's AI readiness and its security and reliability maturity. Back them with clear ownership in your catalog and a recurring OpEx review that asks what's working, what isn't, and what you're changing.

Level 4: Human-assisted, agent-driven

What it looks like: Agents do most of the work, and humans step in at predetermined points. In a typical setup, low-risk PRs merge after automated verification and higher-risk PRs still go to a human reviewer.

What can go wrong: Leaders assume the whole organization is ready because most teams reached level 3, and pull oversight from teams and services that still need it. Teams that don't trust the automation push back, and that fear stalls progress.

How to get it right: Start by asking why you have human-in-the-loop processes like code review in the first place. Reviews exist to catch specific risks, so build automated ways to catch them. At Cortex, we've spent two years building a PR verification system that decides which PRs can ship without human review. Use OpEx signals to find the teams, services, and use cases ready for less oversight. None of this is unique to AI, but agents make these old bottlenecks impossible to ignore.

Level 5: Autonomous software delivery

What it looks like: A spec goes in and production code comes out. Humans design the factory itself: the harnesses, the target architecture, the specs, and the feedback loops that let it improve.

What can go wrong: Trying to build it now. Reaching this level requires a step-function improvement in models, and every existing constraint gets harder once humans leave the loop. We don't think any enterprise can run this way today.

How to get it right: Prepare for it. A level 2 automation that has run reliably on a team already at level 4 overall is a reasonable first candidate for a narrow software factory. Expect to run several factories eventually, since an iOS app, a web app, and embedded firmware each need their own line.

How to tell which level you're at

Different parts of your organization will sit at different levels at the same time, and so will different use cases. A team at level 4 can depend on a code review bot that is itself a level 2 automation. That's expected, because each level teaches you what will break at the next one. Jump from level 1 to level 4 and you discover those failures in production.

Judge your level the way you'd judge a promotion. Engineers usually operate at the next level for a while before their title catches up, and organizations work the same way: if it feels like you've been at level 4 for a year, you've probably just finished level 3. Licensing coding agents for every engineer doesn't put you at level 3, either. You get there when the SDLC around those agents can handle what they produce.

Every move between levels is a cultural change, so each one needs an executive sponsor and a clear read on where you stand. Benchmark against your own organization last quarter.

Where Cortex is today

Cortex is at level 3 overall and pushing into level 4 across several use cases. Our engineering team is about 30 people, which makes every step easier for us than it will be for an organization with thousands of engineers, so treat our experience as a reference point.

To get here, we invested in experimentation budgets for every developer, automated code review, automated risk scoring, and the PR verification system described above. Most of our level 4 work runs on engrams, a self-hosted harness we built to run coding agents in any cloud, one Firecracker microVM per agent. It powers our code review agent, bug-fix flows, dependency upgrades, flaky test fixes, and error triage. We recently open-sourced engrams at EVOLVE. Our upcoming webinar, How to Build Your Own AI Software Factory, covers how we built engrams, how to use it, and how to design a factory of your own.

Measuring progress with DRIVE and the OpEx review

Plenty of leaders measure their AI rollout with developer productivity metrics, but we believe that questions like "Did my engineers open more PRs this quarter?" are the wrong ones to ask. The answer says nothing about cost, quality, or whether the team can sustain the pace, and PR cycle time loses its meaning once background agents open a large share of the PRs.

A better measure starts with a question that predates AI: how effective is the organization at delivering value to its customers? That covers reliability, security, cost, and whether engineers can keep working at this pace without burning out.

The DRIVE framework measures that effectiveness at the organizational level. It assesses five pillars—Delivery, Reliability, Initiatives, Vigilance, and Efficiency—and prescribes a recurring review that turns those signals into action.

Pillar

Leadership question

Recommended signals

Delivery

Are we shipping fast, and is it sustainable?

Deploy frequency, lead time for changes, on-call pager volume

Reliability

Are we delivering on our promises to customers?

Functional SLO status, Sev0 and Sev1 incident count

Initiatives

Are our most important engineering investments making progress?

Tier 1 initiative milestone completion, OpEx action item completion

Vigilance

Is security risk accumulating?

Open critical and high fixable CVEs, assets below the minimum security bar, orphaned assets

Efficiency

Are we spending time and money effectively?

Cloud spend against budget, AI and LLM token costs, share of capacity spent on new work

The signals show direction (healthy, slipping, or ready for more speed), and they're recommendations. Pick the ones that fit your business.

DRIVE only works when it's paired with a recurring Operational Excellence (OpEx) review. Staff-plus engineers and senior leaders, including the CTO or VP of Engineering, meet weekly, biweekly, or monthly to look at the signals and answer three questions: what's going well, what's getting worse, and what needs action right now. We run ours weekly. If the most senior engineering leaders don't attend, the meeting becomes a dashboard walkthrough and nothing changes. As more of the SDLC runs without a human in the loop, this review is where humans govern the system as a whole.

We've run versions of this review on early drafts of DRIVE for more than three years. Over that period, we've more than doubled our PR volume while cutting incidents by 80%. The cultural change mattered as much as the numbers: the team ships AI-assisted code with far less anxiety because everyone can see what's working and what isn't.

At enterprise scale, the review runs into a data problem. Hundreds of teams produce tens of thousands of PRs, hundreds of SLOs, and thousands of open vulnerabilities, far more than any leadership team can read before a meeting. The OpEx Review Agent reasons over the Context Graph in Cortex, evaluates the DRIVE signals you choose, and surfaces the patterns and anomalies by team and product area, so leaders can see where to push harder on AI and where to slow down.

Start with an honest assessment

Most enterprises are two or three years into the largest shift in software engineering since the move to the cloud. If your AI rollout has been messier or more expensive than you planned, you're in the majority. Place each part of your organization honestly on the levels, fix what's breaking at the level you're at, and measure progress against your own baseline.

Take the DRIVE maturity assessment to see where your organization stands across delivery, reliability, initiatives, vigilance, and efficiency. It takes about five minutes.

Anish Dhar

Anish Dhar

CEO & Co-founder

Read next

Start building your AI software factory with Cortex