The AI Factory Has Five Rungs. Most Banks Are on the Second.
· 9 min read

The AI Factory Has Five Rungs. Most Banks Are on the Second.

By Orestes Garcia


Shopify’s CEO told his staff to prove they couldn’t do a job with AI before asking to hire a human for it. Reflexive AI usage, he called it, a baseline expectation. It’s a clean memo. Your bank’s CEO can’t send it.

Not because the idea is wrong. Because in a bank the thing that “AI-first” strains is not overhead. It’s the perimeter.

When Shopify flips the burden of proof, the cost of a bad AI-written change is a bug and an apology. When a regulated institution flips it, the cost is a control failure an examiner writes up. The review gate, the change advisory board, the model-risk validation: at a software company those look like friction to optimize away. At a bank they are the reason the institution is allowed to operate. You cannot delegate accountability to a model, and every AI-first slogan quietly assumes you can.

So the useful question is not “are we AI-first.” It’s narrower and more honest. Which rung of the factory are we actually standing on, and what does the next one cost?

AI-first is an operating model, not a purchase order

The consultancies converged on this before the banks did. BCG Platinion’s framing of the agentic software factory is blunt about where the work goes: “autonomous AI agents build, test, and ship software solutions around the clock, while humans define business intent and review outcomes.” Their sharpest line is the one every engineering leader should tape to the wall. “The bottleneck moves from coding speed to the clarity of organizational intent.”

That reframing matters because most banks are buying tools and calling it a strategy. A license for a coding assistant is not an operating model. It’s rung one. I made the architectural version of this argument in Six Primitives for a Code Factory, which is the how-you-build-it. This is the other axis: how a whole engineering org climbs, and why the climb is steeper inside a regulated wall.

Five rungs. Be honest about which one is yours.

The AI factory maturity ladder: five rungs from Shadow AI to a self-improving factory, wrapped by a governance band, with regulated banks marked at rung two

Rung 0: Shadow AI

Somebody on your team is pasting a stack trace into a consumer chatbot right now. No approval, no data-loss controls, no record. This is not a rung you choose. It’s the one you’re on by default the day the tools got good, whether you sanctioned them or not.

The exposure is the obvious part: proprietary source and material non-public information leaving the building through a browser tab. The subtler cost is that you have no idea how much AI is already in your codebase, which means you can’t answer the first question an auditor will ask. Rung zero is not “we haven’t started.” It’s “we started without noticing.” Getting off it is a governance act before it’s a technology one.

Rung 1: Sanctioned assistants

Here the tools are approved, the data controls are real, and the vendor cleared third-party risk review. Engineers get a coding assistant that keeps prompts inside the tenant. Productivity is individual and it is genuine.

The public numbers land at this rung. One large US bank reported “efficiency gains of over twenty percent” for developers on a GenAI coding tool. Another is rolling GitHub Copilot to roughly forty thousand developers. These are real, and they are also the ceiling of what an assistant alone can do. The AI helps a person type faster. It does not touch the lifecycle around the person.

This is where most regulated engineering orgs actually sit today, and it is a defensible place to sit. The trap is mistaking it for the destination. A faster individual inside an unchanged pipeline is a local optimization. The delivery system still moves at the speed of its slowest gate, and writing code was never the slow gate. I made that case in The Bottleneck Was Never the Code, and it’s the reason rung one plateaus.

Rung 2: The golden-path AI SDLC

Now AI leaves the IDE and enters the pipeline. Test generation, AI code review, security scanning that reasons about the diff, all wired into the paved road every change already travels. The defining move at this rung is that governance stops being a gate bolted on at the end and becomes a stage in the middle.

Two things separate rung two from a pile of point tools. First, provenance: every AI-touched change carries an attestation, so the pipeline can answer who or what wrote this line and who approved it. Second, evals as a pipeline stage, not a vibe: you measure whether the AI’s output actually holds before it moves, the same way you’d gate any other quality signal.

The proof this scales exists. Morgan Stanley’s internal tool reviewed nine million lines of legacy code and saved an estimated 280,000 developer hours translating old languages into plain-English specs. Amazon’s assistant did a Java migration its own CEO priced at roughly 4,500 developer-years and $260 million a year, with most AI-generated reviews shipping unchanged. That is a factory stage doing work, not an engineer with a faster autocomplete.

Rung 3: The governed autonomous factory

Agents run the loop. Spec to code to test to review, kicked off by a ticket or an alert, executing in an isolated environment, handing a human a finished change to approve. This is the six primitives assembled and running. Goldman Sachs put an autonomous coding agent into pilot alongside thousands of engineers, and its CTO described the shift plainly: developers become managers of very smart junior colleagues.

Inside a bank, rung three carries a requirement the software companies don’t share. The agent is an actor now, not a tool, and actors need identity. A governed non-human identity, an audit trail that survives an examiner’s questions, and a clear human it delegates for. The factory does not replace the regulated SDLC. It feeds it: the agent’s output arrives at the change board with its risk assessment attached and its provenance intact, and the whole engineering problem is the reconciliation between the fast track and the controlled one.

Almost no regulated institution is here yet in production, for material systems. That is not timidity. It’s arithmetic on the cost of being wrong.

Rung 4: The self-improving factory

The top rung is the lights-out version, where failing tests, bug reports, and review comments feed back and the factory tunes itself. BCG calls the endgame the “dark factory.” For a narrow, low-risk class of change it’s plausible. For anything material in a bank, it is not, and it should not be.

Here is the part the maturity-model diagrams never say out loud: a regulated institution tops out below rung four on purpose. The ceiling is not a failure of ambition. It is a control. An examiner will not accept “the factory decided” as the answer to who authorized this change to a payment flow. The cap is the feature.

What actually moves you up a rung

Climbing is not about buying the next tool. It’s about paying the specific cost the next rung demands. Four moves do most of the work.

Separate the coding model from the decisioning model. This is the distinction that strands the most teams. In April 2026 the banking regulators reissued their model-risk guidance and explicitly put generative and agentic AI outside the formal framework, calling them “novel and rapidly evolving,” while keeping them under the bank’s broader risk management. Read that carefully. An AI that helps write software is not the same regulated object as an AI that decides who gets a loan. Govern the assistant like it’s a credit model and you’ll never leave rung one. Fail to govern the decisioning code and you have a violation. The whole art is telling them apart.

Make provenance the price of admission. The unlock from rung one to rung two is not a smarter agent. It’s traceability. Attest every AI-touched change so audit can reconstruct authorship and approval after the fact. Counterintuitively, that record is what buys you permission to go faster: the institution will tolerate more autonomy precisely when it can prove what happened.

Pave the road instead of sending the memo. You can’t mandate reflexive AI usage into a system of record. So don’t mandate. Make the governed AI path the easy path, the one with the sandbox pre-cleared and the review baked in, so the compliant route is also the fastest route. Golden paths beat edicts in every org, and they’re the only thing that works where the edict would be a finding.

Measure the loop, not the adoption. The most quoted stat in this space, that AI writes some large share of the code, is the wrong number to chase. Delivery is a system, and the honest research says so.

The part I won’t pretend away

Adoption is not impact. A randomized study from METR put experienced developers on their own mature codebases and found they were 19% slower with AI tools, while forecasting a 24% speedup and believing afterward they’d been 20% faster. The perception gap is the story. The DORA research has repeatedly found AI acting as an amplifier: it makes strong delivery systems stronger and exposes weak ones, and naive adoption can correlate with less stable delivery, not more.

The org-level cautionary tale is Duolingo, which declared itself AI-first, took public heat, and walked the framing back within weeks. Declarations outrun reality. And Gartner’s read on the agentic wave is that a large share of those projects get canceled before they ship, undone by cost and unclear value rather than by the models.

None of that is an argument against the climb. It’s an argument for climbing by verification discipline rather than by press-release percentage. The factory earns velocity through proof, not through vibes. The rungs are real, the ceiling is real, and knowing which of them you’re standing on is worth more than any tool you could buy this quarter.

The engine you run is rented anyway, replaced every few months by a lab you don’t control. What compounds is the factory you build around it, and how honestly you know its shape. I made the rent-versus-own case in You’re Renting the Model. Own the Harness. This is the same argument, one floor up.


The companion read is Six Primitives for a Code Factory, which builds the machine this post grades. And for the individual version of the same maturity question, how you measure your own growth with agents rather than your org’s, see From Threads to Trust.

I write about AI-assisted development, enterprise architecture, and building in regulated environments. If you’re figuring out which rung your own factory is on, I’d like to compare notes. Find me on X @orestesgarcia or LinkedIn /in/setsero.