The discipline of deciding how machine learning enters an institution that already has an operating model, a regulator, and a record it has to be able to produce. Very little of the difficulty is greenfield. Most of it is what is already there.
Enterprise architecture asks how the parts of an institution fit together and who answers for each one. Enterprise AI architecture asks the same question about a component whose behaviour is learned rather than specified.
That difference is the entire discipline. A conventional component does what its specification says, and when it does not, the specification is the place you go to find out why. A model does what its training, its context, and its inputs incline it to do. That is a different kind of claim, and it cannot be discharged by reading the code.
So the accountability has to be carried somewhere other than the component. It is carried by the architecture: by what the model is allowed to reach, what it is allowed to decide alone, what is recorded when it decides, and what an institution can put in front of a regulator eighteen months later. Governance that exists only in a policy document is governance the system never sees.
Architecture practice already has answers to these. Introduce a model and each one degrades in a specific, predictable way. Designing for that degradation is the work.
| Question | Conventional component | Model-backed component |
|---|---|---|
| What does it do? | Read the specification. | Observe the distribution. The specification describes intent, not behaviour. |
| Why did it do that? | Trace the code path. | Reconstruct it from the evidence you decided in advance to keep. If you did not keep it, the answer no longer exists. |
| Will it do it again? | Deterministic within its contract. | Depends on model version, prompt, retrieved context, and the vendor's release schedule. |
| Who answers for it? | The owning team. | Still the owning team, with materially less to point at. |
The fourth row is the one that reaches the board. Accountability does not move when the mechanism becomes statistical. Only the evidence does.
Serving, routing, model selection, tenancy, and cost. Where capability is made available at all, and where a single routing decision quietly changes the risk profile of everything downstream.
Agents, tools, retrieval, and the contracts between them. Where capability becomes behaviour, and where most of the governance surface actually lives.
Delivery, evaluation, incident response, and evidence. Where behaviour is kept accountable over time, which is the only timescale a regulator cares about.
Institutions usually buy the first layer, build the second, and discover the third after an incident. The sequence is expensive in that order, and it is the ordinary one.
These three layers are the structure of the Full-Stack AI Engineering Series, which sets them out in full through a deliberately fictional institution. The infrastructure layer is treated on its own at LLM systems, and the application layer at agentic AI.
A control plane is where the institution's rules become something the system enforces rather than something a committee wrote down. Policy as code, capability contracts that state what a component may reach and decide, trust tiers that make the level of human involvement a property of the deployment rather than a matter of habit, and golden paths that make the compliant route the easy one.
The design test is simple and unforgiving. If a rule cannot be violated by a system that is trying to satisfy its objective, it is a control. If it can, it is a preference. Most AI policy, examined closely, turns out to be preference.
The named patterns behind this are published and can be read before you engage: MESA for scoring where an institution stands, the Five-Gate Deployment Model for what must be true before a model moves, PEVG for the shape of a governed agent, and PARA for running it once it is live.
An engagement produces a current-state architecture of the AI estate as it really is rather than as the inventory claims, a target architecture with the sequence and the dependencies made explicit, a control plane design, an evidence model stating what is recorded at each decision and how long it survives, and the architecture decision records that explain why each choice was made to whoever inherits it.
Deliverables are scoped per engagement and range from reference architecture and governance design through to working implementation. The engagement shapes are set out at engagements, and the fixed-scope entry point is the AI Governance Teardown.
It is the practice of deciding how AI capability enters an existing institutional estate, and who answers for what it decides. It covers the serving and routing layer, the agents and retrieval built on top of it, and the operational discipline that keeps both accountable once they are live.
It differs from ordinary enterprise architecture in one respect that changes everything downstream: the behaviour of a model is learned rather than specified, so it cannot be established by reading the code. The architecture has to carry the accountability the component cannot.
A platform team builds capability. An architect decides what capability is allowed to do, what it must record, and what has to be true before it moves closer to a customer or a balance sheet.
The two are complementary and the work is often done alongside an existing platform team. What is being added is the accountability structure, not the pipeline.
Both, scoped per engagement. Deliverables range from a reference architecture and a governance design through to working implementation of the control plane. The scope is agreed before the work starts, and it is written down.
ISO/IEC 42001 and the NIST AI Risk Management Framework as the general spine, with the jurisdictional obligation layered on top: OSFI Guideline E-23 for Canadian federally regulated institutions, the EU AI Act where it reaches, and the GCC regulators including SAMA, CBUAE, SDAIA and QCB.
The mapping is the point. Standards describe what good governance contains. They do not tell you what your estate is missing, which is what an architecture engagement establishes.
Yes, as scaffolding. Structured EA practice is what makes an AI estate legible to the rest of the institution, and the discipline of ADRs, capability models and traceability is exactly what is missing from most AI programmes.
What conventional practice does not supply is a treatment of a component whose behaviour is statistical. That gap is what this work fills, rather than replacing the practice around it.
With the free self-assessment, which scores you against the four MESA layers in twelve questions and returns a profile per layer rather than a single grade. It runs in your browser and nothing is sent unless you ask for the written interpretation.
If you would rather have it done properly and independently, the fixed-scope entry point is the AI Governance Teardown, which begins with a free thirty-minute Fit Call.
The Fit Call is thirty minutes and free, and it qualifies the work in both directions. If what you need is a governance function rather than an architecture, you will be told that on the call.