A bank is the hardest place to put AI into production and the easiest place to explain why. Every other industry argues about whether a model decision needs to be defensible. In a regulated financial institution that argument was settled decades ago by model risk management, and artificial intelligence simply arrived inside a discipline that already existed.
That is the good news, and it is usually mistaken for bad news. You are not starting from nothing. You are extending a model risk function to cover systems whose behaviour is learned rather than specified, and the gap is narrower and more concrete than a blank-page AI policy makes it look.
Listed by what binds you rather than by what is discussed. Every date here was verified on 9 August 2026; verify again before you rely on it, because these move.
The mistake that costs the most is treating each supervisor as a separate programme. The instruments differ in the evidence they demand, not in what good governance looks like: an inventory, a risk rating per model, a validation record, a named human accountable for each decision, and a way to explain an outcome to the customer it was made about. Build that once as a control plane the platform enforces, tag each artefact with the instrument it answers to, and each regulator becomes a report rather than a project.
The second mistake is subtler and more common. An inventory that lists models is not an inventory. For AI systems the model is rarely the only thing that varies: the prompt, the retrieval corpus and the routing policy all change without a deployment, and an inventory blind to those describes a system that no longer exists. That is the difference between a register that satisfies an examiner and one that merely exists.
More than most institutions expect, and less than the framework claims. The governance skeleton transfers intact: inventory, risk rating, independent validation, approval authority, ongoing monitoring, and a named owner. What does not transfer is the assumption that a model is a fixed artefact you validate and then watch. An AI system changes behaviour when the prompt changes, when the retrieval corpus is reindexed, and when the routing policy sends traffic to a different model, none of which is a deployment in the traditional sense. The honest scoping exercise is to walk the existing MRM function control by control and mark each one as transfers, transfers with modification, or absent. That is what the Teardown produces.
The minimum model inventory standard, and specifically its lineage requirement. Most institutions can produce a list of models. Far fewer can answer what changed, when, who approved it and on what evidence, for every model that carries risk, and that is the question the guideline is actually asking. For AI systems it is harder again because the things that change are not all versioned as code. The second commonly missed item is scope: E-23 covers all models that carry risk to the institution, which is broader than the list of models the MRM function currently governs, and the gap is usually in the business units that built something without calling it a model.
Insurers are explicitly in scope for OSFI E-23, alongside federally regulated deposit-taking institutions, including branches. Federally regulated pension plans are not. In the Gulf the picture depends on the supervisor rather than the sector label, and an insurer operating inside the DIFC is inside Regulation 10 on the same terms as a bank. The underwriting and pricing use cases usually carry the sharpest fairness exposure of anything in a financial institution, so the sector being quieter in the guidance does not make it quieter in risk.
No, and running two is the more expensive answer. Build to the strictest control in each dimension and produce the evidence once, tagged by the instrument it answers to. In practice that means building the inventory to the E-23 minimum standard, the AI register to DIFC Regulation 10, and the human-oversight and customer-review mechanism to the CBUAE expectations, then reporting from the one source. Institutions that run parallel programmes end up with inventories that disagree, and the disagreement surfaces during an examination rather than before it.
The AI Governance Teardown: a fixed-scope, fixed-fee, two-week examination scored against the four MESA layers and delivered board-ready, with the readout on Day 10 and every finding traced to evidence. No production data leaves your environment; it reads governance artefacts, not customer records. For institutions that then want the examiner to stay while they execute, there is a governance retainer with an annual re-score against the Teardown baseline. Everything starts with a free thirty-minute Fit Call that qualifies the work in both directions.
Governance an examiner can follow, in the jurisdiction that actually binds you.