AI GovernanceModel riskSheet 18

AI model risk.

A model the institution cannot classify, validate or stop is not an asset it owns. It is an exposure it operates.

§ 01The obligations

Model risk guidance moved, twice.

Two supervisory changes reset this discipline inside eighteen months, and both widen what counts as a model.

Model risk supervisory instruments and dates
InstrumentDateWhat changed
SR 26-2 (Fed, OCC, FDIC)Issued 17 Apr 2026Revised interagency guidance on model risk management, superseding SR 11-7 after fifteen years. Principles are preserved; the framework becomes more risk-based and scalable. The operational headline is that annual revalidation gives way to oversight tied to model materiality. Most relevant to institutions above roughly thirty billion dollars in assets.
OSFI Guideline E-23Effective 1 May 2027Canadian federally regulated banks and insurers, branches included. Applies to all models carrying risk to the institution, with the revised text giving explicit attention to AI and machine learning.
CBUAE, SAMAIn forceModel management standards and model-risk expectations across the GCC, with Sharia validation as a parallel authority for Islamic finance institutions.

The catch in both is scope, not stringency. Inventories and tiering built for statistical models written by a quant team do not reach the retrieval pipeline, the prompt template that decides an escalation, or the vendor model called through an API. Institutions are discovering that the population of things now meeting the definition of a model is considerably larger than the population they validate.

§ 02The discipline

Six pillars that hold the discipline.

The MESA model-risk discipline integrates CBUAE model management standards, SAMA expectations, the interagency principles now carried by SR 26-2, and Sharia requirements for Islamic finance. It is specified in Chapter 12 of the published Enterprise Playbook.

Pillar 1

Model governance

The strategic layer. Board-approved model risk appetite, a centralised inventory in which every production model is named, owned, classified and current, and committee oversight with real authority.

Pillar 2

Model development

First line. Use case and business justification before the model is built rather than after, data sourcing and quality assessed, algorithm selection reasoned and recorded.

Pillar 3

Independent validation

Second line. Scope set by the model's risk tier. Data validated by reproduction, with tests for leakage and label integrity, and a challenge that is genuinely independent of the builder.

Pillar 4

Deployment

The discipline of the gate. A pre-deployment check that validation findings are addressed and monitoring is configured, then phased rollout with champion-challenger comparison.

Pillar 5

Monitoring and revalidation

Production discipline. Revalidation cadence by risk tier, performance trended against the baseline set at deployment, and drift detected before it is discovered by a customer.

Pillar 6

Sharia integration

Dual validation. The Sharia Supervisory Board engaged at design rather than at validation, model logic tested for Riba, Gharar and Maysir, and Halal data certification in the chain.

A Model Card is not documentation. It is the model's contract with reality.

AI Governance & Compliance Frameworks for the Middle East · Chapter 12
§ 03Classification

Tiering is the load-bearing decision.

Not all models carry equal risk. A recommendation engine suggesting products operates at a different risk altitude from the credit model deciding approvals at the same institution, and classification is what drives everything downstream: validation depth, gate severity, monitoring cadence, and who is allowed to sign.

The classification matrix scores each model on ten dimensions, zero to ten each, for a total of zero to one hundred. Dimensions include regulatory impact, financial materiality and reputational risk among others. Seventy to one hundred is High Risk. Forty to sixty-nine is Medium. Zero to thirty-nine is Low. Generative AI extends the matrix with three further dimensions and expresses the result as tiers.

The failure mode is not choosing the wrong tier once. It is choosing a tier, building every control on top of it, and never revisiting it while the model changes underneath the label. A classification with no review date is a guess with a number attached.

§ 04Questions

What teams ask.

What counts as a model under the new guidance?

Broader than most inventories assume. Supervisors are converging on a functional definition: a quantitative method that processes input to produce an output used in a decision. Under that reading a retrieval pipeline, a prompt template that determines an escalation, a classifier inside a workflow, and a vendor model reached through an API can all qualify, whether or not anyone in the institution calls them models. The practical first step of most engagements is not validation, it is discovering the true population.

What changed with SR 26-2?

The Federal Reserve, OCC and FDIC issued revised interagency guidance on 17 April 2026, superseding SR 11-7 after fifteen years. The foundational principles of sound model risk management are preserved, but the framework becomes more risk-based and scalable, and the most consequential operational change is that fixed annual revalidation gives way to oversight proportionate to model materiality. It is expected to be most relevant to institutions above roughly thirty billion dollars in assets. Confirm applicability with counsel; what an assessment provides is the gap list and the sequence.

How does OSFI E-23 differ from what we already do?

Mostly in reach. E-23 takes effect on 1 May 2027 for Canadian federally regulated banks and insurers including branches, applies to all models that carry risk to the institution, and its revised text gives explicit attention to AI and machine learning. Institutions with a mature MRM function usually find the discipline familiar and the population unfamiliar: the tiering, the inventory and the evidence trail were built for a narrower definition of a model than the one now in force.

Can you validate models independently of the team that built them?

Yes, and that independence is the point of the third pillar. Independent validation means the challenge does not report to the builder: scope set by the risk tier, data validated by reproduction rather than by assurance, tests for leakage and label integrity, and findings that can block a deployment. An engagement can provide that challenge directly, or stand up the internal function that provides it, depending on whether the institution needs an answer now or a capability afterwards.

How do you handle model risk for LLMs and generative systems?

The same discipline with extra dimensions. Generative systems break the classical assumptions in specific ways: the output space is open, the behaviour changes when the vendor updates the model beneath you, and the failure modes include fabrication and prompt manipulation rather than only drift. The classification matrix is extended for that, and controls that matter most shift towards evidence at inference: what was retrieved, what was generated, what a human reviewed, and whether the decision can be reconstructed.

We use vendor models. Is that still our model risk?

It is your decision risk, which is what supervisors examine. Outsourcing the model does not outsource accountability for the outcome, and a vendor that will not disclose enough for you to validate is itself a finding. The vendor risk framework treats this explicitly: what the contract entitles you to know, what happens when the vendor changes the model, and whether you could exit if it became indefensible.

§ 05Start

Find out where you actually stand.

The method is published in full. What an engagement adds is the examination, with the evidence attached.

Fin · Model risk
Book the 30-minute Fit Call →