AI GovernanceModel riskSheet 18

AI model risk.

A model the institution cannot classify, validate or stop is not an asset it owns. It is an exposure it operates.

§ 01The obligations

Model risk guidance moved, twice.

Two supervisory changes reset this discipline inside eighteen months, and both widen what counts as a model.

Model risk supervisory instruments and dates
InstrumentDateWhat changed
SR 26-2 (Fed, OCC, FDIC)Issued 17 Apr 2026Revised interagency guidance on model risk management, superseding SR 11-7 after fifteen years. Principles are preserved; the framework becomes more risk-based and scalable. The operational headline is that annual revalidation gives way to oversight tied to model materiality. Most relevant to institutions above roughly thirty billion dollars in assets.
OSFI Guideline E-23Effective 1 May 2027Canadian federally regulated banks and insurers, branches included. Applies to all models carrying risk to the institution, with the revised text giving explicit attention to AI and machine learning.
CBUAE, SAMAIn forceModel management standards and model-risk expectations across the GCC, with Sharia validation as a parallel authority for Islamic finance institutions.

The catch in both is scope, not stringency. Inventories and tiering built for statistical models written by a quant team do not reach the retrieval pipeline, the prompt template that decides an escalation, or the vendor model called through an API. Institutions are discovering that the population of things now meeting the definition of a model is considerably larger than the population they validate.

§ 02The discipline

Six pillars that hold the discipline.

The MESA model-risk discipline integrates CBUAE model management standards, SAMA expectations, the interagency principles now carried by SR 26-2, and Sharia requirements for Islamic finance. It is specified in Chapter 12 of the published Enterprise Playbook.

Pillar 1

Model governance

The strategic layer. Board-approved model risk appetite, a centralised inventory in which every production model is named, owned, classified and current, and committee oversight with real authority.

Pillar 2

Model development

First line. Use case and business justification before the model is built rather than after, data sourcing and quality assessed, algorithm selection reasoned and recorded.

Pillar 3

Independent validation

Second line. Scope set by the model's risk tier. Data validated by reproduction, with tests for leakage and label integrity, and a challenge that is genuinely independent of the builder.

Pillar 4

Deployment

The discipline of the gate. A pre-deployment check that validation findings are addressed and monitoring is configured, then phased rollout with champion-challenger comparison.

Pillar 5

Monitoring and revalidation

Production discipline. Revalidation cadence by risk tier, performance trended against the baseline set at deployment, and drift detected before it is discovered by a customer.

Pillar 6

Sharia integration

Dual validation. The Sharia Supervisory Board engaged at design rather than at validation, model logic tested for Riba, Gharar and Maysir, and Halal data certification in the chain.

A Model Card is not documentation. It is the model's contract with reality.

AI Governance and Compliance Frameworks for the Middle East · Chapter 12
§ 03Classification

Tiering is the load-bearing decision.

Not all models carry equal risk. A recommendation engine suggesting products operates at a different risk altitude from the credit model deciding approvals at the same institution, and classification is what drives everything downstream: validation depth, gate severity, monitoring cadence, and who is allowed to sign.

The classification matrix scores each model on ten dimensions, zero to ten each, for a total of zero to one hundred. Dimensions include regulatory impact, financial materiality and reputational risk among others. Seventy to one hundred is High Risk. Forty to sixty-nine is Medium. Zero to thirty-nine is Low. Generative AI extends the matrix with three further dimensions and expresses the result as tiers.

The failure mode is not choosing the wrong tier once. It is choosing a tier, building every control on top of it, and never revisiting it while the model changes underneath the label. A classification with no review date is a guess with a number attached.

§ 04The lifecycle

After the tier is set.

A tier is an input to six other decisions. Each of them is where an institution either accumulates evidence continuously or discovers, at examination, that it has been accumulating intent.

01 · Record

Registry and inventory

The model registry and the AI asset inventory are the same question asked by two functions, and institutions routinely maintain neither completely. A tier that is not recorded somewhere the institution can query is a tier nobody can act on, and the first task of most engagements is still discovering the true population rather than validating it.

02 · Permission

Approval workflow

Who may sign at each tier, and what they are signing that the model has cleared. The Five-Gate Deployment Model exists to make that sequence explicit, so approval is a gate with a named owner rather than an email that nobody can find eighteen months later.

03 · Change

Model change management

The behaviour changes when the vendor updates the model beneath you, and a change that does not re-open the model tiering decision is how a High Risk system becomes a Low Risk label by neglect. Change management here has to be triggered by the model moving, not only by someone deciding to revisit it.

04 · Threshold

Performance thresholds

Agreed before deployment, so degradation is a trigger rather than a debate. A threshold set after the numbers are in is not a control, it is a negotiation, and it is the point at which monitoring stops being evidence and becomes commentary.

05 · Acceptance

Risk acceptance

Sometimes the institution runs it anyway, and that is legitimate. What makes it defensible is that the acceptance is written down, owned by someone with the authority to accept it, dated, and given a review point. Residual risk that was never accepted by anyone is the version that surfaces as a finding.

06 · Evidence

Continuous compliance

Regulatory reporting assembled at examination time is a reconstruction, and reconstructions are where gaps are discovered rather than closed. The alternative is continuous compliance: the artifact is produced as the decision happens, because it is a by-product of the control rather than a report about it.

The standards context for this is ordinary and worth naming: ISO/IEC 42001 for the management system, ISO/IEC 23894 for AI risk management specifically, and the NIST AI Risk Management Framework for the function map. Where personal data is in the decision, privacy engineering is part of the same lifecycle rather than a parallel one, because the lawful basis for a decision and the evidence for it are produced by the same architecture.

§ 05Questions

What teams ask.

What counts as a model under the new guidance?

Broader than most inventories assume. Supervisors are converging on a functional definition: a quantitative method that processes input to produce an output used in a decision. Under that reading a retrieval pipeline, a prompt template that determines an escalation, a classifier inside a workflow, and a vendor model reached through an API can all qualify, whether or not anyone in the institution calls them models. The practical first step of most engagements is not validation, it is discovering the true population.

What changed with SR 26-2?

The Federal Reserve, OCC and FDIC issued revised interagency guidance on 17 April 2026, superseding SR 11-7 after fifteen years. The foundational principles of sound model risk management are preserved, but the framework becomes more risk-based and scalable, and the most consequential operational change is that fixed annual revalidation gives way to oversight proportionate to model materiality. It is expected to be most relevant to institutions above roughly thirty billion dollars in assets. Confirm applicability with counsel; what an assessment provides is the gap list and the sequence.

How does OSFI E-23 differ from what we already do?

Mostly in reach. E-23 takes effect on 1 May 2027 for Canadian federally regulated banks and insurers including branches, applies to all models that carry risk to the institution, and its revised text gives explicit attention to AI and machine learning. Institutions with a mature MRM function usually find the discipline familiar and the population unfamiliar: the tiering, the inventory and the evidence trail were built for a narrower definition of a model than the one now in force.

Can you validate models independently of the team that built them?

Yes, and that independence is the point of the third pillar. Independent validation means the challenge does not report to the builder: scope set by the risk tier, data validated by reproduction rather than by assurance, tests for leakage and label integrity, and findings that can block a deployment. An engagement can provide that challenge directly, or stand up the internal function that provides it, depending on whether the institution needs an answer now or a capability afterwards.

How do you handle model risk for LLMs and generative systems?

The same discipline with extra dimensions. Generative systems break the classical assumptions in specific ways: the output space is open, the behaviour changes when the vendor updates the model beneath you, and the failure modes include fabrication and prompt manipulation rather than only drift. The classification matrix is extended for that, and controls that matter most shift towards evidence at inference: what was retrieved, what was generated, what a human reviewed, and whether the decision can be reconstructed.

We use vendor models. Is that still our model risk?

It is your decision risk, which is what supervisors examine. Outsourcing the model does not outsource accountability for the outcome, and a vendor that will not disclose enough for you to validate is itself a finding. The vendor risk framework treats this explicitly: what the contract entitles you to know, what happens when the vendor changes the model, and whether you could exit if it became indefensible.

§ 06Start

Find out where you actually stand.

The method is published in full. What an engagement adds is the examination, with the evidence attached.

§ 07Ask an assistantLive, no key

Ask your AI assistant instead.

This page is a snapshot, accurate at the release it cites. The same corpus is callable, publicly and without a key, so an assistant can query it live and return an answer carrying the source it came from. For this page that is lookup_regulation and identify_relevant_service, which return what is actually in force for a jurisdiction, and map a described problem to an engagement shape with the routing reasoning shown. Sector summaries are where a live obligation and a consultation draft most often get merged into one sentence, so checking the instrument rather than the summary is the point.

01 · Connect
claude mcp add --transport http concylium https://mcp.nabeelkhan.com/api/mcp

Claude Desktop, ChatGPT, Cursor, VS Code and Gemini CLI take the endpoint on its own: https://mcp.nabeelkhan.com/api/mcp. No key, no account, nothing to sign. Setup for every client.

02 · Ask

“Using Concylium, list what binds an AI system deployed by an organisation in this sector in Canada and in the UAE, and tell me which of those are in force today.”

Two jurisdictions at once is where flattening shows. The answer separates in force from guidance and names the instrument behind each.

Fin · Model risk
Book the 30-minute Fit Call →