E-23 Will Not Read Your Policy. It Will Read Your Record.
Guideline E-23 takes effect on 1 May 2027. It does not ask whether your institution has an AI policy. It asks whether you can produce the record. Here is what the record has to contain, and a sequence for building it in the time that is left.
What E-23 actually says
OSFI published the final Guideline E-23, Model Risk Management (2027), on 11 September 2025. It takes effect on 1 May 2027, after what OSFI describes as an eighteen-month transition.
It applies to every federally regulated financial institution: banks, foreign bank branches, life insurance and fraternal companies, property and casualty companies, and trust and loan companies. Federally regulated pension plans are excluded, on the basis of their distinct supervisory framework. It is principles-based and risk-based, so the intensity of what it expects scales with the inherent risk of each model.
Its definition of a model includes artificial intelligence and machine learning methods by name.
Three outcomes sit at the top. Model risk understood and managed across the enterprise. Model risk managed using a risk-based approach. Model governance covering the entire model lifecycle. Twelve principles sit beneath those three, and Appendix 1 sets the minimum content of the model inventory.
That is the guideline. The rest of this page is about what satisfying it actually costs.
Why AI teams should read it as if it were written for them
Three things in the text reach the AI estate directly.
The definition. A model is any application of theoretical, empirical or judgmental assumptions or statistical techniques that processes input data to generate results, and the definition names AI and machine learning methods inside it. A large language model behind a customer-facing assistant is a model. A retrieval pipeline that ranks documents for a credit analyst is a model. A vendor API that scores a claim is a model, and E-23 expects externally developed models to be rated for risk on a standalone basis.
The AI language. The guideline names the level of transparency and explainability required, the alternative controls a black-box or autonomous model needs, the potential for biased outcomes, autonomous decision-making and autonomous re-parametrization as monitoring problems, and model drift as an elevated risk. It expects extensive use of advanced AI and machine learning techniques to be matched by correspondingly mature governance and oversight. Those are not incidental mentions. They are the paragraphs a supervisor reads first when the system under discussion is an AI system.
The inventory. Appendix 1 lists seventeen fields an institution maintains for every model carrying non-negligible risk: version, deployment date, reviewer, approver, dependencies, data sources, approved uses, limitations, last review date, monitoring status, next review date, and the identifying and ownership fields above them. If those fields cannot be produced for the AI systems you run today, you are not ready, whatever the policy says.
Where most institutions will fail
Not on policy. Most organisations will have a policy by May 2027.
They will fail on evidence, and specifically on the distance between the altitude at which governance is written and the altitude at which the decision is made. A model risk framework can be current, approved and filed while the system it describes runs on a substrate that records nothing the framework requires. No lineage captured at training time. No policy version bound to the approval. No monitoring trigger actually armed.
A policy is not evidence. It is a claim that evidence exists somewhere else.
Governance written in a binder is governance the running system never encounters. That is the failure the MESA Framework (Maturity, Evidence, Substrate, Alignment) was built to find. It scores institutional governance at four altitudes, the Regulatory Floor, the Strategic Compass, the Operational Machinery and the Technical Substrate, and it scores them separately, because one combined grade is exactly what lets a strong policy posture conceal a substrate that cannot hold it up.
The twelve principles distribute across all four altitudes, and the pattern is consistent. The principles an institution will satisfy on paper sit at the top two altitudes. The principles that fail an examination sit at the bottom two, where the question stops being what the policy says and becomes whether the evidence existed at the moment of the decision.
The eight-month sequence
Eight months is enough if the sequence is right. It is not enough to do everything at once.
What follows is a Now, Next, Later plan, month by month. Each month produces one artifact a supervisor could be shown.
Month 1, September 2026: the sweep. Identify every model in use or recently decommissioned, including generative AI, retrieval pipelines and vendor APIs (Principle 2.1). Survey the business, not the data science function. Models now sit in areas of the institution that never relied on them, which is why a sweep run through the modelling team finds a fraction of the estate. Triage each for non-negligible inherent risk. Artifact: A candidate inventory with provisional ratings, and a separate list of the AI systems that were on nobody’s list.
Month 2, October 2026: the rating approach. Rewrite the risk rating method so it scores what AI adds: level of autonomy, explainability required, reliability of inputs, customer impact, regulatory exposure (Principle 2.2). Rate externally developed models on a standalone basis. Define the negligible-risk category and the exemption process that governs entry into it, because an undefined exemption is where an estate quietly disappears. Artifact: The rating methodology, approved, and every inventory entry rated under it.
Month 3, November 2026: the framework and the roles. Revise the model risk management framework so it names an AI-specific risk appetite, board reporting, resourcing, and the multi-disciplinary team the guideline expects, including legal or ethics input (Principles 1.1 and 1.2). Name the owner, developer, reviewer and approver for every rated model, in the inventory rather than in a slide. A committee named as accountable is not an accountable person. Artifact: The framework, the accountability map, and the first board report on AI model risk.
Month 4, December 2026: the data. For each rated AI model, establish that development data is accurate, relevant, compliant, traceable and timely (Principle 3.2), with documented lineage and provenance and controls over synthetic and proxy data. This is the month most institutions discover that lineage was never recorded at training time and cannot be reconstructed honestly afterwards. Record that rather than reconstruct it. Artifact: Data lineage records per model, and a gap register for the models where they do not exist.
Month 5, January 2027: independent review. Stand up or contract review capacity that is independent of development (Principle 3.4), with the scope the guideline names for AI: novel methodologies, explainability measured against intended use, third-party components and libraries. Prioritise by rating. The test of independence is narrow and it is not organisational distance. It is whether the reviewer can reject. Artifact: Completed reviews for the highest-rated AI models, with findings and an approval recommendation.
Month 6, February 2027: deployment and monitoring. Bring AI deployment under quality and change control (Principle 3.5): consistency between development and production data, production tests, documented approval hierarchies, tested rollback, and the cyber and operational risk assessments that Guidelines B-13 and E-21 expect alongside. Then define monitoring standards by rating and model type, with thresholds, contingency plans, and the two AI-specific cases the guideline calls out, autonomous re-parametrization and drift (Principle 3.6). Artifact: Deployment procedures, monitoring standards, and armed triggers with evidence that they fire.
Month 7, March 2027: the evidence dry run. Produce every Appendix 1 field for every inventory entry from the systems that hold the evidence, not from someone’s memory. Where a field has to be typed rather than produced, that is a substrate gap and it gets recorded as one. Then run one end-to-end reconstruction. Pick a decision an AI model made, and show who approved the model, under which version, against which rating, with what monitoring status. Artifact: The reconstruction report and the closed gap list.
Month 8, April 2027: attestation and decommission. Adopt the decommission process (Principle 3.6): stakeholder alerting, retention of the retired model as a benchmark, third-party model handling, downstream monitoring. Take the board through the inventory, the ratings, the review status and the residual gaps. Artifact: The board attestation, with residual gaps stated rather than hidden.
On 1 May you are not perfect. You are examinable. Those are different claims, and only one of them is available in eight months.
Five things to do this week
- Ask one question of every business unit: which decisions in your area are supported by something that processes inputs to generate results, and who owns it? Collect the answers in one place. The answers you did not expect are the finding.
- Pull the vendor list. Any third party whose product contains a model is in scope under Guideline B-10 and E-23 both. The AI Vendor Risk Framework questionnaire is fifty-six questions in seven sections, published under CC BY 4.0, and any institution may issue it to any vendor without asking permission.
- Pick your highest-impact AI system and try to fill Appendix 1 for it from records alone. Time how long it takes. That number is your readiness, and it is more honest than a maturity score.
- For each principle, mark whether your evidence would come from a document or from a system. Where a committee sees an inventory, I look for the last hand that typed into it. A field a person maintains is a field that goes stale between reviews.
- Take the free twelve-question readiness self-assessment at the assessment page. It scores in your browser and sends nothing unless you ask. It will not tell you whether you meet E-23 and does not claim to. It will tell you which altitude is your binding constraint, which is the thing worth knowing before you scope anything.
Where the evidence has to live
Each of the twelve principles, and Appendix 1 with it, can be placed at the MESA altitude where its evidence has to exist. Principles 2.1, 3.2 and 3.5, together with the inventory itself, are the substrate rows: if the answer to “where does this come from” is that a person types it, the control exists on paper and not in the system. The Five-Gate Deployment Model supplies the operational spine, and each gate’s passage record produces Appendix 1 fields as a by-product rather than as a task.
Some of the frameworks in that mapping are deposited as self-published technical notes with their own DOI. Others are entries whose authoritative treatment is a chapter of the Enterprise Playbook and which carry no deposit at all. The research index marks which is which, because a reader is entitled to know whether a reference points at a document they can download or at a book they would have to buy.
If you want it done in two weeks
The AI Governance Teardown is a fixed-scope, fixed-fee, two-week examination of how an institution governs its AI, scored against the fifty-question MESA instrument across the four altitudes and delivered as a Governance Gap Report, a Now, Next, Later remediation roadmap, a findings readout and a one-page board summary.
It reads governance artifacts, not customer data. For E-23 it produces the eight-month plan above with your names and your gaps in it.
It does not deliver compliance. Meeting E-23 is a supervisory judgement made by OSFI, and nobody honest sells one. What the examination delivers is the reading of where you actually are, which is the input to an E-23 programme rather than a substitute for one.
It starts with a free thirty-minute fit call. The wider practice, and the other Canadian instruments that sit alongside E-23, are set out on the AI governance practice page and the Canada page.
Questions this page answers
Does OSFI E-23 apply to AI systems?
Yes. The guideline’s definition of a model names artificial intelligence and machine learning methods inside it, and the text addresses explainability, autonomy, biased outcomes, model drift and autonomous re-parametrization as matters the institution has to govern. An AI system that processes inputs to generate results is a model under E-23.
When does E-23 take effect?
On 1 May 2027. OSFI published the final Guideline E-23, Model Risk Management (2027), on 11 September 2025 and describes the period between as an eighteen-month transition. The date is the one fixed point in Canadian AI supervision, and it is the reason a readiness programme has a sequence rather than a wish list.
Who does E-23 apply to?
All federally regulated financial institutions: banks, foreign bank branches, life insurance and fraternal companies, property and casualty companies, and trust and loan companies. Federally regulated pension plans are excluded, on the basis of their distinct supervisory framework. It applies on a risk basis, so intensity scales with size, strategy, risk profile, complexity and interconnectedness.
Does E-23 cover vendor AI models?
Yes. The model risk management framework has to cover models and data sourced externally, and externally developed models are rated for risk on a standalone basis rather than inherited at the vendor’s rating. Guideline B-10 on third-party risk management applies alongside it, not instead of it.
What must the model inventory contain?
Appendix 1 sets seventeen fields for every model carrying non-negligible risk, among them the risk rating, owner, developer, origin, version, deployment date, reviewer, approver, dependencies, data sources, approved uses, limitations, last review date, monitoring status and next review date. The test is whether those fields can be produced by a system or only typed by a person.
Is AIDA relevant to E-23 readiness?
No. The Artificial Intelligence and Data Act died with Bill C-27 on 6 January 2025 and has not been reintroduced, so it carries no obligations, no penalties and no deadline. For a federally regulated financial institution, E-23 is the instrument with a date on it.
What this page cannot do
- This is architecture and governance advisory. It is not legal advice, and it does not substitute for your counsel or your relationship with your regulator. It is an analytical reading of a published guideline, not OSFI guidance.
- Regulatory dates and obligations here were verified against OSFI’s published text on 5 September 2026, and they move. Check the current text at OSFI before you rely on anything above.
- Meeting E-23 is a supervisory judgement made by OSFI, not a certification anyone sells. No engagement described here delivers compliance.
- The framework notes cited are self-deposited technical notes. Zenodo assigns a DOI to what it is given and does not referee it. None of them has been peer reviewed, and each says so on its own first page.
- Engagements are delivered personally, through iSystematic Inc.
Sources: OSFI, Guideline E-23 Model Risk Management (2027), published 11 September 2025, effective 1 May 2027. The MESA Framework, self-deposited technical note, DOI 10.5281/zenodo.22109836. The Five-Gate Deployment Model, self-deposited technical note, DOI 10.5281/zenodo.22170122. The AI Vendor Risk Framework questionnaire, self-deposited technical note, DOI 10.5281/zenodo.22170146. Every DOI above is a concept DOI and resolves to the current version.