§ 01Appendix F

The MESA Self-Assessment

The full instrument. Fifty questions, twelve to thirteen per layer, each with the five-level rubric it is scored against. The shorter twelve-question readiness assessment locates you; this one examines the estate.

It is not a quiz. It is a structured honesty instrument, and three rules matter more than the questions. Score against operating reality, not policy claims. Score conservatively at boundaries, meaning the level you could defend under examination rather than the one you could describe in a leadership memo. And score by triangulation: each layer answered by the layer owner and an independent reviewer, and where they disagree the lower score wins, because the disagreement is the finding.

Scoring runs entirely in your browser. Nothing is transmitted, nothing is stored, and there is no sign-up, so there is nothing to unsubscribe from later.

Companion material to AI Governance & Compliance Frameworks for the Middle East by Nabeel Khan · release v3.2, locked 17 August 2026 · Appendix F · free, no registration
§ 02The instrument

Fifty questions, four layers.

L1 · 13 questions

Regulatory Floor

Non-negotiable compliance: supervisory expectations, data-protection law, and Sharia obligations.

Q1.1Jurisdictional Mapping of the AI Estate

Does the institution maintain a current, written mapping of every production AI system to the jurisdictions and supervisory authorities it is subject to?

Q1.2Sharia Validation Integration (Islamic and Dual-Window Institutions)

How is Sharia validation integrated into AI model governance for Islamic-finance-relevant models?

Q1.3Cross-Border Data Handling

What is the state of governance over cross-border data transfers feeding AI training, inference, or vendor pipelines?

Q1.4Sectoral Regulatory Compliance

For each regulated sector the institution operates in (banking, insurance, healthcare, energy), is sector-specific AI guidance (SAMA MRM, CBUAE C 7/2025, CBB Module RM, AAOIFI standards) documented and operationalized?

Q1.5Personal Data Protection Compliance

For AI systems processing personal data, are PDPL obligations (lawful basis, consent management, data subject rights, breach notification timelines) operationalized rather than only documented?

Q1.6Algorithmic Decision Disclosure (UAE PDPL Article 18, Saudi Implementing Regulations, and Equivalents)

For customer-facing AI decisions subject to the automated-decision-disclosure provisions (UAE PDPL Article 18; the Saudi Implementing Regulations) or equivalents, is per-decision explainability operational?

Q1.7Breach Notification Discipline

Are breach notification procedures (identification, investigation, documentation, regulator notification within the jurisdiction's prescribed window) tested rather than only documented?

Q1.8Records of Processing Activity (ROPA) for AI Systems

For every AI system processing personal data, is a current ROPA entry maintained covering purpose, lawful basis, data sources, recipients, retention, and cross-border transfer mechanisms?

Q1.9Data Protection Impact Assessments (DPIAs) for AI

Are DPIAs conducted for AI systems prior to deployment, covering purpose, lawful basis, risk to data subjects, and proportionality of processing?

Q1.10Supervisor Engagement Posture

Does the institution engage proactively with supervisors (consultations, pre-deployment notifications, sandbox participation) rather than only reactively responding to inquiries?

Q1.11Cross-Border Regulatory Coordination

For multi-jurisdiction institutions, are coordination mechanisms across supervisory authorities (joint notifications, parallel inquiries, conflict-of-laws procedures) documented and tested?

Q1.12Regulatory Penalty Awareness and Control Calibration

Has the institution assessed regulatory sanctions applicable to AI governance violations and calibrated controls to prevent the highest-impact violations?

Q1.13Regulatory Change Management

Is there a documented process for monitoring regulatory developments (new guidance, amended frameworks, draft directives) in every jurisdiction the institution operates in, with assigned ownership and impact assessment?

L2 · 12 questions

Strategic Compass

Board-owned direction: what AI is for, the appetite that bounds it, and what the institution declines.

Q2.1Board-Approved AI Risk Appetite

Does the board approve a written AI risk appetite statement at least annually?

Q2.2National AI Strategy Alignment

Is the institution's AI roadmap mapped to the national AI strategies of the jurisdictions it operates in (UAE National AI Strategy 2031, Saudi Vision 2030 and SDAIA, Qatar National AI Strategy, equivalents elsewhere)?

Q2.3AI Governance as Competitive Capability

Is AI governance treated as a competitive capability or as a compliance cost?

Q2.4Board-Level AI Literacy

Does the board have demonstrable AI literacy sufficient to challenge management on AI strategy, risk, and governance posture?

Q2.5Executive Sponsorship Quality

Is the AI governance program sponsored by a named executive with budget authority and quarterly board reporting?

Q2.6Strategic Use Case Pipeline

Does the institution maintain a strategic pipeline of high-value AI use cases prioritized by business impact and governance feasibility?

Q2.7Governance Investment as a Capital Decision

Is AI governance investment evaluated as a capital decision against documented risk reduction and strategic positioning value?

Q2.8Strategic Communication to External Stakeholders

Does the institution communicate its AI governance posture externally (annual transparency report, governance disclosures, model card publication) in a manner customers, supervisors, and rating agencies can read?

Q2.9Talent Strategy for AI Governance

Is there a documented talent strategy for AI governance, including internal capability building, external hiring, and succession planning?

Q2.10Cross-Functional Operating Model Integration

Is the AI Governance Operating Model from Chapter 10 integrated with the institution's enterprise risk, compliance, and audit functions rather than parallel to them?

Q2.11Strategic Benchmarking

Does the institution benchmark its AI governance posture against peer institutions, regional best practices, and international standards (ISO 42001, NIST AI RMF)?

Q2.12Strategic Risk Integration

Is AI risk integrated into the institution's enterprise risk management framework, with cross-references to financial risk, operational risk, and reputational risk?

L3 · 13 questions

Operational Machinery

The six working pillars: policy, risk, lifecycle, data, third-party, and assurance.

Q3.1Validation Report Coverage

What share of production AI models have current independent validation reports on file?

Q3.2Vendor Risk Classification Coverage

What share of third-party AI vendors have completed vendor risk classification and due diligence per the AVRF from Chapter 14?

Q3.3Incident Response Rehearsal Cadence

Has the institution rehearsed an AI incident response in the past twelve months?

Q3.4AI Data Asset Governance

Are AI training data assets governed under the same data classification, lineage, and quality controls as the institution's other data assets?

Q3.5Material Finding Closure Discipline

What share of material validation findings are closed within their assigned deadline?

Q3.6AI Governance Committee Operating Cadence

Does the AI Governance Committee meet on a documented cadence (monthly is the Chapter 10 default for material institutions) with attendance, decisions, and follow-ups recorded?

Q3.7AI Governance Office Staffing Adequacy

Is the AI Governance Office staffed at the level Chapter 11 specifies for the institution's portfolio size and regulatory perimeter?

Q3.8Model Inventory Completeness

Is the AI model inventory complete, current within 30 days, and reconciled with engineering deployment systems and procurement records?

Q3.9Bias Testing Discipline

Are AI systems handling protected characteristics (gender, nationality, ethnicity, age) tested for bias, with performance consistency measured across demographic groups?

Q3.10Documented Limitation Awareness

Are known limitations, edge cases, and out-of-distribution scenarios documented for each production model, with mitigation strategies?

Q3.11Model Lifecycle Governance

Is there documented governance for model updates, retraining, and replacements, including approval, testing, and stakeholder notification before deployment changes?

Q3.12Human Oversight Mechanism for High-Risk Decisions

For high-risk AI systems affecting legal rights or significant interests, are human-in-the-loop mechanisms operational, with documented qualifications for the human reviewer and audit trail of the review?

Q3.13Post-Incident Review Discipline

Are post-incident reviews conducted analyzing root cause and implementing preventive controls, with findings tracked through the AI Governance Committee?

L4 · 12 questions

Technical Substrate

The instrumented ground: registry, monitoring, lineage, and logs that make every claim auditable.

Q4.1Drift Monitoring Coverage

What share of production AI models have active drift monitoring per Chapter 12.8?

Q4.2Explainability Infrastructure Operability

For customer-facing AI subject to the automated-decision-disclosure provisions, is per-decision explainability operational?

Q4.3Data Residency Enforcement by Architecture

Does the institution's AI technical architecture honor data residency by design or by exception?

Q4.4Adversarial Robustness Testing

Are high-risk models tested against adversarial inputs?

Q4.5Decision Audit Trail Retrievability

Can the institution produce a complete decision audit trail (input data, model version, output, explainability artifact) for any production AI decision within 24 hours?

Q4.6Model Versioning and Rollback Capability

Does the institution maintain immutable model versioning with operational rollback capability tested at least annually?

Q4.7Inference Logging Completeness

Are inference inputs, outputs, confidence scores, and decision metadata logged with cryptographic integrity sufficient for evidentiary use?

Q4.8Model and Data Provenance Tracking

Is provenance tracked end-to-end for each model (training data lineage, code version, hyperparameter set, validation evidence) such that the institution can reproduce any production decision?

Q4.9Privacy-Preserving Computation Infrastructure

Where the regulatory perimeter or institutional risk appetite requires it, does the institution operate privacy-preserving computation (differential privacy, federated learning, secure enclaves) in production?

Q4.10Access Control and Segregation of Duties for AI Systems

Are access controls and segregation of duties enforced across the AI lifecycle (data scientists cannot deploy to production, validators cannot modify model code) with audited evidence?

Q4.11Secure Software Supply Chain for AI Components

Are AI libraries, models, and components governed under the institution's software supply chain controls (signed artifacts, vulnerability scanning, dependency attestation)?

Q4.12Production Observability for AI Workloads

Are AI workloads observable in production at the level of latency, throughput, error rate, output distribution, and resource consumption with alerting tied to the incident response runbooks?

§ 03Method

How the score is produced.

Scoring Methodology

The Self-Assessment produces three artifacts: a per-question score, a per-layer score, and a four-element profile vector. The methodology is deliberately simple because the discipline lives in the three rules, not in the arithmetic.

Per-Question Scoring

Each question is scored Level 1 through Level 5 against the rubric provided. The score is the lowest level the institution can defend under examination by the rules above. Where the institution sits between two levels (the policy describes Level 4 behavior, the operating reality reflects Level 3), the lower level is recorded. Where others see a scoring rubric, I see a contract the institution writes with itself about which version of its own story it is willing to defend.

Per-Layer Scoring

Per-layer scores are the arithmetic mean of the question scores within that layer, rounded to the nearest integer. Half-points round down rather than up. This rounding rule is deliberate. The institution that earns a 3.5 has not yet earned Level 4. It has reached the boundary where Level 4 becomes achievable with sustained discipline. The rounding pulls the institution back to the level it can presently defend.

The Profile Vector

The four per-layer scores form a profile vector of the form (L1, L2, L3, L4). A typical pattern observed across MENA financial-services institutions in the 2024 through 2026 period is (L1: 3, L2: 2, L3: 2, L4: 1). The vector is the institution's MESA fingerprint. It identifies which layer is the binding constraint and therefore which chapter discipline the institution must build first.

Triangulation and Independent Review

Each layer's questions are answered by at least two parties. The layer owner produces a first draft. An independent reviewer drawn from internal audit, risk, or an external assessor produces a second. Where the two parties agree, the score is the agreed score. Where they disagree, the lower score wins and the disagreement is recorded as a finding for the AI Governance Committee. The disagreement is the most valuable output of the Self-Assessment. It surfaces the conversations the institution had been avoiding.

Boundary Discipline

At each level boundary, the conservative reading wins. Level 3 requires the discipline to be operational and reviewed. Level 4 requires it to be measured. The institution that has dashboards but does not act on the readings is at Level 3, not Level 4. The institution that has policies but does not follow them is at Level 1, not Level 2.

Re-Assessment Cadence

The full Self-Assessment is re-administered annually. Layer 3 questions are re-administered quarterly because Layer 3 is the layer that moves fastest. Layer 1 questions are re-administered at any material regulatory event (new framework, amended guidance, jurisdictional expansion). The annual re-assessment produces a year-over-year delta the board reads as the program's signal.

Profile-to-Reading-List Output

The profile vector maps directly to a chapter reading prioritization. The lowest layer is read first because the lowest layer caps everything above it. Within Layer 3, the lowest pillar is read first because the pillars depend on each other in a documented order: data governance is the substrate, MRM the discipline, vendor risk the perimeter, incident response the recovery mechanism, the operating model the cadence, the governance office the staffing.

The reading list is generated by the following logic.

A Level 1 in Layer 4 directs the reader to Chapters 9 (data architecture) and 12 (MRM technical substrate) first, because Layer 4 is the substrate every other layer rests on. A Level 1 in Layer 3 directs the reader to Chapters 12 (MRM) and 13 (data governance) first, because these are the upstream pillars. A Level 1 in Layer 1 directs the reader to Chapters 5 through 8 (the jurisdictional chapters) first, because the regulatory perimeter is the perimeter against which everything else is judged.

The recurring profile vectors map to recurring reading lists.

The regulator-driven profile (3, 2, 2, 1) has Layer 1 strongest because supervisors have been asking direct questions, and the other layers lag because they are not yet under examination pressure. The reading priority is Chapters 9 and 12 first to address Layer 4, then Chapters 13 through 15 to lift Layer 3, then Chapter 10 to integrate the operating model, then Chapter 4 to anchor the program at Layer 2.

The vendor-driven profile (2, 2, 3, 2) has stronger Layer 3 vendor risk because the institution moved through a major vendor transformation. The other Layer 3 pillars lag. The reading priority is Chapters 12 and 13 to strengthen the upstream pillars vendor risk depends on, then Chapter 15 to complete the incident response perimeter, then Chapters 5 through 8 to firm up Layer 1, then Chapter 4 to re-anchor at Layer 2.

The fintech profile (2, 3, 1, 3) has strong Layer 2 (AI is the business model) and strong Layer 4 (the technical substrate is the business) but a weak Layer 3 (the operational machinery has not been built at scale). The reading priority is Chapters 10 and 11 to build the operating model and governance office, then Chapters 12 through 15 to install the Layer 3 disciplines, then Chapters 5 through 8 to confirm jurisdictional coverage.

The reading list is not a syllabus. It is a sequenced intervention. Reading Chapter 4 before reading Chapter 12 produces a strategic frame without operational substance. Reading Chapter 12 before reading Chapter 13 produces an MRM function operating on ungoverned data. The sequence matters because the dependencies matter.

The output the Self-Assessment generates for each institution is a single page with three elements: the profile vector, the prioritized chapter reading list with target completion dates, and the recommended 18-to-24-month MESA roadmap calibrated to the institution's profile per Chapter 4's phase definitions. The page is the institution's working document for the next twelve months. It is reviewed quarterly at the AI Governance Committee and re-baselined annually at the full re-assessment.

This companion appendix is licensed CC BY-NC-ND 4.0, Attribution-NonCommercial-NoDerivatives: share it with credit to the author, but not for commercial use and not as a modified version. The book itself and the named frameworks (the MESA Framework, the Five-Gate Deployment Model, the AI Incident Response Protocol and the others) are © 2026 Nabeel Khan, all rights reserved.

§ 04Stated limits

What this self-assessment does not claim.

Read this before you rely on it

  • It is a self-report. You answered these questions about your own institution, which is the honest limit of every self-assessment and the reason the three rules matter more than the questions do.
  • A profile is a locating aid. It is not an examination, and it is not a record a board or a supervisor can rely on, because nothing here traces a claim to an artifact.
  • Per-layer scores are arithmetic means with half-points rounded down. That rounding is deliberate and conservative, and it means a strong layer average can still sit a level below where a charitable reading would put it.
  • The Sharia validation question is excluded from its layer average when marked not applicable, rather than counted against a conventional institution.
  • Scoring is client-side. Nothing is transmitted and nothing is stored, which also means nothing is saved when you close the tab.
  • This is reference material and advisory practice. It is not legal advice, and it does not substitute for your counsel or your regulator relationship.

The lowest layer is the one that caps the rest.

Fin · Self-assessment