The PEVG and PARA review question set.
Chapter 7 of OSFI E-23 for AI Systems sets out every field of this template in its text. The file is a convenience; the method is in the book. Every field is below, before anything is asked of you.
PEVG is Planner, Executor, Verifier, Generator: four declared contracts inside an agent. PARA is Perception, Action, Reasoning, Adaptation: four faculties, each with its own kind of authority. Both are the author’s own patterns, each specified in its own deposit.
Every field, before the form.
Where a field cites E-23, the guideline states it. Where it says the discipline’s, the handbook’s method chose it, inside what E-23 permits. Open any part to see its fields.
Part A. Four contracts (PEVG)
PEVG requires a declared contract for each of four roles, stating five things: inputs, outputs, authority, prohibitions and evidence. Two roles may run in one model call provided the contracts stay separately stated and separately recorded (Chapter 7).
A.1 The contract for each role
| # | Field | Planner | Executor | Verifier | Generator |
|---|---|---|---|---|---|
| A1 | Held by: a model, a person, or both | ||||
| A2 | Inputs | ||||
| A3 | Outputs | ||||
| A4 | Authority (for the executor, an enumerated set of permitted operations, never a general permission with exceptions) | ||||
| A5 | Prohibitions | ||||
| A6 | Where its evidence is | ||||
| A7 | Contract version, and where the evidence records which versions were in force |
A.2 The validator's question for each role
| # | Role | Purpose | The prohibition that defines it | What the validator asks | Answer, or non-answer verbatim | Artifact, version, record or person named |
|---|---|---|---|---|---|---|
| A8 | Planner | Converts a task into ordered steps with explicit dependencies | Must not perform tool actions | Show me the plan as an artifact the executor consumed without interpreting it, and show me that the planner holds no tool. | ||
| A9 | Executor | Performs the tool actions the plan requires | Must not decide whether its own result is correct | Show me the enumerated set of permitted operations, and the record of one call as performed beside the call as planned. | ||
| A10 | Verifier | Decides what may be believed | Must not produce the response, and must not verify a result it produced | What were the acceptance criteria, per step class, and where in the record is an abstention, as distinct from a rejection? | ||
| A11 | Generator | Produces the response from what survived verification | Must not assert anything the verifier did not pass | Show me the mapping from each claim in the response to the verification decision that supports it. |
A.3 Versions, and the hosted model
E-23's review includes "reviewing third-party models and platforms or sub-components (including data and libraries) used for model development", and its monitoring tracks "external dependencies (for example, version updates)". An agent whose evidence records the prompt template version and not the model version has versioned half of its contract (Chapter 7).
| # | Question | Answer, or non-answer verbatim |
|---|---|---|
| A12 | Does the evidence record the model version as well as the prompt template version? |
A.4 Where the agent runs a function-calling loop
On the book's reading the contracts map onto the loop directly: the model proposing a call is the planner, and its plan artifact is the per-step record of the operation chosen, its parameters and the contract version in force; the runtime performing the call, only from the enumerated list and only under an authorization it checks, is the executor; a check that gates what is believed from each result is the verifier; the model writing the answer is the generator. In a single-model loop the result returns to the model that chose the call, and one call both decides whether to believe it and writes the answer, so the verifier has produced the response (Chapter 7).
| # | Question | Answer, or non-answer verbatim |
|---|---|---|
| A13 | Where does the check that gates belief sit: outside the call that chose the tool, as a deterministic test of the result or a separate call with stated criteria, or inside it? |
Part B. Two boundaries
"The verifier holds the epistemic boundary and does not hold the operational boundary." The epistemic boundary is a question about correctness; the operational boundary is a question about authorization, classification and policy. "A claim can be true, correctly verified, and still forbidden to leave the system." (Chapter 7, quoting the PEVG specification)
The requirement is a policy-enforcement point at the workflow's output, in addition to the verifier, that decides on classification, authorization and policy, does not take correctness as an input, can refuse a claim the verifier accepted, and records its refusal as a policy decision and never as a verification failure (Chapter 7).
| # | Question | Answer, or non-answer verbatim |
|---|---|---|
| B1 | Where is the verifier's decision recorded? | |
| B2 | Where is the policy-enforcement decision recorded? | |
| B3 | Show one recorded refusal of a verified claim, from production or from a test: a claim that is true and forbidden, put to the point at a recorded version and refused in the record |
If neither production nor the tests have such an entry, the second boundary does not exist, whatever the architecture diagram says (Chapter 7).
Part C. Four faculties (PARA)
A faculty is assigned by what a component may read and write, never by what it appears to be doing, because an institution cannot revoke an activity; it can revoke an authority. An agent may hold fewer than four faculties and remain conformant to PARA; an agent that exercises a faculty it has not declared does not (Chapter 7).
C.1 The faculties the agent holds
| # | Faculty | Authority type, in the specification's words | Held (yes / no) |
|---|---|---|---|
| C1 | Perception | "read-only access to system signals and emits structured observations" | |
| C2 | Reasoning | "read access to observations and runbooks, emits a plan, and touches nothing" | |
| C3 | Action | "the sole authority to change production, and only through enumerated, policy-authorized operations" | |
| C4 | Adaptation | "write access to institutional knowledge and no write access to production" |
"The loop runs perception, then reasoning, then action, then adaptation." Reasoning must precede Action for any consequential action (Chapter 7).
| # | Question | Answer, or non-answer verbatim |
|---|---|---|
| C5 | For every consequential action, does the record show the reasoning before the action? |
C.2 The entry for each faculty held
Every agent must have an entry with five fields (Chapter 7). Forbidden actions are named although they are formally the complement of the allowed set, because a reviewer cannot tell a capability deliberately withheld from one nobody thought of. Chapter 5 adds the date each was last set: finding every guardrail dated earlier than its metric is then a query across the entries, not a meeting.
| # | Field | Perception | Reasoning | Action | Adaptation |
|---|---|---|---|---|---|
| C6 | The faculty it holds | ||||
| C7 | Allowed actions, enumerated | ||||
| C8 | Forbidden actions, named explicitly | ||||
| C9 | The governing guardrail, by identifier and version | ||||
| C10 | The date the guardrail was last set (Chapter 5) | ||||
| C11 | Success metrics | ||||
| C12 | The date the success metric was last set (Chapter 5) |
C.3 One question to each faculty
| # | Faculty | Question (Chapter 7) | Answer, or non-answer verbatim |
|---|---|---|---|
| C13 | Perception | Does the observation distinguish what was measured from what was inferred? | |
| C14 | Reasoning | Can the faculty conclude that no action is warranted, and is that conclusion recorded with the same weight as a plan? | |
| C15 | Action | Are the permitted operations enumerated, is the guardrail recorded with each action, and can the agent widen its own guardrail? | |
| C16 | Adaptation | Which classes of knowledge may it write, and is the prior state retained? |
C.4 The guardrail on any write
Adaptation is bounded by the same guardrails as Action, and the guardrail must add three things (Chapter 7). OSFI's letter asks institutions to "establish internal criteria to determine when a self-learning model has materially changed", and E-23 names "autonomous re-parametrization" among the AI/ML challenges an institution should have processes for. In the discipline's terms the internal criteria are the Adaptation guardrail, and the material change is any write that guardrail did not authorize.
| # | The guardrail states | Answer, or non-answer verbatim |
|---|---|---|
| C17 | The classes of knowledge the faculty may write, enumerated | |
| C18 | That an adaptation derived from an incident involving the agent's own action requires human authorization | |
| C19 | That the prior state is retained |
Part D. One human
"Human in the loop" is a claim about topology. Where a role is held by a person, the contract still applies and the prohibitions apply unchanged; a human executor must not verify their own result (Chapter 7).
| # | Question | Answer, or non-answer verbatim |
|---|---|---|
| D1 | Which contract does the person hold? | |
| D2 | What does the record show them doing? | |
| D3 | What were the acceptance criteria, per class? | |
| D4 | Could the person abstain, and did the queue make abstention as cheap as approval? | |
| D5 | Did the person see the executor's record, the call as performed with its parameters, or only the executor's summary of it? | |
| D6 | Was the person's decision recorded per item, with its criterion, or was the queue's clearance the only record? |
A person who cannot abstain, cannot see the executor's record and leaves no per-item decision is not a verifier. E-23's rationale principle asks for alternative controls for autonomous models; a human approval point that cannot reject on a stated criterion supplies no alternative control (Chapter 7).
Part E. The non-answers
Five non-answers recur (Chapter 7). Where you hear one, write it down verbatim and ask the question beside it.
| # | The non-answer | Why it is one | What to ask instead | Heard (verbatim, where) |
|---|---|---|---|---|
| E1 | "The model handles that." | Names no role, no contract, no record. | Which role, under which contract version, and where is the decision recorded? | |
| E2 | "We have guardrails." | A guardrail with no identifier and no version cannot be shown to have been in force. | Which guardrail, by identifier and version, and which action carries it in the record? | |
| E3 | "It is in the system prompt." | A prompt is an instruction to an optimizer, not a boundary the substrate enforces. | What enforces it when the model ignores it, and what does the record show when that happens? | |
| E4 | "The vendor's model card covers it." | A card is the vendor's description of the vendor's model, dated. | Which version of the model does it describe, and is that the version in production today? | |
| E5 | "A human reviews everything." | A topology claim. | Under which contract, with what criteria, with abstention available, and where is each decision recorded? |
E-23 asks that the review be documented and reported "along with an overall recommendation on approval to the model approver", and that it evaluate "the level of explainability for the model workings as per the intended use of the model". A review report that records the engineer's answers, with the artifacts attached, answers the guideline's description in the discipline's form. A review report that records the non-answers, in the engineer's confident phrasing, is a description of the meeting (Chapter 7).
The claim, the test, the artifact (Chapter 7)
The claim. A validator reviews what an agent was permitted to do, and the response is only the last of the four things an agent produces.
The test. Ask for one recorded refusal of a verified claim, in production or in a seeded test; if none exists, neither does the second boundary.
The artifact. The question set answered for one agent, the engineer's answers and non-answers recorded verbatim.
nabeelkhan.com/e-23/review-questions. Questions: nabeelkhan.com/contact.
Your download has started
Next in the book: The AVRF questionnaire, scoring sheet and the Pellbrook worked example (Chapter 9). Or take the E-23 check to see which template matters most for you.
Where this sits.
The Office of the Superintendent of Financial Institutions (OSFI) does not endorse, approve or recommend this book, its author or any framework in it. Conformance with any framework named here is self-declared, by the institution, on its own record. Coldbrook, Thornbury and Pellbrook are fictional institutions, invented for the book.
Ask the engineer what the human in the loop can review in an hour.
Ask your AI assistant instead.
This page is a snapshot, accurate at the release it cites. The same corpus is callable, publicly and without a key, so an assistant can query it live and return an answer carrying the source it came from. For this page that is get_framework, which returns the Defensible AI Framework Registry entry for any framework these templates are built on (the Five-Gate Deployment Model, the AVRF, PEVG, PARA), with its version and the concept DOI of its deposited specification. It does not yet hold the E-23 handbook or the guideline itself; for those, this page and the book are the source.
claude mcp add --transport http concylium https://mcp.nabeelkhan.com/api/mcp
Claude Desktop, ChatGPT, Cursor, VS Code and Gemini CLI take the endpoint on its own: https://mcp.nabeelkhan.com/api/mcp. No key, no account, nothing to sign. Setup for every client.
“Using Concylium, get the Five-Gate Deployment Model and the AVRF from the framework registry, with their versions and DOIs, and tell me which gate a vendor model decision belongs to.”
A framework quoted from memory drifts. One returned from its registry, with the DOI of the deposited specification, does not.