The four-pillar model risk management framework.
The phrase model risk management framework is verbatim regulator language. Read with care, because the count is where secondary writing goes wrong. § III names three elements — development and implementation and use, validation, governance — and § VI carries documentation separately. "Four pillars" is Warrant's grouping of those, not a layout the supervisor named, and nothing in either letter numbers them. What the regulator does require is a framework, not a checklist: an integrated lifecycle it expects to see on every model the bank treats as material. The grouping below is Warrant's way of reading that lifecycle.
Placement matters, and so does the asset threshold, which is routinely misstated. SR 11-7 was issued by the Federal Reserve and the OCC jointly on 4 April 2011 — the attachment names those two agencies only — with the OCC's companion at Bulletin 2011-12, and the FDIC adopted the same guidance six years later through FIL-22-2017. There was no USD 1 billion threshold in the Federal Reserve's letter. The threshold is FDIC-specific and it lives in footnote 1 of the FDIC's version: the guidance is not expected to pertain to FDIC-supervised institutions under USD 1 billion in total assets “unless the institution’s model use is significant, complex, or poses elevated risk to the institution.” The Fed's cover letter says instead that the guidance “should be applied as appropriate to all banking organizations supervised by the Federal Reserve, taking into account each organization's size, nature, and complexity”. SR 26-2 replaced all of that with one figure: most relevant above USD 30 billion in total assets.
The 2011 letter was written when the typical material model was a credit-scoring regression or a Basel III IRB calculation. The text accommodates AI agents only because the model definition — in § III, not § II — is written use-driven rather than technique-driven: “a quantitative method, system, or approach that applies statistical, economic, financial, or mathematical theories, techniques, and assumptions to process input data into quantitative estimates.” If the output of the artefact drives a bank decision, the artefact is a model. SR 26-2 went the other way: it narrowed the definition, moved it into § II, and then carved generative and agentic systems out of it by footnote.
Warrant's four pillars, paraphrased and indexed against SR 11-7's own paragraph layout — the paragraphs are the regulator's, the four-way split is ours:
warrant-v1 carries an officer's name, role or tenant, so the record binds to the submitting system and not to a person. What the receipt does carry is cosign_status, with cosign_signed_actions of cosign_total_actions: a verdict on whether the customer co-signed the trace it submitted, which identifies a key holder rather than an accountable officer.
The current guidance · SR 26-2, and the hole in it.
SR 26-2 was issued jointly by the Federal Reserve Board, the OCC, and the FDIC on 17 April 2026 (OCC Bulletin 2026-13) as the Revised Guidance on Model Risk Management. It supersedes and replaces SR 11-7 (2011) and SR 21-8. The revision keeps the same lifecycle discipline SR 11-7 established — the one Warrant groups as four pillars — but restates it as principles-based and risk-tailored rather than prescriptive, and reads as most relevant to banks above USD 30 billion in assets. Smaller institutions are held to a proportionate standard. Four things in it matter to anyone running an agent, and the first one is not what the market expected.
First, the carve-out. Footnote 3, hung off the model definition in § II, reads in full: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models." That is an exclusion, not a proportionality clause. No transition period, no phase-in, no future guidance promised in the text.
Second, what the guidance does not contain. Across the whole 2026 attachment the phrase artificial intelligence appears 0 times, machine learning 0 times, large language model 0 times, retrieval 0 times, and runtime 0 times. The words generative and agentic appear twice each, both inside footnote 3, both in the sentence that excludes them. Any read of SR 26-2 as an AI-scoping instrument is reading a document that does not exist.
Third, the model definition narrowed. SR 26-2 § II reads the term as "a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates," and expressly excludes simple arithmetic, deterministic rule-based processes, and software with no statistical, economic, or financial theory underpinning it. Compared with 2011 the word complex is new and mathematical is gone. An agent's per-action choice of which tool to call was never a quantitative estimate; on the narrowed definition it is doubly out.
Fourth, what survives the carve-out, which is the part that costs money. Footnote 1 keeps the backstop live: supervisory action may result for any violations of law or unsafe or unsound practices stemming from insufficient management of model risk. The excluded system routes to the bank's general risk management and governance practices, which have no clause structure to comply with. And the documentation bar dropped: SR 11-7's requirement that documentation be sufficiently detailed that parties unfamiliar with a model can understand how it operates did not carry into SR 26-2, whose documentation subsection runs two sentences. The examiner still asks what your monitoring and testing framework for the agent is. The letter that used to answer for you has been withdrawn from the question.
Read commercially, the carve-out is worse for a deployer than in-scope treatment would have been. In-scope means a known bar and a known evidence set. Out-of-scope with a live safety-and-soundness backstop means the bank writes the bar, and the first examiner to read it decides whether it held. The conventional models the agent calls stay fully in scope, so a hybrid system is examined under two regimes at once, one of which does not exist yet.
Governance · the board signoff.
Governance starts at the board. Model risk policy is a board-approved policy. The chief risk officer or a designated risk-management head owns enterprise model risk. An independent model risk management function performs validation and ongoing review. Internal audit reviews the function. SR 11-7 describes that four-tier structure (board, CRO, MRM, audit) as what sound governance looks like; SR 26-2 § I is explicit that its guidance sets no enforceable standards and that non-compliance with it will not result in supervisory criticism, so read it as a supervisory expectation attaching to the conventional models the letter covers — not a rule, and not a duty an agentic system carries. Warrant's inference, not the regulator's statement: it is the structure a supervisor is likeliest to ask about first.
For an AI agent acting inside a bank, the live question is who in this chain holds personal liability when the agent issues a decision. The answer SR 11-7 required was binary: a named human officer signs off on the model and remains accountable for its outputs. SR 26-2 § VI keeps a softer version for models it covers — sound governance practices delineate the individual or individuals responsible for key activities throughout the model lifecycle, from development through validation and ongoing monitoring. For a system footnote 3 excludes, no letter names anyone. The bank does, or nobody does. The named officer's role under the bank's risk policy is what the supervisor binds the model to. The agent does not displace that accountability; it inherits it.
GAO B-331324 is worth naming precisely, because it is often cited as though it were an alternative reference number for SR 11-7. It is not. It is a Comptroller General decision of 22 October 2019, Board of Governors of the Federal Reserve System — Applicability of the Congressional Review Act to Supervision and Regulation Letter 11-7, which concluded that SR 11-7 was a rule for Congressional Review Act purposes and therefore should have been submitted to Congress. That is a decision about the letter's procedural status, not a restatement of its contents.
Warrant does not evidence that chain. The evidence package has no field for an officer's name, role or tenant, and none for a policy version. Its roots are classification, actions, authorizations, obligations, coverage_by_regime, deferred_regimes, risk_tier, refusal_reason and trace_metadata, and the schema is additionalProperties: false, so a package cannot carry a sign-off field the spec does not define. What the record does fix is narrower and real: trace_metadata.timestamp fixes the moment the package was sealed — it is created at attestation, so it is seal time, not decision time, and warrant-v1 has no signed decision-time field at all — trace_metadata.regulations_corpus_sha256 binds the package to the exact corpus of obligations it was read against, and the receipt's cosign_status reports whether the customer co-signed the trace it submitted. That is a system-level and key-holder-level record, independently verifiable without contacting Warrant. It is not the equivalent of a wet-ink CRO sign-off, and the governance leg stays unevidenced until the spec carries a named-owner field.
Development standards · the build phase.
The development pillar sets the build-phase test. Three questions: model selection rationale (why this technique), conceptual soundness (does the technique fit the problem), and development testing (was implementation correctness verified before deployment). The classical 2011 application produced a model methodology document, a development testing report, and a code review. Each was a static artefact lodged before deployment.
For an AI agent, the same three questions apply but the artefacts multiply. The model is not a single estimator. The model is the agent: a foundation model, plus a tool registry, plus a prompt template, plus a retrieval policy, plus a guardrail layer, plus a post-processing rule. Each component is a developmental artefact under SR 11-7. Each version is dated, owned, and tied to a development testing record before the agent is allowed to issue a decision.
The development standard has practical force whichever regime a supervisor reaches it under. What follows is Warrant's inference about examination practice rather than anything either letter says: an examiner walking development documentation for an LLM-driven agent would be expected to pull the prompt template version against the foundation model version, then pull both against the live policy version, and a mismatch (production prompt against a stale foundation, or a current foundation against an undocumented prompt) is where we would expect a gap to surface. On that reading, banks that treat the prompt template as a configuration value rather than a versioned artefact carry the most exposure.
The retrieval policy deserves separate mention, and it is where the carve-out bites hardest. SR 26-2 never uses the word retrieval, and a generative system that selects its own context is out of scope by footnote 3. So the corpus index, the similarity threshold, the chunk window, and the rerank logic are developmental decisions with no clause attached to them. They still change what the agent does. A corpus that ingested a new document set without a testing pass is a material change to behaviour that no letter now obliges the bank to catch, which is precisely why the bank has to catch it and be able to show it did.
Effective validation.
The validation pillar is the one most model-risk writing leads with, and Warrant does not claim it is the most cited in Matters Requiring Attention — those findings are not published, so nobody outside the bank and its examiner can rank them. Validation is independent. Validation evaluates conceptual soundness, including developmental evidence; performs ongoing monitoring, including process verification and benchmarking; and runs outcomes analysis, including back-testing.
The independence test has procedural force. Validation has to be performed by people not involved in development. In practice the bank's MRM function or an external party. Validators must have authority to challenge developers and influence to elevate findings to the CRO and board. SR 11-7 calls this effective challenge. SR 26-2 § III keeps the term and states that it runs throughout the model lifecycle, from model development to ongoing monitoring — for models inside its scope, which an agentic system is not.
Outcomes analysis is the empirical leg. Outputs of the model are compared to actual realised outcomes. For a credit model: predicted probability of default versus actual default rates. For an AI agent: the agent's classification, recommendation, or score versus a gold-labeled ground truth. The Warrant eval suite (see /blog/regulator-grade-evals) operationalises outcomes analysis for an LLM-driven agent: expected outcomes are compared against the agent's actual output, and the harness pins the model ids it ran. Keep the two artefacts apart, because they are not the same record — the pinned model version lives in the eval harness's own output, and the signed evidence package has no model-version field at all.
Sensitivity analysis is the operational leg. The validator perturbs inputs and observes output stability. For an AI agent, sensitivity testing reads as adversarial robustness: prompt injections, paraphrased queries, edge-case populations. A validation report that does not include sensitivity testing is incomplete under the validation standard. Warrant makes no claim about how often that gap appears in Matters Requiring Attention: MRA findings are not published, so any ranking of them is unverifiable from the public record.
Replicate-the-output is the practical test a supervisor can be expected to run at horizontal review. Given a real production decision, can validation reproduce the model's reasoning. For a deterministic linear model, trivial. For a stochastic LLM, the test is non-trivial — and this is the point at which the honest answer about the evidence package is a limit, not a capability. The signed package does not capture the exact inputs, the model version, or the outputs. Those are read from the submitted trace and are not re-emitted, and no model_version property exists in the schema, so replay against the package is verification that a sealed record with these obligation determinations exists — not re-execution of the decision. The material a replication actually needs stays in the bank's own trace store and model inventory.
Comprehensive documentation and audit trails.
The word comprehensive is verbatim regulator language and load-bearing. The documentation pillar requires comprehensive documentation and supplies the third-party replicability test. It sets the standard a supervisor reads on examination: a knowledgeable third party (the examiner) must be able to reconstruct the model's design, intended use, limitations, validation findings, and recent material decisions from the documentation alone.
Three documents satisfy the static portion: the model methodology document, the validation report, and the model card. Together they describe what the model is. None of them, individually or collectively, satisfies the runtime portion: the audit trail of what the model actually did.
The audit trail is where AI agents diverge sharply from 2011-era models. A credit-score regression issues one decision and the audit trail is a row in a database with a model version stamp. An AI agent issues a sequence of tool calls, retrievals, and intermediate reasoning steps before producing the customer-facing decision. An audit trail that answers an examiner has to reconstruct what the model did, when, and why, at the granularity of each tool call and each retrieval. Anything coarser leaves the question unanswered. For a conventional in-scope model that bar is the third-party replicability standard; for an agentic system it is the bank's own, since SR 11-7's replicability sentence did not carry into SR 26-2 and footnote 3 puts the agent outside the guidance in any event. The same audit-trail gap is read against the NYDFS Part 500 rule in why standard logs do not satisfy 23 NYCRR § 500.6.
This is the per-decision evidence question Warrant addresses, and it addresses part of it. No obligation compels it of an agentic system: SR 26-2 sets no enforceable standard and its footnote 3 excludes the agent, so a bank that answers the question answers it as voluntary governance. Per action the package carries actions[*] — action_id, actor, action, subject — plus an authorizations[] row holding within_purpose, preconditions_met, human_oversight_appropriate, reversible, justification and confidence, and the obligations rows those feed, each with a compliance status and an evidence string. trace_metadata carries package_id, timestamp and regulations_corpus_sha256. It does not carry the trace's raw per-action inputs and outputs, which the pipeline reads and does not re-emit, and it has no field for alternatives considered, no model version, no policy version and no accountable officer. So the record fixes the moment it was sealed and is independently verifiable without contacting Warrant — trace_metadata.timestamp is written at attestation, which bounds the record from one side and does not place the action in time. Placing the action itself in time is done from the entity's own trace. There is no per-action timestamp, no model version and no exact input or output anywhere in warrant-v1. On retention, cite the rule that actually applies rather than a generic floor: neither SR 11-7 nor SR 26-2 sets a retention period for model documentation, so the operative figures come from elsewhere — for a NYDFS Covered Entity, 23 NYCRR § 500.6(b) requires five years for § 500.6(a)(1) records and three years for § 500.6(a)(2) audit trails, and § 500.17(b)(3) five years for records supporting the annual certification.
The model inventory question.
The model inventory standard is short and unforgiving. Every model in production. Version. Owner. Last validation date. Residual risk classification. Banks have run model inventories for credit, market, and operational risk models since 2011. AI agents make the inventory question urgent in three ways.
First, when does a prompt change become a new model. No letter answers this for an agent, because footnote 3 removed the agent from the letter. The bank's own policy has to draw the line, and the events that should cross it are not subtle. A prompt-template rewrite that broadens use from internal classification to external customer-facing decisions is one. A guardrail relaxation that admits previously-blocked categories is one. A persona shift that changes the agent's tone in regulated communications is one. Each triggers a fresh inventory entry, a fresh validation pass, and a fresh documentation snapshot.
Second, when does a retrieval policy change require re-validation. Again: not addressed. A new document set in the corpus, a similarity threshold tuned more permissive, a rerank logic swapped in — each changes the agent's output distribution, and none of them trips a clause. Banks that ran retrieval-augmented agents through 2025 on the assumption that model risk management covered them are the population most exposed in 2026, because the framework they were relying on was withdrawn from their system in April.
Third, when does a foundation-model swap reset the clock. A swap from one provider's frontier model to another, or to a newer version of the same model, is a substantial change without exception. Validation repeats. Documentation is re-issued. The inventory entry is updated. Warrant makes no claim about what MRA or MRIA findings say on this: those findings are not published, so nobody outside the bank and its examiner can cite them. What is on the public record is the OCC's Spring 2026 Semiannual Risk Perspective naming "validation challenges where industry approaches are evolving" as one of the unique risks of generative and agentic AI.
Warrant has no field for this, and there is no metadata root in the evidence schema to hang one off. The provenance root is trace_metadata, and it carries package_id, timestamp and regulations_corpus_sha256 — which decision, when it was sealed, and which corpus of obligations it was read against. Nothing in it names an MRM inventory row. The nearest thing in the package is actions[*].actor, which carries whatever actor string the emitting trace supplied: in Warrant's US underwriting sample the actor on every action reads claude-opus-4-7, the foundation model, and the deployment name small_business_underwriter_us_v1 sits at the root of the submitted trace rather than inside the signed package. That sample's fourth step does carry model_id and model_validation_record_id in its raw step inputs — and raw step inputs are read by the pipeline and not re-emitted, so neither reaches the package. A name the emitter chose is not a row in a validated inventory, and no field in the package resolves one against the other. So the walk from one decision back to a model card and an active validation record is not a walk this package supports. It is worth stating plainly, because the bank that cannot make that walk takes weeks of internal investigation to answer the examiner and often produces a partial answer, and that is itself the gap finding.
Recent enforcement signal.
This is the section where model-risk writing usually overreaches, so read it as a correction. Warrant has not found a published US banking enforcement action that cites SR 11-7 or SR 26-2 by number. What the record does contain is narrower, and the difference matters when counsel has to stand behind a citation.
Wells Fargo, 20 April 2018. The OCC assessed a USD 500 million civil money penalty against Wells Fargo Bank, N.A., ordered restitution, and required an effective enterprise-wide compliance risk management program, for "unsafe or unsound practices" connected to collateral protection insurance on auto loans and interest-rate-lock extension fees. The Bureau of Consumer Financial Protection separately assessed USD 1 billion and credited the OCC's amount against its own. The finding was a compliance risk management finding: the word "model" does not appear in the OCC's release. The larger figures often attached to this matter belong to other settlements with other agencies, and no part of the USD 500 million was ordered for model risk.
Citigroup, 7 October 2020 and 10 July 2024. The Federal Reserve's 2020 cease-and-desist order recites "significant ongoing deficiencies in implementation and execution by Citigroup with respect to various areas of risk management and internal controls, including for data quality management and regulatory reporting, compliance risk management, capital planning, and liquidity risk management." Model risk management is not in that list; the word "model" does not appear in the order; and the order carried no monetary penalty — the USD 400 million announced the same day was the OCC's civil money penalty against Citibank, N.A. On 10 July 2024 the Board fined Citigroup USD 60.6 million for insufficient progress on data quality management under the 2020 action, which remains in effect; penalties announced by the Board and the OCC that day totalled approximately USD 135.6 million. The exposure for an agent deployed under a live order is real, but it is a data-quality and controls exposure, not a model-risk citation. Orders retrieved 6 August 2026.
The supervisory signal that does exist is in the OCC's Semiannual Risk Perspective, and it is cooler than the enforcement framing suggests. The Spring 2026 edition records that "Banks are taking a measured approach to the adoption of generative AI (genAI) and agentic AI, with usage generally limited to specific use cases with guardrails and human-in-the-loop accountability to manage risk," that observed use cases are "primarily productivity and customer experience enhancement tools," and that banks "may consider expanding their use of genAI and agentic AI for material financial decisions." It names the unique challenges as "lack of explainability, data privacy and data poisoning issues, cybersecurity threats, and validation challenges where industry approaches are evolving," and restates the carve-out: "GenAI and agentic AI models are novel and rapidly evolving and are not within the scope of the revised model risk management guidance." No 2025 or 2026 edition designates generative AI in lending as a heightened-risk activity, and the Spring and Fall 2025 editions do not use the phrase "model risk" at all. Editions retrieved 6 August 2026.
One forward-looking line in that document is the most consequential sentence a deployer can read today: "The agencies plan to issue in the near future a request for information that addresses model risk management generally and considers, in particular, banks' use of AI, including genAI, agentic AI, and AI-based models." So the gap is acknowledged and a consultation is signalled — but an RFI is not guidance, nothing has issued, and until it does the bank writes the standard. The OCC also issued Bulletin 2025-26, "Model Risk Management: Clarification for Community Banks" (6 October 2025), alongside 2026-13. On enforcement, Warrant makes no claim about what Matters Requiring Attention or MRIA findings cite: those are not published, and a page that characterises them is characterising documents nobody outside the bank and its examiner has read.
Where Warrant maps SR 11-7.
The mapping below names each operative SR 11-7 obligation and what the signed package actually carries against it. Read the third column literally. Two of the six resolve to a real field; four have no field in warrant-v1 and are printed as gaps rather than as shipped behaviour, and because the evidence schema is additionalProperties: false those four are not merely absent but prohibited until a new spec version defines them. The carve-out compounds it: SR 26-2 § II footnote 3 puts generative and agentic AI outside the guidance, so these pillars reach the conventional models an agent calls and not the agent itself, and Warrant does not claim to evidence obligations the guidance excludes. Warrant publishes this as the table it would put in front of an OCC or Federal Reserve examiner on horizontal review — that is our framing of an examination, not a procedure either agency has described.
| SR 11-7 pillar | What AI must evidence | What the package carries |
|---|---|---|
| Governance | Board-policy adherence per decision | authorizations[].preconditions_met — no policy-version field, so the board-policy leg is unevidenced |
| Development | Dev-time documentation per agent change | No field in warrant-v1 |
| Validation | Independent eval results per cohort | No field in warrant-v1 |
| Documentation | Per-decision rationale and uncertainty | authorizations[].justification + authorizations[].confidence + obligations.<action_id>[].evidence — no alternatives-considered field |
| Inventory | Model identifier and version per trace | No field in warrant-v1; actions[].actor carries only the string the emitter chose |
| Use | Use-context check per decision | authorizations[].within_purpose — the run's stated purpose, not a documented intended use |
authorizations[].preconditions_met records whether the action's inputs satisfied its preconditions, as yes, no or uncertain. NO FIELD for the rest of the pillar: warrant-v1 carries no policy identifier, no policy version and no deviation flag, so nothing in the package binds a decision to the board-approved policy that was current when it was made.
warrant-v1 carries no development provenance — no foundation-model version, no prompt-template digest, no tool-registry digest, no retrieval-policy identifier, no dev-test record. classification defines domain, jurisdictions, regimes, risk_tier, risk_tier_justification and two confidence values, and nothing else: it is additionalProperties: false. State one of those precisely rather than counting it as data — classification.regimes is defined in the schema but arrives empty in production, because the classify stage emits its regime list under a different key than the one the aggregate reads, so the producer's value is dropped. The dev-time record stays a bank artefact outside the package either way.
regulator_evidence root in the evidence schema and no eval-suite reference anywhere in it — no pass/fail, no calibration curve, no sensitivity matrix — and nothing in a package attributes a validation to a validator distinct from whoever issued the decision. Independent-validation evidence has no home in warrant-v1 today.
authorizations[].justification carries the per-action reasoning, authorizations[].confidence the stated uncertainty, and obligations.<action_id>[].evidence the per-obligation evidence string — the audit-trail leg of comprehensive documentation. NO FIELD for alternatives: actions[*] carries only action_id, actor, action and subject, so a path the agent weighed and discarded is not recorded.
metadata root in the evidence schema, no MRM inventory identifier and no model-version lineage. trace_metadata carries package_id, timestamp and regulations_corpus_sha256. actions[*].actor carries whatever actor string the emitting trace supplied — claude-opus-4-7 on every action in Warrant's US underwriting sample, with the deployment name small_business_underwriter_us_v1 at the root of the submitted trace and not in the package — which is a name the emitter chose, not a row in a validated inventory, and no field resolves it against one.
authorizations[].within_purpose records whether the action was within the system's stated purpose for the run, as yes, no or uncertain, at decision time rather than at quarterly review. Read its scope precisely: the purpose is the one the submitted trace declares for that run, not a documented intended use held in an inventory, and no field in the package ties the two together.
The comprehensive-documentation clock.
SR 11-7 sets the third-party replicability standard: documentation must allow a knowledgeable third party (an examiner) to reproduce the model's reasoning. For deterministic models the standard is mechanical. For AI agents, the standard meets a hard problem: the model is stochastic, sampling from a distribution at each token, and a literal replay of the same prompt against the same foundation model can produce a different output.
Two artefacts address the reproducibility clock, and neither of them is the whole answer. The deterministic eval suite runs canonical traces against recorded responses; the suite is the model’s behavioural baseline at the version the harness pinned. It is not seeded: the Messages API exposes no seed parameter, so what the harness fixes is the input set and the pinned model version, not the sampling path. The signed evidence package then fixes the obligation determinations for a real production action and verifies the same way across regenerations. What the package does not bind is the decision content: no input, no output, no model version, and no per-action timestamp — the submitted trace carries those and the pipeline does not re-emit them. So the third leg of a replication, the production input-output pair that a validator would replay, comes from the bank's own trace store. Presenting the package as a self-contained replay record would overstate it.
The companion note at /blog/four-layer-evidence-stack sets out the construction in full. The record is independently verifiable without contacting Warrant, and read its time value precisely: trace_metadata.timestamp is written at attestation, so it is the regulator-readable seal time and it bounds the record from one side only. It is not the time the decision ran, and warrant-v1 carries no signed decision-time field. Retention is where this gets cited loosely, so take it from the instruments that actually bind: neither SR 11-7 nor SR 26-2 fixes a retention period for model documentation at all. For a NYDFS Covered Entity, 23 NYCRR § 500.6(b) sets five years for the § 500.6(a)(1) reconstruction records and three years for the § 500.6(a)(2) audit trails, and § 500.17(b)(3) sets five years for the records supporting the annual certification. Longer horizons come from whatever sectoral rule actually applies to the product, which has to be named case by case rather than assumed.
A supervisor asking can you reproduce the model's reasoning on this decision is asking the third-party replicability question. On Warrant's reading, the bank that answers here is the trace, here is a sealed record of the obligation determinations that is independently verifiable without contacting Warrant, and here is the eval calibration profile at this version is answering closer to the standard SR 11-7 set than the bank answering with a screenshot or a quarterly aggregate. Warrant does not claim the package alone discharges the standard: without the input-output pair and the model version, which live in the bank's own records, part of the replication is outside the artefact.
Where SR 26-2 + SR 11-7 fit in the wider US map.
Model risk guidance does not run alone. The same AI agent inside the same US bank is read against an overlapping set of regimes, each from a different supervisor.
NYDFS Part 500 (23 NYCRR § 500). State regulator. Cybersecurity and audit-trail rule for any institution licensed by the New York Department of Financial Services. § 500.6(a)(2) requires audit trails "designed to detect and respond to cybersecurity events that have a reasonable likelihood of materially harming any material part of the normal operations of the covered entity" — the threshold is material harm to operations, not any contact with non-public information. Two limits worth carrying: § 500.6 is one of the sections a limited-exemption entity is exempt from under § 500.19(a), and the 72-hour notice at § 500.17(a)(1) runs on a cybersecurity incident under § 500.1(g), a narrower term than event.
Federal consumer protection. Cite the regulation, not the circulars — the circulars are gone. Consumer Financial Protection Circular 2022-03, "Adverse action notification requirements in connection with credit decisions based on complex algorithms" (87 FR 35864, 14 June 2022), and Circular 2023-03 (89 FR 27361, 17 April 2024) were both withdrawn effective 12 May 2025 by the CFPB's withdrawal notice at 90 FR 20084. What did not move is the enacted rule: Regulation B, 12 CFR § 1002.9(b)(2), still provides that the statement of reasons for adverse action "must be specific and indicate the principal reason(s) for the adverse action" and that "[s]tatements that the adverse action was based on the creditor's internal standards or policies or that the applicant … failed to achieve a qualifying score on the creditor's credit scoring system are insufficient." That sentence, not a withdrawn circular, is what an AI-driven denial has to satisfy. As at the eCFR text current to 1 August 2026.
Third-party risk. The Interagency Guidance on Third-Party Relationships (OCC Bulletin 2023-17; Federal Reserve SR 23-4, 7 June 2023; FDIC) is frequently cited as the model-risk pass-through, and it is not: SR 23-4 does not use the word "model" at all, and it describes the guidance as offering "the agencies' views on sound risk management principles" rather than imposing requirements. The model-risk hook is SR 26-2 § VII, "Vendor and Other Third-Party Products", which states that for vendor and third-party products "the principles of model risk management remain applicable" and that an important element is validation of vendor products. The bank is the supervised entity; the bank's practices are what the supervisor examines; the evidence package is what the bank produces to show them.
The Warrant evidence package satisfies all four overlapping regimes simultaneously. The argument is set out in full at /blog/one-agent-many-jurisdictions. One AI agent. One evidence shape. Four US supervisors reading the same artefact against four different paragraphs of four different regimes. The artefact economy is the point.
Questions a CRO and OCC examiner ask first.
Read the source directly.
- SR 26-2 · Revised Guidance on Model Risk Management · federalreserve.gov · 2026-04-17 (current; supersedes SR 11-7 and SR 21-8)
- OCC Bulletin 2026-13 · companion to SR 26-2 · occ.gov
- SR 11-7 original letter · federalreserve.gov · 2011-04-04 (superseded predecessor)
- Supervision and Regulation Letters archive
- OCC Bulletin 2011-12 · joint guidance
- FDIC FIL-22-2017 · adoption of model risk management guidance
- GAO B-331324 · supervisory consistency reference
- SR 26-2 attachment · Supervisory Guidance on Model Risk Management · federalreserve.gov · 17 April 2026 (source of the § II definition and footnote 3)
- Per-pillar Warrant evidence field mapping (deep-dive)
Authored by Warrant Compliance, the regulatory-analysis function at Warrant. [email protected]. Editorial commentary on regulatory text. Not legal advice. SR 26-2 (17 April 2026, OCC Bulletin 2026-13) is the current interagency guidance and supersedes SR 11-7 (4 April 2011) and SR 21-8. The quotations reflect the lifecycle text SR 11-7 established and that SR 26-2 carries into its principles-based restatement; the four-pillar grouping of that text is Warrant's, since § III names three elements and neither letter numbers them. Footnote 3 and the § II model definition are quoted verbatim from the SR 26-2 attachment as published at federalreserve.gov/supervisionreg/srletters/SR2602a1.pdf, retrieved 28 July 2026; the occurrence counts cited for that attachment were taken from the same file on the same date.