Why fraud models that pass validation in July fail in December, and how to measure the gap before the season measures it for you.
Five findings that carry the argument, each mapped to the section that evidences it. Closes on the report’s thesis: fraud losses are no longer a detection problem—they are an assurance problem.
What peak actually does to a model. Four stressors that arrive at once rather than in sequence, with separate calendars for banking, fintech and insurance.
The technical core. Six failure modes, none of them a modelling error. Includes the label-lag trap and a table pairing each mode with the test that catches it.
Why these tests are rare in production, framed as three mismatches: capacity, cadence and coverage. Includes the two-clocks view.
Six dimensions, one per failure mode, each rated across four maturity levels and resolving to one of four tiers.
Working backwards from the October freeze with four phases: See, Stress, Rehearse and Hold.
Most fraud-AI content stops at “models drift.” This report specifies six mechanisms and the test that detects each.
Six ratings. One score. One maturity tier.
Designed around the models you already have and the freeze you cannot move.
Every statistic is sourced with publisher, publication date and methodology.
Fraud operations, model risk, data science and technology each receive a distinct perspective.
| If you are | What you get |
|---|---|
| Head of Fraud or Financial Crime | The loss mechanics of peak season, and the false-positive economics that turn a stable rate into a revenue problem at volume. |
| Model Risk or Model Validation | Effective challenge expressed as testable engineering, mapped to SR 26-2, with your judgment and sign-off untouched. |
| VP or Director, Data Science / ML | Six failure modes that are not modelling errors, and the instrumentation case that protects your models’ credibility after a hard season. |
| CTO or CIO | A defensible answer to “who owns is-it-still-working?”, and evidence formatted for whoever eventually asks. |
| Head of QA or Engineering | Load, fallback and adversarial testing applied to the production scoring path, not just the application around it. |
Because validation and peak season test different things. Validation examines a model against historical, averaged data, while peak season subjects it to four simultaneous shifts: transaction volume compressed into hours, an attack mix that fraud rings deliberately schedule for the season, a surge of thin-file customers where model confidence is weakest, and alert queues that outrun analyst capacity. The report documents six specific failure modes that emerge under those conditions. None of them is a modelling error, and all of them are structurally invisible to an annual review run on historical averages.
SR 26-2, Revised Guidance on Model Risk Management, was issued jointly by the Federal Reserve, the OCC and the FDIC on 17 April 2026, and supersedes both SR 11-7 (2011) and SR 21-8 (2021). It is the first rewrite of US model risk guidance in fifteen years. Three points matter for fraud models. It retains effective challenge, now defined as critical analysis conducted by objective experts with sufficient independence and the standing to effect change. It sets a risk-based approach tailored to model materiality, and is described as most relevant to banking organizations above $30 billion in total assets, with a stated exception for smaller institutions whose model estates are notable for their prevalence and complexity. And it places generative and agentic AI explicitly outside its scope, while traditional statistical and machine-learning models, including the models scoring fraud, credit and AML decisions, remain covered.
Weekly at minimum, and more often through peak season. Monthly stability reviews are common but too slow for fraud: input distributions can shift materially across a single promotional weekend, and averaging that spike into a monthly figure turns it into a rounding error, reviewed two weeks after scores have already moved. The report recommends weekly drift monitoring with baselines built from the previous year’s peak weeks rather than a calendar-year average, so that December is compared against December.
The institution deploying the model, not the vendor that sold it. SR 26-2 calls for validation of vendor products either by internal or outside parties, with ongoing monitoring and outcomes analysis to confirm they remain fit for purpose. For insurers, the NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers holds companies responsible for AI outcomes whether a model was built in-house or bought, and the NAIC is developing a dedicated third-party data and models framework. In practice, a vendor’s benchmark deck is marketing and a SOC 2 report covers the vendor’s controls rather than your outcomes. What a due-diligence file needs is independent evidence that their model performs on your data, under your peak conditions.