Insight
The confidence gap — why model verification now runs ahead of model creation
A financial model has never been easier to produce — and never harder to be sure about. The two are moving apart, and 2026 is the year the distance became measurable.
Model volume is rising. Model trustworthiness is not.
Financial models are the engine of most investment and lending decisions. The decision is only as good as the model behind it — and the model is only as good as the last person who actually checked it.
Here is the uncomfortable arithmetic. Spreadsheet errors are not historical curiosities. In 2023, Norway’s $1.5tn sovereign fund lost ~$92m from a single wrong date in a benchmark spreadsheet. In 2024, the UK’s Office for Budget Responsibility restated ~£18bn of fiscal headroom after a calculation error in its March forecast. SEC filers in 2026 are still disclosing material weaknesses tied to formula omissions in earnings-per-share workpapers. Decades of research find that most complex models contain at least one error once independently checked (Panko; EuSpRIG). Typical cell-level error rates are often in the 1–5% range — which is not “1% of the deal”; it means hundreds or thousands of defective cells in a large transaction model. The question is not whether the spreadsheet works; it is whether the material cells were found before the decision. That pattern predates AI — but AI has changed the scale and speed of the problem. What changed in 2026 is what happens on the other side of the ledger.
February 2026 was the step change. AI tools became capable of building genuinely strong financial models — not fragments, but complete five-year models with balance sheets, cash flows and drivers. In tested builds, both Claude and Microsoft Copilot produced whole models that contained material errors; in one case the balance sheet did not balance. The tools got dramatically better at producing the model. They did not get better at verifying it — because, structurally, they cannot. A system that predicts the next cell cannot at the same time perform the independent, calculation-level check on the whole workbook. Those are different functions.
The result is a divergence, not an improvement. Model volume rises sharply. Model trustworthiness, on its own, does not. The distance between the two is the problem.
The constraint is not the models. It is the reviewers.
Ask a practitioner and they will say the same thing: the bottleneck was never “we don’t know models can be wrong.” It is “we don’t have enough people qualified to check them, and no time to check them all.”
Two findings from a July 2026 survey of 63 senior financial modellers (Financial Modeling Institute Global Leaders Council — an expert panel, not a random sample) show how deep the scarcity is:
- None said they would rely on an AI-generated model for a high-stakes decision without independent human review.
- 90% said sign-off on model outputs should never be fully delegated to AI.
Both point the same way. Demand for review is growing faster than the supply of reviewers — and the scarcity accrues to whoever can review. A person is a limited review capacity. A product is not.
AI multiplies the number of models that need reviewing without multiplying the number of people qualified to review them. Validation capacity, not model quality, becomes the constraint.
Why independent, and why a human, both matter
Two things must hold for a review to be worth acting on, and they are easy to confuse.
Independence. A model checked by the same system that built it is not independently verified. The builder — human or tool — has a reason to find its own work sound. Verification is only trustworthy when it comes from outside the creation process. That is the principle behind a lender's expectation of an independent reviewer.
A human who accepts it. Even the best verification is evidence, not a decision. A tool can tell you every formula is consistent and every statement ties. It cannot take responsibility for the conclusion, cannot sit across the table, cannot answer the question the client actually asked. Review produces the evidence; a person accepts it, and that acceptance — the professional view — is where responsibility lives.
Put the two together and you get the only combination a third party should rely on: an independent check, produced at scale, with a human who owns the outcome.
The regulatory clock has struck
Regulators are not waiting for the market to sort this out. On 2 August 2026, the EU AI Act's transparency, traceability and human-oversight obligations took effect; and on 6 August 2026, Fannie Mae's LL-2026-04 brought AI/ML governance under lender policy. The direction is consistent: machine-made analysis needs to be documented, traceable and accountable to a human — more institutions will need to show they verified what they relied on.
What closes the gap
The gap does not need a harder model. It needs verification that is independent, exhaustive, repeatable and available before every decision — not only on the handful of models large enough to justify a full external review cycle.
In practice that means an automated, deterministic review: every formula reconstructed from source, every dependency traced, every statement checked, every assumption classified — with cell-level evidence and a ranked action list. Then interpretation: the findings put into the context of the deal, so the reviewer can judge rather than merely inspect. And then the human: the analyst, partner or advisor who accepts the conclusion and owns the answer.
The outcome is not “the model is perfect.” It is stronger: you know where you stand before the decision — because the check was exhaustive, documented and independent, and the judgement was yours.
Sources (selected)
- NBIM / Norwegian Government Pension Fund Global — benchmark calculation error (Feb 2023; ~NOK 980m / ~$92m); EuSpRIG POB2401
- UK Office for Budget Responsibility — PSNFL headroom restatement (Oct 2024; ~£18bn)
- U.S. Global Investors, Inc. — SEC Form 8-K (1 Jun 2026; EPS spreadsheet material weakness)
- Panko, R. — spreadsheet error research (EuSpRIG summaries)
- Financial Modeling Institute — Global Leaders Council survey, The Human Financial Modeller (Jul 2026, n=63)
- ICAEW — How to identify AI errors in financial models (Jun 2026)
- EU AI Act (Regulation (EU) 2024/1689); Fannie Mae LL-2026-04 (effective 6 Aug 2026)
Related
- Independent verification and consulting review — how in-workbook verification complements formal external review
- Verify before high-stakes decisions
- Request beta access