Eighty Minutes Of Elevated Errors And A 57% Wrong-Answer Rate In One Day: Provider Failover And Grounded Calculation For Financial AI That Cannot Afford Either
Monday delivered both halves of the financial-AI reliability problem in a single news cycle. At 00:57 UTC on 22 September Anthropic began investigating elevated errors on Claude Mythos 5.1, Fable 5.1 and Opus 5 across claude.ai, the API, Claude Code and Cowork - an impact window of roughly eighty minutes before the incident was resolved at 02:35. The same day the FT covered Saturn's test of 18 models on 121 debt, mortgage, pension and tax questions, each asked five times: 57% of the 10,000-plus answers were wrong, 88% on complex multi-step questions, one pension answer would have cost a saver a £17,500 HMRC charge, and 57% of US chatbot users say they would act without checking. One is an availability failure, the other a correctness failure, and a financial system needs an engineering answer to both. Here they are, with code.
AlchmAI Engineering15 min read
~80 min
Impact window of the 22 September incident - elevated errors on Mythos 5.1, Fable 5.1 and Opus 5 from 00:57 UTC, resolved 02:35 UTC
57%
Of 10,000-plus answers from 18 models to 121 debt, mortgage, pension and tax questions were wrong, per Saturn's study covered by the FT
88%
Error rate on complex multi-step questions, with some models at 99%; paid tiers averaged 49% wrong, free tiers 63%
£17,500
The HMRC charge one wrong pension-tax answer would have exposed a saver to - and 57% of US chatbot users would act without checking
Reliability for AI in finance has two dimensions that get discussed separately and fail together. The first is availability: is the model there when the workflow calls it? The second is correctness: when it answers, is the answer right? Monday 22 September produced a clean example of each inside twenty-four hours, and the pairing is instructive because the engineering answers are different, both are necessary, and most financial deployments we review have neither.
The availability event was Anthropic's. At 00:57 UTC the status page moved to investigating elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1 and Claude Opus 5, affecting claude.ai, the API, Claude Code and Claude Cowork. The cause was identified at 01:17; Fable and Mythos returned to normal success rates at 01:35 with Opus 5 still degraded; all models were restored at 02:11 and the incident closed at 02:35, for an impact window the page records as roughly one hour twenty minutes. The correctness event was a study. Saturn tested 18 models across ChatGPT, Claude, Copilot, Gemini and Grok on 121 questions about debt, mortgages, pensions and tax, asking each five times for more than 10,000 responses; the FT reported the results on 21 and 22 September. Fifty-seven per cent of answers were inaccurate. On complex multi-step questions the failure rate was 88%, with some models at 99%. Paid tiers averaged 49% wrong, free tiers 63%. One Claude Haiku 4.5 answer on pension tax could have exposed a saver to a £17,500 charge from HMRC; another model invented a student-loan rule about moving abroad.
Half One: Availability Is A Routing Problem
The provider outage is the easier half, because distributed systems have handled dependencies that fail for forty years. The pattern is a circuit breaker per provider, a router that tries providers in a policy order, and - the part financial teams skip - an explicit, tested degraded mode for when every provider is out. The point of the eighty minutes is not that Anthropic had an incident; every provider does. It is that a workflow with one provider and no breaker was down for eighty minutes, and a workflow with a breaker and a second provider was not.
type State = "closed" | "open" | "half_open";
class Breaker {
private state: State = "closed";
private failures = 0;
private openedAt = 0;
constructor(private threshold = 5, private cooldownMs = 30_000) {}
canTry(): boolean {
if (this.state === "open" && Date.now() - this.openedAt > this.cooldownMs) this.state = "half_open";
return this.state !== "open";
}
ok(): void { this.failures = 0; this.state = "closed"; }
fail(): void {
this.failures += 1;
if (this.failures >= this.threshold || this.state === "half_open") { this.state = "open"; this.openedAt = Date.now(); }
}
}
interface Provider {
name: string;
// Snapshot per provider - pinned, never a floating alias, so a failover does
// not silently change the model that answered.
modelSnapshot: string;
call(req: LlmRequest): Promise<LlmResponse>;
breaker: Breaker;
}
/** Retryable = the provider is unhealthy. Not retryable = OUR request is bad. */
const RETRYABLE = new Set([429, 500, 502, 503, 504, 529]);
export async function complete(req: LlmRequest, providers: Provider[], ctx: Ctx): Promise<LlmResponse> {
const attempts: string[] = [];
for (const p of providers) {
if (!p.breaker.canTry()) { attempts.push(p.name + ":open"); continue; }
try {
const res = await withTimeout(p.call(req), ctx.perProviderTimeoutMs);
p.breaker.ok();
audit.write({ kind: "llm.call", provider: p.name, snapshot: p.modelSnapshot, attempts, taskId: ctx.taskId });
return { ...res, provider: p.name, modelSnapshot: p.modelSnapshot };
} catch (e: any) {
const status = e?.status ?? 0;
if (!RETRYABLE.has(status) && status !== 0) throw e; // bad request: do not hop providers
p.breaker.fail(); attempts.push(p.name + ":" + (status || "timeout"));
}
}
// Every provider is out or open. This is the branch that must be DESIGNED,
// not discovered: the workflow degrades to a mode that needs no model.
audit.write({ kind: "llm.all_providers_unavailable", attempts, taskId: ctx.taskId });
return degrade(req, ctx);
}- Pin snapshots per provider and log which one answered. A failover that changes the model is a silent change to a regulated system; the audit record must show it happened.
- Distinguish provider failure from request failure. A 400 or a schema violation on provider A will be a 400 on provider B; hopping just doubles the cost. Only retryable statuses and timeouts open the breaker.
- Keep the breaker per provider per model, not global. Monday's incident restored Fable and Mythos thirty-six minutes before Opus; a per-model breaker would have recovered sooner.
- Put the breaker state on a dashboard with an owner. An open breaker is an incident against an important business service and belongs in the same pipeline as any other.
- Run evals against the fallback provider too. A second provider that has never been evaluated on your golden set is not a fallback; it is an unknown.
Half Two: Correctness Is A Grounding Problem
The 57% is the harder half, and the reason is that it is not a model-quality number - it is an architecture number. Every failure category the study describes is a category a language model should never have been asked to own: miscalculations, overlooked tax changes, invented rules. A model computing a pension annual-allowance charge from memory is doing arithmetic on rules that change every fiscal year. The right architecture does not ask it to. The model's job is to understand the question and explain the answer; a deterministic, versioned calculator's job is to produce it.
from dataclasses import dataclass
from decimal import Decimal
@dataclass(frozen=True)
class RuleSet:
"""Versioned, tested, dated. The model may NOT compute tax; it may only call this."""
version: str # e.g. "uk-tax-2026-27@2026-09-01"
effective_from: str
def pension_annual_allowance_charge(self, contributions: Decimal, income: Decimal) -> dict:
allowance = self.tapered_allowance(income)
excess = max(Decimal(0), contributions - allowance)
charge = excess * self.marginal_rate(income)
return {"allowance": allowance, "excess": excess, "charge": charge,
"rule_version": self.version, "assumptions": ["no carry-forward considered"]}
def tapered_allowance(self, income: Decimal) -> Decimal: ...
def marginal_rate(self, income: Decimal) -> Decimal: ...
# The model gets ONE tool per calculation. Its inputs are validated, its output
# is the only figure that may appear in the answer, and the rule version rides
# along so the answer can be reproduced after the rules change.
TOOLS = [{
"name": "pension_annual_allowance_charge",
"description": "Compute the UK pension annual allowance charge. Use for any question about "
"pension contributions exceeding the allowance. Never estimate this yourself.",
"input_schema": {"type": "object", "properties": {
"contributions_gbp": {"type": "string", "pattern": "^[0-9]+(\\.[0-9]{1,2})?$"},
"income_gbp": {"type": "string", "pattern": "^[0-9]+(\\.[0-9]{1,2})?$"}},
"required": ["contributions_gbp", "income_gbp"]},
}]
def answer(question: str, rules: RuleSet) -> dict:
"""Model understands and narrates; calculator computes; validator checks that
every number in the narrative came from the calculator."""
first = model.complete(system=SYSTEM_GROUNDED, input=question, tools=TOOLS, temperature=0)
if not first.tool_calls:
# No calculation requested: either the question needs none, or the model
# tried to answer a numeric question from memory. Refuse the second.
if looks_numeric(question):
return {"status": "abstained", "reason": "NUMERIC_QUESTION_WITHOUT_CALCULATION"}
return {"status": "narrative_only", "text": first.text, "figures": []}
results = [rules.pension_annual_allowance_charge(Decimal(c.args["contributions_gbp"]),
Decimal(c.args["income_gbp"])) for c in first.tool_calls]
final = model.complete(system=SYSTEM_GROUNDED, input=question, tool_results=results, temperature=0)
# Every currency figure in the narrative must match a calculator output.
quoted = extract_gbp_figures(final.text)
produced = {str(r["charge"]), str(r["allowance"]), str(r["excess"])} | {str(r["charge"].quantize(Decimal("1"))) for r in results}
stray = [q for q in quoted if q not in produced]
if stray:
audit.write(kind="answer.ungrounded_figure", stray=stray, question=question)
return {"status": "abstained", "reason": "UNGROUNDED_FIGURE"}
return {"status": "grounded", "text": final.text, "figures": results,
"rule_version": rules.version, "disclosure": "Calculated under " + rules.version}“A model that is wrong 57% of the time on money questions is a model being asked the wrong question. Ask it to explain a figure a calculator produced, and the failure rate becomes the calculator's - which is a number you can test.”
- One tool per calculation, with validated inputs and a versioned rule set. The £17,500 error was a rule applied wrongly; a rule set with a version and an effective date is a rule set you can test against the fiscal year.
- Refuse numeric questions the model tried to answer from memory. The abstention branch above is the single control that most directly attacks the 57%. It will fire often at first; that is the mechanism working.
- Validate that every figure in the narrative came from the calculator. A model that receives the right numbers can still round, transpose or embellish them. The stray-figure check turns that into an abstention rather than an answer.
- Carry the rule version into the answer and the log. When HMRC changes an allowance, you need to know which answers were computed under the old rules - and the FT's 'overlooked tax changes' category is exactly this failure.
- Treat 'invented rules' as a retrieval problem. The student-loan answer was a hallucinated policy. Ground policy questions in a versioned, cited corpus with the same as-of discipline as market data, and require a citation for any rule the model states.
Putting Both Halves In One Pipeline
- 01Every model call goes through the router: per-provider, per-model breakers; pinned snapshots; retryable-only failover; a designed degraded mode; every attempt logged.
- 02Every numeric answer goes through a grounded path: tool per calculation, versioned rules, abstention on memory arithmetic, stray-figure validation, rule version in the output.
- 03Both emit to the same telemetry: breaker state, provider that answered, abstention rate, ungrounded-figure rate, calculator version. The OpenTelemetry GenAI conventions give you the attribute names.
- 04Both are gated by the same evals: the golden set runs against every provider in the router, and includes the 121-style money questions with the calculator's answer as ground truth.
- 05Both are drilled: open the primary breaker in staging quarterly; replay the golden set after every rule-set update.
The Bottom Line
Twenty-two September put the two reliability failures of financial AI side by side: an eighty-minute window of elevated errors across three frontier models, and a study in which 18 models got 57% of money questions wrong, 88% of the hard ones, with one answer worth a £17,500 tax charge and a majority of users prepared to act on it unchecked. The availability failure is a routing problem with a forty-year-old answer - per-provider circuit breakers, pinned snapshots, retryable-only failover and a degraded mode you designed and drilled before you needed it. The correctness failure is a grounding problem whose answer is to stop asking the model to do arithmetic: one versioned calculator per calculation, abstention when the model tries to answer from memory, validation that every figure in the narrative came from the calculator, and the rule version carried into every answer. Build both and Monday's news becomes an ordinary day. That is the production engineering we do for financial AI in London, and this week supplied the clearest case for it we have seen.
References & Further Reading
- Anthropic - Claude status: elevated errors for multiple models (incident of 22 September 2026). status.claude.com/incidents/7g1qpkyz5gxh
- Silicon UK - AI chatbots get financial queries wrong 'most of the time' (Saturn study, 22 September 2026). silicon.co.uk/fintech/ai-finance-research-631661
- AI Weekly - FT: 18 AI chatbots get financial answers wrong 57% of the time, fail 88% on complex queries. aiweekly.co/alerts/ft-18-ai-chatbots-get-financial-answers-wrong-57-of-the-time-fail-88-on-complex
- Startup Fortune - AI chatbots get financial questions wrong more than half the time, study finds. startupfortune.com/ai-chatbots-get-financial-questions-wrong-more-than-half-the-time-study-finds
- Yahoo Finance - Should you trust a chatbot with your money? A 10,000-answer test has a verdict. finance.yahoo.com/markets/crypto/articles/trust-chatbot-money-10-000-144200180.html
- DeployFlow - Is Claude down? 2026 Anthropic outage and expert failover tips. deployflow.co/blog/claude-anthropic-outage-protect-claude-infrastructure
- Michael Nygard - Release It! (origin of the circuit breaker pattern). pragprog.com/titles/mnee2/release-it-second-edition
- OpenTelemetry - Semantic conventions for generative AI. opentelemetry.io/docs/specs/semconv/gen-ai
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information