Can You Trust A Chatbot With Your Money? A 10,000-Answer Test Says 57% Of The Time, No - Here Is Why, In Plain English
The most shared finance story of the week: Saturn tested 18 AI models from ChatGPT, Claude, Copilot, Gemini and Grok on 121 questions about debt, mortgages, pensions and tax, asking each five times - more than 10,000 answers. Fifty-seven per cent were wrong. On complex, multi-step questions 88% were wrong, some models 99%. Paid versions did better than free ones - 49% wrong against 63% - but not well. One pension answer would have landed a saver with a £17,500 HMRC charge; another model invented a student-loan rule. Meanwhile FCA research finds 26% of UK adults already trust these tools for financial advice. This is the plain-English explainer: why chatbots get money wrong, which questions are safe to ask, and what firms building financial AI do differently.
AlchmAI Editorial12 min read
57%
Of 10,000-plus answers from 18 AI models to 121 money questions were inaccurate (Saturn study, reported by the FT)
88%
Wrong on complex, multi-step questions - with some models failing 99% of the hardest ones
49% vs 63%
Average error rate for paid versus free models: paying helps, but not enough
26%
Of UK adults already trust general-purpose chatbots such as ChatGPT and Claude for financial advice (FCA research)
Every few months a study goes viral because it confirms something people suspected. This week's was about money. Saturn, a UK fintech, asked 18 popular AI models - across ChatGPT, Claude, Copilot, Gemini and Grok - 121 questions about the things people actually worry about: debt, mortgages, pensions and tax. Each question was asked five times, producing more than 10,000 answers, which were checked against the correct response. The Financial Times reported the result: 57% of answers were inaccurate. On the complex questions that need several steps of reasoning, the error rate was 88%, and some models got 99% of the hardest questions wrong.
The errors were not exotic. They were miscalculations, tax rules that had changed and been missed, and rules that do not exist. One answer about pension tax, from Claude Haiku 4.5, could have exposed a saver to a £17,500 charge from HMRC. Another model confidently described a student-loan rule - that repayments stop if you move abroad - that is simply false. And the people asking are acting on the answers: a PensionBee survey found 57% of US chatbot users would act without checking, and FCA research puts the share of UK adults who trust general-purpose chatbots for financial advice at 26%.
Why A Chatbot Gets Money Wrong
- 01It predicts words, it does not look up rules. A language model writes the most plausible answer based on what it learned in training. If the pension allowance or tax band changed after that - or was described inconsistently online - the plausible answer can be last year's answer, stated with complete confidence.
- 02Arithmetic across several steps compounds errors. A question like 'how much tax will I pay if I take this lump sum' is five or six calculations chained together. A small slip in step two becomes a large error by step six. That is why complex questions failed 88% of the time.
- 03It fills gaps instead of asking. Your answer depends on your income, your other pensions, your region, your circumstances. A chatbot usually assumes rather than asks, and the assumption may not be you.
- 04It does not know what it does not know. The invented student-loan rule is the classic example: when the model lacks a fact, it can generate one that sounds right. Nothing in the tone of the answer tells you which parts are real.
- 05Free and paid differ, but both fail. Paid models averaged 49% wrong against 63% for free ones. Better models reduce the problem; they do not remove it, because the problem is the approach rather than the model.
What Is Safe To Ask, And What Is Not
- Reasonably safe: explanations of concepts. What is a SIPP? How does a tracker mortgage differ from a fixed rate? What does 'annual allowance' mean? These are stable, widely documented ideas, and chatbots explain them well.
- Reasonably safe: preparing questions for a professional. Asking a chatbot what you should ask your adviser or your lender is a genuinely good use.
- Use with care: general rules. 'What is the personal allowance?' - check the answer against GOV.UK, because rules change and the model may be out of date.
- Not safe: your specific numbers. Tax due on a withdrawal, whether a contribution breaches an allowance, how much you can borrow. These are exactly the multi-step, rule-dependent questions where the error rate reached 88%.
- Not safe: anything irreversible. Before you move a pension, take a lump sum, or refinance, the answer needs to come from a regulated adviser or the provider's own calculator.
“The dangerous answer is not the one that is obviously wrong. It is the one that is wrong by £17,500 and sounds exactly like the right one.”
What Financial Firms Building AI Do Differently
The study tested general-purpose chatbots answering from memory. That is not how a well-built financial AI system works, and the difference is the whole story for firms deploying AI with customers. We build these systems for financial firms in London, and the design principle is simple: the model should never be the thing that does the sum.
- Calculations come from a calculator. The AI understands the question and explains the answer, but the figure itself comes from a tested, versioned calculation engine using this year's rules. The model is only allowed to repeat numbers the calculator produced.
- Rules come from a maintained source. Tax bands, allowances and product terms live in a versioned rulebook updated when HMRC or the product changes - not in whatever the model remembers.
- It says when it cannot answer. A good system abstains on questions it cannot ground, and hands over to a person rather than guessing. That single behaviour removes most of the damage the study found.
- It shows its working. The answer states which rules and which tax year it used, so both the customer and the firm can check it later.
- It is monitored as a Consumer Duty outcome. The FCA expects firms to monitor customer outcomes, not just processes. An AI assistant's accuracy - and whether it is worse for some customer groups - is exactly such an outcome.
The Bottom Line
A test of 18 AI models on 121 real money questions, more than 10,000 answers in all, found chatbots wrong 57% of the time and 88% of the time on complex questions, with errors like a £17,500 pension tax charge and an invented student-loan rule - while a quarter of UK adults already trust them for financial advice. The reason is structural: language models predict plausible text, and money questions need current rules, exact multi-step arithmetic and personal detail. For individuals the rule is to use chatbots to understand, never to decide. For firms, the lesson is that a financial AI system must never let the model do the sum: calculations from a tested engine, rules from a maintained source, abstention when it cannot ground an answer, and accuracy monitored as a customer outcome. That is how we build client-facing AI tools as a fintech AI agency in London - and it is the difference between the 57% and an answer you can stand behind.
References & Further Reading
- Silicon UK - AI chatbots get financial queries wrong 'most of the time' (Saturn study, 22 September 2026). silicon.co.uk/fintech/ai-finance-research-631661
- AI Weekly - FT: 18 AI chatbots get financial answers wrong 57% of the time, fail 88% on complex queries. aiweekly.co/alerts/ft-18-ai-chatbots-get-financial-answers-wrong-57-of-the-time-fail-88-on-complex
- Yahoo Finance - Should you trust a chatbot with your money? A 10,000-answer test has a verdict. finance.yahoo.com/markets/crypto/articles/trust-chatbot-money-10-000-144200180.html
- Startup Fortune - AI chatbots get financial questions wrong more than half the time, study finds. startupfortune.com/ai-chatbots-get-financial-questions-wrong-more-than-half-the-time-study-finds
- FCA - AI and the FCA: our approach. fca.org.uk/firms/innovation/ai-approach
- Global Regulation Tomorrow - FCA good and poor practice examples on monitoring consumer outcomes under the Consumer Duty. regulationtomorrow.com/2026/07/fca-good-and-poor-practice-examples-in-relation-to-monitoring-consumer-outcomes-under-the-consumer-duty
- GOV.UK - Tax on your private pension contributions (annual allowance). gov.uk/tax-on-your-private-pension/annual-allowance
AlchmAI Editorial
Research and analysis, London
The AlchmAI team writes about the markets, technology and regulation we work with every day. We build trading platforms, real-time charts and AI analysis tools for brokers, prop firms and fintech teams from our office in Mayfair, London. Every article lists its sources. Nothing we publish is investment advice.
This article is general information and commentary. It is not investment advice or a recommendation to buy or sell any investment. Important information