AI Financial Advice Fails: Saturn Report Reveals 57% Error Rate in ChatGPT and Claude
By Lauren Towner · 14 September 2026

New research from fintech firm Saturn reveals that mainstream AI models provide incorrect financial advice 57% of the time, rising to 88% for complex inquiries. With over a quarter of UK consumers already using these tools for money management, the findings highlight a critical regulatory gap and the potential for significant financial loss.
What was announced
Saturn released a report titled Artificial Authority: Should you trust AI to deliver financial advice?, which details a comprehensive analysis of AI-driven financial guidance. The study rigorously tested 18 popular AI models—including ChatGPT, Claude, CoPilot, Grok, and Gemini—against 121 distinct financial questions. To ensure consistency, each question was repeated five times, resulting in a dataset of over 10,000 individual responses across various product types and scenarios.
The results indicate a stark divide between free and paid services. Free AI models failed to provide accurate answers in 63% of cases, while paid-for models saw a slightly lower error rate of 49%. When faced with complex financial scenarios, the failure rate for free models spiked to 93%, with some models failing 99% of the time. The report identified several categories of failure, including calculation errors, the omission of vital risk warnings, ignorance of upcoming tax changes, and the "hallucination" of non-existent financial rules.
Specific examples of harm include a Gemini model incorrectly advising a borrower that a mortgage payment holiday would not impact their credit score. In another instance, Claude Haiku 4.5 provided flawed pension tax advice that could have triggered a £17,500 charge from HMRC. The research also highlighted risks for vulnerable consumers; some models suggested paying off high-interest debt before priority bills like rent or council tax, a move that could lead to eviction, bailiff visits, or legal action. When asked about student loans, Claude invented a rule suggesting graduates could stop repayments if moving abroad, which could lead to higher monthly repayment penalties.
"The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm. Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money."
Amal Jolly, Saturn chief executive.
The companies involved
Saturn is a financial technology firm focused on the intersection of digital services and consumer protection. Operating from its web domain at saturnos.com, the company has positioned itself as a critical observer of how emerging technologies interact with traditional financial frameworks. This latest report follows a period of heightened scrutiny regarding the "advice gap" in the United Kingdom, where many consumers turn to free digital tools in the absence of affordable human financial advisers.
The Financial Conduct Authority (FCA) is the primary regulator for the UK’s financial services industry, maintaining a mandate to protect consumers and enhance market integrity. The FCA’s recent Mills Review underscored the scale of the challenge, noting that 26% of consumers now rely on general-purpose AI tools for financial guidance. OpenAI, the creator of ChatGPT, is a dominant player in this space, having transitioned from a research-focused entity to a major provider of enterprise-grade AI models that are increasingly being integrated into the financial services stack. The FCA is currently considering whether to regulate the financial advice that these AI models provide to ensure consumers have access to protections and compensation.
What FF News has reported before
FF News has closely followed the evolution of generative AI within the banking and investment sectors. We recently covered how OpenAI Debuts ChatGPT for Financial Services with GPT-6 Astra Integration, a move that signaled the tech giant's intent to provide more specialized tools for the industry. While other firms are focusing on the underlying technology, as seen in the launch of PayPal, MoonPay, and M0 Launch PYUSDx to Power Custom Stablecoin Infrastructure, the Saturn report suggests that the consumer-facing side of AI remains fraught with accuracy issues that could undermine these technological advancements.
What this means
The Saturn report is a sobering reality check for the fintech industry’s "AI-first" narrative. While the sector has rushed to integrate large language models to lower operational costs and bridge the advice gap, these findings suggest that the technology is currently unfit for purpose in high-stakes financial environments. The high error rates in tax and debt management are not merely technical glitches; they represent a systemic risk to consumer solvency. This puts immediate pressure on the FCA to move beyond observation and toward establishing clear liability frameworks. For the market, the era of consequence-free AI experimentation in retail finance is likely ending as the demand for verified, regulated AI becomes paramount.
Companies in this story: Saturn, Financial Conduct Authority, OpenAI
People in this story: Tom Ellis, Amal Jolly