Backbase AI Outperforms GPT-4.1 in Banking by Learning to Say 'I Don't Know'
By Lauren Towner · 30 July 2026

Quick Summary
Backbase has developed a 12-billion-parameter banking AI model that outperforms GPT-4.1 by admitting uncertainty. By training the system to say "I don't know" when evidence is missing, Backbase achieved a 7.1% increase in query resolution at 50x lower operational costs compared to frontier models.
How Does Backbase Solve AI Hallucinations in Banking?
Backbase solves the critical issue of AI hallucinations by training its 12B-parameter model, FinRAG-12B, to recognize the limits of its own knowledge. Unlike standard models that are rewarded for confident guessing, this system was trained on a dataset where 22% of examples had no correct answer, forcing the AI to prioritize grounded accuracy over sycophancy.
- The model reached a 12% refusal rate, significantly more honest than the 4.3% rate of untuned base models.
- It outperformed GPT-4.1 on citation grounding by 2.3 points, ensuring answers are tied to source documents.
- Training costs were remarkably low at just $1,800, proving specialized models can beat general-purpose giants.
What Results Has This AI Delivered for Financial Institutions?
In a seven-month live deployment at a large US financial institution, the model demonstrated that honesty drives resolution. By refusing to answer nearly three times as often as its base model, it paradoxically solved more problems because users trusted the validated responses it did provide.
- Query resolution rose by 7.1 percentage points across 3,297 sampled customer queries.
- Operational costs dropped to $0.001 per query, making it 20-50x cheaper than using GPT-4.1.
- Processing speeds were 3-5x faster than frontier models, running efficiently on a single GPU.
FF NEWS TAKE:
This announcement moves the needle by debunking the "bigger is better" myth in fintech AI. While the industry has been obsessed with massive parameter counts, Backbase proves that domain-specific calibration and calibrated refusal are the real keys to ROI. Achieving a 7.1% resolution lift while slashing costs by 50x is a massive win that finally addresses why 95% of AI pilots fail to impact the bottom line.
Companies in this story: MIT, IBM, Kasisto, McKinsey, Google, Backbase
People in this story: Jouk Pleiter, Camille Oster, Denys Katerenchuk