3 Data points that are resetting expectations for AI, treasury, and payments
- Recent studies reveal how financial services are adapting around AI's limits, payments' expanding role, and rising customer expectations.
- The conversation is moving beyond adoption toward execution and measurable outcomes.
Three recent studies reveal meaningful shifts across financial services: where AI still falls short, why payments are becoming strategic, and how customer expectations are being reset.
Here’s the narrative behind those numbers.
26.2% of AI agents completed just one in four real-world work assignments
The number: OpenAI’s Codex, the top-performing system in UC Berkeley’s Agents’ Last Exam benchmark, completed just 26.2% of real-world professional work assignments. On the most complex, multi-step assignments, average success rates across all AI systems fell to 2.6%.
The narrative: The industry conversation around AI agents has largely centered on replacing humans. The assumption has been that if AI keeps getting smarter, it will eventually take over large parts of knowledge work.
The Berkeley benchmark tells a different story, though.
Today’s AI models aren’t necessarily struggling with intelligence; they’re struggling with execution. Many models can answer a question, write code, or analyze a document. But asking them to carry a complex workflow from start to finish – keeping track of multiple decisions, verifying information, correcting mistakes, and producing a finished deliverable – is a very different challenge. This places an even greater premium on keeping humans in the loop.
And financial services is built around exactly those kinds of workflows. Lending decisions, treasury operations, compliance reviews and payment approvals involve a chain of interconnected decisions where accuracy, governance, and accountability matter just as much as speed.
That helps explain a pattern we’ve been seeing across recent Tearsheet coverage.
…
