
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.
#fintech
Financial technology research, platforms, and infrastructure for modern accounting systems

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

FinTrace shows frontier LLMs pick the right financial tools (F1 ~0.9) but score just 3.23/5 on using the results, the step that breaks write-back agents.

FinToolBench finds intent mismatch above 50% for every LLM tested, so aggressive tool calling does not mean better answers on financial tasks.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.