
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under MCP
FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.
#fintech
Изследвания на финансови технологии, платформи и инфраструктура за модерни счетоводни системи

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

FinTrace shows frontier LLMs pick the right financial tools (F1 ~0.9) but score just 3.23/5 on using the results, the step that breaks write-back agents.

FinToolBench съчетава 760 реални финансови API инструмента с 295 изпълними заявки, за да оцени LLM агенти върху реални финансови задачи — разкривайки, че консервативният процент на извикване от 22,7% на GPT-4o води до по-високо качество на отговорите (CSS 0,670) от агресивния TIR от 87,1% на Qwen3-8B, докато несъответствието на намеренията надхвърля 50% при всеки тестван модел.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.