
WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
WildToolBench finds no LLM exceeds 15% session accuracy on 1,024 real-user tasks, with hidden intent and instruction transitions the sharpest failure modes.
#technology
Technology research and software engineering topics relevant to financial AI systems

WildToolBench finds no LLM exceeds 15% session accuracy on 1,024 real-user tasks, with hidden intent and instruction transitions the sharpest failure modes.

LLMs score up to 20 points worse when the answer sits mid-context, so finance RAG pipelines should place the best passages first or last.

OSWorld finds desktop AI agents succeed on 12.24% of real tasks versus 72.36% for humans, with 75% of failures from visuomotor grounding, not reasoning.

StructRAG routes each query to a table, graph, catalogue, algorithm, or chunk structure, beating GraphRAG by 28 points and running 22× faster.

Under equal thinking-token budgets, single-agent LLMs match or beat multi-agent systems on multi-hop reasoning—favor simpler finance agent designs.

Self-RAG trains an LLM to decide when to retrieve and self-grade results, hitting 55.8% on PopQA and 80.2 FactScore — beating ChatGPT on five benchmarks.

AgentBench scored GPT-4 4.01 versus 0.96 for the best open-source LLM. Those failures are exactly what breaks a Beancount write-back agent on a live ledger.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.