
WildToolBench: Why No LLM Exceeds 15% Session Accuracy in Real-World Tool Use
WildToolBench (ICLR 2026) evaluates 57 LLMs on 1,024 tasks drawn from real user behavior — no model exceeds 15% session accuracy, with compositional orchestration, hidden intent, and instruction transitions as the three sharpest failure modes.






