
FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
FinToolBench pairs 760 live financial API tools with 295 executable queries to benchmark LLM agents on real financial tasks — revealing that GPT-4o's conservative 22.7% invocation rate yields higher answer quality (CSS 0.670) than Qwen3-8B's aggressive 87.1% TIR, while intent mismatch exceeds 50% across every model tested.





