
FinToolBench: Evaluating LLM Agents on Real-World Financial Tool Use
FinToolBench finds intent mismatch above 50% for every LLM tested, so aggressive tool calling does not mean better answers on financial tasks.
#compliance
Regulatory compliance, policy enforcement, and audit trail research for financial AI systems

FinToolBench finds intent mismatch above 50% for every LLM tested, so aggressive tool calling does not mean better answers on financial tasks.

STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.

FinAuditing shows top LLMs hit just 13.86% on financial math verification of real SEC XBRL filings, capping what AI accounting tools can automate unaided.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

AuditCopilot cuts journal entry fraud false positives from 942 to 12, but ablation shows the LLM mainly synthesizes Isolation Forest scores.

Constitutional AI replaces hand-labelled harm data with AI feedback, so accounting agents can enforce ledger rules without a reviewer per transaction.