
FinQA: The Benchmark Measuring AI Numerical Reasoning on Financial Reports
FinQA found neural models scored 61% on financial-report math versus 91% for human experts, collapsing to 22% on three-or-more-step programs.
#ai
Artificial intelligence research and applications in finance and accounting

FinQA found neural models scored 61% on financial-report math versus 91% for human experts, collapsing to 22% on three-or-more-step programs.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

DSPy's compiler lifted Llama2-13b from 9.4% to 46.9% on GSM8K, pointing finance AI pipelines toward maintainable declarative LLM calls.

LATS unifies ReAct, Tree of Thoughts, and Reflexion in one MCTS framework, hitting 92.7% pass@1 on HumanEval with GPT-4.

Self-RAG trains an LLM to decide when to retrieve and self-grade results, hitting 55.8% on PopQA and 80.2 FactScore — beating ChatGPT on five benchmarks.

Voyager's persistent code skill library discovers 3.3× more Minecraft items than prior SOTA without fine-tuning—the reuse pattern ledger agents need.

HippoRAG's OpenIE+PageRank memory hits 89.1% Recall@5 on 2WikiMultiHopQA versus 68.2% ColBERTv2—path-aware retrieval for multi-year ledgers.

AgentBench scored GPT-4 4.01 versus 0.96 for the best open-source LLM. Those failures are exactly what breaks a Beancount write-back agent on a live ledger.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

Gorilla's Retriever-Aware Training cuts LLM API hallucination rates from 78% to 11%, making tool calls reliable enough for finance agents that write entries.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.