Skip to main content

Mike Thrift

Marketing Manager

FinQA: The Benchmark Measuring AI Numerical Reasoning on Financial Reports

FinQA (EMNLP 2021) built 8,281 QA pairs from S&P 500 earnings reports requiring multi-step arithmetic programs. Neural models scored 61% at release versus 91% for human experts; accuracy collapses to 22% on three-or-more-step programs. The failure modes — domain constants, cross-modality grounding, chain length — map directly to the challenges Beancount agents face today.

AgentBench: Evaluating LLMs as Agents — Lessons for Financial AI Reliability

AgentBench (Liu et al., ICLR 2024) benchmarks 27 LLMs across 8 interactive environments — GPT-4 scored 4.01 overall versus 0.96 for the best open-source model. The three dominant failure modes (task-limit exceeded in 67.9% of knowledge-graph failures, format errors in 53.3% of database failures, and invalid actions) map directly onto the risks of deploying a Beancount write-back agent on a real ledger.