
Can LLMs Reason Over Tabular Data? What Four Benchmarks Tell Us About Finance AI
Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.
#beancount
Beancount ledger format, tooling, and ecosystem research

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.

Constitutional AI replaces hand-labelled harm data with AI feedback, so accounting agents can enforce ledger rules without a reviewer per transaction.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.

Toolformer teaches a 6.7B model to call APIs via perplexity filtering, beating GPT-3 175B on arithmetic. Its single-step design blocks chained ledger calls.

FinBen finds GPT-4 at 0.63 exact match on FinQA and 0.54 on stock forecasting, barely above random, so accounting agents need validation, not raw LLM math.