
Can your agent balance these books?
One agent run reconciled a three-row ledger to 2958.50 USD checking and 1958.50 USD September profit, recovering from one failure it diagnosed itself.
#reconciliation
Automated ledger reconciliation using language model agents

One agent run reconciled a three-row ledger to 2958.50 USD checking and 1958.50 USD September profit, recovering from one failure it diagnosed itself.

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

Only Qwen3.5-9B survives 80% of EnterpriseArena's 132-month CFO runs, while GPT-5.4 and DeepSeek-V3.1 hit 0%. Skipped ledger reconciliation causes the failures.

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

Subtracting positional bias from LLM attention weights recovers up to 15 points of RAG accuracy when evidence sits mid-context, aiding finance agent pipelines.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

Voyager's persistent code skill library discovers 3.3× more Minecraft items than prior SOTA without fine-tuning—the reuse pattern ledger agents need.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.

CRITIC lifts open-domain QA by 7.7 F1 and cuts toxicity 79.2% by verifying LLM revisions against external tools, not itself.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.