
Can your agent balance these books?
One agent run reconciled a three-row ledger to 2958.50 USD checking and 1958.50 USD September profit, recovering from one failure it diagnosed itself.
#plain-text-accounting
Research grounded in plain-text accounting formats and workflows

One agent run reconciled a three-row ledger to 2958.50 USD checking and 1958.50 USD September profit, recovering from one failure it diagnosed itself.

ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.

τ²-bench finds active users with their own tools cut conversational agent success by 18–25 points. Beancount agents sharing write access lose the same way.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.

τ-bench finds top LLMs fall from pass@1 0.692 to pass@4 0.462 on retail tool-use tasks. Write-back agents on a ledger face the same consistency cliff.

Chain-of-Table hits 67.31% on WikiTQ versus 61.48% for text-only chain-of-thought, and leads by 10.25 points on tables over 4,000 tokens.

TableLlama beats GPT-4 on column type annotation (F1 94 vs 32) but trails by 33 points on WikiTQ compositional reasoning.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.