
OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.
#beancount
Beancount ledger format, tooling, and ecosystem research

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.

FinDER's 5,703 real analyst queries show top RAG recalls only 25.95% of 10-K evidence. Normalize abbreviations first, before swapping embeddings.

CausalTAD reorders table columns by causal dependency before LLM serialization, raising average AUC-ROC from 0.803 to 0.834 over AnoLLM.

AnoLLM beats classical baselines on mixed-type fraud data by scoring rows with LLM negative log-likelihood, but adds no edge on purely numerical tables.

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.

GPT-4 reaches 74.1 mean AUROC on ODDS zero-shot, near the 75.5 ECOD baseline, but fails on high-variance data. Ledger auditing is not solved yet.

DocFinQA swaps FinQA's 700-word passages for full SEC filings, a 175× longer context that nearly halves GPT-4 accuracy on long documents.

τ²-bench finds active users with their own tools cut conversational agent success by 18–25 points. Beancount agents sharing write access lose the same way.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.