
LLMs Score 2.3% on Beancount DSL Generation: The LLMFinLiteracy Benchmark
LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.
#transaction-validation
Validating and verifying financial transactions using language model agents

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.

Multiagent debate gained 14.8 points on arithmetic, but equal-budget single agents match it. That limits debate as a safe check before a ledger commit.

CRITIC lifts open-domain QA by 7.7 F1 and cuts toxicity 79.2% by verifying LLM revisions against external tools, not itself.