
LLMs Score 2.3% on Beancount DSL Generation: The LLMFinLiteracy Benchmark
LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.
#double-entry
Double-entry bookkeeping principles and their application in AI-assisted accounting

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

AuditCopilot cuts journal entry fraud false positives from 942 to 12, but ablation shows the LLM mainly synthesizes Isolation Forest scores.