
LLMs Score 2.3% on Beancount DSL Generation: The LLMFinLiteracy Benchmark
LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.
#financial-literacy
Изследвания на представянето на финансови знания и компетентността на големи езикови модели

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

FinMaster Benchmark: водещите LLM достигат 96% по финансова грамотност, но само 3% при генериране на отчети. Разпространението на грешки струва 21 точки при консултации.