2026年6月23日阅读需 1 分钟LLM 在 Beancount DSL 生成中得分仅为 2.3%:LLMFinLiteracy 基准测试LLMFinLiteracy发现,五个约7B模型正确编写Beancount交易的比例仅为2.3%,失败在于会计推理而非语法。Mike Thriftllmbeancountplain-text-accounting
2026年5月22日阅读需 2 分钟AuditCopilot:大语言模型在复式记账欺诈检测中的应用AuditCopilot 将日记账分录欺诈误报从 942 降至 12,但消融实验显示 LLM 主要在综合 Isolation Forest 的分数。Mike Thriftfraud-detectionllmdouble-entry