
Може ли вашият агент да изравни тези сметки?
Едно изпълнение на агент осчетоводи книга с три реда до 2958.50 USD разплащателна сметка и 1958.50 USD септемврийска печалба, след като сам диагностицира провал.
#beancount
Изследвания на формата Beancount, инструменти и екосистема

Едно изпълнение на агент осчетоводи книга с три реда до 2958.50 USD разплащателна сметка и 1958.50 USD септемврийска печалба, след като сам диагностицира провал.

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

Only Qwen3.5-9B survives 80% of EnterpriseArena's 132-month CFO runs, while GPT-5.4 and DeepSeek-V3.1 hit 0%. Skipped ledger reconciliation causes the failures.

WildToolBench finds no LLM exceeds 15% session accuracy on 1,024 real-user tasks, with hidden intent and instruction transitions the sharpest failure modes.

JSONSchemaBench finds coverage collapses from 86% on simple schemas to 3% on complex ones, so LLM structured output can silently emit non-compliant JSON.

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

FinTrace shows frontier LLMs pick the right financial tools (F1 ~0.9) but score just 3.23/5 on using the results, the step that breaks write-back agents.

FinToolBench съчетава 760 реални финансови API инструмента с 295 изпълними заявки, за да оцени LLM агенти върху реални финансови задачи — разкривайки, че консервативният процент на извикване от 22,7% на GPT-4o води до по-високо качество на отговорите (CSS 0,670) от агресивния TIR от 87,1% на Qwen3-8B, докато несъответствието на намеренията надхвърля 50% при всеки тестван модел.

OmniEval scores the best RAG systems at 36% numerical accuracy across 5 financial task types, so ledger agents need validation before writing entries.

The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.

Subtracting positional bias from LLM attention weights recovers up to 15 points of RAG accuracy when evidence sits mid-context, aiding finance agent pipelines.

ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.