
Може ли вашият агент да изравни тези сметки?
Едно изпълнение на агент осчетоводи книга с три реда до 2958.50 USD разплащателна сметка и 1958.50 USD септемврийска печалба, след като сам диагностицира провал.
#reconciliation
Автоматизирано съгласуване на счетоводни книги с агенти на езикови модели

Едно изпълнение на агент осчетоводи книга с три реда до 2958.50 USD разплащателна сметка и 1958.50 USD септемврийска печалба, след като сам диагностицира провал.

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

Only Qwen3.5-9B survives 80% of EnterpriseArena's 132-month CFO runs, while GPT-5.4 and DeepSeek-V3.1 hit 0%. Skipped ledger reconciliation causes the failures.

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

Subtracting positional bias from LLM attention weights recovers up to 15 points of RAG accuracy when evidence sits mid-context, aiding finance agent pipelines.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

Voyager's persistent code skill library discovers 3.3× more Minecraft items than prior SOTA without fine-tuning—the reuse pattern ledger agents need.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.

CRITIC lifts open-domain QA by 7.7 F1 and cuts toxicity 79.2% by verifying LLM revisions against external tools, not itself.

ReAct редува разсъждение с действия с инструменти, побеждавайки чистия CoT с 34 точки при проверка на факти. Провалите му оформят агентите, пишещи в регистри на Beancount.