
FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.
#machine-learning
Техники за машинно обучение за анализ и автоматизация на финансови данни

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

WildToolBench finds no LLM exceeds 15% session accuracy on 1,024 real-user tasks, with hidden intent and instruction transitions the sharpest failure modes.

Verbalized GPT-4 confidence hits only ~62.7% AUROC, barely above chance. Uncertainty-aware finance agents need better calibration than that.

JSONSchemaBench finds coverage collapses from 86% on simple schemas to 3% on complex ones, so LLM structured output can silently emit non-compliant JSON.

FinMCP-Bench scores the best of six LLMs at just 3.08% exact match on 613 real MCP financial tasks, a 20× collapse from single-tool to multi-turn use.

FinTrace shows frontier LLMs pick the right financial tools (F1 ~0.9) but score just 3.23/5 on using the results, the step that breaks write-back agents.

FinToolBench съчетава 760 реални финансови API инструмента с 295 изпълними заявки, за да оцени LLM агенти върху реални финансови задачи — разкривайки, че консервативният процент на извикване от 22,7% на GPT-4o води до по-високо качество на отговорите (CSS 0,670) от агресивния TIR от 87,1% на Qwen3-8B, докато несъответствието на намеренията надхвърля 50% при всеки тестван модел.

OmniEval scores the best RAG systems at 36% numerical accuracy across 5 financial task types, so ledger agents need validation before writing entries.

The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.

Subtracting positional bias from LLM attention weights recovers up to 15 points of RAG accuracy when evidence sits mid-context, aiding finance agent pipelines.

ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.