
FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.
#data-science
Data science methods applied to financial datasets and accounting workflows

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

WildToolBench finds no LLM exceeds 15% session accuracy on 1,024 real-user tasks, with hidden intent and instruction transitions the sharpest failure modes.

Verbalized GPT-4 confidence hits only ~62.7% AUROC, barely above chance. Uncertainty-aware finance agents need better calibration than that.

FinToolBench finds intent mismatch above 50% for every LLM tested, so aggressive tool calling does not mean better answers on financial tasks.

OmniEval scores the best RAG systems at 36% numerical accuracy across 5 financial task types, so ledger agents need validation before writing entries.

The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.

Subtracting positional bias from LLM attention weights recovers up to 15 points of RAG accuracy when evidence sits mid-context, aiding finance agent pipelines.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

FinDER's 5,703 real analyst queries show top RAG recalls only 25.95% of 10-K evidence. Normalize abbreviations first, before swapping embeddings.

LLMs score up to 20 points worse when the answer sits mid-context, so finance RAG pipelines should place the best passages first or last.

AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.

CausalTAD reorders table columns by causal dependency before LLM serialization, raising average AUC-ROC from 0.803 to 0.834 over AnoLLM.