
FinRAGBench-V: Multimodal RAG with Visual Citations in the Financial Domain
FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.
#financial-reporting
Generating and auditing financial reports with language models

FinRAGBench-V finds top models reach only 20–61% block-level citation recall on financial pages. Multimodal retrieval beats text-only by nearly 50 points.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

FinDER's 5,703 real analyst queries show top RAG recalls only 25.95% of 10-K evidence. Normalize abbreviations first, before swapping embeddings.

DocFinQA swaps FinQA's 700-word passages for full SEC filings, a 175× longer context that nearly halves GPT-4 accuracy on long documents.

FinAuditing shows top LLMs hit just 13.86% on financial math verification of real SEC XBRL filings, capping what AI accounting tools can automate unaided.

TAT-LLM fine-tunes LLaMA 2 7B to 64.60% EM on FinQA, edging GPT-4's 63.91%: an extract-reason-execute pipeline lets a small model do table arithmetic.

MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.

ConvFinQA's best model scores 68.9% execution accuracy versus 89.4% for human experts—a 21-point gap that multi-turn ledger chat still faces.

TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.

FinQA found neural models scored 61% on financial-report math versus 91% for human experts, collapsing to 22% on three-or-more-step programs.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.