
Fusion-in-Decoder: How Multi-Passage Retrieval Improves Generative QA
FiD encodes retrieved passages separately and fuses them in the decoder, beating RAG-Sequence by 4–11 points on NQ and TriviaQA, the synthesis ledger QA needs.
#data-science
Data science methods applied to financial datasets and accounting workflows

FiD encodes retrieved passages separately and fuses them in the decoder, beating RAG-Sequence by 4–11 points on NQ and TriviaQA, the synthesis ledger QA needs.

NeurIPS 2024 ablation: dropping the LLM from Time-LLM and CALF improves accuracy, with up to 1,383× faster training. Use purpose-built models for finance AI.

TAT-LLM fine-tunes LLaMA 2 7B to 64.60% EM on FinQA, edging GPT-4's 63.91%: an extract-reason-execute pipeline lets a small model do table arithmetic.

RAG hits 0.875 accuracy on post-cutoff facts while fine-tuning plateaus at 0.504. For agents needing frequent ledger updates, retrieval beats fine-tuning.

Lewis et al. hit 44.5 EM on Natural Questions with a frozen FAISS index; for Beancount ledgers that means reindex-or-miss when balances change daily.

MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.

ConvFinQA's best model scores 68.9% execution accuracy versus 89.4% for human experts—a 21-point gap that multi-turn ledger chat still faces.

TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

Self-consistency majority-votes many sampled reasoning paths instead of one greedy decode, adding 17.9 points on GSM8K with no extra training.

PAL gains +38.1pp over chain-of-thought on GSM-hard by running Python for arithmetic—the right split for reliable Beancount ledger calculations.

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.