
LLMs Are Not Useful for Time Series Forecasting: What NeurIPS 2024 Means for Finance AI
NeurIPS 2024 ablation: dropping the LLM from Time-LLM and CALF improves accuracy, with up to 1,383× faster training. Use purpose-built models for finance AI.
#machine-learning
Machine learning techniques for financial data analysis and automation

NeurIPS 2024 ablation: dropping the LLM from Time-LLM and CALF improves accuracy, with up to 1,383× faster training. Use purpose-built models for finance AI.

AuditCopilot cuts journal entry fraud false positives from 942 to 12, but ablation shows the LLM mainly synthesizes Isolation Forest scores.

TAT-LLM fine-tunes LLaMA 2 7B to 64.60% EM on FinQA, edging GPT-4's 63.91%: an extract-reason-execute pipeline lets a small model do table arithmetic.

RAG hits 0.875 accuracy on post-cutoff facts while fine-tuning plateaus at 0.504. For agents needing frequent ledger updates, retrieval beats fine-tuning.

IRCoT adds +11.3 recall and +7.1 F1 on HotpotQA over one-step RAG by querying retrieval at every reasoning step. A 3B model can then beat GPT-3 175B.

FLARE hits 51.0 EM on 2WikiMultihopQA versus 39.4 for single-retrieval RAG, but calibration failures in chat models limit it for finance agents.

Lewis et al. hit 44.5 EM on Natural Questions with a frozen FAISS index; for Beancount ledgers that means reindex-or-miss when balances change daily.

MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.

ConvFinQA's best model scores 68.9% execution accuracy versus 89.4% for human experts—a 21-point gap that multi-turn ledger chat still faces.

TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.

FinQA found neural models scored 61% on financial-report math versus 91% for human experts, collapsing to 22% on three-or-more-step programs.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.