
LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.
#analytics
Техники и метрики за анализ на данни за финансови ИИ системи

The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

LLMs score up to 20 points worse when the answer sits mid-context, so finance RAG pipelines should place the best passages first or last.

AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.

τ-bench finds top LLMs fall from pass@1 0.692 to pass@4 0.462 on retail tool-use tasks. Write-back agents on a ledger face the same consistency cliff.

ConvFinQA's best model scores 68.9% execution accuracy versus 89.4% for human experts—a 21-point gap that multi-turn ledger chat still faces.

FinanceBench tests 16 AI setups on 10,231 real SEC filing questions: shared-vector-store RAG answers only 19% right, so retrieval is not the bottleneck.

Self-consistency гласува по мнозинство сред много избрани пътища на разсъждение вместо едно алчно декодиране, добавяйки 17.9 точки на GSM8K без допълнително обучение.