
LLM Anomaly Detection Survey (NAACL 2025): Strong Taxonomy, Absent Tabular Coverage
The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.
#fraud-detection
Detecting anomalous or fraudulent entries in financial data

The NAACL 2025 taxonomy holds, but tabular coverage is absent. Finance AI teams must adapt vision-model methods themselves.

AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.

CausalTAD reorders table columns by causal dependency before LLM serialization, raising average AUC-ROC from 0.803 to 0.834 over AnoLLM.

AnoLLM beats classical baselines on mixed-type fraud data by scoring rows with LLM negative log-likelihood, but adds no edge on purely numerical tables.

GPT-4 reaches 74.1 mean AUROC on ODDS zero-shot, near the 75.5 ECOD baseline, but fails on high-variance data. Ledger auditing is not solved yet.

AuditCopilot cuts journal entry fraud false positives from 942 to 12, but ablation shows the LLM mainly synthesizes Isolation Forest scores.

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.