
AD-LLM Benchmark: GPT-4o Hits 0.93+ AUROC Zero-Shot for Text Anomaly Detection
AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.
#anomaly-detection
Detecting irregularities and duplicate entries in plain-text ledger data

AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.

CausalTAD reorders table columns by causal dependency before LLM serialization, raising average AUC-ROC from 0.803 to 0.834 over AnoLLM.