
AnoLLM: Fine-Tuning LLMs for Tabular Anomaly Detection in Financial Data
AnoLLM beats classical baselines on mixed-type fraud data by scoring rows with LLM negative log-likelihood, but adds no edge on purely numerical tables.
#data-science
Data science methods applied to financial datasets and accounting workflows

AnoLLM beats classical baselines on mixed-type fraud data by scoring rows with LLM negative log-likelihood, but adds no edge on purely numerical tables.

TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.

GPT-4 reaches 74.1 mean AUROC on ODDS zero-shot, near the 75.5 ECOD baseline, but fails on high-variance data. Ledger auditing is not solved yet.

DocFinQA swaps FinQA's 700-word passages for full SEC filings, a 175× longer context that nearly halves GPT-4 accuracy on long documents.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

OSWorld finds desktop AI agents succeed on 12.24% of real tasks versus 72.36% for humans, with 75% of failures from visuomotor grounding, not reasoning.

Chain-of-Table hits 67.31% on WikiTQ versus 61.48% for text-only chain-of-thought, and leads by 10.25 points on tables over 4,000 tokens.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.

Microsoft's GraphRAG posts 72–83% sensemaking wins over vector RAG; a 2025 audit collapses those after correcting judge bias—caution for multi-doc ledger QA.

InvestorBench: Qwen2.5-72B leads stock trading at 46.15% CR; finance-tuned Palmyra-Fin backfires on equities—size beats domain fine-tuning.

M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.

Atlas hits 42.4% accuracy on Natural Questions with 64 examples, beating PaLM 540B by 3 points at 11B parameters via joint retriever-reader pre-training.