
Uncertainty-Aware Deferral for LLM Agents: When to Escalate from Small to Large Models
ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.
#llm
Large language model research with applications in financial tasks

ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.

Fin-RATE shows LLM accuracy collapses 18.60% on longitudinal tracking, with the retrieval pipeline, not the model, as the binding bottleneck.

FinDER's 5,703 real analyst queries show top RAG recalls only 25.95% of 10-K evidence. Normalize abbreviations first, before swapping embeddings.

LLMs score up to 20 points worse when the answer sits mid-context, so finance RAG pipelines should place the best passages first or last.

AD-LLM finds GPT-4o hits 0.93–0.99 AUROC zero-shot for text anomaly detection, but LLM model selection stays unreliable for financial audit AI.

CausalTAD reorders table columns by causal dependency before LLM serialization, raising average AUC-ROC from 0.803 to 0.834 over AnoLLM.

AnoLLM beats classical baselines on mixed-type fraud data by scoring rows with LLM negative log-likelihood, but adds no edge on purely numerical tables.

LLMFinLiteracy finds five ~7B models write correct Beancount transactions just 2.3% of the time, failing on accounting reasoning rather than syntax.

TableMaster hits 78.13% on WikiTQ with GPT-4o-mini, 13 points over Chain-of-Table, via table-of-focus plus adaptive reasoning for agents over Beancount ledgers.

GPT-4 reaches 74.1 mean AUROC on ODDS zero-shot, near the 75.5 ECOD baseline, but fails on high-variance data. Ledger auditing is not solved yet.

DocFinQA swaps FinQA's 700-word passages for full SEC filings, a 175× longer context that nearly halves GPT-4 accuracy on long documents.