
LLM Confidence and Calibration: A Survey of What the Research Actually Shows
Verbalized GPT-4 confidence hits only ~62.7% AUROC, barely above chance. Uncertainty-aware finance agents need better calibration than that.
#trust
Reliability, calibration, and hallucination in financial AI systems

Verbalized GPT-4 confidence hits only ~62.7% AUROC, barely above chance. Uncertainty-aware finance agents need better calibration than that.

ReDAct defers from a small model to a large one only when perplexity signals uncertainty, cutting cost 64% at matching accuracy for agent workflows.

STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.

Huang et al. (ICLR 2024): intrinsic self-correction drops GPT-4 from 95.5% to 91.5% on GSM8K—ledger agents need external validators, not self-review.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.