
LLM Confidence and Calibration: A Survey of What the Research Actually Shows
Verbalized GPT-4 confidence hits only ~62.7% AUROC, barely above chance. Uncertainty-aware finance agents need better calibration than that.
#hallucination-detection
Methods and techniques for detecting factual errors and hallucinations in LLM outputs