
BloombergGPT and the Limits of Domain-Specific LLMs in Finance
BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.
#finance
Financial research, analysis, and domain knowledge for accounting AI

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.

Huang et al. (ICLR 2024): intrinsic self-correction drops GPT-4 from 95.5% to 91.5% on GSM8K—ledger agents need external validators, not self-review.

CRITIC lifts open-domain QA by 7.7 F1 and cuts toxicity 79.2% by verifying LLM revisions against external tools, not itself.

Self-consistency majority-votes many sampled reasoning paths instead of one greedy decode, adding 17.9 points on GSM8K with no extra training.

PAL gains +38.1pp over chain-of-thought on GSM-hard by running Python for arithmetic—the right split for reliable Beancount ledger calculations.

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.

PHANTOM (NeurIPS 2025) measures LLM hallucination detection on real SEC filings. Qwen3-30B leads at F1=0.882, but 7B models guess near random.

FinBen finds GPT-4 at 0.63 exact match on FinQA and 0.54 on stock forecasting, barely above random, so accounting agents need validation, not raw LLM math.