
BloombergGPT and the Limits of Domain-Specific LLMs in Finance
BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.
#plain-text-accounting
Research grounded in plain-text accounting formats and workflows

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.

SWE-bench tests language models on 2,294 real GitHub issues; at publication Claude 2 resolved only 1.96%. Retrieval and patch-length limits shape coding agents.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.

Tree of Thoughts hits 74% on Game of 24 where GPT-4 chain-of-thought gets 4%, by searching and backtracking over steps, a pattern finance agents can borrow.

Reflexion hits 91% on HumanEval with GPT-4 without weight updates, but fails on WebShop. Verbal reinforcement needs a crisp evaluator signal.

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.

Toolformer teaches a 6.7B model to call APIs via perplexity filtering, beating GPT-3 175B on arithmetic. Its single-step design blocks chained ledger calls.