
Can LLM Agents Be CFOs? EnterpriseArena's 132-Month Simulation Reveals a Wide Gap
Only Qwen3.5-9B survives 80% of EnterpriseArena's 132-month CFO runs, while GPT-5.4 and DeepSeek-V3.1 hit 0%. Skipped ledger reconciliation causes the failures.
#forecasting
Financial forecasting and runway modelling with AI agents

Only Qwen3.5-9B survives 80% of EnterpriseArena's 132-month CFO runs, while GPT-5.4 and DeepSeek-V3.1 hit 0%. Skipped ledger reconciliation causes the failures.

InvestorBench: Qwen2.5-72B leads stock trading at 46.15% CR; finance-tuned Palmyra-Fin backfires on equities—size beats domain fine-tuning.

NeurIPS 2024 ablation: dropping the LLM from Time-LLM and CALF improves accuracy, with up to 1,383× faster training. Use purpose-built models for finance AI.

FinBen finds GPT-4 at 0.63 exact match on FinQA and 0.54 on stock forecasting, barely above random, so accounting agents need validation, not raw LLM math.