
TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.
#automation
Automation techniques and tools for financial data processing workflows

TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.

τ²-bench finds active users with their own tools cut conversational agent success by 18–25 points. Beancount agents sharing write access lose the same way.

GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

OSWorld finds desktop AI agents succeed on 12.24% of real tasks versus 72.36% for humans, with 75% of failures from visuomotor grounding, not reasoning.

GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.

τ-bench finds top LLMs fall from pass@1 0.692 to pass@4 0.462 on retail tool-use tasks. Write-back agents on a ledger face the same consistency cliff.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.

MAC-SQL's three-agent design hits 59.59% execution accuracy on BIRD, with the Refiner adding +4.63 points — a template for generating Beancount ledger queries.

STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.

Under equal thinking-token budgets, single-agent LLMs match or beat multi-agent systems on multi-hop reasoning—favor simpler finance agent designs.