
TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.
#llm
Large language model research with applications in financial tasks

TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.

τ²-bench finds active users with their own tools cut conversational agent success by 18–25 points. Beancount agents sharing write access lose the same way.

GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

OSWorld finds desktop AI agents succeed on 12.24% of real tasks versus 72.36% for humans, with 75% of failures from visuomotor grounding, not reasoning.

GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.

τ-bench finds top LLMs fall from pass@1 0.692 to pass@4 0.462 on retail tool-use tasks. Write-back agents on a ledger face the same consistency cliff.

Chain-of-Table hits 67.31% on WikiTQ versus 61.48% for text-only chain-of-thought, and leads by 10.25 points on tables over 4,000 tokens.

TableLlama beats GPT-4 on column type annotation (F1 94 vs 32) but trails by 33 points on WikiTQ compositional reasoning.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.

MAC-SQL's three-agent design hits 59.59% execution accuracy on BIRD, with the Refiner adding +4.63 points — a template for generating Beancount ledger queries.