
TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.
#enterprise-software
Enterprise software automation, web agents, and knowledge work task research

TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.

GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.