
TheAgentCompany: Benchmarking LLM Agents on Real-World Enterprise Tasks
TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.
#productivity
Productivity improvements and automation research for knowledge workers

TheAgentCompany benchmarks 175 enterprise tasks in a simulated intranet, and the best model, Gemini-2.5-Pro, finishes only 30% at $4 each.

GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.