
WorkArena++: The 93% Gap Between Human and AI Agent Performance on Compositional Enterprise Tasks
GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.
#machine-learning
Machine learning techniques for financial data analysis and automation

GPT-4o solves 2.1% of WorkArena++'s 682 compositional enterprise tasks; humans solve 93.9%, showing why agents stall on implicit-goal accounting work.

GAIA's 466 tasks show frontier AI agents at 74.55% versus 92% for humans, with Level 3 coordination still the hardest gap to close.

OSWorld finds desktop AI agents succeed on 12.24% of real tasks versus 72.36% for humans, with 75% of failures from visuomotor grounding, not reasoning.

GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.

WorkArena's 33 ServiceNow tasks show GPT-4o at 42.7% overall but 0% on list filters, a wall structured UI interaction must clear.

τ-bench finds top LLMs fall from pass@1 0.692 to pass@4 0.462 on retail tool-use tasks. Write-back agents on a ledger face the same consistency cliff.

Chain-of-Table hits 67.31% on WikiTQ versus 61.48% for text-only chain-of-thought, and leads by 10.25 points on tables over 4,000 tokens.

TableLlama beats GPT-4 on column type annotation (F1 94 vs 32) but trails by 33 points on WikiTQ compositional reasoning.

TAPAS answers table questions by selecting cells, never generating SQL. It fits small Beancount ledger queries but breaks down at scale.

MAC-SQL's three-agent design hits 59.59% execution accuracy on BIRD, with the Refiner adding +4.63 points — a template for generating Beancount ledger queries.

DIN-SQL lifts GPT-4 from 67.4% to 85.3% Spider execution accuracy via schema-link and self-correct stages—same decomposition fits Beancount BQL.

On BIRD, GPT-4 reaches 54.89% execution accuracy with domain hints and 34.88% without—a 20-point gap any Beancount NL→BQL interface must close.