
WebArena: The 812-Task Benchmark That Measures What Web Agents Actually Can and Cannot Do
GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.
#web-interface
Web-based interfaces and browser agents for financial AI systems