
OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.
#open-source
Open-source tools, frameworks, and research artifacts for financial AI

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.

GPT-4 finishes just 14.41% of WebArena's 812 web tasks versus 78.24% for humans; agents mostly fail by falsely declaring a task impossible.

TableLlama beats GPT-4 on column type annotation (F1 94 vs 32) but trails by 33 points on WikiTQ compositional reasoning.

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.