
OpenHands: Open Platform for AI Software Agents and What It Means for Finance Automation
OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.
#developers
Developer resources, APIs, and integration documentation for finance tools

OpenHands' CodeAct agent scores 26% on SWE-Bench Lite, showing what AI agents reliably do today. Finance automation should start tightly scoped, not autonomous.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

RAG hits 0.875 accuracy on post-cutoff facts while fine-tuning plateaus at 0.504. For agents needing frequent ledger updates, retrieval beats fine-tuning.

Gorilla's Retriever-Aware Training cuts LLM API hallucination rates from 78% to 11%, making tool calls reliable enough for finance agents that write entries.

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.

SWE-bench tests language models on 2,294 real GitHub issues; at publication Claude 2 resolved only 1.96%. Retrieval and patch-length limits shape coding agents.

Toolformer teaches a 6.7B model to call APIs via perplexity filtering, beating GPT-3 175B on arithmetic. Its single-step design blocks chained ledger calls.