
SWE-agent: How Interface Design Unlocks Automated Software Engineering
SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.
#ai
Artificial intelligence research and applications in finance and accounting

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.

SWE-bench tests language models on 2,294 real GitHub issues; at publication Claude 2 resolved only 1.96%. Retrieval and patch-length limits shape coding agents.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.

Huang et al. (ICLR 2024): intrinsic self-correction drops GPT-4 from 95.5% to 91.5% on GSM8K—ledger agents need external validators, not self-review.

Tree of Thoughts hits 74% on Game of 24 where GPT-4 chain-of-thought gets 4%, by searching and backtracking over steps, a pattern finance agents can borrow.

CRITIC lifts open-domain QA by 7.7 F1 and cuts toxicity 79.2% by verifying LLM revisions against external tools, not itself.

Reflexion hits 91% on HumanEval with GPT-4 without weight updates, but fails on WebShop. Verbal reinforcement needs a crisp evaluator signal.

Self-consistency majority-votes many sampled reasoning paths instead of one greedy decode, adding 17.9 points on GSM8K with no extra training.

PAL gains +38.1pp over chain-of-thought on GSM-hard by running Python for arithmetic—the right split for reliable Beancount ledger calculations.

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.

Constitutional AI replaces hand-labelled harm data with AI feedback, so accounting agents can enforce ledger rules without a reviewer per transaction.

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.