
Reflexion: Language Agents That Learn from Mistakes Without Retraining
Reflexion hits 91% on HumanEval with GPT-4 without weight updates, but fails on WebShop. Verbal reinforcement needs a crisp evaluator signal.
#automation
Automation techniques and tools for financial data processing workflows

Reflexion hits 91% on HumanEval with GPT-4 without weight updates, but fails on WebShop. Verbal reinforcement needs a crisp evaluator signal.

Self-consistency majority-votes many sampled reasoning paths instead of one greedy decode, adding 17.9 points on GSM8K with no extra training.

PAL gains +38.1pp over chain-of-thought on GSM-hard by running Python for arithmetic—the right split for reliable Beancount ledger calculations.

Four benchmarks show GPT-4 at 42% on real-world table QA versus 86% for humans, and 19.6% on complex aggregations, so finance agents need clean table input.

Constitutional AI replaces hand-labelled harm data with AI feedback, so accounting agents can enforce ledger rules without a reviewer per transaction.

Chain-of-Thought prompting raises precision but can cut recall on rare financial events, so fraud agents may miss anomalies they should flag.

FinMaster Benchmark: top LLMs hit 96% on financial literacy but only 3% on statement generation. Error propagation costs 21 points on consulting tasks.

ReAct interleaves reasoning with tool actions, beating pure CoT by 34 points on fact verification. Its failure modes shape agents writing to Beancount ledgers.

Toolformer teaches a 6.7B model to call APIs via perplexity filtering, beating GPT-3 175B on arithmetic. Its single-step design blocks chained ledger calls.