
DSPy: Replacing Brittle Prompt Engineering with Compiled LLM Pipelines
DSPy's compiler lifted Llama2-13b from 9.4% to 46.9% on GSM8K, pointing finance AI pipelines toward maintainable declarative LLM calls.
#machine-learning
Machine learning techniques for financial data analysis and automation

DSPy's compiler lifted Llama2-13b from 9.4% to 46.9% on GSM8K, pointing finance AI pipelines toward maintainable declarative LLM calls.

LATS unifies ReAct, Tree of Thoughts, and Reflexion in one MCTS framework, hitting 92.7% pass@1 on HumanEval with GPT-4.

Self-RAG trains an LLM to decide when to retrieve and self-grade results, hitting 55.8% on PopQA and 80.2 FactScore — beating ChatGPT on five benchmarks.

Voyager's persistent code skill library discovers 3.3× more Minecraft items than prior SOTA without fine-tuning—the reuse pattern ledger agents need.

HippoRAG's OpenIE+PageRank memory hits 89.1% Recall@5 on 2WikiMultiHopQA versus 68.2% ColBERTv2—path-aware retrieval for multi-year ledgers.

AgentBench scored GPT-4 4.01 versus 0.96 for the best open-source LLM. Those failures are exactly what breaks a Beancount write-back agent on a live ledger.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

Gorilla's Retriever-Aware Training cuts LLM API hallucination rates from 78% to 11%, making tool calls reliable enough for finance agents that write entries.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.

SWE-bench tests language models on 2,294 real GitHub issues; at publication Claude 2 resolved only 1.96%. Retrieval and patch-length limits shape coding agents.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.