
HippoRAG: 89.1% Recall@5 on 2Wiki vs 68.2% ColBERT
HippoRAG's OpenIE+PageRank memory hits 89.1% Recall@5 on 2WikiMultiHopQA versus 68.2% ColBERTv2—path-aware retrieval for multi-year ledgers.
#beancount
Beancount ledger format, tooling, and ecosystem research

HippoRAG's OpenIE+PageRank memory hits 89.1% Recall@5 on 2WikiMultiHopQA versus 68.2% ColBERTv2—path-aware retrieval for multi-year ledgers.

AgentBench scored GPT-4 4.01 versus 0.96 for the best open-source LLM. Those failures are exactly what breaks a Beancount write-back agent on a live ledger.

BloombergGPT trained on 569B financial tokens, yet GPT-4 matched it with no finance pretraining. For accounting agents, tool-use beats model internals.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.

Gorilla's Retriever-Aware Training cuts LLM API hallucination rates from 78% to 11%, making tool calls reliable enough for finance agents that write entries.

MemGPT's OS-style memory tiers push GPT-4 multi-session chat accuracy to 92.5% versus 32.1% fixed-context—needed for multi-year ledger agents.

SWE-agent's Agent-Computer Interfaces lifted GPT-4 Turbo from raw shell to 12.47% on SWE-bench, a 10.7-point gain from interface design alone.

SWE-bench tests language models on 2,294 real GitHub issues; at publication Claude 2 resolved only 1.96%. Retrieval and patch-length limits shape coding agents.

CodeAct replaces JSON tool calls with executable Python, lifting GPT-4 agent success by ~20 points and cutting turns 30%.

Huang et al. (ICLR 2024): intrinsic self-correction drops GPT-4 from 95.5% to 91.5% on GSM8K—ledger agents need external validators, not self-review.

Reflexion hits 91% on HumanEval with GPT-4 without weight updates, but fails on WebShop. Verbal reinforcement needs a crisp evaluator signal.

PAL gains +38.1pp over chain-of-thought on GSM-hard by running Python for arithmetic—the right split for reliable Beancount ledger calculations.