
GraphRAG: 72–83% wins vanish after LLM-judge bias fix
Microsoft's GraphRAG posts 72–83% sensemaking wins over vector RAG; a 2025 audit collapses those after correcting judge bias—caution for multi-doc ledger QA.
#machine-learning
Machine learning techniques for financial data analysis and automation

Microsoft's GraphRAG posts 72–83% sensemaking wins over vector RAG; a 2025 audit collapses those after correcting judge bias—caution for multi-doc ledger QA.

FinAuditing shows top LLMs hit just 13.86% on financial math verification of real SEC XBRL filings, capping what AI accounting tools can automate unaided.

InvestorBench: Qwen2.5-72B leads stock trading at 46.15% CR; finance-tuned Palmyra-Fin backfires on equities—size beats domain fine-tuning.

StructRAG routes each query to a table, graph, catalogue, algorithm, or chunk structure, beating GraphRAG by 28 points and running 22× faster.

Under equal thinking-token budgets, single-agent LLMs match or beat multi-agent systems on multi-hop reasoning—favor simpler finance agent designs.

M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

Atlas hits 42.4% accuracy on Natural Questions with 64 examples, beating PaLM 540B by 3 points at 11B parameters via joint retriever-reader pre-training.

FiD encodes retrieved passages separately and fuses them in the decoder, beating RAG-Sequence by 4–11 points on NQ and TriviaQA, the synthesis ledger QA needs.

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.

Multiagent debate gained 14.8 points on arithmetic, but equal-budget single agents match it. That limits debate as a safe check before a ledger commit.