
DIN-SQL: GPT-4 jumps 67.4%→85.3% EX on Spider
DIN-SQL lifts GPT-4 from 67.4% to 85.3% Spider execution accuracy via schema-link and self-correct stages—same decomposition fits Beancount BQL.
#llm
Large language model research with applications in financial tasks

DIN-SQL lifts GPT-4 from 67.4% to 85.3% Spider execution accuracy via schema-link and self-correct stages—same decomposition fits Beancount BQL.

On BIRD, GPT-4 reaches 54.89% execution accuracy with domain hints and 34.88% without—a 20-point gap any Beancount NL→BQL interface must close.

STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.

Microsoft's GraphRAG posts 72–83% sensemaking wins over vector RAG; a 2025 audit collapses those after correcting judge bias—caution for multi-doc ledger QA.

FinAuditing shows top LLMs hit just 13.86% on financial math verification of real SEC XBRL filings, capping what AI accounting tools can automate unaided.

InvestorBench: Qwen2.5-72B leads stock trading at 46.15% CR; finance-tuned Palmyra-Fin backfires on equities—size beats domain fine-tuning.

StructRAG routes each query to a table, graph, catalogue, algorithm, or chunk structure, beating GraphRAG by 28 points and running 22× faster.

Under equal thinking-token budgets, single-agent LLMs match or beat multi-agent systems on multi-hop reasoning—favor simpler finance agent designs.

M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

Atlas hits 42.4% accuracy on Natural Questions with 64 examples, beating PaLM 540B by 3 points at 11B parameters via joint retriever-reader pre-training.