
M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.
#automation
Automation techniques and tools for financial data processing workflows

M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

Atlas hits 42.4% accuracy on Natural Questions with 64 examples, beating PaLM 540B by 3 points at 11B parameters via joint retriever-reader pre-training.

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.

Multiagent debate gained 14.8 points on arithmetic, but equal-budget single agents match it. That limits debate as a safe check before a ledger commit.

TAT-LLM fine-tunes LLaMA 2 7B to 64.60% EM on FinQA, edging GPT-4's 63.91%: an extract-reason-execute pipeline lets a small model do table arithmetic.

RAG hits 0.875 accuracy on post-cutoff facts while fine-tuning plateaus at 0.504. For agents needing frequent ledger updates, retrieval beats fine-tuning.

IRCoT adds +11.3 recall and +7.1 F1 on HotpotQA over one-step RAG by querying retrieval at every reasoning step. A 3B model can then beat GPT-3 175B.

FLARE hits 51.0 EM on 2WikiMultihopQA versus 39.4 for single-retrieval RAG, but calibration failures in chat models limit it for finance agents.

DSPy's compiler lifted Llama2-13b from 9.4% to 46.9% on GSM8K, pointing finance AI pipelines toward maintainable declarative LLM calls.

LATS unifies ReAct, Tree of Thoughts, and Reflexion in one MCTS framework, hitting 92.7% pass@1 on HumanEval with GPT-4.