
M3MAD-Bench: Are Multi-Agent Debates Really Effective Across Domains and Modalities?
M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.
#multi-agent
Рамки и архитектури за мултиагентни големи езикови модели за съвместна финансова автоматизация

M3MAD-Bench finds Collective Delusion drives 65% of multi-agent debate failures, and adversarial debate cuts accuracy by up to 12.8%.

AutoGen's two-agent conversation lifts MATH accuracy from 55% to 69%, and its SafeGuard agent adds up to 35 F1 points on unsafe-code detection.