
GuardAgent: Deterministic Safety Enforcement for LLM Agents via Code Execution
GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.
#ai
Artificial intelligence research and applications in finance and accounting

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.

Multiagent debate gained 14.8 points on arithmetic, but equal-budget single agents match it. That limits debate as a safe check before a ledger commit.

NeurIPS 2024 ablation: dropping the LLM from Time-LLM and CALF improves accuracy, with up to 1,383× faster training. Use purpose-built models for finance AI.

AuditCopilot cuts journal entry fraud false positives from 942 to 12, but ablation shows the LLM mainly synthesizes Isolation Forest scores.

TAT-LLM fine-tunes LLaMA 2 7B to 64.60% EM on FinQA, edging GPT-4's 63.91%: an extract-reason-execute pipeline lets a small model do table arithmetic.

RAG hits 0.875 accuracy on post-cutoff facts while fine-tuning plateaus at 0.504. For agents needing frequent ledger updates, retrieval beats fine-tuning.

IRCoT adds +11.3 recall and +7.1 F1 on HotpotQA over one-step RAG by querying retrieval at every reasoning step. A 3B model can then beat GPT-3 175B.

FLARE hits 51.0 EM on 2WikiMultihopQA versus 39.4 for single-retrieval RAG, but calibration failures in chat models limit it for finance agents.

Lewis et al. hit 44.5 EM on Natural Questions with a frozen FAISS index; for Beancount ledgers that means reindex-or-miss when balances change daily.

MultiHiertt shows models score 38% F1 against 87% for humans on 10,440 financial QA pairs, with a 15-point drop on cross-table questions.

ConvFinQA's best model scores 68.9% execution accuracy versus 89.4% for human experts—a 21-point gap that multi-turn ledger chat still faces.

TAT-QA's hybrid table-text questions showed evidence grounding, not arithmetic, is finance AI's bottleneck. Fine-tuned 7B LLMs hit 83% F1 by 2024.