
Verifiably Safe Tool Use for LLM Agents: STPA Meets MCP
STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.
#security
Изследвания на безопасност, сигурност и предпазни механизми за ИИ агенти във финансов контекст

STPA plus capability-enhanced MCP yields formal safety specs for LLM tool use, with Alloy proving no unsafe flows in a calendar case study.

AGrail's two-LLM guardrail cuts prompt injection attack success to 0% while preserving 95.6% of benign agent actions on Safe-OS.

ShieldAgent hits 90.4% accuracy on agent attacks with 64.7% fewer API calls by using probabilistic rule circuits instead of LLM guardrails.

GuardAgent enforces LLM agent policies by running Python code, hitting 98.7% accuracy with no task failures, versus 81% and up to 71% failure for prompt rules.