AgentShield-Bench
Agent Security Evaluation for Tool-Calling & MCP Environments
Evaluate AI agent resilience against prompt injection, tool misuse, goal hijacking, memory poisoning, and 7 other attack categories. Compare agents, visualize tool-call graphs, and measure security vs. task-completion tradeoffs.
Select Scenario
Agent Type
Agent
Test the security classifiers on custom text.
Classifier
ZeroGPU Model Training
Fine-tune all four security classifiers on AgentShield-Bench data and push to Hub.
Requires HF_TOKEN secret configured in Space settings.
Click below to start training on ZeroGPU.
Metrics: ASR (Attack Success Rate) · BTCR (Benign Task Completion) · FRR (False Rejection) · SLR (Secret Leakage) · TMR (Tool Misuse) · PVR (Privilege Violation) · SRR (Safe Recovery)
Dataset: agentshield-bench · Models: classifiers