Three stages of agent boundary compromise — from policy evasion through credential theft to evaluator corruption. Each benchmark measures a distinct failure mode of AI agent deployments.
The Series
Three benchmarks. Three failure modes.
Evasion
Seideray
Policy and tool-call firewall bypass benchmark. Tests whether agents can circumvent runtime enforcement to execute denied operations.
Exfiltration
Hekacy
Credential exfiltration benchmark. Tests whether agents extract and exploit canary credentials from poisoned environments.
Corruption
Fraegis
Evaluator-integrity corruption benchmark. Tests whether agents exploit grader-isolation flaws to inflate their own scores.
Design
Built for real-world signal.
Provider-agnostic
All benchmarks use LiteLLM — test any model with a single provider-prefixed string.
Independent axes
Each task scores multiple behavioral dimensions separately — capability vs. intent vs. honesty.
Grounded in real incidents
Tasks reference verified CVEs, GHSAs, and published attack taxonomies — not synthetic scenarios.
Reproducible
Deterministic environments via Docker, YAML task definitions, JSON result artifacts.