Full-Stack Gauntlet

Smack the whole agent stack in one run.

AgentSmack scores dozens of agent surfaces in isolation — but a CISO buys one answer: is my whole agent stack production-ready? The Full-Stack Gauntlet is that one call. Feed a keyed bundle of per-surface requests (each carries the SAME body that surface’s own /api/v1/* lab accepts, keyed by surface) and AgentSmack runs every present surface’s real scorer once, folds the results into one unified production scorecard, consolidates every finding across surfaces into a single list (which flows through the same OWASP/NIST/ATLAS, remediation, attacker-gauntlet, and audience-report bridges every other surface uses), and emits a single whole-agent verdict. Two non-launderable properties make it trustworthy: the verdict is keyed off the composed scorecard’s blocked dimensions, not an average — so a single critical hard blocker on ANY one surface caps the whole-agent verdict regardless of how many others pass — and honest-empty — an absent surface contributes present:false (never a fabricated pass), and an empty bundle is insufficient_signal. Fill with a sample run to watch the whole stack roll up with no live infra.

Each surface carries the SAME request body its own lab accepts. Want the single worst exploit instead? Proof-of-Exploit Reproducer · How does the score compare? Scorecard percentile.