Probe-Corpus Coverage Auditor
Which dimensions did your run actually test — and which are blind spots?
AgentSmack grades agents across 60 production-readiness dimensions. The unified scorecard does the honest thing and EXCLUDES a dimension nothing tested as no_signal — it never scores “not yet tested” as “safe.” But that leaves a trap: a front-door prompt check exercises only a handful of prompt-mode dimensions, so a buyer can read an apparently-passing report without realizing ~50 dimensions were never probed. This auditor closes that gap. It maps the surfaces your run actually executed onto the dimension lattice and reports, for every dimension, whether it was tested, only partially tested (a surface ran but produced no gradeable signal), or never tested — an honest blind spot. A ran-but-no-signal surface is never laundered into “tested,” and a credential dimension is not counted as tested by a run that fired zero credential-elicitation probes. Each blind spot carries the exact route to run to close it. Honest-empty: a run that tested nothing reports all 60 dimensions untested, verdict minimal — coverage is never fabricated. Load a sample to see the lattice light up with no live infra.
Pair this with the grade it audits: Unified Production Scorecard · Defense-in-Depth Control Coverage.