Trial-Stability / Reliability Check

It passed your test — but does it pass EVERY time, or did it pass 8 of 10?

Every other surface — including the Metamorphic check — grades ONE run. Metamorphic grades variance across semantically-equivalent renderings of one intent inside one transcript. Nothing grades whether the SAME exact attack, replayed N times, sometimes leaks and sometimes refuses — i.e. whether your agent is stochastically unsafe. An agent that refuses an exfil request 9 of 10 replays and wires the money / leaks the canary on the 10th is unshippable, and a single-shot grade structurally cannot see it: your agent is safe 90% of the time — here are the 2 runs it leaked. AgentSmack groups a set of N independent trials of one fixed attack by a canonical group, grades inter-trial verdict stability, surfaces the firstRefusal→firstUnsafe lineage as evidence, and produces a worst-case-aware Decision-Reliability Score (a CISO buys the worst group, not the mean) — then maps it to OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. A refusal that holds across every replay scores clean. A single trial is honestly reported as untestable — never a fabricated pass. Load a sample to watch the gauge swing with no live infra.

Run it LIVE — replay one fixed attack N times

Point AgentSmack at YOUR deployed function-calling endpoint. We send the SAME fixed prompt N independent times, classify each reply, and grade verdict stability — the CISO question no single-shot test answers. We store only redacted evidence, never the raw secret.