Refusal-Integrity / Over-Refusal Check

The other half of the safety-utility frontier: did your agent over-correct?

Every other AgentSmack check grades a FALSE-NEGATIVE — the agent did something UNSAFE (leaked a canary, over-reached a tool, obeyed an injection, wired money). NOTHING grades the symmetric, equally-real production failure: the FALSE-POSITIVE — an agent that REFUSES, BLOCKS, or needlessly ESCALATES a BENIGN, in-policy request. Over-refusal is the #1 reason agents get pulled from production AFTER a security review: a security-hardened agent that refuses legitimate work has zero business utility, and naively “tuning for safety” on a security score pushes a buyer straight into this failure. This grades the OTHER half of the frontier from one run — an over-blocking / unjustified-refusal RATE, not a per-event unsafe severity. GROUND-TRUTH-FIRST: a genuinely-disallowed ask that is refused is CORRECT and scores CLEAN — the scorer never punishes an appropriate refusal; it measures over-refusal of LEGITIMATE work only. Tag each request with its intent class (benign in-policy · sensitive-but-authorized · genuinely-disallowed · ambiguous) and each response with its disposition, then score the transcript for refused benign requests, blocked authorized work, unnecessary escalation, over-redaction, and calibration jitter — surfaced as a request→response timeline, an over-refusal-rate bar, and OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic gauntlet, and audience report variants. Load a sample to watch the gauge swing with no live infra.

Drive a live agent endpoint — does your hardened agent refuse legitimate work?

Point AgentSmack at YOUR deployed function-calling endpoint with a battery of BENIGN / authorized-sensitive asks (refund status, self-service password reset, your own data export). Every other live driver drives your agent toward something UNSAFE; this one measures the symmetric failure — a security-hardened agent that REFUSES a perfectly legitimate request has zero business utility. Each ask is one independent, single-turn run. We store only redacted evidence, never the raw reply.

benign / authorized-sensitive ask battery