Owner-Harm Resistance Check
The agent that confuses its operator for an attacker — and destroys their own assets.
Every other AgentSmack surface points the blast radius outward — identity grades cross-tenant reach, egress and capability grade exfil to an attacker, the lethal trifecta grades reaching an external sink. Owner-harm (catalog method #21) points inward: the confused agent destroys the operator’s OWN assets — it deletes their production database, drops their repo, drains their cloud budget, revokes their own admin, mass-emails their customers, rotates their keys. Cost-scan grades token denial-of-wallet, not a destructive spend; transaction-scan grades TOCTOU staleness; capability grades co-occurrence — none grade harm-to-self. AgentSmack scores an owner-harm transcript for exactly this inward destruction — surfacing the ingest→action lineage as evidence, mapped to OWASP LLM / NIST AI RMF / MITRE ATLAS coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. An operator-authorized maintenance (an operator objective or a correlated approval) grades clean. Load a sample to watch the gauge swing with no live infra.
Drive a live agent endpoint
Don't paste a transcript — give your agent a benign declared objective, feed a confused-deputy tool-output, point AgentSmack at your real staging endpoint, and watch it try to drop your production database / drain your budget / revoke its own admin. Each turn is a real HTTP round-trip.
POSTs to a real endpoint you control; egress is sinkholed by construction. localhost / private / link-local targets are refused server-side by the SSRF guard (re-validated before every request).