Self-Modification / Defense-Evasion Integrity Check
The agent that turns off its own guardrail — then acts.
Mid-run, an agent executes an action that alters its own governing config to slip a constraint: it disables its own guardrail / policy / content-filter, rewrites its own system prompt to embed an injected directive, self-grants a privileged tool or widens its own scope, or lowers its own approval threshold so it can act autonomously. This is the agent analogue of MITRE ATT&CK Impair-Defenses (T1562) — the high-blast-radius failure where the agent reconfigures the safety substrate it runs under to escape it. Persistence-scan grades an EXTERNAL durable mechanism that OUTLIVES the run; 4.SC model-safety-config grades the safety config of a single model CALL; sleeper grades a dormant instruction that detonates later in the SAME transcript; everywhere else a guardrail-disable is graded only as a STATIC pattern in INCOMING text (an MCP registry description, a RAG document, a honeypot tool) — none grade the agent's OWN run EXECUTING the self-reconfiguration. AgentSmack scores a self-modification transcript for exactly this read→modify→action lineage — surfacing the self-disable-then-act path as evidence, mapped to OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. An operator-approved, self-scoped, benign config change scores clean. Load a sample to watch the gauge swing with no live infra.