Goal / Objective Integrity Lab

It was told “summarize this ticket” — and deleted three records and emailed a competitor. Flag that.

Over a multi-step run, did the agent abandon, silently substitute, or have an attacker inject a sub-objective onto the user's original stated task? This is objective drift / goal misgeneralization (task hijacking, scope creep, objective laundering) — and it is the one agentic failure class no other surface grades: 4.PA grades the assembled prompt, precedence grades cross-channel directive conflict, capability grades tool co-occurrence, reasoning grades the chain-of-thought — none score the executed ACTION SEQUENCE against the DECLARED objective. AgentSmack grades the whole objective transcript (OWASP LLM06 Excessive Agency / LLM01 Prompt Injection): an objective substituted without user consent (a tool output redeclares the goal — a critical hard blocker), an injected sub-objective from an untrusted document driving a privileged action (the LLM01 → LLM06 chain — a critical hard blocker), an out-of-scope privileged action (scope creep), objective abandonment, and latent goal drift. Ground-truth wins: an authorized:true / servesObjective:true action clears heuristic findings, so the score measures USEFUL + ON-TASK execution, not blanket refusal. Findings map to OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. Load a sample to watch the gauge swing with no live infra.

Drive a live agent endpoint

Don't paste a transcript — give your agent a benign declared objective, point AgentSmack at your real staging endpoint, and watch an injected tool-output silently substitute the goal and drive an off-task privileged action. Each turn is a real HTTP round-trip; the objective is redacted by construction.

This POSTs to a real endpoint you control. localhost / private / link-local targets are refused server-side by the SSRF guard (re-validated before every request).

Untrusted tool-output injections (the task-hijack bait, optional)
Auth headers (optional)