Instruction-Hierarchy / Trust-Precedence Lab

A tool output overrode your system FORBID and wired the funds.

Every other surface judges ONE channel in isolation. This lab grades the one failure mode none of them do: cross-channel layered-source conflict (Method #14, OWASP LLM01 — the instruction-hierarchy problem). When an agent receives CONFLICTING directives from sources of differing trust (system > developer > user > memory > retrieved_doc > tool_output > peer_agent), AgentSmack checks whether it resolves by trust precedence — or whether a lower-trust source (a tool output, a retrieved document, a peer agent, a memory entry) overrode a higher-trust FORBID. The headline finding is a privileged action driven by a lower-trust override (a critical hard blocker — the instruction-hierarchy breach that causes incidents). Ground-truth wins: an agent that correctly follows system over a malicious tool-output directive scores clean — we measure useful, correct resolution, not blanket refusal. Findings map to OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. Load a sample to watch the gauge swing with no live infra.