Assertion-Integrity / Output-Truthfulness Lab

It said “I issued the $400 refund” — and never did. Catch your agent lying about what it did.

Were the agent's stated claims about what it DID and what is TRUE actually backed by the run's real executed actions and tool results? This is the agent confabulation / output-truthfulness attack class — and it is the one failure no other surface grades: RAG-runtime grades citation-laundering only within a retrieval context, reasoning grades the chain-of-thought, output-handling grades the render channel, goal grades the action SEQUENCE vs the objective — none grade whether the agent's CLAIMS are backed by its ACTUAL execution. AgentSmack grades the whole assertion transcript (OWASP LLM09 Misinformation / LLM06 Excessive Agency): a fabricated / phantom action claimed (“I issued the refund” with no action.executed — the catastrophic phantom-action trust failure, a critical hard blocker), a fabricated tool-result value (an invented balance / order status grounded in a result that does not exist — a critical hard blocker), a failed action claimed successful, an unverified claim asserted as fact, an untrusted-source claim laundered as fact, and a capability overclaim. Ground-truth wins: a claim that correctly matches a real successful action / result scores clean — the scorer measures truthful execution, not blanket suspicion. Findings map to OWASP LLM / NIST AI RMF coverage with paste-able remediation, a synthetic-attacker gauntlet, and audience report variants. Raw claim/value bytes are never stored. Load a sample to watch the gauge swing with no live infra.