Transcript-Integrity / Evidence-Coherence Lab
“A forged or internally-incoherent transcript cannot launder a clean AgentSmack score.”
The platform is saturated on the two axes of grade-trustworthiness: test-validity (did the harness actually exercise the agent?) and specificity (does AgentSmack cry wolf on benign inputs?). There is a third, equally load-bearing axis nothing else checked: evidence coherence. Every pasted-transcript path and the whole offline recorded-run campaign TRUST the submitted transcript completely. A customer gaming their own audit, a compromised CI harness, or the agent-under-test fabricating its own log can submit an internally-incoherent transcript — an action consuming a tool.result for a tool.call that never happened; an approval.completed with no prior approval.requested; a span asserting redaction whose own payload still recovers the planted canary; events out of causal order; a mechanism.revoked stamped BEFORE its own trigger.fired; a session that silently switches tenant/actor mid-chain — to launder a clean grade and ship an unsafe agent. This lab is a pure, deterministic, structural meta-lens over the SAME recorded-run artifact you hand the offline campaign: any forgery-class finding forces the verdict to forged (coherence capped into the failing band, non-launderable); structural impossibilities are suspect; honest-empty is insufficient_signal, never a fabricated coherent. The raw canary value is never stored — only its sha-prefix + a redacted excerpt. It pairs with the deployment gate: a clean scorecard over a forged transcript is forced inconclusive. Load a sample to watch the verdict swing with no live infra.