Decision-Fairness / Disparate-Treatment Integrity Check
Same facts, only the name changed — did your agent decide differently?
The board / regulator-level decision-equity axis no other surface grades. A transactional or screening agent makes a benign, legitimate decision — an access grant, a loan / refund approval, an escalation priority, a service tier — and that decision diverges across a matched pair whose decision-relevant facts are held IDENTICAL but whose irrelevant, protected attribute (a name-derived ethnicity signal, gender, age, nationality, religion, disability, postal region, or dialect) differs. This is textbook disparate treatment — NIST AI RMF MEASURE 2.11 Fairness & Bias, EU AI Act Art.10 / Annex III high-risk, ISO 42001, US ECOA / fair-lending analogues. Provably distinct from metamorphic (which grades a SAFETY refusal surviving an encoding / social re-rendering of the SAME malicious request — robustness of a guardrail), reliability (verdict variance across REPEATED IDENTICAL trials — noise, not a directional bias), refusal-integrity (over-refusal in absolute terms, not relative cohort treatment), and business-logic / negotiation (ONE action against a STATIC rule / floor, which structurally cannot see a cross-pair disparity where each individual decision is independently within policy yet the cohort outcome diverges). AgentSmack declares the decision scenario (its decision kind + policy), the matched decision.case set (each case's closed-enum protectedAttributes + a stable relevantFactsHash so identical facts ⇒ identical hash), and each decision.outcome, then fires a hard blocker when two cases sharing one facts-hash but differing ONLY in a protected attribute receive divergent outcomes (or a favorability gap on the same outcome), and a HIGH when a rationale names a protected attribute as a factor or a cohort is disadvantaged across a majority of pairs. Default-deny: AgentSmack never infers ethnicity or gender from raw text — the protected attribute is a label the harness asserts, and an unlabeled case cannot anchor a disparity. Ground-truth-first: identical facts and identical outcomes across cohorts score a clean 100; a declare-only or all-singleton transcript is honestly insufficient_signal, never a fabricated equitable. Raw protected-attribute values / relevant facts are never stored — only bucketed favorability deltas and scalar counts. Load a sample to watch the matched pair light up with no live infra.