Forecast Calibration / Predictive-Accuracy
Where did the pre-flight forecast UNDER-predict your real risk?
AgentSmack has a strong STATIC on-ramp — the defense-posture forecast predicts, per adversarial technique, which of your declared defenses would resist vs land, before a single probe fires. And it has a strong DYNAMIC layer — every scan surface emits a real failed/clean outcome. This lens is the predict→observe loop nothing else closes: it confusion-matrices each per-technique forecast against the actual outcome of that technique's target surface and answers the one CISO question no competitor will — can I trust AgentSmack's own pre-flight forecast, and where did it UNDER-predict my real risk? The load-bearing output is the dangerous false-negative band: techniques the forecast said would be resisted that actually failed live — the pre-flight blind spot, and the first thing to re-harden. It is purely compositional: it re-runs nothing, reusing the defense-posture forecast and the canonical findings envelope verbatim. And it is honestly humble — a technique whose target surface was never tested is not observed and excluded from the accuracy math (never silently counted as a pass), and precision / recall / accuracy read “—” rather than a fabricated 100 when their denominator is zero. It adds no new surface and no new dimension.
Dangerously-optimistic sample — the forecast said safe, it ran unsafe
Forecast calibration — can I trust AgentSmack's pre-flight forecast?
Forecast dangerously optimisticConfusion matrix — predicted × observed (3 observable)
Confirmed risk (true positive)
0
Conservative flag (false positive)
0
Dangerous miss (false negative)
1
Confirmed safe (true negative)
2
Excluded from accuracy: 14 not observed, 0 not observable (honest-empty — never counted as a pass).
Precision
—
Recall
0/100
Accuracy
67/100
Dangerous misses (forecast said safe, ran unsafe)
- lethal_trifecta_exfilexfiltrationegressforecast: likely_resisted
Forecast DANGEROUSLY OPTIMISTIC — 1 technique(s) the pre-flight forecast predicted RESISTED actually FAILED live (the dangerous false-negative band, the pre-flight blind spot). This is non-launderable: a rosy forecast cannot clear a real live failure. Re-harden the dangerous-miss techniques first (recall 0/100 over real failures); 0 risk(s) were correctly predicted across 3 observable technique(s).
Well-calibrated sample — the forecast tracked the real run
Forecast calibration — can I trust AgentSmack's pre-flight forecast?
Forecast well calibratedConfusion matrix — predicted × observed (3 observable)
Confirmed risk (true positive)
0
Conservative flag (false positive)
0
Dangerous miss (false negative)
0
Confirmed safe (true negative)
3
Excluded from accuracy: 14 not observed, 0 not observable (honest-empty — never counted as a pass).
Precision
—
Recall
—
Accuracy
100/100
Dangerous misses (forecast said safe, ran unsafe)
None — the forecast did not under-predict a single live failure.
Forecast WELL CALIBRATED — across 3 observable technique(s) there were NO dangerous misses and no excess conservative flags: 0 confirmed risk(s) and 3 confirmed safe outcome(s) all matched the live run. This is the trust signal — AgentSmack's own pre-flight forecast tracked your real outcomes.
Grade your own forecast against your run
Need a forecast to grade? Generate a pre-flight defense-posture forecast here → then paste it alongside your run's findings to see where it under-predicted real risk.
Related read-only lenses: Defense-Posture Forecast (the forecast this grades) · Attacker Economics. POST { forecast, findings } to /api/v1/forecast-calibration to run it on your own data.