Forecast Calibration / Predictive-Accuracy

Where did the pre-flight forecast UNDER-predict your real risk?

AgentSmack has a strong STATIC on-ramp — the defense-posture forecast predicts, per adversarial technique, which of your declared defenses would resist vs land, before a single probe fires. And it has a strong DYNAMIC layer — every scan surface emits a real failed/clean outcome. This lens is the predict→observe loop nothing else closes: it confusion-matrices each per-technique forecast against the actual outcome of that technique's target surface and answers the one CISO question no competitor will — can I trust AgentSmack's own pre-flight forecast, and where did it UNDER-predict my real risk? The load-bearing output is the dangerous false-negative band: techniques the forecast said would be resisted that actually failed live — the pre-flight blind spot, and the first thing to re-harden. It is purely compositional: it re-runs nothing, reusing the defense-posture forecast and the canonical findings envelope verbatim. And it is honestly humble — a technique whose target surface was never tested is not observed and excluded from the accuracy math (never silently counted as a pass), and precision / recall / accuracy read “—” rather than a fabricated 100 when their denominator is zero. It adds no new surface and no new dimension.

Dangerously-optimistic sample — the forecast said safe, it ran unsafe

Forecast calibration — can I trust AgentSmack's pre-flight forecast?

Forecast dangerously optimistic

Confusion matrix — predicted × observed (3 observable)

Confirmed risk (true positive)

0

Conservative flag (false positive)

0

Dangerous miss (false negative)

1

Confirmed safe (true negative)

2

Excluded from accuracy: 14 not observed, 0 not observable (honest-empty — never counted as a pass).

Precision

Recall

0/100

Accuracy

67/100

Dangerous misses (forecast said safe, ran unsafe)

  • lethal_trifecta_exfilexfiltrationegressforecast: likely_resisted

Forecast DANGEROUSLY OPTIMISTIC — 1 technique(s) the pre-flight forecast predicted RESISTED actually FAILED live (the dangerous false-negative band, the pre-flight blind spot). This is non-launderable: a rosy forecast cannot clear a real live failure. Re-harden the dangerous-miss techniques first (recall 0/100 over real failures); 0 risk(s) were correctly predicted across 3 observable technique(s).

Well-calibrated sample — the forecast tracked the real run

Forecast calibration — can I trust AgentSmack's pre-flight forecast?

Forecast well calibrated

Confusion matrix — predicted × observed (3 observable)

Confirmed risk (true positive)

0

Conservative flag (false positive)

0

Dangerous miss (false negative)

0

Confirmed safe (true negative)

3

Excluded from accuracy: 14 not observed, 0 not observable (honest-empty — never counted as a pass).

Precision

Recall

Accuracy

100/100

Dangerous misses (forecast said safe, ran unsafe)

None — the forecast did not under-predict a single live failure.

Forecast WELL CALIBRATED — across 3 observable technique(s) there were NO dangerous misses and no excess conservative flags: 0 confirmed risk(s) and 3 confirmed safe outcome(s) all matched the live run. This is the trust signal — AgentSmack's own pre-flight forecast tracked your real outcomes.

Grade your own forecast against your run

Need a forecast to grade? Generate a pre-flight defense-posture forecast here → then paste it alongside your run's findings to see where it under-predicted real risk.

Related read-only lenses: Defense-Posture Forecast (the forecast this grades) · Attacker Economics. POST { forecast, findings } to /api/v1/forecast-calibration to run it on your own data.