GFC / FORECAST INSTRUMENT
gfc-ver-3 // calibration exhibitfield replay
Verification relaunch / five saved packages / archived SPC evidence

Right storm.
Wrong core.

Forecast Grade v3 is built to separate those two truths. It rewards a broad low-probability outlook for finding the event, then asks more of a 15% or 30% core: was the activity actually where the stronger claim was drawn?

The old calculation called the March 10 package 37.7 / F. The revised instrument calls it 76.9 / C: comfortably passing, visibly imperfect, and specific about why.

Exhibit 01 / the instrument

Four signals.
One grade.

The headline score is now an outcome measure with a spatial conscience. Low-risk coverage is allowed to be wide. High-risk cores are not allowed to be vague.

35%

Event capture

How many relevant reports landed inside the 25-mile neighborhood of any non-zero forecast area?

30%

Tier-aware placement

Below 15%, broad coverage earns bounded credit. At 15%+, softened area contingency measures the core.

20%

Event yield

Does a high-probability core produce the amount of activity its size and tier imply?

15%

Significant placement

Significant reports are scored proportionally. One hit is evidence—not perfection.

Exhibit 02 / March 10 replay

Read the map,
then read the math.

This is the saved forecast geometry normalized into a small instrument panel. Colored outlines are real March 10 probability contours; pale points are a sample of the archived SPC reports used in the replay.

evidence field / geometry excerptnot to geographic scale
March 10 forecast geometry and sampled reports Three normalized corridor diagrams for tornado, wind, and hail probability contours, with sampled report points. TORNADO / 46 REPORTS / 44 CAPTURED WIND / 160 REPORTS / 153 CAPTURED HAIL / 402 REPORTS / 402 CAPTURED
2–5% envelope15% core30% coresampled report

Geometry normalized from the saved 2026-03-10 package; report points sampled from the 608-report archived SPC replay. Visual comparison aid, not a geographic map.

Tornado

78.5
capture 44 / 46core placement 48.3%

The broad envelope found nearly every report. The stronger core was less spatially exact, and the significant contour had no significant verification.

Wind

77.3
capture 153 / 160yield .808

Excellent event coverage, but the 15%+ placement and 30% yield claim both leave visible debt.

Hail

75.0
capture 402 / 402sig placement 16 / 117

This is the important correction: the hail outlook saw the event, but its stronger areas did not line up tightly enough to deserve a 100.

Exhibit 03 / calibration set

Not one forecast.
A small family.

Five saved packages, matched to their archived SPC report dates, plus synthetic edge cases. The real set is still small—but it already catches the two failure modes we were worried about: overgrading and overpenalizing.

DateEvidencev3 result
02 / 23no usable severe geometry—
03 / 02no reports / quiet day—
03 / 037 reports / missed hail20.6
03 / 0965 reports / mixed placement74.0
03 / 10608 reports / strong capture76.9

Read the spread: v3 does not pretend a quiet day can yield a meaningful single-case accuracy grade. It does, however, refuse to call a broad successful outlook perfect—and it still fails a genuine miss.

Headline grade / saved replay

03 / 03
20.6
03 / 09
74.0
03 / 10
76.9
The target is not a leaderboard. It is a stable language for “you found it,” “you placed it,” and “you overclaimed it.”
Exhibit 04 / edge behavior

Where the grade
draws a line.

These are the guardrails that keep v3 from becoming either a compactness contest or a participation trophy.

CASE / LOW-RISK CAPTURE

2% envelope, events inside

83.8

Capture earns real credit. Placement is deliberately bounded below perfect because a broad envelope did not localize the threat.

CASE / CORE DEBT

15%+ area, activity clusters elsewhere

70s

The outlook can still pass when the event is found, but core placement and significant coverage pull the grade down.

CASE / OVERCLAIM

Huge 30% core, one report

< 60

High-risk area creates an expectation. When the observed activity cannot support it, event yield and placement fail together.

Exhibit 05 / the ceiling

100 is an exceptional claim.

It is still possible. A forecast with near-total capture, well-placed 15%+ cores, activity that supports its high-risk claims, and proportionally successful significant placement can earn it. But a low-probability wide net cannot reach 100 merely because the storm happened somewhere inside.

What to critique next
Does the core placement penalty feel proportional?
Should a strong low-risk outlook live in the low 80s?
Are the 15% / 30% expectations calibrated across more cases?