v0.4-005-full-panel
A readable record of the completed v0.4 full baseline. Every model answered the same 81 cases, every response has a local automated grade, and the page below shows the case packet and eight scores side by side. This is a hobby comparison, not a verified weather-skill leaderboard.
Run post-mortem
What happened. Eight local models each ran all 81 cases, one at a time. Each model has 81 recorded final responses and 81 scores from the same LM Studio Gemma 3n E4B judge. The run campaign lasted 31h 24m from Gemma 4 E4B starting to Qwen finishing. Grading continued after generation and finished at 5:07 PM on September 23, 2026.
What the score means. The average grade was 81.40, the median was 85, and model averages ranged only 79.27 to 82.86. These are grades for the written forecast, reasoning, hazards, uncertainty, practical value, and communication. They are not 648 independent checks against observed weather.
Why models in the 80s is plausible here. The median case grade is 85, 505 of 648 grades (77.9%) are at least 80, and the judge averages 8.98/10 for communication. The grader often praises answers as clear and concise (447 notes use that exact phrase). That can lift a coherent, cautious narrative even when the forecast is not proven right. The five zeroes and 47 scores below 70 also show that the judge does apply large penalties in some cases.
Most important fairness problem. Seven cases still have draft or provisional inputs. Five zero grades fall on three of them, and the saved answers say the packet lacks the data needed for a forecast. That is a case-authoring failure mixed into the model score. It is not fair to read those zeroes as ordinary model mistakes.
Scoring reliability. The same Gemma 3n judge graded all responses, including its own #254 run. For 165 of 648 grades (25.5%), the judge's stated total did not equal its six category scores. The grading script replaced that total with the sum of the categories and recorded the adjustment in the notes. This keeps arithmetic consistent, but it does not make the judge's category choices objective.
Generation reliability. On first attempts, only 381/648 (58.8%) produced visible forecast text; 267/648 (41.2%) ended in reasoning-only output. Retries brought final response coverage to 648/648. Bonsai 27B and Qwen 3.8 27B were the roughest cases, with 80 and 72 reasoning-only first attempts. Qwen #258 alone took 17h 27m. It used xhigh reasoning with a 1,600-token cap; FB-1116 needed a logged manual recovery with low reasoning and a 3,072-token cap.
Model results
| Run | Model | Mean /100 | Median | Range | Grades | Generation time | Reasoning / cap | Average rubric points |
|---|---|---|---|---|---|---|---|---|
| #257 | prism-ml/bonsai-27b | 82.86 | 85 | 62–90 | 81/81 | 1h 36m | on / 1600 tokens | Weather correctness 19.5/25 · Reasoning 15.9/20 · Hazards 12.3/15 · Uncertainty 13.4/15 · Practical value 12.7/15 · Communication 9.0/10 |
| #256 | nvidia/nemotron-3-nano | 82.35 | 85 | 63–90 | 81/81 | 3h 57m | on / 1600 tokens | Weather correctness 18.8/25 · Reasoning 15.7/20 · Hazards 12.5/15 · Uncertainty 13.5/15 · Practical value 12.8/15 · Communication 9.0/10 |
| #258 | qwen/qwen3.8-27b | 82.02 | 84 | 37–89 | 81/81 | 17h 27m | xhigh / 1600 tokens | Weather correctness 19.5/25 · Reasoning 15.9/20 · Hazards 11.7/15 · Uncertainty 13.4/15 · Practical value 12.5/15 · Communication 9.0/10 |
| #255 | nvidia/nemotron-3-nano-4b | 81.98 | 84 | 63–89 | 81/81 | 19m | on / 1600 tokens | Weather correctness 18.8/25 · Reasoning 15.8/20 · Hazards 12.1/15 · Uncertainty 13.6/15 · Practical value 12.7/15 · Communication 9.0/10 |
| #253 | google/gemma-4-12b | 81.20 | 85 | 0–92 | 81/81 | 1h 20m | on / 1600 tokens | Weather correctness 19.2/25 · Reasoning 15.6/20 · Hazards 11.9/15 · Uncertainty 13.3/15 · Practical value 12.5/15 · Communication 8.8/10 |
| #254 | google/gemma-3n-e4b | 80.81 | 84 | 59–88 | 81/81 | 16m | off / 1600 tokens | Weather correctness 18.9/25 · Reasoning 15.6/20 · Hazards 11.3/15 · Uncertainty 13.5/15 · Practical value 12.5/15 · Communication 9.0/10 |
| #251 | google/gemma-4-e4b | 80.67 | 84 | 54–91 | 81/81 | 28m | on / 1600 tokens | Weather correctness 19.0/25 · Reasoning 15.5/20 · Hazards 11.7/15 · Uncertainty 13.4/15 · Practical value 12.2/15 · Communication 9.0/10 |
| #252 | google/gemma-4-12b-qat | 79.27 | 85 | 0–91 | 81/81 | 1h 13m | on / 1600 tokens | Weather correctness 18.8/25 · Reasoning 15.4/20 · Hazards 11.2/15 · Uncertainty 13.0/15 · Practical value 12.3/15 · Communication 8.7/10 |
The grader was LM Studio's local google/gemma-3n-e4b for every row. Run #254 is self-graded. Scores are rounded only for display; case scores remain available below.
Deeper analysis in plain language
What the score distribution says
- Across all grades, the mean is 81.40/100 and the median is 85. The middle score is higher than the average because low outliers pull the mean down.
- 505 scores (77.9%) are 80 or higher; 47 (7.3%) are below 70; 5 are zero.
- The top model mean and bottom model mean differ by 3.59 points. The median of the eight model means is 81.59. Treat this as a close cluster, not a decisive win.
- The highest share of possible points was Communication (8.93/10, 89.3%). Weather correctness averaged 19.04/25; reasoning averaged 15.67/20.
Score by case module
| Module | Coverage | Mean /100 |
|---|---|---|
| core_weather | 5 cases, 40 grades | 78.7 |
| expert_case_studies | 13 cases, 104 grades | 81.7 |
| intermediate_forecasting | 1 cases, 8 grades | 83.5 |
| operational_forecasting | 20 cases, 160 grades | 80.8 |
| specialty_forecasting | 42 cases, 336 grades | 81.9 |
The intermediate module contains only one case, so its eight scores are too small a sample to compare with the larger modules.
First answer versus final answer
381/648 first attempts showed a forecast. The rest were reasoning-only or otherwise blank in the initial output. Recovery and retry paths filled every final response slot. That is a major reliability improvement over calling the first blank a zero-quality forecast, but the final completion flag only means the model returned visible text. For incomplete draft cases, a request for missing data also counts as visible.
All runs used a 1,600-token output cap. Six used reasoning on, Gemma 3n #254 used it off, and Qwen #258 used xhigh. Qwen's FB-1116 recovery also changed reasoning and token settings. These differences matter when comparing speed and response quality.
Objective probability scores
| Run / model | Parsed forecasts | Scored targets | Brier (lower better) | Brier skill (higher better) | Ranked score (lower better) | Ranked skill (higher better) |
|---|---|---|---|---|---|---|
| #251 google/gemma-4-e4b | 45/81 | 53 | 0.2926 | 0.2884 | 0.5332 | 0.0197 |
| #252 google/gemma-4-12b-qat | 47/81 | 55 | 0.2767 | 0.3995 | 0.4433 | -0.2341 |
| #253 google/gemma-4-12b | 47/81 | 55 | 0.2903 | 0.422 | 0.6143 | -0.1888 |
| #254 google/gemma-3n-e4b | 44/81 | 51 | 0.2688 | 0.3889 | 0.5952 | -0.8386 |
| #255 nvidia/nemotron-3-nano-4b | 46/81 | 54 | 0.3154 | 0.1581 | 0.8154 | -0.9119 |
| #256 nvidia/nemotron-3-nano | 47/81 | 55 | 0.2558 | 0.3641 | 0.6033 | -0.2882 |
| #257 prism-ml/bonsai-27b | 47/81 | 55 | 0.2512 | 0.3912 | 0.6587 | -0.3619 |
| #258 qwen/qwen3.8-27b | 46/81 | 53 | 0.2646 | 0.3826 | 0.3505 | 0.1047 |
These diagnostics cover a different subset of cases for each model, 51 to 55 targets each. The pool mixes synthetic outcomes, observation-backed data, and unresolved draft evidence. The numbers are clues for the audit, not a fair common-sample ranking.
Current project state
The active archive is now the completed v0.4-005 campaign. It preserves each run's metadata, raw generation artifacts, final responses, per-case judge grades, retry history, and reports. The local baseline dashboard and this case guide show only this campaign. Older v0.1-v0.3 runs and cases remain in the repository.
The automation is useful enough to execute and grade a full local panel, but case quality and score interpretation are now the limiting factors. Several start-here documents still describe the pre-v0.4-005 state; the updated project map and status note point to this report.
Case-by-case guide
Each case shows the prompt's setup, data, task, expected concepts, and the eight model grades. This page omits hidden verification truth values. For a hand audit, check whether the input packet gives the model enough information, whether the requested task is clear, and whether the expected-concept list matches what you want to reward.
The observations and model guidance the case gives the model.
A measure of energy that can help thunderstorms grow.
A layer that can stop storms from starting.
How wind speed or direction changes with height.
How much water vapor is moving through the air.
A penalty for probability forecasts. Lower is better when truth is known.
A penalty for probability forecasts with ordered outcomes, such as snow, mixed precipitation, or rain.
How much better or worse a forecast is than a stated reference forecast. Positive is better than that reference.
Showing 81 cases.
core_weather
5 cases, 40 model-case scores
FB-0001 core temperature front
In plain terms. A shallow cold front moved through the area during the morning. Skies have cleared behind the front, but cold advection continues through the afternoon. The forecast problem is whether temperatures recover much after the frontal passage.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 18z-00z
What the model received
- surface_observations: At 15z, temperatures were 58 F ahead of the front and 47 F behind it. Winds behind the front were north at 15-20 kt with gusts near 25 kt. Dewpoints fell from the lower 50s into the upper 30s after frontal passage. By 17z, skies were mostly clear, but temperatures behind the front had only risen to 49-51 F. Source label: Mock author-provided packet
- guidance: The raw model 2-meter temperature forecast warms the area to 59 F by 21z. A bias-corrected blend reaches 54 F. Forecast soundings keep mixing shallow, with the top of the mixed layer near 900 mb and continued low-level cold advection. Source label: Mock author-provided packet
What it had to do
Produce an afternoon high-temperature forecast for the post-frontal area. Explain whether the raw model temperature guidance is believable, and state your confidence.
What the case says a strong answer should cover
- Cold advection limits post-frontal warming
- Clear skies help warming but may not overcome shallow cold advection
- Raw 2-meter guidance may be too warm after frontal passage
- Observed temperature trends should be weighted heavily
Scoring notes in the case file
- Strong answers should forecast a high closer to the lower or middle 50s than near 60 F.
- Credit answers that explain why sunshine alone does not guarantee strong recovery.
- Penalize answers that blindly follow the raw 59 F guidance without discussing frontal timing or cold advection.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 81.2, median 82, range 73–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 80 | 82 | 82 | 73 | 87 | 82 | 82 | 82 |
FB-1001 severe probability
In plain terms. A warm-sector afternoon environment is becoming favorable for organized convection, but the timing of initiation remains uncertain.
Case record. cases/v0.2 · v0.2 · severe_weather · forecast time: Synthetic case issue time: 2026-05-18 18:00 UTC
What the model received
- surface_observations: The area is 78 F with a 70 F dewpoint. Temperatures have risen 4 F since 15z and the surface wind is south at 14 kt. Source label: Synthetic author packet
- sounding_and_guidance: MLCAPE is about 1800 J/kg, 0-6 km bulk shear is 35 kt, and CIN is weakening. Guidance breaks the cap near 21z but differs by two hours on storm initiation. A broken line is possible after initiation. Source label: Synthetic author packet
What it had to do
Forecast the probability of at least one severe thunderstorm in the area during 21z-03z. Give a probability from 0 to 1, the main trigger, the largest uncertainty, and one observation that would change the forecast. End with exactly one fenced JSON block (```json ... ```) using the documented forecasts format, with target_id severe_storm_probability and a numeric value.
What the case says a strong answer should cover
- Separate environmental support from initiation timing
- Represent uncertainty in the probability
- Identify a specific observational trigger
Scoring notes in the case file
- Reasoning is reviewed separately from the structured probability.
- Do not award objective credit for a probability inferred from prose.
Structured target(s). severe_storm_probability (binary, probability)
Across the eight runs. Mean 83.9, median 86, range 76–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 84 | 81 | 76 | 87 | 82 |
FB-1003 temperature point forecast
In plain terms. A dry, mostly sunny day follows a cool morning. The task is a bounded short-range point forecast rather than a discussion of severe weather.
Case record. cases/v0.2 · v0.2 · general · forecast time: Synthetic case issue time: 2026-07-02 12:00 UTC
What the model received
- surface_observations: At 12z the temperature is 58 F with a 42 F dewpoint. The temperature rose 8 F in the last three hours, skies are clearing, and the wind is northwest at 8 kt. Source label: Synthetic author packet
- guidance: Two guidance members predict a high of 30 F and one predicts 34 F. The local persistence baseline is 29 F. Source label: Synthetic author packet
What it had to do
Forecast the maximum 2 m temperature from 12z through 00z in degrees Fahrenheit. Give one numeric value, the main physical reason, and a useful uncertainty range. End with exactly one fenced JSON block (```json ... ```) with target_id maximum_temperature and a numeric value in degrees Fahrenheit.
What the case says a strong answer should cover
- Use the observed warming trend
- Account for clearing and low humidity
- Keep the point forecast within the guidance range unless justified
Scoring notes in the case file
- The objective score uses the numeric JSON value only.
- The narrative range is reviewed separately.
Structured target(s). maximum_temperature (continuous, degF)
Across the eight runs. Mean 72.0, median 71, range 63–82. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 63 | 70 | 72 | 69 | 82 | 82 | 75 | 63 |
FB-1005 canary rain probability
In plain terms. A weak disturbance is approaching, with modest but nonzero rain potential.
Case record. cases/v0.2 · v0.2 · general · forecast time: Synthetic canary issue time: 2026-04-10 12:00 UTC
What the model received
- observations: The area is dry at 12z. Moisture is increasing, but lift remains weak. Source label: Synthetic canary packet
- guidance: Guidance gives measurable rain a 20-40 percent chance during the afternoon. Source label: Synthetic canary packet
What it had to do
Give a probability from 0 to 1 that measurable rain occurs during the afternoon. End with exactly one fenced JSON block (```json ... ```) using target_id rain_probability and a numeric value.
What the case says a strong answer should cover
- Use the weak-lift and increasing-moisture balance
Scoring notes in the case file
- Canary case for structured parsing and target coverage.
Structured target(s). rain_probability (binary, probability)
Across the eight runs. Mean 79.0, median 82, range 73–82. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 82 | 82 | 73 | 82 | 77 | 82 | 73 |
FB-1006 canary temperature
In plain terms. A sunny, dry day supports a straightforward short-range temperature forecast.
Case record. cases/v0.2 · v0.2 · general · forecast time: Synthetic canary issue time: 2026-06-11 12:00 UTC
What the model received
- observations: At 12z the temperature is 60 F, dewpoint is 40 F, and skies are clear. Source label: Synthetic canary packet
- guidance: Guidance clusters at 82-86 F for the afternoon high. Source label: Synthetic canary packet
What it had to do
Forecast the afternoon maximum temperature in degrees Fahrenheit. End with exactly one fenced JSON block (```json ... ```) using target_id maximum_temperature and a numeric value in degrees Fahrenheit.
What the case says a strong answer should cover
- Use the clear-sky and dry-air setup
Scoring notes in the case file
- Canary case for numeric parsing and unit preservation.
Structured target(s). maximum_temperature (continuous, degF)
Across the eight runs. Mean 77.2, median 78, range 70–82. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 70 | 82 | 70 | 82 | 80 | 75 | 77 | 82 |
expert_case_studies
13 cases, 104 model-case scores
FB-0010 forecast bust postmortem
In plain terms. A forecast missed a high-impact precipitation event. The task is to diagnose why the forecast busted and what should have been watched in real time.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid post-event analysis
What the model received
- original_forecast: The forecast issued at 12z called for 0.25-0.75 inches of rain, with the heaviest amounts south of the metro area. Forecasters expected the main forcing to pass too far south and believed dry air would limit rainfall efficiency near the metro. Source label: Mock author-provided packet
- observed_outcome: By 06z, the metro area received 3-5 inches of rain, with localized flooding. Radar showed repeated convective elements moving west-east along a boundary that stalled near the metro for 5 hours. Source label: Mock author-provided packet
- missed_signals: Evening observations showed the boundary had stopped moving south. Surface dewpoints rose 4-6 F north of the expected axis. A 925-mb jet strengthened from the south after 00z, and precipitable water increased to near record values. One model run hinted at the farther-north axis, but it was ignored as an outlier. Source label: Mock author-provided packet
What it had to do
Prepare one forecast-bust analysis package. Include what the corrected forecast should have been, top uncertainties, observations that should have changed the forecast, model guidance that was ignored, confidence, and whether any watch would have been appropriate.
What the case says a strong answer should cover
- Stalled boundary and training convection caused the bust
- Moisture return north of the forecast axis was a key missed signal
- Strengthening low-level jet increased convergence and rainfall efficiency
- Outlier guidance should not be accepted blindly but should trigger monitoring
- A flood watch or escalation would have been appropriate once observations changed
Scoring notes in the case file
- Strong answers should diagnose process failure, not just say totals were too low.
- Credit answers that identify real-time update triggers.
- Penalize answers that rely on hindsight without explaining forecast-time evidence.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 83.4, median 84, range 76–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 76 | 83 | 84 | 85 | 87 | 78 | 87 |
FB-0025 forecast bust missed cap break
In plain terms. A severe forecast busted low when isolated supercells formed despite an expected cap. The task is to diagnose why the cap broke.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, post-event analysis
What the model received
- original_forecast: The forecast called for no surface-based storms because CIN was expected to remain below -100 J/kg through sunset and large-scale ascent was weak. Source label: Mock author-provided packet
- observed_outcome: Two supercells formed near 22z along a dryline/outflow intersection. They produced giant hail and one tornado. Later analysis showed surface temperatures locally reached 92 F, dewpoints pooled to 72 F, and the outflow boundary enhanced convergence. Source label: Mock author-provided packet
- missed_signals: Visible satellite showed persistent towering cumulus for 90 minutes before initiation. Mesonet observations showed a narrow corridor of backed winds and a 4 F dewpoint increase. One CAM run initiated storms in the same corridor, but it was dismissed as an outlier. Source label: Mock author-provided packet
What it had to do
Prepare one forecast-bust analysis package. Include what the corrected forecast should have been, top uncertainties, observations that should have changed the forecast, ignored guidance, confidence, and whether any watch would have been appropriate.
What the case says a strong answer should cover
- Localized boundary processes can break a broader cap
- Towering cumulus persistence was a key warning sign
- Moisture pooling and backed winds reduced inhibition/increased convergence
- Outlier CAM should have triggered monitoring
- Conditional watch may have become appropriate before initiation
Scoring notes in the case file
- Strong answers should diagnose local cap erosion, not just say cap was weaker.
- Credit answers that identify real-time update triggers.
- Penalize pure hindsight without forecast-time observations.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 84.6, median 85, range 82–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 85 | 84 | 84 | 85 | 85 | 82 | 87 |
FB-0026 forecast bust overforecast snow
In plain terms. A snowstorm forecast busted high. The task is to diagnose why expected heavy snow failed to occur.
Case record. cases/v0.1 · v0.1 · winter_weather · forecast time: Mock case, post-event analysis
What the model received
- original_forecast: The forecast called for 8-12 inches of snow along a deformation band. Forecasters expected strong frontogenesis and saturation through the dendritic growth zone. Source label: Mock author-provided packet
- observed_outcome: Only 2-4 inches fell. The heaviest band set up 80 miles north. In the forecast area, radar showed echoes but surface reports often described light snow or drizzle. Temperatures stayed 33-34 F for much of the event. Source label: Mock author-provided packet
- missed_signals: Aircraft soundings showed a dry slot arriving earlier than modeled. The DGZ became unsaturated for several hours, and lift maximized below the DGZ. Mesoscale guidance that shifted the band north was discounted because previous runs were too far north. Source label: Mock author-provided packet
What it had to do
Prepare one forecast-bust analysis package. Include the corrected forecast, top uncertainties, observations that should have changed the forecast, ignored guidance, confidence, and whether any watch would have been appropriate.
What the case says a strong answer should cover
- Dry slot and unsaturated DGZ reduced snow efficiency
- Lift below the DGZ favors poor crystal growth/drizzle
- Band placement error caused high bust
- Marginal surface temperatures lowered accumulation efficiency
- Watch/warning decisions should be revised with band shift
Scoring notes in the case file
- Strong answers should diagnose microphysics and band placement.
- Credit answers that mention downgrading headlines.
- Penalize saying QPF was simply too low without explaining snow efficiency.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 79.2, median 80, range 73–83. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 76 | 73 | 77 | 82 | 82 | 83 | 82 | 79 |
FB-0029 expert anonymized multiscale event
In plain terms. An anonymized high-impact severe-weather case has favorable synoptic support but messy mesoscale details. The model must forecast without event recognition.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 12z-06z
What the model received
- synoptic: A negatively tilted trough ejects into the Plains with a 120-kt upper jet and rapid surface cyclogenesis. A warm front lifts north through the forecast area while a dryline sharpens west of the warm sector. Source label: Mock author-provided packet
- mesoscale: Morning storms leave multiple outflow boundaries. By afternoon, the southern warm sector recovers to 2500 J/kg MLCAPE with dewpoints near 70 F. Effective shear is 60 kt. Low-level shear becomes extreme after 00z, but storm mode may become messy as a cold front accelerates east. Source label: Mock author-provided packet
- guidance: Some CAMs show discrete supercells from 21z-00z. Others quickly merge storms into a line. The highest tornado parameters overlap the warm front and outflow boundaries, not the entire warm sector. Source label: Mock author-provided packet
What it had to do
Prepare one expert forecast package. Include hazard forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Synoptic support is strong but mesoscale boundaries focus greatest risk
- Discrete storm window controls tornado ceiling
- QLCS transition changes hazard profile
- Warm front/outflow intersections are key
- Watch likely needed, but placement/timing matters
Scoring notes in the case file
- Strong answers should identify targeted corridor, not broad-brush all areas.
- Credit answers that handle storm-mode evolution.
- Penalize famous-outbreak style overconfidence.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.8, median 85, range 84–90. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 87 | 85 | 85 | 84 | 90 | 85 | 85 |
FB-0030 expert multi hazard prioritization
In plain terms. Multiple hazards are possible in one forecast area. The task is to prioritize operational decisions rather than list every hazard equally.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 18z-12z
What the model received
- hazards: A deepening cyclone will bring severe thunderstorms south, heavy snow northwest, freezing rain near the transition zone, and high winds across the entire region. The forecast office can issue only one lead briefing emphasis for emergency managers. Source label: Mock author-provided packet
- environment: South: MLCAPE 1500 J/kg, effective shear 55 kt, but storm mode likely linear. Northwest: deformation snow may produce 6-10 inches if the low tracks far enough south. Transition zone: 0.10-0.25 inches of ice is possible near the evening commute. Regionwide: gradient winds gust 45-55 mph overnight. Source label: Mock author-provided packet
- guidance: Model spread is largest in the low track. A north track favors severe storms and wind. A south track favors snow/ice. Ensembles show the highest probability of warning-level impacts in the ice corridor and wind field, while severe probabilities are more conditional. Source label: Mock author-provided packet
What it had to do
Prepare one coordination forecast package. Include prioritized hazard forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Operational prioritization matters in multi-hazard events
- Ice/wind may deserve lead emphasis if probabilities are highest
- Severe threat is conditional and storm-mode limited
- Low track controls snow/ice/severe distribution
- Multiple watch types may be considered, but timing/area should be prioritized
Scoring notes in the case file
- Strong answers should rank hazards, not merely list them.
- Credit answers that explain briefing emphasis.
- Penalize choosing the flashiest hazard without probability/impact reasoning.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 86.1, median 86, range 85–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 86 | 86 | 85 | 86 | 86 | 87 | 87 |
FB-1011 dryline initiation
In plain terms. A dryline is sharpening across the southern Plains during peak heating. The warm sector is moist and unstable, but a capping inversion and modest boundary position errors make convective initiation highly conditional.
Case record. cases/v0.2 · v0.2 · severe_weather · forecast time: Synthetic case issue time: 2026-05-21 18:00 UTC
What the model received
- surface_observations: East of the dryline, temperatures are 91 F with dewpoints of 70-72 F. The dryline has advanced east 20 km since 15z and surface convergence is increasing. West of the boundary, dewpoints are near 35 F. Source label: Synthetic author packet
- sounding_and_guidance: MLCAPE is 2800 J/kg, 0-1 km shear is 18 kt, and 0-6 km bulk shear is 42 kt. CIN is estimated near 50 J/kg at 18z. Guidance differs by 60 km on dryline position and by three hours on whether storms initiate before 23z. A weak short-wave impulse crosses the boundary near 21z. Source label: Synthetic author packet
What it had to do
Forecast the probability that at least one sustained deep convective storm initiates in the target area by 23z. Explain how dryline position, cap erosion, and mesoscale boundaries affect the forecast. Identify the observation that would most increase or decrease the probability. End with exactly one fenced JSON block (```json ... ```) using the documented forecasts format, with target_id convective_initiation_by_23z and a numeric probability.
What the case says a strong answer should cover
- Separate thermodynamic support from the conditional initiation problem
- Recognize dryline position and cap strength as coupled uncertainties
- Use boundary or radar observations as an updating signal
Scoring notes in the case file
- The target is initiation by 23z, not the probability of severe hazards after initiation.
- Discussion of why storms may fail to initiate is scored separately from the probability.
Structured target(s). convective_initiation_by_23z (binary, probability)
Across the eight runs. Mean 83.4, median 84, range 76–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 82 | 84 | 87 | 81 | 76 | 87 | 86 | 84 |
FB-1012 mcs cold pool propagation
In plain terms. A mature nocturnal mesoscale convective system is moving toward a moist unstable corridor. The main forecast question is whether cold-pool-driven propagation will sustain organized convection into the downstream area.
Case record. cases/v0.2 · v0.2 · severe_weather · forecast time: Synthetic case issue time: 2026-06-08 02:00 UTC
What the model received
- radar_and_surface_observations: Radar shows a linear convective system with a trailing stratiform region. A 30-kt cold-pool-relative inflow is measured along the gust front, while surface pressure rises 3 mb behind the line. New cells are repeatedly forming on the southeast flank. Source label: Synthetic author packet
- sounding_and_guidance: Downstream MLCAPE is 1800 J/kg with 0-3 km CAPE near 75 J/kg. A 35-kt low-level jet is nearly perpendicular to the line. Guidance differs on whether the cold pool will outrun the instability corridor after 06z; one solution weakens the system while another maintains a bowing line. Source label: Synthetic author packet
What it had to do
Forecast the probability that organized deep convection persists into the downstream corridor through 08z. Discuss the roles of cold-pool-relative flow, new-cell formation, low-level jet orientation, and downstream instability. State the observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) using the documented forecasts format, with target_id organized_convection_persistence and a numeric probability.
What the case says a strong answer should cover
- Distinguish advection of existing cells from propagation by new cells
- Use cold-pool-relative flow and low-level-jet orientation
- Assess whether downstream instability can support continued organization
Scoring notes in the case file
- The target is persistence into the downstream corridor, not maximum wind magnitude.
- Propagation reasoning is scored separately from the probability.
Structured target(s). organized_convection_persistence (binary, probability)
Across the eight runs. Mean 85.8, median 86, range 81–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 86 | 87 | 85 | 87 | 87 | 87 | 86 |
FB-1101 atmospheric river landfall
In plain terms. A strong west-to-east moisture corridor is approaching a mountainous coast. The event is likely, but the timing, landfall segment, and inland moisture transport depend on a short-wave phase difference.
Case record. cases/v0.3-development · v0.3 · marine · forecast time: Synthetic case issue time: 2026-01-18 12:00 UTC
What the model received
- integrated_vapor_transport: IVT is 650 kg m-1 s-1 just offshore and oriented 255 degrees. The corridor is 250 km wide, with the axis forecast to shift north or south by 120 km between 18z and 06z. Source label: Synthetic author packet
- ensemble_guidance: Twelve of twenty ensemble members cross the target coastal segment between 00z and 12z. Six members are six hours slower and four shift the axis south of the target. Orographic precipitation is stronger in the slower members. Source label: Synthetic author packet
What it had to do
Forecast the probability that the IVT corridor makes landfall in the target coastal segment during 00z-12z. Separate event occurrence from timing and location uncertainty. Identify the observation that would most change the forecast. End with exactly one fenced JSON block using target_id atmospheric_river_landfall_00_12z and a numeric probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Separate corridor existence, landfall timing, and landfall segment
- Use IVT direction and short-wave phase as structural uncertainty
- Do not equate ensemble spread with calibrated probability
Scoring notes in the case file
- This is a development case; the synthetic outcome is not a promotion-quality truth artifact.
- A high event probability with a wide location distribution can be coherent.
Structured target(s). atmospheric_river_landfall_00_12z (binary, probability)
Across the eight runs. Mean 85.8, median 86, range 84–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 84 | 85 | 86 | 85 | 85 |
FB-1103 convective value of information
In plain terms. A warm sector has substantial instability and shear, but convection may be suppressed until a boundary or short-wave impulse erodes the cap. The operational decision is which new observation would most reduce uncertainty.
Case record. cases/v0.3-development · v0.3 · severe_weather · forecast time: Synthetic case issue time: 2026-05-26 18:00 UTC
What the model received
- sounding_and_surface_observations: MLCAPE is 2400 J/kg, CIN is 75 J/kg, 0-6 km shear is 45 kt, and the 850-hPa dewpoint is 14 C. A dryline is 70 km west of the target, while a weak outflow boundary is drifting north from earlier storms. Source label: Synthetic author packet
- ensemble_scenarios: Half the guidance initiates near the dryline by 22z and half keeps the cap intact until after dark. Candidate observations are a special sounding, a profiler, a mesonet transect, a radar boundary scan, or a geostationary satellite trend. Source label: Synthetic author packet
What it had to do
Forecast the probability of sustained convective initiation in the target area by 23z. Rank the five candidate observations by expected decision value, explaining which uncertainty each resolves. End with exactly one fenced JSON block using target_id initiation_by_23z and a numeric probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Separate a favorable environment from the trigger and cap uncertainty
- Choose observations that resolve the controlling uncertainty, not merely add data
- Recognize boundary position and low-level moisture depth as coupled variables
Scoring notes in the case file
- The value-of-information ranking is narrative-only until a formal decision-loss rubric is added.
- The probability target is initiation, not subsequent severe-hazard intensity.
Structured target(s). initiation_by_23z (binary, probability)
Across the eight runs. Mean 84.0, median 85, range 81–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 85 | 87 | 81 | 85 | 85 | 87 | 81 |
FB-1119 forecast the probability of target segment landfall during 00z 12z sep audit flag
In plain terms. A strong atmospheric-river moisture corridor is approaching a mountainous coast, with substantial uncertainty in landfall timing and the coastal segment affected.
Case record. cases/drafts · v0.4 · general · forecast time: Synthetic case issue time: 2026-01-18 12:00 UTC
What the model received
- ensemble_guidance: Twelve of twenty ensemble members cross the target segment between 00z and 12z; six are six hours slower and four shift south. Source label: Draft author packet
- moisture_corridor: IVT is 650 kg m-1 s-1 offshore and orographic precipitation is stronger in slower members. Source label: Draft author packet
- discriminator: A dropsonde and coastal radar update are the most useful observations for separating timing and placement. Source label: Draft author packet
What it had to do
Forecast the probability of target-segment landfall during 00z-12z. Separate event occurrence from timing and location uncertainty, identify the most valuable observation, and describe a bust scenario. End with exactly one fenced JSON block containing target_id atmospheric_river_landfall_00_12z as a scalar probability and target_id atmospheric_river_timing_00_12z as an ordered probability list. The timing list must use this exact order: early_00_06z, late_06_12z, outside_window, and its values must sum to 1.0.
What the case says a strong answer should cover
- atmospheric river landfall timing and location uncertainty
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until the synthetic outcome is replaced or independently archived.
- Reward separating event probability from timing and placement probability.
- Penalize treating the 12-of-20 ensemble count as a calibrated probability, ignoring underdispersion, or claiming that an observation is useful without saying what forecast branch it would discriminate.
Structured target(s). atmospheric_river_landfall_00_12z (binary, probability), atmospheric_river_timing_00_12z (ordered_categorical, probability)
Across the eight runs. Mean 84.9, median 85, range 81–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 85 | 85 | 86 | 85 | 81 | 85 | 87 |
FB-1120 atmospheric river landfall audit flag
In plain terms. A strong atmospheric-river moisture corridor is approaching a mountainous coast with meaningful timing and placement uncertainty.
Case record. cases/drafts · v0.4 · general · forecast time: TODO: YYYY-MM-DD HH:MM UTC
What the model received
- moisture_corridor: Provide IVT magnitude, orientation, corridor width, and terrain relationship. Source label: Draft author packet
- ensemble_guidance: Provide timing and placement scenarios without presenting ensemble membership as calibrated probability. Source label: Draft author packet
- verification_discriminator: Provide a dropsonde, radar, or satellite observation that separates the leading scenarios. Source label: Draft author packet
What it had to do
Forecast the probability of target-segment landfall during the valid window. Separate event occurrence from timing and location uncertainty, identify the most valuable observation, and describe a bust scenario.
What the case says a strong answer should cover
- integrated vapor transport and moisture-corridor structure
- ensemble timing and landfall-location spread
- orographic precipitation sensitivity
- underdispersion and conditional scenario evidence
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- TODO: define the failure modes and what the grader should reward or penalize.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 69.4, median 85, range 0–85. Zero scores: #252.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 57 | 0 | 85 | 85 | 85 | 73 | 85 | 85 |
FB-1121 tropical cyclone rapid intensification audit flag
In plain terms. A developing tropical cyclone may undergo rapid intensification while track and environmental uncertainty remain material.
Case record. cases/drafts · v0.4 · tropical_weather · forecast time: Synthetic case issue time: 2026-08-01 12:00 UTC
What the model received
- storm_structure: Provide generalized intensity, inner-core organization, convection, and recent intensity trend. Source label: Draft author packet
- environment: Provide vertical shear, sea-surface temperature, ocean heat content, moisture, and outflow context. Source label: Draft author packet
- guidance_spread: Provide track and intensity scenario spread without presenting guidance as calibrated truth. Source label: Draft author packet
What it had to do
Forecast the probability of rapid intensification during the valid window. Return an ordered probability vector with exactly these categories: no_RI, RI_20kt_or_more, RI_30kt_or_more. The probabilities must be numeric, non-negative, and sum to 1.0. Give a separate maximum-wind interval, identify the dominant uncertainty source, and state what observation would most change the forecast.
What the case says a strong answer should cover
- initial-intensity and structural uncertainty
- vertical shear, ocean heat content, and environmental coupling
- track-dependent intensity error
- threshold-event calibration and representativeness limits
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- Reward separating the 20-kt threshold probability from the 30-kt tail and identifying the environmental or structural uncertainty that could change the forecast.
- Penalize treating the scenario spread as calibrated probability, collapsing the two rapid-intensification thresholds, or presenting maximum wind as exact truth.
Structured target(s). tropical_cyclone_rapid_intensification_24h (ordered_categorical, probability)
Across the eight runs. Mean 73.5, median 79, range 37–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 69 | 76 | 85 | 82 | 87 | 65 | 87 | 37 |
FB-1122 mjo phase amplitude audit flag
In plain terms. A Madden–Julian Oscillation signal is evolving across the Indian Ocean toward the Maritime Continent with uncertain amplitude and propagation speed.
Case record. cases/drafts · v0.4 · general · forecast time: Synthetic case issue time: 2026-09-01 12:00 UTC
What the model received
- convection_signal: Provide generalized outgoing-longwave-radiation or precipitation-anomaly evolution by basin. Source label: Draft author packet
- dynamical_context: Provide low-level westerly/easterly wind anomalies, moisture structure, and background state. Source label: Draft author packet
- ensemble_evolution: Provide phase and amplitude spread at days 7, 14, and 21 without presenting it as calibrated probability. Source label: Draft author packet
What it had to do
Forecast the MJO phase and amplitude category at days 7, 14, and 21. Separate phase uncertainty from amplitude damping and propagation uncertainty. Return an ordered probability vector for the day-14 phase outcome with exactly these categories: western_indian_ocean, maritime_continent, western_pacific, weak_or_uncertain. The probabilities must be numeric, non-negative, and sum to 1.0. State what evidence would most change the forecast and when no useful signal is justified.
What the case says a strong answer should cover
- MJO phase-space propagation
- amplitude damping and re-emergence
- Maritime Continent propagation barrier
- longer-horizon probabilistic calibration
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- Reward separating phase propagation from amplitude damping and recognizing when the Maritime Continent creates a predictable barrier or a loss of signal.
- Penalize treating day-21 phase as deterministic, confusing amplitude with phase, or claiming a useful signal without packet evidence.
Structured target(s). mjo_day14_phase (ordered_categorical, probability)
Across the eight runs. Mean 75.8, median 76, range 66–84. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 76 | 76 | 81 | 66 | 78 | 84 | 68 | 77 |
intermediate_forecasting
1 cases, 8 model-case scores
FB-0002 intermediate convective initiation
In plain terms. A dryline is sharpening during the afternoon across the southern Plains. Instability is large, but inhibition remains a concern. The main forecast question is whether surface-based storms initiate before sunset.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 18z-03z
What the model received
- surface_observations: At 19z, the dryline is analyzed from northwest to southwest across the forecast area. East of the dryline, temperatures are 82-86 F with dewpoints 64-67 F. West of the dryline, temperatures are in the lower 90s with dewpoints in the 30s. Surface winds east of the dryline are south-southeast at 15 kt. Source label: Mock author-provided packet
- sounding: A modified 18z sounding east of the dryline shows MLCAPE near 2500 J/kg, 700-500 mb lapse rates near 8 C/km, and MLCIN around -125 J/kg. The capping inversion is centered near 800 mb. The LFC remains high unless surface temperatures reach at least 89 F. Source label: Mock author-provided packet
- guidance: One convection-allowing model initiates isolated storms near 22z along the dryline. Two other models remain capped until after sunset. A weak mid-level impulse approaches from the west by 00z, but strongest ascent remains north of the forecast area. Source label: Mock author-provided packet
What it had to do
Forecast whether surface-based thunderstorms are likely before 00z. Explain the main limiting factor, what observation would increase confidence in initiation, and the most likely bust scenario.
What the case says a strong answer should cover
- Large CAPE does not guarantee initiation when CIN remains substantial
- Dryline convergence and surface heating are key to eroding the cap
- High LFC and warm layer near 800 mb limit confidence
- Model disagreement should lower confidence
- A conditional severe threat exists if storms form
Scoring notes in the case file
- Strong answers should describe initiation as conditional or uncertain rather than guaranteed.
- Credit answers that name specific observations, such as deeper cumulus, stronger convergence, or temperatures reaching the convective temperature.
- Penalize answers that equate high CAPE with widespread storms.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 83.5, median 84, range 81–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 81 | 82 | 84 | 84 | 84 | 85 | 84 |
operational_forecasting
20 cases, 160 model-case scores
FB-0003 operational winter ptype
In plain terms. A winter storm is approaching a metro area near a strong low-level thermal gradient. The forecast problem is precipitation type and accumulation.
Case record. cases/v0.1 · v0.1 · winter_weather · forecast time: Mock case, valid 12z-00z
What the model received
- surface_observations: At 12z, the metro area reports 31 F with a northeast wind at 12 kt. Dewpoint is 24 F. A station 40 miles southeast reports 35 F with light rain. A station 50 miles northwest reports 27 F with snow. Road temperatures are near freezing. Source label: Mock author-provided packet
- sounding: The 12z sounding shows a subfreezing surface layer from the ground to 925 mb, a warm nose peaking near +2 C between 850 and 800 mb, and saturated conditions through the dendritic growth zone. Wet-bulb temperature in the lowest 1 km remains below 0 C. Source label: Mock author-provided packet
- guidance: Model A keeps the metro mostly snow with 5 inches of accumulation. Model B brings the warm nose farther northwest and changes precipitation to sleet and freezing rain for 4 hours, with 1 inch of snow and 0.15 inches of ice. Model C is slightly warmer at the surface and changes the southeast half of the metro to rain by 18z. Source label: Mock author-provided packet
- trend: Over the past 3 hours, surface pressure falls have increased southeast of the metro. Northeast winds have strengthened, and temperatures in the northwest half of the metro have fallen 1-2 F. Source label: Mock author-provided packet
What it had to do
Produce a precipitation-type forecast for the metro area through 00z. Identify the most likely zone of highest impact, the biggest forecast uncertainty, and one plausible bust scenario in each direction.
What the case says a strong answer should cover
- Warm nose supports sleet or freezing rain if deep enough
- Wet-bulb cooling and low-level cold air can resist surface warming
- Small thermal shifts can create large accumulation differences
- Metro-scale gradients should be communicated spatially
- Both colder and warmer bust scenarios are plausible
Scoring notes in the case file
- Strong answers should avoid a single deterministic precipitation type for the whole metro.
- Credit answers that split northwest versus southeast metro impacts.
- Penalize answers that ignore the warm nose or treat surface temperature alone as decisive.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.2, median 85, range 84–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 85 | 85 | 85 | 85 | 84 | 87 | 86 |
FB-0008 synoptic cyclogenesis
In plain terms. Medium-range guidance disagrees on the evolution of a deepening mid-latitude cyclone. The forecast problem is choosing the most believable synoptic solution and communicating impacts.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 00z-48h
What the model received
- upper_air: A 500-mb trough is digging into the southern Rockies with a 110-kt jet streak rounding the base. Downstream height falls are spreading into the Plains. A northern-stream shortwave is moving across the northern High Plains and may phase with the southern wave within 24-36 hours. Source label: Mock author-provided packet
- surface: A lee cyclone is developing with a 1004-mb surface low. Gulf moisture is returning northward, with 60 F dewpoints reaching the lower Mississippi Valley. A strong baroclinic zone extends northeast of the low. Source label: Mock author-provided packet
- guidance: Model A phases the streams early and deepens the low to 988 mb, bringing heavy precipitation and strong winds farther northwest. Model B keeps the waves separate longer and tracks a weaker 998-mb low farther southeast. Ensembles favor a middle solution but have large spread in phasing timing. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package for regional coordination. Include the preferred synoptic solution, sensible-weather impacts, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Phasing timing controls cyclone depth and track
- Jet streak dynamics support cyclogenesis
- Baroclinic zone favors precipitation and wind impacts
- Ensemble spread argues against locking onto one deterministic extreme
- Observed shortwave timing should change confidence
Scoring notes in the case file
- Strong answers should choose a solution while preserving uncertainty.
- Credit answers that explain why deterministic extremes are discounted.
- Watch discussion may say none yet or mention likely future winter/wind/severe watches depending on impacts.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.0, median 85, range 82–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 88 | 85 | 85 | 84 | 86 | 82 | 85 |
FB-0009 forecast update radar trends
In plain terms. A morning forecast called for scattered severe storms by late afternoon, but observations are evolving differently than expected. The forecast problem is issuing an update based on new radar, satellite, and surface trends.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 20z update
What the model received
- previous_forecast: The 12z forecast expected storm initiation by 20z along a dryline, with scattered supercells moving east into a moist unstable warm sector. Source label: Mock author-provided packet
- current_observations: At 20z, visible satellite shows a broad field of flat cumulus along the dryline but no sustained towers. Surface temperatures east of the dryline are 84-87 F, dewpoints 65-68 F, and pressure falls are weak. Winds have veered more southwest than expected across the warm sector. Source label: Mock author-provided packet
- radar_and_model_trends: Radar shows only shallow echoes west of the dryline. The latest convection-allowing guidance delays initiation until 23z-01z and shifts the strongest storms north of the original outlook area. Mesoanalysis still shows 2500 J/kg MLCAPE but CIN remains near -100 J/kg. Source label: Mock author-provided packet
What it had to do
Prepare one forecast update package. Include what changes from the previous forecast, top uncertainties, observations that would change the forecast, model guidance you are ignoring, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Observed convective failure should reduce confidence in earlier initiation
- Veered surface winds may reduce convergence and low-level shear
- CIN remains an important limiting factor despite CAPE
- Delayed/north-shifted CAM guidance should be weighed against observations
- Watch decision should consider timing and spatial confidence
Scoring notes in the case file
- Strong answers should explicitly update or downgrade the previous forecast.
- Credit answers that identify observations that would restore severe confidence.
- Penalize answers that repeat the old forecast without using 20z trends.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 83.2, median 84, range 79–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 79 | 82 | 84 | 84 | 84 | 84 | 85 |
FB-0011 severe morning convection recovery
In plain terms. Morning convection has disrupted the warm sector ahead of an afternoon severe setup. The forecast problem is whether recovery is sufficient for renewed severe storms.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 15z-03z
What the model received
- morning_observations: At 15z, a decaying MCS is exiting the eastern forecast area. Temperatures behind the rain shield are 64-68 F with widespread low clouds. Farther southwest, skies are clearing with temperatures rising into the mid 70s and dewpoints near 68 F. An outflow boundary arcs west-east through the central part of the forecast area. Source label: Mock author-provided packet
- environment: Forecast soundings recover to 1500-2500 J/kg MLCAPE where clearing persists for at least 3 hours. Effective shear is 45 kt. Low-level shear is strongest near the outflow boundary, but CIN remains near -75 J/kg beneath the cloud shield. Source label: Mock author-provided packet
- guidance: One CAM rapidly redevelops supercells along the outflow boundary by 20z. Another keeps the boundary rain-cooled and delays storms until evening along the cold front. A third initiates storms only in the clearing southwest counties. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include the severe-weather forecast, top uncertainties, observations that would change the forecast, model guidance you would ignore, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Morning convection can reduce instability and delay recovery
- Outflow boundary may enhance low-level shear if destabilization occurs
- Clearing location controls renewed severe risk
- CAMs should be weighed against real-time cloud/temperature trends
- Watch decision should be conditional and spatially focused
Scoring notes in the case file
- Strong answers should not blindly keep the original severe forecast.
- Credit answers that separate southwest recovery from rain-cooled areas.
- Penalize answers that ignore outflow-boundary potential.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.4, median 85, range 84–89. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 85 | 85 | 85 | 86 | 89 | 85 | 84 |
FB-0022 synoptic model phase error
In plain terms. Recent model runs have trended sharply toward a phased storm, but upstream observations suggest one wave may be slower than initialized. The forecast problem is whether to follow the trend.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 00z-72h
What the model received
- model_trends: The last three deterministic runs deepened a surface cyclone from 1000 mb to 984 mb in the day-3 forecast, shifting heavy precipitation 150 miles northwest. Ensemble means also deepened, but many members remain weaker and southeast. Source label: Mock author-provided packet
- observations: Aircraft and satellite winds upstream show the northern-stream shortwave 6 hours slower and slightly weaker than initialized in the latest deterministic run. The southern wave is well sampled and close to guidance. Source label: Mock author-provided packet
- impacts: The phased solution would produce a major snow/wind event. The unphased solution would produce lighter precipitation mostly southeast of the forecast area. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include preferred synoptic scenario, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Model trend should be checked against upstream observations
- Slower/weaker northern wave reduces early phasing confidence
- Ensemble spread argues for tempered forecast
- High-impact phased solution must remain in risk messaging
- Watch decision depends on confidence and lead time
Scoring notes in the case file
- Strong answers should resist blindly following deterministic trend.
- Credit answers that keep high-impact scenario as plausible but uncertain.
- Penalize ignoring the observed initialization error.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.8, median 87, range 81–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 88 | 84 | 87 | 84 | 88 | 87 | 81 |
FB-0023 forecast update downgrade scary
In plain terms. A high-impact severe forecast was issued in the morning, but afternoon observations are trending less favorable. The forecast problem is whether to downgrade despite scary earlier guidance.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 19z update
What the model received
- morning_forecast: The morning package highlighted numerous supercells and possible tornadoes by 20z in the warm sector. Source label: Mock author-provided packet
- current_observations: At 19z, widespread stratus persists with temperatures only 70-73 F. Dewpoints are still 66-68 F, but objective analysis shows MLCAPE only 800-1200 J/kg. Surface winds have veered southwest, and the boundary is diffuse. No deep cumulus is present. Source label: Mock author-provided packet
- guidance: Older CAMs still show robust 20z-22z storms. The newest runs delay initiation until 01z and weaken storm intensity. The strongest forcing arrives after the low-level jet veers. Source label: Mock author-provided packet
What it had to do
Prepare one forecast update package. Include what changes from the morning forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Observed lack of destabilization should downgrade threat
- Veered winds and diffuse boundary reduce tornado confidence
- Old CAMs should be discounted versus current observations
- Delayed initiation changes hazard and watch timing
- Communication should acknowledge forecast change clearly
Scoring notes in the case file
- Strong answers should explicitly downgrade the morning forecast.
- Credit answers that keep conditional monitoring instead of canceling all risk.
- Penalize clinging to old high-end CAMs.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.6, median 86, range 85–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 86 | 86 | 85 | 86 | 85 | 85 | 86 |
FB-0024 forecast update rapid upgrade
In plain terms. A low-confidence severe forecast is rapidly becoming more concerning based on observations. The forecast problem is whether to upgrade quickly.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 20z emergency update
What the model received
- previous_forecast: The 15z forecast kept severe probabilities low because CAMs showed poor initiation and weak forcing. Source label: Mock author-provided packet
- observations: By 20z, surface temperatures have reached 88-91 F with dewpoints 69-72 F. The dryline has sharpened, pressure falls increased to 3 mb in 3 hours, and visible satellite shows agitated towering cumulus in three clusters. A special sounding shows CIN has eroded to near zero with MLCAPE above 3000 J/kg. Source label: Mock author-provided packet
- wind_profile: Effective shear is 50 kt. Low-level winds are backed southeast near an outflow boundary, giving 0-1 km SRH near 200 m2/s2 in the southern cluster. Source label: Mock author-provided packet
What it had to do
Prepare one urgent forecast update package. Include what changes from the previous forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Observations can override earlier low-CAM initiation signal
- Eroded CIN and towering cumulus sharply increase initiation confidence
- Outflow/backed winds enhance tornado potential locally
- Watch upgrade may be needed quickly
- Uncertainty remains in which cluster dominates
Scoring notes in the case file
- Strong answers should clearly upgrade the forecast.
- Credit answers that discuss urgent watch coordination.
- Penalize waiting on old CAM guidance despite observed initiation signs.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 89.4, median 90, range 87–92. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 91 | 91 | 92 | 87 | 88 | 87 | 90 | 89 |
FB-0027 data conflict obs vs model sounding
In plain terms. Observations disagree with model soundings. The forecast problem is whether to trust the model profile or the real-time data.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 18z-03z
What the model received
- model_profile: Model soundings show a moist boundary layer, weak CIN, and scattered thunderstorms developing after 21z. The model has surface temperatures 86-88 F and dewpoints 66-68 F. Source label: Mock author-provided packet
- observations: At 18z, actual surface observations are 80-82 F with dewpoints 59-62 F because a dry pocket mixed into the warm sector. Aircraft soundings show a deeper mixed layer and a dry layer from 850-700 mb not present in the model. Satellite shows flat cumulus and no vertical growth. Source label: Mock author-provided packet
- guidance: Two CAMs continue to initiate storms. A rapid-refresh run initialized with the dry pocket delays convection until after 02z and keeps coverage isolated. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include convective forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Real observations should override poorly initialized model soundings
- Dry pocket reduces moisture/instability and raises cloud bases
- Dry mid-level layer may enhance evaporative cooling but suppress initiation
- Flat cumulus supports lower initiation confidence
- Watch should be delayed or withheld without better destabilization
Scoring notes in the case file
- Strong answers should explicitly identify model initialization error.
- Credit answers that use aircraft soundings.
- Penalize following CAM storms without reconciling observed dry pocket.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 82.6, median 82, range 81–84. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 82 | 82 | 84 | 84 | 84 | 81 | 83 |
FB-0028 public risk communication uncertainty
In plain terms. A potentially high-impact weather event has large uncertainty. The task is to communicate risk without over-warning or under-warning.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid public briefing
What the model received
- hazard_context: There is a 30 percent chance of a high-impact ice storm and a 50 percent chance of mostly cold rain. The remaining 20 percent is a moderate sleet event. The highest-impact scenario would produce 0.25-0.50 inches of ice during the morning commute. Source label: Mock author-provided packet
- forecast_drivers: The event depends on whether a shallow cold layer remains in place. Cold air is draining south through valleys, but a warm nose aloft is already present. Small surface-temperature errors of 1-2 F change the impact dramatically. Source label: Mock author-provided packet
- guidance: Deterministic guidance has flipped between freezing rain and rain during the last four cycles. Ensembles remain split but show increasing probability of subfreezing valley temperatures. Source label: Mock author-provided packet
What it had to do
Prepare one public-facing forecast package. Include forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Low-probability high-impact scenario must be communicated
- Avoid deterministic ice-storm language when rain remains more likely
- Surface temperatures and cold-air drainage are key
- Ensemble split should be communicated plainly
- Watch may be appropriate if impact threshold risk is sufficient
Scoring notes in the case file
- Strong answers should balance probability and impact.
- Credit clear public messaging.
- Penalize either alarmism or dismissing the ice risk.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.4, median 86, range 84–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 87 | 87 | 86 | 84 | 84 | 86 | 84 |
FB-1002 winter ptype probability
In plain terms. A marginal winter storm is crossing a metro area near a sharp thermal gradient. The precipitation type depends on the depth of the cold layer.
Case record. cases/v0.2 · v0.2 · winter_weather · forecast time: Synthetic case issue time: 2026-01-12 12:00 UTC
What the model received
- surface_observations: The metro is 31 F with a 24 F dewpoint and northeast wind at 12 kt. A station southeast of the metro reports light rain at 35 F. A station northwest reports snow at 27 F. Road temperatures are near freezing. Source label: Synthetic author packet
- sounding_and_guidance: The sounding has a subfreezing layer to 925 mb, a +2 C warm nose from 850-800 mb, and saturation through the dendritic growth zone. Guidance ranges from mostly snow to four hours of sleet and freezing rain. Source label: Synthetic author packet
What it had to do
Forecast the precipitation type through 00z. Provide probabilities in this ordered category list: snow, sleet, freezing rain, rain. Explain the zone of highest impact and the main uncertainty. End with exactly one fenced JSON block (```json ... ```) with target_id precipitation_type and a probability list in that order.
What the case says a strong answer should cover
- Use the low-level thermal profile
- Distinguish the metro gradient from a single-area answer
- State the main precipitation-type uncertainty
Scoring notes in the case file
- The category order is part of the target definition.
- Qualitative discussion and the probability vector are scored separately.
Structured target(s). precipitation_type (ordered_categorical, None)
Across the eight runs. Mean 86.1, median 86, range 85–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 88 | 85 | 85 | 87 | 87 | 87 | 85 | 85 |
FB-1004 wind advisory decision
In plain terms. A tightening pressure gradient is expected to produce a short period of hazardous wind over a coastal operating area.
Case record. cases/v0.2 · v0.2 · marine · forecast time: Synthetic case issue time: 2026-03-04 15:00 UTC
What the model received
- observations: Buoys report sustained winds of 24 kt with gusts to 32 kt. The pressure gradient has increased for six hours and a stronger mixed layer is expected after 18z. Source label: Synthetic author packet
- guidance: Guidance clusters around sustained winds of 28-32 kt and gusts of 38-45 kt from 18z-02z, but differs on how quickly winds ease overnight. Source label: Synthetic author packet
What it had to do
Decide whether to issue a wind advisory for the operating area for 18z-02z. Give a probability from 0 to 1 that advisory-level conditions will occur, the operational consequence, and the observation that would most change the decision. End with exactly one fenced JSON block (```json ... ```) with target_id advisory_conditions and a numeric probability.
What the case says a strong answer should cover
- Use both current observations and the pressure-gradient trend
- Connect the forecast to an operational decision
- State what would invalidate the decision
Scoring notes in the case file
- The binary target is advisory-level conditions, not whether wording sounds cautious.
- The decision discussion is reviewed separately from the probability.
Structured target(s). advisory_conditions (binary, probability)
Across the eight runs. Mean 87.0, median 87, range 85–89. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 87 | 87 | 85 | 89 | 89 | 86 | 87 |
FB-1016 real asos temperature
In plain terms. A cold January morning at Des Moines International Airport is beginning to warm after sunrise. Forecast the airport temperature five hours after the issue time from the observations available at issue time.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-01-15 13:00 UTC
What the model received
- observations: At 12:54 UTC, DSM reported 7 F, dew point 0 F, wind 200 degrees at 7 kt, visibility 10 statute miles, and few clouds. At 13:00 UTC the station reported broken cloud cover, with wind 200 degrees at 4 kt. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1016-DSM-2025-01-15.csv
- context: The target valid time is 18:00 UTC. No numerical guidance is supplied; use the observed temperature trend, moisture, wind, and cloud evolution. Source label: Synthetic task framing around archived observations
What it had to do
Forecast the DSM temperature in degrees Fahrenheit at 18:00 UTC. Explain the physical basis for the warming or lack of warming, identify the observation that would most change the forecast, and distinguish uncertainty in the temperature trend from uncertainty in the exact observation time. End with exactly one fenced JSON block (```json ... ```) using target_id temperature_at_18z_f and a numeric value.
What the case says a strong answer should cover
- Use the observed temperature, dew point, wind, and cloud trend
- Recognize that a persistence forecast is a weak baseline during daytime warming
- Identify cloud evolution or a new temperature observation as an update signal
Scoring notes in the case file
- The objective target is the archived DSM ASOS temperature near 18:00 UTC.
- The source excerpt includes the issue-time observations and the private verification observation.
Structured target(s). temperature_at_18z_f (continuous, degrees Fahrenheit)
Across the eight runs. Mean 72.5, median 73, range 63–81. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 64 | 70 | 81 | 71 | 63 | 76 | 75 | 80 |
FB-1017 real asos ord temperature
In plain terms. A clear summer morning at Chicago O'Hare is warming rapidly. Forecast the airport temperature at 18z from the observations available at issue time.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-07-15 13:00 UTC
What the model received
- observations: At 12:51 UTC, ORD reported 78 F, dew point 66 F, wind 190 degrees at 3 kt, visibility 10 statute miles, and broken clouds. At 13:00 UTC the wind was 200 degrees at 4 kt, visibility 8 miles, and the lowest cloud layer was clear. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1017-ORD-2025-07-15.csv
- context: The target valid time is 18:00 UTC. No numerical guidance is supplied. Source label: Task framing around archived observations
What it had to do
Forecast the ORD temperature in degrees Fahrenheit at 18:00 UTC. Explain the role of solar heating, moisture, wind, and cloud cover. Identify the observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id temperature_at_18z_f and a numeric value.
What the case says a strong answer should cover
- Use the observed temperature trend and clear-to-broken cloud evolution
- Account for humid boundary-layer conditions rather than assuming unlimited warming
- Identify a new temperature or cloud observation as an update signal
Scoring notes in the case file
- Truth is the archived ORD ASOS temperature reported at 17:51 UTC.
Structured target(s). temperature_at_18z_f (continuous, degrees Fahrenheit)
Across the eight runs. Mean 71.6, median 73, range 59–75. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 75 | 70 | 75 | 59 | 73 | 73 | 73 | 75 |
FB-1018 real asos den temperature
In plain terms. A fall morning at Denver International Airport begins with cool, moist air and low wind. Forecast the airport temperature at 18z while accounting for later visibility and wind changes.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- observations: At 12:53 UTC, DEN reported 50 F, dew point 45 F, wind 3 kt, visibility 10 statute miles, and few clouds. At 13:00 UTC the wind was 290 degrees at 2 kt with visibility 10 miles and few clouds. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1018-DEN-2025-10-15.csv
- context: The target valid time is 18:00 UTC. No numerical guidance is supplied. Source label: Task framing around archived observations
What it had to do
Forecast the DEN temperature in degrees Fahrenheit at 18:00 UTC. Explain the role of the dry-air transition, wind changes, visibility, and cloud cover. Identify the observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id temperature_at_18z_f and a numeric value.
What the case says a strong answer should cover
- Use the observed temperature and dew-point spread
- Treat changing wind and visibility as evidence about boundary-layer mixing
- Separate temperature uncertainty from visibility uncertainty
Scoring notes in the case file
- Truth is the archived DEN ASOS temperature reported at 17:53 UTC.
Structured target(s). temperature_at_18z_f (continuous, degrees Fahrenheit)
Across the eight runs. Mean 69.1, median 69, range 63–75. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 75 | 70 | 66 | 63 | 68 | 73 | 68 | 70 |
FB-1019 real asos sea temperature
In plain terms. A damp winter morning at Seattle-Tacoma International Airport has low cloud cover and a small temperature range. Forecast the airport temperature at 18z.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-02-15 13:00 UTC
What the model received
- observations: At 12:53 UTC, SEA reported 37 F, dew point 34 F, wind 170 degrees at 4 kt, visibility 10 statute miles, and broken clouds. At 13:00 UTC the wind was 160 degrees at 2 kt, visibility 10 miles, and overcast. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1019-SEA-2025-02-15.csv
- context: The target valid time is 18:00 UTC. No numerical guidance is supplied. Source label: Task framing around archived observations
What it had to do
Forecast the SEA temperature in degrees Fahrenheit at 18:00 UTC. Explain how cloud cover, moisture, wind, and the small dew-point depression affect the temperature trend. Identify the observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id temperature_at_18z_f and a numeric value.
What the case says a strong answer should cover
- Use the small temperature-dew-point spread and persistent cloud cover
- Avoid importing continental daytime-heating assumptions
- Identify cloud or temperature observations as update signals
Scoring notes in the case file
- Truth is the archived SEA ASOS temperature reported at 17:53 UTC.
Structured target(s). temperature_at_18z_f (continuous, degrees Fahrenheit)
Across the eight runs. Mean 73.6, median 76, range 63–81. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 76 | 81 | 70 | 63 | 80 | 63 | 75 | 81 |
FB-1020 real asos phx temperature
In plain terms. A clear early-summer morning at Phoenix Sky Harbor is warming in a very dry air mass. Forecast the airport temperature at 18z.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-06-15 13:00 UTC
What the model received
- observations: At 12:51 UTC, PHX reported 82 F, dew point 37 F, wind 170 degrees at 3 kt, visibility 10 statute miles, and clear skies. At 13:00 UTC the wind was calm, visibility remained 10 miles, and skies remained clear. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1020-PHX-2025-06-15.csv
- context: The target valid time is 18:00 UTC. No numerical guidance is supplied. Source label: Task framing around archived observations
What it had to do
Forecast the PHX temperature in degrees Fahrenheit at 18:00 UTC. Explain how dry air, clear skies, wind, and solar heating affect the trend. Identify the observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id temperature_at_18z_f and a numeric value.
What the case says a strong answer should cover
- Use the large temperature-dew-point spread and clear skies
- Recognize strong daytime heating in a dry desert boundary layer
- Identify a new wind, cloud, or temperature observation as an update signal
Scoring notes in the case file
- Truth is the archived PHX ASOS temperature reported at 17:51 UTC.
Structured target(s). temperature_at_18z_f (continuous, degrees Fahrenheit)
Across the eight runs. Mean 77.2, median 80, range 66–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 86 | 80 | 66 | 68 | 73 | 80 | 84 |
FB-1104 ensemble reliability regime
In plain terms. Two ensemble forecast clusters disagree about a threshold event. The narrow cluster has a shared physics bias while the broader cluster contains the verifying solution, testing whether spread is being confused with calibrated uncertainty.
Case record. cases/v0.3-development · v0.3 · general · forecast time: Synthetic case issue time: 2026-04-11 00:00 UTC
What the model received
- ensemble_distribution: Twenty members forecast 24-hour precipitation in a 0.25-0.35 inch cluster. Ten members forecast 0.65-0.90 inch, with the second cluster associated with a slower upper trough and stronger moisture transport. The threshold of interest is 0.50 inch. Source label: Synthetic author packet
- regime_context: The current flow regime has historically produced a dry bias in the fast-cluster members. A neighboring observation network shows moisture rising faster than the fast-cluster analysis assumed. Source label: Synthetic author packet
What it had to do
Forecast the probability that 24-hour precipitation exceeds 0.50 inch. State whether the ensemble appears underdispersive, overdispersive, or merely uncertain, and explain how the regime information changes the probability. End with exactly one fenced JSON block using target_id precip_gt_050 and a numeric probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Distinguish ensemble spread from reliability and accuracy
- Use regime-conditioned bias information without treating it as certainty
- Recognize that a tight wrong cluster can be less informative than a broad calibrated one
Scoring notes in the case file
- This case is designed for reliability reasoning; the synthetic event record is provisional.
- A response should not award probability solely from member count without regime context.
Structured target(s). precip_gt_050 (binary, probability)
Across the eight runs. Mean 81.1, median 82, range 73–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 82 | 82 | 73 | 76 | 83 | 85 | 81 |
FB-1112 mountain wave turbulence audit flag
In plain terms. A stable cross-mountain flow may produce clear-air turbulence near a busy flight corridor.
Case record. cases/drafts · v0.3 · aviation · forecast time: TODO: YYYY-MM-DD HH:MM UTC
What the model received
- upper_air_profile: Provide stability, tropopause height, wind direction, and cross-barrier wind by level. Source label: Draft author packet
- terrain_and_route: Provide generalized ridge orientation, route altitude, and distance from the crest. Source label: Draft author packet
- guidance_spread: Provide wave-amplitude and turbulence guidance spread without revealing the answer. Source label: Draft author packet
What it had to do
Forecast the probability, altitude band, and timing of operationally significant turbulence. Separate mountain-wave evidence from uncertainty caused by unresolved wave amplitude.
What the case says a strong answer should cover
- stable stratification and mountain-wave setup
- cross-barrier wind and directional shear
- turbulence altitude-band uncertainty
- operational impact versus physical likelihood
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- TODO: define the failure modes and what the grader should reward or penalize.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 59.1, median 76, range 0–85. Zero scores: #252, #253.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 67 | 0 | 0 | 85 | 84 | 84 | 74 | 79 |
FB-1116 mountain wave turbulence
In plain terms. A strong, stable cross-mountain flow is producing a mountain-wave signal near a busy flight corridor. The wave is physically plausible, but the altitude of the strongest response and whether it becomes operationally significant depend on stability depth, directional shear, and unresolved wave amplitude.
Case record. cases/drafts · v0.3 · aviation · forecast time: Synthetic case issue time: 2026-02-04 18:00 UTC
What the model received
- upper_air_profile: The 700-500 hPa layer is stably stratified with a Brunt-Vaisala frequency near 0.012 s-1. Cross-barrier wind increases from 35 kt at 700 hPa to 55 kt at 500 hPa, with a 25-degree directional change. The tropopause is near 260 hPa and the strongest stability is centered near 600 hPa. Source label: Synthetic author packet
- terrain_and_route: A generalized ridge is oriented nearly perpendicular to the mean flow. The route crosses 40 km downwind of the crest at flight levels 180, 220, and 280. Guidance agrees on wave activity but disagrees on whether the strongest vertical-acceleration layer is near FL200 or FL260. Source label: Synthetic author packet
- guidance_spread: Sixteen ensemble members support moderate wave amplitude. Four members produce a stronger, vertically stacked wave near FL220, while four weaker members keep the response below operational thresholds. The ensemble spread is conditional evidence, not a calibrated probability. Source label: Synthetic author packet
- aircraft_and_profiler_observations: Nearby aircraft report occasional light turbulence below FL200 but no reports near FL240. A profiler shows the cross-barrier jet strengthening faster than guidance, while the mountain crest remains cloud-free. Source label: Synthetic author packet
What it had to do
Forecast the probability of operationally significant turbulence somewhere along the route during 18z-06z. Give a probability for each altitude band (FL180-210, FL220-250, and FL260-290), identify the most decision-relevant observation, and distinguish physical wave likelihood from route-level impact. End with exactly one fenced JSON block using target_id turbulence_operational_18_06z as a binary probability and target_id turbulence_peak_band_18_06z as an ordered probability list. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Stable stratification and cross-barrier flow support mountain-wave formation
- Directional shear and stability depth affect wave altitude and amplitude
- Ensemble spread is not automatically a calibrated probability
- Physical turbulence likelihood and route-level operational impact are separate
- Profiler or targeted aircraft observations can resolve the controlling uncertainty
Scoring notes in the case file
- Development draft. The synthetic outcome is provisional and must be replaced or independently archived before promotion.
- Reward separation of wave presence, peak altitude, and operational significance.
- Do not reward a precise flight-level claim unsupported by the vertical structure.
Structured target(s). turbulence_operational_18_06z (binary, probability), turbulence_peak_band_18_06z (ordered_categorical, probability)
Across the eight runs. Mean 84.6, median 85, range 84–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 85 | 84 | 84 | 85 | 85 | 85 | 85 |
FB-1118 near freezing ptype transition
In plain terms. A shallow subfreezing layer is interacting with a warm nose aloft as a winter storm crosses a metro area. The broad guidance agrees on the storm track but disagrees on how quickly the surface cold layer erodes. The forecast must separate the most likely precipitation type from the timing of the transition and acknowledge that mixed precipitation observations are imperfect.
Case record. cases/drafts · v0.3 · winter_weather · forecast time: Synthetic case issue time: 2026-01-12 12:00 UTC
What the model received
- surface_observations: At issue time the metro is 31 F with a 24 F dewpoint and northeast wind at 12 kt. A nearby southeast station reports light rain at 35 F, while a northwest station reports wet snow at 27 F. Road temperatures are near freezing and the temperature gradient is tightening. Source label: Synthetic author packet
- vertical_thermal_profile: The sounding has a subfreezing layer from the surface to 925 mb, a +2 C warm nose from 850-800 mb, and saturation through the dendritic growth zone. The cold layer is shallow but not yet eroding. Small changes in mixing and precipitation intensity could change whether melting hydrometeors refreeze before reaching the ground. Source label: Synthetic author packet
- model_guidance_spread: Four guidance members keep sleet or freezing rain in the metro through 18z. Three erode the cold layer by 17z and favor rain. Two maintain a colder northwest gradient and keep snow in the first hour. The spread is scenario evidence, not a calibrated probability. Source label: Synthetic author packet
- observation_uncertainty: Precipitation-type reports may be mixed within a station's observation window, and nearby stations sample different elevations and thermal layers. A radar bright band or upstream profiler trend would help resolve the warm-nose depth, but neither observation guarantees a single category at every point in the metro. Source label: Synthetic author packet
What it had to do
Forecast the dominant precipitation type through 18z and the time, in hours after issue, when the metro most likely changes to rain. Explain how the warm-nose depth, shallow cold layer, precipitation rate, and mixing affect the forecast. State whether freezing rain occurs at the metro and identify one observation that would change the forecast most. End with exactly one fenced JSON block containing target_id ptype_distribution as an ordered probability list over snow, sleet, freezing_rain, and rain; target_id changeover_time_hours as a continuous value; and target_id freezing_rain_occurs as a binary probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- The vertical thermal profile controls precipitation type more directly than surface temperature alone
- A shallow cold layer can support sleet or freezing rain even with a warm nose aloft
- Precipitation rate and turbulent mixing can change the timing of cold-layer erosion
- Nearby observations may disagree because of spatial, temporal, and elevation variability
- Scenario spread should be translated into calibrated probabilities, not copied as a member count
Scoring notes in the case file
- Development draft. The synthetic outcome is provisional and must be replaced or independently archived before promotion.
- Score the ordered category distribution separately from transition timing and freezing-rain occurrence.
- Do not reward a confident single-category answer that ignores the vertical profile or observation uncertainty.
Structured target(s). ptype_distribution (ordered_categorical, probability), changeover_time_hours (continuous, hours_after_issue), freezing_rain_occurs (binary, probability)
Across the eight runs. Mean 85.8, median 87, range 82–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 82 | 87 | 87 | 84 | 87 | 87 | 85 |
specialty_forecasting
42 cases, 336 model-case scores
FB-0004 aviation fog stratus
In plain terms. An airport is forecast overnight after evening rain. The main aviation question is whether IFR or LIFR conditions develop before sunrise.
Case record. cases/v0.1 · v0.1 · aviation · forecast time: Mock case, valid 06z-15z
What the model received
- surface_observations: At 06z, the airport reports temperature 53 F, dewpoint 52 F, wind calm, visibility 6 statute miles, and scattered clouds at 500 ft. Rain ended two hours ago. Nearby rural stations already report 2-4 statute miles visibility. Source label: Mock author-provided packet
- satellite: Nighttime satellite imagery shows clearing advancing from the west. Low clouds remain in a narrow band over and east of the airport. Source label: Mock author-provided packet
- guidance: Short-range guidance lowers visibility to 1 mile between 09z and 13z and ceiling to 300 ft. Another model keeps a light 5 kt wind overnight and maintains MVFR ceilings. Source label: Mock author-provided packet
What it had to do
Write a short aviation forecast discussion for ceiling and visibility through 15z. Include the most likely flight category, timing, and confidence.
What the case says a strong answer should cover
- Near-saturated boundary layer favors fog or low stratus
- Calm wind and clearing after rain favor radiational fog
- Residual low clouds may limit cooling but still support IFR ceilings
- Small wind differences can determine fog severity
Scoring notes in the case file
- Strong answers should forecast IFR as plausible or likely, with LIFR possible if winds remain calm and clearing occurs.
- Credit timing focused near the pre-dawn period.
- Penalize answers that ignore the recent rainfall and temperature-dewpoint spread.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 86.2, median 87, range 85–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 85 | 85 | 87 | 85 | 87 |
FB-0005 tropical intensity track
In plain terms. A tropical storm is moving west-northwest toward a weakness in the subtropical ridge. The forecast problem is whether the system strengthens before approaching the coast and how far north the track bends.
Case record. cases/v0.1 · v0.1 · tropical_weather · forecast time: Mock case, valid 12z-36h
What the model received
- current_storm: At 12z, the storm is centered 360 miles southeast of the coastline and moving west-northwest at 12 kt. Maximum sustained winds are estimated at 55 kt. Recent satellite imagery shows deep convection displaced slightly east of the low-level center, but cloud tops have cooled during the past 3 hours. Source label: Mock author-provided packet
- environment: Sea-surface temperatures are 29-30 C along the forecast path. Mid-level relative humidity is near 65 percent. Vertical wind shear is 18-22 kt from the southwest but is forecast to decrease to 10-15 kt in 18-24 hours. Ocean heat content is moderate. Source label: Mock author-provided packet
- guidance: Track guidance clusters near a landfall in 30-36 hours, but the western solutions keep the storm weaker and farther south while the eastern solutions deepen the storm and turn it north sooner. Intensity guidance ranges from 60 kt to 85 kt at landfall. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package for emergency management. Include expected track trend, intensity at closest approach, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Warm water supports strengthening
- Current shear and displaced convection limit immediate intensification
- Decreasing shear raises later intensification risk
- Stronger storms may turn north sooner because they feel deeper steering
- Track and intensity errors are coupled
Scoring notes in the case file
- Strong answers should avoid deterministic rapid intensification.
- Credit answers that connect storm depth to the northward turn.
- Watch discussion should focus on tropical storm or hurricane watches if coastal timing supports it.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.8, median 86, range 85–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 86 | 85 | 85 | 86 | 87 | 85 | 86 |
FB-0006 flash flooding training convection
In plain terms. A slow-moving boundary is expected to focus repeated thunderstorms overnight. The forecast problem is flash-flood potential versus isolated severe wind.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 21z-09z
What the model received
- surface_observations: At 21z, a stationary boundary lies west-east across the forecast area. South of the boundary, dewpoints are 73-76 F with southerly winds at 20 kt. North of it, temperatures are 8-12 F cooler with easterly winds. Source label: Mock author-provided packet
- moisture_instability: Precipitable water is 2.0-2.2 inches, near the 95th percentile for the region. MLCAPE is 1800-2500 J/kg south of the boundary. Warm-cloud depth is deep, and storm motions are expected to be parallel to the boundary. Source label: Mock author-provided packet
- guidance: Several high-resolution models develop multiple rounds of convection after 00z along the boundary. One model places the heaviest rain 50 miles north where the boundary lifts. Three-hour rainfall maxima range from 2 to 5 inches. Flash-flood guidance is 2.5 inches in 3 hours in urban areas and 3.5 inches elsewhere. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package for an overnight operations briefing. Include the rainfall/flash-flood forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Training along a quasi-stationary boundary increases flash-flood risk
- High precipitable water and deep warm-cloud depth favor efficient rainfall
- Boundary placement controls the corridor of heaviest rain
- Urban flash-flood guidance is low enough for concern
- Severe wind may be secondary to flooding
Scoring notes in the case file
- Strong answers should prioritize flash flooding over generic thunderstorm wording.
- Credit answers that discuss a flood watch or flash flood watch.
- Penalize answers that focus mainly on severe wind without addressing training rainfall.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 86.1, median 86, range 85–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 86 | 88 | 87 | 85 | 86 | 85 | 86 |
FB-0007 fire weather red flag
In plain terms. A dry cold front will cross cured grassland fuels during the afternoon. The forecast problem is whether fire weather conditions meet watch or warning thresholds.
Case record. cases/v0.1 · v0.1 · fire_weather · forecast time: Mock case, valid 17z-02z
What the model received
- fuels: Fine fuels are cured after two weeks without measurable rainfall. Energy release component is above the 85th percentile. Recent initial attack activity has increased in adjacent counties. Source label: Mock author-provided packet
- surface_forecast: Temperatures are expected to reach 78-82 F ahead of the front. Minimum relative humidity is forecast at 13-18 percent for 3-5 hours. Sustained southwest winds increase to 20-25 mph with gusts 35-40 mph, then shift west-northwest behind the front. Source label: Mock author-provided packet
- guidance: One model keeps RH above 20 percent because it mixes less deeply. Two models mix to 700 mb and produce RH near 15 percent with gusts above 35 mph. No wetting rain is expected with the front. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package for fire partners. Include the fire-weather forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Low RH, strong winds, and cured fuels overlap
- Dry frontal passage and wind shift increase operational concern
- Mixing depth controls whether RH falls below threshold
- No wetting rain means fuels remain receptive
- Red Flag Warning or Fire Weather Watch logic should be discussed
Scoring notes in the case file
- Strong answers should recommend a fire-weather headline if local criteria are met or nearly met.
- Credit answers that identify wind shift as an operational hazard.
- Penalize answers that evaluate RH or wind alone without fuels.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 86.5, median 86, range 86–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 86 | 88 | 86 | 86 | 86 | 86 | 86 | 88 |
FB-0012 tornado vs hail conditional supercells
In plain terms. A conditional supercell setup has strong instability and shear, but storm coverage is uncertain. The main forecast problem is tornado versus giant hail potential if storms form.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 18z-06z
What the model received
- thermodynamics: By 21z, MLCAPE is forecast to reach 3000-4000 J/kg with 700-500 mb lapse rates near 8.5 C/km. Dewpoints are 67-70 F east of a dryline. CIN is forecast to weaken from -125 J/kg to -50 J/kg by late afternoon. Source label: Mock author-provided packet
- wind_profile: Effective bulk shear is 55-65 kt. The 0-1 km SRH is 100 m2/s2 at 21z but increases to 250 m2/s2 after 00z as the low-level jet strengthens. LCL heights are 1200-1500 m AGL during peak heating, lowering after sunset. Source label: Mock author-provided packet
- guidance: CAMs disagree on initiation. Two runs show isolated supercells near the dryline by 22z. Two runs remain capped until after dark. A synoptic cold front overtakes the dryline after 03z. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include whether the dominant conditional hazard is tornadoes, hail, or both; top uncertainties; observations that would change the forecast; ignored guidance; confidence; and whether any watch would be appropriate.
What the case says a strong answer should cover
- Large CAPE, steep lapse rates, and strong shear support significant hail
- Tornado risk increases later as low-level shear improves and LCLs lower
- Initiation remains conditional because CIN may persist
- Storm timing affects hazard type
- Watch decision may be conditional but high-impact if initiation occurs
Scoring notes in the case file
- Strong answers should distinguish early hail-dominant risk from later tornado risk.
- Credit answers that discuss conditional watch timing.
- Penalize broad outbreak language without initiation confidence.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.4, median 85, range 84–89. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 86 | 85 | 84 | 84 | 86 | 85 | 89 |
FB-0013 qlcs wind embedded tornado
In plain terms. A line of storms will move into a moist but only weakly unstable warm sector overnight. The forecast problem is damaging wind versus embedded tornado risk.
Case record. cases/v0.1 · v0.1 · severe_weather · forecast time: Mock case, valid 00z-09z
What the model received
- radar: At 00z, a mature QLCS is entering the western forecast area. Radar shows several bowing segments and weak mesovortices, but lightning has decreased in the northern half of the line. Source label: Mock author-provided packet
- environment: MLCAPE is 500-1000 J/kg, highest south. The 0-1 km shear is 35 kt and 0-3 km shear is 50 kt. Surface dewpoints are 64-67 F. A stable layer is developing north of a warm front, but the southern counties remain in the warm sector. Source label: Mock author-provided packet
- guidance: HRRR maintains a continuous line with 50-60 mph gusts. Another CAM weakens the line after 03z as instability decreases. Both show strongest rotation near the warm front intersection. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include the main overnight severe threat, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- QLCS damaging wind can persist with modest CAPE if shear is strong
- Embedded tornado risk is favored near warm-front/line intersections
- Northern stable layer should reduce severe risk
- Southern warm sector has higher wind/tornado concern
- Watch decision should be spatially focused
Scoring notes in the case file
- Strong answers should not treat the entire line equally.
- Credit answers that mention embedded tornadoes without overhyping them.
- Penalize ignoring the stable layer north of the warm front.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.1, median 85, range 84–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 85 | 86 | 85 | 85 | 85 | 85 | 86 |
FB-0014 winter evaporative cooling marginal
In plain terms. A marginal winter event is beginning with temperatures above freezing but a dry sub-cloud layer. The forecast problem is whether evaporative cooling changes rain to accumulating snow.
Case record. cases/v0.1 · v0.1 · winter_weather · forecast time: Mock case, valid 09z-21z
What the model received
- surface_observations: At 09z, temperatures across the city are 35-38 F with dewpoints 22-26 F. Northeast winds are 10-15 kt. Light rain is approaching from the south, but the first echoes are virga. Source label: Mock author-provided packet
- sounding: The 09z sounding is above freezing from the surface to 875 mb, but the wet-bulb profile falls to 0 C or slightly below from the surface through 800 mb once saturation occurs. The dendritic growth zone is saturated after 12z. Source label: Mock author-provided packet
- guidance: Model A keeps surface temperatures 36-38 F and produces mostly rain. Model B cools the column by 12z and produces 2-4 inches of wet snow. Model C has a narrow 1-2 inch band on the north side of the city. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include precipitation type and accumulation forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Evaporative cooling can overcome above-freezing initial temperatures
- Wet-bulb profile matters more than raw temperature
- Snow accumulation is sensitive to precipitation intensity and surface temps
- Virga/dry layer delays onset but supports cooling
- Watch decision depends on accumulation confidence and impacts
Scoring notes in the case file
- Strong answers should not use surface temperature alone.
- Credit answers that discuss wet snow and low-ratio accumulation.
- Penalize deterministic 4-inch forecasts without intensity confidence.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 84.1, median 84, range 82–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 85 | 85 | 82 | 84 | 84 | 84 | 84 |
FB-0015 freezing rain accretion surface temps
In plain terms. A freezing rain event is possible, but road and air temperatures are near freezing. The forecast problem is ice accretion versus mainly wet roads.
Case record. cases/v0.1 · v0.1 · winter_weather · forecast time: Mock case, valid 12z-00z
What the model received
- observations: At 12z, air temperatures are 30-32 F in valleys and 33-35 F on ridges. Road temperatures are 31-34 F. Dewpoints are 28-31 F. Light freezing drizzle is reported in two sheltered valley sites. Source label: Mock author-provided packet
- sounding: A warm layer from 850-750 mb peaks at +4 C. The surface cold layer is shallow, only 800-1200 ft deep. Wet-bulb temperatures remain below 0 C in valleys but rise above 0 C on exposed ridges by afternoon. Source label: Mock author-provided packet
- guidance: Model A forecasts 0.25 inches of ice areawide. Model B has 0.05-0.10 inches in valleys only. Model C changes most areas to plain rain by 18z but keeps sheltered valleys below freezing until 21z. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include ice accretion and impact forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Warm nose supports freezing rain rather than snow
- Shallow cold layer and terrain cause large accretion differences
- Road and exposed-surface temperatures control impacts
- Valleys should have higher ice risk than ridges
- Areal ice forecast should not be uniform
Scoring notes in the case file
- Strong answers should avoid areawide 0.25 inch ice unless justified.
- Credit answers that separate valleys/ridges and roads/elevated surfaces.
- Penalize treating all freezing rain as equal accretion.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.1, median 85, range 84–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 85 | 86 | 85 | 84 | 85 | 85 | 84 |
FB-0016 tropical ri false alarm
In plain terms. Some guidance shows rapid intensification, but current structure is poor. The forecast problem is whether to buy the RI signal.
Case record. cases/v0.1 · v0.1 · tropical_weather · forecast time: Mock case, valid 00z-48h
What the model received
- current_structure: The storm has 45 kt winds. Microwave imagery shows a broad, tilted vortex with no closed eyewall. Deep convection pulses near the center at night but weakens during the day. Outflow is restricted on the west side. Source label: Mock author-provided packet
- environment: SSTs are 30 C and ocean heat content is high. Shear is 15 kt now and may fall below 10 kt in 24 hours. Mid-level dry air lies west of the storm. The storm will cross a cool wake in 36-48 hours. Source label: Mock author-provided packet
- guidance: Two statistical models show a 40 percent RI probability. The dynamical hurricane model deepens the storm to 90 kt in 36 hours. Global models keep it 55-65 kt because the core remains broad. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include intensity forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Favorable ocean and decreasing shear support some strengthening
- Poor inner-core structure limits immediate RI confidence
- Dry air and cool wake are negative factors
- RI guidance should be considered but not blindly followed
- Watch decision should depend on land timing and intensity risk
Scoring notes in the case file
- Strong answers should not ignore the RI signal, but should temper it.
- Credit answers that specify structural observations needed for RI.
- Penalize deterministic 90 kt forecasts with poor core structure.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 83.2, median 84, range 78–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 84 | 85 | 78 | 86 | 85 | 82 | 81 |
FB-0017 tropical track ridge weakness
In plain terms. A hurricane is approaching a break in the subtropical ridge. Small timing changes will determine whether it recurves offshore or threatens the coast.
Case record. cases/v0.1 · v0.1 · tropical_weather · forecast time: Mock case, valid 12z-72h
What the model received
- storm: The hurricane is moving west-northwest at 14 kt with 80 kt winds. It is vertically deep and has a compact core. It is 650 miles east-southeast of the coast. Source label: Mock author-provided packet
- synoptic_pattern: A mid-latitude trough is forecast to erode the western edge of the ridge in 48-60 hours. If the storm slows, it may turn north before reaching the coast. If it maintains speed for another 24 hours, it may get close enough for coastal impacts before recurving. Source label: Mock author-provided packet
- guidance: Ensemble tracks are bimodal. Half recurve 150-250 miles offshore. Half bring the hurricane within 50 miles of the coast before turning north. Deterministic Model A is offshore and fast to recurve; Model B is slower to recurve and closer to the coast. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include track and coastal-impact forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Ridge weakness timing controls recurvature
- Forward speed matters for coastal impact risk
- Deep storm responds to deeper-layer steering
- Bimodal ensembles require risk communication
- Watch may be premature at 72h but preparedness messaging is needed
Scoring notes in the case file
- Strong answers should avoid a false deterministic landfall/offshore call.
- Credit answers that communicate coastal risk despite uncertainty.
- Penalize ignoring the bimodal ensemble structure.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.5, median 85, range 85–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 88 | 85 | 85 | 86 | 85 | 85 | 85 | 85 |
FB-0018 aviation taf amendment ceilings
In plain terms. Conditions are deteriorating faster than the current TAF. The forecast problem is whether to amend for IFR/LIFR ceilings and how long they last.
Case record. cases/v0.1 · v0.1 · aviation · forecast time: Mock case, valid 13z TAF amendment
What the model received
- current_taf: The current TAF has MVFR ceilings at 1200 ft through 16z, improving to VFR after 18z. No LIFR is included. Source label: Mock author-provided packet
- observations: At 13z, the airport reports ceiling 700 ft, visibility 3 miles in mist, wind northeast 6 kt. Two upstream airports report ceilings 300-500 ft with visibilities 1-2 miles. Satellite shows the low stratus deck expanding westward rather than eroding. Source label: Mock author-provided packet
- guidance: LAMP keeps MVFR ceilings through 16z. The HRRR and recent observations support IFR through at least 18z, with a 30 percent chance of LIFR before 15z. Mixing improves after 19z. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include the amended aviation forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- TAF should be amended when observations deteriorate faster than forecast
- Upstream IFR/LIFR and satellite expansion support lower ceilings
- MVFR guidance should be discounted if obs are worse
- Timing of improvement depends on mixing/stratus erosion
- Watch is generally not applicable for a TAF amendment
Scoring notes in the case file
- Strong answers should correctly classify 300-700 ft ceilings.
- Credit answers that amend to IFR and mention LIFR tempo/probability.
- Penalize using public watch logic for aviation-only restrictions.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.1, median 85, range 82–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 85 | 82 | 88 | 85 | 86 | 85 | 85 |
FB-0019 fire weather dry thunderstorm
In plain terms. Dry thunderstorms are possible over critically dry fuels. The forecast problem is lightning ignition risk despite limited rainfall.
Case record. cases/v0.1 · v0.1 · fire_weather · forecast time: Mock case, valid 18z-06z
What the model received
- fuels: Fuels are critically dry after a month with less than 20 percent of normal rainfall. ERC is above the 90th percentile. Several large fires are ongoing downwind of the forecast area. Source label: Mock author-provided packet
- environment: Inverted-V soundings show cloud bases near 12000 ft AGL with sub-cloud RH below 20 percent. MUCAPE is 400-800 J/kg above a deep mixed layer. Gusty outflow winds are possible. Precipitable water is only 0.55 inches. Source label: Mock author-provided packet
- guidance: CAMs produce isolated high-based thunderstorms after 22z. Most grid points receive less than 0.05 inches of rain, but lightning density guidance highlights the central mountains. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include dry-thunderstorm/fire forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- High-based storms can produce lightning with little wetting rain
- Critically dry fuels make ignition risk high
- Outflow winds can worsen spread
- Cloud base and PW support dry thunderstorm risk
- Fire Weather Watch/Red Flag logic may apply
Scoring notes in the case file
- Strong answers should emphasize lightning ignition, not rainfall.
- Credit answers that mention outflow wind spread risk.
- Penalize treating isolated storms as low impact because coverage is sparse.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 87.8, median 88, range 84–91. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 88 | 91 | 89 | 86 | 84 | 87 | 88 | 89 |
FB-0020 marine coastal low gales
In plain terms. A coastal low will deepen near a tight pressure gradient. The forecast problem is gale timing, wave growth, and coastal water impacts.
Case record. cases/v0.1 · v0.1 · marine · forecast time: Mock case, valid 06z-00z
What the model received
- marine_observations: Buoys report northeast winds 25 kt gusting 32 kt and seas 7-9 ft at 06z. Coastal tide gauges are running 1.2 ft above predicted tide. Pressure is falling 3 mb per 3 hours offshore. Source label: Mock author-provided packet
- guidance: Model A deepens the coastal low faster and brings 35-40 kt gusts by 12z with seas building to 13 ft. Model B keeps the low farther offshore and delays gales until 18z, with seas near 10 ft. Surge guidance ranges from 1.5 to 2.5 ft above normal during high tide. Source label: Mock author-provided packet
- coastal_context: Minor coastal flooding begins near 1.8 ft above normal at vulnerable roads. The strongest onshore flow overlaps high tide between 14z and 17z if Model A is correct. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include marine/coastal hazards, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Gale timing depends on low track/deepening
- Existing buoy trends support increasing winds/seas
- Onshore flow timing with high tide controls coastal flooding
- Surge range crosses minor flood threshold
- Gale Watch/Warning or coastal flood watch/advisory logic should be discussed
Scoring notes in the case file
- Strong answers should address both marine and coastal hazards.
- Credit answers that connect high tide timing to flooding risk.
- Penalize only forecasting wind without waves/water levels.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 85.6, median 86, range 85–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 86 | 86 | 85 | 87 | 86 | 85 | 85 |
FB-0021 hydrology antecedent wet soils
In plain terms. A moderate rainfall event may produce flooding because soils are already saturated. The forecast problem is hydrologic response rather than only QPF.
Case record. cases/v0.1 · v0.1 · general · forecast time: Mock case, valid 12z-12z
What the model received
- antecedent_conditions: The basin has received 4-6 inches of rain in the past week. Soil moisture is above the 95th percentile. Small streams are already running high, and reservoirs are near seasonal flood-control targets. Source label: Mock author-provided packet
- rainfall_forecast: Widespread rainfall of 1-2 inches is expected, with localized 3-inch amounts if embedded convection develops. Rainfall rates should mostly be 0.25-0.50 inches per hour, but convection could briefly exceed 1 inch per hour. Source label: Mock author-provided packet
- guidance: One hydrologic model shows minor river flooding at two forecast points. Another keeps rivers below flood stage because it uses lower QPF. Flash flood guidance is unusually low: 1.5 inches in 3 hours in the most saturated headwaters. Source label: Mock author-provided packet
What it had to do
Prepare one forecast package. Include flood forecast, top uncertainties, observations that would change the forecast, ignored guidance, confidence, and whether any watch would be appropriate.
What the case says a strong answer should cover
- Antecedent wet soils increase flood response
- Moderate rainfall can flood when FFG is low
- River and flash flood risks may both exist
- Embedded convection controls localized flash-flood risk
- Hydrologic model assumptions should be compared to QPF and soils
Scoring notes in the case file
- Strong answers should not dismiss flooding because QPF is only 1-2 inches.
- Credit answers that separate river flooding from flash flooding.
- Penalize ignoring antecedent conditions.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 86.6, median 86, range 85–89. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 87 | 88 | 85 | 85 | 89 | 86 | 88 |
FB-1007 calibration fire wind
In plain terms. A dry, windy afternoon may support rapid fire spread, but the strongest winds may remain aloft.
Case record. cases/v0.2 · v0.2 · fire_weather · forecast time: Synthetic calibration issue time: 2026-08-21 15:00 UTC
What the model received
- surface_observations: Temperature is 91 F, dewpoint is 39 F, relative humidity is 18 percent, and sustained west wind is 17 kt with gusts to 26 kt. Source label: Synthetic calibration packet
- sounding_and_guidance: The mixed layer reaches 700 mb. Guidance shows 20-25 kt at transport level and uncertainty in afternoon mixing depth. Source label: Synthetic calibration packet
What it had to do
Classify the afternoon fire-weather risk as low, elevated, or red_flag in that order. Explain the main limiting factor and one observation that would change the classification. End with exactly one fenced JSON block (```json ... ```) using target_id fire_weather_category and a three-element probability list in the stated order.
What the case says a strong answer should cover
- Combine humidity, wind, fuels implication, and mixing depth
- Represent uncertainty in the surface wind forecast
Scoring notes in the case file
- Calibration case; independent graders must review narrative and vector separately.
Structured target(s). fire_weather_category (categorical, None)
Across the eight runs. Mean 84.9, median 85, range 82–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 85 | 86 | 84 | 86 | 82 | 85 | 84 |
FB-1008 calibration aviation ceiling
In plain terms. A moist boundary layer is producing low clouds near an airport, with uncertain improvement timing.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: Synthetic calibration issue time: 2026-09-03 09:00 UTC
What the model received
- surface_observations: The airport reports a 700 ft ceiling, 2 statute miles visibility, light southeast wind, and a temperature-dewpoint spread of 1 F. Source label: Synthetic calibration packet
- guidance: Guidance ranges from lifting to 1500 ft by 16z to persistent low ceilings through 20z. Source label: Synthetic calibration packet
What it had to do
Classify the lowest prevailing ceiling category during 16z-20z as VFR, MVFR, IFR, or LIFR in that order. Explain the main timing uncertainty. End with exactly one fenced JSON block (```json ... ```) using target_id ceiling_category and a four-element probability list in the stated order.
What the case says a strong answer should cover
- Use the small temperature-dewpoint spread and current ceiling
- Distinguish improvement timing from eventual improvement
Scoring notes in the case file
- Calibration case; category order must be preserved exactly.
Structured target(s). ceiling_category (ordered_categorical, None)
Across the eight runs. Mean 83.0, median 84, range 77–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 85 | 80 | 84 | 77 | 85 | 85 | 86 | 82 |
FB-1009 provisional marine gust
In plain terms. A tightening pressure gradient is expected across an exposed coastal marine zone, but the timing of the strongest gusts remains uncertain.
Case record. cases/v0.2 · v0.2 · marine · forecast time: Synthetic case issue time: 2026-10-14 12:00 UTC
What the model received
- marine_observations: The buoy reports sustained southwest wind at 24 kt with occasional 31 kt gusts. Seas are 8 ft with a short-period component. Source label: Synthetic provisional holdout packet
- guidance: Guidance ranges from 34 to 42 kt for the maximum gust during the afternoon. The strongest solution brings a narrow convective band through after 19z. Source label: Synthetic provisional holdout packet
What it had to do
Forecast the maximum wind gust in knots from 18z through 00z. Give one numeric value, the main physical driver, and the largest timing uncertainty. End with exactly one fenced JSON block (```json ... ```) using target_id maximum_gust and a numeric value in knots.
What the case says a strong answer should cover
- Use the pressure-gradient signal
- Account for the convective-band timing uncertainty
- Keep the point forecast within or justify departure from guidance
Scoring notes in the case file
- The objective score uses the numeric JSON value only.
- This case is provisional holdout material and is not for prompt tuning.
Structured target(s). maximum_gust (continuous, kt)
Across the eight runs. Mean 84.6, median 87, range 73–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 84 | 73 | 87 | 87 | 85 |
FB-1010 provisional aviation visibility
In plain terms. Low stratus is affecting an airport near a shallow frontal boundary. The main operational question is whether conditions improve enough for a planned afternoon arrival window.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: Synthetic case issue time: 2026-11-06 09:00 UTC
What the model received
- aviation_observations: The latest observation reports a 500 ft ceiling, visibility 2 sm in mist, and a light east wind. The ceiling has lifted 200 ft in the last two hours but the boundary remains nearby. Source label: Synthetic provisional holdout packet
- guidance: Guidance disagrees on the speed of boundary passage. One solution clears the ceiling by 18z, while another keeps a sub-1000 ft ceiling through 21z. Source label: Synthetic provisional holdout packet
What it had to do
Forecast the probability that the ceiling is at least 1000 ft during 18z-21z. Give a probability from 0 to 1, the main improvement signal, and the largest uncertainty. End with exactly one fenced JSON block (```json ... ```) using target_id ceiling_improvement_probability and a numeric value.
What the case says a strong answer should cover
- Use the observed lifting trend
- Recognize uncertainty in boundary timing
- Connect the threshold to aviation operations
Scoring notes in the case file
- The objective score uses the structured probability only.
- This case is provisional holdout material and is not for prompt tuning.
Structured target(s). ceiling_improvement_probability (binary, probability)
Across the eight runs. Mean 82.8, median 83, range 73–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 73 | 81 | 85 | 81 | 81 |
FB-1013 aviation ceiling transition
In plain terms. A moist nocturnal boundary layer is developing near an aviation terminal. The operational question is whether a low ceiling transition will persist through the morning push.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: Synthetic case issue time: 2026-02-14 18:00 UTC
What the model received
- observations: The terminal is currently VFR with a 2500 ft ceiling, falling temperature, and a 3 kt wind. Nearby valley stations have already reached 800-1200 ft ceilings with visibility near 5 statute miles. Source label: Synthetic author packet
- guidance: The boundary layer is nearly saturated below 900 mb. Guidance differs on whether a shallow drainage flow maintains mixing overnight; ceiling forecasts range from 600 ft to 1800 ft at 12z. Source label: Synthetic author packet
What it had to do
Forecast the dominant ceiling category at the terminal at 12z using the ordered categories VFR, MVFR, IFR, LIFR. Explain the boundary-layer mechanism, the main uncertainty, and the observation that would change the decision. End with exactly one fenced JSON block (```json ... ```) using target_id ceiling_category and a probability list in that category order.
What the case says a strong answer should cover
- Use saturation, wind, drainage, and nearby observations together
- Distinguish a temporary low cloud layer from a persistent operational ceiling
- Identify the observation with the greatest value for the ceiling decision
Scoring notes in the case file
- The probability vector is ordered VFR, MVFR, IFR, LIFR.
- Operational usefulness is scored separately from category probability.
Structured target(s). ceiling_category (ordered_categorical, None)
Across the eight runs. Mean 84.2, median 85, range 77–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 77 | 85 | 87 | 84 | 85 | 85 | 85 | 86 |
FB-1014 lake effect band placement
In plain terms. A cold air outbreak is crossing a relatively warm lake, but the snowband placement depends on wind veering and shoreline convergence.
Case record. cases/v0.2 · v0.2 · winter_weather · forecast time: Synthetic case issue time: 2026-01-28 12:00 UTC
What the model received
- observations: Lake-surface temperature is 4 C while 850 mb temperature is -14 C. Radar shows a narrow band forming on the southwest shore. Surface winds are 290 degrees at 18 kt and veering slowly. Source label: Synthetic author packet
- guidance: Guidance differs between a single dominant band, a multi-band regime, and rapid inland drift as winds veer to 320 degrees. Low-level moisture remains abundant for at least six hours. Source label: Synthetic author packet
What it had to do
Forecast the most likely snowband placement category through 18z using the ordered categories southwest shore, central corridor, northeast shore, and diffuse or multi-band. Explain the role of fetch, wind direction, stability, and shoreline convergence. End with exactly one fenced JSON block (```json ... ```) using target_id lake_effect_band_location and a probability list in that category order.
What the case says a strong answer should cover
- Use lake-to-850 mb temperature difference and fetch
- Recognize wind veering as a placement uncertainty
- Separate band intensity from band location
Scoring notes in the case file
- The probability vector is ordered southwest shore, central corridor, northeast shore, diffuse or multi-band.
Structured target(s). lake_effect_band_location (ordered_categorical, None)
Across the eight runs. Mean 85.4, median 86, range 80–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 80 | 87 | 87 | 84 | 85 | 88 | 85 | 87 |
FB-1015 tropical rapid intensification
In plain terms. A tropical cyclone is entering a favorable thermodynamic environment, but recent intensity change and inner-core organization remain uncertain.
Case record. cases/v0.2 · v0.2 · tropical_weather · forecast time: Synthetic case issue time: 2026-09-03 12:00 UTC
What the model received
- satellite_and_guidance: The cyclone has developed a persistent central dense overcast and a partial eyewall. Sea-surface temperatures are 29.5 C, upper-ocean heat content is high, and vertical shear is forecast to fall from 18 kt to 8 kt within 24 hours. Source label: Synthetic author packet
- observations: Microwave imagery shows competing inner-core asymmetries, and aircraft data are sparse. Guidance ranges from 25 kt strengthening to 45 kt strengthening over the next 24 hours. Source label: Synthetic author packet
What it had to do
Forecast the probability of rapid intensification during the next 24 hours. Define the decisive inner-core or shear observation, explain the competing signals, and distinguish structural uncertainty from track uncertainty. End with exactly one fenced JSON block (```json ... ```) using target_id rapid_intensification_24h and a numeric probability.
What the case says a strong answer should cover
- Combine shear, ocean heat content, and inner-core organization
- Treat rapid intensification as conditional on structure, not environment alone
- Identify the observation that would update the intensity forecast
Scoring notes in the case file
- The binary target is rapid intensification during the next 24 hours, not eventual maximum intensity.
Structured target(s). rapid_intensification_24h (binary, probability)
Across the eight runs. Mean 86.5, median 87, range 82–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 86 | 87 | 82 | 88 | 87 | 88 | 87 |
FB-1021 real asos den visibility
In plain terms. Denver International Airport begins the forecast period with good visibility and few clouds, but a shallow moist layer may deepen as the wind changes.
Case record. cases/v0.2 · v0.2 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- observations: At 12:53 UTC, DEN reported 50 F, dew point 45 F, wind 3 kt, visibility 10 statute miles, and few clouds. At 13:00 UTC the wind was 290 degrees at 2 kt, visibility remained 10 miles, and few clouds persisted. Source label: Iowa Environmental Mesonet ASOS archive, excerpt stored in cases/observations/FB-1021-DEN-2025-10-15-visibility.csv
- context: Forecast whether visibility will fall below 1 statute mile at any time before 17:00 UTC. No numerical guidance is supplied. Source label: Task framing around archived observations
What it had to do
Forecast the probability that DEN visibility will fall below 1 statute mile at any time before 17:00 UTC. Explain the competing signals, identify the observation that would most update the probability, and distinguish visibility uncertainty from temperature uncertainty. End with exactly one fenced JSON block (```json ... ```) using target_id visibility_below_one_mile_by_17z and a numeric probability.
What the case says a strong answer should cover
- Use the temperature-dew-point spread, wind shift, and cloud evolution
- Recognize that good visibility at issue time does not rule out later deterioration
- Identify a new visibility, ceiling, or dew-point observation as an update signal
Scoring notes in the case file
- The binary event is any ASOS visibility below 1 statute mile before 17z.
Structured target(s). visibility_below_one_mile_by_17z (binary, probability)
Across the eight runs. Mean 79.1, median 80, range 73–84. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 76 | 82 | 81 | 73 | 76 | 80 | 81 |
FB-1102 precipitation type transition
In plain terms. A shallow subfreezing layer is entrenched near the surface while warm air overruns aloft. The forecast depends on the depth of the warm nose and the rate at which cold air is scoured from the lowest 500 m.
Case record. cases/v0.3-development · v0.3 · winter_weather · forecast time: Synthetic case issue time: 2026-02-07 15:00 UTC
What the model received
- vertical_thermal_profile: Surface temperature is -1 C. A +2.5 C warm nose extends from 925 to 825 hPa, followed by a shallow subfreezing layer from 825 to 780 hPa. Guidance differs by 80 m in warm-nose depth and 2 C in the surface wet bulb after precipitation begins. Source label: Synthetic author packet
- surface_and_radar_trends: Pressure falls slowly, precipitation echoes are strengthening from the southwest, and wet-bulb cooling has lowered the surface temperature by 0.8 C in two hours. A low-level jet is forecast to strengthen overnight. Source label: Synthetic author packet
What it had to do
Provide ordered probabilities for rain, snow, sleet, and freezing rain during 21z-03z, preserving the documented category order and summing to one. Explain how the warm-nose depth and wet-bulb evolution control the transition. End with exactly one fenced JSON block using target_id ptype_transition_21_03z and a probability vector. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Use the complete vertical thermal profile rather than surface temperature alone
- Distinguish melting, refreezing, and supercooled-drop pathways
- Treat wet-bulb cooling and warm-nose depth as competing time-dependent signals
Scoring notes in the case file
- Category order is rain, snow, sleet, freezing_rain. Missing mass is a contract failure.
- This synthetic truth is provisional and must not support a ranking.
Structured target(s). ptype_transition_21_03z (ordered_categorical, probability_vector)
Across the eight runs. Mean 82.1, median 84, range 68–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 68 | 80 | 84 | 82 | 87 | 87 | 84 | 85 |
FB-1105 aviation fog transition
In plain terms. A moist, weak-wind boundary layer is likely to produce a low ceiling overnight, but intermittent mixing could delay or interrupt the transition to instrument flight conditions. The decision depends on saturation depth, cloud-base trend, and whether drainage flow remains decoupled from the surface.
Case record. cases/v0.3-development-next · v0.3 · aviation · forecast time: Synthetic case issue time: 2026-01-18 18:00 UTC
What the model received
- metar_trend: Over the last three hours, temperature and dew point have converged from 5 C and 2 C to 3 C and 2.5 C. The ceiling has lowered from 1800 ft to 900 ft, visibility has fallen from 8 km to 5 km, and the wind has weakened from 8 kt to 3 kt. Source label: Synthetic author packet
- boundary_layer_and_guidance: The lowest 300 m is nearly saturated, while a weak inversion sits near 700 m. One guidance member mixes the surface layer after 06z and keeps the ceiling above 1000 ft; the others retain light winds and lower the ceiling below 500 ft. No frontal passage is expected before 12z. Source label: Synthetic author packet
- aviation_thresholds: Treat the event target as the airport entering IFR conditions at any point during 06z-12z, defined here as ceiling below 1000 ft or visibility below 3 statute miles. Keep the component ceiling and visibility signals separate in the explanation even though the machine-readable target is the combined event. Source label: Synthetic author packet
What it had to do
Forecast the probability that the generalized airport enters IFR conditions at any point during 06z-12z. Explain whether saturation, drainage, or mixing controls the forecast and identify one observation that would most change the call. End with exactly one fenced JSON block using target_id ifr_by_12z and a numeric probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Use the observed temperature-dewpoint convergence and ceiling trend
- Distinguish boundary-layer saturation from intermittent turbulent mixing
- Treat the combined IFR threshold as a decision target while retaining ceiling and visibility mechanisms
Scoring notes in the case file
- The target combines ceiling and visibility for operational usefulness; the narrative must discuss both components separately.
- This synthetic truth is provisional and must not support a ranking.
Structured target(s). ifr_by_12z (binary, probability)
Across the eight runs. Mean 86.6, median 87, range 84–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 88 | 87 | 84 | 86 | 87 | 87 | 87 |
FB-1106 tropical cyclone rapid intensification
In plain terms. A mature tropical cyclone is approaching a favorable oceanic pocket, but rapid intensification depends on whether dry-air entrainment and moderate vertical shear remain separated from the inner core. The useful forecast distinction is between a favorable environment and a realized structural change over the next day.
Case record. cases/v0.3-development-next · v0.3 · tropical_weather · forecast time: Synthetic case issue time: 2026-08-14 12:00 UTC
What the model received
- environmental_analysis: Maximum sustained wind is 80 kt, the central pressure is 970 hPa, sea surface temperature is 29.5 C, and upper-ocean heat content is 85 kJ cm-2. Deep-layer shear is 15 kt from the southwest. Mid-level relative humidity is 62 percent, with a dry-air tongue 250 km west of the center. Source label: Synthetic author packet
- structure_and_guidance: Microwave imagery shows a partially closed eyewall and a small warm core. Eight of twenty ensemble members predict at least a 25 kt wind increase in the next 24 hours, nine predict 10-24 kt, and three predict less than 10 kt. The rapid-intensification members keep the dry-air tongue outside the radius of maximum wind; the weaker members mix it inward as shear increases. Source label: Synthetic author packet
- observation_options: Candidate updates are a rapid-scan microwave pass, a center-fixing reconnaissance mission, a scatterometer wind swath, and a new upper-air analysis. The most valuable update should discriminate the structural pathway that controls dry-air ingestion, not merely provide another estimate of current intensity. Source label: Synthetic author packet
What it had to do
Forecast the probability of rapid intensification, defined here as at least a 25 kt increase in maximum sustained wind during the next 24 hours. Separate environmental favorability from the probability that the inner-core structure actually supports the increase. Explain which observation would most change the call and why. End with exactly one fenced JSON block using target_id rapid_intensification_24h and a numeric probability. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Separate favorable thermodynamic and shear conditions from realized inner-core organization
- Use ensemble probabilities without treating member counts as perfectly calibrated
- Recognize dry-air entrainment and eyewall structure as coupled uncertainty sources
- Choose an observation that resolves structure and dry-air ingestion
Scoring notes in the case file
- This is a development case; the synthetic outcome is not a promotion-quality truth artifact.
- The probability target is rapid intensification, not a categorical intensity forecast.
Structured target(s). rapid_intensification_24h (binary, probability)
Across the eight runs. Mean 85.0, median 85, range 81–88. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 85 | 85 | 84 | 88 | 84 | 86 | 87 |
FB-1107 tropical cyclone intensity change
In plain terms. A mature tropical cyclone is strengthening over open water. Forecast the maximum sustained wind 24 hours after the issue time from the recent intensity trend and the environmental information in the packet.
Case record. cases/v0.3-development-next · v0.3 · tropical_weather · forecast time: 2025-08-14 12:00 UTC
What the model received
- recent_best_track_trend: The storm's maximum sustained wind was 45 kt at 00z and 06z on 14 August, then increased to 50 kt at 12z. Central pressure fell from 1002 hPa to 999 hPa by 12z. The recent signal is strengthening, but six-hourly intensity estimates should not be treated as independent measurements. Source label: NHC HURDAT2 archive excerpt; issue-time information
- environmental_context: Sea-surface temperature is 29.5 C, upper-ocean heat content is high, and deep-layer vertical shear is about 15 kt. Mid-level relative humidity is about 62 percent. A dry-air tongue is located west of the center, so further strengthening depends on whether the inner core remains separated from that dry air. Source label: Historical case packet framing
What it had to do
Forecast the maximum sustained wind in knots at 12z on 15 August 2025. Explain how the recent pressure and wind trend, oceanic support, shear, and dry-air risk affect the forecast. State one observation that would most change the call. End with exactly one fenced JSON block using target_id maximum_wind_24h_kt and a numeric value. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Use the recent wind and pressure trend without treating six-hourly values as independent evidence
- Balance ocean heat content and modest shear against dry-air entrainment risk
- Express uncertainty in the intensity change rather than inventing a track forecast
Scoring notes in the case file
- The target is a continuous maximum-wind forecast, not a rapid-intensification classification.
- The truth is a post-analysis best-track value; the case is staged until hindsight and information-cutoff review is complete.
Structured target(s). maximum_wind_24h_kt (continuous, knots)
Across the eight runs. Mean 83.8, median 84, range 79–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 85 | 85 | 84 | 79 | 81 | 85 | 84 |
FB-1108 visibility below three miles
In plain terms. Archived airport observations begin with good visibility, but the forecast window tests for later visibility deterioration.
Case record. cases/v0.3-development-next · v0.3 · aviation · forecast time: 2025-07-15 13:00 UTC
What the model received
- archived_observations: 2025-07-15 12:51 UTC: tmpf=78.0, dwpf=66.0, drct=190.0, sknt=3.0, vsby=10.0, skyc1=BKN 2025-07-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=200.0, sknt=4.0, vsby=8.0, skyc1=CLR Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1017-ORD-2025-07-15.csv
What it had to do
Forecast the probability that ORD visibility falls below 3 statute miles before 18z. Explain the competing signals, state the largest uncertainty, and identify one observation that would change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id visibility_below_three_miles and a numeric probability.
What the case says a strong answer should cover
- Use the issue-time visibility, temperature-dewpoint spread, and wind evolution.
- Distinguish issue-time conditions from the later verification window.
Scoring notes in the case file
- Development case generated from an archived observation fixture.
- This case is in the v0.3 development-next authoring pool only; do not treat it as calibration or ranking evidence.
Structured target(s). visibility_below_three_miles (binary, probability)
Across the eight runs. Mean 74.6, median 76, range 68–80. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 73 | 68 | 80 | 79 | 77 | 76 | 76 | 68 |
FB-1109 visibility below three miles
In plain terms. Archived DEN observations at issue time show good visibility; the forecast window tests whether visibility deteriorates afterward.
Case record. cases/v0.3-development-next · v0.3 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- archived_observations: 2025-10-15 12:53 UTC: tmpf=50.0, dwpf=45.0, drct=missing, sknt=3.0, vsby=10.00, skyc1=FEW 2025-10-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=290.0, sknt=2.0, vsby=10.00, skyc1=FEW Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1021-DEN-2025-10-15-visibility.csv
What it had to do
Forecast the probability that DEN visibility falls below 3 statute miles before 17z. Explain the competing signals, state the largest uncertainty, and identify one observation that would change the forecast. End with exactly one fenced JSON block (```json ... ```) using target_id visibility_below_three_miles and a numeric probability.
What the case says a strong answer should cover
- Use the issue-time visibility, temperature-dewpoint spread, and wind evolution.
- Distinguish issue-time conditions from the later verification window.
- Treat a rapid visibility deterioration as an aviation forecast signal without assuming a ceiling height that is not provided.
Scoring notes in the case file
- Development case generated from an archived DEN ASOS observation fixture.
- This case is in the v0.3 development-next authoring pool only; do not treat it as calibration or ranking evidence.
Structured target(s). visibility_below_three_miles (binary, probability)
Across the eight runs. Mean 76.1, median 75, range 71–82. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 73 | 73 | 80 | 73 | 77 | 71 | 82 | 80 |
FB-1110 conditional precipitation type
In plain terms. A shallow subfreezing surface layer is being overrun by warmer air aloft. Precipitation occurrence is fairly likely, but the surface type depends on whether the elevated warm layer melts falling ice completely and how long the near-surface cold layer persists.
Case record. cases/v0.3-development-next · v0.3 · winter_weather · forecast time: Synthetic case issue time: 2026-01-18 18:00 UTC
What the model received
- surface_and_radar_observations: The surface temperature is -1.0 C with a wet-bulb temperature of -1.5 C. Radar echoes are increasing west of the target area, and the leading edge is approaching from the southwest. Surface pressure is falling slowly; the low-level wind is northeast at 8 kt. Source label: Synthetic author packet
- vertical_thermal_profile: A representative profile is -1.4 C at 975 hPa, -0.8 C at 950 hPa, +0.8 C at 925 hPa, +2.5 C at 900 hPa, +1.2 C at 850 hPa, and -1.0 C at 800 hPa. The cloud layer is saturated through an ice-producing layer near -12 to -16 C. Guidance differs on the depth of the warm nose and the rate of low-level cold-air erosion. Source label: Synthetic author packet
- ensemble_and_observation_options: Guidance spans a warmer solution with a deeper melting layer and a colder solution with a longer near-surface layer. A targeted sounding, a wet-bulb observation near the surface, and a radar bright-band trend are available as possible updates. The most useful update should resolve the layer controlling the surface phase, not merely confirm that echoes exist. Source label: Synthetic author packet
What it had to do
First forecast the probability that measurable precipitation occurs in the target area during 19z-01z. Conditional on precipitation occurring, provide probabilities in this exact category order: rain, snow, sleet, freezing rain. Explain which layer controls the type, identify the largest uncertainty, and name one observation that would change the forecast. End with exactly one fenced JSON block (```json ... ```) containing forecasts for target_id precipitation_occurrence_19_01z and target_id conditional_ptype_19_01z.
What the case says a strong answer should cover
- Separate precipitation occurrence from the conditional precipitation-type distribution
- Use the vertical warm-nose and near-surface cold-layer structure rather than surface temperature alone
- Distinguish complete melting followed by refreezing from partial melting and sleet formation
- Treat warm-nose depth and low-level cold-air erosion as competing uncertainties
- Choose an observation that resolves the controlling thermal layer
Scoring notes in the case file
- The conditional category order is rain, snow, sleet, freezing_rain.
- Occurrence and conditional type are separate targets by design.
- This synthetic case is development-only and must not support a model ranking.
Structured target(s). precipitation_occurrence_19_01z (binary, probability), conditional_ptype_19_01z (ordered_categorical, probability_vector)
Across the eight runs. Mean 84.2, median 85, range 81–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 81 | 87 | 81 | 85 | 87 | 85 | 87 |
FB-1111 lake effect snowband placement
In plain terms. A cold-air outbreak over Lake Ontario is likely to produce lake-effect convection, but the useful forecast question is where the dominant band will spend its time after landfall. Small changes in low-level wind direction, boundary-layer depth, and band organization can shift the heavy snow corridor by tens of kilometers.
Case record. cases/v0.3-development-next · v0.3 · winter_weather · forecast time: Synthetic case issue time: 2026-12-15 12:00 UTC
What the model received
- lake_and_thermodynamic_environment: Lake surface temperature is 7 C. The 850-hPa temperature is -14 C, giving strong lake-air instability. Moisture extends through roughly 1.6 km, while an inversion rises from 2.0 km to 2.5 km in the latest guidance. Upstream air is becoming drier above 700 hPa. Source label: Synthetic author packet
- wind_profile_and_ensemble: Representative winds are west-northwest at 12 kt near the surface, 290 degrees at 20 kt near 925 hPa, and 310 degrees at 25 kt near 850 hPa. Eighteen ensemble members produce a persistent snow band somewhere east of the lake. Eight place the primary axis over Tug Hill, six place it farther southeast over the Adirondacks, and four keep the heaviest axis offshore or near the western foothills. The members agree on band occurrence more than band placement. Source label: Synthetic author packet
- observation_options: A lake-facing radar scan, an upstream sounding, a profiler wind retrieval, and a satellite cloud-motion product are available. The most valuable update should distinguish a wind-profile or boundary- layer change that moves the band, rather than merely confirm that lake-effect echoes exist. Source label: Synthetic author packet
What it had to do
Forecast the probability that the Tug Hill target area receives at least 6 inches of snow during the next 24 hours. Also provide probabilities for the primary band footprint in this exact category order: tug_hill, adirondacks, offshore_or_western, multiple_bands. Separate confidence in band occurrence from confidence in band placement, explain which part of the wind and thermal profile controls placement, and name one observation that would change the forecast. End with exactly one fenced JSON block (```json ... ```) containing targets heavy_snow_tughill_24h and primary_band_region_24h.
What the case says a strong answer should cover
- Separate high confidence in lake-effect band occurrence from lower confidence in exact placement
- Use the low-level wind profile and boundary-layer depth to reason about band orientation and landfall
- Treat ensemble location spread as spatial uncertainty rather than evidence that the event will not occur
- Choose an observation that resolves the controlling wind or boundary-layer uncertainty
- Avoid treating the most common ensemble location as a calibrated deterministic answer
Scoring notes in the case file
- The categorical order is tug_hill, adirondacks, offshore_or_western, multiple_bands.
- The synthetic truth is intended to test placement reasoning, not to establish a snow climatology.
- This case is development-only and must not support a model ranking.
Structured target(s). heavy_snow_tughill_24h (binary, probability), primary_band_region_24h (categorical, probability_vector)
Across the eight runs. Mean 85.0, median 85, range 84–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 85 | 85 | 85 | 85 | 84 | 87 | 85 |
FB-1113 fire weather plume audit flag
In plain terms. A dry, windy boundary-layer regime may support rapid fire spread and smoke transport toward populated areas.
Case record. cases/drafts · v0.3 · fire_weather · forecast time: TODO: YYYY-MM-DD HH:MM UTC
What the model received
- fire_environment: Provide fuel condition, live fuel moisture, terrain slope, and recent fire behavior. Source label: Draft author packet
- boundary_layer: Provide low-level wind, humidity, mixing height, and frontal or dryline timing. Source label: Draft author packet
- impact_guidance: Provide generalized population sectors and plume-model spread. Source label: Draft author packet
What it had to do
Forecast the probability of rapid fire growth, the most likely smoke-impact sector, and the timing of the worst visibility. Distinguish fuel-driven uncertainty from wind and mixing uncertainty.
What the case says a strong answer should cover
- fuel dryness and available fuel
- boundary-layer wind and alignment with the fire front
- atmospheric stability and smoke mixing depth
- separating fire growth from downwind impact
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- TODO: define the failure modes and what the grader should reward or penalize.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 76.1, median 76, range 67–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 78 | 67 | 75 | 85 | 79 | 67 | 85 | 73 |
FB-1114 marine fog and wind audit flag
In plain terms. A shallow marine boundary layer may produce coastal fog while a pressure gradient strengthens winds near the same waters.
Case record. cases/drafts · v0.3 · marine · forecast time: TODO: YYYY-MM-DD HH:MM UTC
What the model received
- marine_observations: Provide generalized buoy observations, sea-surface temperature, dew point, and visibility trend. Source label: Draft author packet
- synoptic_pattern: Provide pressure pattern, frontal timing, and low-level wind direction. Source label: Draft author packet
- ensemble_spread: Provide fog-coverage and wind-gust spread by forecast window. Source label: Draft author packet
What it had to do
Forecast the probability of dense marine fog, the onset window, and the most likely wind-risk category. Explain whether advection, mixing, or frontal timing controls the uncertainty.
What the case says a strong answer should cover
- sea-surface temperature and dew-point contrast
- marine inversion depth and turbulent mixing
- advection fog versus frontal fog
- joint but separate fog and wind hazards
Scoring notes in the case file
- Development draft. Do not add this case to a benchmark pool until independently checked.
- TODO: define the failure modes and what the grader should reward or penalize.
Structured target(s). No structured objective target is listed.
Across the eight runs. Mean 60.0, median 79, range 0–85. Zero scores: #252, #253.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 67 | 0 | 0 | 85 | 75 | 85 | 83 | 85 |
FB-1115 marine fog wind separation
In plain terms. Warm, moist air is moving across colder coastal water while a tightening pressure gradient is increasing the near-surface wind. Fog formation is plausible, but stronger mixing could prevent dense fog at the same time that it increases wind risk.
Case record. cases/v0.3-development-next · v0.3 · marine · forecast time: Synthetic case issue time: 2026-02-14 18:00 UTC
What the model received
- marine_observations: Over the last three hours, sea-surface temperature has held near 8 C. Air temperature has risen from 9 C to 11 C while dew point has risen from 7 C to 10 C. Visibility has fallen from 8 km to 4 km in the warm-air-advection sector. Winds have increased from 8 kt to 15 kt. Source label: Synthetic author packet
- boundary_layer_structure: A shallow saturated layer is forecast below 150 m, capped by a weak inversion near 300 m. One guidance member mixes through the inversion after 03z and keeps visibility above 1 km; the other members retain a shallow fog layer while the pressure gradient strengthens. Source label: Synthetic author packet
- marine_thresholds: Treat dense marine fog as visibility below 0.5 nautical miles at any point during 03z-09z. Treat the wind target as the maximum sustained wind category during the same window. Keep fog occurrence and wind risk separate in the explanation. Source label: Synthetic author packet
What it had to do
Forecast the probability of dense marine fog during 03z-09z and the probability distribution for the maximum sustained-wind category. Explain whether warm advection, inversion depth, or turbulent mixing controls the uncertainty. Identify one observation that would most change each call. End with exactly one fenced JSON block using target_id dense_marine_fog_by_09z and max_wind_category_by_09z. Use a numeric probability for the fog target and a probability list for the ordered wind categories. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Warm-air advection over colder water supports fog formation
- A shallow inversion and turbulent mixing can oppose each other
- Fog occurrence and wind risk are separate forecast targets
- Operational thresholds must be applied to the declared window
Scoring notes in the case file
- This is a synthetic development case designed to separate fog occurrence from wind risk.
- Do not interpret this case as real marine-fog skill or a stable model ranking.
Structured target(s). dense_marine_fog_by_09z (binary, probability), max_wind_category_by_09z (ordered_categorical, category probability)
Across the eight runs. Mean 84.8, median 85, range 84–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 84 | 85 | 85 | 84 | 84 | 85 | 86 | 85 |
FB-1117 nocturnal mcs persistence
In plain terms. An evening convective line is weakening as the surface boundary layer stabilizes, but a strengthening nocturnal low-level jet may feed elevated instability and new cells along the cold pool. The central question is whether the system decays, persists, or reorganizes overnight.
Case record. cases/drafts · v0.3 · severe_weather · forecast time: Synthetic case issue time: 2026-06-18 00:00 UTC
What the model received
- evening_convective_state: A broken west-to-east line contains several strong cells but has shown a 30-minute decline in lightning density. The cold pool is advancing east at 22 kt, with a surface temperature drop of 7 F behind the line. Surface-based CAPE is falling rapidly after sunset. Source label: Synthetic author packet
- thermodynamic_profile: At 00z, 850 hPa winds are 25 kt from the south and are forecast to increase to 40 kt by 06z. Elevated CAPE above the stable surface layer rises from 900 to 1500 J/kg, while 0-3 km lapse rates remain modest. The warm sector dew point is 70 F, but the cold pool is 3 F cooler than guidance immediately ahead of the line. Source label: Synthetic author packet
- model_guidance_spread: Four members weaken the line east of the initial boundary by 03z. Five members initiate a secondary band on the nose of the low-level jet, producing a persistent heavy-rain and isolated severe-wind signal from 04z-10z. The remaining members maintain a fast-moving line without upscale organization. Spread is conditional evidence, not a calibrated probability. Source label: Synthetic author packet
- observations_and_triggers: A profiler west of the line shows the low-level jet strengthening, but radar trends do not yet show a new band. The most useful discriminator is whether cells redevelop on the cold-pool nose as the jet reaches 35 kt. A rapidly deepening surface pressure rise behind the line would support a stronger cold pool and faster eastward propagation. Source label: Synthetic author packet
What it had to do
Forecast the probability that an organized convective system persists and produces a meaningful heavy-rain or severe-wind threat during 06z-12z. Distinguish surface-based decay from elevated maintenance, explain whether cold-pool propagation or low-level-jet inflow controls the outcome, identify the single most useful next observation, and describe one bust scenario. End with exactly one fenced JSON block using target_id mcs_persists_06_12z as a binary probability and target_id overnight_hazard_06_12z as an ordered probability list over decay, heavy_rain, and severe_wind. Use exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Surface stabilization can reduce surface-based CAPE without eliminating elevated instability
- A strengthening low-level jet can maintain or regenerate convection along a cold-pool boundary
- Cold-pool speed, boundary position, and jet orientation control propagation and organization
- Ensemble spread is conditional scenario evidence, not an automatically calibrated probability
- Profiler and radar redevelopment trends can discriminate persistence from decay
Scoring notes in the case file
- Development draft. The synthetic outcome is provisional and must be replaced or independently archived before promotion.
- Reward separation of persistence, precipitation impact, and severe-wind risk.
- Do not reward a claim that high CAPE alone guarantees organized convection.
Structured target(s). mcs_persists_06_12z (binary, probability), overnight_hazard_06_12z (ordered_categorical, probability)
Across the eight runs. Mean 84.4, median 84, range 82–86. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 82 | 85 | 84 | 85 | 86 | 85 | 84 | 84 |
FB-1123 den visibility below one mile by 17z
In plain terms. Archived Denver ASOS observations show clear visibility at issue time, followed later by a rapid drop to one quarter mile and then three quarters of a mile with an obscured sky.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- archived_observations: 2025-10-15 12:53 UTC: tmpf=50.0, dwpf=45.0, drct=missing, sknt=3.0, vsby=10.00, skyc1=FEW 2025-10-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=290.0, sknt=2.0, vsby=10.00, skyc1=FEW Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1021-DEN-2025-10-15-visibility.csv
What it had to do
Forecast the probability that visibility at Denver will fall below one statute mile by 17Z. Explain whether the initial clear report should be treated as persistent evidence or whether a rapid deterioration scenario deserves weight. Identify the single next observation that would most change the forecast. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Visibility deterioration can occur faster than persistence baselines suggest.
- The target is a threshold event, not a forecast of the exact minimum visibility.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). den_visibility_below_one_mile_by_17z (binary, probability)
Across the eight runs. Mean 83.1, median 84, range 76–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 79 | 84 | 76 | 80 | 87 | 85 |
FB-1124 den sky condition by 17z
In plain terms. Archived Denver ASOS observations begin with FEW clouds and later report an obscured vertical visibility condition as visibility falls.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- archived_observations: 2025-10-15 12:53 UTC: tmpf=50.0, dwpf=45.0, drct=missing, sknt=3.0, vsby=10.00, skyc1=FEW, skyc1=FEW 2025-10-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=290.0, sknt=2.0, vsby=10.00, skyc1=FEW, skyc1=FEW Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1021-DEN-2025-10-15-visibility.csv
What it had to do
Forecast the probability distribution over the dominant Denver sky condition reported during the verification window ending at 17Z. Use the ordered categories exactly as given, explain what observation would most change the forecast, and distinguish cloud-cover persistence from a developing obscuration event. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- A later obscured-sky report can signal a visibility process that is not captured by clear-sky persistence.
- The categorical distribution must follow the supplied category order and sum to one.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). den_sky_condition_by_17z (ordered_categorical, None)
Across the eight runs. Mean 76.2, median 76, range 66–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 66 | 87 | 73 | 76 | 76 | 76 | 75 | 81 |
FB-1125 den clear to obscured transition by 17z
In plain terms. Archived Denver ASOS observations begin with FEW clouds and later report vertical visibility as visibility collapses.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-10-15 13:00 UTC
What the model received
- archived_observations: 2025-10-15 12:53 UTC: tmpf=50.0, dwpf=45.0, drct=missing, sknt=3.0, vsby=10.00, skyc1=FEW, skyc1=FEW 2025-10-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=290.0, sknt=2.0, vsby=10.00, skyc1=FEW, skyc1=FEW Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1021-DEN-2025-10-15-visibility.csv
What it had to do
Forecast the probability distribution over the two possible dominant Denver sky-condition states during the verification window ending at 17Z: FEW or VV. Use exactly this category order, make the two probabilities sum to one, explain what observation would most change the forecast, and distinguish persistence from a developing obscuration event. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- A clear-to-obscured transition is a categorical event distinct from the visibility threshold itself.
- A two-state probability distribution must follow the supplied order and sum to one.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). den_clear_to_obscured_transition_by_17z (ordered_categorical, None)
Across the eight runs. Mean 81.2, median 81, range 71–87. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 87 | 87 | 87 | 71 | 81 | 76 | 81 | 80 |
FB-1126 erin rapid intensification 24h
In plain terms. NHC HURDAT2 best-track observations show Hurricane Erin at 50 kt at the issue time, with later observations reaching 65 kt by 12Z the next day.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-08-14 12:00 UTC
What the model received
- archived_observations: 2025-08-14 00:00 UTC: max_wind_kt=45, central_pressure_hpa=1002 2025-08-14 06:00 UTC: max_wind_kt=45, central_pressure_hpa=1002 2025-08-14 12:00 UTC: max_wind_kt=50, central_pressure_hpa=999 Source label: NHC HURDAT2 AL052025 best-track archive; issue-time excerpt stored in cases/observations/FB-1107-ERIN-2025-08-15.csv
What it had to do
Forecast the probability that Hurricane Erin intensifies by at least 20 kt during the next 24 hours. Explain the difference between a threshold event and the exact intensity track, identify the main uncertainty in the observation record, and state what future observation would most change the forecast. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Rapid-intensification thresholds are changes over time, not point intensity categories.
- Best-track intensity is useful truth for this toy case but still carries representativeness uncertainty.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Change-target baseline uses the latest numeric pre-issue observation at 2025-08-14 12:00.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). erin_rapid_intensification_24h (binary, probability)
Across the eight runs. Mean 78.9, median 80, range 67–85. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 67 | 85 | 85 | 76 | 76 | 85 | 84 | 73 |
FB-1127 ord visibility below three miles by 18z
In plain terms. Archived ORD ASOS observations show visibility at eight to ten statute miles through the available issue and verification observations.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-07-15 13:00 UTC
What the model received
- archived_observations: 2025-07-15 12:51 UTC: tmpf=78.0, dwpf=66.0, drct=190.0, sknt=3.0, vsby=10.0, skyc1=BKN 2025-07-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=200.0, sknt=4.0, vsby=8.0, skyc1=CLR Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1017-ORD-2025-07-15.csv
What it had to do
Forecast the probability that visibility at Chicago O Hare will fall below three statute miles by 18Z. Explain how persistence should be weighted, identify the observation that would most change the forecast, and do not confuse the threshold event with a precise minimum visibility forecast. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- A negative threshold event is still useful calibration evidence.
- The target is a probability of crossing a threshold, not an exact visibility value.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). ord_visibility_below_three_miles_by_18z (binary, probability)
Across the eight runs. Mean 79.2, median 79, range 73–84. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 81 | 78 | 82 | 84 | 73 | 79 | 79 | 78 |
FB-1128 dsm temperature rise at least 10f by 18z
In plain terms. Forecast whether DSM temperature rises at least 10°F by 18Z.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-01-15 13:00 UTC
What the model received
- archived_observations: 2025-01-15 12:54 UTC: tmpf=7.0, dwpf=0.0, drct=200.0, sknt=7.0, vsby=10.0, skyc1=FEW 2025-01-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=200.0, sknt=4.0, vsby=10.0, skyc1=BKN Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1016-DSM-2025-01-15.csv
What it had to do
What is the probability that the 2-m station temperature at Des Moines (DSM) will increase by at least 10°F from the issue-time observation to 18Z? End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Short-fused boundary-layer temperature evolution and calibration under a strong diurnal warming signal.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). dsm_temperature_rise_at_least_10f_by_18z (binary, probability)
Across the eight runs. Mean 66.2, median 66, range 57–75. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 57 | 63 | 63 | 68 | 63 | 68 | 73 | 75 |
FB-1129 phx temperature rise at least 15f by 18z
In plain terms. Forecast whether PHX temperature rises at least 15°F by 18Z.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-06-15 13:00 UTC
What the model received
- archived_observations: 2025-06-15 12:51 UTC: tmpf=82.0, dwpf=37.0, drct=170.0, sknt=3.0, vsby=10.0, skyc1=CLR 2025-06-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=0.0, sknt=0.0, vsby=10.0, skyc1=CLR Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1020-PHX-2025-06-15.csv
What it had to do
What is the probability that the 2-m station temperature at Phoenix (PHX) will increase by at least 15°F from the latest available pre-issue observation to 18Z? End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- Short-fused desert boundary-layer temperature evolution and calibration under strong diurnal warming.
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Change-target baseline uses the latest numeric pre-issue observation at 2025-06-15 12:51.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). phx_temperature_rise_at_least_15f_by_18z (binary, probability)
Across the eight runs. Mean 66.6, median 67, range 54–81. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 54 | 63 | 81 | 66 | 71 | 68 | 62 | 68 |
FB-1130 sea sky condition by 18z
In plain terms. Archived Seattle ASOS observations begin with an overcast marine-layer cloud state at issue time and show clearing to few clouds near the verification cutoff.
Case record. cases/drafts · v0.4 · aviation · forecast time: 2025-02-15 13:00 UTC
What the model received
- archived_observations: 2025-02-15 12:53 UTC: tmpf=37.0, dwpf=34.0, drct=170.0, sknt=4.0, vsby=10.0, skyc1=BKN 2025-02-15 13:00 UTC: tmpf=missing, dwpf=missing, drct=160.0, sknt=2.0, vsby=10.0, skyc1=OVC Source label: IEM ASOS archive; issue-time excerpt stored in cases/observations/FB-1019-SEA-2025-02-15.csv
What it had to do
Forecast the probability distribution for Seattle's dominant sky condition by 18Z. Treat the overcast issue-time report as persistence evidence, but consider whether a marine-layer clearing transition is plausible. Explain which next observation would most change your forecast. End with exactly one fenced JSON block (```json ... ```) for the machine-readable forecast.
What the case says a strong answer should cover
- marine-layer persistence versus breakup
- cloud-regime transition uncertainty
- observation timing and representativeness
Scoring notes in the case file
- Development draft generated from an archived observation fixture.
- Do not promote until the information cutoff, target definition, baseline, and contamination review are checked.
Structured target(s). sea_sky_condition_by_18z (ordered_categorical, None)
Across the eight runs. Mean 76.0, median 78, range 68–81. Zero scores: none.
| Run | #251 | #252 | #253 | #254 | #255 | #256 | #257 | #258 |
|---|---|---|---|---|---|---|---|---|
| Score | 75 | 71 | 81 | 68 | 81 | 81 | 81 | 70 |
Audit findings and proposal for what comes next
Finding. The response runner and grading pipeline completed the job. The main thing stopping this from being a trustworthy comparison is the content being scored, followed by the judge. A model asking for data that the case never supplied should be treated as a case defect, not as a normal weather miss.
My proposal.
- Use this page to hand-audit all 81 cases. Focus first on the seven flagged cases, then on any case whose score row has a wide spread or a zero.
- Keep every existing case and every old run. Repair incomplete packets in place only as a new recorded revision, with source and truth notes. Do not overwrite the input used in this baseline.
- Make one consistent generation setup for the next panel. Give models a real chance to finish visible answers, and report reasoning settings, token use, retries, and elapsed time beside the score.
- Keep the judge score separate from the probability-verification scores. Avoid self-grading in future comparisons, and record when the judge changes a total so users can see how the final number was calculated.
- After the audit direction is settled, decide whether to run a new version with the repaired 81-case pool or a smaller diagnostic subset. I recommend no new full run until the case fixes and score handling are agreed.
This is a proposal, not work already started. The goal is paused here for your case audit and next direction.
runs/analytics/v0.4-005-case-audit.html. Page generated 2026-09-23 17:35 local time.