ForecastBench v0.4-005-full-panel
v0.4 · 81 shared cases · eight local models · started Sep 22, 2026 08:06 AM Central Daylight Time · refreshed 2026-09-23 17:35 Central Daylight Time
Open the full case-by-case guide and run post-mortem · This baseline dashboard
Baseline coverage and completion notes
All eight models answered the same 81-case pool. 648 of 648 model-case responses have an automated score. The table shows each run’s score coverage, mean, median, range, rubric dimensions, generation wall time, mean per-case latency, and tokens per second.
Qwen 3.8 27B initially produced 72 reasoning-only responses at the configured 1,600-token cap. Retries completed the pool; FB-1116 required a manual recovery (attempt 08) with low reasoning, a 3,072-token output cap, and concise-output instruction. The successful recovered answer is scored. This one case differs from the original generation settings.
FB-1112, FB-1113, FB-1114, FB-1119, FB-1120, FB-1121, and FB-1122. Several have placeholder data rather than a complete forecast packet. Five zero scores occur on FB-1112, FB-1114, and FB-1120 for Gemma 4 12B QAT and Gemma 4 12B; their saved answers ask for the missing inputs. Those zeros reflect unusable case inputs as well as model behavior, so close leaderboard differences are not a clean model comparison.Run-by-run results, scoring and elapsed time
| Run | Model | Status | Answers | Grades | Mean | Wall time | Mean / case | Mean tok/s | Retry history |
|---|---|---|---|---|---|---|---|---|---|
| #251 | google/gemma-4-e4b | Complete | 81/81 | 81/81 | 80.7 | 28.4 min | 20.0s | 63.3 | — |
| #252 | google/gemma-4-12b-qat | Complete | 81/81 | 81/81 | 79.3 | 72.5 min | 45.2s | 33.9 | — |
| #253 | google/gemma-4-12b | Complete | 81/81 | 81/81 | 81.2 | 79.8 min | 49.5s | 31.0 | — |
| #254 | google/gemma-3n-e4b | Complete | 81/81 | 81/81 | 80.8 | 16.0 min | 10.1s | 55.3 | 5 retry attempt(s) recorded |
| #255 | nvidia/nemotron-3-nano-4b | Complete | 81/81 | 81/81 | 82.0 | 19.3 min | 13.6s | 77.6 | 1 retry attempt(s) recorded |
| #256 | nvidia/nemotron-3-nano | Complete | 81/81 | 81/81 | 82.3 | 237.3 min | 158.2s | 8.8 | — |
| #257 | prism-ml/bonsai-27b | Complete | 81/81 | 81/81 | 82.9 | 95.6 min | 52.1s | 33.0 | — |
| #258 | qwen/qwen3.8-27b | Complete | 81/81 | 81/81 | 82.0 | 1047.1 min | 520.5s | 3.1 | 8 retry attempt(s) recorded |
Wall time is run start to last recorded state update. Mean per-case time and tokens/s are from saved LM Studio generation artifacts; they exclude grading time. Qwen’s wall time includes prolonged reasoning-only retries and its manual recovery. Retry entries are historical and do not indicate missing answers.
Automated model scores
Mean judge score (coverage shown at right)
| Model | Mean /100 | Median | Range | Graded cases | Rubric means (points earned / max) |
|---|---|---|---|---|---|
| prism-ml/bonsai-27b | 82.9 / 100 | 85.0 | 62–90 | 81/81 | Meteorology 19.5/25 Reasoning 15.9/20 Hazards 12.3/15 Uncertainty 13.4/15 Usefulness 12.7/15 Communication 9.0/10 |
| nvidia/nemotron-3-nano | 82.3 / 100 | 85.0 | 63–90 | 81/81 | Meteorology 18.8/25 Reasoning 15.7/20 Hazards 12.5/15 Uncertainty 13.5/15 Usefulness 12.8/15 Communication 9.0/10 |
| qwen/qwen3.8-27b | 82.0 / 100 | 84.0 | 37–89 | 81/81 | Meteorology 19.5/25 Reasoning 15.9/20 Hazards 11.7/15 Uncertainty 13.4/15 Usefulness 12.5/15 Communication 9.0/10 |
| nvidia/nemotron-3-nano-4b | 82.0 / 100 | 84.0 | 63–89 | 81/81 | Meteorology 18.8/25 Reasoning 15.8/20 Hazards 12.1/15 Uncertainty 13.6/15 Usefulness 12.7/15 Communication 9.0/10 |
| google/gemma-4-12b | 81.2 / 100 | 85.0 | 0–92 | 81/81 | Meteorology 19.2/25 Reasoning 15.6/20 Hazards 11.9/15 Uncertainty 13.3/15 Usefulness 12.5/15 Communication 8.8/10 |
| google/gemma-3n-e4b | 80.8 / 100 | 84.0 | 59–88 | 81/81 | Meteorology 18.9/25 Reasoning 15.6/20 Hazards 11.3/15 Uncertainty 13.5/15 Usefulness 12.5/15 Communication 9.0/10 |
| google/gemma-4-e4b | 80.7 / 100 | 84.0 | 54–91 | 81/81 | Meteorology 19.0/25 Reasoning 15.5/20 Hazards 11.7/15 Uncertainty 13.4/15 Usefulness 12.2/15 Communication 9.0/10 |
| google/gemma-4-12b-qat | 79.3 / 100 | 85.0 | 0–91 | 81/81 | Meteorology 18.8/25 Reasoning 15.4/20 Hazards 11.2/15 Uncertainty 13.0/15 Usefulness 12.3/15 Communication 8.7/10 |
Completion reliability and objective forecast diagnostics
| Run / model | Initially visible | Reasoning-only | Token-cap exhaustion | Final case completion | Forecast JSON parsed | Objective targets | Brier ↓ | Brier skill ↑ | Ranked probability score ↓ | Ranked probability skill ↑ |
|---|---|---|---|---|---|---|---|---|---|---|
| #251 google/gemma-4-e4b | 81/81 | 0/81 | 0.0% | 100.0% | 45/81 | 53 | 0.2926 | 0.2884 | 0.5332 | 0.0197 |
| #252 google/gemma-4-12b-qat | 38/81 | 43/81 | 53.1% | 100.0% | 47/81 | 55 | 0.2767 | 0.3995 | 0.4433 | -0.2341 |
| #253 google/gemma-4-12b | 31/81 | 50/81 | 61.7% | 100.0% | 47/81 | 55 | 0.2903 | 0.422 | 0.6143 | -0.1888 |
| #254 google/gemma-3n-e4b | 81/81 | 0/81 | 0.0% | 100.0% | 44/81 | 51 | 0.2688 | 0.3889 | 0.5952 | -0.8386 |
| #255 nvidia/nemotron-3-nano-4b | 78/81 | 3/81 | 3.7% | 100.0% | 46/81 | 54 | 0.3154 | 0.1581 | 0.8154 | -0.9119 |
| #256 nvidia/nemotron-3-nano | 62/81 | 19/81 | 23.5% | 100.0% | 47/81 | 55 | 0.2558 | 0.3641 | 0.6033 | -0.2882 |
| #257 prism-ml/bonsai-27b | 1/81 | 80/81 | 98.8% | 100.0% | 47/81 | 55 | 0.2512 | 0.3912 | 0.6587 | -0.3619 |
| #258 qwen/qwen3.8-27b | 9/81 | 72/81 | 88.9% | 100.0% | 46/81 | 53 | 0.2646 | 0.3826 | 0.3505 | 0.1047 |
Primary-output visibility and token-cap exhaustion describe first-attempt generation; final completion includes retries and recorded recovery. Objective diagnostics use a variable number of structured targets per run and provisional, mixed/unverified development-pool truth. Compare cautiously; lower Brier/RPS and higher skill scores are generally preferable.
Grades by case module/topic
| Module | Grades | Mean /100 | Meteorology | Reasoning | Hazards | Uncertainty | Usefulness | Communication |
|---|---|---|---|---|---|---|---|---|
| core_weather | 40 | 78.7 | 19.0/25 | 15.7/20 | 8.9/15 | 13.5/15 | 12.6/15 | 9.0/10 |
| expert_case_studies | 104 | 81.7 | 19.0/25 | 15.5/20 | 12.6/15 | 13.2/15 | 12.4/15 | 8.9/10 |
| intermediate_forecasting | 8 | 83.5 | 20.0/25 | 16.0/20 | 11.8/15 | 14.0/15 | 12.8/15 | 9.0/10 |
| operational_forecasting | 160 | 80.8 | 19.1/25 | 15.8/20 | 11.3/15 | 13.3/15 | 12.4/15 | 8.9/10 |
| specialty_forecasting | 336 | 81.9 | 19.0/25 | 15.7/20 | 12.2/15 | 13.4/15 | 12.6/15 | 8.9/10 |
Topic means pool scored model-case rows across the campaign. “Module” comes from each case’s YAML. The count shows grade coverage; no values are imputed.
Campaign facts
Eight models share benchmark run ID v0.4-005-full-panel and case-set fingerprint ec4d4bed…. Each model used fresh context per case and tools disabled. The 81 cases are the combined pool from v0.1, v0.2, v0.3 development folders, and drafts. This report intentionally shows only this campaign.
FB-1116 needed a case-specific recovery, recorded as attempt 08 under the same model run. The successful retry used supported low reasoning, a 3072-token output cap, and a concise-output instruction after repeated reasoning-only or truncated responses. Its score is included, but that one answer did not use the original generation settings.Current elapsed time: 31h 24m total campaign wall time. Campaign start: Sep 22, 2026 08:06 AM Central Daylight Time. Per-model mean generation latency and run wall time are separate: latency excludes the full run’s setup/gaps.