Single-campaign report · full baseline

ForecastBench v0.4-005-full-panel

v0.4 · 81 shared cases · eight local models · started Sep 22, 2026 08:06 AM Central Daylight Time · refreshed 2026-09-23 17:35 Central Daylight Time

Open the full case-by-case guide and run post-mortem · This baseline dashboard

8/8model runs complete
648/648model-case responses recorded
648/648model-case grades completed
31h 24m total campaign wall timecampaign wall time

Baseline coverage and completion notes

All eight models answered the same 81-case pool. 648 of 648 model-case responses have an automated score. The table shows each run’s score coverage, mean, median, range, rubric dimensions, generation wall time, mean per-case latency, and tokens per second.

Qwen 3.8 27B initially produced 72 reasoning-only responses at the configured 1,600-token cap. Retries completed the pool; FB-1116 required a manual recovery (attempt 08) with low reasoning, a 3,072-token output cap, and concise-output instruction. The successful recovered answer is scored. This one case differs from the original generation settings.

Case-quality caveat: the retained full pool includes seven unfinished or provisional draft cases: FB-1112, FB-1113, FB-1114, FB-1119, FB-1120, FB-1121, and FB-1122. Several have placeholder data rather than a complete forecast packet. Five zero scores occur on FB-1112, FB-1114, and FB-1120 for Gemma 4 12B QAT and Gemma 4 12B; their saved answers ask for the missing inputs. Those zeros reflect unusable case inputs as well as model behavior, so close leaderboard differences are not a clean model comparison.

Run-by-run results, scoring and elapsed time

RunModelStatusAnswersGradesMeanWall timeMean / caseMean tok/sRetry history
#251google/gemma-4-e4bComplete81/8181/8180.728.4 min20.0s63.3—
#252google/gemma-4-12b-qatComplete81/8181/8179.372.5 min45.2s33.9—
#253google/gemma-4-12bComplete81/8181/8181.279.8 min49.5s31.0—
#254google/gemma-3n-e4bComplete81/8181/8180.816.0 min10.1s55.35 retry attempt(s) recorded
#255nvidia/nemotron-3-nano-4bComplete81/8181/8182.019.3 min13.6s77.61 retry attempt(s) recorded
#256nvidia/nemotron-3-nanoComplete81/8181/8182.3237.3 min158.2s8.8—
#257prism-ml/bonsai-27bComplete81/8181/8182.995.6 min52.1s33.0—
#258qwen/qwen3.8-27bComplete81/8181/8182.01047.1 min520.5s3.18 retry attempt(s) recorded

Wall time is run start to last recorded state update. Mean per-case time and tokens/s are from saved LM Studio generation artifacts; they exclude grading time. Qwen’s wall time includes prolonged reasoning-only retries and its manual recovery. Retry entries are historical and do not indicate missing answers.

Automated model scores

Mean judge score (coverage shown at right)

prism-ml/bonsai-27b
82.981/81 graded
nvidia/nemotron-3-nano
82.381/81 graded
qwen/qwen3.8-27b
82.081/81 graded
nvidia/nemotron-3-nano-4b
82.081/81 graded
google/gemma-4-12b
81.281/81 graded
google/gemma-3n-e4b
80.881/81 graded
google/gemma-4-e4b
80.781/81 graded
google/gemma-4-12b-qat
79.381/81 graded
ModelMean /100MedianRangeGraded casesRubric means (points earned / max)
prism-ml/bonsai-27b82.9 / 10085.062–9081/81Meteorology 19.5/25 Reasoning 15.9/20 Hazards 12.3/15 Uncertainty 13.4/15 Usefulness 12.7/15 Communication 9.0/10
nvidia/nemotron-3-nano82.3 / 10085.063–9081/81Meteorology 18.8/25 Reasoning 15.7/20 Hazards 12.5/15 Uncertainty 13.5/15 Usefulness 12.8/15 Communication 9.0/10
qwen/qwen3.8-27b82.0 / 10084.037–8981/81Meteorology 19.5/25 Reasoning 15.9/20 Hazards 11.7/15 Uncertainty 13.4/15 Usefulness 12.5/15 Communication 9.0/10
nvidia/nemotron-3-nano-4b82.0 / 10084.063–8981/81Meteorology 18.8/25 Reasoning 15.8/20 Hazards 12.1/15 Uncertainty 13.6/15 Usefulness 12.7/15 Communication 9.0/10
google/gemma-4-12b81.2 / 10085.00–9281/81Meteorology 19.2/25 Reasoning 15.6/20 Hazards 11.9/15 Uncertainty 13.3/15 Usefulness 12.5/15 Communication 8.8/10
google/gemma-3n-e4b80.8 / 10084.059–8881/81Meteorology 18.9/25 Reasoning 15.6/20 Hazards 11.3/15 Uncertainty 13.5/15 Usefulness 12.5/15 Communication 9.0/10
google/gemma-4-e4b80.7 / 10084.054–9181/81Meteorology 19.0/25 Reasoning 15.5/20 Hazards 11.7/15 Uncertainty 13.4/15 Usefulness 12.2/15 Communication 9.0/10
google/gemma-4-12b-qat79.3 / 10085.00–9181/81Meteorology 18.8/25 Reasoning 15.4/20 Hazards 11.2/15 Uncertainty 13.0/15 Usefulness 12.3/15 Communication 8.7/10
All scores use the local LM Studio Gemma 3n E4B judge and are model-based rubric assessments, not observation-verification scores. Gemma 3n E4B graded its own run (#254), so that row is self-judged. Treat close score differences as suggestive, not decisive; the judge can be generous.

Completion reliability and objective forecast diagnostics

Run / modelInitially visibleReasoning-onlyToken-cap exhaustionFinal case completionForecast JSON parsedObjective targetsBrier ↓Brier skill ↑Ranked probability score ↓Ranked probability skill ↑
#251 google/gemma-4-e4b81/810/810.0%100.0%45/81530.29260.28840.53320.0197
#252 google/gemma-4-12b-qat38/8143/8153.1%100.0%47/81550.27670.39950.4433-0.2341
#253 google/gemma-4-12b31/8150/8161.7%100.0%47/81550.29030.4220.6143-0.1888
#254 google/gemma-3n-e4b81/810/810.0%100.0%44/81510.26880.38890.5952-0.8386
#255 nvidia/nemotron-3-nano-4b78/813/813.7%100.0%46/81540.31540.15810.8154-0.9119
#256 nvidia/nemotron-3-nano62/8119/8123.5%100.0%47/81550.25580.36410.6033-0.2882
#257 prism-ml/bonsai-27b1/8180/8198.8%100.0%47/81550.25120.39120.6587-0.3619
#258 qwen/qwen3.8-27b9/8172/8188.9%100.0%46/81530.26460.38260.35050.1047

Primary-output visibility and token-cap exhaustion describe first-attempt generation; final completion includes retries and recorded recovery. Objective diagnostics use a variable number of structured targets per run and provisional, mixed/unverified development-pool truth. Compare cautiously; lower Brier/RPS and higher skill scores are generally preferable.

Grades by case module/topic

ModuleGradesMean /100MeteorologyReasoningHazardsUncertaintyUsefulnessCommunication
core_weather4078.719.0/2515.7/208.9/1513.5/1512.6/159.0/10
expert_case_studies10481.719.0/2515.5/2012.6/1513.2/1512.4/158.9/10
intermediate_forecasting883.520.0/2516.0/2011.8/1514.0/1512.8/159.0/10
operational_forecasting16080.819.1/2515.8/2011.3/1513.3/1512.4/158.9/10
specialty_forecasting33681.919.0/2515.7/2012.2/1513.4/1512.6/158.9/10

Topic means pool scored model-case rows across the campaign. “Module” comes from each case’s YAML. The count shows grade coverage; no values are imputed.

Campaign facts

Eight models share benchmark run ID v0.4-005-full-panel and case-set fingerprint ec4d4bed…. Each model used fresh context per case and tools disabled. The 81 cases are the combined pool from v0.1, v0.2, v0.3 development folders, and drafts. This report intentionally shows only this campaign.

Protocol exception: Qwen’s FB-1116 needed a case-specific recovery, recorded as attempt 08 under the same model run. The successful retry used supported low reasoning, a 3072-token output cap, and a concise-output instruction after repeated reasoning-only or truncated responses. Its score is included, but that one answer did not use the original generation settings.

Current elapsed time: 31h 24m total campaign wall time. Campaign start: Sep 22, 2026 08:06 AM Central Daylight Time. Per-model mean generation latency and run wall time are separate: latency excludes the full run’s setup/gaps.

Generated locally from preserved run states, grading CSVs, and case YAML. This page is a snapshot that refreshes when the report builder runs. No model response text is included.