WEATHER FORECAST COMPARISON RECORD - YENRA Guide: https://yenra.com/a/weather-ai.html Worksheet edition: September 10, 2026 Choose one location, variable, valid period and lead time. Save forecasts before observing the outcome. Duplicate the record for each case. This is a learning and evaluation aid, not a basis for overriding official warnings. COMPARISON DEFINITION Question and intended decision: Location / geographic scale: Variable / units / level: Valid period and time zone: Forecast lead time: Baseline model or method: Model names and exact versions: Observation source and its limitations: Evaluation period / missing-data rule: Metric and reason for choosing it: Any threshold and definition of an event: CASE RECORD Forecast issue or initialization time (with time zone): Valid time or period (with time zone): Forecast A / source / version: Forecast B / source / version: Observed value / source: Absolute error A = absolute value of (A minus observation): Absolute error B = absolute value of (B minus observation): Event occurred? / forecast event probability, if relevant: Data gaps, quality concerns or unusual conditions: FICTIONAL TEMPERATURE EXAMPLE - DAILY MAXIMA IN DEGREES CELSIUS Observations: 20, 24, 28 Model A: 21, 22, 27 Model B: 20, 25, 31 Absolute A: 1, 2, 1 Absolute B: 0, 1, 3 Mean absolute error A = (1 + 2 + 1) / 3 = 1.33 C (rounded). Mean absolute error B = (0 + 1 + 3) / 3 = 1.33 C (rounded). At a fictional 30 C threshold, B predicts an event on day 3; the observed maximum remains below it. Equal average errors can hide different patterns. Three invented days establish no real model ranking or rare-event skill. PROBABILITY REVIEW Event definition / threshold / location / period: Number of comparable cases: Probability group (for example forecasts near 30%): Number of observed events / number of cases in group: Observed event frequency: Sampling limitations and uncertainty: Was the system useful as well as calibrated? REVIEW Main result, metric and sample size: Extreme or high-impact cases considered separately: What the evidence supports: What remains uncertain: Next review trigger, such as a model upgrade: Source context: original GraphCast/AIFS papers and ECMWF operational documentation linked in the guide. Retain the exact versions evaluated.