Weather models and uncertainty: compare like with like

Chapter 8 / 12 · 6 min reading + practice

Understand global and regional forecasts, ensembles, run times and wind definitions; convert disagreement into explicit planning scenarios without inventing safety probabilities.

The purpose of comparing forecasts is not to select the most pleasant number. It is to identify what is robust, what changes between plausible scenarios and which missing information could alter the passage. Numerical models approximate the atmosphere from imperfect observations; uncertainty remains even when the displays agree. An AI explanation can organise this evidence, but it must not invent a probability of a safe voyage or create weather values unsupported by its input.

Global, regional and ensemble forecasts answer different questions

A global model provides the large-scale setting: pressure systems, fronts and evolving flow. Regional higher-resolution models can represent coastlines, terrain and smaller-scale processes in more detail, but rely on boundary information and their own assumptions. More grid points do not guarantee a correctly timed thunderstorm or an accurately modelled harbour gust. Check whether the product actually covers your route and whether a fine-looking app image is displaying native detail or interpolation.

A deterministic forecast is one simulated evolution. An ensemble samples multiple plausible evolutions by varying initial conditions and aspects of the model. Inspect the range and timing of outcomes, not only the ensemble mean, which can smooth important features. An ensemble is not the same as three deterministic forecasts from three applications. Some applications show the same underlying model; counting each screen as independent agreement exaggerates confidence.

Make the comparison table honest

Before comparing, align the location or route segment and valid time, then retain model name and version if available, reference run, received time, lead time, grid or distributed resolution and variable definition. Do not compare an older run’s 48-hour value with a newer run’s 24-hour value and call the difference solely a model disagreement. They started from different information. Conversely, compare successive runs deliberately when evaluating how stable a particular arrival scenario remains.

Run time is not publication time. ECMWF’s current medium-range documentation notes that forecasts become available several hours after their reference times. Avoid telling the crew that a 12:00 run has already been received merely because the clock reads 12:10. Forecast lead time links the reference state to the valid time; report both. Some outputs describe interval maxima or accumulations rather than instantaneous values, so a hourly-looking label does not settle the definition.

Surface wind is not the mast display

A model’s standard 10-metre wind is not automatically a yacht’s masthead measurement. Height, local exposure, small-scale terrain and interpolation affect comparisons. The mast instrument may show apparent wind; reconstructing true wind requires suitable boat-motion and sensor corrections. Gusts and mean wind are also distinct variables with product-specific intervals. Keep their definitions rather than blending them into one wind score. There is no reliable universal percentage that converts every model’s 10 m wind into every yacht’s masthead gust.

Spread does not equal probability

Ensemble spread measures differences among members, commonly with a standard deviation around the ensemble mean. It is not itself the probability of a chosen event. Event probabilities require a specified criterion, time, location and appropriately interpreted member distribution; operational calibration and representativeness still matter. Small spread can coexist with errors shared by all members. Large spread can be practically important even when the mean lies inside a crew’s comfort limits. Agreement is evidence about uncertainty, not certification of safety.

Compare broad pressure and frontal patterns as well as point values. If models place a front at different hours, the relevant issue may be whether it reaches an exposed leg before shelter, not a two-knot difference in their averages. Observations, official warnings and local forecasting discussion can help assess the present pattern. They do not magically eliminate the range of future outcomes. Avoid drawing detailed gust timing from a smoothed pressure chart alone.

Worked example: convert disagreement into a decision problem

These data are invented to demonstrate the method. For an exposed segment at 15:00 UTC, three genuinely different deterministic runs show 10 m mean winds of 18, 22 and 31 knots. Their issue ages, valid times and averaging definitions have been checked. A crew’s selected planning limit is 25 knots mean wind; it is personal to this exercise, not a universal safety threshold. A report of 23.7 knots from averaging the three would conceal that one plausible scenario exceeds the chosen limit.

The better summary is: two scenarios remain below that particular mean-wind limit, one exceeds it; inspect gusts, sea state and frontal timing separately. Three deterministic solutions are not enough to claim a calibrated one-in-three event probability. Suppose an additional hypothetical ensemble of 20 comparable members has six exceeding the same specified criterion at that point and time. The raw member fraction is 30%, not a 30% probability of an unsafe voyage. Report it as a model-derived exceedance fraction with its scope and limitations.

The team next compares an earlier departure, slower boat speed and a sheltered alternative. If a front may arrive three hours earlier than the central scenario, does the yacht still have a practical alternative before the exposed leg? Preserve that question and the source-specific outcomes. Do not average incompatible wave-period definitions, substitute an unmeasured current or rank an alternative without accounting for its approach. A good brief makes the sensitivity visible rather than returning one reassuring verdict.

A fair application comparison

To compare tools, freeze the route, requested date, model runs and parameters. Record what each actually exposes: model provenance, warning coverage, wind/sea components, ensemble information, receipt time, offline export and uncertainty wording. Then run the same task and preserve the output. A feature list or a personal impression is not a measured accuracy test. Accuracy requires a defined target, suitable independent observations, enough cases and a reproducible scoring method. Separate meteorological skill from interface convenience.

Checklist and useful feedback

  1. Align valid times, locations, units, components and variable definitions.
  2. Retain each reference run, lead time, receipt time and actual coverage.
  3. Identify shared model lineage before counting agreement.
  4. Inspect the range, tails and timing, not only the mean.
  5. Distinguish a personal-limit exceedance from an official warning or a safety judgement.
  6. Recompute the route for timing shifts and speed loss, preserving alternatives and data gaps.
  7. Recheck before the next irreversible decision; compare measured outcomes after the trip.

In feedback, save what was forecast before departure, not only the latest retrospective reconstruction. Record measurement quality and the boat’s actual route. A wrong sentence and a wrong weather field are different failures: correcting the language model will not fix missing model data. Keep the calculation, explanation and user experience separable so each can improve honestly.

Quick chapter check

Choose the single best answer for every question. The check shows explanations, important gaps and chapters to revisit. Results stay in this tab: analytics are disabled on this page, answers are not sent and do not survive closing or refreshing it.

1. Three comparable deterministic runs show 18, 22 and 31 knots mean wind. The crew selected a 25-knot planning limit. What is the most useful summary?

2. Three apps display the same underlying model run for the same point and time. How should their matching wind values be counted?

3. Six of 20 comparable ensemble members exceed a specified wind limit at one point and valid time. Which wording is justified without further calibration?

Automatic scoring requires JavaScript. You can still read the questions and check the explanations below without sending data.

1. Show explanation

Two runs are below and one above this selected limit; retain the adverse scenario and examine timing, gusts and sea state separately. Averaging to about 23.7 knots conceals the 31-knot scenario. A personal-limit exceedance is neither an official warning nor a certified safety judgement. Three deterministic runs also do not establish a calibrated voyage probability.

2. Show explanation

As shared model guidance displayed three ways, not three independent confirmations. App names do not establish independent meteorological inputs. Identify model lineage, run and variable definitions before counting agreement. Even genuinely different models can share errors, so agreement never certifies a passage.

3. Show explanation

The raw member exceedance fraction is 30% for that criterion, point and time; it is not a 30% probability of an unsafe voyage. Six divided by 20 is 30%, but the scope is the defined event at a particular point and time. Calibration, representativeness and the full route still matter. A member fraction must not be relabelled as a probability of overall voyage safety.

Check your understanding

First explain your answer in your own words, then open the explanation. These are original Sailing Weather exercises, not official exam questions.

1. Three apps agree. Have three independent weather scenarios been checked?

Not necessarily: they may share a model, run and post-processing. Check provenance before counting agreement.

2. Six of 20 synthetic members exceed a mean-wind criterion. What can be reported?

A raw 30% member exceedance fraction for that criterion, place and time, with interpretation limits—not a 30% probability of an unsafe voyage.

3. Can apparent masthead wind be directly scored against a model's 10 m mean wind?

No. Resolve true/apparent wind, height, motion correction, exposure, variable and averaging interval first.

4. All models have a small spread. Does the report get to state the passage is safe?

No. Shared errors, unresolved local effects, warnings, yacht/crew constraints and route exposure remain.

Sources and verification scope