Helix · /model briefing

What the page is actually saying.

A walk through every block on /model after the this-cycle uncertainty layout: what each piece is for, how it is measured, and the values sitting there now. The page is not a leaderboard. It is a width report, with the line still scored so a tight swarm cannot hide a miss.

Experimental. Helix is not an official forecast. Do not use it for any safety, evacuation, or marine decision. Official information: hurricane.noaa.gov.

Loading live numbers…

On this briefing
  1. The opening claim
  2. The map and the lead scrubber
  3. The sidebar: this cycle’s uncertainty
  4. The line this cycle
  5. Live verification, this storm
  6. Width and the line, this season
  7. The line, every lead
  8. The frozen holdout (how Helix was measured)

1. The opening claim

What it is for

The hero is the argument of the page: “How unsure is this forecast.” NHC’s cone is last year’s typical official miss — the same radius on every Atlantic or East Pacific storm, whether the models agree or split. Helix’s opening is the width of this cycle’s 50-member ECMWF swarm, not a claim that Helix’s center line beats NHC. The three “How to read” cards under the headline are the whole product: official cone, cyan ellipse, then the line as a separate bet.

How it is measured

Three widths can appear and they are not interchangeable.

  • NHC cone — 2026 operational 2/3-probability circles from 2021–2025 official errors. East Pacific 48 h is 56 n mi; Atlantic 48 h is 62 n mi. Same number every storm in that basin. On by default.
  • ECMWF 1σ ellipse — standard deviation of member latitude and longitude around the ensemble mean, drawn as an ellipse (not a circle) on the Helix point at the selected lead. About 68% of the 50 perturbed members plus control sit inside that ellipse. Not a historical error circle, and not the same probability as the official cone.
  • Helix cone (off by default) — still last year’s typical Helix miss (holdout p67). We have not replaced it with the swarm. The page says so.
What it is not claiming

Helix is not “better than NHC.” On the frozen 2024–2025 holdout it is statistically level with OFCL and HCCA through 72 h. The only significant line edge is ~1.5 n mi at 24 h. The 2026 in-season table is mostly v1 forecasts and trails NHC. The thing worth showing is situation-dependent width.

2. The map and the lead scrubber

On /model this is the left card titled “This cycle’s uncertainty,” with the 24 / 48 / 72 / 120 h buttons in the header.

What it is for

Let a desk meteorologist see, at one lead, the official last-year cone next to this run’s ECMWF ellipse — then step the lead without reloading. Named models stay off so the first look is width, not spaghetti. The faint blue cloud is the interpolated aids Helix actually used, not the 50 individual ENS tracks (those are not stored).

How it is measured
  • Default layers — official cone on, cyan 1σ ellipse on, guidance cloud on, Helix cone off, named guidance off. Consensus lines (Helix, OFCL, TVCN, HCCA) stay available; they are not the opening.
  • Scrubber — 24, 48, 72, 120 h. It restyles the ellipse at that lead, updates the sidebar sentence, and keeps the same cycle. It does not change the official cone radii; those are still last year’s numbers for that basin and lead.
  • Ellipse, not a ring — the overlay uses sd_lat_deg and sd_lon_deg from the decoded 50-member swarm. A circle would lie when the swarm is elongated along or across track.
  • Member count on the tooltip — “n of 51.” If members have already dropped the storm, the ellipse is thinning. The sidebar warns below 20 members and refuses to treat it as a full-swarm width below 10.
What it is not claiming

The cyan ellipse is not a watch/warning graphic and is not Helix’s own error cone. Close in spirit to a 2/3 circle, not the same probability. Fifty individual ensemble tracks are not on the map because the live JSON stores the control, the mean, and the per-lead spread statistics — not the raw member polylines.

3. The sidebar: this cycle’s uncertainty

On /model this is the cyan card to the right of the map, plus the first table (“This cycle vs the fixed cone”). The lead scrubber drives both.

What it is for

Give a seasoned weatherman a one-pass read: is this cycle tighter or wider than the official cone, is the disagreement speed or left/right, and is intensity also split. Then keep the older guidance-clustering score, gated so a two-member “100” cannot pose as tight agreement.

How it is measured
  • Air sentence — the first paragraph. At the selected lead: ECMWF 1σ in n mi, the basin’s official cone, tighter/wider and percent of that cone, intensity ± member σ, and control-versus-ensemble along/cross in words. Plus the member-count warning when it applies.
  • Swarm vs cone — the big number is this run’s 1σ. Green/red is that radius minus the 2026 operational cone for the basin. 1σ ≈ 68% of members; the official cone is a historical 2/3 official-error circle.
  • Along-track / cross-track — ECMWF control versus ensemble mean, decomposed relative to the control heading. Along is faster/slower (timing). Cross is left/right of heading (location). The small plot is that vector, not the 50-member swarm.
  • Intensity — Helix vmax at that lead, plus the ECMWF member standard deviation of vmax when stored. A tight track ellipse with a fat intensity σ is still an uncertain forecast.
  • Bifurcation note — if control and mean are farther apart than 75% of the swarm width, or the guidance cloud’s pairwise distances look bimodal, the page says a 1σ ellipse can hide two good solutions. Otherwise it says the cloud looks unimodal. That is a diagnostic, not a cluster analysis of the 50 members.
  • Guidance clustering — 0–100 from mean pairwise distance of the interpolated aids Helix used at 48 h. Withheld if fewer than four aids. Not the 50-member swarm.

The table under the map subtracts 1σ from the official cone at every stored lead. Green = this cycle’s swarm is tighter than last year’s typical miss. The bar chart is the same comparison drawn.

What it found · live now
Loading this-cycle spread…

4. The line this cycle

On /model: “Consensus tracks” and “Guidance this cycle.” Demoted under the width table on purpose.

What it is for

The center-line positions — Helix, NHC official (OFCL), TVCN, HCCA, and the aids that fed this run. Still useful. Not the opening.

How it is measured

Helix is a gradient-boosted residual on a multi-model consensus (v3.0.0). Inputs at advisory time only: public NHC a-deck interpolated aids, ECMWF open-data tracks relocated the way NHC builds EMXI, ECMWF control-versus-ensemble geometry, and SHIPS environment. TVCN is NHC’s track consensus. HCCA is the HFIP corrected consensus. AIFS is shown, not scored, not a trained feature.

What it found · live now
Loading this-cycle tracks…

5. Live verification, this storm

What it is for

Once a best-track fix arrives, score the stored Helix forecast and the OFCL / TVCN / HCCA values from the same a-deck against that fix. A running, public scoreboard — not a recap written after the season.

How it is measured

Every advisory-cycle run is kept under a cycle-keyed JSON and never overwritten. When the operational b-deck has a fix at the verifying time, Helix computes haversine track error (n mi) and absolute intensity error (kt). Homogeneous columns use only cycles where Helix and all three baselines verified, so the four numbers are comparable. “Helix beat NHC” is the share of individual forecasts where Helix’s track error was smaller. In-season b-decks are operational best tracks, not post-season reanalysis.

What it found · active storms
Loading verification…

6. Width and the line, this season

On /model this is the storm-by-storm block with the three tiles on top. It is the seasonal version of the same two bets.

What it is for

Ask, for every 2026 storm: did this cycle’s guidance (or ECMWF swarm) sit inside last year’s official cone — and, separately, was Helix’s center line closer than NHC / TVCN / HCCA? Those are different questions. The page reports both so nobody pretends a tight swarm is a line win.

How it is measured
  • Width — prefer ECMWF 1σ when that cycle stored it. Otherwise the 48 h mean pairwise distance of the aids Helix used (the agreement number). Compare that radius to the NHC cone for that basin and lead. A cycle is “tight” if width < cone, “split” if not.
  • The line — homogeneous 48 h track MAE for Helix, OFCL, TVCN, HCCA, plus the share of cycles Helix beat OFCL.
  • Different bets — Helix-beat-NHC rate on tight cycles versus split cycles. If those rates are the same, calling the width did not buy a line win. The calibration note under the table is that comparison in one paragraph.

Most 2026 cycles predate ECMWF ingest. Open data has no archive, so those width numbers are guidance spread, not the 50-member swarm. The page says that.

What it found · 2026 so far
Loading season width vs line…

7. The line, every lead

On /model this is the last table: season-to-date MAE at 12–120 h, plus the v1 disclosure.

What it is for

The traditional scoreboard. Keep it. Do not lead with it. Most 2026 forecasts were Helix v1, which trained on inputs that were not available live and was withdrawn. That table is not the holdout, and it is not a v3 claim.

How it is measured

Same verifying protocol as §5, pooled across the season. Homogeneous n is the only fair four-way comparison. The note under the table splits v1 / v2 / v3 counts and ECMWF coverage.

What it found · 2026 line (read the version mix)
Loading season scorecard…

8. The frozen holdout — how Helix itself was measured

This is not a /model table. It is why we are allowed to put Helix on the page at all. Full protocol and CIs: whitepaper.

What it is for

A locked test: train 2018–2023, evaluate 2024–2025 without retuning. 64 storms, 1573 cycles. The question is not “did Helix beat NHC this afternoon.” It is “on a season Helix has never seen, is the line even in the same league as OFCL and HCCA?”

How it is measured

Homogeneous track MAE at each lead — same cycles for Helix, OFCL, TVCN, HCCA. Storm-block bootstrap, 2000 resamples, 95% percentile CI of the mean paired difference (Helix minus baseline). Negative favors Helix. A difference whose interval includes 0 is a tie. Intensity is scored the same way; TVCN has no intensity.

What it found · Helix v3.0.0 vs the field
24 h vs OFCL
−1.5 n mi
The only significant line edge. CI 0.3 to 3.0 n mi better than official. n = 1105.
12–72 h beat rate
53.7%
Share of cycles where Helix’s mean 12–72 h track error was smaller than OFCL. n = 1195. A coin flip with a slight lean.
vs HCCA, every lead
tie
Every lead’s CI versus HCCA includes 0. Re-weighting aids NHC already sees is at a ceiling.
LeadnHelixNHCTVCNHCCAHelix − NHCCI vs NHC
12 h118721.321.821.921.4−0.5−1.0 to −0.0
24 h110532.133.734.532.5−1.5−3.0 to −0.3 · sig
36 h101543.544.947.843.7−1.3−3.7 to +0.6 · tie
48 h92155.456.261.355.7−0.8−4.0 to +2.4 · tie
72 h71484.184.393.785.5−0.3−5.4 to +5.0 · tie
96 h564130.2121.3137.7129.7+8.9−0.4 to +19.0 · NHC better
120 h418190.3179.8198.2194.0+10.5NHC better; vs HCCA a tie

Track error, n mi, homogeneous 2024–2025 holdout, Helix v3.0.0. Intensity is a wash after 12 h (Helix is worse at 12 h). SHIPS is missing on the holdout — CIRA files for 2024–25 are not posted — so the unique environmental inputs do not get a clean holdout test. Anyone quoting Helix as better than NHC is misquoting this table.

How to say this in one paragraph

Helix is a public consensus track that is level with NHC and HCCA on a frozen holdout, not ahead of them. The /model page is therefore not a leaderboard. It is a this-cycle uncertainty report: official last-year cone on, this run’s ECMWF 1σ ellipse on, along/cross and intensity spread in the sidebar, named models off until you ask. The line is still scored so the width claim cannot hide a miss. On 2026 so far, about a third of 48 h cycles were tighter than the official cone, Helix’s line beat NHC on about 29% of homogeneous cycles, and those two rates barely move together. Width and the line are different bets. That is the value.

Open /model · Whitepaper + PDF