A walk through every block on /model after the this-cycle uncertainty layout: what each piece is for, how it is measured, and the values sitting there now. The page is not a leaderboard. It is a width report, with the line still scored so a tight swarm cannot hide a miss.
The hero is the argument of the page: “How unsure is this forecast.” NHC’s cone is last year’s typical official miss — the same radius on every Atlantic or East Pacific storm, whether the models agree or split. Helix’s opening is the width of this cycle’s 50-member ECMWF swarm, not a claim that Helix’s center line beats NHC. The three “How to read” cards under the headline are the whole product: official cone, cyan ellipse, then the line as a separate bet.
Three widths can appear and they are not interchangeable.
Helix is not “better than NHC.” On the frozen 2024–2025 holdout it is statistically level with OFCL and HCCA through 72 h. The only significant line edge is ~1.5 n mi at 24 h. The 2026 in-season table is mostly v1 forecasts and trails NHC. The thing worth showing is situation-dependent width.
On /model this is the left card titled “This cycle’s uncertainty,” with the 24 / 48 / 72 / 120 h buttons in the header.
Let a desk meteorologist see, at one lead, the official last-year cone next to this run’s ECMWF ellipse — then step the lead without reloading. Named models stay off so the first look is width, not spaghetti. The faint blue cloud is the interpolated aids Helix actually used, not the 50 individual ENS tracks (those are not stored).
The cyan ellipse is not a watch/warning graphic and is not Helix’s own error cone. Close in spirit to a 2/3 circle, not the same probability. Fifty individual ensemble tracks are not on the map because the live JSON stores the control, the mean, and the per-lead spread statistics — not the raw member polylines.
On /model this is the cyan card to the right of the map, plus the first table (“This cycle vs the fixed cone”). The lead scrubber drives both.
Give a seasoned weatherman a one-pass read: is this cycle tighter or wider than the official cone, is the disagreement speed or left/right, and is intensity also split. Then keep the older guidance-clustering score, gated so a two-member “100” cannot pose as tight agreement.
The table under the map subtracts 1σ from the official cone at every stored lead. Green = this cycle’s swarm is tighter than last year’s typical miss. The bar chart is the same comparison drawn.
On /model: “Consensus tracks” and “Guidance this cycle.” Demoted under the width table on purpose.
The center-line positions — Helix, NHC official (OFCL), TVCN, HCCA, and the aids that fed this run. Still useful. Not the opening.
Helix is a gradient-boosted residual on a multi-model consensus (v3.0.0). Inputs at advisory time only: public NHC a-deck interpolated aids, ECMWF open-data tracks relocated the way NHC builds EMXI, ECMWF control-versus-ensemble geometry, and SHIPS environment. TVCN is NHC’s track consensus. HCCA is the HFIP corrected consensus. AIFS is shown, not scored, not a trained feature.
Once a best-track fix arrives, score the stored Helix forecast and the OFCL / TVCN / HCCA values from the same a-deck against that fix. A running, public scoreboard — not a recap written after the season.
Every advisory-cycle run is kept under a cycle-keyed JSON and never overwritten. When the operational b-deck has a fix at the verifying time, Helix computes haversine track error (n mi) and absolute intensity error (kt). Homogeneous columns use only cycles where Helix and all three baselines verified, so the four numbers are comparable. “Helix beat NHC” is the share of individual forecasts where Helix’s track error was smaller. In-season b-decks are operational best tracks, not post-season reanalysis.
On /model this is the storm-by-storm block with the three tiles on top. It is the seasonal version of the same two bets.
Ask, for every 2026 storm: did this cycle’s guidance (or ECMWF swarm) sit inside last year’s official cone — and, separately, was Helix’s center line closer than NHC / TVCN / HCCA? Those are different questions. The page reports both so nobody pretends a tight swarm is a line win.
Most 2026 cycles predate ECMWF ingest. Open data has no archive, so those width numbers are guidance spread, not the 50-member swarm. The page says that.
On /model this is the last table: season-to-date MAE at 12–120 h, plus the v1 disclosure.
The traditional scoreboard. Keep it. Do not lead with it. Most 2026 forecasts were Helix v1, which trained on inputs that were not available live and was withdrawn. That table is not the holdout, and it is not a v3 claim.
Same verifying protocol as §5, pooled across the season. Homogeneous n is the only fair four-way comparison. The note under the table splits v1 / v2 / v3 counts and ECMWF coverage.
This is not a /model table. It is why we are allowed to put Helix on the page at all. Full protocol and CIs: whitepaper.
A locked test: train 2018–2023, evaluate 2024–2025 without retuning. 64 storms, 1573 cycles. The question is not “did Helix beat NHC this afternoon.” It is “on a season Helix has never seen, is the line even in the same league as OFCL and HCCA?”
Homogeneous track MAE at each lead — same cycles for Helix, OFCL, TVCN, HCCA. Storm-block bootstrap, 2000 resamples, 95% percentile CI of the mean paired difference (Helix minus baseline). Negative favors Helix. A difference whose interval includes 0 is a tie. Intensity is scored the same way; TVCN has no intensity.
| Lead | n | Helix | NHC | TVCN | HCCA | Helix − NHC | CI vs NHC |
|---|---|---|---|---|---|---|---|
| 12 h | 1187 | 21.3 | 21.8 | 21.9 | 21.4 | −0.5 | −1.0 to −0.0 |
| 24 h | 1105 | 32.1 | 33.7 | 34.5 | 32.5 | −1.5 | −3.0 to −0.3 · sig |
| 36 h | 1015 | 43.5 | 44.9 | 47.8 | 43.7 | −1.3 | −3.7 to +0.6 · tie |
| 48 h | 921 | 55.4 | 56.2 | 61.3 | 55.7 | −0.8 | −4.0 to +2.4 · tie |
| 72 h | 714 | 84.1 | 84.3 | 93.7 | 85.5 | −0.3 | −5.4 to +5.0 · tie |
| 96 h | 564 | 130.2 | 121.3 | 137.7 | 129.7 | +8.9 | −0.4 to +19.0 · NHC better |
| 120 h | 418 | 190.3 | 179.8 | 198.2 | 194.0 | +10.5 | NHC better; vs HCCA a tie |
Track error, n mi, homogeneous 2024–2025 holdout, Helix v3.0.0. Intensity is a wash after 12 h (Helix is worse at 12 h). SHIPS is missing on the holdout — CIRA files for 2024–25 are not posted — so the unique environmental inputs do not get a clean holdout test. Anyone quoting Helix as better than NHC is misquoting this table.
Helix is a public consensus track that is level with NHC and HCCA on a frozen holdout, not ahead of them. The /model page is therefore not a leaderboard. It is a this-cycle uncertainty report: official last-year cone on, this run’s ECMWF 1σ ellipse on, along/cross and intensity spread in the sidebar, named models off until you ask. The line is still scored so the width claim cannot hide a miss. On 2026 so far, about a third of 48 h cycles were tighter than the official cone, Helix’s line beat NHC on about 29% of homogeneous cycles, and those two rates barely move together. Width and the line are different bets. That is the value.