Across 5 episodes and 50 district-episode pairs that the cited assessments name, the physics track flagged 50 under some class, 13 under the class that occurred, and 23 scored the occurring class above the alarm band. Those three numbers are not the same number, and the difference between them is the most useful thing on this page.
Each episode is a committed file — a sourced truth set, a driver series and a report — and each report is recomputed in CI from those inputs. The tables on this page are that recomputation, projected, not a re-analysis.
Drivers: Open-Meteo historical archive (ERA5 / ERA5-Land era reanalysis) (https://archive-api.open-meteo.com/v1/archive), temperature_2m_max, temperature_2m_min, precipitation_sum, wind_speed_10m_max, wind_gusts_10m_max, et0_fao_evapotranspiration. Reanalysis knows the weather that occurred inside the window. This measures whether the physics track, given that weather, flags the districts that were hit — a ceiling on detection, not forecast skill. A lead-time hindcast needs archived forecast fields (ECMWF MARS/CDS), which this harness has no credential for.
Alarm band: 0.500 on the class severity score, at 7-day and 15-day horizons. The truth set below names the districts the cited assessments report as affected. Districts nobody named are UNKNOWN, not negatives: the same reports do not claim to have surveyed all 64 districts, and the 2020 reporting was constrained by the COVID-19 response. Counting them as negatives would manufacture a false-alarm rate.
These are five episodes, not a validation set. Detection is counted only over the districts the cited sources name — a district nobody named is unknown, not clear — and the drivers are reanalysis (the weather that occurred), so every number here is a ceiling on detection, not forecast skill.
Each row is one episode, scored against that episode's own truth set. "Flagged any class" counts a district as flagged when any of the eight classes crossed the alarm band; "flagged the class" counts only the class that occurred, which is the strict reading. The last column separates a district the track scored low from a district it scored high but classified under another class.
Totals: 50 of 50 named district-episode pairs carried a scored row, 50 were flagged under some class, and 13 named the class that occurred.
| Episode | Class | Onset | Named districts | With a scored row | Flagged any class | Flagged the class | Class over band |
|---|---|---|---|---|---|---|---|
| Cyclone Amphan | Tropical Cyclone | 2020-05-20 | 14 | 14 | 14 | 0 | 0 |
| Cyclone Yaas | Tropical Cyclone | 2021-05-26 | 9 | 9 | 9 | 0 | 0 |
| Cyclone Mocha | Tropical Cyclone | 2023-05-14 | 4 | 4 | 4 | 0 | 0 |
| Eastern flash floods | Flash Flood | 2024-08-24 | 13 | 13 | 13 | 13 | 13 |
| Northeast and coastal monsoon floods | Flood | 2025-06-01 | 10 | 10 | 10 | 0 | 10 |
These are the standard verification scores, computed over the district-horizon samples in each episode's window. POD is "—" for amphan-2020, yaas-2021, mocha-2023, northeast-flood-2025: the truth set names affected districts but records no dated outcome inside the prediction window, so there is no observed event to divide by. A false alarm ratio of 1.000 in that situation means "no negative sample existed", not "every alarm was wrong" — with no named event there is nothing for an alarm to be right about.
The false alarm ratio is measurable only against districts where an event was recorded as absent. The reports say plainly that no district is treated as a confirmed negative, so treat the FAR column as a bound on the fraction of alarms that hit a district nobody reported as affected — which is the drift signal the district pages also expose.
| Episode | Scored samples | Hits | Misses | False alarms | POD | FAR | CSI |
|---|---|---|---|---|---|---|---|
| Cyclone Amphan | 28 | 0 | 0 | 28 | — | 1.000 | 0.000 |
| Cyclone Yaas | 18 | 0 | 0 | 18 | — | 1.000 | 0.000 |
| Cyclone Mocha | 8 | 0 | 0 | 8 | — | 1.000 | 0.000 |
| Eastern flash floods | 26 | 26 | 0 | 0 | 1.000 | — | — |
| Northeast and coastal monsoon floods | 20 | 0 | 0 | 20 | — | 1.000 | 0.000 |
The product spec bands an alarm as WATCH from 0.40, the harness default is 0.50, and WARNING starts at 0.65. The rows below are the same scoring run at each band, which is how much the published decision depends on where the band is drawn.
In every one of these episodes the three bands produce the identical split: the severity scores are not clustered near the thresholds, so moving the band does not move the alarm set. That is a property of these four windows, not a general result.
| Episode | Band | Threshold | Scored | Hits | Misses | False alarms | POD | FAR |
|---|---|---|---|---|---|---|---|---|
| Cyclone Amphan | WATCH band | 0.400 | 28 | 0 | 0 | 28 | — | 1.000 |
| Cyclone Amphan | harness default | 0.500 | 28 | 0 | 0 | 28 | — | 1.000 |
| Cyclone Amphan | WARNING band | 0.650 | 28 | 0 | 0 | 24 | — | 1.000 |
| Eastern flash floods | WATCH band | 0.400 | 26 | 26 | 0 | 0 | 1.000 | — |
| Eastern flash floods | harness default | 0.500 | 26 | 26 | 0 | 0 | 1.000 | — |
| Eastern flash floods | WARNING band | 0.650 | 26 | 24 | 2 | 0 | 0.923 | — |
| Cyclone Mocha | WATCH band | 0.400 | 8 | 0 | 0 | 8 | — | 1.000 |
| Cyclone Mocha | harness default | 0.500 | 8 | 0 | 0 | 8 | — | 1.000 |
| Cyclone Mocha | WARNING band | 0.650 | 8 | 0 | 0 | 8 | — | 1.000 |
| Northeast and coastal monsoon floods | WATCH band | 0.400 | 20 | 0 | 0 | 20 | — | 1.000 |
| Northeast and coastal monsoon floods | harness default | 0.500 | 20 | 0 | 0 | 20 | — | 1.000 |
| Northeast and coastal monsoon floods | WARNING band | 0.650 | 20 | 0 | 0 | 20 | — | 1.000 |
| Cyclone Yaas | WATCH band | 0.400 | 18 | 0 | 0 | 18 | — | 1.000 |
| Cyclone Yaas | harness default | 0.500 | 18 | 0 | 0 | 18 | — | 1.000 |
| Cyclone Yaas | WARNING band | 0.650 | 18 | 0 | 0 | 18 | — | 1.000 |
The physics cross-check scores the two wind-damage classes — Tropical Cyclone and Severe Local Storm — from the **gust** maximum. That is the correction shipped on 2026-09-18: the forecast unit is a district centroid, and a centroid is not the eyewall, so a sustained 10 m maximum understates what the district actually faced. On Amphan's landfall day the sustained field reached 19–69 km/h where the same archive's gust field reached 51–134 km/h, which is the difference between a class that never left the 0.01–0.25 band and one that reached the coastal alarm band.
Each report therefore re-scores the same rows with the **sustained** maximum — what the track used before the correction — and both rows are below. Where they differ, the driver choice, not the formula or the weather, is what decided whether the district point was detectable.
With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0136–0.2480 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0987–0.5145 and crosses on 2. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.
With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0285–0.2457 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0388–0.3166 and crosses on 0. Even the gust field leaves the episode class below the band, so the wind argument alone does not explain the miss: at this distance from the track the driver the pipeline uses cannot represent the hazard, and the class the formula describes is unreachable for this event. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.
With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0000–0.2173 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0000–0.3773 and crosses on 0. Even the gust field leaves the episode class below the band, so the wind argument alone does not explain the miss: at this distance from the track the driver the pipeline uses cannot represent the hazard, and the class the formula describes is unreachable for this event. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.
With the sustained maximum the shipped pipeline uses, the episode's class scores 0.1175–1.0000 and crosses the 0.5 band on 83 of 128 rows; with the gust field the archive also carries it scores 0.1175–1.0000 and crosses on 83. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Flash Flood under either driver, so the separation between the wind-driven classes is a second, independent defect.
With the sustained maximum the shipped pipeline uses, the episode's class scores 0.2097–1.0000 and crosses the 0.5 band on 101 of 128 rows; with the gust field the archive also carries it scores 0.2097–1.0000 and crosses on 101. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Flash Flood under either driver, so the separation between the wind-driven classes is a second, independent defect.
| Episode | Driver | Rows | Wind (km/h) | Episode-class score | Rows over band | Top class |
|---|---|---|---|---|---|---|
| Cyclone Amphan | gust max (read by the track) | 128 | 51.1–133.9 | 0.0987–0.5145 | 2 | Fire (92), Flash Flood (25), Drought (6), Severe Local Storm (5) |
| Cyclone Amphan | sustained max (pre-correction) | 128 | 19.2–69.1 | 0.0136–0.2480 | 0 | Fire (92), Flash Flood (30), Drought (6) |
| Cyclone Yaas | gust max (read by the track) | 128 | 37.4–84.6 | 0.0388–0.3166 | 0 | Fire (119), Drought (5), Flash Flood (4) |
| Cyclone Yaas | sustained max (pre-correction) | 128 | 14.8–43.3 | 0.0285–0.2457 | 0 | Fire (119), Drought (5), Flash Flood (4) |
| Cyclone Mocha | gust max (read by the track) | 128 | 36.7–85.3 | 0.0000–0.3773 | 0 | Fire (94), Drought (30), Flash Flood (2), Heat Wave (2) |
| Cyclone Mocha | sustained max (pre-correction) | 128 | 18.1–51.0 | 0.0000–0.2173 | 0 | Fire (94), Drought (30), Flash Flood (2), Heat Wave (2) |
| Eastern flash floods | gust max (read by the track) | 128 | 33.1–53.6 | 0.1175–1.0000 | 83 | Flash Flood (78), Fire (49), Drought (1) |
| Eastern flash floods | sustained max (pre-correction) | 128 | 15.0–31.3 | 0.1175–1.0000 | 83 | Flash Flood (78), Fire (49), Drought (1) |
| Northeast and coastal monsoon floods | gust max (read by the track) | 128 | 49.0–85.0 | 0.2097–1.0000 | 101 | Flash Flood (83), Fire (45) |
| Northeast and coastal monsoon floods | sustained max (pre-correction) | 128 | 25.7–44.5 | 0.2097–1.0000 | 101 | Flash Flood (83), Fire (45) |
The physics cross-check feeds each formula an argument taken from the forecast unit. Four of those arguments, measured across all five episodes, used to sit at the top of their formula on every row — so the term could not distinguish one district from another, and any severity difference attributed to it was an artefact of the wiring rather than of the weather. The owner corrected the wiring on 2026-09-18 (the release line on every report says so), and both halves stay published here: what each term reads now, and the argument it used to be handed. Each table row carries the count under the corrected wiring and under the one it replaced. A term that still reaches its ceiling is a measured property of the window — a fortnight whose mean daily drying genuinely reached the fire divisor, a heat spell that really did hold five days past 30 °C — not an artefact of the unit it was passed.
The corrected wiring feeds each formula the quantity it describes: mean daily ET and mean daily maximum wind to the fire terms, the gust maximum to the two wind-damage classes, and the number of days past 30 °C / below 16 °C to the persistence terms. The same rows are scored with the pre-correction wiring beside it, so the change and its size are on the record rather than in a commit message.
| Episode | Rows | fire_wind | fire_drying | heat_persistence | cold_persistence |
|---|---|---|---|---|---|
| Cyclone Amphan | 128 | 0 of 128 | 0 of 128 | 119 of 128 | 0 of 128 |
| Cyclone Amphan — pre-correction | 128 | 116 of 128 | 128 of 128 | 128 of 128 | 128 of 128 |
| Cyclone Yaas | 128 | 1 of 128 | 0 of 128 | 128 of 128 | 0 of 128 |
| Cyclone Yaas — pre-correction | 128 | 86 of 128 | 128 of 128 | 128 of 128 | 128 of 128 |
| Cyclone Mocha | 128 | 1 of 128 | 24 of 128 | 128 of 128 | 0 of 128 |
| Cyclone Mocha — pre-correction | 128 | 74 of 128 | 128 of 128 | 128 of 128 | 128 of 128 |
| Eastern flash floods | 128 | 0 of 128 | 0 of 128 | 94 of 128 | 0 of 128 |
| Eastern flash floods — pre-correction | 128 | 36 of 128 | 128 of 128 | 128 of 128 | 128 of 128 |
| Northeast and coastal monsoon floods | 128 | 2 of 128 | 0 of 128 | 120 of 128 | 0 of 128 |
| Northeast and coastal monsoon floods — pre-correction | 128 | 128 of 128 | 128 of 128 | 128 of 128 | 128 of 128 |
The same rows, scored with the corrected wiring and with the one it replaced. The class a row crowns is what the physics track "would have picked", so this table is the size of the correction: before it, `Fire` — a class that scores high everywhere in the pre-monsoon coastal belt and therefore separates nothing — was the top pick on 127 of the 128 windows of a landfalling cyclone. The rain classes now lead on the two flood episodes and the cyclone windows split between `Fire` and the rain classes, which is the honest limit the reports state: point weather cannot separate a rain class from a wind class it arrives with.
| Episode | Rows | Top class — corrected | Top class — pre-correction |
|---|---|---|---|
| Cyclone Amphan | 128 | Fire 92, Flash Flood 30, Drought 6 | Fire 127, Flash Flood 1 |
| Cyclone Yaas | 128 | Fire 119, Drought 5, Flash Flood 4 | Fire 127, Flash Flood 1 |
| Cyclone Mocha | 128 | Fire 94, Drought 30, Flash Flood 2, Heat Wave 2 | Fire 126, Flash Flood 2 |
| Eastern flash floods | 128 | Flash Flood 78, Fire 49, Drought 1 | Fire 88, Flash Flood 40 |
| Northeast and coastal monsoon floods | 128 | Flash Flood 83, Fire 45 | Fire 75, Flash Flood 53 |
Each episode was scored against the districts the sources below name as affected. They are the same sources recorded in the committed episode files, with the same access date.
Every number here is recomputed in CI from the committed episode files and driver series: `python -m hindcast.cli check --require-reports` re-runs each report and fails if a single value differs, and `node scripts/build_model_performance.mjs --check` fails if the artifact this page is generated from no longer matches those reports. The machine-readable copy is linked below.
No, and this deployment will not publish one. Five episodes are not a validation set, the drivers are reanalysis rather than archived forecast fields, and no district is treated as a confirmed negative. What is published is what the reports actually measured: how many of the named districts were flagged, and POD/FAR/CSI with their denominators stated.
The Amphan truth set names the affected districts but records no dated outcome inside the prediction window, so there is no observed event to divide by and POD cannot be computed. Every alarm then counts as a false alarm because no negative sample exists either. The report says this in place rather than presenting a zero as a score.
No. The class and severity in every episode come from the independent physics cross-check. The CNN was not re-run: its input tensor needs Sentinel-1/2, Landsat and ERA5-Land bands over Earth Engine for the historical window, which this harness has no credential for.
Only when the committed hindcast reports change. The page is generated from an artifact whose own provenance is the reports' hashes, so a model update that is not re-scored on these episodes does not silently rewrite the validation numbers.
Loading the interactive HazardNet application…