What the model did on five historical episodes

Across 5 episodes and 50 district-episode pairs that the cited assessments name, the physics track flagged 50 under some class, 13 under the class that occurred, and 23 scored the occurring class above the alarm band. Those three numbers are not the same number, and the difference between them is the most useful thing on this page.

The five episodes

Each episode is a committed file — a sourced truth set, a driver series and a report — and each report is recomputed in CI from those inputs. The tables on this page are that recomputation, projected, not a re-analysis.

  • Cyclone Amphan — Bangladesh landfall, 20 May 2020 — onset 2020-05-20, 14 districts named as affected, truth completeness: named-affected-only.
  • Cyclone Yaas — Bangladesh coast, 26 May 2021 — onset 2021-05-26, 9 districts named as affected, truth completeness: named-affected-only.
  • Cyclone Mocha — Cox's Bazar coast and Teknaf, 14 May 2023 — onset 2023-05-14, 4 districts named as affected, truth completeness: named-affected-only.
  • Eastern flash floods — Feni, Cumilla, Noakhali, 20–30 August 2024 — onset 2024-08-24, 13 districts named as affected, truth completeness: named-affected-only.
  • Northeast and coastal monsoon floods — Sylhet, Sunamganj and the hill districts, 1 June 2025 — onset 2025-06-01, 10 districts named as affected, truth completeness: named-affected-only.

Read this first

Drivers: Open-Meteo historical archive (ERA5 / ERA5-Land era reanalysis) (https://archive-api.open-meteo.com/v1/archive), temperature_2m_max, temperature_2m_min, precipitation_sum, wind_speed_10m_max, wind_gusts_10m_max, et0_fao_evapotranspiration. Reanalysis knows the weather that occurred inside the window. This measures whether the physics track, given that weather, flags the districts that were hit — a ceiling on detection, not forecast skill. A lead-time hindcast needs archived forecast fields (ECMWF MARS/CDS), which this harness has no credential for.

Alarm band: 0.500 on the class severity score, at 7-day and 15-day horizons. The truth set below names the districts the cited assessments report as affected. Districts nobody named are UNKNOWN, not negatives: the same reports do not claim to have surveyed all 64 districts, and the 2020 reporting was constrained by the COVID-19 response. Counting them as negatives would manufacture a false-alarm rate.

  • The CNN was not evaluated. Severity and class come from the independent physics cross-check (scripts/physics_severity.py) run on reanalysis drivers. The CNN was not re-run: its t2 tensor needs Sentinel-1/2 + Landsat + ERA5-Land bands over Earth Engine for the historical window. These are not model-skill numbers.
  • Scores are shown to three decimals; the unrounded values, the per-district rows and the input hashes are in the machine-readable copy this page is generated from.

These are five episodes, not a validation set. Detection is counted only over the districts the cited sources name — a district nobody named is unknown, not clear — and the drivers are reanalysis (the weather that occurred), so every number here is a ceiling on detection, not forecast skill.

Detection: did the run flag the districts the sources name?

Each row is one episode, scored against that episode's own truth set. "Flagged any class" counts a district as flagged when any of the eight classes crossed the alarm band; "flagged the class" counts only the class that occurred, which is the strict reading. The last column separates a district the track scored low from a district it scored high but classified under another class.

Totals: 50 of 50 named district-episode pairs carried a scored row, 50 were flagged under some class, and 13 named the class that occurred.

Detection per episode, over the districts the cited sources name (class-agnostic and class-strict counts are both shown).
EpisodeClassOnsetNamed districtsWith a scored rowFlagged any classFlagged the classClass over band
Cyclone AmphanTropical Cyclone2020-05-2014141400
Cyclone YaasTropical Cyclone2021-05-2699900
Cyclone MochaTropical Cyclone2023-05-1444400
Eastern flash floodsFlash Flood2024-08-241313131313
Northeast and coastal monsoon floodsFlood2025-06-01101010010

Scores: POD, FAR and CSI — and the rows where they do not exist

These are the standard verification scores, computed over the district-horizon samples in each episode's window. POD is "—" for amphan-2020, yaas-2021, mocha-2023, northeast-flood-2025: the truth set names affected districts but records no dated outcome inside the prediction window, so there is no observed event to divide by. A false alarm ratio of 1.000 in that situation means "no negative sample existed", not "every alarm was wrong" — with no named event there is nothing for an alarm to be right about.

The false alarm ratio is measurable only against districts where an event was recorded as absent. The reports say plainly that no district is treated as a confirmed negative, so treat the FAR column as a bound on the fraction of alarms that hit a district nobody reported as affected — which is the drift signal the district pages also expose.

Verification scores at the shipped alarm band. "Scored samples" is the denominator each row was computed from; "—" is a score the report states cannot be computed.
EpisodeScored samplesHitsMissesFalse alarmsPODFARCSI
Cyclone Amphan2800281.0000.000
Cyclone Yaas1800181.0000.000
Cyclone Mocha80081.0000.000
Eastern flash floods2626001.000
Northeast and coastal monsoon floods2000201.0000.000

Threshold bands: 0.40, 0.50, 0.65

The product spec bands an alarm as WATCH from 0.40, the harness default is 0.50, and WARNING starts at 0.65. The rows below are the same scoring run at each band, which is how much the published decision depends on where the band is drawn.

In every one of these episodes the three bands produce the identical split: the severity scores are not clustered near the thresholds, so moving the band does not move the alarm set. That is a property of these four windows, not a general result.

The same scoring run at the three published alarm bands.
EpisodeBandThresholdScoredHitsMissesFalse alarmsPODFAR
Cyclone AmphanWATCH band0.4002800281.000
Cyclone Amphanharness default0.5002800281.000
Cyclone AmphanWARNING band0.6502800241.000
Eastern flash floodsWATCH band0.4002626001.000
Eastern flash floodsharness default0.5002626001.000
Eastern flash floodsWARNING band0.6502624200.923
Cyclone MochaWATCH band0.40080081.000
Cyclone Mochaharness default0.50080081.000
Cyclone MochaWARNING band0.65080081.000
Northeast and coastal monsoon floodsWATCH band0.4002000201.000
Northeast and coastal monsoon floodsharness default0.5002000201.000
Northeast and coastal monsoon floodsWARNING band0.6502000201.000
Cyclone YaasWATCH band0.4001800181.000
Cyclone Yaasharness default0.5001800181.000
Cyclone YaasWARNING band0.6501800181.000

The wind driver: which field the track reads, and why it decides detection

The physics cross-check scores the two wind-damage classes — Tropical Cyclone and Severe Local Storm — from the **gust** maximum. That is the correction shipped on 2026-09-18: the forecast unit is a district centroid, and a centroid is not the eyewall, so a sustained 10 m maximum understates what the district actually faced. On Amphan's landfall day the sustained field reached 19–69 km/h where the same archive's gust field reached 51–134 km/h, which is the difference between a class that never left the 0.01–0.25 band and one that reached the coastal alarm band.

Each report therefore re-scores the same rows with the **sustained** maximum — what the track used before the correction — and both rows are below. Where they differ, the driver choice, not the formula or the weather, is what decided whether the district point was detectable.

With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0136–0.2480 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0987–0.5145 and crosses on 2. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.

With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0285–0.2457 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0388–0.3166 and crosses on 0. Even the gust field leaves the episode class below the band, so the wind argument alone does not explain the miss: at this distance from the track the driver the pipeline uses cannot represent the hazard, and the class the formula describes is unreachable for this event. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.

With the sustained maximum the shipped pipeline uses, the episode's class scores 0.0000–0.2173 and crosses the 0.5 band on 0 of 128 rows; with the gust field the archive also carries it scores 0.0000–0.3773 and crosses on 0. Even the gust field leaves the episode class below the band, so the wind argument alone does not explain the miss: at this distance from the track the driver the pipeline uses cannot represent the hazard, and the class the formula describes is unreachable for this event. Note what neither number fixes: the track's top class is Fire under either driver, so the separation between the wind-driven classes is a second, independent defect.

With the sustained maximum the shipped pipeline uses, the episode's class scores 0.1175–1.0000 and crosses the 0.5 band on 83 of 128 rows; with the gust field the archive also carries it scores 0.1175–1.0000 and crosses on 83. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Flash Flood under either driver, so the separation between the wind-driven classes is a second, independent defect.

With the sustained maximum the shipped pipeline uses, the episode's class scores 0.2097–1.0000 and crosses the 0.5 band on 101 of 128 rows; with the gust field the archive also carries it scores 0.2097–1.0000 and crosses on 101. The wind driver, not the formula alone, decides whether this event was detectable at the district point. Note what neither number fixes: the track's top class is Flash Flood under either driver, so the separation between the wind-driven classes is a second, independent defect.

Episode-class score range and over-band count under each wind driver, per episode (128 district-horizon rows each).
EpisodeDriverRowsWind (km/h)Episode-class scoreRows over bandTop class
Cyclone Amphangust max (read by the track)12851.1–133.90.0987–0.51452Fire (92), Flash Flood (25), Drought (6), Severe Local Storm (5)
Cyclone Amphansustained max (pre-correction)12819.2–69.10.0136–0.24800Fire (92), Flash Flood (30), Drought (6)
Cyclone Yaasgust max (read by the track)12837.4–84.60.0388–0.31660Fire (119), Drought (5), Flash Flood (4)
Cyclone Yaassustained max (pre-correction)12814.8–43.30.0285–0.24570Fire (119), Drought (5), Flash Flood (4)
Cyclone Mochagust max (read by the track)12836.7–85.30.0000–0.37730Fire (94), Drought (30), Flash Flood (2), Heat Wave (2)
Cyclone Mochasustained max (pre-correction)12818.1–51.00.0000–0.21730Fire (94), Drought (30), Flash Flood (2), Heat Wave (2)
Eastern flash floodsgust max (read by the track)12833.1–53.60.1175–1.000083Flash Flood (78), Fire (49), Drought (1)
Eastern flash floodssustained max (pre-correction)12815.0–31.30.1175–1.000083Flash Flood (78), Fire (49), Drought (1)
Northeast and coastal monsoon floodsgust max (read by the track)12849.0–85.00.2097–1.0000101Flash Flood (83), Fire (45)
Northeast and coastal monsoon floodssustained max (pre-correction)12825.7–44.50.2097–1.0000101Flash Flood (83), Fire (45)

The formula inputs that used to carry no information

The physics cross-check feeds each formula an argument taken from the forecast unit. Four of those arguments, measured across all five episodes, used to sit at the top of their formula on every row — so the term could not distinguish one district from another, and any severity difference attributed to it was an artefact of the wiring rather than of the weather. The owner corrected the wiring on 2026-09-18 (the release line on every report says so), and both halves stay published here: what each term reads now, and the argument it used to be handed. Each table row carries the count under the corrected wiring and under the one it replaced. A term that still reaches its ceiling is a measured property of the window — a fortnight whose mean daily drying genuinely reached the fire divisor, a heat spell that really did hold five days past 30 °C — not an artefact of the unit it was passed.

The corrected wiring feeds each formula the quantity it describes: mean daily ET and mean daily maximum wind to the fire terms, the gust maximum to the two wind-damage classes, and the number of days past 30 °C / below 16 °C to the persistence terms. The same rows are scored with the pre-correction wiring beside it, so the change and its size are on the record rather than in a commit message.

Rows at the term's ceiling, out of the rows scored, per episode, under the corrected wiring and under the one it replaced. Three of the four terms were at their ceiling on every row of every episode before the correction (the fire drying term and both persistence terms); the fire wind term was at its ceiling on 36–128 of 128 rows depending on the episode. What is left after the correction is the windows that genuinely reached the term's own divisor.
EpisodeRowsfire_windfire_dryingheat_persistencecold_persistence
Cyclone Amphan1280 of 1280 of 128119 of 1280 of 128
Cyclone Amphan — pre-correction128116 of 128128 of 128128 of 128128 of 128
Cyclone Yaas1281 of 1280 of 128128 of 1280 of 128
Cyclone Yaas — pre-correction12886 of 128128 of 128128 of 128128 of 128
Cyclone Mocha1281 of 12824 of 128128 of 1280 of 128
Cyclone Mocha — pre-correction12874 of 128128 of 128128 of 128128 of 128
Eastern flash floods1280 of 1280 of 12894 of 1280 of 128
Eastern flash floods — pre-correction12836 of 128128 of 128128 of 128128 of 128
Northeast and coastal monsoon floods1282 of 1280 of 128120 of 1280 of 128
Northeast and coastal monsoon floods — pre-correction128128 of 128128 of 128128 of 128128 of 128

What the wiring correction changed

The same rows, scored with the corrected wiring and with the one it replaced. The class a row crowns is what the physics track "would have picked", so this table is the size of the correction: before it, `Fire` — a class that scores high everywhere in the pre-monsoon coastal belt and therefore separates nothing — was the top pick on 127 of the 128 windows of a landfalling cyclone. The rain classes now lead on the two flood episodes and the cyclone windows split between `Fire` and the rain classes, which is the honest limit the reports state: point weather cannot separate a rain class from a wind class it arrives with.

Top class per district-horizon row, corrected wiring beside the pre-correction wiring, per episode.
EpisodeRowsTop class — correctedTop class — pre-correction
Cyclone Amphan128Fire 92, Flash Flood 30, Drought 6Fire 127, Flash Flood 1
Cyclone Yaas128Fire 119, Drought 5, Flash Flood 4Fire 127, Flash Flood 1
Cyclone Mocha128Fire 94, Drought 30, Flash Flood 2, Heat Wave 2Fire 126, Flash Flood 2
Eastern flash floods128Flash Flood 78, Fire 49, Drought 1Fire 88, Flash Flood 40
Northeast and coastal monsoon floods128Flash Flood 83, Fire 45Fire 75, Flash Flood 53

What is not claimed here

  • Reanalysis is not the forecast that existed on the issue date. This measures whether the physics track, given the weather that actually occurred in the window, flags the districts that were hit — a ceiling on detection, not a forecast skill score.
  • The CNN is NOT re-run: the t2 tensor needs Sentinel-1/2 + Landsat + ERA5-Land bands over GEE for the historical window, which this environment cannot reach (Earth Engine credentials) and which the Phase 9 plan assigns to the archive-loaded run.
  • Peak 24 h precipitation is approximated by the wettest reanalysis day in the window; the live pipeline uses the peak 6-hourly accumulation. A forward-flood term is therefore slightly understated.
  • The affected set is the districts named in the cited assessments. Non-listed districts are unknown, not clear.
  • Identical to the Amphan episode: reanalysis rather than archived forecast fields, CNN not re-run, peak precipitation approximated by the wettest day, unnamed districts unknown.
  • Amphan (2020) had not been recovered from when Yaas struck; the affected set therefore overlaps almost completely, which is a property of the coast, not of the truth set.
  • Reanalysis rather than archived forecast fields: this measures whether the physics track, given the weather that occurred, flags the districts that were hit — a ceiling on detection, not forecast skill (identical to the other three episodes; the harness has no ECMWF MARS/CDS credential).
  • The CNN was not re-run: its input tensor needs Sentinel-1/2, Landsat and ERA5-Land bands over Earth Engine for the historical window.
  • The Bangladesh impact was a near-miss outward wind field rather than a direct eye crossing, so a district-centroid wind from a reanalysis grid will under-represent whatever the coast actually experienced — the same driver-fidelity limit the 2020/2021 cyclone episodes expose, in the opposite direction.
  • A district centroid cannot represent Saint Martin's Island (about 8 km², 120 km south of the Teknaf mainland): the island took the most severe documented damage in Bangladesh and is not a district.
  • The truth set has 4 districts, so the false-alarm denominator stays absent for the same reason as in the other episodes (`absence_means_no_event: false`). Any FAR this episode reports is a reporting boundary, not an operational false-alarm rate.
  • Reanalysis, not archived forecast — the number is a ceiling on detection, and the plan's lead-time requirement needs forecast fields the harness cannot see (same as the cyclone episodes).
  • The physics track was not re-run for this window and the CNN is not evaluated: detection is the independent cross-check track only.
  • Peak precipitation is approximated by the wettest day inside the window, and a district centroid cannot represent the hill-stream catchments that actually flooded — the same driver-fidelity limit the cyclone episodes expose, in the rainfall dimension.
  • `Flash Flood` is the class chosen for an event that was partly riverine: the Feni, Muhuri and Gomti rivers reached record levels, while the Khagrachhari/Rangamati and Sylhet-division flooding was flash and hill-stream in character. Flood and Flash Flood share their rainfall inputs in `physics_severity`, so the distinction is a reporting choice, not a measurement — recorded here so nobody reads a class match as a validation of the distinction.
  • Single-sourced truth set, and an onset date taken from the situation report rather than from a dated event record.
  • Reanalysis, not archived forecast — a ceiling on detection, as with every episode here.
  • CNN not re-run; detection is the physics cross-check track only.
  • Basin flooding depends on upstream rainfall in Meghalaya and Assam that the Bangladesh district centroids do not sample; the driver limitation is structural for the Sylhet/Sunamganj part of this affected set.

How to read the two tracks

  • `evaluation.scores.events` is class-strict: an alarm is a hit only when the predict track *named* the class that occurred. A district correctly flagged under a different class therefore appears there as a false alarm, which is why the class-agnostic `detection.flagged_any_class` exists — read them together.
  • `evaluation.scores.per_class` is one-vs-rest per class. `pod: 0.0` with `misses: N` for the episode class means the track never named that class, not that it scored the district low; `detection.episode_class_over_threshold` separates the two.
  • `detection` counts over the districts the sources name, per horizon.
  • Nothing in this report covers a district nobody named. Unknown is not clear, and the far column is unmeasurable rather than zero when no negative sample exists.

Truth sets and citations

Each episode was scored against the districts the sources below name as affected. They are the same sources recorded in the committed episode files, with the same access date.

  • HCTT Response Plan — Cyclone Amphan, United Nations Bangladesh Coordinated Appeal (June–September 2020), citing MoDMR preliminary reports. — https://reliefweb.int/report/bangladesh/hctt-response-plan-cyclone-amphan-united-nations-bangladesh-coordinated-appeal (accessed 2026-09-18)
  • HCTT Cyclone Amphan Response Plan: Monitoring Dashboard (10 August 2020). — https://reliefweb.int/report/bangladesh/hctt-cyclone-amphan-response-plan-monitoring-dashboard-10-august-2020 (accessed 2026-09-18)
  • Mapping floods in Bangladesh caused by Cyclone Amphan to support humanitarian response (PreventionWeb / UN-SPIDER partner mapping, May 2020). — https://www.preventionweb.net/news/mapping-floods-bangladesh-caused-cyclone-amphan-support-humanitarian-response (accessed 2026-09-18)
  • Cyclone Amphan in Bangladesh: An Overview — Living Deltas Hub (2022). — https://livingdeltas.org/blog/cyclone-amphan-in-bangladesh-an-overview (accessed 2026-09-18)
  • Cyclone YAAS: Light Coordinated Joint Needs Analysis — Needs Assessment Working Group (NAWG) & Information Management Working Group, Bangladesh, 6 June 2021. — https://reliefweb.int/report/bangladesh/cyclone-yaas-light-coordinated-joint-needs-analysis-needs-assessment-working-group (accessed 2026-09-18)
  • Bangladesh: Cyclone YAAS — Final Report, DREF operation MDRBD027 (IFRC, December 2021). — https://reliefweb.int/report/bangladesh/bangladesh-cyclone-yaas-final-report-n-mdrbd027 (accessed 2026-09-18)
  • Bangladesh: Cyclone Mocha Humanitarian Response Situation Report, as of 15 May 2023 — Inter Sector Coordination Group (ISCG) / UN Bangladesh, reporting the Department of Disaster Management (DDM) and Ministry of Disaster Management and Relief (MoDMR) initial damage information. — https://bangladesh.un.org/en/231958-bangladesh-cyclone-mocha-humanitarian-response-situation-report-15-may-2023 (accessed 2026-09-18)
  • Bangladesh and Myanmar: Impact of Cyclone Mocha — ACAPS Briefing Note, 23 May 2023. — https://www.acaps.org/fileadmin/Data_Product/Main_media/20230523_acaps_briefing_note_bangladesh_and_myanmar_impact_of_cyclone_mocha_0.pdf (accessed 2026-09-18)
  • Briefing Note: Cyclone Mocha, Saint Martin Island, 18 May 2023 — Start Fund Bangladesh. — https://reliefweb.int/report/bangladesh/briefing-note-cyclone-mocha-saint-martin-island-18-may-2023 (accessed 2026-09-18)
  • BDRCS Situation Report 3 — Southeastern Flood, August 2024 (Bangladesh Red Crescent Society, 28 August 2024), citing 11 districts: Feni, Cumilla, Chattogram, Khagrachari, Noakhali, Moulvibazar, Habiganj, Brahmanbaria, Sylhet, Lakshmipur, Cox's Bazar. — https://bdrcs.org/wp-content/uploads/2024/08/BDRCS-Sitrep-3-Southeastern-Flood-August-2024.pdf (accessed 2026-09-18)
  • Rising Waters, Rising Challenges: WHO's Response to Severe Flooding in Bangladesh (WHO Bangladesh, 4 November 2024). — https://www.who.int/bangladesh/news/detail/04-11-2024-rising-waters--rising-challenges-who-s-response-to-severe-flooding-in-bangladesh (accessed 2026-09-18)
  • Bangladesh Flooding August 2024 — Disasters activations (NASA Applied Sciences, 28 August 2024). — https://appliedsciences.nasa.gov/what-we-do/disasters/disasters-activations/bangladesh-flooding-august-2024 (accessed 2026-09-18)
  • Bangladesh floods claim 15 lives, affect more than 4.4 million (VOA News, 23 August 2024, reporting the MoDMR briefing). — https://www.voanews.com/a/bangladesh-floods-claim-15-lives-affect-more-than-4-4-million-/7754825.html (accessed 2026-09-18)
  • Global Rapid Post-Disaster Damage Estimation (GRADE) Report — August 2024 Floods, Bangladesh (World Bank, 2024). — https://documents1.worldbank.org/curated/en/099921506192542727/pdf/IDU-3c0e1de6-9af6-40b9-99a3-78397d1041ac.pdf (accessed 2026-09-18)
  • Situation Report 1 — Flood 2025 (Bangladesh Red Crescent Society, 2 June 2025), reporting the situation as of 1 June 2025: Sylhet, Sunamganj, Moulvibazar, Habiganj, Netrokona, Noakhali, Bhola, Khagrachari, Bandarban and Rangamati severely affected. — https://bdrcs.org/situation-report-1-flood-2025/ (accessed 2026-09-18)

Reproducing this page

Every number here is recomputed in CI from the committed episode files and driver series: `python -m hindcast.cli check --require-reports` re-runs each report and fails if a single value differs, and `node scripts/build_model_performance.mjs --check` fails if the artifact this page is generated from no longer matches those reports. The machine-readable copy is linked below.

Questions and answers

Is there an accuracy number for the forecast model?

No, and this deployment will not publish one. Five episodes are not a validation set, the drivers are reanalysis rather than archived forecast fields, and no district is treated as a confirmed negative. What is published is what the reports actually measured: how many of the named districts were flagged, and POD/FAR/CSI with their denominators stated.

Why does Cyclone Amphan show a false alarm ratio of 1.000 and no POD?

The Amphan truth set names the affected districts but records no dated outcome inside the prediction window, so there is no observed event to divide by and POD cannot be computed. Every alarm then counts as a false alarm because no negative sample exists either. The report says this in place rather than presenting a zero as a score.

Was the CNN evaluated on these episodes?

No. The class and severity in every episode come from the independent physics cross-check. The CNN was not re-run: its input tensor needs Sentinel-1/2, Landsat and ERA5-Land bands over Earth Engine for the historical window, which this harness has no credential for.

Does this page change when the forecast model is updated?

Only when the committed hindcast reports change. The page is generated from an artifact whose own provenance is the reports' hashes, so a model update that is not re-scored on these episodes does not silently rewrite the validation numbers.

Content reviewed 2026-09-18. HazardNet is decision support, not an official warning service — see the methodology for scope and limitations.

Loading the interactive HazardNet application…