What the public record says happened
Environment and Climate Change Canada described the broader May 21, 2022 storm as a roughly 1,000-kilometre corridor that crossed Ontario and Quebec over about nine hours. The event affected more than one million electricity customers across the two provinces and caused fatalities and major insured damage. Those facts establish the scale and speed of the event after it occurred. They do not reveal what a particular outage model knew at a chosen forecast cutoff. Source: Environment and Climate Change Canada (opens in a new tab)
Hydro-Québec reported gusts up to 150 km/h in Quebec, 11,254 outages, and 554,649 customers interrupted at 8 p.m. Its recap reported approximately $70 million of work and extensive replacement of damaged grid components. The utility identified Outaouais, Laurentides, Lanaudière, Mauricie, and Capitale-Nationale among affected regions. Every one of these totals and locations is an observed after-event fact. None is used here as a model input, probability, ranking, or validation result. Sources: Hydro-Québec (opens in a new tab) Environment and Climate Change Canada (opens in a new tab)
Evidence: Observed Public Data
Observed public event timeline
Present public facts after the event without implying they were forecast inputs.
-
May 21, 2022
Long-lived storm corridor
ECCC described roughly 1,000 km over about nine hours
-
Quebec impact
Severe wind and widespread interruption
Hydro-Québec reported gusts up to 150 km/h and 11,254 outages
-
8 p.m.
554,649 customers interrupted
Hydro-Québec reported evening customer impact
-
Afterward
Restoration and repair
Approximately $70M of reported work
Observed public facts; not GeoGridIQ predictions or as-issued model inputs.
Figure sources: Environment and Climate Change Canada (2022-12-21) (opens in a new tab) Hydro-Québec (2022-06-14) (opens in a new tab) Environment and Climate Change Canada (2023) (opens in a new tab)
Why the derecho is a useful—but demanding—benchmark
The event combined fast convective development, a long damage corridor, severe wind, leafed-out vegetation, large customer impact, and complex restoration. That makes it useful for asking whether a risk system could preserve issue-time evidence, update as the storm evolved, identify broad operational pressure, and remain honest about local uncertainty. It also makes hindsight especially tempting because the later damage path and affected regions are so visible. Sources: Environment and Climate Change Canada (opens in a new tab) Hydro-Québec (opens in a new tab)
A benchmark should be difficult for the right reasons. The test is not whether today's analyst can draw a corridor around places now known to have been affected. It is whether a frozen pipeline, using only data genuinely available by each issue time, produces stored outputs that match independent observations under rules declared in advance. Broad storm awareness, regional outage risk, feeder impact, and individual asset failure are different targets and require progressively richer evidence. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Keep four evidence layers separate
Layer one is the information available at the prediction cutoff: archived forecast products with issue and valid times, observations received by then, vegetation and infrastructure snapshots whose effective dates precede the event, and historical outage records available to the system. Layer two is the frozen computational environment: source revision, feature definitions, spatial mapping, dependencies, and model artifact. Layer three is the stored output, including status, geography, horizon, probability or class, explanation, and hashes. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Layer four is the outcome revealed after the valid window: observed outages, final station and radar records, damage reports, restoration milestones, and event summaries. A reconstruction becomes invalid when layer-four facts influence the first three. Even a seemingly harmless choice—selecting only the regions later known to be hard hit—can inflate apparent performance by removing the full population of potential false alarms and correct negatives. Sources: World Meteorological Organization (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Evidence: Conceptual Illustration
The evidence boundary
Show which evidence can exist before issuance and which outcomes must remain held back.
- Allowed before cutoff As-issued forecasts and dated context
- Verified availability, source, geometry, and version
- Frozen at issue Features, artifact, and output
- Immutable identities and timestamps
- Held back Observed impact and revisions
- Gust peaks, outages, damage, restoration, final reports
Show which evidence can exist before issuance and which outcomes must remain held back.
The frozen model clock
A defensible run starts with an explicit cutoff such as T−24, T−12, T−6, or T−1 hour relative to a declared Quebec impact window. Those labels are planning checkpoints, not evidence that a corresponding GeoGridIQ product existed. For each cutoff, the archive must show exactly which forecast cycles, warnings, observations, and contextual datasets were ingested by then. The same clock controls feature availability and prevents later revisions from quietly entering an earlier prediction. Sources: Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab)
Current ECCC documentation explains how radar, climate data, and present forecast systems are organized, but it cannot establish the exact May 2022 model configuration or prove that an issued file was archived. Historical climate datasets and retrospective radar help describe the storm. Unless their original availability and revision state are recorded, they cannot stand in for the inputs an operational model had at the cutoff. This case study found no complete as-issued archive for a GeoGridIQ rerun. Sources: Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab)
What genuinely pre-event evidence could contribute
With a verified archive, the weather side could include forecast wind and gust fields, storm timing, precipitation, lightning or convective context, and public alerts issued before the cutoff. Vegetation evidence could include dated seasonal greenness, land cover, or corridor context available before the storm. Historical vulnerability could use prior outage outcomes whose labels and ingestion dates meet the temporal contract. Infrastructure context could support consequence review at an appropriately generalized and protected spatial scale. Sources: Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab)
None of those inputs proves a later outage. Weather forecasts contain location and intensity error. NDVI does not measure clearance or tree failure. Historical outage density may reflect reporting coverage or past network state. Public infrastructure points may be incomplete or sensitive. A useful model would express uncertainty and data gaps, and the evaluation would segment results by cutoff and geography rather than collapse them into one story about the event. Source: National Institute of Standards and Technology (opens in a new tab)
Evidence: Schematic Not Surveyed
Regions named in the observed Hydro-Québec recap
Name affected regions without a probability, confidence score, risk ranking, or feeder geometry.
-
Observed
Outaouais
Named after the event -
Observed
Laurentides
Named after the event -
Observed
Lanaudière
Named after the event -
Observed
Mauricie
Named after the event -
Observed
Capitale-Nationale
Named after the event
Generalized location list from the after-event report; not a network map or forecast ranking.
What must be withheld until validation
Observed peak gusts, the final storm path, customer totals, exact outage locations, damaged components, affected-region lists, vegetation-damage shares, restoration work, and later official summaries are validation evidence. They cannot be used to choose features, tune thresholds, select map labels, or write the prediction before it is scored. Hydro-Québec's finding that vegetation accounted for 90 percent of damage leading to outages is valuable event-specific context; it is not a universal vegetation weight or a feature contribution. Source: Hydro-Québec (opens in a new tab)
The same rule applies to apparently technical datasets. A final quality-controlled station record may differ from what was available in real time. A retrospective radar mosaic may incorporate data after a declared issue time. An outage feed may be revised, deduplicated, or geolocated later. Each outcome needs a provenance and observation window, and any incomplete coverage must be stated before treating a non-match as a false alarm or correct negative. Sources: Environment and Climate Change Canada (opens in a new tab) Environment and Climate Change Canada (opens in a new tab)
What a real backtest bundle would need
The bundle should record event identity; forecast cutoff and valid window; provider, product, issue time, object identity, and hash for every weather file; feature snapshot identity and cutoff; database or dataset fingerprint; source-code revision; environment and dependency lock; model artifact, hash, intended region and horizon, and trust state; region filter; run command; raw and rendered output hashes; and the complete outcome-label population. Without those records, an independent reviewer cannot determine what was run. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
The evaluation plan belongs in the bundle too: event threshold, spatial unit or radius, temporal overlap, handling of duplicate or incomplete outages, lead-time definition, baselines, metrics, uncertainty, and any excluded cases. These choices must be set before outcomes are inspected. Re-running a mutable pipeline today on a convenient historical table can be a useful engineering demonstration, but it is not the same evidence as a preserved as-of backtest. Sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Evidence: Conceptual Illustration
Minimum evidence for an audited backtest
List the records needed for an independent as-of rerun and evaluation.
Scroll horizontally or use the arrow keys to compare every column.
| Layer | Required evidence | Status for this article |
|---|---|---|
| Issued inputs | Provider, issue/valid time, identity, hash, availability | Incomplete / not verified |
| Feature snapshot | Cutoff, code revision, data fingerprint, feature hash | Not available |
| Model | Artifact identity/hash, region/horizon, trust state | No 2022 event artifact bundle |
| Stored output | Timestamp, full prediction population, output hash | Not available |
| Evaluation | Labels, matching rules, denominators, metrics, uncertainty | Not run |
List the records needed for an independent as-of rerun and evaluation.
What the event can illustrate safely
The observed record supports a qualitative discussion of relevant driver families. Severe convective wind can stress conductors, structures, and vegetation. Leafed-out canopy and local corridor conditions can influence tree contact or fall-in. Prior outages may indicate recurring vulnerability but can also reflect historical reporting and network changes. Essential services and access routes affect consequence and preparedness rather than automatically increasing the physical probability of failure. Sources: Environment and Climate Change Canada (opens in a new tab) Hydro-Québec (opens in a new tab)
At conceptual planning checkpoints, a team might review the broad storm corridor, data freshness, vegetation exposure, high-consequence service areas, crews, materials, communications, and access. The appropriate action would depend on forecast confidence, local procedures, and costs of unnecessary preparation. This description teaches a decision framework without assigning probabilities to Outaouais, Laurentides, Lanaudière, Mauricie, Capitale-Nationale, or any other region. Source: National Institute of Standards and Technology (opens in a new tab)
How a completed evaluation should report results
A genuine evaluation would report the full denominator and base rate, not only correctly highlighted regions. For a binary target it would provide the confusion matrix or the data needed to derive precision, recall, specificity, and false-alarm behaviour. A probabilistic target would add calibration and proper scoring. Every result would state cutoff, horizon, geography, threshold, matching contract, sample size, baseline, provider coverage, and uncertainty. Results from different regions or cutoffs would remain separate. Sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Until the evidence bundle exists, the honest value for probability, confidence, coverage, misses, false alarms, and lead time is 'not evaluated.' That is more informative than representative design values because it tells readers exactly which evidence is absent. Current GeoGridIQ model state does not fill the gap: Quebec horizons are suspended, and the BC 24-hour artifact audit found no current forecast batch. Neither state validates a May 2022 scenario. Source: GeoGridIQ
Evidence: Conceptual Illustration
What the evidence currently supports
Replace simulated precision with explicit not-evaluated states.
Scroll horizontally or use the arrow keys to compare every column.
| Question | Current status | What would be required |
|---|---|---|
| Regional probability | Not evaluated | Stored as-of model output |
| Confidence / calibration | Not evaluated | Definition and exact artifact calibration evidence |
| Coverage and misses | Not evaluated | Full prediction and outcome populations plus matching rule |
| False alarms | Not evaluated | Complete observation coverage and declared threshold |
| Lead time | Not evaluated | Stored issue time, valid window, and action deadline |
Replace simulated precision with explicit not-evaluated states.
What useful lead time could support
If a future, validated system identified elevated regional risk early enough, operators could consider intensified monitoring, crew and material readiness, public communications, essential-service coordination, and inspection of known local vulnerabilities. The model should not declare that an outage is guaranteed or promise a restoration outcome. Its role would be to show the evidence and uncertainty behind a priority so qualified people can decide whether preparation is proportionate. Source: National Institute of Standards and Technology (opens in a new tab)
A historical case study can eventually test whether the warning arrived before the decision deadline, whether highlighted areas matched outcomes, which hard-hit areas were missed, and how many alerts did not match a recorded outage. It cannot prove that an action prevented an outage without linked intervention evidence. The value of this conceptual reconstruction is the no-hindsight standard itself: observed drama never substitutes for stored forecast evidence. Sources: National Institute of Standards and Technology (opens in a new tab) World Meteorological Organization (opens in a new tab)
Scope and safeguards
Limitations and responsible use
- No complete as-issued May 2022 forecast archive was verified for this case study.
- No versioned GeoGridIQ run, feature snapshot, model artifact bundle, or output hash exists for the event.
- Public aggregate outage totals are observed facts, not complete spatial labels.
- The affected-region figure is a generalized after-event schematic, not a surveyed network or feeder map.
- Current meteorological documentation cannot be projected backward as proof of the 2022 configuration.
- No probability, confidence, threshold, driver share, hit rate, coverage, or lead-time result is claimed.
Frequently asked questions
Questions this article answers
Did GeoGridIQ predict the derecho in real time?
No. No stored GeoGridIQ real-time forecast or complete reproducible backtest was found for this event.
What does a historical backtest prove?
Only the performance of a frozen target, data snapshot, pipeline, artifact, population, and evaluation contract. It does not automatically prove future deployment performance.
Why must post-event data be withheld?
Using later gusts, outages, damage, regions, or reports in the forecast stage creates hindsight leakage and exaggerates apparent skill.
How far ahead could useful risk have been identified?
That is not evaluated here. It requires archived issue-time inputs and stored outputs at each declared cutoff.
What would count as a miss or false alarm?
The definitions depend on a predeclared event threshold, geography, time window, and observation-coverage contract.
Evidence register
Sources
Sources were reviewed on . Mutable sources are rechecked on the article review schedule.
-
D1 Environment and Climate Change Canada. Canada's Top 10 Weather Stories of 2022 (opens in a new tab). 2022-12-21.
-
D2 Hydro-Québec. A look at the outages caused by the May 21 derecho (opens in a new tab). 2022-06-14.
-
D3 Environment and Climate Change Canada. 2022–23 Departmental Results Report (opens in a new tab). 2023.
-
D4 Environment and Climate Change Canada. About Canadian weather radar (opens in a new tab).
-
D5 Environment and Climate Change Canada. Historical Climate Data (opens in a new tab).
-
D6 Environment and Climate Change Canada. Historical climate data-quality information (opens in a new tab).
-
A1 National Institute of Standards and Technology. AI Risk Management Framework 1.0 (opens in a new tab). 2023-01-26.
-
A2 World Meteorological Organization. Forecast verifications (opens in a new tab).
-
A4 European Centre for Medium-Range Weather Forecasts. Verification metrics guide (opens in a new tab).
-
A6 Kapoor and Narayanan. Leakage and reproducibility failures in machine-learning science (opens in a new tab). 2023-08-04.
-
GGI_MODEL_STATUS GeoGridIQ. Protected-artifact safety and runtime availability audit.