Preserve the forecast before learning what happened
Accountability starts with an immutable as-issued record: target definition, issue time, valid interval, geographic unit, probability or class, full scored population, model artifact identity, feature snapshot, data cutoff, and output hash. Saving only high-risk points hides the denominator and prevents measurement of correct negatives, base rate, and selection effects. Recomputing later with revised weather, outage, or asset data is a different experiment and must not overwrite the original forecast. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Outcome data must also be versioned. Outage feeds can be delayed, updated, merged, removed, or incomplete, and observed absence is meaningful only within a declared coverage contract. A reproducibility bundle should retain the exact label extraction, geographic transformation, exclusions, and observation fingerprint. Without those records, a validation dashboard may look exact while comparing a historical forecast with a later and materially different representation of reality. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Define a match before viewing the outcome
A correct prediction needs a predeclared event threshold, time relation, and geographic relation. The observed outage must fall within the forecast's valid interval or another explicitly justified lead-and-lag window, and its location must match the prediction unit under a fixed rule. A point-radius rule, polygon overlap, service-area join, and feeder match answer different questions. Choosing whichever rule makes an event look successful after the fact is outcome-driven evaluation. Sources: World Meteorological Organization (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Lead time should measure the interval from the stored issue time to the operationally relevant event time, not from a later dashboard refresh. When forecasts overlap, the evaluation must state whether it scores every issue, the earliest qualifying warning, the latest pre-event forecast, or another policy. It should also state how duplicate outage reports and continuing events are handled. These choices determine whether the metric describes early warning, near-real-time detection, or retrospective classification. Sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Evidence: Schematic Not Surveyed
A match requires both time and place
Illustrate matching choices without depicting an actual service territory or claiming a universal radius.
-
Issue
Forecast is stored
Before the outcome and before the valid window closes -
Valid interval
Declared event window
Exact start/end and overlapping-forecast policy -
Prediction unit
Fixed geography
Point, cell, polygon, service area, or feeder -
Observed event
Versioned label
Matches under one predeclared spatial rule
Conceptual evaluation geometry; not a surveyed network and not a GeoGridIQ production matching threshold.
Figure sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Report the full rare-event outcome table
A hit is a qualifying forecast with a matching observed event. A miss is an observed event without a qualifying forecast. A false alarm is a qualifying forecast without a matching observed event under the declared coverage contract. A correct negative is a scored case where neither occurred. These are evaluation labels, not moral judgments. In particular, an observed false alarm is not evidence that crews prevented an outage unless linked, time-stamped intervention records support that separate causal claim. Sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (opens in a new tab)
Outages are usually rare relative to all scored places and times, so accuracy can be high even if a system misses the events users care about. Precision reports the share of positive forecasts that matched an event; recall reports the share of observed events that had a qualifying forecast. False-alarm measures, base rate, event and non-event sample counts, and a relevant baseline complete the picture. Precision-recall analysis is often more informative than ROC summaries for imbalanced classification, but no single metric captures every operational trade-off. Sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (opens in a new tab)
Evidence: Conceptual Illustration
Every scored case belongs in the outcome table
Define rare-event outcomes and show the denominators behind precision and recall.
A two-by-two conceptual outcome table. Precision and recall require their stated denominators; all counts and the observation-coverage contract must be published. A false alarm is not proof of successful intervention.
Scroll horizontally or use the arrow keys to compare every column.
| Forecast / observation | Observed event | No observed event under coverage contract |
|---|---|---|
| Qualifying positive forecast | Hit (true positive) | False alarm (false positive) |
| No qualifying positive forecast | Miss (false negative) | Correct negative (true negative) |
| Derived measures | Recall = hits / (hits + misses) | Precision = hits / (hits + false alarms) |
Define rare-event outcomes and show the denominators behind precision and recall.
Figure sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (2015-03) (opens in a new tab)
Evaluate probabilities before choosing an action threshold
A calibrated forecast aligns predicted probability with observed frequency across comparable cases: among cases assigned a similar probability, the observed event frequency should be similar over a sufficient sample. A reliability diagram groups forecasts into bins and compares forecast probability with outcome frequency. Sample size, dependence among nearby cases, season, region, horizon, and dataset shift all affect interpretation, so a single aggregate curve can conceal important failures. Sources: European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
An operational threshold converts probability into a review or alert decision and therefore trades misses against false alarms. It should be selected against the costs and lead-time needs of a declared use, then evaluated on data not used to tune it. A data-completeness, model-trust, or fallback label is not calibrated probability. Interfaces should show forecast probability, evidence quality, availability, and consequence separately so users do not read an internal status score as statistical certainty. Sources: National Institute of Standards and Technology (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Evidence: Conceptual Illustration
Calibration compares probabilities with observed frequencies
Distinguish statistical calibration from threshold choice and internal evidence-quality labels.
- Forecast bins Group similar probabilities
- Use transparent binning and adequate sample counts
- Observed frequency Measure outcomes in each bin
- Retain geography, horizon, season, and coverage context
- Reliability Compare probability and frequency
- Agreement indicates calibration, not perfect discrimination
- Threshold Choose a decision policy separately
- Trade misses, false alarms, lead time, and action cost
Distinguish statistical calibration from threshold choice and internal evidence-quality labels.
Figure sources: European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Use later holdouts, baselines, and reproducible reruns
Training and threshold selection belong in earlier periods; final evaluation belongs in later, unseen periods or events. Preprocessing, imputation, scaling, feature selection, and label construction must obey the same temporal cutoff. Closely related spatial or storm records also require careful grouping. Otherwise the model may receive information from the event family it is supposed to predict. A versioned data and code bundle lets reviewers rerun the experiment and determine whether reported results survive the intended split. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)
Benchmarks should include useful simple alternatives: base rate, recent-history rules, hazard thresholds, or an existing operational practice where it can be represented fairly. Results should be segmented by region, horizon, season, hazard, data-availability state, and consequential use. Confidence intervals or other uncertainty summaries matter when events are few. A model that improves an aggregate score but fails badly in one operational segment may not be ready for the decision it was designed to support. Sources: National Institute of Standards and Technology (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)
Promotion is a reversible safety decision
A promotion gate should require target and data-contract review, temporal-leakage checks, baseline and segment results, probability calibration, reproducible artifact identity, runtime compatibility, and an approved monitoring and rollback plan. Deployment does not end validation. Data health, prediction distributions, calibration, outcome coverage, and operational feedback need continued review. If identity, temporal lineage, batch generation, or freshness cannot be proven, the public system should fail closed rather than serve an old or fallback result as a trusted forecast. Sources: National Institute of Standards and Technology (opens in a new tab) GeoGridIQ
GeoGridIQ currently demonstrates that pause behaviour. As of July 26, 2026, all Quebec 6-, 24-, 72-, and 168-hour artifacts are suspended and unavailable because their temporal contract is unsafe or unproven. The British Columbia 24-hour artifact passes exact identity and local-cache verification, but no current forecast batch was available, so freshness and public availability remain unverified and unavailable. No production accuracy, precision, recall, calibration, or lead-time metric is claimed here. Source: GeoGridIQ
Evidence: Geogridiq Test Output
Trust can be paused when any protected gate fails
Represent the current fail-closed governance state without presenting performance metrics.
-
Contract
Target and temporal lineage
Fail if unsafe or unproven
-
Evidence
Holdout, baseline, calibration
Fail if incomplete or not decision-relevant
-
Artifact
Exact protected identity
Fail on mismatch or incompatibility
-
Runtime
Current batch and freshness
Fail if missing, stale, or wrong scope
-
Monitor
Outcomes, data health, rollback
Pause when continuing evidence fails
GeoGridIQ audit state as of July 26, 2026. Quebec horizons are suspended; the British Columbia artifact passes identity checks but lacks a current verified batch and freshness evidence.
Figure source: GeoGridIQ
Scope and safeguards
Limitations and responsible use
- Metrics are meaningful only for the declared target, population, geography, horizon, threshold, matching rule, and observation-coverage contract.
- Small event samples, nearby dependence, dataset shift, label revisions, and provider gaps can make estimates unstable.
- A false alarm does not demonstrate intervention, and a hit does not prove that a forecast caused a beneficial action.
- Data-quality or model-status labels must not be presented as calibrated statistical confidence.
- No current GeoGridIQ accuracy, precision, recall, false-alarm, calibration, lead-time, or operational-impact statistic is published.
Frequently asked questions
Questions this article answers
What counts as a correct outage prediction?
A stored positive forecast and observed event must satisfy the event, threshold, valid-time, and geographic matching rules defined before outcomes were inspected.
Is every forecast without an observed outage a useless false alarm?
It is a false alarm under the declared evaluation if observation coverage is adequate. It is not proof of successful intervention without linked intervention evidence.
Why is accuracy insufficient for outage forecasting?
Outages are rare relative to all scored cases, so predicting no event most of the time can appear accurate while missing operationally important outages.
What is forecast calibration?
It is agreement between predicted probabilities and observed event frequencies across comparable cases, assessed with sufficient samples and transparent grouping.
Does GeoGridIQ publish current production performance metrics?
No metrics are claimed here. Quebec horizons are suspended, and British Columbia lacks a current verified forecast batch and freshness evidence.
Evidence register
Sources
Sources were reviewed on . Mutable sources are rechecked on the article review schedule.
-
A1 National Institute of Standards and Technology. AI Risk Management Framework 1.0 (opens in a new tab). 2023-01-26.
-
A2 World Meteorological Organization. Forecast verifications (opens in a new tab).
-
A3 European Centre for Medium-Range Weather Forecasts. Reliability diagram (opens in a new tab).
-
A4 European Centre for Medium-Range Weather Forecasts. Verification metrics guide (opens in a new tab).
-
A5 Saito and Rehmsmeier. Precision-recall is more informative than ROC for imbalanced data (opens in a new tab). 2015-03.
-
A6 Kapoor and Narayanan. Leakage and reproducibility failures in machine-learning science (opens in a new tab). 2023-08-04.
-
GGI_MODEL_STATUS GeoGridIQ. Protected-artifact safety and runtime availability audit.