Charging Grid Intelligence...

Forecast validation

Forecast accountability: how to know whether an outage prediction deserves trust

A persuasive forecast map is not evidence of skill. Forecast accountability preserves what was known at issuance, defines success before outcomes are inspected, reports the complete denominator, and fails closed when model or runtime evidence no longer supports public use.

  • Forecast validation
  • Outage prediction
  • Calibration
  • Model governance

Evidence: Conceptual Illustration

Accountability preserves the path from forecast to evidence

Article overview

Show the auditable sequence and the rule that outcomes arrive only after the forecast is frozen.

  1. Define

    Target and matching contract

    Event, threshold, geography, issue and valid time

  2. Freeze

    Inputs, artifact, and output

    Complete population, identities, cutoffs, hashes

  3. Observe

    Versioned outcomes

    Coverage, revisions, exclusions, event identity

  4. Score

    Outcomes and probabilities

    Denominators, base rate, lead time, calibration

  5. Decide

    Promote, monitor, or pause

    Documented limits and reversible governance

Show the auditable sequence and the rule that outcomes arrive only after the forecast is frozen.

Direct answer

What readers need to know

An outage prediction deserves consideration only when its target, issue and valid times, geography, threshold, and matching rules were fixed in advance; its inputs and outputs are reproducible; and performance is reported across hits, misses, false alarms, correct negatives, lead time, and probability calibration. Validation must use later, unseen periods and must remain separate from internal data-quality labels and operational anecdotes.

Preserve the forecast before learning what happened

Accountability starts with an immutable as-issued record: target definition, issue time, valid interval, geographic unit, probability or class, full scored population, model artifact identity, feature snapshot, data cutoff, and output hash. Saving only high-risk points hides the denominator and prevents measurement of correct negatives, base rate, and selection effects. Recomputing later with revised weather, outage, or asset data is a different experiment and must not overwrite the original forecast. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)

Outcome data must also be versioned. Outage feeds can be delayed, updated, merged, removed, or incomplete, and observed absence is meaningful only within a declared coverage contract. A reproducibility bundle should retain the exact label extraction, geographic transformation, exclusions, and observation fingerprint. Without those records, a validation dashboard may look exact while comparing a historical forecast with a later and materially different representation of reality. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)

Define a match before viewing the outcome

A correct prediction needs a predeclared event threshold, time relation, and geographic relation. The observed outage must fall within the forecast's valid interval or another explicitly justified lead-and-lag window, and its location must match the prediction unit under a fixed rule. A point-radius rule, polygon overlap, service-area join, and feeder match answer different questions. Choosing whichever rule makes an event look successful after the fact is outcome-driven evaluation. Sources: World Meteorological Organization (opens in a new tab) Kapoor and Narayanan (opens in a new tab)

Lead time should measure the interval from the stored issue time to the operationally relevant event time, not from a later dashboard refresh. When forecasts overlap, the evaluation must state whether it scores every issue, the earliest qualifying warning, the latest pre-event forecast, or another policy. It should also state how duplicate outage reports and continuing events are handled. These choices determine whether the metric describes early warning, near-real-time detection, or retrospective classification. Sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

Evidence: Schematic Not Surveyed

A match requires both time and place

Illustrate matching choices without depicting an actual service territory or claiming a universal radius.

  1. Issue

    Forecast is stored

    Before the outcome and before the valid window closes
  2. Valid interval

    Declared event window

    Exact start/end and overlapping-forecast policy
  3. Prediction unit

    Fixed geography

    Point, cell, polygon, service area, or feeder
  4. Observed event

    Versioned label

    Matches under one predeclared spatial rule

Conceptual evaluation geometry; not a surveyed network and not a GeoGridIQ production matching threshold.

Figure sources: World Meteorological Organization (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

Report the full rare-event outcome table

A hit is a qualifying forecast with a matching observed event. A miss is an observed event without a qualifying forecast. A false alarm is a qualifying forecast without a matching observed event under the declared coverage contract. A correct negative is a scored case where neither occurred. These are evaluation labels, not moral judgments. In particular, an observed false alarm is not evidence that crews prevented an outage unless linked, time-stamped intervention records support that separate causal claim. Sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (opens in a new tab)

Outages are usually rare relative to all scored places and times, so accuracy can be high even if a system misses the events users care about. Precision reports the share of positive forecasts that matched an event; recall reports the share of observed events that had a qualifying forecast. False-alarm measures, base rate, event and non-event sample counts, and a relevant baseline complete the picture. Precision-recall analysis is often more informative than ROC summaries for imbalanced classification, but no single metric captures every operational trade-off. Sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (opens in a new tab)

Evidence: Conceptual Illustration

Every scored case belongs in the outcome table

Define rare-event outcomes and show the denominators behind precision and recall.

A two-by-two conceptual outcome table. Precision and recall require their stated denominators; all counts and the observation-coverage contract must be published. A false alarm is not proof of successful intervention.

Scroll horizontally or use the arrow keys to compare every column.

Every scored case belongs in the outcome table. Define rare-event outcomes and show the denominators behind precision and recall.
Forecast / observationObserved eventNo observed event under coverage contract
Qualifying positive forecast Hit (true positive) False alarm (false positive)
No qualifying positive forecast Miss (false negative) Correct negative (true negative)
Derived measures Recall = hits / (hits + misses) Precision = hits / (hits + false alarms)

Define rare-event outcomes and show the denominators behind precision and recall.

Figure sources: World Meteorological Organization (opens in a new tab) Saito and Rehmsmeier (2015-03) (opens in a new tab)

Evaluate probabilities before choosing an action threshold

A calibrated forecast aligns predicted probability with observed frequency across comparable cases: among cases assigned a similar probability, the observed event frequency should be similar over a sufficient sample. A reliability diagram groups forecasts into bins and compares forecast probability with outcome frequency. Sample size, dependence among nearby cases, season, region, horizon, and dataset shift all affect interpretation, so a single aggregate curve can conceal important failures. Sources: European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

An operational threshold converts probability into a review or alert decision and therefore trades misses against false alarms. It should be selected against the costs and lead-time needs of a declared use, then evaluated on data not used to tune it. A data-completeness, model-trust, or fallback label is not calibrated probability. Interfaces should show forecast probability, evidence quality, availability, and consequence separately so users do not read an internal status score as statistical certainty. Sources: National Institute of Standards and Technology (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

Evidence: Conceptual Illustration

Calibration compares probabilities with observed frequencies

Distinguish statistical calibration from threshold choice and internal evidence-quality labels.

Forecast bins Group similar probabilities
Use transparent binning and adequate sample counts
Observed frequency Measure outcomes in each bin
Retain geography, horizon, season, and coverage context
Reliability Compare probability and frequency
Agreement indicates calibration, not perfect discrimination
Threshold Choose a decision policy separately
Trade misses, false alarms, lead time, and action cost

Distinguish statistical calibration from threshold choice and internal evidence-quality labels.

Figure sources: European Centre for Medium-Range Weather Forecasts (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

Use later holdouts, baselines, and reproducible reruns

Training and threshold selection belong in earlier periods; final evaluation belongs in later, unseen periods or events. Preprocessing, imputation, scaling, feature selection, and label construction must obey the same temporal cutoff. Closely related spatial or storm records also require careful grouping. Otherwise the model may receive information from the event family it is supposed to predict. A versioned data and code bundle lets reviewers rerun the experiment and determine whether reported results survive the intended split. Sources: National Institute of Standards and Technology (opens in a new tab) Kapoor and Narayanan (opens in a new tab)

Benchmarks should include useful simple alternatives: base rate, recent-history rules, hazard thresholds, or an existing operational practice where it can be represented fairly. Results should be segmented by region, horizon, season, hazard, data-availability state, and consequential use. Confidence intervals or other uncertainty summaries matter when events are few. A model that improves an aggregate score but fails badly in one operational segment may not be ready for the decision it was designed to support. Sources: National Institute of Standards and Technology (opens in a new tab) European Centre for Medium-Range Weather Forecasts (opens in a new tab)

Promotion is a reversible safety decision

A promotion gate should require target and data-contract review, temporal-leakage checks, baseline and segment results, probability calibration, reproducible artifact identity, runtime compatibility, and an approved monitoring and rollback plan. Deployment does not end validation. Data health, prediction distributions, calibration, outcome coverage, and operational feedback need continued review. If identity, temporal lineage, batch generation, or freshness cannot be proven, the public system should fail closed rather than serve an old or fallback result as a trusted forecast. Sources: National Institute of Standards and Technology (opens in a new tab) GeoGridIQ

GeoGridIQ currently demonstrates that pause behaviour. As of July 26, 2026, all Quebec 6-, 24-, 72-, and 168-hour artifacts are suspended and unavailable because their temporal contract is unsafe or unproven. The British Columbia 24-hour artifact passes exact identity and local-cache verification, but no current forecast batch was available, so freshness and public availability remain unverified and unavailable. No production accuracy, precision, recall, calibration, or lead-time metric is claimed here. Source: GeoGridIQ

Evidence: Geogridiq Test Output

Trust can be paused when any protected gate fails

Represent the current fail-closed governance state without presenting performance metrics.

  1. Contract

    Target and temporal lineage

    Fail if unsafe or unproven

  2. Evidence

    Holdout, baseline, calibration

    Fail if incomplete or not decision-relevant

  3. Artifact

    Exact protected identity

    Fail on mismatch or incompatibility

  4. Runtime

    Current batch and freshness

    Fail if missing, stale, or wrong scope

  5. Monitor

    Outcomes, data health, rollback

    Pause when continuing evidence fails

GeoGridIQ audit state as of July 26, 2026. Quebec horizons are suspended; the British Columbia artifact passes identity checks but lacks a current verified batch and freshness evidence.

Figure source: GeoGridIQ

Scope and safeguards

Limitations and responsible use

  • Metrics are meaningful only for the declared target, population, geography, horizon, threshold, matching rule, and observation-coverage contract.
  • Small event samples, nearby dependence, dataset shift, label revisions, and provider gaps can make estimates unstable.
  • A false alarm does not demonstrate intervention, and a hit does not prove that a forecast caused a beneficial action.
  • Data-quality or model-status labels must not be presented as calibrated statistical confidence.
  • No current GeoGridIQ accuracy, precision, recall, false-alarm, calibration, lead-time, or operational-impact statistic is published.

Frequently asked questions

Questions this article answers

What counts as a correct outage prediction?

A stored positive forecast and observed event must satisfy the event, threshold, valid-time, and geographic matching rules defined before outcomes were inspected.

Is every forecast without an observed outage a useless false alarm?

It is a false alarm under the declared evaluation if observation coverage is adequate. It is not proof of successful intervention without linked intervention evidence.

Why is accuracy insufficient for outage forecasting?

Outages are rare relative to all scored cases, so predicting no event most of the time can appear accurate while missing operationally important outages.

What is forecast calibration?

It is agreement between predicted probabilities and observed event frequencies across comparable cases, assessed with sufficient samples and transparent grouping.

Does GeoGridIQ publish current production performance metrics?

No metrics are claimed here. Quebec horizons are suspended, and British Columbia lacks a current verified forecast batch and freshness evidence.

Evidence register

Sources

Sources were reviewed on . Mutable sources are rechecked on the article review schedule.

  1. A1 National Institute of Standards and Technology. AI Risk Management Framework 1.0 (opens in a new tab). 2023-01-26.

    primary standards guidance · Reviewed 2026-07-26

  2. A2 World Meteorological Organization. Forecast verifications (opens in a new tab).

    primary intergovernmental guidance · Reviewed 2026-07-26 · Mutable source

  3. A3 European Centre for Medium-Range Weather Forecasts. Reliability diagram (opens in a new tab).

    primary operational guidance · Reviewed 2026-07-26 · Mutable source

  4. A4 European Centre for Medium-Range Weather Forecasts. Verification metrics guide (opens in a new tab).

    primary technical guidance · Reviewed 2026-07-26 · Mutable source

  5. A5 Saito and Rehmsmeier. Precision-recall is more informative than ROC for imbalanced data (opens in a new tab). 2015-03.

    peer-reviewed primary methodological · Reviewed 2026-07-26

  6. A6 Kapoor and Narayanan. Leakage and reproducibility failures in machine-learning science (opens in a new tab). 2023-08-04.

    peer-reviewed systematic study · Reviewed 2026-07-26

  7. GGI_MODEL_STATUS GeoGridIQ. Protected-artifact safety and runtime availability audit.

    reproducible internal evidence · Reviewed 2026-07-26 · Root HANDOFF.md and exact runtime audit, 2026-07-26