GeoAI Risks Companion exercises and data

Exercise VI-A. The Practical Uncertainty Stress Test

Part VI. Governing GeoAI Risks

Part VI. Governing GeoAI Risks

Chapters 15 and 16 turn from diagnosis to assurance. Chapter 15 names six functions of uncertainty resilient GeoAI in section 15.3, works through them in 15.4 to 15.6 and 15.8 to 15.10, treats calibration and the communication of uncertainty in 15.7, specifies a practical uncertainty stress test in 15.11, and specifies a GeoAI Uncertainty Assurance Case in 15.17. Chapter 16 develops governance for a system that has become institutional.

Fundamentals in play. These exercises use provenance and lineage tracking, seeded reproducibility, parameter sweeps, and the four measurements the collection has accumulated. Their failure mode is that an assurance artifact can be written to look complete while resting on no measurement at all, which is the specific failure Exercise VI-B is built to expose.

At a glance
Textbook sectionssections 6.8.5, 15.3, 15.4, 15.5, 15.6, 15.7, 15.11
Technical demandTier 1 and Tier 2 required. Python with the packages already in use
Effortfour hours including the write up
PrerequisitesExercise II-A, or Exercise III-A, IV-A or IV-B for the optional Tier 3 wrap

Overview

Section 15.11 specifies a practical uncertainty stress test in ten steps, and section 15.4 asks that every claim be traceable through its provenance and lineage. The harness in this exercise carries parts of three of the ten steps, mapping the evidence chain, perturbing the system and measuring decision stability, and the steps that define the mission consequence and document residual risk return in Exercise VI-B. This exercise wraps one earlier pipeline in a harness that pins the input snapshot by checksum, sets and records a random seed, records the package versions, sweeps the parameters earlier exercises showed to matter, and prints a report carrying the four measurements the collection has accumulated.

The acceptance test is mechanical. The harness runs twice and produces byte identical output, and a student whose second run differs has found an unpinned source of variation and must name it. That hunt teaches more than the harness does, because unpinned variation is invisible until something forces it into the open. Common culprits include an unseeded shuffle, iteration over a set of strings, whose order changes from one interpreter session to the next, and a timestamp written into the report header.

Step by step

  1. Select and wrap. Tier 1. Change to the toolkit folder inside GeoAI_Exercises and run python harness_template.py as shipped, which wraps the need index from Exercise II-A. Tier 3 and optional: wrap a pipeline from Exercise III-A, IV-A or IV-B instead by replacing the body of pipeline and the contents of GRID, normally with an assistant writing the code, and record which you chose and why.
  1. Pin everything. Tier 1. Verify the snapshot checksum on load and fail loudly on mismatch, set and record every seed, and record the Python and package versions. Record the provenance block the harness prints.
Screenshot. The harness provenance block as printed, showing the checksum, the seed and the package versions.
Screenshot. The harness provenance block as printed, showing the checksum, the seed and the package versions.
  1. Sweep what matters. Tier 1. Run the pipeline across the parameter grid the earlier exercise identified as consequential. Record the result under each combination and the decision stability across them.
  1. Report the four measurements. Tier 1. Record decision stability, the reliability screen count, cross validated error where a model is fitted, and the Monte Carlo spread where the pipeline propagates margins of error. Those four state how much the result can be trusted.
  1. Prove reproducibility. Tier 1. Type python harness_template.py --twice, which writes the report two times and compares the files byte for byte, then submit both reports with the comparison. If they differ, find the unpinned source and name it in writing.
  1. Write the model card. Tier 1. Produce a one page card stating what the pipeline does, what data it consumes, what it was tested on, the four measurements, and where it fails.