GeoAI Risks Companion exercises and data

Exercise III-B. Blind-Spot Mapping and the Reliability Screen

Part III. Behavioral and Institutional Risk: Fragility

At a glance
Textbook sectionssections 9.2, 9.3, 9.4, 9.10, 9.11, 9.12, and section 4.5 recommendation 3
Technical demandTier 1 and Tier 2 required. Python with pandas, numpy, scipy and matplotlib
Efforttwo to three hours including the write up
Prerequisitesnone

Overview

Section 9.3 identifies three forms of spatial information asymmetry, and section 9.4 develops the blind-spot problem, defining a blind spot by whether its absence from the information system can change a consequential decision and separating visible missingness from silent missingness. The American Community Survey supplies the cleanest available case, because the Census Bureau publishes a margin of error beside every estimate and almost nobody maps it. A block group choropleth of median household income presents an impression of uniform local precision that the sampling behind it does not support anywhere near uniformly.

Students compute the coefficient of variation for every unit at three geographies, count how many exceed a thirty percent coefficient of variation, a reliability cutoff commonly applied to American Community Survey estimates, count how many estimates were suppressed outright, and produce three maps from one download. The third map suppresses every unit failing the screen, and the difference between the first and third maps marks the silent missingness of section 9.4, units the first map colors with the same confidence as the rest although their evidence fails the screen. Section 9.12 prescribes what an asymmetry-resilient system shows, including visible missingness in section 9.12.1 and the pairing of an estimate with the quality of its evidence in section 9.12.2. The exercise ends by requiring the student to decide what to release, one of the three maps or a combination of them, and to defend the decision against those two subsections.

Step by step

  1. Run the screen. Tier 1. Change to the toolkit folder inside GeoAI_Exercises, type python exercise_3b.py, and record, for block group, tract and county, the number of estimates present, the number suppressed, the percentage exceeding a thirty percent coefficient of variation, and the median coefficient of variation. Describe the direction of the pattern across the three geographies in one sentence.
  1. Map the three products. Tier 1. Record the estimate map, the reliability map and the screened map. The three maps are drawn at tract level, the finest geography for which the package ships boundaries, so they show a smaller share of failing units than the block group screen in Step 1 reports. Then write the caption you would place under the first map so that a reader understands what it does and does not support, in no more than two sentences.
Figure withheld from this edition. This figure prints values the lab asks you to find, so it appears only in the instructor edition. Your own run of the lab's program produces the numbers it summarizes.

Map, three panels at tract level. The estimate, the coefficient of variation, and the estimate with unreliable units suppressed, on a shared extent.

  1. Find what the ranking rested on. Tier 1. Record how many of the ten lowest income block groups fail the reliability screen. State what that means for a program that allocates resources to the ten most distressed block groups in the state.
  1. Locate the loss. Tier 1. The margin of error arrives in the column adjacent to the estimate in the same download. Trace the point in a normal workflow at which that column stops traveling with the estimate, and name the role in an organization who would have had to notice.
  1. Test information parity. Tier 2. Section 9.11 describes information-parity testing, which stratifies validation by variables that influence visibility and asks whether performance degrades as information thins. Apply the evidence half of that test. Block 3 compares the reliability of estimates in the smallest and largest population quintiles of tracts. Report the difference and state which communities are systematically represented by weaker evidence. Then change CV_LIMIT at the top of exercise_3b.py to 20 and then to 40, rerunning after each change, and report how the gap between the smallest and largest quintiles moves.
  1. Apply the assessment. Tier 1. Section 9.10 organizes the Spatial Information Asymmetry Assessment around five dimensions, being Coverage, Currency, Access, Representation and Contestability. Rate the block group income map on each dimension, and cite for each rating a number from Steps 1 through 5 or state that no step measured it.
  1. Read the code and complete the handoff. Tier 1. Annotate the coefficient of variation computation, including the division by 1.645, and answer the decisions file.