GeoAI Risks Companion exercises and data

Exercise II-A. Measuring Aggregation Sensitivity

Part II. Representational Risk: Aggregation and Boundary Distortion

Part II. Representational Risk: Aggregation and Boundary Distortion

Chapters 4, 5 and 6 develop one argument through three lenses. Chapter 4 establishes the modifiable areal unit problem and suggests an Aggregation Sensitivity Index. Chapter 5 extends the argument to zones of confusion, where several spatial frameworks assign one place conflicting, unstable or misleading meanings, and traces how those conflicts reach funding, enforcement and other decisions. Chapter 6 adds scale sensitivity and boundary effects, and supplies the diagnostic pipeline in section 6.7. Chapters 4, 5 and 6 share three tests, being multiscale validation, boundary perturbation and cross-boundary or cross-zonal transfer, and Chapters 5 and 6 each add a fourth test that follows change over time, temporal boundary auditing in section 5.6.4 and spatial drift monitoring in section 6.7.4. The three shared tests let one exercise carry the Part.

Two exercises serve this Part. Exercise II-A runs three of the four tests in the diagnostic pipeline section 6.7 proposes, leaves the temporal stage to the optional Challenge, and computes an Aggregation Sensitivity Index. Exercise II-B hands the student a completed analysis containing a defect and asks the student to find it. Both use the same frozen data and the same toolkit, so an instructor may assign either without additional setup, or assign II-B as a follow up to II-A in a course with the laboratory hours to spend.

Fundamentals in play. These exercises use Pearson correlation at several levels of aggregation, random repartitioning as a reference distribution, percentile ranking and linear rescaling as ways of putting variables on a common scale, Jaccard overlap between ranked sets, and the coefficient of variation computed from a published margin of error. Each fails in a known way. A correlation belongs to the units that carry it, a random partition answers only the question its randomization poses, a rank hides the size of the gap between neighbors, and a median cannot be averaged into a larger unit's median.

At a glance
Textbook sectionssections 4.2, 4.4, 4.5, 5.6.1, 5.6.2, 5.6.3, 5.6.4, 5.6.5, 6.7.1, 6.7.2, 6.7.3, 6.7.5
Technical demandTier 1 and Tier 2 required, Tier 3 optional. Python with pandas and numpy. No GIS, no key, no account
Effortthree hours including the write up
Prerequisitesnone

Overview

Chapter 4 argues that recognizing the modifiable areal unit problem accomplishes nothing by itself, and that analysts need methods for determining whether conclusions survive a change in geographic representation. Section 4.4 names three such methods, being multiscale modeling, boundary perturbation testing and cross-boundary validation, and then suggests an Aggregation Sensitivity Index that would summarize the results into a single reported figure. The chapter states plainly that no universally accepted operational standard exists for deciding how much aggregation sensitivity is acceptable, and it leaves that question as an important research frontier. This exercise takes the chapter at its word and builds one.

The analysis asks whether poverty and household vehicle access are related in Mississippi, which sounds like a question with an answer. The student measures whether the answer depends on the geography chosen to ask it, whether the published administrative boundaries produce a different relationship than random boundaries at the same resolution, and whether an independently defined partition preserves the relationship. Those three questions correspond to the three tests section 4.4 names.

The exercise closes by computing an Aggregation Sensitivity Index from four components and asking the student to defend the weighting. The index has no universal threshold of acceptability, as section 4.4 states, which makes the defense the assessment. A student who reports the number without arguing for the construction has completed the arithmetic and missed the exercise.

What you need before you start

You need Python installed, along with the pandas and numpy libraries. If you have never installed a Python library, open a terminal and type pip install pandas numpy and wait for it to finish. If you cannot install Python on your machine, open Google Colab in a browser, create a new notebook, and everything below works unchanged.

You need the file asi_toolkit.py from the companion repository. Name a folder GeoAI_Exercises, create a folder named toolkit inside it, and place asi_toolkit.py in toolkit. The program writes its snapshots folder beside toolkit, so everything the exercise produces lands inside GeoAI_Exercises.

You need no Census key, no ArcGIS account, no GIS software and no shapefile. The data arrives as plain text tables from a public United States Census Bureau address that requires no registration. That address serves the same American Community Survey estimates the Census application programming interface returns, and it publishes the margin of error in the column adjacent to every estimate. Students who later want the programming interface will find it documented in the companion repository, and nothing in this exercise depends on it.

Step by step

  1. Freeze the data. Tier 1. Open a terminal, change to the toolkit folder inside GeoAI_Exercises, and type python asi_toolkit.py --pull. The program downloads four American Community Survey tables from the Census Bureau summary file service, keeps only the Mississippi rows and only the columns this exercise needs, and writes a single file named acs2023_mississippi_asi.csv into a snapshots folder it creates. It then prints the row count and the full 64 character SHA256 checksum, and writes both into manifest.json. Copy the row count and the checksum onto your provenance card now, along with today's date. Every later step reads the frozen file and never touches the network, so your analysis reproduces even if the Census Bureau changes something tomorrow. That practice carries the versioning discipline of section 6.8.5 from boundary layers to the data themselves, and it is also the reason your numbers and your classmate's numbers can be compared at all.
  1. Run the five diagnostics. Tier 1. Type python asi_toolkit.py --report. The program prints a report with five numbered blocks and a summary. Read it once without analyzing it, then read it again with the book open to sections 4.4 and 4.5, because each block names the book sections it implements.
  1. Record the multiscale result. Tier 1. Block 1 reports the correlation between the percentage of people below poverty and the percentage of households with no vehicle, computed at block group, tract and county. Record all three coefficients, the three sample sizes and the three values of r squared in a table. The coefficient changes as the geography coarsens, and the direction of that change is the scale effect Gehlke and Biehl first reported in 1934. State in one sentence what happens to the apparent strength of the relationship as the units get larger, and state in a second sentence what a decision maker reading only the county figure would conclude that a decision maker reading only the block group figure would not.
  1. Record the boundary perturbation result. Tier 1. Block 2 assigns block groups at random to as many zone labels as the Census table has tract rows, 878, five hundred times, recomputing the correlation under each partition, so some labels receive no block group in a given draw. Record the fifth percentile, the median and the ninety fifth percentile of that distribution, and record where the real tract value falls within it.
Figure withheld from this edition. This figure prints values the lab asks you to find, so it appears only in the instructor edition. Your own run of the lab's program produces the numbers it summarizes.

Figure. Histogram of the correlation across five hundred random partitions, with the observed tract value marked as a vertical line. This is the test section 5.6.2 calls a spatial stress test. Answer the question it raises: if the real boundaries produce a result outside the range that random boundaries of identical resolution produce, what does that tell you about the boundaries? At least two explanations compete, one being that tracts are drawn to be internally homogeneous and therefore preserve relationships that random grouping destroys, and the other being that tracts encode the very socioeconomic sorting the analysis is measuring, and a third follows from the design of the test itself, because random grouping across the whole state ignores the correlation between neighboring block groups that any real zoning keeps. Argue for one explanation and say what evidence would settle it.

  1. Record the cross-zonal transfer result. Tier 1. Block 3 recomputes the same correlation over ZIP Code Tabulation Areas, which the Census Bureau builds from postal delivery geography while it builds tracts to population thresholds. ZCTAs do not nest inside tracts, do not respect county lines, and are not drawn with any statistical purpose, which makes them the independent partition section 5.6.3 asks for. Record the ZCTA coefficient and the ZCTA count. Compare the count against the tract count and notice that ZCTAs are coarser, then compare the coefficient against the tract coefficient. The scale effect you measured in Step 3 predicts a direction for that comparison. State whether the ZCTA result moves in the predicted direction, and if it does not, state what section 6.7.3 calls that outcome and what it implies about whether the relationship belongs to the phenomenon or to the tract system.
  1. Record the rank stability result. Tier 1. Block 4 builds a three variable need index from the percentage over sixty five, the percentage below poverty and the percentage of households without a vehicle, ranks every unit, and reports the Jaccard overlap between the ten highest ranked places under different representations. Record four overlap values. The first compares block group ranking against tract ranking and answers whether the same places surface at two resolutions. The remaining three compare three ways of putting the variables on a common scale before adding them, those being percentile rank, the z score and the min to max rescale. One of those three comparisons returns a very different answer from the other two, and identifying which one and explaining why is worth more than the other three answers combined.
  1. Record the evidence resolution result. Tier 1. Block 5 reports, for each geography, how many median income estimates the Census Bureau suppressed entirely, how many of the survivors carry a coefficient of variation above thirty percent, and the median coefficient of variation. Thirty percent is a threshold many analysts apply, and the Census Bureau itself sets no hard-and-fast rule, asking each user to judge the precision an application needs. Record the three rows. Recommendation 3 in section 4.5 tells organizations to distinguish the resolution at which results are displayed from the resolution of the evidence behind them, so that a detailed map never implies more certainty than the data justify. This block extends that idea to sampling error, which section 6.6.3 names when it warns that finer data can introduce unstable estimates and false precision. State how many block group estimates a Mississippi analyst can actually rely on, and state what a block group choropleth of median income implies to a viewer that the underlying evidence does not support.
  1. Compute and defend the index. Tier 1 and Tier 2. The report closes with four components and an Aggregation Sensitivity Index that averages them with equal weight. Record all five numbers.
Figure withheld from this edition. This figure prints values the lab asks you to find, so it appears only in the instructor edition. Your own run of the lab's program produces the numbers it summarizes.

Figure. The four Aggregation Sensitivity Index components as a bar chart with the composite index marked, so the reader sees which component drives the score. Now open asi_toolkit.py, find the function named asi, and read how each component is constructed. Equal weighting is a choice the authors of this exercise made and did not justify. Change the weights in the dictionary named ASI_WEIGHTS near the top of the file to something you can defend, rerun the report, and record both values. Then apply the three considerations section 4.4 names for interpreting such an index, being the consequences of the decision, the quality of the available evidence and the degree of geographic stability the intended use requires, and state what a decision maker would need to know about each before judging an index of this magnitude acceptable.

  1. Change the study area. Tier 2. Open asi_toolkit.py in any text editor. Near the top you will find three lines reading STATE = "28", STATE_NAME = "Mississippi" and a pair of values ZIP_LO, ZIP_HI. Change them to a different state, using that state's two digit Federal Information Processing Standard code and its ZIP code range, then type python asi_toolkit.py --pull followed by python asi_toolkit.py --report Record the index for your second state beside the index for Mississippi. Two states support no generalization and you should say so plainly. A similar index in both states fits the reading that the sensitivity belongs to the method, and a very different index points toward something particular to one state, though you would need more states before you could claim either reading.
  1. Read the code. Tier 1. The exercise names one block for annotation, being the function boundary_perturbation in asi_toolkit.py. Write a plain language annotation of every line in that function, stating what the line consumes, what it produces, and what would happen to the reported result if the line were removed. Pay particular attention to the line beginning g = bg.groupby(...), because that single line performs the reaggregation the entire test depends on, and to the four column names in the list named cols. Those four columns are counts. Explain why the function sums counts and then divides, and explain what would go wrong if it averaged the percentage columns already present in the data.
  1. Put the assistant to work. Tier 3, optional, ten points. Open a fresh session with any artificial intelligence assistant and give it the following, with no additional context.
I have a table of Mississippi census tracts with columns pct_pov and
pct_nv. Write Python that tells me whether poverty and vehicle access
are related, and give me a number I can put in a report.

Save the response. Then open a second fresh session and give it this instead.

I have American Community Survey data for Mississippi at block group,
tract, county and ZCTA level, with columns pct_pov and pct_nv. Compute
the Pearson correlation between them at every one of the four
geographies and report the unit count at each. Then tell me which of
those four numbers I should put in a report, and what I would need to
know before I could answer that question myself.

Save that response too. Compare the two, then write three paragraphs. The first states what the first response would have led you to report. The second states what the second response added. The third answers the question this exercise exists to raise: if the assistant warned you about aggregation without being asked, which current assistants often do, did it choose the geography for you, and if it did not, who did?

Key takeaways

Aggregation is a variable the analyst selects and tests, which section 4.4 states in one sentence and this exercise states in five numbers. The correlation you computed has no single correct value, and every value you computed is arithmetically correct, which is the uncomfortable part. A GeoAI system trained at one geography inherits that geography's answer and reports it with a precision the evidence never contained.

Administrative boundaries carry information. The boundary perturbation test asks whether real tracts behave differently from random zones of the same count, and a difference would mean the tract system does analytical work that nobody in the pipeline chose or documented. That result cuts both ways, because boundaries that preserve real structure improve an analysis while boundaries that encode the outcome variable contaminate it, and distinguishing the two requires knowing how the boundaries were drawn.

Transfer failure carries information. Section 6.7.3 names transfer failure, the case of a relationship that fails to survive an independent partition, and states that it does not automatically invalidate a model and still matters. The correct response describes the model as specific to the geography it was built on and scopes every claim accordingly. A model that works only inside the census tract system is still useful, provided nobody claims otherwise.

Evidence resolution and output resolution are different quantities, and the gap between them is measurable. Section 4.5 recommends separating them, and Block 5 measures one dimension of the gap, the sampling reliability of the estimates at each geography, in figures any reader can check. A choropleth drawn at block group resolution presents an impression of local precision that the sampling behind it cannot support anywhere near uniformly across the map. The remedy costs almost nothing, since the reliability figures arrive in the same download as the estimates and require one division to compute.

One column, four defensible classifications. 859 of 872 tracts land in a different class under at least one scheme, and exactly one tract sits in the worst class under all four.
One column, four defensible classifications. 859 of 872 tracts land in a different class under at least one scheme, and exactly one tract sits in the worst class under all four.

Questions

  1. Section 4.4 proposes six possible components for an Aggregation Sensitivity Index and this exercise implements four of them. Name the two that were omitted, explain what each would require, and state how including them could change the reported index and why the direction cannot be known before computing them.
  1. The boundary perturbation test destroys spatial contiguity, because random reassignment scatters each zone across the state. Openshaw's original formulation preserved contiguity. Explain how preserving contiguity would change the distribution you measured, and state which version answers the question a decision maker actually asks.
  1. A colleague argues that the block group result should always be preferred because it uses the finest units. Using your Block 5 numbers, explain why that argument fails, and state the conditions under which it would succeed.
  1. Your index reported a single number. Section 4.4 warns against interpreting such an index as a universal threshold. Write the two sentences you would place beneath the number in a report to a decision maker who has never heard of the modifiable areal unit problem.

Challenge, optional, up to ten extra credit points

Section 5.6.4 describes temporal boundary auditing, which compares model output under an old boundary vintage against a new one and identifies locations whose classification changed because the lines moved while conditions on the ground held steady. Census tract boundaries were redrawn between the 2010 and 2020 censuses. Hold the observations fixed by apportioning the 2023 block group counts in the frozen snapshot to 2010 tracts, using the Census Bureau's 2020 to 2010 block relationship file weighted by 2020 block population, then compute the need index on both tract vintages and report how many units changed decile because the lines moved. State what that count means for a program that allocates money by tract decile, and state why an earlier vintage of the survey tables would confound the boundary change with a change in conditions.