GeoAI Risks Companion exercises and data

Exercise II-B. The Analysis That Was Already Wrong

Part II. Representational Risk: Aggregation and Boundary Distortion

At a glance
Textbook sectionssections 4.3, 4.4, 4.5, 5.3, 5.4, 5.6, 6.4
Technical demandTier 1 required, Tier 3 optional. Python with pandas and numpy, plus python-docx to regenerate the memorandum and matplotlib and scipy for the audit figure
Efforttwo hours including the write up
Prerequisitesnone, though students who have completed II-A finish faster

Overview

A consulting firm has delivered a completed analysis recommending which ten Mississippi census tracts should receive a cooling center. The deliverable arrives as a memorandum, a ranked table and the script that produced it, all of which look professional and one of which is wrong. The student's task is to find the defect, quantify what it cost, and write the two paragraph response a client would send back.

This exercise inverts II-A. The student receives a finished output and works backward to the diagnostic that would have caught the defect in it. That difference matters for a course where the assistant may already be good enough to avoid the error, because a defect embedded in a delivered file stays defective no matter how capable the assistant becomes.

The supplied analysis contains one defect of substance and two of presentation, and finding all three is the work of the exercise.

Step by step

  1. Read the deliverable. Tier 1. Open cooling_centers_memo.docx, cooling_centers_ranked.csv and cooling_centers_analysis.py from the companion repository. Read the memorandum first, as a client would, and write down the recommendation and the three sentences that justify it before you look at any code.
Facsimile. The opening of the delivered consulting memorandum, retypeset for this workbook, so students judge the artifact before they audit it. The delivered file carries a date line and a longer table than the facsimile shows.
Facsimile. The opening of the delivered consulting memorandum, retypeset for this workbook, so students judge the artifact before they audit it. The delivered file carries a date line and a longer table than the facsimile shows.
  1. Reproduce the result. Tier 1. Copy cooling_centers_ranked.csv to a file named delivered_ranked.csv before you run anything. Then change to the toolkit folder, type python cooling_centers_analysis.py, and confirm that the table it prints matches delivered_ranked.csv. The script rewrites the ranked table and the memorandum each time it runs, and it needs the python-docx library (pip install python-docx) to write the memorandum. An analysis you cannot reproduce cannot be audited, and confirming reproduction is the first step of any review.
  1. Find the defect of substance. Tier 1. The script builds its need index from four variables and one of them is handled incorrectly. Use the frozen snapshot from Exercise II-A, which carries both the block group values and the published tract values for every variable, and compare what the script computed at tract level against what the Census Bureau publishes at tract level. The audit program performs that comparison without touching the delivered script: type python exercise_2b.py and read Block 1. Report the mean absolute difference across all tracts, the median absolute difference, the largest single difference, and the percentage of tracts where the difference exceeds five thousand dollars. Then name the statistical property that makes this operation invalid, in one sentence, without using the word average.
  1. Quantify the consequence. Tier 1. Block 2 of the audit program corrects the defect with the published tract medians, reruns the ranking, and reports how many of the ten recommended tracts change. Express the result as a Jaccard overlap and also as a plain count, then identify by GEOID and tract number the tract that gains the most positions and the tract that loses the most.
Figure withheld from this edition. This figure prints values the lab asks you to find, so it appears only in the instructor edition. Your own run of the lab's program produces the numbers it summarizes.

Figure, three panels. The dollar error the invalid aggregation introduces, the overlap of the recommended ten as the weight on income rises, and the overlap as the number of funded places changes. The client does not care about the mean absolute difference. The client cares whether the recommendation changes, and the difference between those two framings is the difference between a technical finding and a consequential one.

  1. Find where it matters. Tier 1. Block 3 reruns the delivered and corrected rankings with the income input weighted two, three, five and eight times, and again with the county funding five, ten, fifteen and twenty places, reporting the overlap each time. State the smallest change to the consultant's design at which the defect alters the recommendation, and explain how an invalid method could leave one design untouched while changing its neighbors.
  1. Find the two presentation defects. Tier 1. Section 4.5 makes eight strategic recommendations and the delivered analysis violates several of them. Name the two that the memorandum's missing statements violate most directly, cite each recommendation number, and state in one sentence each what the memorandum would have to add to comply.
  1. Write the response. Tier 1. Write the two paragraph reply a client would send to the consulting firm. The first paragraph states what is wrong and what it cost, in the client's language and with the numbers from Steps 4 and 5. The second states what the firm must deliver to close the finding. Do not exceed two paragraphs, because a review that runs longer than the deliverable it reviews does not get read.
  1. Ask the assistant to audit it. Tier 3, optional, ten points. Give any assistant the script and the single instruction Review this analysis and tell me whether the recommendation is sound. Record whether it found the defect of substance, whether it found either presentation defect, and whether it raised anything the exercise did not anticipate. Then give it the same script with the instruction This analysis aggregates a median. Explain why that is invalid and quantify the error using the tract values in the attached snapshot. Report both outcomes. Say so plainly if the assistant found the defect unprompted, and record which assistant and which version, because that observation is data your instructor can use.

Key takeaways

Medians do not aggregate, and no amount of careful coding around that fact repairs it. The Census Bureau publishes a tract median because computing one requires the underlying distribution, which the block group median does not carry. A script that averages medians runs cleanly, produces plausible numbers, and stays wrong in a way no error message will ever report.

A defect matters in proportion to what it changes. The mean error in dollars is the technical finding and the change in the recommended list is the consequential one, and a reviewer who reports only the former has not finished the review. Section 5.4 traces how a cartographic mismatch becomes an institutional outcome once a model output drives a funding priority, and section 5.4.1 asks whether funding priorities remain stable under alternative spatial representations. A reviewer tests the neighboring designs before declaring any defect harmless, because an invalid step can leave one design untouched while it changes its neighbors.

Reviewing an artifact is a different skill from producing one and the two do not automatically transfer. Most graduates will spend more of their careers evaluating analyses somebody else produced than producing their own, and this exercise practices that skill most directly, with Exercise VI-B returning to it when students mark up an assistant's unaided risk register. The review also has to land in language the client understands, which is why Step 7 caps the response at two paragraphs and asks for the consequence before the mechanism. A finding nobody acts on has the same effect on the decision as a finding nobody made.