GeoAI Risks Companion exercises and data

Start here

Companion exercises for GeoAI Risks: The New Geography of Uncertainty

David J. Alexander and Talbot J. Brooks.

How this workbook is built

Most exercises in this collection implement a diagnostic the book already names, and the three Part 0 exercises teach the spatial statistics the book leaves to its reader. Chapter 4 suggests a standardized Aggregation Sensitivity Index and calls the missing standard for how much aggregation sensitivity an application can accept an important research frontier, Chapter 5 proposes the index as a family of measures in section 5.6.5, Chapter 7 specifies a Spatial Exploitability Assessment, Chapter 9 specifies a Spatial Information Asymmetry Assessment, and Chapter 12 specifies the GeoAI Feedback Reflex Assessment. The chapters describe those procedures and stop short of running them. These exercises run them, on public data, and hand the student the numbers the chapter argues from.

The collection covers the book one Part at a time. A single exercise carries the several perspectives that a Part develops across its chapters, because the chapters within a Part share a diagnostic vocabulary and separating them would make a student repeat the same setup three times. Each Part offers more than one exercise, and the instructor selects. The header block on every exercise states which chapter sections it reaches, what it demands technically and how long it takes, and the instructor edition adds what each exercise does best, so an instructor can choose without reading the whole thing.

Instructors should expect to assign one exercise per Part in a fourteen week term, which leaves room for the reading and the discussion the Reflection and Insight Questions support. A course carrying a heavier laboratory load can assign two exercises within a Part and set the results against each other, since the exercises in a Part reach different sections of its chapters, often from different data, and the comparison produces something neither exercise generates alone.

Coding is a tool here, never the obstacle

Students taking this course need no prior programming experience and will not be asked to produce working code from a blank file. Every exercise ships a complete, tested program that runs on the first attempt, and the intellectual work sits in choosing the inputs, reading what the program did, and defending what the numbers mean. A student who cannot yet write a loop can complete every required step of every exercise in this collection. Instructors should state that in the first class meeting, because students who expect to be graded on code quality will spend their effort in the wrong place. The same meeting should establish that disclosure of assistant use never lowers a grade, because students who believe disclosure carries a penalty will conceal the very information the log of assistant use exists to collect.

Three tiers organize the technical demand, every step in every exercise carries its tier, and the table below defines them once for the whole collection. An instructor teaching a class with no programming background can assign Tier 1 and Tier 2 and lose none of the analytical content, because every Tier 3 step is optional or carries extra credit. The tier also tells a student how long a step should take before they conclude they are stuck.

Technical demand tiers

TierWhat the student doesProgramming requiredWhere it appears
Tier 1Runs the supplied program without modifying it, reads the printed report, and interprets the numbersNone. The program is complete and tested and runs on the first attemptEvery exercise. Always required
Tier 2Changes stated parameters at the top of the supplied file, reruns, and compares the two reportsNone beyond editing a number in a text file and saving itMost exercises. Usually required
Tier 3Extends the program, normally with an assistant writing the code, and evaluates what came backReading code, never writing it from nothingSome exercises. Always optional or extra credit

The technical footprint stays deliberately small for the same reason. Every required exercise needs Python with the pandas, numpy, scipy and matplotlib libraries, and Exercise II-B adds the python-docx library for the one step that regenerates its memorandum. No exercise requires a geographic information system, a shapefile, a projection library, a cloud account or an application programming interface key. That restriction is possible because the data sources this collection uses publish plain text tables, and it removes the installation failures that consume the first two weeks of most spatial analytics courses.

Every program expects one folder layout. A folder named GeoAI_Exercises holds a toolkit folder with the programs, a snapshots folder with the frozen data, and, for Exercise II-B, a lab_artifacts folder with the delivered files under audit. Every command in this workbook runs from inside toolkit, and each program writes its figures and handoff files into figures and handoff folders it creates beside toolkit.

Typographic conventions

This workbook uses four conventions and uses them consistently. Recognizing them saves the reader from guessing whether a word names a file, a button or a variable. The conventions hold across every exercise, every instructor note and every supporting script in the companion repository, so a reader who learns them once carries them through the collection. Where an exercise introduces a convention of its own, it says so at the point of first use.

ConventionApplied toExample
ItalicFile names, folder names, and published dataset or table namesOpen asi_toolkit.py
BoldCommands you type, menu items, and buttons you clickType python asi_toolkit.py --pull
MonospaceCode, column names, variable names, and web addressesThe column pct_pov holds the percentage
Section 4.4A cross reference into the book itselfThe book specifies this in section 4.4

Numbers that a student must reproduce appear in the Instructor Note at the end of each exercise and never in the student text, so that an exercise can be assigned without handing over its answers.

What every exercise requires you to submit

Submissions take one file, formatted as a Word document, a PDF, a slide deck or a workbook, and every submission carries seven components regardless of which exercise produced it.

The provenance card states the data source, the web address it came from, the snapshot date and checksum from the manifest or the program's banner, and the number of rows. The findings are the numbers the exercise asked for, presented in the table the exercise specifies. The reading is your plain language annotation of the one code block the exercise names, stating what each line consumes, what it produces, and what would change if the line were deleted. The decisions are the choices you made that the exercise deliberately left open, each with a sentence of justification. The flow chart and pseudocode show your intended logic, carry no points, and exist so that a wrong answer can still earn partial credit. The artificial intelligence log answers four questions about when you used an assistant, why you used it at that point, where its output appears in your work, and whether you understand and take responsibility for the result. The handoff package, wherever the exercise's program writes one, is the pair of files the program writes for a human reviewer, being the table of every unit and every statistic and the file naming each decision the machine made without being asked, and your written answers to the questions that second file raises.

Disclosure of assistant use never lowers a grade in this collection. Submitting work you cannot explain is the only failure, and a student can commit that failure with no assistant involved at all. The log also gives your instructor a record of which assistants caught which defects unprompted.

The last mile belongs to a person

This workbook takes a position on where machine work stops, and the position runs through every exercise in it. An artificial intelligence assistant can take a first cut at an analysis, a second cut after correction, and often a third that is close to publishable. What it cannot do is carry the result across the final mile, because that mile consists of judgments about human perception, local knowledge and consequence that the machine has no access to. The workflow this collection teaches therefore ends with a handoff, in which the machine turns over its files, its formats and an explicit list of the decisions it made without being asked, and a person decides what survives.

Workflow diagram. The machine takes three cuts at the analysis and turns over the data files, the working formats, every decision it made without being asked and what to check, and a person carries the result across the final mile and decides what survives.
Workflow diagram. The machine takes three cuts at the analysis and turns over the data files, the working formats, every decision it made without being asked and what to check, and a person carries the result across the final mile and decides what survives.

The book makes this argument repeatedly and under several names. Section 8.14 moves from human-in-the-loop to meaningful human authority, section 15.8 puts human review, local knowledge and red teaming together as one function, and the language of human judgment appears as early as section 1.6, local knowledge enters in Chapter 2, and override becomes a working term in Chapter 8 and returns in Chapter 16. The exercises give that argument something to stand on, because a student who has seen the machine make four unjustified choices in a single map understands the handoff differently from a student who has read that it matters. The handoff also gives an instructor something concrete to grade, since the questions in the decisions file have answers that are either supported by the student's own numbers or are not.

The cartography in this workbook was produced by machine and it is imperfect on purpose. Legends collide with map bodies, class counts are unjustified, and color ramps arrive as defaults that no argument supports. Those defects were left in as specimens, and several exercises ask the student to find and repair them. A reader who concludes that the maps in this collection would be better if a cartographer had drawn them has understood the point the collection is making.

Every exercise built on a program of the collection's own design consequently ends with a handoff package, and the three that do not, Exercises II-A, VI-A and VI-B, say what they produce in its place. The machine writes a table carrying every unit and every statistic it computed, and a second file naming each decision it made without being asked alongside the questions a human must answer before the result is published. Students submit both, and the questions in the second file are the ones the write up has to answer. Those questions were written by the same program that made the decisions, which means the list is incomplete by construction, and several exercises ask the student to add the question the machine failed to raise.