October 4, 2026 · Tupll

How Analysts Turn Messy Geographic Data Into Defensible Site Rankings

Geographic data arrives broken. Census tables keyed to boundaries that changed, business listings with duplicate and dead records, demand estimates published at geographies nothing else matches, coordinates that miss the building. Anyone who has actually assembled a site analysis knows the dirty secret: most of the work is not analysis, it is making the data trustworthy enough to analyze.

That work is invisible in the final deliverable, and it is precisely what separates a defensible ranking from a confident-looking guess.

Where the mess comes from

Geographic data has no common grid. Demographics come in census geographies, employment in different ones, spending estimates in vendor polygons, competitors as scraped points of unknown freshness. Combining them means resolving every source to the same places on the map, and every resolution step is a chance to quietly corrupt the result.

The classic failure is aggregation error: assigning a whole zone's characteristics to a site that sits at its edge, where the zone's average describes people who are functionally far away. Rankings built on carelessly joined data can flip entirely when the joins are done right.

The unglamorous pipeline

A defensible analysis runs the boring steps every time, the same way. Normalize geographies to a consistent resolution around each candidate point, in multiple bands, because composition changes with distance. Deduplicate and verify business records rather than trusting counts. Date-stamp everything, since a market read built on stale vintages describes a market that no longer exists. And document every transformation, so any number in the output can be traced to its source.

In our own work, this pipeline is most of the machinery: pulling 40 to 50 variables per zone at a specific point on the map, consistently, is the hard part. The modeling on top is well-established statistics; it is only as good as the discipline underneath.

From clean data to a ranking that holds

Clean inputs make the ranking step almost anticlimactic. Score every candidate with the same variables, weighted by what actually predicts performance for the specific operation, and publish the reasoning next to the number.

Defensibility comes from that combination: consistent inputs, disclosed method, inspectable weights. When someone challenges the ranking, and someone always does, the conversation goes to a specific variable or weight, where evidence can settle it. Rankings without that foundation get challenged at the level of trust, where nothing can.

A quick audit for any analysis you're handed

Four questions expose most weak work. What geography is each input measured at, and how was it resolved to the site? How fresh is each source? How were duplicates and errors handled? Can any number in the summary be traced backward to raw data?

If the answers are vague, the polish of the deliverable is doing the persuading, not the evidence. Messy data is normal. Unexamined data is a choice.


← Back to all posts