How do you find the most polluted street in a city of three million when you cannot afford to put a sensor on every corner? You borrow the eyes of a dozen satellites, and you accept some hard compromises.
This is a build log. Not a polished result, but the actual reasoning behind one: how I built a Composite Air Quality Risk Index for Gurugram and Faridabad, two rapidly growing cities on the edge of Delhi, using Google Earth Engine and a stack of satellite data. The goal was deceptively simple. Produce one map that ranks every square kilometre by how much of an air quality problem it represents, so that a team with limited time can decide where to send people first.
The interesting part is not the final map. It's the decisions along the way, each of which could have gone differently and each of which bakes an assumption into the result. I want to walk through those honestly, including the places where the method is on shakier ground than a clean final visual would suggest. If you have read my piece on geospatial tools or the one on data trust in the field, this is those ideas applied to a concrete problem, with the receipts.
Gurugram and Faridabad sit in one of the most polluted airsheds on Earth. In winter, the whole Indo-Gangetic plain fills with a grey soup of vehicle exhaust, construction dust, industrial emissions, crop-residue smoke drifting in from Punjab, and the output of thousands of brick kilns. Everyone who lives there knows the air is bad. That is not the question. The question a planner actually faces is narrower and harder: where, specifically, and who is most affected?
The official monitoring network cannot answer this. There are a handful of reference-grade CAAQM stations across both cities, which is fine for reporting a citywide number but useless for telling you that the air on one particular road in Sector 7 is far worse than three blocks away. Pollution is intensely local. It spikes near a construction site, an unpaved road, a cluster of kilns, and drops off within hundreds of metres. To see it at that resolution, you need either thousands of sensors, which nobody is funding, or you need to get creative with data that already exists.
So the framing became: build a 1 km grid over both cities, and for every cell, estimate three things. How polluted is the air here? How many pollution sources are packed into this cell? And how many vulnerable people are exposed? Combine those into a single score. It sounds tidy. Each of those three questions turned out to hide a small mountain of compromises.
The index is built from three layers, combined with fixed weights: 40% ambient pollution, 35% source density, 25% exposure. Before defending those numbers, it's worth being clear about what each layer is trying to capture, because they answer genuinely different questions.
Ambient pollution asks what is actually in the air right now. Source density asks where pollution is being produced, which matters because it is more actionable: you cannot regulate a concentration, but you can regulate a crusher or pave a road. Exposure asks who is downwind of all of it, because a pollution hotspot in an empty field matters less than a moderate one next to a school. The whole point of the third layer is to stop the index from simply pointing at the dustiest industrial patch and calling the job done.
Why 40/35/25 and not equal thirds? Because the three layers are not equally trustworthy. The ambient layer is closest to a direct measurement of the thing we actually care about, so it earns the largest share. Source density is a strong signal but it is a proxy, built from indirect evidence, so it gets slightly less. Exposure is important for prioritisation but it is about who is affected rather than how bad the air is, so it plays the role of a weighty tiebreaker rather than a lead. That is the reasoning. It is defensible, not inevitable, and the slider above exists precisely so you can disagree with it.
The ambient layer fuses six satellite and sensor-derived signals. The workhorse is MAIAC AOD, aerosol optical depth from MODIS, which measures how much light is scattered by particles in the atmospheric column. It is the closest thing to a direct read on particulate load, available at roughly 1 km. Around it sit tropospheric NO₂ from TROPOMI, a methane anomaly layer for landfills, an SO₂ signal keyed to kiln season, a UV aerosol index, and carbon monoxide.
| Sub-indicator | Resolution | Weight |
|---|---|---|
| AOD (MODIS MAIAC) | 1 km | 30% |
| Tropospheric NO₂ (TROPOMI) | 3.5 km | 20% |
| CH₄ anomaly / landfill | 7 km | 20% |
| SO₂ kiln season | 3.5 km | 15% |
| UV aerosol index | 3.5 km | 10% |
| Carbon monoxide | 7 km | 5% |
Look at that resolution column, because it contains the first uncomfortable truth. AOD arrives at 1 km. But NO₂ is 3.5 km, and methane and CO are at 7 km. When I say a 1 km grid cell has a certain pollution score, part of that score comes from a methane pixel that is 49 square kilometres in size. I am effectively smearing a coarse measurement across a fine grid and pretending it has detail it does not have. This is standard practice. It is also, if you are honest, a place where the map claims more precision than the underlying physics supports. I resampled everything to the 1 km grid, but resampling does not create information. A 7 km pixel painted onto 49 one-kilometre cells is still one number wearing forty-nine hats.
"A 7 km pixel painted onto forty-nine one-kilometre cells is still one number wearing forty-nine hats. Resampling changes how data looks, not how much it knows."
This is the layer I want to spend the most time on, because it is where the real intellectual work happens and where the method is simultaneously most clever and most fragile. Here is the core problem: you cannot see most pollution sources from space directly. There is no satellite band that says "brick kiln" or "construction dust." So the entire source layer is built from proxies: visible things that correlate with the invisible thing you actually care about.
Consider construction and demolition dust, one of the biggest contributors to particulate pollution in these cities. You cannot detect "dust being kicked up" from orbit. So instead you detect bare ground, using the Dynamic World bare-probability band and a Bare Soil Index computed from Sentinel-2. The logic is that freshly cleared or disturbed land, sitting bare where it was recently vegetated or built, is a strong signal of active construction. It is a good proxy. It is also blind to a fully enclosed site, and it will happily flag a genuinely barren patch of natural ground as if it were a building site.
Brick kilns are the proxy story at its most vivid. Kilns have a distinctive footprint: an oval clamp of fired earth, often with a chimney, sitting in cleared land near a river where the clay is. So the detection stacks several weak signals: a clay index from Landsat, thermal signatures, the shape and bareness of the surrounding ground, and SAR backscatter change from Sentinel-1 that catches the seasonal firing cycle. No single one of these is a kiln detector. Together they produce a "kiln candidate" layer that is right often enough to be useful and wrong often enough that I would never present it as a census.
| Sub-indicator (proxy) | Resolution | Weight |
|---|---|---|
| Dynamic World bare ground | 10 m | 22% |
| GHSL built-up surface | 10 m | 15% |
| BSI temporal variability | 10 m | 13% |
| Weighted road density (OSM) | 30 m | 10% |
| VIIRS nighttime lights | 500 m | 10% |
| SAR backscatter change (S1) | 10 m | 10% |
| Brick kiln candidates (L9) | 30 m | 10% |
| GHSL building volume | 10 m | 5% |
| NDBI / BSI mean (S2) | 10 m | 5% |
This is the deepest idea in the whole project, so it is worth stating plainly. Proxies are individually unreliable but collectively informative. Bare ground alone is a mediocre construction detector. Nighttime lights alone is a mediocre industry detector. Road density alone tells you about traffic but also about ordinary residential streets. But a cell that lights up on bare ground and SAR change and road density and nighttime lights, all at once, is very probably a real source cluster. The proxies are noisy in different directions, so stacking them cancels much of the noise. This is the same logic as the reporting-bias discussion in my data-trust essay, pointed at satellites instead of surveys.
The honest caveat is that the proxies can also share the same blind spot. If a source produces no visible surface disturbance, no heat, no traffic, and no light, it is invisible to all of them at once, and no amount of stacking recovers it. The method sees what leaves a physical fingerprint. Cleaner-looking but genuinely polluting operations can slip through. I would rather say that out loud than let a confident-looking map imply otherwise.
The exposure layer is the most ethically loaded and, in some ways, the most straightforward technically. It combines population density from WorldPop, a deprivation proxy, open-buildings density, and access to schools and hospitals measured as distance and travel time. The point is to weight the final index toward places where pollution meets people, and especially where it meets people with the fewest resources to protect themselves.
The reason this layer matters is that without it, the index would be an environmental map, not a human one. It would faithfully point at the dustiest quarry and ignore the crowded informal settlement beside a moderately polluted road. The exposure weighting is what turns "where is the air worst" into "where should someone go first," which is the question that was actually asked. The deprivation proxy is the part I would most want to validate against ground truth before leaning on it hard, because deprivation estimated from remote data is a proxy on a proxy, and the stakes of getting it wrong are borne by exactly the people the layer is meant to protect.
A 1 km composite score is useful for triage but useless for action. Nobody visits a square kilometre. They visit a street. So the final step drops from the grid down to street level: take the highest-scoring cells, and inside each, use fine-resolution building heights, road classes, and the locations of schools and hospitals to produce a specific, mappable location with a 500 metre context radius. This is where the index stops being an abstraction and becomes a list of places a team can actually walk to.
This descent is honest only because the fine-resolution layers underneath are genuinely fine. The building and road data really is at 10 to 30 metres. What I am not doing is pretending the 1 km pollution score is valid at 10 metres. The hotspot is chosen using the coarse score, but the within-hotspot context, which building, which road, which school, comes from data that legitimately has that resolution. Keeping those two things separate, the coarse ranking and the fine context, is what keeps the final output defensible.
Here is one of the ranked hotspots, Sector 10 in Gurugram, rendered at street level. The composite score flagged this cell; the fine data tells you why it is a problem and who is nearby. Toggle the building heights, highlight the roads by class, and mark the facilities. The expressway cutting through is exactly the kind of source-and-exposure combination the whole index exists to surface.
Zoom out from this single corner and you have the whole point of the project: hundreds of these, ranked, each one a specific place with a specific reason to worry. The interactive atlas lets you explore the composite score at hexagon level across all of Gurugram and Faridabad, drilling from the citywide pattern down to exactly this kind of street scene. What you are looking at above is the last mile of that pipeline: the moment an abstract index becomes an address.
If someone handed you this index and asked you to act on it, here is what you should know. It is a strong relative ranking and a weak absolute one. It is very good at telling you that cell A is a bigger concern than cell B. It is not calibrated to tell you the actual concentration of anything in real units. Treat it as a prioritisation tool, a way to spend a limited field budget well, not as a measurement.
It inherits every bias in its inputs. Where the satellite record is thin, where OpenStreetMap coverage is patchy, where the deprivation proxy is guessing, the index is guessing too, just less visibly. The winter and summer variants matter enormously and a single annual number would have been misleading. And the whole thing is a snapshot of a 2023 to 2024 baseline in cities that are changing month to month. A construction site that closed last year is still in the data.
None of this makes the index useless. It makes it a model, with all that implies. It captures a real and useful slice of a genuinely hard problem, and it is honest about the slice it misses. That combination, useful and clear about its own limits, is the most any model like this can offer. The satellites gave us eyes we could not otherwise afford. The compromises are the price of admission, and the least I can do is show you the receipt.