A Machine Learning Framework for Spatially Honest Benchmarking and Environmental Justice Auditing of Hyperlocal PM 2.5 Predictions – American Journal of Student Research

American Journal of Student Research

A Machine Learning Framework for Spatially Honest Benchmarking and Environmental Justice Auditing of Hyperlocal PM 2.5 Predictions

Publication Date : Aug-20-2026

DOI: 10.70251/HYJR2348.4410241039


Author(s) :

Nathan Tan, Saketh Chebrolu.


Volume/Issue :
Volume 4
,
Issue 4
(Aug - 2026)



Abstract :

Machine learning is increasingly used to map fine particulate matter (PM2.5) between sparse monitors, but models are commonly validated with random splits that place records from the same sensor in both partitions, inflating apparent accuracy. This study assembles a year of quality-assured daily observations from 226 PurpleAir sensors across four Texas metropolitan areas (54,325 sensor-day records) and evaluates four regression algorithms and a tree ensemble under three protocols: a random 80/20 split; leave-one-sensor-out cross-validation (LOSO-CV) with the withheld sensor’s own recent PM2.5 history available (transfer); and LOSO-CV without any site history (cold start), the setting that corresponds to an unmonitored location. When site history is available, random and LOSO scores nearly coincide. Removing site history exposes the true gaps: the ensemble falls from R² = 0.825 (transfer) to 0.662 (cold start; RMSE = 2.37 μg/m³), the linear baseline collapses from 0.469 to 0.037, and random splitting overstates cold-start ensemble skill by 0.176 R². Environmental-justice (EJ) indicators are excluded from the deployed cold-start model so the equity audit is independent of the audited variables. Deployed across 3,758 census tracts with daily 2025 meteorology and averaged to true annual means, predicted annual PM2.5 correlates modestly with the EJScreen EJ index (r = 0.129, p < 10−14), far below the r = 0.875 obtained when EJ variables were model inputs; tract exceedance of the 9 μg/m³ standard rises from 0.2% to 6.0% across EJ quartiles. Spatially honest validation and audit-independent models are prerequisites for credible exposure and equity mapping.