Reassessing the Reliability of WRF-CMAQ-Based PM2.5 Prediction from a Generalization Perspective
Keywords: WRF-CMAQ, Spatial generalization, Temporal generalization, Uncertainty quantification
Abstract. Although data-driven models have been widely used for surface PM2.5 estimation, their reported performance is usually assessed with station-based or sample-based validation schemes, which may not faithfully represent prediction accuracy at unmonitored locations. Building on the numerical-model-informed testbed proposed in the previous work, this study re-examines the accuracy of WRF-CMAQ-based prediction models from the perspective of generalization rather than conventional in-domain fitting metrics. The central concern is that, even when a model performs well at monitored sites or under conventional cross-validation, substantial errors may still exist in station-outside regions where no direct observations are available.
To address this issue, we adopt a generalization-oriented evaluation framework inspired by the proposed idealized testbed and the conceptual design in the project proposal. A physically consistent full-coverage concentration field is treated as an idealized reference, while grids corresponding to monitoring stations are sampled to emulate real-world training conditions. Under this setting, the model is evaluated not only by conventional metrics, but also by its ability to generalize from monitored to unmonitored regions. The study aims to determine whether the apparent accuracy of the WRF-CMAQ-based model is overestimated under conventional evaluation and whether its true predictive precision is limited by generalization error. The results provide a more rigorous basis for diagnosing model bias and for assessing the reliability of data-driven air quality prediction beyond monitoring sites.
