Integrating Multi-Source Agricultural Data with Machine Learning to Improve Crop Mapping Accuracy: A Case Study of the Navajo Nation
Keywords: Crop mapping, Sentinel-2, Cropland Data Layer, Random Forest, Navajo Nation
Abstract. Accurate crop maps are important for agricultural monitoring in water-limited regions because they provide spatial information for crop inventory assessment, land management, and resource planning. In the Navajo Nation, crop classification is challenging because agriculture is influenced by arid environmental conditions, limited water availability, and unevenly distributed cultivated land. This study evaluates a crop-classification workflow for a selected agricultural Region of Interest (ROI) within the Navajo Nation using Sentinel-2 imagery, the USDA/NASS Cropland Data Layer (CDL), and the CDL confidence layer in Google Earth Engine. Highconfidence CDL pixels (confidence ≥ 95%) were used to construct pseudo-reference samples for the 2017 and 2022 growing seasons, and a 3 × 3 neighborhood homogeneity filter was applied to reduce local label uncertainty. Spectral predictors derived from Sentinel-2 imagery included the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Green Chlorophyll Vegetation Index (GCVI), and Land Surface Water Index (LSWI). A Random Forest classifier was implemented separately for each year using an 80% training and 20% testing split. The resulting classifications achieved overall accuracies of 87.30% for 2017 and 90.88% for 2022. These results show that confidence-screened CDL samples combined with multi-temporal Sentinel-2 features can support reliable crop classification within the selected ROI under limited reference-data conditions and provide a practical basis for agricultural monitoring in the Navajo Nation.
