The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
Download
Share
Publications Copernicus
Download
Citation
Share
Articles | Volume XLIX-B2-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1191-2026
https://doi.org/10.5194/isprs-archives-XLIX-B2-2026-1191-2026
23 Jul 2026
 | 23 Jul 2026

Enhancing Vision-Based Perception in Autonomous Driving: YOLO11–DETR Integration with Selection Model

Ahmed M. Reda, Naser El Sheimy, and Adel Moussa

Keywords: Autonomous Driving, Object Detection, Selection Model, Domain Shift, YOLO11, RT-DETR

Abstract. Vision-based object detection is a key component of autonomous driving perception systems; however, models pretrained on large-scale generic datasets usually struggles when implemented in automotive environments due to domain shift. This research introduces a comprehensive evaluation and fusion of YOLO11 and RT-DETR for improving robustness in autonomous driving scenarios using KITTI dataset. Both models are pretrained on COCO dataset and evaluated under a zero-shot transfer setting to assess cross-domain generalization. The results show that RT-DETR-L and RT-DETR-XL experience significant performance degradation, dropping from 53.0 and 54.8 𝑚𝐴𝑃 on COCO to 34.3 and 34.5 on KITTI, respectively. In contrast, YOLO11-Nano and YOLO11-L demonstrate better generalization, achieving 44.0 and 51.3 𝑚𝐴𝑃 on KITTI compared to 40.9 and 55.0 on COCO. Controlled fine-tuning experiments (10 and 100 epochs) are conducted to analyze adaptation dynamics. RT-DETR-L improves to 61.7 and 79.1 𝑚𝐴𝑃, while YOLO11-L reaches 64.9 and 76.2 after 10 and 100 epochs, respectively. To further evaluate robustness under challenging conditions, three degraded data subsets are generated. Building on the strengths of convolutional and transformer-based detectors, this work introduces an image-based selection model that selects the most suitable detector for each input image. Experimental results demonstrate substantial zero-shot degradation, strong recovery after fine-tuning, and consistent performance improvements under degraded conditions using the proposed selection strategy. Our method achieves gains of up to 4 𝑚𝐴𝑃 points over the best standalone detector without incurring the computational overhead. The proposed framework provides a context-aware and computationally efficient perception enhancement strategy suitable for real-world autonomous driving systems.

Share