A Study on 3D Object Detection of Radar-visual Fusion Based on Cross-attention Mechanism
-
Abstract
Extremely high requirements are put forward for the reliability and accuracy of perception algorithms in complex driving scenarios of autonomous driving. The premise for the system to stably complete tasks such as environmental perception, target recognition, and decision-making planning is that the state and attribute information of road traffic targets can be accurately acquired and fully characterized under various complex weather conditions, such as variable illumination, rain, snow, and haze. Aiming at the practical challenges in the field of autonomous driving perception, including large differences in heterogeneous features of multi-modal sensors, weak information correlation, and high fusion difficulty, a radar-vision joint perception method integrated with attention mechanism is proposed in this paper. The four-dimensional millimeter-wave radar point cloud is converted into a pseudo-image representation that retains spatial height dimension information, and channel attention and cross-attention are introduced for collaborative modeling, thereby realizing the adaptive correlation, bidirectional interaction, and weight enhancement of radar and visual heterogeneous features. The results of comparative experiments carried out based on the view-of-Delft dataset show that the average precision of three-dimensional object detection of the proposed method can reach 45.3 %, which is 28.3 % higher than that of the baseline model, and the detection accuracy and robustness of key traffic targets such as pedestrians and vehicles under complex working conditions can be effectively improved.
-
-