| 引用本文: | 蒋留兵,刘仁超,车 俐,吕连辉. GS-CA:基于不确定性驱动的跨模态雷达相机三维目标检测[J]. 雷达科学与技术, 2026, 24(3): 299-307.[点击复制] |
| JIANG Liubing, LIU Renchao, CHE Li, LU Lianhui. GS-CA: Uncertainty-Driven Cross-Modal 3D Object Detection with Radar and Camera[J]. Radar Science and Technology, 2026, 24(3): 299-307.[点击复制] |
|
| 摘要: |
| 本设计为探讨毫米波雷达与相机融合相较纯视觉方案的优势,以及解决动态模态可靠性的挑战,引入Gumbel-Softmax可微分的离散选择机制设计了新的多模态融合框架(GS-CA),一种用于处理BEV视角下的3D感知任务的自适应相机-雷达融合方法。在融合期间增加了特征不确定性的描述,并加入到模型训练当中,增加可解释性的同时为模型的优化提供改进方向。设计相应的可变形交叉注意力机制增加两模态BEV特征的交互。模型评估中,所用方法在nuScenes数据集上呈现可观的性能,目标检测模型除了对普通场景外还引发了对雨天、黑夜等场景的思考。所设计的方法创建了融合新范式的同时,在传感器融合技术发展方面具有一定的指导意义。 |
| 关键词: 多模态融合 3D感知 特征不确定性 可变形交叉注意力 传感器融合 |
| DOI:DOI:10.3969/j.issn.1672-2337.2026.03.007 |
| 分类号:TN957 |
| 基金项目:国家自然科学基金(61561010);广西创新驱动发展专项(桂科AA21077008);广西无线宽带通信与信号处理重点实验室2022年主任基金(GXKL06220102,GXKL06220108);八桂学者专项经费资助(2019A51);桂林电子科技大学研究生教育创新计划资助(2022YXW07,2022YCXS080);2022年广西高等教育本科教学改革工程项目(2022JGB196);桂林电子科技大学学位与研究生教改项目(2022YXW07,2023YXW02);广西研究生教育创新计划资源(YCSW2022271) |
|
| GS-CA: Uncertainty-Driven Cross-Modal 3D Object Detection with Radar and Camera |
|
JIANG Liubing, LIU Renchao, CHE Li, LU Lianhui
|
|
School of Information and Communication, Guilin University of Electronic Technology, Guilin 541004, China
|
| Abstract: |
| This work investigates the advantages of millimeter-wave radar-camera fusion over vision-only approaches and addresses the challenge of dynamic modality reliability. Therefore,a new multimodal fusion framework, GS-CA, an adaptive camera-radar fusion method for handling 3D perception tasks in the BEV perspective, is designed by introducing the discrete selection mechanism of the Gumbel-Softmax differentiable. The description of feature uncertainty is added during fusion and incorporated into the model training to increase the interpretability and provide improvement directions for model optimization. Additionally, a corresponding deformable cross-attention mechanism is designed to increase the interaction of the two-modal BEV features. In the model evaluation, the methods show considerable performance on the nuScenes dataset, and the target detection model demonstrates promising results in rainy and dark scenes in addition to ordinary scenes. Overall, the proposed framework establishes a new paradigm for multimodal fusion and offers valuable insights for the advancement of sensor fusion technologies. |
| Key words: multimodal fusion 3D perception feature uncertainty deformable cross-attention sensor fusion |