← 返回信息流

arXiv cs.AI论文

AffordDrive3D:基于空间理解的具身感知世界-动作建模

arxiv.org作者:Tianhui Cai, Xinglong Sun, Chao Fang, Zhenxin Li, Rui Song, Jose M. Alvarez, Yunxiang Mao, Jiaqi Ma, Langechuan Liu论文AI评分:70/100

针对自动驾驶中现有世界模型主要依赖RGB外观预测、几何信息仅描述场景布局而缺乏驾驶相关性的问题,提出AffordDrive3D。该方法引入具身感知(Affordance-Aware)机制,联合学习未来场景预测与轨迹生成,通过聚焦与自车相关的空间区域提升自动驾驶的决策能力。

这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。

Abstract:World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.

阅读原文