论文解读 · 作者已确认

MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects

作者: , Xiaohan Wang, Linchao Zhu, Yi Yang

IEEE/CVF International Conference on Computer Vision (ICCV 2023) · 2023 · 页 217-227

3D affordance learningarticulated objectsmultimodal learningautoencoderrobotic interaction

论文信息

作者
Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, and Yi Yang
推荐论文引用
Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, and Yi Yang. “MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects.” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023), 217-227. https://doi.org/10.1109/ICCV51070.2023.00027.

版本日期

来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

MAAL 用单阶段 autoencoder pipeline 学习 3D articulated object affordance,其中 MultiModal Energized Encoder 联合建模物体几何、机器人动作及其交互,并只需要少量 positive sample。

研究问题

Affordance prediction 需要从异构的物体与动作信息中判断机器人可以在哪里、以何种方式操作;early fusion 与多阶段 critic pipeline 对这些模态的利用可能不充分且训练低效。

论文贡献

  • 将 3D articulated-object affordance learning 重写为一次训练完成的 multimodality-aware autoencoder pipeline。
  • 提出 MultiModal Energized Encoder,联合建模物体与机器人动作模态。
  • 面向数据效率,只需少量成功交互样本进行训练。

证据与评测范围

PartNet-Mobility 上的实验和可视化支持其多模态学习与 affordance prediction 结论;准确任务定义、actionability 指标和比较应引用 ICCV 正式论文。

适用范围与局限

论文展示范围是 PartNet-Mobility 设置下的 articulated-object affordance learning;摘要并未证明它在未建模的真实传感、操作硬件或物体类别上同样有效。

Related work 定位

MAAL 同时改变了 affordance learning 的训练流程和多模态表示,以一次训练的 autoencoder formulation 替代 early-fusion、多阶段打分。

Related Work 表述

Liang 等提出 MAAL,通过 MultiModal Energized Encoder 联合表示三维物体与机器人动作信息,并以单阶段 autoencoder 框架学习 articulated-object affordance。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to act and how to act. Correspondingly, the task mainly requires producing actionability scores, action proposals, and success likelihood scores according to the given 3D object information and robotic information. Current works usually directly process multi-modal inputs with early fusion and apply critic networks to produce scores, which leads to insufficient multi-modal learning ability and inefficiently iterative training in multiple stages. This paper proposes a novel Multimodality-Aware Autoencoder-based affordance Learning (MAAL) for the 3D object affordance problem. It is an efficient pipeline, trained in one go, and only requires a few positive samples in training data. More importantly, MAAL contains a MultiModal Energized Encoder (MME) for better multi-modal learning. It comprehensively models all multi-modal inputs from 3D objects and robotic actions. Jointly considering information from multiple modalities, the encoder further learns interactions between robots and objects. MME empowers the better multi-modal learning ability for understanding object affordance. Experimental results and visualizations, based on a large-scale dataset PartNet-Mobility, show the effectiveness of MAAL in learning multi-modal data and solving the 3D articulated object affordance problem.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; MAAL and MultiModal Energized Encoder sections
评测结论Abstract; PartNet-Mobility experiments and visualizations

主要核验来源: IEEE version of record (ICCV 2023 version of record; CVF open-access copy cross-checked).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, and Yi Yang. “MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects.” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023), 217-227. https://doi.org/10.1109/ICCV51070.2023.00027.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源