学术研究★★★arXiv · 2026-07-16
SceneBind: Binding What and Where Across Vision, Audio and Language
SceneBind introduces an omni-modal representation that combines semantic and 3D spatial understanding across vision, audio, and language, addressing the gap in existing methods regarding explicit spatial structure.
📌 Key points
- SceneBind represents each scene as a semantic-spatial entity, combining a global
- This representation explicitly captures object-level semantics, spatial attribut
- SceneBind addresses the lack of explicit spatial structure in existing omni-moda
本页为 gitzw.com 基于公开来源的 AI 中文解读,非原文转载。