Multimodal Data
多模态数据CommonImages, depth, joint state, force and touch, language, and other signals recorded in sync while a robot works.
In embodied AI, multimodal data means multiple signals recorded in sync while a robot performs a task: multi-view RGB images, depth maps, proprioception (the robot's own joint angles, end-effector pose, and other internal state), force or tactile readings, language instructions, and sometimes sound. A camera alone can't see a contact point hidden by the hand, or how much force was used, so fine manipulation often needs signals beyond vision recorded together. RoboMIND, from the Beijing Humanoid Robot Innovation Center and Peking University, is an example: each trajectory is stored as one HDF5 file containing multi-view RGB-D, proprioceptive state, end-effector state, and the teleoperator's own body state; RoboMIND 2.0 added 12,000 tactile-augmented segments. The difficulties are synchronizing all these signals in time, storage size, and how to train when one modality is missing.
ExampleIn RoboMIND 2.0's tactile-augmented segments, the same instant carries multi-view RGB-D footage and joint state together with normal and shear forces measured by a Tashan tactile sensor.
- Also called
- Omni-modal Data
- Related
- Multimodal Fusion · Tactile Data · Proprioception · Multi-sensor Time Synchronization / Timestamp Alignment · Hierarchical Data Format version 5 · RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation)
- Sources
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation (arXiv HTML)
RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence - As of
- 2025-12