Embodied AI Glossary中文

Ego4D

Ego4D 数据集Common

A roughly 3,670-hour egocentric daily-video dataset collected by Meta with 13 universities.

Ego4D was collected by Facebook AI (now Meta) together with a consortium of 13 universities; the paper was posted in October 2021 and published at CVPR 2022. More than 900 participants across 9 countries and 74 locations wore head-mounted cameras to record daily activities such as cooking, repairs, shopping, and socializing, totaling 3,670 hours, with some sequences also including audio, eye tracking, 3D scene meshes, and synchronized multi-camera video, plus benchmark tasks for episodic memory, hand-object interaction, social behavior, and future prediction. For robotics, it is one of the largest sources of human egocentric video and is commonly used to pretrain visual representations; but it has no native 3D hand-pose annotation and no robot actions, so it cannot be used for imitation directly. A license agreement is required before use.

ExampleR3M pretrains on Ego4D video using time-contrastive learning and video-language alignment; once frozen, it serves as a visual module for a Franka arm, which learned several manipulation tasks from just 20 demonstrations.

Related
Egocentric Video · Human Video Data · Ego-Exo4D · R3M · EgoDex · Pre-trained Visual Representation
Sources
Ego4D: Around the World in 3,000 Hours of Egocentric Video (arXiv 2110.07058)
Ego4D 官网 (Chinese)
R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)
As of
2022-02

See it in the full glossary →