Embodied AI Glossary中文

Apache Parquet

Parquet 格式Advanced

An open-source column-oriented file format; LeRobot datasets use it to store numeric data like state and actions.

Apache Parquet is an open-source column-oriented data file format, introduced by Twitter and Cloudera in 2013 and made an Apache top-level project in 2015. “Column-oriented” means values from the same column are stored contiguously rather than row by row; grouping similar data together compresses better and lets a reader pull out only the columns it needs. Tools such as Spark, pandas, and DuckDB can all read and write it directly. In embodied AI, Hugging Face's LeRobot dataset format uses Parquet to store low-dimensional, high-frequency data such as joint state, actions, and timestamps, while camera footage is separately encoded as MP4 video, with JSON/Parquet metadata recording where each episode starts and ends within the files. Version 3 packs multiple episodes into the same batch of Parquet files and supports streaming directly from the Hub.

ExampleOpen a LeRobot dataset's data/ directory files with pandas.read_parquet to see columns such as observation.state, action, and timestamp for every frame.

Also called
Parquet
Related
LeRobotDataset · Hierarchical Data Format version 5 · RLDS (Reinforcement Learning Datasets) · Zarr · LeRobot · MCAP
Sources
Apache Parquet 官方文档 Overview (Chinese)
LeRobotDataset v3.0 文档(Hugging Face) (Chinese)
Apache Parquet - Wikipedia
As of
2026-09

See it in the full glossary →