Apache Parquet
Parquet 格式AdvancedAn open-source column-oriented file format; LeRobot datasets use it to store numeric data like state and actions.
Apache Parquet is an open-source column-oriented data file format, introduced by Twitter and Cloudera in 2013 and made an Apache top-level project in 2015. “Column-oriented” means values from the same column are stored contiguously rather than row by row; grouping similar data together compresses better and lets a reader pull out only the columns it needs. Tools such as Spark, pandas, and DuckDB can all read and write it directly. In embodied AI, Hugging Face's LeRobot dataset format uses Parquet to store low-dimensional, high-frequency data such as joint state, actions, and timestamps, while camera footage is separately encoded as MP4 video, with JSON/Parquet metadata recording where each episode starts and ends within the files. Version 3 packs multiple episodes into the same batch of Parquet files and supports streaming directly from the Hub.
ExampleOpen a LeRobot dataset's data/ directory files with pandas.read_parquet to see columns such as observation.state, action, and timestamp for every frame.
- Also called
- Parquet
- Related
- LeRobotDataset · Hierarchical Data Format version 5 · RLDS (Reinforcement Learning Datasets) · Zarr · LeRobot · MCAP
- Sources
- Apache Parquet 官方文档 Overview (Chinese)
LeRobotDataset v3.0 文档(Hugging Face) (Chinese)
Apache Parquet - Wikipedia - As of
- 2026-09