Embodied AI Glossary中文

TFRecord

TFRecord 格式Advanced

TensorFlow's binary file format, storing a sequence of serialized records back to back.

TFRecord is TensorFlow's data storage format: a file consists of binary records placed one after another, each with a length field and a CRC checksum, and the content is usually a tf.train.Example encoded with protobuf (Google's serialization protocol) — essentially a dictionary mapping field names to lists of values. It's well suited to reading large amounts of data sequentially, and the official guidance recommends splitting data into multiple shards, ideally over 100MB each, to support parallel reads. The downside is that it can only be scanned sequentially — it's awkward to randomly fetch one record by index — and parsing depends on TensorFlow. TFDS uses it as the default storage format, so datasets like RLDS and Open X-Embodiment download as a set of TFRecord shards; newer datasets in the PyTorch ecosystem have largely shifted to Parquet plus video, HDF5, or Zarr instead.

ExampleAn RLDS dataset directory typically contains dataset_info.json and features.json, plus shard files named something like xxx-train.tfrecord-00000-of-01024, where each record stores one complete episode.

Also called
.tfrecord, tf.train.Example
Related
TensorFlow Datasets (TFDS) · RLDS (Reinforcement Learning Datasets) · Protocol Buffers · Apache Parquet · Hierarchical Data Format version 5 · WebDataset
Sources
TensorFlow Tutorial: TFRecord and tf.train.Example
TensorFlow Datasets Overview

See it in the full glossary →