PaLI-X
AdvancedA roughly 55-billion-parameter multilingual vision-language model from Google, one of the two backbones behind RT-2.
PaLI-X is a multilingual vision-language model (VLM) that Google Research released in May 2023, a scaled-up version of the earlier PaLI. Its vision encoder is ViT-22B, a 22-billion-parameter Vision Transformer; its language component is a 32-billion-parameter UL2 encoder-decoder; together they total roughly 55 billion parameters. The paper's main finding is that scaling up both the vision and language sides together keeps paying off, and training mixed prefix-completion and masked-token-completion objectives. After fine-tuning, PaLI-X set new state-of-the-art results on more than 15 benchmarks and showed emergent abilities it was never specifically trained for, such as complex object counting and object detection using non-English category names. In embodied AI, PaLI-X is best known as one of the two backbones behind RT-2: RT-2 fine-tuned PaLI-X (55B) and PaLM-E (12B) separately, training each on robot actions represented as text tokens, producing some of the earliest vision-language-action (VLA) models.
ExampleRT-2-PaLI-X-55B: PaLI-X was co-fine-tuned on web-scale image-text data together with robot trajectory data, so it could look at an image, read an instruction, and directly output discretized action tokens.
- Also called
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- Related
- RT-2 · Vision-Language Model · PaLM-E · Vision Transformer · Vision-Language-Action Model · Encoder-Decoder
- Sources
- arXiv 2305.18565: PaLI-X
RT-2 项目主页 (Chinese) - As of
- 2023-07