Downstream Task
下游任务EssentialThe specific task a pretrained model is ultimately meant to solve, usually reached by further adaptation or fine-tuning.
“Downstream” is defined relative to “upstream” pretraining: a foundation model is first trained on massive, general-purpose data, and is then applied to some specific task — that specific task is the downstream task. A 2021 survey on foundation models led by Stanford defines a foundation model as one trained on broad, large-scale data that can be adapted to a wide range of downstream tasks. Adaptation can mean using the model zero-shot, giving it a few examples, or fine-tuning it. In embodied AI, a VLA foundation model's downstream task is usually a concrete manipulation job on a specific robot — folding laundry, clearing a table — or a task from a benchmark like LIBERO; papers commonly report “success rate on the downstream task after fine-tuning” as the measure of how valuable the pretraining was.
Exampleπ0 is first pretrained on over 10,000 hours of data spanning 7 robot configurations, then separately post-trained on downstream tasks like folding laundry, clearing a table, and assembling a cardboard box, using anywhere from 5 to over 100 hours of data per task.
- Also called
- Target Task
- Related
- Pre-training · Post-training · Fine-tuning · Foundation Model · Zero-shot · Benchmark
- Sources
- Bommasani et al. 2021: On the Opportunities and Risks of Foundation Models
Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control