Long-horizon Task
长程任务EssentialA task that requires completing many sub-steps in sequence, over a long stretch of time, to reach the goal.
A long-horizon task requires executing many steps in a row across multiple sub-goals — “clear the table,” for instance, means clearing the plates, dumping the scraps, and loading the sink, in order. There are three main difficulties. First, errors accumulate: a small deviation early on pushes the later state further and further from the training data. Second, the success signal often only appears at the very end, making it hard to tell which intermediate step went wrong. Third, the agent must remember what it has already done and decide what to do next. The CALVIN benchmark (2021) chains 34 sub-tasks into sequences of five instructions each for evaluation; its imitation-learning baseline, trained and tested in the same environment, succeeded on one task in a row about 49% of the time but on all five in a row only 0.08% of the time. Common countermeasures include hierarchical architectures, where a high-level large model breaks down the task while a low-level policy executes atomic skills, along with memory modules and failure recovery.
ExampleClearing a table: put the plates and cups in a bin, throw napkins in the trash, then wipe the surface, in order — if any one step fails, the whole task counts as unfinished.
- Also called
- Multi-stage Task
- Related
- Skill Primitive · Hierarchical Architecture · Compounding Error · Failure Recovery · Embodied Memory · CALVIN Benchmark
- Sources
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv:2112.03227)