Average Length (CALVIN)
平均完成长度Avg. LenCommonIn the CALVIN benchmark, how many consecutive instructions a policy completes on average, out of 5.
Average Length is the core metric of CALVIN's long-horizon evaluation. During evaluation, a policy attempts 1,000 instruction chains in sequence, each made of 5 consecutive language instructions (for example, open a drawer, then push a block into it), with each subtask capped at 360 steps; the chain ends the moment any subtask fails. The number of subtasks completed in a row (0 through 5) is recorded for every chain, and Avg. Len is the average across all chains. It's equal to the sum of the success rates for completing at least 1, at least 2, … up to all 5 tasks in a row, so it captures both single-step competence and whether errors compound over a long horizon. Papers typically report this number under the ABC→D setting (train on environments A, B, and C, test on the unseen environment D), to compare generalization and long-horizon ability.
ExampleThe CALVIN paper's baseline, MCIL, gets success rates of 48.9%, 12.9%, 2.6%, 0.5%, and 0.08% for completing 1 through 5 tasks in a row under the D→D setting — summing to about 0.65, or an average of well under one completed task.
- Also called
- Avg. Len, Average Successful Sequence Length
- Related
- CALVIN Benchmark · Success Rate · Long-horizon Task · Language-conditioned Policy · Closed-Loop Evaluation · Progress Score
- Sources
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv 2112.03227)
CALVIN 官方评测脚本 evaluate_policy.py (Chinese)
CALVIN 评测工具 utils.py(avg_seq_len 计算) (Chinese)