Long-tail Problem
长尾问题CommonThe mass of individually rare situations that training data barely covers, and where models fail most often.
“Long-tail” comes from the long-tail distribution in statistics: a small number of common situations (the head) account for most of the data, while a huge number of rare situations (the tail) each occur only a little but add up to a large total. Autonomous driving was among the first fields to treat this as a central challenge: driving is normal almost all the time, but occasionally the car meets a corner case — a construction detour, debris on the road, a pedestrian suddenly running out — and these scenarios have little data even though they are often the ones that matter most for safety. Embodied AI faces the same issue: the objects, arrangements, and unexpected situations in a home vary endlessly, and demonstration data is naturally skewed toward a handful of common tasks; the ICRA 2026 paper “Beyond the Majority” finds that general-purpose robot policies generalize noticeably worse on tail tasks with sparse data, and that ordinary resampling methods offer only limited help. Countermeasures include continuously collecting failure data from deployment, called a data flywheel, using simulation and synthetic data to fill in the tail, and improving a model's generalization and failure-recovery ability. The long tail is the main source of the gap between “the demo works” and “the product actually ships.”
ExampleA household robot can reliably clear away common bowls and chopsticks, but a spilled bowl of soup, a fork stuck in a crack in the table, or a pet suddenly jumping onto the table are all long-tail situations.
- Also called
- Corner Case, Long-tail Distribution
- Related
- Out-of-Distribution · Generalization · Data Flywheel · Failure Recovery · Robustness · Autonomous Driving
- Sources
- Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation (ICRA 2026)
Dynamically Conservative Self-Driving Planner for Long-Tail Cases - As of
- 2026-02