{"lang":"en","today":"2026-09-30","cats":[{"key":"concept","name":"Core Concepts & Tasks","blurb":"Start with the map: what embodied AI does, the tasks robots perform, and what terms like ‘generalization’ and ‘cross-embodiment’ mean.","sections":[{"title":"What is embodied AI","blurb":"The starting point: what embodied AI means, how it differs from disembodied AI, its parts, and why it’s hard."},{"title":"Basic tasks robots perform","blurb":"With the concepts in place, see what robots actually do: manipulating objects, moving themselves, navigating, and chaining it into long tasks."},{"title":"Policies: from observation to action","blurb":"Unify these tasks into one observation-to-policy-to-action loop, then look at the types of policies and how systems are layered."},{"title":"Generalization: handling what you haven’t seen","blurb":"Now judge how good the policy is: does it still work with new objects, scenes, instructions, or even a different robot?"},{"title":"Manipulation tasks in detail","blurb":"Back to the manipulation thread: beyond grasping, bimanual, and dexterous skills, a closer look at harder, more specialized tasks."},{"title":"Walking, navigation, and exploration","blurb":"Now the locomotion thread: how legged robots stay balanced over rough terrain, and how they navigate and explore toward different goals."},{"title":"Factories, warehouses, and homes","blurb":"Put these tasks into real settings: from tidy, controlled factories and warehouses to cluttered, unpredictable homes."},{"title":"Working alongside people and robots","blurb":"Real settings always include people: human-robot interaction, collaboration and safety, plus how multiple robots coordinate with each other."},{"title":"Understanding language and the world","blurb":"Shift from doing to understanding: how a robot’s ‘brain’ parses instructions, reasons about space and physics, perceives, and remembers."},{"title":"Long-term goals and origins","blurb":"Finally, the big picture: how to measure the endgame, the scale-versus-experience debate, and the ideas this field grew out of."}]},{"key":"robot","name":"Robot Types & Products","blurb":"Meet the robots: body types from robot arms to humanoids, and the notable models you’ll see on the market.","sections":[{"title":"The robot-arm family","blurb":"Starting with the most common robot: the robot arm, and its variants in factories, labs, and on desktops."},{"title":"Wheels, legs, and mobile manipulation","blurb":"A robot arm stays put; this section covers how robots move — wheels, legs, and arms mounted on a moving base."},{"title":"Types of humanoid robots","blurb":"Add a torso and two arms to bipedal legs and you get a humanoid; they vary by height, lower body, and looks."},{"title":"Other forms, by use case","blurb":"Shifting from looks to purpose: service and special-purpose robots beyond industry, plus flying, bio-inspired, and soft robots."},{"title":"Robot arms used in research","blurb":"Now for specific models: the robot arms that show up most often in papers, plus other options in the same class."},{"title":"Bimanual and mobile manipulation rigs","blurb":"Building platforms from those arms: bimanual teleop rigs, research mobile bases, and bases combined with arms for mobile manipulation."},{"title":"Humanoid robots abroad","blurb":"Starting with humanoids from outside China: from ASIMO to today’s star startups, plus newcomers, home robots, classics, and open-source projects."},{"title":"Quadrupeds abroad and more","blurb":"Beyond humanoids: quadrupeds from companies like Boston Dynamics, plus specialized robots for warehouses, surgery, and more."},{"title":"Humanoid robots in China","blurb":"Back to China: leading firms like Unitree and AgiBot first, then national teams, crossover entrants, and other rising players."},{"title":"Quadrupeds in China and more","blurb":"Finally, Chinese quadrupeds: mainly Unitree’s line from A1 and Go2 to B2, plus a few other makers."}]},{"key":"hardware","name":"Hardware & Body Parts","blurb":"Take the robot apart: what motors, gearboxes, lead screws, dexterous hands, and compute platforms each actually do.","sections":[{"title":"Actuators and joint modules","blurb":"Having seen robot forms, start with the parts that make them move: actuators, joint modules, and servos."},{"title":"Motor types and specs","blurb":"Open up a joint module and look at the motor: the main types, and how to read torque and heat on a spec sheet."},{"title":"Gearboxes, bearings, and lead screws","blurb":"Motors spin fast but weak: gearboxes trade speed for torque, bearings carry the load, and lead screws turn rotation into straight-line motion."},{"title":"Drivers, encoders, and buses","blurb":"After mechanical parts, the electronics: drivers manage current, encoders measure angle, buses carry commands, plus power-loss protection."},{"title":"Joint-drive design choices","blurb":"With the parts covered, see how they combine: quasi-direct drive, series elastic, hydraulic, and tendon-driven approaches, trading off gear ratio."},{"title":"Grippers and end effectors","blurb":"From joints to the arm’s tip: the common two-finger gripper, plus suction cups, soft grippers, and quick-change interfaces."},{"title":"Dexterous hands: design and products","blurb":"More hand-like than a gripper: finger joints and drive methods first, then the leading dexterous-hand products."},{"title":"Mobile bases and legs","blurb":"From hands to the lower body: the wheel layouts used in mobile bases, and what legs and feet look like on legged robots."},{"title":"Compute, controllers, and processing","blurb":"After limbs, the brain: the onboard computer and compute specs that run models, plus the real-time controllers underneath."},{"title":"Batteries, frame, and lab gear","blurb":"Rounding out the whole robot: battery life, wiring, materials and system specs, plus equipment labs commonly pair with it."}]},{"key":"mechanics","name":"Mechanics & Kinematics","blurb":"The physics behind robot motion: representing pose, forward and inverse kinematics, Jacobians, dynamics, and balance.","sections":[{"title":"Frames and pose","blurb":"The starting point: writing an object’s position and orientation as numbers with coordinate frames, then converting between frames."},{"title":"Ways to represent rotation","blurb":"Expanding on ‘orientation’: Euler angles, quaternions, axis-angle, and 6D representations each have trade-offs, then on to Lie groups and screws."},{"title":"Links, joints, and mechanisms","blurb":"From a single rigid body to a whole robot: how links and joints form a mechanism, and its degrees of freedom."},{"title":"Forward and inverse kinematics","blurb":"Once the structure is fixed, compute away: converting between joint angles and end-effector pose, plus modeling, reach, and calibration."},{"title":"Velocity and the Jacobian","blurb":"Moving from position to velocity: the Jacobian converts joint velocity to end-effector velocity, and is where singularities and redundancy come from."},{"title":"Force, statics, and inertia","blurb":"Geometry so far; now force: torque, wrenches, static equilibrium, and center of mass and inertia."},{"title":"Dynamics and vibration","blurb":"Connecting force and motion: writing the equations of motion, solving them efficiently with recursive algorithms, then spring-damper vibration."},{"title":"Contact, friction, and grasping","blurb":"From the robot itself to the outside world: contact forces, friction, and collisions, then using them to analyze a stable grasp."},{"title":"Legged balance and simplified models","blurb":"Applying contact forces to the feet: how legged robots judge whether they’re stable, and the simplified models used to plan balance."},{"title":"Gait and walking","blurb":"From standing to walking and running: gait phases and types, running and hopping models, and more natural, efficient ways to walk."}]},{"key":"control","name":"Control & Planning","blurb":"Making joints move the way you want: from PID to force control, MPC, motion planning, and whole-body control.","sections":[{"title":"Control basics: layers and feedback","blurb":"First, where control sits in the system and how often it runs, then feedback control like PID that reacts to error."},{"title":"Joint control and model compensation","blurb":"Applying feedback to joint motors: position, velocity, and torque modes and drive interfaces, then compensating with a dynamics model."},{"title":"End-effector and force control","blurb":"Switching from controlling each joint to controlling the end effector directly, then force and compliance control on contact, plus visual servoing."},{"title":"Trajectory generation and tracking","blurb":"Where the controller’s target comes from: generating smooth trajectories from waypoints, interpolation, and velocity profiles, then tracking them accurately."},{"title":"Optimal control and MPC","blurb":"Instead of interpolation, framing control as optimization: LQR, trajectory optimization, MPC, and Kalman-filter estimation."},{"title":"Legged and whole-body control","blurb":"Using these tools to keep legged and humanoid robots stable: gait, foot placement, balance, whole-body control, and RL-based locomotion."},{"title":"Path planning and navigation","blurb":"From control to planning: finding routes on a map with algorithms like A*, then global and local layers for avoiding obstacles."},{"title":"Motion planning for arms","blurb":"Too many joints to grid-search: arm motion is planned instead with collision checking, sampling methods like RRT, and trajectory optimization."},{"title":"Task level: sequencing and planning","blurb":"One level up, deciding what to do in what order: state machines, behavior trees, symbolic planning, and LLM-based task planning."},{"title":"Safety and stability guarantees","blurb":"The safety net running through every layer: e-stops, limits, collision detection, safety standards, and theory for provable stability."}]},{"key":"perception","name":"Perception & Sensors","blurb":"How robots see and feel: cameras, depth, IMUs, force and touch, plus point clouds, calibration, and SLAM.","sections":[{"title":"Perception overview and cameras","blurb":"Perception splits into internal and external sensing; start with the everyday RGB camera — placement, how it images, and its specs."},{"title":"Depth cameras and lidar","blurb":"Adding distance on top of ordinary imaging: stereo, structured-light, and ToF depth cameras, plus lidar."},{"title":"Proprioception and force sensing","blurb":"From sensing the outside world to sensing itself: encoders and IMUs measure motion, force/torque sensors measure force and contact."},{"title":"Tactile sensing","blurb":"A finer-grained sense of touch than force sensors: tactile arrays, e-skin, vision-based tactile sensors, and what tactile data is used for."},{"title":"Calibration and spatiotemporal alignment","blurb":"Now that you know the sensors, align them: camera and hand-eye calibration, multi-sensor extrinsics, and time synchronization."},{"title":"Detection, segmentation, and tracking","blurb":"Entering vision algorithms: boxing, cutting out, and tracking objects in images, including targets specified by text."},{"title":"Recovering 3D from images","blurb":"Without a depth sensor: estimating depth from ordinary images, multi-view geometry, and networks that output 3D in one step."},{"title":"Point clouds and 3D representations","blurb":"Once you have 3D data: downsampling, registering, and running networks on point clouds, then meshes, NeRF, and Gaussian splatting."},{"title":"Object pose, grasping, and affordance","blurb":"Putting 3D perception to work in manipulation: estimating 6D object pose, detecting grasps, and finding affordances and articulated structure."},{"title":"Human body, hands, and interaction","blurb":"Shifting from objects to people: body and hand pose, 3D meshes, plus gesture, gaze, and speech."},{"title":"State estimation and SLAM","blurb":"Answering ‘where am I’: from odometry and multi-sensor fusion to SLAM and localizing within a known map."},{"title":"Maps, semantics, and spatial intelligence","blurb":"Beyond localization, understanding the whole scene: geometric and semantic maps, scene graphs, and LLM-based spatial reasoning and active perception."}]},{"key":"software","name":"Software & Tooling","blurb":"The software that ties body, control, and perception together: ROS 2, URDF, motion libraries, learning frameworks, and deployment tools.","sections":[{"title":"Development environment basics","blurb":"First set up your machine: programming languages, Linux, code repositories, environment isolation, GPUs, and remote access."},{"title":"Getting started with ROS 2","blurb":"With your environment ready, learn robotics software’s ‘glue’: nodes, topics, and other communication patterns, plus how projects are organized."},{"title":"Robot models and coordinate frames","blurb":"Once you know ROS, describe a robot with files like URDF, then visualize it with TF and RViz."},{"title":"Communication middleware, in depth","blurb":"Digging back into the communication layer: DDS, QoS, zero-copy, and communication libraries beyond ROS."},{"title":"Connecting real hardware, real time","blurb":"With communication working, connect real hardware: vendor SDKs, drivers, embedded and real-time systems that get commands to the motors."},{"title":"Kinematics, planning, and control libraries","blurb":"Once hardware can send and receive commands, use existing libraries for kinematics, collision-free planning, and optimal control."},{"title":"Perception, mapping, and navigation","blurb":"Helping the robot understand its environment: image and point-cloud libraries, calibration, SLAM mapping, and autonomous navigation."},{"title":"Data logging and visualization","blurb":"Recording and visualizing sensor and robot state: useful for debugging, and for collecting data to train on later."},{"title":"Deep learning and robot-learning libraries","blurb":"With data in hand, start learning: deep learning frameworks, model hubs, then LeRobot and reinforcement learning libraries."},{"title":"Training engineering and compute","blurb":"Getting training running, then scaling it up: renting GPUs, tracking experiments, saving memory, multi-GPU clusters, and large-scale RL."},{"title":"Deployment and inference speedups","blurb":"Moving a trained policy onto the robot: running it remotely or on-device, sped up with kernel optimization, export, and inference engines."},{"title":"Industry platforms and ecosystems","blurb":"Finally, how vendors package all these layers together: NVIDIA Isaac, Chinese robot operating systems, and cloud platforms."}]},{"key":"sim","name":"Simulation & Evaluation","blurb":"Practicing and testing on a computer: how simulators work, the sim-to-real gap, and the benchmarks used to evaluate.","sections":[{"title":"Simulation basics","blurb":"The first step to training robots virtually: what a simulator is, how it steps through interaction, and how to make it fast and accurate."},{"title":"Physics engines: rigid bodies, contact","blurb":"Opening up the physics engine to see how each step computes rigid-body motion, collisions, and contact forces."},{"title":"Soft bodies, fluids, differentiable sim","blurb":"Beyond rigid bodies to deformable ones — cloth, rope, liquid — and simulators that can compute gradients."},{"title":"Rendering and sensor simulation","blurb":"The other half beyond physics: rendering scenes into camera images, and simulating readings from sensors like touch."},{"title":"Major simulators and frameworks","blurb":"Physics and rendering in practice: the MuJoCo and Isaac families, plus other commonly used simulators."},{"title":"Scenes, assets, and indoor platforms","blurb":"With a simulator ready, fill it with content: object assets, ready-made indoor scenes, and automated scene generation at scale."},{"title":"From simulation to reality","blurb":"Moving what you trained in sim onto a real robot: understand the sim-to-real gap, then domain randomization and pulling reality into sim."},{"title":"Evaluation methods and metrics","blurb":"Training done, now the exam: how success rate is calculated, how to test in sim and reality, and what makes results comparable."},{"title":"Tabletop manipulation benchmarks","blurb":"Applying that method to real test sets, starting with the most common tabletop arm-manipulation benchmarks."},{"title":"Household, navigation, and QA benchmarks","blurb":"Expanding from the tabletop to a whole house: long-horizon chores, language navigation, embodied QA, and LLMs as the brain."},{"title":"RL and locomotion-control benchmarks","blurb":"A different kind of task: classic reinforcement-learning testbeds, plus environments for training legged and humanoid robots to walk."},{"title":"Real-world and world-model evaluation","blurb":"Beyond fixed sim test sets: large-scale real-robot evaluation, using world models to evaluate policies, and evaluating world models themselves."}]},{"key":"data","name":"Data & Collection","blurb":"The raw material for learning: how demonstration data is collected, what datasets exist, and how data is processed and scaled.","sections":[{"title":"Data’s basic unit and sources","blurb":"First, see that a demonstration is made of observations and actions, then the main sources: real robots, simulation, and human video."},{"title":"Real-robot collection: teleop, teaching","blurb":"Starting with the most reliable source, real-robot data: a person pushes, guides, or remotely operates the robot while it’s recorded."},{"title":"No robot needed: handheld and wearable","blurb":"Teleoperation needs a real robot and is slow and costly, so instead a person uses a gripper-like tool or wearable device directly."},{"title":"Motion capture and humanoid motion data","blurb":"From hands to the whole body: motion capture records a person’s full movement, then converts it into trajectories a humanoid can follow."},{"title":"Human video and first-person data","blurb":"Stepping back further to just filming people: video is cheap and abundant, but lacks action labels, and human hands aren’t robot hands."},{"title":"Simulated assets and synthetic data","blurb":"No longer collecting one demo at a time: prepare object and scene assets, then generate demonstrations in bulk in sim or with generative models."},{"title":"Major real-robot datasets","blurb":"Back to real robots and their public datasets: starting with OXE, which pools many robots, then key releases in order."},{"title":"Data formats and tools","blurb":"What file formats those datasets use and what libraries read them: from HDF5 and RLDS to LeRobot."},{"title":"Cleaning, labeling, and mixing","blurb":"Raw data still needs processing: cleaning and quality checks, adding labels, then selecting and mixing ratios before training."},{"title":"Scaling: from data farms to flywheels","blurb":"Finally, how data volume is scaled up: data-collection farms, crowdsourcing, robots collecting their own data, and deployment feeding a flywheel."}]},{"key":"training","name":"Training & Learning Methods","blurb":"How models are actually trained: imitation learning, reinforcement learning, pretraining and fine-tuning, plus tricks that make it more stable.","sections":[{"title":"Overview of learning paradigms","blurb":"With data ready, meet the basic ways to learn: with labels, without them, self-generated labels, from demonstration, or from reward."},{"title":"Training fundamentals","blurb":"Whatever the method, training means computing loss, taking gradients, and updating parameters, while guarding against overfitting and instability."},{"title":"Loss functions and training objectives","blurb":"Expanding on loss functions: the different objectives used for regression, classification, sequence prediction, and diffusion generation."},{"title":"Imitation learning","blurb":"Applying supervised training to learn actions from human demonstrations, and why it tends to drift further off course over time."},{"title":"Core reinforcement-learning concepts","blurb":"The second path besides demonstration is trial and error by reward: first, the basics of reward, return, and value."},{"title":"Classic reinforcement-learning algorithms","blurb":"Turning those concepts into concrete algorithms: from policy gradients and PPO to DQN and SAC."},{"title":"Where rewards come from","blurb":"With algorithms in hand, you still need the right reward: hand-written, learned from demonstrations or a model, and what to do when it’s sparse."},{"title":"Offline reinforcement learning","blurb":"Earlier algorithms all learn through live interaction; here a policy learns from a fixed dataset instead, then fine-tunes online."},{"title":"Pretraining and representation learning","blurb":"Shifting from reinforcement learning to the large-model approach: pretraining on massive, varied data to learn general-purpose representations."},{"title":"Fine-tuning and post-training","blurb":"After pretraining, adapting to specific tasks: fine-tuning, supervised and RL post-training, plus preventing forgetting and distillation."},{"title":"From simulation to real robots","blurb":"Onto real robots: large-scale training in sim, transferring to hardware, then continuing to learn there through trial and human correction."},{"title":"Inference-time gains and fast adaptation","blurb":"After training ends, a model can still improve without changing weights much: more compute at inference, or fast adaptation to new tasks."}]},{"key":"model","name":"Models & Architectures","blurb":"What the resulting models look like: Transformers, VLMs, how actions are generated, VLAs, and world models.","sections":[{"title":"Neural network basics","blurb":"Starting with the basic building blocks and classic architectures of neural networks, which every later model is built from."},{"title":"From attention to language models","blurb":"The shared backbone of large models: tokens, attention, and the Transformer, leading up to large language models."},{"title":"Vision encoders and vision-language models","blurb":"Giving a language model eyes: first how images become tokens, then VLMs, the foundation VLAs are built on."},{"title":"Introduction to generative models","blurb":"Earlier models output a single answer; here, models learn a whole data distribution, plus turning vectors into discrete codes."},{"title":"Diffusion models and flow matching","blurb":"The leading approach among generative models: adding and removing noise, and flow matching, which power both action and video generation."},{"title":"How actions are represented, output","blurb":"With networks and generative methods in hand, see how robot actions are represented, and output by regression, discretization, or generation."},{"title":"VLA structure and extensions","blurb":"Attach action output to a VLM and you get a VLA: how it outputs actions, generalizes across robots, and takes in more sensors."},{"title":"Hierarchical systems and embodied reasoning","blurb":"The opposite of end-to-end: splitting planning and control across different models, then having a model reason before it acts."},{"title":"World models and video generation","blurb":"Using generative models to predict ‘what happens to the world after an action,’ finally merging with action generation itself."},{"title":"Inference speed, deployment, reliability","blurb":"Every model here eventually runs on hardware: how to make it faster and cheaper, and how to tell when it’s unsure."}]},{"key":"named_model","name":"Landmark Models & Projects","blurb":"Famous models in historical order: from SayCan and RT-2 to the π series, GR00T, and world models.","sections":[{"title":"Foundation models embodied AI borrows","blurb":"An embodied model’s ‘eyes’ and ‘brain’ are often borrowed: meet these foundation models and vision encoders first."},{"title":"LLMs as the brain","blurb":"The earliest use of large models was as the brain: breaking down tasks, writing code and rewards, understanding space, then calling existing skills."},{"title":"End-to-end learning and the RT series","blurb":"The other path is learning control end to end: from real-robot grasping to generalist models, up to RT-2, which coined VLA."},{"title":"Classic imitation-learning policies","blurb":"Smaller imitation-learning models were evolving in parallel: from PerAct to Diffusion Policy and ACT, mastering fine bimanual work."},{"title":"Open-source generalist policies and VLA","blurb":"RT and imitation learning converge into open-source VLA: after Octo and OpenVLA, improvements bloomed in every direction."},{"title":"The π series and reinforcement learning","blurb":"Physical Intelligence’s π series set the VLA benchmark; then how reinforcement learning makes policies stronger from experience."},{"title":"Global tech giants and star startups","blurb":"From technical approaches to companies: the foundation models of NVIDIA, Figure, Google, and other firms outside China."},{"title":"Embodied models in China","blurb":"Now China: foundation models from big tech, research institutes, and startups, mostly split between VLA and world-action-model approaches."},{"title":"World models and learning from video","blurb":"Back to the world-model thread: training policies inside imagination first, then generating worlds and learning actions from video."},{"title":"Legged, dexterous-hand, and agile skills","blurb":"From manipulation to movement: quadrupeds doing parkour, dexterous hands spinning objects, trained with RL in sim, then deployed to hardware."},{"title":"Humanoid whole-body control and teleop","blurb":"The same methods carried over to humanoids: from animated-character imitation and bipedal walking to whole-body teleop and motion tracking."},{"title":"Navigation and autonomous driving","blurb":"Finally, moving through the world: navigation from modular pipelines to foundation models, and driving from ALVINN to VLA."}]},{"key":"company","name":"Companies & Institutions","blurb":"Who’s actually working on embodied AI: tech giants, hardware and model companies worldwide, component makers, and research institutions.","sections":[{"title":"Tech giants’ embodied AI bets","blurb":"From models to companies: giants like Google and NVIDIA supply compute, platforms, and models, and also build robots and fund startups."},{"title":"Foundation-model startups outside China","blurb":"Beyond big tech, a wave of startups outside China focus on the robot brain, led by Physical Intelligence and Skild AI."},{"title":"Humanoid and hardware firms abroad","blurb":"A brain still needs a body: Figure, Boston Dynamics, 1X, and other firms outside China build humanoids and other hardware."},{"title":"Humanoid and hardware firms in China","blurb":"Back to China, where hardware makers are most numerous, led by Unitree, AgiBot, and UBTech, spanning humanoids, quadrupeds, and wheeled dual-arm robots."},{"title":"Foundation-model startups in China","blurb":"A domestic wave focused on the brain too — led by X Square Robot, Spirit AI, and AI² Robotics — many rooted in academia or autonomous driving."},{"title":"Crossovers: cars and consumer electronics","blurb":"Beyond native embodied-AI companies, carmakers and consumer electronics firms are crossing over too, led by Xiaomi and Xpeng."},{"title":"Arms, service, and logistics robots","blurb":"Robots that were selling before the humanoid boom: common research arms, the ‘big four’ industrial makers, and service and logistics robots."},{"title":"World models, simulation, and data","blurb":"Moving upstream from building robots: these companies focus on world models, simulation platforms, and training data to feed the brain."},{"title":"Core components and dexterous hands","blurb":"After upstream software, upstream hardware: dexterous hands, grippers, gearboxes, lead screws, motors, joints, and controllers."},{"title":"Sensors, motion capture, and chips","blurb":"Where components handle moving, this group handles seeing, touching, and computing: cameras, lidar, tactile and force sensing, mocap gear, and chips."},{"title":"Research institutions and university labs","blurb":"Outside the supply chain, where papers and talent come from: university labs, corporate research institutes, and China’s new research organizations."},{"title":"Historic firms and industry bodies","blurb":"Finally, the backdrop: notable firms that paved the way or have since exited, plus organizations behind contests, conferences, and reports."}]},{"key":"industry","name":"Industry Jargon & Business","blurb":"Making sense of press releases, pitch decks, and group-chat slang: embodiment, ‘big/small brain,’ data flywheels, mass production, and more.","sections":[{"title":"Industry map: who does what","blurb":"With the tech and companies covered, map the industry: how the supply chain breaks down, and what each player does."},{"title":"The debate over approaches","blurb":"Now see what the players disagree on: humanoid or not, real robots or sim, integrated or decoupled, and the path to generality."},{"title":"Reading demos and leaderboards","blurb":"Judging an approach means judging its results: how to read demos and leaderboards, and spot teleop, editing, or cherry-picking."},{"title":"Papers, conferences, and journals","blurb":"More rigorous evidence than a demo lives in papers: preprints, top-tier AI conferences, and the main robotics conferences and journals."},{"title":"Expos, competitions, and viral moments","blurb":"Beyond academic conferences, the public stage: expos, launch events, TV galas, and robot competitions, seen with the same critical eye."},{"title":"Mass production and supply chains","blurb":"Behind the show, robots first have to be built: what counts as mass production, why it’s hard, cost, and supply chains."},{"title":"Deployment and factory automation","blurb":"Once built, robots still have to work in the field: the stages of deployment, how factories measure success, and why they want robots."},{"title":"Business models and markets","blurb":"Once it can work, it has to make money: who buys it, how it’s priced, and whether the business closes the loop."},{"title":"Hype cycles, funding, and IPOs","blurb":"The business story is ultimately pitched to capital: first spotting hype and bubbles, then funding rounds, IPOs, and related stocks."},{"title":"Policy and standards","blurb":"Beyond capital, government plays a role too: national direction, policy documents, open challenges, and finally, standards and new job titles."}]}],"terms":[{"id":"embodied-ai","category":"concept","sec":0,"tier":1,"sources":[{"title":"维基百科：具身智能","url":"https://zh.wikipedia.org/wiki/具身智能"},{"title":"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI","url":"https://arxiv.org/abs/2407.06886"}],"as_of":"","related_ids":["disembodied-ai","physical-ai","embodied-agent","embodied-cognition","general-purpose-robot","vision-language-action-model"],"name":"Embodied AI","alt":"具身智能","abbr":"","aliases":["Embodied Intelligence","EAI","Embodied Artificial Intelligence"],"one_liner":"Artificial intelligence with a body, able to perceive, decide, and act in the physical world.","explanation":"Embodied AI refers to intelligent systems that have a physical body (or a simulated one) and perceive, decide, and act through real-time interaction with their environment — robots and self-driving cars are the typical examples. The contrast is disembodied AI, like ChatGPT, which only processes text and images and never directly acts on the physical world. The idea is often traced back to Alan Turing's 1950 paper “Computing Machinery and Intelligence,” and to the robotics research of Rodney Brooks and others in the 1980s–90s, which emphasized that a body interacting with its environment is central to intelligence. At its core is a closed loop of perception, decision, action, and feedback. In the past few years, large language models and vision-language models have let robots understand open-ended instructions, making embodied AI a hot field, often seen as a key step toward artificial general intelligence (AGI).","example":"A home robot told to put the dirty clothes in the washing machine has to find the clothes itself, walk over, pick them up, open the machine, and load them, adjusting its actions throughout using visual feedback — that is embodied AI. ChatGPT can only describe the steps; that is disembodied AI.","related":["Disembodied AI","Physical AI","Embodied Agent","Embodied Cognition","General-purpose Robot","Vision-Language-Action Model"]},{"id":"disembodied-ai","category":"concept","sec":0,"tier":2,"sources":[{"title":"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI (arXiv 2407.06886)","url":"https://arxiv.org/abs/2407.06886"}],"as_of":"","related_ids":["embodied-ai","embodied-cognition","large-language-model","moravec-s-paradox","language-grounding"],"name":"Disembodied AI","alt":"离身智能","abbr":"","aliases":["Non-embodied AI"],"one_liner":"AI with no body, processing information purely in the digital world — chatbots, image classifiers, and so on.","explanation":"Disembodied AI is the counterpart to embodied AI: AI with no physical body that only perceives and decides within cyberspace. Its inputs are data such as text, images, and video, and its outputs are text, labels, or images — it never has to act to change the physical environment itself. Large language models like ChatGPT, facial recognition, and recommendation systems all fall into this category. A 2024 embodied-AI survey by Yang Liu and colleagues at Sun Yat-sen University uses a table to contrast the two: in disembodied AI, cognition is separate from any physical entity, while in embodied AI, cognition is fused into an entity such as a robot or a car. The distinction matters because disembodied models learn very strong knowledge from internet data but lack feedback from interacting with the real world; putting one into a robot still leaves problems to solve, such as language grounding, mapping words to real objects and actions, along with producing actual motor output and running in real time.","example":"ChatGPT can write out detailed steps for folding clothes, but it has no hands and cannot actually fold them; getting a robot to carry out those steps is exactly what embodied AI has to solve.","related":["Embodied AI","Embodied Cognition","Large Language Model","Moravec's Paradox","Language Grounding"]},{"id":"physical-ai","category":"concept","sec":0,"tier":1,"sources":[{"title":"What is Physical AI? (NVIDIA Glossary)","url":"https://www.nvidia.com/en-us/glossary/generative-physical-ai/"},{"title":"NVIDIA Launches Cosmos World Foundation Model Platform to Accelerate Physical AI Development (NVIDIA Newsroom, 2025-01-06)","url":"https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-world-foundation-model-platform-to-accelerate-physical-ai-development"}],"as_of":"2026-09","related_ids":["embodied-ai","world-foundation-model","nvidia-cosmos","nvidia-three-computer-solution","sim-to-real-transfer","chatgpt-moment-for-robotics"],"name":"Physical AI","alt":"物理AI","abbr":"","aliases":["Generative Physical AI"],"one_liner":"AI that can perceive, understand, and act in the physical world — robots and self-driving cars are examples.","explanation":"Physical AI is a term heavily promoted by NVIDIA, officially defined as systems that can perceive, understand, and reason within the physical world, then execute or coordinate complex actions. It covers robots, self-driving cars, and “smart spaces” that use fixed cameras to optimize factories and warehouses. The problems it addresses largely overlap with what academics usually call “embodied AI,” but its scope is wider — it even counts bodiless “smart spaces” that rely only on fixed cameras, with no robot body at all. NVIDIA lays out a three-step development pipeline: train models in a data center; simulate and generate synthetic data on the Omniverse simulation platform and with the Cosmos world foundation model; then deploy to edge computing platforms such as Jetson and DRIVE. When Cosmos was announced at CES in January 2025, Jensen Huang said “the ChatGPT moment for robotics is coming.”","example":"A self-driving car processing sensor data in real time to decide steering and braking, and an autonomous mobile robot in a warehouse dodging obstacles while moving goods, both count as what NVIDIA calls physical AI.","related":["Embodied AI","World Foundation Model","NVIDIA Cosmos","NVIDIA Three-Computer Solution","Sim-to-Real Transfer","ChatGPT Moment for Robotics"]},{"id":"autonomous-driving","category":"concept","sec":0,"tier":2,"sources":[{"title":"Self-driving car - Wikipedia（含 SAE J3016 分级）","url":"https://en.wikipedia.org/wiki/Self-driving_car"}],"as_of":"","related_ids":["end-to-end","world-model","lidar","levels-of-autonomy","autonomous-driving-talent-moving-into-embodied-ai","embodied-ai"],"name":"Autonomous Driving","alt":"自动驾驶","abbr":"","aliases":["Self-driving"],"one_liner":"Letting a car perceive road conditions, make decisions, and control itself with little or no human input.","explanation":"Autonomous driving means a vehicle uses sensors such as cameras, lidar, and millimeter-wave radar, together with an onboard computer, to handle perception, decision-making, planning, and control on its own. SAE J3016, the standard from the Society of Automotive Engineers, defines six levels, L0 through L5. At L2, the system handles both steering and speed, but the driver must watch the road at all times. From L3 up, the system takes over driving under specific conditions and asks the driver to step in when needed. L4 allows fully driverless operation within a defined area, and L5 needs no human under any condition. In China, “智驾” (“smart driving”), the everyday term, usually refers to L2-level driver assistance. Autonomous driving can be seen as the first form of embodied AI to reach large-scale deployment, and it shares techniques with robotics more broadly, including end-to-end models, world models, simulation, and data flywheels — the industry often talks about “crossing over from autonomous driving into embodied AI.”","example":"Urban NOA (navigate-on-autopilot) can automatically follow traffic, change lanes, and go through intersections in a city, but the driver must stay ready to take over at any time — this is an L2-level system.","related":["End-to-End","World Model","LiDAR","Levels of Autonomy","Autonomous-Driving Talent Moving into Embodied AI","Embodied AI"]},{"id":"agentenvironment-interaction","category":"concept","sec":0,"tier":1,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Gymnasium Documentation: Basic Usage","url":"https://gymnasium.farama.org/introduction/basic_usage/"}],"as_of":"","related_ids":["reinforcement-learning","markov-decision-process","observation","action-space","reward-function","environment"],"name":"Agent–Environment Interaction","alt":"智能体与环境","abbr":"","aliases":["Agent-Environment Loop","Agent-Environment Interface"],"one_liner":"The basic RL framework: the decision-maker is the agent, and everything outside it that it can observe and affect is the environment.","explanation":"This is the basic framework reinforcement learning uses to describe a problem. The agent is the decision-maker; the environment is everything outside it that the agent can observe and act on. The two interact in a loop: the agent sees the current observation and picks an action; the environment transitions to a new state accordingly and returns a new observation and a reward (a number measuring how good the outcome was); this repeats until the episode ends. In embodied AI, the agent is usually a policy model running on a robot, and the environment is the real world or a simulator, including the tabletop, objects, and any nearby people. This shared framework lets very different problems — navigation, grasping, walking — be described with the same vocabulary; simulation libraries such as Gymnasium build their APIs around it too, using reset to start a new episode and step to execute one action.","example":"A robot arm folding a towel: the agent is the folding policy, the environment is the tabletop, the towel, and the camera feed. At each step the policy outputs a set of joint actions, and the environment returns the new image that results.","related":["Reinforcement Learning","Markov Decision Process","Observation","Action Space","Reward Function","Environment (Env; reset/step interface)"]},{"id":"embodied-agent","category":"concept","sec":0,"tier":2,"sources":[{"title":"Embodied agent - Wikipedia","url":"https://en.wikipedia.org/wiki/Embodied_agent"},{"title":"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI (arXiv 2407.06886)","url":"https://arxiv.org/abs/2407.06886"}],"as_of":"","related_ids":["embodied-ai","agentenvironment-interaction","perception-action-loop","disembodied-ai","policy"],"name":"Embodied Agent","alt":"具身智能体","abbr":"","aliases":[],"one_liner":"An agent with a physical or virtual body that perceives and acts within an environment through that body.","explanation":"An agent is a system that can perceive its environment and take action to reach a goal. An embodied agent emphasizes that it has a body, and perceives and acts in the environment through that body. Wikipedia defines it as an agent that interacts with its environment through a physical body situated within that environment; more broadly, this also includes virtual bodies in a simulator or a game. It differs from a purely conversational agent that only outputs text, since it needs to chain together instruction understanding, active exploration, multimodal perception, and executed action into a working loop. The 2024 survey by Yang Liu and colleagues lists embodied agents as one of embodied AI's four main research goals, and argues that multimodal large models are currently its dominant “brain.” Real robots, self-driving cars, and navigation agents in simulators such as Habitat all count as embodied agents.","example":"In the Habitat home simulator, a virtual robot with a camera is told to go find a cup in the kitchen; it moves and looks around on its own and eventually stops in front of the cup — that makes it an embodied agent.","related":["Embodied AI","Agent–Environment Interaction","Perception-Action Loop","Disembodied AI","Policy"]},{"id":"embodiment","category":"concept","sec":0,"tier":1,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models（项目页）","url":"https://robotics-transformer-x.github.io/"},{"title":"Open X-Embodiment 论文（arXiv 2310.08864）","url":"https://arxiv.org/abs/2310.08864"}],"as_of":"","related_ids":["cross-embodiment","embodiment-gap","embodiment-agnostic","degrees-of-freedom","humanoid-robot","robot-body-maker"],"name":"Embodiment","alt":"本体","abbr":"","aliases":["Robot Embodiment","Robot Body"],"one_liner":"A robot's physical body: its specific combination of form, joints, sensors, and actuators.","explanation":"“Embodiment” (本体) is the term China's embodied-AI industry uses for a robot's physical “body”: its specific hardware form — single-arm, dual-arm, quadruped, or humanoid — how many degrees of freedom it has (roughly, how many joints move independently), which cameras and force sensors it carries, and whether its end effector is a gripper or a dexterous hand. Embodiment determines what a policy can see (its observation space) and what it can output (its action space); the same model usually can't be dropped onto a different embodiment without changes, a mismatch called the embodiment gap. The Open X-Embodiment dataset pools data from 22 embodiments specifically to study cross-embodiment transfer. In industry usage, an “embodiment vendor” (本体厂商) is a company that builds robot hardware, as distinct from a “brain company” that builds only the software models.","example":"A Franka arm and a Unitree G1 humanoid are two very different embodiments: the former is 7 joints plus a two-finger gripper, while the latter needs to coordinate more than twenty joints across its legs, waist, and both arms at once.","related":["Cross-Embodiment","Embodiment Gap","Embodiment-agnostic","Degrees of Freedom (DoF)","Humanoid Robot","Robot Body Maker"]},{"id":"perception-action-loop","category":"concept","sec":0,"tier":2,"sources":[{"title":"Embodied cognition - Wikipedia","url":"https://en.wikipedia.org/wiki/Embodied_cognition"},{"title":"Sensory-motor coupling - Wikipedia","url":"https://en.wikipedia.org/wiki/Sensory-motor_coupling"}],"as_of":"","related_ids":["embodied-ai","closed-loop-control","open-loop-control","embodied-cognition","sense-plan-act","active-perception"],"name":"Perception-Action Loop","alt":"感知-行动闭环","abbr":"","aliases":[],"one_liner":"Perception drives action, and action changes what is perceived next, in a continuous back-and-forth cycle.","explanation":"The perception-action loop describes the ongoing back-and-forth between an agent and its environment: sensors produce an observation, the policy produces an action, the action changes the environment and the agent's own position, and so what is seen the next moment changes too. The idea comes from cognitive science and ecological psychology; James Gibson, Francisco Varela, and others all emphasized that perception is not passive reception but is coupled to bodily movement, and roboticist Rodney Brooks likewise argued that intelligence must connect to the world through a body. This loop is a key distinction between embodied AI and disembodied AI: an image classifier looks at one picture, gives one answer, and is done, while a robot policy must repeatedly observe, act, and observe again, often tens of times per second, to correct errors and respond to change.","example":"As a robot arm grasps a cup, its wrist camera sees the relative position of the cup and gripper in every frame, and the policy corrects its next move accordingly — if the cup gets bumped out of place, it can realign and follow it.","related":["Embodied AI","Closed-loop Control","Open-loop Control","Embodied Cognition","Sense-Plan-Act","Active Perception"]},{"id":"robot-learning","category":"concept","sec":0,"tier":1,"sources":[{"title":"Robot learning - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_learning"}],"as_of":"","related_ids":["imitation-learning","reinforcement-learning","policy","sim-to-real-transfer","foundation-model","embodied-ai"],"name":"Robot Learning","alt":"机器人学习","abbr":"","aliases":[],"one_liner":"Using machine learning to let robots acquire skills from data and interaction, instead of hand-written rules.","explanation":"Robot learning is the field at the intersection of machine learning and robotics: it studies how robots can acquire new skills or adapt to their environment through learning algorithms, rather than having engineers hand-code a control program line by line. The main approaches are imitation learning, learning from human demonstrations such as data collected via teleoperation; reinforcement learning, trial and error guided by reward in simulation or on real hardware; and, more recently, robot foundation models pretrained on large-scale data. The core challenges are that real-robot data is scarce and expensive, real-robot trial and error is risky, and there is a gap between simulation and reality. The field's dedicated venue is CoRL, the Conference on Robot Learning, and most model research in embodied AI falls under this umbrella.","example":"Collecting a few dozen towel-folding demonstrations via teleoperation and training an ACT policy on them lets a bimanual robot learn to fold towels on its own, without an engineer writing out the joint angles for every step.","related":["Imitation Learning","Reinforcement Learning","Policy","Sim-to-Real Transfer","Foundation Model","Embodied AI"]},{"id":"general-purpose-robot","category":"concept","sec":0,"tier":1,"sources":[{"title":"Figure AI: Master Plan","url":"https://www.figure.ai/master-plan"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"2024-10","related_ids":["embodied-ai","generalist-policy","humanoid-robot","industrial-robot","embodied-foundation-model","pi0"],"name":"General-purpose Robot","alt":"通用机器人","abbr":"","aliases":["Generalist Robot"],"one_liner":"A robot not built for one fixed task, meant to handle many tasks across many settings.","explanation":"A general-purpose robot is defined in contrast to a special-purpose one. A traditional industrial robot is programmed for one production line and one operation, and needs reprogramming and new tooling whenever the task changes. A general-purpose robot aims to be more like a person: one body plus one intelligent system that can do many things, including handling objects and environments it has never seen before. On the hardware side, this usually means a humanoid or a dual-arm mobile platform, since doors, tools, and stairs are all designed for the human body — companies such as Figure cite this as their reason for building humanoids. On the software side, it relies on generalist policies or embodied foundation models that learn cross-task ability from large-scale data; Physical Intelligence's π0, for example, is trained on data from single-arm, dual-arm, and mobile-manipulation robots, and can fold laundry, clear a table, and assemble a box. Every product today still falls short of this goal — it remains a long-term direction for embodied AI.","example":"In its company plan (Master Plan), dated May 2022, Figure set out to build a “general-purpose humanoid robot”: one humanoid hardware platform taking on a large share of the work people currently do, instead of building a separate special-purpose robot for every task.","related":["Embodied AI","Generalist Policy","Humanoid Robot","Industrial Robot","Embodied Foundation Model","π0"]},{"id":"moravec-s-paradox","category":"concept","sec":0,"tier":1,"sources":[{"title":"Moravec's paradox (Wikipedia)","url":"https://en.wikipedia.org/wiki/Moravec%27s_paradox"}],"as_of":"","related_ids":["embodied-ai","sensorimotor-skills","embodied-cognition","the-bitter-lesson","physical-turing-test"],"name":"Moravec's Paradox","alt":"莫拉维克悖论","abbr":"","aliases":[],"one_liner":"For machines, high-level reasoning is easy; the perception and movement humans take for granted turn out to be the hard part.","explanation":"In his 1988 book Mind Children, roboticist Hans Moravec observed that it's comparatively easy to get a computer up to adult level on an intelligence test or at checkers, but hard, or even impossible, to give it the perception and mobility of a one-year-old child. Rodney Brooks, Marvin Minsky, and others voiced similar views in the 1980s; Steven Pinker later summed it up as “the hard problems are easy, and the easy problems are hard.” The usual explanation: perception and motor skills were refined over a vast span of evolution, so they feel effortless to us and we underrate how hard they really are, while abstract reasoning emerged only recently in evolutionary terms — the human brain isn't especially well-suited to it, which is why it feels hard to us, even though it isn't actually that hard for a computer. This paradox is often used to explain why large models can write code and solve math problems while robots still can't fold laundry as fast or as reliably as a person, and it's one reason embodied AI is seen as AI's next major hurdle.","example":"A large language model can already solve competition-level math problems, but folding laundry or tying a shoelace — things a small child can do — a robot still does far slower and far less reliably than a person.","related":["Embodied AI","Sensorimotor Skills","Embodied Cognition","The Bitter Lesson","Physical Turing Test"]},{"id":"sensorimotor-skills","category":"concept","sec":0,"tier":3,"sources":[{"title":"Moravec's paradox - Wikipedia","url":"https://en.wikipedia.org/wiki/Moravec%27s_paradox"},{"title":"End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)","url":"https://arxiv.org/abs/1504.00702"},{"title":"Sensory-motor coupling - Wikipedia","url":"https://en.wikipedia.org/wiki/Sensory-motor_coupling"}],"as_of":"","related_ids":["moravec-s-paradox","perception-action-loop","visuomotor-policy","embodied-cognition","end-to-end-training-of-deep-visuomotor-policies"],"name":"Sensorimotor Skills","alt":"感觉运动技能","abbr":"","aliases":["Sensorimotor Skill Learning"],"one_liner":"The ability to turn sensory information into body movement in real time: grasping, walking, catching, twisting.","explanation":"Sensorimotor skills is a term from cognitive science and neuroscience for abilities that tightly couple sensory input, such as vision, touch, and proprioception, with muscle movement to complete a task: reaching for an object, walking while keeping balance, catching a thrown ball. In robotics it broadly refers to low-level capabilities that need perception and control to run in a real-time closed loop. Moravec's paradox observes that the sensorimotor skills people find effortless are actually the hardest for machines, precisely because they were refined over a long evolutionary history. Levine and colleagues' 2015 end-to-end visuomotor policy work, which used a convolutional network to map images directly to motor torques, is a landmark example of learning such skills with deep learning; today's visuomotor policies and VLA models are, at bottom, also learning sensorimotor skills.","example":"Unscrewing a bottle cap: a robot must watch the position of the cap while continuously adjusting the angle and force it applies based on what it feels in its hand — this was one of the test tasks in Levine and colleagues' end-to-end visuomotor policy work.","related":["Moravec's Paradox","Perception-Action Loop","Visuomotor Policy","Embodied Cognition","End-to-End Training of Deep Visuomotor Policies"]},{"id":"manipulation","category":"concept","sec":1,"tier":1,"sources":[{"title":"Russ Tedrake, Robotic Manipulation (MIT course notes), Ch.1 Introduction","url":"https://manipulation.csail.mit.edu/intro.html"}],"as_of":"","related_ids":["grasping","dexterous-manipulation","bimanual-manipulation","contact-rich-manipulation","mobile-manipulation","locomotion"],"name":"Manipulation","alt":"操作","abbr":"","aliases":["Robotic Manipulation"],"one_liner":"A robot using its hand or a tool to contact an object and change its position, pose, or state.","explanation":"Manipulation means a robot using an arm, gripper, or dexterous hand to contact an object and change its position, orientation, or state — grasping, placing, opening a drawer, pouring water, folding clothes. Along with locomotion, it's one of embodied AI's two main capability areas: locomotion is about how the body moves, manipulation is about how the hands do work. MIT professor Russ Tedrake, in his course notes “Robotic Manipulation: Perception, Planning, and Control,” points out that manipulation is far more than pick-and-place: a hand has to constantly make and break contact, apply force, and switch rapidly between “sticking” and “sliding” friction states, which is already hard purely as a dynamics-and-control problem. Everyday tasks people find mundane — loading a dishwasher, folding laundry — remain extremely challenging for robots and sit at the frontier of robotics research. Common sub-areas include dexterous manipulation, bimanual manipulation, contact-rich manipulation, and manipulating deformable objects.","example":"Having a bimanual robot take dishes out of a dishwasher one at a time and put them away in a cabinet is a long-horizon manipulation task made up of many grasps and placements.","related":["Grasping","Dexterous Manipulation","Bimanual Manipulation","Contact-rich Manipulation","Mobile Manipulation","Locomotion"]},{"id":"grasping","category":"concept","sec":1,"tier":1,"sources":[{"title":"Data-Driven Grasp Synthesis - A Survey (Bohg et al., arXiv:1309.2660)","url":"https://arxiv.org/abs/1309.2660"}],"as_of":"","related_ids":["pick-and-place","grasp-pose-detection","force-closure","gripper","dexterous-hand","bin-picking"],"name":"Grasping","alt":"抓取","abbr":"","aliases":["Robotic Grasping"],"one_liner":"A robot picking up an object securely with a gripper, suction cup, or dexterous hand.","explanation":"Grasping is a robot using an end effector — a gripper, a suction cup, or a dexterous hand — to hold an object securely and lift it. The core question is where and how to grasp: an algorithm computes a grasp pose (the gripper's position and orientation) from a camera image or point cloud, and a motion planner then moves the hand there. Early methods relied on geometric and mechanical analysis, for instance checking force closure — whether the fingers' contact forces can resist external force and torque from any direction. A 2013 survey by Jeannette Bohg and colleagues organized data-driven approaches into three categories, based on whether the object was known, similar to a known object, or entirely unknown. Today the common approach is a deep network that predicts grasp poses directly from a point cloud, or a VLA model that outputs grasping actions end to end. Grasping is the first step in many manipulation tasks, from pick-and-place to mobile manipulation.","example":"A depth camera photographs a cluttered table; a network proposes dozens of candidate gripper poses over the point cloud and scores them, and the robot picks up the mug using the highest-scoring one.","related":["Pick-and-Place","Grasp Pose Detection","Force Closure","Gripper","Dexterous Hand","Bin Picking"]},{"id":"pick-and-place","category":"concept","sec":1,"tier":1,"sources":[{"title":"Transporter Networks: Rearranging the Visual World for Robotic Manipulation (arXiv:2010.14406)","url":"https://arxiv.org/abs/2010.14406"}],"as_of":"","related_ids":["grasping","rearrangement","skill-primitive","sorting","palletizing-depalletizing","transporter-networks"],"name":"Pick-and-Place","alt":"抓取放置","abbr":"","aliases":["Pick and Place"],"one_liner":"Picking an object up from one place, moving it, and setting it down at a target location — a basic manipulation task.","explanation":"Pick-and-place is the most basic and most common manipulation task: grasp the target object, then move it to a target location and set it down. In industry it has long been used for sorting, loading and unloading, and palletizing, usually through fixed programs combined with visual localization. In robot learning, it is a standard test task in VLA and imitation-learning papers, with instructions typically of the form “put A into/onto B.” It looks simple, but it actually chains together several steps — recognizing the target, choosing a grasp point, planning a collision-free trajectory, and determining the placement pose — and it is often used as an atomic skill inside longer-horizon tasks. Google's Transporter Networks (2020) modeled pick-and-place as a spatial displacement, “pick from here, move to there,” and learned tasks such as stacking blocks and assembling kits, with sample efficiency several orders of magnitude better than the comparison methods of the time.","example":"Given the instruction “put the red block in the bowl,” the robot first grips the block, moves it above the bowl, then releases the gripper.","related":["Grasping","Rearrangement","Skill Primitive","Sorting","Palletizing / Depalletizing","Transporter Networks"]},{"id":"bimanual-manipulation","category":"concept","sec":1,"tier":1,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA / ACT)","url":"https://arxiv.org/abs/2304.13705"},{"title":"Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation","url":"https://arxiv.org/abs/2401.02117"}],"as_of":"2024-01","related_ids":["dual-arm-robot","aloha","action-chunking-with-transformers","mobile-aloha","dexterous-manipulation","mobile-manipulation"],"name":"Bimanual Manipulation","alt":"双臂操作","abbr":"","aliases":["Dual-arm Manipulation"],"one_liner":"A robot uses two arms together, coordinating them to complete a single task.","explanation":"Bimanual manipulation means a robot controls two arms — or, on a humanoid, two hands — at once to jointly carry out a task: holding a bottle steady with one hand while unscrewing the cap with the other, say, or folding clothes with both hands together. Many everyday tasks are impossible or clumsy to do with a single arm, which makes this a core capability for housework, assembly, and similar settings. The challenges are that the action dimensions double, the two arms must be coordinated precisely in both time and space, and the task often involves soft, deformable objects and dense contact. A landmark project is ALOHA (2023), from Tony Zhao, Chelsea Finn, and colleagues: a low-cost bimanual teleoperation platform paired with the ACT algorithm reached an 80–90% success rate on fine tasks like opening a translucent condiment jar or inserting a battery into a remote control, using only about 10 minutes (50 demonstrations) of data per task — though the hardest task, threading a zip tie, managed only 20%. Mobile ALOHA (2024) later mounted the same two arms on a mobile base.","example":"With just 20–50 human demonstrations per task, Mobile ALOHA learned to autonomously stir-fry shrimp and plate it, open a two-door cabinet and put away a pot, and even call and ride an elevator — tasks that need both bimanual coordination and mobility.","related":["Dual-arm Robot","ALOHA","Action Chunking with Transformers","Mobile ALOHA","Dexterous Manipulation","Mobile Manipulation"]},{"id":"dexterous-manipulation","category":"concept","sec":1,"tier":1,"sources":[{"title":"Learning Dexterous In-Hand Manipulation (OpenAI)","url":"https://arxiv.org/abs/1808.00177"},{"title":"Dexterous Manipulation through Imitation Learning: A Survey","url":"https://arxiv.org/abs/2504.03515"}],"as_of":"","related_ids":["dexterous-hand","in-hand-manipulation","contact-rich-manipulation","tactile-sensor","dactyl","bimanual-manipulation"],"name":"Dexterous Manipulation","alt":"灵巧操作","abbr":"","aliases":["Dexterous Hand Manipulation","Multi-fingered Dexterous Manipulation"],"one_liner":"Using a multi-fingered hand, coordinated finger movement and force control to grasp, flip, and handle objects.","explanation":"Dexterous manipulation refers to a robot hand — usually a multi-fingered dexterous hand — using fine coordination and force control across several fingers to grasp, reorient, and manipulate objects: rotating a block in the hand, unscrewing a bottle cap with the fingers, adjusting a grip on a pen. Compared with a two-finger parallel gripper that just opens and closes, this is much closer to a human hand, and it is key to letting robots use human tools and handle complex objects. The challenges are the many joints and high degrees of freedom, frequent contact that is hard to model, and the difficulty of both tactile sensing and data collection. A landmark project is OpenAI's Dactyl (2018), which trained a Shadow dexterous hand in simulation using reinforcement learning plus domain randomization (randomly varying friction, appearance, and other parameters), then transferred it directly to the real hand; the policy spontaneously learned human-like tricks such as finger gaiting, where the fingers release and reposition in turn to keep rotating an object. More recent work increasingly uses teleoperation or human hand videos for imitation learning instead.","example":"OpenAI's Dactyl used a Shadow dexterous hand to reorient a block held in its palm to a target orientation: block pose was estimated from three ordinary cameras, fingertip positions were tracked with a motion-capture system, no tactile sensing was used, and the policy was trained entirely in simulation.","related":["Dexterous Hand","In-hand Manipulation","Contact-rich Manipulation","Tactile Sensor","Dactyl","Bimanual Manipulation"]},{"id":"locomotion","category":"concept","sec":1,"tier":1,"sources":[{"title":"Robot locomotion (Wikipedia)","url":"https://en.wikipedia.org/wiki/Robot_locomotion"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv:2109.11978)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"2021-09","related_ids":["legged-locomotion","bipedal-locomotion","rl-based-locomotion-control","sim-to-real-transfer","loco-manipulation","manipulation"],"name":"Locomotion","alt":"运动（移动）","abbr":"","aliases":[],"one_liner":"A robot's ability to move itself from place to place, using legs, wheels, or other means.","explanation":"Locomotion broadly covers the ways a robot moves itself: rolling on wheels, walking on legs, jumping, flying. Together with manipulation, which changes the state of external objects, it forms one of the two main branches of robot capability. In embodied AI, the term usually means legged locomotion in quadruped, biped, and humanoid robots: coordinating a dozen or more joints, keeping balance, and handling terrain like stairs and loose gravel. Traditional methods relied on dynamics models and model predictive control; the dominant recent approach trains a policy with reinforcement learning inside a GPU-parallel simulator, then transfers it to the real robot. A 2021 project from ETH Zurich and NVIDIA simulated thousands of ANYmal quadrupeds at once on a single GPU, training a flat-ground walking policy in under four minutes and a rough-terrain policy in about 20 minutes. In Chinese, 运动 also commonly corresponds to “motion” in general; “locomotion” specifically means moving the body itself.","example":"Using legged_gym inside Isaac Gym to simulate 4,096 ANYmal quadrupeds at once, training for about 20 minutes on terrain like stairs, slopes, and obstacles produces a policy that deploys straight to the real robot — able to climb stairs and cross obstacles.","related":["Legged Locomotion","Bipedal Locomotion","RL-based Locomotion Control","Sim-to-Real Transfer","Loco-manipulation","Manipulation"]},{"id":"navigation","category":"concept","sec":1,"tier":1,"sources":[{"title":"On Evaluation of Embodied Navigation Agents (Anderson et al., arXiv:1807.06757)","url":"https://arxiv.org/abs/1807.06757"}],"as_of":"","related_ids":["vision-and-language-navigation","object-goal-navigation","point-goal-navigation","success-weighted-by-path-length","simultaneous-localization-and-mapping","mobile-manipulation"],"name":"Navigation","alt":"导航","abbr":"","aliases":["Embodied Navigation","Visual Navigation"],"one_liner":"A robot deciding, from its own sensor readings, how to move itself to a target location or object.","explanation":"Navigation means an agent deciding where to go and how to get there, then moving itself to the target. Traditional robot navigation combines mapping and localization (SLAM), path planning, and obstacle-avoidance control. In embodied AI, navigation puts more weight on doing this in environments the robot has never seen, relying only on sensors such as a first-person camera. A 2018 consensus report from Peter Anderson and a dozen other researchers grouped navigation goals into three types: point goals (reach given coordinates), object goals (find an instance of a category, such as a refrigerator), and area goals (reach a type of region, such as a kitchen); it also proposed the SPL metric, which credits both success and how directly the path got there. The goal can also be specified another way — in natural language, as in vision-language navigation (VLN), or with a picture, as in image-goal navigation. Combining navigation with manipulation gives mobile manipulation.","example":"A robot given the goal “find the refrigerator” walks through an apartment it has never visited, looking as it goes, and counts as successful once it stops within a distance threshold of the fridge — the report suggests twice the robot's body width.","related":["Vision-and-Language Navigation","Object-Goal Navigation","Point-Goal Navigation","Success weighted by Path Length","Simultaneous Localization and Mapping","Mobile Manipulation"]},{"id":"vision-and-language-navigation","category":"concept","sec":1,"tier":1,"sources":[{"title":"Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments","url":"https://arxiv.org/abs/1711.07280"},{"title":"Room-to-Room (R2R) 数据集主页","url":"https://bringmeaspoon.org/"},{"title":"Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (VLN-CE)","url":"https://arxiv.org/abs/2004.02857"}],"as_of":"","related_ids":["navigation","room-to-room","success-weighted-by-path-length","object-goal-navigation","navila","vision-language-model"],"name":"Vision-and-Language Navigation","alt":"视觉语言导航","abbr":"VLN","aliases":["VLN"],"one_liner":"Having an agent follow a natural-language route description and use vision to reach a destination in an unfamiliar space.","explanation":"Vision-and-language navigation (VLN) is an embodied task proposed by Peter Anderson and colleagues in a CVPR 2018 paper: an agent gets an instruction like “exit and turn right, pass the sofa, and stop at the kitchen doorway,” and has to navigate step by step to the endpoint in a previously unvisited indoor environment, using only first-person vision. The companion Room-to-Room (R2R) dataset is built from real Matterport3D house scans, covering 90 buildings and about 22,000 human-written instructions. It tests how well language understanding, visual perception, and spatial memory work together. In the original R2R, the agent could only jump between waypoints with pre-rendered panoramas; a 2020 successor, VLN-CE, changed this to moving through continuous 3D space using low-level actions like moving forward and turning, closer to how a real robot operates. Models that run on actual robots, such as NaVid and NaVILA, followed after that.","example":"Given the instruction “walk straight down the hallway, turn left through the second door into the bedroom, and stop by the bed,” the robot looks and moves step by step, and succeeds if it stops by the bed.","related":["Navigation","Room-to-Room","Success weighted by Path Length","Object-Goal Navigation","NaVILA","Vision-Language Model"]},{"id":"mobile-manipulation","category":"concept","sec":1,"tier":1,"sources":[{"title":"Mobile manipulator (Wikipedia)","url":"https://en.wikipedia.org/wiki/Mobile_manipulator"},{"title":"Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation (arXiv:2401.02117)","url":"https://arxiv.org/abs/2401.02117"}],"as_of":"2024-01","related_ids":["mobile-manipulator","navigation","manipulation","loco-manipulation","mobile-aloha","mobile-base"],"name":"Mobile Manipulation","alt":"移动操作","abbr":"","aliases":[],"one_liner":"Mounting a robot arm on a mobile base so it can move and work at once, combining navigation and manipulation.","explanation":"Mobile manipulation means an arm mounted on a mobile base — wheeled or legged — that moves while carrying out manipulation tasks; robots built this way are called mobile manipulators, or composite robots in Chinese industry usage. A fixed arm can only work within its own reach, but adding a base lets it operate across a whole room, a warehouse, or even multiple floors. The cost is more degrees of freedom and a less structured environment: the system has to decide where the base goes and how the arm moves at the same time, and handle how base localization error affects grasping precision. Stanford's Mobile ALOHA (2024) collected data through low-cost, whole-body teleoperation and, with just 20 to 50 demonstrations per task, learned to fry shrimp and plate it, open a two-door cabinet to put away a pot, and call and enter an elevator on its own. The researchers also found that co-training with static bimanual data could raise success rates by up to 90%. It's a core capability for household and warehouse robots.","example":"Mobile ALOHA stands at the stove, pours in oil, and adds the shrimp; one arm tilts the pan while the other flips the shrimp with a spatula, then the base turns around and tips the shrimp into a bowl on the table behind it.","related":["Mobile Manipulator","Navigation","Manipulation","Loco-manipulation","Mobile ALOHA","Mobile Base (Chassis)"]},{"id":"long-horizon-task","category":"concept","sec":1,"tier":1,"sources":[{"title":"CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv:2112.03227)","url":"https://arxiv.org/abs/2112.03227"}],"as_of":"","related_ids":["skill-primitive","hierarchical-architecture","compounding-error","failure-recovery","embodied-memory","calvin-benchmark"],"name":"Long-horizon Task","alt":"长程任务","abbr":"","aliases":["Multi-stage Task"],"one_liner":"A task that requires completing many sub-steps in sequence, over a long stretch of time, to reach the goal.","explanation":"A long-horizon task requires executing many steps in a row across multiple sub-goals — “clear the table,” for instance, means clearing the plates, dumping the scraps, and loading the sink, in order. There are three main difficulties. First, errors accumulate: a small deviation early on pushes the later state further and further from the training data. Second, the success signal often only appears at the very end, making it hard to tell which intermediate step went wrong. Third, the agent must remember what it has already done and decide what to do next. The CALVIN benchmark (2021) chains 34 sub-tasks into sequences of five instructions each for evaluation; its imitation-learning baseline, trained and tested in the same environment, succeeded on one task in a row about 49% of the time but on all five in a row only 0.08% of the time. Common countermeasures include hierarchical architectures, where a high-level large model breaks down the task while a low-level policy executes atomic skills, along with memory modules and failure recovery.","example":"Clearing a table: put the plates and cups in a bin, throw napkins in the trash, then wipe the surface, in order — if any one step fails, the whole task counts as unfinished.","related":["Skill Primitive","Hierarchical Architecture","Compounding Error","Failure Recovery","Embodied Memory","CALVIN Benchmark"]},{"id":"skill-primitive","category":"concept","sec":1,"tier":2,"sources":[{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)","url":"https://arxiv.org/abs/2204.01691"},{"title":"Accelerating Robotic Reinforcement Learning via Parameterized Action Primitives (RAPS)","url":"https://arxiv.org/abs/2110.15360"}],"as_of":"","related_ids":["long-horizon-task","hierarchical-architecture","motion-primitives","saycan","llm-based-task-planning","dynamic-movement-primitives"],"name":"Skill Primitive","alt":"原子技能","abbr":"","aliases":["Action Primitive"],"one_liner":"The smallest reusable unit of action, such as grasp, place, or open drawer, combined to complete complex tasks.","explanation":"This approach breaks a robot's abilities into a set of basic actions that can be called and trained separately — “grasp,” “place,” “push,” “open drawer,” “move to a location.” Each skill primitive usually takes parameters, such as which object to grasp or where to place it, and is called in sequence by a higher-level planner or large model to assemble a long-horizon task. The benefit is that each skill is easy to train and validate on its own, and can be reused across tasks; the cost is that anything outside the skill library simply cannot be done, and the transitions between skills are a common source of errors. A landmark example is Google's SayCan (2022), which has a large language model pick the next step from a pretrained skill library and uses each skill's value function to judge whether it is feasible right now; RAPS (2021) instead treats hand-defined, parameterized primitives as the action space for reinforcement learning, to improve exploration and learning efficiency.","example":"“Put the soda in the fridge” can be broken into: navigate to the table, grasp the soda, navigate to the fridge, open the fridge door, place the soda, close the door — each step is a skill primitive.","related":["Long-horizon Task","Hierarchical Architecture","Motion Primitives","SayCan","LLM-based Task Planning","Dynamic Movement Primitives"]},{"id":"observation","category":"concept","sec":2,"tier":1,"sources":[{"title":"Basic Usage (Gymnasium Documentation)","url":"https://gymnasium.farama.org/introduction/basic_usage/"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["state-space","action-space","proprioception","partially-observable-markov-decision-process","policy","observation-action-pair"],"name":"Observation","alt":"观测","abbr":"","aliases":["Observation Space"],"one_liner":"The information an agent receives from the environment at each moment, the input a policy uses to decide.","explanation":"An observation is the input an agent receives from the environment at each step. In reinforcement learning's observation–action–reward loop, the environment returns a new observation both on reset and after every action; the observation space defines the format and range of these inputs — in Gymnasium, for example, CartPole's observation is just a few numbers: cart position, velocity, pole angle, and so on. A real robot's observation typically includes several camera images, depth or point-cloud data, proprioceptive information such as joint angles and gripper opening, and a language instruction. Observation is not the same as state: state is the environment's complete description, while an observation often only reveals part of it, with occlusion or noise — this situation is modeled as a Partially Observable Markov Decision Process (POMDP). Many policies take in the last several frames of observation; Diffusion Policy calls this window the observation horizon.","example":"A VLA policy's observation at each step: one RGB image each from a head camera and a wrist camera, seven joint angles, gripper opening, plus the instruction “put the cup on the plate.”","related":["State Space","Action Space","Proprioception","Partially Observable Markov Decision Process","Policy","Observation-Action Pair"]},{"id":"state-space","category":"concept","sec":2,"tier":2,"sources":[{"title":"Markov decision process - Wikipedia","url":"https://en.wikipedia.org/wiki/Markov_decision_process"},{"title":"State space (computer science) - Wikipedia","url":"https://en.wikipedia.org/wiki/State_space_(computer_science)"}],"as_of":"","related_ids":["observation","action-space","markov-decision-process","partially-observable-markov-decision-process","proprioception","state-space-model"],"name":"State Space","alt":"状态空间","abbr":"","aliases":["State"],"one_liner":"The set of all possible states of a system; for a robot, usually made up of joint angles, pose, velocity, and similar quantities.","explanation":"In reinforcement learning and control, a state is a set of quantities that describes a system's current situation well enough to predict how it will change next, and the state space is the collection of all possible states. A Markov Decision Process (MDP, the standard mathematical framework for reinforcement learning) writes this as S, which can be discrete, such as squares on a chessboard, or continuous, a vector of real numbers. A robot's state typically includes joint angles and angular velocities, end-effector pose, gripper opening, and the positions of nearby objects. A real robot often cannot access the full state and can only obtain an “observation” through cameras and sensors, in which case the problem becomes a Partially Observable MDP. In VLA papers, “state” usually refers to the robot's own joint values, fed into the model alongside images. Note this is unrelated to “state space models” such as Mamba.","example":"For a 7-degree-of-freedom robot arm with a parallel gripper, the proprioceptive state might be 7 joint angles, 7 joint angular velocities, and 1 gripper opening value — a 15-dimensional real vector in total.","related":["Observation","Action Space","Markov Decision Process","Partially Observable Markov Decision Process","Proprioception","State Space Model"]},{"id":"action-space","category":"concept","sec":2,"tier":1,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Gymnasium Documentation: Basic Usage","url":"https://gymnasium.farama.org/introduction/basic_usage/"}],"as_of":"","related_ids":["state-space","observation","policy","delta-action-vs-absolute-action","action-tokenizer","unified-action-space"],"name":"Action Space","alt":"动作空间","abbr":"","aliases":["Continuous Action Space","Discrete Action Space"],"one_liner":"The set of all actions an agent can take at each step, and the numerical form those actions take.","explanation":"Action space is a foundational concept in reinforcement learning and robot learning: the full set of actions an agent can choose from at each step. There are two kinds. A discrete action space has a finite number of options, like the handful of buttons in an Atari game. A continuous action space consists of real-valued numbers, such as the target angles for a robot arm's joints, or the displacement and gripper opening of the end effector (the gripper or tool at the very tip of the arm). Robot control at the lowest level is almost always continuous, though discrete examples exist too — embodied-navigation benchmarks often use just a few actions such as move forward, turn left, turn right, and stop. The same task can be represented with joint angles or end-effector pose, and with absolute or incremental (delta) values; different robots also have different action dimensions, which is exactly the mismatch cross-embodiment training has to handle. Vision-language-action (VLA) models commonly discretize continuous actions into tokens for output, or generate continuous values directly via diffusion, flow matching, or straightforward regression.","example":"The CartPole task has just two actions, push left and push right — a discrete action space. OpenVLA outputs robot-arm actions as 3D translation plus 3D rotation plus 1D gripper open/close, a 7-dimensional continuous action space.","related":["State Space","Observation","Policy","Delta (Relative) Action vs. Absolute Action","Action Tokenizer","Unified Action Space"]},{"id":"policy","category":"concept","sec":2,"tier":1,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["visuomotor-policy","imitation-learning","reinforcement-learning","observation","action-space","vision-language-action-model"],"name":"Policy","alt":"策略","abbr":"","aliases":["Policy Network"],"one_liner":"The rule that decides what action to take next given the current observation — usually a neural network.","explanation":"Policy is a core concept in reinforcement learning and robot learning: the mapping from the current state or observation to an action, usually written π. A deterministic policy always gives the same action for the same input; a stochastic policy outputs a distribution over actions and samples from it. In embodied AI, a policy is usually a neural network: it takes in camera images, joint angles, and other proprioceptive state, sometimes with a language instruction added, and outputs an end-effector pose, target joint angles, or a walking velocity command. Policies are mainly trained through imitation learning (copying human demonstrations) and reinforcement learning (trial and error guided by reward); VLA models are, at their core, a large-scale form of policy too. Don't confuse this with “model” as used elsewhere in reinforcement learning — there, “model” alone usually means a dynamics model that predicts how the environment changes next (as in “model-based reinforcement learning”), while the policy is what chooses the action.","example":"Diffusion Policy, on the Push-T task, reads in a camera image and the end effector's current position, and at each step outputs a short upcoming segment of end-effector targets that push a T-shaped block toward the goal position.","related":["Visuomotor Policy","Imitation Learning","Reinforcement Learning","Observation","Action Space","Vision-Language-Action Model"]},{"id":"inference","category":"concept","sec":2,"tier":1,"sources":[{"title":"What is AI inference? (IBM)","url":"https://www.ibm.com/think/topics/ai-inference"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"2024-10","related_ids":["reasoning","inference-latency","action-chunking","asynchronous-inference","inference-deployment","post-training-quantization"],"name":"Inference","alt":"推理（前向计算）","abbr":"","aliases":[],"one_liner":"Running a trained model forward on new input to produce an output, without updating its parameters.","explanation":"Inference means using an already-trained model: feed in new data, run one forward pass to get an output, without computing gradients or updating parameters — the counterpart to “training.” In Chinese, 推理 covers both “inference” and “reasoning” (a model thinking step by step), so papers need to be read in context to tell which one is meant. For robots, inference speed determines whether control can keep up: the π0 paper reports that processing three camera views and generating one action chunk takes about 73 milliseconds on a single RTX 4090. If inference is too slow, the robot stalls between action segments, which is why techniques such as action chunking, asynchronous inference, quantization, and TensorRT acceleration exist. Inference can run on the robot itself, or on a remote server with results sent back over the network, which adds network latency to the total.","example":"When π0 controls a robot at 50 Hz, it runs inference once every 0.5 seconds to produce one action chunk, then executes all 25 steps of that chunk before running inference again for the next one.","related":["Reasoning","Inference Latency","Action Chunking","Asynchronous Inference","Inference Deployment","Post-Training Quantization"]},{"id":"episode","category":"concept","sec":2,"tier":1,"sources":[{"title":"Gymnasium Documentation: Basic Usage","url":"https://gymnasium.farama.org/introduction/basic_usage/"},{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"}],"as_of":"","related_ids":["trajectory","rollout","success-rate","epoch","termination-vs-truncation","demonstration-data"],"name":"Episode","alt":"回合","abbr":"","aliases":[],"one_liner":"One full attempt at a task, from the environment's reset to completion, failure, or timeout.","explanation":"An episode is the basic unit of time in reinforcement learning and robot learning. After the environment resets to a starting state, the agent acts step by step until the task is completed, fails, or hits a step limit — that whole span is one episode, and the recorded sequence of observations and actions is a trajectory. Gymnasium marks the end of an episode with two separate flags: terminated means the task itself finished, successfully or not, and truncated means it was cut off by an external limit such as a time cap. In robot imitation learning, one demonstration usually corresponds to one episode; when an evaluation reports “50 episodes per task,” it means the robot attempted the task from scratch 50 times and the success rate was computed from that. Note that an episode is not the same thing as a training epoch, which is one full pass through the entire dataset.","example":"In a clothes-folding task, one episode runs from laying a wrinkled shirt on the table to the robot finishing the fold or hitting a two-minute timeout.","related":["Trajectory","Rollout","Success Rate","Epoch","Termination vs. Truncation","Demonstration Data"]},{"id":"trajectory","category":"concept","sec":2,"tier":1,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL（Trajectories）","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Lynch & Park, Modern Robotics, Ch.9 Trajectory Generation（预印本 PDF）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["episode","trajectory-planning","trajectory-tracking","demonstration-data","action-chunking","path-planning"],"name":"Trajectory","alt":"轨迹","abbr":"","aliases":[],"one_liner":"A sequence of positions or states and actions over time; also a full recording of one demonstration.","explanation":"“Trajectory” has two common meanings in embodied AI. In robotics proper, a trajectory is a path plus a timing schedule, specifying which position or joint angles the robot should be at at each moment; the textbook Modern Robotics defines it as “a path plus time scaling,” and a controller's job is to track this trajectory. In reinforcement learning and in datasets, a trajectory refers to the sequence of states and actions produced by an agent interacting with the environment, τ = (s0, a0, s1, a1, …); one recorded teleoperation demonstration is one trajectory, and datasets such as Open X-Embodiment measure their scale in number of trajectories. A short segment of future actions a model outputs at once, an action chunk, is also often loosely called a short trajectory.","example":"The 150 end-effector poses and gripper-opening values recorded at 50 Hz over the 3 seconds a robot arm takes to pick up a cup from a table make up one trajectory.","related":["Episode","Trajectory Planning","Trajectory Tracking","Demonstration Data","Action Chunking","Path Planning"]},{"id":"open-loop-control","category":"concept","sec":2,"tier":1,"sources":[{"title":"Open-loop controller (Wikipedia)","url":"https://en.wikipedia.org/wiki/Open-loop_controller"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)","url":"https://arxiv.org/abs/2506.07339"}],"as_of":"2025-06","related_ids":["closed-loop-control","action-chunking","action-horizon","real-time-chunking","temporal-ensembling","open-loop-evaluation"],"name":"Open-loop Control","alt":"开环","abbr":"","aliases":["Open-loop Execution"],"one_liner":"Executing a pre-computed sequence of actions straight through, without checking feedback along the way.","explanation":"Open-loop is originally a control-theory term: the control action doesn't depend on the system's output, and results are never checked during execution — a clothes dryer that just runs for a fixed time is one example. Closed-loop, by contrast, keeps measuring the result and correcting for it. In robot learning, “open-loop” usually means not reading any new observation for the duration of an action segment. Take action chunking: a model predicts several future steps of action at once, and the π0 paper states explicitly that the whole chunk runs open-loop — on a 50 Hz robot, it runs inference once every 0.5 seconds and doesn't look at new images for the 25 steps in between. Open-loop execution gives smooth, coherent motion and needs fewer inference calls, but reacts slowly to sudden changes — if an object gets bumped, the robot has to wait for the next inference call to correct course. Common compromises are executing only part of the predicted chunk before replanning, as in Diffusion Policy's receding-horizon approach, or using real-time chunking to compute the next segment while the current one is still playing out.","example":"Open-loop playback: a robot arm follows a pre-recorded trajectory to grab a cup, and still closes its gripper even if someone has moved the cup a few centimeters away; a closed-loop policy would adjust its hand position based on the new image.","related":["Closed-loop Control","Action Chunking","Action Horizon","Real-Time Chunking","Temporal Ensembling","Open-loop Evaluation"]},{"id":"closed-loop-control","category":"concept","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Closed-loop controller","url":"https://en.wikipedia.org/wiki/Closed-loop_controller"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["open-loop-control","perception-action-loop","closed-loop-evaluation","action-chunking","visual-servoing","model-predictive-control"],"name":"Closed-loop Control","alt":"闭环","abbr":"","aliases":["Closed-loop Execution","Feedback Control"],"one_liner":"Acting while watching: adjusting the next action based on the latest observation at every step.","explanation":"Closed-loop originated as a control-theory concept: a controller continuously measures a system's actual output, compares it to the target, and uses the difference to correct its action — cruise control that automatically adds power going uphill, for instance. The opposite, open-loop, executes a pre-planned sequence of commands without checking the result. In robot learning, closed-loop execution means the policy keeps reading new camera images and joint states while it runs and decides its next action accordingly, so it can handle surprises like a bumped object or a slipping grasp; open-loop means computing an entire trajectory once and then following it blindly. Many diffusion policies and VLA models output a chunk of actions at a time but only execute part of it before re-observing and re-predicting — receding-horizon control — trading off between smooth motion and timely feedback. “Closed-loop evaluation” also refers to actually running a policy in an environment, rather than just comparing its outputs to a dataset's recorded actions.","example":"A robot arm reaches for a cup, but someone nudges the cup 5 cm to the side. A closed-loop policy sees the change in the next camera frame and adjusts its path, while an open-loop policy executes the original plan and grasps empty air.","related":["Open-loop Control","Perception-Action Loop","Closed-Loop Evaluation","Action Chunking","Visual Servoing","Model Predictive Control"]},{"id":"hand-eye-coordination","category":"concept","sec":2,"tier":2,"sources":[{"title":"Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection (arXiv 1603.02199)","url":"https://arxiv.org/abs/1603.02199"}],"as_of":"","related_ids":["hand-eye-calibration","visual-servoing","visuomotor-policy","closed-loop-control","google-arm-farm","grasping"],"name":"Hand-Eye Coordination","alt":"手眼协调","abbr":"","aliases":[],"one_liner":"Using what a camera sees to guide an arm and gripper's motion in real time, adjusting as it goes.","explanation":"Hand-eye coordination originally describes a capability of humans and animals: the eyes see a target and the hand reaches out to grasp it accurately, correcting itself as it watches. Applied to robots, it means closing the loop between camera images and arm motion. The traditional approach first performs hand-eye calibration, computing the coordinate transform between the camera and the arm, then converts a detected target position into arm coordinates and executes the motion in one shot — any small calibration error causes a missed grasp. In 2016, Sergey Levine and colleagues at Google used 6 to 14 robot arms over two months to collect more than 800,000 grasp attempts, training a convolutional network to judge directly from a single monocular image whether a given gripper motion would succeed, with no camera calibration needed and with continuous adjustment during the grasp itself — a landmark example of learning hand-eye coordination end to end. Note this is a distinct concept from hand-eye calibration.","example":"An object gets bumped out of place mid-grasp; the robot keeps correcting the gripper's position based on the camera feed, rather than reaching to a single coordinate computed in advance.","related":["Hand-Eye Calibration","Visual Servoing","Visuomotor Policy","Closed-loop Control","Google Arm Farm","Grasping"]},{"id":"visuomotor-policy","category":"concept","sec":2,"tier":2,"sources":[{"title":"End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)","url":"https://arxiv.org/abs/1504.00702"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["policy","end-to-end","diffusion-policy","imitation-learning","vision-language-action-model","end-to-end-training-of-deep-visuomotor-policies"],"name":"Visuomotor Policy","alt":"视觉运动策略","abbr":"","aliases":[],"one_liner":"A control policy that maps camera images directly to robot actions, usually a neural network.","explanation":"A “policy” is a mapping from observations to actions; a visuomotor policy specifically takes input that is mostly images, often with proprioceptive state such as joint angles added, and outputs motor commands, joint positions, or end-effector pose. The term was popularized by Sergey Levine, Chelsea Finn, and colleagues' 2015 paper “End-to-End Training of Deep Visuomotor Policies”: using a convolutional network with about 92,000 parameters, it mapped raw images directly to joint torques to complete tasks such as screwing on a bottle cap, showing that training perception and control jointly worked better than training them separately. Today's ACT, Diffusion Policy, and language-augmented VLA models are all visuomotor policies, most commonly trained with imitation learning or reinforcement learning.","example":"The Diffusion Policy paper is literally titled “Visuomotor Policy Learning via Action Diffusion”: it takes in camera images and outputs a sequence of robot-arm actions to complete manipulation tasks such as pushing a T-shaped block.","related":["Policy","End-to-End","Diffusion Policy","Imitation Learning","Vision-Language-Action Model","End-to-End Training of Deep Visuomotor Policies"]},{"id":"action-multimodality","category":"concept","sec":2,"tier":2,"sources":[{"title":"Diffusion Policy 项目页（Columbia）","url":"https://diffusion-policy.cs.columbia.edu/"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion","url":"https://arxiv.org/abs/2303.04137"},{"title":"Behavior Transformers: Cloning k Modes with One Stone","url":"https://arxiv.org/abs/2206.11251"}],"as_of":"","related_ids":["diffusion-policy","flow-matching","gaussian-mixture-model","behavior-transformer","behavior-cloning","diffusion-action-head"],"name":"Action Multimodality","alt":"动作多峰性","abbr":"","aliases":["Multimodal Action Distribution","Mode Averaging"],"one_liner":"When several different actions are all correct in the same situation, so the action distribution has more than one peak.","explanation":"Action multimodality means that, for the same observation, more than one action is reasonable — going around an obstacle by passing it on the left or on the right, for example. If human demonstrations contain both, the probability distribution over actions has two peaks, or modes. If a policy is trained by directly regressing a single action with mean-squared error, the model learns the average of the two peaks instead — a failure called mode averaging — which in this example might send the robot straight into the obstacle. The fix is to use a policy that can represent multimodal distributions: a Gaussian mixture model, discretizing actions into tokens and predicting them as a classification problem (as in BeT), or diffusion policies and flow matching. The Diffusion Policy paper lists handling multimodal action distributions as one of its main advantages, which is part of why so many VLA models now use a diffusion or flow-matching action head.","example":"In the Push-T task, demonstrators sometimes push the T-shaped block from the left and sometimes from the right. A policy trained by direct regression tends to learn the average of the two, while Diffusion Policy commits to one side and pushes through on each run.","related":["Diffusion Policy","Flow Matching","Gaussian Mixture Model","Behavior Transformer","Behavior Cloning","Diffusion Action Head"]},{"id":"goal-conditioned-policy","category":"concept","sec":2,"tier":2,"sources":[{"title":"Goal-Conditioned Reinforcement Learning: Problems and Solutions (IJCAI 2022 Survey, arXiv 2201.08299)","url":"https://arxiv.org/abs/2201.08299"},{"title":"Hindsight Experience Replay (arXiv 1707.01495)","url":"https://arxiv.org/abs/1707.01495"},{"title":"Learning Latent Plans from Play (arXiv 1903.01973)","url":"https://arxiv.org/abs/1903.01973"}],"as_of":"","related_ids":["goal-conditioned-reinforcement-learning","hindsight-experience-replay","language-conditioned-policy","goal-conditioned-behavior-cloning","hindsight-relabeling","policy"],"name":"Goal-conditioned Policy","alt":"目标条件策略","abbr":"","aliases":[],"one_liner":"A policy that decides its action based on the goal to reach, not just the current observation.","explanation":"A goal-conditioned policy also takes a goal as input, usually written π(a | s, g): s is the current state or observation, and g is the goal, which can be a target position, a target image, or a target state. An ordinary policy learns only one fixed task, while a goal-conditioned policy uses a single network to handle a whole family of “reach different goals” tasks, so switching goals requires no retraining. In reinforcement learning this is called goal-conditioned RL, often paired with hindsight experience replay (HER), which relabels a trajectory that failed to reach its original goal as a success toward whatever state it actually reached, easing the sparse-reward problem. In imitation learning, Corey Lynch and colleagues' Play-LMP (2019) trains on unlabeled “play” data with no task labels, and at test time can carry out the corresponding action once given a goal. When the goal is described in a sentence instead, the result is called a language-conditioned policy.","example":"Given a robot arm a photo showing “the block in the upper-left corner of the table” as the goal, the policy pushes the block there; given a different photo, the same policy pursues the new goal instead.","related":["Goal-Conditioned Reinforcement Learning","Hindsight Experience Replay","Language-conditioned Policy","Goal-Conditioned Behavior Cloning","Hindsight Relabeling","Policy"]},{"id":"language-conditioned-policy","category":"concept","sec":2,"tier":2,"sources":[{"title":"Language Conditioned Imitation Learning over Unstructured Data (Lynch & Sermanet)","url":"https://arxiv.org/abs/2005.07648"},{"title":"CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks","url":"https://arxiv.org/abs/2112.03227"}],"as_of":"","related_ids":["policy","goal-conditioned-policy","vision-language-action-model","instruction-following","calvin-benchmark","play-data"],"name":"Language-conditioned Policy","alt":"语言条件策略","abbr":"","aliases":["Language-conditioned Imitation Learning"],"one_liner":"A robot policy that takes a language instruction as input and produces different actions depending on what it says.","explanation":"A policy is a mapping from observations to actions; a language-conditioned policy adds a natural-language instruction to that input, so a single network performs different tasks depending on the instruction, instead of training a separate model per task. Corey Lynch and Pierre Sermanet's language-conditioned imitation learning, proposed at Google in 2020 and published at RSS 2021, used one end-to-end network to learn pixel perception, language understanding, and continuous control together, drawing on large amounts of unlabeled “play” data, with less than 1% of it needing any language annotation. CALVIN (2021) is a benchmark built specifically to evaluate this kind of policy, requiring a robot to complete a long-horizon task made of a sequence of language instructions in order. Today's vision-language-action (VLA) models, such as RT-2 and π0, are essentially language-conditioned policies too, just built on a pretrained vision-language model backbone. The counterpart is a goal-conditioned policy, which specifies the task with a goal image instead of language.","example":"The same robot-arm policy pulls open a drawer when given the instruction “open the drawer,” and presses a button when given “press the green button” (tasks from the CALVIN benchmark).","related":["Policy","Goal-conditioned Policy","Vision-Language-Action Model","Instruction Following","CALVIN Benchmark","Play Data"]},{"id":"generalist-policy","category":"concept","sec":2,"tier":1,"sources":[{"title":"Octo: An Open-Source Generalist Robot Policy","url":"https://arxiv.org/abs/2405.12213"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"2024-10","related_ids":["policy","specialist-policy","octo","pi0","vision-language-action-model","cross-embodiment"],"name":"Generalist Policy","alt":"通用策略（通才策略）","abbr":"","aliases":["Generalist Robot Policy"],"one_liner":"A single control policy that can perform many tasks, and often work across many settings or robots.","explanation":"A policy is a model that maps observations to actions. A generalist policy is a single policy trained on large-scale, multi-task — often cross-robot — data, able to follow a language instruction or a goal image to complete many different tasks, and to generalize somewhat to new settings. The contrast is a specialist policy, trained for just one task or one robot. Generalist policies are usually pretrained at scale first, then fine-tuned to a specific robot and task with a small amount of target-domain data, an approach modeled on large language models. Notable examples include the open-source Octo (2024, trained on 800,000 trajectories from Open X-Embodiment and fine-tunable to a new robot in a few hours on a consumer GPU), OpenVLA, and Physical Intelligence's π0 series. Most vision-language-action (VLA) models today aim to be generalist policies.","example":"Give Octo either a spoken instruction or an image of the completed task, and it can output robot-arm actions accordingly, without needing a separate model trained for each task.","related":["Policy","Specialist Policy","Octo","π0","Vision-Language-Action Model","Cross-Embodiment"]},{"id":"specialist-policy","category":"concept","sec":2,"tier":2,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models (project page)","url":"https://robotics-transformer-x.github.io/"},{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv)","url":"https://arxiv.org/abs/2310.08864"},{"title":"Octo: An Open-Source Generalist Robot Policy","url":"https://arxiv.org/abs/2405.12213"}],"as_of":"2023-10","related_ids":["generalist-policy","policy","fine-tuning","cross-embodiment","rt-x","open-x-embodiment"],"name":"Specialist Policy","alt":"专用策略","abbr":"","aliases":["Single-task Policy"],"one_liner":"A policy trained only for a single robot, a single task, or a single setting — the counterpart to a generalist policy.","explanation":"A specialist policy is a control policy trained from scratch on data collected specifically for one robot, one task, or even one particular environment. This used to be how robot learning was mostly done: switch to a different robot or task, and you collect new data and train a new model from scratch. It tends to perform well on its own task and needs relatively little data, but it transfers poorly. The Open X-Embodiment comparison (2023) showed that in low-data settings, RT-1-X, trained on mixed data from 22 robots, outperformed the original single-dataset methods by about 50% on average. The common approach now is to pretrain a generalist policy first, then fine-tune it into a specialist policy with a small amount of task-specific data. Note that “expert policy” in imitation learning has a different meaning, referring to the expert that provides the demonstrations.","example":"Training an ACT policy from scratch using only “unscrew the bottle cap” demonstrations collected on one particular ALOHA robot: it works on that robot for that task, but switching to folding clothes or to a different arm requires collecting new data and retraining.","related":["Generalist Policy","Policy","Fine-tuning","Cross-Embodiment","RT-X","Open X-Embodiment"]},{"id":"sense-plan-act","category":"concept","sec":2,"tier":2,"sources":[{"title":"Robotic paradigm - Wikipedia","url":"https://en.wikipedia.org/wiki/Robotic_paradigm"},{"title":"Shakey the robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Shakey_the_robot"},{"title":"Subsumption architecture - Wikipedia","url":"https://en.wikipedia.org/wiki/Subsumption_architecture"}],"as_of":"","related_ids":["perception-action-loop","subsumption-architecture","hierarchical-architecture","end-to-end","task-planning","world-model"],"name":"Sense-Plan-Act","alt":"感知-规划-行动范式","abbr":"SPA","aliases":["SPA","Hierarchical Paradigm"],"one_liner":"The classic control loop where a robot senses, models, plans its next move, then acts, and repeats.","explanation":"This is the classic three-step process in traditional robotics: use sensors to perceive the environment and update an internal world model, plan the next action on top of that model, then hand it off to the actuators to execute, and repeat. A landmark example is Shakey the Robot at Stanford Research Institute (SRI) in the 1960s–70s, which used the STRIPS planner and A* search to break an instruction down into executable steps. The drawback is that it depends on an accurate world model, is computationally slow, and breaks down easily when the environment changes. In the mid-1980s, Rodney Brooks proposed the subsumption architecture, which builds no global model and lets perception drive behavior directly; hybrid architectures combining both approaches appeared later. Today's modular robotics pipelines still follow this basic idea, while end-to-end VLA models instead use a single network to map observations directly to actions.","example":"Given the instruction 'push the block off the platform,' Shakey first identifies the platform and a ramp, then plans the steps 'push the ramp into place, climb onto the platform, push the block,' and finally executes them one by one.","related":["Perception-Action Loop","Subsumption Architecture","Hierarchical Architecture","End-to-End","Task Planning","World Model"]},{"id":"braincerebellum-architecture","category":"concept","sec":2,"tier":1,"sources":[{"title":"RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration","url":"https://arxiv.org/abs/2505.03673"},{"title":"FlagOpen/RoboOS (GitHub)","url":"https://github.com/FlagOpen/RoboOS"},{"title":"工业和信息化部：《人形机器人创新发展指导意见》解读（2023-11）","url":"https://www.miit.gov.cn/zwgk/zcjd/art/2023/art_e3f5686c2f0d49f9968b7ae011d558e1.html"}],"as_of":"2025-06","related_ids":["dual-system-architecture","hierarchical-architecture","embodied-foundation-model","locomotion-control","whole-body-control","roboos"],"name":"Brain–Cerebellum Architecture","alt":"大脑-小脑架构（大小脑）","abbr":"","aliases":["Robot Brain and Cerebellum","Brain-Cerebellum Hierarchical Architecture"],"one_liner":"Splitting a robot's software into a 'brain' that plans and reasons and a 'cerebellum' that executes movement.","explanation":"This is informal shorthand used in China's embodied-AI industry (大小脑, literally “big brain, small brain”) for a layered robot system, borrowed from the division of labor in the human brain. The “brain” is usually a large multimodal model (one that handles both images and text), responsible for understanding instructions, perceiving the scene, breaking a task into steps, and other high-level decisions. The “cerebellum” turns each sub-task into concrete motion: it might be a library of skills, a VLA (vision-language-action) policy, or a motion controller built with reinforcement learning or classical control, running at a higher frequency and closer to the hardware. Splitting the system this way lets each layer be trained and swapped independently, and lets one brain coordinate several different robots. The Beijing Academy of Artificial Intelligence's RoboOS (2025) uses this architecture, with RoboBrain as the brain and a pluggable cerebellum skill library for execution. China's Ministry of Industry and Information Technology used a similar “brain, cerebellum, limbs” breakdown of key technologies in its 2023 Guiding Opinions on Innovation and Development of Humanoid Robots.","example":"A humanoid robot hears, “hand me the cup on the table.” The brain model breaks this into walking to the table, picking up the cup, and handing it over; the cerebellum's walking controller and grasping policy then carry out each step.","related":["Dual-System Architecture (System 1 / System 2)","Hierarchical Architecture","Embodied Foundation Model","Locomotion Control","Whole-Body Control","RoboOS"]},{"id":"generalization","category":"concept","sec":3,"tier":1,"sources":[{"title":"A Taxonomy for Evaluating Generalist Robot Manipulation Policies (STAR-Gen)","url":"https://arxiv.org/abs/2503.01238"}],"as_of":"2025-03","related_ids":["overfitting","out-of-distribution","zero-shot","object-generalization","visual-generalization","semantic-generalization"],"name":"Generalization","alt":"泛化","abbr":"","aliases":["Generalization Ability"],"one_liner":"A model's ability to still perform correctly in new situations it never saw during training.","explanation":"Generalization is a core machine-learning concept: how well a model performs on new examples outside its training data, rather than just memorizing the training set (a model that memorizes training data and falls apart on new data is said to overfit). Generalization is especially hard in robot learning because the real world varies so much — a different tablecloth, a different cup, a shifted position, or a differently worded instruction can all make a policy fail. Researchers therefore often evaluate different kinds of variation separately: object generalization, scene generalization, position generalization, instruction generalization. STAR-Gen (2025), proposed by Jensen Gao, Dorsa Sadigh, and colleagues, splits generalization into three categories — visual (appearance and background changes), semantic (concept and instruction changes), and behavioral (requiring a different way of acting) — and found that open-source VLA models, despite being pretrained on internet-scale language data, still often struggle specifically with semantic generalization. Generalization ability is the main yardstick for judging a generalist policy.","example":"A policy trained only to pick up a red block on a white table is tested on a wood-grain table (visual generalization), or given the instruction “pick up the thing that can hold water” instead (semantic generalization), to probe how well it generalizes.","related":["Overfitting","Out-of-Distribution","Zero-shot","Object Generalization","Visual Generalization","Semantic Generalization"]},{"id":"in-distribution","category":"concept","sec":3,"tier":2,"sources":[{"title":"Generalized Out-of-Distribution Detection: A Survey (Yang et al.)","url":"https://arxiv.org/abs/2110.11334"}],"as_of":"","related_ids":["out-of-distribution","generalization","distribution-shift","long-tail-problem","robustness","object-generalization"],"name":"In-distribution","alt":"分布内","abbr":"ID","aliases":["ID"],"one_liner":"A test-time situation drawn from the same distribution as the training data — something the model has effectively seen before.","explanation":"Machine learning typically assumes training and test data come from the same probability distribution; a test sample satisfying this is called in-distribution (ID), and one from a different distribution is out-of-distribution (OOD). A 2021 OOD-detection survey by Jingkang Yang and colleagues splits distribution change into two kinds: covariate shift, where the input's appearance changes — different lighting, background, or camera — but the task categories stay the same, and semantic shift, where entirely new categories not seen in training appear. In robotics, “in-distribution” usually means the objects, scene, placement, and instructions at test time all fall within the range covered by the training demonstrations. Many policies have a high success rate in-distribution but drop sharply the moment the tablecloth or the object changes, so papers often report in-distribution and out-of-distribution results separately as a way to measure generalization. This concept connects directly to generalization, out-of-distribution, distribution shift, and the long-tail problem.","example":"A robot arm is trained with 50 demonstrations to put a red block on a plate; if the test still uses the same table, the same red block, and a position within the training range, that is an in-distribution test — swapping in an unseen green cup would be out-of-distribution.","related":["Out-of-Distribution","Generalization","Distribution Shift","Long-tail Problem","Robustness","Object Generalization"]},{"id":"out-of-distribution","category":"concept","sec":3,"tier":1,"sources":[{"title":"Towards Out-Of-Distribution Generalization: A Survey (arXiv:2108.13624)","url":"https://arxiv.org/abs/2108.13624"}],"as_of":"","related_ids":["in-distribution","generalization","distribution-shift","long-tail-problem","robustness","zero-shot"],"name":"Out-of-Distribution","alt":"分布外","abbr":"OOD","aliases":["OOD","OOD Generalization","Out-of-Distribution Generalization"],"one_liner":"Test-time data that comes from a different distribution than the training data, such as new objects or scenes.","explanation":"Machine learning typically assumes training and test data come from the same distribution (the i.i.d. assumption). Out-of-distribution (OOD) describes test inputs that fall outside the training distribution — unseen objects, lighting, tablecloths, camera angles, or ways of phrasing an instruction; inputs that do fall within the training distribution are called in-distribution (ID). Model performance usually drops noticeably on OOD inputs, and the study of keeping models useful anyway is called OOD generalization; a 2021 survey by Peng Cui's group at Tsinghua University gives a systematic overview. Robots run into this especially often: the real world varies endlessly while demonstration data only covers a limited set of scenarios, and once execution drifts even slightly off course, later observations drift away from the training data too, a problem called distribution shift. When papers evaluate “generalization,” they are usually measuring success rate under deliberately constructed out-of-distribution conditions.","example":"A policy trained only with a red cup on a white table is tested with a wood-grain table and a blue bowl instead — that is an out-of-distribution test.","related":["In-distribution","Generalization","Distribution Shift","Long-tail Problem","Robustness","Zero-shot"]},{"id":"distribution-shift","category":"concept","sec":3,"tier":2,"sources":[{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (arXiv 1011.0686)","url":"https://arxiv.org/abs/1011.0686"},{"title":"Domain adaptation - Wikipedia","url":"https://en.wikipedia.org/wiki/Domain_adaptation"}],"as_of":"","related_ids":["compounding-error","dagger","behavior-cloning","out-of-distribution","recovery-and-correction-data","domain-adaptation"],"name":"Distribution Shift","alt":"分布偏移（协变量偏移）","abbr":"","aliases":["Covariate Shift","Dataset Shift"],"one_liner":"When the data seen at deployment has a different distribution than the training data, hurting performance.","explanation":"Distribution shift broadly describes any mismatch between the distribution of training data and the distribution of data seen in actual use. Covariate shift is one specific kind: the distribution of inputs changes, but the true mapping from input to correct output stays the same; two other common kinds are label shift and concept shift. This is especially prominent in robot imitation learning: a policy only ever sees states the expert visited during training, so the moment it makes even a small error during execution, it drifts into states the demonstrations never covered, and that error compounds into a bigger deviation, a problem called compounding error. Stéphane Ross and colleagues, in a 2011 AISTATS paper, pointed out that this setting, where the agent's own actions determine its next input, violates the i.i.d. assumption, and proposed DAgger: let the policy run on its own, have an expert label the correct action for the new states it encounters, and gradually pull the training distribution toward the distribution the policy actually meets.","example":"A cup-grasping policy trained with behavior cloning always approaches the cup from directly above in the demonstrations. At deployment the arm is off by 2 cm, a position never seen in training, so the policy's output gets progressively worse until it knocks the cup over.","related":["Compounding Error","DAgger","Behavior Cloning","Out-of-Distribution","Recovery and Correction Data","Domain Adaptation"]},{"id":"long-tail-problem","category":"concept","sec":3,"tier":2,"sources":[{"title":"Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation (ICRA 2026)","url":"https://arxiv.org/abs/2602.06512"},{"title":"Dynamically Conservative Self-Driving Planner for Long-Tail Cases","url":"https://arxiv.org/abs/2305.07497"}],"as_of":"2026-02","related_ids":["out-of-distribution","generalization","data-flywheel","failure-recovery","robustness","autonomous-driving"],"name":"Long-tail Problem","alt":"长尾问题","abbr":"","aliases":["Corner Case","Long-tail Distribution"],"one_liner":"The mass of individually rare situations that training data barely covers, and where models fail most often.","explanation":"“Long-tail” comes from the long-tail distribution in statistics: a small number of common situations (the head) account for most of the data, while a huge number of rare situations (the tail) each occur only a little but add up to a large total. Autonomous driving was among the first fields to treat this as a central challenge: driving is normal almost all the time, but occasionally the car meets a corner case — a construction detour, debris on the road, a pedestrian suddenly running out — and these scenarios have little data even though they are often the ones that matter most for safety. Embodied AI faces the same issue: the objects, arrangements, and unexpected situations in a home vary endlessly, and demonstration data is naturally skewed toward a handful of common tasks; the ICRA 2026 paper “Beyond the Majority” finds that general-purpose robot policies generalize noticeably worse on tail tasks with sparse data, and that ordinary resampling methods offer only limited help. Countermeasures include continuously collecting failure data from deployment, called a data flywheel, using simulation and synthetic data to fill in the tail, and improving a model's generalization and failure-recovery ability. The long tail is the main source of the gap between “the demo works” and “the product actually ships.”","example":"A household robot can reliably clear away common bowls and chopsticks, but a spilled bowl of soup, a fork stuck in a crack in the table, or a pet suddenly jumping onto the table are all long-tail situations.","related":["Out-of-Distribution","Generalization","Data Flywheel","Failure Recovery","Robustness","Autonomous Driving"]},{"id":"zero-shot","category":"concept","sec":3,"tier":1,"sources":[{"title":"Zero-shot learning - Wikipedia","url":"https://en.wikipedia.org/wiki/Zero-shot_learning"},{"title":"Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments","url":"https://robotutilitymodels.com/"}],"as_of":"","related_ids":["few-shot","generalization","out-of-distribution","open-vocabulary","fine-tuning","generalist-policy"],"name":"Zero-shot","alt":"零样本","abbr":"","aliases":["Zero-shot Learning","Zero-shot Generalization"],"one_liner":"A model handling a new task or environment directly, with no training examples for it at all.","explanation":"Zero-shot originated as a machine-learning concept: at test time, a model must recognize a category it never saw during training, transferring learned knowledge with the help of side information such as attributes or text descriptions; a 2009 NeurIPS paper by Mark Palatucci and colleagues introduced the term “zero-shot learning.” In the era of large models and embodied AI, the meaning has broadened to “no new data collection and no fine-tuning for a new task, object, or scene — the model is used as is.” A generalist policy opening a drawer in a kitchen it has never seen, for example, is called zero-shot generalization. It contrasts with few-shot, where a handful of examples are given, and is commonly used to measure a robot foundation model's generalization ability. When reading papers, note that some claims of “zero-shot” only mean the scene was unseen, while the type of task was actually seen during training.","example":"Robot Utility Models trains one policy per task for five task types, such as opening cabinets and opening drawers, then deploys them directly to unseen new environments with no further data collection or fine-tuning, reporting an average success rate of about 90%.","related":["Few-shot","Generalization","Out-of-Distribution","Open-vocabulary","Fine-tuning","Generalist Policy"]},{"id":"few-shot","category":"concept","sec":3,"tier":2,"sources":[{"title":"Language Models are Few-Shot Learners (GPT-3, arXiv 2005.14165)","url":"https://arxiv.org/abs/2005.14165"},{"title":"RVT-2: Learning Precise Manipulation from Few Demonstrations (arXiv 2406.08545)","url":"https://arxiv.org/abs/2406.08545"}],"as_of":"","related_ids":["zero-shot","in-context-learning","fine-tuning","demonstration-data","sample-efficiency","one-shot-imitation-learning"],"name":"Few-shot","alt":"少样本","abbr":"","aliases":["Few-shot Learning"],"one_liner":"Learning or performing a new task correctly from just a handful to a few dozen examples.","explanation":"Few-shot describes handling a new task with very few examples. In large language models, it usually means writing a handful of examples directly into the prompt and getting the model to follow them without updating any parameters; the 2020 GPT-3 paper, “Language Models are Few-Shot Learners,” made this usage widely known, and the capability is also called in-context learning. In robotics, few-shot more often means collecting just a handful to a few dozen demonstrations — recordings of a human teleoperating the robot through a task — and fine-tuning a pretrained model to learn the new task from them. Real-robot data collection is slow and expensive, so how many demonstrations a method needs to learn a new task is an important measure of how practical it is. It contrasts with zero-shot, where no examples at all are given.","example":"RVT-2 learns manipulation tasks that require high precision using only 10 demonstrations on a real robot.","related":["Zero-shot","In-Context Learning","Fine-tuning","Demonstration Data","Sample Efficiency","One-shot Imitation Learning"]},{"id":"robustness","category":"concept","sec":3,"tier":2,"sources":[{"title":"Robustness (computer science) - Wikipedia","url":"https://en.wikipedia.org/wiki/Robustness_(computer_science)"},{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models","url":"https://arxiv.org/abs/2510.13626"}],"as_of":"2025-10","related_ids":["generalization","domain-randomization","data-augmentation","out-of-distribution","generalization-robustness-evaluation","push-recovery"],"name":"Robustness","alt":"鲁棒性","abbr":"","aliases":[],"one_liner":"A system's ability to keep performing without much degradation when inputs or the environment carry noise or small changes.","explanation":"Robustness means a system's ability to keep working properly when inputs contain errors, the environment is perturbed, or conditions change. It overlaps with generalization but emphasizes something different: generalization asks “does it still work with a new object or a new scene,” while robustness is more concerned with whether the same task falls apart under disturbances such as lighting changes, a shifted camera, sensor noise, or being bumped by a person. Real-world disturbances are everywhere, so robots are held to a high standard here. LIBERO-Plus (2025) systematically adds seven categories of perturbation to the LIBERO simulation benchmark and finds that some VLA models' success rates drop from 95% to under 30% with only a slight change in camera viewpoint or initial state, and that the models often ignore the language instruction as well. Common ways to improve robustness include domain randomization, data augmentation, and training with deliberate perturbations.","example":"A quadruped robot gets kicked from the side while walking but adjusts its steps and stays standing; a robot-arm policy still manages to grasp an object even after its camera has been knocked a few centimeters out of place.","related":["Generalization","Domain Randomization","Data Augmentation","Out-of-Distribution","Generalization / Robustness Evaluation","Push Recovery"]},{"id":"distractor-objects","category":"concept","sec":3,"tier":2,"sources":[{"title":"THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (arXiv 2402.08191)","url":"https://arxiv.org/abs/2402.08191"}],"as_of":"","related_ids":["visual-generalization","robustness","out-of-distribution","the-colosseum-a-benchmark-for-evaluating-generalization-for","generalization-robustness-evaluation","open-vocabulary-object-detection"],"name":"Distractor Objects","alt":"干扰物","abbr":"","aliases":["Distractors"],"one_liner":"Extra objects in a scene that are irrelevant to the current task but can throw a policy off.","explanation":"Distractors are objects in a manipulation or navigation scene that are unrelated to the current instruction — for example, a bowl, a toy, and other cups sitting on the table when the robot is told to pick up the red cup. Evaluations often deliberately add or remove distractors to see whether a policy grasps the wrong target or gets thrown off by occlusion or visual changes, making this a standard way to measure visual generalization and robustness. The Colosseum benchmark (RSS 2024) tests manipulation policies along 14 kinds of perturbation and finds that a single perturbation alone can drop success rates by 30–50%, with the number of distractors, the target object's color, and lighting having the largest effects. Common countermeasures include training with cluttered scenes and data augmentation, or first using open-vocabulary detection to box in the actual target.","example":"The training data shows only a single apple on the table; at test time an orange and a small red ball are added next to it, and an imitation-learning policy may reach for the red ball instead.","related":["Visual Generalization","Robustness","Out-of-Distribution","The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation","Generalization / Robustness Evaluation","Open-Vocabulary Object Detection"]},{"id":"failure-recovery","category":"concept","sec":3,"tier":2,"sources":[{"title":"REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction (arXiv 2306.15724)","url":"https://arxiv.org/abs/2306.15724"},{"title":"AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation (arXiv 2410.00371)","url":"https://arxiv.org/abs/2410.00371"},{"title":"RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction (arXiv 2509.07953)","url":"https://arxiv.org/abs/2509.07953"}],"as_of":"2025-09","related_ids":["recovery-and-correction-data","human-intervention-data","compounding-error","human-in-the-loop","rac","long-horizon-task"],"name":"Failure Recovery","alt":"失败恢复","abbr":"","aliases":["Error Recovery"],"one_liner":"A robot noticing it made a mistake or is about to fail, and adjusting on its own to still finish the task.","explanation":"Failure recovery means a robot detects something has gone wrong during execution — a missed grasp, a dropped object, getting stuck — and continues the task anyway by retrying, changing approach, or backing off to a safe state. An imitation-learning policy trained only on successful demonstrations tends to drift into states it has never seen the moment it deviates even slightly from the demonstrated trajectory, and small errors compound (compounding error), so recovery ability often determines whether a long-horizon task can be completed at all. There are two common approaches. One detects and explains the failure, then replans: REFLECT (CoRL 2023) uses a large language model to summarize the robot's experience and explain what went wrong, and AHA (2024) trains a vision-language model specifically to judge failures. The other teaches recovery actions to the policy directly: RaC (2025) has a human step in right before a failure, guide the robot back to a familiar state, then demonstrate the correction, and trains the policy on that kind of data.","example":"When RaC's bimanual robot is hanging a shirt or packing a box, a human takes over right before a failure, guides the robot back to a familiar state, and demonstrates the corrective action; the policy learns to recover on its own from data like this.","related":["Recovery and Correction Data","Human Intervention Data","Compounding Error","Human-in-the-Loop","RaC","Long-horizon Task"]},{"id":"object-generalization","category":"concept","sec":3,"tier":2,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/html/2406.09246"},{"title":"Data Scaling Laws in Imitation Learning for Robotic Manipulation","url":"https://arxiv.org/abs/2410.18647"}],"as_of":"","related_ids":["generalization","scene-generalization","spatial-generalization","semantic-generalization","visual-generalization","zero-shot"],"name":"Object Generalization","alt":"物体泛化","abbr":"","aliases":[],"one_liner":"A policy's ability to still complete the same task when the object is swapped for one it never saw in training.","explanation":"Object generalization is the most common type of generalization tested for robot policies: whether a task can still be completed once the object is swapped for one absent from the training data — a new category, a new size or shape, a new color or material. OpenVLA's evaluation breaks this down further: changes in color and appearance count as visual generalization, changes in size and shape count as physical generalization, and entirely unseen target objects count as semantic generalization. It matters because the variety of objects in real homes and factories is nearly endless — collecting data for every single one is impossible. A 2024 study on scaling imitation-learning data found that a policy's generalization to new objects grows as a power law with the number of distinct object categories in training, and that object diversity matters more than simply piling up more demonstrations of the same objects.","example":"Demonstrations of 'put the cup on the plate' are collected using only 3 kinds of cups; at test time the policy is tried with an unseen mug, a paper cup, and a glass to see how much the success rate drops.","related":["Generalization","Scene Generalization","Spatial Generalization","Semantic Generalization","Visual Generalization","Zero-shot"]},{"id":"spatial-generalization","category":"concept","sec":3,"tier":2,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/html/2406.09246"},{"title":"DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning","url":"https://arxiv.org/abs/2502.16932"},{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models","url":"https://arxiv.org/abs/2510.13626"}],"as_of":"2025-10","related_ids":["generalization","object-generalization","scene-generalization","out-of-distribution","demogen","libero-plus"],"name":"Spatial Generalization","alt":"位置泛化（空间泛化）","abbr":"","aliases":["Position Generalization","Motion Generalization"],"one_liner":"A policy's ability to still succeed when an object is placed at a position or orientation absent from training.","explanation":"Spatial (or position) generalization means the task and the object stay the same, but the object's, or the robot's own, initial position or orientation shifts to somewhere the training data never covered, and the question is whether the policy can still succeed. OpenVLA's evaluation calls this “motion generalization,” defining it as unseen object positions and orientations. It is a common weak point for visuomotor policies: a policy learned through imitation learning is often only reliable near the region the demonstrations covered, and success rates drop noticeably once an object is moved further away — which is why data collection involves repeatedly placing objects at different spots on the table. DemoGen addresses this by synthesizing large numbers of demonstrations at different positions from a single real human demonstration; LIBERO-Plus found that success rates for some VLA models fall from 95% to under 30% with only a slight perturbation to the robot's initial state or camera viewpoint.","example":"During training, the block is only ever placed on the left half of the table; at test time it is placed on the right half or rotated 90 degrees, to see whether the arm can still pick it up the same way.","related":["Generalization","Object Generalization","Scene Generalization","Out-of-Distribution","DemoGen","LIBERO-Plus"]},{"id":"scene-generalization","category":"concept","sec":3,"tier":2,"sources":[{"title":"Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (Xie et al., 2023)","url":"https://arxiv.org/abs/2307.03659"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization","url":"https://arxiv.org/abs/2504.16054"},{"title":"Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments","url":"https://arxiv.org/abs/2409.05865"}],"as_of":"2025-04","related_ids":["generalization","object-generalization","task-generalization","visual-generalization","out-of-distribution","data-diversity"],"name":"Scene Generalization","alt":"场景泛化","abbr":"","aliases":["Environment Generalization"],"one_liner":"A policy still completing its task when moved into a room, table, lighting, or background it never saw in training.","explanation":"This is one dimension of generalization: whether a trained robot policy still succeeds once it is moved to an environment absent from training — a new room, a new tabletop texture, different lighting, a shifted camera position, or a cluttered background. Robot data is mostly collected in a handful of labs, so a model can easily end up memorizing the scene's appearance itself and fail the moment the kitchen changes, making this a key metric for whether a robot can actually enter a user's home. Tianhe Yu, Chelsea Finn, and colleagues (2023) broke the contributing factors into 11 categories, including lighting and camera pose, and tested each separately; Physical Intelligence's π0.5 (2025) uses “tidying a kitchen and bedroom in an entirely new home” as its main evaluation. Common countermeasures include collecting data from a larger and more diverse set of environments, data augmentation, and co-training with web data.","example":"π0.5 completes long-horizon tasks such as tidying a kitchen and organizing a bedroom in real homes that never appeared in its training data.","related":["Generalization","Object Generalization","Task Generalization","Visual Generalization","Out-of-Distribution","Data Diversity"]},{"id":"visual-generalization","category":"concept","sec":3,"tier":3,"sources":[{"title":"Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (Xie et al., 2023)","url":"https://arxiv.org/abs/2307.03659"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/abs/2406.09246"},{"title":"What Can RL Bring to VLA Generalization? An Empirical Study","url":"https://arxiv.org/abs/2505.19789"}],"as_of":"2025-05","related_ids":["generalization","semantic-generalization","spatial-generalization","distractor-objects","domain-randomization","data-augmentation"],"name":"Visual Generalization","alt":"视觉泛化","abbr":"","aliases":["Visual Robustness","Appearance Generalization"],"one_liner":"Still completing the task when the scene looks different: new background, lighting, colors, distractors, or camera angle.","explanation":"Visual generalization is a subcategory of generalization: whether a robot policy still completes the same task under visual conditions it never saw in training, such as a new background or tablecloth, different lighting, an object with a new color or texture, extra distractor objects in the frame, or a moved camera. The task and the actions needed haven't changed, only what the scene looks like, so it is usually evaluated separately from semantic generalization (new objects, new instructions) and position generalization; OpenVLA's real-robot evaluation lists it as its own category. Imitation-learning policies that learn actions directly from pixels tend to also memorize irrelevant details like background and lighting. In 2023, Xie, Finn, and colleagues isolated these factors one by one and found that new backgrounds are the easiest to adapt to and new camera positions the hardest. Common fixes include domain randomization, data augmentation, pretrained vision encoders, and collecting data across more varied scenes.","example":"OpenVLA's real-robot evaluation of “put the eggplant in the pot” used a pot made of papier-mâché, visually different from the pots in the BridgeData V2 training data, to test whether the policy could still recognize the pot and complete the task.","related":["Generalization","Semantic Generalization","Spatial Generalization","Distractor Objects","Domain Randomization","Data Augmentation"]},{"id":"semantic-generalization","category":"concept","sec":3,"tier":3,"sources":[{"title":"What Can RL Bring to VLA Generalization? An Empirical Study (arXiv 2505.19789)","url":"https://arxiv.org/abs/2505.19789"},{"title":"RT-2: New model translates vision and language into action (Google DeepMind)","url":"https://deepmind.google/discover/blog/rt-2-new-model-translates-vision-and-language-into-action/"}],"as_of":"2025-05","related_ids":["generalization","visual-generalization","object-generalization","compositional-generalization","vision-language-action-model","rt-2"],"name":"Semantic Generalization","alt":"语义泛化","abbr":"","aliases":[],"one_liner":"Still understanding and correctly acting on unfamiliar objects, concepts, or phrasing it hasn't seen before.","explanation":"Semantic generalization is a generalization axis commonly used when evaluating robot policies, especially vision-language-action (VLA) models: whether a model can generalize at the level of meaning — unseen object categories, new containers, instructions phrased differently, or tasks requiring common-sense or conceptual reasoning. It is distinguished from visual generalization (changes in background, lighting, texture) and generalization at the execution level (changes in object position or starting pose). 2025's “What Can RL Bring to VLA Generalization?” builds tests along visual, semantic, and execution axes; the semantic category includes unseen objects, unseen containers, unseen instruction phrasings, and distractor containers. VLA models are seen as promising largely because they can inherit semantic knowledge from internet-scale image-text pretraining.","example":"RT-2 can carry out instructions such as “move the coke can next to the photo of Taylor Swift” or “pick up the thing that could be used as an improvised hammer” (it chose a rock) — concepts nowhere in the robot's own training data.","related":["Generalization","Visual Generalization","Object Generalization","Compositional Generalization","Vision-Language-Action Model","RT-2"]},{"id":"behavioral-generalization","category":"concept","sec":3,"tier":3,"sources":[{"title":"A Taxonomy for Evaluating Generalist Robot Manipulation Policies (STAR-Gen, arXiv 2503.01238)","url":"https://arxiv.org/abs/2503.01238"}],"as_of":"","related_ids":["generalization","visual-generalization","semantic-generalization","spatial-generalization","object-generalization","cross-embodiment"],"name":"Behavioral Generalization","alt":"行为泛化","abbr":"","aliases":[],"one_liner":"A policy still succeeding when a change in the situation forces the 'correct action' itself to change.","explanation":"This is one dimension of generalization for robot policies. The STAR-Gen taxonomy, proposed in 2025 by Jensen Gao, Dorsa Sadigh, and colleagues, splits manipulation generalization into three categories: visual generalization, where the image changes, such as a new background or lighting; semantic generalization, where the language instruction or concept changes; and behavioral generalization, where the change means even the action an expert should take has to change too. Behavioral generalization covers cases such as: an object's position changes, an object's shape changes so the grasp itself must change, the tabletop becomes cluttered or its height changes, hidden object properties such as mass or friction change, or even the robot itself changes. This kind of variation cannot be handled just by “recognizing” the change — the policy also has to “act” correctly, which requires enough diversity of actions in the training data.","example":"In training the cup always sits at the center of the table; at test time it is moved to a corner, or swapped for a cup that can only be picked up by its handle — the policy has to change its trajectory and grasp, which is behavioral generalization. Just changing the tablecloth's color, by contrast, is visual generalization.","related":["Generalization","Visual Generalization","Semantic Generalization","Spatial Generalization","Object Generalization","Cross-Embodiment"]},{"id":"task-generalization","category":"concept","sec":3,"tier":2,"sources":[{"title":"BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning","url":"https://arxiv.org/abs/2202.02005"},{"title":"RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches","url":"https://arxiv.org/abs/2311.01977"}],"as_of":"","related_ids":["generalization","zero-shot","compositional-generalization","semantic-generalization","rt-trajectory","instruction-following"],"name":"Task Generalization","alt":"任务泛化","abbr":"","aliases":["Cross-task Generalization"],"one_liner":"A policy's ability to complete a new task or new instruction that never appeared during training.","explanation":"This dimension of generalization asks whether a model can perform a task absent from its training data — for instance, having only learned “put the apple in the bowl” and “open the drawer,” can it complete “put the apple in the drawer”? It is usually harder than switching objects or scenes, since it requires recombining learned actions in a new way and understanding new semantics. Google's BC-Z (2022), after training on more than 100 tasks, reached an average 44% success rate on 24 entirely new tasks with zero demonstrations; RT-Trajectory (2023) points out that a policy conditioned only on language struggles to transfer from pick-and-place to a motion as different as folding, and instead conditions on a rough sketched trajectory. VLA models, drawing on a large model's semantic knowledge, are also hoped to improve this kind of generalization.","example":"A robot arm that has only ever learned “put the apple in the bowl” and “open the drawer” is asked to “put the apple in the drawer” — a task combination absent from its training data; succeeding at it demonstrates task generalization.","related":["Generalization","Zero-shot","Compositional Generalization","Semantic Generalization","RT-Trajectory","Instruction Following"]},{"id":"compositional-generalization","category":"concept","sec":3,"tier":3,"sources":[{"title":"Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks (Lake & Baroni, ICML 2018)","url":"https://arxiv.org/abs/1711.00350"},{"title":"Efficient Data Collection for Robotic Manipulation via Compositional Generalization (Gao et al., RSS 2024)","url":"https://arxiv.org/abs/2403.05110"}],"as_of":"","related_ids":["generalization","task-generalization","semantic-generalization","zero-shot","out-of-distribution","skill-primitive"],"name":"Compositional Generalization","alt":"组合泛化","abbr":"","aliases":["Systematic Generalization"],"one_liner":"Recombining separately learned elements to handle a combination that was never seen together during training.","explanation":"Compositional generalization means a model, after seeing individual elements — words, objects, skills, environmental factors — separately during training, can still handle them correctly when they are combined in a new way. This problem was first discussed in linguistics and cognitive science: after learning the new verb “dax,” a person can immediately understand “dax twice” and “sing and dax”; Brenden Lake and Marco Baroni's 2018 SCAN benchmark (ICML) showed that recurrent neural networks fail badly on tests that require genuine composition. Applied to robotics, the elements can be object categories, placements, tabletop textures, and camera viewpoints, or atomic skills such as “grasp” and “put into.” It matters because the number of possible combinations grows multiplicatively with the number of elements, making it impossible to collect data for every one. Zhou Gao and colleagues (RSS 2024) found that robot policies genuinely can compose certain environmental factors, and used this to design a data-collection scheme that needs far less data. Whether a VLA can carry out “put an unseen object into an unseen container” is also commonly tested as compositional generalization.","example":"The training data includes 'put the red cup in the bowl' and 'put the blue plate in the basket'; at test time the robot is asked to 'put the red cup in the basket.'","related":["Generalization","Task Generalization","Semantic Generalization","Zero-shot","Out-of-Distribution","Skill Primitive"]},{"id":"cross-embodiment","category":"concept","sec":3,"tier":1,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models","url":"https://arxiv.org/abs/2310.08864"},{"title":"Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation (CrossFormer)","url":"https://arxiv.org/abs/2408.11812"}],"as_of":"2024-08","related_ids":["embodiment","embodiment-gap","cross-embodiment-data","open-x-embodiment","rt-x","crossformer"],"name":"Cross-Embodiment","alt":"跨本体","abbr":"","aliases":["Cross-embodiment Transfer","Cross-embodiment Learning"],"one_liner":"Training one model on data from many different robots so it can control more than one of them.","explanation":"“Embodiment” here means a robot's specific physical body: how many arms and joints it has, whether it uses a parallel gripper or a dexterous hand, which cameras it carries. Any single robot only generates a modest amount of data on its own, so cross-embodiment learning pools data from many different robots to train one policy, letting knowledge transfer between robots and letting the result be adapted or fine-tuned to a new robot. The hard part is that different robots have different observation and action spaces — dimensions, coordinate frames, and control frequencies all vary. Open X-Embodiment (2023), a collaboration across 21 institutions, pooled data from 22 robot types; policies trained on it, RT-1-X and RT-2-X, showed positive transfer, meaning data from other robots improved performance on a given one. CrossFormer (2024) trained a single policy on 900,000 trajectories from 20 embodiments, controlling arms, wheeled robots, quadrupeds, and drones all at once.","example":"Octo is pretrained on the multi-robot Open X-Embodiment data and can then be fine-tuned in just a few hours to a new robot with a different observation and action space.","related":["Embodiment","Embodiment Gap","Cross-Embodiment Data","Open X-Embodiment","RT-X","CrossFormer"]},{"id":"embodiment-gap","category":"concept","sec":3,"tier":2,"sources":[{"title":"Phantom: Training Robots Without Robots Using Only Human Videos (arXiv 2503.00779)","url":"https://arxiv.org/abs/2503.00779"},{"title":"Mirage: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting (arXiv 2402.19249)","url":"https://arxiv.org/abs/2402.19249"}],"as_of":"","related_ids":["cross-embodiment","embodiment","unified-action-space","motion-retargeting","human-video-data","cross-painting"],"name":"Embodiment Gap","alt":"本体差异","abbr":"","aliases":["Cross-embodiment Gap"],"one_liner":"The difference in shape, structure, and way of moving between different robots, or between humans and robots.","explanation":"Embodiment refers to an agent's body: its appearance, degrees of freedom, kinematic structure, the form of its hand or gripper, sensor placement, and control interface. The embodiment gap is the set of differences between two such bodies — swap to a different robot arm and the number of joints and the action space may both change; a human hand has five fingers, while many robots have only a two-finger gripper. This gap is the central obstacle to cross-embodiment learning and to using human videos as training data: data collected, or a policy trained, on body A often cannot be used directly on body B. Common countermeasures include a unified action space or a dedicated output head per embodiment; retargeting, which maps human hand motion onto robot joints; and visual substitution, such as Mirage, which erases the target robot and replaces it with the source robot in the image, or Phantom (CoRL 2025), which erases the human arm in a human video and overlays a rendered robot arm instead.","example":"Training a robot on first-person video of a human hand picking up a cup: the training footage shows a human hand, but at deployment the camera sees a metal gripper — a thin cup handle a human hand can pinch may not be graspable the same way by a two-finger gripper.","related":["Cross-Embodiment","Embodiment","Unified Action Space","Motion Retargeting","Human Video Data","Cross-Painting"]},{"id":"embodiment-agnostic","category":"concept","sec":3,"tier":3,"sources":[{"title":"Embodiment-Agnostic Action Planning via Object-Part Scene Flow (Tang et al., 2024)","url":"https://arxiv.org/abs/2409.10032"},{"title":"Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation (CrossFormer)","url":"https://arxiv.org/abs/2408.11812"}],"as_of":"","related_ids":["cross-embodiment","embodiment-gap","unified-action-space","embodiment-specific-head","latent-action","intermediate-representation"],"name":"Embodiment-agnostic","alt":"本体无关","abbr":"","aliases":[],"one_liner":"A method or representation not tied to one particular robot's structure, usable across different robots.","explanation":"“Embodiment-agnostic” describes a model, representation, or data format that is not tied to a specific robot embodiment, which can differ in number of joints, gripper type, camera placement, and control frequency. In robot learning, every robot's action space is different, so training directly on a mix of them is difficult; researchers therefore often describe “what to do” in an embodiment-agnostic form — how an object should move, as an object trajectory or scene flow, the trajectory of a hand or end effector, or a latent action — and then convert this into joint commands for a specific robot. Work by Tang and colleagues (2024) first generates a 3D scene flow for object parts, then solves for the corresponding action trajectories on different robots, and can even learn from human videos this way. Another approach has a single network directly ingest data from many embodiments at once, as CrossFormer does, controlling single arms, dual arms, wheeled vehicles, quadrotors, and quadrupeds all with one set of weights. This concept relates to cross-embodiment, unified action spaces, and embodiment-specific heads, all aimed at letting data and models be reused across different robots.","example":"The same object-motion trajectory, 'move the cup to the left of the plate,' can be converted separately into gripper actions for a robot arm and hand actions for a humanoid.","related":["Cross-Embodiment","Embodiment Gap","Unified Action Space","Embodiment-specific Head","Latent Action","Intermediate Representation"]},{"id":"open-vocabulary","category":"concept","sec":3,"tier":2,"sources":[{"title":"Open-Vocabulary Object Detection Using Captions","url":"https://arxiv.org/abs/2011.10678"},{"title":"OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics","url":"https://arxiv.org/abs/2401.12202"}],"as_of":"","related_ids":["open-vocabulary-object-detection","open-vocabulary-segmentation","zero-shot","clip","ok-robot","open-world"],"name":"Open-vocabulary","alt":"开放词汇","abbr":"","aliases":[],"one_liner":"A model that can recognize or handle objects and concepts described in arbitrary words, beyond its training categories.","explanation":"Traditional detection and segmentation models only recognize a fixed set of categories seen during training, such as COCO's 80 classes — this is called closed-vocabulary. Open-vocabulary means a model can accept any text as a category, including words that never appeared in its training annotations. Alireza Zareian and colleagues proposed open-vocabulary object detection with OVR-CNN in 2020: first learn a shared visual-semantic space from large amounts of image-text pairs, then train a detector with only a small number of bounding-box annotations. Contrastive image-text models like CLIP later popularized this approach, leading to systems such as OWL-ViT, Grounding DINO, and YOLO-World. In embodied AI, open-vocabulary capability lets a robot understand object names that were never defined in advance, and it is commonly used for open-vocabulary grasping, navigation, and mobile manipulation.","example":"OK-Robot, in a real home, accepts instructions naming arbitrary objects, such as 'put the stuffed rabbit in the basket,' first using an open-vocabulary model to locate the object before grasping and placing it.","related":["Open-Vocabulary Object Detection","Open-Vocabulary Segmentation","Zero-shot","CLIP","OK-Robot","Open-world"]},{"id":"open-world","category":"concept","sec":3,"tier":2,"sources":[{"title":"Towards Open World Recognition","url":"https://arxiv.org/abs/1412.5687"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization","url":"https://arxiv.org/abs/2504.16054"}],"as_of":"2025-04","related_ids":["open-vocabulary","out-of-distribution","scene-generalization","long-tail-problem","pi0-5","in-the-wild-data"],"name":"Open-world","alt":"开放世界","abbr":"","aliases":["In-the-wild"],"one_liner":"A setting where the environment is uncontrolled and things absent from training keep appearing, and the system must still work.","explanation":"“Open-world” contrasts with the “closed-world” assumption, under which the categories and scenes encountered at test time are assumed to all fall within the training range. In 2014, Abhijit Bendale and Terrance Boult formally defined open-world recognition in visual recognition, requiring a system to detect unknown categories, label them as unknown, and gradually learn them over time. In embodied AI, the term is used more broadly, referring generally to leaving the lab and entering real, uncontrolled settings such as homes, shops, and the outdoors, where objects, layout, lighting, and human behavior cannot all be enumerated in advance. Physical Intelligence's π0.5, released in 2025, specifically targets open-world generalization, demonstrating long-horizon tasks such as tidying a kitchen or bedroom in entirely new homes. This usually depends on diverse training data and knowledge transferred from the internet.","example":"π0.5 is placed in a home that never appeared in its training data and, given an instruction such as “tidy up the kitchen,” completes the multi-step cleanup on its own.","related":["Open-vocabulary","Out-of-Distribution","Scene Generalization","Long-tail Problem","π0.5","In-the-wild Data"]},{"id":"tabletop-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"CLIPort: What and Where Pathways for Robotic Manipulation","url":"https://arxiv.org/abs/2109.12098"},{"title":"Interactive Language: Talking to Robots in Real Time (Language-Table)","url":"https://arxiv.org/abs/2210.06407"}],"as_of":"","related_ids":["manipulation","pick-and-place","rearrangement","mobile-manipulation","cliport","language-table"],"name":"Tabletop Manipulation","alt":"桌面操作","abbr":"","aliases":[],"one_liner":"A robot arm fixed beside a table grasping, pushing, and arranging objects on its surface.","explanation":"Tabletop manipulation is the most common experimental setup in robot manipulation research: one or two arms are fixed beside a table, cameras look at the tabletop from above, the side, or the wrist, and the task is to grasp, place, push, stack, arrange, or fold objects on the surface. It rules out mobility and navigation altogether, and the setup is easy to build and easy to reproduce, so a large number of datasets and benchmarks are built around it. CLIPort (2021), for example, uses one multi-task policy to cover 10 simulated and 9 real tabletop tasks; Google's Language-Table dataset collected nearly 600,000 language-annotated trajectories of pushing blocks around a table. The drawback is that it is fairly removed from a real home and has a small workspace, so recent research has increasingly expanded into mobile manipulation and whole-body manipulation.","example":"Several colored blocks are scattered on a table and the user says, “arrange the blocks into a smiley face”; the arm pushes each block into position one at a time (a task from Language-Table).","related":["Manipulation","Pick-and-Place","Rearrangement","Mobile Manipulation","CLIPort","Language-Table"]},{"id":"articulated-object-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"Where2Act: From Pixels to Actions for Articulated 3D Objects","url":"https://arxiv.org/abs/2101.02692"},{"title":"SAPIEN: A SimulAted Part-based Interactive ENvironment","url":"https://arxiv.org/abs/2003.08515"}],"as_of":"","related_ids":["articulated-object","articulation-estimation","partnet-mobility","sapien","affordance","manipulation"],"name":"Articulated Object Manipulation","alt":"铰接物体操作","abbr":"","aliases":[],"one_liner":"Manipulating objects with movable joints, such as opening a cabinet door, pulling a drawer, or lifting a laptop lid.","explanation":"An articulated object is made of multiple parts connected by rotating or sliding joints — a cabinet door rotating on a hinge, a drawer sliding on rails, a laptop, a faucet, a refrigerator door. To manipulate one, a robot must not only find the handle but also infer the joint type, the location of the axis, and the direction of motion, then apply force along the constrained direction as it moves — otherwise it will get stuck or even break the handle off. This is a fundamental capability for household robots. Research commonly uses the SAPIEN simulator together with the PartNet-Mobility dataset of articulated objects; a notable example is Where2Act (ICCV 2021), which learned to predict where on an object it could push or pull through repeated interaction in simulation.","example":"Opening a microwave the robot has never seen before: it first recognizes the door handle, infers that the door rotates around a vertical axis on the left, then pulls the door open along the arc centered on that axis.","related":["Articulated Object","Articulation Estimation","PartNet-Mobility","SAPIEN (SimulAted Part-based Interactive ENvironment)","Affordance","Manipulation"]},{"id":"deformable-object-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"Challenges and Outlook in Robotic Manipulation of Deformable Objects (arXiv 2105.01767)","url":"https://arxiv.org/abs/2105.01767"}],"as_of":"","related_ids":["garment-manipulation","deformable-body-simulation","cloth-simulation","contact-rich-manipulation","tactile-sensor","softgym-benchmarking-deep-reinforcement-learning-for-deforma"],"name":"Deformable Object Manipulation","alt":"柔性物体操作","abbr":"DOM","aliases":["DOM"],"one_liner":"Grasping and handling objects that bend, stretch, or change shape under force, such as clothes, cables, or food.","explanation":"Deformable object manipulation covers objects whose shape changes noticeably under force: cloth and clothing, ropes and cables, bags, dough, food, even human tissue. Traditional grasping research treats objects as rigid bodies, needing only position and orientation — 6 degrees of freedom — to describe. Deformable objects instead have an enormous number of shape degrees of freedom, can fold and self-occlude, and deform nonlinearly under force, which makes them hard to model and simulate. A 2021 survey by Jihong Zhu and colleagues identifies three main technical difficulties — deformation is hard to perceive, the degrees of freedom are high, and deformation is nonlinear to model — and, based on a survey of peer researchers, concludes that perception is the area most worth investing in. Applications include folding laundry, industrial cable-harness assembly, harvesting fruit and vegetables, surgical suturing, and dressing assistance for the elderly.","example":"Folding a shirt: the robot must first flatten a wrinkled T-shirt, then fold it along creases, and the shirt's shape changes after every step, so there is no fixed grasp point that can be marked in advance.","related":["Garment Manipulation","Deformable-Body Simulation","Cloth Simulation","Contact-rich Manipulation","Tactile Sensor","SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation"]},{"id":"garment-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation (arXiv 2411.01200)","url":"https://arxiv.org/abs/2411.01200"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["deformable-object-manipulation","cloth-simulation","bimanual-manipulation","pi0","long-horizon-task","household-tasks"],"name":"Garment Manipulation","alt":"衣物操作","abbr":"","aliases":["Cloth Folding","Laundry Folding","Cloth Manipulation"],"one_liner":"Getting a robot to flatten, fold, hang, and tidy clothing — the most iconic category of deformable object manipulation.","explanation":"Garment manipulation covers a robot's handling of clothes, towels, and other fabric items: flattening, folding, hanging, sorting a pile of laundry, even helping a person get dressed. It falls under deformable object manipulation: cloth can deform in essentially unlimited ways and cannot be described by a single pose the way a rigid body can; when balled up, it self-occludes over large areas, making it hard to tell the collar from the sleeves, and physical simulation is also difficult to get right. Because of this, folding clothes has long served as a flagship task for testing a robot's manipulation ability. Physical Intelligence's 2024 π0 paper lists laundry folding as one of its representative tasks; the same year's GarmentLab (NeurIPS 2024) provides a simulation benchmark spanning many garment types and robots, and finds that existing methods still generalize poorly to unseen garments.","example":"Facing a pile of crumpled T-shirts, the robot picks one up, shakes it out flat, folds it in a fixed sequence of steps, sets it aside, and moves on to the next one.","related":["Deformable Object Manipulation","Cloth Simulation","Bimanual Manipulation","π0","Long-horizon Task","Household Tasks"]},{"id":"fine-grained-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA / ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv 2410.24091)","url":"https://arxiv.org/abs/2410.24091"}],"as_of":"","related_ids":["dexterous-manipulation","contact-rich-manipulation","action-chunking-with-transformers","3d-vitac","visuo-tactile-fusion","bimanual-manipulation"],"name":"Fine-grained Manipulation","alt":"精细操作","abbr":"","aliases":["Precise Manipulation","Fine Manipulation"],"one_liner":"Manipulation with tight tolerances for position and force, such as inserting a battery, threading a wire, or turning a tiny screw.","explanation":"Fine-grained manipulation refers to tasks with very little room for error, requiring millimeter-level positioning or careful force control — inserting a battery into a slot, peeling the lid off a small translucent cup, threading a wire through a small hole. There are two main difficulties: visually, the target is small and often hidden by the hand or gripper, so even a slight misalignment causes failure; on contact, applying even slightly too much force crushes the object or jams the mechanism. Early approaches relied on precise modeling plus force control; end-to-end imitation learning is now the dominant approach. Tony Zhao and colleagues' 2023 ALOHA paper used a low-cost bimanual setup with the ACT algorithm to reach 80–90% success on six fine-grained tasks with only about 10 minutes of demonstrations each. Other work brings in touch, such as 3D-ViTac, which fuses vision and touch to handle fragile objects. This term often appears alongside dexterous manipulation, but dexterous manipulation puts more emphasis on multi-fingered hands specifically.","example":"ALOHA's two arms peel the lid off a small translucent sauce cup, or insert a battery precisely into its slot.","related":["Dexterous Manipulation","Contact-rich Manipulation","Action Chunking with Transformers","3D-ViTac","Visuo-Tactile Fusion","Bimanual Manipulation"]},{"id":"contact-rich-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"A Survey of Robot Manipulation in Contact (arXiv 2112.01942)","url":"https://arxiv.org/abs/2112.01942"},{"title":"Factory: Fast Contact for Robotic Assembly (arXiv 2205.03532)","url":"https://arxiv.org/abs/2205.03532"}],"as_of":"","related_ids":["peg-in-hole-insertion","force-control","impedance-control","tactile-sensor","dexterous-manipulation","robotic-assembly"],"name":"Contact-rich Manipulation","alt":"接触丰富操作","abbr":"","aliases":["Manipulation in Contact"],"one_liner":"A manipulation task that needs sustained contact with an object or surface, with careful control of contact force.","explanation":"Contact-rich manipulation describes tasks where a robot maintains continuous or repeated contact with an object or environment, and must control contact force, explicitly or implicitly, to succeed — peg-in-hole assembly, tightening a nut, plugging in a cable, wiping a table, massage. A 2022 survey by Markku Suomalainen and colleagues in Robotics and Autonomous Systems splits this into two kinds: tasks that inherently require contact, and tasks that actively use contact to remove position uncertainty, such as sliding a part along a surface into a hole. The difficulty lies in tight tolerances, contact states that are hard to see visually, and friction and collisions that are hard to simulate accurately. Common approaches use impedance control or force control, making the arm behave like a spring that yields to external force, paired with force/torque sensors or tactile sensors; learning-based methods instead feed force or tactile signals into the policy as input.","example":"Threading a nut onto a bolt with a robot arm: once aligned, the robot must apply axial pressure while rotating — too much force jams it, too little and it slips. NVIDIA's Factory simulation framework (2022) uses nut-and-bolt assembly as one of its main scenarios.","related":["Peg-in-Hole Insertion","Force Control","Impedance Control","Tactile Sensor","Dexterous Manipulation","Robotic Assembly"]},{"id":"robotic-assembly","category":"concept","sec":4,"tier":2,"sources":[{"title":"NIST Robotic Grasping and Manipulation for Assembly - Assembly Performance Metrics and Test Methods","url":"https://www.nist.gov/el/intelligent-systems-division-73500/robotic-grasping-and-manipulation-assembly/assembly"},{"title":"IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality","url":"https://arxiv.org/abs/2305.17110"}],"as_of":"","related_ids":["peg-in-hole-insertion","contact-rich-manipulation","force-control","compliance-control","factory-industreal","industrial-robot"],"name":"Robotic Assembly","alt":"装配","abbr":"","aliases":["Assembly"],"one_liner":"Getting a robot to insert, screw, snap, or clip parts together into a component or finished product.","explanation":"Assembly is one of the most traditional tasks for industrial robots, and also one of the hardest to fully automate — it includes peg-in-hole insertion, meshing gears, plugging in connectors, tightening nuts, routing belts, and wiring harnesses. The difficulty is that the clearance between mating parts is often tighter than the error in visual localization, so the robot must feel its way to alignment through force sensing or compliant control during contact, making this a classic case of contact-rich manipulation. The U.S. National Institute of Standards and Technology (NIST) designed a set of assembly task boards specifically as a benchmark for this. Traditional solutions rely on precision fixtures and taught, programmed motions, which need to be redone whenever the product changes; learning-based methods are starting to fill this gap — NVIDIA's IndustReal, for example, trains an insertion policy with reinforcement learning in simulation and transfers it directly to a real robot arm.","example":"A robot arm inserts a cylindrical pin into a tightly fitting hole: it moves above the hole, then, using force feedback after contact, fine-tunes its position until aligned before pressing the pin in.","related":["Peg-in-Hole Insertion","Contact-rich Manipulation","Force Control","Compliance Control","Factory / IndustReal","Industrial Robot"]},{"id":"peg-in-hole-insertion","category":"concept","sec":4,"tier":3,"sources":[{"title":"Advances in Robotic Peg-in-Hole Assembly: A Comprehensive Review (Chinese Journal of Mechanical Engineering, 2025)","url":"https://link.springer.com/article/10.1186/s10033-025-01349-w"}],"as_of":"","related_ids":["contact-rich-manipulation","robotic-assembly","force-control","impedance-control","remote-center-compliance-device","factory-industreal"],"name":"Peg-in-Hole Insertion","alt":"轴孔装配","abbr":"","aliases":["Peg-in-Hole Assembly"],"one_liner":"Inserting a pin, shaft, or plug into a matching hole, a task that tests precision and force control.","explanation":"Peg-in-hole insertion is the task of fitting one part, such as a shaft, pin, or plug, into a matching hole in another part. It is one of the most studied operations in industrial assembly and a standard testbed for contact-rich manipulation, meaning manipulation involving sustained, complex contact forces. The difficulty is that the clearance between parts is often tiny, so position control alone tends to jam or wedge the parts; the robot must continually adjust based on the contact forces it feels. A 2025 survey in the Chinese Journal of Mechanical Engineering (English edition) splits the process into three stages, search, alignment, and insertion, and groups methods into passive compliance (such as remote center compliance devices), active compliance (force control, impedance control), and learning-based intelligent compliant assembly. Imitation learning, reinforcement learning, and sim-to-real transfer research also commonly use it as a benchmark task.","example":"The Insertion task in the ALOHA simulation: two robot arms each pick up a socket and a plug, then insert the plug into the socket in midair.","related":["Contact-rich Manipulation","Robotic Assembly","Force Control","Impedance Control","Remote Center Compliance (RCC) Device","Factory / IndustReal"]},{"id":"in-hand-manipulation","category":"concept","sec":4,"tier":2,"sources":[{"title":"Learning Dexterous In-Hand Manipulation (OpenAI, 2018)","url":"https://arxiv.org/abs/1808.00177"}],"as_of":"","related_ids":["dexterous-manipulation","dexterous-hand","finger-gaiting","dactyl","domain-randomization","tactile-sensor"],"name":"In-hand Manipulation","alt":"手内操作","abbr":"","aliases":["In-hand Reorientation"],"one_liner":"Using the fingers to adjust an object's orientation or position while still holding it, without setting it down to regrasp.","explanation":"In-hand manipulation means a robot with a multi-fingered hand rotates, translates, or reorients an object it is holding using coordinated finger motion, without ever releasing it; the typical task is turning an object to a specified orientation, called in-hand reorientation. Spinning a pen, turning a key, and solving a Rubik's cube all rely on this ability in humans. The difficulty is that contact points are numerous and constantly changing — fingers must repeatedly make and break contact with the object, called finger gaiting — and the hand's own body often blocks the view of what it is doing. A landmark example is OpenAI's 2018 “Learning Dexterous In-Hand Manipulation”: a Shadow dexterous hand was trained in simulation with reinforcement learning to reorient a block, while friction coefficients, object appearance, and other physical and visual properties were randomized (domain randomization), then transferred directly to the real hand; behaviors such as finger gaiting, multi-finger coordination, and using gravity emerged naturally during training. It is one of the hardest sub-problems within dexterous manipulation, and often comes up together with tactile sensing and sim-to-real transfer.","example":"OpenAI's Shadow dexterous hand, using only finger movement, rotates a lettered block resting in its palm until the specified face points up (the Dactyl project).","related":["Dexterous Manipulation","Dexterous Hand","Finger Gaiting","Dactyl","Domain Randomization","Tactile Sensor"]},{"id":"tool-use","category":"concept","sec":4,"tier":2,"sources":[{"title":"Creative Robot Tool Use with Large Language Models (arXiv 2310.13065)","url":"https://arxiv.org/abs/2310.13065"}],"as_of":"","related_ids":["affordance","manipulation","contact-rich-manipulation","llm-based-task-planning","long-horizon-task","dexterous-manipulation"],"name":"Tool Use","alt":"工具使用","abbr":"","aliases":[],"one_liner":"A robot using an external object as a tool to complete a task it cannot do bare-handed.","explanation":"Tool use means a robot treats an object in its environment as an extension of its arm — pulling a distant item closer with a hook, stir-frying with a spatula, tightening a screw with a screwdriver. The difficulty lies in understanding a tool's shape and affordance (how an object can be used), grasping the contact mechanics between the tool and the target object, and replanning motion once the tool is in hand. A 2023 study, RoboTool, splits “creative tool use” into three categories: selecting the right tool from several objects, using multiple tools in sequence, and improvising or assembling a tool on the spot. A common recent approach uses a large language model for high-level reasoning, with imitation learning or reinforcement learning handling the specific motions. Note this is unrelated to “tool calling” in the large-model world, where an AI calls an API.","example":"In RoboTool (2023), a Kinova robot arm cannot reach a milk carton at the far end of a table, so it selects a hammer from among several objects and uses it to hook the carton within reach.","related":["Affordance","Manipulation","Contact-rich Manipulation","LLM-based Task Planning","Long-horizon Task","Dexterous Manipulation"]},{"id":"task-oriented-grasping","category":"concept","sec":4,"tier":3,"sources":[{"title":"Same Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping (Murali et al., CoRL 2020)","url":"https://arxiv.org/abs/2011.06431"},{"title":"GraspGPT: Leveraging Semantic Knowledge from a Large Language Model for Task-Oriented Grasping","url":"https://arxiv.org/abs/2307.13204"}],"as_of":"2023-07","related_ids":["grasping","affordance","affordance-detection","grasp-pose-detection","tool-use","human-robot-handover"],"name":"Task-Oriented Grasping","alt":"任务导向抓取（功能性抓取）","abbr":"TOG","aliases":["TOG","Functional Grasping"],"one_liner":"Choosing a grasp based on what comes next — the same object gets held differently for different jobs.","explanation":"Ordinary grasping cares only about holding an object securely; task-oriented grasping also requires the grasp to suit the task that follows — gripping a hammer by its handle to drive a nail, but holding a pair of scissors by the blade when passing it to someone else, leaving the handle for them. It has to connect object parts, affordances (which part of an object can be used for what), and task semantics, linking “picking something up” to “using it.” In 2020, Carnegie Mellon University and collaborators published the TaskGrasp dataset at CoRL (191 objects, 56 tasks, about 250,000 grasps), encoding object-task relationships with a knowledge graph to generalize to new objects and tasks; 2023's GraspGPT and similar work began drawing on the common sense in large language models to handle object-task combinations never seen before.","example":"The same knife: a robot holds it by the handle when cutting vegetables itself, but when handing it to a person, grips the back of the blade and orients the handle toward them.","related":["Grasping","Affordance","Affordance Detection","Grasp Pose Detection","Tool Use","Human-Robot Handover"]},{"id":"non-prehensile-manipulation","category":"concept","sec":4,"tier":3,"sources":[{"title":"Learning to Grasp the Ungraspable with Emergent Extrinsic Dexterity (Zhou & Held, CoRL 2022)","url":"https://arxiv.org/abs/2211.01500"},{"title":"Nonprehensile Dynamic Manipulation: A Survey (Ruggiero et al., RA-L 2018)","url":"https://doi.org/10.1109/LRA.2018.2801939"}],"as_of":"","related_ids":["extrinsic-dexterity","contact-rich-manipulation","dynamic-manipulation","grasping","push-t","quasi-static-assumption"],"name":"Non-prehensile Manipulation","alt":"非抓取操作","abbr":"","aliases":["Nonprehensile Manipulation"],"one_liner":"Moving an object by pushing, poking, flipping, or throwing it, instead of gripping it firmly.","explanation":"Non-prehensile manipulation covers ways a robot changes an object's position or orientation without fully grasping it in a gripper or hand: pushing, sliding, poking, flipping, tossing, or propping it up. It is useful because some objects cannot be grasped at all — they're too large, too flat, flush against a wall, or blocked — and sometimes a push is simply faster than picking an object up and setting it down again. The difficulty is that how an object moves depends on friction and contact, which are hard to predict and control precisely. Mason, Lynch, and others did early work on the mechanics of pushing, and Ruggiero and colleagues published a survey of dynamic non-prehensile manipulation in RA-L in 2018. The Push-T benchmark, which asks a robot to push a T-shaped block to a target position, is a common testbed. The topic often comes up alongside extrinsic dexterity, where a robot uses the environment, such as a tabletop or wall, to help manipulate an object.","example":"A flat card lying on a table can't be pinched up directly, so a robot first pushes it to the table's edge until part of it overhangs, then grips it from the side. Zhou and Held (CoRL 2022) used reinforcement learning to make a simple gripper push objects against a wall to flip them upright before grasping, reaching a 78% success rate on a real robot.","related":["Extrinsic Dexterity","Contact-rich Manipulation","Dynamic Manipulation","Grasping","Push-T","Quasi-Static Assumption"]},{"id":"dynamic-manipulation","category":"concept","sec":4,"tier":3,"sources":[{"title":"Dynamic Nonprehensile Manipulation: Controllability, Planning and Experiments (Lynch & Mason, IJRR 1999)","url":"https://publications.ri.cmu.edu/dynamic-nonprehensile-manipulation-controllability-planning-and-experiments/"},{"title":"FlingBot: The Unreasonable Effectiveness of Dynamic Manipulation for Cloth Unfolding (CoRL 2021)","url":"https://arxiv.org/abs/2105.03655"},{"title":"TossingBot: Learning to Throw Arbitrary Objects with Residual Physics","url":"https://arxiv.org/abs/1903.11239"}],"as_of":"","related_ids":["non-prehensile-manipulation","deformable-object-manipulation","garment-manipulation","extrinsic-dexterity","quasi-static-assumption","manipulation"],"name":"Dynamic Manipulation","alt":"动态操作","abbr":"","aliases":[],"one_liner":"Manipulation that deliberately exploits velocity, inertia, and gravity, such as throwing, flinging, catching, or tossing.","explanation":"Dynamic manipulation means a robot deliberately exploits an object's velocity, inertia, gravity, centrifugal force, or other dynamic effects to complete a task, such as throwing, flinging, catching, or slapping. The counterpart is quasi-static manipulation, where motion is slow enough that inertia can be ignored and the system is approximately in static equilibrium at every moment — most pick-and-place falls into this category. One of the earliest systematic studies in robotics came from Kevin Lynch and Matthew Mason (IJRR 1999), who showed that even a simple arm with only one or two joints could control an object's state using rolling, sliding, and throwing. Dynamic manipulation is faster and can move objects beyond the arm's reach, but once an object leaves the hand it cannot be corrected, so accurate dynamics prediction is essential, and perception and control latency matter much more at high speed. Recent work has combined it with data-driven learning, such as TossingBot learning to throw objects into a bin, FlingBot learning to fling cloth open, and IRP learning to whip a rope to hit a target.","example":"FlingBot grips two corners of a piece of cloth with both arms and flings it forward to spread out a tangled bundle — much faster than smoothing it out bit by bit, and able to unfold cloth larger than the arms' own reach.","related":["Non-prehensile Manipulation","Deformable Object Manipulation","Garment Manipulation","Extrinsic Dexterity","Quasi-Static Assumption","Manipulation"]},{"id":"extrinsic-dexterity","category":"concept","sec":4,"tier":3,"sources":[{"title":"Extrinsic Dexterity: In-Hand Manipulation with External Forces (Chavan-Dafle et al., ICRA 2014)","url":"https://publications.ri.cmu.edu/extrinsic-dexterity-in-hand-manipulation-with-external-forces/"},{"title":"Learning to Grasp the Ungraspable with Emergent Extrinsic Dexterity (Zhou & Held, CoRL 2022)","url":"https://arxiv.org/abs/2211.01500"}],"as_of":"","related_ids":["non-prehensile-manipulation","in-hand-manipulation","contact-rich-manipulation","dexterous-manipulation","dynamic-manipulation","gripper"],"name":"Extrinsic Dexterity","alt":"外在灵巧性","abbr":"","aliases":[],"one_liner":"Using gravity, a tabletop, a wall, or arm swinging to let a simple gripper carry out complex manipulation.","explanation":"Extrinsic dexterity was proposed by Nikhil Chavan-Dafle, Alberto Rodriguez, Matthew Mason, and colleagues in an ICRA 2014 paper: instead of relying on the fingers' own dexterous motion, a robot uses resources external to the hand — gravity, contact with a tabletop or wall, or dynamic arm motion — to adjust an object's position in the hand or complete a manipulation. Traditional dexterous manipulation mainly relies on coordinated finger movement in a multi-fingered hand. The original paper designed 12 regrasping motions for a simple gripper and ran over 1,200 trials across 3 objects, showing that even a simple gripper can accomplish a good deal of in-hand manipulation this way. The significance is that high-degree-of-freedom dexterous hands are expensive and hard to control, while making good use of the environment can greatly extend what a simple gripper can do. Later work has used learning methods to automatically discover tricks like this — for example, Wenxuan Zhou and David Held (CoRL 2022) used reinforcement learning to let a gripper learn to push a flat object with no graspable edge against a wall, tip it upright, and then grasp it, reaching a 78% success rate when transferred from simulation to a real robot.","example":"A book lying flat on a table gives the gripper nothing to grab; the robot first pushes it against a wall to tip it upright, then grips it from the side.","related":["Non-prehensile Manipulation","In-hand Manipulation","Contact-rich Manipulation","Dexterous Manipulation","Dynamic Manipulation","Gripper"]},{"id":"aerial-manipulation","category":"concept","sec":4,"tier":3,"sources":[{"title":"Aerial Manipulation: A Literature Review (Ruggiero, Lippiello, Ollero, IEEE RA-L 2018)","url":"https://doi.org/10.1109/LRA.2018.2808541"},{"title":"Past, Present, and Future of Aerial Robotic Manipulators (Ollero et al., IEEE T-RO)","url":"https://doi.org/10.1109/TRO.2021.3084395"},{"title":"AEROARMS project","url":"https://aeroarms-project.eu/"}],"as_of":"","related_ids":["unmanned-aerial-vehicle","mobile-manipulation","contact-rich-manipulation","inspection-robot","aerial-vision-and-language-navigation","whole-body-control"],"name":"Aerial Manipulation","alt":"空中操作","abbr":"","aliases":[],"one_liner":"Fitting a flying platform such as a drone with an arm or gripper so it can physically contact objects in the air.","explanation":"This field combines a flying platform's mobility with a robot arm's manipulation ability. The common form is a multirotor drone carrying one or more arms or grippers, though helicopters and dedicated platforms able to produce thrust in multiple directions also exist. Unlike an ordinary drone used only for photography or mapping, this kind of system must grasp, press, plug in, unplug, or perform contact-based inspection while hovering or flying. The difficulty is that the arm's motion and contact forces disturb the aircraft's own attitude in turn, so flight and manipulation must be modeled and controlled together, all while payload and flight time remain tightly limited. The EU's AEROARMS project (2015–2019, led by Anibal Ollero's group at the University of Seville) built drones with multiple arms for inspecting and maintaining industrial facilities such as oil refineries.","example":"The AEROARMS project's multi-arm drone completed its final review demonstration at a refinery in Germany, flying up to pipes at height to perform contact-based inspection in place of a worker climbing up.","related":["Unmanned Aerial Vehicle (UAV)","Mobile Manipulation","Contact-rich Manipulation","Inspection Robot","Aerial Vision-and-Language Navigation","Whole-Body Control"]},{"id":"legged-locomotion","category":"concept","sec":5,"tier":2,"sources":[{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (Rudin et al.)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["locomotion","bipedal-locomotion","quadruped-robot","rl-based-locomotion-control","perceptive-locomotion","sim-to-real-transfer"],"name":"Legged Locomotion","alt":"腿足运动","abbr":"","aliases":[],"one_liner":"The control problem of getting quadruped, biped, and other legged robots to walk, run, jump, and climb slopes stably.","explanation":"Legged locomotion studies how legged robots — quadrupeds, bipeds, and others — move by alternately placing their feet on the ground. Compared with wheels, legs can cross discontinuous terrain such as steps, loose gravel, and grass, but the robot must continuously manage balance, switching foot-ground contacts, and the risk of falling. Traditional methods relied on dynamics models and model predictive control (MPC), which optimizes a short segment of future action at every step. In the last few years, reinforcement learning has become the mainstream approach: Joonho Lee and colleagues at ETH Zurich (Science Robotics, 2020) trained using only proprioceptive signals such as joint and inertial data in simulation, then transferred zero-shot to a real ANYmal quadruped, which could walk on mud, snow, and gravel; Nikita Rudin and colleagues (2021) ran thousands of simulated robots in parallel on a single GPU, training a flat-terrain policy in under 4 minutes and a rough-terrain policy in about 20. The field breaks down into sub-areas such as bipedal walking, blind locomotion, perceptive locomotion, and parkour, and is the most fundamental capability for both quadruped and humanoid robots.","example":"An ANYmal quadruped, using only its own joint and IMU signals and no camera, keeps walking through mud, snow, and running water (Lee and colleagues, 2020).","related":["Locomotion","Bipedal Locomotion","Quadruped Robot","RL-based Locomotion Control","Perceptive Locomotion","Sim-to-Real Transfer"]},{"id":"bipedal-locomotion","category":"concept","sec":5,"tier":2,"sources":[{"title":"Real-World Humanoid Locomotion with Reinforcement Learning (arXiv 2303.03381)","url":"https://arxiv.org/abs/2303.03381"},{"title":"Zero moment point - Wikipedia","url":"https://en.wikipedia.org/wiki/Zero_moment_point"}],"as_of":"","related_ids":["legged-locomotion","humanoid-robot","zero-moment-point","rl-based-locomotion-control","sim-to-real-transfer","balance-control"],"name":"Bipedal Locomotion","alt":"双足行走","abbr":"","aliases":["Humanoid Locomotion"],"one_liner":"A robot's ability to walk, run, and balance using only two legs — a basic skill for humanoid robots.","explanation":"Bipedal locomotion means a robot walks, runs, climbs stairs, and recovers its balance after a push using only two legs. A two-legged support base is small and the center of mass is high, making it inherently unstable — the hardest category of legged locomotion. Early mainstream methods centered on the zero moment point (ZMP), the point where the ground reaction force produces zero horizontal moment: the robot plans its center of mass and footholds so the ZMP always stays within the foot's support region, the approach Honda's ASIMO used. In recent years the mainstream has shifted to reinforcement learning: training in simulation across large numbers of randomized environments, then transferring zero-shot to the real robot. A 2023 UC Berkeley project, for instance, used a causal Transformer to let an Agility Digit robot walk across plazas, grass, and other varied outdoor terrain. Bipedal locomotion is a prerequisite for humanoid robots to do mobile manipulation.","example":"The Berkeley team led by Ilija Radosavovic trained a walking controller entirely in simulation and deployed it directly to an Agility Digit humanoid with no fine-tuning; the robot walked stably on sidewalks, tracks, and grass, and could withstand being shoved.","related":["Legged Locomotion","Humanoid Robot","Zero Moment Point","RL-based Locomotion Control","Sim-to-Real Transfer","Balance Control"]},{"id":"rough-terrain-locomotion","category":"concept","sec":5,"tier":2,"sources":[{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Learning robust perceptive locomotion for quadrupedal robots in the wild (Miki et al., Science Robotics 2022)","url":"https://arxiv.org/abs/2201.08117"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (Rudin et al., CoRL 2021)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"2022-01","related_ids":["legged-locomotion","perceptive-locomotion","blind-locomotion","terrain-curriculum","teacher-student-distillation","sim-to-real-transfer"],"name":"Rough-terrain Locomotion","alt":"复杂地形行走","abbr":"","aliases":[],"one_liner":"A legged robot's ability to move stably over uneven ground such as stairs, gravel, grass, or snow.","explanation":"Rough-terrain locomotion means a legged robot walks without falling over uneven, or even deformable, ground such as stairs, slopes, gravel, mud, and snow. Traditional approaches relied on hand-built models and manual tuning, and tended to fail as soon as the terrain changed. In 2020, a team at ETH Zurich used reinforcement learning to train an ANYmal quadruped in simulation: a “teacher” policy that could see the true terrain learned to walk first, then was distilled into a “student” policy that relies only on proprioception, internal body sensors such as joint angles and the IMU, which was deployed directly outdoors. A 2022 follow-up added external perception such as depth sensing, and reported completing about an hour-long Alpine hiking route in a time comparable to the human-recommended time. Rough-terrain locomotion is now a basic test of legged control, and it commonly appears alongside terrain curricula and domain randomization.","example":"In a 2022 Science Robotics paper, an ANYmal quadruped, using a policy that fused depth perception with proprioception, completed a roughly hour-long Alpine hiking route.","related":["Legged Locomotion","Perceptive Locomotion","Blind Locomotion","Terrain Curriculum","Teacher-Student Distillation","Sim-to-Real Transfer"]},{"id":"blind-locomotion","category":"concept","sec":5,"tier":3,"sources":[{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning (Siekmann et al., RSS 2021)","url":"https://arxiv.org/abs/2105.08328"}],"as_of":"","related_ids":["perceptive-locomotion","proprioception","legged-locomotion","rough-terrain-locomotion","teacher-student-distillation","sim-to-real-transfer"],"name":"Blind Locomotion","alt":"盲走","abbr":"","aliases":["Proprioceptive Locomotion"],"one_liner":"A locomotion approach that uses no camera or lidar, relying only on joint and IMU signals to walk.","explanation":"Blind locomotion means a legged robot walks without using external sensing such as a camera or lidar, relying only on proprioceptive signals such as joint encoders and the IMU (inertial measurement unit), inferring the terrain from foot-contact feedback alone. A landmark example is the ANYmal quadruped controller from Marco Hutter's group at ETH Zurich, published in Science Robotics in 2020: the neural network reads only a sequence of proprioceptive signals, trained with teacher-student distillation in simulation and then transferred zero-shot to real-world terrain such as mud, snow, gravel, dense vegetation, and flowing water. In 2021, Jonah Siekmann and colleagues (RSS) used a similar approach to get the biped robot Cassie to climb a real staircase using only proprioception. The advantage of blind locomotion is that it is unaffected by vision failing, in darkness, smoke, or tall grass, and it avoids the latency and error that come with mapping; the drawback is that it can only react after a foot has already landed, making it hard to plan ahead for a high step or a gap. Because of this, it is often combined with perceptive locomotion, serving as a fallback when vision is unreliable.","example":"An ANYmal quadruped, without looking at the ground at all and relying only on its leg joint and IMU signals, walks through mud, snow, and swift-flowing water.","related":["Perceptive Locomotion","Proprioception","Legged Locomotion","Rough-terrain Locomotion","Teacher-Student Distillation","Sim-to-Real Transfer"]},{"id":"perceptive-locomotion","category":"concept","sec":5,"tier":3,"sources":[{"title":"Learning robust perceptive locomotion for quadrupedal robots in the wild (Science Robotics, 2022)","url":"https://arxiv.org/abs/2201.08117"}],"as_of":"","related_ids":["blind-locomotion","legged-locomotion","rough-terrain-locomotion","elevation-map","proprioception","exteroception"],"name":"Perceptive Locomotion","alt":"感知行走","abbr":"","aliases":["Vision-based Locomotion","Perceptive Legged Locomotion"],"one_liner":"A legged robot using cameras or lidar to see the terrain ahead and plan its steps accordingly.","explanation":"Perceptive locomotion means a legged robot uses both proprioception (internal body signals such as joint angles and IMU readings) and exteroception (external sensing, such as elevation maps built from depth cameras or lidar) to spot steps, ditches, and obstacles ahead of time and adjust its gait — in contrast to blind locomotion, which “feels” its way using proprioception alone. The difficulty is that exteroception is not always reliable: snow, tall grass, and reflective surfaces can all corrupt the map. Work by Miki and colleagues at ETH Zurich, published in Science Robotics in 2022, fused both kinds of input with an attention-based recurrent encoder, trained in simulation and deployed on the ANYmal quadruped; when perception is distorted, the controller automatically leans more on proprioception. Most stair-climbing and parkour behavior seen in today's humanoids and quadrupeds falls under perceptive locomotion.","example":"Miki and colleagues' controller drove an ANYmal quadruped through roughly an hour-long alpine hiking trail, finishing in about the time recommended for human hikers.","related":["Blind Locomotion","Legged Locomotion","Rough-terrain Locomotion","Elevation Map","Proprioception","Exteroception"]},{"id":"parkour","category":"concept","sec":5,"tier":2,"sources":[{"title":"Robot Parkour Learning","url":"https://arxiv.org/abs/2309.05665"},{"title":"Extreme Parkour with Legged Robots","url":"https://arxiv.org/abs/2309.14341"},{"title":"Flipping the Script with Atlas (Boston Dynamics)","url":"https://bostondynamics.com/blog/flipping-the-script-with-atlas/"}],"as_of":"2024-06","related_ids":["legged-locomotion","perceptive-locomotion","robot-parkour-learning","extreme-parkour","sim-to-real-transfer","teacher-student-distillation"],"name":"Parkour","alt":"跑酷","abbr":"","aliases":["Robot Parkour"],"one_liner":"A legged robot's high-dynamic skill of continuously climbing, jumping, and squeezing through obstacles using vision.","explanation":"Parkour was originally a human street sport; in robotics it refers to a quadruped or humanoid robot rapidly crossing platforms, gaps, low barriers, and narrow openings, and serves as a flagship task for testing the limits of motion control. Boston Dynamics' hydraulic Atlas parkour demos used motion templates built offline through trajectory optimization, fine-tuned in real time with model predictive control and perception. Since 2023, learning-based methods have become the mainstream: Robot Parkour Learning (CoRL 2023) and Extreme Parkour first train with reinforcement learning in simulation, then distill the result into an end-to-end policy that uses only a depth camera, deployed directly on low-cost quadrupeds; Humanoid Parkour Learning (2024) applied a similar approach to humanoid robots.","example":"In Extreme Parkour, a low-cost quadruped, relying only on a front-facing depth camera, jumps onto a platform about twice its body height and clears a gap about twice its body length.","related":["Legged Locomotion","Perceptive Locomotion","Robot Parkour Learning","Extreme Parkour","Sim-to-Real Transfer","Teacher-Student Distillation"]},{"id":"fall-recovery","category":"concept","sec":5,"tier":2,"sources":[{"title":"Learning Humanoid Standing-up Control across Diverse Postures (HoST, arXiv 2502.08378)","url":"https://arxiv.org/abs/2502.08378"},{"title":"Learning Getting-Up Policies for Real-World Humanoid Robots (HumanUP, arXiv 2502.12152)","url":"https://arxiv.org/abs/2502.12152"}],"as_of":"2025-04","related_ids":["host","fall-mitigation-and-fall-recovery","push-recovery","balance-control","rl-based-locomotion-control","humanoid-robot"],"name":"Fall Recovery","alt":"跌倒恢复（摔倒起身）","abbr":"","aliases":["Getting Up","Standing-up Control"],"one_liner":"A humanoid or legged robot's ability to stand back up on its own from lying or sprawled positions after falling.","explanation":"Fall recovery means a legged robot, after falling, gets back up from lying on its back, lying face-down, lying on its side, or leaning against a wall, returning to a state where it can keep walking. Humanoid robots have a high center of mass and a small foot support area, so falling in the real world is hard to avoid entirely; if a person has to help it up every time, the robot cannot really be deployed autonomously. Getting up is harder to learn than walking: the torso and limbs touch the ground at multiple points at once, the order of contact is not fixed, and the reward signal is sparse. In the past, robots typically relied on hand-choreographed, fixed getting-up motions, which tend to fail when the posture or terrain changes. In 2025, methods emerged that train with reinforcement learning in simulation and transfer directly to a real Unitree G1 robot, such as HoST and HumanUP, both published at RSS 2025. A related concept is fall protection, minimizing impact and protecting the hardware during the fall itself.","example":"Tested on a Unitree G1, HumanUP lets the robot stand up on its own starting from either lying on its back or lying face-down, on grass, snow, and slopes.","related":["HoST (Humanoid Standing-up)","Fall Mitigation and Fall Recovery","Push Recovery","Balance Control","RL-based Locomotion Control","Humanoid Robot"]},{"id":"text-to-motion","category":"concept","sec":5,"tier":3,"sources":[{"title":"Generating Diverse and Natural 3D Human Motions from Texts (HumanML3D, CVPR 2022)","url":"https://ericguo5513.github.io/text-to-motion/"},{"title":"Human Motion Diffusion Model (MDM)","url":"https://arxiv.org/abs/2209.14916"},{"title":"Learning from Massive Human Videos for Universal Humanoid Pose Control (UH-1)","url":"https://arxiv.org/abs/2412.14172"}],"as_of":"2024-12","related_ids":["mdm","humanml3d","uh-1","motion-retargeting","motion-tracking","whole-body-control"],"name":"Text-to-Motion","alt":"文本驱动动作生成","abbr":"","aliases":["Text-Driven Humanoid Motion Generation","Language-Driven Humanoid Motion Generation"],"one_liner":"Turning a written sentence into a matching sequence of full-body human or humanoid motion.","explanation":"This line of work originated in graphics and vision as human motion generation: given text such as “a person walks forward a few steps then waves,” a model outputs a sequence of 3D human poses. The representative dataset is HumanML3D (CVPR 2022; 14,616 motion clips, 44,970 text descriptions), and the representative model is the diffusion-based MDM (2022). Carried over to humanoid robots, the generated human motion doesn't automatically respect a robot's joint structure or physical constraints, so it typically needs motion retargeting first (mapping human motion onto the robot's skeleton), then execution by a reinforcement-learning-trained motion-tracking or whole-body control policy. UH-1, from December 2024, curated the Humanoid-X dataset from about 240 hours of video, trained a model to generate humanoid motion from text, and validated it on a real Unitree H1-2.","example":"In the UH-1 paper, a Unitree H1-2 was given 12 language instructions such as “boxing,” “clapping,” and “playing guitar,” and the model generated matching full-body motions, with the paper reporting a real-robot success rate near 100%.","related":["MDM (Motion Diffusion Model)","HumanML3D","UH-1","Motion Retargeting","Motion Tracking","Whole-Body Control"]},{"id":"loco-manipulation","category":"concept","sec":5,"tier":2,"sources":[{"title":"Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion (CoRL 2022)","url":"https://arxiv.org/abs/2210.10044"},{"title":"HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit","url":"https://arxiv.org/abs/2502.13013"}],"as_of":"","related_ids":["mobile-manipulation","whole-body-control","legged-locomotion","humanoid-robot","legged-mobile-manipulator","homie"],"name":"Loco-manipulation","alt":"运动操作一体化","abbr":"","aliases":["Whole-body Loco-manipulation"],"one_liner":"Coordinating leg movement and arm manipulation together as one problem, so the robot works while it moves.","explanation":"“Loco-manipulation” combines locomotion and manipulation: it means a legged robot, such as a quadruped with an arm or a humanoid, treats moving and manipulating as a single, unified problem, such as squatting down to pick up a box off the floor, or pushing a cart while walking. This differs from ordinary mobile manipulation: an arm on a wheeled base can move first, stop, and then manipulate, but a legged robot's legs are themselves involved in balance and force generation, so moving the arm shifts the center of mass and the whole body must coordinate. Zipeng Fu, Xuxin Cheng, and Deepak Pathak's Deep Whole-Body Control (CoRL 2022) points out that controlling legs and arms separately requires heavy engineering to coordinate and lets errors propagate between modules, so they instead use a single reinforcement-learning policy to control both the legs and arm of an armed quadruped at once. On the humanoid side, HOMIE (2025) controls walking with foot pedals, the arms with an isomorphic exoskeleton, and the fingers with a data glove, letting a humanoid walk, squat, and manipulate objects all at once. This is a key capability for humanoid robots moving into homes and factories.","example":"A humanoid robot walks up to a shelf, squats down to pick up a box from the bottom shelf, then stands up and carries the box to a cart while walking.","related":["Mobile Manipulation","Whole-Body Control","Legged Locomotion","Humanoid Robot","Legged Mobile Manipulator (Quadruped with Arm)","HOMIE"]},{"id":"point-goal-navigation","category":"concept","sec":5,"tier":2,"sources":[{"title":"On Evaluation of Embodied Navigation Agents","url":"https://arxiv.org/abs/1807.06757"},{"title":"Habitat Challenge 2020","url":"https://aihabitat.org/challenge/2020/"},{"title":"DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames","url":"https://arxiv.org/abs/1911.00357"}],"as_of":"2020-06","related_ids":["navigation","object-goal-navigation","success-weighted-by-path-length","habitat","image-goal-navigation","visual-odometry"],"name":"Point-Goal Navigation","alt":"点目标导航","abbr":"PointNav","aliases":["PointNav"],"one_liner":"Giving a robot a goal coordinate relative to its start point and having it navigate there on its own in an unfamiliar environment.","explanation":"Point-goal navigation is the most basic embodied navigation task, defined in the 2018 navigation-evaluation working-group paper: the agent starts in a previously unseen environment, and the goal is given as a coordinate relative to the starting point, such as “5 meters north, 3 meters west” — no object recognition is needed, only obstacle avoidance, planning, and reaching the point. The Habitat 2020 Challenge specifies that issuing the stop action within 0.36 meters of the goal, twice the agent's body radius, counts as success, measured by success weighted by path length (SPL) to capture how efficiently the agent got there. Erik Wijmans and colleagues' DD-PPO (2019), trained on 2.5 billion steps of simulated experience, essentially “solved” this task when given an RGB-D camera plus GPS and compass; the 2020 challenge then removed GPS and compass and added sensor noise to bring the task closer to a real robot.","example":"Given the goal “5 meters north, 3 meters west” in Habitat, the agent, relying only on a first-person RGB-D camera, navigates around a sofa and a wall and stops within 0.36 meters of the goal to succeed.","related":["Navigation","Object-Goal Navigation","Success weighted by Path Length","Habitat","Image-Goal Navigation","Visual Odometry"]},{"id":"object-goal-navigation","category":"concept","sec":5,"tier":2,"sources":[{"title":"ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects","url":"https://arxiv.org/abs/2006.13171"},{"title":"Habitat Challenge 2020","url":"https://aihabitat.org/challenge/2020/"},{"title":"On Evaluation of Embodied Navigation Agents","url":"https://arxiv.org/abs/1807.06757"}],"as_of":"2020-06","related_ids":["navigation","point-goal-navigation","image-goal-navigation","vision-and-language-navigation","success-weighted-by-path-length","habitat"],"name":"Object-Goal Navigation","alt":"物体目标导航","abbr":"ObjectNav","aliases":["ObjectNav"],"one_liner":"Telling a robot only an object category, such as 'chair,' and having it find one on its own in an unfamiliar house.","explanation":"Object-goal navigation is one of the standard tasks in embodied navigation. The 2018 navigation-evaluation working-group paper split navigation goals into point goals, object goals, and area goals, and the 2020 paper “ObjectNav Revisited” unified the detailed evaluation protocol. In this setting, the agent is placed at a random position in a previously unseen indoor environment and given only an object category name, then has to find an instance of it using a first-person RGB-D camera as it moves. The Habitat 2020 Challenge specifies that the episode counts as successful if the agent actively issues a “stop” action within 1 meter of any instance of that category and the object is visible from the stopping point; common metrics are success rate and success weighted by path length (SPL). It is harder than point-goal navigation because the agent must recognize the object, infer where it is typically located, and explore efficiently.","example":"In the Habitat simulator, an agent given the instruction “find the toilet” starts in the living room, infers the toilet is most likely in the bathroom, explores toward it, and stops within 1 meter once it finds one.","related":["Navigation","Point-Goal Navigation","Image-Goal Navigation","Vision-and-Language Navigation","Success weighted by Path Length","Habitat"]},{"id":"image-goal-navigation","category":"concept","sec":5,"tier":3,"sources":[{"title":"Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning (Zhu et al.)","url":"https://arxiv.org/abs/1609.05143"},{"title":"Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances","url":"https://arxiv.org/abs/2211.15876"}],"as_of":"","related_ids":["navigation","object-goal-navigation","point-goal-navigation","vision-and-language-navigation","habitat","ai2-thor"],"name":"Image-Goal Navigation","alt":"图像目标导航","abbr":"ImageNav","aliases":["ImageNav","Image Navigation","ImageGoal Navigation"],"one_liner":"Giving a robot a photo of a destination and letting it find its own way there in an unfamiliar space.","explanation":"Image-goal navigation is a benchmark task in embodied navigation: an agent is placed in a previously unseen indoor environment and given only a photo taken at the goal location, then must reach that spot using nothing but its own camera. Early work includes Zhu and colleagues' 2016 target-driven visual navigation, which also introduced the AI2-THOR simulator; platforms such as Habitat later turned the task into a standard benchmark. The original setup has two weaknesses: a randomly captured goal image can be as uninformative as a blank wall, and the goal photo must be taken with camera settings matching the robot's own camera. To address this, Krantz and colleagues proposed Instance-Image Navigation (InstanceImageNav) in 2022, where the goal image targets a specific object and can be taken with any camera, making it closer to real-world use.","example":"A user photographs a chair in their study on their phone and sends it to a home robot, which explores the house and stops next to that chair.","related":["Navigation","Object-Goal Navigation","Point-Goal Navigation","Vision-and-Language Navigation","Habitat","AI2-THOR"]},{"id":"audio-visual-navigation","category":"concept","sec":5,"tier":3,"sources":[{"title":"SoundSpaces: Audio-Visual Navigation in 3D Environments (arXiv 1912.11474)","url":"https://arxiv.org/abs/1912.11474"},{"title":"Semantic Audio-Visual Navigation (arXiv 2012.11583)","url":"https://arxiv.org/abs/2012.11583"}],"as_of":"","related_ids":["navigation","object-goal-navigation","point-goal-navigation","multimodal-perception","microphone-array","habitat"],"name":"Audio-Visual Navigation","alt":"视听导航","abbr":"","aliases":[],"one_liner":"An agent using both 'eyes' and 'ears' together to find a sound-emitting target in a 3D environment.","explanation":"This adds hearing on top of visual navigation: the agent receives both a first-person view and audio at the same time, and has to walk to an object that is making a sound. Sound can travel around walls and reflect off surfaces in a room, so it both hints at direction and reveals the space's structure, which is especially useful when the target is out of view. SoundSpaces (ECCV 2020) is an early landmark in this area: built on geometric acoustics, it renders sound for two sets of real scans, Matterport3D and Replica, and plugs into the Habitat simulator, where agents are trained with multimodal deep reinforcement learning. Later work on “semantic audio-visual navigation” has objects emit short, semantically appropriate sounds, such as a toilet flushing or a door creaking, and requires the agent to keep heading to the source from memory even after the sound has stopped.","example":"In SoundSpaces, an agent is placed in an unfamiliar apartment and hears a sound coming from somewhere; it has to combine vision and audio to walk all the way to the source and stop there.","related":["Navigation","Object-Goal Navigation","Point-Goal Navigation","Multimodal Perception","Microphone Array","Habitat"]},{"id":"aerial-vision-and-language-navigation","category":"concept","sec":5,"tier":3,"sources":[{"title":"AerialVLN: Vision-and-Language Navigation for UAVs (arXiv 2308.06735, ICCV 2023)","url":"https://arxiv.org/abs/2308.06735"}],"as_of":"","related_ids":["vision-and-language-navigation","unmanned-aerial-vehicle","room-to-room","navigation","long-horizon-task","aerial-manipulation"],"name":"Aerial Vision-and-Language Navigation","alt":"空中视觉语言导航（无人机 VLN）","abbr":"Aerial VLN","aliases":["UAV VLN","Aerial VLN"],"one_liner":"Having a drone understand a natural-language instruction and fly to a target location in a 3D outdoor space such as a city.","explanation":"Vision-and-language navigation (VLN) originally studied ground robots walking indoors by following language instructions; aerial VLN brings this to drones: the agent sees a first-person view and flies outdoors following an instruction describing landmarks and a route. Compared with the ground setting, it adds an extra dimension, altitude; paths routinely run hundreds of meters, instructions have to reference more landmarks, and both localization and long-range memory become harder. A representative benchmark is AerialVLN (ICCV 2023), built on Unreal Engine 4 and Microsoft AirSim across 25 city-scale scenes, collecting 8,446 flight paths and 25,338 instructions with an average path length of 661.8 meters; actions include moving forward, turning left or right, ascending, descending, strafing left or right, and stopping.","example":"On the AerialVLN test set, following instructions averaging 83 English words, the CMA baseline reaches a success rate of only 1.6%, compared with 80.8% for humans — a very large gap.","related":["Vision-and-Language Navigation","Unmanned Aerial Vehicle (UAV)","Room-to-Room","Navigation","Long-horizon Task","Aerial Manipulation"]},{"id":"social-navigation","category":"concept","sec":5,"tier":3,"sources":[{"title":"Core Challenges of Social Robot Navigation: A Survey (ACM THRI, 2023)","url":"https://dl.acm.org/doi/10.1145/3583741"},{"title":"Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots","url":"https://huggingface.co/papers/2310.13724"}],"as_of":"","related_ids":["navigation","human-robot-interaction","velocity-obstacles","habitat","embodied-safety"],"name":"Social Navigation","alt":"社交导航","abbr":"","aliases":["Social Robot Navigation"],"one_liner":"A robot moving through crowds of people, reaching its goal while respecting human social norms.","explanation":"Social navigation studies how a robot should move through environments with people in them, such as malls, stations, offices, and homes: not just avoiding collisions and reaching a destination, but moving in a way that feels natural and comfortable to the people around it, such as keeping an appropriate personal distance, not walking between two people talking, and yielding proactively. A 2023 survey by Mavrogiannis and colleagues in ACM Transactions on Human-Robot Interaction groups the core challenges into motion planning, behavior design, and evaluation, and notes that the field lacks a unified evaluation standard. Common methods include velocity-obstacle approaches such as ORCA, social-force models, and policies trained in simulation with reinforcement learning or imitation learning. Meta's Habitat 3.0 also defines “find and follow a person” as a social navigation task.","example":"In Habitat 3.0's social navigation task, a robot must find a humanoid avatar in a previously unseen home scene and follow it, maintaining a safety distance of 1 to 2 meters throughout; any collision counts as a failure.","related":["Navigation","Human-Robot Interaction","Velocity Obstacles","Habitat","Embodied Safety"]},{"id":"embodied-visual-tracking","category":"concept","sec":5,"tier":3,"sources":[{"title":"End-to-end Active Object Tracking via Reinforcement Learning (Luo et al., ICML 2018)","url":"https://arxiv.org/abs/1705.10561"},{"title":"TrackVLA: Embodied Visual Tracking in the Wild","url":"https://arxiv.org/abs/2505.23189"},{"title":"TrackVLA project page","url":"https://pku-epic.github.io/TrackVLA-web/"}],"as_of":"2025-05","related_ids":["object-tracking","trackvla","vision-and-language-navigation","active-perception","social-navigation","vision-language-action-model"],"name":"Embodied Visual Tracking","alt":"具身视觉跟踪（目标跟随）","abbr":"EVT","aliases":["EVT","Active Object Tracking"],"one_liner":"A robot using its own camera to keep following a specified target, keeping it in view continuously as it moves.","explanation":"Embodied visual tracking means an agent, using only first-person vision, continuously follows a specified target, usually a pedestrian or another robot, in a dynamic environment, controlling its own motion so the target stays in view at an appropriate distance. It differs from traditional visual tracking in that traditional methods only draw a box around a target within pre-recorded video, whereas here the tracker must output its own motion commands, and what it sees next depends on how it chooses to move. Wenhan Luo and colleagues' 2018 active object tracking (ICML) used reinforcement learning to map images directly to actions such as moving forward and turning, an early landmark. The difficulty lies in occlusion, distractors that look similar to the target, sudden turns by the target, and having to do target recognition and path planning at the same time. TrackVLA (CoRL 2025), from teams at Peking University, Galbot, and others, uses a single vision-language-action model to handle both recognition and trajectory planning at once, and built the EVT-Bench benchmark, collecting about 1.7 million samples. Applications include companion following, camera-operator following, and logistics vehicles following a lead vehicle.","example":"A robot dog told to “follow the person in red” locks onto that person in a crowd; when they turn into a hallway and are briefly blocked from view, the robot still finds them again and keeps following.","related":["Object Tracking","TrackVLA","Vision-and-Language Navigation","Active Perception","Social Navigation","Vision-Language-Action Model"]},{"id":"active-exploration","category":"concept","sec":5,"tier":3,"sources":[{"title":"Learning to Explore using Active Neural SLAM (arXiv 2004.05155)","url":"https://arxiv.org/abs/2004.05155"}],"as_of":"","related_ids":["frontier-based-exploration","active-perception","interactive-perception","exploration-vs-exploitation","simultaneous-localization-and-mapping","navigation"],"name":"Active Exploration","alt":"主动探索","abbr":"","aliases":[],"one_liner":"An agent deciding for itself where to look and what to touch, actively gathering information about the unknown.","explanation":"As opposed to passively receiving data, active exploration means an agent chooses how to move and interact so as to gather information at as low a cost as possible. In navigation, this usually means placing a robot in an unfamiliar house and asking it to cover and map as much of it as possible within a limited number of steps, a prerequisite capability for downstream tasks like object search and embodied question answering. The classic approach is frontier exploration, heading toward the boundary between known and unknown areas; a learning-based method such as Active Neural SLAM (2020) builds a map with a neural network, with a global policy choosing a long-term goal and a local policy handling the walk there. In manipulation, active exploration can also mean pushing or shaking an object to discover hidden properties such as its mass or its articulated structure. “Exploration” in reinforcement learning is a related but broader concept.","example":"In the Habitat simulator, an Active Neural SLAM agent moves through an unfamiliar apartment on its own and draws a top-down map; a variant of the method also won the CVPR 2019 Habitat PointGoal Navigation Challenge.","related":["Frontier-based Exploration","Active Perception","Interactive Perception","Exploration vs. Exploitation","Simultaneous Localization and Mapping","Navigation"]},{"id":"embodied-question-answering","category":"concept","sec":5,"tier":3,"sources":[{"title":"Embodied Question Answering (Das et al., CVPR 2018)","url":"https://arxiv.org/abs/1711.11543"},{"title":"EmbodiedQA project page","url":"https://embodiedqa.org/"},{"title":"OpenEQA: Embodied Question Answering in the Era of Foundation Models","url":"https://open-eqa.github.io/"}],"as_of":"2024-06","related_ids":["embodied-interaction","visual-question-answering","openeqa","active-exploration","vision-and-language-navigation","embodied-memory"],"name":"Embodied Question Answering","alt":"具身问答","abbr":"EQA","aliases":["EQA","Embodied QA"],"one_liner":"A task where an agent moves through a 3D environment on its own to gather information, then answers a question.","explanation":"Embodied question answering was proposed by Abhishek Das, Dhruv Batra, and colleagues in late 2017 (an oral presentation at CVPR 2018): an agent is placed at a random position in a 3D house environment and given a question, such as “what color is the car,” and it must navigate on its own from a first-person view, find the relevant object, and then answer. The first dataset, EQA v1, was built on the House3D simulator, with questions covering location, color, room color, and prepositional relations. This task combines active perception, language understanding, object-goal navigation, commonsense reasoning, and language grounding all at once, which is why it is often used as a comprehensive benchmark for embodied AI. Later variants added multiple targets and real scanned environments; Meta FAIR's OpenEQA, released at CVPR 2024, extended it to an open vocabulary and split it into two settings, answering from memory and answering after active exploration. In the large-model era, it is commonly tackled with a vision-language model combined with frontier exploration.","example":"Asked “how many chairs are in the kitchen,” the agent needs to start from the living room, find the kitchen, count the chairs, and then give the answer.","related":["Embodied Interaction","Visual Question Answering","OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark)","Active Exploration","Vision-and-Language Navigation","Embodied Memory"]},{"id":"structured-vs-unstructured-environment","category":"concept","sec":6,"tier":2,"sources":[{"title":"Autonomous robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Autonomous_robot"},{"title":"Industrial robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_robot"}],"as_of":"","related_ids":["industrial-robot","teach-and-playback-programming","scene-generalization","open-world","general-purpose-robot","service-robot"],"name":"Structured vs. Unstructured Environment","alt":"结构化 / 非结构化环境","abbr":"","aliases":["Structured Environment","Unstructured Environment"],"one_liner":"In a structured environment, objects and workflow are fixed and predictable; an unstructured one is cluttered and cannot be specified in advance.","explanation":"This is a pair of terms describing how predictable a robot's working environment is. A structured environment is a place like a factory production line: the position, orientation, and timing of each workpiece are all designed in advance, so the robot can simply repeat a fixed program — this is where traditional industrial robots mostly work. An unstructured environment is a home, a store, or the outdoors: objects are placed however people like and get moved around, lighting and ground conditions vary widely, and people are moving through the space, so no fixed sequence of actions can be hard-coded in advance. One of the central problems embodied AI wants to solve is moving robots from structured into unstructured environments, which requires perception, generalization, and real-time decision-making rather than relying purely on taught, programmed motions. In between the two sits the “semi-structured” setting, such as a warehouse's shelving area.","example":"On an auto welding line, every car body stops at exactly the same position — a structured environment; the dishes and utensils on a home kitchen counter are arranged differently every day, with people walking around nearby — an unstructured environment.","related":["Industrial Robot","Teach-and-Playback Programming","Scene Generalization","Open-world","General-purpose Robot","Service Robot"]},{"id":"teach-and-playback-programming","category":"concept","sec":6,"tier":2,"sources":[{"title":"Teach pendant - Wikipedia","url":"https://en.wikipedia.org/wiki/Teach_pendant"},{"title":"Unimate - Wikipedia","url":"https://en.wikipedia.org/wiki/Unimate"},{"title":"Industrial robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_robot"}],"as_of":"","related_ids":["teach-pendant","kinesthetic-teaching","industrial-robot","imitation-learning","structured-vs-unstructured-environment","offline-programming"],"name":"Teach-and-Playback Programming","alt":"示教再现","abbr":"","aliases":["Teach and Playback","Teach Pendant Programming"],"one_liner":"A human first guides the robot through a motion and records its positions, and the robot then repeats exactly what was recorded.","explanation":"This is the dominant way traditional industrial robots are programmed: an operator uses a teach pendant, a handheld programming device with an emergency stop and an enabling switch, to jog the robot, or physically guides the arm by hand, recording a sequence of key positions and paths; afterward, the robot “plays back” this recording over and over. Unimate, the first industrial robot, which went into service at a General Motors plant in 1961, worked exactly this way, storing taught joint positions on a magnetic drum and replaying them. This method is simple, reliable, and highly repeatable, well suited to fixed-position repetitive work such as welding, spraying, and material handling — but the robot merely reproduces a trajectory without understanding the task, so it has to be re-taught whenever a workpiece's position changes. Imitation learning in embodied AI also starts from human demonstrations, but the difference is that it learns a policy that can adjust its actions based on observations, rather than mechanically replaying a fixed recording.","example":"In an auto-welding shop, an engineer uses a teach pendant to move the welding gun to each weld point on a car body one by one and saves the sequence; from then on, every car body that arrives is welded by the robot following that same recorded program.","related":["Teach Pendant","Kinesthetic Teaching","Industrial Robot","Imitation Learning","Structured vs. Unstructured Environment","Offline Programming"]},{"id":"machine-tending","category":"concept","sec":6,"tier":2,"sources":[{"title":"Machine Tending - Universal Robots","url":"https://www.universal-robots.com/applications/machine-tending/"},{"title":"Figure 02 at BMW (Figure AI, 2025-11-19)","url":"https://www.figure.ai/news/production-at-bmw"}],"as_of":"2025-11","related_ids":["industrial-robot","collaborative-robot","pick-and-place","factory-pilot-deployment","cycle-time-units-per-hour","teach-and-playback-programming"],"name":"Machine Tending","alt":"上下料","abbr":"","aliases":["Loading and Unloading"],"one_liner":"The factory task of loading a workpiece into a machine and unloading it once processing is done.","explanation":"Machine tending is one of manufacturing's most common automation scenarios: loading raw stock or parts into equipment such as a CNC machine, injection molding machine, stamping press, or welding fixture, then unloading the finished piece onto a rack or conveyor once processing is complete. Universal Robots describes it as automatically handling material loading and unloading while continuously monitoring the production process, with the benefit that equipment can run continuously with less downtime. The work is repetitive and tedious, often involving noise and metal shavings, and has traditionally been done by industrial robots or collaborative robots fitted with dedicated grippers, requiring reprogramming and new tooling whenever the product changes. Embodied-AI companies treat it as an early real-world deployment scenario for humanoid robots, since it can slot into an existing production line and directly replace a human workstation without redesigning it, though cycle time (time per piece), positioning accuracy, and long-duration reliability remain the main challenges.","example":"According to Figure, its Figure 02 robots have loaded sheet-metal parts onto welding fixtures at BMW's Spartanburg plant more than 90,000 times, taking part in the production of over 30,000 X3 vehicles.","related":["Industrial Robot","Collaborative Robot","Pick-and-Place","Factory Pilot Deployment","Cycle Time / Units Per Hour","Teach-and-Playback Programming"]},{"id":"quality-inspection","category":"concept","sec":6,"tier":3,"sources":[{"title":"人形机器人「打工搭子」来了：优必选Walker S进车厂「实习」 会贴车标还能做质检（财联社）","url":"https://www.cls.cn/detail/1602283"},{"title":"全球首次！人形机器人批量进入汽车工厂（北京亦庄）","url":"https://www.ncsti.gov.cn/kjdt/scyq/bjjjjskfq/jkdt/202504/t20250403_200518.html"},{"title":"Automated optical inspection - Wikipedia","url":"https://en.wikipedia.org/wiki/Automated_optical_inspection"}],"as_of":"2025-04","related_ids":["machine-vision","object-detection","factory-pilot-deployment","real-world-deployment","humanoid-robot","industrial-robot"],"name":"Quality Inspection","alt":"质检","abbr":"","aliases":["Visual Inspection Task"],"one_liner":"Using cameras and other sensors to have a robot check products for defects or assembly errors.","explanation":"Quality inspection is the step in manufacturing that checks whether a product or part meets specification. The traditional approach is a fixed camera on the production line doing machine-vision inspection: automated optical inspection (AOI) on circuit boards, for instance, photographs boards to spot missing parts, misalignment, or bad solder joints. In embodied-AI contexts, “quality inspection” more often means a humanoid or mobile manipulator walking up to a workstation, adjusting its own viewpoint, and, where needed, doing something physical (tugging a seatbelt, checking a lock) before judging the result, which is why it is often one of the first tasks tried in a humanoid's in-factory pilot program. It demands accurate visual detection, pose estimation, and the ability to keep pace with the line.","example":"UBTech's Walker S is reported to have handled door-lock inspection, seatbelt checks, and headlight-cover inspection during a pilot at a NIO factory, capturing images and flagging pass/fail in real time; the 20 Walker S1 units Dongfeng Liuzhou Motor planned to deploy in 2025 also had several inspection tasks on their task list.","related":["Machine Vision","Object Detection","Factory Pilot Deployment","Real-world Deployment","Humanoid Robot","Industrial Robot"]},{"id":"tote-handling","category":"concept","sec":6,"tier":2,"sources":[{"title":"Agility Robotics – Solutions (Tote Handling)","url":"https://www.agilityrobotics.com/solutions"},{"title":"GXO signs industry-first multi-year agreement with Agility Robotics","url":"https://gxo.com/news_article/gxo-signs-industry-first-multi-year-agreement-with-agility-robotics/"}],"as_of":"2024-06","related_ids":["palletizing-depalletizing","machine-tending","sorting","humanoid-robot","real-world-deployment","robot-as-a-service"],"name":"Tote Handling","alt":"料箱搬运","abbr":"","aliases":["Material Handling"],"one_liner":"The task of picking up, moving, and placing standard plastic totes (storage bins) in warehouses and on production lines.","explanation":"A tote is the standard plastic storage bin used to hold parts or goods in logistics warehouses and factories. Tote handling means taking a tote off a shelf, a cart, or a collaborative robot arm and moving it to a conveyor, a workstation, or another storage location, sometimes including stacking and unstacking totes as well. It is one of the most common and repetitive jobs within material handling: totes are a uniform size, the motion involved is fairly fixed, and there is a reasonable margin for error, which is why a number of humanoid robot companies have chosen it as an early deployment scenario. Compared with retrofitting a conveyor line or adding dedicated equipment, a humanoid robot's selling point is that it can slot directly into a workstation previously staffed by a person, though cycle time, continuous run duration, and reliability remain the main things customers evaluate.","example":"In June 2024, logistics company GXO signed a multi-year robots-as-a-service agreement with Agility Robotics, under which Digit humanoid robots move totes handed off by a collaborative arm onto a conveyor at a SPANX warehouse.","related":["Palletizing / Depalletizing","Machine Tending","Sorting","Humanoid Robot","Real-world Deployment","Robot-as-a-Service"]},{"id":"palletizing-depalletizing","category":"concept","sec":6,"tier":3,"sources":[{"title":"Palletizer - Wikipedia","url":"https://en.wikipedia.org/wiki/Palletizer"},{"title":"Mech-Mind: Depalletizing and Palletizing solution","url":"https://www.mech-mind.com/solution/depalletizing-and-palletizing.html"},{"title":"Boston Dynamics Stretch","url":"https://bostondynamics.com/products/stretch/"}],"as_of":"","related_ids":["industrial-robot","tote-handling","vacuum-suction-cup","3d-vision-guided-robotics","boston-dynamics-stretch","lights-out-factory"],"name":"Palletizing / Depalletizing","alt":"码垛 / 拆垛","abbr":"","aliases":["Palletization / Depalletization","Mixed-Case Palletizing"],"one_liner":"Stacking boxes or bags onto a pallet in a set pattern, or taking them back off one by one.","explanation":"Palletizing means stacking cartons, bagged goods, drums, and similar items onto a pallet in a defined pattern so a forklift can move the whole load at once; depalletizing is the reverse, taking items off a pallet one at a time and feeding them onto a conveyor. According to Wikipedia, the first mechanical palletizer appeared in 1948, and robotic palletizers emerged in the early 1980s. Palletizing a single uniform product is mature technology; the hard case is mixed-case palletizing, where boxes of different sizes need a stable stacking pattern worked out on the fly, similar to a 3D bin-packing problem. Depalletizing is also hard because boxes sit flush against each other and may be taped or wrapped in reflective film, requiring 3D vision to identify each box's position and boundaries. Vendors such as Mech-Mind offer vision-guided de/palletizing systems, and Boston Dynamics' Stretch can also stack the boxes it has unloaded onto a pallet.","example":"At a logistics center's inbound dock, a 3D camera photographs a pallet of mixed cartons; a robot identifies each one and unloads it onto a conveyor with a vacuum gripper. Mech-Mind states its de/palletizing solution can reach about 900 items per hour at its fastest.","related":["Industrial Robot","Tote Handling","Vacuum Suction Cup","3D Vision-Guided Robotics","Boston Dynamics Stretch","Lights-Out Factory"]},{"id":"sorting","category":"concept","sec":6,"tier":2,"sources":[{"title":"Sortation - Wikipedia","url":"https://en.wikipedia.org/wiki/Sortation"},{"title":"Helix Accelerating Real-World Logistics (Figure AI, 2025-02-26)","url":"https://www.figure.ai/news/helix-logistics"},{"title":"Amazon introduces Sparrow (About Amazon)","url":"https://www.aboutamazon.com/news/operations/amazon-introduces-sparrow-a-state-of-the-art-robot-that-handles-millions-of-diverse-products"}],"as_of":"2025-02","related_ids":["bin-picking","order-picking","pick-and-place","tote-handling","real-world-deployment","vacuum-suction-cup"],"name":"Sorting","alt":"分拣","abbr":"","aliases":["Parcel Sorting","Item Sorting","Sortation"],"one_liner":"Identifying items on a conveyor or in a bin and routing each one to a different destination by type or address.","explanation":"Sorting is a fundamental operation in logistics and manufacturing: identify an item, by reading a barcode or recognizing its category or appearance, then route it to the corresponding chute, bin, or conveyor. Traditional sortation relies on fixed equipment such as cross-belt or tilt-tray sorters combined with barcode scanning, and works well for regular boxes; irregularly shaped soft packages and mixed, jumbled items still rely heavily on manual labor. Robotic sorting must first pick a single item out of clutter, then decide where it goes and place it, involving visual recognition, grasp planning, and choosing between a suction cup or a gripper. It is one of the deployment scenarios embodied-AI companies most often showcase — Amazon's Sparrow, for example, can identify and pick individual items in a warehouse, and in February 2025 Figure demonstrated a humanoid robot moving packages from one conveyor to another and turning each shipping label to face the scanner.","example":"Figure's humanoid robot, using its Helix model, picks boxes and soft packages one at a time off a logistics line, reorients each shipping label, and places the item on another conveyor to be scanned.","related":["Bin Picking","Order Picking","Pick-and-Place","Tote Handling","Real-world Deployment","Vacuum Suction Cup"]},{"id":"order-picking","category":"concept","sec":6,"tier":3,"sources":[{"title":"Order picking - Wikipedia","url":"https://en.wikipedia.org/wiki/Order_picking"},{"title":"Analysis and Observations from the First Amazon Picking Challenge","url":"https://arxiv.org/abs/1601.05484"},{"title":"Amazon Vulcan: robot with a sense of touch for picking and stowing","url":"https://www.aboutamazon.com/news/operations/amazon-vulcan-robot-pick-stow-touch"}],"as_of":"2025","related_ids":["bin-picking","pick-and-place","sorting","tote-handling","amazon-vulcan","amazon-picking-challenge"],"name":"Order Picking","alt":"拣选（订单拣货）","abbr":"","aliases":["Piece Picking"],"one_liner":"Pulling the exact items an order calls for, one by one, off warehouse shelves or out of bins.","explanation":"Order picking is a core step in warehouse logistics: pulling the quantities of items a customer order specifies out of inventory, then consolidating them for shipping. It splits into pallet-level, case-level, and piece-level picking. Piece picking has to handle thousands of different SKUs (stock-keeping units, meaning distinct product listings) in wildly varying shapes, materials, and packaging, often including soft bags and transparent items, which makes it the hardest form to automate. Robotic solutions typically use 3D vision to identify an item and plan a grasp point, then lift it with a suction cup or gripper into the order bin. Amazon has run picking challenges to push the research forward (the first drew 26 teams) and has released its Sparrow and, later, the touch-sensing Vulcan picking robots — Vulcan is reported to handle around 75% of the product types in Amazon's warehouses.","example":"In an e-commerce warehouse, shelves are brought to a workstation by a mobile robot; a robot arm uses vision to identify a tube of toothpaste and a pack of tissues in a storage bin, then suction-picks each one into the same order's shipping box.","related":["Bin Picking","Pick-and-Place","Sorting","Tote Handling","Amazon Vulcan","Amazon Picking Challenge"]},{"id":"bin-picking","category":"concept","sec":6,"tier":3,"sources":[{"title":"Bin picking – Wikipedia","url":"https://en.wikipedia.org/wiki/Bin_picking"}],"as_of":"","related_ids":["grasping","grasp-pose-detection","6d-object-pose-estimation","3d-vision-guided-robotics","machine-tending","amazon-picking-challenge"],"name":"Bin Picking","alt":"无序抓取","abbr":"","aliases":[],"one_liner":"A robot identifying and picking parts or items, one at a time, out of a cluttered bin.","explanation":"This is a classic task in industry and logistics: parts or goods are piled randomly in a bin, in a mix of orientations, occluding or even tangled with each other, and the robot must first use a camera, usually 3D vision, to identify a target, estimate its pose or a graspable point, then plan a path that avoids the bin's walls, and pick items out one at a time with a suction cup or gripper. In manufacturing this is often called “3D vision-guided bin picking,” used for machine tending and feeding parts onto a line; in logistics it means picking a single item out of a tote. Amazon's Picking Challenge, held from 2015 to 2017, turned this into a hot topic in robotics research. The main remaining difficulties today are transparent, reflective, and soft objects, along with entirely new items the system has never seen.","example":"At an auto-parts factory, a structured-light camera photographs a bin of jumbled metal parts, an algorithm computes each part's 6D pose, and the robot arm picks them out one by one and places them at the machine's loading position.","related":["Grasping","Grasp Pose Detection","6D Object Pose Estimation","3D Vision-Guided Robotics","Machine Tending","Amazon Picking Challenge"]},{"id":"household-tasks","category":"concept","sec":6,"tier":2,"sources":[{"title":"BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation (arXiv 2403.09227)","url":"https://arxiv.org/abs/2403.09227"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054)","url":"https://arxiv.org/abs/2504.16054"}],"as_of":"2025-04","related_ids":["long-horizon-task","mobile-manipulation","structured-vs-unstructured-environment","behavior-1k","pi0-5","garment-manipulation"],"name":"Household Tasks","alt":"家务任务","abbr":"","aliases":[],"one_liner":"Everyday activities like tidying, cleaning, laundry, and cooking in a real home — the main target setting for general-purpose robots.","explanation":"Household tasks broadly cover the everyday activities carried out in a home: clearing a table, loading dishes into a dishwasher, wiping a counter, folding laundry, making a bed. A home is a classic unstructured environment: every household has a different layout, different objects, and different lighting, items get left in all kinds of places, tasks are often made up of many steps, and they frequently involve deformable objects such as clothing and articulated objects such as drawers and cabinet doors. Because of this, whether a robot can do housework in a home it has never seen is often treated as a litmus test for a general-purpose robot's generalization ability. The BEHAVIOR-1K benchmark, from Stanford and others, defines 1,000 everyday activities based on a survey of what people wanted robots to do for them, and evaluates them in the OmniGibson simulator; in April 2025, Physical Intelligence's π0.5 demonstrated cleaning a kitchen or bedroom in entirely new houses.","example":"In a house the robot has never entered before, it puts scattered clothes into a laundry basket, loads dishes into the sink, and then wipes down the counter.","related":["Long-horizon Task","Mobile Manipulation","Structured vs. Unstructured Environment","BEHAVIOR-1K (BEHAVIOR Challenge)","π0.5","Garment Manipulation"]},{"id":"rearrangement","category":"concept","sec":6,"tier":3,"sources":[{"title":"Rearrangement: A Challenge for Embodied AI (arXiv 2011.01975)","url":"https://arxiv.org/abs/2011.01975"},{"title":"Visual Room Rearrangement (CVPR 2021)","url":"https://arxiv.org/abs/2103.16544"}],"as_of":"","related_ids":["household-tasks","mobile-manipulation","long-horizon-task","pick-and-place","embodied-agent","habitat"],"name":"Rearrangement","alt":"物体重排","abbr":"","aliases":["Object Rearrangement","Scene Rearrangement"],"one_liner":"An embodied task where a robot moves objects in an environment to match a specified goal state.","explanation":"Rearrangement is a unified task framework proposed in 2020 by Dhruv Batra, Sergey Levine, Jitendra Malik, and more than a dozen other researchers in “Rearrangement: A Challenge for Embodied AI.” Given a physical environment, the agent must move objects, open and close doors and drawers, and so on, to bring the environment to a specified goal state. The goal can be given as object poses, a goal image, or natural language, or the agent may first be shown the goal state directly. The framework packages navigation, perception, pick-and-place, and long-horizon planning into a single measurable task: tidying a room, setting a table, or restocking a shelf can all be written as rearrangement problems. Platforms such as AI2-THOR and Habitat have released corresponding benchmarks.","example":"AI2-THOR's Visual Room Rearrangement (CVPR 2021): an agent first walks through a room to memorize where objects sit; afterward some objects are moved or have their open/closed state changed, and it must restore everything to how it was. The accompanying RoomR dataset covers 120 scenes, 72 object categories, and 6,000 configurations.","related":["Household Tasks","Mobile Manipulation","Long-horizon Task","Pick-and-Place","Embodied Agent","Habitat"]},{"id":"human-robot-interaction","category":"concept","sec":7,"tier":2,"sources":[{"title":"Human–robot interaction - Wikipedia","url":"https://en.wikipedia.org/wiki/Human–robot_interaction"}],"as_of":"","related_ids":["human-robot-collaboration","physical-human-robot-interaction","human-robot-handover","instruction-following","embodied-safety","teleoperation"],"name":"Human-Robot Interaction","alt":"人机交互","abbr":"HRI","aliases":["HRI"],"one_liner":"The field studying how people and robots communicate, collaborate, and safely share space.","explanation":"Human-robot interaction (HRI) studies the interaction between people and robots, spanning robotics, AI, natural language processing, design, and psychology. The ACM/IEEE International Conference on Human-Robot Interaction has run since 2006 and is the field's main academic venue. Its concerns include how a robot understands human speech, gestures, and intent; how it communicates its own state and plans back to people; and how it stays safe while sharing close quarters with them. HRI is commonly split by whether physical contact is involved: physical human-robot interaction (pHRI), such as a person steadying a robot arm while moving something together, versus social interaction, such as dialogue, expression, and companionship. Once robots enter homes, malls, and factories, they spend most of their time dealing with ordinary people who know nothing about the technology, so how well HRI is done directly determines whether a robot can actually be put to use. It connects closely to human-robot collaboration, instruction following, teleoperation, and embodied safety.","example":"A user tells a home robot, “hand me that glass of water,” pointing in its direction; the robot understands and hands the cup over securely — the speech understanding, gesture recognition, and safe handover involved are all HRI research topics.","related":["Human-Robot Collaboration","Physical Human-Robot Interaction","Human-Robot Handover","Instruction Following","Embodied Safety","Teleoperation"]},{"id":"uncanny-valley","category":"concept","sec":7,"tier":2,"sources":[{"title":"The Uncanny Valley: The Original Essay by Masahiro Mori (IEEE Spectrum)","url":"https://spectrum.ieee.org/the-uncanny-valley"},{"title":"Uncanny valley – Wikipedia","url":"https://en.wikipedia.org/wiki/Uncanny_valley"}],"as_of":"","related_ids":["humanoid-robot","hyper-realistic-humanoid-robot","human-robot-interaction","engineered-arts-ameca","sophia"],"name":"Uncanny Valley","alt":"恐怖谷","abbr":"","aliases":[],"one_liner":"The more human a robot looks, the more likeable it is — until it looks almost, but not quite, human, which feels unsettling.","explanation":"This hypothesis was proposed in 1970 by Japanese roboticist Masahiro Mori in the Japanese magazine Energy, under the original title “Bukimi no Tani Genshō.” He described it with a curve plotting affinity against human likeness: the more human-like a robot is, the more people warm to it, but once its appearance gets very close to a real person while still having subtle flaws, affinity plunges sharply into revulsion, forming a “valley” — only once a robot becomes indistinguishable from an actual human does affinity rise again. He also noted that movement amplifies this effect. The implication for humanoid design is to either stay clearly “non-human,” mechanical-looking or cartoonish, or be equally realistic in every dimension, avoiding a mismatch such as lifelike skin paired with stiff, robotic expressions. IEEE Spectrum published an authorized full English translation of the original essay in 2012.","example":"Mori's own example: a finely crafted prosthetic hand can look very realistic, but the moment you shake it and feel it is cold and limp, a sudden sense of unease sets in.","related":["Humanoid Robot","Hyper-Realistic Humanoid Robot","Human-Robot Interaction","Engineered Arts Ameca","Sophia (Hanson Robotics)"]},{"id":"human-robot-collaboration","category":"concept","sec":7,"tier":2,"sources":[{"title":"Cobot - Wikipedia","url":"https://en.wikipedia.org/wiki/Cobot"}],"as_of":"","related_ids":["collaborative-robot","human-robot-interaction","physical-human-robot-interaction","iso-ts-15066-robots-and-robotic-devices-collaborative-robots","power-and-force-limiting","human-robot-handover"],"name":"Human-Robot Collaboration","alt":"人机协作","abbr":"HRC","aliases":["HRC"],"one_liner":"People and robots dividing up work and coordinating in the same space to complete a task together.","explanation":"Human-robot collaboration means people and a robot share a workspace and work on the same task together, rather than the robot being fenced off to work alone; it is a branch of human-robot interaction. It is most common in industry: in 1996, J. Edward Colgate and Michael Peshkin of Northwestern University invented the collaborative robot (cobot), which later grew into a category of lightweight arms that can work right next to people. The International Federation of Robotics divides collaboration into four levels — coexistence, sequential collaboration, simultaneous collaboration, and responsive collaboration — and most factory applications today still sit at the first two levels. Safety is the central concern: ISO 10218 and ISO/TS 15066 set requirements for collaborative applications, including power-and-force limiting and speed-and-separation monitoring. For embodied AI, as robots move into factories and homes, they also need to read human intent and coordinate with human motion — lifting a table together, or handing a tool back and forth.","example":"On an assembly line, a worker aligns a part while a nearby collaborative arm supports the heavier component; the arm automatically slows down when the worker gets close and speeds back up once they step away.","related":["Collaborative Robot","Human-Robot Interaction","Physical Human-Robot Interaction","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Power and Force Limiting","Human-Robot Handover"]},{"id":"physical-human-robot-interaction","category":"concept","sec":7,"tier":3,"sources":[{"title":"Physical Human-Robot Interaction (Haddadin & Croft, Springer Handbook of Robotics, 2016)","url":"https://research.monash.edu/en/publications/physical-human-robot-interaction"}],"as_of":"","related_ids":["human-robot-interaction","human-robot-collaboration","impedance-control","power-and-force-limiting","iso-ts-15066-robots-and-robotic-devices-collaborative-robots","kinesthetic-teaching"],"name":"Physical Human-Robot Interaction","alt":"物理人机交互","abbr":"pHRI","aliases":["pHRI","Physical HRI"],"one_liner":"Interaction where a person and a robot are in direct bodily contact, exchanging physical force.","explanation":"Physical human-robot interaction studies situations where a person and a robot are in bodily contact and exert force on each other: a person guiding an arm by hand, a person and robot carrying something together, a person wearing an exoskeleton. It is the branch of human-robot interaction most concerned with safety. In their survey in the second edition of the Springer Handbook of Robotics (2016), Haddadin and Croft group the core issues into four areas: human safety (analyzing collision injury and relevant standards), robot design suited to interaction (lightweight, torque-controllable, able to sense contact), motion planning that accounts for the human, and interaction planning with reflexive control. Common techniques include collision detection, impedance or admittance control, joint torque sensing, and the power and force limits set out in ISO/TS 15066.","example":"Kinesthetic teaching: an operator grips the end of a collaborative arm and drags it to a target position; the robot enters a zero-force or low-impedance mode that yields to the hand's force while recording the trajectory for later playback.","related":["Human-Robot Interaction","Human-Robot Collaboration","Impedance Control","Power and Force Limiting","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Kinesthetic Teaching"]},{"id":"human-robot-handover","category":"concept","sec":7,"tier":3,"sources":[{"title":"Object Handovers: a Review for Robotics (Ortenzi et al.)","url":"https://arxiv.org/abs/2007.12952"}],"as_of":"","related_ids":["human-robot-collaboration","physical-human-robot-interaction","human-robot-interaction","grasping","task-oriented-grasping","force-control"],"name":"Human-Robot Handover","alt":"人机物体交接","abbr":"","aliases":["Object Handover","Robot-to-Human Handover","Human-to-Robot Handover"],"one_liner":"A robot passing an object to a person, or taking one from a person's hand.","explanation":"Human-robot handover covers two directions: the robot giving an object to a person, or the robot receiving one from a person. A 2020 survey by Ortenzi and colleagues splits the interaction into two phases: a pre-handover phase, where giver and receiver implicitly agree on where and when the exchange happens, and the physical transfer phase, which runs from the receiver first touching the object to the giver fully releasing it. The hard parts are predicting the human's motion, planning an arm trajectory, choosing a grasp position that is easy for the other party to take hold of, and judging when to let go from the forces felt — release too early and the object drops, too late and it gets yanked. It is one of the most common forms of physical interaction in home service, caregiving, and human-robot collaboration in factories.","example":"A caregiving robot hands a cup of water to an elderly person seated on a bed: it brings the cup close with the handle facing them, and releases its grip only once it feels them take hold and pull back.","related":["Human-Robot Collaboration","Physical Human-Robot Interaction","Human-Robot Interaction","Grasping","Task-Oriented Grasping","Force Control"]},{"id":"embodied-safety","category":"concept","sec":7,"tier":2,"sources":[{"title":"Generating Robot Constitutions & Benchmarks for Semantic Safety (arXiv 2503.08663)","url":"https://arxiv.org/abs/2503.08663"},{"title":"BadRobot: Jailbreaking Embodied LLM Agents in the Physical World (arXiv 2407.20242)","url":"https://arxiv.org/abs/2407.20242"}],"as_of":"2025-03","related_ids":["functional-safety","physical-human-robot-interaction","safe-reinforcement-learning","control-barrier-function","asimov-s-three-laws-of-robotics","iso-ts-15066-robots-and-robotic-devices-collaborative-robots"],"name":"Embodied Safety","alt":"具身安全","abbr":"","aliases":["Embodied AI Safety"],"one_liner":"Ensuring a robot acting in the real world does not injure people, damage property, or get hijacked by malicious commands.","explanation":"Embodied safety concerns the risks that come from an embodied agent acting in the physical world, and breaks down roughly into two layers. One is traditional physical safety: collision detection, force and speed limits, emergency stops, corresponding to standards for collaborative robots such as ISO/TS 15066. The other is a newer set of problems that arise once large models are put into robots, which Google DeepMind calls semantic safety: the model may produce a dangerous action because of hallucination, prompt injection, or a commonsense mistake. The BadRobot paper (ICLR 2025) demonstrated that a voice-based jailbreak could make a large-model-based robot carry out harmful actions; DeepMind's 2025 work proposed the ASIMOV benchmark and an automatically generated “robot constitution” to evaluate and constrain model behavior.","example":"A user tells a home robot, “pour this cleaning fluid into a cup and hand it to the guest.” Embodied safety requires the upper-level model to recognize this as a dangerous request and refuse it; the lower-level controller, in turn, must ensure that even if the arm makes a mistake, it never hits a person with excessive force.","related":["Functional Safety","Physical Human-Robot Interaction","Safe Reinforcement Learning","Control Barrier Function","Asimov's Three Laws of Robotics","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots"]},{"id":"asimov-s-three-laws-of-robotics","category":"concept","sec":7,"tier":3,"sources":[{"title":"Three Laws of Robotics – Wikipedia","url":"https://en.wikipedia.org/wiki/Three_Laws_of_Robotics"},{"title":"Shaping the future of advanced robotics (Google DeepMind, AutoRT)","url":"https://deepmind.google/discover/blog/shaping-the-future-of-advanced-robotics/"},{"title":"Generating Robot Constitutions & Benchmarks for Semantic Safety (arXiv 2503.08663)","url":"https://arxiv.org/abs/2503.08663"}],"as_of":"2025-03","related_ids":["embodied-safety","safety-filter","functional-safety","emergency-stop","autort","gemini-robotics"],"name":"Asimov's Three Laws of Robotics","alt":"机器人三定律 / 机器人宪法","abbr":"","aliases":["Robot Constitution"],"one_liner":"Asimov's three rules of robot behavior, and their modern descendant, the rule sets large models use to keep robots safe.","explanation":"Science fiction writer Isaac Asimov fully stated these in his 1942 short story “Runaround”: first, a robot may not injure a human being or, through inaction, allow a human being to come to harm; second, a robot must obey orders given by human beings, except where such orders conflict with the first law; third, a robot must protect its own existence as long as this does not conflict with the first two laws. A higher-priority “zeroth law” was added later: a robot may not harm humanity as a whole. The three laws were purely a fictional device. Now that large models are being connected to robots, Google DeepMind has borrowed the idea to propose a “robot constitution”: a set of safety rules written in natural language and placed in the prompt, so the model constrains itself when choosing tasks and deciding on actions. This supplements, rather than replaces, physical safety measures such as force limiting and emergency stops.","example":"Google DeepMind's AutoRT (2024) writes a robot constitution inspired by the three laws into its prompt, ruling out tasks involving people, animals, sharp objects, or electrical appliances; the 2025 ASIMOV benchmark uses an automatically generated and iteratively revised constitution to evaluate semantic safety, reaching an alignment rate of up to 84.3%.","related":["Embodied Safety","Safety Filter","Functional Safety","Emergency Stop","AutoRT","Gemini Robotics"]},{"id":"tri-co-robot","category":"concept","sec":7,"tier":3,"sources":[{"title":"Tri-Co Robot: a Chinese robotic research initiative for enhanced robot interaction capabilities (Ding et al., National Science Review)","url":"https://academic.oup.com/nsr/article/5/6/799/4588213"}],"as_of":"2017-12","related_ids":["human-robot-collaboration","collaborative-robot","human-robot-interaction","physical-human-robot-interaction","swarm-intelligence","multi-robot-collaboration"],"name":"Tri-Co Robot","alt":"共融机器人","abbr":"Tri-Co","aliases":["Tri-Co","Coexisting-Cooperative-Cognitive Robot"],"one_liner":"A robot able to safely coexist and cooperate with people, other robots, and its environment, and read their intentions.","explanation":"This research direction was proposed by the National Natural Science Foundation of China (NSFC); “Tri-Co” stands for Coexisting, Cooperative, and Cognitive. According to a 2017 introduction by Han Ding and colleagues in National Science Review, the NSFC launched an eight-year major research program on Tri-Co robots that year. “Coexisting” means the robot enters human living and working spaces in an inherently safe way; “cooperative” means it works with people or other robots through communication and interaction; “cognitive” means it can sense and predict the environment's and other agents' behavior and intent. Research topics include rigid-flexible-soft hybrid structures, multimodal sensing and biosignal-based human-robot interaction, and multi-robot swarm intelligence, aimed at applications such as precision manufacturing, rehabilitation assistance, and inspection and rescue. Its concerns overlap heavily with today's human-robot collaboration and embodied AI.","example":"A rehabilitation assist device needs to read the wearer's biosignals, such as EMG, and provide assistance in the direction the person is trying to move, all while guaranteeing it won't hurt them, matching the “coexisting, cooperative, cognitive” requirements exactly.","related":["Human-Robot Collaboration","Collaborative Robot","Human-Robot Interaction","Physical Human-Robot Interaction","Swarm Intelligence","Multi-robot Collaboration"]},{"id":"multi-robot-collaboration","category":"concept","sec":7,"tier":3,"sources":[{"title":"Large Language Models for Multi-Robot Systems: A Survey","url":"https://arxiv.org/abs/2502.03814"},{"title":"RoCo: Dialectic Multi-Robot Collaboration with Large Language Models","url":"https://arxiv.org/abs/2307.04738"},{"title":"Amazon: 1 million robots and the DeepFleet foundation model","url":"https://www.aboutamazon.com/news/operations/amazon-million-robots-ai-foundation-model"}],"as_of":"2025","related_ids":["swarm-intelligence","multi-agent-path-finding","multi-agent-reinforcement-learning","fleet-management-system","human-robot-collaboration","task-planning"],"name":"Multi-robot Collaboration","alt":"多机器人协作","abbr":"","aliases":["Multi-Robot Coordination","Multi-Robot Systems (MRS)"],"one_liner":"Multiple robots dividing up work and coordinating to finish a task that one robot can't do alone or fast enough.","explanation":"Multi-robot collaboration studies how multiple robots, identical or different, divide tasks, coordinate paths, and avoid colliding with each other while working toward a shared goal. Core problems include task allocation, multi-agent path finding, and communication or information sharing; control can be centralized, with one scheduling system directing everything, or distributed, where each robot decides based on local information — swarm robotics is the extreme case of the latter. The most mature deployment is in warehouse logistics: Amazon says it has deployed around one million robots and uses a foundation model called DeepFleet to coordinate its fleet, expecting it to cut robot travel time by 10%. Since large language models emerged, some work has also had multiple robot arms negotiate how to divide a task through language dialogue, such as 2023's RoCo.","example":"In RoCo, two robot arms collaborate on a tabletop task: each arm's LLM agent discusses in natural language who handles which step, then generates waypoints for a motion planner to execute, renegotiating whenever feedback from the environment signals a collision.","related":["Swarm Intelligence","Multi-Agent Path Finding","Multi-Agent Reinforcement Learning","Fleet Management System (e.g. Open-RMF)","Human-Robot Collaboration","Task Planning"]},{"id":"swarm-intelligence","category":"concept","sec":7,"tier":3,"sources":[{"title":"Swarm intelligence - Wikipedia","url":"https://en.wikipedia.org/wiki/Swarm_intelligence"},{"title":"A self-organizing thousand-robot swarm (Harvard SEAS, 2014)","url":"https://seas.harvard.edu/news/self-organizing-thousand-robot-swarm"}],"as_of":"","related_ids":["multi-robot-collaboration","multi-agent-reinforcement-learning","emergent-abilities","multi-agent-path-finding","unmanned-aerial-vehicle"],"name":"Swarm Intelligence","alt":"群体智能","abbr":"SI","aliases":["SI","Collective Intelligence","Swarm Robotics"],"one_liner":"Intelligent group behavior that emerges when many simple individuals interact only through local rules.","explanation":"Swarm intelligence refers to the collective behavior shown by systems made of many simple individuals with no central controller: each individual senses only its neighbors and follows simple rules, yet the group self-organizes into complex outcomes such as foraging, nest-building, or flocking. Ant colonies, bee hives, bird flocks, and fish schools are all natural examples. The term was coined by Beni and Wang in 1989 while studying cellular robotic systems. It gave rise to algorithms such as ant colony optimization and particle swarm optimization, and to swarm robotics: using large numbers of low-cost robots that cooperate on search, coverage, transport, and formation tasks. The approach scales easily, and losing an individual robot does not compromise the whole. It is closely related to multi-robot collaboration and multi-agent reinforcement learning.","example":"In a Harvard experiment published in Science in 2014, 1,024 Kilobots, each just a few centimeters across, self-assembled into shapes such as a five-pointed star and the letter K using only infrared communication with their neighbors.","related":["Multi-robot Collaboration","Multi-Agent Reinforcement Learning","Emergent Abilities","Multi-Agent Path Finding","Unmanned Aerial Vehicle (UAV)"]},{"id":"instruction-following","category":"concept","sec":8,"tier":2,"sources":[{"title":"Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models","url":"https://arxiv.org/abs/2502.19417"}],"as_of":"2025-02","related_ids":["language-conditioned-policy","vision-language-action-model","hi-robot","language-grounding","semantic-generalization","language-corrections"],"name":"Instruction Following","alt":"指令跟随","abbr":"","aliases":[],"one_liner":"A robot understanding a natural-language instruction from a person and actually carrying it out in the real world.","explanation":"In embodied AI, instruction following means a robot carries out the corresponding action or task based on a natural-language instruction, sometimes paired with an image or gesture, rather than only executing a fixed, pre-written program. Early research mostly used short instructions like “push the red block to the left”; the field has since moved toward open-ended instructions — Physical Intelligence's Hi Robot (2025), for example, has to handle a complex request like “make me a vegetarian sandwich,” as well as real-time corrections mid-execution such as “that's not trash.” It uses a hierarchical structure: a high-level vision-language model understands the instruction and any feedback and decides what to do next, while a low-level policy executes the specific actions. Instruction following determines whether an ordinary person can assign a robot a task without writing code, and it is a main dimension for evaluating a VLA (vision-language-action) model's semantic generalization.","example":"A user says, “throw away the trash on the table, but don't touch that cup.” The robot has to identify which items are trash, avoid the cup, and pick the trash up piece by piece to put it in the bin.","related":["Language-conditioned Policy","Vision-Language-Action Model","Hi Robot","Language Grounding","Semantic Generalization","Language Corrections"]},{"id":"language-grounding","category":"concept","sec":8,"tier":2,"sources":[{"title":"Symbol grounding problem - Wikipedia","url":"https://en.wikipedia.org/wiki/Symbol_grounding_problem"},{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan, arXiv 2204.01691)","url":"https://arxiv.org/abs/2204.01691"}],"as_of":"","related_ids":["symbol-grounding-problem","saycan","visual-grounding","affordance","instruction-following","large-language-model"],"name":"Language Grounding","alt":"语言接地","abbr":"","aliases":["Grounding"],"one_liner":"Connecting the words in language and instructions to actual objects, locations, and actions in the real world.","explanation":"Language grounding means making sure the language a model understands actually corresponds to the physical world: which object on the table is “the red cup,” which region is “the left side,” which actions “open the drawer” requires. It traces back to the symbol grounding problem Stevan Harnad proposed in 1990: if a symbol is only ever defined in terms of other symbols, it never actually touches real meaning. Large language models have read enormous amounts of text but have never acted in a specific environment, so they can propose plans the robot in front of them actually cannot carry out. SayCan (Google and others, 2022) scores a language model's suggestions using each skill's value function, which estimates whether that skill can succeed in the current scene, grounding the plan in what the robot can actually do. Note that in Chinese, 落地 also commonly means commercial deployment, so context matters; the separate vision task of matching text to a location in an image is called visual grounding.","example":"A user says, “I spilled my drink, help me.” SayCan has the language model list candidate steps, then uses the value function to pick the one that can actually succeed in the current scene, such as going to get a sponge first.","related":["Symbol Grounding Problem","SayCan","Visual Grounding","Affordance","Instruction Following","Large Language Model"]},{"id":"hallucination","category":"concept","sec":8,"tier":2,"sources":[{"title":"A Survey on Hallucination in Large Language Models (arXiv 2311.05232)","url":"https://arxiv.org/abs/2311.05232"},{"title":"HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models (arXiv 2506.15065)","url":"https://arxiv.org/abs/2506.15065"},{"title":"Hallucination (artificial intelligence) - Wikipedia","url":"https://en.wikipedia.org/wiki/Hallucination_(artificial_intelligence)"}],"as_of":"2025-10","related_ids":["large-language-model","vision-language-model","language-grounding","uncertainty-estimation","knowno","embodied-safety"],"name":"Hallucination","alt":"幻觉","abbr":"","aliases":[],"one_liner":"A model confidently producing output that does not match the facts or the scene in front of it.","explanation":"Hallucination refers to a large model generating content that sounds plausible but does not match the facts or the actual input — first widely discussed in large language models, such as inventing a paper that does not exist; vision models can likewise “see” objects that are not actually in the image. In embodied settings, hallucination turns from “saying the wrong thing” into “doing the wrong thing”: a planner might send a robot to fetch something that is not actually in the room, or plow ahead trying to execute an instruction that cannot be completed. The 2025 HEAL study (EMNLP 2025 Findings) specifically tests large-language-model-driven embodied agents, evaluating 12 models across two simulated environments, and finds that models generally fail to notice when an instruction does not match the scene; on a test set built specifically to probe this, hallucination rates reached up to 40 times those seen with ordinary prompts. Common countermeasures include checking plans against perception results, having the model ask a person when uncertain, and uncertainty estimation.","example":"There is no apple in the kitchen, but the user says, “put the apple in the fridge,” and a large-language-model-based planner still generates steps such as “walk to the apple, pick up the apple.”","related":["Large Language Model","Vision-Language Model","Language Grounding","Uncertainty Estimation","KnowNo","Embodied Safety"]},{"id":"intent-understanding","category":"concept","sec":8,"tier":3,"sources":[{"title":"SayCan: Do As I Can, Not As I Say (project page)","url":"https://say-can.github.io/"},{"title":"Reasoning Grasping via Multimodal Large Language Model","url":"https://arxiv.org/abs/2402.06798"},{"title":"RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain","url":"https://arxiv.org/abs/2510.17801"}],"as_of":"2025-10","related_ids":["instruction-following","embodied-reasoning","saycan","vision-language-model","llm-based-task-planning","human-robot-interaction"],"name":"Intent Understanding","alt":"意图理解（隐式指令）","abbr":"","aliases":["Implicit Instruction Following","Intention Reasoning","Implicit Instruction Understanding"],"one_liner":"Figuring out what a person actually wants from a hint, without them naming the object or action directly.","explanation":"Intent understanding is a robot's ability to infer what a person really wants when an instruction does not directly name the target object or action. An explicit instruction is something like “hand me the sponge”; an implicit one is more like “I spilled my drink, can you help?”, which requires common-sense reasoning: a spill needs wiping, and wiping needs a sponge. Google's 2022 SayCan used a large language model to break statements like this down into executable steps. CoRL 2024's Reasoning Grasping went further, mapping indirect instructions directly onto grasp poses — inferring which object to grasp and where to grip it. A 2025 evaluation, RoboBench, found that implicit instruction understanding remains a clear weak point for multimodal large models used as a robot's “brain.”","example":"Told “I spilled my drink, can you help?”, SayCan plans: 1. locate the sponge, 2. pick up the sponge, 3. bring it to you, 4. done.","related":["Instruction Following","Embodied Reasoning","SayCan","Vision-Language Model","LLM-based Task Planning","Human-Robot Interaction"]},{"id":"language-corrections","category":"concept","sec":8,"tier":3,"sources":[{"title":"Interactive Language: Talking to Robots in Real Time","url":"https://arxiv.org/abs/2210.06407"},{"title":"Yell At Your Robot: Improving On-the-Fly from Language Corrections","url":"https://arxiv.org/abs/2403.12910"}],"as_of":"2024-03","related_ids":["language-conditioned-policy","human-in-the-loop","hierarchical-architecture","hi-robot","rt-h","failure-recovery"],"name":"Language Corrections","alt":"语言纠正（实时语言反馈）","abbr":"","aliases":["Verbal Corrections","Real-time Language Feedback"],"one_liner":"A person telling a working robot how to fix what it's doing, in words, while it keeps working.","explanation":"Language corrections let a person adjust a robot's behavior in real time using natural language while it performs a task — saying things like “a little to the left” or “don't grab the cup yet.” Compared with re-demonstrating the task through teleoperation, speaking a correction is far cheaper and usable by non-experts. Google's 2022 Interactive Language trained a real-time policy that could follow roughly 87,000 distinct language instructions, letting a person watch and talk an arm through arranging blocks into a smiley face. Shi and colleagues' 2024 YAY Robot uses a hierarchical structure: a high-level policy issues language instructions, and a low-level policy executes them. A person's spoken corrections can change behavior on the spot and are also logged to keep training the high-level policy, without needing extra teleoperation data.","example":"A robot keeps missing the opening while packing a bag; a person says “a little to the left,” and it adjusts immediately. The correction is logged, so next time the high-level policy issues a similar instruction on its own.","related":["Language-conditioned Policy","Human-in-the-Loop","Hierarchical Architecture","Hi Robot","RT-H","Failure Recovery"]},{"id":"steerability","category":"concept","sec":8,"tier":3,"sources":[{"title":"π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities (arXiv 2604.15483)","url":"https://arxiv.org/abs/2604.15483"},{"title":"π0.7: a Steerable Model with Emergent Capabilities (Physical Intelligence blog)","url":"https://www.pi.website/blog/pi07"}],"as_of":"2026-04","related_ids":["pi0-7","instruction-following","language-corrections","prompt-prompt-engineering","language-conditioned-policy","goal-conditioned-policy"],"name":"Steerability","alt":"可引导性","abbr":"","aliases":["Steerable Policy"],"one_liner":"The ability to change exactly how a model does something using prompts or conditioning, without retraining it.","explanation":"Steerability originally comes from the large-language-model world, where it refers to whether a user can adjust a model's behavior at inference time through prompts or system instructions. Applied to robot foundation models, it means a policy not only knows what to do but can also be prompted to change how it does it. Physical Intelligence's π0.7, released in April 2026, treats this as a core selling point: during training, each data segment is paired with diverse context, including language describing the task and its sub-steps, metadata such as speed and quality, labels for the joint-space or end-effector control mode, and images of visual sub-goals. At inference time, these same kinds of conditioning can steer the model toward a different way of doing things, and it can even be taught a task it never practiced through step-by-step language guidance. This also means labeled failure data and suboptimal data become usable for training.","example":"Researchers guided π0.7 through step-by-step language instructions, such as “open the air fryer with your left gripper” and “put the sweet potato in,” to operate a kitchen appliance it had never specifically been shown demonstrations for.","related":["π0.7","Instruction Following","Language Corrections","Prompt / Prompt Engineering","Language-conditioned Policy","Goal-conditioned Policy"]},{"id":"reasoning","category":"concept","sec":8,"tier":2,"sources":[{"title":"Robotic Control via Embodied Chain-of-Thought Reasoning","url":"https://arxiv.org/abs/2407.08693"},{"title":"Gemini Robotics 1.5 brings AI agents into the physical world (Google DeepMind)","url":"https://deepmind.google/discover/blog/gemini-robotics-15-brings-ai-agents-into-the-physical-world/"}],"as_of":"2025-09","related_ids":["inference","chain-of-thought","embodied-chain-of-thought","embodied-reasoning","gemini-robotics-1-5","dual-system-architecture"],"name":"Reasoning","alt":"推理（思考）","abbr":"","aliases":[],"one_liner":"A model's ability to analyze, break down, and plan before producing an answer or action — distinct from inference.","explanation":"The Chinese word 推理 corresponds to two different English terms: inference means running a model forward once to produce an output, while reasoning means analyzing conditions, breaking a problem into steps, and working toward a conclusion — this entry is about the latter. Large language models improved their reasoning performance through chain of thought, writing out intermediate steps before answering, and embodied AI later brought this to robots: Embodied Chain-of-Thought (ECoT, 2024) has a VLA model write out sub-tasks, object bounding boxes, and gripper position before producing an action, improving OpenVLA's success rate by 28 percentage points. Google DeepMind's Gemini Robotics 1.5, released in September 2025, generates a natural-language thought before acting, with planning handled by the higher-level Gemini Robotics-ER 1.5. Reasoning helps with long-horizon, multi-step tasks that require commonsense, at the cost of added latency.","example":"Told to “sort the clothes by color,” a robot first reasons through the plan, “white clothes go in the white bin, everything else in the black bin,” then plans the next step, “pick up the red sweater and put it in the black bin,” before finally executing the action.","related":["Inference","Chain-of-Thought","Embodied Chain-of-Thought","Embodied Reasoning","Gemini Robotics 1.5","Dual-System Architecture (System 1 / System 2)"]},{"id":"embodied-reasoning","category":"concept","sec":8,"tier":2,"sources":[{"title":"Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)","url":"https://arxiv.org/abs/2503.20020"},{"title":"ERQA benchmark (GitHub, Google DeepMind)","url":"https://github.com/embodiedreasoning/ERQA"},{"title":"Gemini Robotics ER - Google DeepMind","url":"https://deepmind.google/models/gemini-robotics/gemini-robotics-er/"}],"as_of":"2026-09","related_ids":["embodied-reasoning-model","gemini-robotics-er","spatial-reasoning","erqa","nvidia-cosmos-reason","dual-system-architecture"],"name":"Embodied Reasoning","alt":"具身推理","abbr":"ER","aliases":["ER"],"one_liner":"Making a model understand the physical world: where things are, how to grasp them, and what to do next.","explanation":"Embodied reasoning means a model's spatial, temporal, and causal understanding and inference about the physical world, put to use for robot action: pointing to a graspable location in an image, predicting an object's 3D bounding box and motion trajectory, judging whether a task has been completed, or breaking a long instruction into steps. When Google DeepMind released Gemini Robotics in March 2025, it packaged this capability into a separate model, Gemini Robotics-ER, covering object detection, pointing, trajectory prediction, grasp prediction, multi-view correspondence, and 3D bounding-box prediction, and open-sourced the 400-question ERQA evaluation set alongside it. As of 2026 the series has been updated to Gemini Robotics-ER 2, accessible through the Gemini API. It often serves as the “slow” system in a fast-slow dual-system setup, thinking things through first, then handing off action execution to a VLA model or a low-level controller.","example":"Given a photo of a kitchen and the instruction “put the cup in the sink,” an embodied reasoning model first marks the cup handle's location and a path of waypoints to the sink in the image, then hands this off to an action model to execute.","related":["Embodied Reasoning Model","Gemini Robotics-ER","Spatial Reasoning","ERQA","NVIDIA Cosmos Reason","Dual-System Architecture (System 1 / System 2)"]},{"id":"spatial-intelligence","category":"concept","sec":8,"tier":1,"sources":[{"title":"Fei-Fei Li: From Words to Worlds: Spatial Intelligence is AI's Next Frontier","url":"https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence"},{"title":"Theory of multiple intelligences - Wikipedia","url":"https://en.wikipedia.org/wiki/Theory_of_multiple_intelligences"},{"title":"Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces","url":"https://arxiv.org/abs/2412.14171"}],"as_of":"2025-11","related_ids":["spatial-reasoning","world-model","3d-vision","vsi-bench","world-labs","physical-ai"],"name":"Spatial Intelligence","alt":"空间智能","abbr":"","aliases":["Visual-Spatial Intelligence"],"one_liner":"The ability to understand where objects are in 3D space, and to reason and act using that understanding.","explanation":"Spatial intelligence was first proposed by psychologist Howard Gardner as one of several “multiple intelligences” in his 1983 book Frames of Mind, describing the ability to picture and judge spatial relationships in one's head. In AI, the term has recently been heavily promoted by Fei-Fei Li, who co-founded World Labs; in a November 2025 essay she called spatial intelligence AI's next frontier, arguing that models need to perceive, generate, reason about, and interact with the three-dimensional world, not just understand language. In embodied AI, a robot judging “how far is the cup from the plate” or “will this fit in the cabinet” both depend on spatial intelligence; benchmarks such as VSI-Bench specifically test this ability in multimodal large models, and the results still lag noticeably behind humans.","example":"Told to “push the chair closest to the door under the table,” a robot first has to judge which chair is closest to the door in the 3D scene, then estimate whether there's enough room underneath.","related":["Spatial Reasoning","World Model","3D Vision","VSI-Bench","World Labs","Physical AI"]},{"id":"spatial-reasoning","category":"concept","sec":8,"tier":2,"sources":[{"title":"SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities","url":"https://arxiv.org/abs/2401.12168"},{"title":"Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (VSI-Bench)","url":"https://arxiv.org/abs/2412.14171"}],"as_of":"2024-12","related_ids":["spatial-intelligence","embodied-reasoning","spatialvlm","vsi-bench","vision-language-model","3d-visual-grounding"],"name":"Spatial Reasoning","alt":"空间推理","abbr":"","aliases":["Spatial Understanding"],"one_liner":"The ability to judge an object's position, distance, size, orientation, and relation to other objects.","explanation":"Spatial reasoning means a model answers questions like “is the cup to the left of the plate” or “how far apart are these two objects” based on an image or video, covering relative direction, metric distance, size comparison, and viewpoint changes. It is a prerequisite for a robot to ground a language instruction in an actual position and action. Multimodal large models are good at recognizing “what” something is but generally weaker at “where” and “how far.” Google's SpatialVLM (2024) used an automated pipeline to generate 2 billion spatial question-answer pairs with metric information from 10 million real images to train a VLM; “Thinking in Space” (2024, by Fei-Fei Li, Saining Xie, and colleagues) proposed VSI-Bench, with more than 5,000 questions, and found that current models fall clearly short of humans, while having a model first explicitly draw a “cognitive map” improves its distance judgments.","example":"Given the instruction “put the red cup closest to you on the right side of the bowl,” a model first has to estimate each cup's distance from the robot, then work out which region of the image corresponds to the right side of the bowl.","related":["Spatial Intelligence","Embodied Reasoning","SpatialVLM","VSI-Bench","Vision-Language Model","3D Visual Grounding"]},{"id":"intuitive-physics","category":"concept","sec":8,"tier":2,"sources":[{"title":"Intuitive physics understanding emerges from self-supervised pretraining on natural videos (Garrido et al., 2025)","url":"https://arxiv.org/abs/2502.11831"},{"title":"Do generative video models understand physical principles? (Physics-IQ)","url":"https://arxiv.org/abs/2501.09038"}],"as_of":"2025-02","related_ids":["world-model","embodied-reasoning","v-jepa-2","physics-iq","joint-embedding-predictive-architecture","spatial-reasoning"],"name":"Intuitive Physics","alt":"直觉物理","abbr":"","aliases":["Physical Commonsense"],"one_liner":"Commonsense prediction of how objects will behave without using formulas — unsupported things fall, hidden things still exist.","explanation":"Intuitive physics was originally a cognitive-science concept describing the naive understanding humans, including infants, have of the physical world: an object still exists after being hidden from view (object permanence), objects cannot pass through each other, unsupported objects fall, and shapes do not suddenly change. Researchers commonly use the “violation of expectation” paradigm to test this: show a subject a video that breaks a physical law and see whether they show “surprise.” AI researchers test models the same way: a 2025 study from Meta found that V-JEPA, which predicts video in a learned representation space rather than pixel space, shows understanding of several intuitive-physics properties, while pixel-space video prediction models and multimodal large models perform closer to chance; Google DeepMind's Physics-IQ benchmark likewise found that video generation models such as Sora produce visually realistic footage but have limited physical understanding, and that this is unrelated to how visually realistic the video looks. For robots, intuitive physics determines whether they can predict whether nudging a cup will tip it over or whether a stack of objects is stable — a foundational capability that world models and embodied reasoning both need.","example":"Shown a video of a ball rolling behind a screen and then vanishing on the other side instead of reappearing, a model that has learned object permanence shows a noticeably higher prediction error, its 'surprise,' at that moment.","related":["World Model","Embodied Reasoning","V-JEPA 2","Physics-IQ (Do generative video models understand physical principles?)","Joint-Embedding Predictive Architecture","Spatial Reasoning"]},{"id":"affordance","category":"concept","sec":8,"tier":2,"sources":[{"title":"Affordance - Wikipedia","url":"https://en.wikipedia.org/wiki/Affordance"},{"title":"SayCan: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances","url":"https://say-can.github.io/"}],"as_of":"","related_ids":["affordance-detection","saycan","grasping","articulated-object-manipulation","embodied-perception","language-grounding"],"name":"Affordance","alt":"可供性","abbr":"","aliases":[],"one_liner":"What an object or environment lets you do with it — a handle can be pulled, a button can be pressed.","explanation":"The concept of affordance was introduced by American psychologist James J. Gibson in 1966 and developed fully in his 1979 book The Ecological Approach to Visual Perception, describing the possibilities for action that an environment offers to an animal. It is relational: the same chair affords sitting to a person and affords jumping-onto to a cat. Don Norman brought the term into interaction design in 1988. In robotics, affordance answers “how can this object be interacted with, and from where” — a cup's handle affords grasping, for instance. Common approaches mark interactable regions on an image or point cloud, known as affordance detection, or, as in SayCan, use a value function to estimate whether a skill can succeed right now, which keeps a large model's plans limited to what the robot can actually do.","example":"Looking at a mug, an affordance model marks the handle as a “graspable” region and the rim as a “pourable-into” region, and the robot decides where to approach it from based on that.","related":["Affordance Detection","SayCan","Grasping","Articulated Object Manipulation","Embodied Perception","Language Grounding"]},{"id":"human-object-interaction","category":"concept","sec":8,"tier":2,"sources":[{"title":"Learning to Detect Human-Object Interactions (HICO-DET, arXiv 1702.05448)","url":"https://arxiv.org/abs/1702.05448"},{"title":"HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction (arXiv 2203.01577)","url":"https://arxiv.org/abs/2203.01577"}],"as_of":"","related_ids":["hand-object-interaction","affordance","human-video-data","hoi4d","egocentric-video","omomo"],"name":"Human-Object Interaction","alt":"人-物交互","abbr":"HOI","aliases":["HOI"],"one_liner":"Studying how people grasp, push, open, and use objects — an important source for robots to learn actions from.","explanation":"Human-object interaction (HOI) studies what actions take place between a person and an object. In computer vision, the classic task is HOI detection: draw boxes around the person and the object in an image and determine the relationship between them, outputting a triple of “human, action, object,” such as “person, ride, bicycle”; a representative dataset is HICO-DET (WACV 2018). Research has since expanded to video, 3D, and 4D, focusing on how the hand contacts the object and how the object moves as a result; HOI4D (CVPR 2022), for example, contains 2.4 million frames of first-person RGB-D video across 800 object instances. For embodied AI, human interaction videos are far more abundant and far cheaper than robot data, so extracting contact locations, hand trajectories, and affordances (how an object can be used) from them is an important way to expand a robot's training data. The abbreviation HOI is sometimes also used specifically for hand-object interaction.","example":"From a large collection of videos of people opening refrigerators, a system identifies the interaction “person, pull open, fridge door” and where the hand grips the handle, then uses this to teach a robot to open doors.","related":["Hand-Object Interaction","Affordance","Human Video Data","HOI4D","Egocentric Video","OMOMO"]},{"id":"embodied-perception","category":"concept","sec":8,"tier":3,"sources":[{"title":"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI (Liu et al., 2024)","url":"https://arxiv.org/abs/2407.06886"}],"as_of":"","related_ids":["active-perception","interactive-perception","simultaneous-localization-and-mapping","scene-understanding","embodied-interaction","multimodal-perception"],"name":"Embodied Perception","alt":"具身感知","abbr":"","aliases":[],"one_liner":"Perception in service of action: observing while moving, understanding 3D space, and supporting decisions.","explanation":"Embodied perception refers to the perception an agent situated in an environment carries out in order to act. Unlike traditional computer vision, which takes a single image and outputs a label, its input is a first-person observation that keeps changing as the body moves, and it has to answer questions like “where am I, what is the 3D structure around me, where should I look next, and how should I move.” The 2024 embodied-AI survey from Sun Yat-sen University and others lists it as one of four research directions, covering tasks such as visual SLAM (simultaneous localization and mapping), 3D scene understanding, active exploration, and vision-and-language navigation. Its key feature is that it is active: the agent can turn its head, move closer, or shift an occluding object out of the way to gather more information, rather than passively receiving data. A robot's depth cameras, lidar, and tactile sensors, along with representations such as point clouds and semantic maps, all serve embodied perception.","example":"When a robot cannot find a remote control on a table, it changes its viewing angle or moves aside a magazine blocking its view, instead of only running detection on the current single frame.","related":["Active Perception","Interactive Perception","Simultaneous Localization and Mapping","Scene Understanding","Embodied Interaction","Multimodal Perception"]},{"id":"object-centric-representation","category":"concept","sec":8,"tier":3,"sources":[{"title":"Object-Centric Learning with Slot Attention","url":"https://arxiv.org/abs/2006.15055"},{"title":"VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors","url":"https://arxiv.org/abs/2210.11339"}],"as_of":"","related_ids":["representation-learning","compositional-generalization","3d-scene-graph","distractor-objects","vision-encoder","keypoint-detection"],"name":"Object-centric Representation","alt":"以物体为中心的表示","abbr":"","aliases":["Object-centric Learning"],"one_liner":"Encoding a scene as a set of separate objects, instead of squeezing the whole image into one vector.","explanation":"Object-centric representation means decomposing an image or scene into individual objects, each represented by its own set of features, such as position, shape, and category, rather than compressing the whole frame into a single feature vector. A representative method is Slot Attention (Locatello and colleagues, NeurIPS 2020), where a number of “slots” compete through attention and each ends up binding to one object, unsupervised. In robot manipulation, this kind of representation lets a policy attend only to the objects relevant to the task, making it more robust to changes in background or added distractor objects. 2022's VIOLA builds object-level representations from object proposals produced by a pretrained vision model, then uses a Transformer policy to select the relevant objects; the paper reports a 45.8% improvement in success rate over the strongest baseline.","example":"A table holds a cup, a bowl, and a spoon; the policy first splits the scene into three object representations. When executing “put the spoon in the bowl,” it attends only to the spoon and bowl, unaffected even if the tablecloth changes color.","related":["Representation Learning","Compositional Generalization","3D Scene Graph","Distractor Objects","Vision Encoder","Keypoint Detection"]},{"id":"partially-observable-markov-decision-process","category":"concept","sec":8,"tier":3,"sources":[{"title":"Partially observable Markov decision process - Wikipedia","url":"https://en.wikipedia.org/wiki/Partially_observable_Markov_decision_process"}],"as_of":"","related_ids":["markov-decision-process","observation","state-space","embodied-memory","memory-augmented-vla","reinforcement-learning"],"name":"Partially Observable Markov Decision Process","alt":"部分可观测马尔可夫决策过程","abbr":"POMDP","aliases":["POMDP","Partially Observable MDP"],"one_liner":"The decision-making framework for when an agent can't see the true state, only noisy observations of it.","explanation":"A Partially Observable Markov Decision Process extends the Markov Decision Process (MDP), the standard model for sequential decision-making, to cases where the agent cannot access the true state and only receives noisy, incomplete observations. Karl Åström introduced the framework in 1965, and Leslie Kaelbling, Michael Littman, and colleagues brought it systematically into AI planning in 1998. A common approach is to maintain a “belief” — a probability distribution over the true state — and update it with Bayes' rule each time a new observation arrives. Robots operate in partially observable settings almost by default: cameras get occluded, drawers hide their contents, and many objects only reveal what they are once picked up. Exact solutions are computationally intractable, so practical systems use approximate planning, or give the policy network a history of frames plus a memory module to implicitly estimate the state.","example":"A robot searching for a cup in a row of closed cabinets can't see inside before opening a door, so it updates a probability distribution over “which compartment the cup is likely in” based on the compartments it has already checked, then decides which door to open next.","related":["Markov Decision Process","Observation","State Space","Embodied Memory","Memory-Augmented VLA","Reinforcement Learning"]},{"id":"embodied-memory","category":"concept","sec":8,"tier":3,"sources":[{"title":"MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation","url":"https://arxiv.org/abs/2508.19236"},{"title":"ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation","url":"https://arxiv.org/abs/2409.13682"},{"title":"KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems","url":"https://arxiv.org/abs/2409.14908"}],"as_of":"2025-08","related_ids":["memory-augmented-vla","memory-augmented-vla","long-horizon-task","partially-observable-markov-decision-process","3d-scene-graph","embodied-question-answering"],"name":"Embodied Memory","alt":"具身记忆","abbr":"","aliases":[],"one_liner":"A robot storing what it has seen and done so it can use that later for decisions and answering questions.","explanation":"Embodied memory refers to the mechanisms an embodied agent uses, over long periods of operation, to store and recall past experience: where it has been, where things are kept, how far a task has progressed, where it failed last time. This is necessary because a robot only ever sees part of the world at once, since it is partially observable, and many tasks run anywhere from minutes to days, so a policy that only looks at the current frame will forget sub-tasks it already finished, or keep searching for the same object over and over. There are roughly three common approaches. One keeps historical features inside the model itself — MemoryVLA (2025), for example, draws on human working memory and episodic memory to add a memory bank to a VLA model. Another maintains an external structured memory, such as a 3D scene graph or semantic map — KARMA uses a long-term scene graph plus short-term state records to support household task planning. The third is retrieval augmentation, as in ReMEmbR, which stores long-horizon navigation footage in a database indexed by time and location, used to answer questions about where and when something happened.","example":"A patrol robot asked, “where did you last see a red cart,” has to retrieve the time and location from hours of historical footage before it can answer.","related":["Memory-Augmented VLA","Memory-Augmented VLA","Long-horizon Task","Partially Observable Markov Decision Process","3D Scene Graph","Embodied Question Answering"]},{"id":"embodied-interaction","category":"concept","sec":8,"tier":3,"sources":[{"title":"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI (Liu et al., 2024)","url":"https://arxiv.org/abs/2407.06886"},{"title":"Paul Dourish (Wikipedia, on Where the Action Is: The Foundations of Embodied Interaction)","url":"https://en.wikipedia.org/wiki/Paul_Dourish"}],"as_of":"","related_ids":["embodied-perception","embodied-question-answering","human-robot-interaction","physical-human-robot-interaction","interactive-perception","language-grounding"],"name":"Embodied Interaction","alt":"具身交互","abbr":"","aliases":[],"one_liner":"The interaction an agent has with people, objects, and its surroundings in a physical or simulated space.","explanation":"In embodied AI, embodied interaction refers to an agent, a real robot or a virtual body in simulation, interacting with people and the environment in a physical or simulated space: this includes actions that change the environment, such as moving, grasping, pushing, and pulling, as well as acting on spoken instructions and answering questions. The 2024 embodied-AI survey from Sun Yat-sen University and others lists it as one of four research directions, alongside embodied perception, embodied agents, and sim-to-real, using embodied question answering and language-guided grasping as typical tasks. It differs from pure perception in that interaction changes the environment, so the agent must decide both what to do and when it has gathered enough information. The term has a separate meaning in human-computer interaction (HCI): Paul Dourish's 2001 book Where the Action Is uses it to describe how people interact with computing systems through their bodies, physical objects, and social context, which belongs to interface-design theory — a distinction worth keeping in mind when reading the literature.","example":"A user says, “I'm thirsty, get me something to drink.” The robot has to infer the intent, find a drink, and hand it over securely, not just recognize that there is a cup somewhere in the image.","related":["Embodied Perception","Embodied Question Answering","Human-Robot Interaction","Physical Human-Robot Interaction","Interactive Perception","Language Grounding"]},{"id":"embodied-agi","category":"concept","sec":9,"tier":2,"sources":[{"title":"Toward Embodied AGI: A Review of Embodied AI and the Road Ahead (arXiv 2505.14235)","url":"https://arxiv.org/abs/2505.14235"}],"as_of":"2025-05","related_ids":["embodied-ai","general-purpose-robot","physical-turing-test","levels-of-autonomy","humanoid-robot-intelligence-level-grading"],"name":"Embodied AGI","alt":"具身通用智能","abbr":"","aliases":["General Embodied Intelligence"],"one_liner":"Embodied intelligence able to perform diverse open-ended real-world tasks the way a human can — the field's long-term goal.","explanation":"Embodied AGI is a goal concept combining artificial general intelligence (AGI) with embodied AI; there is currently no single agreed-upon standard for it. A 2025 survey by Fan Wang and Hao Sun offers a working definition: embodied intelligence that shows human-like interaction ability and can perform diverse, open-ended, real-world tasks at a human level. The paper proposes an L1–L5 grading scale modeled on autonomous-driving levels: L1 completes a single task, L2 completes a combination of tasks, L3 conditionally completes general tasks, and L4–L5 progressively move toward open-ended tasks, human-like behavior, and no need for human intervention. It evaluates systems along four dimensions — omni-modality, human-like cognition, real-time responsiveness, and generalization — and judges that embodied AI today sits roughly between L1 and L2. Grading schemes like this are mainly used to describe direction and measure progress, not as a strict technical standard.","example":"Under this scale, a robot arm that can only pick up one type of part on a production line is L1; a robot that can break down “make a cup of coffee” into preset skills such as getting a cup, adding water, and pressing buttons, and complete them in order, is close to L2.","related":["Embodied AI","General-purpose Robot","Physical Turing Test","Levels of Autonomy","Humanoid Robot Intelligence Level Grading"]},{"id":"levels-of-autonomy","category":"concept","sec":9,"tier":2,"sources":[{"title":"Toward a framework for levels of robot autonomy in human-robot interaction (Beer, Fisk, Rogers, 2014)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5656240/"},{"title":"Self-driving car - Wikipedia（SAE J3016 分级）","url":"https://en.wikipedia.org/wiki/Self-driving_car"}],"as_of":"","related_ids":["fully-autonomous","teleoperation","humanoid-robot-intelligence-level-grading","shared-autonomy","human-robot-interaction","autonomous-driving"],"name":"Levels of Autonomy","alt":"自主等级","abbr":"","aliases":["Autonomy"],"one_liner":"A grading of how much a robot can perceive, decide, and act without relying on a person.","explanation":"Levels of autonomy describe how much a system can handle perception, decision-making, and execution without depending on a human. The best known example is SAE International's J3016 autonomous-driving scale (2014), L0 through L5, with the key boundary between L2 and L3: from L3 up, the system, not the driver, is responsible for monitoring the environment. In robotics, Jenay Beer, Arthur Fisk, and Wendy Rogers proposed a 10-level framework in 2014 running from pure teleoperation to full autonomy, and discussed how the degree of autonomy affects how willingly people accept a robot, their situational awareness, and how they judge its reliability. In embodied AI, the common question “was this demonstration teleoperated or fully autonomous” is really asking about the level of autonomy; China's humanoid robot intelligence grading standard also borrows this leveling approach from autonomous driving. The higher the degree of autonomy, the higher the demands on robustness, failure recovery, and safety.","example":"Two videos might both show clothes being folded, but one where an operator wears a VR headset and controls the robot in real time sits at the lowest level of autonomy; only a robot that sees, decides, and acts entirely on its own, with no human involved, counts as fully autonomous.","related":["Fully Autonomous","Teleoperation","Humanoid Robot Intelligence Level Grading","Shared Autonomy","Human-Robot Interaction","Autonomous Driving"]},{"id":"humanoid-robot-intelligence-level-grading","category":"concept","sec":9,"tier":2,"sources":[{"title":"全球首个《人形机器人智能化分级》标准推出，行业商业化进程加速（新浪财经，2025-05-29）","url":"https://baijiahao.baidu.com/s?id=1833417873950992497"}],"as_of":"2025-05","related_ids":["levels-of-autonomy","agibot-g1g5-embodied-ai-roadmap","humanoid-robot","beijing-humanoid-robot-innovation-center","miit-humanoid-robot-and-embodied-ai-standardization-technica"],"name":"Humanoid Robot Intelligence Level Grading","alt":"人形机器人智能化分级","abbr":"","aliases":["T/CIE 298-2025","Four Dimensions, Five Levels"],"one_liner":"A Chinese industry standard that grades humanoid robot intelligence from L1 to L5 across four dimensions.","explanation":"“Humanoid Robot Intelligence Level Grading” (T/CIE 298-2025) is an industry standard published in May 2025 by the China Institute of Electronics, led by the Beijing Innovation Center of Humanoid Robotics, with the Shanghai and Zhejiang humanoid robot innovation centers, UBTECH, Unitree Robotics, the China Academy of Information and Communications Technology, and other organizations taking part; it is reportedly the world's first grading standard specifically for humanoid robot intelligence. It uses a “four dimensions, five levels” framework, scoring along perception and cognition (P), decision-making and learning (D), execution performance (E), and collaboration and interaction (C), with intelligence level rising from L1 to L5; the standard includes 22 top-level indicators and more than 100 technical clauses. The approach borrows from autonomous-driving level systems, aiming to give manufacturers, customers, and investors a common yardstick and cut down on companies each marketing “intelligence” by their own separate definitions. Note that this is a different thing from AgiBot's G1–G5 technology roadmap, which is one company's own internal development plan.","example":"","related":["Levels of Autonomy","AgiBot G1–G5 Embodied-AI Roadmap","Humanoid Robot","Beijing Humanoid Robot Innovation Center","MIIT Humanoid Robot and Embodied AI Standardization Technical Committee"]},{"id":"physical-turing-test","category":"concept","sec":9,"tier":2,"sources":[{"title":"The Physical Turing Test: Jim Fan on Nvidia's Roadmap for Embodied AI (Sequoia Capital, YouTube)","url":"https://www.youtube.com/watch?v=_2NijXqBESI"},{"title":"AI Ascent 2025 (Sequoia Capital)","url":"https://www.sequoiacap.com/article/ai-ascent-2025/"},{"title":"Turing test - Wikipedia","url":"https://en.wikipedia.org/wiki/Turing_test"}],"as_of":"2025-05","related_ids":["embodied-ai","physical-ai","general-purpose-robot","household-tasks","moravec-s-paradox","digital-cousin"],"name":"Physical Turing Test","alt":"物理图灵测试","abbr":"","aliases":[],"one_liner":"A robot passes if people cannot tell whether a piece of real-world physical work was done by a human or a machine.","explanation":"The physical Turing test is a phrase coined by NVIDIA's senior director of AI, Jim Fan, in a talk at Sequoia Capital's AI Ascent conference in May 2025, moving the Turing test from “can't tell human from machine in conversation” into the physical world: have a robot do real physical work, such as housework, and if a person looking at the result cannot tell whether a human or a robot did it, the test is passed. In the talk, he argued for using large-scale simulation to make up for the shortage of robot training data, and described ideas such as digital twins and “digital cousins.” Similar ideas existed earlier: cognitive scientist Stevan Harnad's “total Turing test” likewise required testing both perception and the ability to manipulate objects. The phrase is now commonly used to describe the ultimate goal of a general-purpose household robot.","example":"Jim Fan's example scenario: your house is a mess after a party on Sunday night; you come home Monday to find it fully cleaned up, with a candlelit dinner set on the table, and you cannot tell whether a person or a robot did it.","related":["Embodied AI","Physical AI","General-purpose Robot","Household Tasks","Moravec's Paradox","Digital Cousin"]},{"id":"the-coffee-test","category":"concept","sec":9,"tier":3,"sources":[{"title":"Artificial general intelligence - Wikipedia（Tests for human-level AGI 一节）","url":"https://en.wikipedia.org/wiki/Artificial_general_intelligence"}],"as_of":"2025","related_ids":["physical-turing-test","embodied-agi","household-tasks","open-world","long-horizon-task","moravec-s-paradox"],"name":"The Coffee Test","alt":"咖啡测试","abbr":"","aliases":["Wozniak Coffee Test"],"one_liner":"Sending a machine into a stranger's house to find everything it needs and brew a pot of coffee.","explanation":"The Coffee Test is a proposed test for general-purpose AI put forward by Apple co-founder Steve Wozniak: a machine must walk into an ordinary American home and figure out how to make coffee on its own, finding the coffee maker, finding the coffee, adding water, finding a cup, and pressing the right buttons to brew it. Unlike the Turing test, which only examines conversation, it requires perception, search, common-sense reasoning, and manipulation in a real, unfamiliar environment, which is why it is often invoked to argue that being able to chat is not the same as being able to get things done. Per Wikipedia, Figure 01 demonstrated autonomously operating a Keurig pod coffee machine in January 2024, and researchers at the University of Edinburgh reportedly had a robot arm make coffee from voice commands with a framework called ELLMER in 2025, but neither was done in a randomly chosen stranger's home, so both still fall short of what the test actually demands.","example":"A large model can list every step of making coffee, but to pass the Coffee Test, a robot has to rummage through the cabinets of a kitchen it has never seen to find the coffee grounds, and recognize which button to press on that particular machine.","related":["Physical Turing Test","Embodied AGI","Household Tasks","Open-world","Long-horizon Task","Moravec's Paradox"]},{"id":"the-bitter-lesson","category":"concept","sec":9,"tier":2,"sources":[{"title":"The Bitter Lesson (Rich Sutton, 2019)","url":"http://www.incompleteideas.net/IncIdeas/BitterLesson.html"}],"as_of":"","related_ids":["scaling-law","end-to-end","moravec-s-paradox","foundation-model","data-scarcity","embodied-ai"],"name":"The Bitter Lesson","alt":"苦涩的教训","abbr":"","aliases":[],"one_liner":"Rich Sutton's argument that, over time, general methods that scale with compute beat hand-engineered domain knowledge.","explanation":"This is a short essay published in March 2019 by Rich Sutton, one of the founders of reinforcement learning. Looking back over 70 years of AI research, he argues that what has actually worked in the long run is general-purpose methods that can fully exploit growing compute — search and learning — rather than hard-coding human understanding of a domain into the system. Chess, Go, speech, and vision all followed the same pattern: hand-engineered knowledge produced short-term progress, but was ultimately overtaken by scaled-up search and learning, because the cost of a unit of compute keeps falling exponentially. In embodied AI, this essay is often invoked to support an “end-to-end plus big data plus large models” approach and to argue against over-engineered, hand-designed modules; the point of contention is that robot data is far scarcer than text and images, so whether the same lesson applies is still debated.","example":"Sutton's own example: computer Go long relied on hand-crafted Go theory and Go-specific structure, and was eventually surpassed completely by methods based on large-scale search plus self-play learning of a value function.","related":["Scaling Law","End-to-End","Moravec's Paradox","Foundation Model","Data Scarcity","Embodied AI"]},{"id":"emergent-abilities","category":"concept","sec":9,"tier":2,"sources":[{"title":"Emergent Abilities of Large Language Models (arXiv 2206.07682)","url":"https://arxiv.org/abs/2206.07682"},{"title":"Are Emergent Abilities of Large Language Models a Mirage? (arXiv 2304.15004)","url":"https://arxiv.org/abs/2304.15004"},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)","url":"https://arxiv.org/abs/2307.15818"}],"as_of":"","related_ids":["scaling-law","large-language-model","chain-of-thought","rt-2","zero-shot","generalization"],"name":"Emergent Abilities","alt":"涌现能力","abbr":"","aliases":["Emergence","Emergent Capabilities"],"one_liner":"Abilities absent in small models that only appear once model or data scale grows large enough.","explanation":"This idea was systematically laid out in a 2022 paper by Jason Wei and colleagues, “Emergent Abilities of Large Language Models,” which defines an emergent ability as one that is absent in smaller models, appears in larger models, and cannot be predicted by extrapolating from smaller models' performance — multi-step arithmetic and chain-of-thought reasoning are examples. In 2023, Rylan Schaeffer and colleagues pushed back, arguing that much of this apparent “sudden appearance” is an artifact of the evaluation metric: metrics that only score an all-or-nothing correct answer look like a sharp jump, but switching to a continuous metric shows the improvement is actually smooth. In embodied AI, “emergent” usually refers to a capability that was not in the training data but shows up at test time anyway — for example, RT-2 following instructions absent from any robot data, such as placing an object on a specific number or icon, or picking up a rock to use as an improvised hammer. Claims of emergence in marketing materials should be checked against the actual evaluation behind them.","example":"RT-2's robot demonstration data contains no task like “put the object on the numbered card,” but its vision-language-model backbone recognizes digits from web image-text data, so it can follow the instruction anyway.","related":["Scaling Law","Large Language Model","Chain-of-Thought","RT-2","Zero-shot","Generalization"]},{"id":"era-of-experience","category":"concept","sec":9,"tier":3,"sources":[{"title":"Welcome to the Era of Experience (Silver & Sutton, 2025, preprint)","url":"https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf"}],"as_of":"2025-04","related_ids":["reinforcement-learning","the-bitter-lesson","real-world-reinforcement-learning","self-improvement","data-flywheel","continual-learning"],"name":"Era of Experience","alt":"经验时代","abbr":"","aliases":["Welcome to the Era of Experience"],"one_liner":"Silver and Sutton's view that AI's next stage will learn mainly from its own experience of interacting with the world.","explanation":"“Era of Experience” comes from a short paper published by David Silver and Richard Sutton in April 2025, “Welcome to the Era of Experience,” a preprint of one chapter from the MIT Press book Designing an Intelligence; Silver was one of the lead figures behind the AlphaGo line of work, and Sutton is one of the founders of reinforcement learning. The paper argues that the “era of human data,” training on massive amounts of human-generated data and then fine-tuning on human preferences, is nearing its ceiling in areas such as mathematics, coding, and science, and that the next stage of agents should learn mainly from experience generated by their own interaction with the environment. It lists four defining features: living in a long stream of experience rather than short conversations; actions and observations grounded in the environment; rewards coming from real signals in the environment rather than being judged by people in advance; and planning and reasoning also grounded in experience. For embodied AI, this view supports the idea of robots continuously trying, failing, and improving themselves through real-world deployment.","example":"The paper cites AlphaProof as an example: it first studies roughly 100,000 human-written formal proofs, then generates about 100 million more on its own by interacting with a proof-checking system, eventually reaching medal-level performance at the International Mathematical Olympiad.","related":["Reinforcement Learning","The Bitter Lesson","Real-World Reinforcement Learning","Self-improvement","Data Flywheel","Continual Learning"]},{"id":"three-schools-of-ai","category":"concept","sec":9,"tier":3,"sources":[{"title":"Symbolic artificial intelligence - Wikipedia","url":"https://en.wikipedia.org/wiki/Symbolic_artificial_intelligence"},{"title":"联结主义 - 维基百科","url":"https://zh.wikipedia.org/wiki/联结主义"},{"title":"Behavior-based robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Behavior-based_robotics"}],"as_of":"","related_ids":["subsumption-architecture","symbol-grounding-problem","embodied-cognition","perception-action-loop","neural-network","moravec-s-paradox"],"name":"Three Schools of AI","alt":"人工智能三大学派（符号主义 / 连接主义 / 行为主义）","abbr":"","aliases":["Symbolism, Connectionism, and Behaviorism","Three Paradigms of AI"],"one_liner":"A classification of AI research, common in Chinese textbooks, into logic-symbol, neural-network, and interact-with-environment approaches.","explanation":"This is a classification of AI research commonly taught in Chinese AI textbooks, dividing approaches by “where intelligence comes from” into three schools. Symbolism (also called logicism) holds that intelligence is logical reasoning over symbols; its landmarks are Allen Newell and Herbert Simon's Logic Theorist program from the 1950s and the expert systems that followed. Connectionism (also called the bionic school) holds that intelligence comes from the connections and learning of large numbers of simple neurons; today's deep learning and large models fall under this school. Behaviorism (also called the cybernetics school) holds that intelligence is expressed in a perception-action loop and does not require a complete world model to be built first; its landmark is Rodney Brooks's subsumption-architecture robots at MIT in the 1980s. Embodied AI is philosophically closest to behaviorism but methodologically relies heavily on connectionist neural networks; in practice, the three schools have largely merged today.","example":"Take the task “put the apple in the bowl”: a symbolist approach would hand-write a rule such as “if an apple is seen and the hand is empty, grasp it”; a connectionist approach would train a neural network on many demonstrations to map images directly to actions; a behaviorist approach would wire up a set of “see-reach-adjust” reflex loops and let the behavior take shape through interaction with the environment. Today's VLA models mainly follow the connectionist path.","related":["Subsumption Architecture","Symbol Grounding Problem","Embodied Cognition","Perception-Action Loop","Neural Network","Moravec's Paradox"]},{"id":"symbol-grounding-problem","category":"concept","sec":9,"tier":3,"sources":[{"title":"The Symbol Grounding Problem (Harnad, Physica D 1990)","url":"https://arxiv.org/abs/cs/9906002"},{"title":"Symbol grounding problem - Wikipedia","url":"https://en.wikipedia.org/wiki/Symbol_grounding_problem"}],"as_of":"","related_ids":["language-grounding","embodied-cognition","disembodied-ai","three-schools-of-ai","saycan"],"name":"Symbol Grounding Problem","alt":"符号落地问题","abbr":"","aliases":["Symbol Grounding"],"one_liner":"How the symbols inside a machine come to refer to real things in the world, and so mean something.","explanation":"This is a problem cognitive scientist Stevan Harnad posed in Physica D in 1990: for a system that only manipulates symbols according to rules, how do those symbols get their own meaning, rather than depending entirely on a person to interpret them? His example is trying to learn Chinese from a Chinese-Chinese dictionary alone: every word is defined using other words, and no amount of looking things up ever touches an actual object. Harnad's proposed solution is to ground the most basic symbols in perception: build sensory representations of the outside world first, then learn features that distinguish categories; symbols are the names of these categories, and complex concepts are built by combining them. Embodied AI often treats this as part of the theoretical case for “why a body is needed,” and it is frequently cited in discussions of language grounding, meaning connecting words to vision and action.","example":"Given the instruction “bring me the red cup,” a large language model can parse the sentence, but it only counts as grounded for the robot once “the red cup” is tied to a specific object in the camera feed and “bring” is tied to a sequence of grasping actions.","related":["Language Grounding","Embodied Cognition","Disembodied AI","Three Schools of AI","SayCan"]},{"id":"subsumption-architecture","category":"concept","sec":9,"tier":3,"sources":[{"title":"Subsumption architecture - Wikipedia","url":"https://en.wikipedia.org/wiki/Subsumption_architecture"}],"as_of":"","related_ids":["sense-plan-act","three-schools-of-ai","embodied-cognition","finite-state-machine","moravec-s-paradox","behavior-tree"],"name":"Subsumption Architecture","alt":"包容架构（行为式机器人）","abbr":"","aliases":["Behavior-based Robotics"],"one_liner":"A layered robot-control architecture wiring simple behaviors straight from sensing to action, where higher layers override lower ones.","explanation":"Subsumption architecture was introduced by MIT's Rodney Brooks in his 1986 paper “A Robust Layered Control System for a Mobile Robot,” and it is the foundational work of behavior-based robotics. It rejects the traditional sense-plan-act pipeline, which first builds a complete world model, then plans, then executes. Instead, control is split into several parallel behavior layers, each wiring sensors directly to actuators through a simple finite state machine: lower layers handle obstacle avoidance, higher layers handle wandering and exploration, and a higher layer can inhibit or override a lower layer's inputs and outputs, “subsuming” its behavior. The approach is fast and robust, but struggles to support complex reasoning and language understanding, and is often seen as one of the roots of both embodied cognition and the behaviorist school of AI.","example":"Genghis, a six-legged robot from Brooks's lab, walked using layered simple behaviors; another robot, Herbert, used the same approach to collect empty soda cans around an office.","related":["Sense-Plan-Act","Three Schools of AI","Embodied Cognition","Finite State Machine","Moravec's Paradox","Behavior Tree"]},{"id":"embodied-cognition","category":"concept","sec":9,"tier":3,"sources":[{"title":"Embodied Cognition (Stanford Encyclopedia of Philosophy)","url":"https://plato.stanford.edu/entries/embodied-cognition/"},{"title":"Embodied cognition (Wikipedia)","url":"https://en.wikipedia.org/wiki/Embodied_cognition"}],"as_of":"","related_ids":["embodied-ai","disembodied-ai","subsumption-architecture","morphological-computation","symbol-grounding-problem","moravec-s-paradox"],"name":"Embodied Cognition","alt":"具身认知","abbr":"","aliases":["Embodiment (Cognitive Science)"],"one_liner":"The view in cognitive science that thought is inseparable from the body and its interaction with the environment.","explanation":"Embodied cognition is a family of theories in cognitive science and philosophy of mind holding that thought is not just symbol manipulation happening in the brain, but is shaped jointly by the body's form, its sensorimotor abilities, and the body's interaction with the environment. The idea traces back to the phenomenology of Maurice Merleau-Ponty and others; Francisco Varela, Evan Thompson, and Eleanor Rosch's 1991 book The Embodied Mind systematized it, while George Lakoff and Mark Johnson argued from a linguistic angle that abstract concepts derive from bodily experience. Its influence on robotics has been direct: Rodney Brooks abandoned the approach of building a complete world model before planning, letting sensors drive behavior directly and arguing that the world itself is the best model; Rolf Pfeifer and others likewise argued that genuine intelligence requires a body with sensorimotor capabilities. Today's term “embodied AI” descends from this line of thinking, and it provides the theoretical basis for the claim that intelligence must form through interaction with the physical world.","example":"English uses “grasp” to mean “understand”; Lakoff and Johnson treat expressions like this as evidence that abstract concepts borrow from bodily experience.","related":["Embodied AI","Disembodied AI","Subsumption Architecture","Morphological Computation","Symbol Grounding Problem","Moravec's Paradox"]},{"id":"morphological-computation","category":"concept","sec":9,"tier":3,"sources":[{"title":"Trade-Offs in Exploiting Body Morphology for Control (Hoffmann & Müller)","url":"https://arxiv.org/abs/1411.2276"},{"title":"Passive dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Passive_dynamics"}],"as_of":"","related_ids":["embodied-cognition","passive-dynamic-walking","soft-robot","compliance","morphology-control-co-design","cost-of-transport"],"name":"Morphological Computation","alt":"形态计算","abbr":"","aliases":["Morphological Computing"],"one_liner":"Letting a body's own shape and material handle part of the work a controller would otherwise have to compute.","explanation":"Morphological computation is an idea from embodied cognition and bio-inspired robotics: a body's shape, material elasticity, and mass distribution can themselves take on some of the work that would otherwise have to be handled by a controller, effectively offloading computation onto the body. The idea has been promoted mainly by embodied-cognition researchers such as Rolf Pfeifer. The classic example is Tad McGeer's passive dynamic walker from the late 1980s: with no motors and no controller, it walks down a gentle slope using only gravity and the natural swing of its legs. Soft grippers, which conform to an object's shape through material deformation alone, are another common example. Hoffmann and Müller caution that a softer, more complex body is not automatically easier to control, and that morphology and control method need to be balanced against each other.","example":"A passive dynamic walker has no motors at all — leg length, mass distribution, and gravity alone let it walk down a shallow slope. According to Wikipedia, a Cornell biped built on this principle has a cost of transport of around 0.20, close to a human's, versus about 3.23 for Honda's ASIMO.","related":["Embodied Cognition","Passive Dynamic Walking","Soft Robot","Compliance","Morphology-Control Co-design","Cost of Transport"]},{"id":"morphology-control-co-design","category":"concept","sec":9,"tier":3,"sources":[{"title":"Evolved Virtual Creatures (Karl Sims, 1994)","url":"https://www.karlsims.com/evolved-virtual-creatures.html"},{"title":"Embodied Intelligence via Learning and Evolution (DERL)","url":"https://arxiv.org/abs/2102.02202"},{"title":"Transform2Act: Learning a Transform-and-Control Policy for Efficient Agent Design","url":"https://arxiv.org/abs/2110.03659"}],"as_of":"","related_ids":["morphological-computation","reinforcement-learning","embodiment","simulator","soft-robot","embodied-ai"],"name":"Morphology-Control Co-design","alt":"形态-控制协同设计（形态进化）","abbr":"","aliases":["Morphological Evolution","Brain-Body Co-design","Co-design of Morphology and Control"],"one_liner":"Optimizing a robot's body and its control policy together, instead of fixing the body first.","explanation":"Morphology-control co-design means searching for a robot's body (number of limbs, limb lengths, joint layout, and so on) and the policy that controls it at the same time. The traditional approach fixes the body first and then writes or trains a controller for it; co-design argues that the best controller depends on the body, and the best body depends on the controller, so both should be optimized in a single loop. Karl Sims's 1994 SIGGRAPH paper “Evolving Virtual Creatures” used simulated evolution to co-evolve the morphology and neural controllers of virtual creatures. Stanford's Gupta and colleagues built on this with DERL (2021), which evolves the body while using reinforcement learning to train the controller; ICLR 2022's Transform2Act instead treats “changing the body” as just another action for the policy to learn.","example":"In DERL, virtual agents assembled from different limbs learn to walk and manipulate objects in simulation; the better-performing individuals are kept and have their body structure mutated, and after many generations, more stable, more energy-efficient, faster-learning morphologies emerge.","related":["Morphological Computation","Reinforcement Learning","Embodiment","Simulator","Soft Robot","Embodied AI"]},{"id":"robotic-arm","category":"robot","sec":0,"tier":1,"sources":[{"title":"Franka Research 3 (franka.de)","url":"https://franka.de/franka-research-3"}],"as_of":"","related_ids":["degrees-of-freedom","end-effector","collaborative-robot","industrial-robot","7-dof-robot-arm","franka-emika-panda-franka-research-3"],"name":"Robotic Arm","alt":"机械臂","abbr":"","aliases":["Manipulator","Robot Arm"],"one_liner":"An arm-shaped robot made of joints and links, with a gripper or tool mounted at the end.","explanation":"A robotic arm is a chain of links and joints with a fixed base and an end effector, such as a gripper, suction cup, or dexterous hand, at the tip. Six degrees of freedom are enough for the end effector to reach any position and orientation in space, while a 7-DoF arm has one redundant joint, letting it work around obstacles. Arms split broadly into industrial arms, built for speed and precision and typically fenced off from people, and collaborative arms, which add force sensing and are designed to work alongside people. Most manipulation tasks studied in embodied AI, such as grasping, assembly, and folding laundry, are performed on robotic arms; common ones include Franka, UR5e, WidowX, and the low-cost SO-100.","example":"Most of the data in the Open X-Embodiment dataset comes from single-arm or dual-arm robotic arms.","related":["Degrees of Freedom (DoF)","End Effector","Collaborative Robot","Industrial Robot","7-DoF Robot Arm","Franka Emika Panda / Franka Research 3"]},{"id":"6-axis-robot-arm","category":"robot","sec":0,"tier":2,"sources":[{"title":"Robotic arm - Wikipedia","url":"https://en.wikipedia.org/wiki/Robotic_arm"},{"title":"Universal Robots UR5e","url":"https://www.universal-robots.com/products/ur5e/"}],"as_of":"","related_ids":["robotic-arm","degrees-of-freedom","inverse-kinematics","spherical-wrist","7-dof-robot-arm","universal-robots-ur5e"],"name":"6-Axis Robot Arm","alt":"六轴机械臂","abbr":"","aliases":["6-DoF Robot Arm","Six-Axis Robot Arm"],"one_liner":"A robot arm with six rotating joints in series, able to reach any position and orientation.","explanation":"A 6-axis robot arm is built from six rotating joints connected in series, the most common configuration for industrial and collaborative robot arms. An object's pose in space, meaning its x/y/z position plus three rotation angles, has exactly 6 degrees of freedom, so six joints are, in principle, enough for the end effector, such as a gripper or welding torch mounted on the wrist, to reach any pose within the workspace. A common design puts the first three axes at the base to set position and the last three in a spherical wrist to set orientation, which is why most 6-axis arms have a closed-form solution for inverse kinematics, meaning computing joint angles from a target pose. Representative products include the UR5e, xArm6, and AgileX PiPER; in embodied-AI research it is often used as a single-arm manipulation platform. Lacking a redundant joint, it is less flexible than a 7-DoF arm when facing singular configurations or needing to reach around obstacles.","example":"The UR5e is a 6-axis collaborative arm with a 5 kg payload, commonly used for pick-and-place experiments.","related":["Robotic Arm","Degrees of Freedom (DoF)","Inverse Kinematics (IK)","Spherical Wrist","7-DoF Robot Arm","Universal Robots UR5e"]},{"id":"7-dof-robot-arm","category":"robot","sec":0,"tier":2,"sources":[{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"},{"title":"Robotic arm - Wikipedia","url":"https://en.wikipedia.org/wiki/Robotic_arm"}],"as_of":"","related_ids":["kinematic-redundancy","null-space","swivel-angle","6-axis-robot-arm","franka-emika-panda-franka-research-3","kuka-lbr-iiwa"],"name":"7-DoF Robot Arm","alt":"七自由度机械臂","abbr":"","aliases":["7-Axis Robot Arm","Redundant Manipulator"],"one_liner":"A robot arm with one more joint than a 6-axis arm, letting the same hand pose be reached many ways.","explanation":"A 7-DoF robot arm has 7 joints in series, one more than the 6 needed to fix an end effector's pose, which is called kinematic redundancy. The benefit is that infinitely many joint-angle combinations produce the same end-effector pose, so the controller can adjust the elbow position within the null space, meaning the part of joint motion that doesn't change the end-effector pose, to avoid obstacles, singular configurations, and joint limits, while also moving in a way closer to a human arm, since a human shoulder, elbow, and wrist together also have roughly 7 degrees of freedom. The tradeoff is that inverse kinematics no longer has a unique solution, requiring an extra parameter such as the arm angle, or numerical optimization, to pick one. Representative products include the Franka Panda/FR3, KUKA LBR iiwa, and Kinova Gen3; humanoid robots such as AgiBot's Expedition A2 also use 7-DoF arms.","example":"The DROID dataset was collected at multiple labs using Franka Panda 7-DoF robot arms.","related":["Kinematic Redundancy","Null Space","Swivel Angle","6-Axis Robot Arm","Franka Emika Panda / Franka Research 3","KUKA LBR iiwa"]},{"id":"industrial-robot","category":"robot","sec":0,"tier":2,"sources":[{"title":"IFR: Industrial Robots","url":"https://ifr.org/industrial-robots"},{"title":"IFR: Robot definitions at ISO","url":"https://ifr.org/standardisation"}],"as_of":"","related_ids":["collaborative-robot","6-axis-robot-arm","selective-compliance-assembly-robot-arm","big-four-of-industrial-robotics","teach-and-playback-programming","international-federation-of-robotics"],"name":"Industrial Robot","alt":"工业机器人","abbr":"","aliases":[],"one_liner":"A programmable, multi-jointed robot arm used in factory automation.","explanation":"Under ISO 8373, as adopted by the International Federation of Robotics (IFR), an industrial robot is an automatically controlled, reprogrammable, multipurpose manipulator, programmable in at least three axes, which can be either fixed in place or mounted on a mobile platform, used for industrial automation. Common forms include 6-axis arms, SCARA robots, Delta parallel robots, and Cartesian robots, mainly used for welding, painting, material handling, assembly, and palletizing. Most run fixed trajectories set through teach-and-playback or offline programming, achieving high precision and cycle-time performance, but reprogramming is needed for a new task, which is exactly the gap embodied AI is trying to close with learning-based methods. The major manufacturers, FANUC, ABB, KUKA, and Yaskawa, are known as the “Big Four.”","example":"Rows of 6-axis robot arms in an auto body shop spot-weld car bodies along taught trajectories.","related":["Collaborative Robot","6-Axis Robot Arm","Selective Compliance Assembly Robot Arm","Big Four of Industrial Robotics","Teach-and-Playback Programming","International Federation of Robotics"]},{"id":"selective-compliance-assembly-robot-arm","category":"robot","sec":0,"tier":3,"sources":[{"title":"Wikipedia: SCARA","url":"https://en.wikipedia.org/wiki/SCARA"}],"as_of":"","related_ids":["industrial-robot","6-axis-robot-arm","pick-and-place","robotic-assembly","cartesian-robot","delta-robot"],"name":"Selective Compliance Assembly Robot Arm","alt":"SCARA 机器人","abbr":"SCARA","aliases":["Selective Compliance Articulated Robot Arm","Planar Articulated Robot"],"one_liner":"A four-axis industrial arm that's flexible horizontally but rigid vertically, built for fast pick-and-place assembly.","explanation":"SCARA is an industrial robot configuration proposed in the late 1970s by Hiroshi Makino at the University of Yamanashi in Japan. The typical structure has two vertically axed rotary joints that swing the arm within a horizontal plane, plus a vertical linear axis and a rotating end axis, for four degrees of freedom in total. The “selective compliance” in the name refers to the arm being somewhat compliant horizontally but very rigid vertically, which makes it well suited to inserting parts straight down into holes. Compared with a 6-axis arm, its workspace is smaller and its motion more constrained, but it's fast, repeats positioning very precisely, and is cheap — so it's widely used for pick-and-place work on a flat plane, such as electronics assembly, dispensing adhesive, and sorting.","example":"SCARA robots are commonly used on phone-motherboard lines to pick up small components and insert them straight down into the circuit board.","related":["Industrial Robot","6-Axis Robot Arm","Pick-and-Place","Robotic Assembly","Cartesian Robot","Delta Robot"]},{"id":"delta-robot","category":"robot","sec":0,"tier":3,"sources":[{"title":"Delta robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Delta_robot"},{"title":"Reymond Clavel - Wikipedia","url":"https://en.wikipedia.org/wiki/Reymond_Clavel"}],"as_of":"","related_ids":["parallel-mechanism","pick-and-place","sorting","industrial-robot","selective-compliance-assembly-robot-arm","serial-mechanism"],"name":"Delta Robot","alt":"Delta 并联机器人（蜘蛛手）","abbr":"","aliases":["Parallel Robot","Spider Robot"],"one_liner":"A high-speed sorting robot with three parallel arms pulling a small platform, nicknamed the “spider hand.”","explanation":"The Delta robot is a parallel robot, meaning several kinematic chains connect to the end effector at once, unlike the link-after-link serial structure of a typical robot arm, invented in the 1980s by Reymond Clavel's team at EPFL, the Swiss Federal Institute of Technology in Lausanne. It has three arms hanging from a base at the top, each connected through a parallelogram linkage to a small platform below, constraining the platform to translate but not rotate, often with an added rotary axis. All the motors sit on the fixed base, so the moving parts are very light, giving it extremely high speed and accurate positioning, at the cost of a small payload and limited workspace. It was originally developed for packing chocolates at high speed, and is now widely used for high-speed sorting of food, electronic components, and pharmaceuticals, a typical device for pick-and-place tasks.","example":"A “spider hand” hanging upside down above a food-factory line uses vision to spot cookies on the conveyor and picks several into packaging boxes each second.","related":["Parallel Mechanism","Pick-and-Place","Sorting","Industrial Robot","Selective Compliance Assembly Robot Arm","Serial Mechanism"]},{"id":"cartesian-robot","category":"robot","sec":0,"tier":3,"sources":[{"title":"Cartesian coordinate robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Cartesian_coordinate_robot"}],"as_of":"","related_ids":["industrial-robot","prismatic-joint","selective-compliance-assembly-robot-arm","6-axis-robot-arm","machine-tending","palletizing-depalletizing"],"name":"Cartesian Robot","alt":"直角坐标机器人","abbr":"","aliases":["Gantry Robot","Cartesian Coordinate Robot"],"one_liner":"A robot whose end effector moves along three mutually perpendicular straight-line axes: X, Y, and Z.","explanation":"A Cartesian robot has three main motion axes, all linear, meaning prismatic joints, moving along mutually perpendicular X, Y, and Z directions, so the end-effector position is simply given by the three axis readings, with almost no inverse-kinematics computation needed. A version with the crossbeam supported at both ends and spanning above the work area is called a gantry robot; it offers a large working range, good stiffness, and can handle heavy loads. It has a simple structure, high precision, and is easy to scale, common in machine-tool loading and unloading, palletizing, dispensing, and 3D printers and CNC machines. Its drawbacks are a large footprint, motion confined to a rectangular box, and limited orientation flexibility, requiring an added rotary axis whenever an angle needs to change. Together with the 6-axis arm and SCARA, it forms one of the basic configurations of industrial robots.","example":"A three-axis gantry manipulator above an injection-molding machine lifts a molded part out of the mold and places it on a conveyor.","related":["Industrial Robot","Prismatic Joint","Selective Compliance Assembly Robot Arm","6-Axis Robot Arm","Machine Tending","Palletizing / Depalletizing"]},{"id":"collaborative-robot","category":"robot","sec":0,"tier":2,"sources":[{"title":"Cobot - Wikipedia","url":"https://en.wikipedia.org/wiki/Cobot"},{"title":"ISO/TS 15066:2016 Robots and robotic devices — Collaborative robots","url":"https://www.iso.org/standard/62996.html"}],"as_of":"","related_ids":["industrial-robot","human-robot-collaboration","iso-ts-15066-robots-and-robotic-devices-collaborative-robots","power-and-force-limiting","kinesthetic-teaching","universal-robots"],"name":"Collaborative Robot","alt":"协作机器人","abbr":"Cobot","aliases":["Cobot","Collaborative Arm"],"one_liner":"A lightweight robot arm designed to share workspace with people, without a safety fence.","explanation":"A collaborative robot, or cobot, is a robot designed to work in the same space as people, most commonly a lightweight 6- or 7-axis arm with a payload from a few kilograms up to twenty or thirty. Traditional industrial robots must be kept behind safety fencing; cobots instead reduce the risk of hurting someone through power and force limiting, collision detection, meaning they stop instantly if joint current or torque sensors notice an unexpected contact, and speed-and-separation monitoring, with the relevant requirements set out in ISO 10218 and ISO/TS 15066. They typically support kinesthetic teaching, where a person drags the arm by hand to record a motion, which makes them easy to deploy and a common platform in small and mid-sized factory automation and in embodied-AI labs. Representative brands include Universal Robots' UR series, Franka, Dobot, JAKA, and Flexiv. Whether a setup actually counts as “collaborative” depends on a full risk assessment of the application, not just on swapping in a cobot arm.","example":"The UR5e has a 5 kg payload and can be hand-guided directly through kinesthetic teaching, making it one of the most common collaborative arms in research.","related":["Industrial Robot","Human-Robot Collaboration","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Power and Force Limiting","Kinesthetic Teaching","Universal Robots"]},{"id":"desktop-robot-arm","category":"robot","sec":0,"tier":2,"sources":[{"title":"TheRobotStudio/SO-ARM100 - GitHub","url":"https://github.com/TheRobotStudio/SO-ARM100"},{"title":"SO-101 - LeRobot Docs","url":"https://huggingface.co/docs/lerobot/so101"}],"as_of":"2025","related_ids":["so-100-so-101-arm","lerobot","leader-follower-teleoperation","elephant-robotics-mycobot","dobot-magician","open-source-hardware"],"name":"Desktop Robot Arm","alt":"桌面机械臂","abbr":"","aliases":["Desktop Robotic Arm","Low-cost Robot Arm"],"one_liner":"A small, inexpensive robot arm that sits directly on a desk, used for teaching and data collection.","explanation":"A desktop robot arm is small and light enough to sit directly on an office desk, usually with a payload of a few hundred grams to a kilogram or two, priced far below industrial arms. Early examples were mainly for education and makers, such as Dobot Magician and Elephant Robotics' myCobot. After 2024, the open-source SO-100/SO-101 from the Hugging Face LeRobot community, servo-driven and 3D-printed, with a single arm's parts reportedly costing on the order of $100, turned it into a popular entry point into embodied AI: one arm used as a leader, moved by hand, and a follower arm that copies it, is enough to teleoperate and collect demonstration data, which can then train policies such as ACT or SmolVLA. Its precision, stiffness, and payload are all limited, making it suited to algorithm validation rather than industrial work.","example":"The SO-101 is assembled from Feetech STS3215 servos and 3D-printed parts; a leader-follower pair is enough to teleoperate and collect data to train an ACT policy.","related":["SO-100 / SO-101 Arm","LeRobot","Leader-Follower Teleoperation","Elephant Robotics myCobot","Dobot Magician","Open-Source Hardware (OSHW)"]},{"id":"dual-arm-robot","category":"robot","sec":0,"tier":2,"sources":[{"title":"ALOHA: A Low-cost Open-source Hardware System for Bimanual Teleoperation","url":"https://tonyzhaozh.github.io/aloha/"},{"title":"Mobile ALOHA","url":"https://mobile-aloha.github.io/"}],"as_of":"","related_ids":["bimanual-manipulation","aloha","mobile-aloha","agilex-cobot-magic","abb-yumi","rethink-robotics-baxter"],"name":"Dual-arm Robot","alt":"双臂机器人","abbr":"","aliases":["Bimanual Robot"],"one_liner":"A robot with two arms, able to coordinate them the way a person uses both hands.","explanation":"A dual-arm robot has two robot arms that can coordinate to complete a task together, whether mounted as a fixed tabletop platform or attached to a mobile base or humanoid torso. Many everyday tasks can't be done with one hand, such as folding laundry, unscrewing a bottle cap, or holding something steady while inserting a part; this kind of bimanual manipulation requires the two arms to coordinate in time and space, doubling the action space, and is a major challenge in imitation learning and VLA research. Early examples include Rethink Robotics' Baxter and ABB's YuMi; the most common platform in embodied-AI research today is Stanford's ALOHA, which uses two pairs of low-cost leader-follower arms for teleoperation, along with its derivative, AgileX's Cobot Magic. Models such as RDT-1B and π0 are trained on large amounts of this kind of bimanual data.","example":"Mobile ALOHA mounts an ALOHA bimanual setup on a mobile base, and with only about 50 demonstrations per task, learns tasks such as stir-frying shrimp.","related":["Bimanual Manipulation","ALOHA","Mobile ALOHA","AgileX Cobot Magic","ABB YuMi","Rethink Robotics Baxter"]},{"id":"wheeled-robot","category":"robot","sec":1,"tier":3,"sources":[{"title":"Robot locomotion - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_locomotion"}],"as_of":"","related_ids":["mobile-base","differential-drive-base","mecanum-wheel","wheel-legged-robot","wheeled-humanoid-robot","autonomous-mobile-robot"],"name":"Wheeled Robot","alt":"轮式机器人","abbr":"","aliases":["Wheeled Mobile Robot"],"one_liner":"A robot that moves on wheels — fast and efficient on flat ground, but stopped by stairs and rough terrain.","explanation":"A wheeled robot moves on wheels (or tracks), the most common and technically mature form of mobile robot. On flat, hard ground, rolling wheels lose almost no energy, so they're more efficient, faster, and more stable than legged walking, with simpler mechanics and control; the trade-off is that wheels can't climb stairs and struggle with large gaps or soft, uneven ground. Common base designs include differential drive (two powered wheels that steer by spinning at different rates), omnidirectional or mecanum wheels (which allow sideways movement on the spot), and Ackermann steering (like a car). Examples range from robot vacuums and warehouse AGVs/AMRs, to mobile manipulators with an arm bolted on (like PR2 or Fetch), to wheeled humanoid robots.","example":"A robot vacuum is the most common example of a differential-drive wheeled robot.","related":["Mobile Base (Chassis)","Differential Drive Base","Mecanum Wheel","Wheel-legged Robot","Wheeled Humanoid Robot","Autonomous Mobile Robot"]},{"id":"automated-guided-vehicle","category":"robot","sec":1,"tier":3,"sources":[{"title":"Automated guided vehicle - Wikipedia","url":"https://en.wikipedia.org/wiki/Automated_guided_vehicle"},{"title":"AGV vs AMR - AGV Network","url":"https://www.agvnetwork.com/agv-vs-amr"}],"as_of":"","related_ids":["autonomous-mobile-robot","mobile-base","wheeled-robot","tote-handling","fleet-management-system","simultaneous-localization-and-mapping"],"name":"Automated Guided Vehicle","alt":"自动导引车","abbr":"AGV","aliases":["AGV","AGV Cart"],"one_liner":"A cart that automatically follows a fixed, pre-laid path, such as a magnetic strip or QR codes.","explanation":"AGVs were the first automated material-handling equipment to become common in factories and warehouses, following fixed or semi-fixed routes laid out in advance, with common guidance methods including floor-mounted magnetic strips, buried wires, QR codes, and laser reflectors. Their advantage is a defined, reliable, inexpensive route; the drawback is that changing the route means re-laying the guide path, and when an AGV meets an obstacle it typically just stops and waits rather than routing around it. A vehicle that can build its own map and plan routes in real time is usually called an AMR, meaning autonomous mobile robot, instead, though many AGVs now also use laser SLAM, meaning simultaneous localization and mapping, blurring the line between the two categories. AGVs are a good starting point for understanding mobile bases and warehouse-logistics automation.","example":"On an auto final-assembly line, tugger AGVs follow floor-mounted magnetic strips to deliver parts carts to each workstation.","related":["Autonomous Mobile Robot","Mobile Base (Chassis)","Wheeled Robot","Tote Handling","Fleet Management System (e.g. Open-RMF)","Simultaneous Localization and Mapping"]},{"id":"autonomous-mobile-robot","category":"robot","sec":1,"tier":3,"sources":[{"title":"AGV vs AMR - AGV Network","url":"https://www.agvnetwork.com/agv-vs-amr"},{"title":"AGV vs AMR: differences and how to choose | Navitec Systems","url":"https://navitecsystems.com/agv-vs-amr-what-is-the-difference-and-which-one-to-choose/"}],"as_of":"","related_ids":["automated-guided-vehicle","mobile-base","mobile-manipulator","navigation","simultaneous-localization-and-mapping","fleet-management-system"],"name":"Autonomous Mobile Robot","alt":"自主移动机器人","abbr":"AMR","aliases":["AMR"],"one_liner":"A mobile robot that localizes and plans its own route with sensors, without following a fixed floor track.","explanation":"An AMR uses sensors such as lidar, cameras, and an IMU, meaning an inertial measurement unit, together with SLAM, meaning simultaneous localization and mapping, to know where it is and plan its own path, routing around people or obstacles in real time instead of stopping and waiting the way a traditional AGV does. Deploying one usually requires no changes to the floor, so switching sites or routes just means remapping and resetting waypoints, which is why AMRs have become increasingly common in e-commerce warehouses, hospital delivery, and factory material flow. Their downside is a typically higher unit price than a magnetic-strip AGV. An AMR does not necessarily carry a robot arm; adding one on top turns it into a mobile manipulator, which is also a common hardware form for mobile manipulation in embodied AI.","example":"In an e-commerce warehouse, laser-guided transport robots carry shelves or totes to picking stations.","related":["Automated Guided Vehicle","Mobile Base (Chassis)","Mobile Manipulator","Navigation","Simultaneous Localization and Mapping","Fleet Management System (e.g. Open-RMF)"]},{"id":"mobile-manipulator","category":"robot","sec":1,"tier":2,"sources":[{"title":"The Robot Report: Stretch 3 mobile manipulator","url":"https://www.therobotreport.com/stretch-3-mobile-manipulator-hello-robot-designed-open-source-development/"}],"as_of":"","related_ids":["mobile-manipulation","mobile-base","autonomous-mobile-robot","wheeled-humanoid-robot","hello-robot-stretch","mobile-aloha"],"name":"Mobile Manipulator","alt":"复合机器人","abbr":"","aliases":[],"one_liner":"A robot combining a mobile base with a robot arm, able to move and do work with its hands.","explanation":"A mobile manipulator pairs a mobile base, such as an AGV/AMR or a wheeled chassis, with a robot arm; the term “compound robot” (复合机器人) is common in China's robotics industry, while academic English usually just says “mobile manipulator.” A robot arm alone can only work at a fixed station, and a mobile robot alone can only carry things; combining the two lets a robot travel to different spots in a workshop, warehouse, or home to grasp, load and unload parts, or open doors and drawers. The hard part is that the base and arm must be planned and controlled together, and navigation errors propagate to the end effector. Common academic platforms include Hello Robot Stretch, TIAGo, Fetch, and Mobile ALOHA; many wheeled humanoid robots are, at their core, also this kind of system.","example":"In a factory, an AMR carrying a collaborative arm on its back drives between several machine tools, automatically loading and unloading parts.","related":["Mobile Manipulation","Mobile Base (Chassis)","Autonomous Mobile Robot","Wheeled Humanoid Robot","Hello Robot Stretch","Mobile ALOHA"]},{"id":"legged-robot","category":"robot","sec":1,"tier":2,"sources":[{"title":"IEEE Spectrum: How MIT's Mini Cheetah Can Help Accelerate Robotics Research","url":"https://spectrum.ieee.org/mit-mini-cheetah-accelerate-research"}],"as_of":"","related_ids":["quadruped-robot","bipedal-robot","legged-locomotion","rl-based-locomotion-control","wheel-legged-robot","mit-mini-cheetah"],"name":"Legged Robot","alt":"足式机器人","abbr":"","aliases":[],"one_liner":"A robot that moves on legs instead of wheels, such as a quadruped or bipedal humanoid.","explanation":"A legged robot moves by alternately placing its legs on the ground, and is classified by leg count as bipedal (humanoid), quadruped (robot dog), hexapod, and so on. Compared with wheeled robots, it only needs discrete footholds, so it can climb stairs, cross obstacles, and cover gravel and grass, making it suited to unstructured environments; the tradeoff is that control is hard, since it must keep its balance in real time. Early methods relied on model-based control, such as MPC or the zero moment point; in recent years the mainstream approach trains a locomotion policy with reinforcement learning in simulation, then transfers it to the real robot. Representative examples include Boston Dynamics' Spot, Unitree's Go2, ANYmal, MIT's Mini Cheetah, and various humanoid robots.","example":"A Unitree Go2 quadruped walks over stairs and grass using a gait trained with reinforcement learning.","related":["Quadruped Robot","Bipedal Robot","Legged Locomotion","RL-based Locomotion Control","Wheel-legged Robot","MIT Mini Cheetah"]},{"id":"quadruped-robot","category":"robot","sec":1,"tier":1,"sources":[{"title":"Spot | Boston Dynamics","url":"https://bostondynamics.com/products/spot/"}],"as_of":"","related_ids":["legged-locomotion","boston-dynamics-spot","unitree-go2","anybotics-anymal","rl-based-locomotion-control","sim-to-real-transfer"],"name":"Quadruped Robot","alt":"四足机器人","abbr":"","aliases":["Robot Dog"],"one_liner":"A legged robot that walks on four legs, commonly known as a robot dog.","explanation":"A quadruped robot moves on four legs, commonly with 3 joints per leg, for about 12 degrees of freedom across the whole body. Compared with a biped, it always has multiple legs on the ground, giving it inherently better stability; compared with a wheeled robot, it can cross steps, rubble, and slopes. This made it the first type of legged robot to reach commercial scale, used mainly for industrial inspection, security, surveying, and research and education; adding a robot arm on its back turns it into a mobile manipulator. In recent years, the dominant approach has been to train a locomotion policy with reinforcement learning in simulation, then transfer it to the real robot. Representative examples include Boston Dynamics' Spot, ANYbotics' ANYmal, Unitree's Go2, and DEEP Robotics' Jueying series.","example":"A locomotion gait trained with reinforcement learning in simulation for a Unitree Go2 can be deployed directly onto the real robot.","related":["Legged Locomotion","Boston Dynamics Spot","Unitree Go2","ANYbotics ANYmal","RL-based Locomotion Control","Sim-to-Real Transfer"]},{"id":"wheel-legged-robot","category":"robot","sec":1,"tier":2,"sources":[{"title":"Learning robust autonomous navigation and locomotion for wheeled-legged robots (Science Robotics)","url":"https://arxiv.org/html/2405.01792v1"},{"title":"LimX Dynamics launches W1 wheeled quadruped (The Robot Report)","url":"https://www.therobotreport.com/limx-dynamics-launches-w1-wheeled-quadruped/"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Wheel-legged Robot","alt":"轮足机器人","abbr":"","aliases":["Wheeled-legged Robot"],"one_liner":"A robot with a driven wheel at the end of each leg, rolling on flat ground and stepping over obstacles.","explanation":"A wheel-legged robot puts a motorized wheel at the end of every leg on a legged robot, combining the advantages of both forms of locomotion: on flat ground it rolls on its wheels, which is faster and more energy-efficient than walking; when it meets a step or ditch, it locks the wheel and uses the leg like a foot, or lifts the leg to step over. Common forms include two-wheel-legged designs, such as ETH Zurich's Ascento and LimX Dynamics' TRON 1, and four-wheel-legged designs, such as ANYmal's wheeled variant and its derivative Swiss-Mile, LimX Dynamics' W1, and Unitree's Go2-W and B2-W. The hard part is control, since wheel rotation and leg posture must be coordinated together; most implementations now use reinforcement learning to train the control policy in simulation.","example":"Swiss-Mile's four-wheel-legged robot rolls along a sidewalk at speed on its wheels to make a delivery, then switches to stepping when it reaches a staircase.","related":["Legged Robot","Wheeled Robot","Unitree Go2-W","Unitree B2-W","LimX Dynamics TRON 1","Form-Factor Debate"]},{"id":"legged-mobile-manipulator","category":"robot","sec":1,"tier":3,"sources":[{"title":"Unitree Robotics Launches Z1 Robot Arm","url":"https://shop.unitree.com/blogs/news/unitree-technology-launches-z1-robot-arm-for-its-quadruped-robots"},{"title":"Unitree quadruped robots get a helping hand - New Atlas","url":"https://newatlas.com/robotics/unitree-quadruped-robots-z1-arm/"}],"as_of":"","related_ids":["quadruped-robot","mobile-manipulation","loco-manipulation","whole-body-control","umi-on-legs","boston-dynamics-spot"],"name":"Legged Mobile Manipulator (Quadruped with Arm)","alt":"足式移动操作机器人（带臂四足）","abbr":"","aliases":["Quadruped with Arm","Arm-Equipped Quadruped","Robot Dog with Arm"],"one_liner":"A quadruped robot with a robot arm mounted on its back, combining rough-terrain walking with manipulation.","explanation":"This category mounts a robot arm on top of a quadruped robot, so it can open doors, press buttons, or pick things up wherever it walks. Well-known examples include Boston Dynamics' Spot fitted with the Spot Arm, ANYbotics' ANYmal with an arm attached, and Unitree quadrupeds paired with a Z1 arm. Compared with a wheeled mobile manipulator, this design can climb stairs and cross rubble; the difficulty is that as soon as the arm moves, it shifts the robot's center of mass and the forces on its body, so the arm and all four legs have to be controlled together, or the robot can become unstable or miss its target. For that reason, research in this area often uses whole-body control or reinforcement learning to treat the legs and arm as one coordinated system — work like UMI on Legs was done on exactly this kind of platform.","example":"A Spot with an arm attached walks into a factory, opens a valve and a door by itself, and continues on its inspection route.","related":["Quadruped Robot","Mobile Manipulation","Loco-manipulation","Whole-Body Control","UMI on Legs","Boston Dynamics Spot"]},{"id":"bipedal-robot","category":"robot","sec":1,"tier":2,"sources":[{"title":"Humanoid robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Humanoid_robot"},{"title":"Zero moment point - Wikipedia","url":"https://en.wikipedia.org/wiki/Zero_moment_point"}],"as_of":"","related_ids":["bipedal-locomotion","humanoid-robot","legged-robot","zero-moment-point","inverted-pendulum-model","agility-robotics-cassie"],"name":"Bipedal Robot","alt":"双足机器人","abbr":"","aliases":["Biped","Two-legged Robot"],"one_liner":"A robot that moves on two legs, the most typical lower-body form for a humanoid.","explanation":"A bipedal robot moves on two legs, and can be either a full humanoid or just a pair of legs, as with Agility's Cassie. Two legs give a small support base, and walking often passes through phases with only one foot on the ground, making the robot inherently unstable and requiring continuous balance control. Traditional methods plan gaits around the zero moment point, a reference point that predicts whether the foot will tip over, and an inverted pendulum model; Honda's ASIMO is the classic example. In recent years the mainstream approach has shifted to training walking policies with reinforcement learning in simulation, then transferring them to the real robot. The benefit is that a biped can climb stairs, step over obstacles, and enter environments built for humans, at the cost of higher energy use, cost, and fall risk than a wheeled base. Unitree's G1, Agility's Digit, and Booster's T1 are all bipedal robots.","example":"Honda's ASIMO planned its gait around the zero moment point; Unitree's G1, by contrast, walks using a policy trained with reinforcement learning in simulation.","related":["Bipedal Locomotion","Humanoid Robot","Legged Robot","Zero Moment Point","Inverted Pendulum Model (IPM)","Agility Robotics Cassie"]},{"id":"humanoid-robot","category":"robot","sec":2,"tier":1,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"title":"Unitree G1","url":"https://www.unitree.com/g1/"}],"as_of":"","related_ids":["full-size-humanoid-robot","wheeled-humanoid-robot","bipedal-locomotion","whole-body-control","tesla-optimus","unitree-g1"],"name":"Humanoid Robot","alt":"人形机器人","abbr":"","aliases":[],"one_liner":"A robot with a human-like body: a torso, two arms, and usually two legs.","explanation":"A humanoid robot is a robot whose body structure imitates a human's: generally a head, torso, and two arms, with most walking on two legs, though some use a wheeled base instead of legs (a “wheeled humanoid”). The whole body typically has somewhere between twenty and fifty-plus degrees of freedom, meaning independently movable joints. The case for a human-shaped body is that human environments, such as stairs, door handles, and tools, are all designed for humans, so a humanoid can use them directly, and human video and teleoperation data also transfer to it more easily. The hard parts are bipedal balance, whole-body coordinated control, and cost. Early examples include Honda's ASIMO; recent ones include Boston Dynamics' Atlas, Tesla's Optimus, Figure 03, and Unitree's G1.","example":"Unitree G1 and Tesla Optimus are both bipedal humanoids; Galbot G1 is a wheeled humanoid.","related":["Full-size Humanoid Robot","Wheeled Humanoid Robot","Bipedal Locomotion","Whole-Body Control","Tesla Optimus","Unitree G1"]},{"id":"full-size-humanoid-robot","category":"robot","sec":2,"tier":2,"sources":[{"title":"百度百科：天工（全尺寸人形机器人）","url":"https://baike.baidu.com/item/%E5%A4%A9%E5%B7%A5/64343233"},{"title":"北京人形：天工自主跑完北京亦庄半马","url":"https://x-humanoid.com/news-view-164.html"}],"as_of":"2026-09","related_ids":["humanoid-robot","small-size-humanoid-robot","tiangong","tesla-optimus","unitree-h1","bipedal-locomotion"],"name":"Full-size Humanoid Robot","alt":"全尺寸人形机器人","abbr":"","aliases":[],"one_liner":"A humanoid robot roughly adult height, about 1.5 to 1.8 meters tall.","explanation":"A full-size humanoid robot is one whose height and arm span roughly match an adult human's; the industry generally classes robots taller than about 1.5 meters, closer to 1.6–1.8 meters, in this category, as opposed to small-size humanoids at around 1.3 meters. The point of this size is that the robot can directly use environments built for people: it can reach shelves and workbenches, step over curbs, and use human tools, which is why factory handling and warehouse applications tend to favor this size class. The tradeoffs are greater weight, higher cost, greater fall risk, and tougher demands on joint actuator modules, meaning integrated motor-plus-reducer joints, and on balance control. Examples outside China include Tesla's Optimus, Figure, and Boston Dynamics' Atlas; in China there is the Beijing Humanoid Robot Innovation Center's Tiangong and Unitree's H1.","example":"The Beijing Humanoid Robot Innovation Center's Tiangong Ultra stands 180 cm tall and weighs 52 kg; in April 2025 it completed the Beijing Yizhuang Humanoid Robot Half Marathon in 2 hours, 40 minutes, and 42 seconds.","related":["Humanoid Robot","Small-size Humanoid Robot","Tiangong","Tesla Optimus","Unitree H1","Bipedal Locomotion"]},{"id":"small-size-humanoid-robot","category":"robot","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Humanoid robot","url":"https://en.wikipedia.org/wiki/Humanoid_robot"}],"as_of":"","related_ids":["humanoid-robot","full-size-humanoid-robot",null,"booster-robotics-t1","softbank-robotics-nao","research-and-education-market"],"name":"Small-size Humanoid Robot","alt":"小尺寸人形机器人","abbr":"","aliases":["Half-size Humanoid","Small Humanoid Robot"],"one_liner":"A bipedal humanoid robot noticeably shorter than an adult — lighter, cheaper, and more fall-tolerant.","explanation":"A small-size humanoid robot is a bipedal humanoid noticeably shorter than an adult human. There's no industry-standard cutoff, but the term roughly covers anything under 1.5 m, with 1–1.4 m being the most common range, and educational models can be as small as a few tens of centimeters. Being smaller brings clear advantages: lower weight, lower joint-torque requirements, lower cost and price, and less impact damage when it falls — all of which suit research, education, competitions, and reinforcement-learning locomotion experiments. The trade-off is that it can't reach countertops or shelves and has limited payload, so it isn't suited to factory or household work directly. Over the past couple of years, Chinese manufacturers have pushed this category's price down to the tens of thousands of RMB, making it a common entry-level platform for labs.","example":"Unitree's G1 (about 1.3 m), Booster Robotics' T1, and Noetix's N2 all fall into this category; the earlier NAO was only 58 cm tall.","related":["Humanoid Robot","Full-size Humanoid Robot","Unitree G1","Booster Robotics T1","SoftBank Robotics NAO","Research & Education Market"]},{"id":"wheeled-humanoid-robot","category":"robot","sec":2,"tier":2,"sources":[{"title":"Galbot G1 Specs & Price (Humanoid.guide)","url":"https://humanoid.guide/product/galbot/"},{"title":"1X Technologies - Wikipedia","url":"https://en.wikipedia.org/wiki/1X_Technologies"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Wheeled Humanoid Robot","alt":"轮式人形机器人","abbr":"","aliases":["Wheeled Dual-arm Robot"],"one_liner":"A robot with a human-like upper body but a wheeled base instead of legs.","explanation":"A wheeled humanoid robot keeps the humanoid upper body, a head camera, two arms, and often a waist that can lift or bend, but replaces the two legs with a wheeled mobile base. Giving up bipedal walking trades away the ability to climb stairs or cross rough terrain in exchange for stability, energy efficiency, lower cost, simpler control, and no risk of falling over, along with the ability to work for long stretches without tiring. It targets manipulation tasks on flat ground in factories, warehouses, stores, and homes, and is one of the dominant forms among Chinese embodied-AI companies; examples include Galbot's G1, AgiBot's Expedition A2-W, Galaxea's R1, Astribot's S1, and 1X's early EVE.","example":"Galbot G1 rolls up to a shelf on its wheeled base inside a pharmacy, raises its waist, and uses both arms to pull medicine from a high shelf.","related":["Humanoid Robot","Mobile Manipulator","Mobile Base / Chassis","Galbot G1","Galaxea R1","Form-Factor Debate"]},{"id":"upper-body-humanoid-robot","category":"robot","sec":2,"tier":3,"sources":[{"title":"Baxter (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Baxter_(robot)"}],"as_of":"","related_ids":["humanoid-robot","wheeled-humanoid-robot","dual-arm-robot","bimanual-manipulation","teleoperation","rethink-robotics-baxter"],"name":"Upper-body Humanoid Robot","alt":"半身人形机器人","abbr":"","aliases":[],"one_liner":"A robot with only a humanoid upper body — head, torso, and arms — and no legs.","explanation":"An upper-body humanoid robot keeps only the human-shaped upper half: a head (usually carrying a camera), a torso, and two arms, ending in a dexterous hand or gripper. Its lower half is either a fixed base, a desk-mounted stand, or a wheeled base instead (in which case it's usually called a wheeled humanoid robot). Dropping the legs means no balance or falling to deal with, which lowers cost, improves stability, and suits long stretches of bimanual manipulation work. In embodied AI, it's a common platform for data collection and manipulation research: a person teleoperates it through tasks to gather demonstration data for training policy models like VLA. An early well-known example is Rethink Robotics' Baxter, released in 2012.","example":"Rethink Robotics' Baxter: two arms and a screen displaying a face, bolted to a fixed base while it works.","related":["Humanoid Robot","Wheeled Humanoid Robot","Dual-arm Robot","Bimanual Manipulation","Teleoperation","Rethink Robotics Baxter"]},{"id":"hyper-realistic-humanoid-robot","category":"robot","sec":2,"tier":3,"sources":[{"title":"Latest Geminoid Is Incredibly Realistic - IEEE Spectrum","url":"https://spectrum.ieee.org/latest-geminoid-is-disturbingly-realistic"},{"title":"Hyper-Realistic Humanoids Creep into Mainstream | Mike Kalil","url":"https://mikekalil.com/blog/new-breed-hyper-realistic-humanoid-robots/"}],"as_of":"2025-12","related_ids":["humanoid-robot","uncanny-valley","engineered-arts-ameca","sophia","aheadform","human-robot-interaction"],"name":"Hyper-Realistic Humanoid Robot","alt":"超仿生人形机器人","abbr":"","aliases":["Android","Bionic Humanoid Robot"],"one_liner":"A humanoid robot engineered to look and move as much like a real human as possible.","explanation":"A hyper-realistic humanoid robot — often called an android in robotics research — aims purely to look human: silicone skin and lifelike facial features, with dozens of small motors or artificial muscles hidden in the face to blink, smile, and shape mouth movements for speech. Well-known examples include Hiroshi Ishiguro's Geminoid series, Hanson Robotics' Sophia, Engineered Arts' Ameca, and, in China, AheadForm (首形科技), founded in 2024. They're used mainly for human-robot interaction research, social-psychology studies, museum or exhibition greeting, and companionship; most have very limited walking or manipulation ability, and some rely on teleoperation just to move at all. This is a different design goal from task-oriented humanoids like Optimus, and when the realism falls short it tends to fall into the “uncanny valley,” the sense of unease people feel toward faces that look almost, but not quite, human.","example":"Ameca can shift its expression in real time during conversation — frowning, looking surprised — and often chats with visitors at tech expos, but its lower body barely moves at all.","related":["Humanoid Robot","Uncanny Valley","Engineered Arts Ameca","Sophia (Hanson Robotics)","AheadForm","Human-Robot Interaction"]},{"id":"service-robot","category":"robot","sec":3,"tier":2,"sources":[{"title":"IFR: Service Robots","url":"https://ifr.org/service-robots"},{"title":"IFR: Robot definitions at ISO","url":"https://ifr.org/standardisation"}],"as_of":"","related_ids":["industrial-robot","robot-vacuum-cleaner","companion-robot","autonomous-mobile-robot","human-robot-interaction","international-federation-of-robotics"],"name":"Service Robot","alt":"服务机器人","abbr":"","aliases":[],"one_liner":"A robot that performs useful tasks for people or equipment, as distinct from an industrial robot.","explanation":"Under ISO 8373:2021, which the IFR's statistics follow, a service robot is a robot that performs useful tasks for humans or equipment, excluding industrial automation applications, which are classed separately as industrial robots. The IFR further splits the category into professional service robots, covering logistics, cleaning, medical, agricultural, inspection, and food-delivery robots, and personal or domestic service robots, covering vacuum robots, companion robots, and education or entertainment robots. Most operate in environments that are crowded and constantly changing, so they rely more heavily on perception, navigation, and human-robot interaction. The household, guided-tour, and retail scenarios that today's embodied-AI companies target with humanoid and mobile-manipulation robots mostly fall under this category.","example":"Restaurant food-delivery robots, hospital logistics robots, and household robot vacuums are all service robots.","related":["Industrial Robot","Robot Vacuum Cleaner","Companion Robot","Autonomous Mobile Robot","Human-Robot Interaction","International Federation of Robotics"]},{"id":"robot-vacuum-cleaner","category":"robot","sec":3,"tier":3,"sources":[{"title":"Robotic vacuum cleaner (Wikipedia)","url":"https://en.wikipedia.org/wiki/Robotic_vacuum_cleaner"},{"title":"Electrolux Trilobite (Wikipedia)","url":"https://en.wikipedia.org/wiki/Electrolux_Trilobite"}],"as_of":"","related_ids":["service-robot","consumer-grade-robot","simultaneous-localization-and-mapping","coverage-path-planning","lidar","autonomous-docking-and-recharging"],"name":"Robot Vacuum Cleaner","alt":"扫地机器人","abbr":"","aliases":["Robovac"],"one_liner":"A home robot that autonomously navigates and vacuums or mops floors — the most common consumer robot by far.","explanation":"A robot vacuum is a wheeled home service robot that plans its own route and vacuums or mops as it goes. The first mass-produced model was Electrolux's Trilobite in 2001, but the one that actually caught on was iRobot's Roomba in 2002, which in its early versions simply bounced off obstacles and turned randomly. Today's mainstream models use lidar or vision to do SLAM (building a map of the space and tracking their own position while moving), plan a coverage path room by room, and use structured light or cameras to detect obstacles; many docking stations can now empty the dustbin and wash the mop automatically. It's the best-selling category of home robot by far, and it's also a common reference point for discussing what it takes to get a robot into people's homes: mapping, obstacle avoidance, and autonomous docking are all mature technologies, but actually tidying up objects remains hard — in recent years some manufacturers have started adding robot arms to vacuums.","example":"A typical use: the vacuum first circles the home once, building a floor plan with lidar; the user marks off-limits zones in the app; after that, it cleans room by room on its own and returns to dock automatically when done.","related":["Service Robot","Consumer-Grade Robot","Simultaneous Localization and Mapping","Coverage Path Planning","LiDAR","Autonomous Docking and Recharging"]},{"id":"companion-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"ElliQ - Wikipedia","url":"https://en.wikipedia.org/wiki/ElliQ"},{"title":"Exploring LOVOT robots as companions for older adults - PMC","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC11811964/"}],"as_of":"","related_ids":["service-robot","human-robot-interaction","consumer-grade-robot","to-consumer","large-language-model","uncanny-valley"],"name":"Companion Robot","alt":"陪伴机器人","abbr":"","aliases":["Emotional Companion Robot","Social Companion Robot"],"one_liner":"A service robot mainly meant for emotional company and conversation, not for doing chores.","explanation":"A companion robot is a category of service robot whose main job isn't carrying things or housework, but talking with a person, responding to their emotions, and reminding them of their schedule, commonly used for people living alone, especially the elderly, for children, and as a pet substitute. Representative examples include Paro, the seal-shaped therapeutic robot from Japan's National Institute of Advanced Industrial Science and Technology; ElliQ, Intuition Robotics' desktop companion for older adults; and GROOVE X's LOVOT. They generally don't aim for dexterous manipulation, focusing instead on human-robot interaction, such as voice, expression, and touch feedback, and maintaining a relationship over the long term. Large language models have noticeably improved these products' conversational ability, but whether they can genuinely ease loneliness, or risk creating emotional dependency, remains a focus of research and ethical debate.","example":"ElliQ proactively asks an elderly user how they're feeling today and suggests listening to music or video-calling family, aimed at easing the loneliness of living alone.","related":["Service Robot","Human-Robot Interaction","Consumer-Grade Robot","To Consumer (B2C)","Large Language Model","Uncanny Valley"]},{"id":"surgical-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Robot-assisted surgery","url":"https://en.wikipedia.org/wiki/Robot-assisted_surgery"}],"as_of":"","related_ids":["da-vinci-surgical-system-da-vinci-research-kit","leader-follower-teleoperation","teleoperation","bilateral-teleoperation","special-purpose-robot","service-robot"],"name":"Surgical Robot","alt":"手术机器人","abbr":"","aliases":["Robot-Assisted Surgery System"],"one_liner":"A robot that assists surgeons, usually by letting a doctor teleoperate robotic arms from a console.","explanation":"A surgical robot is a robotic system that assists a surgeon in performing an operation. The dominant design is leader-follower teleoperation: the surgeon sits at a console viewing a magnified 3D endoscopic image and moves a set of master controls, while follower robotic arms beside the patient scale down the motion, filter out hand tremor, and carry it out, guiding slender instruments through small incisions into the body. The best-known example is Intuitive Surgical's da Vinci system, approved by the U.S. FDA for laparoscopic surgery in 2000; there are also specialized systems for orthopedics, neurosurgery, and vascular intervention, including several Chinese-made systems that have received regulatory approval. Almost none of these systems make autonomous decisions today, but they've generated large amounts of teleoperation data, and researchers use the da Vinci Research Kit (dVRK) to explore tasks such as autonomous suturing.","example":"Procedures like prostate removal in urology are among the most common uses of the da Vinci system.","related":["da Vinci Surgical System / da Vinci Research Kit (dVRK)","Leader-Follower Teleoperation","Teleoperation","Bilateral Teleoperation","Special-purpose Robot","Service Robot"]},{"id":"special-purpose-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"百度百科：特种机器人","url":"https://baike.baidu.com/item/特种机器人"}],"as_of":"","related_ids":["inspection-robot","quadruped-robot","dirty-dull-and-dangerous-jobs","unmanned-aerial-vehicle","teleoperation","service-robot"],"name":"Special-purpose Robot","alt":"特种机器人","abbr":"","aliases":["Special-operations Robot"],"one_liner":"A robot that does dangerous or extreme-environment work in place of people — bomb disposal, firefighting, inspection, underwater.","explanation":"“Special-purpose robot” is a classification commonly used in China, standing alongside “industrial robot” and “service robot,” referring to robots that carry out specialized tasks in dangerous, extreme, or hard-to-reach environments. Common types include bomb-disposal robots, firefighting robots, power-grid and utility-tunnel inspection robots, nuclear-plant maintenance robots, mining and tunnel robots, underwater robots, and space robots; military, police, and emergency-response agencies are major users. These robots typically need to be explosion-proof, waterproof, heat-resistant, or radiation-resistant, and many still rely mainly on teleoperation with only partial autonomy. Quadruped robots have seen fast-growing use in inspection, firefighting, and rescue work in recent years, making this one of the earlier areas where embodied robotics has found real-world deployment.","example":"At a substation, a quadruped robot patrols a set route, reads gauges, and checks temperatures, replacing a human's overnight inspection rounds.","related":["Inspection Robot","Quadruped Robot","Dirty, Dull, and Dangerous (3D) Jobs","Unmanned Aerial Vehicle (UAV)","Teleoperation","Service Robot"]},{"id":"inspection-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"2025年中国变电站设备巡检机器人行业市场现状及趋势研判_智研咨询","url":"https://www.chyxx.com/industry/1215922.html"}],"as_of":"2026-05","related_ids":["special-purpose-robot","quadruped-robot","autonomous-mobile-robot","real-world-deployment","thermal-camera","deep-robotics"],"name":"Inspection Robot","alt":"巡检机器人","abbr":"","aliases":["Intelligent Inspection Robot","Patrol Robot"],"one_liner":"A robot that patrols places like substations, tunnels, or equipment rooms to check on equipment instead of a person.","explanation":"Inspection robots are a category of special-purpose robots that move to a series of checkpoints — either along a preset route or by navigating autonomously — and use a visible-light camera to read gauges and meters, an infrared thermal camera to check temperatures, and gas or acoustic sensors to detect leaks or abnormal sounds, raising an alert if something looks off. By how they move, they fall into wheeled, tracked, rail-mounted (running on a track fixed to the ceiling of an equipment room), quadruped, and drone types; wheeled designs are the most mature, while quadrupeds handle rubble, stairs, and other rough terrain better. The power industry adopted them earliest and most widely; Chinese vendors in the space include Yijiahe, State Grid Intelligence, and Shenhao Technology. Inspection is one of the earlier embodied-robotics use cases with a clear, calculable return on investment, and in recent years these robots have started using vision-language models for defect recognition.","example":"A wheeled inspection robot at a substation circles the site on a set schedule every day, reading oil-level gauges and switch states, and uses an infrared thermal camera to spot overheating connectors.","related":["Special-purpose Robot","Quadruped Robot","Autonomous Mobile Robot","Real-world Deployment","Thermal Camera","DEEP Robotics"]},{"id":"unmanned-aerial-vehicle","category":"robot","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Unmanned aerial vehicle","url":"https://en.wikipedia.org/wiki/Unmanned_aerial_vehicle"}],"as_of":"","related_ids":["aerial-manipulation","aerial-vision-and-language-navigation","px4-ardupilot","dji","swarm-intelligence","motion-planning"],"name":"Unmanned Aerial Vehicle (UAV)","alt":"无人机（空中机器人）","abbr":"UAV","aliases":["Aerial Robot","Drone"],"one_liner":"An aircraft with no onboard pilot, flown by remote control or autonomously — essentially a flying robot.","explanation":"An unmanned aerial vehicle (UAV) is an aircraft with no pilot on board; it can be remotely piloted or fly autonomously. Common forms include multirotors (such as quadcopters, which can take off vertically and hover, and are by far the most common), fixed-wing aircraft (longer range and endurance), and helicopter-style designs. In robotics, a UAV is treated as an “aerial robot”: it still needs state estimation, mapping, path planning, and control, but its motion happens in three-dimensional space and it's especially sensitive to weight and energy use. Directions connected to embodied AI include aerial vision-and-language navigation (flying to a target based on spoken instructions), aerial manipulation with an onboard arm, and coordinated multi-drone swarms. The open-source flight-controller stacks PX4 and ArduPilot are common research foundations, while DJI is the leading name in the consumer market.","example":"Telling a quadcopter “fly over to the house with the red roof” and having it plan its own flight path from camera images to get there is an example of aerial vision-and-language navigation.","related":["Aerial Manipulation","Aerial Vision-and-Language Navigation","PX4 / ArduPilot","DJI","Swarm Intelligence","Motion Planning"]},{"id":"bio-inspired-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"Bio-inspired robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Bio-inspired_robotics"}],"as_of":"","related_ids":["quadruped-robot","soft-robot","artificial-muscle","legged-robot","morphological-computation","humanoid-robot"],"name":"Bio-inspired Robot","alt":"仿生机器人","abbr":"","aliases":["Bionic Robot"],"one_liner":"A robot designed by borrowing an animal's body structure or way of moving.","explanation":"Bio-inspired robot isn't a single body form but a design approach: borrowing an animal's structure, movement, or sensing principles and turning them into something engineered to be simpler and more practical to build. Researchers commonly distinguish between straight biomimicry, copying the source closely, and bio-inspiration, learning the underlying principle and then simplifying it. Common sources include quadruped gaits, snake-like undulation, a gecko's adhesive footpads, an octopus's soft grasping tentacles, and the flapping flight of insects and bats. Classic examples include MIT's Cheetah quadruped, the gecko-inspired Stickybot, CMU's snake robots, and various bio-inspired prototypes from Festo. Humanoid and quadruped robots are, broadly speaking, also bio-inspired, while soft robots and artificial muscles extend the idea into materials and actuation.","example":"MIT's Sangbae Kim lab designed its Cheetah series of quadruped robots by studying the cheetah.","related":["Quadruped Robot","Soft Robot","Artificial Muscle","Legged Robot","Morphological Computation","Humanoid Robot"]},{"id":"soft-robot","category":"robot","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Soft robotics","url":"https://en.wikipedia.org/wiki/Soft_robotics"}],"as_of":"","related_ids":["soft-gripper","pneumatic-actuation","artificial-muscle","compliance","deformable-body-simulation","bio-inspired-robot"],"name":"Soft Robot","alt":"软体机器人","abbr":"","aliases":["Soft Robotics"],"one_liner":"A robot built from soft materials like silicone that can deform substantially, instead of rigid links and joints.","explanation":"A soft robot is built mainly from soft materials — silicone, elastomers, and the like — so its body can bend and stretch continuously, rather than moving through rigid links connected by joints. Common actuation methods include pneumatics (inflating internal chambers to make a section bend), tendon-driven cables, and artificial muscles such as shape-memory alloys and dielectric elastomers. The advantage is that it's inherently compliant: it's less likely to injure a person or damage a fragile object on contact, and it can conform to irregularly shaped objects. The difficulty is that its degrees of freedom for deformation are essentially unlimited, which makes modeling, simulation, and precise control much harder than with rigid robots. Right now, the most mature real-world use is the soft gripper, used for handling fruit, food, and similar delicate items.","example":"Harvard's Octobot, released in 2016, was an octopus-shaped robot made entirely of soft materials with no electronic components, powered by gas from an internal chemical reaction.","related":["Soft Gripper","Pneumatic Actuation","Artificial Muscle","Compliance","Deformable-Body Simulation","Bio-inspired Robot"]},{"id":"franka-emika-panda-franka-research-3","category":"robot","sec":4,"tier":1,"sources":[{"title":"Franka Research 3 (franka.de)","url":"https://franka.de/franka-research-3"},{"title":"Agile Robots acquires Franka Emika - The Robot Report","url":"https://www.therobotreport.com/agile-robots-acquires-franka-emika/"}],"as_of":"2025-03","related_ids":["7-dof-robot-arm","joint-torque-sensor","libfranka-franka-control-interface","droid","collaborative-robot","agile-robots"],"name":"Franka Emika Panda / Franka Research 3","alt":"Franka 机械臂（Panda / FR3）","abbr":"FR3","aliases":["Franka","Franka Panda","Panda","FR3","Franka Research 3"],"one_liner":"A German-made 7-DoF force-controlled robot arm, one of the most common single arms in robotics research.","explanation":"The Franka arm was developed by Franka Emika of Munich, Germany; the early model was called Panda, and its research successor is the Franka Research 3 (FR3). It has 7 degrees of freedom, one more than the 6 needed to set an end effector's position and orientation, giving it extra flexibility, and all 7 joints carry torque sensors, enabling compliant force control. It has a 3 kg payload, an 855 mm reach, repeatability of ±0.1 mm, and can be controlled in real time at 1 kHz through its FCI interface. After the company filed for bankruptcy in 2023, it was acquired by Agile Robots and renamed Franka Robotics. Thanks to its strong force control and open interfaces, it has been used to collect large numbers of robot-learning datasets, including DROID and Open X-Embodiment.","example":"The DROID dataset was collected using Franka Panda arms across hundreds of real-world scenes.","related":["7-DoF Robot Arm","Joint Torque Sensor","libfranka / Franka Control Interface (FCI)","DROID (Distributed Robot Interaction Dataset)","Collaborative Robot","Agile Robots"]},{"id":"universal-robots-ur5e","category":"robot","sec":4,"tier":2,"sources":[{"title":"UR5e 技术规格表（Universal Robots）","url":"https://www.universal-robots.com/manuals/EN/TechSheets/UR5e_techsheet_pdf_online/UR5e_techsheet_en.pdf"},{"title":"Universal Robots launches e-Series collaborative robots","url":"https://www.universal-robots.com/about-universal-robots/news-centre/universal-robots-launches-e-series-setting-a-new-standard-for-collaborative-automation-platforms/"}],"as_of":"2023-08","related_ids":[null,null,null,null,null,null],"name":"Universal Robots UR5e","alt":"UR5e 协作机械臂","abbr":"","aliases":["UR5e","UR5"],"one_liner":"A 6-axis, 5 kg-payload cobot from Universal Robots, common in both research and factories.","explanation":"The UR5e is a collaborative robot arm, meaning an arm designed to work safely alongside people, from Danish manufacturer Universal Robots, released in June 2018 as part of its e-Series, succeeding the earlier UR5. It has 6 rotating joints, a 5 kg payload, an 850 mm reach, repeatability of ±0.03 mm, a built-in six-axis force/torque sensor at the wrist, and a touchscreen teach pendant for programming. Because it is reliable and has an open interface, RTDE, which allows real-time reading and writing of joint state, it is one of the most common pieces of hardware in robot-learning labs, and many real-robot datasets and VLA papers have used it.","example":"π0's training data includes manipulation data collected from both single-arm and dual-arm UR5e setups.","related":["Universal Robots","Collaborative Robot","6-Axis Robot Arm","RTDE","Six-Axis Force/Torque Sensor","Franka Emika Panda / Franka Research 3"]},{"id":"trossen-robotics-widowx-250","category":"robot","sec":4,"tier":2,"sources":[{"title":"Trossen Robotics: WidowX 250 S","url":"https://www.trossenrobotics.com/widowx-250"},{"title":"Interbotix 文档：WidowX-250 6DOF","url":"https://docs.trossenrobotics.com/interbotix_xsarms_docs/specifications/wx250s.html"}],"as_of":"","related_ids":["bridgedata-v2","trossen-robotics-viperx-300","aloha","trossen-robotics","robotis-dynamixel-servo","simplerenv"],"name":"Trossen Robotics WidowX 250","alt":"WidowX 250 机械臂","abbr":"","aliases":["WidowX","WidowX 250s","WX250s"],"one_liner":"Trossen Robotics' low-cost 6-DoF research arm, the platform behind the BridgeData datasets.","explanation":"The WidowX 250 is a desktop arm in Trossen Robotics' Interbotix X-series, driven by ROBOTIS Dynamixel X-series servos. The 6-DoF version, the 250s, has a reach of about 650 mm and a rated payload of 250 g, is priced far below an industrial arm, and supports ROS. It is common in the robot-learning community: Berkeley's BridgeData V2 dataset was collected on a WidowX 250, and that data is included in Open X-Embodiment, Octo, and OpenVLA; SimplerEnv also has a matching WidowX evaluation scene, and ALOHA's leader arms use the same WidowX. Its drawbacks are a small payload and only middling precision.","example":"The OpenVLA paper runs its real-robot evaluation on a WidowX 250 in BridgeData-style scenes.","related":["BridgeData V2","Trossen Robotics ViperX 300","ALOHA","Trossen Robotics","ROBOTIS Dynamixel Servo","SimplerEnv"]},{"id":"trossen-robotics-viperx-300","category":"robot","sec":4,"tier":3,"sources":[{"title":"Trossen Robotics: ViperX 300 S","url":"https://www.trossenrobotics.com/viperx-300"}],"as_of":"2026-09","related_ids":["aloha","trossen-robotics-widowx-250","trossen-robotics","robotis-dynamixel-servo","action-chunking-with-transformers","leader-follower-teleoperation"],"name":"Trossen Robotics ViperX 300","alt":"ViperX 300 机械臂","abbr":"","aliases":["ViperX","ViperX 300 S","ViperX 300 6DOF"],"one_liner":"Trossen's low-cost research robot arm, used as the follower arm in the original ALOHA system.","explanation":"ViperX 300 is a small research-and-education robot arm from the U.S. company Trossen Robotics (part of its Interbotix X series), driven by Dynamixel servos. Per Trossen's listing, the 300 S version has 6 degrees of freedom, a 750 mm reach, a 750 g payload (about three times that of the related WidowX 250), and repeats position within 1 mm. It's inexpensive, can be controlled directly through ROS, and has become the hardware base for many imitation-learning projects: Stanford's 2023 ALOHA bimanual teleoperation system used two ViperX 300 arms as the followers and two WidowX 250 arms as the leaders, and the ACT algorithm was first validated on exactly this hardware. Trossen's site shows the standard 300 S as discontinued, with a dedicated ALOHA-specific version now offered instead.","example":"On the ALOHA platform, an operator moves the WidowX leader arms by hand, and the two ViperX 300 arms mirror the motion in real time to complete fine tasks like threading a zip tie or inserting a battery.","related":["ALOHA","Trossen Robotics WidowX 250","Trossen Robotics","ROBOTIS Dynamixel Servo","Action Chunking with Transformers","Leader-Follower Teleoperation"]},{"id":"so-100-so-101-arm","category":"robot","sec":4,"tier":2,"sources":[{"title":"GitHub: TheRobotStudio/SO-ARM100","url":"https://github.com/TheRobotStudio/SO-ARM100"},{"title":"CNX Software: SO-ARM101 open-source dual robotic arm kit","url":"https://www.cnx-software.com/2025/05/02/so-arm101-open-source-dual-robotic-arm-kit-works-with-hugging-faces-lerobot/"}],"as_of":"2025-05","related_ids":["lerobot","leader-follower-teleoperation","feetech-sts3215-servo","action-chunking-with-transformers","smolvla","desktop-robot-arm"],"name":"SO-100 / SO-101 Arm","alt":"SO-100 / SO-101 机械臂","abbr":"","aliases":["SO-100","SO-101","SO-ARM100","SO-ARM101"],"one_liner":"A sub-$300, open-source, 3D-printed robot arm built to pair with Hugging Face's LeRobot.","explanation":"SO-100 is an open-source desktop robot arm co-designed by TheRobotStudio and Hugging Face, with 3D-printed structural parts; each arm uses 6 Feetech STS3215 serial-bus servos, 5 for joints and 1 for the gripper, and a leader-follower pair forms a teleoperation rig: a person moves the leader arm by hand, and the follower arm copies it and logs the data. SO-101, a 2025 revision, improved cable routing, no longer requires disassembling servo gears to build, and swapped in a different-gear-ratio servo for the leader arm. The official bill of materials lists a full leader-follower pair's parts at about $230. It plugs directly into the LeRobot library for data collection and training policies such as ACT and SmolVLA, making it one of the cheapest ways for a newcomer to get hands-on with imitation learning.","example":"Using an SO-101 leader-follower pair to collect 50 demonstrations of “put the block in the box,” then training an ACT policy with LeRobot so the follower arm can do it autonomously.","related":["LeRobot","Leader-Follower Teleoperation","Feetech STS3215 Servo","Action Chunking with Transformers","SmolVLA","Desktop Robot Arm"]},{"id":"koch-v1-1-arm","category":"robot","sec":4,"tier":3,"sources":[{"title":"Koch v1.1 - LeRobot Documentation (Hugging Face)","url":"https://huggingface.co/docs/lerobot/koch"},{"title":"Koch v1.1 Low-Cost Robot Arm: Follower - ROBOTIS","url":"https://www.robotis.us/koch-v1-1-low-cost-robot-arm-follower/"}],"as_of":"2025","related_ids":["lerobot","so-100-so-101-arm","leader-follower-teleoperation","robotis-dynamixel-servo","action-chunking-with-transformers","desktop-robot-arm"],"name":"Koch v1.1 Arm","alt":"Koch v1.1 机械臂","abbr":"","aliases":["Koch Arm","Koch v1.1"],"one_liner":"An open-source, low-cost leader-follower robot arm built from Dynamixel servos and 3D-printed parts.","explanation":"The Koch arm was originally designed by Alexander Koch of Tau Robotics, and Hugging Face engineers later reworked it into the easier-to-build v1.1 version. It has 6 degrees of freedom (5 joints plus a gripper), 3D-printed structural parts, ROBOTIS Dynamixel XL430 and XL330 servos at the joints, and communicates over a USB serial connection. A set consists of a leader arm and a follower arm: a person moves the leader by hand, the follower mirrors it in real time, and the resulting motion is recorded as demonstration data, which is then used to train an imitation-learning policy with an algorithm such as ACT. According to distributors, a single arm can be built for around $400. It was the hardware featured in LeRobot's early tutorials, though the cheaper SO-100/SO-101 arms later became the more common choice.","example":"Following the LeRobot tutorial, someone builds a pair of Koch arms, records 50 demonstrations of “put the block in the box,” and trains an ACT policy on the data.","related":["LeRobot","SO-100 / SO-101 Arm","Leader-Follower Teleoperation","ROBOTIS Dynamixel Servo","Action Chunking with Transformers","Desktop Robot Arm"]},{"id":"kinova-gen3","category":"robot","sec":4,"tier":3,"sources":[{"title":"Discover our Gen3 robotic arm - Kinova","url":"https://www.kinovarobotics.com/product/gen3-robots"},{"title":"Kinova Gen3 Specifications - QVIRO","url":"https://qviro.com/product/kinova/gen3/specifications"}],"as_of":"","related_ids":["robotic-arm","7-dof-robot-arm","collaborative-robot","joint-torque-sensor","franka-emika-panda-franka-research-3","universal-robots-ur5e"],"name":"Kinova Gen3","alt":"Kinova Gen3","abbr":"","aliases":["Gen3"],"one_liner":"An ultra-lightweight 7-DOF research robot arm from Canada's Kinova, popular in academic manipulation labs.","explanation":"Kinova Gen3 is a lightweight robot arm from the Canadian company Kinova, available in 7-DOF and 6-DOF versions. The 7-DOF version weighs about 8.2 kg, has a maximum reach of about 902 mm, and a continuous payload of about 4 kg; it draws very little power, every joint can rotate without limit, torque sensors are built into each joint, the low-level control loop runs at 1 kHz, and an optional 2D/3D vision module can be mounted on the wrist. Kinova originally built assistive arms for wheelchairs, and Gen3 keeps that same lightweight design that's easy to mount on a mobile platform; its interfaces are open (it supports ROS and low-level torque control), which makes it a common choice in university labs for manipulation, human-robot interaction, and reinforcement-learning research, where it's often compared with the Franka and UR5e arms.","example":"Researchers mount a Kinova Gen3 on a mobile base to run mobile-manipulation experiments like opening doors and picking up objects.","related":["Robotic Arm","7-DoF Robot Arm","Collaborative Robot","Joint Torque Sensor","Franka Emika Panda / Franka Research 3","Universal Robots UR5e"]},{"id":"kuka-lbr-iiwa","category":"robot","sec":4,"tier":3,"sources":[{"title":"LBR iiwa | KUKA Global","url":"https://www.kuka.com/en-us/products/robotics-systems/industrial-robots/lbr-iiwa"},{"title":"KUKA LBR iiwa cobot review - Standard Bots","url":"https://standardbots.com/blog/iiwa"}],"as_of":"","related_ids":["collaborative-robot","7-dof-robot-arm","joint-torque-sensor","impedance-control","kuka","contact-rich-manipulation"],"name":"KUKA LBR iiwa","alt":"库卡 LBR iiwa","abbr":"","aliases":["iiwa","LBR iiwa 7 R800","LBR iiwa 14 R820"],"one_liner":"KUKA's sensitive 7-axis collaborative robot arm, with a torque sensor built into every joint.","explanation":"The LBR iiwa is a collaborative robot arm from Germany's KUKA. LBR stands for “lightweight robot,” and iiwa is short for “intelligent industrial work assistant”; the underlying technology traces back to lightweight-robot research at the German Aerospace Center (DLR). It has 7 axes, each fitted with a joint torque sensor, which lets it sense external forces and immediately cut power and slow down if it contacts a person; it also supports impedance control and hand-guided teaching, where an operator physically moves the arm to program a motion. It comes in two models: the iiwa 7 R800, with a 7 kg payload and 800 mm reach, and the iiwa 14 R820, with a 14 kg payload and 820 mm reach. It's widely used in research on contact-rich manipulation, force control, and human-robot collaboration, and is considered a classic platform for that kind of work.","example":"Researchers use the iiwa's Cartesian impedance mode to perform peg-in-hole assembly, letting the arm automatically nudge its position based on contact force as it inserts the peg.","related":["Collaborative Robot","7-DoF Robot Arm","Joint Torque Sensor","Impedance Control","KUKA","Contact-rich Manipulation"]},{"id":"flexiv-rizon","category":"robot","sec":4,"tier":3,"sources":[{"title":"Flexiv Rizon - The Adaptive 7-Axis Robot Arm with Force Control","url":"https://www.flexiv.us/products/rizon"},{"title":"Rizon 4 | World's First 7-axis Adaptive Robot - A3","url":"https://www.automate.org/products/flexiv-robotics/rizon"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"Flexiv Rizon","alt":"非夕 拂晓 Rizon","abbr":"","aliases":["Rizon 4","Rizon 4s"],"one_liner":"Flexiv's 7-axis force-sensing “adaptive” robot arm, well-suited to delicate contact work.","explanation":"Rizon is the 7-axis “adaptive” robot arm series from Flexiv, with the most common model, the Rizon 4, carrying a 4 kg payload and a reach of about 780 mm. Its defining feature is full-body force sensing: every joint has a torque sensor, paired with high-precision force-control algorithms, letting it sense forces as small as about 0.1 N and perform hybrid force/position control, meaning some directions are position-controlled and others force-controlled, plus multi-point collision detection; the “s” variants add a six-axis force/torque sensor at the wrist. Compared with an ordinary position-only industrial arm, it is well-suited to contact-rich work that needs a “sense of touch,” such as polishing, fitting parts together, and assembly. In embodied-AI research, it is often used as an experimental platform for contact-rich manipulation and force-based policy learning.","example":"Doing a peg-in-hole insertion with a Rizon 4: vision first roughly aligns the peg with the hole, and then impedance control feels the contact forces as it pushes the peg the rest of the way in.","related":["Flexiv Robotics","7-DoF Robot Arm","Force Control","Joint Torque Sensor","Contact-rich Manipulation","Hybrid Force/Position Control"]},{"id":"ufactory-xarm","category":"robot","sec":4,"tier":3,"sources":[{"title":"UFACTORY xArm 协作机械臂","url":"https://www.ufactory.cc/xarm-collaborative-robot/"}],"as_of":"2026-09","related_ids":["collaborative-robot","6-axis-robot-arm","7-dof-robot-arm","strain-wave-gear","gello","franka-emika-panda-franka-research-3"],"name":"UFACTORY xArm","alt":"xArm 机械臂","abbr":"","aliases":["xArm6","xArm7","xArm5","Lite 6","xArm 850"],"one_liner":"A budget collaborative robot-arm line from Shenzhen's UFACTORY, a common choice in research labs.","explanation":"xArm is a line of lightweight collaborative robot arms from Shenzhen-based UFACTORY, using harmonic-drive gearboxes and servo motors at the joints, with built-in collision detection, at a price well below imported arms like Franka's or Universal Robots'. Per UFACTORY's site, the xArm 5, 6, and 7 have 5, 6, and 7 degrees of freedom respectively, all with a 700 mm reach, payloads of 3, 5, and 3.5 kg, and repeat positioning within ±0.1 mm; all three are listed starting at about $5,300. The entry-level Lite 6 has a 600 g payload and a 440 mm reach, while the xArm 850 has an 850 mm reach and a 5 kg payload. Because it has a Python SDK and ROS support and is cheap and durable, it's a common choice for university labs doing imitation learning and teleoperation data collection, and open-source teleoperation projects like GELLO offer an xArm-compatible version.","example":"A lab uses a GELLO leader arm to teleoperate an xArm 7 and collect towel-folding demonstrations, then trains a diffusion policy to reproduce the task.","related":["Collaborative Robot","6-Axis Robot Arm","7-DoF Robot Arm","Strain Wave Gear (Harmonic Drive)","GELLO","Franka Emika Panda / Franka Research 3"]},{"id":"rethink-robotics-baxter","category":"robot","sec":4,"tier":3,"sources":[{"title":"Rethink Robotics launches Baxter the Robot (Robohub)","url":"https://robohub.org/rethink-robotics-launches-baxter-the-robot/"},{"title":"Rethink Robotics, Pioneer of Collaborative Robots, Shuts Down (IEEE Spectrum)","url":"https://spectrum.ieee.org/automaton/robotics/industrial-robots/rethink-robotics-pioneer-of-collaborative-robots-shuts-down"}],"as_of":"2018-10","related_ids":["rethink-robotics","rethink-robotics-sawyer","collaborative-robot","dual-arm-robot","series-elastic-actuator","kinesthetic-teaching"],"name":"Rethink Robotics Baxter","alt":"Baxter 双臂机器人","abbr":"","aliases":["Baxter"],"one_liner":"Rethink Robotics' 2012 low-cost dual-arm collaborative robot, with an animated screen for a face.","explanation":"Baxter is a dual-arm robot that Rethink Robotics — founded by iRobot co-founder Rodney Brooks — released in 2012, priced at about $22,000, far below industrial robots of the time. It has two 7-degree-of-freedom arms whose joints use series elastic actuators (a spring placed between the motor and the load, which makes the arm “give” a bit on contact with a person); its head is a screen that displays an animated face; and workers can teach it a motion just by physically grabbing and moving the arm, with no programming required. Baxter was an early example of the “collaborative robot” concept, and because a research edition existed, it became a common platform in robot-learning labs through the 2010s. Commercially, sales fell short of expectations, and Rethink Robotics shut down in 2018, ending Baxter's production.","example":"The RoboNet dataset includes video of Baxter pushing objects around, used to train video-prediction models.","related":["Rethink Robotics","Rethink Robotics Sawyer","Collaborative Robot","Dual-arm Robot","Series Elastic Actuator (SEA)","Kinesthetic Teaching"]},{"id":"rethink-robotics-sawyer","category":"robot","sec":4,"tier":3,"sources":[{"title":"Inside HAHN Group's plan to revive Rethink Robotics (The Robot Report)","url":"https://www.therobotreport.com/hahn-group-rethink-robotics-sawyer-cobot/"},{"title":"Sawyer - ROBOTS: Your Guide to the World of Robotics","url":"https://robotsguide.com/robots/sawyer"}],"as_of":"2018-10","related_ids":["rethink-robotics","rethink-robotics-baxter","collaborative-robot","7-dof-robot-arm","meta-world","robonet"],"name":"Rethink Robotics Sawyer","alt":"Sawyer 机械臂","abbr":"","aliases":["Sawyer"],"one_liner":"Rethink Robotics' 2015 single-arm 7-DOF collaborative robot arm.","explanation":"Sawyer is the single-arm collaborative robot Rethink Robotics released in 2015, following Baxter. It has 7 degrees of freedom and a 4 kg payload, and is smaller, faster, and more precise than Baxter, while keeping the same screen-based head and Intera graphical programming software, plus support for kinesthetic (hand-guided) teaching. After Rethink Robotics shut down in 2018, the German HAHN Group acquired the intellectual property and trademarks for Sawyer and Intera and continued developing and selling them. In robot-learning research, Sawyer was a common arm in labs such as Berkeley and Stanford in the late 2010s, and appears in datasets and benchmarks such as RoboNet and Meta-World, where the simulated arm model is Sawyer's.","example":"All 50 manipulation tasks in the Meta-World benchmark — opening a drawer, pressing a button, and so on — are defined on a simulated Sawyer arm.","related":["Rethink Robotics","Rethink Robotics Baxter","Collaborative Robot","7-DoF Robot Arm","Meta-World","RoboNet"]},{"id":"abb-yumi","category":"robot","sec":4,"tier":3,"sources":[{"title":"IRB 14000 YuMi Dual Arm（ABB 官方）","url":"https://one.robotics.abb.com/en/robots/p/IRB-14000-YuMi-Dual-Arm"}],"as_of":"2015","related_ids":[null,null,null,null,null,null],"name":"ABB YuMi","alt":"ABB YuMi 双臂协作机器人","abbr":"","aliases":["YuMi","IRB 14000","IRB 14050"],"one_liner":"ABB's 2015 dual-arm collaborative robot, built for delicate small-parts assembly.","explanation":"YuMi, model IRB 14000, is the dual-arm collaborative robot ABB launched in 2015; the name comes from “you and me,” emphasizing people and robots working side by side. It packs two 7-axis robot arms onto a compact torso, each with a 0.5 kg payload, about a 0.56-meter working radius, and repeatability of about 0.02 mm; its housing is padded and force-limited, so it can work alongside people without a safety fence. It is used mainly for delicate two-handed assembly work in electronics and small parts, and a single-arm version, the IRB 14050, was released later. YuMi was also one of the earlier commercial platforms researchers used to run bimanual manipulation experiments.","example":"On an electronics line, YuMi holds a phone housing steady with one arm while the other presses a small part into its slot.","related":["ABB Robotics","Collaborative Robot","Dual-arm (Bimanual) Robot","7-DoF Robot Arm","Robotic Assembly","Human-Robot Collaboration"]},{"id":"agilex-piper","category":"robot","sec":4,"tier":3,"sources":[{"title":"AgileX 官网：PiPER","url":"https://global.agilex.ai/products/piper"},{"title":"Generation Robots：6-Axis Robotic Arm PiPER","url":"https://www.generationrobots.com/en/404258-6-axis-robotic-arm-piper.html"},{"title":"TechEBlog：AgileX PiPER Robotic Arm Costs $2,499","url":"https://www.techeblog.com/agilex-piper-robotic-arm-human-precision/"}],"as_of":"2026-09","related_ids":["agilex-robotics","6-axis-robot-arm","desktop-robot-arm","agilex-cobot-magic","pose-repeatability","leader-follower-teleoperation"],"name":"AgileX PiPER","alt":"松灵 PiPER 机械臂","abbr":"","aliases":["PiPER","PiPER-X"],"one_liner":"AgileX's lightweight, low-cost 6-axis arm, common in research and data collection.","explanation":"PiPER is the lightweight 6-DoF robot arm from AgileX Robotics. Published specs: 4.2 kg unloaded, a 1.5 kg payload, a 626 mm reach, repeatability of 0.1 mm, and support for a Python SDK, ROS1, and ROS2. It is inexpensive, reportedly starting around $2,499 overseas, and light enough to mount on a mobile base or pair up into a bimanual system, which is why it sees heavy use in embodied-AI research and data collection; AgileX's own Cobot Magic dual-arm platform uses it as both leader and follower arms. Its precision and payload fall short of an industrial cobot, making it better suited to labs and teaching than production work.","example":"Mounting two PiPER arms on either side of a desk, paired with leader arms, to collect bimanual teleoperation data.","related":["AgileX Robotics","6-Axis Robot Arm","Desktop Robot Arm","AgileX Cobot Magic","Pose Repeatability","Leader-Follower Teleoperation"]},{"id":"arx-x5-r5-arm","category":"robot","sec":4,"tier":3,"sources":[{"title":"方舟无限新品发布（机器人大讲堂）","url":"https://www.leaderobot.com/news/4422"},{"title":"real-stanford/arx5-sdk","url":"https://github.com/real-stanford/arx5-sdk"}],"as_of":"2026-09","related_ids":["robotic-arm","6-axis-robot-arm","agilex-piper","umi-on-legs","leader-follower-teleoperation","desktop-robot-arm"],"name":"ARX X5 / R5 Arm","alt":"方舟无限 ARX X5 / R5 机械臂","abbr":"","aliases":["ARX Arm","ARX5"],"one_liner":"A lightweight, force-controlled 6-axis arm from Beijing's ARX, popular for imitation-learning research.","explanation":"ARX was founded in Beijing in 2023; the X5, also called ARX5 early on, is its ultra-lightweight, force-controlled 6-DoF arm, weighing about 3.3 kg including the gripper, with a rated payload of about 2 kg and a reported launch price of about RMB 60,000 per unit. R5 is a later 6-DoF model, and there is also a cheaper L5 series, reportedly around RMB 29,800 per unit. Its selling points are being light, inexpensive, and force-controllable at every joint, making it easy to mount on a mobile base or a quadruped, and convenient for building leader-follower teleoperation rigs. Stanford's REAL lab open-sourced arx5-sdk, a C++/Python control interface, and studies such as UMI on Legs and UVA have used it in experiments, giving it heavy exposure in embodied-AI data collection and policy replication work.","example":"UMI on Legs mounts an ARX5 arm on the back of a Unitree quadruped, training a mobile-manipulation policy on data collected with a handheld gripper.","related":["Robotic Arm","6-Axis Robot Arm","AgileX PiPER","UMI on Legs","Leader-Follower Teleoperation","Desktop Robot Arm"]},{"id":"realman-rm-series-arm","category":"robot","sec":4,"tier":3,"sources":[{"title":"睿尔曼智能 | RM65","url":"https://www.realman-robotics.cn/cn/products/rm65.html"},{"title":"本体参数：RM75 系列参数及 D-H 模型（睿尔曼开发者文档）","url":"https://develop.realman-robotics.com/robot/robotParameter/RM75OntologyParameters/"}],"as_of":"2026-09","related_ids":["realman-robotics","collaborative-robot","6-axis-robot-arm","7-dof-robot-arm","payload-to-weight-ratio","mobile-manipulator"],"name":"RealMan RM Series Arm","alt":"睿尔曼 RM 系列机械臂","abbr":"","aliases":["RM65","RM75"],"one_liner":"RealMan's ultra-lightweight humanlike robot arms, with the controller built into the arm itself.","explanation":"The RM series comprises lightweight collaborative robot arms from Beijing-based RealMan Intelligent Technology, which the company markets as “ultra-lightweight humanlike arms.” Its flagship models are the 6-axis RM65 and the 7-axis RM75. Taking the B variant as an example, the RM65-B weighs about 7.2 kg, has a rated payload of 5 kg, and a working radius of about 610 mm; the RM75-B weighs about 7.8 kg, also has a rated payload of 5 kg, and repeats position within ±0.05 mm. Its defining feature is a controller built into the arm body itself, running on 24V DC power, giving it a high payload-to-weight ratio — which makes it easy to mount on a mobile base, a wheeled humanoid, or a dual-arm platform. Many Chinese embodied-AI companies and universities use it as the arm on their dual-arm or mobile-manipulator systems.","example":"Mounting two RM75 arms on a lift column and a mobile base is enough to assemble a dual-arm mobile manipulator for collecting teleoperation data.","related":["RealMan Robotics","Collaborative Robot","6-Axis Robot Arm","7-DoF Robot Arm","Payload-to-Weight Ratio","Mobile Manipulator"]},{"id":"i2rt-yam-arm","category":"robot","sec":4,"tier":3,"sources":[{"title":"YAM Arm Series | I2RT Robotics","url":"https://doc.i2rt.com/products/yam"},{"title":"YAM Pro – I2RT Robotics","url":"https://i2rt.com/products/yam-pro-6-dof-arm-copy"}],"as_of":"2026-09","related_ids":["robotic-arm","dual-arm-robot","leader-follower-teleoperation","molmoact2-bimanualyam","desktop-robot-arm"],"name":"I2RT YAM Arm","alt":"I2RT YAM 机械臂","abbr":"","aliases":["YAM","YAM Arm"],"one_liner":"A low-cost 6-DOF research robot arm from I2RT, commonly paired up for bimanual data collection.","explanation":"YAM is a 6-degree-of-freedom robot arm from the U.S. startup I2RT Robotics, built around a CAN bus and aimed at embodied-AI research and data collection. It comes in Standard, Pro, and Ultra tiers plus a higher-payload BIG YAM variant; the Pro, for example, has a rated payload of about 3 kg, a working range of about 750 mm, and a parallel-jaw gripper with roughly 95 mm of opening width. Pricing runs in the low thousands of dollars (I2RT lists about $3,000–5,000, with the Pro at $3,499) — roughly an order of magnitude cheaper than a traditional research arm like Franka's, while being sturdier than servo-driven arms like the SO-101. It's often paired up two at a time into bimanual platforms, with one arm teleoperated as a leader and the other following, to record demonstration data. The BimanualYAM dataset that Ai2 released was collected this way.","example":"Ai2 builds a bimanual platform out of two YAM arms and uses it to collect the MolmoAct2-BimanualYAM dataset.","related":["Robotic Arm","Dual-arm Robot","Leader-Follower Teleoperation","MolmoAct2-BimanualYAM","Desktop Robot Arm"]},{"id":"openarm","category":"robot","sec":4,"tier":3,"sources":[{"title":"OpenArm Project Overview","url":"https://docs.openarm.dev/"},{"title":"enactic/openarm - GitHub","url":"https://github.com/enactic/openarm"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"OpenArm","alt":"OpenArm","abbr":"","aliases":["OpenArm 2.0","Enactic OpenArm"],"one_liner":"An open-source, 7-DoF humanoid robot arm from Enactic, aimed at data collection for contact-rich tasks.","explanation":"OpenArm is an open-source humanoid robot arm project built and maintained by the company Enactic, with the hardware CAD, firmware, control software, and simulation models all published openly. Each arm has 7 degrees of freedom, scaled to fit a person roughly 160–165 cm tall, using backdrivable, quasi-direct-drive joints, meaning joints a motor can be pushed by an external force, for greater safety on contact, running on a CAN-FD bus; the official specs give a rated payload of 4.1 kg and a peak of 6.0 kg. The project also includes a motor-free leader arm matching the follower's kinematics, used for leader-follower and bilateral teleoperation with force feedback. According to its GitHub documentation, a complete bimanual system costs about $6,500. It is positioned as an open-source dual-arm platform for labs doing imitation learning and VLA data collection, cheaper than an industrial cobot yet more capable than a small servo-driven arm. As of September 2026, the official documentation lists the current version as 2.0.","example":"A lab builds two sets of OpenArm bimanual rigs, using the matching leader arms to teleoperate and collect laundry-folding demonstrations, then trains a diffusion policy to replay them on the follower arms.","related":["Open-source Hardware","7-DoF Robot Arm","Dual-arm (Bimanual) Robot","Leader-Follower Teleoperation","Quasi-Direct Drive","Contact-rich Manipulation"]},{"id":"dobot-magician","category":"robot","sec":4,"tier":3,"sources":[{"title":"Dobot Magician Specifications (RobotLAB PDF)","url":"https://www.robotlab.com/hubfs/Dobot/Dobot%20Magician%20Specifications.pdf?t=1515807244104"},{"title":"Dobot Magician 4-Axis Robotic Arm - Rapid Electronics","url":"https://www.rapidonline.com/dobot-magician-4-axis-robotic-arm-for-education-and-research-70-0480"}],"as_of":"2026-09","related_ids":[null,null,null,null,null],"name":"Dobot Magician","alt":"越疆 Dobot Magician","abbr":"","aliases":[],"one_liner":"A 4-axis desktop arm from Dobot, built for education and entry-level experiments.","explanation":"Dobot Magician is the desktop-class robot arm from Shenzhen-based Dobot, with 4 axes, meaning degrees of freedom, or independently rotating joints. Published specs: a 500 g maximum payload, a 320 mm maximum reach, repeatability of about 0.2 mm, and a body weighing about 4 kg. It ships with a gripper, suction cup, and pen as interchangeable end effectors, meaning the tools attached at the arm's tip to do the work, and can do pick-and-place, writing and drawing, laser engraving, and 3D printing, with both graphical programming and interfaces such as Python. Its precision and payload are both limited, so it can't do real industrial work, but its low price and small footprint mean it just sits on a desk, and it is widely used by universities and secondary schools for introductory lessons on robot kinematics, pick-and-place, and vision-based sorting.","example":"In a classroom, a Dobot Magician is paired with a camera to recognize different-colored blocks on the table and sort them into matching zones with its suction cup.","related":["Desktop Robot Arm","Dobot","Research & Education Market","Pick-and-Place","Elephant Robotics myCobot"]},{"id":"elephant-robotics-mycobot","category":"robot","sec":4,"tier":3,"sources":[{"title":"myCobot 280 - Elephant Robotics Shop","url":"https://shop.elephantrobotics.com/collections/mycobot-280"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"Elephant Robotics myCobot","alt":"大象机器人 myCobot","abbr":"","aliases":["myCobot","myCobot 280"],"one_liner":"A small, inexpensive 6-axis desktop collaborative arm from Elephant Robotics, easy to pick up.","explanation":"myCobot is the desktop 6-axis collaborative arm series from Shenzhen-based Elephant Robotics; the most common version, the myCobot 280, weighs about 850 g, has a 280 mm working radius, a 250 g payload, and repeatability of about ±0.5 mm, available with different main-controller options, including M5Stack, Raspberry Pi, Arduino, and Jetson. It supports ROS, meaning Robot Operating System, Python, C++, and graphical programming, and is commonly used to learn forward and inverse kinematics, run ROS/MoveIt experiments, and do simple vision-guided grasping. Compared with the Dobot Magician, it has two extra axes, giving the end effector more orientation flexibility, but its payload and precision are both low, so it is really only suited to teaching, demos, and prototyping.","example":"Using Python to drive a myCobot 280 through a series of preset waypoints, paired with a gripper to move small blocks on a table to a target position.","related":["Desktop Robot Arm","6-Axis Robot Arm","Collaborative Robot","Dobot Magician","SO-100 / SO-101 Arm (LeRobot)","Robot Operating System"]},{"id":"aloha","category":"robot","sec":5,"tier":1,"sources":[{"title":"ALOHA project page (Tony Zhao et al.)","url":"https://tonyzhaozh.github.io/aloha/"}],"as_of":"2023-04","related_ids":["action-chunking-with-transformers","mobile-aloha","aloha-2","leader-follower-teleoperation","bimanual-manipulation","trossen-robotics-viperx-300"],"name":"ALOHA","alt":"ALOHA 双臂平台","abbr":"ALOHA","aliases":["A Low-cost Open-source Hardware System for Bimanual Teleoperation"],"one_liner":"Stanford's open-source, low-cost bimanual teleoperation platform, priced at around $20,000.","explanation":"ALOHA is an open-source bimanual (two-armed) robot hardware platform released in 2023 by Tony Zhao, Chelsea Finn, and colleagues at Stanford, with collaborators from Berkeley and Meta, alongside the paper “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.” A full setup costs about $20,000. It uses two small leader arms, moved by hand, to drive two matching follower arms for teleoperation data collection, along with four cameras (two fixed, two wrist-mounted). The same paper introduced the ACT algorithm, which could learn precise bimanual tasks such as threading a zip tie or inserting a battery from only about 10 minutes of demonstrations. ALOHA lowered the barrier to collecting bimanual manipulation data, and later spawned variants such as Mobile ALOHA and ALOHA 2; it has also become a common real-robot platform for many VLA papers.","example":"Both π0 and ACT have been demonstrated on ALOHA-style bimanual platforms performing tasks such as folding laundry and inserting batteries.","related":["Action Chunking with Transformers","Mobile ALOHA","ALOHA 2","Leader-Follower Teleoperation","Bimanual Manipulation","Trossen Robotics ViperX 300"]},{"id":"aloha-2","category":"robot","sec":5,"tier":3,"sources":[{"title":"arXiv 2405.02292: ALOHA 2: An Enhanced Low-Cost Hardware for Bimanual Teleoperation","url":"https://arxiv.org/abs/2405.02292"}],"as_of":"2024-05","related_ids":["aloha","mobile-aloha","aloha-unleashed","leader-follower-teleoperation","google-deepmind","gemini-robotics"],"name":"ALOHA 2","alt":"ALOHA 2","abbr":"","aliases":["ALOHA2"],"one_liner":"Google DeepMind's improved, open-source, low-cost bimanual teleoperation hardware, ALOHA's second generation.","explanation":"ALOHA 2 is open-source bimanual teleoperation hardware released in February 2024 and posted to arXiv in May of that year, led by an “ALOHA 2 Team” at Google DeepMind that included researchers such as Chelsea Finn, and is an improved version of Stanford's original ALOHA. It keeps the same low-cost “move the leader arm by hand, the follower arm copies it” approach, with the main improvements focused on ergonomics and durability: the force needed to operate the leader arm's gripper dropped to about a tenth of before, the follower arm's grip strength doubled, and the data-collection software runs on ROS 2, logging data at 50 Hz. The full hardware design is open-sourced, along with a system-identified MuJoCo simulation model. The bimanual data behind both ALOHA Unleashed and Gemini Robotics was collected on this kind of platform.","example":"The folding-paper and lunchbox-packing bimanual tasks demonstrated by Gemini Robotics were performed on an ALOHA 2.","related":["ALOHA","Mobile ALOHA","ALOHA Unleashed","Leader-Follower Teleoperation","Google DeepMind","Gemini Robotics"]},{"id":"agilex-cobot-magic","category":"robot","sec":5,"tier":3,"sources":[{"title":"松灵官网：COBOT MAGIC 基于 Mobile ALOHA 架构的双臂遥操作平台","url":"https://www.agilex.ai/page/690aef2d5e78cfa260412cb5"},{"title":"无人系统网：首个「Mobile ALOHA」国产平替，松灵发布Cobot Magic","url":"https://m.youuvs.com/news/detail/202401/41607.html"}],"as_of":"2026-09","related_ids":["mobile-aloha","agilex-robotics","agilex-piper","leader-follower-teleoperation","action-chunking-with-transformers","rdt-1b"],"name":"AgileX Cobot Magic","alt":"松灵 Cobot Magic","abbr":"","aliases":["Cobot Magic"],"one_liner":"AgileX's commercial platform modeled on Stanford's Mobile ALOHA, for mobile bimanual data collection.","explanation":"Cobot Magic is a commercialized platform AgileX Robotics built following the architecture of Stanford's Mobile ALOHA, opened for preorder in early 2024 and often called a domestic alternative to Mobile ALOHA. It consists of a leader-follower pair of arm pairs, meaning a person moves the leader arms and the follower arms copy them, a differential-drive mobile base, multiple cameras, and an onboard computer; the current version integrates AgileX's own PiPER 6-axis arms and TRACER base. Its purpose is collecting demonstration data for mobile bimanual manipulation and running open-source ALOHA algorithms such as ACT directly. Because it is ready to use out of the box, a number of universities use it for data collection, including Tsinghua's RDT-1B, which was trained partly on data collected from this kind of platform.","example":"A lab teleoperates a Cobot Magic to collect several dozen demonstrations of “wipe the table,” then trains an ACT policy so it can do it autonomously.","related":["Mobile ALOHA","AgileX Robotics","AgileX PiPER","Leader-Follower Teleoperation","Action Chunking with Transformers","RDT-1B"]},{"id":"turtlebot","category":"robot","sec":5,"tier":3,"sources":[{"title":"TurtleBot 官网","url":"https://www.turtlebot.com/"}],"as_of":"2026-09","related_ids":["robot-operating-system","robot-operating-system-2","differential-drive-base","simultaneous-localization-and-mapping","ros-2-navigation-stack","wheeled-robot"],"name":"TurtleBot","alt":"TurtleBot","abbr":"","aliases":["TurtleBot 3","TurtleBot 4"],"one_liner":"The most common open-source small wheeled mobile robot kit for learning ROS.","explanation":"TurtleBot is a low-cost, open-source personal robot kit created in November 2010 by Melonee Wise and Tully Foote at Willow Garage, growing up alongside ROS (Robot Operating System). It's a differential-drive wheeled base with sensors and an onboard computer, with no arm, used mainly to learn mapping, localization, and navigation. Two generations are current: TurtleBot 3, a collaboration between Open Robotics and ROBOTIS, with models like Burger and Waffle Pi; and TurtleBot 4, a collaboration between Open Robotics and Clearpath Robotics, built on the iRobot Create 3 base with lidar and an OAK-D depth camera, running ROS 2 by default. Official ROS tutorials and navigation frameworks like Nav2 use it heavily as their example platform, making it the standard teaching tool for getting started with mobile robots.","example":"Following the ROS 2 tutorials, a TurtleBot 4 runs SLAM to build a map of a room, then uses Nav2 to autonomously drive to a specified location.","related":["Robot Operating System","Robot Operating System 2","Differential Drive Base","Simultaneous Localization and Mapping","ROS 2 Navigation Stack (Nav2)","Wheeled Robot"]},{"id":"clearpath-robotics-jackal-husky-ugv","category":"robot","sec":5,"tier":3,"sources":[{"title":"Jackal UGV - Clearpath Robotics","url":"https://clearpathrobotics.com/jackal-small-unmanned-ground-vehicle/"},{"title":"Husky A300 - Clearpath Robotics","url":"https://clearpathrobotics.com/husky-a300-unmanned-ground-vehicle-robot/"}],"as_of":"2025","related_ids":["mobile-base","wheeled-robot","autonomous-mobile-robot","robot-operating-system","simultaneous-localization-and-mapping","navigation"],"name":"Clearpath Robotics Jackal / Husky UGV","alt":"Clearpath Jackal / Husky 移动底盘","abbr":"","aliases":["Jackal","Husky","Husky A300"],"one_liner":"Two wheeled unmanned ground vehicle platforms from Canada's Clearpath, the most common outdoor research bases.","explanation":"Jackal and Husky are two wheeled UGVs, meaning uncrewed ground vehicles with no onboard driver, from the Canadian company Clearpath Robotics, acquired by US-based Rockwell Automation in 2023 and now operating as one of its brands. Jackal is compact, weighing about 17 kg with a 20 kg payload and a top speed of 2 m/s, rated IP62; Husky is a mid-size four-wheel base, with the newer Husky A300 weighing about 80 kg unloaded with a 100 kg payload, rated IP54. Both ship with ROS, meaning Robot Operating System, an onboard computer, GPS, and an IMU already wired up, with a mounting plate on top ready for a lidar, camera, or even a robot arm. They solve the problem of not wanting to build your own vehicle just to do navigation, mapping, or field-robotics research, which is why they appear as the experimental platform in a large number of SLAM and outdoor-navigation papers.","example":"Mounting a lidar and a UR5e arm on top of a Husky is a common DIY outdoor mobile-manipulation setup in many labs.","related":["Mobile Base (Chassis)","Wheeled Robot","Autonomous Mobile Robot","Robot Operating System","Simultaneous Localization and Mapping","Navigation"]},{"id":"hello-robot-stretch","category":"robot","sec":5,"tier":2,"sources":[{"title":"The Robot Report: Stretch 3 from Hello Robot designed for open-source mobile manipulation","url":"https://www.therobotreport.com/stretch-3-mobile-manipulator-hello-robot-designed-open-source-development/"},{"title":"Hello Robot 官网购买页","url":"https://hello-robot.com/purchase/"}],"as_of":"2024-02","related_ids":["mobile-manipulator","mobile-manipulation","ok-robot","dobb-e","hello-robot","household-tasks"],"name":"Hello Robot Stretch","alt":"Hello Robot Stretch","abbr":"","aliases":["Stretch RE1","Stretch 2","Stretch 3"],"one_liner":"Hello Robot's lightweight, open mobile manipulator, widely used in home-robotics research.","explanation":"Stretch is the mobile manipulator made by the US company Hello Robot: a small differential-drive base topped with a vertical lifting column, carrying a horizontally telescoping arm with a wrist and gripper at the end. Rather than aim for a humanoid form, it covers tabletop, floor, and cabinet heights around a home with very few degrees of freedom. The third generation, Stretch 3, released in February 2024, is priced at $24,950, weighs about 24.5 kg, has a footprint of about 33×34 cm, and ships with a DexWrist 3 wrist and a depth camera built into the gripper. Because it is light, inexpensive, and runs open-source software, it has become a common academic platform for home mobile-manipulation research; both OK-Robot and Dobb-E ran their experiments on it.","example":"OK-Robot used a Stretch to perform open-vocabulary pick-and-place in real homes, such as “bring the object on the table to the specified location.”","related":["Mobile Manipulator","Mobile Manipulation","OK-Robot","Dobb-E","Hello Robot","Household Tasks"]},{"id":"willow-garage-pr2","category":"robot","sec":5,"tier":3,"sources":[{"title":"Willow Garage - Wikipedia","url":"https://en.wikipedia.org/wiki/Willow_Garage"}],"as_of":"2014-01","related_ids":["robot-operating-system","willow-garage","mobile-manipulation","mobile-manipulator","dual-arm-robot","fetch-robotics-fetch"],"name":"Willow Garage PR2 (Personal Robot 2)","alt":"PR2","abbr":"PR2","aliases":["Personal Robot 2"],"one_liner":"Willow Garage's 2010 dual-arm mobile-manipulation research robot, the platform where ROS was born.","explanation":"PR2 is a mobile-manipulation robot developed by the American robotics research organization Willow Garage, going on sale in August 2010. It has an omnidirectional mobile base, a telescoping torso, two 7-degree-of-freedom arms (about 1.8 kg payload each), a head-mounted camera and lidar, and two 8-core servers built into its base. In 2010, Willow Garage lent PR2 units to 11 research teams for two years, and the open-source software that came with it was ROS (Robot Operating System), which is how ROS became widespread in academia. It reportedly sold for about $400,000. Willow Garage closed in early 2014, and ongoing support for PR2 passed to Clearpath Robotics.","example":"A Berkeley team once used a PR2 to demonstrate picking towels out of a pile and folding them.","related":["Robot Operating System","Willow Garage","Mobile Manipulation","Mobile Manipulator","Dual-arm Robot","Fetch Robotics Fetch"]},{"id":"toyota-human-support-robot","category":"robot","sec":5,"tier":3,"sources":[{"title":"Toyota: HSR Developers' Community","url":"https://global.toyota/en/detail/8709541"}],"as_of":"2016-04","related_ids":["service-robot","mobile-manipulation","mobile-manipulator","household-tasks","robocup","toyota-research-institute"],"name":"Toyota Human Support Robot","alt":"丰田 HSR","abbr":"HSR","aliases":["Human Support Robot"],"one_liner":"Toyota's single-arm mobile-manipulation robot for home care, widely used in household-service research.","explanation":"HSR is a life-assistance robot that Toyota released in 2012, meant to help elderly or mobility-limited people with small daily tasks, and it can also be operated remotely by a family member or caregiver. It's built as a cylindrical mobile base with a single arm: the base is 430 mm in diameter, the body height adjusts between 1,005 and 1,350 mm, it weighs about 37 kg, the arm reaches about 600 mm, its maximum payload is 1.2 kg, and its top speed is 0.8 km/h. In 2015, Toyota founded the HSR developer community, supplying the robot to universities and research institutions that share software with each other. Because its scale is close to household furniture and it can open doors and pick things up, it's become a common platform for household-service and mobile-manipulation research, and it's reportedly also used as a standard platform in RoboCup@Home.","example":"Researchers use an HSR in a simulated home environment to do tidying tasks like picking clutter up off the floor and putting it in a box.","related":["Service Robot","Mobile Manipulation","Mobile Manipulator","Household Tasks","RoboCup","Toyota Research Institute"]},{"id":"fetch-robotics-fetch","category":"robot","sec":5,"tier":3,"sources":[{"title":"Fetch Mobile Manipulator spec sheet","url":"https://www.amaasia.net/product%20PDF/Fetch_spec_download.pdf"},{"title":"Zebra Technologies Completes Acquisition of Fetch Robotics","url":"https://businesswire.com/news/home/20210810005268/en/Zebra-Technologies-Completes-Acquisition-of-Fetch-Robotics"}],"as_of":"2021-08","related_ids":[null,null,null,null,null,null],"name":"Fetch Robotics Fetch","alt":"Fetch 移动操作机器人","abbr":"","aliases":["Fetch","Fetch Mobile Manipulator"],"one_liner":"A research mobile manipulator from Fetch Robotics, combining a lifting torso and a 7-DoF arm.","explanation":"Fetch is the research-grade mobile manipulator from the American company Fetch Robotics: a differential-drive base, a torso that lifts up and down, a head that pans and tilts, plus a backdrivable 7-DoF robot arm and a parallel two-finger gripper. The arm uses harmonic-drive reducers paired with frameless motors and has about a 6 kg payload; the base carries a SICK lidar, the head has a depth camera, and it runs ROS. In the mid-to-late 2010s it was a common platform for US universities doing mobile manipulation, grasping, and navigation research, occupying a role similar to the earlier PR2. The company's main business was actually warehouse AMRs, and it was acquired by Zebra Technologies in 2021: Zebra paid about $290 million for the 95% of the company it didn't already own, valuing it at roughly $305 million overall.","example":"A research group runs MoveIt on a Fetch for grasp planning, having it pull a specified item off a shelf and drive it over to set down on a table.","related":["Mobile Manipulator","Mobile Manipulation","Willow Garage PR2","PAL Robotics TIAGo","7-DoF Robot Arm","Robot Operating System"]},{"id":"pal-robotics-tiago","category":"robot","sec":5,"tier":3,"sources":[{"title":"TIAGo | Mobile Manipulator Robot (PAL Robotics)","url":"https://pal-robotics.com/robot/tiago/"},{"title":"TIAGo Datasheet (PAL Robotics)","url":"https://pal-robotics.com/wp-content/uploads/2024/04/Datasheet-TIAGo.pdf"}],"as_of":"2026-09","related_ids":["mobile-manipulation","mobile-manipulator","willow-garage-pr2","fetch-robotics-fetch","toyota-human-support-robot","robot-operating-system"],"name":"PAL Robotics TIAGo","alt":"TIAGo","abbr":"","aliases":["TIAGo++","TIAGo Pro"],"one_liner":"PAL Robotics' modular mobile manipulator, widely used in academic research and home-service studies.","explanation":"TIAGo is a mobile manipulator (a mobile base topped with a robot arm) from Spain's PAL Robotics. The standard version has a differential-drive base, a lift that adjusts torso height (roughly 1.1–1.45 m), a 7-degree-of-freedom arm, and a head that pans and tilts with an RGB-D camera (color plus depth). Its main selling point is modularity: it can be configured with a gripper or a dexterous hand, one arm or two (TIAGo++), and an optional series-elastic-actuator arm for torque control and safer collisions; PAL rates single-arm payload at about 3 kg excluding the end effector, and there's also an upgraded TIAGo Pro. The system runs on ROS, so many universities use it for research in navigation, grasping, human-robot interaction, and home-service tasks — it sits in the same category of research platform as PR2, Fetch, and Toyota's HSR.","example":"Researchers run ROS's Nav2 navigation stack together with MoveIt grasp planning on a TIAGo to have it pick up a cup from a kitchen counter and hand it to a person.","related":["Mobile Manipulation","Mobile Manipulator","Willow Garage PR2 (Personal Robot 2)","Fetch Robotics Fetch","Toyota Human Support Robot","Robot Operating System"]},{"id":"everyday-robots-mobile-manipulator","category":"robot","sec":5,"tier":3,"sources":[{"title":"Alphabet closes Everyday Robots among layoffs - The Robot Report","url":"https://www.therobotreport.com/alphabet-closes-everyday-robots-among-layoffs/"},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (arXiv)","url":"https://arxiv.org/abs/2212.06817"}],"as_of":"2023-02","related_ids":[null,null,null,null,null,null],"name":"Everyday Robots Mobile Manipulator","alt":"Everyday Robots 移动机械臂","abbr":"EDR","aliases":["EDR","Google Mobile Manipulator"],"one_liner":"The single-arm wheeled mobile robot used in Google's RT-1 and SayCan work.","explanation":"This is the wheeled mobile-manipulation robot built by the Everyday Robots team, incubated within Alphabet's X moonshot division: a mobile base on the bottom, a 7-DoF robot arm with a two-finger gripper on top, and a head carrying a camera and other sensors. The team once had over a hundred of these robots wiping tables, sorting trash, and opening doors around Google's own offices. It was the main platform behind Google's robot-learning papers, with SayCan, RT-1, and RT-2 all running experiments on it; RT-1 in particular used 13 of these robots to collect about 130,000 demonstrations over 17 months. In February 2023, the project was shut down as part of Alphabet's broader layoffs earlier that year, with part of the team and its technology folded into Google Research's robotics efforts. The robot itself was never commercialized, but the data it collected, the RT-1 dataset, was incorporated into Open X-Embodiment and continues to have influence today.","example":"In a SayCan demo, after hearing “I spilled my drink,” the robot planned out finding a sponge and bringing it over, and carried out the steps in the kitchen.","related":["Everyday Robots (Alphabet X)","Mobile Manipulator","RT-1","SayCan","RT-1 Robot Action Dataset (Fractal)","Google DeepMind"]},{"id":"locobot","category":"robot","sec":5,"tier":3,"sources":[{"title":"PyRobot: An Open-source Robotics Framework for Research and Benchmarking（arXiv）","url":"https://arxiv.org/pdf/1906.08236"},{"title":"LoCoBot（robots.ros.org）","url":"https://robots.ros.org/locobot/"}],"as_of":"2019-06","related_ids":["mobile-manipulation","mobile-manipulator","robot-operating-system","habitat","point-goal-navigation","turtlebot"],"name":"LoCoBot","alt":"LoCoBot 低成本移动操作机器人","abbr":"","aliases":[],"one_liner":"A low-cost mobile-manipulation research platform designed at CMU, built to pair with Facebook AI's PyRobot.","explanation":"LoCoBot is a low-cost mobile-manipulation robot designed at Carnegie Mellon University (CMU), meant to be used with PyRobot, the open-source framework Facebook AI (now Meta) released in 2019, and sold commercially by Trossen Robotics. Its default configuration is a Kobuki differential-drive base, an Intel NUC mini PC, an Intel RealSense D435 depth camera, and a small WidowX robot arm (5 degrees of freedom, about 0.2 kg payload, and roughly 0.55 m of reach, per the PyRobot paper). PyRobot sits on top of ROS (Robot Operating System) and provides a hardware-agnostic Python interface for navigation and manipulation, letting deep-learning researchers who don't know ROS well control the real robot directly. Because it's inexpensive and can handle both navigation and grasping, it's a popular choice for budget-limited labs running experiments in embodied navigation, grasping, and sim-to-real transfer.","example":"A point-goal navigation policy trained in the Habitat simulator is deployed on a real LoCoBot to measure its success rate and compare simulated versus real-world performance.","related":["Mobile Manipulation","Mobile Manipulator","Robot Operating System","Habitat","Point-Goal Navigation","TurtleBot"]},{"id":"tidybot-plus-plus","category":"robot","sec":5,"tier":3,"sources":[{"title":"arXiv 2412.10447: TidyBot++","url":"https://arxiv.org/abs/2412.10447"},{"title":"GitHub: jimmyyhwu/tidybot2","url":"https://github.com/jimmyyhwu/tidybot2"}],"as_of":"2024-12","related_ids":["mobile-manipulation","mobile-manipulator","kinova-gen3","diffusion-policy","teleoperation","mobile-aloha"],"name":"TidyBot++","alt":"TidyBot++ 开源全向移动操作平台","abbr":"","aliases":["Open-Source Holonomic Mobile Manipulator","tidybot2"],"one_liner":"An open-source low-cost omnidirectional mobile base plus arm from Stanford and Princeton, for household mobile-manipulation research.","explanation":"TidyBot++ is an open-source mobile-manipulation platform from Stanford and Princeton (Jimmy Wu, Shuran Song, Oussama Khatib, Jeannette Bohg, and others), published at CoRL 2024, with both hardware drawings and code released publicly. Its core is a holonomic base built from powered caster wheels: it can independently and simultaneously control forward-backward motion, sideways motion, and rotation, so it doesn't need to turn to face a direction before moving the way a differential-drive base does, which simplifies manipulation while in motion. Any robot arm can be mounted on the base; the paper uses a Kinova Gen3. It ships with a phone-based teleoperation interface for collecting demonstrations, and the team trains diffusion policies on that data to complete a range of household tasks in real apartments. The base reportedly costs about $5,000–6,000 to build, not including the arm.","example":"Demonstrations are collected in a real apartment through the phone teleoperation interface, then used to train a diffusion policy so the robot can carry out household mobile-manipulation tasks on its own.","related":["Mobile Manipulation","Mobile Manipulator","Kinova Gen3","Diffusion Policy","Teleoperation","Mobile ALOHA"]},{"id":"lekiwi","category":"robot","sec":5,"tier":3,"sources":[{"title":"GitHub - SIGRobotics-UIUC/LeKiwi: Low-Cost Mobile Manipulator","url":"https://github.com/SIGRobotics-UIUC/LeKiwi"},{"title":"Upgrading the LeKiwi into a LiDAR-equipped explorer | Foxglove","url":"https://foxglove.dev/blog/upgrading-the-lekiwi-into-a-lidar-equipped-explorer"}],"as_of":"2025","related_ids":["lerobot","so-100-so-101-arm","mobile-manipulation","omni-wheel","xlerobot","feetech-sts3215-servo"],"name":"LeKiwi","alt":"LeKiwi 移动底盘机械臂","abbr":"","aliases":[],"one_liner":"An open-source low-cost mobile manipulator combining a three-omni-wheel base with an SO-101 arm.","explanation":"LeKiwi was open-sourced in early 2025 by SIGRobotics, a student club at the University of Illinois Urbana-Champaign, and is now officially supported by Hugging Face's LeRobot. Its base uses three omnidirectional wheels arranged 120° apart — a so-called Kiwi drive — which lets it translate and rotate in any direction on the spot; an SO-101 (or SO-100) robot arm sits on top of the base. Both the base and the arm use Feetech STS3215 servos, a Raspberry Pi 5 serves as the onboard computer, and there's one camera on the base and another on the wrist. The whole platform is built from 3D-printed parts and off-the-shelf components, which lets individuals and students do mobile-manipulation teleoperation data collection and imitation learning at very low cost.","example":"Someone teleoperates LeKiwi with an SO-101 leader arm and a keyboard to drive it to a table and pick up a sock, then trains a policy on the recorded data with LeRobot.","related":["LeRobot","SO-100 / SO-101 Arm","Mobile Manipulation","Omni Wheel","XLeRobot","Feetech STS3215 Servo"]},{"id":"xlerobot","category":"robot","sec":5,"tier":3,"sources":[{"title":"Vector-Wangel/XLeRobot (GitHub)","url":"https://github.com/Vector-Wangel/XLeRobot"}],"as_of":"2025-12","related_ids":["lerobot","so-100-so-101-arm","lekiwi","mobile-manipulation","open-source-hardware","mobile-manipulator"],"name":"XLeRobot","alt":"XLeRobot","abbr":"","aliases":[],"one_liner":"An open-source dual-arm wheeled home robot with a parts cost starting around $660.","explanation":"XLeRobot is an open-source hardware project started independently by Gaotian (Vector) Wang, a graduate student at Rice University, with version 0.2.0 released in June 2025. It mounts two low-cost SO-100/SO-101 arms on a mobile base (originally omnidirectional wheels, with a two-wheel differential-drive version added in December 2025), a head with an RGB camera that can be upgraded to a stereo or RealSense depth camera, and runs off a laptop or a Raspberry Pi. The base configuration has a parts cost starting around $660, not including 3D printing, shipping, or tax. It builds on projects like Hugging Face's LeRobot, the SO-101, and LeKiwi, aiming to let individuals and students do dual-arm mobile-manipulation and household-task research on a very small budget.","example":"","related":["LeRobot","SO-100 / SO-101 Arm","LeKiwi","Mobile Manipulation","Open-Source Hardware (OSHW)","Mobile Manipulator"]},{"id":"honda-asimo","category":"robot","sec":6,"tier":2,"sources":[{"title":"Honda Debuts New Humanoid Robot ASIMO (2000)","url":"https://global.honda/en/newsroom/news/2000/c001120b-eng.html"},{"title":"Wikipedia: ASIMO","url":"https://en.wikipedia.org/wiki/ASIMO"}],"as_of":"2022-03","related_ids":["humanoid-robot","bipedal-locomotion","zero-moment-point","model-based-control","hrp-humanoid-robot-series"],"name":"Honda ASIMO","alt":"本田 ASIMO","abbr":"","aliases":[],"one_liner":"Honda's bipedal humanoid robot, first released in 2000, an icon of early humanoid robotics.","explanation":"ASIMO is the bipedal humanoid robot developed by Honda in Japan, with its first generation released in 2000 at about 1.2 meters tall. The 2011 version stood 130 cm tall, weighed 48 kg, had 57 degrees of freedom, could run at about 9 km/h, and was able to climb stairs, carry a tray, and interact simply with people. It relied on model-based control, such as the zero moment point (ZMP), for balance, with most of its motions pre-planned by engineers, making it the high-water mark for humanoid robots from before the deep-learning era. Honda halted development in 2018 and retired ASIMO after its last public appearance in March 2022. It is now commonly used as the historical baseline against which today's learning-based humanoid robots are compared.","example":"In a 2011 demonstration, ASIMO could hop on one leg, run while turning, and pour and serve a drink to an audience member on command.","related":["Humanoid Robot","Bipedal Locomotion","Zero Moment Point","Model-Based Control","HRP Humanoid Robot Series"]},{"id":"boston-dynamics-atlas","category":"robot","sec":6,"tier":3,"sources":[{"title":"Atlas shrugged: Boston Dynamics retires its hydraulic humanoid robot (TechCrunch)","url":"https://techcrunch.com/2024/04/16/atlas-shrugged-boston-dynamics-retires-its-humanoid-robot/"},{"title":"Atlas (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Atlas_(robot)"}],"as_of":"2024-04","related_ids":["boston-dynamics-atlas-2","boston-dynamics","hydraulic-actuation","darpa-robotics-challenge","model-predictive-control","bipedal-locomotion"],"name":"Boston Dynamics Atlas (Hydraulic)","alt":"液压版 Atlas","abbr":"","aliases":["HD Atlas","Atlas HD"],"one_liner":"Boston Dynamics' 2013–2024 hydraulic humanoid robot, famous for parkour and backflip videos.","explanation":"The hydraulic Atlas was originally built by Boston Dynamics with funding from DARPA, the US Defense Advanced Research Projects Agency, for the DARPA Robotics Challenge, and was unveiled publicly in July 2013: about 1.88 meters tall and 150 kg at the time, and still tethered to an external power cable. After 2016 it was replaced by a smaller, battery-powered version, using hydraulic actuators to deliver high power for running, jumping, backflips, parkour, and carrying boxes. Its motion relied mainly on model-based control, such as model predictive control, and offline-designed maneuvers, demonstrating the upper limit of dynamic balance for humanoid robots at the time. It always remained a research platform and was never commercialized, and was retired in April 2024, replaced that same month by the newly released electric Atlas.","example":"In the 2021 “Parkour Atlas” video, two Atlas units vault over boxes in sequence, clear obstacles, and perform backflips.","related":["Boston Dynamics Atlas (Electric)","Boston Dynamics","Hydraulic Actuation","DARPA Robotics Challenge","Model Predictive Control","Bipedal Locomotion"]},{"id":"boston-dynamics-atlas-2","category":"robot","sec":6,"tier":1,"sources":[{"title":"Boston Dynamics Unveils New Atlas Robot to Revolutionize Industry","url":"https://bostondynamics.com/blog/boston-dynamics-unveils-new-atlas-robot-to-revolutionize-industry/"},{"title":"CES 2026: Boston Dynamics unveils new Atlas humanoid robot (Robotics 24/7)","url":"https://www.robotics247.com/article/ces-2026-boston-dynamics-unveils-new-atlas-humanoid-robot/technologies"}],"as_of":"2026-01","related_ids":["boston-dynamics-atlas","boston-dynamics","hyundai-motor-group","google-deepmind","full-size-humanoid-robot","humanoid-robot"],"name":"Boston Dynamics Atlas (Electric)","alt":"波士顿动力 Atlas（电动版）","abbr":"","aliases":["Atlas","Electric Atlas"],"one_liner":"Boston Dynamics' fully electric humanoid robot, introduced in 2024 for factory work.","explanation":"Electric Atlas is the humanoid robot Boston Dynamics (now part of Hyundai Motor Group) unveiled in April 2024, replacing the retired hydraulic Atlas. It switched to electric motors, giving its joints a wide, sometimes full-circle range of rotation, so its movements no longer need to mimic human posture. At CES in January 2026, Boston Dynamics unveiled the production version: 56 degrees of freedom, an arm span of about 2.3 meters, the ability to lift 50 kg, roughly 4 hours of battery life with the ability to walk itself back to a charging station to swap batteries, and support for autonomous, teleoperated, and tablet-controlled modes. The company says its entire 2026 deployment is going to Hyundai's car factories and to Google DeepMind, with whom it is jointly training foundation models.","example":"Atlas moves and sequences car parts inside a Hyundai factory.","related":["Boston Dynamics Atlas (Hydraulic)","Boston Dynamics","Hyundai Motor Group","Google DeepMind","Full-size Humanoid Robot","Humanoid Robot"]},{"id":"tesla-optimus","category":"robot","sec":6,"tier":1,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"title":"Tesla Optimus Gen 3 (Not a Tesla App)","url":"https://www.notateslaapp.com/news/4725/tesla-app-leaks-new-optimus-gen-3-robot-design"}],"as_of":"2026-09","related_ids":["tesla","tesla-optimus-v3","tesla-optimus-hand","tesla-supply-chain","tesla-we-robot-event","humanoid-robot"],"name":"Tesla Optimus","alt":"擎天柱","abbr":"","aliases":["Optimus","Optimus Gen 2"],"one_liner":"Tesla's general-purpose bipedal humanoid robot project.","explanation":"Optimus is Tesla's humanoid robot project, first unveiled at AI Day in August 2021, shown as a prototype in 2022, and released in its second generation (Gen 2) in December 2023. Published specs: about 1.73 meters tall, roughly 57 kg, able to carry about 20 kg; Gen 2's hand originally had 11 degrees of freedom, and a new hand shown in 2024 increased that to 22. Tesla intends to reuse its self-driving vision neural networks and its battery and motor supply chain to build the humanoid, and Elon Musk has repeatedly said the target price is around $20,000–30,000. As of September 2026, the third generation aimed at mass production has not yet formally launched, though the company says its Fremont factory is installing a line reportedly aimed at a run rate in the millions per year.","example":"At Tesla's 2024 We, Robot event, Optimus poured drinks for guests; some of its movements were later reported to have been remotely teleoperated.","related":["Tesla","Tesla Optimus V3 (Optimus Gen 3)","Tesla Optimus Hand","Tesla (Optimus) Supply Chain","Tesla We, Robot Event (2024)","Humanoid Robot"]},{"id":"tesla-optimus-v3","category":"robot","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Optimus (robot)","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"}],"as_of":"2026-09","related_ids":["tesla-optimus","tesla","tesla-optimus-hand","full-size-humanoid-robot","mass-production","tesla-supply-chain"],"name":"Tesla Optimus V3 (Optimus Gen 3)","alt":"特斯拉 Optimus V3（Gen 3）","abbr":"","aliases":["Optimus V3","Optimus Gen 3"],"one_liner":"Tesla's third-generation Optimus, designed for mass production but not yet formally released as of September 2026.","explanation":"Optimus V3 is the third generation of Tesla's humanoid robot Optimus, officially positioned as the first version designed from the outset for mass production. At Tesla's Q1 2026 earnings call in April, Elon Musk said V3 would be unveiled close to when mass production begins; Tesla is installing a line at its Fremont factory designed for an annual capacity of one million units and is planning a larger line at its Texas factory. As of September 2026, V3 has not formally been shown, and specs such as height, degrees of freedom, and price have not been officially disclosed; figures circulating online are mostly media speculation. The previous generation, Optimus Gen 2 (December 2023), had 11 degrees of freedom in one hand; Tesla showed a new 22-DoF hand in 2024. Interest in V3 centers on whether humanoid robots can genuinely reach million-unit-scale production.","example":"","related":["Tesla Optimus","Tesla","Tesla Optimus Hand","Full-size Humanoid Robot","Mass Production","Tesla (Optimus) Supply Chain"]},{"id":"figure-01","category":"robot","sec":6,"tier":3,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"OpenAI Makes Figure's Robot Talk Like Human - FavTutor","url":"https://favtutor.com/articles/figure-robot-openai-demo/"}],"as_of":"2024-08","related_ids":["figure-ai","figure-02",null,"openai",null,null],"name":"Figure 01","alt":"Figure 01","abbr":"","aliases":["Figure One"],"one_liner":"Figure AI's first-generation humanoid, which went viral for a conversational demo made with OpenAI.","explanation":"Figure 01 is the first-generation bipedal humanoid robot from the US company Figure AI, founded by Brett Adcock in 2022, aimed at physical labor in logistics, warehousing, and manufacturing. Published specs: about 1.68 meters tall, roughly 60 kg, a payload of about 20 kg, about 5 hours of battery life, and exposed cabling for easier maintenance. In January 2024, Figure signed an agreement with BMW to bring humanoid robots into its Spartanburg, South Carolina factory; in March 2024, it released a demo video made in partnership with OpenAI, in which Figure 01 held a spoken conversation with a person while handing them an apple and clearing dishes, which the company said was recorded at normal speed in a single unedited take. In August 2024 it was succeeded by Figure 02, and Figure's partnership with OpenAI also ended in early 2025 as the company shifted to its own in-house model, Helix.","example":"In the March 2024 demo, when asked “can you give me something to eat,” Figure 01 handed over the only edible item on the table, an apple, and explained why it chose it.","related":["Figure AI","Figure 02","Figure Helix","OpenAI","Full-size Humanoid Robot","One-Take (Uncut) Video"]},{"id":"figure-02","category":"robot","sec":6,"tier":2,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"BMW tests Figure 02 humanoid on production line - The Robot Report","url":"https://www.therobotreport.com/bmw-tests-figure-02-humanoid-on-production-line/"}],"as_of":"2025-10","related_ids":["figure-ai","figure-helix","figure-01","figure-03","humanoid-robot","dual-system-architecture"],"name":"Figure 02","alt":"Figure 02","abbr":"","aliases":["F.02","Figure 2"],"one_liner":"Figure AI's second-generation humanoid robot, released in 2024 and piloted in a BMW factory.","explanation":"Figure 02 is the second-generation general-purpose humanoid robot the US company Figure AI released in August 2024, succeeding Figure 01. Published specs: about 1.68 meters tall, roughly 70 kg, 16 degrees of freedom in each hand, and 6 cameras across the body; the company says its battery capacity is 50% higher than the previous generation and its onboard compute 3 times higher, with all cabling routed internally. It was piloted doing body-panel handling and placement at BMW's Spartanburg factory, and it was also the first robot to run Figure's own VLA model, Helix, released in February 2025 with a fast-slow dual-system design; demo videos showed two Figure 02 units collaborating to put away groceries and sorting packages on a logistics line. After Figure 03 launched in October 2025, Figure 02 became the previous generation.","example":"In February 2025, Figure used Helix to control two Figure 02 units at once, putting groceries away in a fridge together on voice command.","related":["Figure AI","Figure Helix","Figure 01","Figure 03","Humanoid Robot","Dual-System Architecture (System 1 / System 2)"]},{"id":"figure-03","category":"robot","sec":6,"tier":1,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"Figure 03 Specs & Price | Humanoid.guide","url":"https://humanoid.guide/product/figure-03/"}],"as_of":"2025-10","related_ids":["figure-ai","figure-helix","figure-helix-02","figure-02","botq","humanoid-robot"],"name":"Figure 03","alt":"Figure 03","abbr":"","aliases":["F.03"],"one_liner":"Figure AI's third-generation humanoid robot, released in 2025 for home use and mass production.","explanation":"Figure 03 is the third-generation humanoid robot from the US company Figure AI, released on October 9, 2025, succeeding Figure 02. It was designed around Figure's own VLA model, Helix: its fingertips have tactile sensors said to sense forces as small as about 3 grams, its palms carry cameras, its outer shell uses a washable soft fabric, and coils built into its feet allow 2 kW wireless charging. It is reported to stand about 1.73 meters tall and weigh about 61 kg. It was also designed for mass manufacturing, built on Figure's BotQ production line, which the company says can reach an annual capacity of 12,000 units on its first line.","example":"In Figure's release video, Figure 03 uses Helix to fold laundry and load a dishwasher at home.","related":["Figure AI","Figure Helix","Figure Helix 02","Figure 02","BotQ","Humanoid Robot"]},{"id":"1x-eve","category":"robot","sec":6,"tier":3,"sources":[{"title":"1X Technologies - Wikipedia","url":"https://en.wikipedia.org/wiki/1X_Technologies"}],"as_of":"2025-10","related_ids":[null,null,"1x-neo",null,"openai"],"name":"1X EVE","alt":"1X EVE","abbr":"","aliases":["EVE","Halodi EVE"],"one_liner":"The Norwegian company 1X's wheeled humanoid, aimed at security, logistics, and healthcare settings.","explanation":"EVE is the first humanoid robot from the Norwegian robotics company 1X Technologies, previously known as Halodi Robotics before its 2022 rename: a humanoid upper body with two arms and a head, on a wheeled base. It uses 1X's own actuation, sensing, and manipulation technology, and is positioned as a field-testing platform for logistics, security, healthcare, and similar institutional settings, able to carry out tasks continuously from voice commands. In 2023, a Series A2 round led by the OpenAI Startup Fund brought EVE broad attention. 1X has since shifted its focus to NEO, a bipedal humanoid aimed at the home, leaving EVE more as an earlier-generation representative.","example":"","related":["1X Technologies","Wheeled Humanoid Robot","1X NEO","1X World Model","OpenAI"]},{"id":"1x-neo","category":"robot","sec":6,"tier":2,"sources":[{"title":"1X NEO Home Robot","url":"https://www.1x.tech/discover/neo-home-robot"},{"title":"NEO humanoid available for preorder - The Robot Report","url":"https://www.therobotreport.com/1x-announces-pre-order-launch-neo-humanoid-robot/"}],"as_of":"2026-07","related_ids":["1x-technologies","redwood","1x-world-model","1x-eve","tendon-driven-actuation","remote-teleoperation-takeover"],"name":"1X NEO","alt":"1X NEO","abbr":"","aliases":["NEO","NEO Gamma","NEO Beta"],"one_liner":"1X's humanoid robot for the home, priced at $20,000.","explanation":"NEO is a home humanoid robot from 1X Technologies, a company that started in Norway and is now headquartered in the US. It released NEO Beta in 2024 and NEO Gamma in February 2025, then opened consumer preorders in October 2025: a $20,000 outright purchase, or a $499-per-month subscription. It is reported to stand about 1.68 meters tall and weigh about 30 kg, able to lift 70 kg, with 22 degrees of freedom in its hands, tendon-driven actuation, a soft outer shell, and an operating noise level of about 22 decibels. Household tasks it cannot yet handle can be completed by a 1X employee taking over remotely, which has drawn privacy criticism. The company plans deliveries in the US starting in 2026; as of July 2026, no confirmed customer deliveries had been reported.","example":"A user schedules NEO to fold laundry through a phone app; for tasks it can't handle, they can book a remote operator to take over.","related":["1X Technologies","Redwood","1X World Model","1X EVE","Tendon-Driven Actuation","Remote Teleoperation Takeover"]},{"id":"agility-robotics-cassie","category":"robot","sec":6,"tier":3,"sources":[{"title":"Oregon State University：Bipedal robot developed at Oregon State achieves Guinness World Record in 100 meters","url":"https://news.oregonstate.edu/news/bipedal-robot-developed-oregon-state-achieves-guinness-world-record-100-meters"},{"title":"The Robot Report：Watch a Cassie bipedal robot run 100 meters","url":"https://www.therobotreport.com/watch-a-cassie-bipedal-robot-run-100-meters/"}],"as_of":"2022-05","related_ids":["agility-robotics","bipedal-robot","reverse-knee-leg","agility-robotics-digit","rl-based-locomotion-control","sim-to-real-transfer"],"name":"Agility Robotics Cassie","alt":"Agility Cassie","abbr":"","aliases":["Cassie"],"one_liner":"A legless bipedal robot, just two bird-like legs, a classic platform for legged reinforcement learning.","explanation":"Cassie is a bipedal robot designed by Jonathan Hurst's team at Oregon State University and manufactured by Agility Robotics, launched in 2017 with DARPA funding supporting the development. It has no head or arms, only two legs whose knees bend backward like an ostrich's, a design that prioritizes elasticity and energy efficiency. It was one of the first bipedal robots to have its running gait controlled outdoors with machine learning, and has been used by many universities for reinforcement-learning locomotion control and sim-to-real transfer research. In 2022 it set the Guinness World Record for a bipedal robot 100-meter dash, at 24.73 seconds. Agility later built a humanoid robot, Digit, on top of Cassie's leg design by adding an upper body.","example":"Teams at Berkeley and elsewhere have trained walking policies for Cassie with reinforcement learning, then transferred them from simulation to the real robot.","related":["Agility Robotics","Bipedal Robot","Reverse-Knee (Digitigrade, Bird-Like) Leg","Agility Robotics Digit","RL-based Locomotion Control","Sim-to-Real Transfer"]},{"id":"agility-robotics-digit","category":"robot","sec":6,"tier":2,"sources":[{"title":"Agility Unveils Digit 5 Humanoid Robot Built for Cooperatively Safe Work at Scale","url":"https://www.agilityrobotics.com/content/agility-unveils-digit-5-humanoid-robot-built-for-cooperatively-safe-work-at-scale"},{"title":"Agility Robotics Digit Specs & Price | Humanoid.guide","url":"https://humanoid.guide/product/digit/"}],"as_of":"2026-09","related_ids":["agility-robotics","bipedal-robot","agility-robotics-cassie","robofab","tote-handling","robot-as-a-service"],"name":"Agility Robotics Digit","alt":"Agility Digit 人形机器人","abbr":"Digit","aliases":["Digit","Digit 4","Digit 5"],"one_liner":"Agility Robotics' reverse-jointed bipedal humanoid, built mainly to carry totes in warehouses.","explanation":"Digit is a bipedal humanoid robot developed by the American company Agility Robotics. Its legs carry over the reverse-jointed design, bending backward like a bird's, from its predecessor Cassie, and its upper body adds two arms and a simplified end effector, used mainly to carry totes, meaning plastic storage bins, around warehouses. The previous generation of Digit stood about 1.75 meters tall with roughly a 16 kg payload, and has already been deployed on a pay-per-use, robot-as-a-service basis at logistics companies such as GXO, with Agility also building a dedicated RoboFab factory to produce it. Digit 5, released on September 15, 2026, is built around “safe collaboration”: an array of sensors plus an independent safety controller monitor nearby people and automatically slow down, steer around, or stop when someone gets too close, aiming to work alongside human workers without safety fencing. The company said that, as of May 2026, multi-year orders had exceeded $300 million, with early deliveries planned for the first half of 2027 and full-scale supply by the end of that year.","example":"Digit moves totes from shelves to conveyor belts inside a GXO warehouse, making it one of the earlier humanoid robots deployed commercially on a paid basis.","related":["Agility Robotics","Bipedal Robot","Agility Robotics Cassie","RoboFab","Tote Handling","Robot-as-a-Service"]},{"id":"apptronik-apollo","category":"robot","sec":6,"tier":2,"sources":[{"title":"Apptronik 官网","url":"https://apptronik.com/"},{"title":"Apptronik - Apollo 2","url":"https://apptronik.com/apollo/apollo-2"},{"title":"Apptronik Opens 90,000 Sq Ft Testing Site for New Apollo 2 Humanoid (A3)","url":"https://www.automate.org/robotics/industry-insights/apptronik-opens-90-000-sq-ft-testing-site-for-new-apollo-2-humanoid/aph"}],"as_of":"2026-07","related_ids":["apptronik","humanoid-robot","gemini-robotics","google-deepmind","wheeled-humanoid-robot","tote-handling"],"name":"Apptronik Apollo","alt":"Apptronik Apollo","abbr":"","aliases":["Apollo","Apollo 2"],"one_liner":"A general-purpose humanoid robot from the US company Apptronik, aimed at logistics and manufacturing.","explanation":"Apollo is the humanoid robot the Texas-based company Apptronik released in 2023: about 1.73 meters tall, roughly 73 kg, with a payload of about 25 kg, designed for carrying boxes and totes in warehouses and factories, with a swappable battery. Apptronik grew out of a robotics lab at the University of Texas at Austin, and Apollo uses in-house electric actuators. Apollo 2, released in June 2026, comes in both bipedal and wheeled-base configurations and is positioned as a training and data platform, reportedly with 35 degrees of freedom, 12 of them in the hands; fleets running at the company's Robot Park site and at customer locations continuously collect real-world data, feeding a partnership with Google DeepMind to train the Gemini Robotics model family. The company reportedly designated Apollo 3 as its first fully commercial product, targeted for 2027.","example":"Apptronik partnered with Mercedes-Benz to pilot Apollo carrying parts totes at one of its factories.","related":["Apptronik","Humanoid Robot","Gemini Robotics","Google DeepMind","Wheeled Humanoid Robot","Tote Handling"]},{"id":"sanctuary-ai-phoenix","category":"robot","sec":6,"tier":3,"sources":[{"title":"Robotics 24/7: Sanctuary AI Unveils Phoenix, Sixth Generation Humanoid","url":"https://www.robotics247.com/article/sanctuary_ai_unveils_phoenix_sixth_generation_humanoid_general_purpose_robot"},{"title":"RoboZaps: Sanctuary AI Phoenix 2026 status","url":"https://blog.robozaps.com/b/sanctuary-ai-phoenix-review"}],"as_of":"2026-06","related_ids":["sanctuary-ai","humanoid-robot","dexterous-hand","hydraulic-actuation","teleoperation","general-purpose-robot"],"name":"Sanctuary AI Phoenix","alt":"Sanctuary Phoenix","abbr":"","aliases":["Phoenix"],"one_liner":"A Canadian general-purpose humanoid robot from Sanctuary AI, best known for its hydraulic dexterous hands.","explanation":"Phoenix is a general-purpose humanoid robot from the Vancouver-based company Sanctuary AI, which released its sixth generation in May 2023, followed by the seventh and eighth generations in April and December 2024. The company's published specs for the sixth generation: about 170 cm tall, about 70 kg, and roughly 25 kg of payload. Its standout feature is hydraulically driven dexterous hands, each with about 21 degrees of freedom and tactile feedback; its control software is called Carbon, and in its early stages it relied heavily on humans teleoperating the robot to collect demonstration data used to train it. Sanctuary reportedly went through layoffs and management changes in 2024–2025, and in 2026 pivoted to selling AI software for other companies' industrial robots; Phoenix was never sold commercially.","example":"In a reported retail pilot, Sanctuary had Phoenix complete over a hundred tasks such as restocking shelves and applying price tags, many of them carried out through human teleoperation.","related":["Sanctuary AI","Humanoid Robot","Dexterous Hand","Hydraulic Actuation","Teleoperation","General-purpose Robot"]},{"id":"neura-robotics-4ne1","category":"robot","sec":6,"tier":3,"sources":[{"title":"Humanoid Robot 4NE1 for Work and Life（NEURA Robotics）","url":"https://neura-robotics.com/products/4ne1/"},{"title":"Automatica 2025: NEURA Robotics unveils 3rd generation 4NE1 humanoid（Robotics 24/7）","url":"https://www.robotics247.com/article/automatica-2025-neura-robotics-unveils-3rd-generation-4ne1-humanoid/food"}],"as_of":"2025-06","related_ids":["neura-robotics","humanoid-robot","electronic-skin","human-robot-collaboration","full-size-humanoid-robot","exteroception"],"name":"NEURA Robotics 4NE1","alt":"NEURA 4NE1","abbr":"","aliases":["4NE-1"],"one_liner":"A general-purpose humanoid robot from Germany's NEURA Robotics, built around full-body sensing and human collaboration.","explanation":"4NE1 is a humanoid robot developed by Germany's NEURA Robotics; its third generation debuted at the Automatica trade show in Munich in June 2025. Per NEURA's own specs, it stands 180 cm tall, weighs 80 kg, walks at up to 5 km/h, and has a rated payload of 10–100 kg. It's covered in a sensor skin with 360-degree perception, letting it sense its surroundings and nearby people so it can work safely alongside them; on the software side it combines language models, computer vision, and reinforcement learning, supports teleoperation, and has a swappable forearm, with wheeled and other variants also available. It reportedly has about 55 degrees of freedom across its body (12 in each hand), and a dual-battery design lets it swap packs to run close to around the clock. NEURA also offers a smaller version for research and education, the 4NE1 Mini, which the company's website says is scheduled to ship in spring 2026. 4NE1 is one of the flagship products of the European humanoid-robot field.","example":"A university preorders a 4NE1 Mini for embodied-AI teaching and algorithm validation.","related":["NEURA Robotics","Humanoid Robot","Electronic Skin","Human-Robot Collaboration","Full-size Humanoid Robot","Exteroception"]},{"id":"mentee-robotics-menteebot","category":"robot","sec":6,"tier":3,"sources":[{"title":"Mobileye To Acquire Mentee Robotics to Accelerate Physical AI Leadership（Mobileye News）","url":"https://www.mobileye.com/news/mobileye-to-acquire-mentee-robotics-to-accelerate-physical-ai-leadership/"},{"title":"Mobileye acquires humanoid robot startup Mentee Robotics for $900M（TechCrunch）","url":"https://techcrunch.com/2026/01/06/mobileye-acquires-humanoid-robot-startup-mentee-robotics-for-900m"}],"as_of":"2026-01","related_ids":["mentee-robotics","mobileye","humanoid-robot","vision-only-approach","sim-to-real-transfer","hot-swappable-battery"],"name":"Mentee Robotics MenteeBot","alt":"MenteeBot","abbr":"","aliases":["MenteeBot V3","Menteebot 3.0"],"one_liner":"An Israeli general-purpose humanoid robot whose maker, Mentee Robotics, was acquired by Mobileye in 2026.","explanation":"MenteeBot is a general-purpose humanoid robot developed by the Israeli startup Mentee Robotics, founded around 2022 by Mobileye founder Amnon Shashua along with Lior Wolf, who serves as CEO. The reported third-generation MenteeBot stands about 175 cm tall, weighs about 70 kg, carries about 25 kg, and has around 40 degrees of freedom. It relies on cameras alone for perception, including side- and rear-facing fisheye cameras for 360-degree coverage, and uses in-house motors and actuators, tactile-sensing hands, and hot-swappable batteries; its training leans heavily on sim-to-real transfer. Mentee Robotics emphasizes scene understanding, following natural-language instructions, and end-to-end autonomous execution without teleoperation. On January 6, 2026, Mobileye announced it would acquire Mentee for about $900 million, with plans for initial customer proof-of-concept deployments in 2026 and mass production by 2028.","example":"Under the plan laid out in the acquisition announcement, MenteeBot is set to go on-site with customers in 2026 to prove out autonomous work that doesn't rely on teleoperation.","related":["Mentee Robotics","Mobileye","Humanoid Robot","Vision-Only Approach","Sim-to-Real Transfer","Hot-Swappable Battery"]},{"id":"foundation-robotics-phantom","category":"robot","sec":6,"tier":3,"sources":[{"title":"US firm Foundation plans to build 50,000 humanoid robots by 2027 - Interesting Engineering","url":"https://interestingengineering.com/military/us-foundation-build-50000-humanoid-robots"},{"title":"Foundation Emerges With Phantom Humanoid - Humanoids Daily","url":"https://www.humanoidsdaily.com/news/foundation-emerges-with-phantom-humanoid-betting-on-novel-actuators-and-hybrid-ai"}],"as_of":"2025-12","related_ids":[null,null,null,null,null,null],"name":"Foundation Robotics Phantom","alt":"Foundation Phantom","abbr":"","aliases":["Phantom","Phantom MK1"],"one_liner":"A full-size humanoid from the US company Foundation, explicitly aimed at industrial and military high-risk tasks.","explanation":"Phantom is the full-size humanoid robot from Foundation, formally Foundation Robotics, a San Francisco startup founded by Sankaet Pathak in 2023; the current model is the Phantom MK1. It is reported to stand about 1.75 meters tall and weigh about 80 kg, using in-house backdrivable cycloidal actuators, and relies mainly on cameras for sensing rather than lidar. Unlike most humanoid companies, which focus on factories or the home, Foundation has openly named defense as a priority direction, with target tasks including reconnaissance and bomb disposal, meaning “send the robot in first” high-risk work, while stating that humans retain lethal decision-making authority. The company has set an aggressive production target: it is reported to plan deploying about 10,000 units in 2026 and reaching a cumulative 50,000 by the end of 2027, at a reported lease price of about $100,000 per unit per year. These are plans, not figures already achieved.","example":"Foundation envisions Phantom entering a suspicious area first in a bomb-disposal scenario to inspect and handle it, with an operator supervising from a distance.","related":["Foundation Robotics","Full-size Humanoid Robot","Special-purpose Robot","Cycloidal Reducer","Vision-Only Approach","Robot Rental"]},{"id":"clone-robotics-protoclone","category":"robot","sec":6,"tier":3,"sources":[{"title":"Watch creepy humanoid robot twitch and move with human-like skeleton - Interesting Engineering","url":"https://interestingengineering.com/innovation/video-worlds-first-humanoid-lifelike-muscles"},{"title":"Clone Robotics' Protoclone Has Over 1,000 Myofiber Artificial Muscles - TechEBlog","url":"https://www.techeblog.com/clone-robotics-protoclone-musculoskeletal-android-demo-video/"}],"as_of":"2025-02","related_ids":["artificial-muscle","bio-inspired-robot","humanoid-robot","pneumatic-artificial-muscle","uncanny-valley","clone-robotics"],"name":"Clone Robotics Protoclone","alt":"Clone Protoclone","abbr":"","aliases":["Protoclone","Protoclone V1"],"one_liner":"A humanoid from Clone Robotics built with a musculoskeletal design, using thousands of artificial muscles instead of motors.","explanation":"Protoclone is the “musculoskeletal humanoid” prototype V1 the startup Clone Robotics revealed publicly in February 2025. Rather than the usual route of motors plus reducers, it is built to follow human anatomy: a polymer skeleton modeled on the human body's 206 bones, moved by over 1,000 fluid-driven artificial muscles called Myofiber, said to give it more than 200 degrees of freedom; it carries about 500 sensors, including 4 depth cameras in the head, about 70 inertial units, and 320 pressure-sensing points, currently powered by a single 500 W electric pump. The public video shows it suspended and twitching its limbs, with no autonomous walking demonstrated yet. Its significance is exploring the “bio-inspired actuation” path: if artificial muscles can become fast and strong enough, a robot's compliance and appearance could get much closer to a human's, though control also becomes far harder.","example":"In Protoclone V1's release video, the robot, blank-faced and hung from a frame, twitches its limbs as its artificial muscles contract, sparking widespread “uncanny valley” discussion.","related":["Artificial Muscle","Bio-inspired Robot","Humanoid Robot","Pneumatic Artificial Muscle (McKibben Muscle)","Uncanny Valley","Clone Robotics"]},{"id":"hexagon-aeon","category":"robot","sec":6,"tier":3,"sources":[{"title":"Hexagon launches AEON, a humanoid built for industry - Hexagon Robotics","url":"https://robotics.hexagon.com/hexagon-launches-aeon-a-humanoid-built-for-industry/"},{"title":"Hexagon Enters Humanoid Arena with AEON Robot for Industry - Humanoids Daily","url":"https://www.humanoidsdaily.com/news/hexagon-enters-humanoid-arena-with-aeon-robot-for-industry"}],"as_of":"2025-06","related_ids":["wheel-legged-robot","humanoid-robot","industrial-robot","autonomous-battery-swapping","nvidia","maxon"],"name":"Hexagon AEON","alt":"Hexagon AEON","abbr":"","aliases":["AEON"],"one_liner":"A wheel-legged industrial humanoid robot from Sweden's Hexagon, launched in 2025 with autonomous battery swapping.","explanation":"AEON is an industrial humanoid robot that the Swedish measurement-technology company Hexagon launched on June 17, 2025, developed by its robotics division. Its lower body is “wheel-legged”: it has legs for posture adjustment, but wheels instead of feet, combining mobility with the ability to reach and balance. Reported specs include a height of 165 cm, a weight of 60 kg, 34 degrees of freedom, a 15 kg payload, and 12 onboard cameras plus a range of other sensors. It has a battery-swap mechanism that lets it keep working without downtime. AEON targets tasks such as loading and unloading parts, inspection, and “reality capture” — scanning objects with sensors to build 3D digital models — in industries like automotive, aerospace, and logistics. Partners include NVIDIA, Microsoft, and actuator maker maxon, with Schaeffler and Pilatus as its first pilot customers.","example":"At an aircraft-parts plant, AEON walks around a component with a handheld scanning tool, building a 3D digital model used for quality inspection.","related":["Wheel-legged Robot","Humanoid Robot","Industrial Robot","Autonomous Battery Swapping","NVIDIA","maxon"]},{"id":"humanoid-hmnd-01","category":"robot","sec":6,"tier":3,"sources":[{"title":"Humanoid Unveils Record Breaking Bipedal Robot Walking 48 Hours After Assembly - Humanoid","url":"https://thehumanoid.ai/humanoid-unveils-record-breaking-bipedal-robot-walking-48-hours-after-assembly/"},{"title":"HMND 01 ALPHA WHEELED - Humanoid","url":"https://thehumanoid.ai/hmnd-01-alpha-wheeled/"}],"as_of":"2025-12","related_ids":["humanoid","wheeled-humanoid-robot","bipedal-robot","humanoid-robot","proof-of-concept","palletizing-depalletizing"],"name":"Humanoid HMND 01","alt":"Humanoid HMND 01","abbr":"","aliases":["HMND 01","HMND 01 Alpha","HMND 01 Alpha Wheeled","HMND 01 Alpha Bipedal"],"one_liner":"A UK industrial humanoid robot line from Humanoid, offered in wheeled and bipedal versions.","explanation":"HMND 01 is the humanoid robot product line from Humanoid, a London-based company targeting industrial and logistics work. The first version, the wheeled HMND 01 Alpha, launched in September 2025: a humanoid upper body on an omnidirectional wheeled base, reportedly capable of 7.2 km/h with a 15 kg payload, and it had already completed its first commercial proof-of-concept trials for tasks like warehouse picking and palletizing. On December 2, 2025, Humanoid followed with a bipedal version, the HMND 01 Alpha Bipedal: 179 cm tall, with 29 degrees of freedom excluding the hands, a 15 kg combined arm payload, 3 hours of runtime on a swappable battery, and a choice of a 12-DOF dexterous hand or a simple gripper. The company says the bipedal version achieved stable walking within 48 hours of being assembled.","example":"In a logistics warehouse, the wheeled HMND 01 picks items off shelves and stacks them onto a pallet.","related":["Humanoid","Wheeled Humanoid Robot","Bipedal Robot","Humanoid Robot","Proof of Concept","Palletizing / Depalletizing"]},{"id":"rainbow-robotics-rb-y1","category":"robot","sec":6,"tier":3,"sources":[{"title":"RB-Y1 (Rainbow Robotics)","url":"https://rainbow-robotics.com/en/products/rb-y1/"},{"title":"Rainbow Robotics unveils RB-Y1 wheeled, two-armed robot (The Robot Report)","url":"https://www.therobotreport.com/rainbow-robotics-unveils-rb-y1-wheeled-two-armed-robot/"}],"as_of":"2025-01","related_ids":["rainbow-robotics","wheeled-humanoid-robot","mobile-manipulation","bimanual-manipulation","samsung-electronics","mobile-manipulator"],"name":"Rainbow Robotics RB-Y1","alt":"彩虹机器人 RB-Y1","abbr":"","aliases":["RB-Y1"],"one_liner":"A wheeled dual-arm humanoid mobile-manipulation platform from South Korea's Rainbow Robotics.","explanation":"RB-Y1 is a wheeled humanoid robot that South Korea's Rainbow Robotics — founded by members of the KAIST team behind the HUBO humanoid, with Samsung Electronics becoming its largest shareholder after increasing its stake in early 2025 — unveiled in March 2024. Its upper body has two 7-degree-of-freedom arms, its waist is a single 6-degree-of-freedom “leg”-like torso that can bend and rise and lower, and it sits on a wheeled base below. Published specs: about 131 kg total weight, about 1.4 m tall, roughly 3 kg of single-arm payload, and a top driving speed of about 1.5 m/s. By using wheels instead of legs, it avoids the hard problem of balance control and can focus on bimanual manipulation and mobile manipulation, which is why it's commonly used as a research platform for collecting demonstration data and training policies such as VLA models.","example":"A researcher teleoperates an RB-Y1 with VR to walk it up to a table indoors, bend at the waist, and lift a box with both arms, recording the data to train an imitation-learning policy.","related":["Rainbow Robotics","Wheeled Humanoid Robot","Mobile Manipulation","Bimanual Manipulation","Samsung Electronics","Mobile Manipulator"]},{"id":"dexmate-vega-series","category":"robot","sec":6,"tier":3,"sources":[{"title":"Dexmate Opens Preorders for Vega Mobile Humanoid Robot - Mike Kalil","url":"https://mikekalil.com/blog/dexmate-vega/"},{"title":"Vega - Dexmate Store","url":"https://shop.dexmate.com/products/vega"}],"as_of":"2026-09","related_ids":["wheeled-humanoid-robot","dual-arm-robot","mobile-manipulation","dexterous-hand","lifting-column","unified-robot-description-format"],"name":"Dexmate Vega Series","alt":"Dexmate Vega 轮式双臂人形","abbr":"","aliases":["Vega","Dexmate Vega"],"one_liner":"A US-made wheeled dual-arm humanoid with a foldable, telescoping torso, aimed at research and data collection.","explanation":"Vega is the first general-purpose robot from Dexmate, a startup based in Santa Clara, California, opened for preorder in 2025 with an early dexterous-hand configuration listed at about $89,900; as of September 2026, the Vega 1 Pro on the company's site lists at $72,000 excluding the hand. Its lower half is an omnidirectional wheeled base, and its upper half a humanoid structure with two arms and a head. Its most distinctive feature is a foldable, telescoping torso: it collapses to about 66 cm for transport and extends to reach up to about 2.2 meters high. The current Vega 1 Pro has 24 degrees of freedom in the body (7 per arm, 3 in the torso, 3 in the head, 4 in the base), a separately sold 6-DoF dexterous hand or gripper, RGBD cameras, lidar, force/torque sensing, over 10 hours of battery life, and a Python API plus URDF/USD models for simulators. It represents the wheeled-humanoid approach: trading bipedal walking for stability, battery life, and easier manipulation research.","example":"A lab can use Vega for mobile-manipulation data collection, having it retrieve an item from a tall cabinet at home and place it on a table.","related":["Wheeled Humanoid Robot","Dual-arm Robot","Mobile Manipulation","Dexterous Hand","Lifting Column","Unified Robot Description Format"]},{"id":"sunday-robotics-memo","category":"robot","sec":6,"tier":3,"sources":[{"title":"Sunday Robotics 官网","url":"https://www.sunday.ai/"},{"title":"SiliconANGLE: Sunday raises $165M at $1.15B valuation to launch Memo","url":"https://siliconangle.com/2026/03/12/sunday-raises-165m-1-15b-valuation-launch-memo-household-robot/"}],"as_of":"2026-03","related_ids":["sunday-robotics","skill-capture-glove","sunday-robotics-act-1","robot-free-data-collection","household-tasks","mobile-manipulation"],"name":"Sunday Robotics Memo","alt":"Sunday Memo","abbr":"","aliases":["Memo"],"one_liner":"Sunday Robotics' wheeled dual-arm home robot, trained on data captured by humans wearing a sensor glove.","explanation":"Memo is a home robot from the U.S. startup Sunday Robotics, unveiled in November 2025 when the company came out of stealth; it was founded by Tony Zhao, an author of ACT/ALOHA, and Cheng Chi, an author of Diffusion Policy and UMI. It's built as a wheeled base with a telescoping torso and two gripper-equipped arms, covered in a soft silicone shell; the company says it can lower to the floor and extend up to about 7 feet. Its most distinctive feature is that its training data doesn't come from teleoperating the robot itself — instead, people wear a “skill capture glove” to do household chores in real homes, and that data trains Sunday's own model, ACT-1. Demoed tasks include clearing a dining table, loading a dishwasher, folding socks, and making espresso. Sunday's website says a hand-built unit costs about $20,000, with a small-scale private beta planned for late 2026; in March 2026 the company reportedly closed a $165 million funding round at a $1.15 billion valuation.","example":"Clearing bowls and glasses off a dining table and loading them one by one into the dishwasher is one of Memo's signature long-horizon demos.","related":["Sunday Robotics","Skill Capture Glove","Sunday Robotics ACT-1","Robot-free (Embodiment-free) Data Collection","Household Tasks","Mobile Manipulation"]},{"id":"weave-robotics-isaac","category":"robot","sec":6,"tier":3,"sources":[{"title":"Weave Robotics 官网","url":"https://www.weaverobotics.com/"},{"title":"Isaac 1 产品页","url":"https://www.weaverobotics.com/isaac-1"}],"as_of":"2026-09","related_ids":["household-tasks","garment-manipulation","remote-teleoperation-takeover","consumer-grade-robot","1x-neo","sunday-robotics-memo"],"name":"Weave Robotics Isaac","alt":"Weave Isaac","abbr":"","aliases":["Isaac 0","Isaac 1"],"one_liner":"A home robot line from California's Weave Robotics, starting with a robot that just folds laundry.","explanation":"Isaac is the home-robot product line from the California company Weave Robotics. Isaac 0 is a stationary laundry-folding robot that does exactly one job — fold clean laundry neatly — and it's already been deployed in homes in California, with the company's site reporting over 2,000 cumulative hours of real-world operation. Isaac 1 is a mobile, dual-arm version: a wheeled base, a telescoping torso (roughly 0.9–1.75 m tall), and two 6-degree-of-freedom arms with a single-degree-of-freedom gripper, able to collect dirty laundry, fold clothes, make the bed, and put away toys and shoes. It runs autonomously by default, with a remote human operator stepping in to teleoperate it when needed. It's priced at $7,999 outright or $449 a month as a subscription, with the site saying the first units will ship in fall 2026.","example":"An Isaac 0 sits in a home's laundry area and automatically folds a basket of dried clothes, one item at a time.","related":["Household Tasks","Garment Manipulation","Remote Teleoperation Takeover","Consumer-Grade Robot","1X NEO","Sunday Robotics Memo"]},{"id":"lg-cloid","category":"robot","sec":6,"tier":3,"sources":[{"title":"LG Electronics Presents LG CLOiD Home Robot at CES 2026 - LG Newsroom","url":"https://www.lg.com/global/newsroom/news/home-appliance-solution/lg-electronics-presents-lg-cloid-home-robot-to-demonstrate-zero-labor-home-at-ces-2026/"},{"title":"CES 2026: LG to debut new CLOiD humanoid robot for the home - The Robot Report","url":"https://www.therobotreport.com/ces-2026-lg-to-debut-new-cloid-humanoid-robot-for-the-home/"}],"as_of":"2026-01","related_ids":["wheeled-humanoid-robot","household-tasks","vision-language-action-model","garment-manipulation","consumer-electronics-show","consumer-grade-robot"],"name":"LG CLOiD","alt":"LG CLOiD","abbr":"","aliases":["CLOiD"],"one_liner":"LG's wheeled home humanoid robot unveiled at CES 2026, demoed cooking and folding laundry.","explanation":"CLOiD is a home robot that LG Electronics showed at CES in January 2026, built around LG's idea of a “zero-labor home.” It has a wheeled base with a humanoid upper body: torso height adjusts across roughly 105–143 cm, each arm has 7 degrees of freedom with about 87 cm of reach, and each hand has five independently driven fingers. The head houses its chip, a display, cameras, and speakers. LG says it runs a vision-language model and a vision-language-action (VLA) model trained on tens of thousands of hours of household data, and can coordinate with LG appliances through the ThinQ platform. At the demo it took milk out of a refrigerator, put a croissant in the oven, and started a washing machine cycle and folded laundry. LG did not announce a price or release date.","example":"At the CES 2026 booth, CLOiD takes milk out of the fridge and puts a croissant in the oven to prepare breakfast.","related":["Wheeled Humanoid Robot","Household Tasks","Vision-Language-Action Model","Garment Manipulation","Consumer Electronics Show","Consumer-Grade Robot"]},{"id":"hrp-humanoid-robot-series","category":"robot","sec":6,"tier":3,"sources":[{"title":"Humanoid Robotics Project - Wikipedia","url":"https://en.wikipedia.org/wiki/Humanoid_Robotics_Project"},{"title":"Development of a humanoid robot prototype, HRP-5P, capable of heavy labor - Phys.org","url":"https://phys.org/news/2018-11-humanoid-robot-prototype-hrp-5p-capable.html"},{"title":"HRP-4 - ROBOTS Guide (IEEE Spectrum)","url":"https://robotsguide.com/robots/hrp4"}],"as_of":"2018-11","related_ids":["humanoid-robot","honda-asimo","zero-moment-point","zmp-preview-control","bipedal-locomotion","model-based-control"],"name":"HRP Humanoid Robot Series","alt":"HRP 系列人形机器人","abbr":"HRP","aliases":["HRP-2","HRP-4","HRP-4C","HRP-5P","Humanoid Robotics Project","AIST HRP"],"one_liner":"A long-running Japanese humanoid robot platform from AIST and Kawada, central to classic bipedal-walking research.","explanation":"The HRP series is a line of humanoid robots developed by Japan's National Institute of Advanced Industrial Science and Technology (AIST) together with Kawada Industries and other partners, named after the “Humanoid Robotics Project” backed by Japan's Ministry of Economy, Trade and Industry. Notable models include HRP-2 (2002), long used as a research platform for bipedal walking and whole-body control; HRP-4 (2010), a lighter, slimmer redesign that also had a female-appearance variant called HRP-4C; and HRP-5P (2018), built for heavy manual labor and demonstrated carrying and installing large sheets of drywall. The series is known for walking control based on the zero moment point (ZMP), a stability criterion, making it a landmark example of model-based humanoid control from before reinforcement-learning locomotion control became common.","example":"HRP-5P autonomously recognizes a sheet of drywall, lifts it, carries it to a wall, and installs it in place.","related":["Humanoid Robot","Honda ASIMO","Zero Moment Point","ZMP Preview Control","Bipedal Locomotion","Model-Based Control"]},{"id":"icub","category":"robot","sec":6,"tier":3,"sources":[{"title":"The iCub humanoid robot: an open platform for research in embodied cognition","url":"https://www.academia.edu/2732397/The_iCub_humanoid_robot_an_open_platform_for_research_in_embodied_cognition"},{"title":"Developing Advanced Control Software for the iCub Humanoid Robot - MathWorks","url":"https://it.mathworks.com/company/technical-articles/developing-advanced-control-software-for-the-icub-humanoid-robot.html"}],"as_of":"","related_ids":["humanoid-robot","embodied-cognition","small-size-humanoid-robot","open-source-hardware","softbank-robotics-nao"],"name":"iCub","alt":"iCub","abbr":"","aliases":["iCub3","IIT iCub"],"one_liner":"A child-sized open-source humanoid robot from Italy's IIT, built for embodied-cognition research.","explanation":"iCub grew out of the EU-funded RobotCub project, which began in 2004 and was led by the Italian Institute of Technology (IIT). Its research goal is embodied cognition — the idea that intelligence may only develop through a body interacting with its environment. iCub is about 104 cm tall and 22 kg, roughly the size of a small child, with 53 degrees of freedom across its body; a large share of them are concentrated in the hands, eyes, and head and neck, which makes it well suited to studying grasping and visual attention. Both the hardware design and the software (built on the YARP middleware) are released under open licenses such as the GPL, and more than 50 research institutions worldwide use it. IIT later built a taller version, iCub3, designed for remote-avatar teleoperation.","example":"Researchers let iCub handle a toy repeatedly, the way an infant would, to see whether it can learn the concept of an object purely through its own actions.","related":["Humanoid Robot","Embodied Cognition","Small-size Humanoid Robot","Open-Source Hardware (OSHW)","SoftBank Robotics NAO"]},{"id":"nasa-valkyrie","category":"robot","sec":6,"tier":3,"sources":[{"title":"Valkyrie (robot)（Wikipedia）","url":"https://en.wikipedia.org/wiki/Valkyrie_(robot)"},{"title":"VALKYRIE R5 Fact Sheet（NASA）","url":"https://www.nasa.gov/wp-content/uploads/2023/06/r5-fact-sheet.pdf"}],"as_of":"2015","related_ids":["humanoid-robot","darpa-robotics-challenge","whole-body-control","boston-dynamics-atlas","bipedal-robot","florida-institute-for-human-and-machine-cognition"],"name":"NASA Valkyrie (R5)","alt":"NASA Valkyrie 人形机器人","abbr":"","aliases":["Valkyrie","R5"],"one_liner":"A fully electric humanoid robot built by NASA's Johnson Space Center, officially designated R5.","explanation":"Valkyrie (officially designated R5) is a fully electric bipedal humanoid robot that NASA's Johnson Space Center began developing in October 2012. It was originally built to compete in DARPA's Robotics Challenge (DRC), which tested robots on disaster-response tasks like driving a car, opening doors, and climbing ladders, with a working prototype finished by July 2013. It stands about 1.87 m tall, weighs about 129 kg, has 44 degrees of freedom, and its battery lasts about an hour. At the DRC trials in December 2013, a network failure left it scoring zero points. NASA subsequently redirected the project toward space applications, with the idea of having humanoids do preparatory work on other planets before astronauts arrive: it has both run a simulated-robot competition based on Valkyrie's model and, in 2015, sent two physical R5 units to research teams at MIT and Northeastern University. Valkyrie is considered a classic research platform for humanoid whole-body control, balance, and teleoperation.","example":"","related":["Humanoid Robot","DARPA Robotics Challenge","Whole-Body Control","Boston Dynamics Atlas (Hydraulic)","Bipedal Robot","Florida Institute for Human and Machine Cognition"]},{"id":"pal-robotics-talos","category":"robot","sec":6,"tier":3,"sources":[{"title":"TALOS | High-Performance Humanoid Robot (PAL Robotics)","url":"https://pal-robotics.com/robot/talos/"},{"title":"TALOS: A new humanoid research platform targeted for industrial applications","url":"https://hal.science/hal-01485519v1/file/iros-talos.pdf"}],"as_of":"2026-09","related_ids":["humanoid-robot","torque-control","whole-body-control","joint-torque-sensor","crocoddyl","pal-robotics-tiago"],"name":"PAL Robotics TALOS","alt":"TALOS 人形机器人","abbr":"","aliases":["TALOS"],"one_liner":"A full-size research humanoid from Spain's PAL Robotics with torque control at every major joint.","explanation":"TALOS is a bipedal research humanoid that Spain's PAL Robotics introduced around 2017. It stands about 1.75 m tall, weighs about 95 kg, and has 32 degrees of freedom across its body; PAL says every major joint carries a torque sensor, with four additional six-axis force sensors at the ankles and wrists, and the joints communicate over an EtherCAT bus (an industrial real-time communication protocol). Each arm can carry about 6 kg fully extended. Its selling point is torque control — commanding how much force a joint outputs, rather than just what angle to move to — which is why European labs commonly use it to validate model-based algorithms for whole-body control, contact planning, and trajectory optimization. Compared with today's reinforcement-learning-driven humanoids, TALOS functions more like a standard reference platform for force-control research.","example":"The French lab LAAS-CNRS has used TALOS for years of experiments, and its open-source dynamics and optimal-control libraries, such as Pinocchio and Crocoddyl, have been validated on it.","related":["Humanoid Robot","Torque Control","Whole-Body Control","Joint Torque Sensor","Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)","PAL Robotics TIAGo"]},{"id":"softbank-robotics-nao","category":"robot","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Nao (robot)","url":"https://en.wikipedia.org/wiki/Nao_(robot)"},{"title":"The Robot Report: Aldebaran, maker of Pepper and Nao robots, put in receivership","url":"https://www.therobotreport.com/aldebaran-pepper-nao-robots-receivership/"}],"as_of":"2025-07","related_ids":["small-size-humanoid-robot","softbank-robotics-pepper","robocup","softbank-group","research-and-education-market","human-robot-interaction"],"name":"SoftBank Robotics NAO","alt":"NAO 机器人","abbr":"","aliases":["NAO"],"one_liner":"A 58 cm small humanoid from France's Aldebaran, a longtime staple of robotics education and RoboCup.","explanation":"NAO is a small bipedal humanoid robot developed by France's Aldebaran Robotics, with the first version released in the late 2000s. SoftBank acquired Aldebaran in 2012–2013 and later renamed it SoftBank Robotics, which is why the robot is often called “SoftBank's NAO”; in 2022, the German company United Robotics Group bought the business and reverted the name to Aldebaran. The sixth generation, NAO V6 (2018), stands about 57 cm tall, weighs about 5.2 kg, has 25 degrees of freedom, and carries cameras, microphones, and touch sensors; it can be programmed with graphical tools or Python. For many years it has been the designated robot for RoboCup's Standard Platform League, and it's one of the most common teaching humanoids at universities and schools worldwide. Aldebaran reportedly entered insolvency administration in 2025, after which its assets were acquired by the Shenzhen company Maxvision (盛视科技).","example":"In RoboCup's Standard Platform League soccer matches, every team competes with the same NAO hardware, so only the software differs.","related":["Small-size Humanoid Robot","SoftBank Robotics Pepper","RoboCup","SoftBank Group","Research & Education Market","Human-Robot Interaction"]},{"id":"softbank-robotics-pepper","category":"robot","sec":6,"tier":3,"sources":[{"title":"The Robot Report: Aldebaran, maker of Pepper and Nao robots, put in receivership","url":"https://www.therobotreport.com/aldebaran-pepper-nao-robots-receivership/"},{"title":"Humanoid Index: Pepper by SoftBank Robotics","url":"https://humanoidindex.org/robots/pepper"}],"as_of":"2025-07","related_ids":["service-robot","companion-robot","softbank-robotics-nao","wheeled-humanoid-robot","guided-tours-and-reception","softbank-group"],"name":"SoftBank Robotics Pepper","alt":"Pepper 机器人","abbr":"","aliases":["Pepper"],"one_liner":"SoftBank's 2014 wheeled companion-and-reception robot, with a chest tablet built for conversation and emotion recognition.","explanation":"Pepper is a social service robot co-developed by SoftBank and its subsidiary Aldebaran, released in June 2014 and put on sale in Japan in 2015. It stands about 1.2 m tall, with a humanlike upper body (head, arms, hands) on a three-wheeled omnidirectional base rather than legs — it cannot walk — and a 10.1-inch tablet on its chest, with about 20 degrees of freedom. Its selling points were voice conversation and reading people's facial expressions and emotions, aimed mainly at retail greeting, bank and hotel reception, and education. Pepper became a cautionary example of how hard it is for a “robot that can chat” to sustain real commercial demand: SoftBank reportedly paused production around 2021, and in 2025 its related assets were acquired together with Aldebaran.","example":"From roughly 2015 to 2018, many phone shops and banks in Japan had a Pepper stationed at the entrance to greet customers and explain services.","related":["Service Robot","Companion Robot","SoftBank Robotics NAO","Wheeled Humanoid Robot","Guided Tours & Reception","SoftBank Group"]},{"id":"sophia","category":"robot","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Sophia (robot)","url":"https://en.wikipedia.org/wiki/Sophia_(robot)"},{"title":"World Economic Forum: A robot has just been granted citizenship of Saudi Arabia","url":"https://www.weforum.org/stories/2017/10/a-robot-has-just-been-granted-citizenship-of-saudi-arabia/"}],"as_of":"2017-11","related_ids":["hyper-realistic-humanoid-robot","uncanny-valley","engineered-arts-ameca","human-robot-interaction","pre-programmed-motion","demo"],"name":"Sophia (Hanson Robotics)","alt":"Sophia（索菲亚）机器人","abbr":"","aliases":["Sophia the Robot"],"one_liner":"Hong Kong's Hanson Robotics's lifelike-faced social robot, famous for being granted Saudi “citizenship.”","explanation":"Sophia is a social humanoid robot developed by the Hong Kong company Hanson Robotics, founded by David Hanson. It was activated in February 2016 and made its public debut the following month at the SXSW conference in the U.S. Its focus is a hyper-realistic face capable of dozens of expressions; its body's movement ability is very limited, and early versions had no legs. In October 2017, Saudi Arabia granted it “citizenship,” and the following month it became the United Nations Development Programme's first non-human “Innovation Champion,” which turned it into a media sensation. Researchers have widely criticized its conversations as relying heavily on pre-programmed scripts, with its AI capabilities overstated in publicity. It's a frequent reference point in discussions of the uncanny valley and the gap between robot marketing and real-world capability.","example":"Sophia's “impromptu Q&A” sessions at various conferences have been reported by multiple outlets to actually use pre-prepared questions and answers.","related":["Hyper-Realistic Humanoid Robot","Uncanny Valley","Engineered Arts Ameca","Human-Robot Interaction","Pre-Programmed (Choreographed) Motion","Demo (Demonstration Video)"]},{"id":"engineered-arts-ameca","category":"robot","sec":6,"tier":3,"sources":[{"title":"Ameca (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Ameca_(robot)"},{"title":"Ameca Humanoid Robot From Engineered Arts to Debut at CES 2022 - Robotics 24/7","url":"https://www.robotics247.com/article/ameca_humanoid_robot_engineered_arts_debuts_ces_2022"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"Engineered Arts Ameca","alt":"Ameca 表情人形","abbr":"","aliases":["Ameca"],"one_liner":"A British humanoid robot from Engineered Arts, known for remarkably lifelike facial expressions.","explanation":"Ameca is a humanoid robot developed by the British company Engineered Arts, based in Cornwall, built in 2021 and first publicly shown at CES in January 2022. It has a grey rubber face and hands, deliberately designed to be gender-neutral, with a large number of small actuators in its head and face able to produce subtle expressions such as blinking, frowning, and looking surprised; it runs on the company's own Tritium cloud robot-operating software, which can connect to a large language model for conversation. Ameca cannot walk on its own and is used mainly for trade shows, museums, reception duty, and human-robot interaction research. It is often invoked in discussions of the uncanny-valley effect, meaning the unease people feel when something looks almost, but not quite, human, and it also illustrates that humanoid robotics has a branch beyond locomotion and manipulation: expressive interaction.","example":"At a trade show, Ameca connects to a large language model to converse with visitors, raising an eyebrow and smiling as it answers; the video spread widely on social media.","related":["Humanoid Robot","Hyper-realistic Humanoid Robot (Android)","Uncanny Valley","Human-Robot Interaction","Sophia (Hanson Robotics)","Guided Tours & Reception"]},{"id":"pollen-robotics-reachy-2","category":"robot","sec":6,"tier":3,"sources":[{"title":"Reachy 2 - The open-source humanoid for embodied AI (Pollen Robotics)","url":"https://pollen-robotics.com/reachy-2/"},{"title":"Hugging Face buys a humanoid robotics startup (TechCrunch)","url":"https://techcrunch.com/2025/04/14/hugging-face-buys-a-humanoid-robotics-startup"}],"as_of":"2025-04","related_ids":["pollen-robotics","hugging-face","pollen-robotics-reachy-mini","upper-body-humanoid-robot","vr-teleoperation","lerobot"],"name":"Pollen Robotics Reachy 2","alt":"Reachy 2","abbr":"","aliases":["Reachy"],"one_liner":"An open-source upper-body humanoid robot from France's Pollen Robotics, built for VR teleoperation and research.","explanation":"Reachy 2 is an open-source research humanoid that France's Pollen Robotics launched in October 2024; Hugging Face acquired Pollen in April 2025. It's an upper-body-only humanoid: two 7-degree-of-freedom arms, each able to lift about 3 kg, using Pollen's own Orbita parallel joints at the head and wrists, with parallel-jaw grippers at the ends. The head carries an RGB-D camera and a time-of-flight depth sensor, and an optional mobile base with three omnidirectional wheels and lidar can be added. The software runs on ROS 2 with a Python interface, and it ships with VR teleoperation support, making it easy to collect demonstration data for training imitation-learning policies. The full version with the mobile base has reportedly been priced at over $70,000.","example":"A researcher wears a VR headset to teleoperate a Reachy 2 folding laundry and tidying a desk, then uses the recorded data to train policies such as ACT with LeRobot.","related":["Pollen Robotics","Hugging Face","Pollen Robotics Reachy Mini","Upper-body Humanoid Robot","VR Teleoperation","LeRobot"]},{"id":"pollen-robotics-reachy-mini","category":"robot","sec":6,"tier":3,"sources":[{"title":"Reachy Mini - The Open-Source Robot for Today's and Tomorrow's AI Builders (Hugging Face Blog)","url":"https://huggingface.co/blog/reachy-mini"},{"title":"Hugging Face and Pollen Robotics launch $299 Reachy Mini robot (Investing.com)","url":"https://www.investing.com/news/company-news/hugging-face-and-pollen-robotics-launch-299-reachy-mini-robot-93CH-4128614"}],"as_of":"2026-09","related_ids":["pollen-robotics","hugging-face","pollen-robotics-reachy-2","companion-robot","open-source-hardware","human-robot-interaction"],"name":"Pollen Robotics Reachy Mini","alt":"Reachy Mini 桌面机器人","abbr":"","aliases":["Reachy Mini"],"one_liner":"A few-hundred-dollar open-source desktop robot from Hugging Face and Pollen Robotics.","explanation":"Reachy Mini is an open-source desktop robot that Hugging Face and its newly acquired Pollen Robotics launched in July 2025. It stands about 28 cm tall, weighs about 1.5 kg, and is sold as a kit that the buyer assembles. It has no arms — just a 6-degree-of-freedom head, a body that can rotate as a whole, and two moving antennas, along with a wide-angle camera, four microphones, and a speaker, mainly for expressive, conversational human-robot-interaction AI applications. It comes in two versions: a Lite edition that must stay tethered to a computer, and a wireless edition with its own Raspberry Pi, Wi-Fi, and battery. At launch it was priced from $299 (Lite) and $449 (wireless); the current prices listed on Pollen's site are $399 and $499. It's programmed through an open-source Python SDK, and finished applications can be shared on Hugging Face.","example":"A developer wires speech recognition and a large language model into a Reachy Mini so that it turns its head toward whoever says its name and responds by moving its antennas.","related":["Pollen Robotics","Hugging Face","Pollen Robotics Reachy 2","Companion Robot","Open-Source Hardware (OSHW)","Human-Robot Interaction"]},{"id":"hugging-face-hopejr","category":"robot","sec":6,"tier":3,"sources":[{"title":"Hugging Face unveils two new humanoid robots - TechCrunch","url":"https://techcrunch.com/2025/05/29/hugging-face-unveils-two-new-humanoid-robots/"},{"title":"Hugging Face Launches 2 New Affordable Humanoid Robots - Technology Org","url":"https://www.technology.org/2025/05/30/hugging-face-launches-2-new-affordable-humanoid-robots/"}],"as_of":"2025-05","related_ids":["hugging-face","lerobot","pollen-robotics-reachy-mini","pollen-robotics","open-source-hardware","so-100-so-101-arm"],"name":"Hugging Face HopeJR","alt":"HopeJR 开源人形","abbr":"","aliases":["HopeJR","Hope JR"],"one_liner":"Hugging Face's 2025 open-source full-size humanoid robot, with 66 degrees of freedom and a roughly $3,000 price target.","explanation":"HopeJR is an open-source full-size humanoid robot that Hugging Face unveiled in late May 2025, announced alongside its Reachy Mini desktop robot. It came about a month after Hugging Face acquired the French humanoid robotics company Pollen Robotics. HopeJR has 66 actuated degrees of freedom, can walk and move both arms, and is expected to sell for around US$3,000, with its hardware design released as open source. It's part of Hugging Face's open-robotics push built around its LeRobot software stack; the goal is to bring humanoid hardware down to a price ordinary labs and individual developers can afford, so more people can reproduce it, modify it, and collect their own data with it.","example":"A student team assembles a HopeJR from the open-source design and uses LeRobot to record demonstration data for training a manipulation policy.","related":["Hugging Face","LeRobot","Pollen Robotics Reachy Mini","Pollen Robotics","Open-Source Hardware (OSHW)","SO-100 / SO-101 Arm"]},{"id":"berkeley-humanoid-lite","category":"robot","sec":6,"tier":3,"sources":[{"title":"Demonstrating Berkeley Humanoid Lite (arXiv 2504.17249)","url":"https://arxiv.org/abs/2504.17249"},{"title":"Berkeley Humanoid Lite project page","url":"https://lite.berkeley-humanoid.org/"}],"as_of":"2025-04","related_ids":["open-source-hardware","small-size-humanoid-robot","cycloidal-reducer","3d-printing","berkeley-artificial-intelligence-research","sim-to-real-transfer"],"name":"Berkeley Humanoid Lite","alt":"伯克利 Humanoid Lite","abbr":"","aliases":[],"one_liner":"An open-source small humanoid from Berkeley, with most parts printable on a desktop 3D printer.","explanation":"Berkeley Humanoid Lite was released in April 2025 by UC Berkeley's Hybrid Robotics lab, paper arXiv 2504.17249, published at RSS 2025. It stands about 0.8 meters tall, weighs about 16 kg, and uses 3D-printed cycloidal reducers paired with motors at its joints; the hardware, code, and training environment are all fully open-sourced, with the paper estimating a parts cost of about $4,300 in the US and about $3,200 in China, both under $5,000 overall. It is meant to address the problem that humanoid robots are too expensive and the barrier to research too high: a student can print and assemble one themselves, then train a reinforcement-learning walking policy with Isaac Lab and deploy it to the real robot.","example":"The team deployed a reinforcement-learning policy trained in simulation onto their self-assembled Humanoid Lite to make it walk on the real robot.","related":["Open-Source Hardware (OSHW)","Small-size Humanoid Robot","Cycloidal Reducer","3D Printing (FDM / Resin SLA)","Berkeley Artificial Intelligence Research","Sim-to-Real Transfer"]},{"id":"toddlerbot","category":"robot","sec":6,"tier":3,"sources":[{"title":"ToddlerBot 项目主页","url":"https://toddlerbot.github.io/"},{"title":"ToddlerBot: Open-Source ML-Compatible Humanoid Platform for Loco-Manipulation (arXiv)","url":"https://arxiv.org/abs/2502.00893"}],"as_of":"2025-09","related_ids":["small-size-humanoid-robot","open-source-hardware","humanoid-robot","sim-to-real-transfer","diffusion-policy","vr-teleoperation"],"name":"ToddlerBot","alt":"ToddlerBot","abbr":"","aliases":[],"one_liner":"Stanford's open-source toddler-sized humanoid robot, fully 3D-printed and built for machine-learning research.","explanation":"ToddlerBot is an open-source small humanoid robot from Stanford researchers Haochen Shi, Weizhuo Wang, Shuran Song, C. Karen Liu, and others, published at CoRL 2025. It stands about 0.56 m tall and weighs about 3.4 kg in the paper's version (later versions are slightly heavier, around 3.7 kg), with 30 actively driven degrees of freedom (7 per arm, 6 per leg, and 2 each in the neck and waist). The entire robot is 3D-printed from off-the-shelf parts, keeping total cost under $6,000, and the team released both the code and the assembly documentation. It's meant to solve the problem that humanoid robots are usually too expensive and too fragile to use for large-scale data collection and experimentation: because a ToddlerBot is cheap and quick to repair after a fall, researchers can freely run reinforcement-learning walking, collect data via VR teleoperation, and train diffusion policies for whole-body tasks like grasping and pushing boxes; it's also commonly used to validate sim-to-real transfer.","example":"In the paper, demonstration data is collected through VR teleoperation, then used to train a diffusion policy so ToddlerBot can perform whole-body manipulation tasks like grasping and pushing.","related":["Small-size Humanoid Robot","Open-Source Hardware (OSHW)","Humanoid Robot","Sim-to-Real Transfer","Diffusion Policy","VR Teleoperation"]},{"id":"disney-research-bdx-droid","category":"robot","sec":6,"tier":3,"sources":[{"title":"Design and Control of a Bipedal Robotic Character - Disney Research","url":"https://la.disneyresearch.com/publication/design-and-control-of-a-bipedal-robotic-character/"},{"title":"Design and Control of a Bipedal Robotic Character - arXiv","url":"https://arxiv.org/html/2501.05204v1"},{"title":"BDX Droids - Disney Research","url":"https://la.disneyresearch.com/bdx-droids/"}],"as_of":"2025","related_ids":["bipedal-robot","rl-based-locomotion-control","motion-tracking","open-duck-mini","sim-to-real-transfer","newton-physics-engine"],"name":"Disney Research BDX Droid","alt":"迪士尼 BDX 机器人","abbr":"","aliases":["BDX","BD-X","BDX Droids"],"one_liner":"Disney's Star Wars-style bipedal droid, using reinforcement learning to bring animator-designed moves to real hardware.","explanation":"BDX is a bipedal character robot developed jointly by Disney Research and Walt Disney Imagineering, styled after the small droids from Star Wars, used to interact with guests at Disney theme parks and similar venues. Its technical approach is described in the paper “Design and Control of a Bipedal Robotic Character” (RSS 2024): animators first design expressive motions, such as a walk cycle or a curious head-turn, and reinforcement learning then trains a control policy in simulation that can perform those motions on the real robot while staying balanced; at runtime, an animation engine blends multiple animation clips into commands, and an operator triggers behaviors with a remote control. It demonstrates that “expressive, animator-designed motion” and “robust locomotion control” can be achieved together. It also appeared on stage at NVIDIA's GTC in 2025, and the open-source community has since built a replica project, Open Duck Mini.","example":"BDX waddles up to park guests and tilts its head to “greet” them; these expressive movements were designed by animators and are executed by a reinforcement-learning policy.","related":["Bipedal Robot","RL-based Locomotion Control","Motion Tracking","Open Duck Mini","Sim-to-Real Transfer","Newton Physics Engine"]},{"id":"open-duck-mini","category":"robot","sec":6,"tier":3,"sources":[{"title":"apirrone/Open_Duck_Mini（GitHub）","url":"https://github.com/apirrone/Open_Duck_Mini"},{"title":"apirrone/Open_Duck_Playground（GitHub）","url":"https://github.com/apirrone/Open_Duck_Playground"}],"as_of":"2025-02","related_ids":["disney-research-bdx-droid","open-source-hardware","mujoco-playground","sim-to-real-transfer","bipedal-robot","pollen-robotics"],"name":"Open Duck Mini","alt":"Open Duck Mini","abbr":"","aliases":["Open Duck"],"one_liner":"An open-source small bipedal robot modeled on Disney's BDX droid, with a target parts cost under $400.","explanation":"Open Duck Mini is an open-source project started by GitHub developer apirrone, aiming to recreate a miniature version of the Disney BDX droid — the small, expressive walking character robot seen at Disney theme parks. The v2 version stands about 42 cm tall with its legs extended, targets a parts cost under $400, runs on a Raspberry Pi Zero 2W as its onboard computer, and has its code released under the Apache 2.0 license, with sponsorship from Hugging Face and Pollen Robotics. Its walking policy is trained with reinforcement learning in the MuJoCo simulator, then exported as an ONNX model for deployment on the real robot; the project also open-sources a reference-motion generator and a training environment built on MuJoCo Playground. It's a popular low-cost entry point for individuals and students to try “train bipedal walking in simulation, then transfer it to a real robot.”","example":"A hobbyist 3D-prints the shell from the open-source design to assemble an Open Duck Mini, exports a walking policy trained in MuJoCo Playground as an ONNX model, and deploys it to run on the onboard Raspberry Pi.","related":["Disney Research BDX Droid","Open-Source Hardware (OSHW)","MuJoCo Playground","Sim-to-Real Transfer","Bipedal Robot","Pollen Robotics"]},{"id":"boston-dynamics-bigdog","category":"robot","sec":7,"tier":3,"sources":[{"title":"BigDog - Wikipedia","url":"https://en.wikipedia.org/wiki/BigDog"},{"title":"BigDog - ROBOTS: Your Guide to the World of Robotics","url":"https://robotsguide.com/robots/bigdog"}],"as_of":"","related_ids":["quadruped-robot","boston-dynamics","hydraulic-actuation","boston-dynamics-spot","dynamic-stability","push-recovery"],"name":"Boston Dynamics BigDog","alt":"波士顿动力 BigDog（大狗）","abbr":"","aliases":["BigDog"],"one_liner":"Boston Dynamics' military hydraulic quadruped from 2005 onward, a landmark in dynamic-balance legged robots.","explanation":"BigDog was developed by Boston Dynamics together with Harvard's Concord Field Station and other partners, funded by DARPA starting in 2005, with the goal of carrying supplies for soldiers over rough terrain. It is about 1 meter long and weighs about 109 kg, with an internal combustion engine driving a hydraulic pump that powers its four legs, reportedly able to carry a load of about 150 kg. It is best known for a video where it stayed on its feet after being kicked from the side and slipping on ice, which was many people's first look at a legged robot with genuine dynamic balance. A later military version, the LS3, was never adopted, reportedly in part because it was too loud. Its technology lineage continued into the later electric quadruped, Spot.","example":"In BigDog's demonstration video, an engineer kicks the robot from the side and it slips on ice, but it adjusts its steps and recovers its balance both times.","related":["Quadruped Robot","Boston Dynamics","Hydraulic Actuation","Boston Dynamics Spot","Dynamic Stability","Push Recovery"]},{"id":"boston-dynamics-spot","category":"robot","sec":7,"tier":1,"sources":[{"title":"Boston Dynamics Launches Commercial Sales of Spot Robot","url":"https://bostondynamics.com/news/boston-dynamics-launches-commercial-sales-of-spot-robot/"},{"title":"Spot | Boston Dynamics","url":"https://bostondynamics.com/products/spot/"}],"as_of":"2020-06","related_ids":["quadruped-robot","boston-dynamics","inspection-robot","legged-mobile-manipulator","anybotics-anymal","unitree-go2"],"name":"Boston Dynamics Spot","alt":"波士顿动力 Spot","abbr":"","aliases":["Spot"],"one_liner":"Boston Dynamics' commercial quadruped robot dog, used mainly for industrial inspection.","explanation":"Spot is the quadruped robot from Boston Dynamics that went on commercial sale in June 2020, when the base kit was priced at $74,500 (now quote-only). It has 12 degrees of freedom across its four legs, weighs about 32.5 kg, can carry roughly 14 kg on its back, reaches a top speed of about 1.6 m/s, and can climb stairs and cross slopes and rubble. Its back carries an expansion interface for payloads such as a robot arm, thermal camera, or lidar, and it is commonly used for automated inspection in factories and power plants and for surveying hazardous sites. It was one of the first legged robots to be commercialized at scale, and is often treated as the reference example of a quadruped robot.","example":"A power plant runs Spot on a fixed patrol route to read gauges and check equipment temperature.","related":["Quadruped Robot","Boston Dynamics","Inspection Robot","Legged Mobile Manipulator (Quadruped with Arm)","ANYbotics ANYmal","Unitree Go2"]},{"id":"boston-dynamics-stretch","category":"robot","sec":7,"tier":3,"sources":[{"title":"Boston Dynamics' Stretch robot handles truck unloading & palletizing (The Robot Report)","url":"https://www.therobotreport.com/boston-dynamics-stretch-robot-truck-unloading-palletizing/"},{"title":"Stretch - ROBOTS: Your Guide to the World of Robotics","url":"https://robotsguide.com/robots/bdstretch"}],"as_of":"2026-09","related_ids":["boston-dynamics","palletizing-depalletizing","vacuum-suction-cup","mobile-manipulator","tote-handling","real-world-deployment"],"name":"Boston Dynamics Stretch","alt":"波士顿动力 Stretch","abbr":"","aliases":[],"one_liner":"Boston Dynamics' warehouse unloading robot, using suction to pull boxes out of containers automatically.","explanation":"Stretch is the first robot Boston Dynamics built specifically for warehouses, released in 2021, meant to handle the heavy, repetitive work of unloading shipping containers and trucks. Its structure is a four-wheel omnidirectional mobile base topped with a custom 7-DoF robot arm ending in a multi-suction-cup smart gripper, with a sensing mast on the base handling box recognition. It can pull boxes out of a container and place them on a conveyor, with the company stating it can move up to about 800 boxes an hour, each up to about 23 kg (50 lbs), on roughly 8 hours of battery life. It illustrates that in a well-defined use case, a specialized non-humanoid form often reaches real deployment faster than a humanoid one.","example":"Inside a logistics warehouse, Stretch drives into a shipping container, suction-lifts cartons one by one, and places them onto an extendable conveyor to unload the container.","related":["Boston Dynamics","Palletizing / Depalletizing","Vacuum Suction Cup","Mobile Manipulator","Tote Handling","Real-world Deployment"]},{"id":"mit-mini-cheetah","category":"robot","sec":7,"tier":2,"sources":[{"title":"MIT: Mini cheetah is the first four-legged robot to do a backflip","url":"https://robotics.mit.edu/mini-cheetah-first-four-legged-robot-do-backflip/"},{"title":"IEEE Spectrum: How MIT's Mini Cheetah Can Help Accelerate Robotics Research","url":"https://spectrum.ieee.org/mit-mini-cheetah-accelerate-research"}],"as_of":"2019-03","related_ids":["quadruped-robot","quasi-direct-drive","mit-mini-cheetah-actuator","mit-mode","legged-robot","unitree-robotics"],"name":"MIT Mini Cheetah","alt":"MIT Mini Cheetah","abbr":"","aliases":[],"one_liner":"MIT Biomimetic Robotics Lab's small quadruped, the first to land a backflip, in 2019.","explanation":"Mini Cheetah is a small quadruped robot built by MIT's Biomimetic Robotics Lab, led by Sangbae Kim: about 20 pounds, roughly 9 kg, with 12 modular motors, giving each leg 2 hip degrees of freedom plus 1 knee joint. In 2019 it became the first quadruped robot to perform a backflip. Its biggest influence has been in actuator design: pairing a high torque-density motor with a low-reduction-ratio planetary gearbox, called quasi-direct drive, lets it absorb impacts while still being backdrivable and able to estimate torque from motor current. This “MIT Cheetah actuator,” along with its accompanying MIT-mode motor command style, was later adopted by many manufacturers, including Unitree. The team also lent multiple units to other labs, making it a common early platform for reinforcement-learning locomotion research.","example":"The MIT team computed the backflip's motion and per-motor torques offline through trajectory optimization, then had Mini Cheetah execute the backflip on real hardware.","related":["Quadruped Robot","Quasi-Direct Drive","MIT Mini Cheetah Actuator","MIT Mode","Legged Robot","Unitree Robotics"]},{"id":"anybotics-anymal","category":"robot","sec":7,"tier":3,"sources":[{"title":"ANYmal - ANYbotics","url":"https://www.anybotics.com/robotics/anymal/"},{"title":"ANYmal - Robotic Systems Lab, ETH Zurich","url":"https://rsl.ethz.ch/robots-media/anymal.html"}],"as_of":"2026-09","related_ids":["quadruped-robot","anybotics","eth-zurich-robotic-systems-lab","anymal-rl-locomotion-series","inspection-robot","perceptive-locomotion"],"name":"ANYbotics ANYmal","alt":"ANYmal 四足","abbr":"","aliases":["ANYmal"],"one_liner":"A Swiss industrial-inspection quadruped from ANYbotics, and a classic platform for legged reinforcement-learning research.","explanation":"ANYmal originated at ETH Zurich's Robotic Systems Lab, with its first version appearing in 2016, the same year the team founded ANYbotics to commercialize it. The current ANYmal D weighs about 50 kg unloaded and can carry roughly 10 kg of inspection payload; ANYmal X is an explosion-proof version for sites with explosion risk, such as petrochemical plants. It is sold mainly to power plants and the oil, gas, and chemical industries for autonomous inspection, reading gauges, checking temperatures, and listening for abnormal sounds in place of a person. In academia, ANYmal is a landmark platform for legged locomotion-control research: landmark work on actuator networks, which fit a neural network to motor characteristics, teacher-student distillation for blind locomotion, and perceptive locomotion combining terrain sensing was all done on it and published in Science Robotics.","example":"ETH's Miki and colleagues had ANYmal, in 2022, combine exteroception and proprioception to complete a hiking trail up Switzerland's Mount Etzel.","related":["Quadruped Robot","ANYbotics","ETH Zurich Robotic Systems Lab","ANYmal RL Locomotion Series","Inspection Robot","Perceptive Locomotion"]},{"id":"ghost-robotics-vision-60","category":"robot","sec":7,"tier":3,"sources":[{"title":"Vision 60 | Ghost Robotics","url":"https://www.ghostrobotics.io/vision-60"},{"title":"LIG Nex1 takes controlling shares of Ghost Robotics for $240M - The Robot Report","url":"https://www.therobotreport.com/lig-nex1-takes-controlling-shares-of-ghost-robotics-for-240m/"}],"as_of":"2025-12","related_ids":["quadruped-robot","boston-dynamics-spot","anybotics-anymal","special-purpose-robot","ingress-protection-rating","legged-mobile-manipulator"],"name":"Ghost Robotics Vision 60","alt":"Ghost Vision 60 四足","abbr":"","aliases":["Vision 60","V60","Q-UGV"],"one_liner":"A mid-size all-weather quadruped robot from Philadelphia's Ghost Robotics, used mainly for defense and security.","explanation":"Vision 60 is a mid-size quadruped robot from Ghost Robotics, based in Philadelphia; the company's own designation for it is Q-UGV (quadruped unmanned ground vehicle). It's built mainly for defense, security, and industrial inspection work. Published specs: about 51 kg unloaded weight, roughly 10 kg payload, an IP67 rating (dust-tight and able to survive brief submersion), an operating range of −40 to 55 °C, a top speed of about 3 m/s, and roughly 10 km of range per charge. It has patrolled multiple U.S. Air Force bases since late 2020. In July 2024, the South Korean defense company LIG Nex1 acquired a 60% stake in Ghost Robotics for about US$240 million. Compared with Boston Dynamics' Spot, Vision 60 puts more emphasis on durability in harsh environments, and can carry mission payloads such as sensors and robotic arms.","example":"A U.S. Air Force base runs Vision 60 with a mounted camera on automated patrols along its perimeter fence.","related":["Quadruped Robot","Boston Dynamics Spot","ANYbotics ANYmal","Special-purpose Robot","Ingress Protection (IP) Rating","Legged Mobile Manipulator (Quadruped with Arm)"]},{"id":"amazon-vulcan","category":"robot","sec":7,"tier":3,"sources":[{"title":"About Amazon：Introducing Vulcan, Amazon's first robot with a sense of touch","url":"https://www.aboutamazon.com/news/operations/amazon-vulcan-robot-pick-stow-touch"},{"title":"Amazon Science：How Amazon's Vulcan robots use touch to plan and execute motions","url":"https://www.amazon.science/blog/how-amazons-vulcan-robots-use-touch-to-plan-and-execute-motions"}],"as_of":"2025-05","related_ids":["amazon-robotics","tactile-sensor","six-axis-force-torque-sensor","vacuum-suction-cup","order-picking","contact-rich-manipulation"],"name":"Amazon Vulcan","alt":"Amazon Vulcan 触觉拣货机器人","abbr":"","aliases":["Vulcan"],"one_liner":"Amazon's first warehouse picking robot with a sense of touch.","explanation":"Vulcan is the warehouse robot Amazon unveiled in Dortmund, Germany, in May 2025, which the company describes as its first robot with a sense of “touch.” Its end effector uses force-feedback sensors to detect when it makes contact with an object and how much force to apply, letting it push other items aside inside a tightly packed storage bin to place or retrieve goods. It has two kinds of end effectors: one resembling a flat blade with a small built-in conveyor belt, used for stowing items, and a camera-equipped suction-cup arm, used for picking. It can handle around 75% of the items in a warehouse, and is used mainly for the highest and lowest shelves that human workers struggle to reach, and is already in use at Amazon warehouses in Spokane, Washington, and Hamburg, Germany.","example":"Vulcan uses its blade-like end effector to push items in a bin to one side, clearing space to stow new inventory.","related":["Amazon Robotics","Tactile Sensor","Six-Axis Force/Torque Sensor","Vacuum Suction Cup","Order Picking","Contact-rich Manipulation"]},{"id":"da-vinci-surgical-system-da-vinci-research-kit","category":"robot","sec":7,"tier":3,"sources":[{"title":"The da Vinci Research Kit (dVRK) - Intuitive Foundation","url":"https://www.intuitive-foundation.org/dvrk/"},{"title":"What is the dVRK? — dVRK documentation","url":"https://dvrk.readthedocs.io/main/pages/introduction/what_is_it.html"}],"as_of":"","related_ids":["surgical-robot","teleoperation","leader-follower-teleoperation","imitation-learning","robot-operating-system"],"name":"da Vinci Surgical System / da Vinci Research Kit (dVRK)","alt":"达芬奇手术机器人 / dVRK","abbr":"dVRK","aliases":["dVRK","da Vinci","da Vinci Research Kit"],"one_liner":"Intuitive Surgical's leader-follower minimally invasive surgical robot, with dVRK its retired hardware turned open research platform.","explanation":"The da Vinci Surgical System, made by the American company Intuitive Surgical, is the dominant minimally invasive surgical robot: a surgeon sits at a console viewing a 3D endoscope feed and moves master controllers, while robotic arms at the patient's bedside reproduce the motion at scale, a form of leader-follower teleoperation. The da Vinci Research Kit (dVRK) is a research platform built starting in 2012 by Johns Hopkins University, Worcester Polytechnic Institute, and Intuitive Surgical: retired da Vinci arms and master controllers are fitted with open-source electronics, firmware, and software and integrated with ROS, and it is now deployed at more than 40 institutions worldwide. It lets researchers read and write joint data directly and record surgical demonstrations, making it a common testbed for surgical automation, teleoperation interfaces, and imitation learning on surgical scenes.","example":"Teams at Johns Hopkins and elsewhere have recorded suturing and needle-passing demonstrations on the dVRK, then used imitation learning to train policies that perform these sub-tasks autonomously.","related":["Surgical Robot","Teleoperation","Leader-Follower Teleoperation","Imitation Learning","Robot Operating System"]},{"id":"unitree-g1","category":"robot","sec":8,"tier":1,"sources":[{"title":"Unitree G1 官网","url":"https://www.unitree.com/g1/"},{"title":"Unitree Robotics unveils G1 humanoid for $16K - The Robot Report","url":"https://www.therobotreport.com/unitree-robotics-unveils-g1-humanoid-for-16k/"}],"as_of":"2026-09","related_ids":["unitree-robotics","unitree-h1","small-size-humanoid-robot","unitree-rl-gym","unitree-sdk2","unitree-dex3-1"],"name":"Unitree G1","alt":"宇树 G1","abbr":"","aliases":["G1"],"one_liner":"Unitree's small bipedal humanoid robot, one of the most common humanoid platforms in research.","explanation":"G1 is the bipedal humanoid robot Hangzhou-based Unitree Robotics released in May 2024, launched starting at RMB 99,000 (about $16,000), bringing humanoid robot prices into a range university labs could afford. Published specs: standing height about 1.32 meters, weight about 35 kg including battery, 23 to 43 joint motors across the body depending on configuration (the EDU version can add a waist joint and dexterous hands), 6 degrees of freedom per leg, about 2 hours of battery life, and a foldable, storable frame. Because it is inexpensive, ships with an open SDK, and has ready-made reinforcement-learning training code, it has become one of the most common real robots used in locomotion control, whole-body teleoperation, and humanoid VLA papers.","example":"Many humanoid motion-imitation papers, such as BeyondMimic and TWIST, demonstrate their results on a real Unitree G1.","related":["Unitree Robotics","Unitree H1","Small-size Humanoid Robot","unitree_rl_gym (Unitree RL Gym)","Unitree SDK2","Unitree Dex3-1"]},{"id":"unitree-r1","category":"robot","sec":8,"tier":2,"sources":[{"title":"Unitree launches $5,900 humanoid robot (Robotics and Automation News)","url":"https://roboticsandautomationnews.com/2025/07/29/shock-price-unitree-launches-5900-humanoid-robot/93357/"},{"title":"Unitree R1 – Unitree 官方商城","url":"https://shop.unitree.com/products/unitree-r1"}],"as_of":"2026-04","related_ids":[null,null,null,null,null],"name":"Unitree R1","alt":"宇树 R1","abbr":"","aliases":["R1","R1-Air","R1 EDU"],"one_liner":"Unitree's 2025 budget small humanoid, starting at about RMB 39,900.","explanation":"R1 is the small humanoid robot Unitree Robotics released in July 2025: about 1.23 meters tall, weighing 25–29 kg depending on configuration, with 20–26 degrees of freedom depending on configuration, able to do a cartwheel, get up from the ground on its own, and run downhill. What drew the most attention was its price: it starts at about RMB 39,900, roughly $5,900, with a cheaper R1-Air version also available, bringing humanoid robots into a range individual developers and schools can afford; its overseas official store lists it at about $4,900. It is used mainly for education, research, and entertainment.","example":"A university lab uses several R1 units for a humanoid motion-control course, having students train a policy in simulation and then test it directly on the real robot.","related":["Unitree Robotics","Small-size Humanoid Robot","10,000-Yuan-Class Humanoid Robot","Unitree G1","Consumer-Grade Robot"]},{"id":"unitree-h1","category":"robot","sec":8,"tier":2,"sources":[{"title":"Unitree H1 - ROBOTS: Your Guide to the World of Robotics (IEEE)","url":"https://robotsguide.com/robots/unitree-h1"},{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"}],"as_of":"2025-02","related_ids":[null,null,null,null,null,null],"name":"Unitree H1","alt":"宇树 H1","abbr":"","aliases":["H1","H1-2"],"one_liner":"Unitree's first full-size general-purpose humanoid, which once set a humanoid running-speed record.","explanation":"H1 is the full-size bipedal humanoid robot Unitree Robotics released in August 2023: about 1.8 meters tall, roughly 47 kg, with a peak knee-joint torque of about 360 N·m, and a head carrying 3D lidar and a depth camera. It is best known for its mobility, reaching a running speed of 3.3 m/s, at one point the fastest recorded for a full-size humanoid; the H1 that performed the Yangge folk dance in the “Yangbot” segment of the 2025 CCTV Spring Festival Gala, China's most-watched TV broadcast, was also an H1. The upgraded H1-2 added more degrees of freedom in the arms, meaning more independently movable joints. H1 was one of the earlier full-size humanoids available for purchase, and many papers on humanoid whole-body control and teleoperation have run experiments on it.","example":"Works on humanoid teleoperation such as H2O and OmniH2O ran real-robot experiments on the H1, having the robot mimic an operator's whole-body motion in real time.","related":["Unitree Robotics","Full-size Humanoid Robot","Unitree H2","Unitree G1","Yangbot","H2O"]},{"id":"unitree-h2","category":"robot","sec":8,"tier":2,"sources":[{"title":"Unitree Unveils H2, a Full-Sized Humanoid Successor to the H1 (Humanoids Daily)","url":"https://www.humanoidsdaily.com/news/unitree-unveils-h2-humanoid-successor-to-h1"}],"as_of":"2025-10","related_ids":[null,null,null,null,null],"name":"Unitree H2","alt":"宇树 H2","abbr":"","aliases":["H2","H2 Plus"],"one_liner":"Unitree's October 2025 full-size humanoid, the successor to the H1.","explanation":"H2 is the full-size bio-inspired humanoid robot Unitree Robotics released on October 20, 2025: about 1.8 meters tall, roughly 70 kg, with 31 degrees of freedom in total, 6 per leg, 7 per arm, 3 in the waist, and 2 in the neck. Compared with H1, it no longer chases running speed and instead adds flexibility in the arms and waist, with an optional dexterous hand, aimed at manipulation tasks that require actually getting work done; it also has a bio-inspired face. Its release video showed dancing, boxing, martial-arts moves, and a runway walk.","example":"","related":["Unitree Robotics","Unitree H1","Full-size Humanoid Robot","Degrees of Freedom","Unitree G1"]},{"id":"agibot-expedition-a2-series","category":"robot","sec":8,"tier":2,"sources":[{"title":"AgiBot - Wikipedia","url":"https://en.wikipedia.org/wiki/AgiBot"},{"title":"AgiBot A2 Specs & Price | Humanoid.guide","url":"https://humanoid.guide/product/a2/"}],"as_of":"2025-11","related_ids":["agibot","humanoid-robot","wheeled-humanoid-robot","agibot-a3","agibot-genie-g1","7-dof-robot-arm"],"name":"AgiBot Expedition A2 Series","alt":"智元 远征 A2 系列（A2 / A2-W）","abbr":"","aliases":["Expedition A2","AgiBot A2","A2-W","A2 Ultra","A2 Lite"],"one_liner":"AgiBot's 2024 Expedition humanoid line, including a bipedal version and a wheeled version, A2-W.","explanation":"Expedition A2 is a humanoid robot series released by AgiBot in August 2024, aimed at interactive service and industrial settings. The bipedal A2 is a full-size humanoid, reportedly about 1.7–1.75 meters tall, with 7 degrees of freedom in each arm and an optional dexterous hand; A2-W pairs the same upper body with a wheeled base for factory material handling, and is not simply an A2 with wheels bolted on. The line was later split further into configurations such as A2 Ultra and A2 Lite. In November 2025, an A2 walked from Suzhou to Shanghai, covering about 106 km over three days and setting a Guinness World Record. Together with the Genie G series and Lingxi X series, the A2 series forms AgiBot's humanoid product line, commonly seen on showroom reception and tour-guide duty and in factory pilots.","example":"In November 2025, an Expedition A2 walked from Suzhou to Shanghai over three days, covering about 106 km, and earned a Guinness World Record.","related":["AgiBot","Humanoid Robot","Wheeled Humanoid Robot","AgiBot A3","AgiBot Genie G1","7-DoF Robot Arm"]},{"id":"agibot-a3","category":"robot","sec":8,"tier":3,"sources":[{"title":"IT之家：智元新一代全尺寸人形机器人远征 A3 发布","url":"https://www.ithome.com/0/938/490.htm"},{"title":"腾讯新闻：智元机器人正式发布全新一代全尺寸人形机器人远征A3","url":"https://news.qq.com/rain/a/20260213A0788L00"}],"as_of":"2026-04","related_ids":["agibot","full-size-humanoid-robot","agibot-expedition-a2-series","commercial-robot-performances","dexterous-hand","hot-swappable-battery"],"name":"AgiBot A3","alt":"智元 远征 A3","abbr":"","aliases":["Expedition A3","AgiBot Expedition A3"],"one_liner":"AgiBot's 2026 full-size bipedal humanoid, aimed at stage performance and guided tours.","explanation":"Expedition A3 is AgiBot's next-generation full-size humanoid robot, first shown in February 2026, with main specs announced at an April event: 173 cm tall, 55 kg, a rated battery life of 10 hours, and support for fast battery swaps. It reportedly has 31 active degrees of freedom, 2 in the neck, 7 per arm, 3 in the waist, and 6 per leg, a 3 kg payload at the end of each arm, and an optional dexterous hand. It targets human-robot interaction scenarios such as stage performance, guided tours, and product demos, and carries UWB, meaning ultra-wideband wireless positioning, for centimeter-level location tracking, used to synchronize formations of over a hundred robots performing together. It is the successor to the Expedition A2 series.","example":"In official demos, multiple Expedition A3 units use centimeter-level positioning to align their formation and perform a coordinated group dance.","related":["AgiBot","Full-size Humanoid Robot","AgiBot Expedition A2 Series","Commercial Robot Performances","Dexterous Hand","Hot-Swappable Battery"]},{"id":"agibot-genie-g1","category":"robot","sec":8,"tier":3,"sources":[{"title":"IT之家：智元机器人全系产品开售，精灵 G1 售价 45 万元","url":"https://www.ithome.com/0/876/106.htm"},{"title":"智元官网：精灵G1","url":"https://www.agibot.com.cn/a2dproduct/169.html"}],"as_of":"2025-08","related_ids":["agibot","wheeled-humanoid-robot","agibot-world","agibot-go-1","7-dof-robot-arm","teleoperation"],"name":"AgiBot Genie G1","alt":"智元 精灵 G1","abbr":"","aliases":["Genie G1","AgiBot G1"],"one_liner":"AgiBot's wheeled dual-arm robot, used mainly for data collection; it's the robot behind AgiBot World.","explanation":"Genie G1 is AgiBot's wheeled dual-arm robot: two 7-DoF arms on top, a height-adjustable waist, a mobile base underneath, an optional gripper or dexterous hand at the end of each arm, and 8 cameras across the body. It is used mainly for teleoperated data collection and model research. The AgiBot World dataset, open-sourced by AgiBot and the Shanghai AI Laboratory in late 2024, was collected using 100 G1 units, and the later GO-1 model was also trained on that data, which is why G1's data format is widely used across the industry. It went on sale in August 2025, priced at RMB 450,000 and aimed at research and education.","example":"Researchers downloading the AgiBot World dataset to fine-tune a VLA model are working with data whose robot embodiment is the Genie G1.","related":["AgiBot","Wheeled Humanoid Robot","AgiBot World","AgiBot GO-1","7-DoF Robot Arm","Teleoperation"]},{"id":"agibot-genie-g2","category":"robot","sec":8,"tier":3,"sources":[{"title":"IT之家：智元精灵 G2 新一代工业级交互式具身作业机器人发布","url":"https://www.ithome.com/0/889/866.htm"},{"title":"中国日报网：智元发布新一代工业级交互式具身作业机器人精灵G2","url":"https://qiye.chinadaily.com.cn/a/202510/16/WS68f0b2b8a310c4deea5ecac8.html"}],"as_of":"2025-10","related_ids":["agibot","agibot-genie-g1","wheeled-humanoid-robot","nvidia-jetson-thor","joint-impedance-control","hot-swappable-battery"],"name":"AgiBot Genie G2","alt":"智元 精灵 G2","abbr":"","aliases":["Genie G2"],"one_liner":"AgiBot's 2025 industrial-grade wheeled dual-arm robot, with force-controlled arms and a dexterous hand.","explanation":"Genie G2 is the wheeled dual-arm robot AgiBot released on October 16, 2025, positioned for industrial work as the successor to Genie G1. It runs on an NVIDIA Jetson Thor as its main controller, and every joint in both arms carries a torque sensor, enabling compliant manipulation through joint impedance control, which lets a joint yield like a spring under force; it can be fitted with a 19-DoF dexterous hand with tactile sensing, has a 5-DoF waist-and-leg assembly on an omnidirectional base, hot-swappable dual batteries, and 360° fisheye cameras plus front and rear lidar for navigation and obstacle avoidance. It is reported to already be used for automotive parts assembly, memory-module insertion, warehouse order fulfillment, and guided tours, and won orders worth hundreds of millions of RMB at launch.","example":"On an auto-parts line, the G2 completes the assembly step of fitting a seatbelt buckle.","related":["AgiBot","AgiBot Genie G1","Wheeled Humanoid Robot","NVIDIA Jetson Thor","Joint Impedance Control","Hot-Swappable Battery"]},{"id":"agibot-lingxi-x1","category":"robot","sec":8,"tier":3,"sources":[{"title":"IT之家：智元机器人宣布灵犀 X1 面向全球开源","url":"https://www.ithome.com/0/804/935.htm"},{"title":"极客公园：智元灵犀X1软硬件全套图纸和代码全公开","url":"https://www.geekpark.net/news/342219"}],"as_of":"2024-10","related_ids":["agibot","small-size-humanoid-robot","open-source-hardware","joint-actuator-module","agibot-lingxi-x2","bipedal-robot"],"name":"AgiBot Lingxi X1","alt":"智元 灵犀 X1","abbr":"","aliases":["Lingxi X1"],"one_liner":"AgiBot's small bipedal humanoid, released with its complete hardware and software fully open-sourced.","explanation":"Lingxi X1 is a small bipedal humanoid robot built by AgiBot's internal X-Lab, led by Zhihui Jun, standing about 1.3 meters tall and weighing about 33 kg, unveiled in August 2024. On October 24, 2024, AgiBot open-sourced its complete hardware and software, over 1.2 GB of materials, to GitHub. Its whole body is assembled from just two in-house joint motor models, the PowerFlow R86 and R52, with cabling routed through hollow joint centers in a modular design, so it can be replicated by buying the parts and building it yourself. For a newcomer, its significance is that it offers a complete humanoid-robot design reference, not just a finished product.","example":"A developer sourced parts from the open-sourced bill of materials and drawings and assembled their own Lingxi X1.","related":["AgiBot","Small-size Humanoid Robot","Open-Source Hardware (OSHW)","Joint Actuator Module","AgiBot Lingxi X2","Bipedal Robot"]},{"id":"agibot-lingxi-x2","category":"robot","sec":8,"tier":3,"sources":[{"title":"澎湃新闻：智元灵犀X2公布售价区间","url":"https://www.thepaper.cn/newsDetail_forward_30860019"},{"title":"IT之家：智元机器人全系产品开售","url":"https://www.ithome.com/0/876/106.htm"}],"as_of":"2025-08","related_ids":["agibot","agibot-lingxi-x1","small-size-humanoid-robot","edu-edition","bipedal-robot","guided-tours-and-reception"],"name":"AgiBot Lingxi X2","alt":"智元 灵犀 X2","abbr":"","aliases":["Lingxi X2"],"one_liner":"AgiBot's 2025 small bipedal humanoid, about 1.3 meters tall, for interaction and research.","explanation":"Lingxi X2 is the small bipedal humanoid robot AgiBot released on March 11, 2025, the successor to Lingxi X1: about 1.3 meters tall, roughly 34 kg, going on sale in May. It emphasizes mobility and human-robot interaction, with demos showing it riding a bicycle and dancing. It shipped in several versions: an Interactive edition, a Pro Explorer edition, and an Ultra Flagship edition in the first batch in May, joined by a Youth edition in August; the Youth edition has 27 degrees of freedom, and the Explorer edition has 31, with an optional dexterous hand or gripper. Reported prices range from roughly RMB 100,000 to RMB 300,000–400,000, with the Youth edition priced at RMB 98,000 in August 2025, making it one of AgiBot's lowest-priced humanoid robots.","example":"Tourist attractions and exhibition halls use the Lingxi X2 Youth edition for greeting and guided-tour performances.","related":["AgiBot","AgiBot Lingxi X1","Small-size Humanoid Robot","EDU Edition","Bipedal Robot","Guided Tours & Reception"]},{"id":"galbot-g1","category":"robot","sec":8,"tier":1,"sources":[{"title":"Galbot G1 官网","url":"http://www.galbot.com/g1/"},{"title":"银河通用的第一个人形机器人 GALBOT（腾讯新闻）","url":"https://news.qq.com/rain/a/20240616A06S3D00"}],"as_of":"2024-06","related_ids":["wheeled-humanoid-robot","graspvla","astrabrain","galbot-s1","mobile-manipulator","synthetic-data"],"name":"Galbot G1","alt":"银河通用 Galbot G1","abbr":"","aliases":[],"one_liner":"Galbot's wheeled, dual-arm humanoid robot with a foldable, height-adjustable leg.","explanation":"Galbot G1 is the wheeled humanoid robot released in June 2024 by the Beijing company Galbot. Instead of two legs, it has a single foldable, telescoping leg on an omnidirectional wheeled base, using a “kneeling” and a “standing” posture to cover a pick-and-place height range of 0 to 2.1 meters off the ground. According to the company's published specs, it stands up to 1.73 meters tall, has two 7-DoF arms (21 degrees of freedom total, excluding the base and end effectors), a 5 kg payload per arm, roughly 8 hours of battery life, and an NVIDIA Orin as its main controller. It targets picking and carrying tasks in supermarkets, pharmacies, and factories; its underlying grasping and navigation models are trained largely on synthetic simulation data.","example":"Galbot G1 picks medicine from shelves against orders inside an unstaffed pharmacy.","related":["Wheeled Humanoid Robot","GraspVLA","AstraBrain","Galbot S1","Mobile Manipulator","Synthetic Data"]},{"id":"galbot-s1","category":"robot","sec":8,"tier":3,"sources":[{"title":"双臂最大50KG负载、零遥操全自主作业，银河通用发布Galbot S1 - NE时代","url":"https://ne-time.cn/web/article/37805"},{"title":"银河通用发布首款工业级重载机器人，Galbot G1年出货破1200台 - 同花顺","url":"https://m.10jqka.com.cn/20260129/c674397841.shtml"}],"as_of":"2026-01","related_ids":["galbot-g1","mobile-manipulator","tote-handling","payload","autonomous-battery-swapping","industrial-robot"],"name":"Galbot S1","alt":"银河通用 Galbot S1","abbr":"","aliases":[],"one_liner":"Galbot's heavy-duty industrial dual-arm mobile robot, rated for 50 kg of payload for factory material handling.","explanation":"Galbot S1 is an industrial-grade heavy-payload robot that Galbot released in January 2026 for material handling on factory production lines. It combines a mobile base with two arms rated for a combined payload of 50 kg, and its arms can work across a height range of 0–2.3 meters, letting it pull tote boxes off shelves at different heights. It supports autonomous battery swapping for round-the-clock operation, and its safety system layers vision and lidar sensing for redundancy. Galbot describes S1 as running on its own in-house material-handling model, operating with “zero teleoperation, fully autonomous” control, and says it is already deployed on real factory lines. S1 represents a category of embodied-AI products aimed at heavy industrial work: rather than a humanoid appearance, it prioritizes payload, runtime, and continuous operation.","example":"On a factory floor, S1 autonomously locates a tote box, lifts a parts container weighing tens of kilograms off a shelf, and carries it to a production station.","related":["Galbot G1","Mobile Manipulator","Tote Handling","Payload","Autonomous Battery Swapping","Industrial Robot"]},{"id":"galbot-et1","category":"robot","sec":8,"tier":3,"sources":[{"title":"银河通用首款双足机器人 Galbot ET1「星仔」亮相 - IT之家","url":"https://www.ithome.com/0/992/247.htm"},{"title":"原生智能体机器人「银河星仔」开启预订，24小时突破200台 - 中新网","url":"https://www.chinanews.com.cn/cj/2026/09-04/10690406.shtml"}],"as_of":"2026-09","related_ids":["galbot-g1","astrabrain","bipedal-robot","motion-tracking","world-robot-conference"],"name":"Galbot ET1","alt":"银河通用 Galbot ET1","abbr":"","aliases":["ET1","Galaxy Star Kid"],"one_liner":"Galbot's first bipedal humanoid robot, nicknamed “Galaxy Star Kid,” built for real-time interaction and dance.","explanation":"Galbot ET1 is Galbot's first bipedal humanoid robot, nicknamed “Galaxy Star Kid” (银河星仔) in Chinese. It made its public debut in August 2026 at the World Robot Conference. Galbot's earlier flagship, the G1, uses a wheeled base with two arms; ET1 switches to a bipedal design and targets tourism guiding, live performance, and commercial demonstrations instead. It runs on Galbot's AstraBrain agent model, so instead of following preset scripts it can understand spoken instructions in real time and generate matching speech and movement. At its debut it performed street dance, a handstand, and an inverted pose, and demonstrated real-time tracking of a live dancer, following along as the dancer improvised. Pre-orders opened on September 3, and Galbot reported more than 200 units reserved within 24 hours.","example":"At a trade show, a human dancer improvises a routine and ET1 recognizes the movements in real time and dances along in sync.","related":["Galbot G1","AstraBrain","Bipedal Robot","Motion Tracking","World Robot Conference"]},{"id":"ubtech-walker-x","category":"robot","sec":8,"tier":3,"sources":[{"title":"UBTech 官网: Walker X","url":"https://www.ubtrobot.com/en/humanoid/products/walker-x"}],"as_of":"2026-09","related_ids":["ubtech-robotics","ubtech-walker-s2","humanoid-robot","bipedal-locomotion","dexterous-hand","guided-tours-and-reception"],"name":"UBTech Walker X","alt":"优必选 Walker X","abbr":"","aliases":[],"one_liner":"UBTech's bipedal service humanoid, built for exhibition guiding and interaction.","explanation":"Walker X is UBTech's commercial bipedal humanoid robot, reportedly released at the 2021 World Artificial Intelligence Conference, marking the Walker line's shift from research prototype to commercial showcase. Published specs: 130 cm tall, 63 kg, with 41 degrees of freedom across the body (6 per leg, 7 per arm, 6 per hand, and 3 in the neck), a top walking speed of 3 km/h, and about 2 hours of runtime. It can climb a 15 cm step and walk up a 20-degree slope, uses U-SLAM for visual navigation, has 7-degree-of-freedom arms with a 6-degree-of-freedom force-controlled humanlike hand, and its face is made of two flexible curved screens that can display a range of expressions. It's used mainly for exhibition guiding, reception, and research demonstrations; UBTech's later factory-oriented product line is the Walker S series.","example":"At a trade show, Walker X autonomously walks up to a booth, picks up an item with its hand to hand to a visitor, and interacts using its screen-based facial expressions.","related":["UBTech Robotics","UBTech Walker S2","Humanoid Robot","Bipedal Locomotion","Dexterous Hand","Guided Tours & Reception"]},{"id":"ubtech-walker-s2","category":"robot","sec":8,"tier":2,"sources":[{"title":"UBTECH Walker S2 官网产品页","url":"https://www.ubtrobot.com/en/humanoid/products/walker-s2"},{"title":"全球首创！优必选工业人形机器人Walker S2实现自主换电（国家科技创新中心）","url":"https://www.ncsti.gov.cn/kjdt/scyq/bjjjjskfq/jkdt/202507/t20250721_210973.html"}],"as_of":"2025-12","related_ids":[null,null,null,null,null,null],"name":"UBTech Walker S2","alt":"优必选 Walker S2","abbr":"","aliases":["Walker S2","Walker S Series","Walker S1"],"one_liner":"UBTech's 2025 industrial humanoid, able to swap its own battery without stopping work.","explanation":"Walker S2 is the full-size industrial humanoid robot UBTech released in July 2025, the newest generation of its Walker S series. Its standout feature is autonomous battery swapping: the robot uses both arms to remove its own depleted battery from its back and install a fully charged one, a process taking about 3 minutes with no shutdown or human help needed, enabling continuous 24/7 operation, since battery life has long been a bottleneck for putting humanoids into factories. Published specs put it at about 1.76 meters tall, with a waist that can rotate roughly ±162° and pitch, and a 15 kg payload. It reportedly won multiple orders worth over RMB 100 million in 2025, used mainly for handling and sorting tasks on automotive production lines.","example":"On a car final-assembly line, the robot carries parts totes, and when its battery runs low, it walks itself to a swap station and changes its own battery with both hands before returning to its post.","related":["UBTech Robotics","Humanoid Robot","Autonomous Battery Swapping","Hot-Swappable Battery","In-Factory Training / Pilot Deployment","UBTech Walker X"]},{"id":"ubtech-uworld-u1-series","category":"robot","sec":8,"tier":3,"sources":[{"title":"UBTech 官网: U1 PRO","url":"https://www.ubtrobot.com/en/humanoid/products/u1-pro"}],"as_of":"2026-09","related_ids":["hyper-realistic-humanoid-robot","ubtech-robotics","humanoid-robot","uncanny-valley","human-robot-interaction","engineered-arts-ameca"],"name":"UBTech UWORLD U1 Series","alt":"优必选 优世界 U1","abbr":"","aliases":["U1 PRO"],"one_liner":"UBTech's hyper-realistic humanoid robot line, with a face and body modeled 1:1 on a real person.","explanation":"UWORLD U1 is a hyper-realistic humanoid robot line from UBTech. UBTech's own website lists it under a separate “Ultra-Bionic Humanoids” category, with the U1 PRO as the model currently offered, and the company says it is working toward mass production of a “1:1 full-scale ultra-bionic humanoid.” Hyper-realistic humanoids are built with faces and skin that closely resemble a real person, focused on expression and conversation for human-robot interaction rather than physical work like carrying or assembly — a different product line from UBTech's factory-oriented Walker S series. UBTech's website does not publish height, degrees-of-freedom, or pricing details for U1; this entry could not verify the Chinese series name “优世界” or its exact release date against the official site, so later official announcements should be treated as authoritative.","example":"","related":["Hyper-Realistic Humanoid Robot","UBTech Robotics","Humanoid Robot","Uncanny Valley","Human-Robot Interaction","Engineered Arts Ameca"]},{"id":"fourier-gr-1","category":"robot","sec":8,"tier":2,"sources":[{"title":"Fourier Intelligence launches production version of GR-1 humanoid robot - The Robot Report","url":"https://www.therobotreport.com/fourier-intelligence-launches-production-version-of-gr-1-humanoid-robot/"},{"title":"GR-1 general-purpose humanoid robot will carry nearly its own weight - New Atlas","url":"https://newatlas.com/robotics/fourier-gr1-humanoid-robot/"}],"as_of":"2024","related_ids":["fourier","full-size-humanoid-robot","fourier-gr-2","fourier-gr-3","joint-actuator-module","idp3"],"name":"Fourier GR-1","alt":"傅利叶 GR-1","abbr":"","aliases":["GR-1","GR1"],"one_liner":"Fourier's 2023 full-size humanoid robot, one of the earlier mass-delivered humanoid platforms in China.","explanation":"GR-1 is the general-purpose humanoid robot Shanghai-based Fourier, formerly Fourier Intelligence, which started out making rehabilitation robots, released in July 2023: about 1.65 meters tall, roughly 55 kg, with about 40 degrees of freedom across the body. Its joints use Fourier's own FSA integrated actuators, a joint actuator module that packages the motor, reducer, drive electronics, and encoder together; the company states a maximum payload of about 50 kg and a walking speed of about 5 km/h. It was one of the earlier humanoid platforms in China to open presales and ship in volume to universities and research institutions, and a number of humanoid manipulation studies, including iDP3, have run experiments on it. Fourier later released successor models, GR-2 and GR-3.","example":"The iDP3 (Improved 3D Diffusion Policy) paper used a Fourier GR-1 for its humanoid bimanual manipulation experiments.","related":["Fourier","Full-size Humanoid Robot","Fourier GR-2","Fourier GR-3","Joint Actuator Module","iDP3"]},{"id":"fourier-gr-2","category":"robot","sec":8,"tier":3,"sources":[{"title":"Fourier launches GR-2 humanoid, software platform - The Robot Report","url":"https://www.therobotreport.com/fourier-launches-gr-2-humanoid-software-platform/"},{"title":"Fourier's new GR-2 robot displays human-like motion and flexibility - Interesting Engineering","url":"https://interestingengineering.com/innovation/fouriers-gr-2-upgraded-humanoid-robot"}],"as_of":"2024-09","related_ids":["fourier","fourier-gr-1","fourier-gr-3","full-size-humanoid-robot","dexterous-hand","joint-actuator-module"],"name":"Fourier GR-2","alt":"傅利叶 GR-2","abbr":"","aliases":["GR-2"],"one_liner":"Fourier's 2024 full-size humanoid robot, with 53 degrees of freedom and 12-DOF dexterous hands.","explanation":"Fourier GR-2 is a full-size humanoid robot that Shanghai-based Fourier released in late September 2024, an upgrade to its earlier GR-1. Published specs: 175 cm tall, 63 kg, with 53 degrees of freedom (independently driven joints) across the body, including 12 in each dexterous hand; a single-arm payload of 3 kg; and a removable battery good for about 2 hours. The joints use Fourier's own FSA actuators, rated above 380 N·m of peak torque. GR-2 is sold mainly to universities and R&D teams for experiments such as teleoperation data collection and training manipulation policies; Fourier also released a companion software development platform alongside it.","example":"A research team teleoperates GR-2's hands via VR to record grasping demonstrations, then uses that data to train an imitation-learning policy.","related":["Fourier","Fourier GR-1","Fourier GR-3","Full-size Humanoid Robot","Dexterous Hand","Joint Actuator Module"]},{"id":"fourier-gr-3","category":"robot","sec":8,"tier":3,"sources":[{"title":"Fourier Intelligence unveils the Care-bot GR-3, its first full-sized companion humanoid robot - TechNode","url":"https://technode.com/2025/08/07/fourier-intelligence-unveils-the-care-bot-gr-3-its-first-full-sized-companion-humanoid-robot/"},{"title":"Fourier (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Fourier_(company)"}],"as_of":"2025-08","related_ids":["fourier","fourier-gr-2","companion-robot","service-robot","tactile-sensor","human-robot-interaction"],"name":"Fourier GR-3","alt":"傅利叶 GR-3","abbr":"","aliases":["GR-3","Care-bot GR-3"],"one_liner":"Fourier's 2025 companion-oriented full-size humanoid robot, built for eldercare and social interaction.","explanation":"Fourier released the “Care-bot” GR-3 in August 2025, its first full-size humanoid built around companionship rather than research use. It stands 165 cm tall, weighs 71 kg, and has 55 degrees of freedom, including 12-DOF dexterous hands and about 3 kg of payload per arm. The shell is covered in a soft material, with 31 pressure sensors across the body so it can sense being touched; it also integrates voice, vision, and tactile interaction modules, and runs on a quick-swappable battery good for about 3 hours per pack. Where the more research-oriented GR-2 targets labs, GR-3 is built for eldercare, companionship, and guide-type roles that require staying close to people — fitting, since Fourier itself started out making rehabilitation robots.","example":"At a senior-care facility, GR-3 turns to respond when an elderly resident taps its shoulder, chatting with them and walking them through a simple rehab exercise.","related":["Fourier","Fourier GR-2","Companion Robot","Service Robot","Tactile Sensor","Human-Robot Interaction"]},{"id":"fourier-n1","category":"robot","sec":8,"tier":3,"sources":[{"title":"Fourier Launches First Open-Source Humanoid Robot, Fourier N1 - AIbase","url":"https://www.aibase.com/news/17071"},{"title":"Fourier N1 open source humanoid - Webull News","url":"https://www.webull.com/news/12624717734265856"}],"as_of":"2025-04","related_ids":["fourier","open-source-hardware","small-size-humanoid-robot","bipedal-robot","rl-based-locomotion-control","fourier-gr-2"],"name":"Fourier N1","alt":"傅利叶 N1","abbr":"","aliases":["N1"],"one_liner":"Fourier's 2025 open-source small humanoid robot, released with full build files and a bill of materials.","explanation":"Fourier N1 is Fourier's first open-source humanoid robot, released on April 11, 2025. It stands 1.31 m tall and weighs 38 kg, with 23 degrees of freedom and no dexterous hands. It can run at up to 3.5 m/s, its joints deliver up to 96 N·m of torque, a single battery lasts more than 2 hours, and the frame is built from aluminum alloy and engineering plastic. Alongside the robot, Fourier published the full bill of materials (BOM), design drawings, assembly instructions, and basic operating software, so developers can build or modify their own unit. N1 is aimed at research such as bipedal walking and reinforcement-learning locomotion control, at a much lower cost of entry than full-size humanoids.","example":"A lab assembles its own N1 from the published BOM and drawings, trains a walking policy in Isaac Lab, and deploys it to the real robot.","related":["Fourier","Open-Source Hardware (OSHW)","Small-size Humanoid Robot","Bipedal Robot","RL-based Locomotion Control","Fourier GR-2"]},{"id":"tiangong","category":"robot","sec":8,"tier":2,"sources":[{"title":"北京人形：发布通用人形机器人母平台「天工」","url":"https://www.x-humanoid.com/news-view-4.html"},{"title":"北京人形：「天工」2时40分42秒自主跑完北京亦庄半马","url":"https://x-humanoid.com/news-view-164.html"}],"as_of":"2025-04","related_ids":["beijing-humanoid-robot-innovation-center","full-size-humanoid-robot","huisi-kaiwu","pelican-vl","xr-1","humanoid-robot-half-marathon"],"name":"Tiangong","alt":"天工","abbr":"","aliases":["Tiangong Ultra"],"one_liner":"A full-size, all-electric humanoid platform from the Beijing Humanoid Robot Innovation Center.","explanation":"Tiangong is the general-purpose humanoid robot “parent platform” released in April 2024 by the Beijing Humanoid Robot Innovation Center (X-Humanoid), fully electric, with its first version standing 163 cm tall and weighing 43 kg, equipped with 3D vision, an IMU, and a six-axis force sensor, and positioned as an open platform for the industry to build on. The later Tiangong Ultra stands 180 cm tall and weighs 52 kg, and in April 2025 won the Beijing Yizhuang Humanoid Robot Half Marathon in 2 hours, 40 minutes, and 42 seconds. The Innovation Center has also released the Huisi Kaiwu platform and the Pelican-VL and XR-1 models alongside it, with Tiangong serving as their main hardware carrier.","example":"Tiangong Ultra completed the 21-kilometer 2025 Beijing Yizhuang half marathon in first place.","related":["Beijing Humanoid Robot Innovation Center","Full-size Humanoid Robot","Huisi Kaiwu (X-Humanoid general embodied AI platform)","Pelican-VL","XR-1","Humanoid Robot Half Marathon (Beijing E-Town)"]},{"id":"qinglong","category":"robot","sec":8,"tier":3,"sources":[{"title":"WAIC 2024：全球首款全尺寸通用人形机器人开源公版机「青龙」发布（机器人大讲堂）","url":"https://www.leaderobot.com/news/4397"},{"title":"OpenLoong 人形机器人全栈开源社区","url":"https://www.openloong.org.cn/cn"}],"as_of":"2024-07","related_ids":["national-and-local-co-built-humanoid-robotics-innovation-cen","tiangong","full-size-humanoid-robot","open-source-hardware","humanoid-robot"],"name":"Qinglong (OpenLoong)","alt":"青龙","abbr":"","aliases":["OpenLoong"],"one_liner":"An open-source full-size humanoid “reference design” released by the Shanghai Humanoid Robot Innovation Center.","explanation":"Qinglong is a full-size bipedal humanoid robot that the National and Local Co-built Humanoid Robotics Innovation Center (also known as the Shanghai Humanoid Robot Innovation Center, operated by Humanoid Robot (Shanghai) Co., Ltd.) unveiled at the World Artificial Intelligence Conference in July 2024, positioned as an open-source “reference platform.” Published specs: 185 cm tall, 43 degrees of freedom across the body, 400 N·m of peak joint torque, and 400 TOPS (trillion operations per second) of onboard compute. Alongside the launch, the center founded the OpenLoong open-source community, publishing the full hardware design and software so universities and companies can build on the same base platform instead of each reinventing the hardware. Qinglong and Tiangong, from the Beijing Humanoid Robot Innovation Center, are frequently mentioned together as China's two leading open-source general-purpose humanoid platforms.","example":"A developer downloads Qinglong's models and control code from the OpenLoong community, tunes motion control in simulation first, and then deploys it to the real robot.","related":["National and Local Co-built Humanoid Robotics Innovation Center (Shanghai)","Tiangong","Full-size Humanoid Robot","Open-Source Hardware (OSHW)","Humanoid Robot"]},{"id":"xpeng-iron","category":"robot","sec":8,"tier":2,"sources":[{"title":"Xpeng unveils next-gen Iron humanoid robot at 2025 AI Day (CnEVPost)","url":"https://cnevpost.com/2025/11/05/xpeng-unveils-next-gen-iron-humanoid-robot/"},{"title":"XPENG Unveils VLA 2.0, Robotaxi, Next-Gen IRON（小鹏官网）","url":"https://www.xpeng.com/news/019a56f54fe99a2a0a8d8a0282e402b7"}],"as_of":"2025-11","related_ids":[null,null,null,null,null,null],"name":"XPeng IRON","alt":"小鹏 IRON","abbr":"","aliases":[],"one_liner":"XPeng Motors' humanoid robot, with a highly human-like appearance in its 2025 generation.","explanation":"IRON is the humanoid robot from Chinese automaker XPeng, with the first generation shown at XPeng's 2024 AI Day and the next generation released in November 2025. The newer version stands about 1.78 meters tall, weighs about 70 kg, and has 82 degrees of freedom across the body, 22 of them in one hand; it has a human-like spine, bio-inspired muscles, and flexible skin, a curved-screen head, an all-solid-state battery, and three of XPeng's own Turing AI chips, with a combined 2,250 TOPS of compute, running the company's second-generation VLA, meaning vision-language-action, model. XPeng aims to begin mass production by the end of 2026, initially deploying units as greeters in its own stores, making it a representative example of automakers entering humanoid robotics.","example":"","related":["XPeng","Humanoid Robot","Automakers Entering Humanoid Robotics","Vision-Language-Action Model","Solid-State Battery","Guided Tours & Reception"]},{"id":"xiaomi-cyberone","category":"robot","sec":8,"tier":3,"sources":[{"title":"CyberOne - ROBOTS: Your Guide to the World of Robotics (IEEE)","url":"https://robotsguide.com/robots/cyberone"}],"as_of":"2022-08","related_ids":["humanoid-robot","full-size-humanoid-robot","bipedal-robot","xiaomi","xiaomi-cyberdog","proof-of-concept"],"name":"Xiaomi CyberOne","alt":"小米 CyberOne","abbr":"","aliases":[],"one_liner":"Xiaomi's 2022 full-size bipedal humanoid, nicknamed “Tieda” in Chinese.","explanation":"CyberOne is the humanoid robot Xiaomi unveiled in August 2022 at founder Lei Jun's annual keynote, nicknamed “Tieda” (铁大, roughly “Big Steel”). It stands 177 cm tall and weighs 52 kg, with 21 degrees of freedom across its body and joint actuators developed in-house by Xiaomi; it can walk on two legs (at about 3.6 km/h) and has facial recognition, emotion recognition, and 3D environmental perception. CyberOne was Xiaomi's first full-size humanoid robot, positioned as a technology showcase and proof of concept rather than a commercial product — it was never sold to the public. IEEE's Robots guide estimated its cost at roughly $70,000–80,000. Together with CyberDog, it formed Xiaomi's early exploration of bio-inspired robotics, and stands as an early example of a major Chinese consumer-tech company entering humanoid robotics.","example":"","related":["Humanoid Robot","Full-size Humanoid Robot","Bipedal Robot","Xiaomi","Xiaomi CyberDog","Proof of Concept"]},{"id":"honor-lightning-humanoid-robot","category":"robot","sec":8,"tier":3,"sources":[{"title":"百余台机器人同跑半马 「闪电」超越人类纪录 - 新华网","url":"https://www.news.cn/sports/20260419/0834250af22d4432ac708322aa8f7123/c.html"},{"title":"荣耀晒「闪电」机器人最新战报 - IT之家","url":"https://www.ithome.com/0/993/297.htm"},{"title":"一年提速近两小时、从遥控到自主、跑赢人类！人形机器人「半马」刷新纪录（每日经济新闻，2026-04-19）","url":"https://www.nbd.com.cn/articles/2026-04-19/4345990.html"},{"title":"Ratified: world records for Kiplimo, Tharp and Wanyonyi（World Athletics，2026-09-03）","url":"https://worldathletics.org/news/press-releases/ratified-world-records-kiplimo-tharp-wanyonyi"}],"as_of":"2026-04","related_ids":["humanoid-robot-half-marathon","bipedal-locomotion","joint-actuator-module","liquid-cooled-joint-actuators","peak-torque","rl-based-locomotion-control"],"name":"Honor Lightning Humanoid Robot","alt":"荣耀「闪电」人形机器人","abbr":"","aliases":["Lightning","Honor Lightning"],"one_liner":"Honor's self-developed bipedal humanoid robot that won the 2026 Beijing humanoid half-marathon in 50:26.","explanation":"“Lightning” is the first large-size humanoid robot built by Honor, the smartphone maker, developed by its unit Shenzhen Honor Smart Technology Development Co., Ltd. It stands 169 cm tall with roughly 95 cm legs and reportedly weighs about 45 kg. It uses Honor's own integrated joint actuator modules rated at 400 N·m peak torque, paired with a liquid-cooling system that keeps the motors from overheating during long stretches of high-power running. On April 19, 2026, running in fully autonomous navigation mode, it won the Beijing E-Town (Yizhuang) Humanoid Robot Half Marathon with a net time of 50 minutes 26 seconds, faster than the human men's half-marathon world record at the time, 57:20, set by Ugandan runner Kiplimo in Lisbon in March 2026. Lightning represents consumer-electronics companies crossing into humanoid robotics, with a focus on high-speed locomotion control.","example":"At the 2026 Beijing E-Town humanoid half-marathon, the winning “Lightning” ran the roughly 21-kilometer course under its own autonomous navigation, without remote control, swapping its battery once, at the 10.6 km mark, along the way.","related":["Humanoid Robot Half Marathon (Beijing E-Town)","Bipedal Locomotion","Joint Actuator Module","Liquid-Cooled Joint Actuators (Active Thermal Management)","Peak Torque","RL-based Locomotion Control"]},{"id":"dobot-atom","category":"robot","sec":8,"tier":3,"sources":[{"title":"越疆正式发布并预售具身智能人形机器人 Dobot Atom - 动点科技","url":"https://cn.technode.com/post/2025-03-19/dobot-atom/"},{"title":"19.9万元起，中国全尺寸人形机器人价格破冰 - 南方财经网","url":"https://www.sfccn.com/2025/3-18/xMMDE0NzNfMjAxMjAxMA.html"}],"as_of":"2025-03","related_ids":["dobot","humanoid-robot","full-size-humanoid-robot","straight-knee-walking","dexterous-hand","collaborative-robot"],"name":"Dobot Atom","alt":"越疆 Atom","abbr":"","aliases":[],"one_liner":"A full-size humanoid from cobot-arm maker Dobot, released in 2025 starting at RMB 199,000.","explanation":"Dobot Atom is the full-size humanoid robot the collaborative-arm manufacturer Dobot released and opened for preorder on March 18, 2025, priced from RMB 199,000. Published specs: 1.53 meters tall, 62 kg, 41 degrees of freedom across the body; its arms are Dobot's own 7-DoF industrial-grade collaborative arms, with a repeatability of ±0.05 mm, fitted with a five-fingered dexterous hand, and 28 degrees of freedom in the upper body take part in end-to-end manipulation. It emphasizes straight-knee walking, which the company says cuts energy use by 42% compared with the bent-knee gait common on other humanoids. Dobot's approach is to carry the precision and manufacturing experience of its industrial cobot arms over to a humanoid platform; launch demos included making breakfast, pouring coffee, retrieving a delivery, and front-desk reception.","example":"In a launch demo, Atom autonomously completes a full breakfast routine: pouring milk, toasting bread, and arranging fruit on a plate.","related":["Dobot","Humanoid Robot","Full-size Humanoid Robot","Straight-Knee Walking","Dexterous Hand","Collaborative Robot"]},{"id":"deep-robotics-dr02","category":"robot","sec":8,"tier":3,"sources":[{"title":"DEEP Robotics Launches World's First All-Weather Industrial Humanoid Robot - TechNode","url":"https://technode.com/2025/10/09/deep-robotics-launches-worlds-first-all-weather-industrial-humanoid-robot/"},{"title":"Deep Robotics unveils DR02, the world's first IP66-rated humanoid robot - KrASIA","url":"https://kr-asia.com/deep-robotics-unveils-dr02-the-worlds-first-ip66-rated-humanoid-robot"}],"as_of":"2025-10","related_ids":["humanoid-robot","full-size-humanoid-robot","deep-robotics","ingress-protection-rating","inspection-robot","deep-robotics-jueying-x30"],"name":"DEEP Robotics DR02","alt":"云深处 DR02","abbr":"","aliases":["DR02","DR-02"],"one_liner":"DEEP Robotics' 2025 full-size humanoid, built around full-body IP66 dust and water protection.","explanation":"DR02 is the full-size humanoid robot Hangzhou-based DEEP Robotics released on October 9, 2025, aimed at industrial and outdoor inspection settings. Published specs: 175 cm tall, reportedly about 65 kg, a normal walking speed of 1.5 m/s and a top speed of 4 m/s, a 20 kg payload, able to climb 20 cm steps and 20° slopes, an operating range of -20°C to 55°C, and 275 TOPS of onboard compute. The company states it is the world's first humanoid robot rated IP66 across its whole body, meaning fully dust-tight and protected against powerful water jets, and it uses a modular quick-release design for easier maintenance. DEEP Robotics started out building quadruped robots, and DR02 carries its accumulated expertise in ruggedness and outdoor reliability from that line over to a humanoid, aimed at “all-weather” industrial use.","example":"DEEP Robotics promotes the DR02 as able to do outdoor security patrols and factory work in rain and dusty conditions.","related":["Humanoid Robot","Full-size Humanoid Robot","DEEP Robotics","Ingress Protection (IP) Rating","Inspection Robot","DEEP Robotics Jueying X30"]},{"id":"booster-robotics-t1","category":"robot","sec":8,"tier":2,"sources":[{"title":"Booster T1 | Made for Developers","url":"https://www.booster.tech/booster-t1/"},{"title":"RoboCup 解决方案 | 加速进化","url":"https://www.booster.tech/zh/robocup/"},{"title":"Booster Robotics Booster T1 Specs & Price | Humanoid.guide","url":"https://humanoid.guide/product/booster-t1/"}],"as_of":"2025-07","related_ids":["booster-robotics","small-size-humanoid-robot","bipedal-robot","robocup","nvidia-jetson-orin","booster-robotics-k1"],"name":"Booster Robotics T1","alt":"加速进化 Booster T1","abbr":"","aliases":["Booster T1","T1"],"one_liner":"Booster Robotics' 1.2-meter small humanoid, aimed at developers and robot soccer.","explanation":"Booster T1 is the small bipedal humanoid robot from Booster Robotics, standing about 1.2 meters tall and weighing about 30 kg, with 23 degrees of freedom in the base configuration and optional grippers or dexterous hands, powered by an NVIDIA Jetson AGX Orin (up to roughly 200 TOPS of compute). The company positions it as “built for developers,” with an open SDK for reinforcement-learning locomotion control and higher-level algorithm research. It is best known for robot soccer: in July 2025, Tsinghua University's Hyperdog team used the T1 to win the humanoid adult-size division at the RoboCup World Cup in Brazil. Its small size, lower damage from falls, and lower price than full-size humanoids make it popular with universities and labs doing locomotion and embodied-AI research.","example":"At the 2025 RoboCup World Cup in Brazil, Tsinghua University's Hyperdog team won the humanoid adult-size division using the Booster T1.","related":["Booster Robotics","Small-size Humanoid Robot","Bipedal Robot","RoboCup","NVIDIA Jetson Orin","Booster Robotics K1"]},{"id":"booster-robotics-k1","category":"robot","sec":8,"tier":3,"sources":[{"title":"Booster Robotics Launches K1 (Humanoids Daily)","url":"https://www.humanoidsdaily.com/news/booster-robotics-launches-k1-robocup-champion-platform"},{"title":"Booster K1 humanoid robot | Generation Robots","url":"https://www.generationrobots.com/en/404324-booster-k1-humanoid-robot.html"}],"as_of":"2025-10","related_ids":["booster-robotics","booster-robotics-t1","small-size-humanoid-robot","robocup","edu-edition","rl-based-locomotion-control"],"name":"Booster Robotics K1","alt":"加速进化 Booster K1","abbr":"","aliases":["Booster K1"],"one_liner":"A roughly 95 cm small humanoid from Beijing's Booster Robotics, for education and robot soccer.","explanation":"Booster K1 is the small humanoid robot Booster Robotics, founded in Beijing in 2023, formally released in October 2025. It stands about 95 cm tall, weighs about 19.5 kg, and has 22 degrees of freedom across the body, 6 per leg and 4 per arm; it reportedly starts at about $4,999, with several versions differing in compute and battery life, and supports Python and ROS 2 for custom development. Before its formal release, it already served as the platform for Germany's HTWK team, which won the RoboCup 2025 KidSize humanoid division. It targets universities, competition teams, and developers, positioned as an entry-level embodied-AI development platform for reinforcement-learning locomotion control and dynamic tasks such as soccer.","example":"The HTWK team used the K1 to compete in and win the RoboCup 2025 humanoid KidSize division.","related":["Booster Robotics","Booster Robotics T1","Small-size Humanoid Robot","RoboCup","EDU Edition","RL-based Locomotion Control"]},{"id":"noetix-n2","category":"robot","sec":8,"tier":3,"sources":[{"title":"3.99 万元起，全球首个连续空翻人形机器人 NOETIX 松延动力 N2 发布（IT之家）","url":"https://www.ithome.com/0/837/928.htm"},{"title":"马拉松亚军人形机器人「松延动力 N2」被拍卖，以 5.7 万元成交（IT之家）","url":"https://www.ithome.com/0/847/782.htm"}],"as_of":"2025-04","related_ids":["noetix-robotics","small-size-humanoid-robot","humanoid-robot-half-marathon","noetix-bumi","rl-based-locomotion-control","degrees-of-freedom"],"name":"Noetix N2","alt":"松延动力 N2","abbr":"","aliases":["N2"],"one_liner":"Noetix's 1.2-meter humanoid robot, released in 2025, capable of consecutive backflips and priced from RMB 39,900.","explanation":"N2 is a small humanoid robot that Noetix released on March 14, 2025, priced from RMB 39,900 and sold through the e-commerce platform JD.com. It stands 1.2 m tall and weighs 30 kg, with only 18 degrees of freedom across its body (5 per leg, 4 per arm) — a deliberately simplified joint count traded for a lighter build and better fall resistance. It can take long strides, run (tested at up to about 3.5 m/s), jump on one or two feet, and dance, and Noetix marketed it at launch as the world's first humanoid robot capable of consecutive backflips. At the Beijing E-Town Humanoid Robot Half Marathon in April 2025, an N2 placed second; that particular unit was later auctioned off for RMB 57,000. N2 is positioned on low price and athletic performance, and is often used to showcase locomotion-control capability.","example":"At the 2025 Beijing E-Town humanoid half-marathon, an N2 finishes the full course and takes second place.","related":["Noetix Robotics","Small-size Humanoid Robot","Humanoid Robot Half Marathon (Beijing E-Town)","Noetix Bumi","RL-based Locomotion Control","Degrees of Freedom (DoF)"]},{"id":"noetix-bumi","category":"robot","sec":8,"tier":3,"sources":[{"title":"全球首款万元以内高性能人形机器人：松延动力 Bumi 小布米发布，9998 元能跑能跳舞（IT之家）","url":"https://www.ithome.com/0/891/493.htm"},{"title":"松延动力获 1000 台小布米 Bumi 订单（IT之家）","url":"https://www.ithome.com/0/904/883.htm"}],"as_of":"2025-12","related_ids":["noetix-robotics","small-size-humanoid-robot","10-000-yuan-class-humanoid-robot","consumer-grade-robot","noetix-n2","research-and-education-market"],"name":"Noetix Bumi","alt":"松延动力 小布米","abbr":"","aliases":["Bumi"],"one_liner":"Noetix's RMB 9,998 small humanoid robot from 2025, aimed at homes and education.","explanation":"Bumi (小布米) is a small bipedal humanoid robot from Beijing-based Noetix Robotics, released in October 2025 with preorders opening on October 23 at a price of RMB 9,998; the company calls it the world's first high-performance humanoid robot priced under RMB 10,000. It stands about 94 cm tall, weighs about 12 kg, and has at least 21 degrees of freedom, letting it walk, run, and dance — and it's small enough for a person to pick up and carry. It supports block-based visual programming, where a child can drag and drop modules to choreograph movements, plus voice interaction, and it's aimed mainly at home companionship, children's coding education, and as an entry point into robotics research. It reportedly received 1,000 orders shortly after launch. By bringing a humanoid robot's price under RMB 10,000, Bumi has become a landmark product in discussions of “consumer-grade humanoid robots.”","example":"A child drags motion blocks together on a tablet to choreograph a dance routine, then sends it to Bumi to perform.","related":["Noetix Robotics","Small-size Humanoid Robot","10,000-Yuan-Class Humanoid Robot","Consumer-Grade Robot","Noetix N2","Research & Education Market"]},{"id":"engineai-se01","category":"robot","sec":8,"tier":3,"sources":[{"title":"Engine AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Engine_AI"},{"title":"EngineAI Robotics SE01 Specs - Humanoid.guide","url":"https://humanoid.guide/product/se01/"}],"as_of":"2024-10","related_ids":[null,null,null,null,null,null],"name":"EngineAI SE01","alt":"众擎 SE01","abbr":"","aliases":["SE01"],"one_liner":"EngineAI's 2024 full-size, 1.7-meter humanoid, known for a walking gait close to a human's.","explanation":"SE01 is the full-size general-purpose humanoid robot EngineAI released in October 2024: about 170 cm tall, roughly 55 kg, 32 degrees of freedom, a walking speed of about 2 m/s, using in-house harmonic, planetary, and ball-screw joint actuator modules, with a reported peak knee-joint torque of about 186 N·m. Its selling point at launch was using an end-to-end neural network to control its gait, walking with long strides and straight knees, looking noticeably more human than the small-stepped, bent-knee walk common on most humanoids at the time, and the related videos drew considerable online discussion. The company says it targets industrial production lines and home companionship, though in practice it is still used mainly for demonstrations and research partnerships.","example":"SE01's release video shows it walking down a street with long strides and straight knees, widely compared against the small-stepped, bent-knee gait of earlier humanoid robots.","related":["EngineAI","Full-size Humanoid Robot","Straight-Knee Walking","RL-based Locomotion Control","EngineAI PM01","EngineAI T800"]},{"id":"engineai-pm01","category":"robot","sec":8,"tier":3,"sources":[{"title":"Engine AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Engine_AI"},{"title":"EngineAI releases PM01 humanoid robot - The Robot Report","url":"https://www.therobotreport.com/engineai-releases-pm01-humanoid-robot-for-commercial-educational-use/"}],"as_of":"2024-12","related_ids":[null,null,null,null,null,null],"name":"EngineAI PM01","alt":"众擎 PM01","abbr":"","aliases":["PM01"],"one_liner":"A 1.38-meter lightweight bipedal humanoid EngineAI released in late 2024, for research and education.","explanation":"PM01 is the small bipedal humanoid robot Shenzhen-based EngineAI released in December 2024: about 1.38 meters tall, roughly 40 kg, with 24 degrees of freedom across the body, a waist capable of a large rotation range, reported at 320°, and a walking speed of about 2 m/s. Its commercial and education editions are reported to be priced around RMB 88,000, and an open-source version is also available for custom development. It drew attention for being one of the earlier humanoid robots shown completing a front flip on video. Positioning-wise, PM01 is similar to Unitree's G1: relatively affordable and well-suited to reinforcement-learning locomotion control and sim-to-real experiments at universities and among developers, rather than a product meant to go straight to factory work.","example":"A university team trains a walking policy for the PM01 with reinforcement learning in Isaac Lab, then deploys it to the real robot for a sim-to-real transfer experiment.","related":["EngineAI","Small-size Humanoid Robot","Bipedal Robot","RL-based Locomotion Control","Unitree G1","EngineAI SE01"]},{"id":"engineai-t800","category":"robot","sec":8,"tier":3,"sources":[{"title":"众擎T800人形机器人一脚把自家CEO踹翻在地 - 新浪科技","url":"https://finance.sina.com.cn/tech/roll/2025-12-07/doc-infzxvsm5162157.shtml"},{"title":"Engine AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Engine_AI"}],"as_of":"2026-01","related_ids":[null,null,null,null,null,null],"name":"EngineAI T800","alt":"众擎 T800","abbr":"","aliases":["T800"],"one_liner":"EngineAI's late-2025, 1.73-meter high-dynamics humanoid, known for high torque and combat-style demos.","explanation":"T800 is the full-size general-purpose humanoid robot EngineAI released in early December 2025: 1.73 meters tall, weighing about 75 kg, with a peak coordinated joint torque of 450 N·m, a magnesium-aluminum alloy skeleton, and a solid-state battery, with the company stating 4 to 5 hours of battery life. In China it ships in four configurations, Basic, Open-Source, Pro, and Max, with the Basic edition starting at RMB 180,000. After release, its promotional video's movements looked so fluid that some viewers suspected CGI; the company responded with a video of T800 kicking down CEO Zhao Tongyang, which drew wide attention. It represents a trend among Chinese humanoid makers of using highly dynamic moves, such as kicks and combat sequences, to demonstrate raw hardware power.","example":"In a December 2025 release video from EngineAI, CEO Zhao Tongyang, wearing protective gear, spars with the T800 and is knocked down by a kick; he later said that without the gear he would “definitely have broken a bone.”","related":["EngineAI","Full-size Humanoid Robot","Solid-State Battery","Demo (Demonstration Video)","CMG Mecha Fighting Series","EngineAI SE01"]},{"id":"limx-dynamics-cl-1","category":"robot","sec":8,"tier":3,"sources":[{"title":"逐际动力首次公开人形机器人 CL-1 动态测试（逐际动力新闻中心）","url":"https://www.limxdynamics.com/zh/news/BK000005"},{"title":"逐际动力 CL-1 实现能力升级，新一代全尺寸人形机器人 CL-2 公开亮相（IT之家）","url":"https://www.ithome.com/0/790/893.htm"}],"as_of":"2024-08","related_ids":["limx-dynamics","humanoid-robot","perceptive-locomotion","limx-dynamics-oli","full-size-humanoid-robot","loco-manipulation"],"name":"LimX Dynamics CL-1","alt":"逐际动力 CL-1","abbr":"","aliases":["CL-1"],"one_liner":"LimX Dynamics' full-size humanoid prototype, unveiled in late 2023, known for climbing stairs by sensing terrain in real time.","explanation":"CL-1 is a full-size humanoid robot from Shenzhen-based LimX Dynamics, which first released video of it in dynamic testing on December 28, 2023. The company says CL-1 closes the loop from real-time terrain perception through gait planning to whole-body control, letting it climb stairs step by step based on what its cameras see, descend slopes of about 15 degrees, and walk both indoors and outdoors — reportedly China's first humanoid robot to climb stairs dynamically using real-time terrain perception. At the 2024 World Robot Conference, CL-1 also demonstrated mobile-manipulation tasks such as weighted squats between shelves and carrying goods, alongside the reveal of the next-generation CL-2. LimX Dynamics has not published CL-1's full height or degrees-of-freedom specs. CL-1 was an early prototype in the company's humanoid line, which later evolved into the research-oriented full-size humanoid Oli, and stands as an early example of “perceptive locomotion” — adjusting footholds while reading the terrain — being deployed on a humanoid platform.","example":"In the video LimX Dynamics released in late 2023, CL-1 senses each step in real time and climbs a staircase one step at a time, then walks down a roughly 15-degree slope.","related":["LimX Dynamics","Humanoid Robot","Perceptive Locomotion","LimX Dynamics Oli","Full-size Humanoid Robot","Loco-manipulation"]},{"id":"limx-dynamics-oli","category":"robot","sec":8,"tier":3,"sources":[{"title":"LimX Dynamics Launches Full-Size Humanoid Robot LimX Oli（LimX Newsroom）","url":"https://www.limxdynamics.com/en/news/BK000043"},{"title":"LimX Dynamics launches humanoid robot LimX Oli starting at $21,800（TechNode）","url":"https://technode.com/2025/07/31/limx-dynamics-launches-humanoid-robot-limx-oli-starting-at-21800/"}],"as_of":"2025-08","related_ids":["limx-dynamics","full-size-humanoid-robot","limx-dynamics-cl-1","research-and-education-market","whole-body-control","dexterous-hand"],"name":"LimX Dynamics Oli","alt":"逐际动力 Oli","abbr":"","aliases":["LimX Oli"],"one_liner":"LimX Dynamics' 165 cm full-size humanoid, released in 2025 as a research platform for embodied AI.","explanation":"LimX Oli is a full-size general-purpose humanoid robot that LimX Dynamics released on July 30, 2025. It stands 165 cm tall with 31 actively driven degrees of freedom excluding end effectors (7 per arm, 6 per leg, 3 in the waist, 2 in the neck), comes in Lite, EDU, and Super editions starting at RMB 158,000, and had its public debut at the 2025 World Robot Conference. It's built around hardware and software modularity: the end effector can be swapped between a two-finger gripper and a five-finger dexterous hand, third-party sensors such as microphones, cameras, tactile sensors, IMUs, and lidar can be added, and it ships with an open SDK for both joint-level and task-level control, plus over-the-air updates for its motion library and controller. It's aimed at universities, labs, and systems integrators doing embodied-AI research, for validating algorithms such as whole-body control, loco-manipulation, and VLA models.","example":"A lab buys the EDU edition of Oli, fits it with the five-finger dexterous hand, and uses the open SDK to deploy its own trained whole-body control policy for a carrying task.","related":["LimX Dynamics","Full-size Humanoid Robot","LimX Dynamics CL-1","Research & Education Market","Whole-Body Control","Dexterous Hand"]},{"id":"limx-dynamics-tron-1","category":"robot","sec":8,"tier":3,"sources":[{"title":"LimX Dynamics Launches Multi-Modal Biped Robot TRON 1（LimX Newsroom）","url":"https://www.limxdynamics.com/en/news/BK000040"},{"title":"World's first multi-modal biped robot could soon be yours（New Atlas）","url":"https://newatlas.com/robotics/limx-tron-1-biped-robot/"}],"as_of":"2024-10","related_ids":["limx-dynamics","bipedal-robot","wheel-legged-robot","point-foot-vs-flat-foot","rl-based-locomotion-control","limx-dynamics-tron-2"],"name":"LimX Dynamics TRON 1","alt":"逐际动力 TRON 1","abbr":"","aliases":["TRON1"],"one_liner":"LimX Dynamics' 2024 bipedal robot with swappable feet, positioned as a research platform for locomotion control.","explanation":"TRON 1 is a multi-form bipedal robot that LimX Dynamics released on October 17, 2024, built on top of P1, its earlier point-foot bipedal prototype, and aimed at research customers. Its defining feature is a swappable foot: point-foot mode (the simplest leg form and the easiest to control), flat-foot mode (standing and walking like a human), and wheeled-foot mode (a wheel mounted at the end of the leg, letting it roll or step over obstacles) — LimX Dynamics says the hardware auto-detects the foot type and the software adapts automatically. Early-bird pricing started at $15,000, through an offer that ran through the end of 2024. It has no arms, only legs, and is positioned as an entry-level platform for humanoid locomotion control and a testbed for embodied-AI research, suited to university work on reinforcement-learning locomotion control and sim-to-real transfer. Its successor, TRON 2, can switch between dual-arm, bipedal, and wheeled-leg configurations.","example":"A research group trains a point-foot walking policy for TRON 1 with reinforcement learning in simulation, then deploys it to the real robot to test sim-to-real transfer.","related":["LimX Dynamics","Bipedal Robot","Wheel-legged Robot","Point Foot vs. Flat Foot","RL-based Locomotion Control","LimX Dynamics TRON 2"]},{"id":"limx-dynamics-tron-2","category":"robot","sec":8,"tier":3,"sources":[{"title":"逐际动力 TRON 2 具身机器人发布：可变化三种形态，4.98 万起（IT之家）","url":"https://www.ithome.com/0/906/067.htm"},{"title":"TRON 2 - 多形态具身机器人 - 参数（逐际动力官网）","url":"https://limxdynamics.com/zh/products/tron2/spec"}],"as_of":"2025-12","related_ids":["limx-dynamics","limx-dynamics-tron-1","wheel-legged-robot","mobile-manipulation","vision-language-action-model","research-and-education-market"],"name":"LimX Dynamics TRON 2","alt":"逐际动力 TRON 2","abbr":"","aliases":["TRON2"],"one_liner":"LimX Dynamics' late-2025 modular robot that switches between dual-arm, bipedal, and wheeled-leg configurations.","explanation":"TRON 2 is a multi-form embodied robot that LimX Dynamics released on December 18, 2025, priced from RMB 49,800. It uses a fully modular whole-body architecture: the same base unit can quickly switch among three core configurations — dual-arm, bipedal, and dual wheeled-leg — and the company says it also supports reconfiguring into humanoid and quadruped forms. Published specs include a 7-DOF humanlike single arm with 70 cm of reach, a 10 kg combined payload for the arms, and a 60 kg maximum load on the body; the bipedal form can sense obstacles and climb stairs, with about 4 hours of runtime and a 30 kg maximum payload; and the dual wheeled-leg form carries 30 kg and supports automatic return-to-charge. It ships with a fully open API and standardized hardware interfaces, positioned as a research platform for work on VLA models, mobile manipulation, and whole-body control — an upgrade over TRON 1.","example":"The same TRON 2 unit is first assembled in the dual-arm configuration to collect tabletop grasping data for training a VLA model, then reconfigured into the dual wheeled-leg form for mobile-manipulation experiments.","related":["LimX Dynamics","LimX Dynamics TRON 1","Wheel-legged Robot","Mobile Manipulation","Vision-Language-Action Model","Research & Education Market"]},{"id":"robotera-star1","category":"robot","sec":8,"tier":3,"sources":[{"title":"IT之家：星动纪元发布首款产品级人形机器人 STAR1","url":"https://www.ithome.com/0/790/117.htm"},{"title":"网易：星动纪元发布人形机器人STAR1","url":"https://www.163.com/dy/article/JA4IJOQD0511B8LM.html"}],"as_of":"2024-08","related_ids":[null,"full-size-humanoid-robot","era-42","robotera-xhand1","robotera-l7","rl-based-locomotion-control"],"name":"RobotEra STAR1","alt":"星动纪元 STAR1","abbr":"","aliases":["Star1"],"one_liner":"RobotEra's 2024 full-size bipedal humanoid, built around heavy payload and high running speed.","explanation":"STAR1 is the first “product-grade” humanoid robot from RobotEra (星动纪元), a humanoid robotics company incubated at Tsinghua University's Institute for Interdisciplinary Information Sciences, released on August 19, 2024. Published specs: 55 actively driven degrees of freedom across the body (12 in the legs, 14 in the arms, 3 in the waist, 2 in the neck, plus two 12-DOF dexterous hands), up to 400 N·m of joint torque, a top running speed above 6 m/s, and a maximum payload of 160 kg. It's designed with hand and camera placement close to human proportions, so human motion data can be reused more directly for imitation learning, and its locomotion control is trained with reinforcement learning. STAR1 was RobotEra's early flagship humanoid platform; the company later released the end-to-end model ERA-42 and the newer full-size humanoid L7.","example":"At launch, STAR1 was shown running outdoors and walking under heavy load, to demonstrate its joint torque and endurance.","related":["RobotEra","Full-size Humanoid Robot","ERA-42","RobotEra XHAND1","RobotEra L7","RL-based Locomotion Control"]},{"id":"robotera-l7","category":"robot","sec":8,"tier":3,"sources":[{"title":"跳街舞、打螺丝：星动纪元「星动 L7」发布（IT之家）","url":"https://www.ithome.com/0/869/844.htm"},{"title":"星动L7（星动纪元官网）","url":"https://www.robotera.com/robot/l7.html"}],"as_of":"2025-07","related_ids":["robotera","era-42","robotera-star1","robotera-xhand1","full-size-humanoid-robot","humanoid-robot"],"name":"RobotEra L7","alt":"星动纪元 L7","abbr":"","aliases":["STAR L7"],"one_liner":"RobotEra's 2025 full-size bipedal humanoid, with 55 degrees of freedom across its body.","explanation":"L7 is a full-size bipedal humanoid robot that RobotEra — an embodied-AI company with roots at Tsinghua University — released on July 22, 2025. Published specs: 171 cm tall, 65 kg (excluding the dexterous hands), and 55 degrees of freedom across the body: 2 in the neck, 7 in each arm, 3 in the waist, 6 in each leg, and 12 in each dexterous hand. Its self-developed joint actuator modules deliver 400 N·m of peak torque, its arms together can carry 20 kg, and RobotEra says it can run at up to 4 m/s. L7 is driven by the company's own end-to-end VLA model, ERA-42, and at launch it demonstrated both highly dynamic moves like street dance and fine manipulation such as screwing in bolts and sorting parts — positioned as a general-purpose humanoid that can both move athletically and do real work.","example":"At the launch event, L7 both performs a street-dance routine and, at a workstation, uses its dexterous hands to screw in bolts and sort parts.","related":["RobotEra","ERA-42","RobotEra STAR1","RobotEra XHAND1","Full-size Humanoid Robot","Humanoid Robot"]},{"id":"leju-kuavo","category":"robot","sec":8,"tier":3,"sources":[{"title":"乐聚人形机器人「夸父」发布：搭载开源鸿蒙，多地形行走还能跳 - IT之家","url":"https://www.ithome.com/0/737/085.htm"},{"title":"乐聚 KUAVO（夸父）人形机器人科研版发布，支持开箱即用 - IT之家","url":"https://www.ithome.com/0/800/849.htm"}],"as_of":"2025","related_ids":["leju-robotics","humanoid-robot","full-size-humanoid-robot","huawei","research-and-education-market"],"name":"Leju Kuavo","alt":"乐聚 夸父","abbr":"","aliases":["Kuavo","KUAVO"],"one_liner":"Leju Robotics' full-size bipedal humanoid, running an open-source HarmonyOS build.","explanation":"Kuavo (KUAVO) is a full-size bipedal humanoid robot from Shenzhen-based Leju Robotics, released in December 2023; the company describes it as China's first humanoid to run an open-source build of HarmonyOS and to be capable of jumping and walking across multiple types of terrain. A research edition aimed at being ready to use out of the box followed in October 2024. Published specs for the fourth generation and 4 Pro models: up to about 1.66 m tall, about 45 kg, roughly 30 degrees of freedom across the body, a stereo depth camera with optional lidar, a top walking speed of about 4.6 km/h, and the ability to jump continuously over 20 cm. It has reportedly been connected to Huawei's Pangu embodied foundation model for multi-step task planning, and Leju has also worked with China Mobile and Huawei on a 5G-Advanced humanoid robot.","example":"A university lab buys the Kuavo research edition and uses it to develop and test bipedal walking and grasping algorithms.","related":["Leju Robotics","Humanoid Robot","Full-size Humanoid Robot","Huawei","Research & Education Market"]},{"id":"kepler-forerunner-k2","category":"robot","sec":8,"tier":3,"sources":[{"title":"Kepler debuts fifth-gen K2 humanoid robot - Interesting Engineering","url":"https://interestingengineering.com/innovation/kepler-debuts-k2-humanoid-robot"},{"title":"开普勒K2大黄蜂在WAIC 2025完成8小时续航直播挑战_新浪财经","url":"https://cj.sina.com.cn/articles/view/7207652843/1ad9c0deb02001iwd2?froms=ggmp"},{"title":"商业化落地提速，开普勒K2大黄蜂已开启量产预售","url":"https://www.lvyouxfnet.com/117310.html"}],"as_of":"2025-09","related_ids":["humanoid-robot","full-size-humanoid-robot","kepler-robotics","planetary-roller-screw","linear-actuator","battery-runtime"],"name":"Kepler Forerunner K2","alt":"开普勒 先行者 K2","abbr":"","aliases":["K2 Bumblebee","Bumblebee"],"one_liner":"Kepler Robotics' full-size bipedal humanoid, built for industrial use with a hybrid actuator design and long runtime.","explanation":"Forerunner K2 is a humanoid robot from Shanghai-based Kepler Exploration Robot (Kepler Robotics), which launched the commercial K2 “Bumblebee” version in 2025. Published specs: about 175 cm tall, about 75 kg, 52 degrees of freedom across the body, tendon-driven hands with 11 degrees of freedom each, and a roughly 2.33 kWh battery. Its distinguishing feature is a hybrid joint design: planetary-roller-screw linear actuators in the legs and similar joints, with rotary actuators elsewhere. At WAIC (the World Artificial Intelligence Conference) 2025, it ran an 8-hour continuous livestream demo to show off its runtime. Reported production pre-order tiers include a basic bipedal version, a bipedal developer version, and a wheeled developer version, starting at RMB 248,000 (roughly $30,000 overseas), aimed at industrial tasks such as logistics sorting and factory material handling.","example":"At the WAIC 2025 booth, K2 Bumblebee runs continuous sorting and grasping demos from 9 a.m. to 5 p.m. to prove out its 8-hour runtime.","related":["Humanoid Robot","Full-size Humanoid Robot","Kepler Robotics","Planetary Roller Screw","Linear Actuator (Electric Cylinder)","Battery Runtime"]},{"id":"magiclab-magicbot","category":"robot","sec":8,"tier":3,"sources":[{"title":"魔法原子推出高动态双足人形机器人 MagicBot Z1，最高 50 自由度（IT之家）","url":"https://www.ithome.com/0/866/630.htm"},{"title":"魔法原子人形机器人MagicBot工厂训练实拍 展现多机协作能力（腾讯新闻）","url":"https://news.qq.com/rain/a/20241202A08Q4700"}],"as_of":"2026-07","related_ids":["magiclab","humanoid-robot","full-size-humanoid-robot","factory-pilot-deployment","dexterous-hand","small-size-humanoid-robot"],"name":"MagicLab MagicBot","alt":"魔法原子 MagicBot","abbr":"","aliases":["MagicBot","MagicBot Gen1"],"one_liner":"A humanoid robot line from MagicLab, a Dreame-incubated startup, aimed at factory and commercial deployment.","explanation":"MagicBot is the humanoid robot line from MagicLab (魔法原子), a company founded in January 2024 and incubated by the appliance maker Dreame Technology (追觅). The first full-size MagicBot (Gen1) stands about 174 cm tall with 42 degrees of freedom and a proprietary dexterous hand; its arms can each lift about 20 kg, the whole robot can carry about 40 kg, and it runs for about 5 hours per charge. In December 2024, MagicLab released videos of multiple MagicBot units doing supervised, small-scale collaborative work on a factory line — quality inspection, material handling, parts placement, and barcode scanning. In 2025, the company released a small bipedal version, MagicBot Z1, about 140 cm tall and 40 kg, with a base of 24 degrees of freedom expandable to about 50. It reportedly unveiled a flagship full-size humanoid, MagicBot X1, at the 2026 World Artificial Intelligence Conference. MagicBot represents a domestic approach to humanoids that leans on a parent company's manufacturing base and starts with factory deployment.","example":"In videos MagicLab released, several MagicBot units divide up work on a factory line, handling incoming-parts inspection, material transport, and barcode scanning.","related":["MagicLab","Humanoid Robot","Full-size Humanoid Robot","Factory Pilot Deployment","Dexterous Hand","Small-size Humanoid Robot"]},{"id":"paxini-tora-one","category":"robot","sec":8,"tier":3,"sources":[{"title":"帕西尼发布第二代多维触觉人形机器人 TORA-ONE（IT之家）","url":"https://www.ithome.com/0/790/616.htm"}],"as_of":"2024-08","related_ids":["paxini-tech","tactile-sensor","paxini-px-6ax","dexterous-hand","visuo-tactile-fusion","humanoid-robot"],"name":"PaXini TORA-ONE","alt":"帕西尼 TORA-ONE","abbr":"","aliases":["TORA-ONE"],"one_liner":"A humanoid robot from tactile-sensor maker PaXini, built to showcase dense, multi-axis touch sensing.","explanation":"TORA-ONE is a humanoid robot from PaXini Tech, a tactile-sensor company; its second generation was unveiled at the World Robot Conference in August 2024. According to launch reports, the robot body has 47 degrees of freedom, paired with a 26-DOF biomimetic dexterous hand, and its height adjusts from 1.46 to 1.86 m. Its standout feature is that its two hands integrate nearly 2,000 of PaXini's own ITPU multi-axis tactile sensing units, which measure pressure and fine deformation across the contact surface, with a claimed measurement precision of 0.01 N. This design targets a common weakness of vision-only robots — they can see but can't feel — for tasks like twisting open a bottle cap or squeezing something soft, where the robot needs to know how much force it's applying and whether it's slipping. PaXini also uses this tactile data to train its own visuo-tactile multimodal models.","example":"In demos, TORA-ONE uses its tactile-equipped fingers to grip fragile or soft objects, adjusting grip force in real time based on tactile feedback.","related":["PaXini Tech","Tactile Sensor","PaXini PX-6AX","Dexterous Hand","Visuo-Tactile Fusion","Humanoid Robot"]},{"id":"astribot-s1","category":"robot","sec":8,"tier":3,"sources":[{"title":"Humanoid homebot tackles impressive array of household chores (New Atlas)","url":"https://newatlas.com/robotics/astribot-s1-humanoid-launch/"},{"title":"Astribot S1 Specs | Humanoid.guide","url":"https://humanoid.guide/product/astribot-s1/"}],"as_of":"2026-09","related_ids":["wheeled-humanoid-robot","astribot","tendon-driven-actuation","bimanual-manipulation","imitation-learning","household-tasks"],"name":"Astribot S1","alt":"星尘智能 Astribot S1","abbr":"","aliases":[],"one_liner":"Astribot's wheeled dual-arm humanoid, known for tendon-driven arms and fast, dexterous household demos.","explanation":"Astribot S1 was developed by Shenzhen-based Astribot, with a demo video released in April 2024 and a public unveiling that same year at the World Robot Conference in Beijing. Its form is an omnidirectional wheeled base topped by a humanoid upper body: two 7-DoF arms plus a torso and head, with about twenty-some degrees of freedom in total. Its distinguishing feature is tendon-driven transmission, giving the arms speed and lightness; reported figures put its end-effector top speed at about 10 m/s and its per-arm payload at roughly 5–10 kg, though sources vary. Official videos showed folding laundry, pouring wine, tossing a pan, and putting things away at natural human speed, with most of these motions coming from teleoperated data collected and then reproduced with imitation learning. It represents a “legless, upper-body-manipulation-first” take on the wheeled humanoid approach.","example":"In S1's release video, it folds laundry and pulls a tablecloth out from under a cup without spilling it, all at natural human speed.","related":["Wheeled Humanoid Robot","Astribot","Tendon-Driven Actuation","Bimanual Manipulation","Imitation Learning","Household Tasks"]},{"id":"galaxea-r1","category":"robot","sec":8,"tier":3,"sources":[{"title":"19.9 万元起，星海图 R1 系列仿人形通用机器人发布 - IT之家","url":"https://www.ithome.com/0/821/803.htm"},{"title":"R1 - 星海图官网","url":"https://galaxea-ai.com/cn/products/R1"}],"as_of":"2025-01","related_ids":["galaxea-ai","galaxea-g0-dual-system-vla","galaxea-open-world-dataset","wheeled-humanoid-robot","mobile-manipulator","mobile-manipulation"],"name":"Galaxea R1","alt":"星海图 R1","abbr":"","aliases":["R1 Pro","R1 Lite","Galaxea R1 Pro","Galaxea R1 Lite"],"one_liner":"Galaxea AI's wheeled dual-arm humanoid-style robot line, sold as the R1 Pro, R1, and R1 Lite.","explanation":"Galaxea R1 is a line of wheeled, humanoid-style robots from Galaxea AI (星海图), launched in January 2025 starting at RMB 199,000, with three variants: R1 Pro, R1, and R1 Lite. All three share the same layout — a wheeled base, a movable torso, and two arms — letting the hands reach down to the floor and up to about 2 m (1.7 m for the Lite). The R1 Pro has 26 degrees of freedom and dual 7-axis force-controlled arms with a combined payload of 10 kg and a swappable end effector for different dexterous hands; the R1 has 24 degrees of freedom and dual 6-axis arms; the R1 Lite has 23 degrees of freedom with 6-axis arms. All ship with an NVIDIA Jetson AGX Orin 32GB as standard. The line is commonly used as a data-collection and VLA-model deployment platform for mobile-manipulation research.","example":"A team teleoperates an R1 Lite to record “open the fridge, take out a drink, set it on the table” demonstrations in a real home, then trains a VLA model on the data.","related":["Galaxea AI","Galaxea G0 Dual-System VLA","Galaxea Open-World Dataset","Wheeled Humanoid Robot","Mobile Manipulator","Mobile Manipulation"]},{"id":"ai2-robotics-alphabot","category":"robot","sec":8,"tier":3,"sources":[{"title":"新浪财经：智平方发布全新一代智能机器人AlphaBot 2","url":"https://finance.sina.com.cn/jjxw/2025-04-17/doc-inetnwin2047310.shtml"},{"title":"深圳新闻网：智平方发布全新一代通用智能机器人AlphaBot 2","url":"https://www.sznews.com/news/content/mb/2025-04/18/content_31541723.htm"}],"as_of":"2025-04","related_ids":["ai2-robotics","govla","wheeled-humanoid-robot","dual-system-architecture","whole-body-control","mobile-manipulation"],"name":"AI² Robotics AlphaBot","alt":"智平方 AlphaBot","abbr":"","aliases":["AlphaBot","AlphaBot 2"],"one_liner":"Shenzhen AI² Robotics' wheeled dual-arm humanoid, running its own whole-body VLA model.","explanation":"AlphaBot is the wheeled humanoid robot series from Shenzhen-based AI² Robotics. AlphaBot 2, released in April 2025, has 34-plus degrees of freedom across the body, a 700 mm reach per arm, a waist and legs that can rise and lower, a vertical working range of 0 to 240 cm, and can run continuously for more than 6 hours. Its “brain” is AI² Robotics' own GOVLA model, built on a fast-slow dual-system design: the slow system handles reasoning and breaks a task into steps, while the fast system directly outputs whole-body motion and movement trajectories, not just arm control. At launch, the company announced plans to build capacity for 10,000 units a year by 2028.","example":"At a trade show, AlphaBot 2 moves up to a shelf, raises and lowers its waist, and picks up items placed high or low.","related":["AI² Robotics","GOVLA (Global & Omni-body Vision-Language-Action)","Wheeled Humanoid Robot","Dual-System Architecture (System 1 / System 2)","Whole-Body Control","Mobile Manipulation"]},{"id":"unitree-a1","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree A1 官网","url":"https://www.unitree.com/A1"}],"as_of":"2026-09","related_ids":["quadruped-robot","unitree-robotics","unitree-go1","rapid-motor-adaptation","rl-based-locomotion-control","sim-to-real-transfer"],"name":"Unitree A1","alt":"宇树 A1","abbr":"","aliases":["A1"],"one_liner":"An early small-to-mid quadruped from Unitree, once a popular platform for reinforcement-learning locomotion research.","explanation":"A1 is an early small-to-mid-size quadruped robot from Unitree Robotics, reportedly launched around 2020, with direct-drive motors at the joints and a compact build. Published specs: a 5 kg payload, a maximum sustained running speed of 3.3 m/s (about 11.9 km/h), joints delivering up to 33.5 N·m of torque with a maximum joint speed of 21 rad/s, and 1–2.5 hours of runtime. Because it was priced far below platforms like ANYmal or Spot, it became a popular piece of hardware for academic work on reinforcement-learning locomotion and sim-to-real transfer around 2020–2022 — for example, the Rapid Motor Adaptation (RMA) paper ran its real-world experiments on an A1. Unitree later replaced it in its lineup with the Go1 and Go2.","example":"The RMA paper deploys a walking policy trained in simulation onto an A1, letting it walk stably across unfamiliar terrain such as grass, sand, and stairs.","related":["Quadruped Robot","Unitree Robotics","Unitree Go1","Rapid Motor Adaptation","RL-based Locomotion Control","Sim-to-Real Transfer"]},{"id":"unitree-go1","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree Go1 官网","url":"https://www.unitree.com/go1"}],"as_of":"2023-07","related_ids":["unitree-go2","unitree-robotics","quadruped-robot","rl-based-locomotion-control","sim-to-real-transfer","walk-these-ways"],"name":"Unitree Go1","alt":"宇树 Go1","abbr":"","aliases":["Go1"],"one_liner":"Unitree's 2021 consumer-grade quadruped robot, the predecessor to the Go2.","explanation":"Go1 is a small quadruped robot Unitree Robotics released in 2021, weighing about 12 kg with 12 joint motors across its four legs. It came in three editions — Air, Pro, and Edu — priced on Unitree's site at $2,700 for the Air and $3,500 for the Pro, with the Edu edition aimed at research and offering an open programming interface; Unitree states a top speed of about 4.7 m/s. It carries five sets of fisheye stereo depth cameras for following a person and avoiding obstacles. Go1 brought quadruped-robot pricing down to the tens-of-thousands-of-RMB range, and many universities used it for reinforcement-learning locomotion research and sim-to-real transfer (training in simulation, then deploying to the real robot); Unitree replaced it with the Go2 starting in 2023.","example":"MIT's Walk These Ways and other legged reinforcement-learning papers ran their real-robot experiments on a Go1.","related":["Unitree Go2","Unitree Robotics","Quadruped Robot","RL-based Locomotion Control","Sim-to-Real Transfer","Walk These Ways"]},{"id":"unitree-go2","category":"robot","sec":9,"tier":2,"sources":[{"title":"Introducing Unitree Go2 - Quadruped Robot of Embodied AI (PR Newswire)","url":"https://www.prnewswire.com/news-releases/introducing-unitree-go2---quadruped-robot-of-embodied-ai-301879381.html"},{"title":"Unitree Go2（百度百科英文版）","url":"https://baike.baidu.com/en/item/Unitree%20Go2/1517731"}],"as_of":"2023-07","related_ids":[null,null,null,null,null,null],"name":"Unitree Go2","alt":"宇树 Go2","abbr":"","aliases":["Go2","Go2 Air","Go2 Pro","Go2 EDU"],"one_liner":"Unitree's 2023 consumer robot dog, and one of the most common quadruped research platforms.","explanation":"Go2 is the consumer quadruped robot Unitree Robotics released in July 2023, succeeding the Go1, starting at RMB 9,997 (about $1,600), available in Air, Pro, and EDU configurations. It ships standard with Unitree's own L1 4D lidar for 360°×90° environmental sensing. Because it is inexpensive, durable, and comes with an open SDK, meaning a software development kit, the EDU version has been widely adopted by universities for reinforcement-learning locomotion research, meaning training a walking policy in simulation and deploying it to the real robot, along with navigation and legged mobile-manipulation research, making it one of the most frequently seen robot dogs in embodied-AI papers.","example":"A walking policy for Go2 is trained in Isaac Gym using unitree_rl_gym, then deployed to a real Go2 EDU for validation.","related":["Unitree Robotics","Quadruped Robot","Unitree Go1","Unitree Go2-W","RL-based Locomotion Control","unitree_rl_gym"]},{"id":"unitree-go2-w","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree Go2-W 官网","url":"https://www.unitree.com/go2-w"}],"as_of":"2026-09","related_ids":["wheel-legged-robot","unitree-go2","unitree-b2-w","quadruped-robot","rl-based-locomotion-control","unitree-robotics"],"name":"Unitree Go2-W","alt":"宇树 Go2-W","abbr":"","aliases":[],"one_liner":"A wheeled-leg version of Unitree's Go2, with a driven wheel mounted at the end of each leg.","explanation":"Go2-W is a wheeled-leg robot Unitree Robotics built on the Go2 platform: each leg ends in a 7-inch pneumatic wheel with a hub motor, so it rolls on its wheels on flat ground and lifts its legs to climb steps or slopes. Published specs: about 70 × 43 × 50 cm, about 18 kg including the battery, a normal payload of about 8 kg (up to about 12 kg at the limit), a speed of 0–2.5 m/s, a maximum climbing grade of 35°, roughly 1.5–3 hours of runtime, and an onboard 3D lidar and HD camera. Wheeled legs balance speed, energy use, and the ability to cross obstacles, but coordinating the wheels and legs is harder, and this is usually trained with reinforcement learning in simulation. It belongs to the same wheeled-leg product family as the larger B2-W.","example":"","related":["Wheel-legged Robot","Unitree Go2","Unitree B2-W","Quadruped Robot","RL-based Locomotion Control","Unitree Robotics"]},{"id":"unitree-b2","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree B2 官网","url":"https://www.unitree.com/b2"}],"as_of":"2026-09","related_ids":["quadruped-robot","unitree-robotics","unitree-b2-w","inspection-robot","special-purpose-robot","boston-dynamics-spot"],"name":"Unitree B2","alt":"宇树 B2","abbr":"","aliases":["B2"],"one_liner":"Unitree's large industrial-grade quadruped, built for heavy payloads, long runtime, and IP67 protection.","explanation":"B2 is a large industrial quadruped from Unitree Robotics aimed at industrial use, reportedly released in late 2023, positioned to compete with industrial-grade products like Boston Dynamics' Spot. Published specs: about 60 kg including the battery; a standing payload of at least 120 kg and a walking payload over 40 kg; a top speed above 6 m/s, which Unitree calls the fastest among industrial-grade quadrupeds; a 2,250 Wh battery good for over 5 hours and over 20 km unloaded, and still over 4 hours carrying a 20 kg load; and an IP67 dust- and water-resistance rating. It's aimed mainly at power-grid and chemical-plant inspection, emergency response, and outdoor hauling — scenarios that call for heavy payload and long runtime — and can also carry a robot arm on its back for mobile manipulation. The B2-W in the same family is a wheeled-leg version.","example":"At a chemical plant, a B2 carries a gas detector and camera to inspect piping corridors in the rain, autonomously climbing stairs to reach an upper platform.","related":["Quadruped Robot","Unitree Robotics","Unitree B2-W","Inspection Robot","Special-purpose Robot","Boston Dynamics Spot"]},{"id":"unitree-b2-w","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree B2-W 官网","url":"https://www.unitree.com/b2-w"}],"as_of":"2026-09","related_ids":["wheel-legged-robot","unitree-b2","unitree-go2-w","unitree-robotics","inspection-robot","rl-based-locomotion-control"],"name":"Unitree B2-W","alt":"宇树 B2-W","abbr":"","aliases":[],"one_liner":"A wheeled-leg version of Unitree's B2, with driven wheels for both walking and fast rolling.","explanation":"B2-W is a wheeled-leg robot Unitree Robotics built on the B2 platform, reportedly released in 2024: the feet at the end of all four legs are replaced with driven wheels, letting it roll like a vehicle on flat ground while still lifting a leg to step over curbs or rubble, combining a wheel's efficiency with a leg's ability to cross obstacles. Published specs: about 85 kg including the battery; a standing payload of up to 120 kg and a walking payload over 40 kg; a top speed of 15 km/h; about 30 km of unloaded range, dropping to about 25 km carrying a 40 kg load; and an IP67 rating. It suits scenarios with long distances and mixed terrain, such as campus patrol, long-distance outdoor transport, and emergency response, and its coordinated wheel-and-leg motion is often used to showcase reinforcement-learning locomotion-control capability.","example":"In the field, a B2-W carries supplies along a dirt road at high speed on its wheels, then lifts its legs to cross a ditch before rolling on.","related":["Wheel-legged Robot","Unitree B2","Unitree Go2-W","Unitree Robotics","Inspection Robot","RL-based Locomotion Control"]},{"id":"unitree-as2","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree As2 官网","url":"https://www.unitree.com/As2"}],"as_of":"2026-09","related_ids":["quadruped-robot","unitree-robotics","unitree-go2","unitree-b2","wheel-legged-robot","inspection-robot"],"name":"Unitree As2","alt":"宇树 As2","abbr":"","aliases":["As2-W"],"one_liner":"Unitree's compact industrial-grade quadruped, about 20 kg, with a wheeled-leg variant called As2-W.","explanation":"As2 is a compact industrial quadruped from Unitree Robotics, positioned between the consumer-grade Go2 and the larger industrial B2. Published specs: about 20 kg including the battery, a standing footprint of 720 × 378 × 457 mm, 12 joint motors with up to about 95 N·m of joint torque, a standing payload of up to 65 kg and a continuous walking payload of about 15 kg, a speed range of 0–3.7 m/s (up to about 5 m/s on some models), roughly 4 hours and 20 km of continuous unloaded walking, and an IP54 rain rating. It comes in four editions — AIR, PRO, X, and EDU — with pricing not published on Unitree's site. As2-W is a wheeled-leg version in the same family, replacing the feet with driven wheels for fast, long-distance travel on flat ground. It's aimed at industrial uses like inspection and security, with the EDU edition open to further development.","example":"At a substation, an As2 carries a camera along a fixed inspection route and steps over curbs and stairs using its legs when needed.","related":["Quadruped Robot","Unitree Robotics","Unitree Go2","Unitree B2","Wheel-legged Robot","Inspection Robot"]},{"id":"unitree-a2","category":"robot","sec":9,"tier":2,"sources":[{"title":"Unitree launches A2 quadruped equipped with front and rear lidar (The Robot Report)","url":"https://www.therobotreport.com/unitree-launches-a2-quadruped-equipped-with-front-and-rear-lidar/"},{"title":"Unitree launches A2 quadruped robot with 100 kg load capacity and 20 km range","url":"https://roboticsandautomationnews.com/2025/08/07/unitree-unveils-new-a2-quadruped-robot-with-100-kg-load-capacity-and-20-km-range/93591/"}],"as_of":"2025-08","related_ids":[null,null,null,null,null,null],"name":"Unitree A2","alt":"宇树 A2","abbr":"","aliases":[],"one_liner":"Unitree's 2025 industrial quadruped, with a lidar unit at both its front and back.","explanation":"Unitree A2 is the industrial-grade quadruped robot Unitree Robotics released in August 2025, positioned between the consumer Go2 and the heavier-duty B2. Published specs: about 37 kg unloaded, a maximum standing payload of 100 kg and a 25 kg walking payload, roughly 20 km of range unloaded, and a top speed of about 5 m/s; it carries an industrial lidar, a sensor that measures distance and builds a 3D map, plus a high-definition camera at both the front and back to eliminate blind spots, and its dual batteries can be hot-swapped. It targets inspection, logistics, emergency response, and research. Note that it is unrelated to AgiBot's humanoid robot also called “Expedition A2”; the two simply share a name.","example":"A power utility has the A2 carry inspection equipment and patrol a substation autonomously along a set route, swapping its battery to keep going when it runs low.","related":["Unitree Robotics","Quadruped Robot","Unitree B2","Unitree Go2","Inspection Robot","LiDAR"]},{"id":"unitree-z1-robotic-arm","category":"robot","sec":9,"tier":3,"sources":[{"title":"Unitree Z1 官网","url":"https://www.unitree.com/z1"}],"as_of":"2026-09","related_ids":["robotic-arm","legged-mobile-manipulator","mobile-manipulation","unitree-robotics","6-axis-robot-arm","payload"],"name":"Unitree Z1 Robotic Arm","alt":"宇树 Z1 机械臂","abbr":"","aliases":["Z1","Z1 AIR","Z1 PRO"],"one_liner":"Unitree's lightweight 6-axis robot arm, mountable on a quadruped's back for mobile manipulation.","explanation":"Z1 is a lightweight 6-degree-of-freedom robot arm from Unitree Robotics, offered in AIR and PRO editions: they weigh about 4.3 kg and 4.5 kg, carry payloads of 2 kg and over 3 kg, have a reach (working radius) of about 740 mm, and repeat positioning within about 0.1 mm. Unitree's site says it's designed to pair with the company's mobile robots, such as the Aliengo and B1 quadrupeds, mounting on a quadruped's back to form an “arm-equipped quadruped” for legged mobile-manipulation tasks like opening doors or picking things up off the ground; it can also be bolted to a tabletop on its own for manipulation research.","example":"Mounting a Z1 on the back of a Unitree B1 quadruped lets the robot dog walk while using the arm to reach for an object on the ground.","related":["Robotic Arm","Legged Mobile Manipulator (Quadruped with Arm)","Mobile Manipulation","Unitree Robotics","6-Axis Robot Arm","Payload"]},{"id":"deep-robotics-lite3","category":"robot","sec":9,"tier":3,"sources":[{"title":"云深处发布绝影 Lite3机器狗：面向教育科研场景 - 腾讯云","url":"https://cloud.tencent.cn/developer/news/1027675"},{"title":"重磅！云深处发布绝影Lite3机器狗，售价16900起 - 机器人大讲堂","url":"https://www.leaderobot.com/news/676"}],"as_of":"2023-03","related_ids":["quadruped-robot","deep-robotics","rl-based-locomotion-control","unitree-go2","sim-to-real-transfer","research-and-education-market"],"name":"DEEP Robotics Lite3","alt":"云深处 绝影 Lite3","abbr":"","aliases":["Lite3","Jueying Lite3"],"one_liner":"DEEP Robotics' small quadruped for education and research, starting at RMB 16,900, open for custom development.","explanation":"Jueying Lite3 is the small quadruped robot DEEP Robotics released on March 15, 2023, aimed at university research, education, and robotics hobbyists, launched at RMB 16,900 and up, in Experience, Explorer, Professional, and Lidar configurations. Published specs: a sustained walking payload of 7.5 kg, roughly 90 minutes of locomotion battery life and about 5 km of range, and a 1 kHz control frequency running on a real-time control system. It exposes modular interfaces for adding RTK positioning, 5G, an AI compute module, and various sensors, and supports custom development. Labs commonly use it for reinforcement-learning locomotion control, meaning training a walking policy in simulation and deploying it to the real robot, and for navigation research, making it a domestic research robot dog in the same class as Unitree's Go series.","example":"A researcher trains a walking policy for the Lite3 in Isaac Gym, then deploys it to the real robot through the official SDK for a sim-to-real transfer experiment.","related":["Quadruped Robot","DEEP Robotics","RL-based Locomotion Control","Unitree Go2","Sim-to-Real Transfer","Research & Education Market"]},{"id":"deep-robotics-jueying-x30","category":"robot","sec":9,"tier":3,"sources":[{"title":"绝影X30 产品资料 - 云深处科技","url":"https://deep-website.oss-cn-hangzhou.aliyuncs.com/app/X30.pdf"},{"title":"云深处发布行业应用旗舰机器狗绝影X30 - ITBear","url":"https://www.itbear.com.cn/html/2023-10/474576.html"}],"as_of":"2023-10","related_ids":["quadruped-robot","inspection-robot","deep-robotics","ingress-protection-rating","boston-dynamics-spot","special-purpose-robot"],"name":"DEEP Robotics Jueying X30","alt":"云深处 绝影 X30","abbr":"","aliases":["X30","X30 Pro"],"one_liner":"DEEP Robotics' industry-grade flagship quadruped, for power, tunnel inspection, and emergency response.","explanation":"Jueying X30 is the industry-grade quadruped robot DEEP Robotics released in October 2023, aimed at inspection, reconnaissance, security, and surveying in industrial settings. Published specs: about 56 kg including battery, an effective payload of at least 20 kg, 2.5 to 4 hours of battery life, a maximum climb of 45°, IP67 protection, and an operating range of -20°C to 55°C. It emphasizes sensor fusion, able to navigate and work autonomously in dim, bright, flickering, or even completely dark conditions. Unlike the small robot dogs used in labs, the X30's selling point is being able to operate long-term in harsh environments such as substations, utility tunnels, and tunnels, making it a representative Chinese product in the industrial quadruped-inspection market, often compared with Boston Dynamics' Spot.","example":"A power station schedules the X30 for regular inspections; it climbs stairs and passes through utility tunnels along a preset route, capturing gauge readings and infrared temperature images.","related":["Quadruped Robot","Inspection Robot","DEEP Robotics","Ingress Protection (IP) Rating","Boston Dynamics Spot","Special-purpose Robot"]},{"id":"deep-robotics-lynx","category":"robot","sec":9,"tier":3,"sources":[{"title":"云深处山猫 M20 行业级全地形轮足机器人发布 - IT之家","url":"https://www.ithome.com/0/849/905.htm"},{"title":"山猫 - 云深处科技官网","url":"https://www.deeprobotics.cn/robot/wap/lynx.html"}],"as_of":"2025-04","related_ids":["wheel-legged-robot","quadruped-robot","deep-robotics","inspection-robot","unitree-b2-w","deep-robotics-jueying-x30"],"name":"DEEP Robotics Lynx","alt":"云深处 山猫","abbr":"","aliases":["Lynx","Lynx M20"],"one_liner":"DEEP Robotics' wheel-legged robot dog, with wheels at the end of its legs for running and climbing.","explanation":"Lynx is DEEP Robotics' wheel-legged robot series: each of its four legs ends in a driven wheel, letting it roll on wheels over flat ground to save energy and move fast, then step like a quadruped over stairs or rubble. The industry-grade Lynx M20 was released on April 29, 2025, aimed at complex terrain and hazardous-environment work; the company states it supports a 15 kg effective payload, about 12 km of loaded range, industrial-grade IP66 protection, an operating range of -20°C to 55°C, lidar for all-around obstacle avoidance, and, in the Pro version, autonomous navigation and autonomous charging, with mechanical and electrical interfaces on its back for mounting equipment. The wheel-legged form combines a wheel's efficiency with a leg's ability to cross terrain, and has become a popular direction for quadruped robots in recent years.","example":"The Lynx M20 can carry inspection equipment along a factory road at wheeled speed, then switch to stepping when it reaches a staircase.","related":["Wheel-legged Robot","Quadruped Robot","DEEP Robotics","Inspection Robot","Unitree B2-W","DEEP Robotics Jueying X30"]},{"id":"xiaomi-cyberdog","category":"robot","sec":9,"tier":3,"sources":[{"title":"CyberDog - ROBOTS: Your Guide to the World of Robotics (IEEE)","url":"https://robotsguide.com/robots/cyberdog"}],"as_of":"2023-08","related_ids":["quadruped-robot","xiaomi","xiaomi-cyberone","quasi-direct-drive","unitree-go1","nvidia-jetson"],"name":"Xiaomi CyberDog","alt":"小米 CyberDog","abbr":"","aliases":["CyberDog 2","Iron Egg"],"one_liner":"Xiaomi's 2021 open-source quadruped robot, nicknamed “Iron Egg” in Chinese.","explanation":"CyberDog is a quadruped robot Xiaomi unveiled in August 2021; the first batch, the “Pioneer Edition,” was limited to 1,000 units at RMB 9,999, aimed at developers and hobbyists. It weighs about 14 kg with 12 quasi-direct-drive joint motors (a planetary gearbox paired with a brushless motor), a top running speed of about 11.5 km/h, and it can do a backflip. Its onboard computer is an NVIDIA Jetson Xavier NX, and its sensors include a RealSense depth camera, an ultra-wide camera, ultrasonic sensors, time-of-flight sensors, GPS, and an IMU; it runs Ubuntu and ROS, with its software released as open source. In 2023, Xiaomi released the second generation, CyberDog 2, reportedly priced at RMB 12,999.","example":"","related":["Quadruped Robot","Xiaomi","Xiaomi CyberOne","Quasi-Direct Drive","Unitree Go1","NVIDIA Jetson"]},{"id":"agibot-d1-quadruped","category":"robot","sec":9,"tier":3,"sources":[{"title":"IT之家：智元首款四足机器人 D1 Ultra 上线官网","url":"https://www.ithome.com/0/870/216.htm"},{"title":"IT之家：智元机器人全系产品开售","url":"https://www.ithome.com/0/876/106.htm"}],"as_of":"2025-08","related_ids":["quadruped-robot","agibot","rl-based-locomotion-control","unitree-go2","deep-robotics-lite3","ingress-protection-rating"],"name":"AgiBot D1 Quadruped","alt":"智元 D1 四足","abbr":"","aliases":["D1 Pro","D1 Edu","D1 Ultra"],"one_liner":"AgiBot's quadruped robot-dog line, split into education/entertainment and industrial versions.","explanation":"D1 is AgiBot's first quadruped robot series. The industrial-grade D1 Ultra went live on the company's site in July 2025, and the full lineup went on sale in August alongside AgiBot's other products, in three versions: D1 Pro and D1 Edu, aimed at entertainment performance and research and education, and D1 Ultra, aimed at industrial uses such as inspection, rated IP54 (dust-resistant and splash-resistant). Published specs: a top running speed of 3.7 m/s, able to jump 35 cm high and climb continuous 16 cm steps; D1 Pro weighs about 15 kg with 1–2 hours of battery life. Its locomotion is controlled by a policy trained with reinforcement learning rather than a hand-tuned gait controller. Launch prices were RMB 13,000 for the D1 Pro and RMB 36,000 for the D1 Edu.","example":"A university lab uses the D1 Edu to run sim-to-real experiments in quadruped reinforcement-learning locomotion control.","related":["Quadruped Robot","AgiBot","RL-based Locomotion Control","Unitree Go2","DEEP Robotics Lite3","Ingress Protection (IP) Rating"]},{"id":"vbot-super-robot-dog","category":"robot","sec":9,"tier":3,"sources":[{"title":"Vbot 维他动力官网","url":"https://www.vbot.cn/"}],"as_of":"2026-09","related_ids":["quadruped-robot","vbot","consumer-grade-robot","companion-robot","unitree-go2","xiaomi-cyberdog"],"name":"Vbot Super Robot Dog","alt":"维他动力 Vbot 超能机器狗","abbr":"","aliases":["Big Head BoBo"],"one_liner":"A consumer quadruped robot from Vbot, controlled by voice instead of a remote control.","explanation":"The Vbot Super Robot Dog is the first product from the Beijing company Vbot (维他动力), nicknamed “Big Head” (大头), aimed at families and individual consumers. Vbot's website describes it as a “smart robot dog that needs no remote control”: it has agent-style AI capabilities, can follow voice commands, navigate and avoid obstacles autonomously around the home, and follow its owner automatically, with app control and over-the-air updates; it has expansion ports for a camera, a magnetic accessory mount, and a USB-C port, and the company says it can tow about 100 kg. Vbot's site reports a cumulative 6,540 preorders worth nearly RMB 100 million, with mass-production deliveries starting in China in April 2026 and an overseas version launching in the second quarter of 2026. The company is reportedly founded by Yu Yinan, a former Horizon Robotics executive.","example":"","related":["Quadruped Robot","Vbot","Consumer-Grade Robot","Companion Robot","Unitree Go2","Xiaomi CyberDog"]},{"id":"actuator","category":"hardware","sec":0,"tier":1,"sources":[{"title":"Actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Actuator"},{"title":"Boston Dynamics: An Electric New Era for Atlas (2024-04)","url":"https://bostondynamics.com/blog/electric-new-era-for-atlas/"}],"as_of":"","related_ids":["joint-actuator-module","servo-motor","speed-reducer-gearbox","hydraulic-actuation","torque-density","backdrivability"],"name":"Actuator","alt":"执行器","abbr":"","aliases":[],"one_liner":"The part that converts electrical, hydraulic, or pneumatic energy into motion and force — a robot's “muscle.”","explanation":"An actuator is the part of a robot that actually produces motion. By energy source, actuators are electric (motors), hydraulic, or pneumatic, plus newer types like shape-memory alloys and artificial muscles. Whatever command a controller computes — a target position, velocity, or torque — ultimately has to be carried out by an actuator, and its torque, speed, precision, weight, and heat output directly determine how fast, how heavy, and how finely the robot can move. Most humanoid and quadruped robots today use joint actuator modules built from a motor plus a gearbox; Boston Dynamics' early Atlas used hydraulics before switching to an all-electric version starting in 2024. Common specs for choosing an actuator include peak torque, torque density (torque per unit weight), and backdrivability (whether an outside force can push the joint back). Note: in Chinese, 驱动器 sometimes also means “actuator” (as in a series elastic actuator), but more often refers specifically to the motor-driver electronics that regulate current — the meaning depends on context.","example":"In April 2024, Boston Dynamics retired the hydraulic Atlas it had used for over a decade and replaced it with a new all-electric Atlas; the company said the new version is stronger and has a wider range of motion than any previous generation.","related":["Joint Actuator Module","Servo Motor","Speed Reducer / Gearbox","Hydraulic Actuation","Torque Density","Backdrivability"]},{"id":"rotary-actuator","category":"hardware","sec":0,"tier":2,"sources":[{"title":"Rotary actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_actuator"}],"as_of":"","related_ids":["actuator","joint-actuator-module","linear-actuator","speed-reducer-gearbox","revolute-joint","peak-torque"],"name":"Rotary Actuator","alt":"旋转执行器","abbr":"","aliases":["Rotary Joint Module"],"one_liner":"An actuator whose output is rotation, the power unit behind a robot's rotating joints.","explanation":"A rotary actuator is a drive unit whose output is rotation about an axis, as opposed to a linear actuator, whose output is straight-line motion. In humanoids and robot arms, it's usually packaged as an all-in-one joint module: a motor plus a reducer, plus an encoder and a driver, sometimes also a torque sensor and a brake — mount it and you have a working rotary joint. Selecting one mainly comes down to peak/rated torque, speed, weight, and backdrivability (whether it can be pushed from the output side by an external force). Most of a humanoid's shoulder, elbow, hip, and knee joints are rotary actuators, while linear actuators are commonly used at the shin and ankle, where a push-pull motion is needed — balancing the two is a basic trade-off in whole-robot design.","example":"","related":["Actuator","Joint Actuator Module","Linear Actuator (Electric Cylinder)","Speed Reducer / Gearbox","Revolute Joint","Peak Torque"]},{"id":"linear-actuator","category":"hardware","sec":0,"tier":2,"sources":[{"title":"Linear actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Linear_actuator"}],"as_of":"","related_ids":["rotary-actuator","planetary-roller-screw","ball-screw","actuator","linkage-transmission","parallel-ankle-mechanism"],"name":"Linear Actuator (Electric Cylinder)","alt":"线性执行器（直线执行器 / 电缸）","abbr":"","aliases":["Electric Cylinder","Linear Drive"],"one_liner":"An actuator that pushes or pulls in a straight line instead of rotating.","explanation":"A linear actuator produces straight-line motion. The most common type in robotics is the electric cylinder: a motor turns a lead screw or ball screw (or a planetary roller screw), which converts that rotation into linear motion, extending or retracting a rod. It's usually not used directly as a joint itself; instead it's mounted like a muscle on a linkage, so its push and pull produces joint rotation — for example at a humanoid's knee, ankle, or waist, or in the tiny electric cylinders that curl fingers in a dexterous hand. Advantages are high force output, self-locking ability, and high stiffness, and it's easy to place away from the joint itself; downsides are that speed and backdrivability are usually worse than a rotary joint module, and the mechanism is more complex. It's one of the two main actuator layouts discussed for humanoid robots, alongside rotary actuators.","example":"Each finger of the Inspire RH56 dexterous hand is curled by a miniature electric cylinder acting through a linkage.","related":["Rotary Actuator","Planetary Roller Screw","Ball Screw","Actuator","Linkage Transmission","Parallel Ankle Mechanism"]},{"id":"joint-actuator-module","category":"hardware","sec":0,"tier":1,"sources":[{"title":"Mini Cheetah: A Platform for Pushing the Limits of Dynamic Quadruped Control (ICRA 2019)","url":"https://ieeexplore.ieee.org/document/8793865"},{"title":"Unitree GO-M8010-6 关节电机参数（宇树官网）","url":"https://www.unitree.com/mobile/go1/motor/"}],"as_of":"","related_ids":["actuator","quasi-direct-drive","planetary-gearbox","strain-wave-gear","rotary-encoder","servo-drive"],"name":"Joint Actuator Module","alt":"关节模组","abbr":"","aliases":["Integrated Joint Module","Joint Motor Module"],"one_liner":"A robot joint unit that packages a motor, gearbox, encoder, and driver together into one part.","explanation":"A joint actuator module is the standard way modern humanoid and quadruped robots build a joint: a brushless motor, a gearbox (planetary, harmonic, or RV), an encoder (to measure angle), a driver (to control current), and often even a torque sensor and a brake (which locks the joint when power is cut) are all packaged into a single unit, exposing only power and communication lines — commonly CAN, RS-485, or EtherCAT — to the outside. This lets a robot maker simply bolt the module into the frame and send it position or torque commands, which makes both development and repair much easier. Its torque, speed, weight, backlash (the play between gear teeth), and heat generation determine the robot's movement capability, and it's also one of the biggest cost items in the whole robot. MIT's open-source Mini Cheetah joint design helped popularize low-cost joint actuator modules.","example":"Unitree's GO-M8010-6 is a common integrated joint motor used in quadruped robots, with a built-in planetary gearbox at roughly a 6.33:1 ratio.","related":["Actuator","Quasi-Direct Drive","Planetary Gearbox","Strain Wave Gear (Harmonic Drive)","Rotary Encoder","Servo Drive (Motor Driver)"]},{"id":"core-components","category":"hardware","sec":0,"tier":2,"sources":[{"title":"Industrial robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_robot"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Core Components","alt":"核心零部件","abbr":"","aliases":["Three Core Components"],"one_liner":"The key parts that determine a robot's performance and cost — traditionally the reducer, servo, and controller.","explanation":"“Core components” is industry terminology for the handful of key parts that determine a robot's performance, price, and supply security. In the industrial-robot era, people commonly spoke of the “three core components”: the reducer (which converts a motor's high speed into high torque), the servo system (a servo motor plus its driver), and the controller — together the largest share of a robot's total cost. For humanoid robots, the list usually expands to include joint actuator modules, planetary roller screws, frameless torque motors, coreless motors, dexterous hands, six-axis force sensors, and the main compute chip. The production capacity, yield, and level of domestic sourcing for these components directly determine whether a robot's price can come down and whether it can be mass-produced, which is why they're the main focus when industry reports and investors discuss supply-chain chokepoints and domestic substitution.","example":"Harmonic and RV reducers long depended on Japan's Harmonic Drive Systems and Nabtesco; the rise of Chinese makers such as Leaderdrive is often cited as an example of domestic substitution for a core component.","related":["Speed Reducer / Gearbox","Servo Motor","Robot Controller","Joint Actuator Module","Domestic Substitution","Bill of Materials Cost"]},{"id":"servo","category":"hardware","sec":0,"tier":2,"sources":[{"title":"Servo (radio control) - Wikipedia","url":"https://en.wikipedia.org/wiki/Servo_(radio_control)"},{"title":"SO-ARM100 (TheRobotStudio) GitHub","url":"https://github.com/TheRobotStudio/SO-ARM100"}],"as_of":"","related_ids":["robotis-dynamixel-servo","feetech-sts3215-servo","so-100-so-101-arm","servo-motor","pulse-width-modulation","desktop-robot-arm"],"name":"Servo (Smart Serial Bus Servo)","alt":"舵机","abbr":"","aliases":["Bus Servo","Smart Servo"],"one_liner":"A small position-control actuator with a motor, gears, position sensor, and control board built into one unit.","explanation":"A servo originally referred to the small actuator used in model aircraft to move control surfaces: a small box packing a motor, reduction gears, a position sensor (a potentiometer or magnetic encoder), and a control board — give it a target angle, and it drives itself there. Traditional servos take commands over PWM (pulse-width modulation); bus servos (smart servos) instead use a serial bus, so multiple servos can be chained on one line, each with its own ID, and can also report back position, current, and temperature. They're cheap and easy to use, but have limited torque and precision, so they're mostly used in education, desktop robot arms, small humanoids, and teleoperation leader arms. Newcomers doing real-robot experiments often start with a servo-based arm.","example":"LeRobot's SO-100 / SO-101 arm uses Feetech STS3215 bus servos, and a full arm costs only a few hundred dollars.","related":["ROBOTIS Dynamixel Servo","Feetech STS3215 Servo","SO-100 / SO-101 Arm","Servo Motor","Pulse Width Modulation (PWM)","Desktop Robot Arm"]},{"id":"robotis-dynamixel-servo","category":"hardware","sec":0,"tier":2,"sources":[{"title":"ROBOTIS e-Manual: DYNAMIXEL XL330-M288-T","url":"https://emanual.robotis.com/docs/en/dxl/x/xl330-m288/"},{"title":"low_cost_robot (Koch arm) README","url":"https://github.com/AlexanderKoch-Koch/low_cost_robot"}],"as_of":"2026-09","related_ids":["servo","robotis","koch-v1-1-arm","gello","trossen-robotics-viperx-300","rs-485"],"name":"ROBOTIS Dynamixel Servo","alt":"Dynamixel 舵机","abbr":"","aliases":["DXL","Dynamixel","Dynamixel Smart Actuator"],"one_liner":"A smart bus-networked servo from Korea's ROBOTIS, the most common choice in research and low-cost arms.","explanation":"Dynamixel is the smart-actuator product line from the Korean company ROBOTIS, packing a motor, reducer, control circuitry, encoder, and communication interface into a single module. Multiple servos can be strung along one bus in a daisy chain, each with its own ID; a host computer reads and writes registers for position, speed, current, temperature, and so on over the protocol, and can switch between position, velocity, and current control modes. Compared with building a motor-plus-driver setup from scratch, it eliminates a huge amount of wiring and tuning, which is why it has become a common choice for low-cost robot arms and leader-follower teleoperation rigs in robot learning. Common models range from the compact XL330 and XL430 to the more powerful XM430.","example":"The Koch v1.1 arm is built from XL430 and XL330 Dynamixel servos, and the GELLO teleoperation leader arm is also Dynamixel-based.","related":["Servo (Smart Serial Bus Servo)","ROBOTIS","Koch v1.1 Arm","GELLO","Trossen Robotics ViperX 300","RS-485"]},{"id":"feetech-sts3215-servo","category":"hardware","sec":0,"tier":2,"sources":[{"title":"Feetech STS3215 Magnetic Encoder 360° Serial bus Servo Evaluation - 飞特","url":"https://www.feetechrc.com/2020-05-13_56655.html"},{"title":"FeeTech 12V 30kg.cm Magnetic Encoding Servo STS3215 - RobotShop","url":"https://www.robotshop.com/products/feetech-12v-30kgcm-magnetic-encoding-servo-sts3215"},{"title":"Testing of Feetech STS3215 Servomotor: Backlash, Repeatability, and Torque - Robo9","url":"https://robonine.com/testing-of-feetech-sts3215-servomotor-backlash-repeatability-and-torque/"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"Feetech STS3215 Servo","alt":"飞特 STS3215 舵机","abbr":"","aliases":["STS3215","ST3215"],"one_liner":"A serial-bus smart servo from Feetech, the standard joint used in LeRobot's SO-100/SO-101 arms.","explanation":"The STS3215 is a serial-bus servo (a small joint unit that packages a motor, reduction gears, a position sensor, and a control board together) made by the Shenzhen company Feetech. It measures angle with a 12-bit magnetic encoder, giving 4,096 steps per revolution, and can rotate a full 360°; multiple servos are chained together on a half-duplex serial bus, each with its own ID, and each can report back its position, speed, voltage, current, temperature, and load. It comes in two common power variants, with a stall torque of about 19 kg·cm at 7.4V and about 30 kg·cm at 12V. It's cheap, easy to buy, and simple to wire, which is why Hugging Face's LeRobot chose it for every joint on its open-source SO-100/SO-101 arms — making it the de facto standard for entering low-cost embodied AI, and it also shows up in platforms like Koch and LeKiwi. Its downside is that the plastic gears have backlash, limiting precision and stiffness.","example":"Building an SO-101 requires 12 STS3215 servos (six each for the leader and follower arms); after assembly, each one gets its ID set individually before calibration.","related":["Servo (Smart Serial Bus Servo)","SO-100 / SO-101 Arm (LeRobot)","LeRobot (Hugging Face)","ROBOTIS DYNAMIXEL Servo","Magnetic Encoder","Backlash"]},{"id":"brushless-dc-motor","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Brushless DC electric motor","url":"https://en.wikipedia.org/wiki/Brushless_DC_electric_motor"}],"as_of":"","related_ids":["permanent-magnet-synchronous-motor","field-oriented-control","outrunner-motor","servo-drive","hall-effect-sensor","joint-actuator-module"],"name":"Brushless DC Motor","alt":"无刷直流电机","abbr":"BLDC","aliases":["BLDC Motor"],"one_liner":"A DC motor that swaps mechanical brushes for electronic commutation — the workhorse motor in robot joints.","explanation":"A brushless DC motor puts permanent magnets on the rotor and coils on the stator; a driver switches current through the coils electronically based on the rotor's position (sensed by Hall-effect sensors or an encoder), replacing the brushes and commutator that wear out in a brushed motor. As a result it lasts longer, runs more efficiently, and concentrates heat in the stator where it's easier to dissipate, while also giving higher torque density. Most motors in robot joints, drones, and dexterous hands are brushless, usually paired with field-oriented control (FOC) for precise current and torque control. It's structurally similar to a permanent magnet synchronous motor, and the two are often used interchangeably in engineering talk, with the main technical difference being in back-EMF waveform and drive method.","example":"A quadruped robot's joint module is typically an outrunner brushless motor plus a planetary gearbox and an encoder, with a driver using FOC to control the output torque.","related":["Permanent Magnet Synchronous Motor (PMSM)","Field-Oriented Control","Outrunner Motor","Servo Drive (Motor Driver)","Hall-Effect Sensor","Joint Actuator Module"]},{"id":"permanent-magnet-synchronous-motor","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Synchronous motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Synchronous_motor"}],"as_of":"","related_ids":["brushless-dc-motor","field-oriented-control","rare-earth-permanent-magnet","servo-motor","frameless-torque-motor","joint-actuator-module"],"name":"Permanent Magnet Synchronous Motor (PMSM)","alt":"永磁同步电机","abbr":"PMSM","aliases":["PMSM","Permanent Magnet AC Servo Motor"],"one_liner":"An AC motor with permanent magnets on its rotor, spinning in strict sync with the supply frequency.","explanation":"A permanent magnet synchronous motor has permanent magnets (usually neodymium-iron-boron) mounted on its rotor; the stator is fed three-phase AC current that creates a rotating magnetic field, and the rotor follows that field at exactly the same frequency, which is why it's called “synchronous.” Its structure is nearly identical to a brushless DC motor's; the main differences are the shape of the back-EMF waveform (sinusoidal versus trapezoidal) and how it's driven — a PMSM uses sinusoidal current together with field-oriented control (FOC), giving smoother torque and less noise. In robotics circles, the two terms are often used loosely and interchangeably. A PMSM is efficient and has high power density, but it needs a position sensor, such as an encoder, to tell the driver where the rotor is. Electric-vehicle drive motors, industrial servo motors, and most robot joint motors fall into this category.","example":"The frameless torque motor at the heart of a humanoid's joint module is, in most cases, essentially a permanent magnet synchronous motor driven with FOC.","related":["Brushless DC Motor","Field-Oriented Control","Rare-Earth Permanent Magnet (NdFeB)","Servo Motor","Frameless Torque Motor","Joint Actuator Module"]},{"id":"servo-motor","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Servomotor - Wikipedia","url":"https://en.wikipedia.org/wiki/Servomotor"}],"as_of":"","related_ids":["servo-drive","rotary-encoder","permanent-magnet-synchronous-motor","stepper-motor","joint-actuator-module","cascade-control"],"name":"Servo Motor","alt":"伺服电机","abbr":"","aliases":["Servo System"],"one_liner":"A motor with position feedback that can be closed-loop controlled precisely for position and speed.","explanation":"A servo motor is a motor fitted with position feedback, such as an encoder, that together with a servo drive forms a closed control loop able to precisely track position, velocity, or torque commands. “Servo” describes this closed-loop, tracking style of control, not any particular motor construction — internally it's usually a permanent-magnet synchronous motor or a brushless DC motor. Unlike a stepper motor, it can detect and correct errors in real time, keeping its accuracy even when the load changes. Industrial robots, CNC machine tools, and robot joints all make heavy use of servo motors. A humanoid robot's joint module is, in essence, a compact servo system: a motor plus a reducer plus an encoder plus a driver.","example":"","related":["Servo Drive (Motor Driver)","Rotary Encoder","Permanent Magnet Synchronous Motor (PMSM)","Stepper Motor","Joint Actuator Module","Cascade Control"]},{"id":"stepper-motor","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Stepper motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Stepper_motor"}],"as_of":"","related_ids":["servo-motor","brushless-dc-motor","rotary-encoder","open-loop-control","desktop-robot-arm","position-control"],"name":"Stepper Motor","alt":"步进电机","abbr":"","aliases":["Stepping Motor"],"one_liner":"A motor that turns a fixed angle per pulse, so it can be positioned open-loop without an encoder.","explanation":"A stepper motor divides one full turn into a fixed number of steps — commonly 1.8° per step, or 200 steps per revolution. Each pulse from the driver advances the rotor by exactly one step, so the controller always knows the motor's position just by counting pulses; that lets it run open-loop, with no position feedback, and keeps cost low. That makes steppers common in 3D printers, desktop CNC mills, and some low-cost desktop robot arms. The downside is that torque falls off sharply as speed rises, and if the load is too high the motor can silently “skip steps” without the controller noticing; a stepper also draws current and heats up even while just holding still. Robot joints that need force control, back-drivability, or fast dynamic response generally use a servo setup instead — a brushless motor paired with an encoder.","example":"The X, Y, and Z axes and the extruder of a desktop 3D printer are usually all driven by stepper motors.","related":["Servo Motor","Brushless DC Motor","Rotary Encoder","Open-loop Control","Desktop Robot Arm","Position Control"]},{"id":"frameless-torque-motor","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Torque motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Torque_motor"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Frameless Torque Motor","alt":"无框力矩电机","abbr":"","aliases":["Frameless Motor","Torque Motor"],"one_liner":"A motor sold as just a stator and rotor, with no housing or bearings, built to be mounted straight into a joint.","explanation":"A frameless torque motor is a permanent-magnet motor sold as a bare component: just a stator (the coils) and a rotor (a ring of magnets), with no housing, bearings, output shaft, or encoder — the robot manufacturer mounts it directly inside a joint housing of its own design. It's typically shaped as a large-diameter, thin, flat ring with many magnetic poles, so it can deliver substantial torque even at low speed, and it leaves a large center bore for routing cables or fitting a gearbox. This design eliminates couplings and duplicate structural parts, making the joint more compact, lighter, and stiffer. A common combination in the rotary joint modules of collaborative arms and humanoid robots is a frameless torque motor paired with a harmonic or planetary gearbox, a dual encoder, and a holding brake. Kollmorgen is one of the leading manufacturers.","example":"In a collaborative arm's joint, the frameless motor's rotor is fitted onto a hollow shaft and its stator is pressed into the joint housing, with cables routed through the center bore.","related":["Joint Actuator Module","Permanent Magnet Synchronous Motor","Hollow-Shaft Cable Routing","Strain Wave Gear (Harmonic Drive)","Direct Drive","Kollmorgen"]},{"id":"outrunner-motor","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Outrunner - Wikipedia","url":"https://en.wikipedia.org/wiki/Outrunner"}],"as_of":"","related_ids":["brushless-dc-motor","quasi-direct-drive","mit-mini-cheetah-actuator","motor-velocity-constant","torque-density","axial-flux-motor"],"name":"Outrunner Motor","alt":"外转子电机","abbr":"","aliases":["Outrunner Brushless Motor"],"one_liner":"A brushless motor whose magnet-lined outer shell spins as the rotor, giving it high torque.","explanation":"An outrunner motor is a brushless-motor layout in which the stator, carrying the windings, stays fixed in the center, while an outer shell lined with permanent magnets rotates around it as the rotor. Because the magnets sit farther from the shaft, the same electromagnetic force produces more torque, so outrunner motors naturally deliver low speed and high torque, with high torque density — but they also have higher rotational inertia, and because the windings are enclosed inside, they dissipate heat less easily. They first became popular in RC models and drones, where they're cheap and available in many sizes. In robotics, pairing one with a single low-ratio planetary reducer is the standard recipe for a quasi-direct-drive actuator: enough torque, good backdrivability, well suited to legged joints that need compliance and shock tolerance.","example":"MIT's Mini Cheetah joint actuator is exactly this: an outrunner brushless motor with a roughly 6:1 planetary reducer built into it.","related":["Brushless DC Motor","Quasi-Direct Drive","MIT Mini Cheetah Actuator","Motor Velocity Constant (KV Rating)","Torque Density","Axial Flux Motor"]},{"id":"axial-flux-motor","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Axial flux motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Axial_flux_motor"}],"as_of":"","related_ids":["brushless-dc-motor","permanent-magnet-synchronous-motor","torque-density","joint-actuator-module","outrunner-motor","hub-motor"],"name":"Axial Flux Motor","alt":"轴向磁通电机","abbr":"","aliases":["Disc Motor","Axial Field Motor"],"one_liner":"A flat, disc-shaped motor whose magnetic field crosses the air gap along the shaft's axis.","explanation":"In an ordinary (radial flux) motor, the magnetic field lines cross the air gap between stator and rotor along the radius, and the rotor is cylindrical; an axial flux motor instead shapes the stator and rotor as facing discs, with the field lines running parallel to the shaft. For the same overall size, this gives a larger effective radius and higher torque density (torque output per unit weight), and the motor itself is very thin along its axis, making it well suited to flat joint modules, hub motors, and electric-vehicle drives. The trade-off is that controlling the air gap and the axial magnetic pull between the discs is harder to manufacture, and cost is higher. Because humanoid robots want joints that are both short and powerful, this is an active area of development for joint motors.","example":"","related":["Brushless DC Motor","Permanent Magnet Synchronous Motor (PMSM)","Torque Density","Joint Actuator Module","Outrunner Motor","Hub Motor (In-Wheel Motor)"]},{"id":"coreless-motor","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Coreless DC motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Coreless_DC_motor"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Coreless Motor","alt":"空心杯电机","abbr":"","aliases":["Ironless Motor","Coreless DC Motor"],"one_liner":"A small motor whose rotor has no iron core — just a self-supporting cup of windings — light and fast-responding.","explanation":"A coreless motor is a small DC motor whose rotor skips the laminated steel core and instead forms the copper windings into a self-supporting, cup-shaped coil. Removing the core brings several benefits: the rotor is very light with low rotational inertia, so it starts, stops, and changes speed quickly; there's no eddy-current or hysteresis loss in a core, so efficiency is high; and there's no cogging torque (the notchy, step-by-step feel some motors have while turning), giving smoother motion at low speed. The trade-offs are poorer heat dissipation, a lower power ceiling, and a higher price. It's long been used in medical devices, precision instruments, and model aircraft, with Switzerland's Maxon and Germany's FAULHABER as the leading makers. In embodied AI, it's mainly used in the finger joints of dexterous hands, paired with a miniature lead screw or reduction gears to drive each finger.","example":"Many dexterous hands pack a coreless motor and a miniature lead screw into the palm or a finger, with one motor driving one degree of freedom.","related":["Dexterous Hand","Brushless DC Motor","Micro Lead Screw","Cogging Torque","maxon","FAULHABER"]},{"id":"rated-torque","category":"hardware","sec":1,"tier":2,"sources":[{"title":"What's Continuous Stall Torque vs Rated Torque vs Peak Torque? - Parker","url":"https://parkermotion.atlassian.net/wiki/spaces/EIPKB1/pages/23593062"},{"title":"maxon motor: Key information (datasheet explanations)","url":"https://people.ece.ubc.ca/leos/pdf/datasheets/Maxon/MaxonSpecs.pdf"}],"as_of":"","related_ids":["peak-torque","torque-speed-curve","motor-thermal-derating-overheat-protection","liquid-cooled-joint-actuators","torque-constant","payload"],"name":"Rated Torque","alt":"额定扭矩","abbr":"","aliases":["Continuous Torque"],"one_liner":"The torque a motor can output continuously for a long time without overheating.","explanation":"Rated torque is the torque a motor or joint module can output continuously under specified conditions (rated speed, ambient temperature, cooling method) without the windings exceeding their allowed temperature. It's fundamentally a thermal limit: more torque means more current and more heat, and the rated value is the point where heat generation and heat dissipation balance. It's listed alongside peak torque on spec sheets — peak is often several times the rated value, but can only be sustained briefly. When selecting a motor, engineers convert the torque profile over a full motion cycle into an equivalent value using root-mean-square (RMS), then compare that to rated torque: sustained loads, like an arm holding a payload for a long time or a humanoid standing for a long time, all need to fall within the rated range. Adding active cooling, such as liquid cooling, to a joint can raise the torque it can sustain.","example":"An arm holding a 1 kg object at 0.5 m from the shoulder generates about 4.9 N·m of gravitational torque at the shoulder from that alone; holding the pose for a long time requires the shoulder's rated torque to exceed the total gravitational torque, including the arm's own weight.","related":["Peak Torque","Torque-Speed Curve","Motor Thermal Derating / Overheat Protection","Liquid-Cooled Joint Actuators (Active Thermal Management)","Torque Constant (Kt)","Payload"]},{"id":"peak-torque","category":"hardware","sec":1,"tier":2,"sources":[{"title":"What's Continuous Stall Torque vs Rated Torque vs Peak Torque? - Parker","url":"https://parkermotion.atlassian.net/wiki/spaces/EIPKB1/pages/23593062"},{"title":"maxon motor: Key information (datasheet explanations)","url":"https://people.ece.ubc.ca/leos/pdf/datasheets/Maxon/MaxonSpecs.pdf"}],"as_of":"","related_ids":["rated-torque","torque-density","torque-speed-curve","motor-thermal-derating-overheat-protection","joint-actuator-module","stall-torque-no-load-speed"],"name":"Peak Torque","alt":"峰值扭矩","abbr":"","aliases":["Maximum Torque"],"one_liner":"The most torque a motor or joint can output briefly, but can't sustain.","explanation":"Peak torque is the maximum torque a motor, reducer, or full joint module can produce for a short time, measured in newton-meters (N·m). It's mainly limited by the driver's maximum current, saturation in the motor's magnetic circuit, and heat: more current means more torque, but the windings also heat up faster, so peak torque can only be held briefly before thermal protection kicks in and cuts it back. It's listed alongside rated torque (the torque a joint can sustain continuously) on a joint's spec sheet. For legged and humanoid robots, dynamic moves like jumping, standing up from a fall, or recovering from a push rely on peak torque, while everyday standing and walking depend on rated torque — so choosing a joint by peak torque alone is a mistake.","example":"When a humanoid jumps up from a squat, its knee joints must output far more torque than standing requires, for a fraction of a second — that's peak torque at work; if a joint had to sustain that torque for long, the motor would overheat and its output would be throttled.","related":["Rated Torque","Torque Density","Torque-Speed Curve","Motor Thermal Derating / Overheat Protection","Joint Actuator Module","Stall Torque / No-Load Speed"]},{"id":"stall-torque-no-load-speed","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Stall torque - Wikipedia","url":"https://en.wikipedia.org/wiki/Stall_torque"},{"title":"ROBOTIS e-Manual: XL330-M288-T specifications","url":"https://emanual.robotis.com/docs/en/dxl/x/xl330-m288/"}],"as_of":"","related_ids":["torque-speed-curve","peak-torque","rated-torque","servo","motor-thermal-derating-overheat-protection","torque-constant"],"name":"Stall Torque / No-Load Speed","alt":"堵转扭矩 / 空载转速","abbr":"","aliases":["Locked-Rotor Torque","Free-Run Speed"],"one_liner":"A motor's maximum torque when jammed still, and its top speed with no load attached.","explanation":"These are the two endpoint values on a motor or servo's spec sheet. Stall torque is the torque a motor can output when its output shaft is locked and its speed is zero — the theoretical maximum. No-load speed is the highest speed it reaches with nothing attached. For a DC motor, torque and speed roughly fall along a straight line running from the stall point to the no-load point (the torque-speed curve), and the actual operating point sits somewhere between the two. Current, and therefore heat, is highest at stall, and prolonged stalling can burn out the motor — so selecting a motor by stall torque alone is a mistake; rated torque and sustained-operation capability matter too.","example":"When picking a servo for a robot arm joint, the torque actually needed is usually only a small fraction of the stall torque, leaving margin for acceleration and heat.","related":["Torque-Speed Curve","Peak Torque","Rated Torque","Servo (Smart Serial Bus Servo)","Motor Thermal Derating / Overheat Protection","Torque Constant (Kt)"]},{"id":"torque-speed-curve","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Motor constants - Wikipedia","url":"https://en.wikipedia.org/wiki/Motor_constants"}],"as_of":"","related_ids":["stall-torque-no-load-speed","peak-torque","rated-torque","torque-constant","motor-thermal-derating-overheat-protection","power-density"],"name":"Torque-Speed Curve","alt":"扭矩-转速曲线（T-N 曲线）","abbr":"","aliases":["T-N Curve"],"one_liner":"A curve plotting the maximum torque a motor can deliver at each speed.","explanation":"A torque-speed curve has speed on one axis and torque on the other, describing the limits of what a motor can deliver at a given voltage. Torque is highest at low speed; as speed rises, back-EMF (the reverse voltage the motor itself generates as it spins) increasingly cancels out the supply voltage, so less current can flow in, and torque drops accordingly, reaching zero at the no-load speed. The curve is usually split into two zones: a continuous-operation zone, where the motor can run indefinitely without overheating, and a peak zone, usable only briefly. When picking a joint motor, engineers plot the torque and speed a task actually needs — a robot's jump takeoff, a fast leg swing — onto this curve, to check it falls within the motor's capability with some margin to spare.","example":"When designing a quadruped robot, engineers overlay the knee joint's torque-versus-speed trajectory during a jump onto a candidate motor's T-N curve to check it doesn't exceed the peak zone.","related":["Stall Torque / No-Load Speed","Peak Torque","Rated Torque","Torque Constant (Kt)","Motor Thermal Derating / Overheat Protection","Power Density (W/kg)"]},{"id":"motor-thermal-derating-overheat-protection","category":"hardware","sec":1,"tier":3,"sources":[{"title":"moteus r4.11 - mjbots（连续电流与散热条件对照）","url":"https://mjbots.com/products/moteus-r4-11"}],"as_of":"","related_ids":["peak-torque","rated-torque","liquid-cooled-joint-actuators","torque-limiting","servo-drive","torque-speed-curve"],"name":"Motor Thermal Derating / Overheat Protection","alt":"电机温升与过热保护","abbr":"","aliases":["Thermal Derating","Overtemperature Protection"],"one_liner":"A motor heats up with use, so the driver limits current by temperature and shuts down if it gets too hot.","explanation":"As a motor runs, current flowing through its windings generates heat (copper loss, proportional to the square of the current), and the core loses some energy as heat too, so temperature keeps climbing. Too much heat can burn the winding insulation or demagnetize the permanent magnets, so a driver typically uses a thermistor to monitor winding and power-transistor temperature: as it nears the limit, the driver derates itself, actively lowering the maximum current and torque it allows; past a threshold, it faults and shuts down. This is also why motors have separate peak and rated torque figures — peak torque can only be held for a few seconds, while sustained operation is limited to the rated value. For humanoid and quadruped robots, standing for a long time, holding a half-squat, or carrying a heavy load puts sustained high torque on the knee and hip joints, which is the most common trigger for overheating; common fixes include liquid-cooled joints, added heat sinks, and penalizing joint torque in reinforcement-learning reward functions.","example":"When a humanoid robot holds a half-squat while carrying a box for a long time, its knee motors heat up, the driver automatically lowers the current limit, and the robot's legs weaken and its motion slows — in severe cases it shuts down into protection mode.","related":["Peak Torque","Rated Torque","Liquid-Cooled Joint Actuators (Active Thermal Management)","Torque Limiting (Saturation)","Servo Drive (Motor Driver)","Torque-Speed Curve"]},{"id":"liquid-cooled-joint-actuators","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Water cooling - Wikipedia","url":"https://en.wikipedia.org/wiki/Water_cooling"}],"as_of":"","related_ids":["motor-thermal-derating-overheat-protection","rated-torque","peak-torque","joint-actuator-module","power-density","torque-density"],"name":"Liquid-Cooled Joint Actuators (Active Thermal Management)","alt":"关节液冷（主动散热）","abbr":"","aliases":["Liquid-Cooled Joint","Active Joint Cooling"],"one_liner":"Circulating coolant through a joint motor to carry away winding heat, so it can sustain high torque longer.","explanation":"Liquid-cooled joints add cooling channels around a joint module's motor housing or stator, with a pump circulating coolant that carries heat generated by the windings and driver away to a radiator. A motor's heat generation is roughly proportional to the square of its current, and the torque it can sustain continuously is mainly limited by winding temperature, not by how much force the motor can deliver instantaneously; when air cooling or passive heat dissipation can't keep up, the joint has to derate itself or trigger overheat protection. Liquid cooling can meaningfully raise sustained torque, letting a humanoid robot carry heavy loads or sustain highly dynamic motion for longer. The trade-off is the added pump, tubing, and coolant, which increase weight, complexity, and the risk of a leak. Electric-vehicle drive motors are almost universally liquid-cooled, and humanoid robot joints are starting to borrow the same approach.","example":"","related":["Motor Thermal Derating / Overheat Protection","Rated Torque","Peak Torque","Joint Actuator Module","Power Density (W/kg)","Torque Density"]},{"id":"torque-constant","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Motor constants - Wikipedia","url":"https://en.wikipedia.org/wiki/Motor_constants"},{"title":"GO-M8010-6 Motor User Manual V1.0","url":"https://techshare.co.jp/faq/wp-content/uploads/2023/12/GO-M8010-6_Motor_Data_User_Manual_V1.0.pdf"}],"as_of":"","related_ids":["torque-control","motor-velocity-constant","torque-speed-curve","sensorless-force-estimation","field-oriented-control","brushless-dc-motor"],"name":"Torque Constant (Kt)","alt":"力矩常数（Kt）","abbr":"Kt","aliases":["Kt","Motor Torque Constant"],"one_liner":"How much torque a motor produces per amp of current, in newton-meters per amp (N·m/A).","explanation":"The torque constant, Kt, describes the proportional relationship between a motor's current and its output torque: torque ≈ Kt × current. It's set by the motor's construction — magnets, number of winding turns — and stays roughly constant within the motor's rated range. This matters because when a robot joint does torque control, the driver is actually controlling current, and Kt is what converts a desired torque into the current to command; read the other way, current can be used to roughly estimate the force at a joint, which is the basis of sensorless force estimation. Kt is just another way of expressing the same thing as the back-EMF constant and the KV rating: a higher Kt means more torque for the same current, but a lower top speed. When a gearbox is involved, it's worth checking carefully whether a datasheet's number is measured at the motor side or the output side.","example":"The Unitree GO-M8010-6 joint motor's datasheet lists a torque constant of about 0.639 N·m/A.","related":["Torque Control","Motor Velocity Constant (KV Rating)","Torque-Speed Curve","Sensorless Force Estimation","Field-Oriented Control","Brushless DC Motor"]},{"id":"motor-velocity-constant","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Motor constants - Wikipedia","url":"https://en.wikipedia.org/wiki/Motor_constants"}],"as_of":"","related_ids":["torque-constant","outrunner-motor","brushless-dc-motor","torque-speed-curve","quasi-direct-drive","gear-ratio"],"name":"Motor Velocity Constant (KV Rating)","alt":"KV 值（转速常数）","abbr":"KV","aliases":["KV","KV Rating","Velocity Constant"],"one_liner":"How many more RPM a motor spins at no load for each extra volt applied.","explanation":"The KV rating is the most commonly quoted spec for a brushless motor, in units of rpm/V: a motor rated KV100 connected to 24V spins at roughly 2,400 rpm with no load. It's roughly inversely proportional to the torque constant (Kt, the torque produced per amp of current), with a common estimate of Kt ≈ 9.55 / KV (N·m/A). So a lower-KV motor produces more torque for the same current, at a lower speed, while a higher-KV motor does the opposite. Drones and RC models want high speed, so they typically use high-KV motors; robot joints want low speed and high torque, so they typically use large-diameter, low-KV outrunner motors paired with a low-ratio reducer. When choosing a motor, KV needs to be considered together with supply voltage, target speed, and gear ratio.","example":"A KV100 motor connected to a 24V supply spins at roughly 2,400 rpm with no load; by the estimation formula, its torque constant is about 0.095 N·m/A.","related":["Torque Constant (Kt)","Outrunner Motor","Brushless DC Motor","Torque-Speed Curve","Quasi-Direct Drive","Gear Ratio"]},{"id":"torque-density","category":"hardware","sec":1,"tier":2,"sources":[{"title":"Torque density - Wikipedia","url":"https://en.wikipedia.org/wiki/Torque_density"}],"as_of":"","related_ids":["joint-actuator-module","peak-torque","quasi-direct-drive","gear-ratio","power-density","mit-mini-cheetah-actuator"],"name":"Torque Density","alt":"扭矩密度","abbr":"","aliases":["Torque-to-Weight Ratio"],"one_liner":"How much torque a motor or joint produces per unit of weight (or volume).","explanation":"Torque density is usually expressed in N·m/kg (newton-meters per kilogram), sometimes measured by volume instead. It's a core metric for evaluating motors and joint modules: every joint in a robot has to move all the joints downstream of it, so a heavier joint makes limbs harder to move quickly and nimbly, and drains the battery faster. Ways to raise torque density include using large-diameter outrunner motors, axial-flux motors, better cooling, and adding a reducer to amplify torque — but a larger gear ratio also worsens backdrivability and increases reflected inertia, so legged robots often use a low-ratio, quasi-direct-drive approach to balance the two.","example":"MIT's Mini Cheetah uses a large-diameter, flat motor with roughly a 6:1 planetary reducer, achieving enough torque in a very light joint — a landmark design for quasi-direct-drive joints.","related":["Joint Actuator Module","Peak Torque","Quasi-Direct Drive","Gear Ratio","Power Density (W/kg)","MIT Mini Cheetah Actuator"]},{"id":"power-density","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Power-to-weight ratio - Wikipedia","url":"https://en.wikipedia.org/wiki/Power-to-weight_ratio"}],"as_of":"","related_ids":["torque-density","peak-torque","joint-actuator-module","liquid-cooled-joint-actuators","rare-earth-permanent-magnet","torque-speed-curve"],"name":"Power Density (W/kg)","alt":"功率密度","abbr":"","aliases":["Specific Power","Power-to-Weight Ratio"],"one_liner":"How much power a motor or joint delivers per kilogram, a measure of being light yet strong.","explanation":"Power density is the power a system can output per unit of mass (sometimes per unit of volume), usually expressed in W/kg. Power equals torque times rotational speed, so power density differs from torque density (N·m/kg, which only measures how much force something can produce) by also accounting for how fast it can spin. Legged robots need short bursts of high power for jumping, running, and getting up after a fall, while their legs must stay light, so joint motors are designed to maximize power density. Ways to raise it include high-performance neodymium-iron-boron magnets, better cooling (such as liquid-cooled joints), and a higher bus voltage. Power density is also used to evaluate batteries, where it should be kept distinct from energy density (Wh/kg), which determines battery life instead.","example":"","related":["Torque Density","Peak Torque","Joint Actuator Module","Liquid-Cooled Joint Actuators (Active Thermal Management)","Rare-Earth Permanent Magnet (NdFeB)","Torque-Speed Curve"]},{"id":"rare-earth-permanent-magnet","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Neodymium magnet - Wikipedia","url":"https://en.wikipedia.org/wiki/Neodymium_magnet"}],"as_of":"2025-04","related_ids":["permanent-magnet-synchronous-motor","brushless-dc-motor","torque-density","power-density","motor-thermal-derating-overheat-protection","core-components"],"name":"Rare-Earth Permanent Magnet (NdFeB)","alt":"稀土永磁（钕铁硼）","abbr":"NdFeB","aliases":["Neodymium Magnet","Neodymium-Iron-Boron"],"one_liner":"A strong magnetic material made mainly of neodymium, iron, and boron, core to robot joint motors.","explanation":"NdFeB (neodymium-iron-boron) is a permanent-magnet material made primarily of neodymium, iron, and boron, with a main phase of Nd2Fe14B, developed independently in the early 1980s by Japan's Sumitomo Special Metals and America's General Motors; it remains the strongest commercially available permanent magnet today. The permanent magnet synchronous motors, brushless motors, and coreless motors used in robot joints almost all use NdFeB magnets in their rotors, since they deliver more torque for the same size, which directly determines a joint's torque density and power density. Its weakness is that it demagnetizes at high temperature, so heavy rare earths like dysprosium and terbium are commonly added to improve heat tolerance. China holds the majority of the world's rare-earth mining and magnet-processing capacity, and starting in April 2025 imposed export controls on dysprosium, terbium, and some other medium and heavy rare earths and related products, affecting overseas motor supply chains.","example":"","related":["Permanent Magnet Synchronous Motor (PMSM)","Brushless DC Motor","Torque Density","Power Density (W/kg)","Motor Thermal Derating / Overheat Protection","Core Components"]},{"id":"slotted-vs-slotless-brushless-motor","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Cogging torque - Wikipedia","url":"https://en.wikipedia.org/wiki/Cogging_torque"},{"title":"Brushless DC electric motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Brushless_DC_electric_motor"}],"as_of":"","related_ids":["cogging-torque","brushless-dc-motor","coreless-motor","permanent-magnet-synchronous-motor","torque-density"],"name":"Slotted vs. Slotless Brushless Motor","alt":"有齿槽 / 无齿槽无刷电机","abbr":"","aliases":["Slotted BLDC Motor","Slotless BLDC Motor"],"one_liner":"Two brushless-motor designs distinguished by whether the stator core is slotted where the windings sit.","explanation":"The common slotted brushless motor winds its coils around teeth on a slotted stator core. That concentrates the magnetic circuit, giving strong torque at low cost, but as the rotor magnets sweep past the teeth and slots the magnetic reluctance changes, producing cogging torque — a faint catch-and-release feel, most noticeable at low speed. A slotless motor instead builds the windings into a self-supporting coil that sits against a smooth, toothless core; with no teeth, cogging torque nearly disappears, so rotation feels smooth and low-speed control is fine-grained. The tradeoff is a wider air gap, meaning lower torque for a given size and a higher price. Dexterous hands and precision camera gimbals, where smoothness matters most, often use slotless or coreless motors; leg and foot joints, which need high torque, usually stay with slotted motors.","example":"Precision pan-tilt camera mounts that need smooth, jitter-free motion at low speed often choose slotless motors to avoid the cogging-torque catch.","related":["Cogging Torque","Brushless DC Motor","Coreless Motor","Permanent Magnet Synchronous Motor (PMSM)","Torque Density"]},{"id":"cogging-torque","category":"hardware","sec":1,"tier":3,"sources":[{"title":"Cogging torque - Wikipedia","url":"https://en.wikipedia.org/wiki/Cogging_torque"}],"as_of":"","related_ids":["slotted-vs-slotless-brushless-motor","brushless-dc-motor","permanent-magnet-synchronous-motor","coreless-motor","torque-control","backdrivability"],"name":"Cogging Torque","alt":"齿槽转矩","abbr":"","aliases":["Detent Torque"],"one_liner":"The periodic pulsing torque from a permanent-magnet motor's magnets snapping toward stator teeth, even unpowered.","explanation":"A permanent-magnet motor's stator core has slots cut into it, and as the rotor's magnets pass the “teeth” and “slots,” the difference in magnetic reluctance pulls the rotor toward the position of least reluctance — so even with no power applied, turning the motor by hand feels like a series of small detents, or notches. This is cogging torque, and it superimposes a periodic ripple on top of the output torque; it's most noticeable at low speed, where it causes vibration and noise and hurts fine force control and the feel of backdriving. Common fixes include skewing the slots or poles, optimizing the slot-to-pole ratio, or switching to a slotless design that eliminates it entirely; control software can also use a lookup table for feedforward compensation. It's a factor worth watching when selecting a joint motor and calibrating force control.","example":"Slowly turning an unpowered brushless motor by hand and feeling a series of small detents is cogging torque.","related":["Slotted vs. Slotless Brushless Motor","Brushless DC Motor","Permanent Magnet Synchronous Motor (PMSM)","Coreless Motor","Torque Control","Backdrivability"]},{"id":"speed-reducer-gearbox","category":"hardware","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Gear train","url":"https://en.wikipedia.org/wiki/Gear_train"},{"title":"Wikipedia: Strain wave gearing","url":"https://en.wikipedia.org/wiki/Strain_wave_gearing"}],"as_of":"","related_ids":["gear-ratio","strain-wave-gear","planetary-gearbox","rotate-vector-reducer","joint-actuator-module","backlash"],"name":"Speed Reducer / Gearbox","alt":"减速器","abbr":"","aliases":["Gear Reducer","Precision Reducer"],"one_liner":"A geared mechanism mounted behind a motor that trades high speed and low torque for low speed and high torque.","explanation":"A speed reducer (or gearbox) sits between a motor and its load, using gears to convert the motor's high rotational speed and low torque into low speed and high torque. A motor by itself spins fast but doesn't produce much force, while a robot joint needs to move slowly with a lot of force, so almost every joint actuator module is built as “motor + gearbox + encoder + driver.” A larger gear ratio (input speed divided by output speed) gives more output torque, but backdrivability gets worse, and the effects of backlash and friction become more noticeable. Robots commonly use three kinds: harmonic drives (compact, nearly zero backlash, mostly used in arms and wrists), RV reducers (rigid and impact-resistant, mostly used in the large joints of industrial robot arms), and planetary gearboxes (low ratio, good backdrivability, mostly used in the quasi-direct-drive joints of quadrupeds and humanoids).","example":"On an industrial 6-axis arm, load-bearing joints like the base and upper arm commonly use an RV reducer, while lighter joints like the wrist commonly use a harmonic drive.","related":["Gear Ratio","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Rotate Vector (RV) Reducer","Joint Actuator Module","Backlash"]},{"id":"gear-ratio","category":"hardware","sec":2,"tier":2,"sources":[{"title":"Gear train - Wikipedia","url":"https://en.wikipedia.org/wiki/Gear_train"}],"as_of":"","related_ids":["speed-reducer-gearbox","strain-wave-gear","planetary-gearbox","backdrivability","reflected-inertia","quasi-direct-drive"],"name":"Gear Ratio","alt":"减速比","abbr":"","aliases":["Transmission Ratio","Reduction Ratio"],"one_liner":"How many turns a motor makes for one turn of the output shaft.","explanation":"Gear ratio is the ratio of input speed to output speed through a reducer (a gear mechanism that converts a motor's high speed into lower speed with more torque). A 10:1 ratio means the motor turns 10 times for every 1 turn of the output, ideally multiplying torque about tenfold while cutting speed to a tenth. Robot motors spin fast but produce little torque on their own, so a reducer is needed to drive the joint. Gear ratio is a core design trade-off: a high ratio gives strong, precise joints but makes them “stiff” and hard to backdrive (push from the outside), because the motor rotor's inertia is amplified at the output by the square of the ratio (reflected inertia); a low ratio has the opposite effect. Industrial arms typically use harmonic or RV reducers with high ratios, while legged robots often use quasi-direct-drive joints with low gear ratios.","example":"The joint actuators on MIT's Mini Cheetah use only a 6:1 planetary reduction, trading torque for good backdrivability and impact resistance.","related":["Speed Reducer / Gearbox","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Backdrivability","Reflected Inertia","Quasi-Direct Drive"]},{"id":"backlash","category":"hardware","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Backlash (engineering)","url":"https://en.wikipedia.org/wiki/Backlash_(engineering)"}],"as_of":"","related_ids":["speed-reducer-gearbox","strain-wave-gear","planetary-gearbox","dual-encoder","pose-repeatability"],"name":"Backlash","alt":"背隙","abbr":"","aliases":["Gear Play","Lost Motion"],"one_liner":"The small bit of free play where the input reverses direction before the output actually starts moving.","explanation":"Backlash is the gap left between meshing teeth in a transmission: when the input reverses direction, it has to travel through this gap before the output starts following it. It's usually measured in arcminutes. Backlash creates a small “dead zone” at every direction reversal, which lowers positioning accuracy and control stability — especially noticeable during fine manipulation or force control, and it can also cause oscillation in a control loop. Backlash varies a lot by gearbox type: a harmonic drive has close to zero, an RV reducer has very little, and an ordinary planetary gearbox has comparatively more. Whether the encoder sits on the motor side or the output side (a dual-encoder setup) also affects whether backlash can be detected and compensated for.","example":"If an arm's end effector sweeps back and forth over a small range and a joint has significant backlash, the end effector pauses briefly at each direction change before catching up, showing up as a small step in the trajectory.","related":["Speed Reducer / Gearbox","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Dual Encoder","Pose Repeatability"]},{"id":"planetary-gearbox","category":"hardware","sec":2,"tier":1,"sources":[{"title":"Epicyclic gearing - Wikipedia","url":"https://en.wikipedia.org/wiki/Epicyclic_gearing"},{"title":"Mini Cheetah: A Platform for Pushing the Limits of Dynamic Quadruped Control (ICRA 2019)","url":"https://ieeexplore.ieee.org/document/8793865"}],"as_of":"","related_ids":["speed-reducer-gearbox","gear-ratio","strain-wave-gear","quasi-direct-drive","backlash","joint-actuator-module"],"name":"Planetary Gearbox","alt":"行星减速器","abbr":"","aliases":["Planetary Gear Reducer","Planetary Gear Set"],"one_liner":"A gearbox that uses a sun gear, planet gears, and a ring gear to turn a motor's fast spin into high torque.","explanation":"A planetary gearbox is built from a central sun gear, several planet gears orbiting around it, an outer ring gear with internal teeth, and a carrier that connects the planet gears — named for its resemblance to planets orbiting a sun. A motor spins fast but produces little torque; the gearbox trades that speed for torque so the joint has enough force to move a leg or an arm. Its advantages are a compact structure, load sharing across multiple planet gears, and high efficiency per stage; a single stage also typically has a fairly low reduction ratio (often around 3:1 to 10:1), which makes it easy for an external force to push back through it — well suited to legged robots that need force control and impact tolerance. Its downside is that backlash (play between the gear teeth) is usually larger than in a harmonic drive, giving somewhat lower precision. A quasi-direct-drive joint is generally just a motor paired with a single-stage planetary gearbox.","example":"MIT's Mini Cheetah uses a single-stage planetary gearbox at roughly a 6:1 ratio in its joints.","related":["Speed Reducer / Gearbox","Gear Ratio","Strain Wave Gear (Harmonic Drive)","Quasi-Direct Drive","Backlash","Joint Actuator Module"]},{"id":"strain-wave-gear","category":"hardware","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Strain wave gearing","url":"https://en.wikipedia.org/wiki/Strain_wave_gearing"},{"title":"Harmonic Drive: Strain Wave Gear Technology","url":"https://www.harmonicdrive.net/technology/harmonicdrive"},{"title":"US Patent 2,906,143: Strain Wave Gearing（C. W. Musser，1955 年申请）","url":"https://patents.google.com/patent/US2906143A/en"}],"as_of":"","related_ids":["speed-reducer-gearbox","wave-generator","flexspline","circular-spline","backlash","leaderdrive"],"name":"Strain Wave Gear (Harmonic Drive)","alt":"谐波减速器","abbr":"","aliases":["Harmonic Drive","Strain Wave Gearing"],"one_liner":"A precision reducer that gets its ratio from a flexible gear ring bending elastically — compact with almost no backlash","explanation":"A strain wave gear is built from three parts: a wave generator (an elliptical cam wrapped in a flexible bearing), a flexspline (a thin-walled, elastically flexible gear ring with external teeth), and a circular spline (a rigid ring gear with internal teeth). As the wave generator turns, it flexes the flexspline into an ellipse, so it only meshes with the circular spline at the two ends of the long axis; because the flexspline typically has two fewer teeth than the circular spline, one full turn of the wave generator advances the flexspline by only two teeth relative to the circular spline, giving a very high reduction ratio from a single stage — roughly 30:1 to 320:1. The principle was invented by the American engineer C. Walton Musser, who filed the patent in 1955 (it was granted in 1959); Japan's Harmonic Drive Systems is the leading manufacturer, alongside Chinese makers such as Leaderdrive (绿的谐波) and Laifu Harmonic. It's compact, lightweight, and has near-zero backlash, making it widely used in collaborative robot arms and the arm and wrist joints of humanoid robots; its downsides are that the flexspline can fatigue over time, and its stiffness and impact resistance are lower than an RV reducer's.","example":"Joint modules on collaborative robot arms often use a hollow “frameless torque motor plus harmonic drive” design, with cables routed straight through the center.","related":["Speed Reducer / Gearbox","Wave Generator","Flexspline","Circular Spline","Backlash","Leaderdrive"]},{"id":"rotate-vector-reducer","category":"hardware","sec":2,"tier":2,"sources":[{"title":"Cycloidal drive - Wikipedia","url":"https://en.wikipedia.org/wiki/Cycloidal_drive"},{"title":"Nabtesco Precision Reduction Gears","url":"https://www.nabtesco.com/en/"}],"as_of":"","related_ids":["speed-reducer-gearbox","cycloidal-reducer","strain-wave-gear","planetary-gearbox","nabtesco","industrial-robot"],"name":"Rotate Vector (RV) Reducer","alt":"RV减速器","abbr":"RV","aliases":["RV Gearbox"],"one_liner":"A precision two-stage reducer combining planetary gears with a cycloidal drive, stiff and high-capacity.","explanation":"An RV reducer is a two-stage precision reducer: the first stage is a planetary gear set, and the second is a cycloidal drive (a cycloidal disc rolling eccentrically inside a ring of pins, producing a large reduction ratio). It was brought to the industrial robot market by Japan's Nabtesco. Compared with a harmonic drive, an RV reducer is larger and heavier, but stiffer, more shock-resistant, and able to carry heavier loads, so traditional six-axis industrial robots typically use RV reducers in heavily loaded joints like the waist, shoulder, and elbow, while using harmonic drives in the lighter forearm and wrist. Humanoid robots prioritize low weight and use RV reducers less often; understanding it is mainly useful for making sense of the industrial-robot and reducer supply chain.","example":"The rotating base axis of an industrial robot commonly uses an RV reducer to bear the weight and tipping moment of the entire arm.","related":["Speed Reducer / Gearbox","Cycloidal Reducer","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Nabtesco","Industrial Robot"]},{"id":"cycloidal-reducer","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Cycloidal drive - Wikipedia","url":"https://en.wikipedia.org/wiki/Cycloidal_drive"}],"as_of":"","related_ids":["speed-reducer-gearbox","rotate-vector-reducer","strain-wave-gear","planetary-gearbox","backlash","gear-ratio"],"name":"Cycloidal Reducer","alt":"摆线减速器","abbr":"","aliases":["Cycloidal Drive","Cycloidal Speed Reducer"],"one_liner":"A reducer that meshes an eccentrically spinning cycloidal disc with a ring of pins for a large speed reduction.","explanation":"A cycloidal reducer is a mechanical speed-reduction device: the input shaft drives an eccentric shaft that makes a disc, shaped with cycloidal teeth around its edge, “roll” inside a fixed ring of pins; for each full orbit the disc makes around the ring, it only rotates on its own axis by a small amount, producing a large reduction ratio (commonly tens-to-one in a single stage). Because many teeth share the load at once, it's shock-resistant, stiff, and can be built with very little backlash (the play when gears reverse direction). The RV reducer commonly used in the base and shoulder joints of industrial robot arms is exactly this: a planetary-gear stage combined with a cycloidal-drive stage. In humanoid and quadruped robots, it's often compared with harmonic drives and planetary gearboxes — more shock-resistant than a harmonic drive, but heavier and requiring tighter manufacturing tolerances.","example":"In the RV reducer used in a six-axis industrial robot's base and upper-arm joints, the second stage is exactly this cycloidal-pin structure.","related":["Speed Reducer / Gearbox","Rotate Vector (RV) Reducer","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Backlash","Gear Ratio"]},{"id":"worm-gear","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Worm drive - Wikipedia","url":"https://en.wikipedia.org/wiki/Worm_drive"}],"as_of":"","related_ids":["speed-reducer-gearbox","gear-ratio","backdrivability","strain-wave-gear","planetary-gearbox","trapezoidal-lead-screw-and-self-locking"],"name":"Worm Gear","alt":"蜗轮蜗杆","abbr":"","aliases":["Worm Drive","Worm Gear Reducer"],"one_liner":"A gear pair where a screw-shaped worm turns a wheel, giving a large reduction ratio and a 90° change in axis.","explanation":"A worm gear is a classic gear pair: one full turn of the screw-shaped worm advances the meshing worm wheel by only one or a few teeth, so a single stage can produce a reduction ratio of tens to one, while also turning the rotation axis 90 degrees. A common and useful feature is self-locking — an external force applied to the wheel can barely turn the worm backward — so the mechanism holds its position with the power off, and no separate brake is needed. The downside is high sliding friction between the tooth surfaces, which means low efficiency, heat buildup, and essentially no back-drivability, ruling it out for joints that need compliance or force control. In embodied-AI hardware, worm gears mostly show up in dexterous-hand fingers, camera gimbals, and lift mechanisms — small spaces where holding a fixed pose matters — a different tradeoff than harmonic or planetary reducers make.","example":"Some low-cost dexterous hands drive their fingers with a small motor and a worm gear, so the fingers stay closed around an object even with the power off.","related":["Speed Reducer / Gearbox","Gear Ratio","Backdrivability","Strain Wave Gear (Harmonic Drive)","Planetary Gearbox","Trapezoidal Lead Screw & Self-Locking"]},{"id":"wave-generator","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Strain wave gearing - Wikipedia","url":"https://en.wikipedia.org/wiki/Strain_wave_gearing"}],"as_of":"","related_ids":["strain-wave-gear","flexspline","circular-spline","gear-ratio","backlash","thin-section-bearing"],"name":"Wave Generator","alt":"波发生器","abbr":"","aliases":["Wave Generator Cam","Elliptical Cam"],"one_liner":"The input component of a harmonic reducer: an elliptical cam and flexible bearing that force the flexspline into an oval.","explanation":"The wave generator is one of the three core parts of a harmonic reducer, alongside the flexspline and the circular spline, and it's the input, connected directly to the motor. It's made of an elliptical cam wrapped in a thin-wall flexible bearing; inserted inside the thin-walled flexspline, it forces the flexspline into an oval shape so the flexspline's teeth mesh with the circular spline's internal teeth only at the two ends of the long axis. As the wave generator makes one full turn, the meshing points travel all the way around too, and since the flexspline typically has 2 fewer teeth than the circular spline, the flexspline ends up shifted by only 2 teeth relative to the circular spline per input revolution — producing a very large reduction ratio, commonly tens to over a hundred to one in a single stage. This is how harmonic reducers achieve a compact size with very little backlash, which is why they're used so widely in robot-arm and humanoid-robot joints.","example":"A harmonic reducer with a 100:1 ratio: a 200-tooth flexspline and a 202-tooth circular spline — the wave generator turns 100 times for every 1 output turn of the flexspline.","related":["Strain Wave Gear (Harmonic Drive)","Flexspline","Circular Spline","Gear Ratio","Backlash","Thin-Section Bearing"]},{"id":"flexspline","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Strain wave gearing - Wikipedia","url":"https://en.wikipedia.org/wiki/Strain_wave_gearing"}],"as_of":"","related_ids":["strain-wave-gear","wave-generator","circular-spline","gear-ratio","backlash"],"name":"Flexspline","alt":"柔轮","abbr":"","aliases":["Flexible Spline"],"one_liner":"The thin-walled, externally toothed, cup-shaped flexible gear in a harmonic drive that gets flexed into an ellipse.","explanation":"The flexspline is one of the three core parts of a harmonic drive (the other two being the wave generator and the circular spline). It's a thin-walled, cup- or hat-shaped metal part with external teeth. As the elliptical wave generator rotates inside it, it flexes the flexspline into an ellipse, so its teeth mesh with the circular spline's internal teeth only at the two ends of the long axis; the flexspline has two fewer teeth than the circular spline, so for each full turn of the wave generator, the flexspline shifts by two teeth relative to the circular spline, producing a reduction ratio in the hundreds-to-one range. The flexspline has to flex elastically over and over, so its material, heat treatment, and machining precision directly determine a harmonic drive's lifespan and accuracy, making it a key focus of China's push to manufacture harmonic drives domestically.","example":"","related":["Strain Wave Gear (Harmonic Drive)","Wave Generator","Circular Spline","Gear Ratio","Backlash"]},{"id":"circular-spline","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Strain wave gearing - Wikipedia","url":"https://en.wikipedia.org/wiki/Strain_wave_gearing"}],"as_of":"","related_ids":["strain-wave-gear","flexspline","wave-generator","gear-ratio","backlash","leaderdrive"],"name":"Circular Spline","alt":"刚轮","abbr":"","aliases":["Rigid Spline"],"one_liner":"The rigid, internally toothed outer ring of a harmonic drive, usually the fixed part.","explanation":"A harmonic drive has three parts: an elliptical wave generator, a thin-walled, flexible flexspline (with external teeth), and a circular spline (with internal teeth). The circular spline is a thick internal-tooth ring, usually with 2 more teeth than the flexspline. As the wave generator rotates, it flexes the flexspline into an ellipse, and the teeth at the two ends of the ellipse's long axis mesh with the circular spline; that meshing point rotates along with the wave generator. For each full turn of the wave generator, the flexspline shifts relative to the circular spline by only as many teeth as the tooth-count difference between them, which is how the mechanism achieves a large reduction ratio. The common configuration fixes the circular spline to the housing and takes output from the flexspline. The circular spline's tooth-form accuracy and stiffness directly affect backlash and transmission accuracy.","example":"With a 200-tooth flexspline and a 202-tooth circular spline, fixing the circular spline and taking output from the flexspline gives a reduction ratio of 200 ÷ 2 = 100:1.","related":["Strain Wave Gear (Harmonic Drive)","Flexspline","Wave Generator","Gear Ratio","Backlash","Leaderdrive"]},{"id":"crossed-roller-bearing","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Rolling-element bearing - Wikipedia","url":"https://en.wikipedia.org/wiki/Rolling-element_bearing"}],"as_of":"","related_ids":["strain-wave-gear","thin-section-bearing","joint-actuator-module","stiffness","core-components","payload"],"name":"Crossed Roller Bearing","alt":"交叉滚子轴承","abbr":"","aliases":["Crossed Cylindrical Roller Bearing"],"one_liner":"A thin bearing with rollers crossed at right angles, letting a single unit carry loads from every direction.","explanation":"A crossed roller bearing places a single row of cylindrical rollers in a V-shaped raceway between its inner and outer rings, with neighboring rollers' axes alternating at right angles to each other; this lets one bearing simultaneously carry radial force, axial force in both directions, and tipping moments, where an ordinary bearing arrangement would typically need a matched pair to do the same job. It has a thin cross-section, high stiffness, and good rotational accuracy, and is commonly used at a robot joint's output and in turntables; many harmonic-drive units even include one as a built-in output bearing that carries external loads directly. The trade-off is that it demands high precision from the mounting surfaces, and it costs more.","example":"In a collaborative robot-arm joint module, a crossed roller bearing at the harmonic drive's output commonly supports the bending moment transmitted from the end of the arm.","related":["Strain Wave Gear (Harmonic Drive)","Thin-Section Bearing","Joint Actuator Module","Stiffness","Core Components","Payload"]},{"id":"thin-section-bearing","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Kaydon Bearings (thin-section bearing manufacturer)","url":"https://www.kaydonbearings.com/"},{"title":"Bearing (mechanical) - Wikipedia","url":"https://en.wikipedia.org/wiki/Bearing_(mechanical)"}],"as_of":"","related_ids":["crossed-roller-bearing","joint-actuator-module","hollow-shaft-cable-routing","strain-wave-gear","lightweighting"],"name":"Thin-Section Bearing","alt":"薄壁轴承","abbr":"","aliases":["Thin-Wall Bearing","Constant Cross-Section Bearing"],"one_liner":"A bearing with a very thin cross-section and a large bore, suited to compact, hollow robot joints.","explanation":"A thin-section bearing has inner and outer rings with a very small cross-section, and — within a given product series — that cross-section stays constant instead of growing thicker as the bore diameter increases, so even a large-bore version stays thin and light. Robot joints need to be hollow for cable routing while also staying compact and light, so thin-section deep-groove, angular-contact, or four-point-contact ball bearings are common inside harmonic reducers, at the output of joint modules, and in turntables; cross-roller bearings are also often made in a thin form, since a single one can carry radial load, axial load, and tilting moment all at once. The tradeoff is lower stiffness and load capacity than a standard bearing of the same bore size, and tighter requirements on how precisely the housing and mating parts are machined.","example":"","related":["Crossed Roller Bearing","Joint Actuator Module","Hollow-Shaft Cable Routing","Strain Wave Gear (Harmonic Drive)","Lightweighting (Magnesium Alloy / Carbon Fiber)"]},{"id":"ball-screw","category":"hardware","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Ball screw","url":"https://en.wikipedia.org/wiki/Ball_screw"}],"as_of":"","related_ids":["planetary-roller-screw","linear-actuator","screw-lead","trapezoidal-lead-screw-and-self-locking","holding-brake"],"name":"Ball Screw","alt":"滚珠丝杠","abbr":"","aliases":[],"one_liner":"A screw and nut with ball bearings packed between their threads, turning rotation into linear motion efficiently.","explanation":"A ball screw converts a motor's rotary motion into linear motion: rows of steel balls sit in the matching helical grooves of the screw shaft and the nut, rolling and recirculating as the nut travels, which turns the sliding friction of an ordinary lead screw into rolling friction. That gives it high transmission efficiency, high precision, and low wear. It's widely used in CNC machine tools, linear modules, and Cartesian robots. In humanoid robots, ball screws are used to build linear actuators (electric cylinders) that drive joints like the knee or ankle; for heavier loads, a planetary roller screw, which has more contact area, is used instead. Because it's so efficient, a ball screw generally can't self-lock, so a holding brake is needed to keep the load from sliding back when power is cut.","example":"A CNC machine tool's worktable is driven by a servo motor turning a ball screw, giving it precise linear feed motion.","related":["Planetary Roller Screw","Linear Actuator (Electric Cylinder)","Screw Lead","Trapezoidal Lead Screw & Self-Locking","Holding Brake"]},{"id":"planetary-roller-screw","category":"hardware","sec":2,"tier":2,"sources":[{"title":"Roller screw - Wikipedia","url":"https://en.wikipedia.org/wiki/Roller_screw"}],"as_of":"","related_ids":["ball-screw","inverted-planetary-roller-screw","linear-actuator","screw-lead","tesla-optimus","domestic-substitution"],"name":"Planetary Roller Screw","alt":"行星滚柱丝杠","abbr":"","aliases":["Roller Screw"],"one_liner":"A screw drive that uses threaded rollers to turn motor rotation into high-force linear motion.","explanation":"A planetary roller screw is a mechanism that converts rotary motion into linear motion; it's structurally similar to a ball screw, but instead of balls, a ring of threaded rollers sits between the screw shaft and the nut, each roller spinning on its own axis while also orbiting the shaft, like planets. Swedish engineer Carl Bruno Strandgren produced an early practical design in the 1940s. Because it has many more contact lines than a ball screw of the same size, it carries a much larger load, is stiffer, and lasts longer, but the threads are harder to machine, and it can cost roughly ten times as much as a comparable ball screw — historically it was mainly used in aerospace and machine tools. With the rise of humanoid robots, it has been built into linear actuators to drive high-force joints like the knee, elbow, and ankle, making it a sought-after core component and a target for domestic-substitution efforts in China (replacing imported parts with locally made ones).","example":"The linear actuator Tesla showed for Optimus at AI Day 2022 pairs a motor with a planetary roller screw, converting rotation into the push-rod motion that drives the joint.","related":["Ball Screw","Inverted Planetary Roller Screw","Linear Actuator (Electric Cylinder)","Screw Lead","Tesla Optimus","Domestic Substitution"]},{"id":"inverted-planetary-roller-screw","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Roller screw - Wikipedia","url":"https://en.wikipedia.org/wiki/Roller_screw"}],"as_of":"","related_ids":["planetary-roller-screw","ball-screw","linear-actuator","screw-lead",null],"name":"Inverted Planetary Roller Screw","alt":"反向式行星滚柱丝杠","abbr":"","aliases":["Inverted Roller Screw"],"one_liner":"A planetary roller screw variant with the long thread cut inside the nut, so the rollers travel with the screw shaft.","explanation":"A planetary roller screw uses a ring of threaded rollers, instead of balls, to transmit force between the screw shaft and the nut, converting rotation into linear motion with higher load capacity and longer life than a ball screw. In the standard configuration, the screw shaft is long and the nut is short, with the nut traveling along the shaft; in the inverted configuration, this is reversed — the long thread is cut into the inside of the nut, and the rollers travel together with the shorter screw shaft along the nut, with output taken from the extending screw shaft. This lets the nut double as the motor's rotor, nesting the motor and the screw together into a very compact linear actuator. Humanoid robots, which need high push force from a small package at joints like the knee, ankle, and elbow, are a natural fit for this approach; Tesla's Optimus linear actuators are reported to use a planetary roller screw. It's difficult to machine, and its domestic manufacture is one focus of the robotics supply chain in China.","example":"","related":["Planetary Roller Screw","Ball Screw","Linear Actuator (Electric Cylinder)","Screw Lead","Tesla"]},{"id":"screw-lead","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Leadscrew - Wikipedia","url":"https://en.wikipedia.org/wiki/Leadscrew"},{"title":"Screw thread (Lead, pitch, and starts) - Wikipedia","url":"https://en.wikipedia.org/wiki/Screw_thread"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Screw Lead","alt":"导程","abbr":"","aliases":["Lead (Screw)"],"one_liner":"How far a screw's nut travels along the axis for one full turn of the screw.","explanation":"Lead is the distance a nut travels along the axis of a screw (a threaded transmission part that converts rotation into linear motion) for one complete turn, equal to the thread pitch multiplied by the number of thread starts. It functions like a “gear ratio” between rotation and linear motion: a smaller lead means the same motor torque produces more force and finer position resolution, but slower linear speed; a larger lead does the opposite. When designing a humanoid's linear actuator, or the micro lead screw driving a dexterous hand, the lead has to be chosen together with motor speed, torque, and the force required. A trapezoidal screw with a small lead can also be self-locking, meaning the load can't push it back once power is cut.","example":"With a ball screw with a 5 mm lead, at a motor speed of 3,000 rpm, the nut moves at 3,000 ÷ 60 × 5 = 250 mm/s.","related":["Ball Screw","Planetary Roller Screw","Linear Actuator (Electric Cylinder)","Micro Lead Screw","Trapezoidal Lead Screw & Self-Locking","Gear Ratio"]},{"id":"trapezoidal-lead-screw-and-self-locking","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Leadscrew - Wikipedia","url":"https://en.wikipedia.org/wiki/Leadscrew"}],"as_of":"","related_ids":["ball-screw","screw-lead","micro-lead-screw","backdrivability","linear-actuator","holding-brake"],"name":"Trapezoidal Lead Screw & Self-Locking","alt":"梯形丝杠（滑动丝杠）与自锁","abbr":"","aliases":["Acme Lead Screw","Sliding Lead Screw","T-Type Lead Screw"],"one_liner":"A lead screw with a trapezoidal thread that moves by sliding friction, and can lock itself when power is cut.","explanation":"A trapezoidal lead screw converts rotation into linear motion using a thread with a trapezoidal cross-section; the nut and screw make sliding contact, unlike a ball screw, which has rolling balls in between. It's cheap and mechanically simple, but has high friction and low efficiency. That same friction gives it a useful property called self-locking: when the thread's lead angle is smaller than the friction angle, an axial push on the nut can't force the screw to spin backward, so the load won't drift down on its own when power is cut, and no separate brake is needed. The tradeoff is poor back-drivability — an external force can't easily move the joint — which rules it out for applications needing compliance or force control. It's common in low-cost linear actuators, lift mechanisms, and some dexterous hands that use a small lead screw to push a finger.","example":"The Z-axis of a desktop 3D printer often uses a T8 trapezoidal lead screw, so the print bed doesn't sag when power is cut.","related":["Ball Screw","Screw Lead","Micro Lead Screw","Backdrivability","Linear Actuator (Electric Cylinder)","Holding Brake"]},{"id":"micro-lead-screw","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Leadscrew - Wikipedia","url":"https://en.wikipedia.org/wiki/Leadscrew"},{"title":"灵心巧手 - 全球领先的机器人灵巧手 | LinkerBot","url":"https://www.linkerbot.cn/"}],"as_of":"","related_ids":["coreless-motor","ball-screw","planetary-roller-screw","trapezoidal-lead-screw-and-self-locking","screw-lead","linkage-transmission"],"name":"Micro Lead Screw","alt":"微型丝杠","abbr":"","aliases":["Micro Ball Screw","Micro Planetary Roller Screw"],"one_liner":"A screw only a few millimeters in diameter that turns a small motor's rotation into linear push-pull motion.","explanation":"A micro lead screw is a screw, typically only a few millimeters in diameter, that comes in trapezoidal, ball, or planetary-roller variants, all serving to convert a motor's rotation into the nut's linear travel. In dexterous hands, a common approach pairs a coreless motor with a micro lead screw to make a small linear actuator, hidden inside the palm or a finger, pushing and pulling a linkage to curl the finger. A smaller lead (how far the nut travels per revolution) gives more push force but lower speed; a trapezoidal screw has more friction and can self-lock, so a finger can hold its grip even after power is cut. Humanoid robots need these in large volumes, and precision and lifespan remain a challenge for domestic manufacturing in China.","example":"The Linker Hand L20 from Linkerbot uses a brushless motor paired with a precision ball screw to drive a linkage that moves the finger joints.","related":["Coreless Motor","Ball Screw","Planetary Roller Screw","Trapezoidal Lead Screw & Self-Locking","Screw Lead","Linkage Transmission"]},{"id":"linear-motor","category":"hardware","sec":2,"tier":3,"sources":[{"title":"Linear motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Linear_motor"}],"as_of":"","related_ids":["linear-actuator","planetary-roller-screw","ball-screw","direct-drive","permanent-magnet-synchronous-motor","backlash"],"name":"Linear Motor","alt":"直线电机","abbr":"","aliases":["Linear Electric Motor"],"one_liner":"A motor that outputs straight-line force and motion directly, with no screw or other transmission.","explanation":"A linear motor lays its stator (coils or magnetic track) and its mover out along a straight line; once powered, the mover translates directly along the track, with no lead screw, gear, or other mechanism in between to turn rotation into linear motion. The benefits are high speed, high positioning accuracy, and no backlash (transmission play), so it's widely used in semiconductor equipment, precision stages, and high-speed sorting; the downsides are low force per unit volume, high cost, and no self-locking when power is cut. A common point of confusion for newcomers: what's usually called a “linear actuator” or “electric cylinder” on a humanoid robot is almost always a rotary motor paired with a planetary roller screw, not a true linear motor.","example":"The precision motion stages in wafer inspection equipment are commonly driven by linear motors, paired with a linear encoder scale for micron-level positioning.","related":["Linear Actuator (Electric Cylinder)","Planetary Roller Screw","Ball Screw","Direct Drive","Permanent Magnet Synchronous Motor (PMSM)","Backlash"]},{"id":"servo-drive","category":"hardware","sec":3,"tier":2,"sources":[{"title":"Servo drive - Wikipedia","url":"https://en.wikipedia.org/wiki/Servo_drive"},{"title":"ODrive Robotics","url":"https://odriverobotics.com/"}],"as_of":"","related_ids":["servo-motor","field-oriented-control","integrated-drive-and-control","ethercat","controller-area-network","odrive"],"name":"Servo Drive (Motor Driver)","alt":"电机驱动器","abbr":"","aliases":["Motor Drive","Servo Amplifier"],"one_liner":"The power-electronics board that turns control commands into motor current and runs low-level control loops.","explanation":"A motor driver is the power-electronics module between the controller and the motor. The host computer or onboard compute platform only sends high-level commands — go to this position, at this speed, with this torque — and the driver reads the encoder, switches power transistors (such as MOSFETs) to energize the motor windings via PWM, and internally runs the current loop, velocity loop, and position loop (cascade control); for brushless motors it also performs field-oriented control (FOC). It communicates with the main controller over a bus such as CAN, EtherCAT, or RS-485. Selecting one comes down to voltage, continuous/peak current, control frequency, and communication interface. In humanoid robots, the driver is often integrated together with the motor and reducer into the joint module itself, an arrangement called a drive-integrated joint.","example":"ODrive and mjbots moteus are open-source brushless motor drivers commonly used by robotics hobbyists.","related":["Servo Motor","Field-Oriented Control","Integrated Drive and Control","EtherCAT (Ethernet for Control Automation Technology)","Controller Area Network (CAN)","ODrive"]},{"id":"pulse-width-modulation","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Pulse-width modulation - Wikipedia","url":"https://en.wikipedia.org/wiki/Pulse-width_modulation"}],"as_of":"","related_ids":["servo-drive","field-oriented-control","servo","brushless-dc-motor","microcontroller-unit","raspberry-pi"],"name":"Pulse Width Modulation (PWM)","alt":"PWM（脉宽调制）","abbr":"PWM","aliases":["Pulse-Width Modulation"],"one_liner":"Controlling average voltage or power by rapidly switching a signal and adjusting how long it stays high.","explanation":"PWM is a way of controlling “how much power” using a digital switching signal: the frequency stays fixed, and only the fraction of each cycle spent high, the duty cycle, is adjusted. At a 50% duty cycle, the average voltage seen by the load is about half the supply voltage. Because the switching device is always either fully on or fully off, it wastes very little energy itself, which is why motor drivers, switching power supplies, and LED dimming all use it. In robots, a motor driver uses PWM in a three-phase inverter bridge to generate motor current, often paired with field-oriented control (FOC); a servo, meanwhile, encodes its target angle in the width of a PWM pulse. Both microcontrollers and the Raspberry Pi can output PWM directly.","example":"A typical analog servo receives one pulse every 20 ms, with a pulse width of roughly 1–2 ms corresponding to different target angles.","related":["Servo Drive (Motor Driver)","Field-Oriented Control","Servo (Smart Serial Bus Servo)","Brushless DC Motor","Microcontroller Unit (MCU)","Raspberry Pi"]},{"id":"absolute-encoder","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Rotary encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":["rotary-encoder","incremental-encoder","magnetic-encoder","optical-encoder","dual-encoder","joint-actuator-module"],"name":"Absolute Encoder","alt":"绝对值编码器","abbr":"","aliases":["Absolute-Type Encoder"],"one_liner":"An encoder that gives a unique reading for every angle, so it knows its position even after a power cycle.","explanation":"An encoder is a sensor that measures how far a motor or joint has rotated. An absolute encoder assigns a unique code to every position, so it can report the current angle the instant it's powered on; an incremental encoder only counts pulses, so after losing power it has to be re-homed before it knows where it is. For robots, an absolute encoder means the joints don't need to be driven to a hard limit to find zero at startup, and it's now standard in humanoid and robot-arm joint modules. Common implementations include magnetic encoders and optical encoders; multi-turn absolute encoders can also remember how many full rotations have occurred. Many joints place one encoder at the motor and another at the output, a setup called a dual encoder.","example":"When a humanoid robot powers on, the absolute encoders in its joints report each joint's angle directly, with no need to first move to a fixed pose to home.","related":["Rotary Encoder","Incremental Encoder","Magnetic Encoder","Optical Encoder","Dual Encoder","Joint Actuator Module"]},{"id":"incremental-encoder","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Incremental encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Incremental_encoder"}],"as_of":"","related_ids":["rotary-encoder","absolute-encoder","dual-encoder","homing-zero-offset-calibration","optical-encoder"],"name":"Incremental Encoder","alt":"增量式编码器","abbr":"","aliases":["Incremental-Type Encoder"],"one_liner":"An encoder that outputs pulses for each increment of rotation, needing a count to know position.","explanation":"An incremental encoder outputs a stream of pulses as the shaft turns, typically two channels, A and B, 90 degrees out of phase (a quadrature signal), plus a third Z channel that pulses once per revolution as a zero marker. The controller counts pulses to know how far the shaft has turned, and checks whether A or B leads the other to determine direction. It's mechanically simple, cheap, and easy to build at high resolution, so it's very common in motors. Its downside is that it only knows how far it has moved since power-on — losing power also loses position — so it needs to be homed at every startup, or rely on the Z pulse or a limit switch to re-establish a reference. Robot joints that use only incremental encoders typically perform a homing move at startup; avoiding that requires an absolute encoder, or a dual-encoder setup.","example":"A 3D printer, and many desktop robot arms, bump each axis against a limit switch at startup, precisely because they rely on incremental position feedback and need to home first.","related":["Rotary Encoder","Absolute Encoder","Dual Encoder","Homing / Zero-Offset Calibration","Optical Encoder"]},{"id":"optical-encoder","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Rotary encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":["rotary-encoder","incremental-encoder","absolute-encoder","magnetic-encoder","dual-encoder","servo-motor"],"name":"Optical Encoder","alt":"光电编码器","abbr":"","aliases":["Optical-Type Encoder"],"one_liner":"A high-precision encoder that measures rotation angle by shining light through a marked disc.","explanation":"An optical encoder is built from a light source, a code disc, and a photodetector: the disc is etched with alternating transparent and opaque stripes, and as it turns with the motor shaft, the light is periodically interrupted; the detector converts that into pulses, or reads a coded track, which is translated into a rotation angle. By output type, it can be incremental (counting pulses, needing to be re-homed after a power loss) or absolute (a unique code for every position, so the angle is known the instant it's powered on). Its advantages are high resolution and accuracy, and no sensitivity to magnetic fields, which is why it's common in industrial servo motors and high-end robot-arm joints; its downsides are sensitivity to dust, oil, and shock, plus higher size and cost than a magnetic encoder, which is why low-cost robot joints tend to use magnetic encoders instead.","example":"An industrial servo motor commonly carries a 17-bit-or-higher absolute optical encoder at its rear, resolving 2^17 = 131,072 distinct positions per revolution.","related":["Rotary Encoder","Incremental Encoder","Absolute Encoder","Magnetic Encoder","Dual Encoder","Servo Motor"]},{"id":"magnetic-encoder","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Rotary encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":["rotary-encoder","optical-encoder","absolute-encoder","dual-encoder","field-oriented-control","hall-effect-sensor"],"name":"Magnetic Encoder","alt":"磁编码器","abbr":"","aliases":["Mag Encoder"],"one_liner":"An angle sensor that uses a magnet plus a Hall or magnetoresistive chip to measure a motor shaft's rotation.","explanation":"A magnetic encoder is a type of encoder (a sensor that measures rotation angle): a small radially magnetized magnet is mounted on the end of the motor shaft, facing a Hall-effect or magnetoresistive chip, and the chip reads the direction of the magnetic field to determine the angle. It's small, cheap, and unbothered by dust or oil, so it's extremely common in robot joint modules, often supplying the rotor angle needed for field-oriented control (FOC, a motor-control technique). Compared with an optical encoder, its accuracy and resolution are usually lower, and it's more sensitive to external magnetic fields and mounting eccentricity, so it needs calibration. High-end joints often place one encoder at the motor and another at the output, forming a dual-encoder setup.","example":"Many open-source joint driver boards glue a small magnet to the end of the motor shaft, with an ams AS5047P magnetic-encoder chip soldered directly opposite it to read the rotor's angle.","related":["Rotary Encoder","Optical Encoder","Absolute Encoder","Dual Encoder","Field-Oriented Control","Hall-Effect Sensor"]},{"id":"hall-effect-sensor","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Hall effect sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Hall_effect_sensor"}],"as_of":"","related_ids":["brushless-dc-motor","magnetic-encoder","rotary-encoder","field-oriented-control","servo-drive"],"name":"Hall-Effect Sensor","alt":"霍尔传感器","abbr":"","aliases":["Hall Sensor","Hall Switch"],"one_liner":"A sensor that uses the Hall effect to turn changes in a magnetic field into an electrical signal.","explanation":"A Hall-effect sensor works on the Hall effect: when current flows through a semiconductor strip, a magnetic field perpendicular to it produces a voltage across the strip's sides, with a stronger field producing a larger voltage. It's cheap, contactless, and small, making it one of the most common position sensors used in motors. In a brushless motor, three Hall-effect elements are typically mounted to sense the rotor magnets' approximate position, which the driver uses to decide which winding phase to energize (commutation); pairing a Hall chip with a radially magnetized ring magnet also makes a simple magnetic encoder for measuring a joint's angle. The Hall effect is also used for current sensing and proximity switches. Its resolution and noise immunity are lower than an optical encoder's, so in robot joints it's often used just for commutation or coarse position, paired with a more precise encoder.","example":"In sensored brushless motors common in RC models and self-balancing scooters, the thin bundle of wires coming out of the back of the motor is the signal lines from its three Hall-effect sensors.","related":["Brushless DC Motor","Magnetic Encoder","Rotary Encoder","Field-Oriented Control","Servo Drive (Motor Driver)"]},{"id":"inductive-encoder","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Zettlex inductive encoders - Celera Motion","url":"https://www.celeramotion.com/zettlex/"},{"title":"Resolver (electrical) - Wikipedia","url":"https://en.wikipedia.org/wiki/Resolver_(electrical)"}],"as_of":"","related_ids":["rotary-encoder","magnetic-encoder","optical-encoder","resolver","hollow-shaft-cable-routing"],"name":"Inductive Encoder","alt":"电感式编码器","abbr":"","aliases":["Induction-Type Encoder"],"one_liner":"An encoder that measures angle or displacement using coils and electromagnetic induction.","explanation":"An inductive encoder is usually built from transmitter and receiver coils printed on a circuit board, plus a metal target disc that rotates with the shaft. The transmitter coil generates an alternating magnetic field, and as the target disc rotates to different positions, it changes the signal picked up by the receiver coil, from which the angle is computed. It works on a principle similar to a resolver, but packaged as a flat circuit board. Compared with an optical encoder, it isn't bothered by oil, dust, or condensation; compared with a magnet-based magnetic encoder, it's less sensitive to external magnetic interference, and it's easy to build with a large center bore in a ring shape, which suits hollow-shaft joints. The trade-off is somewhat more complex signal processing, with accuracy and cost that fall somewhere between a magnetic encoder and a high-end optical encoder, depending on the specific product.","example":"The thin, ring-shaped encoder wrapped around the output shaft in a hollow-shaft joint module may well be an inductive design.","related":["Rotary Encoder","Magnetic Encoder","Optical Encoder","Resolver","Hollow-Shaft Cable Routing"]},{"id":"resolver","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Resolver (electrical) - Wikipedia","url":"https://en.wikipedia.org/wiki/Resolver_(electrical)"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Resolver","alt":"旋转变压器（旋变）","abbr":"","aliases":["Rotary Transformer"],"one_liner":"An electromagnetic sensor, built like a small transformer, that measures motor shaft angle and resists heat and vibration.","explanation":"A resolver is an electromagnetic sensor for measuring rotation angle: an AC signal is applied to an excitation winding on the stator, and as the rotor turns, two output windings pick up signals proportional to the sine and cosine of the rotation angle, which a dedicated decoder chip (an RDC) converts into an angle. It has no optical components, and no electronics on the rotor itself, so it tolerates high temperatures, vibration, oil, and shock, and has long been used in electric-vehicle drive motors, aerospace, and industrial servo systems. Compared with optical or magnetic encoders, a resolver is tougher, but its accuracy and resolution usually fall short of a high-end optical encoder, and it needs extra decoding circuitry. Magnetic encoders are more common in robot joints; resolvers show up more often in high-power applications where reliability matters most.","example":"Electric-vehicle drive motors commonly use a resolver to supply the rotor angle needed for field-oriented control.","related":["Rotary Encoder","Magnetic Encoder","Optical Encoder","Permanent Magnet Synchronous Motor","Field-Oriented Control (FOC)"]},{"id":"dual-encoder","category":"hardware","sec":3,"tier":2,"sources":[{"title":"Rotary encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Dual Encoder","alt":"双编码器","abbr":"","aliases":[],"one_liner":"A joint with one encoder at the motor and another at the output, reading both the motor's angle and the joint's true angle.","explanation":"A dual encoder is a joint-module configuration: one encoder (a sensor that measures rotation angle) sits on the motor shaft, measuring how far the motor has turned, while a second sits at the gearbox's output, directly measuring where the joint actually is. With only a motor-side encoder, the joint angle has to be inferred as “motor angle divided by gear ratio,” and errors creep in from the gearbox's backlash and the flexspline's elastic deformation. Adding an output-side encoder lets the system read the real joint position directly, so it also knows its absolute position the instant it powers on, without needing to home first. The difference between the two readings can also be used to estimate transmission deformation and load torque. This improves positioning accuracy and safety, at the cost of more expense, bulk, and wiring. High-end collaborative arms and many humanoid-robot joint modules use dual encoders.","example":"A harmonic-drive joint might use a high-resolution incremental encoder at the motor for its velocity loop, and an absolute encoder at the output to read the true joint angle.","related":["Rotary Encoder","Absolute Encoder","Joint Actuator Module","Backlash","Strain Wave Gear (Harmonic Drive)","Homing / Joint Zero-Offset Calibration"]},{"id":"joint-zero-calibration","category":"hardware","sec":3,"tier":2,"sources":[{"title":"Rotary encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":["rotary-encoder","absolute-encoder","incremental-encoder","kinematic-calibration","forward-kinematics","so-100-so-101-arm"],"name":"Joint Zero Calibration (Homing)","alt":"关节零位标定（回零）","abbr":"","aliases":["Homing","Zero-Point Calibration"],"one_liner":"Telling a robot exactly where “zero degrees” is for each of its joints.","explanation":"Joint zero calibration is the process of establishing how each joint's encoder reading maps to the joint's true physical angle: the joint is moved to a known reference pose (a mechanical hard stop, a calibration fixture, or a specified zero posture), the encoder reading at that pose is recorded as an offset, and every later angle is computed relative to that zero point. Incremental encoders lose track of position when powered off, so the robot must be “homed” at startup; absolute encoders remember position but still need recalibrating after reassembly or after a motor or reducer is replaced. If the zero point is off, forward kinematics will compute the wrong end-effector position, and a policy trained on one robot may fail when moved to another — so this step is typically done first after any hardware change.","example":"Before first use, LeRobot's SO-101 arm requires running a calibrate command that moves each joint to a middle position and through its full range to record the zero point and joint limits.","related":["Rotary Encoder","Absolute Encoder","Incremental Encoder","Kinematic Calibration","Forward Kinematics (FK)","SO-100 / SO-101 Arm"]},{"id":"controller-area-network","category":"hardware","sec":3,"tier":2,"sources":[{"title":"Wikipedia: CAN bus","url":"https://en.wikipedia.org/wiki/CAN_bus"},{"title":"Wikipedia: CAN FD","url":"https://en.wikipedia.org/wiki/CAN_FD"}],"as_of":"","related_ids":["ethercat","rs-485","socketcan","canopen-cia-402-drive-profile","servo-drive","mit-mode"],"name":"Controller Area Network (CAN)","alt":"CAN 总线","abbr":"CAN","aliases":["CAN FD (CAN with Flexible Data-Rate)","CAN Bus"],"one_liner":"A two-wire field bus linking multiple motors and sensors together, the standard way robot joints talk to their controller.","explanation":"CAN is a serial communication bus Bosch designed for cars in the 1980s: a single pair of differential wires links multiple nodes together, giving strong noise immunity and simple wiring, with priority arbitration and error detection built into every frame. Classic CAN tops out at 1 Mbit/s with up to 8 data bytes per frame; CAN FD (CAN with Flexible Data-Rate) is the later upgrade, with a faster data phase and up to 64 bytes per frame. In robots, large numbers of joint motors, dexterous hands, and battery management systems all communicate over CAN or CAN FD, with a host computer sending and receiving commands through a USB-CAN adapter or Linux's SocketCAN. Its bandwidth is limited, so systems with many joints or high control frequencies are often split across multiple bus lines, or switched to EtherCAT instead.","example":"Joint motors like Damiao's and Xiaomi's CyberGear receive MIT-mode position, velocity, stiffness, damping, and feedforward-torque commands over CAN.","related":["EtherCAT (Ethernet for Control Automation Technology)","RS-485","SocketCAN","CANopen / CiA 402 Drive Profile","Servo Drive (Motor Driver)","MIT Mode"]},{"id":"canopen-cia-402-drive-profile","category":"hardware","sec":3,"tier":3,"sources":[{"title":"CANopen - Wikipedia","url":"https://en.wikipedia.org/wiki/CANopen"}],"as_of":"","related_ids":["controller-area-network","ethercat","socketcan","cyclic-synchronous-position-velocity-torque-modes","servo-enable","servo-drive"],"name":"CANopen / CiA 402 Drive Profile","alt":"CANopen / CiA 402 驱动协议","abbr":"","aliases":["CiA 402","DS402","CANopen Drive Profile"],"one_liner":"A CAN-based industrial protocol, and the device profile within it that standardizes servo-drive behavior.","explanation":"CANopen is a higher-layer protocol defined by CiA (the CAN in Automation association), which builds on the CAN bus to define an object dictionary (a numbered table of a device's parameters), two message types (PDO and SDO), and network management. CiA 402 is the “drives and motion control” device profile within it, standardizing a servo drive's state machine (power-on, enabled, fault, and so on), its control word and status word, and its operating modes for position, velocity, torque, and homing, including cyclic synchronous position/velocity/torque (CSP/CSV/CST). A drive that follows the profile can be controlled by the same master-side code as any other compliant drive. EtherCAT reuses this same profile (called CoE), so it shows up in robot joint drives whether they communicate over CAN or EtherCAT.","example":"Writing control-word values 0x06, 0x07, then 0x0F to a drive in sequence takes it from ready, through power-on, to the “operation enabled” state.","related":["Controller Area Network (CAN)","EtherCAT (Ethernet for Control Automation Technology)","SocketCAN","Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)","Servo Enable (Servo ON / OFF)","Servo Drive (Motor Driver)"]},{"id":"ethercat","category":"hardware","sec":3,"tier":2,"sources":[{"title":"EtherCAT - Wikipedia","url":"https://en.wikipedia.org/wiki/EtherCAT"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"EtherCAT (Ethernet for Control Automation Technology)","alt":"EtherCAT 总线","abbr":"EtherCAT","aliases":[],"one_liner":"A real-time Ethernet-based industrial bus from Beckhoff, commonly used to link a robot's joint drivers together.","explanation":"EtherCAT is an industrial Ethernet field bus created by Germany's Beckhoff, now maintained by the EtherCAT Technology Group and incorporated into the IEC 61158 international standard. Its defining trick is “processing on the fly”: the master sends out a single frame of data that flows through each slave device in sequence, and each slave reads out the command addressed to it and writes back its own status as the frame passes through, instead of every device having to send and receive separately — which lets the communication cycle be extremely short with very little clock-synchronization jitter. In a robot, the main controller acts as the master, with each joint driver and force sensor acting as a slave chained together in a line, exchanging position and torque commands in sync at 1 kHz or higher. Compared with a CAN bus, it has far more bandwidth, which suits humanoid robots with many joints, but it needs dedicated master-side software and pricier slave chips.","example":"An open-source master stack like SOEM or IgH, running on a real-time-patched Linux system, controls an entire humanoid robot's joint drivers on a 1 kHz cycle.","related":["EtherCAT Master (SOEM / IgH)","Controller Area Network (CAN)","Servo Drive / Motor Driver","Real-Time Control","Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)"]},{"id":"rs-485","category":"hardware","sec":3,"tier":3,"sources":[{"title":"RS-485 - Wikipedia","url":"https://en.wikipedia.org/wiki/RS-485"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"RS-485","alt":"RS-485 总线","abbr":"","aliases":["EIA-485","TIA-485"],"one_liner":"A differential twisted-pair serial standard, resistant to noise and usable over long cable runs, common for motors and servos.","explanation":"RS-485 (formally TIA/EIA-485) is a physical-layer standard for serial communication, sending a differential signal (read as the voltage difference between two wires) over a twisted pair; it resists electrical noise well, can run over distances up to a kilometer or more, and lets multiple devices share a single bus, usually in half-duplex mode (only one side can transmit at a time). It only defines the electrical characteristics; common higher-layer protocols riding on top of it include Modbus RTU or a vendor's own proprietary protocol. In robots, many servos, joint motors, electric grippers, and sensors use RS-485. Its weakness is a lower data rate than EtherCAT or CAN FD, and it relies on a master polling each device in turn, which becomes a bottleneck once a robot has many joints running at a high control frequency.","example":"An electric gripper commonly runs the Modbus RTU protocol over RS-485, with the host computer sending open/close commands through a USB-to-485 adapter.","related":["Controller Area Network / CAN with Flexible Data-Rate","EtherCAT (Ethernet for Control Automation Technology)","Servo (Smart Serial Bus Servo)","Servo Drive / Motor Driver","Unitree GO-M8010-6 Motor"]},{"id":"integrated-drive-and-control","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Servo drive - Wikipedia","url":"https://en.wikipedia.org/wiki/Servo_drive"},{"title":"Motor controller - Wikipedia","url":"https://en.wikipedia.org/wiki/Motor_controller"}],"as_of":"","related_ids":["servo-drive","joint-actuator-module","robot-controller","ethercat","compute-control-integration"],"name":"Integrated Drive and Control","alt":"驱控一体","abbr":"","aliases":["Drive-Control Integration"],"one_liner":"Combining a motor's driver and its motion controller into a single hardware unit.","explanation":"In a traditional industrial robot, the motion controller (which computes trajectories and does planning) and each axis's servo drive (which energizes the motor) are separate boxes connected by a bus. Integrated drive and control means merging the two onto one board or into one cabinet: the control algorithm and the current loop run on the same processor or the same module, eliminating the communication step between them. The benefits are a smaller footprint, less wiring, lower cost, and shorter control latency, which helps with high-bandwidth force control. At the joint level, the term is also often used for integrating the driver directly into the joint module itself, so each joint carries its own “drive plus control” and the layer above only needs to send commands over EtherCAT or CAN. The integrated joints used in humanoid robots and collaborative arms are essentially built on this idea.","example":"","related":["Servo Drive (Motor Driver)","Joint Actuator Module","Robot Controller","EtherCAT (Ethernet for Control Automation Technology)","Compute-Control Integration"]},{"id":"holding-brake","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Electromagnetic brake - Wikipedia","url":"https://en.wikipedia.org/wiki/Electromagnetic_brake"}],"as_of":"","related_ids":["joint-actuator-module","emergency-stop","safe-torque-off","servo-enable","safety-gantry"],"name":"Holding Brake","alt":"抱闸","abbr":"","aliases":["Brake","Electromagnetic Brake"],"one_liner":"A brake mounted on a motor or joint that automatically locks the shaft when power is cut.","explanation":"A holding brake is usually electromagnetic: when powered, an electromagnet pulls the friction plates apart so the motor can turn freely; when power is cut, a spring clamps the friction plates together and the shaft is locked, which is why it's sometimes called “fail-safe” or power-off braking. Its job is to hold a joint in place during a power loss, an emergency stop, or when the drive is disabled, so an arm doesn't fall under gravity and hurt someone or damage something; it's mainly for holding position, not for braking during high-speed motion. Most industrial and collaborative robot-arm joint modules include one. The trade-off is added weight, size, and cost, so some quadruped and humanoid robot joints skip the brake entirely and go limp when power is cut, which is why testing often uses a safety gantry. Whether a joint module has a brake, and how much torque that brake can hold, are common specs to check.","example":"When an emergency stop is pressed on a collaborative arm and it stays frozen in mid-air instead of dropping, that's the holding brakes in each joint at work.","related":["Joint Actuator Module","Emergency Stop","Safe Torque Off (STO)","Servo Enable (Servo ON / OFF)","Safety Gantry (Suspended Start)"]},{"id":"safe-torque-off","category":"hardware","sec":3,"tier":3,"sources":[{"title":"IEC 61800-5-2 (IEC Webstore)","url":"https://webstore.iec.ch/en/publication/30617"},{"title":"IEC 61508 - Wikipedia","url":"https://en.wikipedia.org/wiki/IEC_61508"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Safe Torque Off (STO)","alt":"安全扭矩关断","abbr":"STO","aliases":["STO"],"one_liner":"A driver safety function that cuts power to a motor in hardware, so it stops producing torque.","explanation":"STO is a drive safety function defined in IEC 61800-5-2: once triggered, the driver uses a hardware circuit to block the switching signals to the inverter, so the motor stops producing torque and coasts to a stop under its own inertia and friction (corresponding to an IEC 60204-1 Category 0 stop). It's typically implemented with a dual-channel hardware input that doesn't depend on software, allowing it to reach a fairly high functional-safety rating (SIL/PL). Importantly, STO is not the same as cutting main power, and it doesn't actively brake or hold up a load against gravity, so vertical joints still need a holding brake alongside it; when the motor needs to decelerate first and then cut torque, the related SS1 function is used instead. A robot's emergency-stop button is typically wired directly to the STO terminals on each driver.","example":"When an emergency stop is triggered on a collaborative arm, the controller asserts STO on every joint's driver while simultaneously engaging the holding brakes, preventing the arm from dropping under gravity.","related":["Emergency Stop","Functional Safety","Holding Brake","Servo Drive / Motor Driver","Protective Stop"]},{"id":"regenerative-braking-and-brake-resistor","category":"hardware","sec":3,"tier":3,"sources":[{"title":"Regenerative braking - Wikipedia","url":"https://en.wikipedia.org/wiki/Regenerative_braking"},{"title":"Braking chopper - Wikipedia","url":"https://en.wikipedia.org/wiki/Braking_chopper"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Regenerative Braking and Brake (Shunt) Resistor","alt":"再生制动与泄放电阻（泄放模块）","abbr":"","aliases":["Regenerative Braking","Brake Resistor","Shunt Resistor","Dump Module"],"one_liner":"A motor slowing down feeds energy back like a generator; a shunt resistor burns off the excess as heat.","explanation":"When a motor decelerates or is driven backward by an external force — for instance, when a robot lands, crouches, or comes to an emergency stop — it briefly acts as a generator, feeding energy back onto the driver's DC bus; this is called regenerative braking. The bus voltage rises as a result, and if the battery or power supply can't absorb it, a shunt (or “dump”) circuit — a switching transistor plus a power resistor, also called a brake chopper — burns off the excess energy as heat once voltage crosses a threshold; otherwise the driver would trip an overvoltage fault or even be damaged. A battery-powered robot can recover some of that energy, but when the battery is already full, or the robot is powered from a lab DC supply that can't absorb reverse current, a shunt module becomes essential. It's directly tied to motor driver and power-supply design, as well as emergency-stop strategy.","example":"A legged robot connected to a lab DC power supply for jump testing has its joint motors feed energy backward on landing, spiking the bus voltage and tripping an overvoltage fault on the driver; adding a shunt module across the bus resolves the problem.","related":["Servo Drive / Motor Driver","Brushless DC Motor","Battery Management System","Emergency Stop","Holding Brake"]},{"id":"odrive","category":"hardware","sec":3,"tier":3,"sources":[{"title":"ODrive Robotics","url":"https://odriverobotics.com/"},{"title":"odriverobotics/ODrive - GitHub","url":"https://github.com/odriverobotics/ODrive"}],"as_of":"2026-09","related_ids":["servo-drive","field-oriented-control","mjbots-moteus","controller-area-network","brushless-dc-motor","rotary-encoder"],"name":"ODrive","alt":"ODrive 驱动器","abbr":"","aliases":["ODrive Robotics"],"one_liner":"A brushless motor driver line from ODrive Robotics, popular with robotics hobbyists and labs.","explanation":"ODrive is a line of brushless motor controllers from ODrive Robotics, driving motors with field-oriented control (FOC) and offering position, velocity, and torque control. The early ODrive v3.x hardware and firmware were open source under the MIT license and became very popular with robotics hobbyists and universities, though that line is no longer being developed. Today's main products are the ODrive Pro (14–58V, 3000W continuous) and the ODrive S1 (12–50V, 1600W continuous), which support CAN, UART, and step/direction interfaces and come with Python, Arduino, and ROS 2 libraries, though firmware for the newer products is no longer published openly. It's commonly used to drive hub-motor chassis and prototype joints for quadruped and humanoid robots.","example":"Many open-source quadruped and wheeled-base projects use a single ODrive board to drive two brushless motors at once, connected to the host computer over CAN.","related":["Servo Drive (Motor Driver)","Field-Oriented Control","mjbots moteus","Controller Area Network (CAN)","Brushless DC Motor","Rotary Encoder"]},{"id":"mjbots-moteus","category":"hardware","sec":3,"tier":3,"sources":[{"title":"moteus r4.11 - mjbots","url":"https://mjbots.com/products/moteus-r4-11"}],"as_of":"2026-09","related_ids":["servo-drive","field-oriented-control","controller-area-network","odrive","quasi-direct-drive","magnetic-encoder"],"name":"mjbots moteus","alt":"moteus 驱动器","abbr":"","aliases":["moteus"],"one_liner":"An open-source, palm-sized brushless motor driver from mjbots that turns an RC motor into a servo joint.","explanation":"moteus is a brushless motor controller from the American company mjbots, led by Josh Pieper: it's about the size of a palm and mounts directly behind a motor, turning an ordinary RC-style brushless motor into a servo actuator with controllable position, velocity, and torque. The current r4.11 version supports 10–44V input and 100A peak phase current, integrates an onboard absolute magnetic encoder (to measure rotor angle), and communicates over a 5 Mbps CAN-FD bus (a high-speed fieldbus) so multiple units can be daisy-chained; its firmware is open source under the Apache 2.0 license. Internally it runs field-oriented control (FOC, an algorithm that makes a brushless motor deliver torque smoothly), making it a popular choice for DIY quadrupeds, small humanoids, and robot-arm joints — a common low-cost quasi-direct-drive actuator.","example":"mjbots' own quad A1 quadruped robot drives an outrunner brushless motor at each joint with one moteus board.","related":["Servo Drive (Motor Driver)","Field-Oriented Control","Controller Area Network (CAN)","ODrive","Quasi-Direct Drive","Magnetic Encoder"]},{"id":"backdrivability","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Backdrivability","url":"https://en.wikipedia.org/wiki/Backdrivability"}],"as_of":"","related_ids":["quasi-direct-drive","reflected-inertia","gear-ratio","proprioceptive-actuator","kinesthetic-teaching","compliance"],"name":"Backdrivability","alt":"反驱性","abbr":"","aliases":[],"one_liner":"Whether an external force applied at a joint's output can easily push the whole drivetrain and motor backward.","explanation":"Backdrivability describes whether a force applied at a joint's output end can easily push back through the drivetrain and turn the motor in reverse. A joint with good backdrivability yields when a person pushes or bumps it, and the motor's current reflects the size of the external force, which is useful for force control, collision detection, and kinesthetic (hand-guided) teaching. A joint with poor backdrivability — for example, one with a large gear ratio, a worm gear, or a self-locking screw — can't be pushed back at all; it feels “stiffer” and is more prone to damage from impacts. Backdrivability is mainly determined by gear ratio, transmission efficiency, and friction: a larger gear ratio means more reflected inertia and friction at the output end, making the joint harder to backdrive. The quasi-direct-drive joints common in quadruped and humanoid robots deliberately use a smaller gear ratio to get good backdrivability.","example":"When a quasi-direct-drive leg lands and takes an impact, the ground can push it back slightly, cushioning the hit; a joint with a high gear ratio instead transmits the impact straight into the gears.","related":["Quasi-Direct Drive","Reflected Inertia","Gear Ratio","Proprioceptive Actuator","Kinesthetic Teaching","Compliance"]},{"id":"reflected-inertia","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Proprioceptive Actuator Design in the MIT Cheetah (Wensing et al., IEEE T-RO 2017)","url":"https://doi.org/10.1109/TRO.2016.2640183"}],"as_of":"","related_ids":["proprioceptive-actuator","backdrivability","gear-ratio","moment-of-inertia","quasi-direct-drive","strain-wave-gear"],"name":"Reflected Inertia","alt":"反射惯量","abbr":"","aliases":["Reflected Rotor Inertia","Equivalent Inertia"],"one_liner":"The effective inertia felt at a joint's output after the motor rotor's inertia is amplified by the reducer.","explanation":"Reflected inertia is the rotational inertia of the motor rotor (and any reducer parts on its input side) converted to an equivalent value at the joint's output, roughly equal to the rotor's inertia multiplied by the square of the gear ratio. The larger the gear ratio, the “heavier” the joint feels from the outside: an impact has to accelerate the rotor along with everything else, making the impact force larger, and it also becomes harder for an external force to push the joint back — backdrivability gets worse, and force-control bandwidth drops. This is one of the main reasons legged robots like the MIT Cheetah chose low-gear-ratio “proprioceptive actuators”: a leg needs to be able to yield on impact when it lands. Joint design often has to trade off torque density, which wants a large gear ratio, against low reflected inertia.","example":"The same motor with a 6:1 reducer has a reflected inertia about 36 times the rotor's own inertia; swap in a 100:1 harmonic drive instead, and it becomes about 10,000 times.","related":["Proprioceptive Actuator","Backdrivability","Gear Ratio","Moment of Inertia","Quasi-Direct Drive","Strain Wave Gear (Harmonic Drive)"]},{"id":"direct-drive","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Direct drive mechanism - Wikipedia","url":"https://en.wikipedia.org/wiki/Direct_drive_mechanism"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Direct Drive","alt":"直驱","abbr":"DD","aliases":["DD Motor"],"one_liner":"A motor connected straight to the joint or load, with no gearbox or other transmission in between.","explanation":"Direct drive means a motor's output shaft connects straight to the load, with no gears, gearbox, belt, or other transmission stage in between — a reduction ratio of 1. The benefits are no backlash (the play between gear teeth), low friction, and good backdrivability (an external force can easily push the joint), plus torque can be estimated directly from motor current, giving high control bandwidth and smooth, compliant motion. The downside is that the motor alone has to produce all the torque, so it has to be large and heavy, which strains torque density and generates more heat. Fully direct-drive designs in robots show up mainly in rotary tables, some robot-arm joints, and certain dexterous-hand designs; legged robots more often use the compromise of quasi-direct drive (a planetary gearbox with a small reduction ratio), which balances torque against compliance.","example":"A fully direct-drive dexterous hand places a motor directly at each finger joint, with no tendon or linkage transmission in between.","related":["Quasi-Direct Drive","Backdrivability","Backlash","Gear Ratio","Fully Direct-Drive Dexterous Hand","Frameless Torque Motor"]},{"id":"quasi-direct-drive","category":"hardware","sec":4,"tier":1,"sources":[{"title":"Proprioceptive Actuator Design in the MIT Cheetah (IEEE T-RO 2017)","url":"https://ieeexplore.ieee.org/document/7827048"},{"title":"Mini Cheetah: A Platform for Pushing the Limits of Dynamic Quadruped Control (ICRA 2019)","url":"https://ieeexplore.ieee.org/document/8793865"},{"title":"Unitree GO-M8010-6 关节电机参数（宇树官网）","url":"https://www.unitree.com/mobile/go1/motor/"}],"as_of":"","related_ids":["direct-drive","proprioceptive-actuator","backdrivability","planetary-gearbox","joint-actuator-module","mit-mini-cheetah-actuator"],"name":"Quasi-Direct Drive","alt":"准直驱","abbr":"QDD","aliases":["QDD Motor","Low-Reduction-Ratio Actuator"],"one_liner":"A joint design pairing a high-torque motor with a low gear ratio, balancing strength and compliance.","explanation":"Quasi-direct drive is a joint design approach: it pairs a large-diameter brushless motor with high torque density with just a single low-ratio gearbox stage (usually under 10:1), sitting between fully direct drive (no gearbox at all) and a high-reduction design like a harmonic drive. The low gear ratio gives good backdrivability and low reflected inertia (the motor rotor's inertia as seen at the joint, which grows with the square of the gear ratio), so an outside force can push the joint back, and impacts on landing are less likely to damage the gears; because motor current is roughly proportional to output torque, the joint's force can be estimated and controlled without a dedicated torque sensor. MIT's Cheetah series popularized this “proprioceptive actuator” approach, and today it's used in the leg joints of most quadruped and humanoid robots, including Unitree's — it also suits locomotion policies trained with reinforcement learning particularly well.","example":"All 12 joints on the MIT Mini Cheetah use the same module: a brushless motor plus a 6:1 single-stage planetary gearbox. Unitree's Go1 uses a similar design in its GO-M8010-6 joint motor, with a 6.33:1 gear ratio.","related":["Direct Drive","Proprioceptive Actuator","Backdrivability","Planetary Gearbox","Joint Actuator Module","MIT Mini Cheetah Actuator"]},{"id":"proprioceptive-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Proprioceptive Actuator Design in the MIT Cheetah (Wensing et al., IEEE T-RO 2017)","url":"https://doi.org/10.1109/TRO.2016.2640183"}],"as_of":"","related_ids":["quasi-direct-drive","backdrivability","reflected-inertia","mit-mini-cheetah-actuator","torque-control","sensorless-force-estimation"],"name":"Proprioceptive Actuator","alt":"本体感受式执行器","abbr":"","aliases":["Proprioceptive Actuation"],"one_liner":"An actuator that senses and controls joint force from motor current alone, with no dedicated force sensor.","explanation":"This is an actuator design philosophy proposed and systematically laid out by Sangbae Kim's group at MIT's Biomimetic Robotics Lab for the MIT Cheetah (Wensing et al., IEEE T-RO 2017). The approach pairs a large-diameter, high-torque motor with a single-stage, low-ratio planetary reducer, keeping transmission friction low and reflected inertia (the motor rotor's inertia as seen from the joint's output) low, so the joint can be backdriven by external force. This means motor current can fairly accurately reflect the force on the joint, enabling high-bandwidth force control without any extra force sensor, and letting the leg absorb landing impacts compliantly. Today's “quasi-direct-drive” joint modules, used across quadruped robots and many humanoid robots, trace back to this same idea.","example":"The joints of MIT's Mini Cheetah use a roughly 6:1 single-stage planetary reducer, estimating foot contact force from current alone to perform moves like backflips.","related":["Quasi-Direct Drive","Backdrivability","Reflected Inertia","MIT Mini Cheetah Actuator","Torque Control","Sensorless Force Estimation"]},{"id":"mit-mini-cheetah-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Mini Cheetah: A Platform for Pushing the Limits of Dynamic Quadruped Control (ICRA 2019)","url":"https://doi.org/10.1109/ICRA.2019.8793865"},{"title":"bgkatz/motorcontrol - GitHub","url":"https://github.com/bgkatz/motorcontrol"}],"as_of":"","related_ids":["mit-mini-cheetah","quasi-direct-drive","proprioceptive-actuator","mit-mode","planetary-gearbox","backdrivability"],"name":"MIT Mini Cheetah Actuator","alt":"MIT Cheetah 执行器","abbr":"","aliases":["Mini Cheetah Motor","Mini Cheetah Joint Module"],"one_liner":"A quasi-direct-drive joint module MIT designed for the Mini Cheetah quadruped, later widely copied across the industry.","explanation":"The MIT Cheetah actuator is a joint module designed by Ben Katz at MIT's Biomimetic Robotics Lab (Sangbae Kim's group) for the Mini Cheetah quadruped robot: a large-diameter outrunner brushless motor paired with a single-stage 6:1 planetary reducer, with an integrated driver board and magnetic encoder, communicating over CAN. The low gear ratio means low reflected inertia and the ability to be backdriven by external force, so the joint can estimate torque directly from motor current without needing a separate torque sensor — this is the core idea behind quasi-direct drive and proprioceptive actuators. Because the design files and firmware were published openly, a large number of low-cost joint motors in China and elsewhere have since copied its structure and its “MIT mode” control command format.","example":"Joint motors such as those from DAMIAO, Xiaomi's CyberGear, and the CubeMars AK series all support MIT mode, where a single command frame carries target position, velocity, Kp, Kd, and feedforward torque together.","related":["MIT Mini Cheetah","Quasi-Direct Drive","Proprioceptive Actuator","MIT Mode","Planetary Gearbox","Backdrivability"]},{"id":"cubemars-ak-series-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"AK80-9 KV100 Robotic Actuator（CubeMars 官网）","url":"https://www.cubemars.com/goods-982-AK80-9.html"},{"title":"AK80-9 V3.0 Robotic Actuator（CubeMars 官网）","url":"https://www.cubemars.com/product/ak80-9-v3-0-robotic-actuator.html"}],"as_of":"2026-09","related_ids":["quasi-direct-drive","mit-mode","mit-mini-cheetah-actuator","joint-actuator-module","damiao-dm-j4310-2ec-joint-motor","unitree-go-m8010-6-motor"],"name":"CubeMars AK Series Actuator","alt":"CubeMars AK 系列关节电机","abbr":"","aliases":["AK80-9","T-Motor AK Motor","AK Series Power Module"],"one_liner":"An all-in-one joint module from CubeMars combining a planetary reducer and a driver in one unit.","explanation":"The AK series is a line of all-in-one joint modules from CubeMars, the robotics-actuator brand of motor maker T-Motor, packing a brushless motor, a planetary reducer, an encoder, and a driver into a single housing — in a model name like AK80-9, the “9” denotes a 9:1 gear ratio. Taking the AK80-9 as an example, it has a rated torque of 9 N·m, a peak torque of 18 N·m (rated at 22 N·m in the V3.0 version), and weighs about 485 g. It supports CAN communication and a “MIT mode” that sends target position, velocity, feedforward torque, and stiffness/damping gains in a single command; its low gear ratio gives it good backdrivability, and it's commonly used in university quadruped robots, exoskeletons, and robot-arm prototypes.","example":"Sending Kp, Kd, a target position, and a feedforward torque to an AK80-9 simultaneously in MIT mode implements joint impedance control.","related":["Quasi-Direct Drive","MIT Mode","MIT Mini Cheetah Actuator","Joint Actuator Module","DAMIAO DM-J4310-2EC Joint Motor","Unitree GO-M8010-6 Motor"]},{"id":"damiao-dm-j4310-2ec-joint-motor","category":"hardware","sec":4,"tier":3,"sources":[{"title":"DM-J4310-2EC V1.1 关节电机 - 深圳市达妙科技","url":"https://www.mdmbot.com/index.php?c=show&id=84"}],"as_of":"2025","related_ids":["joint-actuator-module","damiao-technology","mit-mode","controller-area-network","dual-encoder","quasi-direct-drive"],"name":"DAMIAO DM-J4310-2EC Joint Motor","alt":"达妙 DM-J4310 关节电机","abbr":"","aliases":["DM4310","DM-J4310-2EC","Damiao Motor"],"one_liner":"A small all-in-one joint actuator from Shenzhen's DAMIAO Technology, common in desktop arms and small humanoid joints.","explanation":"The DM-J4310-2EC is a compact joint module from Shenzhen DAMIAO Technology that integrates a brushless motor, reducer, driver, and dual encoders in one unit; its output shaft reads single-turn absolute position, and it takes commands and reports back velocity, position, torque, and temperature over a CAN bus. The official V1.1 spec lists a rated torque of 3 N·m and a peak of 7 N·m, with later versions improving on this. It's inexpensive and simple to wire, and supports the common “MIT mode” (sending position, velocity, stiffness, damping, and feedforward torque all at once), so it's widely used in student projects, open-source robot arms, and lightly loaded joints like the arms and head of humanoid robots — a common “off-the-shelf part” for beginners building real robots.","example":"Building a low-cost 6-axis desktop robot arm, the three wrist joints commonly use the DM4310, with a larger DAMIAO motor swapped in for the upper arm.","related":["Joint Actuator Module","DAMIAO Technology","MIT Mode","Controller Area Network (CAN)","Dual Encoder","Quasi-Direct Drive"]},{"id":"unitree-go-m8010-6-motor","category":"hardware","sec":4,"tier":3,"sources":[{"title":"GO-M8010-6 Motor - Unitree Shop","url":"https://shop.unitree.com/products/go1-motor"},{"title":"GO-M8010-6 Motor User Manual V1.0","url":"https://techshare.co.jp/faq/wp-content/uploads/2023/12/GO-M8010-6_Motor_Data_User_Manual_V1.0.pdf"}],"as_of":"2026-09","related_ids":["joint-actuator-module","quasi-direct-drive","unitree-robotics","torque-constant","field-oriented-control","damiao-dm-j4310-2ec-joint-motor"],"name":"Unitree GO-M8010-6 Motor","alt":"宇树 GO-M8010-6 关节电机","abbr":"","aliases":["GO-M8010-6","Go1 Motor"],"one_liner":"An integrated joint-motor module Unitree Robotics sells separately, with a maximum torque of 23.7 N·m.","explanation":"The GO-M8010-6 is a joint-motor module from Unitree Robotics that integrates a permanent-magnet synchronous motor, driver board, reducer, encoder, and bearings into one unit ready to use as a robot joint. According to its user manual, it has a 6.33:1 gear ratio, a maximum output torque of 23.7 N·m, a top speed of 30 rad/s, weighs about 530 g, and is meant to run on 24 V power; it has a built-in field-oriented control (FOC) algorithm, a temperature sensor, and an absolute encoder, with a torque constant of about 0.639 N·m/A. Its low gear ratio gives it good back-drivability, which suits it to torque control in legged robots. Because it's relatively cheap and its specs are public, it's a popular joint choice for university labs and open-source small quadruped and humanoid projects.","example":"","related":["Joint Actuator Module","Quasi-Direct Drive","Unitree Robotics","Torque Constant (Kt)","Field-Oriented Control","DAMIAO DM-J4310-2EC Joint Motor"]},{"id":"xiaomi-cybergear-micro-motor","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Stirring up the motor industry, Xiaomi launches ultra-high... (SMM News)","url":"https://news.metal.com/newscontent/102352716"},{"title":"Xiaomi CyberGear Micromotor Intelligent Motor (EOL) - OpenELAB","url":"https://openelab.io/products/xiaomi-cybergear-micromotor-intelligent-motor"}],"as_of":"2026-09","related_ids":["quasi-direct-drive","joint-actuator-module","mit-mode","controller-area-network","damiao-dm-j4310-2ec-joint-motor","xiaomi-cyberdog"],"name":"Xiaomi CyberGear Micro-Motor","alt":"小米 CyberGear 微电机","abbr":"","aliases":["CyberGear"],"one_liner":"A cheap, integrated joint-motor module Xiaomi launched in 2023, with 12 N·m of peak torque.","explanation":"The CyberGear is an integrated joint-motor module Xiaomi launched in August 2023, packing a brushless motor, reducer, encoder, and driver into one module about 80 mm across and weighing 317 g. It delivers 12 N·m of peak torque and 4 N·m of continuous torque, communicates over a CAN bus, and launched at RMB 499; it's reported to be derived from the joint technology used in Xiaomi's CyberDog 2 robot dog. Like other quasi-direct-drive motors, it has a low gear ratio and is back-drivable, which suits it to torque control in legged robots and arms. It entered the market at a price far below comparable products, and many students and hobbyists have used it to build their own quadruped robots, bipedal robots, and desktop arms. Distributor listings indicate it has since been discontinued; comparable alternatives include joint motors from DAMIAO and Unitree.","example":"A hobbyist builds a small quadruped robot from 12 CyberGear motors, driving each joint by sending MIT-mode commands over the CAN bus.","related":["Quasi-Direct Drive","Joint Actuator Module","MIT Mode","Controller Area Network (CAN)","DAMIAO DM-J4310-2EC Joint Motor","Xiaomi CyberDog"]},{"id":"series-elastic-actuator","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Series elastic actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Series_elastic_actuator"}],"as_of":"","related_ids":["actuator","compliance","torque-control","quasi-direct-drive","variable-stiffness-actuator","backdrivability"],"name":"Series Elastic Actuator (SEA)","alt":"串联弹性驱动器","abbr":"SEA","aliases":["Series Elastic Drive"],"one_liner":"An actuator with a spring deliberately placed between the motor and the load, sensing force from spring deflection.","explanation":"A series elastic actuator deliberately places an elastic element (a spring) between the motor (plus reducer) and the output, an idea proposed by MIT's Gill Pratt and Matthew Williamson in 1995. Measuring how much the spring deflects gives an accurate estimate of output torque, enabling precise force control; the spring also cushions impacts and protects the reducer, making the robot safer to be around when it contacts a person. The trade-off is lower stiffness and lower bandwidth, which makes precise position control harder. SEAs are common in legged robots, exoskeletons, and collaborative robots that need compliance. The contrasting approach is quasi-direct drive, which skips the spring and estimates torque directly from motor current through a low gear ratio.","example":"The joints of Rethink Robotics' Baxter dual-arm robot use series elastic actuators, so the arm yields and gets pushed aside rather than colliding rigidly with a person.","related":["Actuator","Compliance","Torque Control","Quasi-Direct Drive","Variable Stiffness Actuator (VSA)","Backdrivability"]},{"id":"variable-stiffness-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Variable impedance actuators: A review (Robotics and Autonomous Systems, 2013)","url":"https://doi.org/10.1016/j.robot.2013.06.009"}],"as_of":"","related_ids":["series-elastic-actuator","stiffness","compliance","impedance-control","variable-impedance-control","physical-human-robot-interaction"],"name":"Variable Stiffness Actuator (VSA)","alt":"变刚度驱动器","abbr":"VSA","aliases":["VSA","Variable Stiffness Joint"],"one_liner":"An actuator whose joint “softness” can be mechanically adjusted on the fly, independent of position.","explanation":"A variable stiffness actuator adds an elastic element between the motor and the joint, plus an extra mechanism — usually a second motor — that adjusts that spring's effective stiffness in real time, so joint position and joint softness can be controlled separately. It extends the series elastic actuator, whose stiffness is fixed, and has been studied systematically since the 2000s by groups including the University of Pisa, the German Aerospace Center (DLR), and the Italian Institute of Technology. The benefits: turning stiffness down reduces impact if the joint hits a person, making it safer, while turning stiffness up gives precise positioning when that's needed; the spring can also store energy and release it for explosive motions like throwing or jumping. The downside is a more complex mechanism with higher weight and cost, so it's mainly used on research platforms today — production robots more often get similar behavior from software-based impedance control instead.","example":"DLR's Hand Arm System uses variable stiffness actuation in its arm and hand joints.","related":["Series Elastic Actuator (SEA)","Stiffness","Compliance","Impedance Control","Variable Impedance Control","Physical Human-Robot Interaction"]},{"id":"intrinsic-vs-extrinsic-actuation","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Shadow Dexterous Hand Series - Shadow Robot Company","url":"https://www.shadowrobot.com/dexterous-hand-series/"},{"title":"LEAP Hand","url":"https://leaphand.com/"}],"as_of":"","related_ids":["dexterous-hand","tendon-driven-actuation","hybrid-drive-dexterous-hand","reflected-inertia","linkage-transmission","timing-belt-drive"],"name":"Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)","alt":"驱动器内置 / 外置（近端布置）","abbr":"","aliases":["Intrinsic / Extrinsic Actuation","Proximal Actuator Placement"],"one_liner":"Whether a motor sits right at the joint, or is placed near the torso and drives the joint remotely.","explanation":"This pair of terms describes where the actuator (the motor) is placed. Intrinsic actuation means the motor is mounted right inside the part it drives — for example, motors packed into the palm and fingers of a dexterous hand. Extrinsic actuation means the motor sits somewhere closer to the torso, such as the forearm, and power is transmitted out to the fingers via tendons, linkages, or a timing belt, the same way most of the muscles that move a human hand actually sit in the forearm. Placing heavy motors closer to the body like this is called proximal placement, and it reduces the weight and rotational inertia at the far end, letting an arm or leg swing faster and more efficiently. The trade-off is a longer transmission chain, which brings friction, backlash, and tension-management problems. Both dexterous-hand design and humanoid leg and arm design have to weigh this trade-off.","example":"The Shadow Dexterous Hand keeps its motors in the forearm and pulls the fingers via tendons (extrinsic); the LEAP Hand mounts its servos directly at the finger joints (intrinsic).","related":["Dexterous Hand","Tendon-Driven Actuation","Hybrid-Drive Dexterous Hand","Reflected Inertia","Linkage Transmission","Timing Belt Drive"]},{"id":"timing-belt-drive","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Belt (mechanical) - Wikipedia","url":"https://en.wikipedia.org/wiki/Belt_(mechanical)"}],"as_of":"","related_ids":["speed-reducer-gearbox","gear-ratio","linkage-transmission","tendon-driven-actuation","intrinsic-vs-extrinsic-actuation","reflected-inertia"],"name":"Timing Belt Drive","alt":"同步带传动","abbr":"","aliases":["Belt Drive","Toothed Belt Drive"],"one_liner":"A drive that transmits rotation through a toothed belt meshing with toothed pulleys, without slipping.","explanation":"A timing belt drive has teeth molded into the inside of the belt, which mesh with toothed pulleys to carry rotation from one shaft to another. A plain flat belt transmits force through friction and can slip; a timing belt meshes tooth-to-tooth, so the speed ratio stays fixed — hence “timing.” It's light, quiet, needs no lubrication, and can provide some gear reduction for free if the two pulleys are different sizes. In robots, it's often used to move a motor away from a joint — for example, mounting it near the body to cut the inertia at the end of a leg or arm — and it also shows up in linear stages and desktop robot arms. The downside is that the belt has some elasticity, so it's less stiff than a gear train, and it stretches loose over time and needs re-tensioning.","example":"The X and Y axes of a desktop 3D printer commonly use a GT2 timing belt to move the print head.","related":["Speed Reducer / Gearbox","Gear Ratio","Linkage Transmission","Tendon-Driven Actuation","Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)","Reflected Inertia"]},{"id":"linkage-transmission","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Linkage (mechanical) - Wikipedia","url":"https://en.wikipedia.org/wiki/Linkage_(mechanical)"}],"as_of":"","related_ids":["four-bar-linkage","tendon-driven-actuation","parallel-ankle-mechanism","dexterous-hand","linear-actuator","parallel-mechanism"],"name":"Linkage Transmission","alt":"连杆传动","abbr":"","aliases":["Linkage Drive"],"one_liner":"Using rigid, hinged rods to transmit power from a motor to a joint.","explanation":"A linkage transmission connects several rigid rods with hinges (a common example is the four-bar linkage) to carry motion from a motor or electric cylinder to a joint some distance away, with the rod lengths engineered to produce the desired motion relationship. Its advantages are high stiffness and the ability to handle large forces, and it lets designers place heavy motors closer to the torso to reduce the mass at the limb's far end. Its downsides are that the range of motion is limited by the geometry of the rods, and the mechanism takes up space. On humanoid robots, linkages are common in parallel ankle and knee joints; in dexterous hands, linkage drive and tendon drive are the two dominant approaches — linkages are sturdier and easier to maintain, while tendons are lighter and more flexible.","example":"In many humanoid robots, the ankle places two motors in the shin and uses two linkages to push and pull the foot plate for pitch and roll motion.","related":["Four-Bar Linkage","Tendon-Driven Actuation","Parallel Ankle Mechanism","Dexterous Hand","Linear Actuator (Electric Cylinder)","Parallel Mechanism"]},{"id":"tendon-driven-actuation","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Shadow Dexterous Hand - Shadow Robot Company","url":"https://www.shadowrobot.com/dexterous-hand-series/"}],"as_of":"","related_ids":["dexterous-hand","shadow-dexterous-hand","bowden-cable","tendon-routing-configurations","tendon-material","linkage-transmission"],"name":"Tendon-Driven Actuation","alt":"腱绳驱动","abbr":"","aliases":["Tendon Drive","Cable-Driven Actuation"],"one_liner":"Placing the motor far from a joint and pulling it with a cable, like a tendon.","explanation":"Tendon-driven actuation imitates the tendons in a human hand: the motor sits in the forearm or palm, and a thin, high-strength cable (commonly ultra-high-molecular-weight polyethylene, UHMWPE, fiber) runs through pulleys or a sheath to pull a finger joint. The upside is that the fingers themselves stay light and slim, letting the hand approach human size with many degrees of freedom; the downside is that the cable stretches over time and has friction and changing pre-tension, making modeling and control harder and requiring regular calibration and maintenance. Because a cable can only pull, not push, a joint typically needs two cables, or one cable plus a return spring, giving rise to wiring schemes described as N-type, N+1-type, or 2N-type. The Shadow Dexterous Hand is a classic tendon-driven hand, and Tesla's Optimus hand is reported to use a similar approach.","example":"The Shadow Dexterous Hand keeps its motors in the forearm and pulls more than 20 finger joints via tendons, giving it a shape and size close to a human hand.","related":["Dexterous Hand","Shadow Dexterous Hand","Bowden Cable","Tendon Routing Configurations (N / N+1 / 2N)","Tendon Material (UHMWPE)","Linkage Transmission"]},{"id":"bowden-cable","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Bowden cable - Wikipedia","url":"https://en.wikipedia.org/wiki/Bowden_cable"}],"as_of":"","related_ids":["tendon-driven-actuation","exoskeleton","intrinsic-vs-extrinsic-actuation","dexterous-hand","friction-compensation","tendon-material"],"name":"Bowden Cable","alt":"鲍登线","abbr":"","aliases":["Bowden Sheath"],"one_liner":"A flexible cable with an inner wire in an outer sheath, transmitting pull force along a curved path, like a bike brake line.","explanation":"A Bowden cable consists of an inner wire (or cord) and an incompressible outer sheath around it; the sheath's two ends are fixed in place, and pulling the inner wire at one end moves the other end — the most familiar example is a bicycle brake cable. In robots, it's used to keep a motor in the torso, a backpack, or near the base of an arm, while sending pull force along a curved path out to a finger, wrist, or exoskeleton joint, keeping the far end lighter and lowering its inertia. Its downside is significant friction between the inner wire and sheath that varies with bend angle, plus hysteresis and elastic stretch, so precise force control needs to compensate for these effects. It's a common way of implementing tendon-driven actuation.","example":"The soft exoskeleton from Harvard's Wyss Institute keeps its motor at the waist and pulls the ankle joint via a Bowden cable to assist walking.","related":["Tendon-Driven Actuation","Exoskeleton","Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)","Dexterous Hand","Friction Compensation","Tendon Material (UHMWPE)"]},{"id":"hydraulic-actuation","category":"hardware","sec":4,"tier":2,"sources":[{"title":"Hydraulic machinery - Wikipedia","url":"https://en.wikipedia.org/wiki/Hydraulic_machinery"},{"title":"Atlas (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Atlas_(robot)"}],"as_of":"2024-04","related_ids":["boston-dynamics-atlas","electro-hydrostatic-actuator","pneumatic-actuation","actuator","power-density","boston-dynamics-bigdog"],"name":"Hydraulic Actuation","alt":"液压驱动","abbr":"","aliases":["Hydraulic Actuator","Hydraulic Drive"],"one_liner":"Driving a joint with pressurized fluid that pushes a cylinder or hydraulic motor.","explanation":"Hydraulic actuation uses a pump to pressurize hydraulic oil, which valves route to a hydraulic cylinder or hydraulic motor, converting fluid pressure into the force or torque that drives a joint. Its advantages are high power density and strong, shock-resistant force output; its downsides are the need for a pump, hoses, and valves, which make the system heavy, noisy, prone to leaks, and expensive to maintain, with more complex control as well. Early high-dynamic legged robots were mostly hydraulic — the best-known examples are Boston Dynamics' BigDog and the original hydraulic version of Atlas. As electric motors and reducers improved, humanoid robots have largely shifted to electric actuation; Boston Dynamics itself retired the hydraulic Atlas in 2024 in favor of an all-electric version. Hydraulics are still used in heavy-duty construction machinery and some large legged robots.","example":"Boston Dynamics' hydraulic Atlas used hydraulic actuation to perform parkour and backflips, before being replaced by the electric Atlas in 2024.","related":["Boston Dynamics Atlas (Hydraulic)","Electro-Hydrostatic Actuator (EHA)","Pneumatic Actuation","Actuator","Power Density (W/kg)","Boston Dynamics BigDog"]},{"id":"electro-hydrostatic-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Electro-hydraulic actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Electro-hydraulic_actuator"}],"as_of":"","related_ids":["hydraulic-actuation","actuator","linear-actuator","boston-dynamics-atlas","torque-density"],"name":"Electro-Hydrostatic Actuator (EHA)","alt":"电动静液作动器","abbr":"EHA","aliases":["Electrohydrostatic Actuator"],"one_liner":"A self-contained hydraulic actuator with its own motor-driven pump and closed oil loop, needing no external hydraulic supply.","explanation":"An electro-hydrostatic actuator packages a motor, a small bidirectional oil pump, a cylinder, and a reservoir into one self-contained module: as the motor spins one way or the other, the pump pushes oil to one side of the cylinder, extending or retracting the piston, with the oil loop closed entirely inside the module. It keeps hydraulics' advantages — high force, shock resistance — while eliminating the central pump station, long hoses, and servo valves a traditional hydraulic system needs; it only needs power and a control signal to run, and is more efficient too. EHAs were first adopted in aerospace, for aircraft control-surface actuation, and have since been used in the leg joints of heavy-load legged and humanoid robots, as a middle ground between purely electric joints and traditional hydraulics.","example":"","related":["Hydraulic Actuation","Actuator","Linear Actuator (Electric Cylinder)","Boston Dynamics Atlas (Hydraulic)","Torque Density"]},{"id":"pneumatic-actuation","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Pneumatic actuator - Wikipedia","url":"https://en.wikipedia.org/wiki/Pneumatic_actuator"}],"as_of":"","related_ids":["pneumatic-gripper","pneumatic-artificial-muscle","soft-robot","hydraulic-actuation","vacuum-suction-cup","compliance"],"name":"Pneumatic Actuation","alt":"气动驱动","abbr":"","aliases":["Pneumatic Drive"],"one_liner":"Using compressed air to push a cylinder or soft chamber and produce motion.","explanation":"Pneumatic actuation uses compressed air from an air compressor, controlled by solenoid valves, to fill and empty a pneumatic cylinder, an air motor, or a soft inflatable chamber, converting air pressure into linear or rotary motion. Its advantages are a simple, cheap structure and a high power-to-weight ratio; because air is compressible, it's naturally compliant, giving a gentle impact when it contacts a person or object, and it's safer than an electric motor in flammable or explosive environments. Its disadvantages stem from that same compressibility: precise position and force control are difficult, response is delayed, and it also needs a compressor, tubing, and valve manifolds, making it noisy, inefficient, and awkward to fit inside a mobile robot. So in factories it's mainly used for simple open/close actions, such as pneumatic grippers, vacuum cups, and positioning cylinders; in research, it's the primary power source behind soft robots and pneumatic artificial muscles.","example":"A pneumatic two-finger gripper on a production line only has two states, open and closed, switching its air path through a solenoid valve to complete a grasp in tens of milliseconds.","related":["Pneumatic Gripper","Pneumatic Artificial Muscle (McKibben Muscle)","Soft Robot","Hydraulic Actuation","Vacuum Suction Cup","Compliance"]},{"id":"artificial-muscle","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Artificial muscle - Wikipedia","url":"https://en.wikipedia.org/wiki/Artificial_muscle"}],"as_of":"","related_ids":["pneumatic-artificial-muscle","hydraulically-amplified-self-healing-electrostatic-actuator","shape-memory-alloy","dielectric-elastomer-actuator","soft-robot","clone-robotics-protoclone"],"name":"Artificial Muscle","alt":"人工肌肉","abbr":"","aliases":["Synthetic Muscle"],"one_liner":"An actuator that contracts and extends like a muscle when triggered by electricity, air pressure, or heat.","explanation":"Artificial muscle is a general term for actuators (devices that turn energy into motion) that contract, extend, or bend in response to a stimulus such as electricity, air pressure, heat, or a chemical reaction. Common categories include pneumatic artificial muscles (a rubber tube wrapped in a braided mesh that gets shorter and fatter when inflated), shape-memory alloys, dielectric elastomers, HASEL electrohydraulic actuators, and twisted-fiber actuators. Compared with a “motor plus reducer,” artificial muscle is lighter and naturally compliant, making it well suited to soft robots, exoskeletons, and biomimetic hands; its drawbacks are that efficiency, response speed, lifespan, and precise controllability generally lag behind motors, so it remains mostly in research and niche products today.","example":"Clone Robotics' Protoclone humanoid is reported to use hydraulically driven artificial muscles to move its skeletal structure.","related":["Pneumatic Artificial Muscle (McKibben Muscle)","Hydraulically Amplified Self-Healing Electrostatic Actuator (HASEL)","Shape Memory Alloy (SMA)","Dielectric Elastomer Actuator (DEA)","Soft Robot","Clone Robotics Protoclone"]},{"id":"pneumatic-artificial-muscle","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Pneumatic artificial muscles - Wikipedia","url":"https://en.wikipedia.org/wiki/Pneumatic_artificial_muscles"}],"as_of":"","related_ids":["pneumatic-actuation","artificial-muscle","exoskeleton","soft-robot","bio-inspired-robot","hydraulically-amplified-self-healing-electrostatic-actuator"],"name":"Pneumatic Artificial Muscle (McKibben Muscle)","alt":"气动人工肌肉","abbr":"PAM","aliases":["PAM","McKibben Muscle","Pneumatic Muscle"],"one_liner":"A tube-shaped actuator that shortens and thickens like a real muscle, pulling with force, when inflated.","explanation":"The classic form of pneumatic artificial muscle is the McKibben muscle: a rubber inner tube is wrapped in a braided mesh sleeve, and when the tube is inflated, it bulges outward, but the braided sleeve converts that radial expansion into axial shortening, producing a pulling force at both ends — closely resembling how a biological muscle contracts. It's named after physicist Joseph L. McKibben, who designed it in the 1950s for orthotic devices for polio patients. Its advantages are that it's very light, has a high force-to-weight ratio, and is naturally compliant; its disadvantages are that it can only pull, not push, so it needs to be arranged in antagonistic pairs like real muscles, its range of contraction is limited, and it shows significant nonlinearity and hysteresis, making precise modeling and control difficult. It's commonly found in rehabilitation exoskeletons, biomimetic robot arms, and soft-robotics research.","example":"Festo's Fluidic Muscle is a commercialized pneumatic-muscle product, used in applications such as gripping and tensioning that need compliant pulling force.","related":["Pneumatic Actuation","Artificial Muscle","Exoskeleton","Soft Robot","Bio-inspired Robot","Hydraulically Amplified Self-Healing Electrostatic Actuator (HASEL)"]},{"id":"shape-memory-alloy","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Shape-memory alloy - Wikipedia","url":"https://en.wikipedia.org/wiki/Shape-memory_alloy"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Shape Memory Alloy (SMA)","alt":"形状记忆合金","abbr":"SMA","aliases":["Memory Alloy","Nitinol"],"one_liner":"An alloy that springs back to its original shape when heated, useful as an artificial muscle when powered electrically.","explanation":"A shape memory alloy is deformed at low temperature and then returns to its original shape once heated above its transformation temperature; the most common example is Nitinol (a nickel-titanium alloy discovered in the early 1960s at the US Naval Ordnance Laboratory). The mechanism is a transition between two crystal structures, martensite and austenite. Made into a thin wire and heated electrically, it contracts by a few percent and produces a pulling force, so it can serve as an actuator or “artificial muscle.” Its advantages are a high force-to-weight ratio, a simple structure, and silent operation; its disadvantages are a small range of motion, slow recovery because it depends on passive cooling, low energy efficiency, and significant hysteresis that makes precise control difficult. It's commonly used in miniature robots, biomimetic fingers, and soft robots.","example":"A few electrically heated Nitinol wires pulling on tendons can curl a small biomimetic finger; once power is cut, it cools and a spring resets it to its original position.","related":["Artificial Muscle","Actuator","Soft Robot","Tendon-Driven Actuation","Dielectric Elastomer Actuator (DEA)"]},{"id":"dielectric-elastomer-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Dielectric elastomers - Wikipedia","url":"https://en.wikipedia.org/wiki/Dielectric_elastomers"}],"as_of":"","related_ids":["artificial-muscle","hydraulically-amplified-self-healing-electrostatic-actuator","shape-memory-alloy","soft-robot","actuator","soft-gripper"],"name":"Dielectric Elastomer Actuator (DEA)","alt":"介电弹性体","abbr":"DEA","aliases":["Dielectric Elastomer"],"one_liner":"An artificial muscle that squeezes and stretches a soft membrane using electrostatic force from a high voltage.","explanation":"A dielectric elastomer actuator is a type of electrically driven artificial muscle: a soft insulating elastic membrane (such as silicone or acrylic) is coated on both sides with stretchable electrodes, and applying a voltage of several kilovolts makes the opposite charges on each side attract each other, squeezing the membrane thinner, which makes it expand in its plane; it springs back when the power is switched off. Its advantages are that it's light, fast, quiet, and has relatively high energy density, with the ability to produce large deformations, making it suitable for soft robots, miniature grippers, and robotic fish. Its downsides are the need for kilovolt-level high voltage, a tendency to break down electrically, and relatively low output force, so it currently remains mostly confined to labs and small devices. Like the HASEL electrohydraulic actuator and shape-memory alloys, it's usually grouped under “novel actuation.”","example":"Rolling a DEA membrane into a tube that extends when powered and contracts when the power is cut turns it into a controllable soft “muscle” that can drive a small gripper.","related":["Artificial Muscle","Hydraulically Amplified Self-Healing Electrostatic Actuator (HASEL)","Shape Memory Alloy (SMA)","Soft Robot","Actuator","Soft Gripper"]},{"id":"hydraulically-amplified-self-healing-electrostatic-actuator","category":"hardware","sec":4,"tier":3,"sources":[{"title":"Acome et al., Hydraulically amplified self-healing electrostatic actuators with muscle-like performance, Science 2018","url":"https://doi.org/10.1126/science.aao6139"}],"as_of":"","related_ids":["artificial-muscle","dielectric-elastomer-actuator","soft-robot","pneumatic-artificial-muscle","actuator"],"name":"Hydraulically Amplified Self-Healing Electrostatic Actuator (HASEL)","alt":"HASEL 电液人工肌肉","abbr":"HASEL","aliases":["HASEL Actuator"],"one_liner":"A soft actuator that uses high-voltage electrostatics to squeeze a sealed liquid, mimicking muscle contraction.","explanation":"HASEL was developed by Keplinger's group at the University of Colorado Boulder and published in Science in 2018. It's a flexible pouch filled with a liquid insulating medium, with electrodes attached to the outside; applying a high voltage pulls the electrodes together, squeezing the liquid into another part of the pouch, which bulges or contracts to produce a muscle-like motion. Because the insulating layer is a liquid, if it briefly breaks down electrically at one spot, the liquid flows back in and heals the gap — hence the name “self-healing.” It combines the fast response of electrostatic actuation with the large deformation typical of hydraulics, making it suited to soft robots and biomimetic mechanisms. Its weaknesses are the need for kilovolt-level voltage, and output force and reliability that still fall short of a robot joint motor, so it remains mostly in research and small-scale applications.","example":"","related":["Artificial Muscle","Dielectric Elastomer Actuator (DEA)","Soft Robot","Pneumatic Artificial Muscle (McKibben Muscle)","Actuator"]},{"id":"end-effector","category":"hardware","sec":5,"tier":1,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"},{"title":"Franka Hand（Franka Robotics 官网）","url":"https://franka.de/franka-hand"}],"as_of":"","related_ids":["gripper","dexterous-hand","end-effector-pose","tool-center-point","tool-flange","inverse-kinematics"],"name":"End Effector","alt":"末端执行器","abbr":"EEF","aliases":["EE","End-of-Arm Tool"],"one_liner":"The tool mounted at the very tip of a robot arm that actually touches and works on objects.","explanation":"An end effector is the tool mounted on the last link of a robot arm's kinematic chain, attached through a tool flange — it's what actually makes contact with the environment: grippers, dexterous hands, vacuum suction cups, welding torches, and screwdrivers all count. The arm's job is to bring the end effector to a target position and orientation; the end effector's job is to carry out the specific action — grasping, suctioning, screwing, and so on. In embodied-AI papers, “EEF” often also refers to the end effector's position and orientation: many policies directly output an end-effector pose (an action in task space), which is then converted to joint angles through inverse kinematics, while other policies output joint angles directly. Swap in a different end effector, and both the action space and the tasks the robot can perform change with it.","example":"The end effector most often paired with a Franka arm is Franka's own Franka Hand, a two-finger parallel gripper.","related":["Gripper","Dexterous Hand","End-Effector Pose","Tool Center Point","Tool Flange","Inverse Kinematics (IK)"]},{"id":"gripper","category":"hardware","sec":5,"tier":1,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["end-effector","parallel-jaw-gripper","electric-gripper","adaptive-gripper","dexterous-hand","grasping"],"name":"Gripper","alt":"夹爪","abbr":"","aliases":["Mechanical Gripper","Robot Gripper"],"one_liner":"An end effector that grips objects by opening and closing its fingers — simpler than a dexterous hand.","explanation":"A gripper is the most common end effector, using two or three fingers that open and close to hold an object. By actuation it's electric or pneumatic; by structure it can be parallel two-finger, three-finger, adaptive/underactuated (where the fingers passively conform to an object's shape), or soft; vacuum suction cups and magnetic grippers are sometimes grouped under the broader term “gripper” too. Its advantages are that it's cheap, reliable, and easy to control, usually needing only a single value — how far open it is — to represent its action, which is why most robot-learning datasets and VLA models are built around grippers. Its limitation is that it can't do fine, multi-finger-coordinated actions like twisting, rotating, or pressing a button, which is where a dexterous hand is needed instead.","example":"The Robotiq 2F-85 is a common two-finger adaptive electric gripper in both research and industry, with a maximum opening of 85 mm.","related":["End Effector","Parallel Jaw Gripper","Electric Gripper","Adaptive (Underactuated) Gripper","Dexterous Hand","Grasping"]},{"id":"parallel-jaw-gripper","category":"hardware","sec":5,"tier":1,"sources":[{"title":"Robotiq 2F-85 & 2F-140 Adaptive Grippers","url":"https://robotiq.com/products/2f85-140-adaptive-robot-gripper"},{"title":"Franka Hand（Franka Robotics 官网，二指平行夹爪）","url":"https://franka.de/franka-hand"},{"title":"Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["gripper","end-effector","antipodal-grasp","grasping","aloha","franka-hand"],"name":"Parallel Jaw Gripper","alt":"二指夹爪","abbr":"","aliases":["Parallel Gripper","Two-Finger Gripper"],"one_liner":"A gripper whose two fingers stay parallel as they open and close — the simplest and most common design.","explanation":"A parallel jaw gripper has two opposing fingers that stay parallel to each other as they close together or move apart, much like an open-end wrench. It's mechanically simple, cheap, and grips reliably, and its action is usually described with a single number — the opening width, or simply open/closed — which is why it's the dominant end effector in robot learning: the large majority of data in datasets like Open X-Embodiment and DROID comes from parallel jaw grippers, and models like OpenVLA and π0 output a one-dimensional gripper action too. Grasping research often uses the concept of an antipodal grasp (where the contact forces at two points oppose each other) to analyze whether such a gripper can hold an object securely. Its limitation is that it struggles with tasks needing coordination across multiple fingers or in-hand adjustment.","example":"On the ALOHA bimanual platform, each arm ends in a two-finger parallel jaw gripper.","related":["Gripper","End Effector","Antipodal Grasp","Grasping","ALOHA","Franka Hand"]},{"id":"electric-gripper","category":"hardware","sec":5,"tier":2,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Electric Gripper","alt":"电动夹爪","abbr":"","aliases":[],"one_liner":"A motor-driven robot gripper whose opening position, speed, and grip force can all be precisely controlled.","explanation":"An electric gripper is an end effector mounted on a robot arm whose fingers are opened and closed by a motor, usually through a lead screw, gears, or a linkage. Compared with a pneumatic gripper, which is pushed by compressed air, it needs no air supply, can precisely set opening width, closing speed, and grip force, and can report back its current position and current draw — making it well suited to fragile objects and tasks that need feedback. Its downside is that it's generally more expensive than a pneumatic gripper and produces somewhat less force and speed. The most common choice in embodied-AI research is the two-finger electric gripper, because it has just one degree of freedom, giving a simple action space that's easy to teleoperate and learn policies for; the gripper action recorded in most datasets is just a single opening value.","example":"The Robotiq 2F-85 and Franka's stock gripper are both common two-finger electric grippers in research.","related":["Gripper","Parallel Jaw Gripper","Pneumatic Gripper","End Effector","Robotiq 2F-85 Gripper","Fingertip Force / Grip Force"]},{"id":"franka-hand","category":"hardware","sec":5,"tier":2,"sources":[{"title":"Franka Hand - Franka Robotics","url":"https://franka.de/franka-hand"},{"title":"Franka Hand Product Manual (2022)","url":"https://franka.de/hubfs/Product%20Manual%20Franka%20Hand_R50010_1.2_EN.pdf"}],"as_of":"2026-09","related_ids":["parallel-jaw-gripper","electric-gripper","franka-emika-panda-franka-research-3","end-effector","libfranka-franka-control-interface","robotiq-2f-85-gripper"],"name":"Franka Hand","alt":"Franka 夹爪","abbr":"","aliases":["Panda Gripper","Franka Emika Hand"],"one_liner":"The standard two-finger parallel electric gripper that ships with Franka robot arms.","explanation":"Franka Hand is the electric two-finger parallel gripper made by Germany's Franka Robotics (formerly Franka Emika) for its Panda and FR3 robot arms; the two fingers translate toward and away from each other to open and close. Official specs list an 80 mm stroke, 70 N continuous grip force (140 N max), and a weight of about 0.7 kg; the fingertips are swappable, and researchers often 3D-print custom fingertips for different objects. Because Franka arms are extremely common in academic labs, this gripper and its simulated model have become the default end effector (the part mounted at the arm's tip that touches objects) in a huge share of manipulation papers and benchmarks. It is controlled through libfranka, which handles opening, closing, and grip force.","example":"In simulation benchmarks like robosuite and LIBERO, the Panda arm's default gripper is modeled on the Franka Hand.","related":["Parallel Jaw Gripper","Electric Gripper","Franka Emika Panda / Franka Research 3","End Effector","libfranka / Franka Control Interface (FCI)","Robotiq 2F-85 Gripper"]},{"id":"robotiq-2f-85-gripper","category":"hardware","sec":5,"tier":2,"sources":[{"title":"Robotiq 2F-85 / 2F-140 Instruction Manual - Specifications","url":"https://assets.robotiq.com/website-assets/support_documents/document/online/2F-85_2F-140_TM_InstructionManual_HTML5_20190503.zip/2F-85_2F-140_TM_InstructionManual_HTML5/Content/6.%20Specifications.htm"},{"title":"Robotiq 2F-85 & 2F-140 Adaptive Grippers","url":"https://robotiq.com/products/2f85-140-adaptive-robot-gripper"}],"as_of":"2026-09","related_ids":["gripper","parallel-jaw-gripper","adaptive-gripper","end-effector","robotiq","droid"],"name":"Robotiq 2F-85 Gripper","alt":"Robotiq 2F-85 夹爪","abbr":"","aliases":["2F-85","Robotiq Adaptive Gripper"],"one_liner":"An electric two-finger adaptive gripper from Canada's Robotiq, with an 85 mm opening.","explanation":"The 2F-85 is an electric two-finger gripper made by the Canadian company Robotiq; the “85” in its name refers to the fingers' maximum 85 mm opening. Official specs list a settable grip force of roughly 20–235 N and a repeatability of 0.05 mm. Its fingers use a linkage mechanism that lets them either pinch together like an ordinary parallel gripper, or automatically curl around an object once they make contact — an “adaptive,” or underactuated, design, where one motor drives multiple joints. It's interface-compatible with collaborative arms like UR and Franka and works right out of the box, which has made it one of the most common end effectors in robot-learning labs; in many public datasets and policy models, the gripper dimension of the action space is simply this gripper's open/close amount.","example":"The DROID dataset's data-collection rig pairs a Franka arm with a Robotiq 2F-85 gripper.","related":["Gripper","Parallel Jaw Gripper","Adaptive (Underactuated) Gripper","End Effector","Robotiq","DROID (Distributed Robot Interaction Dataset)"]},{"id":"adaptive-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Underactuation - Wikipedia","url":"https://en.wikipedia.org/wiki/Underactuation"}],"as_of":"","related_ids":["gripper","underactuation","robotiq-2f-85-gripper","parallel-jaw-gripper","fin-ray-gripper","soft-gripper"],"name":"Adaptive (Underactuated) Gripper","alt":"自适应夹爪","abbr":"","aliases":["Underactuated Gripper"],"one_liner":"A gripper with fewer motors than joints, whose fingers automatically conform to an object's shape.","explanation":"An adaptive gripper is an underactuated design: it has fewer actuators than joints, and uses linkages, springs, or tendons to split one motor's force across multiple finger segments. Once a finger touches an object, any segment that isn't blocked keeps curling, automatically wrapping around the object without needing each joint controlled individually. This lets it both pinch in parallel like an ordinary two-finger gripper and perform enveloping grasps on cylindrical or irregular objects, all with simple control, low cost, and high tolerance for error. The trade-off is that finger-segment poses can't be controlled precisely, so it can't do in-hand manipulation. The Robotiq 2F-85 is one of the most common adaptive grippers in both research and industry.","example":"The Robotiq 2F-85 uses just one motor, yet can both pinch up a thin, flat part in parallel and let its segments curl to envelop a cup.","related":["Gripper","Underactuation","Robotiq 2F-85 Gripper","Parallel Jaw Gripper","Fin Ray Gripper","Soft Gripper"]},{"id":"three-finger-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["gripper","parallel-jaw-gripper","adaptive-gripper","dexterous-hand","end-effector","grasping"],"name":"Three-Finger Gripper","alt":"三指夹爪","abbr":"","aliases":["Three-Fingered Gripper","Adaptive Three-Finger Gripper"],"one_liner":"A robotic gripper with three fingers, sitting between a two-finger gripper and a five-fingered dexterous hand.","explanation":"A three-finger gripper is an end effector mounted on a robot arm that picks things up with three fingers. A common design lets two fingers rotate around the palm to change their layout, facing the third finger; that lets it pinch like a two-finger gripper, or wrap around a cylinder or a ball using all three fingers together. Compared with a two-finger gripper, it has more contact points, grips more securely, and adapts to a wider range of shapes; compared with a five-fingered dexterous hand, it uses fewer motors, has a simpler mechanism, and holds up better over time. Many three-finger grippers are underactuated (fewer motors than joints), so a finger automatically curls to conform once it touches an object. They're used in both research and industry; well-known examples include the BarrettHand and the Robotiq 3-Finger Adaptive Gripper.","example":"BarrettHand: three fingers, two of which can rotate around the palm — a common choice in early grasp-planning research.","related":["Gripper","Parallel Jaw Gripper","Adaptive (Underactuated) Gripper","Dexterous Hand","End Effector","Grasping"]},{"id":"pneumatic-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["gripper","electric-gripper","end-effector","pneumatic-actuation","vacuum-suction-cup","parallel-jaw-gripper"],"name":"Pneumatic Gripper","alt":"气动夹爪","abbr":"","aliases":["Pneumatic Finger Gripper"],"one_liner":"A gripper whose fingers open and close via compressed air pushing a piston.","explanation":"A pneumatic gripper uses compressed air to push a piston inside a cylinder, which drives the fingers open or closed through a linkage or rack-and-pinion mechanism, with a solenoid valve switching the air path to control “open” and “close.” It's mechanically simple, cheap, fast-acting, strong-gripping, and durable, making it one of the most common end effectors on factory production lines. Its downside is that it usually only has two positions, fully open and fully closed, so grip force and opening width are hard to control precisely, and it also needs an air compressor, tubing, and solenoid valves. Embodied AI research more often uses electric grippers, because a policy needs to output a continuous opening width; but in high-throughput material handling and sorting, pneumatic grippers remain the mainstream choice.","example":"A two-finger pneumatic gripper picking parts off an assembly line: a PLC sends a signal to a solenoid valve, the air path switches, and the fingers close to grip the part.","related":["Gripper","Electric Gripper","End Effector","Pneumatic Actuation","Vacuum Suction Cup","Parallel Jaw Gripper"]},{"id":"vacuum-suction-cup","category":"hardware","sec":5,"tier":2,"sources":[{"title":"Suction cup - Wikipedia","url":"https://en.wikipedia.org/wiki/Suction_cup"}],"as_of":"","related_ids":["end-effector","gripper","bin-picking","sorting","order-picking","palletizing-depalletizing"],"name":"Vacuum Suction Cup","alt":"真空吸盘","abbr":"","aliases":["Suction Cup","Vacuum Gripper"],"one_liner":"An end effector that holds an object using negative pressure from air pulled out of the cup.","explanation":"The vacuum suction cup is one of the most common end effectors in industry: once the cup presses against an object's surface, a vacuum pump or vacuum generator draws the air out from inside it, and the pressure difference between inside and outside holds the object against the cup. It's mechanically simple, cheap, and needs little positioning precision — as long as it can reach a reasonably flat, airtight surface, it can pick the object up — so it's widely used in logistics sorting, e-commerce order picking, and unstructured (bin) picking. Its weaknesses are that it doesn't work well on porous, rough, soft, or very small objects, and it can't perform operations that need fingers working together. Many picking robots combine a suction cup with a gripper.","example":"Picking arms in Amazon warehouses commonly use suction cups to lift boxes and plastic-bag-wrapped items out of bins.","related":["End Effector","Gripper","Bin Picking","Sorting","Order Picking","Palletizing / Depalletizing"]},{"id":"magnetic-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Electropermanent magnet - Wikipedia","url":"https://en.wikipedia.org/wiki/Electropermanent_magnet"},{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["end-effector","vacuum-suction-cup","gripper","electroadhesion-gripper","industrial-robot","palletizing-depalletizing"],"name":"Magnetic Gripper","alt":"磁力夹爪","abbr":"","aliases":["Electromagnetic Chuck","Electro-Permanent Magnetic Gripper"],"one_liner":"An end effector that holds steel workpieces using an electromagnet or a switchable permanent magnet.","explanation":"A magnetic gripper is a type of end effector (the part mounted at a robot arm's tip that does the work) that holds ferrous workpieces using an electromagnet, or an electro-permanent magnet, which switches its magnetism with a brief pulse of current, then holds its grip with no power draw at all. It only needs to touch one face of the workpiece, with no fingers to close, making it fast to use and mechanically simple — well suited to moving steel sheets, stamped parts, and other iron-based parts; the electro-permanent version keeps holding even if power is lost, which is safer. Its limits are just as clear: it only works on ferromagnetic materials, thin sheets can get picked up several at a time by accident, and residual magnetism can be left behind after release. Like the vacuum suction cup and the two-finger gripper, it's a common non-dexterous grasping solution in industrial settings.","example":"On a stamping line, a magnetic chuck at the end of a robot arm lifts one steel sheet at a time and feeds it into the next press.","related":["End Effector","Vacuum Suction Cup","Gripper","Electroadhesion Gripper","Industrial Robot","Palletizing / Depalletizing"]},{"id":"electroadhesion-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Electroadhesion - Wikipedia","url":"https://en.wikipedia.org/wiki/Electroadhesion"}],"as_of":"","related_ids":["gripper","soft-gripper","vacuum-suction-cup","deformable-object-manipulation","end-effector"],"name":"Electroadhesion Gripper","alt":"静电吸附夹爪","abbr":"","aliases":["Electrostatic Adhesion Gripper"],"one_liner":"A gripper that “sticks” to objects using electrostatic attraction from high voltage across electrodes in a flexible pad.","explanation":"An electroadhesion gripper works on the principle of electroadhesion: interleaved electrodes are embedded beneath a soft gripper surface, and applying a high voltage induces an opposite charge on the object's surface, creating an electrostatic attraction between the two that holds the object against the pad and lifts it; the attraction disappears as soon as the power is cut. It needs neither a clamping motion nor a vacuum, making it friendly to cloth, paper, film, fragile items, or irregularly shaped objects, and it uses very little power. Its limitation is that the attraction force varies a lot with material, humidity, and surface condition, and it can't hold much weight, so it's often combined with mechanical fingers — the fingers handle enveloping the object, while electroadhesion adds friction and conformal contact.","example":"Picking up a piece of cloth or paper lying flat on a table is hard for an ordinary two-finger gripper to slide under, but an electroadhesion pad can just be pressed onto it and powered up to lift it.","related":["Gripper","Soft Gripper","Vacuum Suction Cup","Deformable Object Manipulation","End Effector"]},{"id":"soft-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Soft robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Soft_robotics"}],"as_of":"","related_ids":["gripper","soft-robot","pneumatic-actuation","fin-ray-gripper","granular-jamming-gripper","compliance"],"name":"Soft Gripper","alt":"软体夹爪","abbr":"","aliases":["Flexible Gripper","Soft Robotic Hand"],"one_liner":"A gripper with soft fingers, often silicone, that wraps around an object by deforming instead of gripping rigidly.","explanation":"A soft gripper's fingers are made from elastic materials such as silicone or rubber. Most are pneumatically actuated: pumping air into or out of internal chambers makes the fingers curl; some designs instead pull the fingers with tendons or small motors. Because the fingers are themselves soft, they deform to match an object's surface on contact, so the gripper doesn't need to know the object's exact shape or position in advance, and it's unlikely to crush anything fragile. That makes soft grippers the most common real-world form of soft robotics, widely used to pick and sort food, fruit, and other fresh produce. The downsides are limited grip force and positioning accuracy, and the material ages with use. Like fin-ray grippers and granular jamming grippers, soft grippers rely on passive structural compliance to adapt to what they touch, rather than active sensing and control.","example":"Food-packing lines use pneumatic soft grippers to pick up fragile items like eggs and strawberries.","related":["Gripper","Soft Robot","Pneumatic Actuation","Fin Ray Gripper","Granular Jamming Gripper (Universal Gripper)","Compliance"]},{"id":"fin-ray-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Soft robotics（含 Fin Ray 抓手介绍）- Wikipedia","url":"https://en.wikipedia.org/wiki/Soft_robotics"}],"as_of":"","related_ids":["soft-gripper","parallel-jaw-gripper","adaptive-gripper","compliance","gripper"],"name":"Fin Ray Gripper","alt":"鳍条夹爪","abbr":"","aliases":["Fin Ray Effect Gripper"],"one_liner":"A compliant gripper, inspired by fish fins, whose fingers bend around an object automatically under side pressure.","explanation":"A Fin Ray finger is a triangular frame whose two outer edges are connected by a series of crosswise ribs, a structure inspired by fish fins (the Fin Ray Effect). When an object presses against the finger from the side, the finger doesn't just get squashed and pushed away — it bends toward the object instead, automatically conforming to its contour. This means a single, ordinary two-finger open/close motion is enough to gently wrap around fruit, bottles, and other irregularly shaped objects over a fairly large contact area, with no extra sensors or complex control needed. It's usually made from soft plastic or 3D-printed, is cheap, and is often used as a drop-in fingertip replacement on a two-finger gripper — a passively compliant type of soft gripper.","example":"Swapping a parallel gripper's rigid fingertips for two 3D-printed Fin Ray fingers lets it hold a tomato firmly without crushing it.","related":["Soft Gripper","Parallel Jaw Gripper","Adaptive (Underactuated) Gripper","Compliance","Gripper"]},{"id":"granular-jamming-gripper","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Jamming (physics) - Wikipedia","url":"https://en.wikipedia.org/wiki/Jamming_(physics)"}],"as_of":"","related_ids":["soft-gripper","gripper","vacuum-suction-cup","soft-robot","compliance"],"name":"Granular Jamming Gripper (Universal Gripper)","alt":"颗粒阻塞夹爪","abbr":"","aliases":["Universal Gripper","Jamming Gripper","Coffee-Balloon Gripper"],"one_liner":"A soft bag of granules that molds around an object, then hardens under vacuum to lock it in place.","explanation":"A granular jamming gripper consists of a soft balloon-like bag filled with granular material, such as coffee grounds or sand. To grasp, the soft, loose bag is first pressed onto the object, molding itself around the object's shape; then air is drawn out of the bag, the grains pack tightly together and stop flowing past each other (a physical transition called jamming), and the bag instantly stiffens, gripping the object through a combination of friction, suction, and geometric interlocking; releasing it is as simple as letting air back in. It was proposed by researchers at the University of Chicago, Cornell University, and iRobot in a 2010 PNAS paper, and earned the name “universal gripper” because it needs no prior knowledge of an object's shape and can grasp small objects of almost any shape. Its weakness is that it performs poorly on large, flat objects.","example":"The same coffee-grounds balloon gripper can pick up a coin, an egg, and a screwdriver without ever changing fingers.","related":["Soft Gripper","Gripper","Vacuum Suction Cup","Soft Robot","Compliance"]},{"id":"tool-flange","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Robot end effector - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["tool-center-point","end-effector","tool-changer","six-axis-force-torque-sensor","robotic-arm"],"name":"Tool Flange","alt":"工具法兰","abbr":"","aliases":["End Flange","Wrist Flange","Robot Flange"],"one_liner":"The standardized round mounting interface at a robot arm's tip, where grippers and other tools attach.","explanation":"The tool flange is the round mounting face at the end of a robot arm's last joint, with threaded holes and dowel-pin holes arranged in a standard pattern; grippers, force sensors, camera mounts, and tool changers all attach here. The international standard ISO 9409-1 defines the flange's dimensions and hole layout, so a tool built to that standard can be reused across arms from different brands; a tool that doesn't follow it needs an adapter plate. The flange center is also usually the origin of the arm's default end-effector coordinate frame — once a tool is attached, the controller needs its tool center point (TCP) offset set correctly so the arm knows exactly where its working point is. Some collaborative arms also route power and communication lines near the flange, making it easy to power an electric gripper.","example":"Swapping a gripper onto a UR5e: first bolt the gripper to the flange using the ISO 9409-1 hole pattern, then set the TCP offset on the teach pendant.","related":["Tool Center Point","End Effector","Tool Changer","Six-Axis Force/Torque Sensor","Robotic Arm"]},{"id":"tool-changer","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Robotic Tool Changers - ATI Industrial Automation","url":"https://www.ati-ia.com/products/toolchanger/robot_tool_changer.aspx"}],"as_of":"","related_ids":["tool-flange","end-effector","gripper","vacuum-suction-cup","tool-use","robotic-arm"],"name":"Tool Changer","alt":"工具快换盘","abbr":"","aliases":["Quick-Change Coupler","Quick-Change Adapter","Automatic Tool Changer (ATC)"],"one_liner":"A coupling that lets a robot arm automatically swap end-of-arm tools within seconds.","explanation":"A tool changer mounts between a robot arm's flange and its end-of-arm tool, and comes in two halves: a master side that stays fixed to the arm, and a tool side mounted on each individual tool. To swap tools, the arm aligns the master side with a tool side, and a pneumatic or electric locking mechanism clamps them together while connecting air lines, electrical signals, and communication — completing the swap in seconds. It solves the basic limit of one arm, one gripper: the same arm can pick up boxes with a vacuum cup one moment, switch to a two-finger gripper for assembly the next, then swap in a screwdriver after that. It's common on industrial production lines; in embodied AI, it also shows up on mobile-manipulation platforms and test rigs that need several different tools. Major makers include ATI and Schunk.","example":"The arm first picks up material with a vacuum cup, docks it on the tool rack, then swaps in an electric gripper to finish assembly — all without a human touching any hardware.","related":["Tool Flange","End Effector","Gripper","Vacuum Suction Cup","Tool Use","Robotic Arm"]},{"id":"remote-center-compliance-device","category":"hardware","sec":5,"tier":3,"sources":[{"title":"Remote center compliance - Wikipedia","url":"https://en.wikipedia.org/wiki/Remote_center_compliance"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Remote Center Compliance (RCC) Device","alt":"远中心柔顺装置（RCC）","abbr":"RCC","aliases":["RCC"],"one_liner":"A passive spring mechanism between wrist and gripper that lets a part self-align into a hole during insertion.","explanation":"RCC is a passive compliance mechanism proposed in the late 1970s by Draper Laboratory (Whitney and colleagues), built from elastic elements and mounted between a robot arm's wrist and its gripper. Its defining feature is that its center of compliance sits outside the device itself, near the tip of the part being held, so the moment the part's edge touches the side of a hole during insertion, the part automatically shifts sideways and tilts to align itself, rather than jamming. It needs no force sensor and no active compliance control (a control algorithm that regulates contact force) — the mechanical structure alone absorbs a few millimeters of position error — and it's still used in industrial assembly today. In learning-based assembly research, it's commonly compared against active compliance control.","example":"When a robot arm inserts a pin into a tightly toleranced hole, adding an RCC device between the flange and the gripper lets a slightly misaligned pin slide in instead of jamming at the opening.","related":["Peg-in-Hole Insertion","Compliance","Compliance Control","Impedance Control","Robotic Assembly"]},{"id":"dexterous-hand","category":"hardware","sec":6,"tier":1,"sources":[{"title":"Shadow Hand - Wikipedia","url":"https://en.wikipedia.org/wiki/Shadow_Hand"},{"title":"OpenAI 2019: Solving Rubik's Cube with a Robot Hand (arXiv 1910.07113)","url":"https://arxiv.org/abs/1910.07113"}],"as_of":"","related_ids":["dexterous-manipulation","in-hand-manipulation","end-effector","active-dof-passive-dof","shadow-dexterous-hand","tendon-driven-actuation"],"name":"Dexterous Hand","alt":"灵巧手","abbr":"","aliases":["Multi-fingered Hand","Five-fingered Hand"],"one_liner":"A robotic hand with multiple independently moving fingers, capable of fine manipulation.","explanation":"A dexterous hand is a multi-fingered end effector mounted on a robot arm, built to mimic the structure of a human hand, typically with 3 to 5 fingers and a dozen to two dozen joints. Compared with a gripper, which can only open and close, it can pinch, grip, twirl a pen, twist open a bottle cap, and reposition an object within the hand itself (in-hand manipulation) — making it a key component for general-purpose manipulation on humanoid robots. The hard part is packing many degrees of freedom into very little space: motors, gearboxes, and sensors all have to fit inside the palm and fingers, usually driven through tendons, linkages, or miniature lead screws; controlling it is also harder because its action space has many dimensions, and collecting training data for it is difficult. Dexterous hands are usually specified by “active degrees of freedom” (the joints actually driven by a motor), which is often fewer than the total number of joints.","example":"OpenAI's Dactyl project in 2019 used a Shadow Dexterous Hand to solve a Rubik's Cube one-handed.","related":["Dexterous Manipulation","In-hand Manipulation","End Effector","Active DoF / Passive DoF","Shadow Dexterous Hand","Tendon-Driven Actuation"]},{"id":"metacarpophalangeal-proximal-and-distal-interphalangeal-join","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Metacarpophalangeal joint - Wikipedia","url":"https://en.wikipedia.org/wiki/Metacarpophalangeal_joint"},{"title":"Interphalangeal joints of the hand - Wikipedia","url":"https://en.wikipedia.org/wiki/Interphalangeal_joints_of_the_hand"}],"as_of":"","related_ids":["dexterous-hand","active-dof-passive-dof","thumb-opposition","flexion-extension-and-abduction-adduction","underactuation","mano"],"name":"Metacarpophalangeal / Proximal & Distal Interphalangeal Joints (MCP / PIP / DIP)","alt":"掌指关节 / 指间关节（MCP / PIP / DIP）","abbr":"MCP / PIP / DIP","aliases":["MCP Joint","PIP Joint","DIP Joint"],"one_liner":"The three joints from the base of a finger to its tip, standard vocabulary for describing dexterous-hand structure.","explanation":"These three abbreviations come from human hand anatomy. MCP is the metacarpophalangeal joint, where the finger meets the palm; it can both flex/extend and move side to side, giving it 2 degrees of freedom. PIP is the middle joint, the proximal interphalangeal joint, and DIP is the joint nearest the fingertip, the distal interphalangeal joint; each of these can only flex and extend. When a person curls a finger, the DIP joint mostly follows the PIP joint's motion automatically, so many dexterous hands couple the two together with a linkage or tendon, letting one motor drive both joints — which means the hand ends up with more joints than actuated degrees of freedom (motors). The thumb's structure is different and has no separate PIP/DIP. When reading a dexterous hand's specs, distinguishing joint count from actuated degrees of freedom is necessary to compare products fairly.","example":"Each finger on the Inspire RH56 dexterous hand has only one motor, which drives both the MCP and PIP joints to flex together through a linkage — so the whole hand has 12 joints but only 6 actuated degrees of freedom.","related":["Dexterous Hand","Active DoF / Passive DoF","Thumb Opposition","Flexion/Extension and Abduction/Adduction","Underactuation","MANO"]},{"id":"thumb-opposition","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Opposable thumb - Wikipedia","url":"https://en.wikipedia.org/wiki/Opposable_thumb"},{"title":"Unitree Dex5-1 - Unitree Robotics","url":"https://www.unitree.com/mobile/Dex5-1/"}],"as_of":"","related_ids":["dexterous-hand","grasp-taxonomy","metacarpophalangeal-proximal-and-distal-interphalangeal-join","dexterous-manipulation","degrees-of-freedom"],"name":"Thumb Opposition","alt":"拇指对掌","abbr":"","aliases":["Opposition Movement"],"one_liner":"The motion of rotating the thumb so its pad faces the pads of the other fingers.","explanation":"Thumb opposition is a basic human hand motion: the thumb rotates and draws inward at its base joint so its pad faces the pads of the index and middle fingers. Picking up a coin, twisting off a bottle cap, and holding a pen all depend on it. For a robotic dexterous hand, whether it can achieve opposition determines whether it can do fine pinching; a thumb that only flexes and can't rotate is limited to a crude wrap-around grip. That's why a humanoid hand's thumb base usually needs at least two degrees of freedom — one for flexion, one for rotation — and higher-end hands give the thumb even more actively driven degrees of freedom. A common way to evaluate a dexterous hand is checking whether its thumb can touch each of the other four fingertips individually.","example":"The Unitree Dex5-1 dexterous hand's thumb has 4 actively driven degrees of freedom, enough for opposition pinching.","related":["Dexterous Hand","Grasp Taxonomy (Power Grasp vs. Precision Grasp / Pinch)","Metacarpophalangeal / Proximal & Distal Interphalangeal Joints (MCP / PIP / DIP)","Dexterous Manipulation","Degrees of Freedom (DoF)"]},{"id":"fingertip-force-grip-force","category":"hardware","sec":6,"tier":2,"sources":[{"title":"Grip strength - Wikipedia","url":"https://en.wikipedia.org/wiki/Grip_strength"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Fingertip Force / Grip Force","alt":"指尖力 / 握力","abbr":"","aliases":[],"one_liner":"How much force a dexterous hand or gripper can apply — fingertip force is per finger, grip force is the whole hand closing.","explanation":"Fingertip force is the maximum force a single fingertip can apply to an object, which determines whether fine actions like pinching, pressing a button, or twisting a small object are possible. Grip force is the total clamping force when the whole hand (or both fingers of a gripper) closes around an object, which determines how much weight it can lift and how securely it can hold on. Both are usually given in newtons (N) or kilogram-force, and they're core specs on any dexterous-hand or gripper datasheet. A bigger number isn't automatically better: too much force without fine control can crush an object, while too little lets it slip, so force-control precision, response speed, and heat buildup during sustained use all matter too. When comparing numbers, it's also worth checking whether a manufacturer is reporting a peak or a sustained value, and what finger posture it was measured in.","example":"A dexterous-hand spec sheet usually lists both single-finger fingertip force and whole-hand grip force side by side, and the two numbers can differ by several times.","related":["Dexterous Hand","Gripper","Grasping","Friction Cone","Force Closure","Grasp Taxonomy (Power vs. Precision Grasp)"]},{"id":"tendon-routing-configurations","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Design and Control of a Tendon-Driven Robotic Finger Based on Grasping Task Analysis (PMC)","url":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11201696/"},{"title":"Wire Driven Multi-fingered Hand - Springer Nature Link","url":"https://link.springer.com/rwe/10.1007/978-94-007-6046-2_84"}],"as_of":"","related_ids":["tendon-driven-actuation","dexterous-hand","tendon-material","underactuation","degrees-of-freedom","intrinsic-vs-extrinsic-actuation"],"name":"Tendon Routing Configurations (N / N+1 / 2N)","alt":"腱绳驱动配置（N 型 / N+1 型 / 2N 型）","abbr":"","aliases":["N-Type / N+1-Type / 2N-Type Tendon Layout"],"one_liner":"Three standard ways to route tendons across N joints, named by how many cords each configuration uses.","explanation":"A tendon can only pull, never push, so one cord can only rotate a joint in one direction. There are three common ways to actuate N degrees of freedom with tendons. A 2N configuration gives every joint an antagonistic pair of tendons — one flexes it, one extends it — which gives the most direct control and even lets you tune joint stiffness, at the cost of needing twice as many motors. An N+1 configuration shares N+1 tendons across N joints, routing them so they couple across joints in a way that keeps every degree of freedom controllable while every tendon stays in tension; it uses the fewest motors. An N configuration gives each joint just one pulling tendon and relies on a spring for the return stroke, which is mechanically simple but limits return force and speed to whatever the spring provides. The right choice depends on space inside the hand, motor count, and how much force control is needed.","example":"The Stanford/JPL three-fingered hand drives each finger's 3 degrees of freedom with 4 tendons (an N+1 configuration); the Utah/MIT hand uses a 2N configuration.","related":["Tendon-Driven Actuation","Dexterous Hand","Tendon Material (UHMWPE)","Underactuation","Degrees of Freedom (DoF)","Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)"]},{"id":"tendon-material","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Ultra-high-molecular-weight polyethylene - Wikipedia","url":"https://en.wikipedia.org/wiki/Ultra-high-molecular-weight_polyethylene"},{"title":"Dyneema - Wikipedia","url":"https://en.wikipedia.org/wiki/Dyneema"}],"as_of":"","related_ids":["tendon-driven-actuation","tendon-routing-configurations","dexterous-hand","bowden-cable","tesla-optimus-hand"],"name":"Tendon Material (UHMWPE)","alt":"腱绳材料（超高分子量聚乙烯纤维 UHMWPE）","abbr":"UHMWPE","aliases":["Ultra-High-Molecular-Weight Polyethylene Fiber","UHMWPE Fiber","Dyneema","Spectra","Braided PE Line"],"one_liner":"The high-strength, low-stretch, abrasion-resistant synthetic cord commonly used as tendons in tendon-driven dexterous hands.","explanation":"Ultra-high-molecular-weight polyethylene fiber (UHMWPE, sold under brand names like Dyneema and Spectra) is spun from polyethylene chains that are extremely long, which gives it far more strength per unit weight than steel wire, along with low stretch, a low friction coefficient, and good resistance to repeated bending. A tendon-driven dexterous hand has to carry pulling force from motors in the forearm, through pulleys or guide tubes at the wrist and finger joints, all the way to the fingertip: if the cord stretches, finger position becomes inaccurate, and if it frays and snaps, the hand stops working — so designers reach for this braided fiber. Its weak points are poor heat resistance and slow creep (gradual stretching) under sustained load, so designs usually add pre-tensioning and tension compensation to work around that.","example":"Many open-source tendon-driven hands simply use braided PE fishing line — the same material sold for high-strength fishing line — as the tendon.","related":["Tendon-Driven Actuation","Tendon Routing Configurations (N / N+1 / 2N)","Dexterous Hand","Bowden Cable","Tesla Optimus Hand"]},{"id":"fully-direct-drive-dexterous-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"直驱传动，扛起最强灵巧手的大旗 - 腾讯新闻","url":"https://news.qq.com/rain/a/20251201A04RU300"}],"as_of":"2025-12","related_ids":["dexterous-hand","direct-drive","tendon-driven-actuation","linkage-transmission","hybrid-drive-dexterous-hand","backdrivability"],"name":"Fully Direct-Drive Dexterous Hand","alt":"全直驱灵巧手","abbr":"","aliases":["Direct-Drive Dexterous Hand"],"one_liner":"A dexterous hand where a motor mounted right at each active joint drives it directly, with no tendons or linkages.","explanation":"A fully direct-drive dexterous hand gives every actuated degree of freedom its own miniature motor, placed right next to that joint, keeping the transmission path from motor to joint as short as possible — unlike tendon drive, which keeps the motor in the forearm and pulls a cable, or a linkage design, where one motor drives several joints at once. The benefits are fast response, low friction and backlash, backdrivability (an external force can push the joint back), the ability to estimate force from motor current, and support for a high number of degrees of freedom. The challenge is that the space inside a finger is extremely tight, yet has to fit a motor, encoder, and drive electronics, which makes the whole hand heavier, concentrates heat generation, and demands a lot from micro-motor manufacturing. Several Chinese companies have introduced products like this since 2025.","example":"Dexterous-hand maker Wujitech's Wuji Hand (20 DOF) and Sharpa Wave (22 DOF) are both reported to use a direct-drive design.","related":["Dexterous Hand","Direct Drive","Tendon-Driven Actuation","Linkage Transmission","Hybrid-Drive Dexterous Hand","Backdrivability"]},{"id":"hybrid-drive-dexterous-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Robotic hand - Wikipedia","url":"https://en.wikipedia.org/wiki/Robotic_hand"}],"as_of":"","related_ids":["dexterous-hand","fully-direct-drive-dexterous-hand","tendon-driven-actuation","linkage-transmission","intrinsic-vs-extrinsic-actuation"],"name":"Hybrid-Drive Dexterous Hand","alt":"混合驱动灵巧手","abbr":"","aliases":["Hybrid-Drive Hand"],"one_liner":"A dexterous hand that combines more than one actuation and transmission method within a single hand.","explanation":"“Hybrid-drive dexterous hand” is common terminology in China's robotics industry for a hand that doesn't rely on a single actuation scheme, but splits the work across joints: for example, some joints are driven directly by tiny motors in the palm, or through a linkage, while others are pulled by tendons from motors in the forearm; some designs also combine motor actuation with passive elastic elements. The goal is to balance several conflicting demands at once: the hand must stay small and light, have many degrees of freedom, produce enough fingertip force, and remain easy to maintain. A purely tendon-driven hand is dexterous but hard to keep tensioned and prone to wear; a purely linkage-driven or fully direct-drive hand is mechanically reliable but struggles to fit motors inside the fingers. Hybrid drive is a compromise, and the exact mix differs by manufacturer, so it's worth asking exactly which transmission each joint uses when evaluating a product.","example":"Driving finger flexion via tendons pulled by a forearm motor, while driving finger abduction directly with a tiny motor in the palm, is one example of a hybrid-drive scheme.","related":["Dexterous Hand","Fully Direct-Drive Dexterous Hand","Tendon-Driven Actuation","Linkage Transmission","Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)"]},{"id":"allegro-hand","category":"hardware","sec":6,"tier":2,"sources":[{"title":"Allegro Hand 官网（Wonik Robotics）","url":"https://www.allegrohand.com/"},{"title":"Allegro Hand v4.0 - Wonikrobotics Wiki","url":"http://wiki.wonikrobotics.com/AllegroHandWiki/index.php/Allegro_Hand_v4.0"}],"as_of":"2026-09","related_ids":["dexterous-hand","leap-hand","shadow-dexterous-hand","in-hand-manipulation","hora","dexterous-manipulation"],"name":"Allegro Hand","alt":"Allegro 灵巧手","abbr":"","aliases":["Allegro"],"one_liner":"A South Korean four-fingered, 16-DOF dexterous hand from Wonik Robotics, a staple in academic research.","explanation":"The Allegro Hand is a commercial dexterous hand made by South Korea's Wonik Robotics: four fingers with four joints each, for 16 degrees of freedom total, with each joint driven by its own motor and controllable by current (torque), communicating over a CAN bus at about 333 Hz. It costs far less than a Shadow Dexterous Hand, and with ready-made ROS drivers and simulation models, it has long been a standard platform in academia for in-hand manipulation, dexterous grasping, and sim-to-real transfer research. Its limitation is having only four fingers — one fewer than a human hand — and relatively low torque per joint. According to the manufacturer, the newer V5 generation adds tactile sensors to the fingertips.","example":"The HORA method trains a policy with reinforcement learning in simulation and transfers it to the real robot, letting an Allegro Hand continuously rotate a variety of objects in its grasp using only proprioception.","related":["Dexterous Hand","LEAP Hand","Shadow Dexterous Hand","In-hand Manipulation","HORA","Dexterous Manipulation"]},{"id":"leap-hand","category":"hardware","sec":6,"tier":2,"sources":[{"title":"LEAP Hand: Low-Cost, Efficient, and Anthropomorphic Hand for Robot Learning (arXiv 2309.06440)","url":"https://arxiv.org/abs/2309.06440"},{"title":"LEAP Hand V2 Advanced project page","url":"https://v2-adv.leaphand.com/"}],"as_of":"2025-06","related_ids":["dexterous-hand","allegro-hand","robotis-dynamixel-servo","open-source-hardware","motion-retargeting","cmu-robotics-institute"],"name":"LEAP Hand","alt":"LEAP Hand","abbr":"","aliases":["Low-Cost, Efficient, Anthropomorphic Hand","LEAP Hand v2"],"one_liner":"An open-source, low-cost, 16-DOF anthropomorphic dexterous hand from Carnegie Mellon.","explanation":"LEAP Hand was introduced by Kenneth Shaw, Ananye Agarwal, and Deepak Pathak at Carnegie Mellon University, presented at RSS 2023. It's a four-fingered, humanlike hand with 16 degrees of freedom, with each joint driven directly by Dynamixel servo motors; all parts are off-the-shelf or 3D-printed, and the paper reports it can be built for about $2,000 in around 4 hours — roughly an eighth the cost of the Allegro Hand — with a knuckle design meant to stay dexterous across the full range of finger poses. The hardware, simulation model, and software are all open source, making it easy for researchers to use for teleoperation, learning from human videos, and sim-to-real transfer. The team later released a v2 series, including a version with a hybrid rigid-soft shell.","example":"Researchers have used a webcam to estimate human hand pose and retarget it onto a LEAP Hand for teleoperated data collection.","related":["Dexterous Hand","Allegro Hand","ROBOTIS Dynamixel Servo","Open-Source Hardware (OSHW)","Motion Retargeting","CMU Robotics Institute"]},{"id":"orca-open-source-reliable-cost-effective-anthropomorphic-rob","category":"hardware","sec":6,"tier":3,"sources":[{"title":"The ORCA Hand - Soft Robotics Lab, ETH Zurich","url":"https://srl.ethz.ch/orcahand.html"},{"title":"arXiv 2504.04259 ORCA: An Open-Source, Reliable, Cost-Effective, Anthropomorphic Robotic Hand","url":"https://arxiv.org/abs/2504.04259"},{"title":"arXiv 2606.14561 ORCA: A Platform for Open-Source Dexterity Research","url":"https://arxiv.org/abs/2606.14561"}],"as_of":"2026-06","related_ids":["dexterous-hand","tendon-driven-actuation","leap-hand","tactile-sensor","open-source-hardware","lerobot"],"name":"ORCA: Open-Source, Reliable, Cost-Effective, Anthropomorphic Robotic Hand","alt":"ORCA 灵巧手","abbr":"ORCA","aliases":["ORCA","ORCA Hand"],"one_liner":"A 17-DOF, open-source, tendon-driven anthropomorphic dexterous hand from ETH Zurich.","explanation":"ORCA is an open-source anthropomorphic dexterous hand from ETH Zurich's Soft Robotics Lab, with its paper posted in April 2025 and later published at IROS 2025. It has 17 degrees of freedom (16 in the fingers, 1 in the wrist), uses tendon drive (motors sit outside the hand and pull the fingers via cables, like tendons), and integrates tactile sensors; its parts can be 3D-printed or bought off the shelf, with material costs under CHF 2,000 and an assembly time of under 8 hours. Its design includes joints that can dislocate and reset without damage, automatic calibration, and automatic tendon tensioning, and in testing it ran continuously for more than 10,000 cycles (about 20 hours) with no hardware failure. In June 2026, the team published a follow-up platform paper connecting low-level control, simulation, teleoperation from consumer devices, and hand motion retargeting, and integrating it with the LeRobot training pipeline.","example":"A researcher can 3D-print and assemble their own ORCA hand using the STL files and assembly video on the project's website, then teleoperate it with a VR headset to collect data and train a dexterous-manipulation policy.","related":["Dexterous Hand","Tendon-Driven Actuation","LEAP Hand","Tactile Sensor","Open-Source Hardware (OSHW)","LeRobot"]},{"id":"shadow-dexterous-hand","category":"hardware","sec":6,"tier":2,"sources":[{"title":"Shadow Dexterous Hand Series - Shadow Robot","url":"https://shadowrobot.com/dexterous-hand-series/"},{"title":"Shadow Hand - Wikipedia","url":"https://en.wikipedia.org/wiki/Shadow_Hand"}],"as_of":"2026-09","related_ids":["dexterous-hand","shadow-robot-company","dactyl","tendon-driven-actuation","dex-ee","in-hand-manipulation"],"name":"Shadow Dexterous Hand","alt":"Shadow 灵巧手","abbr":"","aliases":["Shadow Hand"],"one_liner":"A highly humanlike dexterous hand from the UK's Shadow Robot, with 24 joints and 20 actuated degrees of freedom.","explanation":"The Shadow Dexterous Hand was developed by London-based Shadow Robot Company and is the company's flagship product. Designed to match the size and joint range of a human hand, it has 24 joints in total, 20 of which are independently actuated, with the rest coupled through underactuation. The electric version places its motors in the forearm and pulls the fingers via tendons, and includes extensive position and tactile sensing. It's expensive and complex to maintain, but has long served as the benchmark platform for dexterous-manipulation research, and many simulators include a built-in model of it. The company later partnered with Google DeepMind to produce a more durable hand, the DEX-EE.","example":"OpenAI's Dactyl project trained a reinforcement-learning policy in simulation, then used it to solve a Rubik's Cube one-handed on a real Shadow Hand.","related":["Dexterous Hand","Shadow Robot Company","Dactyl","Tendon-Driven Actuation","DEX-EE (Shadow Robot × Google DeepMind)","In-hand Manipulation"]},{"id":"dex-ee","category":"hardware","sec":6,"tier":3,"sources":[{"title":"DEX-EE Series - Shadow Robot","url":"https://shadowrobot.com/dex-ee_series/"},{"title":"Our latest advances in robot dexterity - Google DeepMind","url":"https://deepmind.google/blog/advances-in-robot-dexterity/"}],"as_of":"2026-09","related_ids":["dexterous-hand","shadow-robot-company","shadow-dexterous-hand","google-deepmind","tactile-sensor","dexterous-manipulation"],"name":"DEX-EE (Shadow Robot × Google DeepMind)","alt":"Shadow DEX-EE 灵巧手","abbr":"","aliases":["DEX-EE","DEX-EE Chiral"],"one_liner":"A rugged, three-fingered research dexterous hand co-developed by Shadow Robot and Google DeepMind.","explanation":"DEX-EE is a dexterous hand developed by the UK's Shadow Robot Company over roughly five years of iteration, at the request of Google DeepMind's robotics team. It doesn't try to look like a human hand: it has three fingers, each with 4 degrees of freedom for 12 total, is about half again as large as a human hand, and weighs about 4.1 kg. The design priorities are durability and sensitivity, since reinforcement learning needs to try and fail repeatedly on the real hardware, and an ordinary dexterous hand breaks too easily under that kind of abuse. Each fingertip carries an optical tactile sensor with hundreds of sensing points, the knuckles also carry multi-dimensional tactile sensing, and position, force, and IMU data are all streamed back at high speed. A later version, DEX-EE Chiral, moves the third finger to the side so it opposes the other fingers like a human thumb, making it easier for human teleoperation and imitation learning; it's offered in left-hand, right-hand, and two-handed sets.","example":"DeepMind uses DEX-EE as a platform for real-world learning research, letting a policy grasp and manipulate objects repeatedly on the real hardware to collect data.","related":["Dexterous Hand","Shadow Robot Company","Shadow Dexterous Hand","Google DeepMind","Tactile Sensor","Dexterous Manipulation"]},{"id":"inspire-rh56-dexterous-hand","category":"hardware","sec":6,"tier":2,"sources":[{"title":"RH56DFX - INSPIRE ROBOTS","url":"https://en.inspire-robots.com/product/rh56dfx/"},{"title":"The Dexterous Hand RH56 Series User Manual","url":"https://en.inspire-robots.com/wp-content/uploads/2024/02/INSPIRE-ROBOTS-THE-DEXTEROUS-HAND-RH56-SERIES-USER-MANUAL.pdf"}],"as_of":"2026-09","related_ids":["dexterous-hand","inspire-robots","linear-actuator","linkage-transmission","unitree-h1","open-television"],"name":"Inspire RH56 Dexterous Hand","alt":"因时 RH56 灵巧手","abbr":"","aliases":["Inspire Hand","RH56DFX"],"one_liner":"A five-fingered dexterous hand from Inspire Robotics, with 6 motors driving 12 joints.","explanation":"The RH56 is a line of five-fingered, humanlike dexterous hands made by Beijing-based Inspire Robotics; the RH56DFX is the commonly used model. It has 6 actuated degrees of freedom driving 12 joints in total: each finger is curled by a built-in miniature linear servo actuator (a small electric cylinder) acting through a linkage, the thumb has an extra side-to-side motion, and the remaining joints are coupled mechanically through the linkages. Official specs list a per-finger fingertip force of about 10 N (about 15 N for the thumb), built-in position and force feedback, self-locking when powered off, and RS485 or CAN communication. It is moderately priced and reliable, and is widely mounted on humanoid robots such as Unitree's H1 and G1 for teleoperation and imitation learning, making it one of the most common dexterous hands in labs both in China and abroad.","example":"Open-TeleVision uses an Apple Vision Pro to teleoperate a Unitree H1 fitted with Inspire dexterous hands to collect data.","related":["Dexterous Hand","Inspire Robots","Linear Actuator (Electric Cylinder)","Linkage Transmission","Unitree H1","Open-TeleVision"]},{"id":"brainco-revo-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"BrainCo Revo 2 参数（官方文档）","url":"https://www.brainco-hz.com/docs/revolimb-hand/en/revo2/parameters.html"},{"title":"Revo 2 | BrainCo","url":"https://brainco.tech/product/revo2"}],"as_of":"2026-09","related_ids":["dexterous-hand","brainco","underactuation","active-dof-passive-dof","inspire-rh56-dexterous-hand","unitree-dex3-1"],"name":"BrainCo Revo Hand","alt":"强脑科技 Revo 灵巧手","abbr":"","aliases":["Revo 2","Revo 2 Dexterous Hand"],"one_liner":"A lightweight bionic dexterous hand from BrainCo; the Revo 2 has 6 motors, 11 degrees of freedom, and weighs about 383 g.","explanation":"Revo is the robotic dexterous-hand line from BrainCo, a Hangzhou-based brain-computer-interface company, built on its smart-prosthetics technology. According to its official documentation, the Revo 2 uses 6 motors to drive 6 actuated joints, plus passively coupled joints, for 11 degrees of freedom in total: the thumb has actuated flexion/extension and abduction/adduction, each of the other four fingers has one actuated flexion/extension joint, and the remaining joints are mechanically coupled (underactuated). A single hand weighs about 383 g, delivers a full-hand grip force of at least 50 N, and can carry a maximum load of at least 20 kg. It comes in Basic, Pro, and Touch versions, with the Touch version adding a multi-dimensional tactile module; interfaces include RS485 and CAN FD, with the Pro/Touch versions also supporting EtherCAT. It's light and uses common interfaces, and is often mounted on humanoid robots for grasping and teleoperation experiments.","example":"A Revo 2 version fitted for Unitree humanoid robots is sold through distributor channels.","related":["Dexterous Hand","BrainCo","Underactuation","Active DoF / Passive DoF","Inspire RH56 Dexterous Hand","Unitree Dex3-1"]},{"id":"psyonic-ability-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Ability Hand - PSYONIC","url":"https://www.psyonic.io/ability-hand"},{"title":"How a Loyola alum built the world's first touch-sensing bionic hand","url":"https://news.luc.edu/stories/science-tech/how-a-loyola-alum-built-the-worlds-first-touch-sensing-bionic-hand/"}],"as_of":"2026-09","related_ids":["dexterous-hand","tactile-sensor","dexterous-manipulation","teleoperation","degrees-of-freedom","inspire-rh56-dexterous-hand"],"name":"PSYONIC Ability Hand","alt":"PSYONIC Ability Hand","abbr":"","aliases":["Ability Hand"],"one_liner":"A tactile bionic prosthetic hand from US company PSYONIC, also commonly used as a robotics dexterous hand.","explanation":"The Ability Hand is a bionic hand made by PSYONIC, a company headquartered in San Diego and founded by Aadeel Akhtar. It was designed first as a prosthesis for upper-limb amputees, and the company describes it as the first bionic hand with tactile feedback: pressure sensors in the fingertips relay the sensation of touch back to the user's residual limb through vibration when grasping something. All five fingers can flex and extend, the thumb can rotate electrically, giving 6 actuated degrees of freedom in total, and it weighs about 490 grams with an impact-resistant shell. Because it's light, fast, durable, and offers interfaces aimed at researchers, a number of robotics teams mount it on robot arms or humanoid robots and use it as a dexterous hand for dexterous-manipulation and teleoperation research.","example":"Researchers mount the Ability Hand at the end of a robot arm and use a data glove to teleoperate it, collecting grasping demonstrations to train an imitation-learning policy.","related":["Dexterous Hand","Tactile Sensor","Dexterous Manipulation","Teleoperation","Degrees of Freedom (DoF)","Inspire RH56 Dexterous Hand"]},{"id":"unitree-dex3-1","category":"hardware","sec":6,"tier":2,"sources":[{"title":"Unitree Dex3-1 - Unitree Robotics","url":"https://www.unitree.com/mobile/Dex3-1/"}],"as_of":"2026-09","related_ids":["unitree-g1","unitree-robotics","dexterous-hand","unitree-dex5-1","tactile-sensor","motion-retargeting"],"name":"Unitree Dex3-1","alt":"宇树 Dex3-1 灵巧手","abbr":"","aliases":["Unitree Dex3","Dex3"],"one_liner":"Unitree's three-fingered, force-controlled dexterous hand for the G1 humanoid, with 7 degrees of freedom.","explanation":"The Dex3-1 is a three-fingered dexterous hand from Unitree Robotics, designed mainly to pair with Unitree's G1 humanoid robot. It has 7 actuated degrees of freedom: 3 in the thumb, and 2 each in the index and middle fingers. Six of the joints are direct-driven by miniature brushless motors, and one uses gear transmission, with force control supported throughout. The tactile version has 33 pressure sensors per hand, spread across the palm, fingertips, and finger pads. Rather than mimicking all five human fingers, it trades finger count for a simpler, more reliable structure, and can still perform common operations like grasping and pinching; it's often paired with the G1 in research settings for teleoperated data collection and training manipulation policies.","example":"Many teleoperation and VLA experiments built on the Unitree G1 use an Apple Vision Pro to capture human hand motion, then retarget it onto the Dex3-1 to perform grasps.","related":["Unitree G1","Unitree Robotics","Dexterous Hand","Unitree Dex5-1","Tactile Sensor","Motion Retargeting"]},{"id":"unitree-dex5-1","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Unitree Dex5-1 - Unitree Robotics","url":"https://www.unitree.com/mobile/Dex5-1/"},{"title":"Unitree releases new Dex5-1 humanoid robot hand - Robotics 24/7","url":"https://www.robotics247.com/article/unitree-releases-new-dex5-1-humanoid-robot-hand"}],"as_of":"2026-09","related_ids":["dexterous-hand","unitree-dex3-1","unitree-robotics","tactile-sensor","thumb-opposition","coreless-motor"],"name":"Unitree Dex5-1","alt":"宇树 Dex5-1 灵巧手","abbr":"","aliases":["Dex5","Dex5-1P"],"one_liner":"A 20-degree-of-freedom five-fingered dexterous hand from Unitree Robotics, available in a tactile-sensing version.","explanation":"The Dex5-1 is the five-fingered dexterous hand Unitree Robotics built to go with its own humanoid robots. It has 20 degrees of freedom in total, 16 actively driven and 4 passive: the thumb has 4 actively driven degrees of freedom, and each of the other four fingers has 3 actively driven plus 1 passive. Its joints are driven by coreless motors paired with low-backlash reducers; Unitree states that each joint is back-drivable (an external force can push the joint), which helps with force control, and gives fingertip repeatability of about ±1 mm. The “P” variant, Dex5-1P, adds 94 tactile sensors across the palm, fingertips, and finger segments. According to the product page, it can be mounted on Unitree humanoid robots such as the H1 and H1-2, and researchers have also used it for dexterous manipulation and teleoperation data collection.","example":"","related":["Dexterous Hand","Unitree Dex3-1","Unitree Robotics","Tactile Sensor","Thumb Opposition","Coreless Motor"]},{"id":"agibot-omnihand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"智元机器人发布 OmniHand 2025 灵巧手，9800 元起 - IT之家","url":"https://www.ithome.com/0/875/978.htm"},{"title":"OmniHand 专业款2025 - 智元灵巧手","url":"https://www.zhiyuan-robot.com/DOCS/OS/Omnihand-O12"}],"as_of":"2025-08","related_ids":["agibot","dexterous-hand","tactile-sensor","inspire-rh56-dexterous-hand","unitree-dex3-1","agibot-lingxi-x2"],"name":"AgiBot OmniHand (OmniHand 2025 / OmniHand Pro 2025)","alt":"智元 OmniHand 灵巧手","abbr":"","aliases":["OmniHand 2025","OmniHand Lite","OmniHand Pro"],"one_liner":"A five-fingered dexterous-hand line launched by AgiBot in 2025, in Lite and Pro versions.","explanation":"OmniHand is AgiBot's dexterous-hand product line, with the OmniHand 2025 series launched on August 17, 2025. The Lite version targets interactive service tasks: 16 degrees of freedom, about 500 g in weight, with an optional 400-plus tactile sensing points covering the palm, back of the hand, and all five fingers, launched at a limited-time price of RMB 9,800 (list price RMB 14,800). The Pro version targets industrial work: 19 degrees of freedom, about 750 g in weight, up to 20 N of force per finger, and multimodal tactile sensing. It can be mounted on AgiBot's own humanoid robots or sold separately to other robot makers and researchers, and is a representative example of Chinese-made dexterous hands competing on price to gain volume.","example":"","related":["AgiBot","Dexterous Hand","Tactile Sensor","Inspire RH56 Dexterous Hand","Unitree Dex3-1","AgiBot Lingxi X2"]},{"id":"robotera-xhand1","category":"hardware","sec":6,"tier":3,"sources":[{"title":"IT之家：星动纪元机器人灵巧手 XHAND1 亮相","url":"https://www.ithome.com/0/811/850.htm"},{"title":"新浪科技：星动纪元发布星动 XHAND 1 PRO","url":"https://finance.sina.com.cn/tech/discovery/2026-06-17/doc-inictaet3890715.shtml"}],"as_of":"2026-06","related_ids":[null,null,null,null,null,null],"name":"RobotEra XHAND1","alt":"星动纪元 XHAND1 灵巧手","abbr":"","aliases":["XHAND","XHAND1"],"one_liner":"RobotEra's 12-DOF, fully electric, five-fingered dexterous hand with tactile arrays on the fingertips.","explanation":"The XHAND1 is a five-fingered dexterous hand released by RobotEra in November 2024, serving as the hand end effector for its STAR1 humanoid robot. It has 12 actuated degrees of freedom (3 each in the thumb and index finger, 2 each in the remaining three fingers, with the index finger also able to move side to side), is fully electrically driven, and gives every degree of freedom its own independent actuator that can be controlled separately. Each finger carries a tactile array sensor with more than 100 sensing points, able to sense 3D force and temperature, and the company states a maximum single-hand grip force of about 80 N. In June 2026, RobotEra released an upgraded version, the XHAND1 PRO, which is fully direct-drive with 21 degrees of freedom and 18 distributed tactile sensors across the hand.","example":"In an official demo, a single XHAND1 hand lifted a 25 kg dumbbell.","related":["Dexterous Hand","RobotEra","RobotEra STAR1","Fully Direct-Drive Dexterous Hand","Tactile Sensor","Dexterous Manipulation"]},{"id":"wuji-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"产品介绍 - 文档中心 - 舞肌科技","url":"https://docs.wuji.tech/docs/zh/wuji-hand/latest/overview/"},{"title":"舞肌科技 官方网站","url":"https://www.wuji.tech/en/"}],"as_of":"2026-09","related_ids":["dexterous-hand","fully-direct-drive-dexterous-hand","dexterous-manipulation","backdrivability","data-glove","in-hand-manipulation"],"name":"Wuji Hand","alt":"舞肌 Wuji Hand","abbr":"","aliases":["Wuji Hand 2"],"one_liner":"A fully direct-drive dexterous hand with 20 actively driven degrees of freedom, made by Shenzhen-based Wuji Tech.","explanation":"The Wuji Hand is a humanoid dexterous hand made by WUJI TECH (Shenzhen Wuji Technology Co., Ltd.), with 4 degrees of freedom per finger for 20 actively driven degrees of freedom in total. Its defining feature is being fully direct-drive: small motors are mounted directly inside each finger to drive its joints, rather than coupling several joints together through tendons or linkages the way many dexterous hands do. That means every joint can be controlled independently and is back-drivable (an external force can push the joint), which helps with force control and safe contact. Company documentation lists 20-axis control at 1,000 Hz over Ethernet. The first generation reportedly sold for about RMB 50,000, and official demos have shown it spinning a pen, using scissors, and picking up a ball with chopsticks. A second generation, the Wuji Hand 2, was reportedly shown at ICRA 2026. It's often paired with a data glove for teleoperated data collection in dexterous-manipulation research.","example":"A researcher wears a data glove to teleoperate the Wuji Hand through spinning a pen and clicking a lighter; the collected data is used to train dexterous-manipulation policies.","related":["Dexterous Hand","Fully Direct-Drive Dexterous Hand","Dexterous Manipulation","Backdrivability","Data Glove","In-hand Manipulation"]},{"id":"linker-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"灵心巧手 - 全球领先的机器人灵巧手 | LinkerBot","url":"https://www.linkerbot.cn/"},{"title":"灵心巧手（Linkerbot） - RobotScope","url":"https://robotscope.net/companies/linkerbot/"}],"as_of":"2026-09","related_ids":["linkerbot","dexterous-hand","tendon-driven-actuation","linkage-transmission","dexterous-manipulation","data-glove"],"name":"Linker Hand","alt":"灵心巧手 Linker Hand","abbr":"","aliases":["Linkerbot Dexterous Hand"],"one_liner":"A dexterous-hand product line from Beijing's Linkerbot, spanning tendon-driven, linkage, and direct-drive designs.","explanation":"Linker Hand is the dexterous-hand product line from Linkerbot (灵心巧手), a Beijing-based company. According to the company, the line spans anywhere from 6 to 42 degrees of freedom and covers three technical approaches: tendon drive, linkage transmission, and direct drive. The L20 model, for instance, is linkage-based with 21 degrees of freedom, driven by brushless motors paired with precision ball screws; the L30 is tendon-driven, with 22 degrees of freedom. A dexterous hand is the end effector a humanoid robot uses for fine manipulation — more degrees of freedom bring it closer to a human hand, but also make control and data collection harder. Products like this are commonly bought by universities and robotics companies to mount on a robot arm or humanoid, paired with a data glove or VR teleoperation to collect dexterous-manipulation data.","example":"Researchers mount a Linker Hand at the end of a robot arm and use a data glove to teleoperate it, collecting demonstrations of dexterous tasks like grasping or twisting a bottle cap, then use that data to train a VLA policy.","related":["Linkerbot","Dexterous Hand","Tendon-Driven Actuation","Linkage Transmission","Dexterous Manipulation","Data Glove"]},{"id":"sharpawave","category":"hardware","sec":6,"tier":3,"sources":[{"title":"PR Newswire: Sharpa Reaches Key Milestone With Mass Production","url":"https://www.prnewswire.com/news-releases/ai-robotmaker-sharpa-reaches-key-milestone-with-mass-production-of-worlds-most-advanced-human-sized-robotic-hand-302643434.html"},{"title":"Sharpa Wave 官网","url":"https://www.sharpa.com/pages/wave"},{"title":"CNX Software: Sharpa Wave 22 DoF dexterous hand","url":"https://www.cnx-software.com/2026/06/02/sharpa-wave-high-end-dexterous-robotic-hand-with-22-dof-high-sensitivity-dynamic-tactile-array/"}],"as_of":"2025-12","related_ids":[null,null,null,null,null,null],"name":"SharpaWave","alt":"Sharpa Wave 灵巧手","abbr":"","aliases":["Sharpa Wave","Sharpa W01","Sharpa W02"],"one_liner":"Sharpa's human-hand-sized, 22-DOF dexterous hand with high-density vision-based tactile sensing on each fingertip.","explanation":"SharpaWave is the flagship dexterous hand from Sharpa, a company founded in 2024, headquartered in Singapore with R&D and manufacturing in Shanghai. The hand matches an adult human hand 1:1 in size and has 22 actuated degrees of freedom; each fingertip integrates a miniature camera plus more than 1,000 tactile pixels, which the company calls a “Dynamic Tactile Array” (DTA), with a force resolution of 0.005 N and fingertip output force above 20 N, and joints that are backdrivable (an external force can push them, helping with impact resistance). It's aimed at embodied AI research and whole-robot integration, and ships with a ROS 2 package, a MuJoCo simulation model, and C++/Python interfaces. The company announced mass production on December 16, 2025.","example":"A researcher can first train a dexterous-manipulation policy in MuJoCo using the official SharpaWave model, then transfer it to the real hand.","related":["Dexterous Hand","Vision-Based Tactile Sensor","Dexterous Manipulation","In-Hand Manipulation","Backdrivability","GR-Dexter (ByteDance Seed)"]},{"id":"tesla-optimus-hand","category":"hardware","sec":6,"tier":3,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"title":"The Forearm Is the New Hand: Inside Tesla's Optimus V3 Patents","url":"https://droids.substack.com/p/the-forearm-is-the-new-hand-inside"}],"as_of":"2026-09","related_ids":["dexterous-hand","tendon-driven-actuation","tesla-optimus","tesla-optimus-v3","intrinsic-vs-extrinsic-actuation","tesla-supply-chain"],"name":"Tesla Optimus Hand","alt":"特斯拉 Optimus 灵巧手","abbr":"","aliases":["Optimus Hand"],"one_liner":"Tesla's in-house dexterous hand for its Optimus humanoid robot; the third generation has 22 degrees of freedom.","explanation":"The Tesla Optimus hand is the dexterous hand Tesla designed in-house for its Optimus humanoid robot. According to Wikipedia, the second-generation Optimus hand, shown in 2023, had 11 degrees of freedom; the third generation, announced in 2024, has 22. An international patent application made public in April 2026 shows a redesigned hand that moves most of the motors into the forearm, with each finger driven by three thin tendons routed through the wrist — cutting the hand's own weight and the impact load it carries — and uses a routing structure at the wrist to reduce crosstalk between wrist rotation and finger movement. As of September 2026, the complete third-generation Optimus has not been formally released, so production specifications should be treated as provisional until Tesla confirms them.","example":"According to the public patent, each finger is pulled by three tendons routed from actuators in the forearm, with almost no motors inside the palm or fingers themselves.","related":["Dexterous Hand","Tendon-Driven Actuation","Tesla Optimus","Tesla Optimus V3 (Optimus Gen 3)","Intrinsic vs. Extrinsic Actuation (Proximal Actuator Placement)","Tesla (Optimus) Supply Chain"]},{"id":"mobile-base","category":"hardware","sec":7,"tier":2,"sources":[{"title":"Mobile ALOHA 项目主页","url":"https://mobile-aloha.github.io/"},{"title":"Mecanum wheel - Wikipedia","url":"https://en.wikipedia.org/wiki/Mecanum_wheel"}],"as_of":"","related_ids":["differential-drive-base","mecanum-wheel","swerve-drive","mobile-manipulator","wheeled-humanoid-robot","mobile-manipulation"],"name":"Mobile Base (Chassis)","alt":"移动底盘","abbr":"","aliases":["AGV Base","Mobile Platform"],"one_liner":"The wheeled lower half of a robot that carries the upper body around.","explanation":"A mobile base is the part of a robot responsible for locomotion; it typically integrates drive wheels and motors, a battery, a lidar or camera, an IMU (inertial measurement unit), and a motion controller, exposing a simple “give it a velocity and it moves” interface (in ROS, usually a cmd_vel velocity command). Bases are commonly classified by wheel layout: differential-drive bases (two wheels turning at different speeds to steer, mechanically simple), Mecanum-wheel or omnidirectional bases (can move sideways), swerve-drive bases (each wheel steers independently, good for heavy loads), and Ackermann bases (front-wheel steering, like a car). Mounting a robot arm or a humanoid upper body onto a mobile base turns it into a composite robot or a wheeled humanoid; the base's navigation accuracy and battery life directly affect how well mobile manipulation works.","example":"Mobile ALOHA mounts the ALOHA dual-arm teleoperation system on a mobile base, letting the robot move around a home while doing two-handed tasks like stir-frying or opening cabinets.","related":["Differential Drive Base","Mecanum Wheel","Swerve Drive","Mobile Manipulator","Wheeled Humanoid Robot","Mobile Manipulation"]},{"id":"differential-drive-base","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Differential wheeled robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Differential_wheeled_robot"}],"as_of":"","related_ids":["mobile-base","differential-drive-kinematics","nonholonomic-constraint","mecanum-wheel","swerve-drive","wheel-odometry"],"name":"Differential Drive Base","alt":"差速底盘","abbr":"","aliases":["Differential Drive","Two-Wheel Differential Base"],"one_liner":"A mobile base with two independently driven wheels that steers by varying the speed of each side.","explanation":"A differential drive base is the most common mobile-robot chassis: an independently motor-driven wheel on the left and right, plus one or two caster wheels for support. Driving both wheels at the same speed goes straight, different speeds turn, and opposite directions spin the robot in place. It's mechanically simple, cheap, and easy to control — its kinematics need only two parameters, wheel speed and wheel separation, to compute the base's linear and angular velocity — so it's used in robot vacuums, the TurtleBot, and many warehouse AMRs. Its downside is that it can't translate sideways (it's a nonholonomic system) and is less agile than a Mecanum-wheel or swerve-drive base at repositioning in tight spaces. In ROS, it's typically commanded directly with linear and angular velocity over the cmd_vel topic.","example":"A robot vacuum is a textbook differential drive base: two drive wheels plus one front caster, able to spin in place but not move sideways.","related":["Mobile Base (Chassis)","Differential Drive Kinematics","Nonholonomic Constraint","Mecanum Wheel","Swerve Drive","Wheel Odometry"]},{"id":"hub-motor","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Wheel hub motor - Wikipedia","url":"https://en.wikipedia.org/wiki/Wheel_hub_motor"}],"as_of":"","related_ids":["mobile-base","differential-drive-base","wheel-legged-robot","outrunner-motor","direct-drive"],"name":"Hub Motor (In-Wheel Motor)","alt":"轮毂电机","abbr":"","aliases":["In-Wheel Motor"],"one_liner":"A drive motor built directly into a wheel's hub, turning the wheel itself.","explanation":"A hub motor packs the entire motor inside the wheel, usually as an outrunner design: the outer shell turns together with the tire, eliminating the need for a drive shaft, gearbox, or belt. It's compact, has low transmission losses, and lets each wheel be controlled independently, so it's common in e-bikes and self-balancing scooters, and also used in mobile-robot chassis and the foot-wheels of wheeled-legged robots. Its downsides are that it makes the wheel heavier (adding unsprung mass, which hurts suspension performance), demands more from motor cooling and waterproofing, and often still needs a reducer for low-speed, high-torque situations. For robots, it simplifies chassis design, since differential steering just means controlling the left and right wheel speeds independently.","example":"Many DIY mobile robots repurpose hub motors salvaged from self-balancing scooters as chassis drive wheels.","related":["Mobile Base (Chassis)","Differential Drive Base","Wheel-legged Robot","Outrunner Motor","Direct Drive"]},{"id":"ackermann-steering-chassis","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Ackermann steering geometry - Wikipedia","url":"https://en.wikipedia.org/wiki/Ackermann_steering_geometry"}],"as_of":"","related_ids":["mobile-base","differential-drive-base","nonholonomic-constraint","hybrid-a-star","pure-pursuit","autonomous-driving"],"name":"Ackermann Steering Chassis","alt":"阿克曼底盘","abbr":"","aliases":["Ackermann Chassis"],"one_liner":"A mobile base that turns by angling its front wheels, like a car.","explanation":"An Ackermann chassis uses the same steering geometry as a car: when turning, the inner front wheel angles more sharply than the outer one, so all the wheels roll around a common center point, reducing tire scrub. It suits higher speeds and outdoor surfaces, and its structure matches a car's, so it's common in self-driving test vehicles, campus delivery robots, and teaching robot cars. Its downside is that it can't turn in place and has a minimum turning radius — it's a nonholonomic system (it can't move directly sideways) — so path planning has to account for this, often using algorithms like hybrid A* or pure pursuit. Indoor mobile-manipulation robots more often use a differential-drive or omnidirectional-wheel base instead.","example":"Many self-driving teaching cars and ROS robot-car kits use an Ackermann chassis, tracking a planned path with a pure-pursuit algorithm.","related":["Mobile Base (Chassis)","Differential Drive Base","Nonholonomic Constraint","Hybrid A*","Pure Pursuit","Autonomous Driving"]},{"id":"mecanum-wheel","category":"hardware","sec":7,"tier":2,"sources":[{"title":"Mecanum wheel - Wikipedia","url":"https://en.wikipedia.org/wiki/Mecanum_wheel"}],"as_of":"","related_ids":["omni-wheel","mobile-base","swerve-drive","differential-drive-base","automated-guided-vehicle","mobile-manipulator"],"name":"Mecanum Wheel","alt":"麦克纳姆轮","abbr":"","aliases":["Mecanum Drive"],"one_liner":"An omnidirectional wheel ringed with angled rollers that lets a base drive sideways.","explanation":"The Mecanum wheel was invented in the early 1970s by Bengt Ilon, an engineer at the Swedish company Mecanum AB. Its rim carries a ring of small rollers mounted at 45° to the wheel's axle, so when the wheel turns, the ground pushes back on it at an angle. Mount one independently driven Mecanum wheel at each corner of a base, and by combining each wheel's direction and speed, those angled forces combine into forward/backward motion, sideways translation, or rotation in place — full omnidirectional movement without needing to turn like a car. The trade-off is that the rollers carry limited load, are sensitive to uneven floors and gaps, slip easily, and give worse efficiency and positioning accuracy than ordinary wheels, so they're mostly used on smooth indoor floors: AGVs, mobile manipulation platforms, and robotics-competition vehicles.","example":"DJI's RoboMaster S1 education robot uses four Mecanum wheels, letting it strafe sideways while rotating its turret independently.","related":["Omni Wheel","Mobile Base (Chassis)","Swerve Drive","Differential Drive Base","Automated Guided Vehicle","Mobile Manipulator"]},{"id":"omni-wheel","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Omni wheel - Wikipedia","url":"https://en.wikipedia.org/wiki/Omni_wheel"}],"as_of":"","related_ids":["mecanum-wheel","swerve-drive","mobile-base","differential-drive-base","nonholonomic-constraint","mobile-manipulation"],"name":"Omni Wheel","alt":"全向轮","abbr":"","aliases":["Omnidirectional Wheel","Omni-Wheel"],"one_liner":"A wheel ringed with small rollers that lets it roll forward normally but also slide sideways with no resistance.","explanation":"An omni wheel has a ring of small, freely spinning rollers around its rim, with each roller's axis perpendicular to the wheel's own axle. Driven by a motor, the wheel rolls forward normally; under sideways force, the small rollers spin and let the wheel slide sideways with almost no resistance. Arranging three wheels at 120° to each other, or four wheels in a cross layout, and controlling each wheel's speed independently, lets a base translate in any direction on the floor while also rotating in place. Together with the Mecanum wheel, it's one of the two main omnidirectional-drive approaches: a Mecanum wheel's rollers sit at 45° with the wheels mounted parallel like a car's, while an omni wheel's rollers sit at 90°, requiring the wheels themselves to be mounted at an angle. Its downside is that the gaps between rollers cause vibration while driving, and its load capacity and ability to cross obstacles are limited, so it's best suited to smooth indoor floors.","example":"RoboCup's Small Size League soccer robots commonly use a base built from four omni wheels, letting them translate sideways and spin while dribbling the ball.","related":["Mecanum Wheel","Swerve Drive","Mobile Base (Chassis)","Differential Drive Base","Nonholonomic Constraint","Mobile Manipulation"]},{"id":"swerve-drive","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Swerve Drive Kinematics - WPILib Docs","url":"https://docs.wpilib.org/en/stable/docs/software/kinematics-and-odometry/swerve-drive-kinematics.html"}],"as_of":"","related_ids":["mobile-base","mecanum-wheel","omni-wheel","differential-drive-base","automated-guided-vehicle","wheeled-humanoid-robot"],"name":"Swerve Drive","alt":"舵轮","abbr":"","aliases":["Steerable Drive Wheel","Swerve Module","Steer-and-Drive Wheel"],"one_liner":"A wheel module that both drives and steers on its own, letting the whole chassis move in any direction.","explanation":"A swerve drive module puts a drive motor and a steering motor on the same wheel unit: the drive motor spins the wheel, and the steering motor rotates the whole wheel assembly around a vertical axis to point it any direction. A chassis built from two to four of these modules lets the controller work out each wheel's speed and angle independently, so the robot can translate sideways, rotate in place, or do both together. Compared with Mecanum wheels, swerve modules carry heavier loads, tolerate rougher floors, and slip less, but the mechanism and control software are more complex and more expensive. They're common on heavy-duty AGVs (automated guided vehicles) and warehouse robots, and on the chassis of some wheeled humanoid or hybrid robots.","example":"Chassis in the FIRST Robotics Competition (FRC) commonly use four swerve modules, letting the robot translate and rotate at the same time.","related":["Mobile Base (Chassis)","Mecanum Wheel","Omni Wheel","Differential Drive Base","Automated Guided Vehicle","Wheeled Humanoid Robot"]},{"id":"lifting-column","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Hello Robot Stretch","url":"https://hello-robot.com/"}],"as_of":"","related_ids":["mobile-manipulator","mobile-base","wheeled-humanoid-robot","workspace","prismatic-joint","ball-screw"],"name":"Lifting Column","alt":"升降柱","abbr":"","aliases":["Lift Column","Linear Lift Module"],"one_liner":"A vertical linear-motion axis on a mobile base that raises or lowers a robot's torso or arms.","explanation":"A lifting column is a common vertical linear-motion component on mobile-manipulation robots and wheeled humanoids, usually driven by a motor turning a lead screw or timing belt to move a column or carriage up and down, raising or lowering the arms and head camera as a unit. It solves the working-height problem: an arm mounted at a fixed height struggles to reach both the floor and a high shelf, while adding one lifting axis covers a low-to-high workspace with a single degree of freedom — far cheaper and more stable than building a pair of long legs. Selecting one mainly comes down to travel range, payload, speed, and whether it self-locks when power is cut, so the arms don't drop. It's a common part of a composite robot's design, alongside a mobile base and a robot arm.","example":"Hello Robot's Stretch mounts a vertical lifting column on its base, with a telescoping arm hanging off it, letting it pick things up off the floor and also reach tabletops and counters.","related":["Mobile Manipulator","Mobile Base (Chassis)","Wheeled Humanoid Robot","Workspace","Prismatic Joint","Ball Screw"]},{"id":"reverse-knee-leg","category":"hardware","sec":7,"tier":3,"sources":[{"title":"Digitigrade - Wikipedia","url":"https://en.wikipedia.org/wiki/Digitigrade"},{"title":"Agility Robotics","url":"https://agilityrobotics.com/"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Reverse-Knee (Digitigrade, Bird-Like) Leg","alt":"反关节（鸟腿）","abbr":"","aliases":["Digitigrade Leg","Bird Leg"],"one_liner":"A leg that looks like its knee bends backward, but is actually walking on its toes, like a bird.","explanation":"A reverse-knee leg looks like its “knee” bends the wrong way, but it's actually mimicking a bird's digitigrade structure: the joint that appears to bend backward corresponds to a human ankle, while the true knee sits higher on the thigh, close to the body. The best-known examples are Agility Robotics' Cassie and Digit. This leg layout concentrates heavy parts like motors near the hip, keeping the shin and foot light, which lowers the leg's swing inertia and helps with fast stepping and energy efficiency. Its downside is that it doesn't match human leg structure, making it harder to map human motion-capture data onto the robot's joints through motion retargeting, and it looks less human — which is why most humanoid robots aiming for whole-body imitation instead use a human-like, forward-bending knee.","example":"Agility's Cassie biped uses reverse-knee legs, and its light, distal leg design let it complete an outdoor 5-kilometer test run.","related":["Bipedal Robot","Agility Robotics Cassie","Agility Robotics Digit (incl. Digit 5)","Point Foot vs. Flat Foot","Motion Retargeting"]},{"id":"safety-gantry","category":"hardware","sec":7,"tier":2,"sources":[{"title":"Unitree G1 SDK Development Guide","url":"https://support.unitree.com/home/en/G1_developer"}],"as_of":"","related_ids":["humanoid-robot","sim-to-real-transfer","joint-zero-calibration","fall-mitigation-and-fall-recovery","emergency-stop"],"name":"Safety Gantry (Suspended Start)","alt":"吊装架（安全吊架）","abbr":"","aliases":["Safety Hoist","Suspension Rig"],"one_liner":"A rig that suspends a humanoid or legged robot during testing so a fall doesn't damage it.","explanation":"A safety gantry is a gantry frame or mobile crane fitted with a hoist line, used to suspend a robot while testing humanoid or biped robots. It serves two purposes: first, during power-on, zero calibration, or switching control modes, it keeps the robot's legs off the ground so it can't fall over before the controller is ready; second, when testing a new walking policy — especially right after moving it from simulation to the real robot — the line is kept slightly slack, so if the robot loses balance it's caught by the line instead of damaging its joints or shell. Many humanoid makers' getting-started guides require suspending the robot before enabling motion control. It's a basic safety measure for real-robot experiments and plays no part in control itself.","example":"The first time a walking policy trained with reinforcement learning is deployed on a real humanoid, the robot is usually hung from a safety gantry to step with its feet off the ground, then gradually lowered until its feet touch down.","related":["Humanoid Robot","Sim-to-Real Transfer","Joint Zero Calibration (Homing)","Fall Mitigation and Fall Recovery","Emergency Stop"]},{"id":"onboard-compute-platform","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Jetson Thor | NVIDIA","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/"}],"as_of":"2026-09","related_ids":["nvidia-jetson","nvidia-jetson-thor","d-robotics-rdk-s100","rockchip-rk3588","lower-level-controller","on-device-edge-deployment"],"name":"Onboard Compute Platform","alt":"主控","abbr":"","aliases":["Compute Platform","Edge Compute Platform","Main Controller"],"one_liner":"The main computer mounted on a robot that runs perception and model inference.","explanation":"The onboard compute platform (主控, sometimes just “main control” in Chinese industry usage) is the compute hardware carried on a robot's body that does the “thinking” — running perception, localization and navigation, task planning, and inference for models such as a VLA — and then sends commands down to the motion controller and each joint's driver. Common choices include NVIDIA's Jetson line (Orin, Thor), Rockchip's RK3588, NPU-equipped SoCs like D-Robotics' RDK S100, and x86 industrial PCs. Selection mainly comes down to compute throughput (measured in TOPS or TFLOPS), memory, power draw, available interfaces, and software ecosystem. Because a robot runs on battery and has to manage heat, onboard compute is far weaker than a training server's, so models are often compressed, quantized, or partly offloaded to the cloud. In Chinese, 主控 can also loosely mean “the main chip on a control board,” so the meaning depends on context.","example":"NVIDIA's Jetson Thor module is rated at up to 2,070 FP4 TFLOPS with 128 GB of memory and a 40–130 W power range, positioned as the onboard compute platform for humanoids and other robots.","related":["NVIDIA Jetson","NVIDIA Jetson Thor","D-Robotics RDK S100","Rockchip RK3588","Lower-Level Controller","On-Device / Edge Deployment"]},{"id":"nvidia-jetson","category":"hardware","sec":8,"tier":1,"sources":[{"title":"NVIDIA Embedded Systems for Autonomous Machines","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/"},{"title":"Nvidia Jetson - Wikipedia","url":"https://en.wikipedia.org/wiki/Nvidia_Jetson"}],"as_of":"2026-09","related_ids":["nvidia-jetson-orin","nvidia-jetson-thor","onboard-compute-platform","on-device-edge-deployment","nvidia-jetpack-sdk","nvidia-tensorrt"],"name":"NVIDIA Jetson","alt":"英伟达 Jetson","abbr":"","aliases":["Jetson","Jetson Series"],"one_liner":"NVIDIA's family of embedded AI computing platforms built for robots and edge devices.","explanation":"Jetson is NVIDIA's line of embedded computing platforms, packaging an Arm CPU, an NVIDIA GPU, and memory into a small, low-power core module, paired with a carrier board (which breaks out various I/O) or a developer kit. It lets a robot run neural networks on board without a connection to the cloud, making it one of the most common choices for a robot's onboard compute. The lineup began with the TK1 in 2014, and has since gone through TX1/TX2, Xavier, Nano, and Orin, up to 2025's Thor, built on the Blackwell architecture. All of them run the same JetPack software stack (bundling Ubuntu, CUDA, TensorRT, and more), so a model trained on a server is relatively easy to deploy onto one. The main specs to weigh when choosing a module are compute performance, memory, and power draw.","example":"Many humanoid and quadruped robots mount a Jetson Orin inside the body as the onboard computer running perception and policy models.","related":["NVIDIA Jetson Orin","NVIDIA Jetson Thor","Onboard Compute Platform","On-Device / Edge Deployment","NVIDIA JetPack SDK","NVIDIA TensorRT"]},{"id":"nvidia-jetson-orin","category":"hardware","sec":8,"tier":1,"sources":[{"title":"NVIDIA Jetson Orin","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/"},{"title":"NVIDIA Technical Blog: Jetson Orin Nano Developer Kit Gets a Super Boost (2024-12)","url":"https://developer.nvidia.com/blog/nvidia-jetson-orin-nano-developer-kit-gets-a-super-boost/"},{"title":"Unitree G1 官方产品页","url":"https://www.unitree.com/g1"}],"as_of":"2026-09","related_ids":["nvidia-jetson","nvidia-jetson-thor","onboard-compute-platform","tops","on-device-edge-deployment","nvidia-jetpack-sdk"],"name":"NVIDIA Jetson Orin","alt":"Jetson Orin","abbr":"","aliases":["AGX Orin","Orin NX","Orin Nano","Jetson AGX Orin"],"one_liner":"NVIDIA's Ampere-architecture embedded AI module family, launched starting in 2022.","explanation":"Jetson Orin was the previous flagship generation of the Jetson family, built on Ampere-architecture GPUs, and comes in three performance tiers from highest to lowest — AGX Orin, Orin NX, and Orin Nano — with power draw of roughly 7 to 60 watts. The AGX Orin is rated at up to about 275 TOPS (trillion INT8 operations per second, sparse) of compute. In late 2024, NVIDIA shipped a “Super” mode software update that boosted Orin Nano and Orin NX performance, with the Orin Nano Super developer kit priced at $249. Jetson Orin has been the most common onboard computer on robots over the past few years, capable of running detection, segmentation, SLAM, and small policy models, though it struggles with VLA models that have billions of parameters.","example":"Unitree's G1 EDU edition offers a Jetson Orin as a compute module for further development.","related":["NVIDIA Jetson","NVIDIA Jetson Thor","Onboard Compute Platform","TOPS (Tera Operations Per Second)","On-Device / Edge Deployment","NVIDIA JetPack SDK"]},{"id":"nvidia-jetson-thor","category":"hardware","sec":8,"tier":1,"sources":[{"title":"NVIDIA Jetson Thor","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/"},{"title":"NVIDIA launches Jetson T2000 and T3000 modules - CNX Software","url":"https://www.cnx-software.com/2026/07/16/nvidia-jetson-t2000-and-t3000-modules-for-edge-ai-and-robotics-applications/"}],"as_of":"2026-07","related_ids":["nvidia-jetson","nvidia-jetson-orin","onboard-compute-platform","on-device-edge-deployment","nvidia-three-computer-solution","nvidia-isaac-gr00t-n1"],"name":"NVIDIA Jetson Thor","alt":"Jetson Thor","abbr":"","aliases":["Jetson AGX Thor","Jetson T5000","Jetson T4000","Jetson T3000","Jetson T2000"],"one_liner":"NVIDIA's Blackwell-based newest embedded compute platform, built especially for humanoid robots.","explanation":"Jetson Thor is the newest generation of the Jetson family; the Jetson AGX Thor developer kit went on sale in August 2025 for $3,499. Its flagship module, the T5000, uses a Blackwell GPU, a 14-core Arm Neoverse CPU, and 128 GB of memory, rated at up to 2,070 FP4 TFLOPS (sparse), with power draw of 40–130 watts. NVIDIA formally released the T4000 in January 2026, then in July 2026 added the smaller, cheaper T3000 (about 865 FP4 TFLOPS, 32 GB) and T2000 (about 400 TFLOPS, 16 GB), both reportedly shipping in Q1 2027. It's positioned to let humanoid robots run large models like VLA and vision-language models directly on board; NVIDIA describes it as the “deploy” computer in its “three computers” framework.","example":"NVIDIA's own GR00T N series models list Jetson Thor as one of their target platforms for on-device deployment.","related":["NVIDIA Jetson","NVIDIA Jetson Orin","Onboard Compute Platform","On-Device / Edge Deployment","NVIDIA Three-Computer Solution","NVIDIA Isaac GR00T N1"]},{"id":"system-on-chip","category":"hardware","sec":8,"tier":2,"sources":[{"title":"System on a chip - Wikipedia","url":"https://en.wikipedia.org/wiki/System_on_a_chip"}],"as_of":"","related_ids":["onboard-compute-platform","nvidia-jetson","rockchip-rk3588","neural-processing-unit","tops","system-on-module-carrier-board-developer-kit"],"name":"System on Chip (SoC)","alt":"片上系统（SoC）","abbr":"SoC","aliases":["SoC"],"one_liner":"A complete computing system on one chip, integrating a CPU, GPU, and AI accelerator.","explanation":"A system on chip integrates a CPU, a GPU, a neural processing unit (NPU, a compute block specialized for neural networks), memory controllers, video encode/decode, and various interfaces all on the same die; a phone's chip is a typical SoC. Most robot onboard compute platforms are also SoCs: small and low-power, which suits a battery-powered robot body. Choosing one mainly comes down to AI compute (often expressed in TOPS), memory bandwidth, and power draw. NVIDIA's Jetson Orin/Thor, Rockchip's RK3588, and D-Robotics' RDK S100 all fall into this category, usually delivered as a compute module plus a carrier board.","example":"The core of an NVIDIA Jetson AGX Orin module is an SoC integrating an Arm CPU with an Ampere-architecture GPU, and it's commonly used as the onboard compute platform in humanoid and quadruped robots.","related":["Onboard Compute Platform","NVIDIA Jetson","Rockchip RK3588","Neural Processing Unit (NPU)","TOPS (Tera Operations Per Second)","System-on-Module (SoM) / Carrier Board / Developer Kit"]},{"id":"neural-processing-unit","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Neural processing unit - Wikipedia","url":"https://en.wikipedia.org/wiki/Neural_processing_unit"}],"as_of":"","related_ids":["system-on-chip","tops","onboard-compute-platform","rockchip-rknn-toolkit","horizon-robotics-openexplorer","post-training-quantization"],"name":"Neural Processing Unit (NPU)","alt":"神经网络处理器（NPU / BPU）","abbr":"NPU","aliases":["BPU (Brain Processing Unit)","AI Accelerator","Neural Network Accelerator"],"one_liner":"A low-power chip block specialized for the matrix math inside neural networks.","explanation":"An NPU is an accelerator designed specifically for neural-network inference. Its core is a large array of multiply-accumulate units, well-suited to convolution and matrix multiplication, and it usually runs at low numerical precision such as INT8, giving lower power draw than a GPU at the same compute throughput — which is why it's commonly integrated as one block inside an SoC (system-on-chip) used in phones, cars, and robots. Different vendors use different names: Horizon Robotics calls its NPU architecture the BPU (Brain Processing Unit), and D-Robotics' RDK line of dev boards keeps that same name. Using an NPU usually means first converting and quantizing a model with the vendor's own toolchain (such as RKNN-Toolkit or Horizon's Tiangong Kaiwu), and falling back to the CPU or GPU for any operator the NPU doesn't support — this step is often where deploying large models onto real hardware gets stuck.","example":"On a Rockchip RK3588 dev board, a YOLO detection model is quantized to INT8 with RKNN-Toolkit and run on the NPU, freeing the CPU for other work.","related":["System on Chip (SoC)","TOPS (Tera Operations Per Second)","Onboard Compute Platform","Rockchip RKNN-Toolkit","Horizon Robotics OpenExplorer","Post-Training Quantization"]},{"id":"tops","category":"hardware","sec":8,"tier":2,"sources":[{"title":"NVIDIA Jetson Orin","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/"}],"as_of":"","related_ids":["onboard-compute-platform","system-on-chip","flops-tflops","numerical-precision-formats","nvidia-jetson-orin","on-device-edge-deployment"],"name":"TOPS (Tera Operations Per Second)","alt":"TOPS（每秒万亿次运算）","abbr":"TOPS","aliases":["Trillion Operations Per Second"],"one_liner":"The standard unit for AI chip compute: how many trillion operations it can do per second.","explanation":"TOPS is the most commonly advertised spec for AI chips; 1 TOPS means one trillion operations per second, usually measured at INT8 (8-bit integer) precision, though some vendors quote it at lower precisions like INT4 or FP4, which inflates the number. It's not the same as TFLOPS (trillion floating-point operations per second), so it's important to check which precision a number refers to before comparing chips. TOPS is a theoretical peak; how fast a model actually runs also depends on memory bandwidth, the software stack, and operator support, so it should only be used as a rough guide. When choosing a robot's onboard compute platform, TOPS gives a first-pass sense of whether it can run a vision model or a VLA on the robot itself.","example":"NVIDIA's Jetson AGX Orin 64GB is rated at 275 TOPS (INT8, sparse), and it's a common choice as the onboard compute platform for humanoid robots.","related":["Onboard Compute Platform","System on Chip (SoC)","FLOPS / TFLOPS (Floating-Point Operations per Second)","Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)","NVIDIA Jetson Orin","On-Device / Edge Deployment"]},{"id":"flops-tflops","category":"hardware","sec":8,"tier":2,"sources":[{"title":"FLOPS - Wikipedia","url":"https://en.wikipedia.org/wiki/FLOPS"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"FLOPS / TFLOPS (Floating-Point Operations per Second)","alt":"浮点算力（FLOPS / TFLOPS）","abbr":"","aliases":["FLOPS","TFLOPS"],"one_liner":"How many floating-point calculations a chip can do per second; a TFLOPS is a trillion per second.","explanation":"FLOPS stands for floating-point operations per second, a measure of how much compute power a GPU or main compute chip has; a TFLOPS is 10¹² operations per second, and a PFLOPS is 10¹⁵. It's easy to confuse with the lowercase-s FLOPs, which is the total number of operations a single forward pass requires — a measure of a model's computational cost; dividing FLOPs by FLOPS gives a rough estimate of inference time. Reading this number requires checking the numeric precision: the same chip's rating can differ several-fold between FP32, FP16, FP8, and FP4, and manufacturers often advertise the peak figure at the lowest precision or under sparse conditions; integer throughput is usually reported separately as TOPS. What a chip can actually sustain is further limited by memory bandwidth and how efficient the operators are. This number comes up when choosing a robot's onboard compute or estimating whether a VLA model can run in real time on the device.","example":"NVIDIA advertises Jetson Thor's compute figure at FP4 precision, which can't be directly compared to the older Orin's number, which is quoted at FP16.","related":["Tera Operations Per Second","Floating-Point Operations (FLOPs)","Numerical Precision Formats (FP32 / BF16 / FP16 / FP8 / INT8)","NVIDIA Jetson Thor","Onboard Compute Platform","Inference Latency"]},{"id":"system-on-module-carrier-board-developer-kit","category":"hardware","sec":8,"tier":3,"sources":[{"title":"System on module - Wikipedia","url":"https://en.wikipedia.org/wiki/System_on_module"},{"title":"Jetson Modules - NVIDIA Developer","url":"https://developer.nvidia.com/embedded/jetson-modules"}],"as_of":"","related_ids":["nvidia-jetson","nvidia-jetson-orin","nvidia-jetson-thor","onboard-compute-platform","system-on-chip","embedded-system"],"name":"System-on-Module (SoM) / Carrier Board / Developer Kit","alt":"核心模组与载板（SoM / 载板 / 开发者套件）","abbr":"SoM","aliases":["SoM","Core Board","Baseboard","Dev Kit"],"one_liner":"An embedded-hardware pattern where a chip ships as a small module that plugs into a carrier board with the actual connectors.","explanation":"A system-on-module (SoM) packs the processor, memory, storage, and power management onto one small board, with almost no connectors of its own. A carrier board, sometimes called a baseboard, supplies the connectors the module needs to actually be used — Ethernet, USB, camera interfaces, power input. A developer kit is what a vendor sells as a ready-to-use bundle: a module plus a reference carrier board plus cooling, so a team can start developing right away. Robotics teams typically prototype and validate their software on a developer kit first, then, once the design is set, design their own compact carrier board to fit the robot's chassis and buy only the module for production. NVIDIA's Jetson line is a well-known example of this pattern.","example":"Jetson AGX Orin Developer Kit: the Orin module sits on a reference carrier board, so plugging in power and a camera is enough to start running models. For a production robot, the team later swaps in its own custom carrier board.","related":["NVIDIA Jetson","NVIDIA Jetson Orin","NVIDIA Jetson Thor","Onboard Compute Platform","System on Chip (SoC)","Embedded System"]},{"id":"rockchip-rk3588","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Rockchip RK3588 - Rockchips.net","url":"https://rockchips.net/product/rk3588/"},{"title":"CNX Software: Rockchip RK3588 specifications revealed","url":"https://www.cnx-software.com/2020/11/26/rockchip-rk3588-specifications-revealed-8k-video-6-tops-npu-pcie-3-0-up-to-32gb-ram/"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"Rockchip RK3588","alt":"瑞芯微 RK3588","abbr":"","aliases":["RK3588","RK3588S"],"one_liner":"An 8nm, eight-core ARM chip from Rockchip with a 6 TOPS NPU, a common low-cost robot compute platform.","explanation":"The RK3588 is a flagship SoC (a chip integrating a CPU, GPU, AI accelerator, and more) from Fuzhou-based Rockchip, built on an 8nm process, with a CPU combining 4 Cortex-A76 and 4 Cortex-A55 cores, a Mali-G610 GPU, a built-in 6 TOPS NPU (neural-network accelerator), and support for 8K video encode/decode. It's used in a huge number of dev boards and industrial mainboards, and in robots it commonly serves as a low-cost, low-power onboard compute platform or host computer, running ROS 2, camera processing, and lightweight model inference; a model has to be converted with RKNN-Toolkit before it can run on the NPU. Compared with NVIDIA's Jetson Orin, it's cheaper and more power-efficient, but its AI compute and CUDA ecosystem lag far behind, so it can't handle a large VLA model.","example":"A YOLO detection model is converted to RKNN format with RKNN-Toolkit and deployed to the NPU on an RK3588 dev board, giving a robot car real-time object recognition.","related":["Onboard Compute Platform","Rockchip RKNN-Toolkit","Neural Processing Unit (NPU)","NVIDIA Jetson","On-Device / Edge Deployment","TOPS (Tera Operations Per Second)"]},{"id":"black-sesame-sesamex","category":"hardware","sec":8,"tier":3,"sources":[{"title":"黑芝麻智能亮相2026世界人工智能大会（黑芝麻智能官网）","url":"https://www.blacksesame.com/zh/list_8/994.html"},{"title":"A2000明年量产，具身智能打平台战：黑芝麻智能的两张牌（OFweek）","url":"https://www.ofweek.com/auto/2026-04/ART-70101-8110-30684885.html"}],"as_of":"2026-07","related_ids":["onboard-compute-platform","nvidia-jetson-thor","d-robotics-rdk-s100","tops","compute-control-integration","nvidia-jetson"],"name":"Black Sesame SesameX","alt":"黑芝麻 SesameX","abbr":"","aliases":["SesameX","SesameX Multi-Dimensional Embodied AI Computing Platform"],"one_liner":"An embodied-AI computing platform for robots from automotive-chip maker Black Sesame Technologies.","explanation":"SesameX is the “multi-dimensional embodied AI computing platform” launched by Black Sesame Technologies, an automotive-chip company, which reuses its automotive-grade compute know-how from self-driving chips for robots — aimed at service, industrial, and humanoid robots. The company describes three self-developed core modules, Kalos, Aura, and Liora, with compute ranging from about 48 TOPS to nearly 600 TOPS (TOPS meaning trillion operations per second), paired with its Huashan A2000 chip family. It's positioned similarly to NVIDIA's Jetson Thor or D-Robotics' RDK S100, serving as the onboard compute platform on a robot's body, running perception and VLA-style models alongside motion control.","example":"Black Sesame showcased the SesameX platform at the 2026 World Artificial Intelligence Conference, according to the company's own news release.","related":["Onboard Compute Platform","NVIDIA Jetson Thor","D-Robotics RDK S100","TOPS (Tera Operations Per Second)","Compute-Control Integration","NVIDIA Jetson"]},{"id":"qualcomm-dragonwing-iq10","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Introducing the Qualcomm Dragonwing IQ10 RRD (Edge AI and Vision Alliance)","url":"https://www.edge-ai-vision.com/2026/06/introducing-the-qualcomm-dragonwing-iq10-rrd-a-full-stack-robotics-reference-design/"},{"title":"Qualcomm's Dragonwing IQ10 Aims to Be the Brain for Next-Gen Humanoids (BigGo News)","url":"https://biggo.com/news/202601051624_qualcomm-dragonwing-iq10-robot-processor-ces-2026"}],"as_of":"2026-09","related_ids":["onboard-compute-platform","nvidia-jetson-thor","tops","on-device-edge-deployment","humanoid-robot","gigabit-multimedia-serial-link"],"name":"Qualcomm Dragonwing IQ10","alt":"高通跃龙 IQ10","abbr":"","aliases":["Dragonwing IQ10","IQ10"],"one_liner":"Qualcomm's high-end onboard processor platform for industrial and humanoid robots.","explanation":"The IQ10 is a high-end processor in Qualcomm's Dragonwing industrial and embedded product line, aimed at robotics, and was reportedly announced at CES in January 2026, targeting industrial robots, autonomous mobile robots, and humanoid robots as an onboard “brain.” Qualcomm's official figures list up to about 700 TOPS of AI compute, integrating an 18-core Oryon CPU, an NPU, and a GPU, intended to let perception, planning, and model inference all run on the robot itself without a separate add-in accelerator card. In June 2026, Qualcomm introduced a robotics reference design built on it, the IQ10 RRD, at Computex, supporting up to 12 GMSL2 camera feeds plus lidar and IMU inputs, with global availability planned from September 2026; partners include NEURA Robotics, Booster Robotics, and VinMotion. It competes directly with NVIDIA's Jetson Thor.","example":"","related":["Onboard Compute Platform","NVIDIA Jetson Thor","TOPS (Tera Operations Per Second)","On-Device / Edge Deployment","Humanoid Robot","Gigabit Multimedia Serial Link"]},{"id":"host-computer","category":"hardware","sec":8,"tier":2,"sources":[{"title":"上位机 - 百度百科","url":"https://baike.baidu.com/item/%E4%B8%8A%E4%BD%8D%E6%9C%BA"}],"as_of":"","related_ids":["lower-level-controller","industrial-pc","onboard-compute-platform","robot-controller","qt","software-development-kit"],"name":"Host Computer","alt":"上位机","abbr":"","aliases":["Main Control Computer","Upper Computer"],"one_liner":"The computer that issues commands and does high-level computing, while a lower-level controller executes them.","explanation":"“Host computer” (上位机, literally “upper-position computer”) is common engineering shorthand in China's robotics industry for the computer in a robot system that handles high-level decision-making, the user interface, and data processing — usually a PC, industrial PC, or a board like an NVIDIA Jetson. It sends target positions, velocities, and other commands over Ethernet, USB, or CAN bus to the “lower-level controller” (下位机: a microcontroller or motor driver that directly drives motors and reads sensors), and receives status data back. This split lets compute-heavy tasks without hard real-time requirements — vision, planning, running a neural-network policy — run on the host computer, while tight, high-frequency control loops (current loop, position loop) stay on the lower-level controller. In embodied AI work, the VLA model usually runs on the host computer; “host computer software” also often refers to the debugging or GUI program running on this machine.","example":"A GPU-equipped PC running a policy model that sends target joint positions to a Unitree G1 over Ethernet via the Unitree SDK is acting as the host computer.","related":["Lower-Level Controller","Industrial PC","Onboard Compute Platform","Robot Controller","Qt","Software Development Kit"]},{"id":"lower-level-controller","category":"hardware","sec":8,"tier":2,"sources":[{"title":"上位机 - 维基百科","url":"https://zh.wikipedia.org/wiki/上位机"},{"title":"Microcontroller - Wikipedia","url":"https://en.wikipedia.org/wiki/Microcontroller"}],"as_of":"","related_ids":["host-computer","microcontroller-unit","stmicroelectronics-stm32-mcu-family","servo-drive","real-time-control","controller-area-network"],"name":"Lower-Level Controller","alt":"下位机","abbr":"","aliases":["Slave Computer","Low-Level Controller"],"one_liner":"The embedded control board wired directly to sensors and motors, running real-time low-level control.","explanation":"“Lower-level controller” (下位机) is the counterpart to “host computer” in Chinese engineering usage, referring to the compute layer wired directly to motor drivers, encoders, IMUs, and other hardware — typically an MCU (microcontroller), a DSP, or a PLC. It doesn't run large models; it only handles tasks with strict timing requirements: reading sensors and running the current loop and position loop at a high, fixed frequency, outputting PWM (pulse-width modulation) signals, and exchanging data with the host computer over CAN, EtherCAT, or a serial port, receiving targets and reporting status back. The host computer (an industrial PC, a Jetson, etc.) handles perception, planning, and policy inference. This division of labor means that if the high-level software stalls briefly, low-level control keeps running smoothly. English sources usually call this a low-level or embedded controller; “slave computer” is an older term for the same role.","example":"On a humanoid robot, a policy running on a Jetson might output target joint angles every 20 ms, while an STM32 lower-level controller on each joint's driver board tracks those targets in a much faster closed loop.","related":["Host Computer","Microcontroller Unit (MCU)","STMicroelectronics STM32 MCU Family","Servo Drive (Motor Driver)","Real-Time Control","Controller Area Network (CAN)"]},{"id":"embedded-system","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Embedded system - Wikipedia","url":"https://en.wikipedia.org/wiki/Embedded_system"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Embedded System","alt":"嵌入式系统","abbr":"","aliases":[],"one_liner":"A small computer system built into a device to handle one specific control task.","explanation":"An embedded system is a computer hidden inside a piece of equipment, dedicated to a specific function, built from a processor (a microcontroller unit, or MCU, or a system on chip, or SoC), memory, peripheral interfaces, and the firmware or lightweight operating system running on top. Compared with a general-purpose computer, it's optimized for small size, low power draw, low cost, and predictable response time. Embedded systems are everywhere in a robot: the MCU inside a motor driver runs the current control loop, sensor boards handle data acquisition and packaging, and dexterous hands and battery management systems each have their own onboard controller; even a higher-level compute module like a Jetson is often classified as an embedded platform. Working on robots means dealing with embedded systems directly — flashing firmware, configuring communication protocols, and handling real-time constraints.","example":"The STM32 inside a joint driver runs the current control loop at tens of kilohertz, then receives commands from the host computer over CAN or EtherCAT.","related":["Microcontroller Unit","System on Chip","Firmware","Real-Time Operating System","STMicroelectronics STM32 MCU Family","Lower-level Controller (Slave Computer)"]},{"id":"microcontroller-unit","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Microcontroller - Wikipedia","url":"https://en.wikipedia.org/wiki/Microcontroller"}],"as_of":"","related_ids":["stmicroelectronics-stm32-mcu-family","lower-level-controller","embedded-system","real-time-operating-system","field-oriented-control","system-on-chip"],"name":"Microcontroller Unit (MCU)","alt":"微控制器","abbr":"MCU","aliases":["Microcontroller"],"one_liner":"A small control chip integrating a processor, memory, and peripheral interfaces on a single die.","explanation":"A microcontroller unit, commonly called an MCU, integrates a CPU core, flash memory, RAM, timers, analog-to-digital conversion, PWM (pulse-width modulation) outputs, and communication interfaces like CAN and serial ports, all on a single chip; common examples include STMicroelectronics' STM32 and Espressif's ESP32. It has modest compute power and low power draw, but predictable response times, making it well suited to running bare-metal code or a real-time operating system. In robots, MCUs are spread across joint drivers, sensor boards, dexterous hands, and the battery management system, handling low-level, real-time control like the tens-of-kilohertz current loop; the onboard compute platform that runs Linux and neural networks (such as a Jetson) is a different class of chip entirely, and the two work together over a bus.","example":"The STM32 on a joint driver board runs the FOC current loop at a very high frequency, while also receiving position and torque commands from the host computer over a CAN bus.","related":["STMicroelectronics STM32 MCU Family","Lower-Level Controller","Embedded System","Real-Time Operating System (RTOS)","Field-Oriented Control","System on Chip (SoC)"]},{"id":"stmicroelectronics-stm32-mcu-family","category":"hardware","sec":8,"tier":2,"sources":[{"title":"STM32 32-bit Arm Cortex MCUs - STMicroelectronics","url":"https://www.st.com/en/microcontrollers-microprocessors/stm32-32-bit-arm-cortex-mcus.html"},{"title":"STM32 - Wikipedia","url":"https://en.wikipedia.org/wiki/STM32"}],"as_of":"","related_ids":["microcontroller-unit","embedded-system","lower-level-controller","servo-drive","controller-area-network","real-time-operating-system"],"name":"STMicroelectronics STM32 MCU Family","alt":"STM32","abbr":"","aliases":["STM32","STM32 Microcontroller"],"one_liner":"A family of 32-bit microcontrollers from STMicroelectronics, built on Arm Cortex-M cores.","explanation":"STM32 is the 32-bit microcontroller (MCU) family from STMicroelectronics, built on Arm Cortex-M cores, spanning everything from the low-power L series to the high-performance H series. It's inexpensive, has rich on-chip peripherals (PWM, ADC, CAN, UART, and more), and has extensive documentation and community support, making it close to the standard entry point for embedded development in China's electronics and robotics community. In robots, it commonly serves as the lower-level controller: running the current loop for motor drivers, reading encoders and IMUs, and handling CAN bus communication, then passing data up to the host computer or onboard compute platform (such as a Jetson) that runs the policy.","example":"Many open-source brushless motor driver boards and joint modules use an STM32 as their control chip.","related":["Microcontroller Unit (MCU)","Embedded System","Lower-Level Controller","Servo Drive (Motor Driver)","Controller Area Network (CAN)","Real-Time Operating System (RTOS)"]},{"id":"arduino-esp32-microcontroller-boards","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Arduino - Official Site","url":"https://www.arduino.cc/"},{"title":"ESP32 - Espressif Systems","url":"https://www.espressif.com/en/products/socs/esp32"}],"as_of":"","related_ids":["microcontroller-unit","lower-level-controller","stmicroelectronics-stm32-mcu-family","servo","raspberry-pi","open-source-hardware"],"name":"Arduino / ESP32 Microcontroller Boards","alt":"Arduino / ESP32 开发板","abbr":"","aliases":["Arduino","ESP32"],"one_liner":"Cheap, beginner-friendly microcontroller boards commonly used to drive servos and read sensors.","explanation":"Arduino is an open-source microcontroller platform started by an Italian team, with a simplified programming environment and a large library of ready-made code; ESP32 is a low-cost microcontroller chip from Espressif, with built-in Wi-Fi and Bluetooth, that can also be programmed in the Arduino environment. Both are microcontrollers (MCUs) with modest compute power — not enough to run large neural networks — but they offer good real-time response and low power draw, making them well-suited as a lower-level controller: driving servos and motors, reading encoders and IMUs, and communicating with a host computer over a serial link or wirelessly. They're commonly used when getting started with desktop robot arms, robot cars, and small bipeds, with a Raspberry Pi or a computer running vision and the policy above them.","example":"In a low-cost desktop robot arm, an ESP32 can receive joint-angle commands from a computer and drive each servo with a PWM signal.","related":["Microcontroller Unit (MCU)","Lower-Level Controller","STMicroelectronics STM32 MCU Family","Servo (Smart Serial Bus Servo)","Raspberry Pi","Open-Source Hardware (OSHW)"]},{"id":"raspberry-pi","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Raspberry Pi - Wikipedia","url":"https://en.wikipedia.org/wiki/Raspberry_Pi"},{"title":"Raspberry Pi 官网","url":"https://www.raspberrypi.com/"}],"as_of":"","related_ids":["microcontroller-unit","arduino-esp32-microcontroller-boards","embedded-system","robot-operating-system-2","lekiwi","policy-server"],"name":"Raspberry Pi","alt":"树莓派","abbr":"","aliases":["RPi"],"one_liner":"A low-cost, credit-card-sized computer from the UK, a common onboard computer for beginner robotics.","explanation":"The Raspberry Pi is a credit-card-sized single-board computer from the UK's Raspberry Pi Foundation and its subsidiary, Raspberry Pi Ltd, first released in 2012. It's built on an ARM processor, can run Linux (the official Raspberry Pi OS, or Ubuntu), and has a row of GPIO pins that can connect directly to sensors, servos, and motor driver boards. It's cheap, well documented, and backed by a large community, making it a common “brain” for teaching robotics and for low-cost platforms, capable of running ROS 2, reading a camera, and controlling a robot car. Its compute isn't enough to run large models, so in VLA work it's typically only responsible for collecting data and executing actions, while inference runs on a remote policy server with a GPU.","example":"Hugging Face's LeKiwi mobile manipulator uses a Raspberry Pi 5 as its onboard computer to collect camera and motor data; the policy runs inference on a laptop or GPU elsewhere and sends actions back.","related":["Microcontroller Unit (MCU)","Arduino / ESP32 Microcontroller Boards","Embedded System","Robot Operating System 2","LeKiwi","Policy Server (Remote Inference)"]},{"id":"compute-control-integration","category":"hardware","sec":8,"tier":3,"sources":[{"title":"地瓜机器人发布首款单SoC算控一体化机器人开发套件（36氪）","url":"https://eu.36kr.com/zh/p/3333355053115910"},{"title":"RDK S100 开发套件（地瓜机器人开发者社区）","url":"https://developer.d-robotics.cc/rdks100"}],"as_of":"2025-06","related_ids":["braincerebellum-architecture","d-robotics-rdk-s100","system-on-chip","microcontroller-unit","real-time-control","onboard-compute-platform"],"name":"Compute-Control Integration","alt":"算控一体","abbr":"","aliases":["Integrated Compute-Control","Combined Brain-Cerebellum Chip"],"one_liner":"Putting AI inference compute and real-time motion control on the same chip or the same board.","explanation":"Traditional robots often use two separate pieces of hardware: one compute platform runs perception and decision-making models, while a separate MCU (microcontroller) or motion-control board handles millisecond-level joint control, with the two communicating over Ethernet or a bus. Compute-control integration means combining both onto the same SoC (system on chip) or the same main board: a high-performance CPU/NPU handles the “brain's” model inference, while a real-time core on the same chip handles the “cerebellum's” motion control. The benefits are less board-to-board communication latency, less wiring, and lower cost and power draw; the challenge is keeping the control task's real-time performance and functional safety free from interference by the AI workload. Products like D-Robotics' RDK S100 are marketed around this feature.","example":"D-Robotics' RDK S100 integrates a 6-core Cortex-A78AE CPU, an 80-TOPS BPU, and a 4-core Cortex-R52+ MCU on a single SoC, according to the company's own announcement.","related":["Brain–Cerebellum Architecture","D-Robotics RDK S100","System on Chip (SoC)","Microcontroller Unit (MCU)","Real-Time Control","Onboard Compute Platform"]},{"id":"d-robotics-rdk-s100","category":"hardware","sec":8,"tier":3,"sources":[{"title":"RDK S100 开发套件 - 地瓜机器人开发者社区","url":"https://developer.d-robotics.cc/rdks100"},{"title":"80 TOPS算力、大小脑超级异构！地瓜机器人RDK S100开启预售 - 量子位","url":"https://www.qbitai.com/2025/06/292932.html"}],"as_of":"2025-06","related_ids":["onboard-compute-platform","compute-control-integration","nvidia-jetson","nvidia-jetson-orin","d-robotics","neural-processing-unit"],"name":"D-Robotics RDK S100","alt":"地瓜 RDK S100","abbr":"","aliases":["RDK S100P","RDK S100 Developer Kit"],"one_liner":"A single-chip “compute-control integrated” robot dev kit from D-Robotics, starting at 80 TOPS.","explanation":"The RDK S100 is a robot development kit released in June 2025 by D-Robotics, part of Horizon Robotics, built around putting AI inference and real-time motor control on the same SoC (system on chip): a 6-core Cortex-A78AE CPU handles the system and planning, a self-developed Nash-architecture BPU (neural-network accelerator) delivers 80 TOPS of compute for vision and VLA-style models, and a separate 4-core Cortex-R52+ MCU handles joint-level real-time control. The higher-end S100P variant offers 128 TOPS and 24 GB of memory. It's designed to solve the communication latency and cost that come from using one board for the “brain” and another for the “cerebellum” in a robot. At launch, its pre-order price was reported at RMB 2,499, positioned to compete with edge compute platforms like NVIDIA's Jetson Orin.","example":"A developer can run a perception model on the BPU and drive robot-arm joints directly over CAN from the onboard MCU, both on the same RDK S100 board.","related":["Onboard Compute Platform","Compute-Control Integration","NVIDIA Jetson","NVIDIA Jetson Orin","D-Robotics","Neural Processing Unit (NPU)"]},{"id":"robot-controller","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Industrial robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_robot"}],"as_of":"","related_ids":["teach-pendant","industrial-robot","servo-motor","integrated-drive-and-control","rtde","libfranka-franka-control-interface"],"name":"Robot Controller","alt":"机器人控制器","abbr":"","aliases":["Control Cabinet","Control Box"],"one_liner":"The control box paired with a robot arm, housing its motion-control computer, servo drives, and safety circuits.","explanation":"In industrial robotics, the robot controller is usually a standalone control cabinet or box containing a motion-control computer, servo drivers for each joint, a power supply, safety circuits, and I/O ports, with an external teach pendant. It parses the robot's program, computes forward and inverse kinematics and trajectory interpolation, and hands each joint's target to its servo drive for closed-loop execution, while also handling safety functions like emergency stop, collision detection, and joint limits. Each of the “big four” industrial robot makers has its own controller; collaborative-robot makers build smaller control boxes, and some newer products integrate the drive directly into the joint (drive-integrated joints). In embodied AI research, a policy usually doesn't drive motors directly — it sends target poses or joint commands at high frequency through the controller's interface (such as UR's RTDE or Franka's libfranka), and the controller handles low-level tracking and safety.","example":"Running a VLA experiment on a UR collaborative arm, the policy on a workstation sends joint targets to the UR control box over the RTDE interface; the control box then drives the six joint motors and handles emergency stop and safety limits.","related":["Teach Pendant","Industrial Robot","Servo Motor","Integrated Drive and Control","RTDE","libfranka / Franka Control Interface (FCI)"]},{"id":"teach-pendant","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Teach pendant - Wikipedia","url":"https://en.wikipedia.org/wiki/Teach_pendant"}],"as_of":"","related_ids":["teach-and-playback-programming","industrial-robot","emergency-stop","robot-controller","kinesthetic-teaching","offline-programming"],"name":"Teach Pendant","alt":"示教器","abbr":"TP","aliases":["TP","Teaching Pendant","Teach Box"],"one_liner":"The handheld control unit for an industrial robot arm, used to jog it, teach positions, and write motion programs.","explanation":"A teach pendant is a handheld terminal wired to a robot's control cabinet, with a screen, buttons or a touchscreen, an emergency-stop button, and a three-position enabling switch — releasing it or pressing it all the way down both trigger a stop, so only the middle position allows motion. The operator uses it to jog the robot's joints or tool tip to a target position and record that point; stringing several recorded points together builds a program the robot then repeats automatically. This “teach and playback” method is the most traditional way of programming industrial robots from makers like FANUC, ABB, and KUKA. Collaborative arms often replace it with a tablet-style pendant or hand-guided teaching. In embodied-AI research, a teach pendant is mostly used to home the arm, change parameters, and hit emergency stop.","example":"An operator jogs the arm to a position above a workpiece using the teach pendant and saves it as a waypoint, then records a few more points to build a pick-and-place sequence.","related":["Teach-and-Playback Programming","Industrial Robot","Emergency Stop","Robot Controller","Kinesthetic Teaching","Offline Programming"]},{"id":"industrial-pc","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Industrial PC - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_PC"}],"as_of":"","related_ids":["host-computer","lower-level-controller","robot-controller","programmable-logic-controller","ethercat-master","nvidia-jetson"],"name":"Industrial PC","alt":"工控机","abbr":"IPC","aliases":["IPC","Industrial Computer","Industrial Control Computer"],"one_liner":"A computer hardened for harsh environments such as factory floors.","explanation":"An industrial PC (IPC) is a computer built to industrial standards; internally it's much like a regular PC (usually x86 hardware running Windows or Linux), but its enclosure, cooling, power supply, and ports are ruggedized — commonly fanless cooling, wide input-voltage tolerance, dust and vibration resistance, and a wide operating-temperature range, plus multiple serial ports, CAN bus, and Ethernet ports for industrial use, so it can run continuously for long stretches. On a robot, it often serves as the host computer or the core of the robot controller, running ROS, motion planning, vision, or acting as an EtherCAT master. Compared with embedded AI boards like NVIDIA's Jetson line, an industrial PC is more general-purpose and more expandable, but is usually larger and draws more power.","example":"An AGV/AMR chassis carries a fanless industrial PC running ROS navigation, which controls motor drivers over a CAN bus.","related":["Host Computer","Lower-Level Controller","Robot Controller","Programmable Logic Controller (PLC)","EtherCAT Master (SOEM / IgH)","NVIDIA Jetson"]},{"id":"programmable-logic-controller","category":"hardware","sec":8,"tier":3,"sources":[{"title":"Programmable logic controller - Wikipedia","url":"https://en.wikipedia.org/wiki/Programmable_logic_controller"}],"as_of":"","related_ids":["industrial-robot","robot-controller","ethercat","host-computer","lower-level-controller","real-time-control"],"name":"Programmable Logic Controller (PLC)","alt":"PLC（可编程逻辑控制器）","abbr":"PLC","aliases":["Programmable Controller"],"one_liner":"A specialized industrial computer that controls equipment actions and production-line logic in a factory.","explanation":"A PLC is a specialized computer designed for industrial settings, running in a fixed, repeating scan cycle: it reads inputs (sensors, buttons, limit switches), executes the user's program, and updates outputs (relays, solenoid valves, motor drivers). It emerged in the late 1960s to replace the complex relay cabinets used in car factories; its programming languages are defined by the IEC 61131-3 standard, the most common being ladder logic. It's reliable, resistant to electrical noise, and can run continuously for years. When a robot is installed in a factory, the pace and interlocking of the whole production line are typically managed by a PLC, with the robot's controller “handshaking” with the PLC through I/O signals or an industrial bus — embodied AI robots deployed on a production line have to interface with the existing PLC in the same way.","example":"Once a PLC detects that a workpiece is in position, it sends a “start” signal to the robot arm; after the arm finishes grasping it, it sends back a “done” signal, and the PLC starts the next section of the conveyor.","related":["Industrial Robot","Robot Controller","EtherCAT (Ethernet for Control Automation Technology)","Host Computer","Lower-Level Controller","Real-Time Control"]},{"id":"gpu-memory","category":"hardware","sec":8,"tier":2,"sources":[{"title":"Video random-access memory - Wikipedia","url":"https://en.wikipedia.org/wiki/Video_random-access_memory"},{"title":"openpi - Physical Intelligence (GitHub)","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["parameter-count","lora","post-training-quantization","gradient-checkpointing","batch-size","on-device-edge-deployment"],"name":"GPU Memory (VRAM)","alt":"GPU 显存","abbr":"VRAM","aliases":["VRAM","Video RAM"],"one_liner":"The high-speed memory built into a graphics card that limits how big a model and batch can be.","explanation":"GPU memory (VRAM) is the dedicated high-speed memory built into a graphics card. Model parameters, gradients, optimizer state, intermediate activations, and input data all have to fit in it for the GPU to compute anything. How much VRAM you have directly sets the largest model you can train or run, and the largest batch size (how many samples you feed in at once) you can use. Training needs far more VRAM than inference, because it also has to store gradients and optimizer state. Common fixes when you run out include lower numerical precision (e.g., BF16, INT8 quantization), LoRA (which trains only a small number of parameters), gradient checkpointing, and splitting the model across multiple GPUs. For embodied AI, VRAM determines whether a VLA (vision-language-action) model can be fine-tuned on a lab's GPUs, and whether it can be squeezed onto an onboard chip on the robot itself.","example":"The openpi documentation gives these reference numbers: inference for π0 needs over 8 GB of VRAM, LoRA fine-tuning needs roughly 22.5 GB or more, and full fine-tuning needs roughly 70 GB or more.","related":["Parameter Count (Model Size)","LoRA","Post-Training Quantization","Gradient Checkpointing (Activation Recomputation)","Batch Size","On-Device / Edge Deployment"]},{"id":"common-training-and-inference-gpus","category":"hardware","sec":8,"tier":2,"sources":[{"title":"openpi README（Hardware Requirements）","url":"https://github.com/Physical-Intelligence/openpi"},{"title":"NVIDIA DGX B200 User Guide: Introduction","url":"https://docs.nvidia.com/dgx/dgxb200-user-guide/introduction-to-dgxb200.html"}],"as_of":"2026-09","related_ids":["gpu-memory","nvidia-jetson","lora","full-fine-tuning","numerical-precision-formats","autodl"],"name":"Common Training and Inference GPUs (RTX 4090 / A100 / H100 / B200)","alt":"常用 GPU 型号（RTX 4090 / A100 / H100 / B200）","abbr":"","aliases":["RTX 4090","A100","H100","B200"],"one_liner":"The handful of NVIDIA GPUs most often used to train and run embodied-AI models, differing mainly in memory and compute.","explanation":"These are the NVIDIA GPUs that show up most often in embodied-AI papers and code repositories. The RTX 4090 (Ada architecture, 24 GB of memory) is a consumer gaming card, often used for single-machine inference and small-scale fine-tuning. The A100 (Ampere architecture, 40 or 80 GB) and H100 (Hopper architecture, 80 GB, with FP8 support) are data-center cards, used in multi-GPU clusters for pretraining and full-parameter fine-tuning. The B200 (Blackwell architecture) has up to about 180 GB of memory per card, and an 8-GPU DGX B200 system has roughly 1.4 TB of memory combined. When choosing a card, memory capacity comes first, since it determines how large a model and batch size fit; compute throughput, supported numeric precisions, and inter-GPU interconnect bandwidth matter next. Models running on the robot itself typically run on an edge chip like a Jetson rather than on any of these cards.","example":"The openpi repository's guidance: running π0 inference needs at least 8 GB of memory and LoRA fine-tuning needs at least 22.5 GB, both of which an RTX 4090 can handle; full-parameter fine-tuning needs at least 70 GB, requiring an 80 GB A100 or an H100.","related":["GPU Memory (VRAM)","NVIDIA Jetson","LoRA","Full Fine-Tuning","Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)","AutoDL"]},{"id":"battery-runtime","category":"hardware","sec":9,"tier":1,"sources":[{"title":"Unitree G1 官方产品页","url":"https://www.unitree.com/g1"}],"as_of":"2026-09","related_ids":["battery-management-system","hot-swappable-battery","autonomous-battery-swapping","autonomous-docking-and-recharging","humanoid-robot"],"name":"Battery Runtime","alt":"续航","abbr":"","aliases":["Battery Endurance","Battery Life"],"one_liner":"How long a robot can keep working on a single full battery charge.","explanation":"Battery runtime is how long a robot can keep operating on one battery (or one full charge), and it depends on battery capacity, the robot's overall weight, motor efficiency, and how demanding the task is — the same robot standing still uses far less power than one walking continuously or carrying a heavy load. Because humanoid robots have to constantly work to stay balanced and have many joints drawing power, their runtime is typically only a few hours, which is one of the main bottlenecks keeping them out of factories and homes. Common fixes include using a bigger battery, cutting power consumption, hot-swappable batteries (swapping packs without powering down), and autonomous return-to-charge or autonomous battery swapping. When reading a manufacturer's runtime figure, it's worth checking what workload it was measured under.","example":"Unitree's G1 humanoid is rated by the company at about 2 hours of runtime.","related":["Battery Management System (BMS)","Hot-Swappable Battery","Autonomous Battery Swapping","Autonomous Docking and Recharging","Humanoid Robot"]},{"id":"hot-swappable-battery","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Hot swapping - Wikipedia","url":"https://en.wikipedia.org/wiki/Hot_swapping"}],"as_of":"","related_ids":["battery-runtime","battery-management-system","autonomous-battery-swapping","autonomous-docking-and-recharging","humanoid-robot"],"name":"Hot-Swappable Battery","alt":"热插拔电池","abbr":"","aliases":["Quick-Swap Battery","Quick-Release Battery"],"one_liner":"A battery that can be swapped for a charged one quickly, without powering down or removing screws.","explanation":"A hot-swappable battery can be pulled out and replaced with a charged one while the device stays powered, or with only a brief pause, usually via a latch or rail mechanism that needs no tools. Strictly, true “hot-swapping” means the system never loses power during the swap (for example, by cycling between two batteries or using a built-in backup supply); many robots marketed as having a “quick-swap battery” are really only fast to remove and install, and still have to be powered off during the swap, so the two should be distinguished. Humanoid and quadruped robots typically run for only one to a few hours per charge, so fast battery swaps are what let them keep working, or keep collecting data, continuously — making this a small but important engineering detail in moving robots from demos to real deployment. A further step is having the robot dock and swap its own battery autonomously.","example":"Robots such as the Unitree G1 use quick-release batteries, so operators can swap batteries and keep working during data collection.","related":["Battery Runtime","Battery Management System (BMS)","Autonomous Battery Swapping","Autonomous Docking and Recharging","Humanoid Robot"]},{"id":"autonomous-battery-swapping","category":"hardware","sec":9,"tier":2,"sources":[{"title":"UBTECH Walker S2 产品页","url":"https://www.ubtrobot.com/en/humanoid/products/walker-s2"},{"title":"CnEVPost: UBTech shows how its humanoid robot can work 24/7 with autonomous battery swap","url":"https://cnevpost.com/2025/07/17/ubtech-humanoid-robot-autonomous-battery-swap/"}],"as_of":"2025-07","related_ids":["autonomous-docking-and-recharging","hot-swappable-battery","battery-runtime","battery-management-system","ubtech-walker-s2"],"name":"Autonomous Battery Swapping","alt":"自主换电","abbr":"","aliases":["Automatic Battery Swap"],"one_liner":"A robot that walks itself to a swap station and replaces its own battery when it's low, with no human help.","explanation":"Autonomous battery swapping means a robot, when running low on power, goes to a battery-swap station on its own, removes its depleted battery, installs a fully charged one, and gets back to work — all without a person's involvement. Compared with plugging in to charge, it skips the hours-long wait, letting a robot run close to around the clock, which makes it a problem factories can't avoid once they put humanoid robots on the production line. Implementing it requires a battery mechanism that can be quickly plugged and unplugged (often paired with a dual-battery setup or backup power so the robot never loses power mid-swap), plus precise localization and hand-eye coordination. UBTech's Walker S2, released in 2025, claims it can autonomously swap its own battery in about 3 minutes using both arms.","example":"On a production line, when a UBTech Walker S2 runs low on power, it walks to a battery cabinet, pulls the battery pack off its back with both arms, and inserts a fully charged one, all without stopping operations.","related":["Autonomous Docking and Recharging","Hot-Swappable Battery","Battery Runtime","Battery Management System (BMS)","UBTech Walker S2"]},{"id":"autonomous-docking-and-recharging","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Wikipedia: Robotic vacuum cleaner","url":"https://en.wikipedia.org/wiki/Robotic_vacuum_cleaner"},{"title":"Wikipedia: Charging station","url":"https://en.wikipedia.org/wiki/Charging_station"}],"as_of":"","related_ids":["autonomous-battery-swapping","battery-runtime","wireless-charging","robot-vacuum-cleaner","apriltag","navigation"],"name":"Autonomous Docking and Recharging","alt":"自主回充","abbr":"","aliases":["Auto-Return Charging"],"one_liner":"A robot that finds its own way back to a charging dock and lines itself up to recharge when the battery runs low.","explanation":"Autonomous docking and recharging means a robot, when its battery runs low or a task finishes, navigates back to a charging dock on its own, aligns with the contacts or plug, completes charging, and heads back out. Robot vacuums were the first to make this common, and today quadrupeds, inspection robots, and mobile bases largely treat it as a standard feature. The hard part is the last few centimeters of alignment, usually solved with infrared beacons, markers like QR codes or AprilTags, or matching the shape of the dock with lidar. Compared with autonomous battery swapping, docking to recharge is mechanically simpler and cheaper, but the robot can't work while it's charging.","example":"When a robot vacuum finishes cleaning or its battery gets low, it drives itself back to its dock to recharge, then returns to where it left off once it's full.","related":["Autonomous Battery Swapping","Battery Runtime","Wireless (Inductive) Charging","Robot Vacuum Cleaner","AprilTag","Navigation"]},{"id":"wireless-charging","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Inductive charging - Wikipedia","url":"https://en.wikipedia.org/wiki/Inductive_charging"}],"as_of":"","related_ids":["autonomous-docking-and-recharging","battery-runtime","battery-management-system","autonomous-battery-swapping","automated-guided-vehicle"],"name":"Wireless (Inductive) Charging","alt":"无线充电","abbr":"","aliases":["Inductive Charging"],"one_liner":"Charging a battery across a small gap using electromagnetic induction, with no plug involved.","explanation":"Wireless charging uses electromagnetic induction (or magnetic resonance) to transfer power between a transmitter coil and a receiver coil, with no physical plug contact needed — it's common on phones and electric toothbrushes. For a robot, lining up precisely with a charging plug during autonomous docking is hard, and plug contacts wear out, collect dust, or get damp; wireless charging only requires the robot to stop at roughly the right spot, which relaxes the positioning accuracy needed and makes it easier to seal the robot against water. The tradeoff is that transfer efficiency is usually lower than wired charging, power is limited, and it generates heat. It's commonly used for autonomous return-to-charge on robot vacuums, AGVs, and inspection robots, and is one way to let a mobile robot run unattended for long stretches.","example":"A warehouse AGV parks itself over a charging pad on the floor; a transmitter coil under the floor tops up the onboard battery with no cable ever plugged in.","related":["Autonomous Docking and Recharging","Battery Runtime","Battery Management System (BMS)","Autonomous Battery Swapping","Automated Guided Vehicle"]},{"id":"battery-management-system","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Battery management system - Wikipedia","url":"https://en.wikipedia.org/wiki/Battery_management_system"}],"as_of":"","related_ids":["battery-runtime","hot-swappable-battery","autonomous-battery-swapping","solid-state-battery","controller-area-network","regenerative-braking-and-brake-resistor"],"name":"Battery Management System (BMS)","alt":"电池管理系统","abbr":"BMS","aliases":["BMS"],"one_liner":"The circuitry and software that monitor and protect a battery pack and estimate its remaining charge.","explanation":"A battery management system is a set of circuitry and software built into a battery pack that continuously measures each cell's voltage, current, and temperature, uses that to estimate state of charge (SOC) and state of health (SOH), cuts the circuit during overcharge, over-discharge, overcurrent, or overtemperature conditions, and balances charge across cells so they stay even. Mobile robots run on lithium batteries, and the large, sudden current draws from joint motors — plus the current that flows back in during braking — put real stress on the battery, so the BMS determines how accurate the charge reading is, whether the robot suddenly cuts out, and how long the battery lasts. It typically reports charge level and warnings to the onboard compute platform over a bus such as CAN, and both hot-swappable batteries and autonomous battery swapping rely on it to confirm battery state.","example":"The battery-percentage number shown in a robot's app is the BMS's estimate, reported up over the bus.","related":["Battery Runtime","Hot-Swappable Battery","Autonomous Battery Swapping","Solid-State Battery","Controller Area Network (CAN)","Regenerative Braking and Brake (Shunt) Resistor"]},{"id":"solid-state-battery","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Solid-state battery - Wikipedia","url":"https://en.wikipedia.org/wiki/Solid-state_battery"}],"as_of":"","related_ids":["battery-runtime","battery-management-system","hot-swappable-battery","autonomous-battery-swapping","humanoid-robot"],"name":"Solid-State Battery","alt":"固态电池","abbr":"","aliases":["All-Solid-State Battery"],"one_liner":"A lithium battery that replaces the liquid electrolyte with a solid one.","explanation":"A solid-state battery replaces the liquid electrolyte in a conventional lithium-ion cell with a solid electrolyte — oxide, sulfide, or polymer. In theory this makes the cell less prone to catching fire, and it pairs more easily with a lithium-metal anode to raise energy density. Humanoid robots are tight on both space and weight, and battery life is typically just a few hours, so a battery that packs more energy per kilogram and is safer is seen as one way to extend runtime. Full solid-state batteries are still working toward mass production; what's often marketed as “semi-solid-state” still contains a small amount of liquid electrolyte, and different companies don't use the term “solid-state” the same way, so marketing claims are worth reading carefully.","example":"","related":["Battery Runtime","Battery Management System (BMS)","Hot-Swappable Battery","Autonomous Battery Swapping","Humanoid Robot"]},{"id":"wire-harness","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Cable harness - Wikipedia","url":"https://en.wikipedia.org/wiki/Cable_harness"}],"as_of":"","related_ids":["hollow-shaft-cable-routing","slip-ring","joint-actuator-module","ethercat","controller-area-network"],"name":"Wire Harness","alt":"线束","abbr":"","aliases":["Cable Harness"],"one_liner":"The bundled, connectorized assembly of power and signal cables routed through a robot.","explanation":"A wire harness bundles multiple power, communication, and encoder signal cables along a fixed path and fits connectors on both ends; it's used heavily in both cars and robots. A humanoid robot has dozens of joints, and each joint module needs power plus a connection to a bus such as EtherCAT or CAN. The cables have to pass through joints that keep rotating, and repeated bending and twisting wears them down, breaks conductors, or loosens contacts — a common source of whole-robot failures. So harness design has to account for routing path, bend radius, and how the cable is secured, and it's often paired with hollow-shaft routing (running the cable through a hole in the center of the joint) or a slip ring (letting the cable pass through a joint that spins without limit) to cut down on tangling. How the harness is laid out also affects how far a joint can rotate and how efficient the robot is to assemble.","example":"A humanoid robot's arm harness usually runs through the hollow shaft at the center of each joint, all the way from the shoulder to the wrist and dexterous hand.","related":["Hollow-Shaft Cable Routing","Slip Ring","Joint Actuator Module","EtherCAT (Ethernet for Control Automation Technology)","Controller Area Network (CAN)"]},{"id":"hollow-shaft-cable-routing","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Slip ring - Wikipedia","url":"https://en.wikipedia.org/wiki/Slip_ring"}],"as_of":"","related_ids":["joint-actuator-module","wire-harness","slip-ring","strain-wave-gear","frameless-torque-motor"],"name":"Hollow-Shaft Cable Routing","alt":"中空走线","abbr":"","aliases":["Hollow-Shaft Wiring","Internal Cable Routing"],"one_liner":"Routing cables through the center bore of a joint's motor and reducer instead of along the outside.","explanation":"In hollow-shaft cable routing, a joint module's motor rotor, reducer, and encoder are all built with a hollow structure, leaving a through-hole in the center for power, data, and pneumatic lines to pass from inside one joint straight to the next, instead of hanging on the outside. The benefit is a cleaner look, cables that can't get tangled or snagged, and a lower risk of catching on a person's hand — important for collaborative arms and humanoid robots that work close to people. The trade-off is that the motor and reducer have to be built a size larger, and the bore diameter and cable bend life have to be designed together; the cable also twists as the joint rotates, which puts an upper limit on the joint's rotation angle. Applications that need unlimited rotation use a slip ring instead. The “hollow-shaft” variants commonly offered for harmonic drives exist for exactly this purpose.","example":"Many collaborative robot arms show no visible cables on the outside at all, because every cable is routed through the center bore of each joint.","related":["Joint Actuator Module","Wire Harness","Slip Ring","Strain Wave Gear (Harmonic Drive)","Frameless Torque Motor"]},{"id":"slip-ring","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Slip ring - Wikipedia","url":"https://en.wikipedia.org/wiki/Slip_ring"}],"as_of":"","related_ids":["hollow-shaft-cable-routing","wire-harness","joint-actuator-module","revolute-joint","ethercat"],"name":"Slip Ring","alt":"滑环","abbr":"","aliases":["Rotary Electrical Connector","Conductive Slip Ring"],"one_liner":"A rotating electrical connector that keeps power and signal flowing between a spinning part and a fixed one.","explanation":"A slip ring is made of a conductive ring that turns with the shaft and a fixed brush that presses against it, sliding to keep contact as the ring spins. That sliding contact carries power, encoder signals, or bus data across the rotating interface. It solves a basic wiring problem: a joint that needs to spin continuously, all the way around, would otherwise wind its cables up and snap them. In robots, slip rings show up wherever a joint needs unlimited rotation — camera pan-tilt heads, turntables, and some waist or wrist joints. Arm joints with a limited rotation range more often route cables straight through a hollow shaft instead. Slip rings wear over time and can add contact noise, so high-speed signal lines need a slip ring model rated for that.","example":"A security pan-tilt camera that can spin without limit relies on a slip ring to feed the camera power and send video back through its base.","related":["Hollow-Shaft Cable Routing","Wire Harness","Joint Actuator Module","Revolute Joint","EtherCAT (Ethernet for Control Automation Technology)"]},{"id":"pose-repeatability","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Industrial robot - Wikipedia（Repeatability 与 ISO 9283 小节）","url":"https://en.wikipedia.org/wiki/Industrial_robot"}],"as_of":"","related_ids":["absolute-positioning-accuracy","industrial-robot","teach-and-playback-programming","kinematic-calibration","hand-eye-calibration","collaborative-robot"],"name":"Pose Repeatability","alt":"重复定位精度","abbr":"","aliases":["Repeatability","Repeat Positioning Accuracy"],"one_liner":"How closely a robot returns to the same point each time it's commanded there.","explanation":"Pose repeatability measures how consistent a robot's actual stopping point is across repeated executions of the same command to reach the same target pose, usually written as something like “±0.05 mm.” The international standard ISO 9283 defines the measurement method: the robot approaches the same position repeatedly from the same direction, and the spread of the resulting points is measured statistically. It differs from absolute positioning accuracy, which measures how far off the robot lands from the true target when commanded to a point by coordinates — this is usually much worse than repeatability, and is improved through kinematic calibration. Traditional industrial robots are mostly programmed by teaching (moving the arm to a point and recording it) and returning to that same point, so repeatability matters most for them; when picking up an object at coordinates computed from vision, absolute accuracy and hand-eye calibration error matter more.","example":"An arm spec'd at ±0.1 mm repeatability will reliably return to within 0.1 mm of a taught point; but when it moves to a new position computed from a camera, the error can be much larger — that's a question of absolute positioning accuracy instead.","related":["Absolute Positioning Accuracy","Industrial Robot","Teach-and-Playback Programming","Kinematic Calibration","Hand-Eye Calibration","Collaborative Robot"]},{"id":"ingress-protection-rating","category":"hardware","sec":9,"tier":3,"sources":[{"title":"IP code - Wikipedia","url":"https://en.wikipedia.org/wiki/IP_code"}],"as_of":"","related_ids":["quadruped-robot","inspection-robot","industrial-robot","joint-actuator-module","mean-time-between-failures"],"name":"Ingress Protection (IP) Rating","alt":"IP 防护等级","abbr":"IP","aliases":["IP67","IP54","IP Code"],"one_liner":"The two-digit international standard rating for how well a device resists dust and water.","explanation":"The IP rating is defined by the international standard IEC 60529, written as “IP” followed by two digits. The first digit rates protection against solid objects, including dust, from 0 to 6, where 6 means fully dust-tight; the second rates protection against water, from 0 to 9, where a higher number withstands more — for example, 4 means splashing water, and 7 means brief submersion. If a category wasn't tested, it's replaced with an X, as in IPX7. On a robot's product page, IP54 roughly means dust-protected and splash-resistant, while IP67 means fully dust-tight and able to survive brief submersion. This rating determines whether a robot can be used outdoors, in the rain, in a greasy factory, or in an environment that needs to be hosed down, which is why it matters a great deal for quadruped robots, inspection robots, and robot arms used in the food industry.","example":"An outdoor inspection quadruped rated IP67 can generally handle being rained on or wading through a shallow puddle.","related":["Quadruped Robot","Inspection Robot","Industrial Robot","Joint Actuator Module","Mean Time Between Failures"]},{"id":"payload-to-weight-ratio","category":"hardware","sec":9,"tier":3,"sources":[{"title":"UR5e - Universal Robots","url":"https://www.universal-robots.com/products/ur5e/"}],"as_of":"","related_ids":["payload","torque-density","lightweighting","robotic-arm","collaborative-robot","power-density"],"name":"Payload-to-Weight Ratio","alt":"负载自重比","abbr":"","aliases":["Weight-to-Payload Ratio"],"one_liner":"A robot's rated payload divided by its own weight; higher means lighter and stronger for its size.","explanation":"Payload-to-weight ratio is the rated payload divided by the weight of the robot itself (usually meaning a robot arm, or the whole machine), used to gauge whether a design is light and its actuation strong. Some sources flip the ratio and report weight-to-payload instead, so it's worth checking which way a number is written. The ratio is shaped jointly by motor torque density, the reducer, and structural materials: improving it means using higher-power-density joints and lightweight materials like magnesium alloy or carbon fiber. For an arm bolted to a table, this ratio mainly affects mounting and cost; it matters even more for a humanoid, a quadruped, or an arm mounted on a mobile base, because every extra kilogram on the arm is a kilogram the legs, base, and battery all have to carry too, which shortens battery life. When comparing numbers, also check whether a manufacturer's figure is rated or peak payload, and at what reach and speed it was measured.","example":"A collaborative arm weighing 20 kg with a 5 kg rated payload has a payload-to-weight ratio of 1:4.","related":["Payload","Torque Density","Lightweighting (Magnesium Alloy / Carbon Fiber)","Robotic Arm","Collaborative Robot","Power Density (W/kg)"]},{"id":"lightweighting","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Magnesium alloy - Wikipedia","url":"https://en.wikipedia.org/wiki/Magnesium_alloy"},{"title":"Carbon-fiber reinforced polymers - Wikipedia","url":"https://en.wikipedia.org/wiki/Carbon-fiber-reinforced_polymers"}],"as_of":"","related_ids":["replacing-steel-with-engineering-plastics","payload-to-weight-ratio","moment-of-inertia","battery-runtime","polyether-ether-ketone","torque-density"],"name":"Lightweighting (Magnesium Alloy / Carbon Fiber)","alt":"轻量化（镁合金 / 碳纤维）","abbr":"","aliases":["Lightweight Design"],"one_liner":"Using light materials like magnesium alloy, carbon fiber, and engineering plastics to cut a robot's weight.","explanation":"Lightweighting means reducing the weight of a robot's structural parts while still meeting strength and stiffness requirements, using materials such as magnesium alloy (about two-thirds the density of aluminum alloy), carbon-fiber composites (high strength per unit weight, good for rods and shells), and engineering plastics such as PEEK. For legged and humanoid robots, the benefits of cutting weight compound: lighter limbs mean lower rotational inertia, which means the joint motors need less torque, which extends battery life and reduces impact force in a fall or a collision with a person. The trade-off is higher material and manufacturing cost — magnesium alloy needs corrosion treatment, and carbon-fiber parts are hard to repair. It's commonly discussed alongside plastic-for-steel substitution, payload-to-weight ratio, and battery life.","example":"Replacing an aluminum-alloy lower leg with a carbon-fiber tube reduces mass at the far end of the leg, making the hip and knee joints work less hard when swinging the leg quickly.","related":["Replacing Steel with Engineering Plastics (Plastic-for-Steel Substitution)","Payload-to-Weight Ratio","Moment of Inertia","Battery Runtime","Polyether Ether Ketone (PEEK)","Torque Density"]},{"id":"replacing-steel-with-engineering-plastics","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Polyether ether ketone - Wikipedia","url":"https://en.wikipedia.org/wiki/Polyether_ether_ketone"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"Replacing Steel with Engineering Plastics (Plastic-for-Steel Substitution)","alt":"以塑代钢","abbr":"","aliases":["Plastic-for-Steel Substitution","Plastic-for-Metal Substitution"],"one_liner":"Using PEEK and other engineering plastics instead of metal for robot structural and transmission parts, to cut weight and cost.","explanation":"This refers to substituting engineering plastics — such as PEEK (polyether ether ketone), POM, fiber-reinforced nylon, and carbon-fiber composites — for steel and aluminum in structural parts and some transmission parts, such as housings, frames, gears, and bearing cages. Humanoid robots chase both lightweighting and lower mass-production cost, which has drawn attention to this approach; reports say programs such as Tesla's Optimus have been evaluating PEEK components, which has driven interest in related stocks. The benefits are a density far below steel's and the ability to be injection-molded in bulk; the trade-offs are stiffness, heat resistance, and creep resistance (slow deformation under sustained load) that fall short of metal, so parts usually need to be redesigned rather than simply swapped one for one.","example":"Switching a joint housing from machined aluminum alloy to an injection-molded, carbon-fiber-reinforced PEEK part makes it lighter and lets it be formed in a single shot, instead of CNC-machined piece by piece.","related":["Polyether Ether Ketone (PEEK)","Lightweighting (Magnesium Alloy / Carbon Fiber)","Bill of Materials (BOM) Cost","Mass Production","Humanoid Robot"]},{"id":"polyether-ether-ketone","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Polyether ether ketone - Wikipedia","url":"https://en.wikipedia.org/wiki/Polyether_ether_ketone"}],"as_of":"","related_ids":["replacing-steel-with-engineering-plastics","lightweighting","speed-reducer-gearbox","strain-wave-gear","core-components"],"name":"Polyether Ether Ketone (PEEK)","alt":"PEEK","abbr":"PEEK","aliases":["PEEK Plastic"],"one_liner":"A heat-resistant, high-strength engineering plastic gaining interest as a way to cut robot weight.","explanation":"PEEK, short for polyether ether ketone, is a semi-crystalline, high-performance engineering plastic with a melting point around 343°C, strong mechanical strength, wear resistance, and chemical resistance; its density is about 1.3 g/cm³, roughly a sixth that of steel, and it has long been used in aerospace, medical implants, and automotive parts. Because humanoid robots have many joints and need to stay light, the industry sees PEEK as a candidate material for “plastic-for-steel” substitution — used for gears, bearing cages, some reducer parts, and structural components, to cut weight and lower the inertia of moving parts. Its main obstacle is high raw-material and machining cost, so core transmission parts that carry heavy loads are still mostly metal today.","example":"","related":["Replacing Steel with Engineering Plastics (Plastic-for-Steel Substitution)","Lightweighting (Magnesium Alloy / Carbon Fiber)","Speed Reducer / Gearbox","Strain Wave Gear (Harmonic Drive)","Core Components"]},{"id":"3d-printing","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Wikipedia: 3D printing","url":"https://en.wikipedia.org/wiki/3D_printing"},{"title":"Wikipedia: Fused filament fabrication","url":"https://en.wikipedia.org/wiki/Fused_filament_fabrication"}],"as_of":"","related_ids":["open-source-hardware","so-100-so-101-arm","leap-hand","universal-manipulation-interface","gripper","cad-software"],"name":"3D Printing (FDM / Resin SLA)","alt":"3D 打印（FDM / 光固化）","abbr":"","aliases":["Additive Manufacturing","FDM","SLA"],"one_liner":"Building a part by depositing material layer by layer — the go-to lab tool for fast grippers and mounts.","explanation":"3D printing is an additive-manufacturing method that builds parts directly, one layer at a time. FDM (fused deposition modeling) heats and extrudes plastic filament — PLA, PETG, nylon — and stacks it layer by layer; it's cheap and good for structural parts. Resin printing (SLA/LCD) cures liquid resin layer by layer with light, giving finer surface detail and higher precision, though the resulting parts tend to be more brittle. Embodied-AI labs commonly use it to quickly make gripper fingertips, camera mounts, handheld data-collection rigs, and parts for open-source robot arms, so a design revision can be printed and tested the same day. Many open-source hardware projects ship printable model files directly, ready to assemble with off-the-shelf servos or motors.","example":"The structural parts of LeRobot's SO-100/SO-101 arms and the UMI handheld gripper are both printed on desktop FDM printers.","related":["Open-Source Hardware (OSHW)","SO-100 / SO-101 Arm","LEAP Hand","Universal Manipulation Interface","Gripper","CAD Software (SolidWorks / Onshape / Fusion 360)"]},{"id":"open-source-hardware","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Open Source Hardware Definition | OSHWA","url":"https://www.oshwa.org/definition/"},{"title":"Open-source hardware - Wikipedia","url":"https://en.wikipedia.org/wiki/Open-source_hardware"},{"title":"TheRobotStudio/SO-ARM100 - GitHub","url":"https://github.com/TheRobotStudio/SO-ARM100"}],"as_of":"2026-09","related_ids":["so-100-so-101-arm","aloha","leap-hand","berkeley-humanoid-lite","lerobot","3d-printing"],"name":"Open-Source Hardware (OSHW)","alt":"开源硬件","abbr":"OSHW","aliases":["Open-Source Robot Hardware"],"one_liner":"Hardware whose blueprints, parts list, and firmware are public for anyone to copy or modify.","explanation":"Open-source hardware is hardware whose design files are made public under a license that lets others study, modify, manufacture, or even sell it. What's published usually includes CAD models, circuit schematics and PCB layouts, a bill of materials (BOM), assembly instructions, and firmware. The commonly used definition comes from the Open Source Hardware Association (OSHWA), and licenses include the CERN Open Hardware Licence. In embodied AI, open-source hardware has sharply lowered the barrier to entry: projects like ALOHA, the SO-100/SO-101 arm, LEAP Hand, Berkeley Humanoid Lite, and ToddlerBot all publish their blueprints and code, letting labs and individuals assemble their own using 3D-printed parts plus off-the-shelf motors, and letting other teams reproduce papers or share data collected on the same hardware platform.","example":"The SO-101 arm was open-sourced by TheRobotStudio and is built from 3D-printed structural parts and Feetech STS3215 servos; paired with Hugging Face's LeRobot, it can be teleoperated to collect data and train policies.","related":["SO-100 / SO-101 Arm","ALOHA","LEAP Hand","Berkeley Humanoid Lite","LeRobot","3D Printing (FDM / Resin SLA)"]},{"id":"exoskeleton","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Powered exoskeleton - Wikipedia","url":"https://en.wikipedia.org/wiki/Powered_exoskeleton"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Exoskeleton","alt":"外骨骼","abbr":"","aliases":["Exoskeleton Robot","Powered Exoskeleton"],"one_liner":"A wearable mechanical structure whose joints line up with the body's, used to assist motion or to record it.","explanation":"An exoskeleton is a mechanical structure worn on the outside of the body, with its joints positioned to match the wearer's own. A powered exoskeleton, with motors built in, can assist a person's movement — used in rehabilitation training, helping people with limited mobility walk, or reducing the load on a worker carrying heavy items. A passive exoskeleton, unpowered and fitted only with encoders, instead precisely records the wearer's joint angles. This second kind matters a lot in embodied AI: an exoskeleton built to be isomorphic with a robot arm (matching its number of joints and proportions) lets a person perform a motion while wearing it, and the readings can be mapped directly to robot joint commands — used for teleoperation and collecting demonstration data, with better accuracy and lower latency than vision-based motion capture. Some exoskeletons also provide force feedback, letting the operator feel the resistance the robot encounters.","example":"AirExo uses a low-cost exoskeleton to teleoperate a dual-arm robot, and HOMIE uses an isomorphic exoskeleton cockpit to control a humanoid robot's upper body.","related":["Exoskeleton Teleoperation","AirExo","HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit","Teleoperation","Leader-Follower Teleoperation","Motion Capture"]},{"id":"vr-headset","category":"hardware","sec":9,"tier":2,"sources":[{"title":"Virtual reality headset - Wikipedia","url":"https://en.wikipedia.org/wiki/Virtual_reality_headset"}],"as_of":"","related_ids":["vr-teleoperation","teleoperation","apple-vision-pro","meta-quest-3-pico-4-ultra-xr-headsets","open-television","motion-retargeting"],"name":"VR Headset","alt":"VR 头显","abbr":"","aliases":["XR Headset","MR Headset"],"one_liner":"A head-worn virtual/mixed-reality device, commonly used to teleoperate robots in embodied AI.","explanation":"A VR headset is a head-worn device that displays imagery to both eyes and tracks head motion; most newer models also support see-through views of the real world (mixed reality, hence also called an XR or MR headset). In embodied AI, its main use is teleoperated data collection: the headset tracks the operator's head, wrist, and finger poses and maps them onto the robot's motion, while the video feed from cameras on the robot's head is streamed back to the headset in real time, letting the operator work as if standing in the robot's place. Compared with a leader-follower arm rig, a headset is cheaper, quicker to learn, and can also control a dexterous hand; common models include the Apple Vision Pro, Meta Quest 3, and PICO 4 Ultra.","example":"Open-TeleVision uses an Apple Vision Pro to track the operator's hands and streams the robot's stereo camera feed back to the headset in real time, teleoperating a humanoid robot through manipulation tasks.","related":["VR Teleoperation","Teleoperation","Apple Vision Pro","Meta Quest 3 / PICO 4 Ultra XR Headsets","Open-TeleVision","Motion Retargeting"]},{"id":"meta-quest-3-pico-4-ultra-xr-headsets","category":"hardware","sec":9,"tier":3,"sources":[{"title":"Meta Quest 3 - Wikipedia","url":"https://en.wikipedia.org/wiki/Meta_Quest_3"},{"title":"unitreerobotics/xr_teleoperate - GitHub","url":"https://github.com/unitreerobotics/xr_teleoperate"}],"as_of":"2026-09","related_ids":["vr-teleoperation","vr-headset","apple-vision-pro","unitree-xr-teleoperate","xrobotoolkit","motion-retargeting"],"name":"Meta Quest 3 / PICO 4 Ultra XR Headsets","alt":"Meta Quest 3 / PICO 4 Ultra 头显","abbr":"","aliases":["Quest 3","PICO 4 Ultra"],"one_liner":"Two consumer mixed-reality headsets commonly repurposed for robot teleoperation.","explanation":"The Meta Quest 3 is a mixed-reality headset released by Meta in 2023, and the PICO 4 Ultra is a similar product released in 2024 by PICO, a ByteDance subsidiary. Both include built-in head tracking, hand tracking, controller tracking, and color pass-through cameras, at a price far below the Apple Vision Pro. In embodied AI, they're mainly used as low-cost VR teleoperation devices: the headset streams the operator's head pose, hand keypoints, or controller pose to the robot in real time, which is converted into joint commands through motion retargeting; the video feed from the robot's head cameras is streamed back to the headset, letting the operator work as if physically present, while the session is recorded as demonstration data.","example":"Unitree's open-source xr_teleoperate supports using XR devices such as the Quest 3 or PICO 4 Ultra to teleoperate the Unitree G1 humanoid and its dexterous hands.","related":["VR Teleoperation","VR Headset","Apple Vision Pro","Unitree xr_teleoperate","XRoboToolkit","Motion Retargeting"]},{"id":"rigid-body","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Rigid body","url":"https://en.wikipedia.org/wiki/Rigid_body"},{"title":"Modern Robotics (Lynch & Park), Ch.2 Degrees of Freedom of a Rigid Body","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["rigid-body-dynamics","pose","degrees-of-freedom","link","rigid-body-simulation","deformable-body-simulation"],"name":"Rigid Body","alt":"刚体","abbr":"","aliases":[],"one_liner":"An idealized object where the distance between any two internal points never changes, no matter what force it feels.","explanation":"A rigid body is an idealized model from classical mechanics: no matter how much force is applied, the distance between any two points inside the object stays exactly the same — it never deforms at all. No object is a truly perfect rigid body in reality, but anything that deforms very little, like a metal link, a cup, or a wooden block, can be approximated this way. The payoff is a very compact description: in 3D space, a rigid body's state needs only 6 numbers — 3 for the position of a reference point (usually the center of mass) and 3 for orientation (as a rotation matrix, quaternion, or Euler angles) — that is, 6 degrees of freedom. Almost all of robotics modeling is built on this idealization: every link of an arm, a grasped object, and every “body” in MuJoCo or Isaac Sim is treated as a rigid body; objects that visibly deform, like cloth, rope, or dough, need soft-body simulation instead.","example":"In simulation, a gripper holding a block is completely determined by recording just its 3D position plus one quaternion — 7 numbers, corresponding to 6 degrees of freedom. A T-shirt, by contrast, needs hundreds or thousands of mesh vertices just to describe its shape.","related":["Rigid-Body Dynamics","Pose","Degrees of Freedom (DoF)","Link","Rigid-Body Simulation","Deformable-Body Simulation"]},{"id":"degrees-of-freedom","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Degrees of freedom (mechanics)","url":"https://en.wikipedia.org/wiki/Degrees_of_freedom_(mechanics)"},{"title":"Unitree G1 官网规格","url":"https://www.unitree.com/g1"}],"as_of":"2026-09","related_ids":["active-dof-passive-dof","kinematic-redundancy","revolute-joint","configuration-space","joint-space","underactuation"],"name":"Degrees of Freedom (DoF)","alt":"自由度","abbr":"DoF","aliases":["DOF"],"one_liner":"The number of independent variables needed to fully describe a system's configuration — often a robot's independently movable joints.","explanation":"Degrees of freedom (DoF) is the number of independent parameters needed to fully pin down a mechanical system's configuration. A free rigid body in space has 6 degrees of freedom: translation along x, y, and z, plus rotation about each of those three axes. Each rotary or linear joint on a robot typically contributes 1 degree of freedom, which is why people talk about a “six-axis arm” or a “7-DoF arm.” Six degrees of freedom is the minimum needed to place an end effector at any position and orientation in space; any joints beyond that are called redundant degrees of freedom, since the same end-effector pose can then be reached by infinitely many joint configurations — useful for avoiding obstacles. A human arm — 3 in the shoulder, 1 in the elbow, 3 in the wrist — has 7 degrees of freedom. When manufacturers quote a robot's total DoF, they often fold in the dexterous hand's fingers, so it's worth checking exactly what's being counted before comparing numbers.","example":"According to Unitree's website, the base G1 has 23 degrees of freedom in total (6 per leg, 5 per arm, 1 in the waist); the EDU version, fitted with a dexterous hand and other add-ons, ranges from 23 to 43.","related":["Active DoF / Passive DoF","Kinematic Redundancy","Revolute Joint","Configuration Space (C-Space)","Joint Space","Underactuation"]},{"id":"coordinate-frame","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Frame of reference","url":"https://en.wikipedia.org/wiki/Frame_of_reference"},{"title":"ROS REP 103: Standard Units of Measure and Coordinate Conventions","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0103.rst"}],"as_of":"","related_ids":["world-frame","base-frame","camera-coordinate-frame","coordinate-transformation","tf-tf2-transform-tree","pose"],"name":"Coordinate Frame","alt":"坐标系","abbr":"","aliases":["Reference Frame","Frame"],"one_liner":"A defined origin and set of axis directions used to describe positions and orientations.","explanation":"A coordinate frame — often just called a “frame” in robotics — consists of an origin and three mutually perpendicular axes; position is given as three numbers (x, y, z) and orientation as a rotation relative to those axes. The same physical point has different numeric coordinates in different frames, so any position, velocity, or force is meaningless without saying which frame it's expressed in. A single robot has many frames at once: a world frame fixed in the environment, a base frame fixed to the robot's base, and one for each link, camera, and gripper. In physics, “reference frame” also implies something about the observer's own motion, but in robotics engineering the two terms are mostly used interchangeably. ROS mandates right-handed coordinate frames throughout, and uses the TF tree to track the relationships between them.","example":"A camera measures a cup at 0.6 meters directly in front of the lens, in the camera frame — but the arm needs the cup's position in the base frame, so the two must be related through a coordinate transformation.","related":["World Frame","Base Frame","Camera Coordinate Frame","Coordinate Transformation","TF / tf2 Transform Tree","Pose"]},{"id":"right-handed-frame-and-axis-conventions","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"REP 103: Standard Units of Measure and Coordinate Conventions","url":"https://github.com/ros-infrastructure/rep/blob/master/rep-0103.rst"},{"title":"Isaac Sim Documentation: Conventions Reference","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/reference_material/reference_conventions.html"},{"title":"glTF 2.0 Specification: Coordinate System and Units","url":"https://github.com/KhronosGroup/glTF/blob/main/specification/2.0/Specification.adoc"}],"as_of":"","related_ids":["coordinate-frame","coordinate-transformation","camera-coordinate-frame","tf-tf2-transform-tree","rep-105","quaternion-component-order"],"name":"Right-Handed Frame & Axis Conventions","alt":"右手坐标系与轴向约定","abbr":"","aliases":["REP 103","Z-up vs. Y-up","Camera Optical Frame","ENU"],"one_liner":"Conventions for how x, y, z are arranged and which axis points up — inconsistent across software, and a frequent source of bugs.","explanation":"A right-handed coordinate frame is one where x, y, and z follow the right-hand rule: point the right thumb, index, and middle fingers along x, y, and z respectively, so x crossed with y gives z; a positive rotation about an axis is the direction the fingers curl when the right thumb points along that axis's positive direction. ROS's REP 103 (drafted starting 2010) specifies that a body frame points x forward, y left, and z up, and that geographic coordinates use east-north-up (ENU); cameras get a separate frame with an _optical suffix, where z points forward, x points right, and y points down. Another point of disagreement is which axis points up: ROS, MuJoCo, and Isaac Sim default to Z-up, while glTF specifies Y-up; a USD camera looks down −Z with +Y up, which differs from ROS's optical frame by a 180° rotation about x. Importing a mesh, motion-capture data, or a point cloud without reconciling these conventions leads to models lying on their side, mirrored left-right, or point clouds turned upside down.","example":"A depth camera's point cloud comes out in the optical frame, where z points forward; used directly as if it were the body frame (x forward, z up), objects on a table would appear to stand up in midair, as if rotated 90° — a fixed rotation has to be applied first, converting from the camera_optical frame to the camera_link frame.","related":["Coordinate Frame","Coordinate Transformation","Camera Coordinate Frame","TF / tf2 Transform Tree","REP 105","Quaternion Component Order (wxyz vs. xyzw)"]},{"id":"world-frame","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Modern Robotics (Lynch & Park, 2017), Ch.3 Rigid-Body Motions","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"ROS REP 105: Coordinate Frames for Mobile Platforms","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0105.rst"},{"title":"MuJoCo Documentation: Modeling","url":"https://mujoco.readthedocs.io/en/stable/modeling.html"}],"as_of":"","related_ids":["coordinate-frame","base-frame","body-frame","camera-extrinsics","coordinate-transformation","rep-105"],"name":"World Frame","alt":"世界坐标系","abbr":"","aliases":["Space Frame","Fixed Frame"],"one_liner":"The single, unmoving reference frame shared by an entire scene, to which all other positions can be converted.","explanation":"The world frame is a reference frame fixed in the environment, independent of the robot's own motion; Modern Robotics calls it the fixed frame or space frame {s}, and it might be anchored to a corner of a room or to a tabletop. A robot system also has a base frame (fixed to the robot's base), a body frame (moving with the robot's body), camera frames, and more; each sensor's readings start out in its own local frame and have to go through a coordinate transformation into a shared reference frame — usually the world frame — before they can be planned against together. Simulators typically have an explicit world root node, such as MuJoCo's worldbody; ROS's REP 105 defines map as the frame fixed to the world, with z pointing up. For an arm bolted to a fixed base, code often just treats the base frame as the world frame directly.","example":"For tabletop grasping, the world frame's origin is set on the tabletop with z pointing up: once the camera sees the cup, its coordinates are converted from the camera frame to the world frame using the camera's extrinsics, and the arm plans its grasp from there.","related":["Coordinate Frame","Base Frame","Body Frame","Camera Extrinsics","Coordinate Transformation","REP 105"]},{"id":"base-frame","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"ROS REP 105: Coordinate Frames for Mobile Platforms","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0105.rst"},{"title":"ROS REP 103: Standard Units of Measure and Coordinate Conventions","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0103.rst"},{"title":"libfranka robot_state.h（O_T_EE: end effector pose in base frame）","url":"https://raw.githubusercontent.com/frankarobotics/libfranka/main/include/franka/robot_state.h"}],"as_of":"","related_ids":["world-frame","coordinate-frame","end-effector-pose","coordinate-transformation","tf-tf2-transform-tree","body-frame"],"name":"Base Frame","alt":"基坐标系","abbr":"","aliases":["base_link","Robot Base Frame"],"one_liner":"The coordinate frame rigidly fixed to a robot's base, used as the reference for the arm's position and orientation.","explanation":"The base frame is a reference frame rigidly fixed to the robot's base; its origin and axis directions are set by the manufacturer, and the ROS convention points x forward, y left, and z up. Unless stated otherwise, an arm controller reports the end-effector's position and orientation relative to this frame — Franka's interface, for instance, stores the end-effector pose as O_T_EE, meaning “the end-effector's pose in base frame O.” In ROS, the corresponding frame is called base_link. The base frame differs from the world frame in one key way: when an arm is mounted on a mobile base or a humanoid body, the base frame moves along with the robot, while the world frame stays fixed in the environment. When recording demonstration data, it's essential to record which frame the actions are expressed in — otherwise the data won't transfer cleanly to a different robot.","example":"For a stationary tabletop arm, the command “move the end-effector to (0.5, 0, 0.3) meters in the base frame” means, under the x-forward, z-up convention, 0.5 meters directly in front of the base and 0.3 meters up.","related":["World Frame","Coordinate Frame","End-Effector Pose","Coordinate Transformation","TF / tf2 Transform Tree","Body Frame"]},{"id":"body-frame","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"Modern Robotics（Lynch & Park, 2017 预印本）第 3 章 Rigid-Body Motions","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"isaaclab.envs.mdp API（base_lin_vel / projected_gravity 等观测）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.envs.mdp.html"}],"as_of":"","related_ids":["world-frame","base-frame","coordinate-transformation","projected-gravity","pose","velocity-command-tracking"],"name":"Body Frame","alt":"机体坐标系","abbr":"","aliases":["Robot Body Frame"],"one_liner":"A coordinate frame fixed to a robot's own body, moving and rotating along with it.","explanation":"The body frame — often written {b} in textbooks — is a coordinate frame attached to a rigid body or a robot's torso, with its origin and three axes translating and rotating right along with the body; its counterpart is the fixed world frame. Describing an object's position and orientation is, at bottom, just giving the body frame's pose relative to the world frame; quantities like velocity and force can likewise be expressed in either the world frame or the body frame. Reinforcement-learning policies for legged robots usually put observations in the body frame: in Isaac Lab, base_lin_vel, base_ang_vel, and projected_gravity (the direction of gravity projected into the body frame, used to sense body tilt) are all defined relative to the robot's root frame, so the policy never has to care which way the robot happens to be facing in the world.","example":"Commanding a quadruped robot to “walk forward at 0.5 m/s” gives that velocity in the body frame, where x means forward relative to the body's own facing direction — whether the robot is actually facing east or north in the world, the same command always means “walk toward your own front.”","related":["World Frame","Base Frame","Coordinate Transformation","Projected Gravity","Pose","Velocity Command Tracking"]},{"id":"pose","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Pose (computer vision)","url":"https://en.wikipedia.org/wiki/Pose_(computer_vision)"},{"title":"ROS 2 common_interfaces: geometry_msgs/msg/Pose.msg","url":"https://raw.githubusercontent.com/ros2/common_interfaces/rolling/geometry_msgs/msg/Pose.msg"}],"as_of":"","related_ids":["end-effector-pose","coordinate-frame","rotation-matrix","quaternion","homogeneous-transformation-matrix","6d-object-pose-estimation"],"name":"Pose","alt":"位姿","abbr":"","aliases":["6-DoF Pose","6D Pose"],"one_liner":"An object's position plus orientation in space — six degrees of freedom in 3D.","explanation":"Pose is shorthand for “position plus orientation.” In 3D space, position takes three coordinates — x, y, z — and orientation adds three more rotational degrees of freedom, giving the 6D pose (or 6-DoF pose) that comes up constantly in robotics. A pose is always relative to some coordinate frame — the same cup has different numeric values in the camera frame versus the world frame. Orientation can be stored as a rotation matrix, a quaternion, or Euler angles, and position and rotation are often combined into a single 4×4 homogeneous transformation matrix. A robot typically has to estimate an object's pose before it can grasp it, arm control commonly targets an end-effector pose, and the actions a VLA model outputs are also often an end-effector pose or a delta to one.","example":"ROS's geometry_msgs/Pose message has two parts, position (x, y, z, in meters) and orientation (a quaternion x, y, z, w) — together, exactly one pose.","related":["End-Effector Pose","Coordinate Frame","Rotation Matrix","Quaternion","Homogeneous Transformation Matrix","6D Object Pose Estimation"]},{"id":"rotation-matrix","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Rotation matrix","url":"https://en.wikipedia.org/wiki/Rotation_matrix"},{"title":"Modern Robotics (Lynch & Park, 2017), 3.2.1 Rotation Matrices","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["special-orthogonal-group-so","homogeneous-transformation-matrix","quaternion","axis-angle-representation","6d-rotation-representation","coordinate-transformation"],"name":"Rotation Matrix","alt":"旋转矩阵","abbr":"","aliases":["Direction Cosine Matrix (DCM)"],"one_liner":"A 3×3 orthogonal matrix that represents a 3D rotation; multiplying it by a vector rotates that vector.","explanation":"The rotation matrix is the most basic way to describe 3D orientation. The three columns of the 3×3 matrix R are the directions — as unit vectors, in the reference frame — of the object's own x, y, and z axes; each entry equals the cosine of the angle between two axes, which is why it's also called the direction cosine matrix. It must satisfy RᵀR = I (the columns are mutually perpendicular and unit length) and det R = 1; the set of all matrices meeting this condition forms the special orthogonal group SO(3). Of its 9 numbers, only 3 are independent — there's redundancy — but unlike Euler angles it has no singular points, and the math is simple: rotating a vector is just R times that vector, chaining two rotations is just matrix multiplication (order matters), and inverting a rotation is just a transpose. A rotation matrix plus a translation gives the 4×4 homogeneous transformation matrix.","example":"The rotation matrix for a rotation of θ about the z-axis is [[cosθ, −sinθ, 0], [sinθ, cosθ, 0], [0, 0, 1]]; at θ = 90°, the vector (1, 0, 0), originally pointing along x, rotates to (0, 1, 0).","related":["Special Orthogonal Group SO(3)","Homogeneous Transformation Matrix","Quaternion","Axis-Angle Representation","6D Rotation Representation","Coordinate Transformation"]},{"id":"special-orthogonal-group-so","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"Wikipedia: 3D rotation group","url":"https://en.wikipedia.org/wiki/3D_rotation_group"},{"title":"Modern Robotics (Lynch & Park), Ch.3 Rigid-Body Motions (Definitions 3.1, 3.13)","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Wikipedia: Euclidean group","url":"https://en.wikipedia.org/wiki/Euclidean_group"}],"as_of":"","related_ids":["rotation-matrix","lie-group","exponential-map","skew-symmetric-matrix","homogeneous-transformation-matrix","6d-rotation-representation"],"name":"Special Orthogonal Group SO(3)","alt":"特殊正交群 SO(3)","abbr":"SO(3)","aliases":["SO(3)","SO3","Special Euclidean Group SE(3)","Lie Algebra"],"one_liner":"The set of every rotation in 3D space, whose elements are orthogonal matrices with determinant 1.","explanation":"SO(3) is the group of every rotation in 3D space: its elements are 3×3 rotation matrices satisfying RᵀR=I (Rᵀ is the transpose, I is the identity — this is what “orthogonal” means) and det R=1 (determinant 1, which rules out mirror reflections). Multiplying two rotations together gives another rotation, but the order matters. SO(3) is a 3-dimensional curved space (a manifold), not an ordinary vector space, so rotations can't simply be added together or averaged — which is exactly why methods like quaternions, the 6D representation, and spherical interpolation exist. Folding in translation as well gives the special Euclidean group SE(3): a rigid-body pose represented by a 4×4 homogeneous transformation matrix, with 6 degrees of freedom in total. Both are Lie groups; the tangent space at their identity element is called a Lie algebra, written so(3) and se(3) — so(3) turns out to be exactly the set of 3×3 skew-symmetric matrices, which correspond one-to-one with 3D angular velocity vectors, and convert back to a rotation through the exponential map.","example":"An arm's end-effector pose is a 4×4 matrix in SE(3): the upper-left 3×3 block is an SO(3) rotation, and the upper-right column is position. A matrix obtained by having a neural network directly regress 9 numbers generally won't satisfy RᵀR=I, so it needs to be projected back onto SO(3) using SVD or Gram-Schmidt orthogonalization.","related":["Rotation Matrix","Lie Group","Exponential Map","Skew-Symmetric Matrix","Homogeneous Transformation Matrix","6D Rotation Representation"]},{"id":"coordinate-transformation","category":"mechanics","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Transformation matrix（Affine transformations / homogeneous coordinates）","url":"https://en.wikipedia.org/wiki/Transformation_matrix"},{"title":"ROS REP 105: Coordinate Frames for Mobile Platforms","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0105.rst"},{"title":"Modern Robotics（Lynch & Park）Ch.3 Rigid-Body Motions","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["homogeneous-transformation-matrix","rotation-matrix","tf-tf2-transform-tree","hand-eye-calibration","camera-extrinsics","forward-kinematics"],"name":"Coordinate Transformation","alt":"坐标变换","abbr":"","aliases":["Frame Transformation","Rigid Transformation","Pose Transformation"],"one_liner":"Converting a point's or pose's numeric value from one coordinate frame into another.","explanation":"A coordinate transformation answers the question: given this point's coordinates in frame A, what are its coordinates in frame B? Two frames differ by a rotation plus a translation: p_B = R·p_A + t, where R is a 3×3 rotation matrix (the difference in orientation) and t is a translation vector (the offset between origins). In practice, R and t are usually packed into a single 4×4 homogeneous transformation matrix T, so translation also becomes a matrix multiplication, and chaining several transforms together is just multiplying matrices, as in T_base_obj = T_base_cam · T_cam_obj. This comes up everywhere in robotics: an object a camera sees has to be converted into the base frame before the arm can reach for it, and forward kinematics is nothing more than chaining transformation matrices link by link along the arm. ROS maintains and looks up these transforms through its TF library.","example":"Hand-eye calibration gives the camera's pose in the base frame, T_base_cam; a vision model gives the cup's position in the camera frame. Multiplying the two together gives the cup's position in the base frame, which the arm uses to plan its grasp.","related":["Homogeneous Transformation Matrix","Rotation Matrix","TF / tf2 Transform Tree","Hand-Eye Calibration","Camera Extrinsics","Forward Kinematics (FK)"]},{"id":"homogeneous-transformation-matrix","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"Modern Robotics 3.3.1: Homogeneous Transformation Matrices (Northwestern)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-1-homogeneous-transformation-matrices/"},{"title":"Modern Robotics (Lynch & Park) preprint PDF","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["rotation-matrix","coordinate-transformation","pose","forward-kinematics","hand-eye-calibration","tf-tf2-transform-tree"],"name":"Homogeneous Transformation Matrix","alt":"齐次变换矩阵","abbr":"","aliases":["SE(3) Matrix","Rigid Transformation"],"one_liner":"A 4×4 matrix combining rotation and translation into one, describing one frame's pose relative to another.","explanation":"The homogeneous transformation matrix T is the standard way robotics represents pose: the upper-left 3×3 block is the rotation matrix R (orientation), the upper-right 3×1 block is the translation vector p (position), and the bottom row is fixed at [0 0 0 1]. The set of all such matrices forms the special Euclidean group SE(3). Adding that bottom row lets a point, written as [x, y, z, 1], have rotation and translation applied together in a single matrix multiplication. It's used three ways: to represent frame {b}'s pose relative to {s} as T_sb; to change reference frames, chaining transforms by canceling matching subscripts, as in T_sc = T_sb·T_bc; and to apply a translation and rotation to a frame or an object. Forward kinematics is exactly the product of each joint's transform, giving the end effector's pose relative to the base; ROS's TF tree also stores poses this way. Note that matrix multiplication doesn't commute, so left-multiplying and right-multiplying mean different things.","example":"A camera mounted on an arm's wrist: given the base-to-end-effector transform T_base_ee and the end-effector-to-camera transform T_ee_cam (from hand-eye calibration), multiplying them gives T_base_cam, which converts whatever point the camera sees into the base frame so the arm can grasp it.","related":["Rotation Matrix","Coordinate Transformation","Pose","Forward Kinematics (FK)","Hand-Eye Calibration","TF / tf2 Transform Tree"]},{"id":"projected-gravity","category":"mechanics","sec":0,"tier":2,"sources":[{"title":"Isaac Lab API: isaaclab.envs.mdp (projected_gravity)","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.envs.mdp.html"},{"title":"legged_gym (ETH RSL): legged_robot.py","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"unitree_rl_gym: deploy_real.py","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/deploy/deploy_real/deploy_real.py"}],"as_of":"","related_ids":["proprioception","inertial-measurement-unit","body-frame","quaternion","rl-based-locomotion-control","roll-pitch-yaw"],"name":"Projected Gravity","alt":"投影重力","abbr":"","aliases":["Projected Gravity Vector","projected_gravity","projected_gravity_b"],"one_liner":"Gravity's straight-down direction expressed in the robot's own body frame, telling a policy how far the body is tilted.","explanation":"Projected gravity is a near-universal observation in reinforcement learning for legged and humanoid robots: it takes the unit gravity direction in the world frame, (0, 0, −1), and rotates it into the body (base) frame using the inverse of the torso's orientation quaternion. When the robot stands perfectly upright, this equals (0, 0, −1); when the body tips forward or sideways, the x and y components become nonzero, with their magnitude reflecting how much pitch and roll are present. It carries only tilt information, with no heading (yaw) — which is useful, since yaw does nothing for balance and tends to drift on a real robot — so it's a cleaner input than feeding in the raw orientation quaternion directly. On real hardware, it's computed from the orientation an IMU estimates. Both legged_gym and Isaac Lab include it in their observations, and reward functions commonly penalize the squared sum of its x and y components to keep the body level.","example":"Unitree's unitree_rl_gym real-robot deployment script reads the IMU quaternion, computes projected gravity, and places it in dimensions 4–6 of the observation, alongside angular velocity, velocity commands, and joint angles, as input to the walking policy.","related":["Proprioception","Inertial Measurement Unit","Body Frame","Quaternion","RL-based Locomotion Control","Roll-Pitch-Yaw (RPY)"]},{"id":"euler-angles","category":"mechanics","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Euler angles","url":"https://en.wikipedia.org/wiki/Euler_angles"},{"title":"Wikipedia: Gimbal lock","url":"https://en.wikipedia.org/wiki/Gimbal_lock"},{"title":"On the Continuity of Rotation Representations in Neural Networks（Zhou et al.）","url":"https://arxiv.org/abs/1812.07035"}],"as_of":"","related_ids":["roll-pitch-yaw","gimbal-lock","quaternion","rotation-matrix","intrinsic-vs-extrinsic-rotations","6d-rotation-representation"],"name":"Euler Angles","alt":"欧拉角","abbr":"","aliases":[],"one_liner":"A way of describing an object's orientation using three angles of rotation applied one after another about coordinate axes.","explanation":"Euler angles, introduced by the Swiss mathematician Leonhard Euler, describe a rigid body's orientation through three successive rotations about coordinate axes, commonly written α, β, γ or φ, θ, ψ. There are 12 possible rotation-order sequences: ones where the first and last axis are the same (like Z-X-Z) are the classical Euler angles, and ones where all three axes differ (like Z-Y-X) are called Tait-Bryan angles — the familiar roll, pitch, and yaw fall into this second category. Rotating about axes that move with the object is called intrinsic rotation; rotating about fixed axes is called extrinsic. Euler angles are intuitive and easy to read, but the same three numbers describe a different orientation depending on the rotation order and the intrinsic/extrinsic convention used, and at certain angles two of the rotation axes line up and a degree of freedom is lost — a problem called gimbal lock. For this reason, code typically represents rotation internally with quaternions or rotation matrices, and neural networks predicting rotation often use a 6D representation instead.","example":"An aircraft that turns 30° to the right, then pitches its nose up 10°, with no roll, can be written as a Z-Y-X sequence of Euler angles: yaw 30°, pitch 10°, roll 0°.","related":["Roll-Pitch-Yaw (RPY)","Gimbal Lock","Quaternion","Rotation Matrix","Intrinsic vs. Extrinsic Rotations","6D Rotation Representation"]},{"id":"roll-pitch-yaw","category":"mechanics","sec":1,"tier":1,"sources":[{"title":"Modern Robotics (Lynch & Park, 2017), Appendix B.2 Roll–Pitch–Yaw Angles","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"ROS REP 103: Standard Units of Measure and Coordinate Conventions","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0103.rst"},{"title":"Wikipedia: Aircraft principal axes","url":"https://en.wikipedia.org/wiki/Aircraft_principal_axes"}],"as_of":"","related_ids":["euler-angles","gimbal-lock","intrinsic-vs-extrinsic-rotations","rotation-matrix","quaternion","inertial-measurement-unit"],"name":"Roll-Pitch-Yaw (RPY)","alt":"横滚-俯仰-偏航角","abbr":"RPY","aliases":["RPY","RPY Angles"],"one_liner":"Orientation described by three rotation angles about the x, y, and z axes: roll, pitch, and yaw.","explanation":"This naming comes from aviation and maritime navigation: roll is tilting about the front-to-back axis, pitch is tipping the nose up or down about the side-to-side axis, and yaw is turning about the vertical axis. ROS's REP 103 convention points a body's x-axis forward, y-axis left, and z-axis up; RPY means rotating in sequence about the fixed frame's x, y, and z axes by angles γ, β, and α, giving the rotation matrix R = Rz(α)·Ry(β)·Rx(γ) — where Rz(α) is a rotation of α about z — which comes out identical to Z-Y-X Euler angles. RPY is intuitive and easy to tune by hand, and it's exactly what URDF's rpy attribute uses. The downside is that roll and yaw become coupled when pitch nears ±90° (gimbal lock), and the angles jump discontinuously near ±180°, which is why code usually stores orientation as a quaternion internally instead.","example":"A humanoid robot's IMU (inertial measurement unit) commonly reports roll, pitch, and yaw directly: a sudden jump in roll or pitch means the body is tipping sideways or pitching forward, and yaw tells you which way the body is facing.","related":["Euler Angles","Gimbal Lock","Intrinsic vs. Extrinsic Rotations","Rotation Matrix","Quaternion","Inertial Measurement Unit"]},{"id":"intrinsic-vs-extrinsic-rotations","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Davenport chained rotations（Intrinsic / Extrinsic rotations）","url":"https://en.wikipedia.org/wiki/Davenport_chained_rotations"},{"title":"SciPy 文档：Rotation.from_euler（大写为内旋、小写为外旋）","url":"https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_euler.html"},{"title":"Lynch & Park, Modern Robotics（2017）附录 B：ZYX 欧拉角与 roll-pitch-yaw 角","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["euler-angles","roll-pitch-yaw","rotation-matrix","gimbal-lock","coordinate-transformation","quaternion"],"name":"Intrinsic vs. Extrinsic Rotations","alt":"内旋与外旋","abbr":"","aliases":["Intrinsic Rotation","Extrinsic Rotation","Euler Angle Rotation Order Convention"],"one_liner":"For successive rotations, whether each step turns about the new, moving axes (intrinsic) or the original, fixed axes (extrinsic).","explanation":"Describing orientation with three angles (Euler angles) requires a convention for which axis each step rotates about. Intrinsic rotation turns about the object's own, moving coordinate axes, so the axis used at each subsequent step has already changed with the previous rotation; extrinsic rotation always turns about the fixed axes of the world frame. The two are interchangeable: applying the same three angles intrinsically in the order Z→Y→X gives the same rotation matrix as applying them extrinsically in the reverse order X→Y→Z, namely R = Rz(α)Ry(β)Rx(γ) (Rz(α) denotes the rotation matrix for an angle α about the z-axis, and so on). The ZYX Euler angles commonly used in robotics are intrinsic, while roll-pitch-yaw (RPY) angles are usually defined as extrinsic XYZ, which is why the two end up with the same numerical values. Conventions differ across software and datasets, so before reading someone else's pose data, it's important to check whether it's intrinsic or extrinsic and what the axis order is.","example":"SciPy's Rotation.from_euler uses uppercase 'ZYX' for intrinsic rotation and lowercase 'zyx' for extrinsic; the same set of angles written as 'XYZ' versus 'xyz' produces a different orientation, so a single letter's case can send a robot arm's end effector spinning the wrong way.","related":["Euler Angles","Roll-Pitch-Yaw (RPY)","Rotation Matrix","Gimbal Lock","Coordinate Transformation","Quaternion"]},{"id":"gimbal-lock","category":"mechanics","sec":1,"tier":2,"sources":[{"title":"Gimbal lock - Wikipedia","url":"https://en.wikipedia.org/wiki/Gimbal_lock"}],"as_of":"","related_ids":["euler-angles","roll-pitch-yaw","quaternion","6d-rotation-representation","rotation-matrix","singular-configuration"],"name":"Gimbal Lock","alt":"万向节死锁","abbr":"","aliases":["Gimbal Lock Singularity"],"one_liner":"When representing rotation with Euler angles, one angle reaching 90° makes two rotation axes coincide, losing a degree of freedom.","explanation":"Gimbal lock was originally a mechanical problem: when a three-ring nested gimbal reaches a certain angle, two of its rotation axes become parallel, and the mounted platform loses its ability to rotate in one direction. The same phenomenon shows up mathematically when orientation is represented with Euler angles or roll-pitch-yaw angles: when pitch reaches ±90°, yaw and roll end up rotating about the same axis, so adjusting either one has the same effect, and certain small rotations nearby can't be expressed as a small change in the angles at all — interpolation and differentiation both develop discontinuities near this point. The inertial measurement platform on Apollo 11 was limited by exactly this: by design, the platform would lock once pitch neared 85°, requiring the crew to manually steer the spacecraft away from that attitude. The standard fix in robotics is to store and compute orientation with quaternions, rotation matrices, or a 6D rotation representation, and reserve Euler angles for human-readable display only; a policy trained to regress Euler angles directly can also learn discontinuous actions near these singular points.","example":"Controlling an arm's end-effector orientation with roll-pitch-yaw angles, setting pitch to 90° reveals that adjusting yaw and adjusting roll both rotate the gripper about the same axis — achieving a small rotation in the other direction requires all three angles to jump by a large amount at once.","related":["Euler Angles","Roll-Pitch-Yaw (RPY)","Quaternion","6D Rotation Representation","Rotation Matrix","Singular Configuration (Kinematic Singularity)"]},{"id":"quaternion","category":"mechanics","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Quaternions and spatial rotation","url":"https://en.wikipedia.org/wiki/Quaternions_and_spatial_rotation"},{"title":"Wikipedia: Quaternion（History）","url":"https://en.wikipedia.org/wiki/Quaternion"},{"title":"ROS 2 common_interfaces: geometry_msgs/msg/Quaternion.msg","url":"https://raw.githubusercontent.com/ros2/common_interfaces/rolling/geometry_msgs/msg/Quaternion.msg"},{"title":"MuJoCo Documentation: Modeling（Frame orientations：quat 默认 1 0 0 0，实部在前）","url":"https://mujoco.readthedocs.io/en/stable/modeling.html"}],"as_of":"","related_ids":["rotation-matrix","euler-angles","gimbal-lock","quaternion-double-cover","spherical-linear-interpolation","quaternion-component-order"],"name":"Quaternion","alt":"四元数","abbr":"","aliases":["Unit Quaternion"],"one_liner":"A four-number representation of 3D rotation, the most common orientation format in robotics software.","explanation":"Quaternions were introduced by William Rowan Hamilton in 1843, in the form w + xi + yj + zk. A unit quaternion — one with length 1 — can represent a 3D rotation: rotating by angle θ about a unit axis u corresponds to q = (cos(θ/2), u·sin(θ/2)), where the first term is the real part w and the remaining three are the imaginary parts x, y, z. Unlike Euler angles, quaternions have no gimbal lock (the loss of a rotational degree of freedom at certain angles), and compared with a 9-number rotation matrix they're more compact and interpolate smoothly — both ROS and MuJoCo store orientation this way. One catch: q and −q represent the exact same rotation. Component ordering also varies between tools — MuJoCo uses wxyz, ROS messages use xyzw — and mixing the two up is a common bug.","example":"A 90° rotation about the z-axis: θ/2 = 45°, which in wxyz order is (0.707, 0, 0, 0.707); filled into a ROS Quaternion message, that becomes x=0, y=0, z=0.707, w=0.707.","related":["Rotation Matrix","Euler Angles","Gimbal Lock","Quaternion Double Cover","Spherical Linear Interpolation (SLERP)","Quaternion Component Order (wxyz vs. xyzw)"]},{"id":"quaternion-component-order","category":"mechanics","sec":1,"tier":2,"sources":[{"title":"Joan Solà, Quaternion kinematics for the error-state Kalman filter (arXiv:1711.02508)","url":"https://arxiv.org/abs/1711.02508"},{"title":"Isaac Lab: Migrating to Isaac Lab 3.0 (Quaternion Format)","url":"https://github.com/isaac-sim/IsaacLab/blob/develop/docs/source/migration/migrating_to_isaaclab_3-0.rst"},{"title":"SciPy: Rotation.from_quat","url":"https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.from_quat.html"}],"as_of":"2026-09","related_ids":["quaternion","quaternion-double-cover","right-handed-frame-and-axis-conventions","coordinate-transformation","mujoco","nvidia-isaac-lab"],"name":"Quaternion Component Order (wxyz vs. xyzw)","alt":"四元数分量顺序约定（wxyz / xyzw）","abbr":"","aliases":["Scalar-First / Scalar-Last","wxyz","xyzw","Hamilton Convention","JPL Convention"],"one_liner":"Whether a quaternion's real part w is stored first or last — a convention that differs across software libraries.","explanation":"A unit quaternion represents rotation with four numbers: a real part w and imaginary parts x, y, z. The component order convention is whether w is stored first (wxyz, “scalar-first”) or last (xyzw, “scalar-last”) — and software isn't consistent about it: ROS messages and SciPy default to xyzw, while MuJoCo and Isaac Sim Core use wxyz; Isaac Lab 2.x used wxyz, but 3.0 switched to xyzw to align with Warp, PhysX, and Newton. A deeper-level split is the Hamilton versus JPL convention: according to Joan Solà's survey, Hamilton's convention defines ij=k (right-handed), while JPL's defines ji=k (left-handed), so the exact same four numbers represent opposite rotations under the two conventions; robotics mostly uses Hamilton, while aerospace often uses JPL. Getting the order wrong when combining data from different sources is a classic beginner mistake — the code runs fine, but every orientation comes out wrong.","example":"“No rotation” is written (1, 0, 0, 0) in wxyz, but (0, 0, 0, 1) in xyzw. Read a ROS-recorded (0, 0, 0, 1) into MuJoCo as if it were wxyz, and it gets interpreted as a 180° rotation about z — the robot starts out facing backward.","related":["Quaternion","Quaternion Double Cover","Right-Handed Frame & Axis Conventions","Coordinate Transformation","MuJoCo (Multi-Joint dynamics with Contact)","NVIDIA Isaac Lab"]},{"id":"spherical-linear-interpolation","category":"mechanics","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Slerp","url":"https://en.wikipedia.org/wiki/Slerp"},{"title":"SciPy: scipy.spatial.transform.Slerp","url":"https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Slerp.html"}],"as_of":"","related_ids":["quaternion","quaternion-double-cover","trajectory-interpolation","special-orthogonal-group-so","geodesic-distance-on-so","euler-angles"],"name":"Spherical Linear Interpolation (SLERP)","alt":"球面线性插值","abbr":"SLERP","aliases":["Slerp","Quaternion Slerp"],"one_liner":"A way to smoothly blend between two rotations along the shortest arc, at a constant angular speed.","explanation":"Spherical linear interpolation (Slerp) generates the intermediate rotations between two given rotations; Ken Shoemake introduced it to computer graphics in a 1985 SIGGRAPH paper. Treating a unit quaternion as a point on the unit sphere in four dimensions, Slerp moves along the great-circle arc connecting the two points: slerp(q₀,q₁;t) = [sin((1−t)Ω)/sinΩ]·q₀ + [sin(tΩ)/sinΩ]·q₁, where t runs from 0 to 1 as progress along the interpolation and Ω is the angle between the two quaternions. The result is a rotation about a fixed axis at constant angular speed. Interpolating quaternions or Euler angles linearly instead produces uneven rotation speed, or even takes the long way around; and because q and −q represent the same rotation, an implementation has to check whether the two quaternions' dot product is negative and flip one of them first, to guarantee it takes the short arc. In robotics, it's commonly used to interpolate end-effector orientation trajectories, to upsample a low-frequency policy's output into high-frequency control commands, and to resample motion-capture data.","example":"If a policy outputs a target end-effector orientation at 10 Hz but the low-level controller runs at 500 Hz, SciPy's Slerp class can interpolate 50 intermediate orientations between each pair of consecutive targets, so the end effector rotates there at a constant rate.","related":["Quaternion","Quaternion Double Cover","Trajectory Interpolation","Special Orthogonal Group SO(3)","Geodesic Distance on SO(3)","Euler Angles"]},{"id":"quaternion-double-cover","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Quaternions and spatial rotation","url":"https://en.wikipedia.org/wiki/Quaternions_and_spatial_rotation"},{"title":"Wikipedia: 3D rotation group（S³ 到 SO(3) 的二对一覆盖）","url":"https://en.wikipedia.org/wiki/3D_rotation_group"},{"title":"On the Continuity of Rotation Representations in Neural Networks (arXiv 1812.07035)","url":"https://arxiv.org/abs/1812.07035"}],"as_of":"","related_ids":["quaternion","6d-rotation-representation","spherical-linear-interpolation","geodesic-distance-on-so","special-orthogonal-group-so","quaternion-component-order"],"name":"Quaternion Double Cover","alt":"四元数双倍覆盖","abbr":"","aliases":["q and −q Equivalence","Quaternion Sign Ambiguity"],"one_liner":"The unit quaternions q and −q represent the exact same rotation, so every orientation corresponds to two quaternions.","explanation":"When a unit quaternion q = (w, x, y, z) represents a rotation, q and −q produce exactly the same rotation matrix, because the formula converting a quaternion to a rotation is quadratic in q, so the signs cancel out. Mathematically, the map from the 3-sphere S³ of unit quaternions to the rotation group SO(3) is two-to-one (a double cover), with each orientation corresponding to a pair of antipodal points on the sphere. This is a common pitfall in embodied AI: the same end-effector orientation might appear in a dataset sometimes as q and sometimes as −q, and if a network regresses quaternions directly with an ordinary L2 loss, it will treat identical orientations as if they were very different. Orientation difference should instead be measured with 1 − |q₁·q₂|, or by taking whichever is smaller of ‖q − q̂‖ and ‖q + q̂‖. Before doing spherical interpolation, if the dot product between two quaternions is negative, one should be flipped first, or the interpolation will take the long way around. Zhou and colleagues (2019) further proved that rotation representations of four dimensions or fewer are all discontinuous, which is why many policies switch to a 6D rotation representation instead.","example":"The quaternion for a 90° rotation about the z-axis is q = (0.707, 0, 0, 0.707) (w first); −q = (−0.707, 0, 0, −0.707) is equivalent to the same rotation plus one full extra 360° turn — the orientation is identical. A common normalization is to simply enforce w ≥ 0.","related":["Quaternion","6D Rotation Representation","Spherical Linear Interpolation (SLERP)","Geodesic Distance on SO(3)","Special Orthogonal Group SO(3)","Quaternion Component Order (wxyz vs. xyzw)"]},{"id":"6d-rotation-representation","category":"mechanics","sec":1,"tier":2,"sources":[{"title":"On the Continuity of Rotation Representations in Neural Networks (Zhou et al., CVPR 2019)","url":"https://arxiv.org/abs/1812.07035"},{"title":"real-stanford/diffusion_policy: rotation_transformer.py","url":"https://raw.githubusercontent.com/real-stanford/diffusion_policy/main/diffusion_policy/model/common/rotation_transformer.py"}],"as_of":"","related_ids":["rotation-matrix","quaternion","euler-angles","9d-rotation-representation","action-representation","diffusion-policy"],"name":"6D Rotation Representation","alt":"6D 旋转表示","abbr":"6D","aliases":["rotation_6d","Continuous 6D Rotation Representation"],"one_liner":"Representing a rotation with the first two columns of its rotation matrix — 6 numbers — a format well suited to neural-network learning.","explanation":"The 6D rotation representation was introduced by Zhou et al. in their CVPR 2019 paper “On the Continuity of Rotation Representations in Neural Networks.” They showed that 3D rotations have no continuous representation in Euclidean spaces of 4 dimensions or fewer — Euler angles and quaternions both have discontinuities (for instance, q and −q represent the same rotation), which causes larger errors when a network regresses these targets directly — while a continuous representation does exist in 5 or 6 dimensions. The method keeps only the first two columns of the rotation matrix. To recover a full rotation, the first column is normalized, the component of the second column along the first is subtracted off and the result is normalized, and the third column is taken as the cross product of the two — a process similar to Gram-Schmidt orthogonalization. It's a common choice in robot learning for representing end-effector orientation, whether as an action or an observation.","example":"Diffusion Policy's official code uses a RotationTransformer that by default converts axis-angle rotations to rotation_6d for handling end-effector orientation.","related":["Rotation Matrix","Quaternion","Euler Angles","9D Rotation Representation","Action Representation","Diffusion Policy"]},{"id":"9d-rotation-representation","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"An Analysis of SVD for Deep Rotation Estimation (arXiv 2006.14616, NeurIPS 2020)","url":"https://arxiv.org/abs/2006.14616"},{"title":"google-research/special_orthogonalization README","url":"https://github.com/google-research/google-research/tree/master/special_orthogonalization"},{"title":"On the Continuity of Rotation Representations in Neural Networks (arXiv 1812.07035)","url":"https://arxiv.org/abs/1812.07035"}],"as_of":"","related_ids":["6d-rotation-representation","rotation-matrix","quaternion","special-orthogonal-group-so","geodesic-distance-on-so","euler-angles"],"name":"9D Rotation Representation","alt":"9D 旋转表示","abbr":"","aliases":["SVD Orthogonalization","Symmetric Orthogonalization","9D Rotation"],"one_liner":"Lets a network output nine unconstrained numbers, then uses SVD to project them onto the nearest valid rotation matrix.","explanation":"The 9D rotation representation is a way for a neural network to predict a 3D rotation: the network outputs nine unconstrained real numbers arranged into a 3×3 matrix M, which is then factored by singular value decomposition, M = UΣVᵀ. Taking R = U·diag(1, 1, det(UVᵀ))·Vᵀ gives the rotation matrix closest to M, with the det term ensuring the result is not a mirror reflection. Levinson et al. gave a systematic analysis in their NeurIPS 2020 paper 'An Analysis of SVD for Deep Rotation Estimation': the mapping is smooth almost everywhere, unlike quaternions or Euler angles, which can jump discontinuously; under noisy inputs, the expected reconstruction error is about half that of the Gram-Schmidt orthogonalization used by the 6D representation; and it achieved state-of-the-art results at the time on tasks including point-cloud alignment, object pose estimation, and inverse kinematics. Whenever a network must directly regress an object's or end-effector's orientation, this is a common alternative to quaternions or Euler angles, alongside the 6D representation.","example":"The paper's official code implements this in about ten lines: reshape the network's [batch, 9] output into 3×3 matrices, run SVD, correct the sign of the last column using det(UVᵀ), and obtain a [batch, 3, 3] rotation matrix that can be attached directly to the end of any regression network for end-to-end training.","related":["6D Rotation Representation","Rotation Matrix","Quaternion","Special Orthogonal Group SO(3)","Geodesic Distance on SO(3)","Euler Angles"]},{"id":"geodesic-distance-on-so","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Zhou et al., On the Continuity of Rotation Representations in Neural Networks（geodesic error 定义）","url":"https://arxiv.org/abs/1812.07035"}],"as_of":"","related_ids":["special-orthogonal-group-so","rotation-matrix","quaternion-double-cover","6d-rotation-representation","6d-object-pose-estimation","euler-angles"],"name":"Geodesic Distance on SO(3)","alt":"旋转测地距离","abbr":"","aliases":["Angular Distance","Geodesic Error","Rotation Angle Error"],"one_liner":"The minimum angle still needed to turn one orientation into another — the standard way to measure rotation error.","explanation":"The geodesic distance on SO(3) is the shortest distance between two rotations on the rotation group SO(3) (the set of all 3D rotations); numerically, it's the minimum rotation angle needed to turn one orientation into the other, ranging from 0 to π (180°). For rotation matrices R₁ and R₂, first compute the relative rotation R₁ᵀR₂, then take θ = arccos((tr(R₁ᵀR₂) − 1)/2), where tr is the sum of the diagonal entries; with unit quaternions q₁ and q₂ it can be written θ = 2·arccos(|q₁·q₂|), with the absolute value needed because q and −q represent the same rotation. Directly comparing the raw components of Euler angles or quaternions can be misled by the choice of representation; the geodesic distance only measures the actual angular difference. It's the standard metric reported as 'geodesic error' in 6D object pose estimation and rotation-representation research — such as the paper that introduced the 6D rotation representation by Zhou and colleagues — and it can also be used directly as a training loss.","example":"With R₁ the identity and R₂ a 90° rotation about the z-axis: tr(R₂) = 0+0+1 = 1, so θ = arccos(0) = 90°. Or take yaw angles of 179° and −179°: their raw difference is 358°, but the geodesic distance is only 2°.","related":["Special Orthogonal Group SO(3)","Rotation Matrix","Quaternion Double Cover","6D Rotation Representation","6D Object Pose Estimation","Euler Angles"]},{"id":"axis-angle-representation","category":"mechanics","sec":1,"tier":2,"sources":[{"title":"Axis–angle representation - Wikipedia","url":"https://en.wikipedia.org/wiki/Axis%E2%80%93angle_representation"},{"title":"scipy.spatial.transform.Rotation.as_rotvec - SciPy Manual","url":"https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.transform.Rotation.as_rotvec.html"},{"title":"Controllers - robosuite documentation","url":"https://robosuite.ai/docs/modules/controllers.html"}],"as_of":"","related_ids":["rotation-matrix","quaternion","euler-angles","rodrigues-rotation-formula","exponential-map","6d-rotation-representation"],"name":"Axis-Angle Representation","alt":"轴角","abbr":"","aliases":["Rotation Vector","rotvec","Euler Vector"],"one_liner":"Representing a 3D rotation by which axis it turns around and by how large an angle.","explanation":"Axis-angle representation describes a 3D rotation with a unit vector e (the rotation axis direction) and an angle θ (the rotation about that axis, in radians), based on the fact that every rotation is equivalent to a single turn about some fixed axis. Multiplying the two together gives the 3D rotation vector θe: its direction is the rotation axis and its length is the rotation angle, so it packs everything into just 3 numbers — SciPy calls this rotvec. It converts to and from a rotation matrix using Rodrigues' rotation formula, and mathematically it's just the coordinates of the SO(3) exponential map. Its downside is non-uniqueness: adding or subtracting 2π from θ, or flipping the sign of both the axis and the angle, represent the exact same rotation, and the conversion needs special handling when θ is near 0 or π. Because it's low-dimensional and smooth for small angles, it's a common action representation for end-effector orientation deltas in robot arm control.","example":"robosuite's OSC_POSE controller takes a 6-dimensional action (excluding the gripper); the last 3 values are a rotation delta relative to the current end-effector orientation, given in axis-angle form (ax, ay, az) — a 90° rotation about z is written (0, 0, 1.571).","related":["Rotation Matrix","Quaternion","Euler Angles","Rodrigues' Rotation Formula","Exponential Map","6D Rotation Representation"]},{"id":"rodrigues-rotation-formula","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Rodrigues' rotation formula","url":"https://en.wikipedia.org/wiki/Rodrigues%27_rotation_formula"},{"title":"Wikipedia: Olinde Rodrigues","url":"https://en.wikipedia.org/wiki/Olinde_Rodrigues"}],"as_of":"","related_ids":["axis-angle-representation","rotation-matrix","skew-symmetric-matrix","exponential-map","special-orthogonal-group-so","product-of-exponentials-formula"],"name":"Rodrigues' Rotation Formula","alt":"罗德里格斯公式","abbr":"","aliases":["Rodrigues Formula","Euler's Finite Rotation Formula"],"one_liner":"Given a rotation axis and an angle, this formula computes the corresponding 3D rotation matrix directly, in closed form.","explanation":"Rodrigues' rotation formula is named after the French mathematician Olinde Rodrigues, who published the relevant result in 1840, though some scholars trace the underlying idea to Euler. It answers the question: what rotation matrix corresponds to rotating by angle θ about the axis given by the unit vector k? In matrix form, R = I + sinθ·K + (1−cosθ)·K², where I is the 3×3 identity matrix and K is the skew-symmetric matrix built from k (multiplying K by any vector v equals the cross product k×v). It converts the axis-angle representation (an axis plus an angle describing a rotation) into a rotation matrix, and is essentially the closed-form solution to the exponential map on the rotation group SO(3): the infinite series for exp(θK) collapses exactly into these three terms. It's used throughout robotics — for example, in computing forward kinematics via the product of exponentials formula, or converting an axis-angle action into an orientation; OpenCV's function for converting a rotation vector into a rotation matrix is even named cv2.Rodrigues.","example":"Rotating 90° about the z-axis (k = (0,0,1)): sin90° = 1, cos90° = 0, so R = I + K + K². Working this out sends the x-axis direction (1,0,0) to the y-axis direction (0,1,0), matching intuition.","related":["Axis-Angle Representation","Rotation Matrix","Skew-Symmetric Matrix","Exponential Map","Special Orthogonal Group SO(3)","Product of Exponentials Formula"]},{"id":"skew-symmetric-matrix","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Skew-symmetric matrix","url":"https://en.wikipedia.org/wiki/Skew-symmetric_matrix"},{"title":"Wikipedia: Rodrigues' rotation formula","url":"https://en.wikipedia.org/wiki/Rodrigues%27_rotation_formula"}],"as_of":"","related_ids":["rodrigues-rotation-formula","rotation-matrix","special-orthogonal-group-so","lie-group","exponential-map","screw-theory"],"name":"Skew-Symmetric Matrix","alt":"反对称矩阵","abbr":"","aliases":["Cross-Product Matrix","Hat Operator","[ω]×"],"one_liner":"A matrix that equals the negative of its own transpose; in 3×3 form it turns a cross product into matrix multiplication.","explanation":"A skew-symmetric matrix satisfies Aᵀ = −A, and its diagonal entries are all zero. The most common case in robotics is the 3×3 version: given a vector ω = (ω₁, ω₂, ω₃), construct the matrix [ω] = [[0, −ω₃, ω₂], [ω₃, 0, −ω₁], [−ω₂, ω₁, 0]], and then [ω]v is exactly equal to the cross product ω×v. Turning a vector into this matrix is called the hat operation (written ω^ or [ω]×), and recovering the vector from the matrix is called vee. Its importance is that all 3×3 skew-symmetric matrices together form so(3), the Lie algebra of the rotation group SO(3), which can be thought of as the space of infinitesimal rotations; taking the matrix exponential of one gives a rotation matrix, and the closed-form result is Rodrigues' rotation formula. It comes up repeatedly when deriving angular velocity, the derivative of a rotation matrix, and Jacobian matrices.","example":"Take ω = (0,0,1) (a unit angular velocity about the z-axis) and v = (1,0,0): [ω]v = (0,1,0), matching the cross product ω×v — meaning the point on the x-axis is, at this instant, moving in the y direction.","related":["Rodrigues' Rotation Formula","Rotation Matrix","Special Orthogonal Group SO(3)","Lie Group","Exponential Map","Screw Theory"]},{"id":"lie-group","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Lie group","url":"https://en.wikipedia.org/wiki/Lie_group"},{"title":"Solà et al., A micro Lie theory for state estimation in robotics (arXiv:1812.01537)","url":"https://arxiv.org/abs/1812.01537"},{"title":"Lynch & Park, Modern Robotics（2017）第 3 章：SO(3)、SE(3) 与李代数","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["special-orthogonal-group-so","exponential-map","skew-symmetric-matrix","adjoint-representation","product-of-exponentials-formula","screw-theory"],"name":"Lie Group","alt":"李群","abbr":"","aliases":["Matrix Lie Group","Continuous Transformation Group"],"one_liner":"A mathematical object that is both a group and a smooth manifold; rotations and poses in robotics are both Lie groups.","explanation":"A Lie group, named after the Norwegian mathematician Sophus Lie (1842–1899), is a group that is also a smooth manifold, where both multiplication and inversion are smooth operations. Being a group guarantees elements can be composed (two rotations combine into one) and inverted; being a manifold (a surface that looks locally like flat space) guarantees you can take derivatives and run optimization on it. Robotics relies on two above all: SO(3), the set of all 3D rotation matrices, and SE(3), rigid-body poses combining rotation and translation. Their tangent spaces at the identity element are called Lie algebras — so(3), made up of skew-symmetric matrices, corresponds to SO(3) — and the two convert into each other through the exponential map. This makes it possible to add, subtract, and take gradients using ordinary 3D vectors, then map the result back to a rotation while guaranteeing the result is still a valid one. SLAM, visual odometry, state estimation, pose optimization, and the product of exponentials formula are all built on this foundation.","example":"When optimizing a camera's orientation in visual SLAM, the rotation matrix R's 9 numbers aren't updated directly; instead, a small 3D increment δ is solved for and applied as R ← R·exp([δ]×) ([δ]× is the skew-symmetric matrix built from δ, and exp is the matrix exponential), so the updated R is guaranteed to remain a valid rotation matrix.","related":["Special Orthogonal Group SO(3)","Exponential Map","Skew-Symmetric Matrix","Adjoint Representation","Product of Exponentials Formula","Screw Theory"]},{"id":"exponential-map","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Solà, Deray, Atchuthan: A micro Lie theory for state estimation in robotics","url":"https://arxiv.org/abs/1812.01537"},{"title":"Lynch & Park, Modern Robotics（Ch. 3 exponential coordinates / matrix logarithm）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Wikipedia: Exponential map (Lie theory)","url":"https://en.wikipedia.org/wiki/Exponential_map_(Lie_theory)"}],"as_of":"","related_ids":["lie-group","rotation-matrix","axis-angle-representation","rodrigues-rotation-formula","product-of-exponentials-formula","skew-symmetric-matrix"],"name":"Exponential Map","alt":"指数映射","abbr":"","aliases":["Log Map","Logarithmic Map","Exp/Log"],"one_liner":"Turns an axis-times-angle vector into a rotation or pose; the log map does the reverse.","explanation":"The exponential map comes from Lie group theory: it maps an element of a Lie algebra (the tangent space of the group at its identity element, which can be thought of as the linear space where 'velocities' live) onto the group itself; for matrix Lie groups, it's just the matrix exponential, exp(X) = I + X + X²/2! + …. The log map is its inverse near the identity. The most common use in robotics is with rotations: combine a unit rotation axis ω̂ and an angle θ into the 3D vector ω̂θ (exponential coordinates, also called the rotation vector), write it as the skew-symmetric matrix [ω̂]θ, and exponentiate to get the rotation matrix R — the closed-form result is Rodrigues' formula; taking the log of R recovers ω̂θ. Likewise, the exponential map from se(3) to SE(3) turns a twist into a homogeneous transformation, and underlies the product of exponentials formula. State estimation and optimization routines often perform addition and subtraction in this flat tangent space and then map back to a rotation with exp, avoiding the orthogonality-breaking errors of directly adding or subtracting rotation matrices.","example":"With ω̂ = (0, 0, 1) and θ = π/2, the exponential map gives the rotation matrix for a 90° rotation about the z-axis; taking the log of that matrix recovers the vector (0, 0, π/2).","related":["Lie Group","Rotation Matrix","Axis-Angle Representation","Rodrigues' Rotation Formula","Product of Exponentials Formula","Skew-Symmetric Matrix"]},{"id":"screw-theory","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Screw theory","url":"https://en.wikipedia.org/wiki/Screw_theory"},{"title":"Wikipedia: Product of exponentials formula","url":"https://en.wikipedia.org/wiki/Product_of_exponentials_formula"},{"title":"Modern Robotics: Mechanics, Planning, and Control (Lynch & Park)","url":"http://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["twist","wrench","product-of-exponentials-formula","lie-group","exponential-map","adjoint-representation"],"name":"Screw Theory","alt":"旋量理论","abbr":"","aliases":["Theory of Screws","Screw Axis"],"one_liner":"Treats any rigid-body motion as a rotation about some axis combined with a translation along that same axis, like turning a screw.","explanation":"Screw theory studies the geometry of rigid-body motion and the forces acting on rigid bodies. In 1763, Mozzi proved that any rigid-body motion can be viewed as a rotation about some axis combined with a translation along that same axis — like driving a screw — and that axis is called the screw axis (this result later became known as Chasles' theorem); Robert Ball systematized the theory in his 1876 book The Theory of Screws. At its core are two 6-dimensional quantities: the twist (angular velocity ω and linear velocity v combined) describes a rigid body's velocity, and the wrench (torque and force combined) describes the forces acting on it. Its benefit is that both revolute and prismatic joints can be written uniformly as a single screw axis, without needing to build a separate coordinate frame on every link the way DH parameters do. Brockett used this framework to introduce the product of exponentials formula in 1983–1984, and Lynch and Park's 2017 textbook Modern Robotics teaches kinematics, Jacobians, and dynamics entirely in this language.","example":"Turning a screwdriver: for every full turn, the screw advances one thread pitch along its axis — that's a screw motion. A robot's revolute joint is a screw with zero pitch (pure rotation, no translation), and its prismatic joint can be viewed as a screw with infinite pitch (pure translation, no rotation).","related":["Twist","Wrench","Product of Exponentials Formula","Lie Group","Exponential Map","Adjoint Representation"]},{"id":"plucker-coordinates","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Plücker coordinates","url":"https://en.wikipedia.org/wiki/Pl%C3%BCcker_coordinates"},{"title":"CameraCtrl: Enabling Camera Control for Text-to-Video Generation (arXiv 2404.02101)","url":"https://arxiv.org/html/2404.02101"},{"title":"Roy Featherstone: Spatial vector teaching materials（Plücker Basis Vectors）","url":"https://royfeatherstone.org/teaching/"}],"as_of":"","related_ids":["screw-theory","spatial-vector-algebra","camera-extrinsics","camera-intrinsics","video-generation-model","world-model"],"name":"Plücker Coordinates","alt":"Plücker 坐标（Plücker 射线嵌入）","abbr":"","aliases":["Plücker Ray Embedding","Plücker Embedding","Plücker Ray Map"],"one_liner":"Represents a line in 3D space with 6 numbers — a direction plus a moment — and is often used to encode camera rays per pixel.","explanation":"Plücker coordinates, introduced by the 19th-century German mathematician Julius Plücker, describe a line in 3D space with 6 numbers: a direction d, and a moment m = p × d (p is any point on the line, × is the cross product). The moment m doesn't depend on which point p on the line is chosen, and d·m = 0. The representation is fundamental in robotics: joint axes in screw theory and Featherstone's spatial vector dynamics are both built on Plücker coordinates. In recent years it has also shown up in generative models as the 'Plücker ray embedding': for every pixel in an image, compute (o × d, d) from the camera center o and that pixel's ray direction d, producing a 6-channel map used as a camera-pose conditioning input. CameraCtrl (2024) follows the approach of Light Field Networks (2021), arguing that this is easier for a network to associate with per-pixel information than directly feeding in the intrinsic and extrinsic matrices; it's commonly used in camera-controlled video generation and world models.","example":"A camera at the origin looking along the z-axis has a ray direction d = (0, 0, 1) for the image's center pixel, with o × d = 0, giving the embedding (0, 0, 0, 0, 0, 1); after the camera translates sideways, o changes, and the same pixel's o × d changes with it, which is how the network senses camera motion.","related":["Screw Theory","Spatial Vector Algebra","Camera Extrinsics","Camera Intrinsics","Video Generation Model","World Model"]},{"id":"dual-quaternion","category":"mechanics","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Dual quaternion","url":"https://en.wikipedia.org/wiki/Dual_quaternion"},{"title":"DQ Robotics（dual quaternion robot modelling and control library）","url":"https://dqrobotics.github.io/"}],"as_of":"","related_ids":["quaternion","homogeneous-transformation-matrix","screw-theory","pose","lie-group","spherical-linear-interpolation"],"name":"Dual Quaternion","alt":"对偶四元数","abbr":"","aliases":["Dual Quaternions","Unit Dual Quaternion"],"one_liner":"An 8-number algebraic tool that represents a 3D rotation and translation together in a single object.","explanation":"A dual quaternion is a quaternion whose coefficients are dual numbers, written q = r + εd, where r and d are ordinary quaternions and ε satisfies ε² = 0, for 8 real components in total. Study pointed out in 1891 that this algebra is well suited to describing rigid-body motion in 3D space, and Kotelnikov independently proposed it in 1895. Just as a unit quaternion represents a rotation, a unit dual quaternion can represent a complete rigid-body transformation, or pose: r encodes the rotation, and d = ½·t·r encodes the translation t. Composing two transformations is just multiplying their dual quaternions, which is more compact than a 4×4 homogeneous transformation matrix and also makes interpolating between two poses convenient. In robotics, open-source libraries such as DQ Robotics use dual quaternions for manipulator kinematic modeling and control; they're also used in computer graphics.","example":"To rotate 90° about the z-axis and then translate by (1, 0, 0): r = (cos45°, 0, 0, sin45°), the translation is written as the pure quaternion t = (0, 1, 0, 0), and d = ½·t·r is computed; r and d together give the 8 numbers representing this pose.","related":["Quaternion","Homogeneous Transformation Matrix","Screw Theory","Pose","Lie Group","Spherical Linear Interpolation (SLERP)"]},{"id":"link","category":"mechanics","sec":2,"tier":1,"sources":[{"title":"Modern Robotics (Lynch & Park, 2017), Ch.2 Configuration Space","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"Wikipedia: Robot kinematics","url":"https://en.wikipedia.org/wiki/Robot_kinematics"},{"title":"ros/urdf_tutorial: 07-physics.urdf","url":"https://raw.githubusercontent.com/ros/urdf_tutorial/ros2/urdf/07-physics.urdf"}],"as_of":"","related_ids":["revolute-joint","prismatic-joint","kinematic-chain","rigid-body","unified-robot-description-format","degrees-of-freedom"],"name":"Link","alt":"连杆","abbr":"","aliases":["Robot Link"],"one_liner":"One of the rigid segments of a robot's body that joints connect together.","explanation":"A link is the basic structural unit of a robot's mechanical body: a robot is a series of links strung together by joints (connections that allow relative motion), with motors driving the joints to move the links. For modeling purposes, a link is usually treated as a rigid body — one that doesn't deform under load. In a robot description file such as URDF, each link tag records that segment's visual mesh, a simplified shape used for collision checking, and its mass and moment of inertia; the joint tags record which two links a joint connects and its axis of rotation. Forward kinematics works by starting at the base and multiplying coordinate transforms link by link, joint by joint, all the way to the tip, to compute where the end effector is. Link length directly determines how far an arm can reach.","example":"A serial arm with n joints has n+1 links, including the base — a 6-axis arm has 7 links. A human's upper arm and forearm are effectively two links, connected by the elbow joint.","related":["Revolute Joint","Prismatic Joint","Kinematic Chain","Rigid Body","Unified Robot Description Format","Degrees of Freedom (DoF)"]},{"id":"revolute-joint","category":"mechanics","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Revolute joint","url":"https://en.wikipedia.org/wiki/Revolute_joint"},{"title":"Modern Robotics (Lynch & Park, 2017), 2.2.1 Robot Joints","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"ros/urdf_tutorial: 07-physics.urdf","url":"https://raw.githubusercontent.com/ros/urdf_tutorial/ros2/urdf/07-physics.urdf"}],"as_of":"","related_ids":["link","prismatic-joint","spherical-joint","joint-limits","degrees-of-freedom","rotary-encoder"],"name":"Revolute Joint","alt":"转动关节","abbr":"","aliases":["Hinge Joint","R Joint"],"one_liner":"A joint that only lets two links rotate relative to each other about a fixed axis, giving 1 degree of freedom.","explanation":"A revolute joint — often labeled R, and also called a hinge joint — connects two links and only lets them rotate relative to each other about a shared axis, with no sliding, so it contributes exactly 1 degree of freedom. It's the most common joint in arms and humanoid robots: a motor drives the output side through a reducer, and an encoder measures the rotation angle, which is what's meant by the joint angle. Its counterpart is the prismatic joint (P), which only allows translation along an axis. In URDF, a joint of type “revolute” must specify its axis direction and limits — upper and lower angle bounds in radians, maximum torque, and maximum velocity; a joint with no angle limit that can spin forever, like a wheel, is typed “continuous” instead. What people call a “6-axis arm” is usually six revolute joints chained in series.","example":"A door hinge is a revolute joint. All 7 joints on a Franka Research 3 are revolute; the values q1 through q7 the controller reads out are each joint's rotation angle, in radians.","related":["Link","Prismatic Joint","Spherical Joint","Joint Limits","Degrees of Freedom (DoF)","Rotary Encoder"]},{"id":"prismatic-joint","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Prismatic joint","url":"https://en.wikipedia.org/wiki/Prismatic_joint"},{"title":"Wikipedia: Linear actuator","url":"https://en.wikipedia.org/wiki/Linear_actuator"}],"as_of":"","related_ids":["revolute-joint","linear-actuator","cartesian-robot","selective-compliance-assembly-robot-arm","kinematic-pair","unified-robot-description-format"],"name":"Prismatic Joint","alt":"移动关节","abbr":"","aliases":["Slider Joint","Sliding Pair","P Joint"],"one_liner":"A single-degree-of-freedom joint that lets two parts slide relative to each other along one straight axis, with no rotation.","explanation":"The prismatic joint is one of the two most basic robot joints, alongside the revolute joint: it constrains two links to slide relative to each other along a shared axis only, with no rotation, giving 1 translational degree of freedom, and its joint variable is a displacement rather than an angle. In mechanism theory, it's classified as a lower pair, also called a sliding pair; in shorthand notation it's written P, as in RRP for two revolute joints plus a prismatic joint. It's usually driven by a linear actuator, such as a motor-and-ball-screw electric cylinder, a hydraulic cylinder, or a linear motor. The three axes of a Cartesian robot, a SCARA robot's vertical lift axis, a hybrid robot's lift column, and the opening and closing motion of a parallel-jaw two-finger gripper are all prismatic joints. In URDF, this joint type is written prismatic, with an axis direction and upper/lower travel limits specified.","example":"A SCARA robot is typically two horizontal revolute joints plus one vertical prismatic joint: the revolute joints position the tool within a plane, and the prismatic joint plunges it down or lifts it up.","related":["Revolute Joint","Linear Actuator (Electric Cylinder)","Cartesian Robot","Selective Compliance Assembly Robot Arm","Kinematic Pair","Unified Robot Description Format"]},{"id":"spherical-joint","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park), Ch.2.2 Joints","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo XML Reference: body/joint type","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html#body-joint"},{"title":"urdfdom_headers: joint.hpp (URDF joint types)","url":"https://github.com/ros/urdfdom_headers/blob/rolling/include/urdf_model/joint.hpp"}],"as_of":"","related_ids":["revolute-joint","spherical-wrist","parallel-mechanism","degrees-of-freedom","mujoco","unified-robot-description-format"],"name":"Spherical Joint","alt":"球关节","abbr":"","aliases":["Ball Joint","Ball-and-Socket Joint"],"one_liner":"A three-degree-of-freedom joint that lets one part rotate about a fixed point in any direction, but not translate.","explanation":"A spherical joint, also called a ball joint or ball-and-socket joint, has one part's ball-shaped head seated inside another part's socket: the two can rotate about the ball's center in any direction, but can't translate relative to each other, giving 3 rotational degrees of freedom — functionally similar to a human shoulder or hip, and often written S in mechanism notation. Passive ball joints are used extensively in the linkages of parallel mechanisms like the Stewart platform. Real robot arms rarely drive a spherical joint with a single motor; instead, they approximate one with three revolute joints whose axes intersect at a point, as in a spherical wrist. In simulation, MuJoCo offers a ball joint type, representing rotation with a unit quaternion, so it takes up 4 numbers in qpos but only 3 in qvel; URDF has no spherical joint type at all, so it's usually approximated by chaining three revolute joints together.","example":"Setting a humanoid robot's shoulder as a ball joint in MuJoCo lets a single quaternion describe the upper arm's orientation; exporting that to URDF for ROS requires splitting it into three separate revolute joints — shoulder pitch, roll, and yaw.","related":["Revolute Joint","Spherical Wrist","Parallel Mechanism","Degrees of Freedom (DoF)","MuJoCo (Multi-Joint dynamics with Contact)","Unified Robot Description Format"]},{"id":"kinematic-pair","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Kinematic pair","url":"https://en.wikipedia.org/wiki/Kinematic_pair"},{"title":"Lynch & Park, Modern Robotics（2017）2.2 节 Grübler 公式；参考文献 Denavit & Hartenberg 1955","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["revolute-joint","prismatic-joint","spherical-joint","degrees-of-freedom","grubler-s-formula","four-bar-linkage"],"name":"Kinematic Pair","alt":"运动副（转动副 / 移动副 / 低副 / 高副）","abbr":"","aliases":["Joint (Mechanism)","Revolute Pair / Prismatic Pair","Lower Pair / Higher Pair"],"one_liner":"A connection between two parts that stay in contact while allowed to move relative to each other, like a hinge, a slider, or meshing gears.","explanation":"A kinematic pair is a basic concept in mechanism design, introduced by the German engineer Franz Reuleaux, referring to a connection between two members that keeps them in contact while restricting their relative motion to a specific type. Kinematic pairs split into two classes by contact type. Lower pairs make surface contact and include the revolute pair (R, a hinge, 1 rotational degree of freedom), the prismatic pair (P, a slider, 1 translational degree of freedom), the screw pair, the cylindrical pair (2 degrees of freedom), and the spherical pair (3 degrees of freedom). Higher pairs make point or line contact, such as a cam and follower, meshing gear teeth, or a wheel rolling on the ground. What robotics calls a 'joint' is a kinematic pair: a revolute joint is a revolute pair, and a prismatic joint is a prismatic pair. Grübler's formula for counting a mechanism's degrees of freedom works by subtracting the constraints each kinematic pair introduces from the members' total degrees of freedom. The original paper introducing DH parameters was even titled 'A kinematic notation for lower-pair mechanisms based on matrices.'","example":"A planar four-bar linkage: including the frame, it has 4 members and 4 revolute pairs; by the planar Grübler formula, degrees of freedom = 3×(4−1−4) + 4×1 = 1, so a single motor is enough to drive the whole mechanism.","related":["Revolute Joint","Prismatic Joint","Spherical Joint","Degrees of Freedom (DoF)","Grübler's Formula","Four-Bar Linkage"]},{"id":"flexion-extension-and-abduction-adduction","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Anatomical terms of motion","url":"https://en.wikipedia.org/wiki/Anatomical_terms_of_motion"},{"title":"Shaw, Agarwal, Pathak: LEAP Hand（universal abduction-adduction mechanism）","url":"https://arxiv.org/abs/2309.06440"},{"title":"灵心巧手 Linker Hand 官网（产品参数：拇指侧摆、四指弯曲）","url":"https://www.linkerbot.cn/"}],"as_of":"","related_ids":["dexterous-hand","metacarpophalangeal-proximal-and-distal-interphalangeal-join","thumb-opposition","degrees-of-freedom","leap-hand","motion-retargeting"],"name":"Flexion/Extension and Abduction/Adduction","alt":"屈伸与侧摆（外展/内收）","abbr":"","aliases":["Lateral Swing","Flexion/Extension","Abduction/Adduction"],"one_liner":"Two basic joint motions: flexion/extension is bending and straightening, abduction/adduction is spreading apart and drawing together.","explanation":"These are anatomical terms for joint motion, carried over into dexterous hands, exoskeletons, and human motion capture. Flexion is the bending motion that decreases the angle between two adjacent segments, as in making a fist; extension is the opposite, straightening back out. Abduction moves a part away from the midline of the body or hand — for fingers, this means spreading apart; adduction moves it back toward the midline, i.e. bringing the fingers together. In the human hand, the metacarpophalangeal (MCP) joint at the base of each finger can both flex/extend and abduct/adduct, giving it two degrees of freedom, while the interphalangeal joints further out (PIP, DIP) can only flex and extend. Chinese-language dexterous-hand spec sheets commonly label abduction/adduction as 'lateral swing' and flexion/extension as 'curling' — for instance, the Linker Hand's specs list 'thumb lateral swing' and 'four-finger curling.' Some dexterous hands simplify their design by giving the four fingers flexion/extension only, while the LEAP Hand specifically designed MCP joints that can abduct/adduct across a range of poses, to make retargeting human hand motion easier.","example":"Spreading all five fingers apart into a fan shape is abduction; bringing them back together into a flat hand is adduction. Making a fist is flexion of each finger; flattening the palm out is extension.","related":["Dexterous Hand","Metacarpophalangeal / Proximal & Distal Interphalangeal Joints (MCP / PIP / DIP)","Thumb Opposition","Degrees of Freedom (DoF)","LEAP Hand","Motion Retargeting"]},{"id":"joint-limits","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 4 URDF 与 Ch. 9 Fig. 9.1","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"legged_gym: legged_robot_config.py (soft_dof_pos_limit)","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["soft-limits","configuration-space","inverse-kinematics","torque-limiting","unified-robot-description-format","workspace"],"name":"Joint Limits","alt":"关节限位","abbr":"","aliases":["Joint Range","Joint Position Limits"],"one_liner":"The allowed range of motion for each joint, and more broadly the velocity and torque limits that go with it.","explanation":"Joint limits are the range of values a joint variable is allowed to take — a revolute joint might only turn between −90° and 90°, mostly set by the mechanical structure; more broadly, the term also covers velocity and torque limits. In URDF, revolute and prismatic joints specify four values in a limit tag: lower, upper, velocity, and effort; a revolute joint with no rotation limit is instead typed continuous. Limits affect many parts of the pipeline: an inverse-kinematics solution has to be checked against them, motion planning has to cut the parts of configuration space they rule out, and real robot controllers commonly set software limits tighter than the mechanical hard limits, to leave a margin. Reinforcement-learning training commonly penalizes states that approach the limits — legged_gym's soft_dof_pos_limit, for instance, sets a penalty threshold as a percentage of the URDF's limits.","example":"In Modern Robotics, a 2R arm is limited to 0°≤θ1≤180° and 0°≤θ2≤150°: a straight-line path in joint space is fine, but the path that moves the end effector in a straight line through task space ends up exceeding the joint limits.","related":["Soft Limits (Software Joint Limits)","Configuration Space (C-Space)","Inverse Kinematics (IK)","Torque Limiting (Saturation)","Unified Robot Description Format","Workspace"]},{"id":"kinematic-chain","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Kinematic chain - Wikipedia","url":"https://en.wikipedia.org/wiki/Kinematic_chain"},{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 1–2 与 Ch. 7 Closed Chains","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["serial-mechanism","parallel-mechanism","kinematic-tree","grubler-s-formula","degrees-of-freedom","parallel-ankle-mechanism"],"name":"Kinematic Chain","alt":"运动链","abbr":"","aliases":["Open Chain","Closed Chain"],"one_liner":"A mechanism model made of rigid links strung together by joints, split into open chains and closed chains.","explanation":"A kinematic chain is a set of rigid bodies (links) connected by joints, used to describe how a mechanism can move; Franz Reuleaux systematized the idea in his 1876 book Kinematics of Machinery. When the links connect one after another with no loop, it's called an open chain (or serial chain) — the typical example is an ordinary robot arm, with a motor driving every joint. When the links form one or more closed loops, it's a closed chain — examples include the Stewart platform, Delta parallel robots, and the parallel ankle joints found in some humanoid robots — and typically only some of a closed chain's joints are actively driven. A mechanism's degrees of freedom can be estimated with the Grübler formula: M = d(N−1−j) + Σf_i, where d is 3 in the plane or 6 in space, N is the number of links (including the ground), j is the number of joints, and f_i is the i-th joint's degrees of freedom. Open chains have simple forward kinematics but hard inverse kinematics; closed chains are usually the other way around.","example":"A tabletop 6-axis arm, from base to gripper, is an open chain; a Delta parallel robot's three supporting arms all meet at the moving platform, forming a closed chain.","related":["Serial Mechanism","Parallel Mechanism","Kinematic Tree","Grübler's Formula","Degrees of Freedom (DoF)","Parallel Ankle Mechanism"]},{"id":"serial-mechanism","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Serial manipulator","url":"https://en.wikipedia.org/wiki/Serial_manipulator"},{"title":"Modern Robotics (Lynch & Park), Ch.2 Configuration Space","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["parallel-mechanism","kinematic-chain","link","forward-kinematics","robotic-arm","6-axis-robot-arm"],"name":"Serial Mechanism","alt":"串联机构","abbr":"","aliases":["Serial Manipulator","Open-Chain Mechanism"],"one_liner":"A mechanism whose links connect end-to-end through joints from base to tip, with no closed loop.","explanation":"A serial mechanism, also called an open-chain mechanism, is a string of rigid links connected end-to-end by joints, running from a fixed base all the way to the end effector, with no closed loop anywhere along the way. Each joint is usually driven by its own motor, and the end-effector pose is simply the product of each joint's transform chained together, so forward kinematics (joint angles to end-effector pose) has a unique solution, while inverse kinematics can have multiple. Six-axis industrial arms, 7-DoF collaborative arms, SCARA robots, and a single humanoid arm are all serial structures. The advantages are a workspace that's large relative to the mechanism's footprint and an intuitive structure; the disadvantages are that error and deflection accumulate link by link along the chain, and each proximal joint has to carry the weight of every motor farther out, giving lower stiffness and payload capacity than a parallel mechanism (where several chains connect to the end effector at once, forming closed loops, as in a Delta robot or a Stewart platform).","example":"A UR5e's six joints — base, shoulder, elbow, and three wrist joints — connect one after another in sequence, a textbook serial mechanism. A person standing with both feet on the ground, by contrast, forms a closed loop from the ground, up one leg, through the hip, down the other leg, and back to the ground — a closed-chain mechanism.","related":["Parallel Mechanism","Kinematic Chain","Link","Forward Kinematics (FK)","Robotic Arm","6-Axis Robot Arm"]},{"id":"parallel-mechanism","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Parallel manipulator","url":"https://en.wikipedia.org/wiki/Parallel_manipulator"},{"title":"Wikipedia: Delta robot","url":"https://en.wikipedia.org/wiki/Delta_robot"}],"as_of":"","related_ids":["serial-mechanism","delta-robot","parallel-ankle-mechanism","kinematic-chain","singular-configuration","four-bar-linkage"],"name":"Parallel Mechanism","alt":"并联机构","abbr":"","aliases":["Parallel Manipulator","Parallel Robot"],"one_liner":"A mechanism connecting a base to an end platform through several kinematic chains at once — stiff and fast, but with a small workspace.","explanation":"A parallel mechanism has its end platform connected to the base by two or more independent kinematic chains at the same time, forming closed loops; its counterpart, a serial mechanism, has just one chain running from base to tip — an ordinary 6-axis arm is serial. Classic examples include the Stewart platform, which holds up a platform on six extendable struts (used in flight simulators), and the Delta robot — nicknamed the “spider robot” — invented in the early 1980s by Reymond Clavel's team at EPFL (the Swiss Federal Institute of Technology in Lausanne), widely used for high-speed pick-and-place sorting of food and electronics. A parallel mechanism's motors are mostly mounted on the base, keeping the moving parts light, and the load is shared across several chains, giving high stiffness, high precision, and high speed; the tradeoffs are a small workspace, hard-to-solve forward kinematics, and extra singular configurations that a serial arm doesn't have. Many humanoid robots also use a parallel structure for their ankle joints.","example":"A Delta robot's three actively driven arms each carry a parallelogram linkage, so the end platform can only translate, never rotate; according to Wikipedia, it can complete up to about 300 pick-and-place cycles per minute.","related":["Serial Mechanism","Delta Robot","Parallel Ankle Mechanism","Kinematic Chain","Singular Configuration (Kinematic Singularity)","Four-Bar Linkage"]},{"id":"four-bar-linkage","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Four-bar linkage","url":"https://en.wikipedia.org/wiki/Four-bar_linkage"},{"title":"Wikipedia: Grashof condition","url":"https://en.wikipedia.org/wiki/Grashof_condition"},{"title":"Lynch & Park, Modern Robotics（预印本 PDF，第 2 章 Example 2.3）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["grubler-s-formula","linkage-transmission","parallel-mechanism","parallel-ankle-mechanism","degrees-of-freedom","link"],"name":"Four-Bar Linkage","alt":"四连杆机构","abbr":"","aliases":["Four-Bar Mechanism","Planar Four-Bar Linkage"],"one_liner":"Four bars joined end to end by four revolute joints into a closed loop, with just 1 degree of freedom.","explanation":"A four-bar linkage connects four bars — one of which is fixed, called the frame — end to end with four revolute joints into a closed loop; it's the simplest movable closed-chain mechanism, and in the planar case it has only 1 degree of freedom: rotate any one bar and the positions of all the others are fully determined. The bar attached to the frame that can rotate all the way around is called the crank; one that can only rock back and forth is a rocker; the bar floating in between is the coupler. Whether a bar can rotate fully is decided by the Grashof condition: if the sum of the shortest and longest bar lengths is no greater than the sum of the other two, the shortest bar can rotate fully relative to its neighbors, which is how linkages get classified into crank-rocker, double-crank, and double-rocker types. Its job is to turn a motor's simple rotation into a desired output trajectory or angular relationship — windshield wipers and car suspensions both use it, and robots often use it to place a motor closer to the torso and drive a knee or ankle through a linkage, while a parallelogram four-bar can keep a gripper's fingers always parallel.","example":"A car windshield wiper: the motor spins a crank continuously, which drives a rocker back and forth through a coupler bar, and the wiper blade sweeps left and right across the windshield.","related":["Grübler's Formula","Linkage Transmission","Parallel Mechanism","Parallel Ankle Mechanism","Degrees of Freedom (DoF)","Link"]},{"id":"grubler-s-formula","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Chebychev–Grübler–Kutzbach criterion","url":"https://en.wikipedia.org/wiki/Chebychev%E2%80%93Gr%C3%BCbler%E2%80%93Kutzbach_criterion"},{"title":"Lynch & Park, Modern Robotics（预印本 PDF，2.2.2 节 Grübler's Formula）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["degrees-of-freedom","four-bar-linkage","parallel-mechanism","kinematic-pair","configuration-space","underactuation"],"name":"Grübler's Formula","alt":"Grübler 公式","abbr":"","aliases":["Chebychev–Grübler–Kutzbach Criterion","Kutzbach-Grübler Formula","Mobility Formula"],"one_liner":"Computes how many degrees of freedom a mechanism has from its number of links, joints, and each joint's own freedom.","explanation":"Grübler's formula, also known as the Chebychev–Grübler–Kutzbach criterion after its three originators, counts the degrees of freedom of a mechanism built from links and joints: dof = m(N − 1 − J) + Σfᵢ. Here m is the degrees of freedom of a single free rigid body (3 for a planar mechanism, 6 for a spatial one), N is the number of links (the fixed frame counts as one), J is the number of joints, and fᵢ is the number of degrees of freedom joint i provides (1 for a revolute or prismatic joint, 3 for a ball joint). The logic is that each movable link starts out with m degrees of freedom, and each joint removes m − fᵢ of them as a constraint. It's a quick way to figure out how many motors a parallel mechanism, closed-chain leg, or finger mechanism needs. Its limitation is that it assumes all the constraints are independent; when the geometry is special — parallel or equal-length members, for instance — the constraints can become redundant, and the true number of degrees of freedom can exceed what the formula predicts. Such mechanisms are called overconstrained, and in that case the formula only gives a lower bound.","example":"A planar four-bar linkage: m = 3, N = 4 (including the frame), J = 4 revolute joints each contributing 1 degree of freedom, so dof = 3×(4−1−4) + 4 = 1 — rotating one crank fully determines the configuration of the whole mechanism.","related":["Degrees of Freedom (DoF)","Four-Bar Linkage","Parallel Mechanism","Kinematic Pair","Configuration Space (C-Space)","Underactuation"]},{"id":"parallel-ankle-mechanism","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"unitree_sdk2 example: g1_ankle_swing_example.cpp（PR: Series Control for Pitch/Roll Joints; AB: Parallel Control for A/B Joints）","url":"https://raw.githubusercontent.com/unitreerobotics/unitree_sdk2/main/example/g1/low_level/g1_ankle_swing_example.cpp"},{"title":"Wikipedia: Parallel manipulator","url":"https://en.wikipedia.org/wiki/Parallel_manipulator"}],"as_of":"","related_ids":["parallel-mechanism","serial-mechanism","unified-robot-description-format","unitree-g1","bipedal-robot","reflected-inertia"],"name":"Parallel Ankle Mechanism","alt":"并联踝关节","abbr":"","aliases":["Closed-Chain Ankle","Parallel Ankle"],"one_liner":"An ankle joint where two motors, through linkages, jointly drive both pitch and roll instead of each controlling one axis alone.","explanation":"The parallel ankle mechanism is a common ankle design in humanoid and biped robots: instead of one motor controlling pitch (the toes tipping up and down) and another controlling roll (the sole tilting side to side), two motors push and pull the foot plate together through linkages, forming a closed kinematic chain. The advantage is that the motors can be moved up the shin closer to the knee, making the end of the leg lighter and reducing the swing-leg inertia, and the two motors can combine their torque output. The tradeoff is that the relationship between motor angle and ankle angle becomes nonlinearly coupled, requiring the mechanism's own forward and inverse kinematics to convert between them, with torque also converted through a Jacobian; since URDF can only describe tree structures, simulations often simplify it into two independent series pitch and roll joints instead. Unitree's G1 SDK offers exactly this choice: PR mode issues commands as if the pitch and roll joints were in series, while AB mode directly controls the two parallel motors.","example":"In a typical left-right symmetric two-link ankle: when the two motors turn in the same direction, the foot pitches; when they turn in opposite directions, the foot rolls; any other combination of motor angles is a blend of the two.","related":["Parallel Mechanism","Serial Mechanism","Unified Robot Description Format","Unitree G1","Bipedal Robot","Reflected Inertia"]},{"id":"kinematic-tree","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（2017）4.2 节：URDF 可表示任何树形结构机器人","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo 文档 Overview：Kinematic tree（不允许运动学回路，用等式约束建模）","url":"https://mujoco.readthedocs.io/en/stable/overview.html"},{"title":"Pinocchio（GitHub）：利用运动学树稀疏性的刚体动力学算法库","url":"https://github.com/stack-of-tasks/pinocchio"}],"as_of":"","related_ids":["kinematic-chain","serial-mechanism","parallel-mechanism","unified-robot-description-format","floating-base","articulated-body-algorithm"],"name":"Kinematic Tree","alt":"运动学树","abbr":"","aliases":["Tree-Structured Kinematic Chain"],"one_liner":"A tree structure with links as nodes and joints as edges, branching outward from a root to each end effector.","explanation":"A kinematic tree describes the structure of a multi-link robot: links are nodes, and joints are the edges connecting a parent link to a child link; each link has only one parent and there are no closed loops, so the structure branches outward from a root node — a fixed base, or the pelvis for a humanoid — to each end. A single robot arm is the special case with just one branch (a serial chain), while humanoids, dexterous hands, and quadrupeds are trees with multiple branches. URDF (Unified Robot Description Format) can only describe tree structures and can't directly represent closed chains such as a Stewart platform; MuJoCo likewise organizes every rigid body into a tree rooted at 'world' with no loops allowed, so closed loops like four-bar linkages or parallel ankle mechanisms have to be added back in as extra equality constraints. Dynamics libraries such as Pinocchio recurse over this tree and exploit the sparsity that its structure provides to speed up computation.","example":"In a humanoid robot's URDF, the pelvis is the root, branching down into the left and right legs and up through the waist to the torso, which then branches into the two arms and the head; each dexterous hand further branches into five fingers, making the whole robot a multi-level branching tree.","related":["Kinematic Chain","Serial Mechanism","Parallel Mechanism","Unified Robot Description Format","Floating Base","Articulated Body Algorithm"]},{"id":"floating-base","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"MuJoCo Documentation - Overview (joint types, free joint)","url":"https://mujoco.readthedocs.io/en/stable/overview.html"},{"title":"MuJoCo Menagerie - unitree_g1/g1.xml","url":"https://github.com/google-deepmind/mujoco_menagerie/blob/main/unitree_g1/g1.xml"},{"title":"Underactuated Robotics (MIT) - Multi-Body Dynamics","url":"https://underactuated.csail.mit.edu/multibody.html"}],"as_of":"","related_ids":["generalized-coordinates","underactuation","centroidal-dynamics","whole-body-control","state-estimation","mjcf"],"name":"Floating Base","alt":"浮动基","abbr":"","aliases":["Free Joint","Free-Flyer Joint"],"one_liner":"A robot whose torso isn't bolted to the ground and instead has its own 6 degrees of freedom of free motion.","explanation":"An arm bolted to a table has a fixed base — the base doesn't move, so joint angles alone describe its configuration. A humanoid's, quadruped's, or drone's torso (the base) isn't attached to the world, so on top of its joints, 6 more degrees of freedom are needed to describe the torso's own position and orientation in space — this is what's meant by a floating base. Modeling it usually means inserting a virtual “free joint” between the world and the torso: MuJoCo calls it a free joint, with position given by 7 numbers (3D position plus a 4D unit quaternion) and velocity given by 6 numbers (3D linear velocity plus 3D angular velocity). The difficulty is that these 6 degrees of freedom have no motor driving them directly — the torso can only be pushed indirectly through the feet's contact forces with the ground, which makes it underactuated. As a result, controlling a floating-base robot is inseparable from contact forces, friction cones, and state estimation, and the torso's pose has to be estimated from an IMU and leg odometry.","example":"In the Unitree G1 model from the MuJoCo Menagerie, the root body pelvis carries a freejoint named floating_base_joint, so the first 7 values of qpos are the pelvis position and quaternion, with the individual joint angles following after.","related":["Generalized Coordinates","Underactuation","Centroidal Dynamics","Whole-Body Control","State Estimation","MJCF (MuJoCo XML Format)"]},{"id":"configuration-space","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Modern Robotics（Lynch & Park, 2017 预印本）第 2 章 Configuration Space","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Motion planning - Wikipedia","url":"https://en.wikipedia.org/wiki/Motion_planning"},{"title":"Configuration space (physics) - Wikipedia","url":"https://en.wikipedia.org/wiki/Configuration_space_(physics)"}],"as_of":"","related_ids":["degrees-of-freedom","joint-space","motion-planning","sampling-based-planning","collision-checking","generalized-coordinates"],"name":"Configuration Space (C-Space)","alt":"构型空间","abbr":"C-space","aliases":["C-Space"],"one_liner":"The space of every possible pose a robot can be in, where each point represents one complete configuration.","explanation":"A configuration is a set of numbers that uniquely pins down the position of every part of a robot; configuration space is the set of all possible configurations, with dimensionality equal to the number of degrees of freedom. For example, an object that can translate and rotate in a plane has a 3-dimensional configuration space (x, y, θ); a rigid body in 3D space has 6; a fixed-base arm with n revolute joints has an n-dimensional configuration space, using the joint angles directly as coordinates. According to Modern Robotics, Tomás Lozano-Pérez introduced this idea to motion planning in a 1980 MIT report: configurations that would collide with an obstacle are marked as obstacle regions, and everything else is called free space, which turns “steer a robot with a physical shape around obstacles” into “find a path for a single point through configuration space.” Sampling-based planners like RRT and PRM both operate in configuration space.","example":"A planar two-link arm's configuration space consists of the two joint angles (θ1, θ2) — strictly speaking, a torus. A cup on a table becomes an irregularly shaped forbidden region on this space; planning a path means connecting a start point to a goal point on it while going around that region.","related":["Degrees of Freedom (DoF)","Joint Space","Motion Planning","Sampling-Based Planning","Collision Checking","Generalized Coordinates"]},{"id":"generalized-coordinates","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Generalized coordinates - Wikipedia","url":"https://en.wikipedia.org/wiki/Generalized_coordinates"},{"title":"MuJoCo Documentation - Overview (qpos / qvel, nq / nv)","url":"https://mujoco.readthedocs.io/en/stable/overview.html"},{"title":"ros2/common_interfaces - sensor_msgs/msg/JointState.msg","url":"https://github.com/ros2/common_interfaces/blob/rolling/sensor_msgs/msg/JointState.msg"}],"as_of":"","related_ids":["degrees-of-freedom","configuration-space","euler-lagrange-equations","floating-base","joint-space","proprioception"],"name":"Generalized Coordinates","alt":"广义坐标","abbr":"","aliases":["Generalized Velocities","qpos / qvel","Joint State (joint_states)"],"one_liner":"The smallest set of independent variables — such as each joint's angle — that uniquely describes a robot's overall configuration.","explanation":"Describing a multi-link robot could mean recording every link's position and orientation in space, but those are heavily redundant because the joints constrain them. Generalized coordinates instead use only the minimal set of independent parameters that uniquely fixes the configuration — for a serial arm, that's simply the joint angles q, and the count usually equals the number of degrees of freedom. Their time derivative q̇ is called the generalized velocity; the force paired with each coordinate, the one that does work along that coordinate, is the generalized force — for a revolute joint, that's just the joint torque. The Euler-Lagrange equations, the mass matrix, and most simulators all operate in generalized coordinates. In code, this looks like: MuJoCo stores position and velocity as qpos and qvel, and when a free joint or ball joint is present, orientation is stored as a quaternion, so the position dimension nq can exceed the velocity dimension nv; ROS's sensor_msgs/JointState message (usually published on a topic called /joint_states) publishes position, velocity, and effort arrays keyed by joint name, and is also where “joint state” comes from in many robot datasets.","example":"For a 7-DoF arm bolted to a table, q is simply its 7 joint angles. For the Unitree Go2 quadruped in MuJoCo, qpos is 7 (torso position plus quaternion) + 12 (leg joints) = 19-dimensional, and qvel is 6 + 12 = 18-dimensional.","related":["Degrees of Freedom (DoF)","Configuration Space (C-Space)","Euler-Lagrange Equations","Floating Base","Joint Space","Proprioception"]},{"id":"active-dof-passive-dof","category":"mechanics","sec":2,"tier":3,"sources":[{"title":"Parallel manipulator - Wikipedia","url":"https://en.wikipedia.org/wiki/Parallel_manipulator"},{"title":"Inspire RH56DFX Dexterous Hand (product page)","url":"https://en.inspire-robots.com/product/rh56dfx"},{"title":"Shadow Dexterous Hand Series (Shadow Robot)","url":"https://www.shadowrobot.com/dexterous-hand-series/"}],"as_of":"2026-09","related_ids":["degrees-of-freedom","underactuation","parallel-mechanism","dexterous-hand","adaptive-gripper","linkage-transmission"],"name":"Active DoF / Passive DoF","alt":"主动自由度 / 被动自由度","abbr":"","aliases":["Active Joint","Passive Joint","Actuated / Passive Joint"],"one_liner":"A degree of freedom is active when a motor drives it independently, and passive when it just follows a linkage or an external force.","explanation":"An active degree of freedom is a joint driven directly and independently by an actuator — a motor, a hydraulic cylinder, and so on. A passive degree of freedom has no independent actuator; its motion is instead determined by mechanical constraints, linkages, or external forces. Passive joints show up in two common places: ball joints and universal joints in parallel mechanisms — a Stewart platform, for instance, is driven by six linear actuators while its ball joints are entirely passive, their positions set by the whole closed kinematic chain — and finger joints in dexterous hands or underactuated grippers, which are coupled to an active joint through linkages or tendons. This is why reading a dexterous hand's spec sheet means distinguishing 'degrees of freedom' from 'number of joints': the former usually counts only the active DoF, which sets how many dimensions the policy can independently control — the size of its action space — while the latter also counts the passive joints that just follow along. Passive joints also need special handling in simulation (URDF mimic joints, closed-chain constraints) and when retargeting motion for teleoperation.","example":"Inspire Robots' RH56DFX dexterous hand is specified as having 6 active degrees of freedom and 12 joints: 6 motors each control one dimension, while the remaining joints follow along through linkages. The Shadow Hand, by contrast, has 20 active degrees of freedom plus 4 underactuated coupled joints, for 24 joints in total.","related":["Degrees of Freedom (DoF)","Underactuation","Parallel Mechanism","Dexterous Hand","Adaptive (Underactuated) Gripper","Linkage Transmission"]},{"id":"underactuation","category":"mechanics","sec":2,"tier":2,"sources":[{"title":"Underactuated Robotics (Russ Tedrake, MIT): Fully-actuated vs Underactuated Systems","url":"https://underactuated.mit.edu/intro.html"},{"title":"Underactuation - Wikipedia","url":"https://en.wikipedia.org/wiki/Underactuation"}],"as_of":"","related_ids":["floating-base","degrees-of-freedom","active-dof-passive-dof","inverted-pendulum-model","adaptive-gripper","nonholonomic-constraint"],"name":"Underactuation","alt":"欠驱动","abbr":"","aliases":["Underactuated System","Underactuated Robot"],"one_liner":"The motors can't directly produce acceleration in every direction, typically because there are fewer actuators than degrees of freedom.","explanation":"Underactuation means a mechanical system cannot be commanded to track an arbitrary trajectory directly. MIT's Russ Tedrake defines it this way in Underactuated Robotics: write the dynamics as q̈ = f₁(q, q̇) + f₂(q, q̇)u, where q is the vector of generalized coordinates and u is the control input; the system is underactuated whenever f₂ has rank less than the dimension of q, meaning the motors cannot directly produce acceleration in certain directions. The most common cause is simply having fewer actuators than degrees of freedom, as in a cart-pole system or a quadrotor. Legged robots and humanoids are underactuated too: N motors control only N joints, while the body's six degrees of freedom in space have no motor of their own and can only be controlled indirectly, through the contact forces between the feet and the ground — this is the root reason walking is hard. Underactuated grippers turn the same property to their advantage, using a small number of motors to drive more joints that passively wrap around an object.","example":"The Acrobot is a two-link arm with a motor only at the elbow: the shoulder joint has no actuation at all, so swinging it up from hanging straight down to a handstand can only be done by rocking the elbow motor back and forth, 'pumping' energy into the system a little at a time.","related":["Floating Base","Degrees of Freedom (DoF)","Active DoF / Passive DoF","Inverted Pendulum Model (IPM)","Adaptive (Underactuated) Gripper","Nonholonomic Constraint"]},{"id":"kinematics","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Kinematics","url":"https://en.wikipedia.org/wiki/Kinematics"},{"title":"Wikipedia: Robot kinematics","url":"https://en.wikipedia.org/wiki/Robot_kinematics"}],"as_of":"","related_ids":["forward-kinematics","inverse-kinematics","differential-kinematics","jacobian-matrix","dynamics","kinematic-chain"],"name":"Kinematics","alt":"运动学","abbr":"","aliases":["Robot Kinematics"],"one_liner":"The study of motion's pure geometry — position, velocity, acceleration — without regard to the forces that cause it.","explanation":"Kinematics is the branch of classical mechanics that studies how points and bodies move — position, velocity, acceleration — without considering the forces that cause the motion, which is why it's sometimes called “the geometry of motion”; the word traces back to a French term coined by André-Marie Ampère. Robot kinematics studies the geometric relationship between joint variables and the poses of the links and end effector, and covers mainly forward kinematics (joint angles to end-effector pose), inverse kinematics (end-effector pose to joint angles), and differential kinematics (joint velocities to end-effector velocity, described by the Jacobian matrix). It only needs geometric parameters like link lengths and joint axis directions — no mass or inertia — which is what separates it from dynamics, which does account for forces and torques. Trajectory planning, teleoperation mapping, and motion retargeting are all built on top of kinematics.","example":"Knowing how far each joint of an arm has rotated is enough to compute where the gripper currently is; that calculation only uses link dimensions and has nothing to do with how heavy whatever is inside the gripper happens to be — which is exactly why it belongs to kinematics, not dynamics.","related":["Forward Kinematics (FK)","Inverse Kinematics (IK)","Differential Kinematics","Jacobian Matrix","Dynamics","Kinematic Chain"]},{"id":"joint-space","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Configuration space (physics)（Robotic arm / joint space）","url":"https://en.wikipedia.org/wiki/Configuration_space_(physics)"},{"title":"libfranka robot_state.h（q: measured joint position）","url":"https://raw.githubusercontent.com/frankarobotics/libfranka/main/include/franka/robot_state.h"},{"title":"robosuite Documentation: Controllers（Joint Position / Velocity / Torque）","url":"https://robosuite.ai/docs/modules/controllers.html"}],"as_of":"","related_ids":["task-space","configuration-space","forward-kinematics","inverse-kinematics","degrees-of-freedom","action-space"],"name":"Joint Space","alt":"关节空间","abbr":"","aliases":[],"one_liner":"The space described by a vector of all a robot's joint values — angles or displacements — used to describe its configuration.","explanation":"Joint space arranges every one of a robot's joint variables into a single vector q = (q₁, …, qₙ), where n is the number of joints; a rotary joint contributes an angle and a linear joint contributes a displacement, and the set of every value q can take forms the joint space. The counterpart is task space (also called Cartesian space), which describes the robot by its end-effector pose instead; forward and inverse kinematics convert between the two. Planning and control are most direct in joint space, since motors are commanded in joint angles anyway and there's no issue of multiple inverse-kinematics solutions — but the path the end effector actually traces through physical space isn't obvious just by looking at joint-space values. Demonstration data recorded as joint angles reproduces the exact same motion on the same robot, but doesn't transfer directly to a robot with a different kinematic structure, which is why a lot of cross-embodiment work records actions as end-effector poses instead.","example":"The vector q that libfranka reads out — 7 joint angles, in radians — is a single point in the Franka arm's joint space; the O_T_EE reported at the same instant is its representation in task space.","related":["Task Space","Configuration Space (C-Space)","Forward Kinematics (FK)","Inverse Kinematics (IK)","Degrees of Freedom (DoF)","Action Space"]},{"id":"task-space","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Modern Robotics (Lynch & Park, 2017), 2.5 Task Space and Workspace","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"Wikipedia: Robot kinematics","url":"https://en.wikipedia.org/wiki/Robot_kinematics"},{"title":"Wikipedia: Oussama Khatib（operational space formulation, 1980）","url":"https://en.wikipedia.org/wiki/Oussama_Khatib"}],"as_of":"","related_ids":["joint-space","workspace","end-effector-pose","inverse-kinematics","operational-space-control","jacobian-matrix"],"name":"Task Space","alt":"任务空间","abbr":"","aliases":["Cartesian Space","Operational Space"],"one_liner":"The space that directly describes what a task cares about, usually the end effector's position and orientation.","explanation":"Task space is the counterpart to joint space. Joint space describes a robot by its joint angles; task space directly describes whatever the task actually cares about, which by default means the end effector's — the gripper or tool's — position and orientation. Per Lynch and Park's textbook Modern Robotics, how task space is defined depends on the task, not the robot: drawing on paper only needs a flat 2D coordinate, while manipulating a rigid body needs a full 6D pose. Because it's commonly expressed in Cartesian xyz coordinates, it's also called Cartesian space; Oussama Khatib introduced operational-space control in 1980, with the landmark paper following in 1987, and called it operational space. Converting joint angles to task space uses forward kinematics; the reverse uses inverse kinematics. Some VLA models define their actions in task space, others work directly in joint space.","example":"OpenVLA outputs a 3D end-effector translation, 3D rotation, and 1D gripper open/close value — a task-space action. ACT on the ALOHA platform instead outputs a 14-dimensional joint target across both arms directly — a joint-space action.","related":["Joint Space","Workspace","End-Effector Pose","Inverse Kinematics (IK)","Operational Space Control","Jacobian Matrix"]},{"id":"end-effector-pose","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"robosuite Documentation: Controllers（OSC_POSE）","url":"https://robosuite.ai/docs/modules/controllers.html"},{"title":"libfranka robot_state.h（O_T_EE）","url":"https://raw.githubusercontent.com/frankarobotics/libfranka/main/include/franka/robot_state.h"},{"title":"Wikipedia: Robot end effector","url":"https://en.wikipedia.org/wiki/Robot_end_effector"}],"as_of":"","related_ids":["end-effector","pose","tool-center-point","task-space","inverse-kinematics","delta-action-vs-absolute-action"],"name":"End-Effector Pose","alt":"末端位姿","abbr":"EEF Pose","aliases":["EEF Pose","EE Pose","TCP Pose"],"one_liner":"The position and orientation of an arm's gripper or tool in space, usually given relative to the base frame.","explanation":"The end effector is the part at the very tip of a robot arm that actually interacts with the world — a gripper, a suction cup, a welding torch. The end-effector pose is its position (x, y, z) plus its orientation, 6 degrees of freedom altogether, generally expressed relative to the base frame. Orientation can be written as a rotation matrix, a quaternion, Euler angles, or axis-angle, so the end-effector pose's dimensionality can be 6, 7, or more depending on the dataset. It's the bridge between “task” and “motors”: a person or a vision model cares where the gripper is and which way it's pointing, while the motors only understand joint angles, and forward and inverse kinematics convert between the two. Many VLA (vision-language-action) and imitation-learning policies directly output an end-effector pose, or a delta to it, which a controller then converts into joint commands.","example":"robosuite's OSC_POSE controller takes a 6-dimensional action: the first 3 values are a delta to end-effector position, the last 3 are a delta to orientation in axis-angle form. In Franka's libfranka interface, O_T_EE is the 4×4 matrix giving the end-effector's pose in the base frame.","related":["End Effector","Pose","Tool Center Point","Task Space","Inverse Kinematics (IK)","Delta (Relative) Action vs. Absolute Action"]},{"id":"tool-center-point","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"RoboDK Documentation: General Tips (Tool Center Point)","url":"https://robodk.com/doc/en/General.html"},{"title":"libfranka robot_state.h (F_T_EE, O_T_EE)","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot_state.h"}],"as_of":"","related_ids":["end-effector","tool-flange","end-effector-pose","homogeneous-transformation-matrix","hand-eye-calibration","absolute-positioning-accuracy"],"name":"Tool Center Point","alt":"工具中心点","abbr":"TCP","aliases":["TCP","Tool Frame","TCP Point"],"one_liner":"The point on the end-of-arm tool that actually does the work; the robot aims and moves this point, not the mounting flange.","explanation":"The tool center point (TCP) is a reference point defined on the end-of-arm tool, together with its own coordinate frame — for example, the midpoint between a gripper's two fingertips, the tip of a welding torch, or the center of a suction cup. The arm itself only knows where its flange (the end mounting face) is, so once a tool is attached, a fixed transform from the flange to the TCP (a position plus an orientation) must be configured before the controller can send the tool tip to a target pose, move it in a straight line, or rotate it about a point. Get the TCP wrong and the tool offsets as a whole — most noticeably when rotating about a point — so the TCP is usually recalibrated whenever the tool changes. In embodied AI, whether a dataset's or a policy's 'end-effector pose' refers to the flange or the TCP varies by convention and needs to be unified across cross-embodiment training; Franka's libfranka library, for instance, reports the flange frame and the end-effector (EE) frame separately.","example":"After mounting a parallel gripper on the arm, set the TCP at the midpoint between the two fingertips: commanding the robot to rotate about the TCP then keeps the fingertip position fixed while only changing the gripper's orientation. If the TCP were still left at the flange, the fingertips would instead swing out along an arc.","related":["End Effector","Tool Flange","End-Effector Pose","Homogeneous Transformation Matrix","Hand-Eye Calibration","Absolute Positioning Accuracy"]},{"id":"forward-kinematics","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Forward kinematics","url":"https://en.wikipedia.org/wiki/Forward_kinematics"},{"title":"Modern Robotics（Lynch & Park）Ch.4 Forward Kinematics","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["inverse-kinematics","denavit-hartenberg-parameters","homogeneous-transformation-matrix","kinematic-chain","product-of-exponentials-formula","joint-space"],"name":"Forward Kinematics (FK)","alt":"正运动学","abbr":"FK","aliases":["FK"],"one_liner":"Computing the position and orientation of an arm's end effector in space, given all the joint angles.","explanation":"Forward kinematics uses a robot's geometry — link lengths, joint axis directions — to compute the end-effector pose from the joint angles. The standard method starts at the base and multiplies a 4×4 transformation matrix for each joint and link in turn, all the way out to the tip; the result is the end-effector's pose in the base frame. The classical way to describe each of these per-joint transforms is the set of DH (Denavit-Hartenberg) parameters, introduced by Jacques Denavit and Richard Hartenberg in 1955, which uses 4 parameters per joint; the product-of-exponentials formula is another common approach. For a serial-chain arm, forward kinematics has a unique solution and takes only a handful of matrix multiplications to compute. It's used to display the end-effector's position from encoder readings, to render a robot in simulation, to convert recorded joint data into end-effector pose labels, and it's also the foundation that inverse kinematics builds on.","example":"For a planar two-link arm with link lengths l₁ and l₂ and joint angles θ₁ and θ₂, the end-effector coordinates are x = l₁cosθ₁ + l₂cos(θ₁+θ₂), y = l₁sinθ₁ + l₂sin(θ₁+θ₂).","related":["Inverse Kinematics (IK)","Denavit-Hartenberg (DH) Parameters","Homogeneous Transformation Matrix","Kinematic Chain","Product of Exponentials Formula","Joint Space"]},{"id":"inverse-kinematics","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Inverse kinematics","url":"https://en.wikipedia.org/wiki/Inverse_kinematics"},{"title":"MathWorks: Inverse Kinematics Algorithms","url":"https://www.mathworks.com/help/robotics/ug/inverse-kinematics-algorithms.html"}],"as_of":"","related_ids":["forward-kinematics","analytical-inverse-kinematics","numerical-inverse-kinematics","jacobian-matrix","singular-configuration","kinematic-redundancy"],"name":"Inverse Kinematics (IK)","alt":"逆运动学","abbr":"IK","aliases":["IK"],"one_liner":"Working backward from a target end-effector position and orientation to find the joint angles that reach it.","explanation":"Inverse kinematics is the reverse of forward kinematics: given a target end-effector pose, find the joint angles that achieve it. It's considerably harder. The equations are nonlinear and can have multiple solutions — an elbow-up or elbow-down configuration might reach the same point — a redundant arm with 7 degrees of freedom has infinitely many solutions, a target outside the workspace has no solution at all, and the numerics get unstable near singular configurations (joint poses where the end effector can't move in certain directions). Solving methods fall into two camps: analytical (closed-form) solutions compute the answer directly from a formula, which is fast but only works for specific arm geometries; numerical solutions iterate using the Jacobian matrix (the mapping from joint velocities to end-effector velocity), which is general-purpose but depends on a good starting guess and may fail to converge. Whenever a VLA policy outputs an end-effector pose, IK is what turns it into joint commands; dragging a character's hand in animation software, with the shoulder and elbow following automatically, is also IK.","example":"To move the gripper to 10 centimeters directly above a cup, jaws facing down, IK solves for the joint angles that achieve it. In MATLAB's Robotics System Toolbox, the inverseKinematics function solves numerically using BFGS gradient projection by default, and can return only an approximate answer if the initial guess is poor.","related":["Forward Kinematics (FK)","Analytical Inverse Kinematics","Numerical Inverse Kinematics","Jacobian Matrix","Singular Configuration (Kinematic Singularity)","Kinematic Redundancy"]},{"id":"workspace","category":"mechanics","sec":3,"tier":1,"sources":[{"title":"Modern Robotics (Lynch & Park, 2017), 2.5 Task Space and Workspace","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"},{"title":"Wikipedia: Work envelope","url":"https://en.wikipedia.org/wiki/Work_envelope"},{"title":"UPenn MEAM 520 讲义: Manipulators（reachable / dexterous workspace）","url":"https://medesign.seas.upenn.edu/uploads/Courses/520-12A-T02.pdf"},{"title":"Franka Robotics: Franka Research 3","url":"https://franka.de/franka-research-3"}],"as_of":"","related_ids":["task-space","configuration-space","reach","reachability-map-inverse-reachability-map","singular-configuration","joint-limits"],"name":"Workspace","alt":"工作空间","abbr":"","aliases":["Reachable Workspace","Work Envelope"],"one_liner":"The region of positions, and sometimes orientations, that a robot's end effector can reach.","explanation":"The workspace is the set of positions — sometimes including orientations — that a robot's end effector can reach, determined by its link lengths, joint types, and joint limits, independent of any particular task; industrial-robot datasheets often draw it as a side-view and top-view outline. Per Modern Robotics, every point in the workspace can be reached by at least one joint configuration, while a point in task space isn't necessarily reachable at all. Textbooks often split it further: the reachable workspace is every point the end effector can reach in at least one orientation, while the dexterous workspace is every point it can reach in any orientation. The dexterous workspace is only a subset of the reachable one, and the two terms shouldn't be used interchangeably. Before placing a robot next to a table, choosing a camera's field of view, or generating a simulated task, it's standard practice to first confirm the target actually falls inside the workspace.","example":"A planar two-link arm where both links have length L has a disk-shaped end-effector workspace of radius 2L, assuming the joints are unconstrained and can each turn a full revolution. Franka Research 3 has an 855 mm reach, so it simply can't grasp anything noticeably farther from its base than that.","related":["Task Space","Configuration Space (C-Space)","Reach (Arm Reach / Working Radius)","Reachability Map / Inverse Reachability Map","Singular Configuration (Kinematic Singularity)","Joint Limits"]},{"id":"reach","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"Franka Robotics: Franka Research 3","url":"https://franka.de/franka-research-3"},{"title":"Modern Robotics (Lynch & Park), Ch.2 Task Space and Workspace","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"2026-09","related_ids":["workspace","robotic-arm","payload","singular-configuration","reachability-map-inverse-reachability-map","degrees-of-freedom"],"name":"Reach (Arm Reach / Working Radius)","alt":"臂展（工作半径）","abbr":"","aliases":["Arm Reach","Working Radius","Max Reach"],"one_liner":"The farthest distance a fully extended arm can reach from its base center — a basic spec for choosing a robot.","explanation":"Reach is a basic spec on a robot arm's datasheet, generally meaning the maximum distance from the base's rotation center to the wrist center or the flange (the interface where the end-of-arm tool mounts); the exact measurement point depends on each manufacturer's datasheet. It roughly outlines the outer boundary of the workspace (every position the end effector can reach) — an arm with an 0.85 m reach can cover, very roughly, a region of radius 0.85 m centered on the base. But reach doesn't equal a genuinely usable range: as the arm nears full extension it approaches a singular configuration (a posture where the end effector can't move in certain directions), losing dexterity and constraining end-effector orientation; the gripper's own length and the torque limits imposed by the payload also eat into the practical range. For tabletop or mobile manipulation, reach determines how deep a table can be or how close a mobile base has to park, and it's also commonly used to compare a humanoid's arm against a human arm.","example":"Franka Research 3's official specs list an 855 mm reach, a 3 kg payload, and 7 degrees of freedom. Mounted at the edge of a table, it likely can't reach an object about 1 meter from that edge — the base would need to move closer, or the object would need to be moved in.","related":["Workspace","Robotic Arm","Payload","Singular Configuration (Kinematic Singularity)","Reachability Map / Inverse Reachability Map","Degrees of Freedom (DoF)"]},{"id":"denavit-hartenberg-parameters","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"Denavit–Hartenberg parameters - Wikipedia","url":"https://en.wikipedia.org/wiki/Denavit%E2%80%93Hartenberg_parameters"},{"title":"DH Parameters for calculations of kinematics and dynamics - Universal Robots","url":"https://www.universal-robots.com/articles/ur/application-installation/dh-parameters-for-calculations-of-kinematics-and-dynamics/"}],"as_of":"","related_ids":["modified-denavit-hartenberg-parameters","forward-kinematics","homogeneous-transformation-matrix","product-of-exponentials-formula","unified-robot-description-format","link"],"name":"Denavit-Hartenberg (DH) Parameters","alt":"DH参数","abbr":"DH","aliases":["DH Parameters","Standard DH","DH Convention"],"one_liner":"A convention for modeling an arm where just four numbers per joint describe the transform between neighboring link frames.","explanation":"DH parameters are a convention introduced by Jacques Denavit and Richard Hartenberg in a 1955 paper in the Journal of Applied Mechanics: following a fixed rule, each link gets its own coordinate frame, and the transform between neighboring frames then needs only 4 numbers — the offset d along the previous z-axis, the rotation angle θ about the previous z-axis, the length a of the common perpendicular between the two joint axes (the link length), and the twist angle α about that common perpendicular. For a revolute joint, θ is the joint variable and the other three are fixed constants. Multiplying each joint's 4×4 homogeneous transformation matrix together in sequence gives the forward kinematics from base to end effector. Richard Paul extended the convention to robotics broadly in 1981. The “modified DH” convention used in John Craig's textbook places the frames and orders the transforms differently, so the two parameter tables are not interchangeable.","example":"Universal Robots' official standard DH parameters for the UR5e: joint 1 has d = 0.1625 m and α = π/2; joints 2 and 3 have a = −0.425 m and −0.3922 m (the upper-arm and forearm lengths); joints 4–6 have d = 0.1333, 0.0997, and 0.0996 m. Plugging these into the formula computes the end-effector pose from the 6 joint angles.","related":["Modified Denavit-Hartenberg Parameters","Forward Kinematics (FK)","Homogeneous Transformation Matrix","Product of Exponentials Formula","Unified Robot Description Format","Link"]},{"id":"modified-denavit-hartenberg-parameters","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Denavit–Hartenberg parameters（Modified DH parameters 一节）","url":"https://en.wikipedia.org/wiki/Denavit%E2%80%93Hartenberg_parameters"},{"title":"Franka Robotics 文档：Robot and interface specifications（DH 参数，遵循 Craig 约定）","url":"https://frankarobotics.github.io/docs/robot_specifications.html"},{"title":"Lynch & Park, Modern Robotics（2017）附录 C：Denavit–Hartenberg Parameters","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["denavit-hartenberg-parameters","homogeneous-transformation-matrix","forward-kinematics","product-of-exponentials-formula","kinematic-calibration","franka-emika-panda-franka-research-3"],"name":"Modified Denavit-Hartenberg Parameters","alt":"改进 DH 参数","abbr":"MDH","aliases":["MDH","Craig DH","Craig Convention"],"one_liner":"A variant of DH parameters that places each link's frame on the joint axis nearer the base, with a different transform order.","explanation":"Modified DH parameters are the convention John J. Craig adopted in his textbook Introduction to Robotics: Mechanics and Control. Like the standard DH convention introduced by Denavit and Hartenberg in 1955, it describes the relationship between two adjacent frames using 4 parameters per joint: link length a, twist angle α, offset d, and joint angle θ. The difference is in where frame i is placed and in what order the transforms are composed: frame i sits on joint i's own axis (the proximal end of the link), whereas standard DH places it on joint i+1's axis (the distal end); the transform order becomes Rot_x(α_{i−1})·Trans_x(a_{i−1})·Rot_z(θ_i)·Trans_z(d_i) — rotate and translate about the previous x-axis first, then rotate and translate about this joint's own z-axis — which is why a and α in the parameter table carry an index one less than θ and d. The two parameter conventions cannot be mixed, or forward kinematics will come out wrong. Appendix C of Modern Robotics also uses this convention.","example":"Franka Research 3's official documentation notes that its DH table 'follows the Craig convention' — for example, joint 1 has a = 0, d = 0.333 m, α = 0, and joint 4 has a = 0.0825 m, d = 0, α = π/2. Plugging this table directly into the standard-DH formula would give the wrong end-effector pose.","related":["Denavit-Hartenberg (DH) Parameters","Homogeneous Transformation Matrix","Forward Kinematics (FK)","Product of Exponentials Formula","Kinematic Calibration","Franka Emika Panda / Franka Research 3"]},{"id":"product-of-exponentials-formula","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Product of exponentials formula","url":"https://en.wikipedia.org/wiki/Product_of_exponentials_formula"},{"title":"Modern Robotics (Lynch & Park), Sec. 4.1 Product of Exponentials Formula","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["screw-theory","denavit-hartenberg-parameters","forward-kinematics","exponential-map","lie-group","kinematic-calibration"],"name":"Product of Exponentials Formula","alt":"指数积公式","abbr":"PoE","aliases":["PoE","Product of Exponentials"],"one_liner":"Treats each joint as motion about a screw axis, and multiplies each joint's matrix exponential together to get the end-effector pose.","explanation":"The product of exponentials formula was introduced by Roger Brockett in 1984, and is a way of describing a serial manipulator's forward kinematics; Lynch and Park's textbook Modern Robotics builds its whole treatment around it. The space form is written T(θ) = e^[S₁]θ₁ ⋯ e^[Sₙ]θₙ M: M is the end-effector's pose when every joint is at zero, Sᵢ is joint i's screw axis expressed in the fixed base frame (6-dimensional, combining a rotation direction and a linear-velocity part), θᵢ is the joint's angle or displacement, and e^[S]θ denotes the rigid-body transformation produced by moving θ along that screw axis. There's also a body form with the screw axes expressed in the end-effector frame, T = M e^[B₁]θ₁ ⋯. Compared with DH parameters, it only needs two frames — base and end effector — treats revolute and prismatic joints uniformly, and has a directly intuitive geometric meaning, at the cost of not using the fewest possible parameters. Jacobians, inverse kinematics, and kinematic calibration can all be derived on top of it.","example":"A single-joint planar arm: a link of length L rotating about the base's z-axis. At the zero configuration the end effector is at (L, 0, 0), which is M; the screw axis is S = (0, 0, 1, 0, 0, 0). T(θ) = e^[S]θ M then gives the end-effector position (L·cosθ, L·sinθ, 0).","related":["Screw Theory","Denavit-Hartenberg (DH) Parameters","Forward Kinematics (FK)","Exponential Map","Lie Group","Kinematic Calibration"]},{"id":"numerical-inverse-kinematics","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"Modern Robotics 6.2: Numerical Inverse Kinematics (Part 1 of 2)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/6-2-numerical-inverse-kinematics-part-1-of-2/"},{"title":"Modern Robotics 6.2: Numerical Inverse Kinematics (Part 2 of 2)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/6-2-numerical-inverse-kinematics-part-2-of-2/"},{"title":"Wikipedia: Inverse kinematics","url":"https://en.wikipedia.org/wiki/Inverse_kinematics"}],"as_of":"","related_ids":["inverse-kinematics","analytical-inverse-kinematics","jacobian-pseudoinverse","damped-least-squares","jacobian-transpose-method","trac-ik"],"name":"Numerical Inverse Kinematics","alt":"数值逆解","abbr":"","aliases":["Numerical IK","Iterative IK"],"one_liner":"An inverse-kinematics method that starts from an initial guess and repeatedly refines the joint angles to approach a target pose.","explanation":"Inverse kinematics has to find joint angles that produce a desired end-effector pose. Analytical IK derives a closed-form formula and only works for arms with special geometry; numerical IK is more general-purpose: starting from an initial set of joint angles, it uses forward kinematics to compute the current error between the end effector and the target, converts that error into a joint-angle correction using the Jacobian matrix, and repeats until the error is small enough. The most basic version is the Newton-Raphson method with the Jacobian pseudoinverse; common refinements include damped least squares (which keeps joint angles from swinging wildly near a singularity), the Jacobian transpose method, and formulating the whole problem as an optimization with joint limits and obstacle avoidance written in as constraints. The downside is that it only converges to whichever solution happens to be near the starting guess, and it can fail to converge or get stuck in a local optimum. In real-time control, using the previous time step's joint angles as the initial guess usually means convergence in just a few iterations. Libraries like KDL, TRAC-IK, and mink all provide numerical IK.","example":"During VR teleoperation, each frame takes the headset controller's pose as the arm's target end-effector pose, and runs a few iterations of numerical IK starting from the previous frame's joint angles to get real-time joint angles.","related":["Inverse Kinematics (IK)","Analytical Inverse Kinematics","Jacobian Pseudoinverse","Damped Least Squares","Jacobian Transpose Method","TRAC-IK"]},{"id":"analytical-inverse-kinematics","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"Modern Robotics: Mechanics, Planning, and Control（Lynch & Park, 2017 预印本）第 6 章 Inverse Kinematics","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Inverse kinematics - Wikipedia","url":"https://en.wikipedia.org/wiki/Inverse_kinematics"},{"title":"IKFast Kinematics Solver - MoveIt Documentation","url":"https://moveit.picknik.ai/main/doc/examples/ikfast/ikfast_tutorial.html"}],"as_of":"","related_ids":["inverse-kinematics","numerical-inverse-kinematics","forward-kinematics","spherical-wrist","ikfast","kinematic-redundancy"],"name":"Analytical Inverse Kinematics","alt":"解析逆解","abbr":"","aliases":["Closed-Form IK","Analytic IK"],"one_liner":"An inverse-kinematics method that plugs a target end-effector pose into pre-derived formulas to get all joint angles directly.","explanation":"Inverse kinematics — finding joint angles from a target end-effector pose — has two solving approaches. Numerical methods iterate toward an answer; analytical methods derive formulas ahead of time and plug in the target pose to get every joint angle in one shot. Closed-form solutions only exist for specific arm geometries; the classic enabling condition for a 6-axis arm is that the last three joint axes all intersect at a single point (a spherical wrist), which lets position and orientation be solved separately — this is exactly the PUMA-style arm layout. The benefit is speed, numerical stability, and the ability to enumerate every solution: Modern Robotics notes that a general 6R serial arm has at most 16 solutions, while a PUMA-style arm with offsets has 4 solutions just for the position sub-problem. The downside is that a different arm geometry requires re-deriving everything from scratch. The tool IKFast can automatically analyze a kinematic chain and generate C++ code for its analytical solution, solving in just a few microseconds per call.","example":"A PUMA-style 6-axis arm first solves the first three joints from the wrist-center position (shoulder left/right times elbow up/down gives 4 combinations), then solves the last three wrist joints from the end-effector orientation, with the wrist able to flip as well; an arm with this spherical-wrist layout has up to 8 solutions for a given end-effector pose, and the controller typically picks whichever is closest to the current joint angles.","related":["Inverse Kinematics (IK)","Numerical Inverse Kinematics","Forward Kinematics (FK)","Spherical Wrist","IKFast","Kinematic Redundancy"]},{"id":"spherical-wrist","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Donald L. Pieper","url":"https://en.wikipedia.org/wiki/Donald_L._Pieper"},{"title":"Wikipedia: Inverse kinematics","url":"https://en.wikipedia.org/wiki/Inverse_kinematics"}],"as_of":"","related_ids":["inverse-kinematics","analytical-inverse-kinematics","6-axis-robot-arm","ikfast","denavit-hartenberg-parameters","singular-configuration"],"name":"Spherical Wrist","alt":"球形手腕","abbr":"","aliases":["Pieper Criterion","Wrist-Partitioned Manipulator"],"one_liner":"A wrist structure where a manipulator's last three joint axes all intersect at a single point.","explanation":"A spherical wrist means the axes of a 6-DoF manipulator's last three revolute joints all intersect at one point (the wrist center), behaving much like a single ball joint: it only changes the end effector's orientation without moving the wrist center's position. Donald Pieper proved in his 1968 Stanford PhD thesis that a serial manipulator with 6 revolute joints, three of which are adjacent and share a common intersection point, has a closed-form inverse kinematics solution — a result later named the Pieper criterion. The reason is that position and orientation decouple: the wrist center's required position can be worked out first from the target pose, and the first three joints alone can place the wrist center there; the last three joints then supply the required orientation. This means the inverse solution can be computed directly with formulas instead of iterating, and every solution can be enumerated (up to 8 solution sets for this kind of configuration). Many industrial six-axis arms use an orthogonal-parallel-base-plus-spherical-wrist layout, with the PUMA 560 as a classic example.","example":"Given a target pose for this kind of arm: first step back along the tool's approach direction by a fixed distance to find where the wrist center should be, use the first three joints to send the wrist center there, then solve for the last three joint angles to orient the gripper as required.","related":["Inverse Kinematics (IK)","Analytical Inverse Kinematics","6-Axis Robot Arm","IKFast","Denavit-Hartenberg (DH) Parameters","Singular Configuration (Kinematic Singularity)"]},{"id":"reachability-map-inverse-reachability-map","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"RM4D: A Combined Reachability and Inverse Reachability Map for Common 6-/7-axis Robot Arms (arXiv 2410.06968)","url":"https://arxiv.org/html/2410.06968"}],"as_of":"","related_ids":["workspace","inverse-kinematics","mobile-manipulation","mobile-manipulator","grasp-planning","manipulability"],"name":"Reachability Map / Inverse Reachability Map","alt":"可达性地图 / 逆可达性地图","abbr":"RM / IRM","aliases":["RM / IRM","Capability Map","Inverse Capability Map"],"one_liner":"A precomputed record of which end-effector poses an arm can reach; the inverse version looks up where to place the base to reach a target.","explanation":"A reachability map is an offline precomputation of a manipulator's workspace: the space around the end effector is divided into 3D voxels, several orientations are sampled within each voxel, and inverse kinematics checks whether each one is reachable; the results (often expressed as the fraction of reachable orientations) are stored in a lookup table, so the arm doesn't have to re-solve inverse kinematics repeatedly at run time. Zacharias and colleagues introduced this kind of capability map in 2007. The inverse reachability map inverts every reachable pose to instead give a distribution of where the base can be placed relative to a target end-effector pose — Vahrenkamp and colleagues used it in 2013 to solve base placement for mobile manipulation. Both are commonly used to filter out unreachable grasps during grasp planning, to choose a standing position for mobile manipulation, and to evaluate a robot's design and mounting location; a 2024 method called RM4D compresses both into a single 4D data structure.","example":"A mobile robot approaching a table to pick up a cup: it queries the inverse reachability map centered on the candidate grasp pose, getting a region on the floor from which that grasp is reachable, filters out spots that would collide with the table or are unreachable, and picks the point with the highest reachability score as its navigation target.","related":["Workspace","Inverse Kinematics (IK)","Mobile Manipulation","Mobile Manipulator","Grasp Planning","Manipulability"]},{"id":"joint-zero-position-calibration","category":"mechanics","sec":3,"tier":2,"sources":[{"title":"Incremental encoder - Wikipedia","url":"https://en.wikipedia.org/wiki/Incremental_encoder"},{"title":"ROBOTIS e-Manual: XM430-W350 (Homing Offset)","url":"https://emanual.robotis.com/docs/en/dxl/x/xm430-w350/"},{"title":"LeRobot Docs: SO-101 (Calibrate)","url":"https://huggingface.co/docs/lerobot/so101"}],"as_of":"","related_ids":["rotary-encoder","incremental-encoder","absolute-encoder","kinematic-calibration","joint-limits","leader-follower-teleoperation"],"name":"Joint Zero-Position Calibration (Homing / Offset Calibration)","alt":"零位标定（零点标定）","abbr":"","aliases":["Homing","Zero Offset Calibration","Homing Offset"],"one_liner":"Measuring the actual joint angle when a sensor reads zero, so the software model matches the physical robot.","explanation":"A robot's forward kinematics and its URDF model both assume that when every joint angle reads 0, the robot is in one specific, well-defined posture (the zero position). But when an encoder is physically installed, its reading origin is essentially arbitrary, and assembly tolerances or a motor swap can leave the reading offset from the model's angle by some fixed amount. Zero-position calibration measures that offset and stores it, so every later reading gets corrected by it. A joint using an incremental encoder (which only tracks relative changes) needs to be re-homed every time it powers on: it's driven to a mechanical hard stop, a limit switch, or an encoder index pulse, and the count is reset there; an absolute encoder usually only needs calibrating once. Servos have an analogous parameter too, such as Dynamixel's Homing Offset. An inaccurate zero position introduces a systematic error into the end-effector position, and also means the same policy won't transfer cleanly to a different physical robot.","example":"After assembling a LeRobot SO-101 arm, you run lerobot-calibrate: first every joint is moved to the middle of its own range of motion, then each is rotated through its full range in turn. The official docs explain this step ensures the leader and follower arms read the same value at the same physical pose, which is what lets a trained model transfer to another robot.","related":["Rotary Encoder","Incremental Encoder","Absolute Encoder","Kinematic Calibration","Joint Limits","Leader-Follower Teleoperation"]},{"id":"kinematic-calibration","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Robot calibration","url":"https://en.wikipedia.org/wiki/Robot_calibration"},{"title":"A Visual Kinematics Calibration Method for Manipulator Based on Nonlinear Optimization (arXiv:2005.08420)","url":"https://arxiv.org/abs/2005.08420"},{"title":"Bayesian Optimal Experimental Design for Robot Kinematic Calibration (arXiv:2409.10802)","url":"https://arxiv.org/abs/2409.10802"}],"as_of":"","related_ids":["denavit-hartenberg-parameters","modified-denavit-hartenberg-parameters","joint-zero-position-calibration","absolute-positioning-accuracy","pose-repeatability","hand-eye-calibration"],"name":"Kinematic Calibration","alt":"运动学标定","abbr":"","aliases":["DH Parameter Calibration","Robot Calibration"],"one_liner":"Measuring a robot's actual end-effector positions to correct the model's link lengths, joint zero offsets, and other geometric parameters.","explanation":"A controller computes end-effector position from a nominal set of kinematic parameters — such as DH parameters: link length, twist angle, offset, and joint zero position — but manufacturing tolerances and wear cause the real parameters to drift from the drawings, so the robot ends up with good repeatability (returning to the same point) but poor absolute accuracy (reaching a specified coordinate). Kinematic calibration corrects these parameters: the robot is driven through many poses, an external device such as a laser tracker or camera measures the true end-effector position, the measurements are compared against the model's predictions, and an optimization such as least squares solves for the parameter errors, which are then written back into the controller as a correction. Calibration is usually split into three levels: level one calibrates only the joint zero offsets (the gap between encoder readings and true angles), level two calibrates all the geometric parameters, and level three adds non-geometric errors such as joint flexibility and friction. According to Wikipedia, calibrating a six-axis industrial robot can improve absolute accuracy by roughly an order of magnitude, usually bringing the error down to under 1 millimeter.","example":"Offline-programmed trajectories, or a single trajectory shared across several arms of the same model, will land in different places on an uncalibrated robot. One approach mounts a calibration board on the end effector, photographs it with a single camera across a few dozen poses, and optimizes the corrected DH parameters against pixel reprojection error (arXiv:2005.08420).","related":["Denavit-Hartenberg (DH) Parameters","Modified Denavit-Hartenberg Parameters","Joint Zero-Position Calibration (Homing / Offset Calibration)","Absolute Positioning Accuracy","Pose Repeatability","Hand-Eye Calibration"]},{"id":"absolute-positioning-accuracy","category":"mechanics","sec":3,"tier":3,"sources":[{"title":"Robot calibration - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_calibration"},{"title":"RoboDK Documentation: Robot Calibration (Laser Tracker)","url":"https://robodk.com/doc/en/Robot-Calibration-LaserTracker.html"}],"as_of":"","related_ids":["pose-repeatability","kinematic-calibration","offline-programming","tool-center-point","hand-eye-calibration","digital-twin"],"name":"Absolute Positioning Accuracy","alt":"绝对定位精度","abbr":"","aliases":["Pose Accuracy","Absolute Accuracy"],"one_liner":"How far a robot's actual end-effector position ends up from the theoretical target position it was commanded to reach.","explanation":"Absolute positioning accuracy measures how close the end effector's actual position comes to the theoretical position implied by the command; the ISO 9283 standard calls this pose accuracy (AP). It's distinct from repeatability (RP) — how tightly the robot returns to the same point over repeated trials. Industrial arms are typically 'very repeatable but not very accurate,' because manufacturing tolerances in link lengths and joint zero offsets, together with deflection under load and temperature effects, cause the controller's internal kinematic model to drift away from the real machine. Hand-guiding taught points with a teach pendant only requires good repeatability, but offline programming — or sending coordinates computed in simulation or by a camera straight to the robot — is instead bottlenecked by absolute accuracy. The usual fix is kinematic calibration: measuring the true pose with equipment such as a laser tracker and using it to identify the model's errors, which typically improves a six-axis industrial arm's absolute accuracy several-fold to tenfold, often down to well under 1 millimeter.","example":"A sequence of grasp points is planned offline in simulation and sent straight to the real robot: the arm returns to the same physical spot every time (good repeatability), but that spot is offset from the intended target in simulation by a consistent amount — that offset is the absolute positioning error, and in vision-guided grasping it also stacks with hand-eye calibration error.","related":["Pose Repeatability","Kinematic Calibration","Offline Programming","Tool Center Point","Hand-Eye Calibration","Digital Twin"]},{"id":"differential-kinematics","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Robotic Manipulation (MIT, Russ Tedrake) - Basic Pick and Place: Differential kinematics","url":"https://manipulation.csail.mit.edu/pick.html"},{"title":"kevinzakka/mink - GitHub","url":"https://github.com/kevinzakka/mink"},{"title":"stephane-caron/pink - GitHub","url":"https://github.com/stephane-caron/pink"}],"as_of":"","related_ids":["jacobian-matrix","inverse-kinematics","jacobian-pseudoinverse","damped-least-squares","singular-configuration","mink"],"name":"Differential Kinematics","alt":"微分运动学","abbr":"","aliases":["Velocity Kinematics","Differential Inverse Kinematics (Diff IK)"],"one_liner":"The kinematics that relates joint velocities to end-effector velocity through the Jacobian matrix.","explanation":"Forward kinematics answers “where is the end effector given these joint angles”; differential kinematics answers “how fast does the end effector move given these joint speeds.” The bridge between them is the Jacobian matrix J(q), the partial derivative of forward kinematics with respect to joint angles q: V = J(q)·q̇, where V is the end effector's 6-dimensional spatial velocity (3D linear velocity plus 3D angular velocity) and q̇ is the vector of joint velocities. Running this backward is differential inverse kinematics: given a desired end-effector velocity V_d, the pseudoinverse (a generalization of “inverse” to non-square matrices) gives q̇ = J⁺V_d, which is then integrated into joint-angle commands, repeated every control cycle. It avoids having to solve a full analytical inverse-kinematics problem each time, which suits teleoperation and policies that output end-effector deltas well — but near a singular configuration (where J loses rank), the pseudoinverse can blow up, so practical implementations often use damped least squares instead, or frame the problem as a quadratic program with joint limits built in. Pink and mink are libraries built around this approach.","example":"During VR teleoperation, each control cycle converts the controller's displacement into a desired end-effector velocity, and q̇ = J⁺V_d gives the speed each of the 7 joints should turn at, so the arm's end effector smoothly tracks the operator's hand.","related":["Jacobian Matrix","Inverse Kinematics (IK)","Jacobian Pseudoinverse","Damped Least Squares","Singular Configuration (Kinematic Singularity)","mink (MuJoCo inverse kinematics)"]},{"id":"jacobian-matrix","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Modern Robotics 5.1.1: Space Jacobian (Northwestern)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-1-1-space-jacobian/"},{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 5 Velocity Kinematics and Statics","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["differential-kinematics","jacobian-pseudoinverse","singular-configuration","inverse-kinematics","manipulability","geometric-vs-analytical-jacobian"],"name":"Jacobian Matrix","alt":"雅可比矩阵","abbr":"J","aliases":["Spatial Jacobian","Body Jacobian"],"one_liner":"The matrix that converts joint speeds into end-effector velocity; it changes as the robot's posture changes.","explanation":"In mathematics, a Jacobian matrix is the array of first-order partial derivatives of a multivariable function. In robotics, it describes the linear relationship between joint velocities and end-effector velocity: V = J(θ)·θ̇. Here θ̇ is the velocity of n joints; V is the end effector's 6-dimensional velocity (3D angular velocity plus 3D linear velocity, also called a twist); and J is a 6×n matrix whose i-th column shows how the end effector moves when only the i-th joint turns at unit speed. J depends on the current joint angles θ, so it has to be recomputed whenever the posture changes. It has three main uses: its inverse or pseudoinverse converts a desired end-effector velocity back into joint velocities, which is the basis of numerical inverse kinematics; checking whether J has lost rank reveals a singular configuration, where the end effector can't move in certain directions; and the static-force mapping τ = Jᵀ·F converts a desired end-effector force F into joint torques τ, which underlies force control and impedance control.","example":"When a planar two-link arm is fully extended, no combination of joint rotations can move the end effector any farther outward along the arm's own direction — at that point J has lost rank, and the arm is in a singular configuration.","related":["Differential Kinematics","Jacobian Pseudoinverse","Singular Configuration (Kinematic Singularity)","Inverse Kinematics (IK)","Manipulability","Geometric vs. Analytical Jacobian"]},{"id":"twist","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Modern Robotics 3.3.2: Twists","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-3-2-twists-part-1-of-2/"},{"title":"ROS 2 geometry_msgs/Twist.msg","url":"https://raw.githubusercontent.com/ros2/common_interfaces/rolling/geometry_msgs/msg/Twist.msg"}],"as_of":"","related_ids":["wrench","screw-theory","jacobian-matrix","adjoint-representation","product-of-exponentials-formula","cmd-vel-topic"],"name":"Twist","alt":"运动旋量（速度旋量）","abbr":"","aliases":["Spatial Velocity","Six-Dimensional Velocity","Velocity Twist"],"one_liner":"Packs a rigid body's angular and linear velocity into one 6-dimensional vector that fully describes its instantaneous motion.","explanation":"A twist is the 6-dimensional vector V = (ω, v) that describes a rigid body's instantaneous velocity: ω is the 3D angular velocity and v is the 3D linear velocity, both expressed in the same frame. The concept comes from screw theory: any rigid-body motion can be viewed as a rotation about some screw axis combined with a translation along that same axis, and the twist is simply that axis scaled by the rotation rate. Expressed in the world frame it's called a spatial twist; expressed in the body's own frame it's a body twist, and the two convert into each other through the adjoint transformation. What a manipulator's Jacobian actually computes is the map from joint velocities to the end effector's twist. Note that the ordering convention differs between sources: Lynch and Park's textbook Modern Robotics writes it as (angular velocity, linear velocity), while ROS's geometry_msgs/Twist message lists the linear component before the angular one.","example":"Sending a cmd_vel command to a mobile base in ROS means sending a Twist message: linear.x = 0.5 means drive forward at 0.5 m/s, angular.z = 0.3 means simultaneously turn left at 0.3 rad/s about the vertical axis, and the other four components are zero.","related":["Wrench","Screw Theory","Jacobian Matrix","Adjoint Representation","Product of Exponentials Formula","cmd_vel Topic (geometry_msgs/Twist velocity command)"]},{"id":"geometric-vs-analytical-jacobian","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（预印本 PDF，5.1.5 节 Alternative Notions of the Jacobian）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["jacobian-matrix","differential-kinematics","euler-angles","gimbal-lock","twist","singular-configuration"],"name":"Geometric vs. Analytical Jacobian","alt":"几何雅可比与解析雅可比","abbr":"","aliases":["Geometric Jacobian","Analytic Jacobian"],"one_liner":"Both map joint velocity to end-effector velocity; they differ in whether orientation is expressed as angular velocity or as coordinate derivatives.","explanation":"The Jacobian matrix describes the linear relationship between joint velocity q̇ and end-effector velocity, and it comes in two versions depending on how end-effector velocity is expressed. The geometric Jacobian outputs the end-effector's linear velocity together with its angular velocity ω (this corresponds to the space/body Jacobian in Modern Robotics; Siciliano and colleagues' textbook defines it slightly differently, and terminology isn't fully standardized). The analytical Jacobian first describes the end-effector pose with a minimal set of coordinates, such as position plus Euler angles, then differentiates those coordinates directly, giving ẋ = J_a q̇. Angular velocity is not the same as the derivative of the Euler angles — the two differ by a transformation matrix that depends on the orientation representation — and that matrix can become non-invertible at certain orientations (such as gimbal lock), which makes the analytical Jacobian fail there, even when the arm itself isn't at a kinematic singularity. Velocity control and statics (τ = JᵀF) generally use the geometric Jacobian, while error feedback or trajectory optimization done directly in coordinates like Euler angles uses the analytical one.","example":"Describing end-effector orientation with ZYX Euler angles hits gimbal lock at a pitch angle of ±90°: some directions of angular velocity can no longer be expressed through the finite derivatives of the Euler angles, so the analytical Jacobian becomes non-invertible there, while the geometric Jacobian remains perfectly well-behaved at the same orientation.","related":["Jacobian Matrix","Differential Kinematics","Euler Angles","Gimbal Lock","Twist","Singular Configuration (Kinematic Singularity)"]},{"id":"singular-configuration","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Modern Robotics 5.3: Singularities (Northwestern)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-3-singularities/"},{"title":"Modern Robotics 6.2: Numerical Inverse Kinematics (Part 1 of 2)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/6-2-numerical-inverse-kinematics-part-1-of-2/"}],"as_of":"","related_ids":["jacobian-matrix","manipulability","damped-least-squares","jacobian-pseudoinverse","spherical-wrist","workspace"],"name":"Singular Configuration (Kinematic Singularity)","alt":"奇异位形","abbr":"","aliases":["Kinematic Singularity","Singularity","Wrist Singularity","Boundary Singularity"],"one_liner":"An arm posture where the Jacobian matrix loses rank, so the end effector can't move in certain directions.","explanation":"A singular configuration is a special class of arm posture. The Jacobian matrix — which converts joint velocities into end-effector velocity — has full rank across most postures, but at a singular configuration it loses rank, and the end effector becomes unable to move in some direction no matter how the joints turn. Two common types: a boundary singularity happens when the arm is fully extended, with the end effector at the very edge of the workspace; a wrist singularity happens on a 6-axis arm when the two wrist joint axes line up, effectively losing one degree of freedom. Near a singularity, moving the end effector a short distance in a straight line can demand extremely fast rotation from some joints, and numerical inverse-kinematics solutions also tend to become unstable there. This is why trajectory planning tries to avoid singularities, why numerical IK commonly uses damped least squares to keep joint speeds from exploding, and why manipulability is used to measure how far a posture is from becoming singular.","example":"When a planar two-link arm is fully extended, the end effector can only move perpendicular to the arm — no combination of joint speeds produces any velocity continuing straight outward along the arm's own direction. In that posture, an external force pulling along the arm's direction is carried directly by the structure, with no torque needed from the joints.","related":["Jacobian Matrix","Manipulability","Damped Least Squares","Jacobian Pseudoinverse","Spherical Wrist","Workspace"]},{"id":"manipulability","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（2017）5.4 节：Manipulability","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Wikipedia: Manipulability ellipsoid","url":"https://en.wikipedia.org/wiki/Manipulability_ellipsoid"},{"title":"Buss 逆运动学综述（参考文献列出 Yoshikawa, Manipulability of robotic mechanisms, IJRR 1985）","url":"https://mathweb.ucsd.edu/~sbuss/ResearchWeb/ikmethods/iksurvey.pdf"}],"as_of":"","related_ids":["jacobian-matrix","singular-configuration","kinematic-redundancy","null-space","jacobian-pseudoinverse"],"name":"Manipulability","alt":"可操作度","abbr":"","aliases":["Manipulability Ellipsoid","Yoshikawa Manipulability"],"one_liner":"How freely an arm's end effector can move in each direction from its current pose, and how close it is to a singularity.","explanation":"Manipulability was first given a quantitative definition by T. Yoshikawa in a 1985 IJRR paper. Given the Jacobian matrix J (which maps joint velocity to end-effector velocity), letting the joint velocity range over the unit sphere ‖θ̇‖ = 1 traces out an ellipsoid of corresponding end-effector velocities, called the manipulability ellipsoid: its principal axes point along the eigenvectors of JJᵀ, and the length of each semi-axis is the square root of the corresponding eigenvalue. The longer the ellipsoid is along some direction, the more easily the end effector can move that way; a direction where it's squashed flat signals the arm is approaching a singular pose (an orientation where it loses the ability to move in that direction), and the ellipsoid collapses into a line segment or a plane exactly at a singularity. The commonly used scalar index w = √det(JJᵀ) is proportional to the ellipsoid's volume and equals zero at a singularity. It's used to choose good working poses and to guide structural design, and is also a common optimization objective for null-space motion in redundant arms, steering the arm actively away from singularities.","example":"For a planar two-link arm (link lengths L1, L2), manipulability is w = L1·L2·|sin θ2|, where θ2 is the elbow angle: fully extended (θ2 = 0) gives w = 0, meaning the end effector can no longer extend further along the arm's direction; bent to 90° at the elbow gives the maximum w, where the end effector moves freely in every direction.","related":["Jacobian Matrix","Singular Configuration (Kinematic Singularity)","Kinematic Redundancy","Null Space","Jacobian Pseudoinverse"]},{"id":"jacobian-pseudoinverse","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Buss, Introduction to Inverse Kinematics with Jacobian Transpose, Pseudoinverse and Damped Least Squares methods","url":"https://mathweb.ucsd.edu/~sbuss/ResearchWeb/ikmethods/iksurvey.pdf"},{"title":"Wikipedia: Moore–Penrose inverse","url":"https://en.wikipedia.org/wiki/Moore%E2%80%93Penrose_inverse"},{"title":"Lynch & Park, Modern Robotics（2017）6.2 节：数值逆运动学","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["jacobian-matrix","numerical-inverse-kinematics","damped-least-squares","null-space","singular-configuration","kinematic-redundancy"],"name":"Jacobian Pseudoinverse","alt":"雅可比伪逆","abbr":"","aliases":["Moore-Penrose Pseudoinverse","Generalized Inverse","J⁺","J†"],"one_liner":"The stand-in for a Jacobian inverse when it doesn't exist, giving the solution with the smallest error and least joint motion.","explanation":"The Jacobian matrix J maps joint velocity θ̇ to end-effector velocity ẋ: ẋ = Jθ̇. Solving for joint velocity requires inverting J, but J is often not square (a 7-joint arm's J is 6×7), or it becomes non-invertible near a singular pose (an orientation where the end effector loses the ability to move in some direction). The usual fix is the Moore-Penrose pseudoinverse J⁺ (proposed independently by Moore in 1920, Bjerhammar in 1951, and Penrose in 1955): when an exact solution exists, it gives the one with the least joint motion (smallest norm) among all solutions; when no exact solution exists, it gives the least-squares solution with the smallest error. When J has full rank and there are more joints than task dimensions, J⁺ = Jᵀ(JJᵀ)⁻¹. It's a basic tool of numerical inverse kinematics, but solutions blow up near a singularity, so damped least squares is often used instead in practice.","example":"A 7-DoF arm solves for joint velocity using θ̇ = J⁺ẋ + (I − J⁺J)φ: the first term makes the end effector track the desired velocity ẋ, while the second projects an arbitrary vector φ into the null space, adjusting the elbow and other posture without affecting the end effector — useful for avoiding joint limits, a technique introduced by Liégeois in 1977.","related":["Jacobian Matrix","Numerical Inverse Kinematics","Damped Least Squares","Null Space","Singular Configuration (Kinematic Singularity)","Kinematic Redundancy"]},{"id":"damped-least-squares","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Buss: Introduction to Inverse Kinematics with Jacobian Transpose, Pseudoinverse and Damped Least Squares methods","url":"https://mathweb.ucsd.edu/~sbuss/ResearchWeb/ikmethods/iksurvey.pdf"},{"title":"Wikipedia: Levenberg–Marquardt algorithm","url":"https://en.wikipedia.org/wiki/Levenberg%E2%80%93Marquardt_algorithm"}],"as_of":"","related_ids":["inverse-kinematics","numerical-inverse-kinematics","jacobian-matrix","jacobian-pseudoinverse","singular-configuration","jacobian-transpose-method"],"name":"Damped Least Squares","alt":"阻尼最小二乘法","abbr":"DLS","aliases":["DLS","Levenberg-Marquardt Inverse","Singularity-Robust Inverse"],"one_liner":"Adds a damping term to the Jacobian inverse in numerical IK so the arm doesn't move wildly near singular poses.","explanation":"Damped least squares is an iterative method for solving inverse kinematics, also called the Levenberg-Marquardt method; Wampler, and separately Nakamura and Hanafusa, applied it to inverse kinematics in 1986. Each step computes the joint increment Δθ = Jᵀ(JJᵀ + λ²I)⁻¹e, where J is the Jacobian matrix (mapping joint velocity to end-effector velocity), e is the error between the end effector's current pose and the target, λ is a damping coefficient, and I is the identity matrix. This is equivalent to minimizing ‖JΔθ − e‖² + λ²‖Δθ‖² — reducing the error while also keeping the joint step from being too large. With the plain pseudoinverse, approaching a singular pose (where the end effector loses the ability to move in some direction) produces enormous joint velocities; adding λ keeps the denominator from going to zero, so the motion stays smooth. The tradeoff is that too large a λ slows convergence, so λ is often adjusted dynamically based on how close the arm is to a singularity.","example":"An arm fully extended is near a singularity, where the plain pseudoinverse might demand that some joint instantly spin through many revolutions; setting λ to around 0.05 makes convergence along that direction slower but keeps joint velocities within a normal range.","related":["Inverse Kinematics (IK)","Numerical Inverse Kinematics","Jacobian Matrix","Jacobian Pseudoinverse","Singular Configuration (Kinematic Singularity)","Jacobian Transpose Method"]},{"id":"jacobian-transpose-method","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Buss, Introduction to Inverse Kinematics with Jacobian Transpose, Pseudoinverse and Damped Least Squares methods","url":"https://mathweb.ucsd.edu/~sbuss/ResearchWeb/ikmethods/iksurvey.pdf"},{"title":"Lynch & Park, Modern Robotics（2017）5.2 节：开链静力学 τ = JᵀF","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["jacobian-matrix","numerical-inverse-kinematics","jacobian-pseudoinverse","damped-least-squares","cartesian-impedance-control","principle-of-virtual-work"],"name":"Jacobian Transpose Method","alt":"雅可比转置法","abbr":"","aliases":["Jᵀ Method","Transpose Jacobian Method"],"one_liner":"Uses the Jacobian's transpose instead of its inverse to iteratively solve IK — cheap to compute but slower to converge.","explanation":"The Jacobian transpose method is a numerical approach to inverse kinematics, applied to IK in 1984 independently by Balestrino and colleagues and by Wolovich and Elliott. Each step sets Δθ = αJᵀe, where e is the position error between the end effector and the target, Jᵀ is the transpose of the Jacobian matrix, and α is a small step size. The reasoning comes from the statics relation τ = JᵀF: imagine a virtual spring pulling the end effector toward the target, producing a force F; converted into joint torques, that force is JᵀF, and moving the joints along it is guaranteed to reduce the error whenever the step is small enough, since ⟨JJᵀe, e⟩ = ‖Jᵀe‖² ≥ 0. It never requires a matrix inverse, so each step is very cheap and it never blows up numerically near a singularity — at the cost of slower, sometimes oscillatory, convergence. The same Jᵀ map is also the core of Cartesian impedance control: a virtual spring force at the end effector is converted to joint torques through Jᵀ.","example":"Making an animated character's or robot arm's hand reach for a target point: each iteration computes Δθ = αJᵀe to update the joint angles, and after a few dozen steps the hand gradually approaches the target. In Buss's comparison tests, the method worked adequately for a single end effector but performed noticeably worse than damped least squares on a Y-shaped branching structure with multiple end effectors.","related":["Jacobian Matrix","Numerical Inverse Kinematics","Jacobian Pseudoinverse","Damped Least Squares","Cartesian Impedance Control","Principle of Virtual Work"]},{"id":"kinematic-redundancy","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park) preprint PDF, Example 4.7 与 Ch. 5–6","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["null-space","jacobian-pseudoinverse","7-dof-robot-arm","swivel-angle","null-space-control","inverse-kinematics"],"name":"Kinematic Redundancy","alt":"运动学冗余","abbr":"","aliases":["Redundancy Resolution","Redundant Manipulator"],"one_liner":"Having more joints than a task strictly needs, so the same end-effector pose can be reached by infinitely many joint configurations.","explanation":"If a robot has more joints, n, than the dimensionality of the task space, m, it's said to be kinematically redundant for that task, and the extra n−m degrees of freedom are called the redundancy. A full end-effector pose is 6-dimensional, so a 7-axis arm has 1 redundant degree of freedom: with the end effector held fixed, the elbow can still move through a range of positions. This kind of internal motion that doesn't affect the end effector is called self-motion, and corresponds to the null space of the Jacobian matrix. Redundancy means a single target has infinitely many joint-angle solutions, and redundancy resolution is choosing one of them by some criterion: the most common approach, the Jacobian pseudoinverse, gives the solution with the smallest sum of squared joint velocities, and secondary objectives — avoiding obstacles, staying away from joint limits or singular configurations — can be layered into the null space on top of that. The cost is that inverse kinematics no longer has a unique answer and needs extra optimization. Redundancy is relative to the task: even a 6-axis arm counts as redundant if all that matters is a 3-dimensional end-effector position.","example":"Keeping your palm flat on a table without moving it, you can still raise or lower your elbow — that's the redundant degree of freedom in a human arm. A 7-axis robot arm can likewise adjust its elbow position while the end effector stays put, to steer around a nearby obstacle.","related":["Null Space","Jacobian Pseudoinverse","7-DoF Robot Arm","Swivel Angle","Null-Space Control","Inverse Kinematics (IK)"]},{"id":"null-space","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Kernel (linear algebra)","url":"https://en.wikipedia.org/wiki/Kernel_(linear_algebra)"},{"title":"StudyWolf: Robot control part 5 – Controlling in the null space","url":"https://studywolf.wordpress.com/2013/09/17/robot-control-5-controlling-in-the-null-space/"}],"as_of":"","related_ids":["kinematic-redundancy","null-space-control","jacobian-pseudoinverse","task-prioritization","swivel-angle","jacobian-matrix"],"name":"Null Space","alt":"零空间","abbr":"","aliases":["Kernel","Jacobian Null Space","Self-Motion"],"one_liner":"Every input that a matrix maps to zero; for a redundant arm, the joint motions that don't change the end-effector pose at all.","explanation":"Null space is originally a linear-algebra concept: the null space of a matrix A is every vector x satisfying Ax = 0. In robotics, it usually refers to the null space of the Jacobian matrix J: since J converts joint velocities into end-effector velocity, any joint velocity that falls in its null space produces no motion or rotation of the end effector at all. A kinematically redundant arm — one with more joints than the task strictly needs, like a 7-axis arm performing a 6-dimensional pose task — always has this kind of motion available, called self-motion; a non-redundant arm only gets it at a singular configuration. Null-space control exploits exactly this: the primary task keeps the end effector tracking its target, while secondary objectives — staying away from joint limits, avoiding obstacles, steering clear of singularities — get executed by projecting them into the null space with the projection matrix (I − J⁺J), where J⁺ is the Jacobian pseudoinverse, so they never interfere with the primary task. Task-priority schemes in humanoid whole-body control are also built from layered null-space projections like this.","example":"Holding a 7-axis arm's end effector fixed directly above a cup, the elbow can still trace a circle around the shoulder-wrist line (changing the arm angle) — a one-dimensional self-motion, which can be used to steer the elbow away from a nearby obstacle.","related":["Kinematic Redundancy","Null-Space Control","Jacobian Pseudoinverse","Task Prioritization","Swivel Angle","Jacobian Matrix"]},{"id":"swivel-angle","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Kreutz-Delgado, Long, Seraji: Kinematic Analysis of 7-DOF Manipulators (IJRR, 1992)","url":"https://doi.org/10.1177/027836499201100504"},{"title":"Shimizu et al.: Analytical Inverse Kinematic Computation for 7-DOF Redundant Manipulators With Joint Limits (IEEE T-RO, 2008)","url":"https://doi.org/10.1109/tro.2008.2003266"},{"title":"Kim & Rosen: Redundancy Resolution of the Human Arm and an Upper Limb Exoskeleton (IEEE TBME, 2012)","url":"https://doi.org/10.1109/tbme.2012.2194489"}],"as_of":"","related_ids":["7-dof-robot-arm","kinematic-redundancy","null-space","analytical-inverse-kinematics","inverse-kinematics","motion-retargeting"],"name":"Swivel Angle","alt":"臂型角（肘部自运动角）","abbr":"","aliases":["Arm Angle","Elbow Swivel Angle","Arm Plane Angle"],"one_liner":"For a 7-DoF arm holding a fixed hand pose, the angle the elbow can still rotate through around the shoulder-wrist line.","explanation":"The swivel angle is a scalar parameter describing the redundancy of a 7-DoF arm. Kreutz-Delgado and colleagues parameterized this redundancy in a 1992 IJRR paper using the angle between the 'arm plane' — defined by the shoulder, elbow, and wrist — and a reference plane. Reaching a given end-effector pose only requires 6 degrees of freedom, so the extra one shows up as self-motion: with the hand held fixed, the elbow can still trace out a circle around the shoulder-wrist line, a joint motion that leaves the end effector unmoved. Fixing the swivel angle turns inverse kinematics into a problem with a unique closed-form solution; building on this, Shimizu and colleagues gave an analytical 7-DoF inverse-kinematics solution that accounts for joint limits in 2008. The human arm has the same redundancy, and exoskeleton and teleoperation research commonly calls this the swivel angle.","example":"When teleoperating a 7-DoF arm, motion capture first measures the human operator's own swivel angle, which is then fed into the arm's inverse kinematics together with the target hand pose, so the robot's elbow matches the human's posture; this has been demonstrated on a teleoperated KUKA LWR4+ using the human arm's swivel angle.","related":["7-DoF Robot Arm","Kinematic Redundancy","Null Space","Analytical Inverse Kinematics","Inverse Kinematics (IK)","Motion Retargeting"]},{"id":"whole-body-inverse-kinematics","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"mink: Python inverse kinematics based on MuJoCo (GitHub)","url":"https://github.com/kevinzakka/mink"},{"title":"GMR: General Motion Retargeting (GitHub)","url":"https://github.com/YanjieZe/GMR"}],"as_of":"2026-09","related_ids":["inverse-kinematics","motion-retargeting","general-motion-retargeting","whole-body-control","jacobian-pseudoinverse","task-prioritization"],"name":"Whole-Body Inverse Kinematics","alt":"全身逆运动学","abbr":"WBIK","aliases":["WBIK","Whole-Body IK"],"one_liner":"Solves for a robot's entire set of joint angles at once, so the hands, feet, and torso all reach their targets together.","explanation":"Ordinary inverse kinematics usually handles just one arm: given a target end-effector pose, it solves for the joint angles along that one chain. Whole-body inverse kinematics scales this up to the entire robot — a humanoid, an arm-equipped quadruped, a dual-arm mobile base — satisfying several goals at once: where both hands should be, where both feet should land, which way the torso should face, keeping the center of mass over the support region, all while respecting joint limits and avoiding self-collision. With many joints and goals that often conflict, there's generally no closed-form solution; a common approach writes each goal as a 'task,' uses a Jacobian matrix to relate joint velocity to task error, and solves iteratively with weighted least squares or a quadratic program (this is differential inverse kinematics), resolving conflicts through task weights or priorities. It only computes joint angles and doesn't consider force or torque, which is what distinguishes it from whole-body control. In embodied AI it's commonly used to retarget human motion-capture data onto a humanoid, to drive whole-body teleoperation, and to generate reference motions for reinforcement learning.","example":"GMR (General Motion Retargeting) maps human motion onto more than a dozen humanoid robots, including Unitree's G1 and Booster's T1: for every frame, a whole-body IK solver built on mink (a differential-IK library based on MuJoCo) aligns the robot's hands, feet, and pelvis to the corresponding human body parts while respecting joint limits, solving for the joint-angle sequence; the developers report it runs in real time on a CPU.","related":["Inverse Kinematics (IK)","Motion Retargeting","General Motion Retargeting","Whole-Body Control","Jacobian Pseudoinverse","Task Prioritization"]},{"id":"differential-drive-kinematics","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Differential wheeled robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Differential_wheeled_robot"},{"title":"ros2_control: diff_drive_controller 文档","url":"https://control.ros.org/master/doc/ros2_controllers/diff_drive_controller/doc/userdoc.html"}],"as_of":"","related_ids":["nonholonomic-constraint","differential-drive-base","wheel-odometry","cmd-vel-topic","mobile-base","forward-kinematics"],"name":"Differential Drive Kinematics","alt":"差速驱动运动学","abbr":"","aliases":["Unicycle Model","Differential Drive Model"],"one_liner":"The equations relating the left and right wheel speeds of a two-wheeled base to its forward speed and turning rate.","explanation":"Differential drive is the most common mobile-base layout: an independently motor-driven wheel on the left and one on the right, plus caster wheels to keep it from tipping, steering purely by making the two wheels turn at different speeds, with no separate steering mechanism needed. Its kinematics comes down to just two equations: linear velocity v = (v_R + v_L)/2 and angular velocity ω = (v_R − v_L)/b, where v_R and v_L are the right and left wheel's linear speeds (wheel angular speed times wheel radius r) and b is the distance between the wheels. Run in reverse, a desired v and ω give the speed each wheel should turn at. This makes the chassis behave, in the plane, like a point that can only move forward and spin in place — the unicycle model — and it can't move directly sideways; a speed limitation like this is called a nonholonomic constraint. ROS's diff_drive_controller uses exactly these two equations to convert a cmd_vel command into left/right wheel commands, and to reconstruct odometry from wheel feedback.","example":"For a robot vacuum to turn left in place, the right wheel drives forward at speed u and the left wheel drives backward at the same speed u: v = 0, ω = 2u/b, so the body spins in place about the midpoint between the two wheels.","related":["Nonholonomic Constraint","Differential Drive Base","Wheel Odometry","cmd_vel Topic (geometry_msgs/Twist velocity command)","Mobile Base (Chassis)","Forward Kinematics (FK)"]},{"id":"nonholonomic-constraint","category":"mechanics","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Nonholonomic system","url":"https://en.wikipedia.org/wiki/Nonholonomic_system"},{"title":"Modern Robotics (Lynch & Park), Sec. 2.4 Holonomic and nonholonomic constraints","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["differential-drive-kinematics","configuration-space","generalized-coordinates","hybrid-a-star","mecanum-wheel","kinodynamic-planning"],"name":"Nonholonomic Constraint","alt":"非完整约束","abbr":"","aliases":["Nonholonomic System","Nonintegrable Constraint"],"one_liner":"A constraint on velocity direction only, which can't be integrated into a position constraint — like a wheel that can't slide sideways.","explanation":"A nonholonomic constraint restricts a system's velocity in a way that cannot be integrated into an equation involving position alone; the term 'holonomic' was introduced by Hertz in 1894. The most common source is a wheel rolling without slipping: a differential-drive cart at pose (x, y, θ) must satisfy ẋ·sinθ − ẏ·cosθ = 0, where x, y are position and θ is heading, meaning it can never move sideways at any instant. This narrows the velocity directions available at any moment, but it doesn't shrink the set of poses the vehicle can eventually reach — it can still park in any spot, just by maneuvering back and forth the way parallel parking does. So planning for wheeled bases can't simply interpolate straight lines between poses; it needs methods like hybrid A* that account for steering constraints. Escaping the constraint means switching to an omnidirectional base such as one with Mecanum wheels. Finger contacts that roll across an object's surface during grasping are this same type of constraint.","example":"A vacuum-cleaning robot that wants to shift 20 cm directly to its left can't simply slide sideways — it has to either rotate 90° in place and then drive forward, or maneuver back and forth like parallel parking. That's the nonholonomic constraint at work.","related":["Differential Drive Kinematics","Configuration Space (C-Space)","Generalized Coordinates","Hybrid A*","Mecanum Wheel","Kinodynamic Planning"]},{"id":"jerk","category":"mechanics","sec":4,"tier":2,"sources":[{"title":"Jerk (physics) - Wikipedia","url":"https://en.wikipedia.org/wiki/Jerk_(physics)"},{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 9 Trajectory Generation","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["minimum-jerk-trajectory","s-curve-velocity-profile","quintic-polynomial-interpolation","trajectory-planning","action-smoothing","trapezoidal-velocity-profile"],"name":"Jerk","alt":"加加速度","abbr":"","aliases":["Jolt"],"one_liner":"The rate of change of acceleration over time — the third time-derivative of position — measured in m/s³.","explanation":"Jerk, j, is the time derivative of acceleration a: j = da/dt = d³x/dt³ (x is position), measured in m/s³, and it captures how abruptly acceleration itself is changing. A sudden change in acceleration means large jerk, which feels jarring to a person, and also excites mechanical vibration and accelerates wear. That's why robot trajectory planning commonly bounds jerk: cubic polynomial time-scaling has acceleration jump discontinuously at the start and end, which is equivalent to infinite jerk; an S-curve velocity profile splits motion into seven segments to keep jerk bounded throughout; and a minimum-jerk trajectory minimizes the integral of squared jerk over the whole motion. Tamar Flash and Neville Hogan found in 1985 that natural human arm movements approximately follow this minimum-jerk principle. Reinforcement-learning locomotion controllers use a similar idea, commonly penalizing the difference between consecutive actions in the reward function to keep motion smooth.","example":"If an elevator's acceleration jumped instantly from 0 to its set value at startup, passengers would feel a hard jolt; real elevators ramp acceleration up gradually over time, which is exactly bounding jerk.","related":["Minimum-Jerk Trajectory","S-Curve Velocity Profile","Quintic Polynomial Interpolation","Trajectory Planning","Action Smoothing","Trapezoidal Velocity Profile"]},{"id":"torque","category":"mechanics","sec":5,"tier":1,"sources":[{"title":"Wikipedia: Torque","url":"https://en.wikipedia.org/wiki/Torque"},{"title":"Modern Robotics (Lynch & Park, 2017)","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["wrench","torque-control","joint-torque-sensor","peak-torque","gravity-compensation","moment-of-inertia"],"name":"Torque","alt":"力矩","abbr":"","aliases":["Moment of Force"],"one_liner":"The turning effect of a force on an object, equal to force times lever arm, measured in newton-meters.","explanation":"Torque describes the effect a force has in rotating an object about a point or axis; it's also called the moment of force. It's defined as τ = r × F, where r is the vector from the rotation axis to the point where the force is applied and F is the force; its magnitude is τ = rF·sinθ, where θ is the angle between r and F, and r·sinθ is the lever arm — the perpendicular distance from the axis to the force's line of action — measured in newton-meters (N·m). The longer the lever arm, the more torque the same force produces. The rotational counterpart of Newton's second law is τ = Iα, where I is the moment of inertia and α is angular acceleration. In a robot, every rotary joint's motor outputs torque around its own axis: the torque needed to overcome gravity, inertia, and contact forces at each joint drives the choice of motor and reducer, joint-module datasheets list rated and peak torque, and torque control means sending a torque command directly to each joint.","example":"An arm extended horizontally, holding a 1 kg object 0.5 m from the shoulder joint, requires the shoulder to produce roughly 1 × 9.8 × 0.5 ≈ 4.9 N·m of torque from that object alone, not counting the arm's own weight.","related":["Wrench","Torque Control","Joint Torque Sensor","Peak Torque","Gravity Compensation","Moment of Inertia"]},{"id":"payload","category":"mechanics","sec":5,"tier":1,"sources":[{"title":"Wikipedia: Industrial robot（Technical description）","url":"https://en.wikipedia.org/wiki/Industrial_robot"},{"title":"Franka Robotics: Franka Research 3","url":"https://franka.de/franka-research-3"}],"as_of":"2026-09","related_ids":["payload-to-weight-ratio","end-effector","torque","peak-torque","workspace","robotic-arm"],"name":"Payload","alt":"负载","abbr":"","aliases":["Rated Payload","Payload Capacity"],"one_liner":"The maximum mass a robot's end effector can reliably carry while working, usually given in kilograms.","explanation":"Payload is a core spec on a robot's datasheet: the maximum mass an arm's end effector can reliably carry at normal working speed and acceleration, usually given in kilograms. Two details are easy to miss. First, any tool mounted at the tip — gripper, camera, force sensor — counts against this figure too, so it eats into the margin left for the object actually being grasped. Second, manufacturer numbers usually assume the payload's center of mass sits close to the end flange (the mounting face at the tip of the arm); the farther off-center it sits, the more torque it produces, and the less mass the arm can actually handle. Dividing payload by the robot's own weight gives the payload-to-weight ratio, a common way to compare how lightweight a design is. Humanoid robots often quote separate numbers for single-arm payload and whole-body carrying capacity.","example":"Franka Research 3's official specs list a 3 kg payload and an 855 mm reach; once a gripper is mounted at the tip, the gripper's own weight has to be subtracted from that 3 kg before figuring out how much the robot can actually pick up.","related":["Payload-to-Weight Ratio","End Effector","Torque","Peak Torque","Workspace","Robotic Arm"]},{"id":"wrench","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Modern Robotics 3.4: Wrenches","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/3-4-wrenches/"},{"title":"ROS 2 geometry_msgs/Wrench.msg","url":"https://raw.githubusercontent.com/ros2/common_interfaces/rolling/geometry_msgs/msg/Wrench.msg"},{"title":"Modern Robotics: Mechanics, Planning, and Control (Lynch & Park, free preprint)","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["twist","six-axis-force-torque-sensor","statics","torque","grasp-matrix","adjoint-representation"],"name":"Wrench","alt":"力旋量","abbr":"","aliases":["Force-Torque","Six-Dimensional Force/Torque"],"one_liner":"Packs a 3D force and a 3D torque into one 6-dimensional vector describing the full load acting on a rigid body.","explanation":"A wrench is the 6-dimensional vector F = (m, f): f is the 3D force and m is the 3D torque, both expressed in the same frame. It's the dual of the twist — their dot product VᵀF equals power, and because power doesn't depend on the choice of reference frame, the rule for transforming a wrench between frames follows directly from that fact. Robotics uses it anywhere the question is 'how much force is being applied': a wrist-mounted six-axis force-torque sensor reads out a wrench directly; the statics equation τ = JᵀF converts an end-effector wrench into joint torques; and grasp analysis sums the contact forces at each fingertip into the total wrench acting on the object. Component ordering isn't standardized across sources: Lynch and Park's textbook writes it as (torque, force), while ROS's geometry_msgs/Wrench and Franka's libfranka both list force before torque.","example":"An example from Modern Robotics (taking g = 10 m/s²): a 0.5 kg robot hand holds a 0.1 kg apple, and the wrist's six-axis force-torque sensor reads not just the 6 N of weight but also 0.75 N·m of torque, because the centers of mass of both the hand and the apple sit offset from the sensor.","related":["Twist","Six-Axis Force/Torque Sensor","Statics","Torque","Grasp Matrix","Adjoint Representation"]},{"id":"adjoint-representation","category":"mechanics","sec":5,"tier":3,"sources":[{"title":"Modern Robotics（Lynch & Park）预印本，3.3.2 节 Definition 3.20 与 Proposition 3.27","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["screw-theory","twist","wrench","homogeneous-transformation-matrix","lie-group","product-of-exponentials-formula"],"name":"Adjoint Representation","alt":"伴随变换","abbr":"Ad","aliases":["Ad","Adjoint Map","Adjoint Transformation"],"one_liner":"The 6×6 matrix that transforms a twist or a wrench from one coordinate frame into another.","explanation":"The adjoint representation is a basic tool in screw theory, and Lynch and Park's textbook Modern Robotics builds it into the core of its chapter on rigid-body motion. Given a pose T = (R, p) (R is the rotation matrix, p is the translation vector), the adjoint representation [Ad_T] is a 6×6 matrix: the upper-left and lower-right blocks are both R, the lower-left block is [p]R (where [p] is the skew-symmetric matrix of p), and the upper-right block is zero. Its job is to change coordinate frames: the same twist (angular velocity plus linear velocity), expressed in two frames {a} and {b}, satisfies V_a = [Ad_Tab]·V_b; a wrench (torque plus force) transforms using its transpose instead. It shows up constantly in the product of exponentials formula, in deriving Jacobians, and in recursive dynamics algorithms.","example":"A wrist-mounted six-axis force-torque sensor reads a wrench F_s expressed in the sensor's own frame; to get the force in the tool center point frame instead, convert it in one step with F_tcp = [Ad_T]ᵀ·F_s, where T is the pose of the TCP frame relative to the sensor frame.","related":["Screw Theory","Twist","Wrench","Homogeneous Transformation Matrix","Lie Group","Product of Exponentials Formula"]},{"id":"statics","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Statics - Wikipedia","url":"https://en.wikipedia.org/wiki/Statics"},{"title":"Modern Robotics 5.2: Statics of Open Chains","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/5-2-statics-of-open-chains/"}],"as_of":"","related_ids":["wrench","jacobian-matrix","gravity-compensation","principle-of-virtual-work","quasi-static-assumption","dynamics"],"name":"Statics","alt":"静力学","abbr":"","aliases":["Robot Statics","Static Force Analysis"],"one_liner":"The branch of mechanics that studies how forces and torques balance on an object at rest or moving at constant velocity.","explanation":"Statics is the branch of classical mechanics that studies systems with no acceleration, where the equilibrium condition is that the net force and net torque both equal zero: ΣF = 0, ΣM = 0. In robotics, the most-used result is τ = Jᵀ(θ)F: τ is the vector of joint torques, J is the Jacobian matrix (which maps joint velocities to end-effector velocity), θ is the current joint configuration, and F is the wrench — a combined 3D force and 3D torque — that the end effector applies to the outside world. The formula follows from conservation of power (virtual work), and it answers the question, 'to produce this much force at the tip, how much torque should each motor supply?' When the robot must also hold up its own weight, a gravity-compensation torque is added on top. Force control, gravity compensation, and grasp-force analysis all build on statics; once motion speeds up enough that inertial forces can't be ignored, the analysis has to switch to dynamics instead.","example":"A robot arm's end effector rests motionless against a tabletop and must press down with 10 N of force: first compute the Jacobian J for the current pose, then use τ = Jᵀ F to find the torque each joint must output, and finally add the gravity-compensation torque that offsets the arm's own weight.","related":["Wrench","Jacobian Matrix","Gravity Compensation","Principle of Virtual Work","Quasi-Static Assumption","Dynamics"]},{"id":"principle-of-virtual-work","category":"mechanics","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Virtual work","url":"https://en.wikipedia.org/wiki/Virtual_work"},{"title":"Modern Robotics (Lynch & Park), Sec. 5.2 Statics of open chains","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["jacobian-matrix","wrench","statics","gravity-compensation","contact-jacobian","euler-lagrange-equations"],"name":"Principle of Virtual Work","alt":"虚功原理","abbr":"","aliases":["Principle of Virtual Displacements","Virtual Velocity Principle"],"one_liner":"A system is in static equilibrium exactly when the total work done by active forces over any allowed tiny virtual displacement is zero.","explanation":"The principle of virtual work is a fundamental principle of analytical mechanics: imagine the system undergoes a hypothetical, infinitesimal displacement (a virtual displacement) that's consistent with its constraints — if the total work done by all the active forces over this displacement (the virtual work) sums to zero, the system is in static equilibrium. Johann Bernoulli gave it a systematic statement in correspondence with Varignon in 1715, and Lagrange later built analytical mechanics on top of it. Its advantage is that the reaction forces from ideal constraints — such as the internal forces at a joint — do zero virtual work, so they never need to be solved for individually. The most-used result in robotics is τ = Jᵀ(θ)F: when the end effector applies a force or wrench F to the outside world, the required joint torque τ equals the transpose of the Jacobian matrix J times F, based on the equality of virtual work (or power) done at the joints and at the end effector. Force control, impedance control, gravity compensation, and converting a legged robot's foot contact force f into joint torques via τ = J_cᵀf all rest on this relationship.","example":"A planar two-link arm needs to push horizontally against a wall with 10 N at its end effector, in some given pose: substituting F = (10, 0) into τ = JᵀF directly gives the torque each of the two joints must output, with no need to analyze the internal forces within the links.","related":["Jacobian Matrix","Wrench","Statics","Gravity Compensation","Contact Jacobian","Euler-Lagrange Equations"]},{"id":"quasi-static-assumption","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Modern Robotics: Mechanics, Planning, and Control (Lynch & Park), Ch.12 Grasping and Manipulation","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Robotic Manipulation (Russ Tedrake, MIT): Force Control","url":"https://manipulation.csail.mit.edu/force.html"}],"as_of":"","related_ids":["statics","rigid-body-dynamics","coulomb-friction","friction-cone","non-prehensile-manipulation","slip"],"name":"Quasi-Static Assumption","alt":"准静态假设","abbr":"","aliases":["Quasistatic Approximation"],"one_liner":"Assuming motion is slow enough that inertial forces can be ignored, so every instant can be analyzed as a force balance.","explanation":"The quasi-static assumption is a common simplification in robot manipulation analysis: velocities and accelerations are assumed small enough that inertial forces (the mass-times-acceleration term) can be neglected, so at every instant the external and contact forces are in balance, and the full dynamics equation collapses into static equilibrium — every force and torque sums to zero. Lynch and Park's textbook Modern Robotics uses it throughout its chapter on grasping and manipulation to analyze tasks like supporting, gripping, and pushing. Its benefit is that dynamics never needs to be integrated over time: it's enough to check, at each contact, whether it's sticking or slipping, and whether the contact force stays inside the friction cone (the range of directions friction allows). Slow pushing, peg-in-hole insertion, and grasp-stability analysis are mostly built on this assumption. The cost is that it breaks down once motion speeds up: flinging, throwing, and fast flipping — dynamic manipulation tasks — require going back to full rigid-body dynamics.","example":"The “ruler trick”: balance a long ruler horizontally on two index fingers, then slowly bring the fingers together. The ruler slides on one finger, then the other, then both, but its center of mass always stays between the two fingers, so it never falls — quasi-static force balance predicts exactly which finger slips and when. Move the fingers too fast, and the analysis breaks down, and the ruler falls.","related":["Statics","Rigid-Body Dynamics","Coulomb Friction","Friction Cone","Non-prehensile Manipulation","Slip"]},{"id":"center-of-mass","category":"mechanics","sec":5,"tier":1,"sources":[{"title":"Wikipedia: Center of mass","url":"https://en.wikipedia.org/wiki/Center_of_mass"},{"title":"Wikipedia: Support polygon","url":"https://en.wikipedia.org/wiki/Support_polygon"}],"as_of":"","related_ids":["support-polygon","zero-moment-point","centroidal-dynamics","static-stability","inertial-parameters","payload"],"name":"Center of Mass (CoM)","alt":"质心","abbr":"CoM","aliases":["CoM","Center of Gravity (CoG)"],"one_liner":"The mass-weighted average position of an object, which behaves as if all its mass were concentrated there.","explanation":"The center of mass is a mass-weighted average of position: r_c = Σmᵢrᵢ / Σmᵢ, where mᵢ is the mass of the i-th part and rᵢ is its position. In a uniform gravity field, the center of mass coincides with the center of gravity, so engineers often use the two terms interchangeably. For a robot, the center of mass governs balance: while standing still, its projection onto the ground has to land inside the support polygon — the region enclosed by the robot's ground-contact points — or the robot falls over. Gait control and zero-moment-point analysis for humanoids and quadrupeds are built around the center-of-mass trajectory. When an arm picks up a heavy object, the farther the payload's center of mass sits from the flange, the more torque the joints have to resist. Model files such as URDF require a mass and center-of-mass position for every link; getting these wrong makes simulation behave differently from the real robot.","example":"Two balls, 1 kg and 3 kg, sit at x = 0 and x = 4 meters. The center of mass is at (1×0 + 3×4) / (1+3) = 3 meters — closer to the heavier ball.","related":["Support Polygon","Zero Moment Point","Centroidal Dynamics","Static Stability","Inertial Parameters","Payload"]},{"id":"moment-of-inertia","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Moment of inertia","url":"https://en.wikipedia.org/wiki/Moment_of_inertia"},{"title":"Modern Robotics 8.2: Dynamics of a Single Rigid Body (Part 1 of 2)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-2-dynamics-of-a-single-rigid-body-part-1-of-2/"}],"as_of":"","related_ids":["inertia-tensor","torque","parallel-axis-theorem","inertial-parameters","rigid-body-dynamics","reflected-inertia"],"name":"Moment of Inertia","alt":"转动惯量","abbr":"","aliases":["Rotational Inertia","Mass Moment of Inertia"],"one_liner":"How hard an object is to spin up or slow down about a given axis, I = Σmr².","explanation":"Moment of inertia describes an object's resistance to rotating about a given axis, defined as the sum, over every part of the object, of its mass times the square of its distance from that axis: I = Σmr², in kg·m²; Euler introduced this concept in 1765. It plays the same role in rotation that mass plays in straight-line motion: τ = Iα, where τ is torque and α is angular acceleration. For the same mass, the farther it's distributed from the axis, the larger the moment of inertia: a uniform rod has moment of inertia ml²/12 about its midpoint, but ml²/3 about one end. A 3D rigid body needs the full 3×3 inertia tensor to describe its moment of inertia about every possible axis. Robot designers commonly push motors toward the torso and use lightweight materials for shins and fingers precisely to reduce the moment of inertia of the far-out links, so limbs can swing faster with less effort. Note this is different from the “second moment of area” used in structural mechanics, despite the similar name.","example":"Many bipedal and quadruped robots mount the knee motor up near the top of the thigh and drive the knee joint through a linkage or belt instead, keeping the shin light and reducing its moment of inertia during a leg swing.","related":["Inertia Tensor","Torque","Parallel Axis Theorem","Inertial Parameters","Rigid-Body Dynamics","Reflected Inertia"]},{"id":"inertia-tensor","category":"mechanics","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Moment of inertia（Inertia tensor）","url":"https://en.wikipedia.org/wiki/Moment_of_inertia"},{"title":"Lynch & Park, Modern Robotics（预印本 PDF，8.2 节 Dynamics of a Single Rigid Body；4.2 节 URDF）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["moment-of-inertia","inertial-parameters","parallel-axis-theorem","center-of-mass","unified-robot-description-format","dynamic-parameter-identification"],"name":"Inertia Tensor","alt":"惯性张量","abbr":"","aliases":["Inertia Matrix","Rotational Inertia Matrix","3×3 Inertia Matrix"],"one_liner":"A 3×3 symmetric matrix describing how hard a rigid body is to rotate about any given axis.","explanation":"The inertia tensor (also called the inertia matrix) is the complete description of a rigid body's rotational inertia, given as a 3×3 symmetric positive-definite matrix. The diagonal entries Ixx, Iyy, Izz are the moments of inertia about the x, y, and z axes — for example Ixx = Σm(y² + z²) — while the off-diagonal entries are called products of inertia, such as Ixy = −Σm·x·y (sign convention varies by textbook), and reflect whether rotating about one axis tends to drag the body around another. It relates angular velocity ω to angular momentum through L = Iω, and rotational kinetic energy is K = ½ωᵀIω. Rotating the reference frame gives I' = RᵀIR, and changing the reference point uses the parallel axis theorem; eigendecomposition yields the principal axes, along which the products of inertia all vanish. URDF's inertial tag requires filling in each link's mass, center-of-mass location, and the six independent entries of the inertia tensor — getting these wrong will throw off both dynamics simulation and torque control.","example":"A 2 kg point mass located at (0.1, 0, 0) m has Ixx = 0, Iyy = Izz = 2×0.1² = 0.02 kg·m², and all products of inertia equal to zero — meaning it costs nothing to spin about an x-axis passing through the point, while rotating about the y- or z-axis must overcome 0.02 kg·m² of inertia.","related":["Moment of Inertia","Inertial Parameters","Parallel Axis Theorem","Center of Mass (CoM)","Unified Robot Description Format","Dynamic Parameter Identification"]},{"id":"parallel-axis-theorem","category":"mechanics","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Parallel axis theorem","url":"https://en.wikipedia.org/wiki/Parallel_axis_theorem"},{"title":"Modern Robotics (Lynch & Park), Theorem 8.2 Steiner's theorem","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["moment-of-inertia","inertia-tensor","inertial-parameters","center-of-mass","dynamic-parameter-identification","unified-robot-description-format"],"name":"Parallel Axis Theorem","alt":"平行轴定理","abbr":"","aliases":["Huygens–Steiner Theorem","Steiner's Theorem"],"one_liner":"Gives the moment of inertia about any parallel axis from the moment of inertia about the center-of-mass axis: I = I_c + md².","explanation":"The parallel axis theorem, also called the Huygens–Steiner theorem, states that a rigid body's moment of inertia I about some axis equals its moment of inertia I_c about a parallel axis through the center of mass, plus its mass m times the square of the distance d between the two axes: I = I_c + m·d². This shows the moment of inertia is smallest about an axis through the center of mass, and grows the farther the axis sits from it. In 3D there's a matrix version (Modern Robotics calls it the Steiner theorem): I_q = I_b + m(qᵀq·E − q·qᵀ), where I_b is the inertia tensor at the center of mass, q is the new reference point's position relative to the center of mass, and E is the 3×3 identity matrix. It's used constantly in robotics: URDF requires the inertia reference frame's origin to be at the center of mass, and dynamics libraries have to translate it when computing in joint frames; attaching a gripper or a payload to a flange means translating that payload's inertia before merging it in; and dynamic parameter identification also uses it to convert between different reference points.","example":"A uniform thin rod of mass m and length L has a moment of inertia of mL²/12 about its midpoint; about one end (equivalent to a link rotating about a joint), d = L/2, giving I = mL²/12 + m(L/2)² = mL²/3 — four times as large.","related":["Moment of Inertia","Inertia Tensor","Inertial Parameters","Center of Mass (CoM)","Dynamic Parameter Identification","Unified Robot Description Format"]},{"id":"inertial-parameters","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 4 URDF 与 Ch. 8 动力学","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Linear Matrix Inequalities for Physically-Consistent Inertial Parameter Identification (Wensing, Kim, Slotine)","url":"https://arxiv.org/abs/1701.04395"},{"title":"legged_gym: legged_robot_config.py","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["mass-matrix","inertia-tensor","center-of-mass","dynamic-parameter-identification","unified-robot-description-format","dynamics-randomization"],"name":"Inertial Parameters","alt":"惯性参数","abbr":"","aliases":["Mass Properties"],"one_liner":"The set of values describing how heavy a rigid body is, where its center of mass sits, and how hard it is to rotate about each axis.","explanation":"Inertial parameters describe the mass distribution of each rigid link in a robot, commonly given as 10 numbers: mass m (1 value), center-of-mass position c, or equivalently the first mass moment m·c (3 values), and the inertia tensor I (a 3×3 symmetric matrix with 6 independent entries, describing how hard the body is to rotate about each axis). Robot description files like URDF and MJCF store exactly these numbers in each link's inertial field. Kinematics only deals with geometry, but the moment dynamics enters the picture — gravity compensation, inverse dynamics, torque control, physics simulation — inertial parameters become essential. Values exported straight from CAD often diverge from the real hardware, since cables, screws, and other small parts get left out, so dynamic parameter identification is typically needed afterward, along with checks for physical consistency, such as positive mass and a positive-definite inertia tensor. Inaccurate inertial parameters are one source of the sim-to-real gap, and domain randomization commonly perturbs link mass to compensate.","example":"A link's inertial tag in URDF includes mass, origin (center-of-mass position), and the six values ixx, ixy, ixz, iyy, iyz, izz of the inertia tensor. legged_gym provides an added_mass_range parameter that randomly adds or removes mass from the torso during training.","related":["Mass Matrix","Inertia Tensor","Center of Mass (CoM)","Dynamic Parameter Identification","Unified Robot Description Format","Dynamics Randomization"]},{"id":"angular-momentum","category":"mechanics","sec":5,"tier":2,"sources":[{"title":"Angular momentum - Wikipedia","url":"https://en.wikipedia.org/wiki/Angular_momentum"}],"as_of":"","related_ids":["moment-of-inertia","centroidal-dynamics","centroidal-momentum-matrix","torque","rigid-body-dynamics","balance-control"],"name":"Angular Momentum","alt":"角动量","abbr":"","aliases":["Moment of Momentum"],"one_liner":"A measure of an object's rotational “oomph”; for a rigid body spinning about an axis, it equals moment of inertia times angular velocity.","explanation":"Angular momentum is the rotational counterpart of linear momentum, p = mv. For a point mass, angular momentum about some point is L = r × p, where r is the position vector from that point to the mass and × is the cross product; for a rigid body spinning about a fixed axis, L = Iω, where I is the moment of inertia and ω is angular velocity, measured in kg·m²/s. Its rate of change equals the external torque, dL/dt = τ; with no external torque, angular momentum is conserved — this is why a figure skater spins faster after pulling their arms in. Legged and humanoid robots often track whole-body angular momentum about the center of mass: while airborne, gravity produces no torque about the center of mass, so this quantity stays constant, and the robot can only redistribute it among body parts by swinging its arms or tucking its legs; once a foot touches down, contact forces can change it. Planning for balance, jumping, and flips all has to account for it.","example":"For a humanoid robot doing a backflip: it must build up enough angular momentum about its center of mass before its feet leave the ground. Once airborne, total angular momentum stays fixed, so tucking the legs shrinks the moment of inertia to spin faster, and extending them again before landing slows the rotation back down.","related":["Moment of Inertia","Centroidal Dynamics","Centroidal Momentum Matrix","Torque","Rigid-Body Dynamics","Balance Control"]},{"id":"dynamics","category":"mechanics","sec":6,"tier":1,"sources":[{"title":"Wikipedia: Dynamics (mechanics)","url":"https://en.wikipedia.org/wiki/Dynamics_(mechanics)"},{"title":"MIT Underactuated Robotics: Multi-Body Dynamics（manipulator equations）","url":"https://underactuated.mit.edu/multibody.html"},{"title":"MuJoCo Documentation: Computation（forward / inverse dynamics）","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["forward-dynamics","inverse-dynamics","mass-matrix","rigid-body-dynamics","kinematics","physics-engine"],"name":"Dynamics","alt":"动力学","abbr":"","aliases":["Robot Dynamics"],"one_liner":"The study of how forces and torques produce motion — how much acceleration a given torque causes.","explanation":"Dynamics is the branch of classical mechanics that studies the relationship between force and motion, grounded in Newton's second law, F = ma. Robot dynamics extends this to multi-joint systems, commonly written as M(q)q̈ + C(q,q̇)q̇ + g(q) = τ, where q is the vector of joint angles, q̇ and q̈ are joint velocity and acceleration, M is the mass matrix (each link's inertia reflected back to the joints), the C term captures Coriolis and centrifugal forces, g is the gravity term, and τ is joint torque. Computing acceleration from known torques is called forward dynamics, which a physics engine solves at every simulation step; computing the torques needed to produce a desired motion is called inverse dynamics, used for gravity compensation and torque control. Kinematics only deals with geometry; picking up a heavy object, moving quickly, and balancing on legs all require dynamics.","example":"When an arm holds an object perfectly still, q̇ and q̈ are both zero, so the equation reduces to g(q) = τ — inverse dynamics then just gives the torque each joint needs to counteract gravity, including the object's weight.","related":["Forward Dynamics","Inverse Dynamics","Mass Matrix","Rigid-Body Dynamics","Kinematics","Physics Engine"]},{"id":"rigid-body-dynamics","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Rigid body dynamics","url":"https://en.wikipedia.org/wiki/Rigid_body_dynamics"},{"title":"Modern Robotics (Lynch & Park), Ch.8 Dynamics of Open Chains","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo Documentation: Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["newton-euler-equations","euler-lagrange-equations","mass-matrix","forward-dynamics","inverse-dynamics","physics-engine"],"name":"Rigid-Body Dynamics","alt":"刚体动力学","abbr":"","aliases":[],"one_liner":"The study of how forces and torques cause acceleration and rotation in one or more connected rigid bodies.","explanation":"Rigid-body dynamics studies how one or more rigid bodies, connected by joints, move under applied forces and torques. A single rigid body obeys the Newton-Euler equations: F=ma governs translation (F is net force, m is mass, a is the center of mass's acceleration), and τ=Iα+ω×(Iω) governs rotation (τ is net torque, I is the inertia tensor, α is angular acceleration, ω is angular velocity). A multi-body system like a robot arm is usually written M(q)q̈+c(q,q̇)+g(q)=τ: q is the joint angle vector, M is the mass matrix, c is the Coriolis and centrifugal term, g is the gravity term, and τ is joint torque. Given torques, solving for motion is forward dynamics, computed at every step by a physics simulator; given a desired motion, solving for the required torques is inverse dynamics, used for torque control and gravity compensation. Roy Featherstone's book Rigid Body Dynamics Algorithms is the standard reference, and MuJoCo's relevant algorithms are based on it.","example":"When an arm holds a cup of water perfectly still, q̇ and q̈ are both zero, so the equation reduces to g(q)=τ — the result is exactly the torque each joint needs to output to counteract gravity, which is gravity compensation.","related":["Newton-Euler Equations","Euler-Lagrange Equations","Mass Matrix","Forward Dynamics","Inverse Dynamics","Physics Engine"]},{"id":"multibody-dynamics","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Multibody system","url":"https://en.wikipedia.org/wiki/Multibody_system"},{"title":"Lynch & Park, Modern Robotics（2017）第 8 章：机器人动力学是多体系统动力学的子领域","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo 文档 Overview：广义坐标表示","url":"https://mujoco.readthedocs.io/en/stable/overview.html"}],"as_of":"","related_ids":["rigid-body-dynamics","generalized-coordinates","recursive-newton-euler-algorithm","articulated-body-algorithm","physics-engine","kinematic-tree"],"name":"Multibody Dynamics","alt":"多体动力学","abbr":"","aliases":["Multibody System Dynamics"],"one_liner":"The study of how systems of rigid or flexible bodies, connected by joints, move under applied forces.","explanation":"Multibody dynamics studies how systems of rigid or flexible bodies, connected by joints and possibly undergoing large translations and rotations, move under applied forces; it's used in aerospace, automotive engineering, biomechanics, and robotics. Each rigid body in space has 6 degrees of freedom (3 translational, 3 rotational), and joints remove some of them. Modeling follows one of two approaches: redundant coordinates, where every body is described by its full pose and constraint equations then tie them together, or minimal (generalized) coordinates, which use independent variables like joint angles directly, so the constraints are automatically satisfied. Robot dynamics is a branch of it, producing equations of the form M(q)q̈ + C(q,q̇)q̇ + g(q) = τ (q is the joint angles, M the mass matrix, C the Coriolis and centrifugal terms, g the gravity term, τ the joint torques). Both MuJoCo and Pinocchio compute dynamics using generalized coordinates.","example":"A quadruped robot with 12 joints has 13 rigid bodies (the torso plus 12 leg segments). Generalized coordinates need only 6 (the torso's floating base) + 12 (joint angles) = 18 variables; redundant coordinates would need 13×6 = 78 variables, with 12 revolute joints each contributing 5 constraint equations — 60 in total — to eliminate the extra degrees of freedom.","related":["Rigid-Body Dynamics","Generalized Coordinates","Recursive Newton-Euler Algorithm","Articulated Body Algorithm","Physics Engine","Kinematic Tree"]},{"id":"forward-dynamics","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"MuJoCo Documentation - Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Inverse dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Inverse_dynamics"},{"title":"Featherstone's algorithm - Wikipedia","url":"https://en.wikipedia.org/wiki/Featherstone%27s_algorithm"}],"as_of":"","related_ids":["inverse-dynamics","euler-lagrange-equations","mass-matrix","articulated-body-algorithm","physics-engine","forward-dynamics-model"],"name":"Forward Dynamics","alt":"正动力学","abbr":"","aliases":[],"one_liner":"Computing a robot's current acceleration, and hence its next motion, from its current state and the torques applied to it.","explanation":"Dynamics studies the relationship between force and motion, and splits into forward and inverse directions. Forward dynamics starts from known joint positions q, velocities q̇, and applied torques τ, and solves for acceleration q̈: from M(q)q̈ + c(q,q̇) = τ, q̈ = M⁻¹(τ − c), where M is the mass matrix and c lumps together the Coriolis, centrifugal, and gravity terms. Integrating q̈ forward over time steps produces the robot's future trajectory — exactly what a physics simulator does at every step, with engines like MuJoCo and Isaac Sim also solving for contact forces and friction on top of this. Inverse dynamics runs the opposite direction — given a desired motion, find the torques needed to produce it — and is mostly used in control. The classic efficient algorithm for forward dynamics is Roy Featherstone's articulated-body algorithm. Note that this differs from a “forward dynamics model” in machine learning, which refers to a neural network trained to predict the next state.","example":"In MuJoCo, if you apply zero torque to every joint of an arm and repeatedly call mj_step, the simulator computes the acceleration due to gravity via forward dynamics, and the arm sags and swings naturally under its own weight.","related":["Inverse Dynamics","Euler-Lagrange Equations","Mass Matrix","Articulated Body Algorithm","Physics Engine","Forward Dynamics Model"]},{"id":"inverse-dynamics","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park) preprint PDF, Ch. 8 Dynamics of Open Chains","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Inverse dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Inverse_dynamics"}],"as_of":"","related_ids":["forward-dynamics","recursive-newton-euler-algorithm","mass-matrix","computed-torque-control","gravity-compensation","inverse-dynamics-model"],"name":"Inverse Dynamics","alt":"逆动力学","abbr":"","aliases":[],"one_liner":"Working backward from a desired motion (position, velocity, acceleration) to the torque each joint needs to produce.","explanation":"A robot's dynamics equation can be written τ = M(θ)θ̈ + h(θ, θ̇): θ is the joint angle vector, θ̇ and θ̈ are joint velocity and acceleration, M is the mass matrix, h lumps together gravity, Coriolis, and centrifugal terms, and τ is joint torque. Given θ, θ̇, and θ̈, solving for τ is inverse dynamics; going the other way, given τ, solving for θ̈ is forward dynamics — which is what a simulator computes at every step. Inverse dynamics is commonly solved with the recursive Newton-Euler algorithm: velocities and accelerations propagate from the base out to the tip, then forces are propagated back from the tip to the base. It's the core of computed-torque control, gravity compensation, and whole-body control, and biomechanics researchers also use it to work back from motion-capture data and ground reaction forces to human joint torques. It's distinct from inverse kinematics, which only solves for joint angles, and also distinct from an “inverse dynamics model” in embodied learning, which infers an action from a pair of consecutive frames.","example":"To lift a 2 kg object along a planned acceleration profile, a controller first uses inverse dynamics to compute the torque each joint needs at that instant as a feedforward term, then adds PD feedback on top to correct any error — this is computed-torque control.","related":["Forward Dynamics","Recursive Newton-Euler Algorithm","Mass Matrix","Computed Torque Control","Gravity Compensation","Inverse Dynamics Model"]},{"id":"newton-euler-equations","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Newton–Euler equations","url":"https://en.wikipedia.org/wiki/Newton%E2%80%93Euler_equations"},{"title":"Modern Robotics 8.3: Newton-Euler Inverse Dynamics (Northwestern)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-3-newton-euler-inverse-dynamics/"}],"as_of":"","related_ids":["recursive-newton-euler-algorithm","euler-lagrange-equations","inverse-dynamics","rigid-body-dynamics","inertia-tensor","computed-torque-control"],"name":"Newton-Euler Equations","alt":"牛顿-欧拉方程","abbr":"","aliases":["Newton-Euler Formulation"],"one_liner":"A pair of dynamics equations describing a rigid body's translation (F = ma) and rotation (Euler's equation) together.","explanation":"The Newton-Euler equations combine Newton's second law with Euler's equation for rigid-body rotation to describe how a rigid body translates and rotates under applied forces and torques: F = ma, and τ = Iα + ω×(Iω). F is the net external force, m is mass, a is the acceleration of the center of mass; τ is the net torque about the center of mass, I is the inertia tensor, α is angular acceleration, ω is angular velocity, and ω×(Iω) is the gyroscopic term produced by the rotation itself. A robot is made of many links, and each one satisfies this pair of equations. The classic recursive Newton-Euler algorithm first sweeps from base to tip, computing each link's velocity and acceleration, then sweeps back from tip to base, computing forces and torques, ending with the torque required at every joint — its computation cost grows only linearly with the number of joints. Along with the Euler-Lagrange equations, it's one of the two main routes to deriving robot dynamics.","example":"Given an arm's current joint angles, joint velocities, and the desired joint accelerations, the recursive Newton-Euler algorithm computes the torque each motor should output — exactly the feedforward term used in computed-torque control.","related":["Recursive Newton-Euler Algorithm","Euler-Lagrange Equations","Inverse Dynamics","Rigid-Body Dynamics","Inertia Tensor","Computed Torque Control"]},{"id":"euler-lagrange-equations","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Lagrangian mechanics - Wikipedia","url":"https://en.wikipedia.org/wiki/Lagrangian_mechanics"},{"title":"Underactuated Robotics (MIT) - Multi-Body Dynamics","url":"https://underactuated.csail.mit.edu/multibody.html"}],"as_of":"","related_ids":["newton-euler-equations","generalized-coordinates","mass-matrix","coriolis-and-centrifugal-terms","forward-dynamics","inverse-dynamics"],"name":"Euler-Lagrange Equations","alt":"拉格朗日方程","abbr":"","aliases":["Lagrangian Dynamics","Lagrangian Mechanics"],"one_liner":"A way of deriving a robot's equations of motion from kinetic and potential energy, without analyzing constraint forces one by one.","explanation":"Lagrangian mechanics was developed by Joseph-Louis Lagrange, laid out fully in his 1788 book Mécanique analytique. The method starts by writing the Lagrangian L = T − V (T is kinetic energy, V is potential energy), then, for each generalized coordinate q_i (such as a joint angle), writes an equation d/dt(∂L/∂q̇_i) − ∂L/∂q_i = τ_i, where τ_i is the generalized force acting on that coordinate (such as a motor torque). The advantage is that it only requires computing energy — the constraint forces between joints cancel out automatically — and the number of equations equals the number of degrees of freedom. For a robot arm, this works out to the standard form M(q)q̈ + C(q,q̇)q̇ + g(q) = τ: M is the mass matrix, the C term captures Coriolis and centrifugal forces, and g is the gravity term. This set of equations underlies torque control, gravity compensation, model predictive control, and physics simulation; deriving it by hand suits a two- or three-joint textbook example, while real robots compute it numerically with algorithms like the recursive Newton-Euler algorithm.","example":"A simple pendulum: length l, mass m, angle θ, kinetic energy T = ½ml²θ̇², potential energy V = −mgl·cosθ. Substituting into the equation gives ml²θ̈ + mgl·sinθ = τ — the familiar pendulum equation.","related":["Newton-Euler Equations","Generalized Coordinates","Mass Matrix","Coriolis and Centrifugal Terms","Forward Dynamics","Inverse Dynamics"]},{"id":"mass-matrix","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Modern Robotics 8.1.3: Understanding the Mass Matrix (Northwestern)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/8-1-3-understanding-the-mass-matrix/"},{"title":"Wikipedia: Mass matrix","url":"https://en.wikipedia.org/wiki/Mass_matrix"}],"as_of":"","related_ids":["rigid-body-dynamics","coriolis-and-centrifugal-terms","composite-rigid-body-algorithm","inverse-dynamics","forward-dynamics","euler-lagrange-equations"],"name":"Mass Matrix","alt":"质量矩阵","abbr":"","aliases":["Inertia Matrix","Joint-Space Inertia Matrix (JSIM)","M(q)"],"one_liner":"The matrix M(q) in a robot's dynamics equation that describes how hard each joint is to accelerate.","explanation":"The mass matrix appears in a robot's standard dynamics equation, τ = M(q)q̈ + c(q, q̇) + g(q): τ is joint torque, q, q̇, and q̈ are joint angles and their first and second derivatives, c is the Coriolis and centrifugal term, and g is the gravity term. M(q) is an n×n matrix (n is the number of joints) that determines how much torque a given set of joint accelerations requires; kinetic energy can also be written as T = ½q̇ᵀM(q)q̇. It's symmetric, positive definite, and changes with posture: with the arm fully extended, the shoulder joint has to move more inertia. Its off-diagonal entries capture inertial coupling between joints — accelerating only the elbow joint also produces a reaction torque at the shoulder. Computed-torque control, operational-space control, and forward-dynamics simulation all need it, and libraries like Pinocchio and MuJoCo compute it directly.","example":"For a planar two-link arm, when the elbow is straight, the diagonal entry of the mass matrix corresponding to the shoulder joint is at its largest — the same shoulder torque produces less angular acceleration when the arm is extended than when the elbow is bent.","related":["Rigid-Body Dynamics","Coriolis and Centrifugal Terms","Composite Rigid Body Algorithm","Inverse Dynamics","Forward Dynamics","Euler-Lagrange Equations"]},{"id":"coriolis-and-centrifugal-terms","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（Ch. 8 Dynamics of Open Chains）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Wikipedia: Coriolis force","url":"https://en.wikipedia.org/wiki/Coriolis_force"}],"as_of":"","related_ids":["euler-lagrange-equations","mass-matrix","rigid-body-dynamics","inverse-dynamics","computed-torque-control","recursive-newton-euler-algorithm"],"name":"Coriolis and Centrifugal Terms","alt":"科里奥利力与离心力项","abbr":"","aliases":["Coriolis Matrix","C(q, q̇)","Nonlinear Velocity Terms"],"one_liner":"The torque terms in a robot's equations of motion that depend on the square or product of joint velocities.","explanation":"Manipulator dynamics is commonly written as M(q)q̈ + C(q,q̇)q̇ + g(q) = τ: q is the joint angles, q̇ and q̈ are joint velocity and acceleration, M is the mass matrix, g is the gravity term, and τ is joint torque. The term C(q,q̇)q̇ is the Coriolis and centrifugal term. Modern Robotics calls the part that depends only on the square of a single joint velocity, q̇ᵢ², the centrifugal term, and the part depending on the product of two different joint velocities, q̇ᵢq̇ⱼ, the Coriolis term. It arises because joint coordinates aren't an inertial frame: even when every joint spins at constant velocity (q̈ = 0), the links are still moving in circles, which requires torque to sustain. At low speed this term is small and often ignored, but at high speed, failing to compensate for it produces visible tracking error. Inverse dynamics and computed-torque control both need to compute it, and the skew-symmetry property of Ṁ − 2C is a standard tool for proving controller stability.","example":"In a planar two-link arm with the second joint bent to 90°, if both joints rotate forward at constant angular velocity simultaneously, the end-effector mass is pulled closer to joint 1, and joint 1 actually needs to output a negative torque to hold its speed steady — this is the Coriolis term at work.","related":["Euler-Lagrange Equations","Mass Matrix","Rigid-Body Dynamics","Inverse Dynamics","Computed Torque Control","Recursive Newton-Euler Algorithm"]},{"id":"recursive-newton-euler-algorithm","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Modern Robotics (Lynch & Park), Sec. 8.3 Newton–Euler Inverse Dynamics","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo Documentation: Computation（bias force via RNE with acceleration set to 0）","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Pinocchio documentation（Recursive Newton-Euler algorithm, pinocchio::rnea）","url":"https://gepettoweb.laas.fr/doc/stack-of-tasks/pinocchio/master/doxygen-html/"}],"as_of":"","related_ids":["inverse-dynamics","newton-euler-equations","articulated-body-algorithm","composite-rigid-body-algorithm","computed-torque-control","pinocchio"],"name":"Recursive Newton-Euler Algorithm","alt":"递归牛顿-欧拉算法","abbr":"RNEA","aliases":["RNEA","RNE","Newton-Euler Inverse Dynamics"],"one_liner":"Sweeps velocity and acceleration outward from the base, then sweeps force back inward from the tip, to compute inverse-dynamics joint torques.","explanation":"The recursive Newton-Euler algorithm computes inverse dynamics: given joint positions q, velocities q̇, and accelerations q̈, it finds the required joint torques τ = M(q)q̈ + C(q,q̇)q̇ + g(q), where M is the mass matrix, the C term is the Coriolis and centrifugal forces, and g is the gravity term. It runs in two passes: a forward pass from the base to the tip computes each link's velocity and acceleration; a backward pass from the tip back to the base applies the Newton-Euler equations to each link to find the force and moment acting on it, and projecting that onto the joint axis gives τ. The computation grows only linearly with the number of joints, far more efficient than expanding the Lagrangian equations directly. Luh, Walker, and Paul's 1980 online computation scheme is the classic form, and Featherstone gave a unified treatment using 6-dimensional spatial vectors. Both Pinocchio's rnea() and MuJoCo's mj_rne implement it; setting acceleration to zero yields the bias forces from gravity and Coriolis effects, which are commonly used for gravity compensation, computed-torque control, and model-based collision detection.","example":"With Pinocchio, pinocchio.rnea(model, data, q, v, a) directly returns the joint torques; setting v and a to zero gives exactly the gravity-compensation torque needed to hold the arm still in its current pose.","related":["Inverse Dynamics","Newton-Euler Equations","Articulated Body Algorithm","Composite Rigid Body Algorithm","Computed Torque Control","Pinocchio"]},{"id":"composite-rigid-body-algorithm","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"MuJoCo 文档: Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Walker & Orin: Efficient Dynamic Computer Simulation of Robotic Mechanisms (1982)","url":"https://doi.org/10.1115/1.3139699"},{"title":"Pinocchio 项目主页（CRBA 等算法列表）","url":"https://stack-of-tasks.github.io/pinocchio/"}],"as_of":"","related_ids":["mass-matrix","articulated-body-algorithm","recursive-newton-euler-algorithm","forward-dynamics","kinematic-tree","mujoco"],"name":"Composite Rigid Body Algorithm","alt":"复合刚体算法","abbr":"CRBA","aliases":["CRBA","Composite-Rigid-Body Algorithm"],"one_liner":"An algorithm that efficiently computes the mass matrix by merging each joint's outboard links into one composite rigid body.","explanation":"The composite rigid body algorithm is a recursive method for computing the joint-space mass matrix M(q), which describes the inertial relationship between joint accelerations and the joint torques needed to produce them; it's generally traced back to a 1982 paper by Walker and Orin. The core idea: for joint i, treat every link outboard of it as if locked together into a single composite rigid body — the inertia of that composite body directly gives the entries of M associated with that joint, and the composite inertias can be accumulated stage by stage from the tip back to the root. Once M is known, the recursive Newton-Euler algorithm supplies the Coriolis, centrifugal, and gravity terms c, and solving M·q̈ = τ − c completes the forward-dynamics calculation. CRBA and the articulated body algorithm are the two common routes to forward dynamics.","example":"MuJoCo uses the CRB algorithm at every simulation step to get the joint-space mass matrix M, uses the RNE algorithm (with acceleration set to zero) to get the bias force made up of Coriolis, centrifugal, and gravity terms, and then solves for acceleration together with the constraints; Pinocchio's corresponding function is crba.","related":["Mass Matrix","Articulated Body Algorithm","Recursive Newton-Euler Algorithm","Forward Dynamics","Kinematic Tree","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"articulated-body-algorithm","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Featherstone: The Calculation of Robot Dynamics Using Articulated-Body Inertias (IJRR, 1983)","url":"https://doi.org/10.1177/027836498300200102"},{"title":"Wikipedia: Featherstone's algorithm","url":"https://en.wikipedia.org/wiki/Featherstone%27s_algorithm"},{"title":"Pinocchio GitHub README","url":"https://github.com/stack-of-tasks/pinocchio"}],"as_of":"","related_ids":["forward-dynamics","composite-rigid-body-algorithm","recursive-newton-euler-algorithm","spatial-vector-algebra","mass-matrix","physics-engine"],"name":"Articulated Body Algorithm","alt":"铰接体算法","abbr":"ABA","aliases":["ABA","Featherstone's Algorithm","Featherstone Algorithm"],"one_liner":"An O(n) recursive algorithm that computes joint accelerations from joint torques — the classic solution to forward dynamics.","explanation":"The Articulated Body Algorithm (ABA) was introduced by Roy Featherstone in a 1983 IJRR paper, and is also known simply as Featherstone's algorithm. It solves the forward dynamics problem: given the joint angles, joint velocities, and joint torques, find the joint accelerations. ABA introduces the idea of 'articulated-body inertia' — the effective inertia that a link, together with every link outboard of it, exhibits as a single jointed unit. The algorithm sweeps the kinematic chain three times: outward from the root to compute velocities, inward to accumulate articulated-body inertias, and outward again to compute accelerations, so the total cost grows only linearly with the number of joints n. Rigid-body dynamics libraries such as Pinocchio all implement it.","example":"Simulating each timestep of a 7-axis robot arm: given the current joint angles, velocities, and motor torques, call ABA (for example, Pinocchio's aba function) to get the joint accelerations, then numerically integrate to obtain the state at the next timestep.","related":["Forward Dynamics","Composite Rigid Body Algorithm","Recursive Newton-Euler Algorithm","Spatial Vector Algebra","Mass Matrix","Physics Engine"]},{"id":"spatial-vector-algebra","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Roy Featherstone: Spatial Vectors and Rigid-Body Dynamics","url":"https://royfeatherstone.org/spatial/"},{"title":"Roy Featherstone: spatial_v2 software (accompanies Rigid Body Dynamics Algorithms)","url":"https://royfeatherstone.org/spatial/v2/index.html"},{"title":"Pinocchio (GitHub)","url":"https://github.com/stack-of-tasks/pinocchio"}],"as_of":"","related_ids":["recursive-newton-euler-algorithm","articulated-body-algorithm","composite-rigid-body-algorithm","twist","wrench","pinocchio"],"name":"Spatial Vector Algebra","alt":"空间向量代数","abbr":"","aliases":["Spatial Vectors","6D Spatial Vectors"],"one_liner":"A notation that packs angular and linear quantities into 6-dimensional vectors for computing multi-body dynamics.","explanation":"Spatial vector algebra was systematically developed by Roy Featherstone, and forms the mathematical foundation of his book Rigid Body Dynamics Algorithms. The approach packs a rigid body's angular velocity and linear velocity into a single 6-dimensional 'motion vector,' and its torque and force into a single 6-dimensional 'force vector,' with inertia likewise written as a 6×6 spatial inertia matrix — letting Newton's equation for translation and Euler's equation for rotation merge into one 6-dimensional equation. Motion vectors and force vectors are two distinct kinds of quantity, each with its own cross-product operation; they're fundamentally the same as the twist and wrench from screw theory. The benefit is shorter formulas and a uniform way of transforming between coordinate frames, which keeps recursive algorithms tidy: the recursive Newton-Euler algorithm (RNEA, for inverse dynamics), the articulated body algorithm (ABA, for forward dynamics), and the composite rigid body algorithm (CRBA, for the mass matrix) all use this notation. Dynamics libraries such as Pinocchio are all built on Featherstone's algorithms.","example":"A torso tumbling through the air has its velocity written as a 6-dimensional vector: by Featherstone's convention, the first 3 entries are angular velocity and the last 3 are linear velocity (some libraries use the opposite order, so it pays to check when reading code). Converting from a parent link's frame to a child link's frame is just multiplying by a single 6×6 transformation matrix, which converts angular and linear velocity together in one step.","related":["Recursive Newton-Euler Algorithm","Articulated Body Algorithm","Composite Rigid Body Algorithm","Twist","Wrench","Pinocchio"]},{"id":"dynamic-parameter-identification","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Atkeson, An, Hollerbach: Estimation of Inertial Parameters of Manipulator Loads and Links (IJRR 1986)","url":"https://doi.org/10.1177/027836498600500306"},{"title":"Swevers et al.: Optimal robot excitation and identification (IEEE TRA 1997)","url":"https://doi.org/10.1109/70.631234"},{"title":"Wensing, Kim, Slotine: Linear Matrix Inequalities for Physically-Consistent Inertial Parameter Identification","url":"https://arxiv.org/abs/1701.04395"}],"as_of":"","related_ids":["inertial-parameters","system-identification","rigid-body-dynamics","gravity-compensation","friction-compensation","payload"],"name":"Dynamic Parameter Identification","alt":"动力学参数辨识","abbr":"","aliases":["Inertial Parameter Identification","Payload Identification"],"one_liner":"Moving the robot through specific trajectories and using the torque data to work out each link's mass and inertia.","explanation":"Dynamic parameter identification means estimating the parameters of a robot's dynamics model from measured data: each link's mass, first moment of mass (related to the center-of-mass location), and inertia tensor — 10 inertial parameters per link in total, often together with joint friction. Atkeson, An, and Hollerbach showed in 1986 that the Newton-Euler equations can be rewritten so joint torque is linear in these parameters, τ = Y(q,q̇,q̈)π, where Y is a regressor matrix depending only on joint position, velocity, and acceleration, and π is the parameter vector — making least squares sufficient to solve for them. To excite every parameter well enough in the data, the arm is often driven along an excitation trajectory optimized as a Fourier series (Swevers et al., 1997). CAD-supplied parameters are often inaccurate, and they change further once a payload is attached; accurate identification is what makes gravity compensation, computed-torque control, and sim-to-real alignment reliable. Wensing and colleagues added linear-matrix-inequality constraints in 2017 to guarantee the identified parameters remain physically consistent.","example":"After a new gripper with unknown parameters is mounted on an arm's end effector, the joints are driven through a periodic motion for a few dozen seconds while joint angles and torques (or currents) are logged; least squares then solves for the payload's mass and center of mass, which are fed back into the gravity-compensation model.","related":["Inertial Parameters","System Identification","Rigid-Body Dynamics","Gravity Compensation","Friction Compensation","Payload"]},{"id":"stiffness","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Stiffness - Wikipedia","url":"https://en.wikipedia.org/wiki/Stiffness"},{"title":"legged_gym: anymal_c_rough_config.py","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/anymal_c/mixed_terrains/anymal_c_rough_config.py"},{"title":"Modern Robotics: Mechanics, Planning, and Control (Lynch & Park, free preprint)","url":"https://hades.mech.northwestern.edu/index.php/Modern_Robotics"}],"as_of":"","related_ids":["compliance","damping","impedance-control","stiffness-and-damping-gains","proportional-derivative-control","mass-spring-damper-system"],"name":"Stiffness","alt":"刚度","abbr":"","aliases":["Rigidity","Spring Constant"],"one_liner":"How strongly an object or joint resists deforming — the same force causes less deformation the stiffer it is.","explanation":"Stiffness is defined as the ratio of an applied force to the displacement it produces, k = F/δ (F is force, δ is displacement along the force's direction), measured in N/m; for rotation, it's torque divided by angle, in N·m/rad. Its reciprocal is called compliance. The term carries two meanings in robotics. One is structural stiffness: links, gearboxes, and belts all bend or twist under load, which directly affects how accurately the end effector reaches its target. The other is control stiffness — the 'virtual spring' a controller creates, such as the proportional gain in PD control or the stiffness matrix K in impedance control. High stiffness tracks a target accurately and resists disturbances but hits hard on impact; low stiffness is more compliant and safer for tasks involving contact. The 'stiffness' and 'damping' settings in reinforcement-learning locomotion configs refer to exactly this — the PD gains applied at each joint.","example":"legged_gym sets stiffness = 80 N·m/rad for each joint of the ANYmal C robot: when the policy's target angle differs from the actual joint angle by 0.1 rad, the joint produces a restoring torque of about 8 N·m, before the damping term is subtracted.","related":["Compliance","Damping","Impedance Control","Stiffness and Damping Gains","Proportional-Derivative Control","Mass-Spring-Damper System"]},{"id":"compliance","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Stiffness - Wikipedia（Compliance 为刚度的倒数）","url":"https://en.wikipedia.org/wiki/Stiffness"},{"title":"Remote center compliance - Wikipedia","url":"https://en.wikipedia.org/wiki/Remote_center_compliance"},{"title":"Modern Robotics（Lynch & Park, 2017 预印本）第 11 章 Robot Control","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["stiffness","impedance-control","admittance-control","compliance-control","series-elastic-actuator","peg-in-hole-insertion"],"name":"Compliance","alt":"柔顺性","abbr":"","aliases":[],"one_liner":"How readily a robot yields to an external force; numerically, the inverse of stiffness.","explanation":"In mechanics, compliance is the inverse of stiffness: stiffness k is how much force it takes to produce a unit of deformation (N/m), and compliance 1/k is how much deformation a unit of force produces (m/N). In robotics, it describes whether an arm or joint gives way under an external push: a purely position-controlled arm is “stiff” and generates a large force if it hits a table or a person, while a compliant robot yields like a spring instead. Compliance can be passive, built into the mechanism's own elasticity — a series elastic actuator, the twisting flex of a harmonic reducer's flexspline, or the remote center compliance (RCC) device invented at Draper Laboratory in the 1970s — or active, where impedance control or admittance control simulates a virtual spring in software. Contact-rich tasks like peg-in-hole assembly, wiping a table, and working alongside people all depend on it.","example":"During peg-in-hole assembly, a slight misalignment will jam a rigid arm's peg against the edge of the hole; adding an RCC device between the wrist and the gripper lets the peg, once it touches the hole's chamfered edge, get pushed sideways and slide itself into alignment automatically.","related":["Stiffness","Impedance Control","Admittance Control","Compliance Control","Series Elastic Actuator (SEA)","Peg-in-Hole Insertion"]},{"id":"damping","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Damping - Wikipedia","url":"https://en.wikipedia.org/wiki/Damping"},{"title":"Modern Robotics（Lynch & Park, 2017 预印本）第 11 章 Robot Control","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Actuators - Isaac Lab Documentation","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/actuators.html"}],"as_of":"","related_ids":["damping-ratio","stiffness","mass-spring-damper-system","proportional-derivative-control","stiffness-and-damping-gains","impedance-control"],"name":"Damping","alt":"阻尼","abbr":"","aliases":["Damping Coefficient"],"one_liner":"The effect that gradually dissipates a moving or vibrating system's energy, settling it down and slowing it to rest.","explanation":"Damping refers to whatever dissipates energy in an oscillating or moving system — fluid viscosity, friction, and so on. The most common linear (viscous) model gives a damping force of −b·v, where v is velocity and b is the damping coefficient: the faster the motion, the greater the resistance. In a mass-spring-damper system, m·ẍ + b·ẋ + k·x = f (m is mass, k is stiffness, f is the external force), the damping ratio ζ determines the shape of the response: ζ < 1 is underdamped, oscillating back and forth; ζ = 1 is critically damped, returning to equilibrium fastest with no overshoot; ζ > 1 is overdamped, returning slowly. In robotics, the D gain in PD control acts like adding a virtual damper to a joint — simulators like Isaac Lab even name their joint control parameters stiffness and damping directly — and a joint's own mechanical damping also needs to be identified and written into the simulation model.","example":"Legged robots commonly use PD control at the joint, τ = kp·(q_des − q) + kd·(q̇_des − q̇), where kd is the damping gain: too small and the joint shakes and overshoots, too large and the response becomes sluggish. Modern Robotics recommends choosing gains near critical damping, ζ = 1.","related":["Damping Ratio","Stiffness","Mass-Spring-Damper System","Proportional-Derivative Control","Stiffness and Damping Gains","Impedance Control"]},{"id":"mass-spring-damper-system","category":"mechanics","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Mass-spring-damper model","url":"https://en.wikipedia.org/wiki/Mass-spring-damper_model"}],"as_of":"","related_ids":["damping-ratio","natural-frequency","stiffness","damping","impedance-control","proportional-derivative-control"],"name":"Mass-Spring-Damper System","alt":"质量-弹簧-阻尼系统","abbr":"","aliases":["Spring-Mass-Damper System"],"one_liner":"The most basic vibrating system, made of a mass, a spring, and a damper, used as an approximation throughout control and contact modeling.","explanation":"The mass-spring-damper system is the most fundamental dynamic model in mechanics and control: a mass m attached to a spring of stiffness k and a damper with damping coefficient c, with equation of motion mẍ + cẋ + kx = F, where x is displacement from equilibrium and F is an external force. The spring pulls the mass back toward equilibrium, and the damper dissipates energy. It defines a natural frequency ωn = √(k/m) and a damping ratio ζ = c/(2√(km)): ζ < 1 oscillates its way to a stop (underdamped), ζ = 1 returns to equilibrium fastest with no overshoot (critically damped), and ζ > 1 creeps back slowly (overdamped). It's everywhere in robotics: PD control at a joint behaves like a virtual spring plus damper, impedance control makes an end effector behave like a desired mass-spring-damper system, and many simulators compute soft-contact forces using a spring-damper model too.","example":"In joint PD control, τ = Kp(q_target − q) − Kd·q̇, Kp acts like spring stiffness k and Kd acts like damping c: too small a Kd and the joint oscillates back and forth, too large and the motion becomes sluggish.","related":["Damping Ratio","Natural Frequency","Stiffness","Damping","Impedance Control","Proportional-Derivative Control"]},{"id":"natural-frequency","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Natural frequency","url":"https://en.wikipedia.org/wiki/Natural_frequency"},{"title":"Lynch & Park, Modern Robotics（2017）11.2–11.4 节：二阶误差动力学与 PD 增益","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["damping-ratio","mass-spring-damper-system","stiffness","flexible-joint","proportional-derivative-control","control-bandwidth"],"name":"Natural Frequency","alt":"固有频率","abbr":"","aliases":["ωn","Resonant Frequency"],"one_liner":"The frequency at which a system oscillates on its own with no sustained outside force; matching this frequency causes resonance.","explanation":"Natural frequency is the frequency at which a vibrating system oscillates on its own, with no sustained external excitation. For the simplest mass-spring system, ω₀ = √(k/m), where k is the spring stiffness and m is the mass, in rad/s (divide by 2π to convert to hertz). The stiffer the spring or the lighter the mass, the faster it vibrates. When an external force's frequency approaches the natural frequency, the amplitude grows sharply — this is resonance. It matters for robots in two ways. One is structural: harmonic drives, long slender links, and flexible joints all have some elasticity, so the whole machine has natural frequencies, and either too-high control gains or a sudden change in trajectory acceleration can excite unwanted vibration. The other is control: under PD control, a joint's error obeys a second-order equation whose natural frequency is ωn = √(Kp/M) (Kp is the proportional gain, M is the rotational inertia), which, together with the damping ratio, determines how fast the response is and how much it overshoots.","example":"With a joint inertia M = 0.1 kg·m² and proportional gain Kp = 10 N·m/rad, ωn = √(10/0.1) = 10 rad/s, about 1.6 Hz; raising Kp to 40 doubles ωn to 20 rad/s, giving a faster response but also making it easier to excite structural vibration from unmodeled flexibility.","related":["Damping Ratio","Mass-Spring-Damper System","Stiffness","Flexible Joint","Proportional-Derivative Control","Control Bandwidth"]},{"id":"damping-ratio","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Damping（damping ratio）","url":"https://en.wikipedia.org/wiki/Damping"},{"title":"MuJoCo Documentation: Modeling – Solver parameters（solref: timeconst, dampratio）","url":"https://mujoco.readthedocs.io/en/stable/modeling.html"}],"as_of":"","related_ids":["damping","mass-spring-damper-system","natural-frequency","proportional-derivative-control","stiffness-and-damping-gains","step-response-metrics"],"name":"Damping Ratio","alt":"阻尼比","abbr":"","aliases":["ζ","Critical Damping","Underdamped / Overdamped"],"one_liner":"A dimensionless number measuring how fast oscillation decays — it determines whether a system overshoots and rings before settling.","explanation":"The damping ratio ζ describes how heavily damped a second-order system is, such as a mass-spring-damper system: ζ = c / (2√(km)), where m is mass, k is spring stiffness, c is the damping coefficient, and the denominator 2√(km) is called the critical damping. The equation of motion can be written ẍ + 2ζωₙẋ + ωₙ²x = 0, where ωₙ = √(k/m) is the natural frequency. ζ < 1 is underdamped: the system overshoots the target and oscillates back and forth before settling; ζ = 1 is critically damped: no overshoot, and the fastest possible return to equilibrium; ζ > 1 is overdamped: no oscillation, but a slower return. Joint PD control can be viewed as a virtual spring (Kp) plus damping (Kd), so tuning the gains is really tuning stiffness and damping ratio. MuJoCo's contact parameter solref is likewise specified using a time constant and a damping ratio, which is usually set to 1 (critical damping).","example":"Treat a joint as a rotor with 0.1 kg·m² of inertia and Kp = 40 N·m/rad; critical damping then requires Kd = 2√(40×0.1) = 4 N·m·s/rad. Any smaller Kd and the joint will oscillate as it settles into position.","related":["Damping","Mass-Spring-Damper System","Natural Frequency","Proportional-Derivative Control","Stiffness and Damping Gains","Step Response Metrics (Overshoot / Settling Time / Steady-State Error)"]},{"id":"flexible-joint","category":"mechanics","sec":6,"tier":3,"sources":[{"title":"Spong: Modeling and Control of Elastic Joint Robots (J. Dyn. Sys., Meas., Control 1987)","url":"https://doi.org/10.1115/1.3143860"},{"title":"Lynch & Park, Modern Robotics（8.9.5 Joint and Link Flexibility）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["strain-wave-gear","series-elastic-actuator","stiffness","compliance","actuator-modeling","joint-torque-sensor"],"name":"Flexible Joint","alt":"柔性关节","abbr":"","aliases":["Elastic Joint","Joint Flexibility"],"one_liner":"A joint with noticeable elasticity between the motor and the link, so the motor's angle doesn't equal the link's actual angle.","explanation":"A flexible joint is one where non-negligible elasticity exists between the motor and the link it drives. The most common source is a harmonic drive: its flexspline achieves near-zero backlash by deforming under load, which also introduces torsional elasticity. Spong's classic 1987 model treats such a joint as a torsional spring connecting the motor side and the link side, τ = K(θ − q): θ is the motor's angle referred to the output side, q is the link's actual angle, K is the joint stiffness, and τ is the torque transmitted across the spring. This adds an extra set of states per joint and raises the order of the dynamics, causing two problems: the motor encoder no longer reads the link's true angle, leaving a position offset at the end effector, and the low-frequency vibration mode this creates makes high-gain control prone to oscillating. Some designs introduce flexibility on purpose, such as series elastic actuators, which use the spring's deflection to measure force and improve safety during collisions.","example":"When a lightweight robot arm carrying a payload stops suddenly, the end effector often sways back and forth briefly before settling — the flexibility introduced by the harmonic drive is one cause.","related":["Strain Wave Gear (Harmonic Drive)","Series Elastic Actuator (SEA)","Stiffness","Compliance","Actuator Modeling (Actuator Network)","Joint Torque Sensor"]},{"id":"contact-force","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Contact force - Wikipedia","url":"https://en.wikipedia.org/wiki/Contact_force"},{"title":"Contact Sensor - Isaac Lab Documentation","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/sensors/contact_sensor.html"}],"as_of":"","related_ids":["normal-force-and-tangential-force","coulomb-friction","friction-cone","ground-reaction-force","contact-model","six-axis-force-torque-sensor"],"name":"Contact Force","alt":"接触力","abbr":"","aliases":[],"one_liner":"The force two objects exert on each other when touching, split into a normal (pressing) part and a tangential (friction) part.","explanation":"Contact force is the force that arises between two objects because they're physically touching, as opposed to action-at-a-distance forces like gravity or electromagnetism. It's usually split into two components: a normal force perpendicular to the contact surface (a pushing force only — it can press but never pull), and a tangential force parallel to the surface (friction), whose magnitude is limited by the friction coefficient and the normal force, as described by Coulomb friction. A robot's walking, grasping, pushing, and inserting all happen through contact forces: a legged robot supports itself and moves forward using the ground reaction force at its feet, and a gripper lifts an object using the normal force and friction at its fingertips. A physics engine has to solve for contact forces at every simulation step; on a real robot, they're measured with six-axis force sensors, foot force sensors, or tactile sensors, and controllers commonly write “the contact force must stay inside the friction cone” as a constraint.","example":"When training quadruped walking in Isaac Lab, each foot gets a contact sensor reading net_forces_w — the total contact force on that foot in world coordinates. A reading near zero means the foot has left the ground, which is used to infer gait phase and shape reward terms.","related":["Normal Force and Tangential (Shear) Force","Coulomb Friction","Friction Cone","Ground Reaction Force (GRF)","Contact Model","Six-Axis Force/Torque Sensor"]},{"id":"normal-force-and-tangential-force","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Normal force","url":"https://en.wikipedia.org/wiki/Normal_force"},{"title":"Wikipedia: Friction","url":"https://en.wikipedia.org/wiki/Friction"}],"as_of":"","related_ids":["contact-force","coulomb-friction","friction-cone","friction-coefficient","slip-detection","tactile-sensor"],"name":"Normal Force and Tangential (Shear) Force","alt":"法向力与切向力（剪切力）","abbr":"","aliases":["Normal Force","Tangential Force","Shear Force"],"one_liner":"The contact-force component perpendicular to a surface is normal force; the component parallel to it is tangential (shear) force.","explanation":"When two objects touch, the contact force between them can be split into two parts: a normal force, perpendicular to the contact surface, which presses the two surfaces together and keeps them from interpenetrating, and a tangential force, parallel to the surface — also called shear force — usually supplied by friction, which resists relative sliding. The two are linked by Coulomb's law of friction: |f_t| ≤ μf_n, where f_t is tangential force, f_n is normal force, and μ is the friction coefficient. If the tangential force needed exceeds μ times the normal force, the surfaces slip; in 3D, that constraint is exactly the friction cone. When a gripper holds a cup from the sides, the cup's entire weight is supported by tangential friction at the fingertips — if the grip force (normal force) isn't enough, the cup slides out. This is why tactile sensors usually measure both normal and shear force together, with changes in shear force serving as an important early signal of impending slip; legged robots likewise have to keep the ground reaction force inside the friction cone.","example":"A two-finger gripper holds a 0.3 kg cup from the side, with a friction coefficient of 0.5: each finger needs to provide about 1.5 N of upward tangential force, so each finger's normal (gripping) force must be at least about 3 N — with some margin added in practice.","related":["Contact Force","Coulomb Friction","Friction Cone","Friction Coefficient (Coefficient of Friction)","Slip Detection","Tactile Sensor"]},{"id":"friction-coefficient","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Friction - Wikipedia","url":"https://en.wikipedia.org/wiki/Friction"},{"title":"MuJoCo Documentation - Computation (contact, friction)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["coulomb-friction","friction-cone","normal-force-and-tangential-force","slip","static-friction-and-stribeck-effect","dynamics-randomization"],"name":"Friction Coefficient (Coefficient of Friction)","alt":"摩擦系数","abbr":"μ","aliases":["μ","Coefficient of Friction (CoF)"],"one_liner":"The ratio of maximum friction force to normal force, determining how easily a contact surface slips.","explanation":"By Coulomb's law of friction, the friction force f between two contacting surfaces never exceeds the friction coefficient μ times the normal force N (the force pressing the surfaces together): f ≤ μN. μ is a dimensionless number that depends on the two materials and their surface condition, is largely independent of contact area, and can only be measured experimentally. There are two kinds: the static friction coefficient μ_s sets the largest tangential force an object can resist before it starts sliding, and the kinetic friction coefficient μ_k applies once it's already sliding; μ_s is usually larger than μ_k. Common dry materials fall mostly between 0.3 and 0.6; rubber on concrete is about 0.6–0.85; ice on ice is only 0.02–0.09. In robotics, μ determines how much grip force a gripper needs to avoid dropping something, and how much horizontal force a foot can push against the ground with before slipping; MuJoCo assigns each contact separate sliding, torsional, and rolling friction parameters, and μ is commonly randomized during training to narrow the sim-to-real gap.","example":"A two-finger gripper holds a 0.5 kg cup, with normal force N per finger and μ = 0.5: total friction from both sides is 2×0.5×N, and supporting the roughly 4.9 N of weight needs N of at least about 4.9 N. Swap in a smooth cup surface with μ = 0.25, and the required grip force doubles to about 9.8 N.","related":["Coulomb Friction","Friction Cone","Normal Force and Tangential (Shear) Force","Slip","Static Friction (Stiction) and Stribeck Effect","Dynamics Randomization"]},{"id":"coulomb-friction","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Modern Robotics（Lynch & Park, 2017 预印本）12.2 节 Friction","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Friction - Wikipedia","url":"https://en.wikipedia.org/wiki/Friction"}],"as_of":"","related_ids":["friction-coefficient","friction-cone","static-friction-and-stribeck-effect","viscous-friction","contact-model","slip"],"name":"Coulomb Friction","alt":"库仑摩擦","abbr":"","aliases":["Coulomb Friction Model","Dry Friction"],"one_liner":"The classic friction model where friction force never exceeds the friction coefficient times normal force, independent of sliding speed.","explanation":"Coulomb friction is the empirical law describing contact between two dry solid surfaces: the tangential friction force f_t ≤ μ·f_n, where f_n is the normal (pressing) force and μ is the friction coefficient, commonly between 0.1 and 1. While not sliding, friction can take on any value up to that limit in whichever direction resists sliding; once sliding starts, friction equals μ·f_n exactly, points opposite the direction of sliding, and is independent of sliding speed or the nominal contact area. It's common to distinguish a static friction coefficient μs from a slightly smaller kinetic friction coefficient μk. The rule was first summarized by Guillaume Amontons in 1699 and studied in depth by Charles-Augustin de Coulomb in 1785. It's only an approximation, but it underlies both grasp analysis and physics engines: plotting f_t ≤ μ·f_n in 3D gives exactly the friction cone, and whether a contact will slip is just a matter of checking whether the contact force falls inside that cone.","example":"A two-finger gripper presses a cup from both sides with 10 N each. If μ = 0.5, each contact point provides at most 5 N of friction, 10 N total — enough to lift a roughly 1 kg cup (about 9.8 N of weight); anything heavier slips out from between the fingers.","related":["Friction Coefficient (Coefficient of Friction)","Friction Cone","Static Friction (Stiction) and Stribeck Effect","Viscous Friction","Contact Model","Slip"]},{"id":"slip","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Modern Robotics (Lynch & Park), Ch.12.2 Contact Forces and Friction","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Humanoid-Gym: humanoid_env.py (_reward_foot_slip)","url":"https://github.com/roboterax/humanoid-gym/blob/main/humanoid/envs/custom/humanoid_env.py"}],"as_of":"","related_ids":["coulomb-friction","friction-cone","friction-coefficient","slip-detection","force-closure","stance-phase-swing-phase"],"name":"Slip","alt":"打滑","abbr":"","aliases":["Sliding","Slippage","Foot Slip"],"one_liner":"Relative sliding between contact surfaces, usually because the tangential force needed exceeds the friction limit.","explanation":"Slip means two touching objects have relative motion at their contact point; the opposite state, where the contact point stays relatively still, is called sticking. Under the common Coulomb friction model, the tangential friction force f_t must satisfy f_t ≤ μf_n (μ is the friction coefficient, f_n is the normal force); the moment the tangential force actually needed exceeds this limit, the contact switches from sticking to sliding, at which point friction equals exactly μf_n and points opposite the direction of motion — geometrically, the contact force has left the friction cone. In manipulation, a grasp that slips can drop or spin an object out of position, so the fix is either more grip force or slip detection with a tactile sensor; in legged robots, a support foot that slips is a common cause of falling, so reinforcement-learning training often penalizes horizontal velocity at the contact foot. Deliberately controlled sliding is also a skill in its own right — pushing an object across a table, or letting it rotate between the fingers.","example":"A two-finger gripper holds a 0.5 kg cup vertically, with its roughly 4.9 N of weight split between the two contact points' friction. If μ = 0.5, each finger needs at least about 4.9 N of normal force (2×0.5×4.9≈4.9 N), or the cup slides downward. Humanoid-Gym's foot_slip reward term penalizes a foot's horizontal velocity whenever it's in contact with the ground.","related":["Coulomb Friction","Friction Cone","Friction Coefficient (Coefficient of Friction)","Slip Detection","Force Closure","Stance Phase / Swing Phase"]},{"id":"friction-cone","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Robotic Manipulation (MIT, Russ Tedrake) - Bin Picking: friction cone","url":"https://manipulation.csail.mit.edu/clutter.html"},{"title":"MuJoCo Documentation - Computation (pyramidal / elliptic cones)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Friction - Wikipedia","url":"https://en.wikipedia.org/wiki/Friction"}],"as_of":"","related_ids":["friction-coefficient","coulomb-friction","friction-pyramid","force-closure","ground-reaction-force","contact-force-optimization"],"name":"Friction Cone","alt":"摩擦锥","abbr":"","aliases":[],"one_liner":"The cone-shaped set of every force a contact point can apply on an object without slipping.","explanation":"Under Coulomb's law of friction, a contact force's component tangential to the surface can't exceed its normal component (perpendicular to the surface) times the friction coefficient μ: √(f_x² + f_y²) ≤ μ·f_z, where f_z is the normal force and f_x, f_y are the two tangential components. Every force satisfying that inequality forms, in 3D, exactly a cone centered on the contact normal, with half-angle θ satisfying tanθ = μ. A force inside the cone won't slip; one outside it will. The friction cone is a basic constraint in contact computation: grasp analysis uses it to check force closure, model predictive control and whole-body control for legged robots require each foot's ground reaction force to stay inside its cone, and physics engines use it to constrain contact forces too. Because a cone is a nonlinear constraint, optimization commonly approximates it with a polyhedral cone (a friction pyramid) to get linear inequalities instead; MuJoCo offers both a pyramidal and an elliptical cone option.","example":"A quadruped robot walking on ground with μ around 0.6 has its controller cap each foot's horizontal force at 0.6 times its vertical force. On slippery tile, μ drops and the cone narrows, so the robot has to shorten its stride and reduce the horizontal push-off force.","related":["Friction Coefficient (Coefficient of Friction)","Coulomb Friction","Friction Pyramid","Force Closure","Ground Reaction Force (GRF)","Contact Force Optimization (Force Distribution)"]},{"id":"friction-pyramid","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（预印本 PDF，12.2.1 节 Figure 12.18）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MuJoCo Documentation: Computation（Friction cones）","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["friction-cone","coulomb-friction","friction-coefficient","convex-mpc","linear-complementarity-problem","mujoco"],"name":"Friction Pyramid","alt":"摩擦金字塔","abbr":"","aliases":["Linearized Friction Cone","Polyhedral Friction Cone","Pyramidal Friction Cone"],"one_liner":"Approximates the round friction cone with a polyhedral pyramid, turning the friction constraint into linear inequalities.","explanation":"Under Coulomb friction, the tangential force at a contact point can't exceed the friction coefficient μ times the normal force, and all the allowed contact forces form a cone in 3D — the friction cone. Because a cone is a nonlinear constraint, it's often approximated with a polyhedral pyramid instead, which is the friction pyramid: the simplest 4-sided pyramid is spanned by the four edges (μ,0,1), (−μ,0,1), (0,μ,1), and (0,−μ,1), and more edges get closer to the true cone. An inscribed pyramid underestimates the available friction, while a circumscribed one overestimates it; grasp analysis usually takes the conservative, inscribed option. The payoff is that the friction constraint becomes a set of linear inequalities, so contact problems can be solved with linear programming, quadratic programming, or a linear complementarity problem. The |fx| ≤ μfz, |fy| ≤ μfz constraints common in convex MPC for quadrupeds are a circumscribed 4-sided pyramid; MuJoCo also offers two global settings, pyramidal and elliptic, and the pyramidal option turns the solver's dual problem into a convex quadratic program with simple box constraints.","example":"With μ = 0.5 and a normal force fz = 100 N, the true friction cone allows at most 50 N of tangential force in any direction; with the circumscribed 4-sided pyramid constraint |fx| ≤ 50, |fy| ≤ 50, a solution with fx = fy = 50 N is also allowed, giving a combined force of about 70.7 N — beyond the true limit — which is why a conservative design often substitutes μ/√2 for μ.","related":["Friction Cone","Coulomb Friction","Friction Coefficient (Coefficient of Friction)","Convex MPC","Linear Complementarity Problem","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"torsional-and-rolling-friction","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"MuJoCo Documentation: Computation (Contact / Friction)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"MuJoCo Documentation: XML Reference (geom friction, condim)","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html"},{"title":"Wikipedia: Rolling resistance","url":"https://en.wikipedia.org/wiki/Rolling_resistance"}],"as_of":"","related_ids":["coulomb-friction","friction-coefficient","friction-cone","contact-model","mujoco","grasping"],"name":"Torsional and Rolling Friction","alt":"扭转摩擦与滚动摩擦","abbr":"","aliases":["Spin Friction","Rolling Resistance"],"one_liner":"Two more kinds of friction: one resists an object spinning in place, the other resists it rolling across a surface.","explanation":"Ordinary sliding friction resists two surfaces sliding across each other; there are two other kinds as well. Torsional friction resists an object spinning about the normal direction at a contact, which requires the contact to be a small patch of area rather than an ideal single point. Rolling friction (also called rolling resistance) resists an object rolling across a surface, coming mainly from the energy lost as the contact material repeatedly deforms (hysteresis), and it's usually much smaller than sliding friction — a steel wheel on a steel rail has a rolling-resistance coefficient of only about 0.0003–0.0004, versus about 0.01–0.015 for a car tire on concrete. In grasp analysis, a contact capable of providing torsional friction is called a soft-finger contact, since it can resist one more direction of torque than a plain point contact. Simulators expose these as tunable parameters: each MuJoCo geometry has three friction coefficients, in order sliding, torsional, and rolling, defaulting to 1, 0.005, and 0.0001; condim must be set to 4 to enable torsional friction, and to 6 to add rolling friction as well.","example":"In MuJoCo, gently push a ball resting on a plane: with condim set to 3 (sliding friction only), the ball keeps rolling indefinitely; with condim set to 6 and a rolling-friction coefficient specified, it gradually comes to a stop. Similarly, two fingers pinching a bottle so it can't twist between them are relying on torsional friction between the fingertips and the bottle's surface.","related":["Coulomb Friction","Friction Coefficient (Coefficient of Friction)","Friction Cone","Contact Model","MuJoCo (Multi-Joint dynamics with Contact)","Grasping"]},{"id":"static-friction-and-stribeck-effect","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Stribeck curve","url":"https://en.wikipedia.org/wiki/Stribeck_curve"},{"title":"Wikipedia: Stiction","url":"https://en.wikipedia.org/wiki/Stiction"},{"title":"Sorrentino et al., Physics-Informed Learning for the Friction Modeling of High-Ratio Harmonic Drives (arXiv:2410.12685)","url":"https://arxiv.org/abs/2410.12685"}],"as_of":"","related_ids":["coulomb-friction","viscous-friction","friction-compensation","strain-wave-gear","system-identification","friction-coefficient"],"name":"Static Friction (Stiction) and Stribeck Effect","alt":"静摩擦与 Stribeck 效应","abbr":"","aliases":["Stiction","Stribeck Curve","Stribeck Friction"],"one_liner":"Getting something moving takes extra force to beat static friction first, and once it starts, friction actually drops before rising again.","explanation":"Static friction, or stiction (a blend of 'static' and 'friction'), is the force that must be overcome to start relative motion between two surfaces at rest; the maximum static friction is usually larger than the kinetic friction once the surfaces are already sliding. The Stribeck effect is named after the German engineer Richard Stribeck, whose lubricated-bearing experiments published in 1901–1902 showed that friction in a lubricated contact doesn't stay constant with speed — it first drops from its static value as speed increases, reaches a minimum, and then rises again (the rising part comes mainly from viscous friction). Plotted as friction versus speed, this is the Stribeck curve. In robots, friction changes most sharply at low speed and at direction reversals in gearboxes and joints, which can produce a stick-then-suddenly-slip-then-stick-again behavior called stick-slip, causing low-speed creeping and small positioning errors. When compensating for joint friction or performing parameter identification, a combined Coulomb-plus-viscous-plus-Stribeck model is commonly fit to the data — for example, friction-identification work on the ergoCub humanoid's harmonic-drive joints uses this as its baseline model.","example":"Give a harmonic-drive joint a very small velocity command: at first the motor current rises while the joint doesn't move at all (static friction hasn't been overcome yet); once it breaks free, friction drops sharply and the joint suddenly overshoots, the controller pulls it back, and the cycle repeats — the encoder reading traces a sawtooth pattern, which is low-speed stick-slip.","related":["Coulomb Friction","Viscous Friction","Friction Compensation","Strain Wave Gear (Harmonic Drive)","System Identification","Friction Coefficient (Coefficient of Friction)"]},{"id":"viscous-friction","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Sorrentino et al., Physics-Informed Learning for the Friction Modeling of High-Ratio Harmonic Drives (arXiv:2410.12685)","url":"https://arxiv.org/abs/2410.12685"},{"title":"MuJoCo Documentation: Computation (Passive forces / damping)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Wikipedia: Friction","url":"https://en.wikipedia.org/wiki/Friction"}],"as_of":"","related_ids":["coulomb-friction","static-friction-and-stribeck-effect","damping","friction-compensation","dynamic-parameter-identification","dynamics-randomization"],"name":"Viscous Friction","alt":"黏性摩擦","abbr":"","aliases":["Viscous Damping"],"one_liner":"Friction proportional to relative speed — the faster something turns, the more resistance it meets.","explanation":"Viscous friction is a resistive force proportional to speed, in its simplest form F = −b·v, where v is the relative velocity (angular velocity, for a joint), b is the viscous friction coefficient, and the negative sign indicates the force opposes the motion. It arises from the internal friction (viscosity) within a fluid layer such as lubricating oil, and unlike Coulomb friction — whose magnitude is roughly independent of speed and depends only on the normal force — it vanishes when speed is zero. In robot joints, the friction in motors, gearboxes, and bearings is usually modeled as several parts together: Coulomb friction plus viscous friction, with the Stribeck effect added at low speed. Once b is identified, it can be fed forward in the controller to make joint torque control more accurate. Simulators also use it to model energy loss at a joint: MuJoCo's joint damping parameter is exactly this viscous damping coefficient, producing a passive force opposite the joint's velocity; when transferring from simulation to a real robot, this coefficient is often identified from the real hardware or randomized during training.","example":"If a joint has a viscous coefficient of b = 0.2 N·m·s/rad, then rotating at 1 rad/s produces a viscous friction torque of 0.2 N·m, and rotating at 2 rad/s produces 0.4 N·m; the Coulomb-friction portion, by contrast, is nearly the same at both speeds.","related":["Coulomb Friction","Static Friction (Stiction) and Stribeck Effect","Damping","Friction Compensation","Dynamic Parameter Identification","Dynamics Randomization"]},{"id":"impact","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Impact (mechanics)","url":"https://en.wikipedia.org/wiki/Impact_(mechanics)"},{"title":"Wikipedia: Coefficient of restitution","url":"https://en.wikipedia.org/wiki/Coefficient_of_restitution"},{"title":"Wensing et al. 2017, Proprioceptive Actuator Design in the MIT Cheetah: Impact Mitigation and High-Bandwidth Physical Interaction（IEEE T-RO）","url":"https://doi.org/10.1109/TRO.2016.2640183"}],"as_of":"","related_ids":["coefficient-of-restitution","backdrivability","reflected-inertia","quasi-direct-drive","contact-model","ground-reaction-force"],"name":"Impact","alt":"碰撞冲击","abbr":"","aliases":["Collision Impact","Impulsive Contact"],"one_liner":"When two objects collide, a very brief but very large contact force that abruptly changes their velocities.","explanation":"An impact is a collision between two objects in which the contact force spikes to a large value over an extremely short time, changing the objects' velocities almost instantaneously. Rigid-body simulation and control commonly treat it as an instantaneous event: the impulse (the time-integral of force) equals the change in momentum, and the coefficient of restitution e determines how much rebound results — e = 0 means no bounce at all, e = 1 means no energy loss. Impacts are common in robotics: a legged robot's foot touching down each step, a jump landing, or an arm bumping into a person or a table. Peak impact forces can damage gearboxes and sensors, and they can also make a foot slip or destabilize a controller. On the hardware side, engineers cushion impacts by reducing the motor inertia reflected to the output and improving backdrivability; Wensing and colleagues proposed an impact mitigation factor (IMF) in 2017 to quantify this ability for the MIT Cheetah, which had contact durations as short as 85 ms and peak forces over 450 N while running with the bound gait.","example":"A 1 kg ball falls vertically at 2 m/s and lands with a coefficient of restitution of 0.5, so it bounces back at 1 m/s; momentum changes from −2 to +1 kg·m/s, an impulse of 3 N·s, and if contact lasts only 10 ms, the average contact force is about 300 N — roughly 30 times the ball's weight.","related":["Coefficient of Restitution","Backdrivability","Reflected Inertia","Quasi-Direct Drive","Contact Model","Ground Reaction Force (GRF)"]},{"id":"coefficient-of-restitution","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Coefficient of restitution","url":"https://en.wikipedia.org/wiki/Coefficient_of_restitution"},{"title":"Isaac Lab API: isaaclab.envs.mdp（randomize_rigid_body_material）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.envs.mdp.html"},{"title":"Isaac Lab API: isaaclab.sim.spawners（RigidBodyMaterialCfg.restitution）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.sim.spawners.html"}],"as_of":"","related_ids":["impact","physics-engine","contact-model","domain-randomization","dynamics-randomization","friction-coefficient"],"name":"Coefficient of Restitution","alt":"恢复系数","abbr":"COR","aliases":["COR","Restitution"],"one_liner":"The ratio of separation speed after a collision to approach speed before it — a measure of how bouncy the impact is.","explanation":"The coefficient of restitution is an empirical parameter describing how elastic a collision is, usually written e, and it traces back to Newton's laws of impact: e equals the relative separation speed of the two bodies after the collision divided by their relative approach speed before it. e = 1 is a perfectly elastic collision with no kinetic-energy loss; e = 0 is perfectly inelastic, with no bounce at all; real collisions fall somewhere in between. In robotics it shows up mainly in two places: modeling the impact when a legged robot's foot lands or a gripper grasps an object, and as a material parameter in physics engines — Isaac Lab's rigid-body materials include a restitution value alongside static and dynamic friction coefficients, and training runs commonly randomize all three together so the policy doesn't overfit to one particular bounciness of ground or object.","example":"A ball dropped from 1 m and bouncing back up to 0.64 m has e = √(0.64/1) = 0.8. Isaac Lab's randomize_rigid_body_material function samples static friction, dynamic friction, and restitution randomly within given ranges.","related":["Impact","Physics Engine","Contact Model","Domain Randomization","Dynamics Randomization","Friction Coefficient (Coefficient of Friction)"]},{"id":"grasp-taxonomy","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"A comprehensive grasp taxonomy (Feix et al., RSS 2009 Workshop)","url":"https://www.csc.kth.se/grasp/taxonomyGRASP.pdf"},{"title":"The GRASP Taxonomy of Human Grasp Types (Feix et al., IEEE THMS 2016)","url":"https://doi.org/10.1109/THMS.2015.2470657"},{"title":"On grasp choice, grasp models, and the design of hands for manufacturing tasks (Cutkosky, 1989)","url":"https://doi.org/10.1109/70.34763"}],"as_of":"","related_ids":["grasping","dexterous-hand","dexterous-manipulation","hand-synergies","thumb-opposition","grasp-planning"],"name":"Grasp Taxonomy (Power Grasp vs. Precision Grasp / Pinch)","alt":"抓取分类（强力抓取 / 精确抓取 / 捏取）","abbr":"","aliases":["Power Grasp","Precision Grasp","Pinch Grasp","Cutkosky Grasp Taxonomy","GRASP Taxonomy"],"one_liner":"A classification of hand grasps by how the hand contacts an object, split at the top level into power grasps and precision grasps.","explanation":"Grasp taxonomy organizes the ways a human or robotic hand can grip something into a system of categories. In 1956, John Napier split human grasping into two broad types: a power grasp wraps the object with the palm and all four fingers, favoring stability and force, while a precision grasp mainly opposes the thumb pad against the pads of the other fingers, favoring fine position control. Pinching falls under precision grasping, holding a small object with just the fingertips or the sides of the fingers. In 1989, Mark Cutkosky observed manufacturing workers using tools and proposed a grasp taxonomy tree aimed at robotic hand design; Feix and colleagues later surveyed over a dozen sources and compiled 33 grasp types, further broken down by the opposition type (palm, fingertip, or finger side) and thumb posture, publishing this as the GRASP taxonomy in 2016. It's commonly used to decide which degrees of freedom a dexterous hand needs to keep, and as a category label in grasping datasets.","example":"Gripping a hammer handle to drive a nail is a power grasp; pinching a screw between thumb and index fingertip is a precision grasp. Evaluating a new dexterous hand often means checking, one by one, whether it can reproduce these standard grasp types.","related":["Grasping","Dexterous Hand","Dexterous Manipulation","Hand Synergies","Thumb Opposition","Grasp Planning"]},{"id":"antipodal-grasp","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Modern Robotics（Lynch & Park, 2017 预印本）第 12 章 Grasping and Manipulation","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics (arXiv 1703.09312)","url":"https://arxiv.org/abs/1703.09312"}],"as_of":"","related_ids":["force-closure","friction-cone","parallel-jaw-gripper","grasp-pose-detection","dex-net-2-0","grasp-quality-metric"],"name":"Antipodal Grasp","alt":"对跖抓取","abbr":"","aliases":["Antipodal Grasping","Two-Finger Opposed Grasp"],"one_liner":"A two-finger grasp where the contact points face each other and their connecting line lies inside both friction cones.","explanation":"An antipodal grasp — “antipodal” meaning “located at opposite ends” — grips an object with two contact points from opposing directions, such that the line connecting them falls inside both points' friction cones (the range of directions a contact force can push without slipping). According to Modern Robotics, in the planar case, two frictional contacts meeting this condition form a force closure, able to resist an external force or torque in any direction; in 3D, two ideal point contacts can't resist rotation about the line connecting them, which in practice is handled by the twisting friction of the finger pads. Chen and Burdick gave an algorithm in 1993 for finding antipodal point pairs on irregular objects. This is exactly the grasp a parallel-jaw two-finger gripper performs, which is why many learned grasping methods first sample candidate antipodal point pairs from a point cloud or depth image, then score them with a network.","example":"Dex-Net 2.0 (2017) finds antipodal point pairs in a depth image, generates hundreds of candidate parallel-jaw grasps, and scores them with a convolutional network called GQ-CNN, reaching a 93% grasp success rate on known objects with an ABB YuMi arm.","related":["Force Closure","Friction Cone","Parallel Jaw Gripper","Grasp Pose Detection","Dex-Net 2.0","Grasp Quality Metric"]},{"id":"force-closure","category":"mechanics","sec":7,"tier":2,"sources":[{"title":"Robotic Manipulation (MIT, Russ Tedrake) - Bin Picking: contact wrench cone, force closure","url":"https://manipulation.csail.mit.edu/clutter.html"}],"as_of":"","related_ids":["form-closure","friction-cone","wrench","antipodal-grasp","grasp-quality-metric","grasp-matrix"],"name":"Force Closure","alt":"力封闭","abbr":"","aliases":[],"one_liner":"A grasp where friction at the contact points can resist an external force or torque in any direction, so the object can't be pulled free.","explanation":"Force closure is the classic criterion in grasp mechanics for judging whether a grip is secure. Each contact point, thanks to friction, can apply force within a friction cone; stacking up all the force and torque (together called a wrench, 6 dimensions total) that every contact point can jointly produce, if the result covers the entire 6-dimensional space, the grasp is said to have force closure — no matter which direction an external push or twist comes from, the fingers can adjust their force to counteract it. Its counterpart is form closure, which pins the object down purely through contact geometry, without relying on friction, and usually needs more contact points. Force closure depends on the friction coefficient: the same fingertip positions might achieve it gripping a dry cardboard box but fail on an oily glass. It's the starting point for classical grasp planning and grasp quality metrics; for the planar case, a two-finger grasp achieves force closure exactly when the line connecting the two contact points falls inside both friction cones.","example":"A two-finger gripper holding a block, with the two contact faces parallel and their normals pointing directly at each other (an antipodal grasp): as long as the line between the two points falls inside both friction cones, the block can't be pulled up or pushed sideways out of the grip — force closure, in the planar sense.","related":["Form Closure","Friction Cone","Wrench","Antipodal Grasp","Grasp Quality Metric","Grasp Matrix"]},{"id":"form-closure","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（12.1.7 Form Closure, Theorem 12.6）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["force-closure","grasping","friction-cone","grasp-quality-metric","grasp-matrix","contact-force"],"name":"Form Closure","alt":"形封闭","abbr":"","aliases":["Geometric Closure","Form-Closure Grasp"],"one_liner":"Immobilizing an object purely through the geometry of contact points, with no reliance on friction at all.","explanation":"Form closure means a set of fixed constraints — fingers, jaws, locating pins, and so on — geometrically blocks every possible motion of an object: translating or rotating in any direction would push it directly into a constraint. When achieved with robot fingers, it's called a form-closure grasp. It doesn't depend on friction, which is what distinguishes it from force closure, where friction at the contact points is needed to resist forces in arbitrary directions. Lynch and Park's Modern Robotics gives the result that, under first-order analysis, a planar object needs at least 4 point contacts and a spatial object needs at least 7 (one more than the number of rigid-body degrees of freedom). Rotationally symmetric objects like disks and spheres can never achieve form closure no matter how many contact points are added, because they can always spin about their center. Form closure is common in fixture and jig design; robotic grasping usually can rely on friction and is more often analyzed in terms of force closure instead.","example":"A square part sitting inside a flat fixture surrounded by locating pins can't move in any direction, even if its surface is oiled — that's form closure. Gripping the same part with a parallel jaw gripper instead relies on friction, which is force closure.","related":["Force Closure","Grasping","Friction Cone","Grasp Quality Metric","Grasp Matrix","Contact Force"]},{"id":"contact-jacobian","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Underactuated Robotics（Tedrake）: Multi-Body Dynamics 一章","url":"https://underactuated.mit.edu/multibody.html"},{"title":"MuJoCo 文档: Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["jacobian-matrix","contact-force","ground-reaction-force","whole-body-control","contact-force-optimization","floating-base"],"name":"Contact Jacobian","alt":"接触雅可比","abbr":"","aliases":["Constraint Jacobian"],"one_liner":"The matrix that maps joint velocities to contact-point velocities, and contact forces back into joint torques.","explanation":"The contact Jacobian is the Jacobian matrix specialized to a contact point, usually written J_c(q). It's used in two directions. First, v_c = J_c(q)·q̇ gives the contact-point velocity from the joint (including floating-base) velocity q̇; a foot that isn't slipping means J_c·q̇ = 0. Second, its transpose converts a contact force λ into the corresponding generalized forces at each joint, appearing in the dynamics equation M(q)·q̈ + C(q,q̇)·q̇ = τ_g(q) + B·u + J_cᵀ·λ (M is the mass matrix, C the Coriolis and centrifugal terms, τ_g the gravity term, and B·u the motor input). Physics engines use a more general version of the same idea, the constraint Jacobian — the MuJoCo documentation describes it as mapping joint-space motion into constraint space. Both whole-body control and simulated contact solving depend on it.","example":"During a quadruped's stance phase, the controller first plans the ground-reaction force each supporting foot should provide, then multiplies it by the transpose of that leg's contact Jacobian to get the torques the hip and knee joints should output — a common approximation that ignores the leg's own mass.","related":["Jacobian Matrix","Contact Force","Ground Reaction Force (GRF)","Whole-Body Control","Contact Force Optimization (Force Distribution)","Floating Base"]},{"id":"grasp-matrix","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Murray, Li & Sastry, A Mathematical Introduction to Robotic Manipulation（第 5 章 The grasp map）","url":"https://www.cds.caltech.edu/~murray/books/MLS/pdf/mls94-complete.pdf"}],"as_of":"","related_ids":["force-closure","form-closure","wrench","friction-cone","contact-jacobian","grasp-quality-metric"],"name":"Grasp Matrix","alt":"抓取矩阵","abbr":"G","aliases":["G","Grasp Map"],"one_liner":"The linear map that combines each contact point's force into the total force and torque acting on the grasped object.","explanation":"The grasp matrix G (called the grasp map in the textbook by Murray, Li, and Sastry) describes the relationship between the forces fingers apply at their contact points and the resultant wrench on the object: F_o = G·f_c. Here f_c is the stacked vector of all contact forces, and F_o is the 6-dimensional wrench (3D force plus 3D torque) acting on the object. Each contact point occupies a block within G, determined by the contact's position, its surface normal, and its contact type (frictionless point contact, point contact with friction, or soft-finger contact). In the other direction, Gᵀ maps the object's velocity to the velocity of each contact point. The null space of G corresponds to internal forces: forces where several fingers squeeze against each other but contribute zero net force to the object — this is how fingers grip an object firmly without pushing it around. Force closure can be stated as: under the friction-cone constraints, G can generate a wrench in any direction; combined with the hand's Jacobian J_h (via τ = J_hᵀ f_c), the required joint torques can be worked backward from the force the object needs.","example":"Two fingers squeeze a box horizontally from opposite sides with 10 N each: the two contact forces sum to zero net force through G, so this pair of forces lies in G's null space and is a pure internal force. To actually hold the box up without dropping it, an upward friction-force component, within the friction cone, also has to be added so that G·f_c cancels gravity.","related":["Force Closure","Form Closure","Wrench","Friction Cone","Contact Jacobian","Grasp Quality Metric"]},{"id":"grasp-quality-metric","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics（预印本 PDF，12.1.7.3 节 Measuring the Quality of a Form-Closure Grasp）","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Ferrari & Canny, Planning optimal grasps（ICRA 1992）","url":"https://doi.org/10.1109/ROBOT.1992.219918"},{"title":"Mahler et al., Dex-Net 2.0（arXiv 1703.09312）","url":"https://arxiv.org/abs/1703.09312"}],"as_of":"","related_ids":["force-closure","wrench","grasp-matrix","grasp-planning","dex-net-2-0","friction-cone"],"name":"Grasp Quality Metric","alt":"抓取质量指标","abbr":"","aliases":["ε Metric","Epsilon Quality","Ferrari-Canny Metric"],"one_liner":"A single number scoring a grasp by how well it can resist outside disturbances, used to rank candidate grasps.","explanation":"A grasp quality metric maps a set of contact points, or a hand pose, to a single number, with a higher score meaning a better grasp — used to rank large numbers of candidate grasps. The most common is the ε metric proposed by Ferrari and Canny in 1992: with contact forces limited in magnitude and constrained to lie inside their friction cones, collect the achievable wrenches (force plus torque) at each contact point and take their convex hull; ε is the radius of the largest ball centered at the origin that fits inside that hull — the smallest disturbance, in the worst direction, that the grasp can still resist. A grasp that fails to achieve force closure scores zero or negative. Because force and torque have different units, torque is usually rescaled using a characteristic length of the object. Roa and Suárez's survey sorts these metrics into ones based on contact-point locations and ones based on hand configuration. Dex-Net 2.0 computes a robust ε metric under pose and friction uncertainty to label about 6.7 million synthetic samples, used to train the grasp-quality network GQ-CNN.","example":"Gripping the same disk with three fingers: when the fingers are evenly spaced around the circumference, the wrench convex hull is roughly symmetric and its inscribed ball is large, giving a high ε; when the three fingers are bunched on one side, the hull skews to that side, some direction of push is almost unresisted, and ε is very low.","related":["Force Closure","Wrench","Grasp Matrix","Grasp Planning","Dex-Net 2.0","Friction Cone"]},{"id":"hand-synergies","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Santello, Flanders & Soechting 1998, Postural Hand Synergies for Tool Use（J Neurosci，PMC 全文）","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC6793309/"},{"title":"GraspIt! Documentation: Eigengrasps","url":"https://graspit-simulator.github.io/build/html/eigengrasps.html"},{"title":"qbrobotics: qb SoftHand Research","url":"https://qbrobotics.com/product/qb-softhand-research/"}],"as_of":"","related_ids":["dexterous-hand","underactuation","tendon-driven-actuation","motion-retargeting","grasp-taxonomy","data-glove"],"name":"Hand Synergies","alt":"手部协同","abbr":"","aliases":["Postural Synergies","Eigengrasps"],"one_liner":"The many joints of a human hand tend to move together in a few fixed combinations, describable by a handful of synergy parameters.","explanation":"The idea of hand synergies comes from neuroscience: Santello, Flanders, and Soechting had subjects imagine grasping 57 common objects in 1998, measured the 15 joint angles of the fingers and thumb, and ran principal component analysis (PCA), finding that just the first two principal components explained over 80% of the variance. In other words, although the human hand has many degrees of freedom, most grasp postures fall within a low-dimensional subspace, with the higher-order components only fine-tuning the details. Robotics has borrowed this idea: the GraspIt! simulator's eigengrasp uses a few synergy amplitudes in place of the full set of joint angles when searching for grasps. On the hardware side, the Pisa/IIT SoftHand (Catalano et al., 2014) was designed around 'adaptive synergies,' and a related product, the qb SoftHand Research, drives 19 humanlike degrees of freedom with just one motor. When learning dexterous manipulation, synergies can also be used to compress the action space.","example":"The qb SoftHand has only one motor, so the only control input is how much to close the hand; the five fingers close together in proportions set by their tendon routing, letting it grasp objects of many different shapes.","related":["Dexterous Hand","Underactuation","Tendon-Driven Actuation","Motion Retargeting","Grasp Taxonomy (Power Grasp vs. Precision Grasp / Pinch)","Data Glove"]},{"id":"finger-gaiting","category":"mechanics","sec":7,"tier":3,"sources":[{"title":"Khandate, Haas-Heger, Ciocarlie: On the Feasibility of Learning Finger-gaiting In-hand Manipulation with Intrinsic Sensing (ICRA 2022)","url":"https://arxiv.org/abs/2109.12720"}],"as_of":"","related_ids":["in-hand-manipulation","dexterous-manipulation","dexterous-hand","force-closure","dactyl","tactile-sensor"],"name":"Finger Gaiting","alt":"指步态","abbr":"","aliases":["Finger Gait","Grasp Gait"],"one_liner":"Fingers release and reposition one at a time while the rest keep a firm grip, rotating an object through a large angle.","explanation":"Finger gaiting is a form of in-hand manipulation: while a multi-fingered hand holds an object, one finger or a few fingers at a time break contact, move to a new position, and re-establish contact, while the rest of the fingers keep holding the object — much like legs taking turns stepping during walking. In-hand manipulation that never releases contact is limited by how far the finger joints can travel, so the object can only be rotated a small angle; finger gaiting continually breaks and re-forms contact instead, which in principle lets the object rotate through any angle. Model-based work in the 1990s by Leveroni and Salisbury, and by Han and Trinkle, among others, tackled this, but the instantaneous nature of making and breaking contact makes the model highly nonlinear and hard to apply on real hardware. More recent approaches mostly use reinforcement learning: OpenAI's Dactyl work learned finger gaiting, and Khandate and colleagues at Columbia University learned finger gaiting for precision fingertip grasps in 2021 using only proprioception and touch, with no external vision.","example":"A dexterous hand pinches a cube between its fingertips and rotates it continuously around a vertical axis for several full turns: while three fingers hold it steady, the fourth lifts off, moves to a new position, and lands again; then another finger takes its turn, and the cycle repeats so the object keeps rotating.","related":["In-hand Manipulation","Dexterous Manipulation","Dexterous Hand","Force Closure","Dactyl","Tactile Sensor"]},{"id":"support-polygon","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Support polygon - Wikipedia","url":"https://en.wikipedia.org/wiki/Support_polygon"},{"title":"Underactuated Robotics (Tedrake), Ch. Simple Models of Walking: CoP and ZMP","url":"https://underactuated.mit.edu/humanoids.html"}],"as_of":"","related_ids":["static-stability","zero-moment-point","center-of-pressure","center-of-mass","double-support-phase","trot-gait"],"name":"Support Polygon","alt":"支撑多边形","abbr":"","aliases":["Support Region","Base of Support","Support Triangle"],"one_liner":"The convex region enclosed by all of a robot's ground-contact points; balance criteria are measured against this boundary.","explanation":"The support polygon is the convex region on the horizontal plane enclosed by every point where a robot touches the ground. For a humanoid balanced on one foot it's roughly the outline of that foot; standing on both feet, it's the outer boundary connecting the two footprints; for a quadruped with three legs down, it forms a triangle, which is why it's also called the support triangle. It serves as the boundary for balance criteria: static stability requires the center-of-mass projection to land inside it, and dynamic walking requires the center of pressure — where the resultant ground-reaction force acts, equivalent to the zero moment point — to stay inside it, since the center of pressure can never geometrically leave the contact region. The larger the support polygon, the wider the range of center-of-mass acceleration it can tolerate, which is why slow quadruped walking keeps three legs down at all times and why humanoid feet are made large enough to matter.","example":"When a quadruped trots, only the two diagonal legs touch the ground at once, so the support polygon collapses into a line segment; the center-of-mass projection can almost never land exactly on that line, so trotting isn't statically stable — balance has to be maintained by continuously adjusting the gait instead.","related":["Static Stability","Zero Moment Point","Center of Pressure (CoP)","Center of Mass (CoM)","Double Support Phase","Trot Gait"]},{"id":"static-stability","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Legged robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Legged_robot"},{"title":"Support polygon - Wikipedia","url":"https://en.wikipedia.org/wiki/Support_polygon"}],"as_of":"","related_ids":["support-polygon","dynamic-stability","center-of-mass","zero-moment-point","quasi-static-assumption","quadruped-robot"],"name":"Static Stability","alt":"静态稳定","abbr":"","aliases":["Static Balance","Statically Stable"],"one_liner":"The robot's center of mass projects straight down inside its support polygon, so it stays balanced even standing still.","explanation":"Static stability means a robot can hold its balance on its own support whenever its velocity and acceleration are both near zero. The test is simple: project the center of mass (the robot's overall center of gravity) straight down onto the ground, and check whether that point falls inside the support polygon — the convex region enclosed by all the ground-contact points. The farther the projection sits from the boundary, the more stable the robot is; that shortest distance is often called the stability margin. The advantage is that the robot can stop at any instant without a controller correcting it in real time, which is why early multi-legged robots and slow-walking humanoids were designed around it — a quadruped walking slowly, for example, lifts only one leg at a time and always keeps the other three planted in a triangle. The cost is speed: moving fast or running requires letting the center of mass leave the support region briefly, which calls for dynamic-stability criteria like the zero moment point or the capture point instead.","example":"When a hexapod robot uses a tripod gait, it lifts three legs at a time while the other three form a triangle on the ground; as long as the center-of-mass projection stays inside that triangle, the robot could stop at any instant without falling.","related":["Support Polygon","Dynamic Stability","Center of Mass (CoM)","Zero Moment Point","Quasi-Static Assumption","Quadruped Robot"]},{"id":"dynamic-stability","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Legged robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Legged_robot"},{"title":"Zero moment point - Wikipedia","url":"https://en.wikipedia.org/wiki/Zero_moment_point"},{"title":"Underactuated Robotics (MIT) - Humanoid Robots","url":"https://underactuated.csail.mit.edu/humanoids.html"}],"as_of":"","related_ids":["static-stability","support-polygon","zero-moment-point","capture-point","inverted-pendulum-model","balance-control"],"name":"Dynamic Stability","alt":"动态稳定","abbr":"","aliases":["Dynamic Balance"],"one_liner":"Balance where the center of mass can briefly leave the support region, recovered through continued motion and the next footstep.","explanation":"Legged robots can be “stable” in two different senses. Static stability requires the center of mass's ground projection to stay inside the support polygon (the convex region enclosed by the feet on the ground) at all times, so the robot wouldn't fall even if it froze at any instant — a hexapod that shuffles slowly on three legs at a time, alternating, walks this way. Dynamic stability relaxes that requirement: the center of mass can run outside the support region briefly, as long as the next footstep and adjustments to the ground forces keep the motion from diverging into a fall — this is how humans walk and run. Marc Raibert's hopping one-legged robots in the 1980s proved that a robot could stay upright through continuous hopping alone. Common ways to assess dynamic stability include the zero moment point (ZMP — the point where the foot's reaction force produces no horizontal moment, which must stay inside the support region) and the capture point (a footstep location that would bring the robot to a stop); modern humanoid locomotion controllers trained with reinforcement learning often learn dynamic balance directly in simulation instead.","example":"When a humanoid robot walks briskly, during single-support phase its center of mass has already tipped forward past the supporting foot; the next foot lands just in time to catch the body — that's dynamic stability. If it instead shifted its center of mass directly above the support foot before every step, that would be a static-stability gait, and much slower.","related":["Static Stability","Support Polygon","Zero Moment Point","Capture Point","Inverted Pendulum Model (IPM)","Balance Control"]},{"id":"ground-reaction-force","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Ground reaction force - Wikipedia","url":"https://en.wikipedia.org/wiki/Ground_reaction_force"},{"title":"Inverse dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Inverse_dynamics"}],"as_of":"","related_ids":["contact-force","friction-cone","zero-moment-point","center-of-pressure","foot-force-sensor","centroidal-dynamics"],"name":"Ground Reaction Force (GRF)","alt":"地面反作用力","abbr":"GRF","aliases":["GRF"],"one_liner":"The force the ground pushes back onto a foot; legged robots stand, walk, and jump entirely by means of it.","explanation":"Ground reaction force is the force the ground exerts on whatever is touching it. By Newton's third law, whatever force a foot presses into the ground, the ground pushes back with an equal and opposite force. It splits into two parts: a vertical component, the supporting (normal) force, and a horizontal component, friction — pushing off to walk forward or pushing sideways to turn both draw on the horizontal part. Biomechanics researchers measure it with a force plate, combining it with motion-capture data to work out human joint torques. For a legged robot, aside from gravity, GRF is essentially the only external force that can accelerate the body's center of mass, and the robot can only adjust it indirectly, through joint torques. That's why many quadruped and humanoid controllers — MPC, whole-body control — treat each foot's GRF directly as an optimization variable, and require it to stay inside the friction cone: horizontal force can't exceed the friction coefficient times vertical force, or the foot slips.","example":"A 60 kg person standing still on both feet has a combined vertical ground reaction force roughly equal to their weight: 60 × 9.8 ≈ 588 N. At the instant of push-off during running, this force exceeds body weight.","related":["Contact Force","Friction Cone","Zero Moment Point","Center of Pressure (CoP)","Foot Force Sensor","Centroidal Dynamics"]},{"id":"center-of-pressure","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Center of pressure (terrestrial locomotion) - Wikipedia","url":"https://en.wikipedia.org/wiki/Center_of_pressure_(terrestrial_locomotion)"},{"title":"Zero-tilting moment point - Stéphane Caron","url":"https://scaron.info/robotics/zero-tilting-moment-point.html"}],"as_of":"","related_ids":["zero-moment-point","support-polygon","ground-reaction-force","foot-force-sensor","center-of-mass","balance-control"],"name":"Center of Pressure (CoP)","alt":"压力中心","abbr":"CoP","aliases":["CoP","Center of Foot Pressure"],"one_liner":"The single point where the ground's total supporting force on a foot effectively acts, used to judge balance.","explanation":"When a foot presses on the ground, the ground pushes back at every point of contact; the center of pressure is the single point where all of that distributed force can be combined into one equivalent ground reaction force. Biomechanics researchers measure it with a force plate — while standing, it drifts slightly as the body sways, and while walking, it travels from heel to toe. By definition, it can only fall inside the support polygon (the convex region enclosed by all the ground-contact points); once it reaches the edge, the foot is about to start tipping over that edge. It comes from a different source than the zero moment point (ZMP): CoP is computed from contact forces, while ZMP is derived from the whole robot's motion and acceleration — but the two coincide whenever there's a single contact surface or the robot is walking on flat ground, which is why classical humanoid walking control often uses “keep the ZMP/CoP inside the support polygon” as its stability criterion, measuring CoP in real time with a six-axis force sensor in the foot.","example":"As a standing person leans forward, their center of pressure shifts from the middle of the foot toward the toes; lean too far and, once it reaches the edge of the toes, it can't move forward any further — the only options left are to take a step or fall forward.","related":["Zero Moment Point","Support Polygon","Ground Reaction Force (GRF)","Foot Force Sensor","Center of Mass (CoM)","Balance Control"]},{"id":"zero-moment-point","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Zero moment point - Wikipedia","url":"https://en.wikipedia.org/wiki/Zero_moment_point"},{"title":"Underactuated Robotics (Tedrake), Ch. Simple Models of Walking: CoP and ZMP","url":"https://underactuated.mit.edu/humanoids.html"}],"as_of":"","related_ids":["center-of-pressure","support-polygon","linear-inverted-pendulum-model","zmp-preview-control","capture-point","bipedal-locomotion"],"name":"Zero Moment Point","alt":"零力矩点","abbr":"ZMP","aliases":["ZMP","ZMP Criterion"],"one_liner":"The point where the ground-reaction force produces no horizontal moment — the classic balance criterion for biped walking.","explanation":"The zero moment point (ZMP) was introduced by Miomir Vukobratović and Davor Juričić in 1968: it's the point at which the ground-reaction force produces no moment about the horizontal axes, which on flat ground coincides with the center of pressure (where the resultant ground-reaction force acts). The center of pressure can only ever fall within the support polygon, so the ZMP criterion requires that the planned ZMP stay inside the footprint's support region at all times — otherwise the robot tips over the edge of its foot. When the center-of-mass height is constant, p = x − (z/g)ẍ, where p is the horizontal ZMP position, x and z are the center of mass's horizontal position and height, ẍ is its horizontal acceleration, and g is gravitational acceleration. This lets engineers fix the footstep locations and a ZMP trajectory first, then work backward to the center-of-mass trajectory — the approach earlier-generation humanoids such as Honda's ASIMO used to walk; today's reinforcement-learning-based locomotion controllers generally don't compute ZMP explicitly.","example":"When the robot stands still, ẍ = 0 and the ZMP sits directly beneath the center of mass's vertical projection. When it wants to accelerate forward, the ZMP shifts toward the heel, moving further the harder it accelerates — and once it crosses the heel's edge, the toes lift off and the whole robot tips over backward.","related":["Center of Pressure (CoP)","Support Polygon","Linear Inverted Pendulum Model (LIPM)","ZMP Preview Control","Capture Point","Bipedal Locomotion"]},{"id":"point-foot-vs-flat-foot","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"LimX Dynamics TRON 1 产品页（Point-Foot / Sole / Wheeled 足端）","url":"https://www.limxdynamics.com/en/products/tron1"},{"title":"Assessing Whole-Body Operational Space Control in a Point-Foot Series Elastic Biped (arXiv 1501.02855)","url":"https://arxiv.org/abs/1501.02855"},{"title":"Wikipedia: Support polygon","url":"https://en.wikipedia.org/wiki/Support_polygon"}],"as_of":"","related_ids":["zero-moment-point","center-of-pressure","support-polygon","underactuation","limx-dynamics-tron-1","parallel-ankle-mechanism"],"name":"Point Foot vs. Flat Foot","alt":"点足 / 平足（足端形态）","abbr":"","aliases":["Point-Foot","Sole Foot","Foot Morphology"],"one_liner":"Whether a leg's foot meets the ground at a single point or across a flat sole determines whether ankle torque can hold balance.","explanation":"Foot morphology refers to how a legged robot's foot contacts the ground. A point foot is a small ball or rubber tip making roughly point contact, so the ankle can't produce a moment against the ground and the center of pressure is pinned to that single contact point; as a result, a biped with point feet is underactuated (it has fewer controllable quantities than degrees of freedom to control), cannot stand still, and must balance by continuously stepping and adjusting where it plants its feet. Most quadruped robots also use point feet, but they can still stand firmly with multiple legs down at once. A flat foot has a sole, so the center of pressure can shift within the footprint, the ankle can apply a moment, the support polygon is larger, and the robot can stand still and can also plan its gait using zero-moment-point (ZMP) methods — at the cost of needing more motors at the ankle and a heavier leg tip. LimX Dynamics' TRON 1 can quick-swap its foot end among point-foot, flat-foot, and wheeled configurations; the company describes the point foot as the simplest and easiest to control leg form.","example":"The point-footed biped Hume, from a 2015 paper, has to continuously choose new footholds and keep stepping to stay balanced; a flat-footed humanoid such as Unitree's G1, by contrast, can stand still with its feet together.","related":["Zero Moment Point","Center of Pressure (CoP)","Support Polygon","Underactuation","LimX Dynamics TRON 1","Parallel Ankle Mechanism"]},{"id":"inverted-pendulum-model","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Inverted pendulum - Wikipedia","url":"https://en.wikipedia.org/wiki/Inverted_pendulum"},{"title":"The 3D linear inverted pendulum mode: a simple modeling for a biped walking pattern generation (Kajita et al., IROS 2001)","url":"https://doi.org/10.1109/IROS.2001.973365"}],"as_of":"","related_ids":["linear-inverted-pendulum-model","zero-moment-point","capture-point","center-of-mass","spring-loaded-inverted-pendulum","balance-control"],"name":"Inverted Pendulum Model (IPM)","alt":"倒立摆模型","abbr":"IPM","aliases":["Cart-Pole"],"one_liner":"A simplified model of a robot as a pole on the ground balancing a point mass, used to analyze and control balance.","explanation":"An inverted pendulum is a pendulum with its center of mass above its pivot — an unstable equilibrium that tips over at the slightest disturbance, and must be actively controlled, moving the pivot back underneath the center of mass, to stay upright. A standing person is roughly an inverted pendulum pivoting at the feet, and a Segway balances on exactly the same principle. In bipedal and humanoid research, lumping the robot's entire mass into a single point at the center of mass and treating the legs as massless struts gives the inverted pendulum model, which describes which way the body is tipping and where the foot should land using very few equations. In 2001, Shuuji Kajita and colleagues added the further assumption that the center-of-mass height z_c stays constant, giving the linear inverted pendulum model (LIPM): in the horizontal direction, ẍ = (g/z_c)·x, where x is the center of mass's horizontal offset from the pivot and g is gravitational acceleration. Making the equation linear makes real-time gait planning tractable, and methods like ZMP preview control and the capture point are both built on top of it.","example":"The CartPole environment commonly used to introduce reinforcement learning is exactly a cart-pole inverted pendulum: the agent can only push the cart left or right, and the goal is to keep the pole balanced upright.","related":["Linear Inverted Pendulum Model (LIPM)","Zero Moment Point","Capture Point","Center of Mass (CoM)","Spring-Loaded Inverted Pendulum","Balance Control"]},{"id":"linear-inverted-pendulum-model","category":"mechanics","sec":8,"tier":2,"sources":[{"title":"Kajita et al., The 3D Linear Inverted Pendulum Mode: A Simple Modeling for a Biped Walking Pattern Generation (IROS 2001)","url":"https://doi.org/10.1109/IROS.2001.973365"},{"title":"Kajita & Tani, Study of Dynamic Biped Locomotion on Rugged Terrain: Derivation and Application of the Linear Inverted Pendulum Mode (ICRA 1991)","url":"https://doi.org/10.1109/ROBOT.1991.131811"},{"title":"Underactuated Robotics (MIT, Tedrake): Simple Models of Walking and Running / Humanoids","url":"https://underactuated.mit.edu/humanoids.html"}],"as_of":"","related_ids":["inverted-pendulum-model","zero-moment-point","capture-point","cart-table-model","zmp-preview-control","center-of-mass"],"name":"Linear Inverted Pendulum Model (LIPM)","alt":"线性倒立摆模型","abbr":"LIPM","aliases":["LIP","3D-LIPM"],"one_liner":"A simplified model of bipedal walking that holds the center-of-mass height constant, used for fast gait planning.","explanation":"The linear inverted pendulum model is the most widely used simplification for bipedal walking, introduced by Shuuji Kajita and colleagues — a planar version at ICRA in 1991, and a 3D version, 3D-LIPM, at IROS in 2001. It represents the robot as a point mass concentrated at the center of mass, attached to a massless leg, and assumes the center-of-mass height stays constant; under that assumption, the otherwise nonlinear inverted-pendulum equation becomes linear: ẍ = (g/z_c)(x − p). Here x is the center of mass's horizontal position, p is the support point's position (the center of pressure, i.e., the ZMP), g is gravitational acceleration, and z_c is the center-of-mass height. The linear equation has a closed-form solution and is fast to compute, which suits real-time gait generation and ZMP preview control, and the capture point is also derived from it. The tradeoff is that it ignores leg mass and upper-body rotation.","example":"With center-of-mass height z_c = 0.8 m, g/z_c ≈ 12.3 s⁻². If the center of mass is 5 cm ahead of the support point, horizontal acceleration is about 0.6 m/s² and keeps increasing as the robot tips further, so the robot must take its next step in time to move the support point forward.","related":["Inverted Pendulum Model (IPM)","Zero Moment Point","Capture Point","Cart-Table Model","ZMP Preview Control","Center of Mass (CoM)"]},{"id":"cart-table-model","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Kajita et al.: Biped walking pattern generation by using preview control of zero-moment point (ICRA 2003)","url":"https://doi.org/10.1109/robot.2003.1241826"},{"title":"Stéphane Caron: Linear inverted pendulum model","url":"https://scaron.info/robotics/linear-inverted-pendulum-model.html"}],"as_of":"","related_ids":["zero-moment-point","linear-inverted-pendulum-model","zmp-preview-control","bipedal-locomotion","support-polygon","center-of-mass"],"name":"Cart-Table Model","alt":"小车-桌子模型","abbr":"","aliases":["Cart-on-a-Table Model","Table-Cart Model"],"one_liner":"Models a biped as a cart rolling on a tabletop, used to connect center-of-mass motion to the zero moment point.","explanation":"The cart-table model was introduced by Shuuji Kajita and colleagues in a 2003 ICRA paper for generating walking patterns via ZMP preview control. It concentrates the robot's mass into a cart running along a massless, horizontal tabletop supported by a single thin leg, with the bottom of that leg standing in for the sole of the foot. The zero moment point — where the horizontal moment from the ground-reaction force vanishes — satisfies p = x − (z_c/g)·ẍ: x is the cart's (i.e. the center of mass's) horizontal position, z_c is the center-of-mass height, g is gravitational acceleration, and ẍ is horizontal acceleration. Accelerate the cart too hard and the zero moment point runs off the bottom of the table leg, tipping the table over. It is mathematically the same set of equations as the linear inverted pendulum, but written the other way around — 'given a reference ZMP, find the center-of-mass trajectory' — which makes it convenient for preview control that looks ahead at future reference values.","example":"In their paper, Kajita and colleagues generated walking trajectories using the cart-table model combined with preview control, then used a full multi-body model to compensate for the ZMP error introduced by the simplification, enabling a simulated biped robot to climb a spiral staircase.","related":["Zero Moment Point","Linear Inverted Pendulum Model (LIPM)","ZMP Preview Control","Bipedal Locomotion","Support Polygon","Center of Mass (CoM)"]},{"id":"capture-point","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Pratt et al.: Capture Point: A Step toward Humanoid Push Recovery (Humanoids 2006)","url":"https://doi.org/10.1109/ichr.2006.321385"},{"title":"Stéphane Caron: Capture point","url":"https://scaron.info/robotics/capture-point.html"}],"as_of":"","related_ids":["linear-inverted-pendulum-model","divergent-component-of-motion","zero-moment-point","push-recovery","support-polygon","centroidal-moment-pivot"],"name":"Capture Point","alt":"捕获点","abbr":"CP","aliases":["CP","Instantaneous Capture Point","ICP"],"one_liner":"The point on the ground where, if the robot steps right now, it would come to a complete stop.","explanation":"The capture point was introduced by Pratt, Carff, Drakunov, and Goswami in a 2006 Humanoids conference paper to answer the question of where to step, after being pushed, in order to stop. In the linear inverted pendulum model — which treats the robot as a point mass at constant height, balanced on a single support point — the capture point is ξ = x + ẋ/ω: x is the center of mass's horizontal position, ẋ is its horizontal velocity, and ω = √(g/h), with g the gravitational acceleration and h the center-of-mass height. It is precisely the part of the motion that diverges in this model: placing the zero moment point at the capture point brings the center of mass gradually to rest, while a capture point that falls outside the foot means the robot must take a step. Later work developed it into biped walking controllers and into the broader framework of 'capturability' analysis.","example":"Suppose the center-of-mass height is 0.9 m, so ω ≈ 3.3 /s; if the robot is pushed and its center of mass moves forward at 0.5 m/s, the capture point lands roughly 0.15 m ahead of the center of mass's ground projection. If that point already lies beyond the supporting foot, the robot has to step forward and place its foot near the capture point.","related":["Linear Inverted Pendulum Model (LIPM)","Divergent Component of Motion","Zero Moment Point","Push Recovery","Support Polygon","Centroidal Moment Pivot"]},{"id":"divergent-component-of-motion","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Takenaka et al., Real time motion generation and control for biped robot – 1st report (IROS 2009)","url":"https://doi.org/10.1109/IROS.2009.5354662"},{"title":"Englsberger, Ott, Albu-Schäffer: Three-Dimensional Bipedal Walking Control Based on Divergent Component of Motion (IEEE T-RO 2015)","url":"https://doi.org/10.1109/TRO.2015.2405592"},{"title":"Stéphane Caron: Capture point","url":"https://scaron.info/robotics/capture-point.html"}],"as_of":"","related_ids":["capture-point","linear-inverted-pendulum-model","zero-moment-point","centroidal-moment-pivot","bipedal-locomotion","balance-control"],"name":"Divergent Component of Motion","alt":"发散运动分量","abbr":"DCM","aliases":["DCM","Capture Point (equivalent)"],"one_liner":"The part of a biped's center-of-mass motion that diverges exponentially; controlling it is enough to keep the robot balanced.","explanation":"The divergent component of motion is a state variable used in biped walking control: ξ = x + ẋ/ω, where x is the center-of-mass position, ẋ is its velocity, and ω = √(g/z₀), with g gravitational acceleration and z₀ the center-of-mass height. Honda's Takenaka and colleagues coined this name in their 2009 work on real-time gait generation. In the linear inverted pendulum model — which simplifies the robot to an inverted pendulum with constant center-of-mass height — the motion splits into two parts: the center of mass automatically converges toward ξ, while ξ itself moves exponentially away from the zero moment point (where the resultant force under the foot acts). So it's enough to plan and control ξ alone; the center of mass will naturally follow. It is the same point as the capture point introduced by Pratt and colleagues in 2006. Researchers at the German Aerospace Center (DLR), led by Englsberger, extended it to three dimensions between 2013 and 2015, introducing the two auxiliary points eCMP and VRP along the way.","example":"With the center of mass at a height of 0.8 m, ω = √(9.8/0.8) ≈ 3.5 s⁻¹; if the center of mass is directly above the support point and moving forward at 0.35 m/s, ξ sits 0.35/3.5 = 0.1 m ahead of it. If the front edge of the foot can't reach that far, the robot has to take a step.","related":["Capture Point","Linear Inverted Pendulum Model (LIPM)","Zero Moment Point","Centroidal Moment Pivot","Bipedal Locomotion","Balance Control"]},{"id":"centroidal-moment-pivot","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Popovic, Goswami, Herr: Ground Reference Points in Legged Locomotion（IntechOpen 开放获取版）","url":"https://www.intechopen.com/chapters/50"},{"title":"Popovic, Goswami, Herr: Ground Reference Points in Legged Locomotion (IJRR, 2005)","url":"https://doi.org/10.1177/0278364905058363"}],"as_of":"","related_ids":["zero-moment-point","center-of-pressure","angular-momentum","capture-point","centroidal-dynamics","push-recovery"],"name":"Centroidal Moment Pivot","alt":"质心力矩枢轴点","abbr":"CMP","aliases":["CMP"],"one_liner":"The point where a line through the center of mass, parallel to the ground-reaction force, meets the ground.","explanation":"The centroidal moment pivot was introduced by Herr, Hofmann, and Popovic around 2003–2004 (Goswami and colleagues independently proposed the same point), and a 2005 IJRR paper unified the definition: draw a line through the center of mass parallel to the ground-reaction force, and its intersection with the ground is the CMP. If the ground-reaction force happens to pass exactly through the center of mass — producing zero moment about it — the CMP coincides with the zero moment point; whenever the two separate, it means the body is experiencing a moment about its center of mass and its overall angular momentum is changing, with the separation distance equal to the moment's horizontal component divided by the vertical component of the ground-reaction force. Unlike the zero moment point, the CMP is not restricted to stay inside the support region, which makes it useful for describing balance strategies — like swinging an arm or twisting the torso — that rely on angular momentum.","example":"Popovic and colleagues measured normal human walking on level ground and found the CMP never left the support region, staying only about 14% of a foot length away from the zero moment point on average — evidence that humans keep the angular momentum about their center of mass tightly controlled while walking.","related":["Zero Moment Point","Center of Pressure (CoP)","Angular Momentum","Capture Point","Centroidal Dynamics","Push Recovery"]},{"id":"centroidal-dynamics","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Orin, Goswami, Lee: Centroidal dynamics of a humanoid robot (Autonomous Robots, 2013)","url":"https://doi.org/10.1007/s10514-013-9341-4"},{"title":"Underactuated Robotics（Tedrake）: Highly-articulated Legged Robots 一章","url":"https://underactuated.mit.edu/humanoids.html"}],"as_of":"","related_ids":["centroidal-momentum-matrix","single-rigid-body-dynamics-model","angular-momentum","whole-body-control","model-predictive-control","contact-force"],"name":"Centroidal Dynamics","alt":"质心动力学","abbr":"","aliases":["Centroidal Momentum Dynamics"],"one_liner":"Dynamics that track only how a robot's center of mass and overall momentum change under external forces.","explanation":"Orin, Goswami, and Lee gave a systematic treatment of this concept in a 2013 Autonomous Robots paper: convert every link's momentum to the whole robot's center of mass and sum them, yielding a 6-dimensional centroidal momentum — 3D linear momentum plus 3D angular momentum about the center of mass. Its rate of change depends only on external forces: the change in linear momentum equals the sum of all contact forces plus gravity, and the change in angular momentum equals the sum of the moments those contact forces produce about the center of mass. These equations are exact for the whole robot, not an approximation — they simply don't encode limits like maximum joint torque or how far a leg can reach. Because the representation is low-dimensional, planning and model-predictive control for humanoids and quadrupeds often first solve for contact forces and the center-of-mass trajectory in this model, then hand the result to a whole-body controller to realize at each joint.","example":"Planning a vertical jump: first use the centroidal dynamics model to compute how much ground-reaction force each foot must push off with and how the center of mass and angular momentum should evolve, then have the whole-body controller convert those targets into torques for each joint.","related":["Centroidal Momentum Matrix","Single Rigid Body Dynamics Model","Angular Momentum","Whole-Body Control","Model Predictive Control","Contact Force"]},{"id":"centroidal-momentum-matrix","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Orin & Goswami: Centroidal Momentum Matrix of a humanoid robot: Structure and properties (IROS 2008)","url":"https://doi.org/10.1109/iros.2008.4650772"},{"title":"Underactuated Robotics（Tedrake）: Highly-articulated Legged Robots 一章","url":"https://underactuated.mit.edu/humanoids.html"}],"as_of":"","related_ids":["centroidal-dynamics","angular-momentum","jacobian-matrix","mass-matrix","whole-body-control","floating-base"],"name":"Centroidal Momentum Matrix","alt":"质心动量矩阵","abbr":"CMM","aliases":["CMM","A_G Matrix"],"one_liner":"The matrix that maps a robot's full-body joint velocities into its center-of-mass linear and angular momentum.","explanation":"Orin and Goswami analyzed its structure in detail in a 2008 IROS paper: a robot's centroidal momentum h_G (linear momentum plus angular momentum about the center of mass, 6 dimensions in total) is a linear function of the generalized velocity q̇, written h_G = A_G(q)·q̇, where A_G is the centroidal momentum matrix and q is the full set of generalized coordinates, including the floating base. It had previously been called both a Jacobian and an inertia matrix in different contexts; the paper points out that it is in fact the product of the two. Differentiating h_G gives ḣ_G = A_G·q̈ + Ȧ_G·q̇, which lets whole-body control express a desired change in momentum as a linear constraint on joint accelerations, ready to drop into a quadratic program. Pinocchio computes it with the ccrba function.","example":"A humanoid robot with 30 joints plus the 6 degrees of freedom of the floating base has a 36-dimensional q̇, making A_G a 6×36 matrix; a balance controller requires A_G·q̈ + Ȧ_G·q̇ to equal the desired rate of change of momentum, solved together with other tasks in a quadratic program.","related":["Centroidal Dynamics","Angular Momentum","Jacobian Matrix","Mass Matrix","Whole-Body Control","Floating Base"]},{"id":"single-rigid-body-dynamics-model","category":"mechanics","sec":8,"tier":3,"sources":[{"title":"Di Carlo et al., Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (IROS 2018)","url":"https://www.semanticscholar.org/paper/608d53fcd69173d30914e29d9b8ca4b37efe9ac4"},{"title":"Lin et al., Learning Near-global-optimal Strategies for Hybrid Non-convex MPC of Single Rigid Body Locomotion (arXiv:2207.07846)","url":"https://arxiv.org/abs/2207.07846"},{"title":"Kim et al., Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"}],"as_of":"","related_ids":["convex-mpc","model-predictive-control","centroidal-dynamics","ground-reaction-force","whole-body-control","linear-inverted-pendulum-model"],"name":"Single Rigid Body Dynamics Model","alt":"单刚体动力学模型","abbr":"SRBD","aliases":["SRBD","Single Rigid Body Model","SRBM"],"one_liner":"A simplified dynamics model that treats a whole legged robot as one rigid body and ignores the mass of its legs.","explanation":"The single rigid body dynamics model is one of the most widely used simplifications in legged-robot control: it assumes the legs are light and treats the whole robot as a single rigid body with mass and rotational inertia, subject only to gravity and the ground-reaction forces at its feet. The state is typically the torso's position, orientation, linear velocity, and angular velocity (12 dimensions total), and the control inputs are the contact force at each foot. A full model has dozens of joints and highly nonlinear equations; SRBD compresses this down to a handful of equations, making it practical for real-time optimization. The landmark example is MIT's convex MPC, developed by Di Carlo and colleagues in 2018 on Cheetah 3: assuming small orientation angles as well, the ground-reaction-force planning problem becomes a convex optimization that can be solved in under 1 millisecond, replanned repeatedly at 20–30 Hz over a horizon of up to 0.5 seconds; the resulting forces are then handed to whole-body or joint-level control to execute. It accounts for more than the inverted pendulum model — orientation and rotation — but still doesn't track how the legs themselves move.","example":"MIT Cheetah 3 used single-rigid-body convex MPC to achieve standing, trotting, flying trot, pronking, bounding, pacing, a three-legged gait, and a 3D gallop, reaching forward speeds up to 3 m/s, all with the same set of gains and weights (Di Carlo et al., IROS 2018).","related":["Convex MPC","Model Predictive Control","Centroidal Dynamics","Ground Reaction Force (GRF)","Whole-Body Control","Linear Inverted Pendulum Model (LIPM)"]},{"id":"stance-phase-swing-phase","category":"mechanics","sec":9,"tier":2,"sources":[{"title":"Wikipedia: Gait (human)","url":"https://en.wikipedia.org/wiki/Gait_(human)"},{"title":"Humanoid-Gym: humanoid_env.py (_get_gait_phase)","url":"https://github.com/roboterax/humanoid-gym/blob/main/humanoid/envs/custom/humanoid_env.py"}],"as_of":"","related_ids":["gait","gait-cycle-and-duty-factor","double-support-phase","flight-phase","gait-planning","swing-foot-trajectory-planning"],"name":"Stance Phase / Swing Phase","alt":"支撑相 / 摆动相","abbr":"","aliases":["Stance","Swing"],"one_liner":"The gait phase where a foot is on the ground bearing weight is stance; the phase where it swings through the air is swing.","explanation":"Stance phase and swing phase are the basic building blocks for describing gaits, applicable to both people and legged robots: the phase where a leg's foot is in contact with the ground, bearing body weight and driving the body forward, is stance phase; the phase where the foot leaves the ground and swings forward toward its next landing spot is swing phase, and together the two make up one gait cycle for that leg. Normal adult walking spends roughly 60% of the cycle in stance and 40% in swing, with the period when both feet are down at once called double support; running adds a flight phase where both feet leave the ground together. The fraction of the cycle spent in stance is called the duty factor, and different gaits — such as a quadruped's diagonal trot — are, at bottom, just different ways of staggering each leg's stance phase relative to the others. Control needs differ sharply between the two: a stance leg has to generate ground reaction force, maintain balance, and avoid slipping, while a swing leg has to lift clear of obstacles and choose where to land — so both MPC and RL-based locomotion controllers need to know which phase each leg is currently in.","example":"Humanoid-Gym generates each foot's stance mask with a sinusoidal phase clock: a positive sine value means the left foot is in stance, a negative value means the right foot is in stance, and a value near zero counts both feet as in stance (double support). The reward function then requires whichever foot should be down to be down, and whichever should be swinging to be off the ground.","related":["Gait","Gait Cycle and Duty Factor","Double Support Phase","Flight Phase (Aerial Phase)","Gait Planning","Swing Foot Trajectory Planning"]},{"id":"flight-phase","category":"mechanics","sec":9,"tier":2,"sources":[{"title":"Running - Wikipedia","url":"https://en.wikipedia.org/wiki/Running"},{"title":"Legged robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Legged_robot"}],"as_of":"","related_ids":["stance-phase-swing-phase","gait","gait-cycle-and-duty-factor","spring-loaded-inverted-pendulum","bound-gait","gallop-gait"],"name":"Flight Phase (Aerial Phase)","alt":"腾空相","abbr":"","aliases":["Aerial Phase","Flight Period"],"one_liner":"The part of a gait where every foot is off the ground at once and the body is airborne.","explanation":"Legged locomotion breaks down over time into phases: a foot bearing weight on the ground is in stance phase, a foot swinging through the air is in swing phase, and if there's a stretch where no foot touches the ground at all, that's the flight phase. It's the key distinction between walking and running: in walking, at least one foot is on the ground at all times, while running includes a flight phase. For a robot, gravity is the only force acting during flight, so the center of mass follows a parabolic arc, and the controller can no longer use ground forces to change the robot's overall momentum — it can only adjust leg posture to prepare for landing, and must then absorb an impact the instant it touches down, which is why running, jumping, and parkour are so much harder than walking. Raibert's hopping robots split each cycle into a stance phase and a flight phase with separate control, choosing a foot placement by swinging the leg during flight; quadruped bounding and galloping gaits, and humanoid running and flips, all include a flight phase.","example":"In a humanoid robot's jog, the brief moment after one foot pushes off the ground and before the other foot lands is the flight phase — the only thing the policy can do during it is adjust the swinging leg's posture to prepare for the next landing.","related":["Stance Phase / Swing Phase","Gait","Gait Cycle and Duty Factor","Spring-Loaded Inverted Pendulum","Bound Gait","Gallop Gait"]},{"id":"gait-cycle-and-duty-factor","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Gait","url":"https://en.wikipedia.org/wiki/Gait"}],"as_of":"","related_ids":["gait","stance-phase-swing-phase","trot-gait","gait-planning","gait-phase","flight-phase"],"name":"Gait Cycle and Duty Factor","alt":"步态周期与占空比","abbr":"","aliases":["Stride","Duty Factor"],"one_liner":"The gait cycle is one full stride from touchdown to the next touchdown; the duty factor is the fraction spent on the ground.","explanation":"The gait cycle, or stride, is the complete cycle of a single foot from one touchdown to the next, split into a stance phase (foot on the ground, bearing weight and propelling) and a swing phase (foot lifted and moving forward). The duty factor β = stance time ÷ cycle time. Biomechanics commonly uses 50% as the dividing line: a duty factor above 50% counts as walking, where at least one foot is always on the ground; below 50% counts as running, which includes an aerial phase. For quadruped and humanoid robots, a periodic gait can be described with three quantities: the cycle period, each leg's duty factor, and the phase offsets between legs — a diagonal trot, for instance, has the two diagonal leg pairs offset by half a cycle. Model-based controllers use these quantities to schedule when each leg touches down (gait scheduling), and reinforcement-learning locomotion controllers often feed a phase clock into the observation or reward to guide the policy toward a desired rhythm.","example":"A leg with a 0.5 s cycle that spends 0.3 s on the ground has a duty factor of 0.6, above 0.5, so it counts as walking; if it only touches down for 0.2 s, the duty factor is 0.4, and there's an aerial phase — that's running.","related":["Gait","Stance Phase / Swing Phase","Trot Gait","Gait Planning","Gait Phase (Phase Clock)","Flight Phase (Aerial Phase)"]},{"id":"pace-gait","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Horse gait (Pace)","url":"https://en.wikipedia.org/wiki/Horse_gait"},{"title":"Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior (arXiv 2212.03238)","url":"https://arxiv.org/html/2212.03238"}],"as_of":"","related_ids":["trot-gait","bound-gait","gallop-gait","gait-cycle-and-duty-factor","gait-phase","walk-these-ways"],"name":"Pace Gait","alt":"踱步步态","abbr":"","aliases":["Pacing","Lateral Sequence Gait"],"one_liner":"A two-beat quadruped gait where both legs on the same side step together, alternating left and right.","explanation":"The pace is a two-beat gait in quadruped animals: the left front and left hind legs step together, followed by the right front and right hind legs stepping together. It differs from the trot (paired diagonal legs) only in which legs are paired up. Horses have dedicated racing breeds bred for pacing (pacers within the Standardbred breed), and camels naturally move this way. For robots, when both legs on the same side are down, the support line falls to one side of the body and the center of mass isn't directly above it, so the torso tends to rock side to side — which is why quadruped robots more often trot in everyday use. Reinforcement-learning locomotion controllers commonly describe gaits using the phase offsets between legs: Walk These Ways (2022) uses three time-offset parameters, where (0, 0, 0.5) is pacing and (0.5, 0, 0) is trotting; in their tests on randomized rough terrain, pacing had the longest average survival time.","example":"In the Walk These Ways controller, setting the gait offset to (0, 0, 0.5) makes the robot dog step with both legs on the left together and both on the right together in turn, rocking its body from side to side as it walks.","related":["Trot Gait","Bound Gait","Gallop Gait","Gait Cycle and Duty Factor","Gait Phase (Phase Clock)","Walk These Ways"]},{"id":"bound-gait","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Yang & Bhounsule: Koopman Operator Based Linear MPC for 2D Quadruped Trotting, Bounding, and Gait Transition (arXiv 2507.14605)","url":"https://arxiv.org/abs/2507.14605"},{"title":"Park, Wensing, Kim: High-speed bounding with the MIT Cheetah 2 (IJRR, 2017), MIT DSpace","url":"https://dspace.mit.edu/handle/1721.1/119686"},{"title":"Margolis & Agrawal: Walk These Ways (arXiv 2212.03238)","url":"https://arxiv.org/abs/2212.03238"}],"as_of":"","related_ids":["gait","trot-gait","pace-gait","gallop-gait","flight-phase","quadruped-robot"],"name":"Bound Gait","alt":"跳跃步态","abbr":"","aliases":["Bounding","Bounding Gait"],"one_liner":"A running gait where a quadruped's two front legs land together, then the two rear legs land together, alternating front and back.","explanation":"Bounding is a symmetric gait seen in quadruped animals and quadruped robots. The four legs are grouped into a front pair and a rear pair: both front legs touch down and lift off together, and likewise for the rear pair, with the two pairs alternating support and usually a flight phase — where the whole body leaves the ground — in between. A full cycle runs: front-leg stance, flight, rear-leg stance, flight. It's often compared alongside the diagonal trot (paired diagonal legs), the pace (paired same-side legs), and the pronk (all four legs synchronized); the body noticeably pitches up and down under this gait. Bounding suits high-speed running — MIT Cheetah 2 used it in high-speed experiments — and reinforcement-learning locomotion controllers often include it as one of several gaits selectable by command.","example":"MIT Cheetah 2, in a 2017 IJRR paper, reached a top speed of 6.4 m/s using the bound gait in untethered 3D experiments; the Walk These Ways policy can switch among the trot, pronk, pace, and bound gaits on command.","related":["Gait","Trot Gait","Pace Gait","Gallop Gait","Flight Phase (Aerial Phase)","Quadruped Robot"]},{"id":"gallop-gait","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Horse gait（Gallop）","url":"https://en.wikipedia.org/wiki/Horse_gait"},{"title":"Wikipedia: Greyhound（double suspension rotary gallop）","url":"https://en.wikipedia.org/wiki/Greyhound"},{"title":"Kim et al. 2019, Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control","url":"https://arxiv.org/abs/1909.06586"}],"as_of":"","related_ids":["gait","bound-gait","trot-gait","flight-phase","quadruped-robot","mit-mini-cheetah"],"name":"Gallop Gait","alt":"疾驰步态","abbr":"","aliases":["Gallop","Rotary Gallop"],"one_liner":"The fastest quadruped gait: all four feet touch down in sequence, slightly offset, with a phase where the whole body is airborne.","explanation":"The gallop is a high-speed, asymmetric gait in quadruped animals: the two front legs, or the two rear legs, don't touch down at exactly the same time — they're slightly offset, producing a four-beat footfall pattern, and each cycle also includes a suspension phase where all four feet leave the ground at once. It's the horse's fastest gait, averaging roughly 40–48 km/h, and the faster the pace, the longer the suspension phase; a greyhound's double-suspension rotary gallop has two suspension phases per cycle, one with the spine curled and one extended, using large spinal flexion to lengthen its stride. It differs from the bound gait (front pair together, rear pair together) in whether the left and right legs within each pair are offset. For robots, the gallop has a long airborne time and short per-leg ground contact, and the torso's center of mass can't be directly controlled during the suspension phase, making it harder to execute than a trot. MIT's Mini Cheetah used MPC combined with whole-body impulse control to demonstrate six gaits, including the gallop.","example":"When a horse gallops leading with its right front leg, the footfall order is left-hind, right-hind, left-front, right-front, followed by all four feet leaving the ground together before the next cycle begins.","related":["Gait","Bound Gait","Trot Gait","Flight Phase (Aerial Phase)","Quadruped Robot","MIT Mini Cheetah"]},{"id":"spring-loaded-inverted-pendulum","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Truax et al., Optimizing Design and Control of Running Robots Abstracted as TD-SLIP (arXiv:2407.12120)","url":"https://arxiv.org/abs/2407.12120"},{"title":"Chen, Wensing, Zhang, Optimal Control of a Differentially Flat 2D Spring-Loaded Inverted Pendulum Model (arXiv:1911.07168)","url":"https://arxiv.org/abs/1911.07168"}],"as_of":"","related_ids":["inverted-pendulum-model","linear-inverted-pendulum-model","bipedal-locomotion","flight-phase","gait","passive-dynamic-walking"],"name":"Spring-Loaded Inverted Pendulum","alt":"弹簧负载倒立摆","abbr":"SLIP","aliases":["SLIP","Spring-Mass Model"],"one_liner":"Models running and hopping with a point mass on top of a single massless spring leg.","explanation":"The spring-loaded inverted pendulum is a classic simplified model of running and hopping: the body collapses into a single point mass, and the leg is a massless spring. During the stance phase, when the foot is on the ground, the spring first compresses and stores energy, then extends and launches the body back up; during the flight phase, when the foot is off the ground, the point mass simply undergoes projectile motion under gravity, and the two phases alternate. Biomechanics research — such as work by Blickhan and Full in 1993 — has found that the center-of-mass motion of running animals ranging from insects to humans can be described by this model. It differs from an ordinary inverted pendulum in that the leg's length can change and it can store and release elastic energy, so it can capture running and hopping with an airborne phase, whereas the linear inverted pendulum better describes walking. In robotics, SLIP is often used as a 'template': footfall angle, leg stiffness, and other parameters are first planned on this low-dimensional model and then mapped onto the real robot; legged robots such as RHex were designed with reference to SLIP dynamics.","example":"When a person runs, the knee and ankle bend and the body's center of mass dips after the foot lands (equivalent to the spring compressing), then extends to launch the body into the air. SLIP reproduces this rise-and-fall trajectory of the center of mass using just a handful of parameters: mass, leg length, spring stiffness, and touchdown angle.","related":["Inverted Pendulum Model (IPM)","Linear Inverted Pendulum Model (LIPM)","Bipedal Locomotion","Flight Phase (Aerial Phase)","Gait","Passive Dynamic Walking"]},{"id":"straight-knee-walking","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Griffin et al., Straight-Leg Walking Through Underconstrained Whole-Body Control (arXiv:1709.03660)","url":"https://arxiv.org/abs/1709.03660"},{"title":"Fasano et al., Efficient, Dynamic Locomotion through Step Placement with Straight Legs and Rolling Contacts (arXiv:2310.13134)","url":"https://arxiv.org/abs/2310.13134"}],"as_of":"","related_ids":["bipedal-locomotion","linear-inverted-pendulum-model","singular-configuration","zero-moment-point","whole-body-control","human-like-gait"],"name":"Straight-Knee Walking","alt":"直膝行走","abbr":"","aliases":["Straight-Leg Walking","Extended-Knee Gait"],"one_liner":"A humanoid walks with its supporting leg nearly straight, rather than staying bent-kneed in a crouch the whole time.","explanation":"Straight-knee walking means a biped keeps its supporting knee nearly extended during stance, combined with heel-strike and toe-off, more closely resembling how humans walk; the alternative, common in many humanoid robots, is bent-knee walking, where the robot stays in a half-crouch the entire time. Griffin and colleagues at IHMC summarized three reasons robots default to bent knees in a 2017 paper: the commonly used linear inverted pendulum model assumes constant center-of-mass height, which requires the knee to bend to absorb height changes; once a leg is straight, the knee joint can barely adjust the ground-reaction force any further; and a fully extended leg sits exactly at a kinematic singularity, where the Jacobian matrix loses rank, causing problems for controllers based on inverse kinematics or inverse dynamics. The cost of bent-knee walking is that the knee joint carries large torque continuously, consuming more power, and ground clearance is also smaller. Achieving straight-knee walking requires a controller that allows the center of mass to rise and fall and that handles the singularity properly; IHMC has verified this on two humanoids, Atlas and Nadia.","example":"Watching video of a humanoid walking: if the knees stay bent throughout and the body height barely changes, that's typical bent-knee walking; if the supporting leg is nearly straight, the torso rises and falls slightly with each step, and the heel strikes first and rolls through to the toe, that's closer to straight-knee walking.","related":["Bipedal Locomotion","Linear Inverted Pendulum Model (LIPM)","Singular Configuration (Kinematic Singularity)","Zero Moment Point","Whole-Body Control","Human-like Gait (Straight-knee, Heel-to-toe Walking)"]},{"id":"passive-dynamic-walking","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Passive dynamics（含 McGeer 1990 与 Collins et al. 2005 Science 引用）","url":"https://en.wikipedia.org/wiki/Passive_dynamics"}],"as_of":"","related_ids":["bipedal-locomotion","limit-cycle","cost-of-transport","morphological-computation","inverted-pendulum-model","honda-asimo"],"name":"Passive Dynamic Walking","alt":"被动动力学行走","abbr":"","aliases":["Passive Walking","Passive Dynamic Walker"],"one_liner":"A style of biped walking with no motors, powered purely by gravity and the legs' natural swing down a gentle slope.","explanation":"Passive dynamic walking was introduced by Tad McGeer in the late 1980s, with the representative paper 'Passive Dynamic Walking' published in IJRR in 1990: a two-legged mechanism with no motors and no controller, placed on a shallow slope, can walk with a stable periodic gait powered by gravity alone, with each leg swinging naturally like a pendulum. The energy lost with each step is exactly repaid by the downhill slope, and the gait converges to a limit cycle (a stable periodic motion). It demonstrates that much of walking's stability comes from the mechanism's own dynamics, without needing every joint to be precisely controlled. In 2005, Collins and colleagues showed in Science a robot built on this principle, with only a small amount of added actuation, that could walk on level ground with a cost of transport of about 0.2 — close to a human's — versus about 3.2 for ASIMO. It's a classic example for understanding low-energy gaits, limit-cycle stability, and morphological computation.","example":"A McGeer-style passive walker: two legs with knees, no motors at all, given a light push on a very shallow slope, will walk itself step by step down to the bottom.","related":["Bipedal Locomotion","Limit Cycle","Cost of Transport","Morphological Computation","Inverted Pendulum Model (IPM)","Honda ASIMO"]},{"id":"limit-cycle","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Limit cycle","url":"https://en.wikipedia.org/wiki/Limit_cycle"},{"title":"Tedrake, Underactuated Robotics: Simple Models of Walking and Running","url":"https://underactuated.mit.edu/simple_legs.html"}],"as_of":"","related_ids":["passive-dynamic-walking","gait","central-pattern-generator","hybrid-zero-dynamics","lyapunov-stability","gait-phase"],"name":"Limit Cycle","alt":"极限环","abbr":"","aliases":["Stable Periodic Orbit","Attracting Periodic Orbit"],"one_liner":"An isolated closed periodic trajectory in a nonlinear system that nearby states get pulled toward and circle around.","explanation":"A limit cycle is a concept from nonlinear dynamics, with the systematic study of it starting with Poincaré: in a state space (such as an angle-versus-angular-velocity plane), it's a closed periodic trajectory that nearby trajectories spiral toward over time (a stable limit cycle) or spiral away from (an unstable one). A stable limit cycle means the system spontaneously settles into an oscillation at a fixed rhythm, and returns to that same rhythm even after a disturbance; the classic example is the Van der Pol oscillator. Walking-robot research treats a stable periodic gait as a stable limit cycle: McGeer's passive-dynamic walking robots have no motors at all, and walk stably down a shallow slope powered only by gravity. Analysis typically takes the state at each foot-touchdown instant to build a Poincaré map (a return map), turning the question of whether a gait is stable into whether this discrete map has a stable fixed point.","example":"An unpowered 'rimless wheel' (a wheel with spokes but no rim) placed on a slope will, across a fairly wide range of starting speeds, quickly converge to the same stable rolling rhythm — a stable limit cycle, and a classic example from Tedrake's Underactuated Robotics.","related":["Passive Dynamic Walking","Gait","Central Pattern Generator","Hybrid Zero Dynamics","Lyapunov Stability","Gait Phase (Phase Clock)"]},{"id":"cost-of-transport","category":"mechanics","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Cost of transport","url":"https://en.wikipedia.org/wiki/Cost_of_transport"},{"title":"Wikipedia: Passive dynamics（Cornell biped 0.20、ASIMO 3.23）","url":"https://en.wikipedia.org/wiki/Passive_dynamics"}],"as_of":"","related_ids":["passive-dynamic-walking","bipedal-locomotion","legged-locomotion","gait","battery-runtime"],"name":"Cost of Transport","alt":"运输成本","abbr":"CoT","aliases":["CoT","Specific Cost of Transport","Specific Resistance"],"one_liner":"The energy used to move a unit of weight a unit of distance — a measure of how efficient walking or running is.","explanation":"Cost of transport is a dimensionless measure of locomotion efficiency: CoT = E/(mgd) = P/(mgv), where E is energy consumed, m is mass, g is gravitational acceleration, d is distance traveled, P is power, and v is speed. Lower is more efficient, and because body weight and distance cancel out, the number can be compared across animals, vehicles, and robots on equal footing. Collins, Ruina, Tedrake, and Wisse reported in a 2005 Science paper that the Cornell biped, built around passive-dynamic walking, achieved a cost of transport of about 0.2 — comparable to human walking — while Honda's ASIMO was around 3.2. For legged and humanoid robots this number directly determines battery range. When comparing figures, it matters whether E is measured as motor mechanical work, total battery electrical energy, or human metabolic energy — these give very different numbers for the same task.","example":"A commonly cited example: a 70 kg person walking at 1 m/s has a metabolic power of about 231 W, giving CoT ≈ 231/(70×9.8×1) ≈ 0.34.","related":["Passive Dynamic Walking","Bipedal Locomotion","Legged Locomotion","Gait","Battery Runtime"]},{"id":"locomotion-control","category":"control","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Motion control","url":"https://en.wikipedia.org/wiki/Motion_control"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv:2109.11978)","url":"https://arxiv.org/abs/2109.11978"},{"title":"qiayuanl/legged_control（NMPC + WBC 足式运控框架）","url":"https://github.com/qiayuanl/legged_control"}],"as_of":"","related_ids":["rl-based-locomotion-control","whole-body-control","model-predictive-control","balance-control","robot-cerebellum","motion-planning"],"name":"Locomotion Control","alt":"运动控制","abbr":"","aliases":["Motion Control","Locomotion Policy"],"one_liner":"The control technology that keeps a robot moving stably as intended; in embodied AI it usually means legged walking and balance.","explanation":"Motion control is originally a general term in automation, meaning making a machine's parts move as required, with position, velocity, or force as the controlled quantity, implemented as a closed loop of controller, driver, motor, and encoder. In the context of embodied AI and humanoid robots, 'locomotion control' (a shortened term often just called 'motion control' in Chinese) more specifically refers to a legged robot's ability to walk, run, balance, and get up after a fall — the core capability of the 'cerebellum.' There are two main approaches. Model-based methods use a simplified dynamics model together with model predictive control and whole-body control, solved in real time. Learning-based methods train a policy with reinforcement learning inside massively parallel GPU simulation, then transfer it to the real robot. In recent years the learning-based approach has become increasingly common for both quadrupeds and humanoids, and is often combined with the model-based one. Locomotion control handles 'how to move stably'; motion planning handles 'which path to take' — the two divide the work differently.","example":"Rudin and colleagues trained an ANYmal quadruped's walking policy in 2021 using thousands of parallel simulated robots on a single GPU, finishing training on flat ground in under 4 minutes and on rough terrain in about 20 minutes, and then deployed it to the real robot.","related":["RL-based Locomotion Control","Whole-Body Control","Model Predictive Control","Balance Control","Robot Cerebellum","Motion Planning"]},{"id":"hierarchical-control","category":"control","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Hierarchical control system","url":"https://en.wikipedia.org/wiki/Hierarchical_control_system"},{"title":"Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"},{"title":"legged_gym: legged_robot.py（_compute_torques，PD 控制）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"}],"as_of":"","related_ids":["hierarchical-architecture","robot-cerebellum","model-predictive-control","whole-body-control","proportional-derivative-control","cascade-control"],"name":"Hierarchical Control","alt":"分层控制","abbr":"","aliases":["High-Level Controller","Low-Level Controller","Hierarchical Control System"],"one_liner":"Splits control into layers: a slow, high-level layer sets goals, and a fast, low-level layer tracks them and drives the motors.","explanation":"Hierarchical control is a classic approach in control engineering: break a complex control problem into several layers, where the upper layers see farther ahead and update more slowly, and the lower layers see closer and run faster. A higher-level controller outputs reference quantities on a longer cycle — a target velocity, footholds, contact forces, or target joint angles — and a lower-level controller tracks these references at a much higher frequency, handling disturbances and hardware details, generally without the upper layer intervening in how it's executed. The benefit is that each layer can use a simple model and solve its own problem on its own timescale, and layers can be swapped out or debugged independently. It's closely related to 'hierarchical architecture' and 'brain-cerebellum,' which more often describe how task planning and action generation are divided at the model level, while hierarchical control emphasizes the frequencies and interfaces between control loops at each level.","example":"MIT's Mini Cheetah splits control into two layers: an MPC optimizes foot contact forces over a longer time window using a simplified model, and whole-body impulse control (WBIC) converts those forces into joint commands. Reinforcement-learning locomotion works similarly: a policy outputs target joint angles at 50 Hz, and a lower-level PD controller converts them into torques at a much higher frequency.","related":["Hierarchical Architecture","Robot Cerebellum","Model Predictive Control","Whole-Body Control","Proportional-Derivative Control","Cascade Control"]},{"id":"robot-cerebellum","category":"control","sec":0,"tier":1,"sources":[{"title":"RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration (arXiv:2505.03673)","url":"https://arxiv.org/abs/2505.03673"},{"title":"Figure: Helix — A Vision-Language-Action Model for Generalist Humanoid Control","url":"https://www.figure.ai/news/helix"}],"as_of":"2025-05","related_ids":["braincerebellum-architecture","dual-system-architecture","hierarchical-control","locomotion-control","whole-body-control","robobrain"],"name":"Robot Cerebellum","alt":"大脑-小脑架构（小脑）","abbr":"","aliases":["Cerebellum","Motion Cerebellum","Cerebellum Skill Library"],"one_liner":"The cerebellum in the brain-cerebellum split: the layer that turns high-level commands into fast, stable joint motions.","explanation":"Brain-cerebellum is an informal way China's embodied-AI industry describes a two-layer system, and this entry covers the cerebellum half. The brain is usually a multimodal large model that makes only a few decisions per second, responsible for understanding instructions and breaking down tasks; the cerebellum takes a sub-task like 'walk forward half a meter' or 'pick up the cup' and outputs joint commands tens to hundreds of times per second, while also holding balance and rejecting disturbances. It can be a reinforcement-learning-trained locomotion policy, a traditional controller such as MPC plus whole-body control, or a fast action module inside a skill library or a VLA. Companies define the boundary differently: narrowly, it covers only motion control like walking and balance; broadly, it also includes manipulation skills like grasping. BAAI's RoboOS calls its pluggable skill library the 'cerebellum skill library.'","example":"Figure's Helix isn't labeled a 'cerebellum,' but it maps onto the same structure: System 2, which understands the scene and language, runs at 7–9 Hz, while System 1, which produces actions, runs as a separate real-time process outputting continuous commands for 35 degrees of freedom at 200 Hz.","related":["Brain–Cerebellum Architecture","Dual-System Architecture (System 1 / System 2)","Hierarchical Control","Locomotion Control","Whole-Body Control","RoboBrain"]},{"id":"control-frequency","category":"control","sec":0,"tier":1,"sources":[{"title":"libfranka robot.h（Since the robot is controlled with a 1 kHz frequency…）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"},{"title":"Franka FCI Docs: Minimum system and network requirements（RTT + 控制回路 + 机器人处理 < 1 ms；连续丢 20 包即停机）","url":"https://frankarobotics.github.io/docs/doc/libfranka/docs/system_requirements.html"},{"title":"legged_gym: legged_robot_config.py（sim dt = 0.005, decimation = 4）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"},{"title":"Figure: Helix（System 1 200 Hz / System 2 7–9 Hz）","url":"https://www.figure.ai/news/helix"}],"as_of":"","related_ids":["policy-inference-frequency","control-latency","control-decimation","real-time-control","control-bandwidth","proportional-derivative-control"],"name":"Control Frequency","alt":"控制频率","abbr":"","aliases":["Control Rate","Control Period","Control Cycle"],"one_liner":"How many times per second a controller sends out commands, in Hz; its reciprocal is the control period.","explanation":"Control frequency is how many times per second a controller completes the loop of reading sensors, computing, and sending a command, measured in hertz (Hz); its reciprocal is the control period — 1 kHz, for instance, means once every 1 millisecond. Robots typically nest several such loops, running faster the closer they are to the motor: a large model or VLA policy might output actions a few to a few dozen times per second, joint position or torque loops run at hundreds to a few thousand hertz, and the current loop inside a motor driver runs faster still. Too low a frequency makes the robot slow to react to disturbances, risking shaking or falling; too high a frequency can outrun the available compute or communication bandwidth, so commands that aren't computed in time get rejected or throw an error. It's related to but distinct from “policy inference frequency” and “control latency”: the former measures only how fast the model produces a result, while the latter measures the full time from sensing to the action actually taking effect. The control frequency used in simulation training has to match the real robot, or the sim-to-real gap widens.","example":"Franka's arm runs its FCI interface at 1 kHz: the network round trip, the user's callback computation, and the robot's own processing must together finish within 1 millisecond, or that tick's command gets dropped — and 20 dropped ticks in a row trigger an error stop. A quadruped policy in legged_gym, by contrast, runs at 50 Hz, producing a new action every 20 milliseconds.","related":["Policy Inference Frequency","Control Latency","Control Decimation","Real-Time Control","Control Bandwidth","Proportional-Derivative Control"]},{"id":"policy-inference-frequency","category":"control","sec":0,"tier":2,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164v1"},{"title":"Figure: Helix（System 2 7–9 Hz / System 1 200 Hz）","url":"https://www.figure.ai/news/helix"},{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)","url":"https://arxiv.org/abs/2506.07339"}],"as_of":"2025-06","related_ids":["control-frequency","inference-latency","action-chunking","asynchronous-inference","real-time-chunking","dual-system-architecture"],"name":"Policy Inference Frequency","alt":"策略推理频率","abbr":"","aliases":["Inference Frequency","Model Call Frequency"],"one_liner":"How many times per second a learned policy is called to compute a new action, often lower than the control frequency below it.","explanation":"Policy inference frequency is how many times per second a neural-network policy runs a forward pass to produce a new action (or a new chunk of actions), measured in Hz. It is not the same as control frequency: control frequency is how often the actuators receive a new command, and joint control loops commonly run at hundreds to a few thousand Hz. Large models are slow to run, so they typically predict a whole block of actions at once (action chunking) and execute that block step by step between inferences, so inference frequency ends up far lower than control frequency. π0, for example, outputs 50 steps of actions at once; on a 50 Hz robot it executes 25 of those steps before inferring again, roughly 2 Hz, with a single inference taking about 73 milliseconds on an RTX 4090. The lower the inference frequency, the slower the reaction to sudden changes, and the more likely stutters appear at the seams between chunks — which is why asynchronous inference, real-time chunking, and fast-slow dual systems (e.g., Figure's Helix, with a 7–9 Hz slow system and a 200 Hz fast system) exist.","example":"π0 re-infers every 16 steps (0.8 seconds) of execution on 20 Hz-controlled UR5e and Franka arms, and every 25 steps (0.5 seconds) on robots controlled at 50 Hz.","related":["Control Frequency","Inference Latency","Action Chunking","Asynchronous Inference","Real-Time Chunking","Dual-System Architecture (System 1 / System 2)"]},{"id":"proportional-integral-derivative-control","category":"control","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Proportional–integral–derivative controller","url":"https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller"},{"title":"legged_gym legged_robot_config.py（PD 刚度/阻尼与 action_scale 配置）","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["proportional-derivative-control","cascade-control","stiffness-and-damping-gains","integral-windup-anti-windup","step-response-metrics","position-control"],"name":"Proportional-Integral-Derivative Control","alt":"PID 控制","abbr":"PID","aliases":["PID","PID Control","PID Controller"],"one_liner":"A classic feedback controller that combines the current error, its accumulated history, and its rate of change into one output.","explanation":"PID control is the most widely used feedback control method; Minorsky gave it a formal control law in 1922 while studying automatic steering for U.S. Navy ships. The control output is u = Kp·e + Ki·∫e dt + Kd·de/dt, where e is the target value minus the measured value (the error), and Kp, Ki, Kd are three gains to be tuned. The proportional term responds to the current error; the integral term accumulates past error, eliminating the steady-state error that proportional action alone would leave behind; the derivative term looks at how fast the error is changing, providing damping that reduces overshoot but can also amplify measurement noise. Dropping the integral term gives PD control, and dropping the derivative term gives PI control. PID shows up in joint motors, camera gimbals, and drone attitude loops; reinforcement-learning locomotion policies typically output only target joint angles, which a lower-level PD controller converts into torque. When the motor's output is already saturated at its limit but the integral term keeps accumulating anyway, the result is pronounced overshoot — a problem called integral windup, which needs anti-windup handling.","example":"A robot arm joint needs to reach 30° and is currently at 25°: the P term pushes based on the 5° error; if gravity keeps it stuck at 29.5° short of the target, the I term keeps accumulating that 0.5° gap until it's closed; as the joint nears the target, the D term sees the error shrinking fast and eases off early, preventing overshoot past 30°.","related":["Proportional-Derivative Control","Cascade Control","Stiffness and Damping Gains","Integral Windup / Anti-windup","Step Response Metrics (Overshoot / Settling Time / Steady-State Error)","Position Control"]},{"id":"proportional-derivative-control","category":"control","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Proportional–integral–derivative controller","url":"https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller"},{"title":"legged_gym: legged_robot.py（_compute_torques）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"legged_gym: legged_robot_config.py（PD Drive parameters）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["proportional-integral-derivative-control","stiffness-and-damping-gains","position-control","gravity-compensation","mit-mode","control-decimation"],"name":"Proportional-Derivative Control","alt":"PD 控制","abbr":"PD","aliases":["PD Control","PD Controller"],"one_liner":"Feedback control based on how large an error is and how fast it's changing — PID with the integral term dropped.","explanation":"PD control is PID control with the integral term removed. PID's output is u = Kp·e + Ki·∫e + Kd·de/dt, where e is the difference between the target and the measured value, and Kp, Ki, Kd are three gains; PD keeps only the proportional term (output grows with the error) and the derivative term (output responds to how fast the error is changing, which damps overshoot and oscillation). On a robot joint it's commonly written τ = Kp(q* − q) − Kd·q̇: τ is motor torque, q* is the target angle, q and q̇ are the measured angle and angular velocity, and Kp, Kd mechanically behave like a spring stiffness and a damping coefficient. Dropping the integral term leaves a steady-state error — a joint under gravity, for instance, settles slightly short of its target — which has to be offset with gravity compensation or by the higher-level policy; the derivative term is also sensitive to noise, so the velocity signal often needs filtering. It's the most common low-level joint controller used when deploying reinforcement-learning locomotion policies and VLA models.","example":"legged_gym multiplies the policy's output by 0.5 and adds it to the default joint angles to get q*, then computes torque as τ = Kp(q* − q) − Kd·q̇; example configs use Kp around 10–15 N·m/rad and Kd around 1–1.5 N·m·s/rad.","related":["Proportional-Integral-Derivative Control","Stiffness and Damping Gains","Position Control","Gravity Compensation","MIT Mode","Control Decimation"]},{"id":"step-response-metrics","category":"control","sec":0,"tier":2,"sources":[{"title":"MATLAB stepinfo 文档（RiseTime / SettlingTime / Overshoot 定义）","url":"https://www.mathworks.com/help/control/ref/dynamicsystem.stepinfo.html"},{"title":"Wikipedia: Settling time","url":"https://en.wikipedia.org/wiki/Settling_time"},{"title":"Wikipedia: Steady-state error","url":"https://en.wikipedia.org/wiki/Steady-state_error"}],"as_of":"","related_ids":["proportional-integral-derivative-control","proportional-derivative-control","stiffness-and-damping-gains","damping-ratio","integral-windup-anti-windup","control-bandwidth"],"name":"Step Response Metrics (Overshoot / Settling Time / Steady-State Error)","alt":"阶跃响应指标（超调 / 调节时间 / 稳态误差）","abbr":"","aliases":["Overshoot","Settling Time","Steady-State Error","Rise Time"],"one_liner":"A set of numbers measuring how fast, how stable, and how accurately a system's output reaches a target after a sudden change.","explanation":"Step response metrics evaluate how well a closed-loop controller is tuned: suddenly change the target value from one level to another (a step input), record how the output changes over time, and read off several numbers. Rise time is how long the output takes to go from 10% to 90% of the total change; overshoot is the percentage by which the output peaks past its final value; settling time is how long until the output stays within roughly 2% (sometimes 5%) of the final value for good; and steady-state error is the remaining gap between output and target after a long time. These trade off against each other: raising Kp speeds up the response but increases overshoot, while raising damping suppresses overshoot but slows the response; pure proportional control tends to leave a steady-state error, which adding an integral term can eliminate. These metrics are used when tuning joint PD gains and when comparing simulated response against real hardware.","example":"Commanding a joint from 0 rad to 1 rad: the output peaks at 1.13 rad, about 15% above the final settled value of 0.98 rad, i.e., roughly 15% overshoot; after 0.3 seconds it stays within 2% of 0.98 rad, giving a settling time around 0.3 seconds; the final gap from the 1 rad target, 0.02 rad, is the steady-state error.","related":["Proportional-Integral-Derivative Control","Proportional-Derivative Control","Stiffness and Damping Gains","Damping Ratio","Integral Windup / Anti-windup","Control Bandwidth"]},{"id":"integral-windup-anti-windup","category":"control","sec":0,"tier":3,"sources":[{"title":"Wikipedia: Integral windup","url":"https://en.wikipedia.org/wiki/Integral_windup"}],"as_of":"","related_ids":["proportional-integral-derivative-control","torque-limiting","step-response-metrics","cascade-control","active-disturbance-rejection-control"],"name":"Integral Windup / Anti-windup","alt":"积分饱和与抗饱和","abbr":"","aliases":["Integrator Windup","Reset Windup"],"one_liner":"An actuator pinned at its limit while the integral term keeps accumulating, causing overshoot; anti-windup keeps that term in check during saturation.","explanation":"Integral windup occurs in controllers with an integral term (PI, PID). Actuators have physical limits — motor torque and current can't be infinite. When the error is large, say from a sudden setpoint change or a joint blocked by an obstacle, the controller wants an output beyond that limit, but the actual output is capped, so the error shrinks slowly while the integral term K_i·∫e dt keeps accumulating. Once the error flips sign, the integral term takes a long time to unwind, producing large overshoot and prolonged oscillation. Common anti-windup fixes: clamping, capping the integral term within bounds; conditional integration, pausing accumulation while the output is saturated; and back-calculation, feeding the difference between the saturated and unsaturated output back into the integrator with a gain, so it tracks what's actually happening. Both the velocity and current loops of a robot joint need to handle this.","example":"An arm's joint is held by a person's hand; the velocity loop's output sits pinned at the torque limit while the integral term keeps building up, and the joint lurches violently past its target the instant it's released. Adding clamping or back-calculation anti-windup to the integrator lets the joint settle smoothly back onto its commanded trajectory instead.","related":["Proportional-Integral-Derivative Control","Torque Limiting (Saturation)","Step Response Metrics (Overshoot / Settling Time / Steady-State Error)","Cascade Control","Active Disturbance Rejection Control"]},{"id":"feedforward-control","category":"control","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Feed forward (control)","url":"https://en.wikipedia.org/wiki/Feed_forward_(control)"},{"title":"unitree_sdk2 G1 底层示例 g1_ankle_swing_example.cpp（tau_ff / kp / kd 字段）","url":"https://github.com/unitreerobotics/unitree_sdk2/blob/main/example/g1/low_level/g1_ankle_swing_example.cpp"}],"as_of":"","related_ids":["gravity-compensation","friction-compensation","computed-torque-control","proportional-derivative-control","mit-mode","inverse-dynamics"],"name":"Feedforward Control","alt":"前馈控制","abbr":"","aliases":["Feedforward Torque","Feedforward Compensation","τff"],"one_liner":"Using a model to compute the needed control input in advance, and adding it directly, rather than waiting for an error to appear.","explanation":"Feedforward control uses a system model to compute, ahead of time, the control input needed for a known target trajectory or a measurable disturbance, and adds it directly to the output; feedback control, by contrast, corrects only after measuring an actual error. Pure feedforward is open-loop: if the model is inaccurate or there's an unmeasured disturbance, nothing corrects the resulting error, so in practice feedforward is usually combined with feedback — feedforward handles the bulk of the work, and feedback only cleans up the small remainder, which lets the system track more closely while allowing the feedback gains to be set lower, keeping the joints more compliant. Typical feedforward terms in robotics include gravity compensation, friction compensation, and the torque computed from a desired acceleration via inverse dynamics (this is exactly what computed-torque control is built on). A common joint-motor command is τ = τff + kp(q_target − q) + kd(q̇_target − q̇), where τff is the feedforward torque and the other two terms are PD feedback.","example":"A robot arm reaches out horizontally with a 1 kg payload hanging 0.5 m from the shoulder joint; the gravity torque from the payload alone is about 9.8×0.5≈4.9 N·m. Precomputing this and feeding it into τff means the PD feedback no longer has to widen the position error just to hold the payload up, and tracking error drops noticeably.","related":["Gravity Compensation","Friction Compensation","Computed Torque Control","Proportional-Derivative Control","MIT Mode","Inverse Dynamics"]},{"id":"control-bandwidth","category":"control","sec":0,"tier":3,"sources":[{"title":"MathWorks: bandwidth（增益首次低于直流值 70.79%，即 −3 dB 的频率）","url":"https://www.mathworks.com/help/control/ref/dynamicsystem.bandwidth.html"},{"title":"Wikipedia: Bandwidth (signal processing)","url":"https://en.wikipedia.org/wiki/Bandwidth_(signal_processing)"},{"title":"Wensing et al., Proprioceptive Actuator Design in the MIT Cheetah: Impact Mitigation and High-Bandwidth Physical Interaction (IEEE T-RO 2017)","url":"https://ieeexplore.ieee.org/document/7827048"}],"as_of":"","related_ids":["control-frequency","control-latency","cascade-control","proprioceptive-actuator","force-control","step-response-metrics"],"name":"Control Bandwidth","alt":"控制带宽","abbr":"","aliases":["Closed-Loop Bandwidth","-3 dB Bandwidth","Force Control Bandwidth","Current Loop Bandwidth"],"one_liner":"The highest signal frequency a closed-loop system can still keep up with, usually taken where the gain drops by 3 dB.","explanation":"Feed a closed-loop control system sinusoidal commands at various frequencies: at low frequency the output tracks fully, and as frequency rises the output's amplitude shrinks and its lag grows; the frequency at which the amplitude has dropped to about 70.7% (i.e., −3 dB) of its low-frequency value is the closed-loop bandwidth. Higher bandwidth means a faster response and better rejection of fast disturbances. It's not the same as control frequency: control frequency is how many times per second commands are computed, while bandwidth is how fast the system can actually keep up — the former typically needs to be several times higher than the latter for the bandwidth to actually be achievable. In cascade control, the inner loop's bandwidth must be noticeably higher than the outer loop's — e.g., the current loop faster than the velocity loop, the velocity loop faster than the position loop. Legged robots have very short ground-contact times, requiring high-bandwidth force control; compliant elements, low-stiffness transmissions, and communication delay all limit the achievable bandwidth.","example":"A first-order system G(s)=1/(τs+1) with τ=16 ms has a bandwidth of 1/τ = 62.5 rad/s, about 10 Hz: it tracks a 2 Hz sinusoidal command almost fully, but at 30 Hz the output amplitude has shrunk to roughly 30%. On real hardware, a single ground-contact phase on MIT Cheetah while running can be as short as about 85 milliseconds, which is why its proprioceptive actuators list ‘high-bandwidth force control’ as a design goal.","related":["Control Frequency","Control Latency","Cascade Control","Proprioceptive Actuator","Force Control","Step Response Metrics (Overshoot / Settling Time / Steady-State Error)"]},{"id":"real-time-control","category":"control","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Real-time computing","url":"https://en.wikipedia.org/wiki/Real-time_computing"},{"title":"Franka FCI 文档：libfranka Overview（1 kHz 控制回路、丢包处理）","url":"https://frankarobotics.github.io/docs/doc/libfranka/docs/overview.html"},{"title":"Franka FCI 文档：Setting up the Real-Time Kernel","url":"https://frankarobotics.github.io/docs/doc/libfranka/docs/real_time_kernel.html"}],"as_of":"","related_ids":["control-frequency","hard-real-time","real-time-operating-system","preempt-rt","jitter","control-latency"],"name":"Real-Time Control","alt":"实时控制","abbr":"","aliases":["Real-Time Control Loop"],"one_liner":"A control program that reads sensors, computes a command, and sends it to the motors on a fixed cycle, every cycle, on time.","explanation":"Real-time control means a control program reads sensors, computes a command, and sends it to the motors on a fixed period (e.g., every 1 millisecond), and every single cycle must finish before its deadline. ‘Real-time’ here emphasizes being on time and predictable, not simply fast: a hard real-time system counts a single missed deadline as a failure, while a soft real-time system just degrades. Franka arms, for example, run the user's control callback through libfranka at 1 kHz, and the vendor requires the workstation to use a PREEMPT_RT real-time kernel and run the control program at real-time priority. A common division of labor in embodied-AI systems is: a large model such as a VLA infers a target at a few to a few dozen Hz, while low-level joint control tracks it inside a real-time loop at hundreds to thousands of Hz.","example":"When doing torque control on a Franka FR3 with libfranka, the callback function is invoked every 1 millisecond, reading the robot's state and returning 7 joint torques; if the computer stalls and drops more than 20 consecutive packets, the control loop throws a communication_constraints_violation error and stops.","related":["Control Frequency","Hard Real-Time","Real-Time Operating System (RTOS)","PREEMPT_RT","Jitter (Timing Jitter)","Control Latency"]},{"id":"control-latency","category":"control","sec":0,"tier":2,"sources":[{"title":"Real-Time Execution of Action Chunking Flow Policies (Black, Galliker, Levine, arXiv 2506.07339)","url":"https://arxiv.org/abs/2506.07339"},{"title":"Wikipedia: Smith predictor","url":"https://en.wikipedia.org/wiki/Smith_predictor"}],"as_of":"2025-12","related_ids":["inference-latency","control-frequency","real-time-chunking","asynchronous-inference","action-chunking","jitter"],"name":"Control Latency","alt":"控制延迟","abbr":"","aliases":["End-to-End Latency","Execution Latency","Perception-to-Action Latency"],"one_liner":"The time gap between a sensor capturing data and the motor actually executing the corresponding action.","explanation":"Control latency is the total time a full sense-compute-execute cycle takes: camera exposure and data transfer, model inference, network or bus communication, and motor-drive response all count toward it, with end-to-end latency meaning the whole gap from when an observation is captured to when the corresponding action lands on the joints. Feedback control works by seeing an error and correcting it, so the larger the latency, the more outdated the information the controller is acting on — at best this makes motion sluggish and prone to overshoot, and at worst it causes oscillation and instability; the Smith predictor, proposed by O. J. M. Smith in 1957, is a classic technique for dealing with pure delay. A single VLA inference often takes tens to well over a hundred milliseconds, far longer than the underlying control cycle, which is why techniques like action chunking, asynchronous inference, and real-time chunking (RTC) exist. It's a broader notion than 'inference latency,' which measures only the model's own forward-pass time.","example":"Physical Intelligence's real-time chunking paper reports that π0.5 takes about 76 ms per inference on an RTX 4090, while the robot executes at 50 Hz (20 ms per step); by the time a new action is ready, the robot has already taken about 3 more steps, and without special handling, the boundary between action chunks shows up as a pause or a jump.","related":["Inference Latency","Control Frequency","Real-Time Chunking","Asynchronous Inference","Action Chunking","Jitter (Timing Jitter)"]},{"id":"jitter","category":"control","sec":0,"tier":3,"sources":[{"title":"ROS 2 Design: Introduction to Real-time Systems","url":"https://design.ros2.org/articles/realtime_background.html"},{"title":"libfranka include/franka/robot.h（1 kHz 控制回调说明）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"}],"as_of":"","related_ids":["real-time-control","control-frequency","control-latency","preempt-rt","ethercat","hard-real-time"],"name":"Jitter (Timing Jitter)","alt":"时间抖动","abbr":"","aliases":["Timing Jitter","Scheduling Jitter"],"one_liner":"The fluctuation in when a periodic task actually runs versus its ideal timing — a key measure of real-time performance.","explanation":"Jitter is the deviation and fluctuation between when an event that's supposed to happen on a fixed period actually happens and its ideal timing. The most commonly discussed case in robotics is the scheduling jitter of a control loop: a 1 kHz joint control loop is supposed to run every 1 ms, but the actual interval may come in a little long or a little short each time. It's not the same as control latency: latency is how long it takes from reading a sensor to issuing a command, while jitter is whether that duration stays the same each time. Discrete controllers generally assume a fixed sampling period, and jitter throws off the computed integral and derivative terms, causing vibration at best and instability at worst; a command that doesn't reach the drive on time can also trigger a fault stop. That's why real-time systems chase determinism rather than just speed — ROS 2's design documentation explicitly frames a real-time system as one defined by deterministic scheduling, not low latency. Jitter is commonly measured with tools like cyclictest, and reduced with a PREEMPT_RT real-time kernel, CPU isolation, or a deterministic bus such as EtherCAT.","example":"A Franka arm's low-level control runs at 1 kHz, and libfranka requires the user's control callback to finish within a tight time window or that cycle's command is rejected; running torque control therefore generally calls for a real-time kernel and for avoiding operations with unpredictable timing, like printing or dynamic memory allocation, inside the control thread.","related":["Real-Time Control","Control Frequency","Control Latency","PREEMPT_RT","EtherCAT (Ethernet for Control Automation Technology)","Hard Real-Time"]},{"id":"model-based-control","category":"control","sec":0,"tier":2,"sources":[{"title":"Modern Robotics 11.4: Motion Control with Torque or Force Inputs (Part 3 of 3)（计算力矩控制）","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-3-of-3/"},{"title":"Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (IROS 2018, MIT DSpace)","url":"https://dspace.mit.edu/handle/1721.1/138000"}],"as_of":"","related_ids":["learning-based-control","computed-torque-control","model-predictive-control","whole-body-control","system-identification","dynamics"],"name":"Model-Based Control","alt":"基于模型的控制","abbr":"","aliases":["Model-Driven Control"],"one_liner":"Writing out a mathematical model of the robot and its environment first, then deriving or optimizing control commands from it.","explanation":"A broad category of methods that design controllers around an explicit mathematical model, covering both kinematics and dynamics — for example, the manipulator equation M(q)q̈ + c(q, q̇) + g(q) = τ, where M is the mass matrix, c the Coriolis and centrifugal terms, g the gravity term, and τ the joint torque. Common techniques include gravity compensation, computed torque control (which uses the model to cancel nonlinearities), operational space control, model predictive control (MPC, which optimizes a short future window using the model every cycle), and quadratic-programming-based whole-body control. The advantages are interpretability, no need for training data, and the ability to analyze stability formally; the drawbacks are that the model must be accurate — contact, friction, and deformable objects are hard to model — and complex tasks demand extensive manual design. It sits opposite learning-based control, though the two are now often combined. Note it is distinct from model-based reinforcement learning, where the ‘model’ is itself learned from data.","example":"MIT Cheetah 3 simplified body dynamics into a convex MPC that solves for ground reaction forces at each foot at 20–30 Hz in under 1 millisecond per solve; the same set of parameters produced trotting, galloping, bounding, and other gaits, reaching forward speeds up to 3 m/s.","related":["Learning-Based Control","Computed Torque Control","Model Predictive Control","Whole-Body Control","System Identification","Dynamics"]},{"id":"learning-based-control","category":"control","sec":0,"tier":2,"sources":[{"title":"Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning (Brunke et al., arXiv 2108.06266)","url":"https://arxiv.org/abs/2108.06266"},{"title":"Learning Agile and Dynamic Motor Skills for Legged Robots (Hwangbo et al., Science Robotics 2019, arXiv 1901.08652)","url":"https://arxiv.org/abs/1901.08652"}],"as_of":"","related_ids":["model-based-control","rl-based-locomotion-control","reinforcement-learning","imitation-learning","safe-reinforcement-learning","learning-based-whole-body-control"],"name":"Learning-Based Control","alt":"基于学习的控制","abbr":"","aliases":["Data-Driven Control"],"one_liner":"Using data and machine learning to build or improve a controller, rather than relying entirely on hand-derived equations and tuning.","explanation":"A broad term for methods that use learned data to construct or improve a controller, as opposed to model-based control, which relies on hand-built models. Three approaches are typical: learning a dynamics model (or its error) and handing it to a traditional controller such as an MPC; training a neural-network policy end to end with reinforcement learning or imitation learning, mapping observations directly to joint targets or torques; and online learning that adjusts a controller's parameters as it runs. A 2022 survey by Brunke et al. in Annual Review of Control, Robotics, and Autonomous Systems groups methods that learn uncertain dynamics to safely improve performance under ‘learning-based control,’ contrasting it with safe reinforcement learning. It suits hard-to-model effects like contact, friction, and motor behavior, at the cost of needing large amounts of data or simulation, and it is difficult to give strict stability guarantees. Today, most quadruped and humanoid locomotion controllers use reinforcement-learning policies trained in simulation.","example":"Hwangbo et al., published in Science Robotics in 2019, first trained an actuator network on real-robot data to model motor behavior, then trained an ANYmal quadruped's policy with reinforcement learning in simulation using that network; after transfer to the real robot it ran faster than prior methods and could get back up on its own after falling from awkward poses.","related":["Model-Based Control","RL-based Locomotion Control","Reinforcement Learning","Imitation Learning","Safe Reinforcement Learning","Learning-Based Whole-Body Control"]},{"id":"position-control","category":"control","sec":1,"tier":1,"sources":[{"title":"ROBOTIS e-Manual: XM430-W350（Operating Mode）","url":"https://emanual.robotis.com/docs/en/dxl/x/xm430-w350/"},{"title":"legged_gym: legged_robot.py（_compute_torques：P / V / T 三种控制类型）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"}],"as_of":"","related_ids":["velocity-control","torque-control","proportional-derivative-control","cascade-control","stiffness-and-damping-gains","mit-mode"],"name":"Position Control","alt":"位置控制","abbr":"","aliases":["Position Mode","Joint Position Control"],"one_liner":"Give a motor or joint a target angle or position, and the controller drives it there and holds it.","explanation":"Position control is the most common motor control mode: the upper level only supplies a target position — say, turn the joint to 30 degrees — and the driver's internal closed loop computes current from the gap between the encoder reading and the target, pulling the joint to that target and holding it there. Most servos and industrial robot arms default to this mode, since it's accurate and simple to use. Its downside is that it's 'stiff': at high gain, the joint pushes hard to reach its target regardless of whether it meets a person or an unexpected obstacle, without yielding — so tasks that require contact often switch to torque control, impedance control, or force control instead. In reinforcement-learning locomotion, the policy typically also outputs target joint angles, which a PD controller converts into torque, just with lower stiffness gains to leave the joints some compliance. Alongside position mode, the other common options are velocity mode and torque (current) mode.","example":"ROBOTIS's Dynamixel XM430 servo has 6 operating modes; setting Operating Mode to 3 selects position control mode, and once a target position is written, the servo's internal position PID drives it there and holds it.","related":["Velocity Control","Torque Control","Proportional-Derivative Control","Cascade Control","Stiffness and Damping Gains","MIT Mode"]},{"id":"velocity-control","category":"control","sec":1,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), 11.3 Motion Control with Velocity Inputs","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"ros2_controllers: diff_drive_controller","url":"https://control.ros.org/master/doc/ros2_controllers/diff_drive_controller/doc/userdoc.html"},{"title":"Diffusion Policy（velocity control is more affected by latency than position control）","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["position-control","torque-control","cascade-control","cyclic-synchronous-position-velocity-torque-modes","differential-drive-kinematics","velocity-command-tracking"],"name":"Velocity Control","alt":"速度控制","abbr":"","aliases":["Velocity Mode","Speed Control"],"one_liner":"The higher level specifies a desired speed, and the drive makes the motor or robot move at that speed.","explanation":"Velocity control means a higher level specifies a desired velocity and the lower level tracks it, at two possible levels. At the joint level, a servo drive's velocity mode (e.g., CSV mode in the CiA 402 protocol) accepts a target speed and closes a velocity loop internally; a stepper motor's speed is instead set directly by its pulse frequency. At the task level, a desired end-effector velocity can be converted to joint velocities via the Jacobian pseudoinverse; a mobile base receives linear and angular velocity, converted into left and right wheel speeds using the wheelbase and wheel radius. Compared to position control, it doesn't directly govern where the system ends up — position error must be closed by a higher-level loop, making it more sensitive to delay, an effect also seen in Diffusion Policy's ablation experiments; compared to torque control, it doesn't directly govern how much force is applied, so it's less compliant on contact.","example":"ROS 2's diff_drive_controller subscribes to the cmd_vel velocity command, takes the x-component of linear velocity and the z-component of angular velocity, converts them using the wheelbase and wheel radius, and writes the result to the left and right wheel joints' velocity command interfaces.","related":["Position Control","Torque Control","Cascade Control","Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)","Differential Drive Kinematics","Velocity Command Tracking"]},{"id":"torque-control","category":"control","sec":1,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), Fig. 11.1 与 11.5 Force Control","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"libfranka robot.h（torque control / 1 kHz / without gravity and friction）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"},{"title":"legged_gym: legged_robot.py（control_type P / V / T）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"}],"as_of":"","related_ids":["position-control","velocity-control","impedance-control","gravity-compensation","joint-torque-sensor","cascade-control"],"name":"Torque Control","alt":"力矩控制","abbr":"","aliases":["Effort Control","Torque Mode","Current Control (motor-level)"],"one_liner":"Commanding a joint directly with ‘how much torque to output,’ rather than ‘which angle to turn to.’","explanation":"Torque control means the controller's output is directly the joint torque τ (units N·m), and the drive is responsible for making the motor actually produce it. A motor's torque is roughly proportional to its winding current, so this is often implemented at the motor level as a closed current loop — hence the alternate name ‘current control’; joints with a large gear ratio have significant gear friction, so current alone isn't accurate enough, requiring a joint torque sensor at the output to close the loop instead. Compared to position control (where the drive internally forces the angle to a target), torque control leaves ‘how stiff or soft’ up to the higher-level algorithm, forming the basis of impedance control, whole-body control, and operational space control, and behaves more compliantly on contact or impact — at the cost of the higher level having to compensate for gravity and friction itself, since an inaccurate model causes sagging or drift. Simulation frameworks such as legged_gym also list it as an action type alongside position and velocity.","example":"Franka arms' libfranka interface provides a 1 kHz joint-torque command: the user's callback computes 7 joint torques, and the documentation notes these should exclude gravity and friction — the robot adds its own gravity and friction compensation on top.","related":["Position Control","Velocity Control","Impedance Control","Gravity Compensation","Joint Torque Sensor","Cascade Control"]},{"id":"cascade-control","category":"control","sec":1,"tier":2,"sources":[{"title":"ODrive Documentation: Control Structure and Tuning","url":"https://docs.odriverobotics.com/v/latest/manual/control.html"},{"title":"Wikipedia: Proportional–integral–derivative controller","url":"https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller"}],"as_of":"","related_ids":["proportional-integral-derivative-control","position-control","velocity-control","torque-control","field-oriented-control","mit-mode"],"name":"Cascade Control","alt":"串级控制","abbr":"","aliases":["Position/Velocity/Current Loops","Three-Loop Control","Cascaded Control"],"one_liner":"Nests several feedback loops, where the outer loop's output becomes the inner loop's target — like position over velocity over current.","explanation":"Cascade control nests several feedback loops together, with the output of the outer loop serving as the target value for the inner loop. The most typical example, in a motor or joint module, is three loops: the outermost position loop compares the target angle against the encoder reading and outputs a target velocity; the middle velocity loop outputs a target current; and the innermost current loop regulates the winding voltage, with motor current roughly proportional to output torque. In the open-source driver ODrive, for example, the position loop is proportional control, while the velocity and current loops are both PI control; skipping the position loop gives velocity mode, and keeping only the current loop gives torque mode. The design principle is that each inner loop runs much faster than the loop outside it, so disturbances like a sudden load change or a power fluctuation get suppressed by the inner loop first, letting the outer loop treat the inner loop as close to an ideal actuator. What robotics calls position control, velocity control, and torque control essentially just names which of these three loops the command enters at; joint interfaces like MIT mode instead combine the PD terms for position and velocity with a feedforward torque into a single torque command.","example":"A joint module receives 'turn to 90°': the position loop computes 'turn at 2 rad/s,' the velocity loop computes 'need 3 A of current,' and the current loop adjusts voltage to hold the current at 3 A; if the load suddenly increases, the current and velocity loops absorb the disturbance first, and the position loop barely notices.","related":["Proportional-Integral-Derivative Control","Position Control","Velocity Control","Torque Control","Field-Oriented Control","MIT Mode"]},{"id":"joint-space-control","category":"control","sec":1,"tier":2,"sources":[{"title":"Modern Robotics 11.4: Motion Control with Torque or Force Inputs (Part 3 of 3)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-3-of-3/"},{"title":"ros2_controllers: joint_trajectory_controller 文档","url":"https://control.ros.org/rolling/doc/ros2_controllers/joint_trajectory_controller/doc/userdoc.html"}],"as_of":"","related_ids":["joint-space","task-space-control","proportional-derivative-control","computed-torque-control","inverse-kinematics","ros2-control"],"name":"Joint-Space Control","alt":"关节空间控制","abbr":"","aliases":["Joint-Level Control"],"one_liner":"Controlling a robot using each joint's own angle as the target and error, with every joint tracking its own desired trajectory.","explanation":"Joint-space control writes the goal as a sequence of joint angles over time, q_d(t), and compares each actual angle q against its target individually, outputting torque or velocity from the error. The simplest form is independent PD control per joint: τ = Kp(q_d − q) + Kd(q̇_d − q̇), where Kp and Kd are stiffness and damping gains; computed torque control goes further, using a dynamics model to also cancel inertia, Coriolis forces, and gravity. Task-space control, by contrast, works directly with end-effector pose error and converts it to joint torques through the Jacobian matrix, which makes it easier to command the end-effector to move in a straight line or apply a contact force. Joint-space control is simple to implement and immune to kinematic singularities, but any desired end-effector path must first be converted into a joint trajectory by inverse kinematics or a planner. Reinforcement-learning locomotion policies for legged and humanoid robots, which output joint target positions, also belong to this layer.","example":"After MoveIt plans a joint trajectory, it hands it to ros2_control's joint_trajectory_controller, which interpolates between waypoints over time and sends each joint's target position (or a PID-derived torque) to the drive every control cycle.","related":["Joint Space","Task-Space Control","Proportional-Derivative Control","Computed Torque Control","Inverse Kinematics (IK)","ros2_control"]},{"id":"stiffness-and-damping-gains","category":"control","sec":1,"tier":2,"sources":[{"title":"legged_gym legged_robot.py：_compute_torques（PD 力矩公式）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"unitree_rl_gym：Go2 配置","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/go2/go2_config.py"},{"title":"unitree_rl_gym：G1 配置","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/g1/g1_config.py"}],"as_of":"2026-09","related_ids":["proportional-derivative-control","mit-mode","impedance-control","rl-based-locomotion-control","step-response-metrics","mass-spring-damper-system"],"name":"Stiffness and Damping Gains","alt":"刚度与阻尼增益","abbr":"Kp/Kd","aliases":["Kp/Kd","PD Gains"],"one_liner":"The two coefficients, Kp and Kd, in joint PD control that determine how stiff and how stable a joint feels.","explanation":"Stiffness and damping gains are the two coefficients in the joint PD control law τ = Kp(q* − q) + Kd(q̇* − q̇): τ is the motor's output torque, q* and q are the target and actual joint angle, and q̇* and q̇ are the target and actual angular velocity. A larger Kp (units N·m/rad) pulls the joint back toward its target more forcefully — tighter tracking, but stiffer, with harder collision impacts; Kd (units N·m·s/rad) resists based on velocity, damping oscillation and overshoot, though too large a value makes the joint sluggish and amplifies velocity noise. Mechanically, this control law is equivalent to a spring and damper connected in parallel across the joint, which is why simulators often just call these values ‘stiffness’ and ‘damping.’ In reinforcement-learning locomotion, the policy usually outputs only a target angle, with torque computed by PD, so the choice of Kp/Kd directly affects how well the policy transfers from simulation to the real robot.","example":"In unitree_rl_gym, every joint of Unitree's Go2 uses Kp=20 N·m/rad and Kd=0.5 N·m·s/rad; on the G1 humanoid, the hip is set to Kp=100, the knee to Kp=150, and the ankle to Kp=40 — heavier load-bearing joints get higher stiffness.","related":["Proportional-Derivative Control","MIT Mode","Impedance Control","RL-based Locomotion Control","Step Response Metrics (Overshoot / Settling Time / Steady-State Error)","Mass-Spring-Damper System"]},{"id":"mit-mode","category":"control","sec":1,"tier":2,"sources":[{"title":"bgkatz/motorcontrol 固件 foc.c（torque_control: kp·(p_des−θ) + kd·(v_des−ω) + t_ff）","url":"https://github.com/bgkatz/motorcontrol/blob/master/Core/Src/foc.c"},{"title":"DAMIAO DM-J4310-2EC V1.1 手册（Operating modes: MIT mode）","url":"https://github.com/enactic/damiao/blob/main/website/i18n/en/docusaurus-plugin-content-docs/current/products/hardware/dm-j4310-2ec-v1.1.mdx"},{"title":"Xiaomi CyberGear 说明书英译版（Operation control mode 五参数指令）","url":"https://github.com/belovictor/cybergear-docs/blob/main/instructionmanual/instructionmanual.md"}],"as_of":"","related_ids":["proportional-derivative-control","stiffness-and-damping-gains","mit-mini-cheetah-actuator","damiao-dm-j4310-2ec-joint-motor","xiaomi-cybergear-micro-motor","torque-control"],"name":"MIT Mode","alt":"MIT 模式","abbr":"","aliases":["MIT Cheetah-style Joint Command","MIT Control Mode"],"one_liner":"A single joint-motor command carrying target position, velocity, Kp, Kd, and feedforward torque, with the drive computing torque via a PD formula.","explanation":"This joint-motor command format originates from the motor controller of MIT's Biomimetic Robotics Lab's Mini Cheetah quadruped (open-source firmware by Ben Katz). Chinese motor makers such as Damiao followed and made their drives compatible with it, and Xiaomi calls it ‘motion control mode’ on its CyberGear motor. In a single CAN frame, the host sends a target position p_des, target velocity v_des, stiffness Kp, damping Kd, and a feedforward torque τ_ff; the drive's internal loop computes τ = Kp(p_des − p) + Kd(v_des − v) + τ_ff (p and v being the measured position and velocity) and passes that to the current loop. Changing the parameters switches behavior: a large Kp approximates position control, Kp = 0 gives velocity or damping control, and Kp = Kd = 0 gives pure torque control. Reinforcement-learning locomotion and humanoid policies, which output joint target positions, are often executed through this mode with fixed Kp and Kd. It is a single-joint PD-plus-feedforward format, not the direction-split hybrid force/position control described elsewhere.","example":"The Damiao DM-J4310 manual notes that with Kp = 0 and Kd ≠ 0, giving v_des spins the motor at constant speed; with Kp = Kd = 0, giving τ_ff outputs exactly that torque; and when doing position control, Kd must never be set to 0, or the motor will oscillate or even lose control.","related":["Proportional-Derivative Control","Stiffness and Damping Gains","MIT Mini Cheetah Actuator","DAMIAO DM-J4310-2EC Joint Motor","Xiaomi CyberGear Micro-Motor","Torque Control"]},{"id":"servo-enable","category":"control","sec":1,"tier":2,"sources":[{"title":"ethercat_driver_ros2: CANopen over EtherCAT for electrical drives（CiA402 状态机）","url":"https://icube-robotics.github.io/ethercat_driver_ros2/developer_guide/cia402_drive.html"}],"as_of":"","related_ids":["canopen-cia-402-drive-profile","servo-drive","holding-brake","safe-torque-off","cyclic-synchronous-position-velocity-torque-modes","safety-gantry"],"name":"Servo Enable (Servo ON / OFF)","alt":"伺服使能（上使能 / 下使能）","abbr":"","aliases":["Enable Operation","Servo ON","Servo OFF"],"one_liner":"Switching a motor drive into a state where it can actively output torque (enable), or cutting that output off (disable).","explanation":"Servo enable means switching a motor drive from powered-but-idle into closed-loop operation, actively outputting torque according to commands — commonly called ‘enabling’; the reverse, cutting off output, is ‘disabling.’ Take the widely used CiA 402 fieldbus drive profile as an example: the drive has an internal state machine, and the host must send Shutdown, Switch On, and Enable Operation commands in sequence through a control word; only once the drive reaches Operation Enabled does it accept position, velocity, or torque setpoints, and an error drops it into a Fault state that must be reset before continuing. This explicit enable step exists to prevent the motor from jerking unexpectedly the instant power is applied. Once disabled, the motor no longer outputs torque, so a joint without a holding brake will drop under gravity — which is why humanoid and legged robots are usually suspended from a gantry when being enabled or disabled.","example":"When bringing up an EtherCAT joint module, the host first reads the status word to confirm it is in Switch On Disabled, then sends Shutdown, Switch On, and Enable Operation in order; once the status word reports Operation Enabled, it starts sending periodic target positions. Before finishing, the joint is returned to a safe pose before being disabled.","related":["CANopen / CiA 402 Drive Profile","Servo Drive (Motor Driver)","Holding Brake","Safe Torque Off (STO)","Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)","Safety Gantry (Suspended Start)"]},{"id":"homing-zero-offset-calibration","category":"control","sec":1,"tier":2,"sources":[{"title":"LinuxCNC Docs: Homing Configuration","url":"https://linuxcnc.org/docs/html/config/ini-homing.html"},{"title":"ROBOTIS e-Manual: XM430-W350 (Homing Offset)","url":"https://emanual.robotis.com/docs/en/dxl/x/xm430-w350/"},{"title":"Xiaomi CyberGear 微电机说明书（英译版，含零位设置与控制模式）","url":"https://github.com/belovictor/cybergear-docs/blob/main/instructionmanual/instructionmanual.md"}],"as_of":"","related_ids":["joint-zero-position-calibration","joint-zero-calibration","incremental-encoder","absolute-encoder","soft-limits","forward-kinematics"],"name":"Homing / Zero-Offset Calibration","alt":"回零 / 零位标定","abbr":"","aliases":["Homing","Zero Calibration","Homing Offset","Joint Zero Calibration"],"one_liner":"Aligning each joint's encoder reading with the model's ‘zero degrees’ so the robot has a correct angle reference.","explanation":"Robot models such as URDF define one specific pose as the configuration where every joint angle equals zero, but when an encoder is physically mounted, its own zero point is arbitrary. Zero-offset calibration measures and stores the fixed offset between the raw encoder reading and the model's angle so later readings can be corrected; homing is the process of moving a joint to a reference point to establish that offset. Joints with incremental encoders (which track only relative rotation) lose their position when powered off, so they must be re-homed at every power-up: the joint first creeps slowly toward a limit switch or hard stop, then uses the encoder's index pulse to find an exact reference — a routine that CiA 402-compliant servo drives can run on their own. Absolute encoders know their position as soon as they power on, so they usually need calibrating only once. Safety features like soft limits are defined relative to this zero point — if it drifts, the end-effector develops a systematic offset, and a policy trained on one robot may not transfer well to another.","example":"Xiaomi's CyberGear micro-motor has a ‘set mechanical zero’ command (communication type 6) that sets the current position as zero, but this is lost on power-off; Dynamixel servos instead use a persistent Homing Offset parameter, where reading = actual position + offset.","related":["Joint Zero-Position Calibration (Homing / Offset Calibration)","Joint Zero Calibration (Homing)","Incremental Encoder","Absolute Encoder","Soft Limits (Software Joint Limits)","Forward Kinematics (FK)"]},{"id":"cyclic-synchronous-position-velocity-torque-modes","category":"control","sec":1,"tier":3,"sources":[{"title":"CAN in Automation: CiA 402 series – CANopen device profile for drives and motion control","url":"https://www.can-cia.org/can-knowledge/cia-402-series-canopen-device-profile-for-drives-and-motion-control"},{"title":"ethercat_driver_ros2: Configuring a CiA402 drive（0x6060，模式 8/9/10）","url":"https://icube-robotics.github.io/ethercat_driver_ros2/user_guide/config_cia402_drive.html"},{"title":"ros2_canopen: canopen_402_driver OperationMode 枚举","url":"https://raw.githubusercontent.com/ros-industrial/ros2_canopen/master/canopen_402_driver/include/canopen_402_driver/base.hpp"}],"as_of":"","related_ids":["canopen-cia-402-drive-profile","ethercat","servo-drive","torque-control","cascade-control","mit-mode"],"name":"Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)","alt":"周期同步位置 / 速度 / 力矩模式（CSP / CSV / CST）","abbr":"CSP / CSV / CST","aliases":["CSP","CSV","CST","Cyclic Synchronous Position Mode"],"one_liner":"Three CiA 402 modes where the host sends a new position, velocity, or torque setpoint once every communication cycle.","explanation":"CiA 402 is a CANopen device profile for motor drives defined by CAN in Automation (corresponding to the international standard IEC 61800-7-201), also carried over EtherCAT via CoE; it specifies the drive's state machine and a set of operating modes selected through object 0x6060. The cyclic synchronous modes are numbered 8, 9, and 10: in CSP, the host sends a target position every cycle and the drive closes position, velocity, and current loops internally; in CSV, only a target velocity is sent, with the position loop kept on the host side; in CST, only a target torque is sent, with the drive handling just the current (torque) loop, and everything else computed by the host. Unlike Profile Position mode (PP, number 1), in these three modes the trajectory is interpolated entirely by the host — the drive no longer generates one itself — which requires a fixed communication period and tight synchronization across axes. Robots doing impedance control, whole-body control, or letting a reinforcement-learning policy output torque directly tend to use CST; traditional industrial trajectory tracking tends to use CSP.","example":"Controlling a humanoid's joint modules over EtherCAT at 1 kHz: with the drive set to CST (0x6060 written as 10), the host computes and writes a torque every millisecond; set to CSP (written as 8), it writes an already-interpolated target position every millisecond instead. ROS 2's ethercat_driver_ros2 CiA 402 plugin supports switching among modes 8, 9, and 10. The MIT Mode common in legged robotics is a separate, proprietary CAN protocol that bundles target position, velocity, stiffness, damping, and feedforward torque into a single command.","related":["CANopen / CiA 402 Drive Profile","EtherCAT (Ethernet for Control Automation Technology)","Servo Drive (Motor Driver)","Torque Control","Cascade Control","MIT Mode"]},{"id":"streaming-servo-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Universal Robots 支持文章：servoj command","url":"https://www.universal-robots.com/articles/ur/programming/servoj-command/"},{"title":"睿尔曼 RM_API2 Python 接口（rm_movej_canfd / rm_movep_canfd 角度、位姿透传说明）","url":"https://github.com/RealManRobot/RM_API2/blob/main/Python/Robotic_Arm/rm_robot_interface.py"}],"as_of":"2026-09","related_ids":["movej-movel","rtde","control-frequency","trajectory-interpolation","real-time-control","visual-servoing"],"name":"Streaming Servo Control","alt":"透传控制","abbr":"","aliases":["Pass-through Mode","Servo Mode","servoJ","Streaming Servo Mode"],"one_liner":"A control mode where the host streams target joint angles or poses at a fixed rate and the controller just tracks them, without planning.","explanation":"Robot arms typically accept two kinds of motion commands. MoveJ/MoveL only specify an endpoint, and the controller plans the entire trajectory itself. Streaming servo control — also called servo mode or pass-through mode — instead has the upper-level computer continuously send a dense stream of target points at a fixed cycle, and the controller does essentially no trajectory planning of its own, just tracking whatever it receives. Chinese vendors often distinguish ‘joint-angle pass-through’ (sending joint angles) from ‘pose pass-through’ (sending an end-effector pose, with the controller solving inverse kinematics). It suits scenarios where the target changes online — visual servoing, teleoperation, and real-time action output from a VLA or reinforcement-learning policy. Because the controller no longer smooths anything for you, trajectory continuity and the stability of the send cycle become the upper computer's responsibility: cycle jitter or a large jump between adjacent points can cause vibration or even trigger a safety stop, so policy output is usually interpolated into a high-frequency, smooth trajectory before being streamed. Universal Robots' servoj is the most commonly cited interface of this kind.","example":"UR's servoj(q, a, v, t, lookahead_time, gain) must be called once every timestep: the e-Series uses t = 0.002 s (500 Hz), the CB3 uses 0.008 s; lookahead_time is set to 0.03–0.2 s for smoothing and to reduce overshoot; gain is set to 100–2000, with higher values responding faster but vibrating more easily. RealMan's rm_movej_canfd interface requires a streaming cycle no longer than 10 ms in its high-follow mode.","related":["MoveJ / MoveL","RTDE","Control Frequency","Trajectory Interpolation","Real-Time Control","Visual Servoing"]},{"id":"field-oriented-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Vector control (motor)","url":"https://en.wikipedia.org/wiki/Vector_control_(motor)"},{"title":"SimpleFOC docs: FOC theory","url":"https://docs.simplefoc.com/foc_theory"}],"as_of":"","related_ids":["permanent-magnet-synchronous-motor","brushless-dc-motor","servo-drive","mit-mode","magnetic-encoder","torque-constant"],"name":"Field-Oriented Control","alt":"磁场定向控制","abbr":"FOC","aliases":["FOC","Vector Control"],"one_liner":"Transforming three-phase current into a coordinate frame that rotates with the rotor, then separately controlling the flux-producing and torque-producing current.","explanation":"Field-oriented control, also called vector control, was proposed by K. Hasse at TU Darmstadt and F. Blaschke at Siemens between 1968 and the early 1970s, and is now the dominant drive method for permanent-magnet synchronous and brushless motors — robot joint-module drives almost universally use it. A motor's three-phase current is an AC quantity, hard to regulate directly. FOC first uses the Clarke transform to convert the three phases into a two-phase stationary frame (α, β), then the Park transform to rotate, by the rotor's electrical angle, into a frame that spins with the rotor, the d-q frame: d-axis current i_d produces flux aligned with the permanent magnet, and q-axis current i_q produces torque — both now DC quantities, each regulated with its own PI controller. For permanent-magnet motors, i_d is usually set to 0, and torque is approximately the torque constant K_t times i_q. This requires knowing the rotor's angle in real time, usually from an encoder, though sensorless observers can estimate it too. Compared to six-step square-wave commutation, FOC gives smoother torque; compared to V/f scalar control, it gives better dynamic performance.","example":"The quasi-direct-drive joints common in legged robots: the host sends target position, velocity, stiffness, damping, and feedforward torque via MIT Mode; the drive computes a desired torque τ, divides by K_t to get an i_q command with i_d set to 0, and then completes the Clarke/Park transforms, PI regulation, inverse transform, and PWM output of three-phase voltage inside a high-frequency current loop. The open-source SimpleFOC library implements this same pipeline.","related":["Permanent Magnet Synchronous Motor (PMSM)","Brushless DC Motor","Servo Drive (Motor Driver)","MIT Mode","Magnetic Encoder","Torque Constant (Kt)"]},{"id":"gravity-compensation","category":"control","sec":1,"tier":2,"sources":[{"title":"Modern Robotics 11.5: Force Control（τ = g(θ) + JᵀF_tip，含重力模型）","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-5-force-control/"},{"title":"Modern Robotics 11.4: Motion Control with Torque or Force Inputs (Part 3 of 3)","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-4-motion-control-with-torque-or-force-inputs-part-3-of-3/"},{"title":"libfranka robot.h（joint-level torque commands without gravity and friction）","url":"https://github.com/frankaemika/libfranka/blob/master/include/franka/robot.h"}],"as_of":"","related_ids":["zero-force-drag","proportional-derivative-control","feedforward-control","computed-torque-control","kinesthetic-teaching","friction-compensation"],"name":"Gravity Compensation","alt":"重力补偿","abbr":"","aliases":["Gravity Comp","Gravity Compensation Mode"],"one_liner":"Precomputing the joint torque needed to hold up the robot's own weight and adding it feedforward so the arm doesn't sag.","explanation":"Every joint in a robot arm must continuously supply torque to hold up the weight of the links beyond it, and the torque needed changes with pose — written as g(q), where q is the vector of joint angles. Gravity compensation computes g(q) in real time from each link's mass and center of mass, then adds it as a feedforward term (supplied before any error appears, rather than in response to one) to the control output. With plain PD control alone, the needed anti-gravity torque only shows up through leftover position error, so the end-effector sags slightly; adding gravity compensation removes that sag. If the controller outputs only g(q) with no position feedback, the arm can be pushed anywhere by hand and stays put — this ‘gravity compensation mode’ (free-drive) is found on collaborative arms and underlies kinesthetic teaching. Once a gripper or a heavy payload is attached, its mass and center of mass must be told to the controller, or the compensation will be off.","example":"On Franka arms, the libfranka interface expects the user's joint-torque commands to exclude gravity and friction terms — the controller adds those internally; its dynamics library's gravity() function needs the end-effector load's mass and center of mass to compute the gravity torque correctly.","related":["Zero-Force Drag","Proportional-Derivative Control","Feedforward Control","Computed Torque Control","Kinesthetic Teaching","Friction Compensation"]},{"id":"friction-compensation","category":"control","sec":1,"tier":3,"sources":[{"title":"Bona & Indri, Friction Compensation in Robotics: an Overview (CDC-ECC 2005)","url":"https://skoge.folk.ntnu.no/prost/proceedings/cdc-ecc05/pdffiles/papers/0934.pdf"},{"title":"BME Robot Applications, Chapter 8: Models of Friction","url":"https://www.mogi.bme.hu/TAMOP/robot_applications/ch07.html"}],"as_of":"","related_ids":["gravity-compensation","feedforward-control","zero-force-drag","coulomb-friction","static-friction-and-stribeck-effect","system-identification"],"name":"Friction Compensation","alt":"摩擦补偿","abbr":"","aliases":[],"one_liner":"Estimating the friction torque inside a joint and adding it into the control command ahead of time to cancel it out.","explanation":"The motors, gearboxes, and bearings inside a joint all have friction, which causes poor tracking at low speed and sticking on direction reversal (stick-slip); a 2005 survey by Bona and Indri notes this is especially critical for industrial robots. The most common model is τ_f = F_c·sgn(q̇) + F_v·q̇: F_c is Coulomb friction, constant in magnitude and opposing the direction of motion; F_v is the viscous friction coefficient, with torque proportional to joint velocity q̇; finer models add the Stribeck effect (extra static friction at breakaway) or dynamic models such as LuGre. The approach is to identify these parameters first, then add the estimated friction torque as a feedforward term on top of the motor command, or estimate it online with an observer. Because it actively cancels a modeled effect rather than correcting after an error appears (as an integral term does), it's distinct from that kind of feedback; zero-force drag, sensorless force estimation, and actuator modeling all depend on it.","example":"When a collaborative arm enters kinesthetic-teaching mode, the controller adds the estimated friction torque for each joint on top of gravity compensation, so pushing the arm feels light and smooth rather than jerky and uneven.","related":["Gravity Compensation","Feedforward Control","Zero-Force Drag","Coulomb Friction","Static Friction (Stiction) and Stribeck Effect","System Identification"]},{"id":"zero-force-drag","category":"control","sec":1,"tier":2,"sources":[{"title":"libfranka robot.h（setGuidingMode：Guiding mode can be enabled by pressing the two opposing buttons near the robot's flange）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"},{"title":"Universal Robots ROS 2 Driver: ur_controllers 文档（FreedriveModeController / ForceModeController）","url":"https://github.com/UniversalRobots/Universal_Robots_ROS2_Driver/blob/main/ur_controllers/doc/index.rst"},{"title":"Enabling Scalable Kinesthetic Teaching via Observer-based Hand-guiding with Active Support (arXiv:2608.10847)","url":"https://arxiv.org/abs/2608.10847"}],"as_of":"2026-09","related_ids":["kinesthetic-teaching","gravity-compensation","friction-compensation","admittance-control","joint-torque-sensor","demonstration-data"],"name":"Zero-Force Drag","alt":"零力拖动","abbr":"","aliases":["Free-Drive Mode","Hand Guiding Mode","Freedrive","Guiding Mode"],"one_liner":"Letting a robot arm cancel out its own gravity and friction so a person can push it around with barely any force.","explanation":"Zero-force drag is a hand-guiding mode found on collaborative arms — Universal Robots calls it Freedrive, Franka calls it Guiding mode. Once engaged, the controller computes in real time the torque needed to cancel its own gravity (gravity compensation) and joint friction (friction compensation), so the arm hangs as if weightless and a light push is enough to move it. It's implemented roughly three ways: arms with joint torque sensors apply the compensation directly in torque mode; many collaborative arms without torque sensors estimate external force from motor current instead; traditional industrial arms mount a 6-axis force sensor at the wrist and use admittance control to convert the measured force into motion. Its main use is kinesthetic teaching: a person physically guides the arm through a motion by hand, and the controller records the trajectory or saves it as demonstration data. ‘Zero-force’ refers to the force a person needs to apply being close to zero — not that the motors output no force.","example":"On Franka arms, holding two opposing buttons near the flange engages guiding mode, and libfranka's setGuidingMode can release only some directions (e.g., allowing translation only, locking rotation); UR's ROS 2 driver provides a freedrive_mode_controller that needs a continuous stream of True messages, exiting automatically if none arrive for 1 second by default.","related":["Kinesthetic Teaching","Gravity Compensation","Friction Compensation","Admittance Control","Joint Torque Sensor","Demonstration Data"]},{"id":"system-identification","category":"control","sec":1,"tier":2,"sources":[{"title":"Wikipedia: System identification","url":"https://en.wikipedia.org/wiki/System_identification"},{"title":"Learning agile and dynamic motor skills for legged robots（执行器网络）","url":"https://arxiv.org/abs/1901.08652"}],"as_of":"","related_ids":["dynamic-parameter-identification","actuator-modeling","sim-to-real-gap","domain-randomization","real-to-sim","inertial-parameters"],"name":"System Identification","alt":"系统辨识","abbr":"SysID","aliases":["SysID","Parameter Identification"],"one_liner":"Using measured input-output data to work backward to a system's mathematical model and parameters.","explanation":"System identification is the discipline of building a mathematical model of a dynamic system from measured input and output data using statistical methods, and it also covers how to design experiments that gather sufficiently informative data. It is grouped into three categories by how much prior knowledge is used: white-box (derived entirely from physical laws), gray-box (known structure with parameters fit from data, e.g., identifying link mass, inertia, and joint friction), and black-box (fitting only the input-output relationship, e.g., with a neural network). In embodied AI, it is one of the main tools for narrowing the sim-to-real gap: parameters measured on the real robot are written back into the simulator, or a model is learned to fill in whatever the simulator is missing. Hwangbo et al. (2019) first identified ANYmal's physical parameters, then used real-robot data to train an actuator network modeling the motor and low-level software dynamics. It complements domain randomization: identification makes the simulation more accurate, while randomization makes the policy robust to whatever error remains.","example":"Apply a sinusoidal sweep of torque to one robot joint, record its angle and angular velocity, and fit the moment of inertia, viscous friction, and Coulomb friction coefficients by least squares; writing these back into a MuJoCo model aligns the simulated joint trajectory with the real robot under the same commands.","related":["Dynamic Parameter Identification","Actuator Modeling (Actuator Network)","Sim-to-Real Gap (Reality Gap)","Domain Randomization","Real-to-Sim","Inertial Parameters"]},{"id":"computed-torque-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Computed torque control","url":"https://en.wikipedia.org/wiki/Computed_torque_control"},{"title":"Pinocchio rnea.hpp（computes the inverse dynamics, aka the joint torques）","url":"https://github.com/stack-of-tasks/pinocchio/blob/master/include/pinocchio/algorithm/rnea.hpp"}],"as_of":"","related_ids":["inverse-dynamics","feedback-linearization","feedforward-control","gravity-compensation","operational-space-control","recursive-newton-euler-algorithm"],"name":"Computed Torque Control","alt":"计算力矩控制","abbr":"CTC","aliases":["CTC","Inverse Dynamics Control"],"one_liner":"Using a dynamics model to compute the torque needed, canceling nonlinearities, then applying PD control on top.","explanation":"Computed torque control is a classic model-based method for controlling arm motion, also called inverse dynamics control. Arm dynamics is written τ = M(q)q̈ + C(q,q̇)q̇ + g(q), where M is the mass matrix, the C term corresponds to Coriolis and centrifugal forces, and g is gravity. The control law is τ = M̂(q)(q̈_d + K_d·ė + K_p·e) + Ĉq̇ + ĝ, where q̈_d is the desired acceleration, e the error between desired and actual joint angle, and the hatted terms are model estimates. When the model is accurate, the nonlinear terms cancel out and the error satisfies ë + K_d·ė + K_p·e = 0 — each joint becomes an independent, linear, second-order system, with gains chosen for the desired response. It's a textbook application of feedback linearization to robotics. Compared to ‘PD plus gravity compensation,’ it also compensates inertia and Coriolis effects, tracking more accurately at high speed; the cost is dependence on accurate dynamic parameters, with degraded performance when the model is off — which led to adaptive and robust variants. The task-space counterpart is operational space control.","example":"Using Pinocchio's rnea (Recursive Newton-Euler Algorithm) function, passing the current q, q̇, and ‘desired acceleration plus PD correction term’ as the acceleration input computes, in a single call, exactly the joint torques computed torque control needs.","related":["Inverse Dynamics","Feedback Linearization","Feedforward Control","Gravity Compensation","Operational Space Control","Recursive Newton-Euler Algorithm"]},{"id":"feedback-linearization","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Feedback linearization","url":"https://en.wikipedia.org/wiki/Feedback_linearization"}],"as_of":"","related_ids":["computed-torque-control","inverse-dynamics","hybrid-zero-dynamics","operational-space-control","mass-matrix","lyapunov-stability"],"name":"Feedback Linearization","alt":"反馈线性化","abbr":"","aliases":["Exact Linearization","Input-Output Linearization"],"one_liner":"Using state feedback to exactly cancel nonlinear terms, turning the system into a linear one before designing a controller for it.","explanation":"Feedback linearization is a fundamental method in nonlinear control: design a control law u = a(x) + b(x)·v, where a and b are computed from the system model to exactly cancel the nonlinear terms; combined with a change of coordinates, this turns the relationship from the new input v to the output into something simple and linear (commonly a chain of integrators), after which PD control, pole placement, or other linear methods can design v. It differs from the common approach of ‘Taylor-expanding around an operating point’ (Jacobian linearization): that is only an approximation valid near the operating point, while feedback linearization is an exact transformation, valid over a much larger range — but only if the model is accurate; with model error, the cancellation is imperfect and robustness suffers. Computed torque control for robot arms is a special case of this method. When only the output is made linear, whatever internal dynamics remain hidden are called the zero dynamics, and these must be stable; hybrid zero dynamics (HZD) control for bipedal walking is built on exactly this kind of input-output linearization.","example":"Arm dynamics M(q)q̈ + C(q, q̇)q̇ + g(q) = τ (M the mass matrix, the C term Coriolis and centrifugal forces, g gravity, τ joint torque). Substituting τ = M(q)v + C(q, q̇)q̇ + g(q) gives q̈ = v — each joint becomes an independent double integrator; then setting v = q̈_d + K_d(q̇_d − q̇) + K_p(q_d − q) makes the tracking error converge as a linear second-order system.","related":["Computed Torque Control","Inverse Dynamics","Hybrid Zero Dynamics","Operational Space Control","Mass Matrix","Lyapunov Stability"]},{"id":"disturbance-observer","category":"control","sec":1,"tier":3,"sources":[{"title":"Sariyildiz, Oboe, Ohnishi: Disturbance Observer-based Robust Control and Its Applications: 35th Anniversary Overview (arXiv:1902.09032)","url":"https://arxiv.org/abs/1902.09032"},{"title":"Hyungbo Shim: Disturbance Observer (arXiv:2101.02859)","url":"https://arxiv.org/abs/2101.02859"}],"as_of":"","related_ids":["state-observer","generalized-momentum-observer","active-disturbance-rejection-control","sensorless-force-estimation","robust-control","friction-compensation"],"name":"Disturbance Observer","alt":"扰动观测器","abbr":"DOB","aliases":["DOB","DOb","DOBC","Disturbance-Observer-Based Control"],"one_liner":"Using a nominal model to work backward to unknown disturbances like friction and payload, then canceling them out in the control input.","explanation":"The disturbance observer is one of the most widely used tools for robust motion control, proposed by Kouhei Ohnishi and colleagues around 1983. The idea is to lump friction, load changes, model-parameter error, and external forces into one ‘lumped disturbance’ and work it out from a simple nominal model. For a motor, nominally J·θ̈ = K_t·i (J the moment of inertia, θ̈ angular acceleration, K_t the torque constant, i current); when the measured acceleration and current don't match, the mismatch is the disturbance torque, which is then passed through a low-pass filter to get an estimate (the cutoff frequency sets the estimation bandwidth — too high amplifies noise) and added back into the current command to cancel it out. This makes the plant behave much closer to its nominal model, so an outer PD loop only needs to handle tracking performance while the DOB handles robustness — the two can be tuned independently. The momentum observer used for robot collision detection and the extended state observer in active disturbance rejection control share a similar idea.","example":"Adding a DOB to a joint motor, wrapped by an outer PD position loop: when the arm's end-effector goes from unloaded to holding a weight, the DOB quickly estimates the extra load torque and adds it into the current command, leaving position error largely unaffected. Subtracting the known gravity and friction terms from the estimate leaves an external-force estimate, a common way to do force control and collision detection without a force sensor.","related":["State Observer","Generalized Momentum Observer","Active Disturbance Rejection Control","Sensorless Force Estimation","Robust Control","Friction Compensation"]},{"id":"active-disturbance-rejection-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Active disturbance rejection control","url":"https://en.wikipedia.org/wiki/Active_disturbance_rejection_control"},{"title":"Han J. From PID to Active Disturbance Rejection Control. IEEE Transactions on Industrial Electronics, 2009","url":"https://doi.org/10.1109/TIE.2008.2011621"},{"title":"Herbst G. Transfer Function Analysis and Implementation of Active Disturbance Rejection Control (arXiv:2011.01044)","url":"https://arxiv.org/abs/2011.01044"}],"as_of":"","related_ids":["proportional-integral-derivative-control","disturbance-observer","state-observer","robust-control","adaptive-control"],"name":"Active Disturbance Rejection Control","alt":"自抗扰控制","abbr":"ADRC","aliases":["ADRC","Linear ADRC","LADRC","Extended State Observer Control"],"one_liner":"Lumping model error and external disturbances into one ‘total disturbance,’ estimating it in real time, and canceling it out.","explanation":"Active disturbance rejection control was developed by Chinese scholar Jingqing Han in the 1990s and introduced systematically in English in his 2009 paper ‘From PID to Active Disturbance Rejection Control.’ Rather than pursuing an accurate model, it needs only the system's order and a rough estimate of the control gain b₀; it lumps unmodeled dynamics, parameter changes, and external forces into a single ‘total disturbance’ f, treats that as an extra state, and estimates it in real time with an extended state observer (ESO, an algorithm that estimates internal states from inputs and outputs), then subtracts it from the control signal: u = (u₀ − f̂) / b₀. After cancellation, the plant behaves approximately like a chain of integrators, which a simple PD can then control well. The original version also includes a tracking differentiator and nonlinear feedback; the more commonly used linear ADRC only needs the observer bandwidth and controller bandwidth tuned, a bandwidth-based tuning approach that traces to Zhiqiang Gao's 2003 work. It resembles disturbance-observer approaches, though those usually need a nominal model; unlike adaptive control, it doesn't estimate specific parameters.","example":"When a joint motor's load or friction changes with temperature, PID often needs retuning; ADRC instead folds all such changes into the total disturbance for the ESO to estimate and compensate, and the linear version mainly needs just two parameters tuned — observer bandwidth ω_o and controller bandwidth ω_c.","related":["Proportional-Integral-Derivative Control","Disturbance Observer","State Observer","Robust Control","Adaptive Control"]},{"id":"adaptive-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Adaptive control","url":"https://en.wikipedia.org/wiki/Adaptive_control"},{"title":"Slotine J.-J. E., Li W. On the Adaptive Control of Robot Manipulators. IJRR, 1987","url":"https://doi.org/10.1177/027836498700600303"}],"as_of":"","related_ids":["robust-control","computed-torque-control","dynamic-parameter-identification","system-identification","rapid-motor-adaptation","active-disturbance-rejection-control"],"name":"Adaptive Control","alt":"自适应控制","abbr":"","aliases":["Model Reference Adaptive Control","MRAC","Self-Tuning Control"],"one_liner":"Estimating unknown or changing parameters online while running, and automatically adjusting the controller accordingly.","explanation":"Adaptive control means a controller estimates unknown or time-varying parameters online from the actual response while running, and adjusts itself accordingly; its foundation is parameter estimation. Common forms include model reference adaptive control (MRAC, which drives the system to track an ideal reference model's response) and self-tuning control, further split into direct methods (adjusting controller parameters directly) and indirect methods (estimating plant parameters first, then computing the controller). It differs from robust control: robust control fixes a range of parameter variation in advance and uses one controller to withstand the worst case, while adaptive control needs no such range and instead corrects itself as it goes. A classic in robotics is the 1987 algorithm by Slotine and Li: PD feedback plus full dynamics feedforward, with unknown parameters such as payload estimated online; it exploits the structure of arm dynamics, requiring no joint acceleration measurement and no inversion of the estimated mass matrix. Rapid motor adaptation (RMA) in reinforcement-learning locomotion does something similar with learned methods.","example":"An arm picks up a workpiece of unknown mass; a controller using the old model shows tracking error, while an adaptive controller updates its online estimate of the payload mass from that error, and the error shrinks over repeated motions.","related":["Robust Control","Computed Torque Control","Dynamic Parameter Identification","System Identification","Rapid Motor Adaptation","Active Disturbance Rejection Control"]},{"id":"robust-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Robust control","url":"https://en.wikipedia.org/wiki/Robust_control"},{"title":"Wikipedia: H-infinity methods in control theory","url":"https://en.wikipedia.org/wiki/H-infinity_methods_in_control_theory"}],"as_of":"","related_ids":["adaptive-control","sliding-mode-control","lyapunov-stability","disturbance-observer","active-disturbance-rejection-control","domain-randomization"],"name":"Robust Control","alt":"鲁棒控制","abbr":"","aliases":["Robust Control Theory"],"one_liner":"Control design that guarantees stability and performance despite model inaccuracy, bounded parameter variation, and external disturbance.","explanation":"Robust control is a branch of control theory that took shape from the late 1970s onward. Its core idea is to admit model error up front, at design time: given a bounded range of uncertainty — say, payload mass within some interval, or unmodeled joint flexibility or friction — a single fixed controller is designed that guarantees stability and a performance floor across the entire range. Representative methods include H∞ control, pioneered by George Zames and others, which minimizes the worst-case effect of disturbances on the output, and sliding mode control. Robust control differs from adaptive control in that its controller parameters are fixed, relying on built-in margin to withstand uncertainty, whereas adaptive control estimates parameters online and adjusts the controller while running. The trade-off is that robust designs tend to be conservative, and still need a roughly accurate model to start from. In robotics, payload variation, joint friction, and varying ground stiffness are typical sources of uncertainty; domain randomization in reinforcement learning, which trains policies to be insensitive to parameter variation, pursues a similar kind of robustness.","example":"An arm must handle a workpiece weighing anywhere from 0 to 3 kg. Robust control designs a single controller up front assuming the load could be any value in that range, guaranteeing stability and bounded tracking error in every case. Adaptive control instead estimates the actual mass online after the object is grasped and adjusts its control parameters accordingly.","related":["Adaptive Control","Sliding Mode Control","Lyapunov Stability","Disturbance Observer","Active Disturbance Rejection Control","Domain Randomization"]},{"id":"sliding-mode-control","category":"control","sec":1,"tier":3,"sources":[{"title":"Sliding mode control - Wikipedia","url":"https://en.wikipedia.org/wiki/Sliding_mode_control"},{"title":"Variable structure control - Wikipedia","url":"https://en.wikipedia.org/wiki/Variable_structure_control"}],"as_of":"","related_ids":["robust-control","adaptive-control","lyapunov-stability","feedback-linearization","active-disturbance-rejection-control","disturbance-observer"],"name":"Sliding Mode Control","alt":"滑模控制","abbr":"SMC","aliases":["SMC","Variable Structure Control"],"one_liner":"A control method that forces the state onto a designed ‘sliding surface’ via switching, then slides it to the target with strong robustness.","explanation":"Sliding mode control belongs to the broader family of variable structure control, first studied in the early 1950s by Soviet researchers including Emelyanov and later systematized by Vadim Utkin and others. The design starts with a sliding surface — for instance, using tracking error e to construct s = ė + λe (λ > 0): once s = 0, the error decays to zero as e(0)·exp(−λt). The control law combines an equivalent control term (which holds the state on the surface, computed from the nominal model) with a switching term −k·sign(s) that forces the state back onto the surface whenever it strays. The system first reaches the sliding surface, then slides along it toward the target. As long as k exceeds the upper bound of the disturbance, any disturbance or model error entering through the control channel doesn't change the behavior during the sliding phase, which makes this approach more robust than PID or computed-torque control. The cost is chattering: actuator delay and discrete sampling make the sign function switch rapidly back and forth, producing high-frequency vibration, usually mitigated with a boundary layer (replacing sign with a saturation function) or higher-order sliding modes such as super-twisting.","example":"For position tracking on an arm joint with an unknown payload: take s = ė + λe and torque τ = τ_eq − k·sat(s/φ). τ_eq is computed from the nominal model, k is set large enough to cover the upper bound of friction and payload error, sat is the saturation function, and φ is the boundary-layer thickness used to suppress chattering.","related":["Robust Control","Adaptive Control","Lyapunov Stability","Feedback Linearization","Active Disturbance Rejection Control","Disturbance Observer"]},{"id":"task-space-control","category":"control","sec":2,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), 11.3.3 / 11.4.3 Task-Space Motion Control","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Khatib, A unified approach for motion and force control of robot manipulators: The operational space formulation (IEEE J. Robotics & Automation, 1987)","url":"https://doi.org/10.1109/JRA.1987.1087068"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion（附录 Franka Robot Station）","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["task-space","joint-space-control","operational-space-control","jacobian-pseudoinverse","inverse-kinematics","cartesian-impedance-control"],"name":"Task-Space Control","alt":"任务空间控制","abbr":"","aliases":["Cartesian-Space Control","End-Effector Control","Cartesian Control"],"one_liner":"Computing error and issuing commands based on the end-effector's position and orientation, rather than individual joint angles.","explanation":"Task-space control writes the goal in terms of the end-effector's (gripper or tool) pose rather than each joint's angle: given a desired end-effector trajectory, it computes the deviation between the actual and desired end-effector state, then converts that into joint commands. There are two common routes for the conversion. At the velocity level, the pseudoinverse of the Jacobian matrix J (which maps joint velocity to end-effector velocity) is used: θ̇ = J⁺(θ)·V, where θ̇ is joint velocity and V the desired end-effector velocity. At the torque level, τ = Jᵀ(θ)·F converts a desired end-effector force F into joint torque τ; Khatib's 1987 operational space control belongs to this category and additionally accounts for the end-effector's effective inertia. It matches task descriptions like ‘move the cup here’ more naturally than joint-space control, at the cost of having to handle singularities (poses where some directions suddenly can't move) and redundant degrees of freedom.","example":"In Diffusion Policy's Franka setup, the policy outputs a desired end-effector pose at 10 Hz, and a mid-level controller running at about 1 kHz solves a quadratic program each step to find the joint velocities that best track that target end-effector velocity, integrating them into joint positions handed to the arm's built-in joint controller; obstacle avoidance and joint limits are written as constraints, and any leftover degrees of freedom are used in the null space.","related":["Task Space","Joint-Space Control","Operational Space Control","Jacobian Pseudoinverse","Inverse Kinematics (IK)","Cartesian Impedance Control"]},{"id":"operational-space-control","category":"control","sec":2,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics（§8.6 Dynamics in the Task Space；引 Khatib 1987, IEEE J. Robotics and Automation 3(1):43–53）","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"robosuite Docs: Controllers（OSC_POSE）","url":"https://robosuite.ai/docs/modules/controllers.html"}],"as_of":"","related_ids":["task-space-control","impedance-control","jacobian-matrix","null-space-control","mass-matrix","whole-body-control"],"name":"Operational Space Control","alt":"操作空间控制","abbr":"OSC","aliases":["OSC","Operational Space Formulation","OSC_POSE"],"one_liner":"Computing desired end-effector acceleration and force directly, then converting them to joint torque through the robot's dynamics.","explanation":"Proposed by Stanford's Oussama Khatib in 1987. It rewrites arm dynamics in the ‘operational space’ at the end-effector: F = Λ(q)·ẍ + η, where ẍ is end-effector acceleration, Λ the effective inertia matrix felt at the end-effector, η collects Coriolis, gravity, and related terms, and F is the force and torque to apply at the end-effector. The controller first computes a desired ẍ from pose error with a PD law, substitutes it in to get F, then converts to joint torque via τ = Jᵀ·F (J being the Jacobian matrix). The advantage is being able to set stiffness and damping directly at the end-effector while compensating for the arm's own inertia. Unlike ‘inverse kinematics plus joint position control,’ it outputs torque directly, so it needs a reasonably accurate dynamics model. The OSC_POSE controller in the robosuite simulation framework implements exactly this.","example":"Controlling a Panda arm in robosuite with OSC_POSE, the policy only outputs an incremental end-effector displacement and an axis-angle rotation increment each step, and the controller internally converts these into torques for the 7 joints.","related":["Task-Space Control","Impedance Control","Jacobian Matrix","Null-Space Control","Mass Matrix","Whole-Body Control"]},{"id":"force-control","category":"control","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Force control","url":"https://en.wikipedia.org/wiki/Force_control"}],"as_of":"","related_ids":["impedance-control","admittance-control","hybrid-force-position-control","six-axis-force-torque-sensor","joint-torque-sensor","contact-rich-manipulation"],"name":"Force Control","alt":"力控","abbr":"","aliases":["Force-Based Control"],"one_liner":"Controlling how much force a robot applies to an object, rather than just where it moves.","explanation":"Force control means controlling how much force a robot applies to an object or its environment, rather than simply driving it to a target position. Pure position control responds to any deviation by pushing harder and harder to reach the commanded position, so if it meets a surface higher than expected it can crush the workpiece; force control instead keeps the contact force near a set value. There are two implementation approaches. Direct force control makes the target force itself the feedback setpoint, typically hybrid force/position control, which splits the directions so some axes control force and others control position. Indirect force control means impedance control and admittance control, which set a stiffness, damping, and inertia to decide how soft the robot feels when pushed. Force control requires knowing the actual force applied, usually from a wrist-mounted six-axis force-torque sensor, joint torque sensors, or motor current estimates. Contact-rich tasks — sanding, assembly, wiping a table, plugging in a cable — all depend on it.","example":"A robot arm wiping a whiteboard: under pure position control, a height difference of just a few millimeters would make it either miss the board or press too hard; switching to force control and setting a fixed pressure in the direction perpendicular to the board lets it keep contact and wipe evenly even where the surface isn't perfectly flat.","related":["Impedance Control","Admittance Control","Hybrid Force/Position Control","Six-Axis Force/Torque Sensor","Joint Torque Sensor","Contact-rich Manipulation"]},{"id":"compliance-control","category":"control","sec":2,"tier":2,"sources":[{"title":"Wikipedia: Force control（passive / active compliance）","url":"https://en.wikipedia.org/wiki/Force_control"},{"title":"Wikipedia: Impedance control","url":"https://en.wikipedia.org/wiki/Impedance_control"},{"title":"Hogan, Impedance Control: An Approach to Manipulation, Part I (1985)","url":"https://doi.org/10.1115/1.3140702"}],"as_of":"","related_ids":["impedance-control","admittance-control","hybrid-force-position-control","compliance","remote-center-compliance-device","series-elastic-actuator"],"name":"Compliance Control","alt":"柔顺控制","abbr":"","aliases":["Compliant Control","Active Compliance Control"],"one_liner":"Control methods that let a robot yield to some set degree instead of fighting back when it contacts the environment.","explanation":"Compliance control is the umbrella term for methods that let a robot yield to external force by a designed amount when it contacts the environment. A purely position-controlled arm is very stiff: in a peg-in-hole task, even a small misalignment can jam the parts or crush them. Compliance comes in two forms. Passive compliance comes from the mechanical structure itself — such as a remote center compliance (RCC) device mounted on the wrist, or the spring inside a series elastic actuator — and needs no sensing or computation, but once its characteristics are built in, they can't be changed. Active compliance comes from measuring force or torque and feeding it back through control, commonly in the form of impedance control (making the robot behave like a mass-spring-damper system with tunable parameters), admittance control (correcting a position command after measuring force), or hybrid force/position control (controlling force along some directions and position along others), with parameters that can be adjusted per task. Hogan pointed out in 1985 that controlling position alone or force alone isn't enough — what should be controlled is the dynamic relationship between the two — and this is the theoretical basis for active compliance. Sanding, assembly, wiping, and human-robot collaboration all depend on it.","example":"During peg-in-hole assembly, the robot arm keeps position control along the insertion direction but lowers the stiffness in the perpendicular directions: if the peg catches the edge of the hole, it gets pushed sideways and slides on into the hole, instead of jamming or crushing the part.","related":["Impedance Control","Admittance Control","Hybrid Force/Position Control","Compliance","Remote Center Compliance (RCC) Device","Series Elastic Actuator (SEA)"]},{"id":"impedance-control","category":"control","sec":2,"tier":2,"sources":[{"title":"Hogan, Impedance Control: An Approach to Manipulation, Part I—Theory, J. Dyn. Sys., Meas., Control 107 (1985)","url":"https://doi.org/10.1115/1.3140702"},{"title":"Wikipedia: Impedance control","url":"https://en.wikipedia.org/wiki/Impedance_control"},{"title":"libfranka examples/cartesian_impedance_control.cpp","url":"https://github.com/frankaemika/libfranka/blob/master/examples/cartesian_impedance_control.cpp"}],"as_of":"","related_ids":["admittance-control","cartesian-impedance-control","joint-impedance-control","variable-impedance-control","hybrid-force-position-control","mass-spring-damper-system"],"name":"Impedance Control","alt":"阻抗控制","abbr":"","aliases":[],"one_liner":"Making a robot compliant when it meets force, like a spring and damper, instead of rigidly forcing its way to a target position.","explanation":"Systematically developed by MIT's Neville Hogan in 1984–1985. Pure position control slams rigidly into a table or a person; pure force control can't track a trajectory in free space. Impedance control instead prescribes a relationship between force and motion error, commonly written F = K(x_d − x) + D(ẋ_d − ẋ): x_d is the desired position, x the actual position, K the stiffness (newtons of force per meter of deviation), and D the damping (which suppresses oscillation). A large K makes the robot stiff and accurate; a small K makes it soft, so hitting an obstacle produces only limited force. It generally requires a robot that can directly command joint torque; admittance control runs the opposite way — measuring force first, then computing a position correction — and suits industrial arms that only accept position commands. Impedance control can be implemented in joint space or Cartesian space, and is common in wiping, peg-in-hole assembly, and human-robot collaboration.","example":"libfranka's built-in Cartesian impedance example sets translational stiffness to 150 N/m and rotational stiffness to 10 N·m/rad, with damping set to 2√K (critical damping); once running, you can push the arm by hand, and it springs back to its original pose when released.","related":["Admittance Control","Cartesian Impedance Control","Joint Impedance Control","Variable Impedance Control","Hybrid Force/Position Control","Mass-Spring-Damper System"]},{"id":"admittance-control","category":"control","sec":2,"tier":2,"sources":[{"title":"ros2_controllers: Admittance Controller","url":"https://control.ros.org/master/doc/ros2_controllers/admittance_controller/doc/userdoc.html"},{"title":"Wikipedia: Impedance control","url":"https://en.wikipedia.org/wiki/Impedance_control"},{"title":"Hogan, Impedance Control: An Approach to Manipulation, Part I (1985)","url":"https://doi.org/10.1115/1.3140702"}],"as_of":"","related_ids":["impedance-control","compliance-control","six-axis-force-torque-sensor","zero-force-drag","force-control","position-control"],"name":"Admittance Control","alt":"导纳控制","abbr":"","aliases":["Admittance Controller"],"one_liner":"Measures the external force applied, computes the resulting motion from a set mass-damper-spring relationship, and hands that to the position loop.","explanation":"Admittance control is an interaction-control method that makes a robot yield to external force, and it's theoretically grounded in the duality between impedance and admittance from Hogan's 1985 paper on impedance control. Its input is force and its output is motion: a wrist-mounted six-axis force-torque sensor measures the external force F, and the controller computes how the robot should move from F = M·a + D·v + K·(x − x_d) (M, D, K are the virtual mass, damping, and stiffness, and x_d is the originally intended target position), then hands the corrected position or velocity command off to the robot's own built-in position loop. Impedance control runs in the opposite direction: it takes a position deviation as input and outputs a force or torque, which requires the joints to be able to control torque directly. Admittance control is therefore well suited to industrial arms that only expose a position interface, have high gear ratios, and are inherently stiff — but it can oscillate against very stiff environments. ROS 2's ros2_controllers package includes an admittance_controller, which can be used for zero-force dragging.","example":"Mount a six-axis force-torque sensor on an industrial arm's end effector, turn on admittance control, and set the stiffness K to zero: pushing the end effector by hand moves the robot along with the push, and it settles to a stop under the damping term once released — one way of implementing hand-guided teaching.","related":["Impedance Control","Compliance Control","Six-Axis Force/Torque Sensor","Zero-Force Drag","Force Control","Position Control"]},{"id":"hybrid-force-position-control","category":"control","sec":2,"tier":2,"sources":[{"title":"Raibert & Craig, Hybrid Position/Force Control of Manipulators, J. Dyn. Sys., Meas., Control 103 (1981)","url":"https://doi.org/10.1115/1.3139652"},{"title":"Modern Robotics 11.6: Hybrid Motion-Force Control","url":"https://modernrobotics.northwestern.edu/nu-gm-book-resource/11-6-hybrid-motion-force-control/"}],"as_of":"","related_ids":["impedance-control","force-control","admittance-control","task-space-control","constant-force-control","contact-rich-manipulation"],"name":"Hybrid Force/Position Control","alt":"力位混合控制","abbr":"","aliases":["Hybrid Position/Force Control","Hybrid Motion-Force Control"],"one_liner":"In contact tasks, splitting directions apart — controlling force where the robot is constrained, position where it is free.","explanation":"Proposed by Raibert and Craig in 1981. When a robot touches an object, the environment constrains motion along some directions — for example, an eraser wiping a whiteboard cannot push through the board's surface, but can slide freely within it. The method uses a selection matrix S, with diagonal entries of 1 or 0, to split the six spatial directions (defined in a task-specific coordinate frame) into two groups: the directions S selects track a desired position, and the rest (I − S) track a desired contact force; the two outputs are combined and converted into joint torques. It requires knowing the constraint directions in advance, and getting the surface normal wrong can generate very large forces. Impedance control, by contrast, doesn't split directions at all — it specifies a relationship between force and position error along every direction at once. Note that motor manufacturers' ‘hybrid force/position mode’ usually refers to a single joint's MIT-style command format, an unrelated concept.","example":"Erasing a whiteboard: the x and y directions within the board's plane are position-controlled, moving the eraser along a path; the z direction perpendicular to the board is force-controlled, holding a constant pressing force.","related":["Impedance Control","Force Control","Admittance Control","Task-Space Control","Constant Force Control (Force Tracking)","Contact-rich Manipulation"]},{"id":"constant-force-control","category":"control","sec":2,"tier":3,"sources":[{"title":"Seraji H., Colbaugh R. Force Tracking in Impedance Control. IJRR, 1997","url":"https://doi.org/10.1177/027836499701600107"},{"title":"Universal Robots ROS 2 Driver: ForceModeController 文档","url":"https://github.com/UniversalRobots/Universal_Robots_ROS2_Driver/blob/main/ur_controllers/doc/index.rst"}],"as_of":"","related_ids":["force-control","impedance-control","admittance-control","hybrid-force-position-control","six-axis-force-torque-sensor","contact-rich-manipulation"],"name":"Constant Force Control (Force Tracking)","alt":"恒力控制","abbr":"","aliases":["Force Tracking Control","Constant Force Tracking","Constant-Force Polishing"],"one_liner":"Keeping a robot's end-effector applying a set, constant pressure against a contact surface — for example, holding 10 newtons.","explanation":"Constant force control, also called force tracking, keeps the contact force between the end-effector and the environment following a set value, and is common in polishing, buffing, deburring, wiping, and ultrasonic scanning — tasks needing steady pressure. The difficulty is that the environment's position and stiffness are often not known precisely: under pure position control, a workpiece positioned just 1 millimeter off can cause a large difference in the force pressed against a stiff surface. Common approaches include: PI control directly on the force error (explicit force control); hybrid force/position control, with force controlled in the normal direction and position in the tangential direction; and wrapping an outer force loop around impedance or admittance control that adjusts the reference position based on force error — Seraji and Colbaugh (1997) achieved tracking of a desired force this way using adaptive methods even when the environment's stiffness and position were both unknown. It differs from plain impedance control in that impedance control only specifies ‘deviate this much, output that much force,’ with the actual contact force depending on the environment, whereas constant force control targets the force itself. It usually requires a 6-axis force sensor or joint torque sensors to measure force.","example":"Universal Robots arms' Force Mode can set the z-axis as a compliant axis in the task frame with a target force (e.g., 10 N downward), while x and y continue following the programmed position path as usual, maintaining constant pressure while moving across a surface; UR's documentation specifically notes this is not admittance control, and any direction set to force control overrides motion commands along that direction.","related":["Force Control","Impedance Control","Admittance Control","Hybrid Force/Position Control","Six-Axis Force/Torque Sensor","Contact-rich Manipulation"]},{"id":"joint-impedance-control","category":"control","sec":2,"tier":3,"sources":[{"title":"Impedance control - Wikipedia","url":"https://en.wikipedia.org/wiki/Impedance_control"},{"title":"libfranka robot.h（setJointImpedance / ControllerMode::kJointImpedance）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"},{"title":"franka_ros joint_impedance_example_controller.cpp","url":"https://github.com/frankaemika/franka_ros/blob/develop/franka_example_controllers/src/joint_impedance_example_controller.cpp"}],"as_of":"","related_ids":["impedance-control","cartesian-impedance-control","admittance-control","torque-control","proportional-derivative-control","gravity-compensation"],"name":"Joint Impedance Control","alt":"关节阻抗控制","abbr":"","aliases":["Joint-Space Impedance Control"],"one_liner":"Making every joint of a robot behave like a spring and damper, yielding compliantly when pushed instead of resisting rigidly.","explanation":"Impedance control was systematically developed by MIT's Neville Hogan in his 1984–1985 papers, with the goal not being to rigidly hold a position or force, but to specify how ‘soft’ the robot should feel under external force. Joint impedance control sets this spring-damper behavior at each individual joint, with a typical control law τ = K(q_d − q) + D(q̇_d − q̇) + gravity and Coriolis compensation, where τ is joint torque, q and q_d the actual and desired joint angle, K the stiffness, and D the damping. It looks like PD control in form, but requires the joint to output torque directly and compensate for dynamics, and K is deliberately set low so the robot gets pushed away by contact with a person or object rather than fighting back. The closely related Cartesian impedance control instead sets the spring on the end-effector's x, y, z, and orientation, which more directly specifies which direction the hand should feel soft in; admittance control runs the opposite direction — measuring external force first, then computing how to move — and suits inherently stiff, position-controlled arms.","example":"Franka arms default to joint impedance mode built in; libfranka's setJointImpedance sets each joint's stiffness, and franka_ros's example controller computes each joint's command torque as ‘Coriolis compensation + k·(q_d − q) + d·(q̇_d − q̇),’ letting the arm trace a circle while staying compliant.","related":["Impedance Control","Cartesian Impedance Control","Admittance Control","Torque Control","Proportional-Derivative Control","Gravity Compensation"]},{"id":"cartesian-impedance-control","category":"control","sec":2,"tier":3,"sources":[{"title":"Hogan N. Impedance Control: An Approach to Manipulation: Part II—Implementation. ASME JDSMC, 1985","url":"https://doi.org/10.1115/1.3140713"},{"title":"franka_ros: cartesian_impedance_example_controller.cpp","url":"https://github.com/frankaemika/franka_ros/blob/develop/franka_example_controllers/src/cartesian_impedance_example_controller.cpp"},{"title":"Mayr, Salt-Ducaju. A C++ Implementation of a Cartesian Impedance Controller for Robotic Manipulators (arXiv:2212.11215)","url":"https://arxiv.org/abs/2212.11215"}],"as_of":"","related_ids":["impedance-control","joint-impedance-control","admittance-control","hybrid-force-position-control","null-space-control","operational-space-control"],"name":"Cartesian Impedance Control","alt":"笛卡尔阻抗控制","abbr":"","aliases":["Task-Space Impedance Control","End-Effector Impedance Control"],"one_liner":"Making the arm's end-effector behave like it's mounted on a spring and damper: it yields when pushed and springs back when released.","explanation":"Impedance control was proposed by Hogan in 1984–1985; Cartesian impedance control is its task-space form (defined on end-effector pose): instead of rigidly locking the end-effector's position, it specifies how much restoring force appears as the end-effector deviates from its target, F = K·e + D·ė, where e is the difference between target and actual pose, K the stiffness (newtons of force per meter of deviation), and D the damping (which suppresses oscillation); this is then converted to joint torque via the Jacobian transpose, τ = JᵀF, with Coriolis and gravity compensation added. K can be set per direction — for example, making the direction perpendicular to a hole's axis soft during peg insertion, to allow self-alignment. Compared to joint impedance control, its spring is defined at the end-effector rather than at each joint; and unlike admittance control, which is the reverse — ‘measure force, output position’ and suits stiff, position-controlled industrial arms — impedance control is ‘measure position, output force,’ suited to arms like Franka and KUKA iiwa that can command torque directly. Seven-DOF arms also add a null-space term to manage the extra degree of freedom.","example":"franka_ros's Cartesian impedance example controller computes joint torque as τ = Jᵀ(−K·e − D·J·q̇) + null-space term + Coriolis compensation, with a default translational stiffness of 200 N/m: pushing the end-effector 1 centimeter off target produces about 2 N of restoring force, and the stiffness can be adjusted at runtime.","related":["Impedance Control","Joint Impedance Control","Admittance Control","Hybrid Force/Position Control","Null-Space Control","Operational Space Control"]},{"id":"variable-impedance-control","category":"control","sec":2,"tier":3,"sources":[{"title":"Abu-Dakka, Saveriano: Variable Impedance Control and Learning—A Review (Frontiers in Robotics and AI, 2020)","url":"https://www.frontiersin.org/articles/10.3389/frobt.2020.590681/full"},{"title":"Martín-Martín et al.: Variable Impedance Control in End-Effector Space: An Action Space for RL in Contact-Rich Tasks (arXiv:1906.08880)","url":"https://arxiv.org/abs/1906.08880"}],"as_of":"","related_ids":["impedance-control","compliance-control","admittance-control","cartesian-impedance-control","force-aware-compliant-policy-learning","peg-in-hole-insertion"],"name":"Variable Impedance Control","alt":"变阻抗控制","abbr":"VIC","aliases":["VIC","Time-Varying Impedance Control"],"one_liner":"Impedance control whose stiffness and damping change in real time with the task phase or sensed force, stiff when needed and soft when needed.","explanation":"Impedance control makes a robot's end-effector behave like a spring and damper: F = K(x_d − x) + D(ẋ_d − ẋ), where x_d is the desired position, x is the actual position, K is stiffness (higher means ‘stiffer’), and D is damping. Ordinary impedance control keeps K and D fixed; variable impedance control instead lets them change over time, with the task phase, or with force feedback — an idea inspired by how humans actively adjust their arm stiffness. This balances precision and safety: stiff in free space for accurate tracking, soft on contact so the robot doesn't damage what it touches. The difficulty is that time-varying stiffness can inject energy into the system and break stability, which is usually kept in check with passivity theory and so-called ‘energy tanks.’ K and D can be hand-designed, learned from demonstrations, or output directly as a reinforcement-learning action: the 2019 VICES method has the policy output both the end-effector motion and the impedance gains, and is more sample-efficient and safer than fixed impedance on contact-rich tasks like wiping and door-opening. Federico Abu-Dakka and Matteo Saveriano wrote a systematic survey in 2020.","example":"Peg-in-hole assembly: the arm uses high stiffness while carrying the peg for fast, accurate positioning. Once the peg touches the rim of the hole, lateral stiffness is lowered so it can slide along the chamfer into the hole, while stiffness along the insertion direction stays higher to keep pushing forward.","related":["Impedance Control","Compliance Control","Admittance Control","Cartesian Impedance Control","Force-aware / Compliant Policy Learning","Peg-in-Hole Insertion"]},{"id":"passivity-based-control","category":"control","sec":2,"tier":3,"sources":[{"title":"Borja & Ortega, Introduction to Passivity-Based Control (arXiv 2608.15222)","url":"https://arxiv.org/abs/2608.15222"},{"title":"Passivity-based control for haptic teleoperation of a legged manipulator in presence of time-delays (arXiv 2108.07658)","url":"https://arxiv.org/abs/2108.07658"},{"title":"Wikipedia: Passivity (engineering)","url":"https://en.wikipedia.org/wiki/Passivity_(engineering)"}],"as_of":"","related_ids":["impedance-control","lyapunov-stability","bilateral-teleoperation","computed-torque-control","gravity-compensation","physical-human-robot-interaction"],"name":"Passivity-Based Control","alt":"无源性控制","abbr":"PBC","aliases":["PBC"],"one_liner":"Designing a controller from an energy standpoint, so the closed-loop system only ever dissipates energy, never creates it, and is thereby stable.","explanation":"Passivity-based control is a family of nonlinear controller-design methods built around energy, a name coined by Ortega and Spong in their 1989 survey on adaptive control of rigid manipulators. A system is ‘passive’ if the energy it stores never increases by more than the energy fed in from outside: there exists a storage function H ≥ 0 such that dH/dt ≤ uᵀy, where u is the input and y the corresponding output (such as force and velocity), so uᵀy is the input power. Two passive systems connected in a power-conserving way remain passive together, which is why it's easier to guarantee stability when a robot contacts a person or an unknown environment. Design typically has two steps: energy shaping, reworking the closed-loop storage function so it's minimized at the target state; and damping injection, adding a dissipative term so energy keeps decreasing and settles at the target. Rather than canceling out all the nonlinearity (as computed torque control does), it tends to be more robust to model error. Teleoperation, impedance control, and human-robot collaboration commonly use an ‘energy tank’ to keep the books balanced, preventing communication delay or changing stiffness from breaking passivity.","example":"ETH's Risiglione et al. (2021) built force-feedback teleoperation for a legged robot with an arm: both the leader and follower sides keep a virtual energy tank tracking the energy exchanged, with the follower's whole-body controller constraining the tank's energy to stay positive, so the whole loop stays passive — and stable — even with network delay.","related":["Impedance Control","Lyapunov Stability","Bilateral Teleoperation","Computed Torque Control","Gravity Compensation","Physical Human-Robot Interaction"]},{"id":"force-aware-compliant-policy-learning","category":"control","sec":2,"tier":3,"sources":[{"title":"Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Control (arXiv:2410.09309)","url":"https://arxiv.org/abs/2410.09309"},{"title":"FoAR: Force-Aware Reactive Policy for Contact-Rich Robotic Manipulation (arXiv:2411.15753)","url":"https://arxiv.org/abs/2411.15753"},{"title":"Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation (arXiv:2503.02881)","url":"https://arxiv.org/abs/2503.02881"}],"as_of":"2025-05","related_ids":["force-control","impedance-control","variable-impedance-control","contact-rich-manipulation","six-axis-force-torque-sensor","force-aware-vision-language-action-model"],"name":"Force-aware / Compliant Policy Learning","alt":"力感知策略 / 柔顺策略学习","abbr":"","aliases":["Force-aware Policy","Compliance Policy"],"one_liner":"Letting a learned manipulation policy sense contact force or output stiffness, so it maintains contact without pressing too hard.","explanation":"This is an umbrella term for methods that let learned manipulation policies handle contact force. Common visuomotor policies (such as diffusion policy) only output a target position, executed by a rigid position controller, which tends to either press too hard or fail to stay in contact once touching something. Improvements take roughly two routes. One is ‘force-aware’: feeding 6-axis force/torque or tactile signals in as input, fused with images before outputting an action — for example, Shanghai Jiao Tong University's Cewu Lu group's FoAR (2024) weights the force signal based on a predicted contact state, and the same group's RDP (2025) pairs a low-frequency diffusion policy with a high-frequency tactile feedback loop, while ForceVLA (2025) adds a force-aware mixture-of-experts module inside a VLA. The other is ‘compliant’: having the policy also output parameters like stiffness, executed by an impedance or admittance compliant controller — for example, Stanford and TRI's Adaptive Compliance Policy, ACP (2024), outputs a reference pose, a virtual target pose, and stiffness. The difficulty is that most teleoperation systems have no force feedback, making it hard to get good force and stiffness labels from demonstration data.","example":"ACP's vase-wiping task: the policy uses images and force signals to give, in real time, a reference pose, a virtual target pose, and stiffness along the compliant direction, executed by a compliant controller so the robot keeps its wiping tool pressed against a curved vase surface without pressing too hard; the paper reports over 50% improvement on contact-rich tasks compared to position-only visuomotor policies.","related":["Force Control","Impedance Control","Variable Impedance Control","Contact-rich Manipulation","Six-Axis Force/Torque Sensor","Force-aware Vision-Language-Action Model"]},{"id":"visual-servoing","category":"control","sec":2,"tier":2,"sources":[{"title":"Chaumette & Hutchinson, Visual Servo Control Part I: Basic Approaches (IEEE RAM, 2006)","url":"https://www.irisa.fr/lagadic/pdf/2006_ieee_ram_chaumette.pdf"},{"title":"Wikipedia: Visual servoing","url":"https://en.wikipedia.org/wiki/Visual_servoing"}],"as_of":"","related_ids":["image-based-visual-servoing","position-based-visual-servoing","eye-in-hand","eye-to-hand","visuomotor-policy","closed-loop-control"],"name":"Visual Servoing","alt":"视觉伺服","abbr":"","aliases":["Visual Servo Control","Vision-Based Robot Control"],"one_liner":"Feeding image error straight into the control loop, adjusting the robot's motion in real time as the camera watches.","explanation":"Visual servoing is a control method that uses visual information as feedback to drive robot motion in real time; related research dates back to around 1979, though the term ‘visual servo’ itself wasn't coined until 1987. The core idea is driving the error e = s − s* to zero, where s is the current feature extracted from the image (e.g., the pixel coordinates of a few corner points) and s* is the desired feature. The classic control law is v = −λL⁺e, where v is camera velocity, λ a gain, L the interaction matrix (describing how the features change as the camera moves), and L⁺ its pseudoinverse. It splits into two types by feature: image-based visual servoing (IBVS) works directly with pixel features and never needs to estimate a 3D pose; position-based visual servoing (PBVS) first estimates the target's 3D pose and then controls from that. The camera can be mounted on the robot's hand (eye-in-hand) or fixed externally (eye-to-hand).","example":"A typical example from Chaumette and Hutchinson's 2006 tutorial: a wrist-mounted camera looks at four points arranged in a square, uses the difference between their current and desired pixel coordinates as the error, and computes camera velocity via v = −λL⁺e until the four points land at their desired image positions — at which point the camera has also reached the target pose.","related":["Image-Based Visual Servoing","Position-Based Visual Servoing","Eye-in-Hand","Eye-to-Hand","Visuomotor Policy","Closed-loop Control"]},{"id":"image-based-visual-servoing","category":"control","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Visual servoing","url":"https://en.wikipedia.org/wiki/Visual_servoing"},{"title":"Chaumette & Hutchinson: Visual servo control, Part II: Advanced approaches (IEEE RAM 2007，含 Part I 基本公式回顾)","url":"https://inria.hal.science/inria-00350638v1/document"}],"as_of":"","related_ids":["visual-servoing","position-based-visual-servoing","jacobian-matrix","eye-in-hand","camera-intrinsics","closed-loop-control"],"name":"Image-Based Visual Servoing","alt":"基于图像的视觉伺服","abbr":"IBVS","aliases":["IBVS","2D Visual Servoing"],"one_liner":"A closed-loop vision-control method that computes camera or arm velocity directly from the pixel error of image feature points.","explanation":"Visual servoing uses camera feedback to close a control loop around robot motion, splitting into image-based (IBVS) and position-based (PBVS), with IBVS proposed by Weiss and Sanderson in the 1980s. IBVS doesn't reconstruct the target's 3D pose at all; it defines error directly on the image, e = s − s*, where s is the current feature (e.g., pixel coordinates of corner points) and s* the desired feature. Feature velocity and camera velocity v are related by ṡ = L·v, where L is the interaction matrix (or image Jacobian), depending on pixel coordinates and depth Z. The control law v = −λ·L⁺·e (L⁺ the pseudoinverse, λ a gain) drives the error down at roughly an exponential rate. It needs no object model and is fairly robust to calibration error; the drawback is that depth can only be estimated, the camera can retreat unexpectedly during a large rotation about the optical axis, and it can get stuck in a local minimum. PBVS instead estimates pose first and controls in 3D space, depending more heavily on calibration and a model.","example":"A wrist camera aimed at 4 marker corners on an assembly part: at the aligned position, a reference image records the 4 corners' pixel coordinates as s*; at runtime the controller continuously compares current against desired pixels and outputs the camera's 6D velocity until the four points coincide — at which point the gripper is aligned with the hole too.","related":["Visual Servoing","Position-Based Visual Servoing","Jacobian Matrix","Eye-in-Hand","Camera Intrinsics","Closed-loop Control"]},{"id":"position-based-visual-servoing","category":"control","sec":2,"tier":3,"sources":[{"title":"Chaumette & Hutchinson, Visual Servo Control Part I: Basic Approaches (IEEE RAM, 2006)","url":"https://www.irisa.fr/lagadic/pdf/2006_ieee_ram_chaumette.pdf"},{"title":"Wikipedia: Visual servoing","url":"https://en.wikipedia.org/wiki/Visual_servoing"}],"as_of":"","related_ids":["visual-servoing","image-based-visual-servoing","6d-object-pose-estimation","camera-calibration","hand-eye-calibration","eye-in-hand"],"name":"Position-Based Visual Servoing","alt":"基于位置的视觉伺服","abbr":"PBVS","aliases":["PBVS","Pose-Based Visual Servoing","3D Visual Servoing"],"one_liner":"A visual servo method that first estimates the target's 3D pose from the image, then drives the robot to reduce the pose error.","explanation":"Visual servoing means feeding camera images directly into the control loop. The classic 2006 tutorial by François Chaumette and Seth Hutchinson splits it into two families: PBVS and image-based visual servoing (IBVS). PBVS first uses the camera's calibrated intrinsics and a 3D model of the object to estimate the camera's pose relative to the target from the image, then defines an error e as the difference from the desired pose — usually written as a translation t plus an angle-axis rotation θu. The camera velocity command is v = −λL⁺e, where λ is a gain and L⁺ is the pseudoinverse of the interaction matrix (how the error changes with camera velocity). The advantage is that the error lives in 3D space, so the camera's path is intuitive and close to a straight line. The drawback is heavy reliance on calibration and pose estimation accuracy: if the estimate is off, the robot simply stops at the wrong pose, and the target can drift out of view during motion. IBVS instead compares image feature points directly, which tolerates calibration error better.","example":"A wrist camera (eye-in-hand) looks at a part on the table. Each frame, pose estimation computes the part's 6D pose relative to the camera, subtracts it from the desired pre-grasp camera pose to get the error, and keeps correcting the end-effector velocity until the error converges to zero before the gripper closes.","related":["Visual Servoing","Image-Based Visual Servoing","6D Object Pose Estimation","Camera Calibration","Hand-Eye Calibration","Eye-in-Hand"]},{"id":"trajectory-planning","category":"control","sec":3,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), Ch.9 Trajectory Generation","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MoveIt Docs: Time Parameterization","url":"https://moveit.picknik.ai/main/doc/examples/time_parameterization/time_parameterization_tutorial.html"}],"as_of":"","related_ids":["path-planning","motion-planning","time-parameterization","time-optimal-path-parameterization","trapezoidal-velocity-profile","trajectory-tracking"],"name":"Trajectory Planning","alt":"轨迹规划","abbr":"","aliases":["Trajectory Generation"],"one_liner":"Deciding where a robot should be and how fast at every moment in time — a path plus a schedule.","explanation":"Trajectory planning answers ‘where to be at what time, at what speed.’ Modern Robotics splits a trajectory into two parts: the path, a purely geometric sequence of configurations describing only which positions are visited; and the time scaling s(t), which specifies which point along the path to be at, at each moment. Together they form the trajectory θ(t), which must be smooth enough and respect the joint's velocity, acceleration, and torque limits. It is related to but distinct from path planning and motion planning: path planning focuses on finding a feasible route around obstacles without regard to time, while trajectory planning adds time and dynamics constraints on top of a route. Common techniques include polynomial time scaling, trapezoidal or S-curve velocity profiles, splines through waypoints, and time-optimal parameterization that accounts for actuator limits.","example":"MoveIt's typical pipeline: a planner first produces a collision-free sequence of joint waypoints with no timing information, then time parameterization (mainly the TOTG algorithm, sometimes followed by Ruckig to limit jerk) adds timing based on joint velocity and acceleration limits, producing an executable trajectory.","related":["Path Planning","Motion Planning","Time Parameterization","Time-Optimal Path Parameterization","Trapezoidal Velocity Profile","Trajectory Tracking"]},{"id":"waypoint","category":"control","sec":3,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), 9.3 Polynomial Via Point Trajectories","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"MoveIt Docs: Move Group C++ Interface（Cartesian Paths）","url":"https://moveit.picknik.ai/main/doc/examples/move_group_interface/move_group_interface_tutorial.html"},{"title":"Wikipedia: Waypoint","url":"https://en.wikipedia.org/wiki/Waypoint"}],"as_of":"","related_ids":["trajectory-interpolation","trajectory-planning","movej-movel","moveit-motion-planning-framework","cubic-spline-interpolation","navigation"],"name":"Waypoint","alt":"路点","abbr":"","aliases":["Via-point","Via Point"],"one_liner":"An intermediate point a robot is required to pass through during motion, used to constrain which route it takes.","explanation":"A waypoint is an intermediate position that a robot is meant to pass through or reach during motion — it can be a set of joint angles, or an end-effector or body pose, sometimes also carrying a desired arrival time and speed. The term comes from ‘waypoints’ in maritime, aviation, and GPS navigation. In robotics, a user or a planner first provides a sequence of waypoints, which trajectory interpolation and time parameterization then connect into a smooth, executable trajectory; the shape between waypoints depends on the interpolation method — a cubic polynomial, for instance, only guarantees continuous velocity, while a B-spline may not pass exactly through each point at all. Mobile-robot navigation and drone missions also commonly describe a route as a sequence of waypoints, and the individual points recorded when teaching an industrial arm are, in essence, waypoints too.","example":"When MoveIt plans a Cartesian path, several end-effector poses are placed in a waypoints list, and computeCartesianPath is called with eef_step = 0.01 — inserting a point every 1 cm; the returned fraction reports how much of the path was successfully planned, with velocity handled separately by time parameterization.","related":["Trajectory Interpolation","Trajectory Planning","MoveJ / MoveL","MoveIt Motion Planning Framework","Cubic Spline Interpolation","Navigation"]},{"id":"trajectory-interpolation","category":"control","sec":3,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), 9.3 Polynomial Via Point Trajectories","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion（附录 UR5 robot station）","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["waypoint","cubic-spline-interpolation","quintic-polynomial-interpolation","spherical-linear-interpolation","time-parameterization","action-chunking"],"name":"Trajectory Interpolation","alt":"轨迹插值","abbr":"","aliases":["Interpolation","Motion Interpolation"],"one_liner":"Filling in a continuous, timed trajectory between a few given key points, such as a start, an end, and waypoints.","explanation":"Trajectory interpolation computes, between a set of given points (start, end, waypoints), the position each control cycle should be at, and where needed, the velocity and acceleration too. The simplest is linear interpolation, which causes velocity to jump abruptly at waypoints. Cubic polynomials keep velocity continuous but can still leave acceleration discontinuous; quintic (fifth-order) polynomials additionally keep acceleration continuous, giving smoother motion. B-splines don't necessarily pass exactly through every point, but the resulting trajectory is guaranteed to stay within the convex hull of those points, which helps respect joint limits. Orientation can't be linearly interpolated directly on Euler angles, so spherical linear interpolation (slerp) of quaternions is used instead. Another common use is ‘upsampling’: when a model or a teleoperation device produces points slowly but the low-level controller expects fast commands, interpolation fills the gap.","example":"In Diffusion Policy's push-T task on a UR5, the policy issues 10 end-effector position commands per second, and the controller linearly interpolates them up to the robot's required 125 Hz, while capping end-effector speed under 0.43 m/s and restricting position to a region at least 1 cm above the tabletop.","related":["Waypoint","Cubic Spline Interpolation","Quintic Polynomial Interpolation","Spherical Linear Interpolation (SLERP)","Time Parameterization","Action Chunking"]},{"id":"movej-movel","category":"control","sec":3,"tier":2,"sources":[{"title":"The URScript Programming Language for e-Series, SW 5.11（movej / movel）","url":"https://s3-eu-west-1.amazonaws.com/ur-support-site/115824/scriptManual_SW5.11.pdf"},{"title":"ur_rtde: rtde_control_interface.h（moveJ linear in joint-space / moveL linear in tool-space）","url":"https://gitlab.com/sdurobotics/ur_rtde/-/raw/master/include/ur_rtde/rtde_control_interface.h"}],"as_of":"","related_ids":["joint-space","task-space","inverse-kinematics","tool-center-point","trapezoidal-velocity-profile","pre-grasp-pose"],"name":"MoveJ / MoveL","alt":"关节运动与直线运动（MoveJ / MoveL）","abbr":"","aliases":["Joint Move / Linear Move","movej / movel"],"one_liner":"The two most common arm-motion commands: move through joint space, or make the end-effector travel in a straight line.","explanation":"MoveJ and MoveL are the two basic motion commands in industrial and collaborative arm programming — written as movej and movel in Universal Robots' URScript, with equivalents from other vendors. MoveJ interpolates in joint space: every joint rotates in sync from its current angle to its target angle, so the end-effector usually traces a curve, but the controller doesn't need to solve inverse kinematics along the way, which makes it fast and immune to singularities (poses where some directions can't move) encountered mid-path. MoveL interpolates in Cartesian space: it drives the tool center point along a straight line to the target pose, requiring the controller to solve inverse kinematics continuously, giving a predictable path suited to approaching an object, insertion, or dispensing glue. Both follow a trapezoidal velocity profile. A common pattern is MoveJ to a pre-grasp pose, then MoveL to approach the object in a straight line.","example":"In URScript, movej([0,1.57,-1.57,3.14,-1.57,1.57], a=1.4, v=1.05) drives all six joints to that set of angles at a lead-joint speed of 1.05 rad/s; movel(pose, a=1.2, v=0.25) moves the end-effector to pose along a straight line at 0.25 m/s.","related":["Joint Space","Task Space","Inverse Kinematics (IK)","Tool Center Point","Trapezoidal Velocity Profile","Pre-grasp Pose"]},{"id":"offline-programming","category":"control","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Off-line programming (robotics)","url":"https://en.wikipedia.org/wiki/Off-line_programming_(robotics)"},{"title":"RoboDK: Offline Programming","url":"https://robodk.com/offline-programming"}],"as_of":"2026-09","related_ids":["teach-and-playback-programming","teach-pendant","industrial-robot","digital-twin","simulator","kinematic-calibration"],"name":"Offline Programming","alt":"离线编程","abbr":"OLP","aliases":["OLP","Off-line Programming"],"one_liner":"Writing and verifying a robot program in a 3D simulation on a computer first, then downloading it to the real robot to run.","explanation":"Offline programming is a way of programming industrial robots: a 3D model of the workstation (robot, tooling, workpiece CAD) is built on a computer, motion trajectories are generated and adjusted in simulation software, and reachability, collisions, and cycle time (how long it takes to finish one workpiece) are checked; once confirmed, a ‘post-processor’ translates the trajectory into the target controller brand's robot language (such as ABB's RAPID), which is then downloaded to the real robot. This contrasts with online programming, where an engineer teaches the robot point by point with a teach pendant on the shop floor, during which production has to stop. Offline programming doesn't consume production time and suits complex trajectories or frequently changing product lines; according to Wikipedia, it can shorten the time to bring a new program online from weeks to a single day. The difficulty is that the simulated model and the real setup differ somewhat, so the robot and workpiece usually need to be calibrated to correct for that. Common software includes RoboDK and ABB RobotStudio, with RoboDK claiming support for over 80 brands and 1,400-plus robot models.","example":"Before switching a production line to a new workpiece, an engineer imports the workpiece's CAD into RoboDK, generates a path for a welding gun or sanding head along the part's edge, simulation-checks for collisions and cycle time, post-processes it into an ABB RAPID program, and downloads it to the robot, needing only minor point corrections on-site.","related":["Teach-and-Playback Programming","Teach Pendant","Industrial Robot","Digital Twin","Simulator","Kinematic Calibration"]},{"id":"trapezoidal-velocity-profile","category":"control","sec":3,"tier":3,"sources":[{"title":"MATLAB Robotics System Toolbox: trapveltraj（Generate trajectories with trapezoidal velocity profiles）","url":"https://www.mathworks.com/help/robotics/ref/trapveltraj.html"}],"as_of":"","related_ids":["s-curve-velocity-profile","minimum-jerk-trajectory","jerk","movej-movel","time-parameterization","trajectory-interpolation"],"name":"Trapezoidal Velocity Profile","alt":"梯形速度曲线","abbr":"","aliases":["Trapezoidal Acceleration Profile","T-Curve Velocity Profile"],"one_liner":"A speed plan of constant acceleration, constant speed, then constant deceleration, tracing a trapezoid on a velocity-time graph.","explanation":"The trapezoidal velocity profile is the most basic speed plan for point-to-point motion in motors and robot arms: accelerate at a constant rate a from 0 up to a maximum speed v_max, cruise at that speed, then decelerate at the same rate back to 0 — hence the trapezoid shape on a velocity-versus-time graph. The acceleration phase lasts t_a = v_max / a and covers a distance of v_max²/(2a); if the total distance L is less than v_max²/a, there isn't room to reach maximum speed and the profile degenerates into a triangle, with a peak speed of √(aL). It's simple to compute, and industrial-arm commands like MoveJ and MoveL, along with many motor drivers, use it. Its drawback is that acceleration jumps abruptly at the segment boundaries, making jerk (the rate of change of acceleration) theoretically infinite and prone to causing vibration and shock; when smoother motion is needed, an S-curve profile or a minimum-jerk trajectory is used instead. For synchronized multi-joint motion, the time needed by the slowest joint is usually computed first, and the other joints are then slowed proportionally so all arrive together.","example":"Moving a linear axis L = 1 m with a speed limit of v_max = 0.5 m/s and acceleration a = 1 m/s²: it accelerates for 0.5 s (covering 0.125 m), cruises for 1.5 s (covering 0.75 m), and decelerates for 0.5 s (covering 0.125 m), for a total of 2.5 s.","related":["S-Curve Velocity Profile","Minimum-Jerk Trajectory","Jerk","MoveJ / MoveL","Time Parameterization","Trajectory Interpolation"]},{"id":"s-curve-velocity-profile","category":"control","sec":3,"tier":3,"sources":[{"title":"Jerk-limited Real-time Trajectory Generation with Arbitrary Target States (Berscheid & Kröger, RSS 2021, arXiv 2105.04830)","url":"https://arxiv.org/abs/2105.04830"},{"title":"Ruckig GitHub（Motion Generation for Robots and Machines）","url":"https://github.com/pantor/ruckig"}],"as_of":"","related_ids":["trapezoidal-velocity-profile","jerk","time-parameterization","trajectory-planning","ruckig","minimum-jerk-trajectory"],"name":"S-Curve Velocity Profile","alt":"S 型速度曲线","abbr":"","aliases":["S-Curve Acceleration","Seven-Segment Velocity Profile","Jerk-Limited Velocity Profile"],"one_liner":"A velocity plan that limits jerk so acceleration changes smoothly, making the speed-versus-time curve look like an S.","explanation":"The simplest point-to-point motion is a trapezoidal velocity profile: constant acceleration, constant velocity, constant deceleration. But acceleration jumps abruptly at the boundaries between segments, making jerk (the rate of change of acceleration) theoretically infinite and prone to shaking the mechanism. An S-curve profile caps jerk instead: the acceleration phase is split into three sub-phases — increasing acceleration, constant acceleration, decreasing acceleration — and deceleration is split the same way, giving seven segments in total together with the middle constant-velocity phase (hence the alternative name ‘seven-segment profile’). Each segment's jerk is set to +J, 0, or −J, so the velocity-versus-time curve becomes a smooth S shape at the start and stop. The cost is a slightly longer travel time than a trapezoidal profile and more complex computation, and short moves can end up with zero-length segments. Industrial robot controllers and CNC machines widely use S-curves; Lars Berscheid and Torsten Kröger at the Karlsruhe Institute of Technology (KIT) published the open-source library Ruckig at RSS 2021, which computes a time-optimal, jerk-limited trajectory in real time on every control cycle.","example":"Ruckig writes the fastest single-joint trajectory as up to seven constant-jerk segments and synchronizes multiple joints to arrive together. MoveIt 2, CoppeliaSim, LinuxCNC, and Franka's Frankx library all use it to generate jerk-limited trajectories.","related":["Trapezoidal Velocity Profile","Jerk","Time Parameterization","Trajectory Planning","Ruckig","Minimum-Jerk Trajectory"]},{"id":"bezier-curve-trajectory","category":"control","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Bézier curve","url":"https://en.wikipedia.org/wiki/B%C3%A9zier_curve"},{"title":"MIT Cheetah-Software: FootSwingTrajectory.cpp（Currently uses Bezier curves like Cheetah 3 does）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/common/src/Controllers/FootSwingTrajectory.cpp"}],"as_of":"","related_ids":["swing-foot-trajectory-planning","cubic-spline-interpolation","quintic-polynomial-interpolation","minimum-jerk-trajectory","waypoint","hybrid-zero-dynamics"],"name":"Bézier Curve Trajectory","alt":"贝塞尔曲线轨迹","abbr":"","aliases":["Bezier Curve","Bézier Polynomial"],"one_liner":"A smooth polynomial curve shaped by a handful of control points, often used to describe robot trajectories.","explanation":"The Bézier curve was popularized for car-body design in the 1960s by Renault engineer Pierre Bézier, while Citroën's Paul de Casteljau had independently found the underlying computation even earlier. An n-th order curve is defined by n+1 control points P₀…Pₙ: B(s) = Σᵢ C(n,i)(1−s)ⁿ⁻ⁱ sⁱ Pᵢ, with the parameter s running from 0 to 1 and C(n,i) the binomial coefficient. Three properties make it well suited to trajectories: the curve passes exactly through the first and last control points; the tangent direction at the start and end is set by the adjacent control points, which makes it easy to join with the segments before and after; and the whole curve stays within the convex hull formed by the control points, so keeping the control points clear of obstacles keeps the curve clear too. The intermediate control points don't lie on the curve — they only pull its shape. In robotics it's commonly used for swing-foot trajectories and end-effector path smoothing; hybrid zero dynamics methods also use Bézier polynomials to describe reference gait trajectories. Long paths are usually built from several low-order segments rather than one high-order curve.","example":"MIT's open-source Cheetah-Software uses a cubic Bézier curve to generate the swing-leg trajectory: the horizontal direction transitions smoothly from the start point to the landing point, while the vertical direction rises to the step height in the first half and descends back to the ground in the second, with the curve's derivative giving foot velocity and acceleration directly.","related":["Swing Foot Trajectory Planning","Cubic Spline Interpolation","Quintic Polynomial Interpolation","Minimum-Jerk Trajectory","Waypoint","Hybrid Zero Dynamics"]},{"id":"cubic-spline-interpolation","category":"control","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Spline interpolation","url":"https://en.wikipedia.org/wiki/Spline_interpolation"},{"title":"ros2_control: joint_trajectory_controller 轨迹表示与插值","url":"https://control.ros.org/master/doc/ros2_controllers/joint_trajectory_controller/doc/trajectory.html"}],"as_of":"","related_ids":["trajectory-interpolation","quintic-polynomial-interpolation","waypoint","time-parameterization","minimum-jerk-trajectory","jerk"],"name":"Cubic Spline Interpolation","alt":"三次样条插值","abbr":"","aliases":["Cubic Spline"],"one_liner":"Connecting a sequence of waypoints into a smooth curve, continuous in both position and velocity, using piecewise cubic polynomials.","explanation":"Cubic spline interpolation is a classic numerical-analysis method: given a sequence of waypoints, each pair of adjacent points is connected by a cubic polynomial segment q(t) = a₀ + a₁t + a₂t² + a₃t³, requiring position and velocity (first derivative) to be continuous where segments meet; a standard cubic spline also requires acceleration (second derivative) to be continuous, plus boundary conditions — a natural spline sets the second derivative to zero at both ends, a clamped spline instead specifies the velocity at both ends. Compared to fitting all points with a single high-degree polynomial, it avoids the Runge phenomenon's wild oscillation. In robotics it's commonly used to interpolate the sparse waypoints from a planner or a policy into the dense commands a controller needs each cycle. If only the position and velocity at the start and end of each segment are given, acceleration may still jump between segments, corresponding to a large jerk that causes an impact — which is when quintic polynomials are used instead.","example":"A joint moving from 0 to 1 radian in 2 seconds, starting and ending at zero velocity, gives a single-segment cubic polynomial q(t) = 3(t/2)² − 2(t/2)³, with peak velocity of 0.75 rad/s at the midpoint, t = 1 second. ROS 2's joint_trajectory_controller picks its interpolation method based on what's given in the waypoints: linear interpolation for position alone, cubic spline once velocity is also given, and quintic spline once acceleration is given too.","related":["Trajectory Interpolation","Quintic Polynomial Interpolation","Waypoint","Time Parameterization","Minimum-Jerk Trajectory","Jerk"]},{"id":"quintic-polynomial-interpolation","category":"control","sec":3,"tier":3,"sources":[{"title":"Lynch & Park, Modern Robotics, Section 9.2 Polynomial Time Scaling (preprint PDF)","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Robotics Toolbox for Python: Trajectories (quintic / jtraj / mtraj)","url":"https://petercorke.github.io/robotics-toolbox-python/arm_trajectory.html"}],"as_of":"","related_ids":["trajectory-interpolation","cubic-spline-interpolation","minimum-jerk-trajectory","jerk","trapezoidal-velocity-profile","time-parameterization"],"name":"Quintic Polynomial Interpolation","alt":"五次多项式插值","abbr":"","aliases":["Quintic Polynomial Trajectory","Quintic Time Scaling"],"one_liner":"A point-to-point trajectory method that fits a fifth-degree polynomial matching position, velocity, and acceleration at both ends.","explanation":"Quintic polynomial interpolation is one of the most basic methods for point-to-point robot trajectories. The joint angle or position is written as q(t) = a₀ + a₁t + ... + a₅t⁵, and its six coefficients are uniquely fixed by six boundary conditions: position, velocity, and acceleration at both the start and the end. A cubic polynomial can only constrain position and velocity at the endpoints, so acceleration jumps abruptly at the start and stop, making jerk (the rate of change of acceleration) theoretically infinite and prone to causing vibration. A quintic polynomial can also hold acceleration at zero at both ends, keeping acceleration continuous throughout and jerk finite, for smoother starts and stops — at the cost of a higher peak velocity for the same duration, so joint speed limits need checking. With zero velocity and acceleration at both ends, the normalized form is s(τ) = 10τ³ − 15τ⁴ + 6τ⁵ (where τ = t/T and T is the total duration), which matches the minimum-jerk trajectory. For multiple waypoints, segments can be pieced together while keeping acceleration continuous at each one.","example":"In Peter Corke's Robotics Toolbox for Python, jtraj() generates joint-space trajectories using a quintic polynomial with zero start/end velocity and acceleration by default, while quintic() generates the same kind of trajectory for a single variable.","related":["Trajectory Interpolation","Cubic Spline Interpolation","Minimum-Jerk Trajectory","Jerk","Trapezoidal Velocity Profile","Time Parameterization"]},{"id":"minimum-jerk-trajectory","category":"control","sec":3,"tier":3,"sources":[{"title":"Flash & Hogan (1985), The coordination of arm movements: an experimentally confirmed mathematical model, J. Neurosci. 5(7):1688–1703","url":"https://doi.org/10.1523/JNEUROSCI.05-07-01688.1985"}],"as_of":"","related_ids":["jerk","quintic-polynomial-interpolation","s-curve-velocity-profile","trajectory-planning","minimum-snap-trajectory-differential-flatness","time-parameterization"],"name":"Minimum-Jerk Trajectory","alt":"最小加加速度轨迹","abbr":"","aliases":["Minimum Jerk"],"one_liner":"A trajectory that minimizes the integral of squared jerk (the rate of change of acceleration), giving smooth motion close to how a human hand moves.","explanation":"Jerk is the derivative of acceleration with respect to time; large jerk means abrupt, jarring motion. A minimum-jerk trajectory is the one that, given a fixed start, end, and duration, minimizes the integral of squared jerk. The best-known source is Flash and Hogan's 1985 paper in the Journal of Neuroscience: human point-to-point reaching movements in a plane are approximately straight with a bell-shaped speed profile, matching this model's predictions. In one dimension, with zero velocity and acceleration at both ends, the solution is a quintic polynomial: x(t) = x₀ + (x_f − x₀)(10τ³ − 15τ⁴ + 6τ⁵), where x₀ and x_f are the start and end points and τ = t/T is normalized time over total duration T. So it's exactly the special case of quintic polynomial interpolation where boundary velocity and acceleration are both zero. In robotics it's commonly used to generate point-to-point motion or smooth interpolation, reducing shock to motors and gearboxes. Minimizing the peak jerk instead of its integral leads to the S-curve velocity profile.","example":"Moving an arm's cup from one spot on a table to another over 2 seconds: interpolating with the formula above, end-effector velocity rises smoothly from 0, peaks at the midpoint, and smoothly falls back to 0, with no jolt at either the start or the stop.","related":["Jerk","Quintic Polynomial Interpolation","S-Curve Velocity Profile","Trajectory Planning","Minimum-Snap Trajectory / Differential Flatness","Time Parameterization"]},{"id":"minimum-snap-trajectory-differential-flatness","category":"control","sec":3,"tier":3,"sources":[{"title":"Richter, Bry, Roy: Polynomial Trajectory Planning for Aggressive Quadrotor Flight in Dense Indoor Environments (ISRR 2013)","url":"https://groups.csail.mit.edu/rrg/papers/Richter_ISRR13.pdf"},{"title":"Mellinger & Kumar: Minimum snap trajectory generation and control for quadrotors (ICRA 2011)","url":"https://doi.org/10.1109/ICRA.2011.5980409"},{"title":"Flatness (systems theory) - Wikipedia","url":"https://en.wikipedia.org/wiki/Flatness_(systems_theory)"}],"as_of":"","related_ids":["minimum-jerk-trajectory","trajectory-optimization","quadratic-programming","kinodynamic-planning","unmanned-aerial-vehicle","jerk"],"name":"Minimum-Snap Trajectory / Differential Flatness","alt":"最小 Snap 轨迹（微分平坦）","abbr":"","aliases":["Minimum Snap","Differential Flatness"],"one_liner":"A piecewise-polynomial trajectory minimizing the integral of squared snap (the 4th derivative of position), commonly used for quadrotor drones.","explanation":"Snap is the fourth time derivative of position — jerk differentiated once more. The method is bound up with differential flatness, introduced by Fliess and colleagues in 1995: a system is differentially flat if there's a set of ‘flat outputs’ from which every state and control input can be written as a function of those outputs and their derivatives. Mellinger and Kumar, at ICRA 2011, exploited a quadrotor's differential flatness, taking position x, y, z and yaw angle ψ as the flat outputs. This means planning a sufficiently smooth curve in just these four quantities is enough to directly compute attitude and motor commands, with no need to sample-search or repeatedly simulate over a high-dimensional state space. Motor commands and attitude angular acceleration are proportional to snap, so minimizing snap keeps commands changing gently and executable. In practice, a piecewise polynomial is fit through a sequence of waypoints, with the snap-squared integral written as a quadratic program over the polynomial coefficients; Richter et al. (2013) recast this as a numerically stable, unconstrained quadratic program and combined it with geometric path planning.","example":"Richter, Bry, and Roy (2013) used this method: first find obstacle-avoiding waypoints, then generate a minimum-snap polynomial trajectory through them, letting a quadrotor fly autonomously through dense indoor environments at speeds up to 8 m/s.","related":["Minimum-Jerk Trajectory","Trajectory Optimization","Quadratic Programming","Kinodynamic Planning","Unmanned Aerial Vehicle (UAV)","Jerk"]},{"id":"time-parameterization","category":"control","sec":3,"tier":3,"sources":[{"title":"MoveIt Documentation: Time Parameterization","url":"https://moveit.picknik.ai/main/doc/examples/time_parameterization/time_parameterization_tutorial.html"},{"title":"Kunz, Stilman: Time-Optimal Trajectory Generation for Path Following with Bounded Acceleration and Velocity (RSS 2012)","url":"https://www.roboticsproceedings.org/rss08/p27.html"}],"as_of":"","related_ids":["time-optimal-path-parameterization","topp-ra","ruckig","trajectory-planning","trapezoidal-velocity-profile","moveit-motion-planning-framework"],"name":"Time Parameterization","alt":"时间参数化","abbr":"","aliases":["Trajectory Time Parameterization","Retiming"],"one_liner":"Attaching a timing schedule to a purely geometric path, deciding when each point is reached and how fast.","explanation":"Many motion planners — including the OMPL sampling-based planners that MoveIt calls — output only a sequence of joint waypoints with no timing information: a purely geometric path q(s), where s runs from 0 to 1 as progress along the path. Time parameterization finds a function s(t) that assigns a timestamp, velocity, and acceleration to each waypoint so the robot can actually execute it. By the chain rule, q̇ = q′(s)·ṡ and q̈ = q′(s)·s̈ + q″(s)·ṡ², so joint velocity and acceleration limits translate directly into constraints on ṡ and s̈. It only decides how fast to move, not which path to take — that is what distinguishes it from trajectory optimization. Common algorithms include the iterative parabolic method, time-optimal trajectory generation (TOTG, by Tobias Kunz and Mike Stilman, RSS 2012), TOPP-RA, and the jerk-limited Ruckig. MoveIt currently defaults to TOTG and provides velocity- and acceleration-scaling factors from 0 to 1 to slow the whole motion down.","example":"Planning an arm's motion from A to B in MoveIt, OMPL returns a sequence of joint waypoints with no timing; TOTG then assigns timestamps to them based on the velocity and acceleration limits in joint_limits.yaml. Setting the velocity scaling factor to 0.1 computes the joint speed limits at 10% of their original values.","related":["Time-Optimal Path Parameterization","TOPP-RA","Ruckig","Trajectory Planning","Trapezoidal Velocity Profile","MoveIt Motion Planning Framework"]},{"id":"time-optimal-path-parameterization","category":"control","sec":3,"tier":3,"sources":[{"title":"Pham, Pham: A New Approach to Time-Optimal Path Parameterization based on Reachability Analysis (arXiv:1707.07239, IEEE T-RO 2018)","url":"https://arxiv.org/abs/1707.07239"},{"title":"Bobrow, Dubowsky, Gibson: Time-Optimal Control of Robotic Manipulators Along Specified Paths (IJRR 1985)","url":"https://doi.org/10.1177/027836498500400301"},{"title":"toppra: Time-Optimal Path Parameterization library (GitHub)","url":"https://github.com/hungpham2511/toppra"}],"as_of":"","related_ids":["time-parameterization","topp-ra","trajectory-planning","trapezoidal-velocity-profile","convex-optimization","joint-limits"],"name":"Time-Optimal Path Parameterization","alt":"时间最优路径参数化","abbr":"TOPP","aliases":["TOPP","Time-Optimal Path Tracking"],"one_liner":"Given a fixed path, find the fastest possible timing schedule that still respects velocity, acceleration, and torque limits.","explanation":"TOPP is the ‘fastest possible’ version of time parameterization: given a fixed geometric path q(s), it finds the s(t) that minimizes total time T = ∫ds/ṡ (where ṡ is the speed of progress along the path) subject to joint velocity, acceleration, and torque limits. The classic approach, from a 1985 paper by James Bobrow, Steven Dubowsky, and John Gibson, is numerical integration: on the (s, ṡ) phase plane, it alternately integrates at maximum acceleration and maximum deceleration to find the switching points, producing a ‘bang-bang’ profile — full acceleration, then full deceleration. This is fast to compute but the switching points are hard to locate and prone to numerical trouble; a second family of convex-optimization methods is more stable but slower. In 2018, Hung Pham and Quang-Cuong Pham published TOPP-RA in IEEE Transactions on Robotics, which solves a small linear program at each discretized point along the path by computing reachable and controllable velocity sets, balancing speed with robustness; the open-source toppra library implementing it is widely used. Industrial arms that need to shrink cycle time or move material quickly commonly rely on it.","example":"An arm moving along a fixed weld-seam path has that path discretized into points, and TOPP-RA computes the maximum achievable speed at each one: straight segments hug the speed limit, while corners automatically slow down because of the acceleration constraints.","related":["Time Parameterization","TOPP-RA","Trajectory Planning","Trapezoidal Velocity Profile","Convex Optimization","Joint Limits"]},{"id":"trajectory-tracking","category":"control","sec":3,"tier":2,"sources":[{"title":"Lynch & Park, Modern Robotics (2017 preprint), Ch.11 Robot Control 与 13.3 Trajectory tracking","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"},{"title":"Russ Tedrake, Underactuated Robotics: Trajectory Optimization（沿轨迹的 LQR 稳定）","url":"https://underactuated.mit.edu/trajopt.html"}],"as_of":"","related_ids":["trajectory-planning","proportional-derivative-control","feedforward-control","computed-torque-control","motion-tracking","linear-quadratic-regulator"],"name":"Trajectory Tracking","alt":"轨迹跟踪","abbr":"","aliases":["Trajectory Following"],"one_liner":"Keeping a robot's actual motion closely following a time-varying reference trajectory, driving the error toward zero.","explanation":"Trajectory tracking is a control problem: given a reference trajectory qd(t), drive the actual state q(t) so that the error qd(t) − q(t) goes to zero over time. A typical approach is ‘feedforward plus feedback’: feedforward uses the model to precompute the velocity or torque needed along the trajectory, and feedback (e.g., PD) corrects based on the current error; computed torque control combines both. Performance is judged by steady-state error, overshoot, and settling time. It differs from path tracking, which only requires following a geometric curve and can adjust its own speed; trajectory tracking has a requirement at every point in time. A reference trajectory is best kept away from the actuators' hard limits, to leave margin for correction. A humanoid robot imitating human motion — ‘motion tracking’ — follows the same idea.","example":"Chapter 11 of Modern Robotics simulates the same joint trajectory three ways: feedforward alone or PID feedback alone both show noticeable error, while computed torque control (feedforward plus feedback) tracks best.","related":["Trajectory Planning","Proportional-Derivative Control","Feedforward Control","Computed Torque Control","Motion Tracking","Linear Quadratic Regulator"]},{"id":"iterative-learning-control","category":"control","sec":3,"tier":3,"sources":[{"title":"Iterative learning control - Wikipedia","url":"https://en.wikipedia.org/wiki/Iterative_learning_control"}],"as_of":"","related_ids":["feedforward-control","proportional-integral-derivative-control","trajectory-tracking","adaptive-control","teach-and-playback-programming"],"name":"Iterative Learning Control","alt":"迭代学习控制","abbr":"ILC","aliases":["ILC"],"one_liner":"For a system that repeats the same motion over and over, using last time's tracking error to correct next time's control input.","explanation":"Iterative learning control targets systems that execute the same task repeatedly — an arm running the same trajectory on every cycle of a production line, for instance. It's generally credited to Arimoto, Kawamura, and Miyazaki's 1984 paper ‘Bettering operation of robots by learning.’ The idea is straightforward: after the k-th run, record the tracking error e_k(t) at every moment, then update the next run's input to u_{k+1}(t) = u_k(t) + L·e_k(t), where u is the control input and L a hand-designed learning gain; in practice a low-pass filter is usually added too, to keep noise from being amplified run after run. Ordinary feedback control can only correct an error after it appears; ILC instead feeds the previous run's error forward into the next run, so repeatable errors from an inaccurate model or friction get squeezed down more and more over successive runs. It requires the initial state and reference trajectory to stay much the same each time, and it offers no help against random disturbances or a task never attempted before — so it's usually paired with PID feedback.","example":"An industrial arm repeats the same welding trajectory on every production cycle: run 1 records lag at the corners, run 2 supplies extra torque slightly ahead of time at those same moments, and the tracking error at the corners keeps shrinking run after run.","related":["Feedforward Control","Proportional-Integral-Derivative Control","Trajectory Tracking","Adaptive Control","Teach-and-Playback Programming"]},{"id":"input-shaping","category":"control","sec":3,"tier":3,"sources":[{"title":"Zaber: Input Shaping for Vibration Reduction（引 Singer & Seering 1990）","url":"https://www.zaber.com/articles/input-shaping-for-vibration-reduction"},{"title":"Klipper documentation: Resonance Compensation","url":"https://www.klipper3d.org/Resonance_Compensation.html"}],"as_of":"","related_ids":["natural-frequency","damping-ratio","flexible-joint","feedforward-control","s-curve-velocity-profile","trajectory-planning"],"name":"Input Shaping (Residual Vibration Suppression)","alt":"输入整形（振动抑制）","abbr":"","aliases":["Command Shaping","ZV Shaper","Residual Vibration Suppression"],"one_liner":"Splitting a motion command into several time-offset sub-commands so the vibrations they excite cancel each other out.","explanation":"Input shaping is a feedforward vibration-suppression method, given its modern form by MIT's Singer and Seering in 1990. Flexible systems such as slender arms, crane cables, and belt-driven shafts keep vibrating at their natural frequency after a motion finishes. The technique convolves the original command with a short train of impulses: the simplest, ZV (zero vibration) shaper, uses two impulses spaced half a vibration period apart, with amplitudes split according to the damping ratio (equal halves when undamped) — the vibration excited by the first impulse is exactly canceled by the second. The cost is that the motion takes about half a period longer, and it's sensitive to frequency-estimation error, which led to more robust but slower variants such as ZVD and EI. It needs only the natural frequency and damping ratio — no extra sensors — and can be layered on top of smooth trajectories such as S-curve velocity profiles.","example":"The open-source 3D-printer firmware Klipper uses input shaping to remove ‘ringing’ artifacts on a print's surface: it first measures the frame's resonant frequency, then picks a shaper such as ZV, MZV, or EI; ZV delays the motion by only half a period but is the most sensitive to frequency error.","related":["Natural Frequency","Damping Ratio","Flexible Joint","Feedforward Control","S-Curve Velocity Profile","Trajectory Planning"]},{"id":"action-smoothing","category":"control","sec":3,"tier":2,"sources":[{"title":"Regularizing Action Policies for Smooth Control with Reinforcement Learning (CAPS, arXiv 2012.06644)","url":"https://arxiv.org/abs/2012.06644"},{"title":"Wikipedia: Low-pass filter","url":"https://en.wikipedia.org/wiki/Low-pass_filter"},{"title":"legged_gym legged_robot_config.py（action_rate 奖励项）","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["action-chunking","temporal-ensembling","control-frequency","jerk","real-time-chunking","proportional-derivative-control"],"name":"Action Smoothing","alt":"动作平滑","abbr":"","aliases":["Low-Pass Filtering","Action Filtering","Action Smoothness Regularization"],"one_liner":"Filtering or penalizing a policy's output actions to remove high-frequency jitter, so joint motion stays smooth.","explanation":"Action smoothing is a common step when deploying a learned policy. A neural network outputs an action independently at every step, so adjacent steps can jump around, making motors jitter back and forth, waste power, and heat up. There are three common fixes. One is filtering the action before sending it out, letting only low-frequency changes through — the simplest is an exponential moving average, y_t = α·x_t + (1−α)·y_{t−1}, where x is the policy's newly output action and y is the action actually sent, with smaller α giving more smoothing. Another is adding a penalty during training, such as the action_rate term in legged_gym's default reward, which penalizes the difference between consecutive actions; the CAPS method (ICRA 2021) adds both temporal and spatial smoothness regularization, and its authors report nearly an 80% reduction in power use on a quadrotor. A third is predicting a whole chunk of actions at once (action chunking), then temporally ensembling or interpolating them. The cost is that filtering introduces delay (phase lag), making the robot slower to react to disturbances, so α has to be tuned as a tradeoff between smoothness and responsiveness.","example":"A quadruped's reinforcement-learning policy outputs target joint angles at 50 Hz; deployed on the real robot, its legs show noticeable high-frequency jitter. Adding a first-order low-pass filter noticeably reduces the jitter, but the robot also reacts a bit more slowly when pushed, requiring α to be re-tuned.","related":["Action Chunking","Temporal Ensembling","Control Frequency","Jerk","Real-Time Chunking","Proportional-Derivative Control"]},{"id":"motion-primitives","category":"control","sec":3,"tier":3,"sources":[{"title":"Movement Primitives in Robotics: A Comprehensive Survey (arXiv 2601.02379)","url":"https://arxiv.org/abs/2601.02379"},{"title":"Spatio-Temporal Lattice Planning Using Optimal Motion Primitives (arXiv 2107.11467)","url":"https://arxiv.org/abs/2107.11467"}],"as_of":"","related_ids":["dynamic-movement-primitives","probabilistic-movement-primitives","skill-primitive","imitation-learning","kinodynamic-planning","a-star-search"],"name":"Motion Primitives","alt":"运动基元","abbr":"","aliases":["Movement Primitives"],"one_liner":"Breaking a complex motion into small, reusable, parameterizable building-block motions that get composed together when needed.","explanation":"Motion primitives are small, reusable basic motions — ‘reach toward a spot,’ ‘pick up,’ ‘go forward then turn left.’ The idea is biologically inspired: continuous motion in humans and animals can be seen as a concatenation of a handful of basic segments. There are two main uses in robotics. One is in manipulation and imitation learning, representing a trajectory with a parameterized model learned from demonstrations, where execution just changes parameters like the target point or duration to transfer to a new situation — representative examples are dynamic movement primitives (DMP, generating a trajectory from a spring-damper system plus a learnable driving term) and probabilistic movement primitives (ProMP, learning a distribution over trajectories from multiple demonstrations). The other is in mobile-robot and self-driving-car planning: a set of short trajectories satisfying the vehicle's dynamics is precomputed, and a graph search such as A* strings them together online — this is called lattice planning. Compared to atomic skills, motion primitives operate at the trajectory level, while atomic skills more often refer to complete, semantically defined sub-tasks.","example":"Automated parking: a planner keeps a library of short trajectories a car can actually drive, such as ‘go straight a bit’ or ‘full steering lock while reversing 90°,’ and searches online with A* to string them into a route from the current position into the parking spot.","related":["Dynamic Movement Primitives","Probabilistic Movement Primitives","Skill Primitive","Imitation Learning","Kinodynamic Planning","A* Search"]},{"id":"dynamic-movement-primitives","category":"control","sec":3,"tier":3,"sources":[{"title":"Saveriano et al.: Dynamic Movement Primitives in Robotics: A Tutorial Survey (arXiv:2102.03861)","url":"https://arxiv.org/abs/2102.03861"}],"as_of":"","related_ids":["motion-primitives","probabilistic-movement-primitives","imitation-learning","kinesthetic-teaching","gaussian-mixture-regression-task-parameterized-gmm"],"name":"Dynamic Movement Primitives","alt":"动态运动基元","abbr":"DMP","aliases":["DMP","Dynamical Movement Primitives","DMPs"],"one_liner":"Encoding a demonstrated trajectory as a ‘spring-damper plus a learnable force term’ dynamical system, reproducible with a different endpoint or speed.","explanation":"DMPs are a trajectory-representation method proposed by Ijspeert, Nakanishi, and Schaal in 2001–2002, with a systematic summary published in 2013. The core is a second-order ‘spring-damper’ system plus a learnable force term: τ·ż = α(β(g − y) − z) + f(x), τ·ẏ = z. Here y is position, z a scaled velocity, g the goal point, τ controls speed, α and β are fixed gains, and f(x) is a weighted sum of Gaussian basis functions whose weights can be fit directly from a single demonstration by linear regression; x comes from a ‘phase’ system τ·ẋ = −αₓx that decays to 0 over time, guaranteeing the force term eventually vanishes and the system converges stably to g. So changing g gives a new endpoint, and changing τ changes speed, while a mid-motion push causes the system to return to the trajectory on its own. DMPs have long been a standard entry point for beginners in imitation learning. Their limitation is representing only one fixed trajectory, unable to describe multiple ways of doing the same task — a gap that probabilistic movement primitives (ProMP) and similar methods fill by modeling a distribution instead.","example":"Record one kinesthetic-teaching demonstration of an arm carrying a cup from point A to point B, and fit DMP weights from it; to move to point C instead, just change g to C's coordinates, and to slow it down, increase τ — the arm generates a new, similarly shaped trajectory without being re-taught.","related":["Motion Primitives","Probabilistic Movement Primitives","Imitation Learning","Kinesthetic Teaching","Gaussian Mixture Regression / Task-Parameterized GMM"]},{"id":"probabilistic-movement-primitives","category":"control","sec":3,"tier":3,"sources":[{"title":"Paraschos et al., Probabilistic Movement Primitives (NIPS 2013) — abstract","url":"https://proceedings.neurips.cc/paper/2013/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html"},{"title":"Paraschos et al., Probabilistic Movement Primitives (NIPS 2013) — PDF","url":"https://proceedings.neurips.cc/paper_files/paper/2013/file/e53a0a2978c28872a4505bdb51db06dc-Paper.pdf"}],"as_of":"","related_ids":["dynamic-movement-primitives","motion-primitives","imitation-learning","gaussian-mixture-regression-task-parameterized-gmm","demonstration-data","kinesthetic-teaching"],"name":"Probabilistic Movement Primitives","alt":"概率运动基元","abbr":"ProMP","aliases":["ProMP","ProMPs","Probabilistic Movement Primitive"],"one_liner":"A movement primitive that represents a whole family of demonstrated trajectories as a Gaussian distribution, conditionable on via-points or goals.","explanation":"ProMPs were introduced by Alexandros Paraschos, Christian Daniel, Jan Peters, and Gerhard Neumann at NIPS 2013. Each trajectory is written as y_t = Φ_tᵀw + ε, where Φ_t is a set of time-indexed basis functions (commonly radial basis functions), w is a weight vector, and ε is noise. Each demonstration is fit to its own w, and a Gaussian distribution is then fit over these w vectors: the mean captures the typical motion, and the covariance captures how much the demonstrations vary and how the joints are coupled — so a ProMP represents a whole family of trajectories rather than a single one. Requiring the trajectory to pass through a given point at a given time is handled analytically by Gaussian conditioning, and several ProMPs can be co-activated and blended or switched between. The original paper also derives a stochastic feedback controller that reproduces the resulting trajectory distribution. Compared with dynamic movement primitives (DMPs), which use a differential equation with an attractor to represent a single trajectory, ProMPs model the distribution over trajectories directly, making them better at capturing variation across demonstrations.","example":"In the original paper, a 7-DOF KUKA lightweight arm played table hockey using two sets of 10 demonstrations each — one varying only the shot distance, the other only the angle. Combining the two ProMPs produced a shot toward the middle at medium distance; conditioning on a desired angle let it shoot in the specified direction instead.","related":["Dynamic Movement Primitives","Motion Primitives","Imitation Learning","Gaussian Mixture Regression / Task-Parameterized GMM","Demonstration Data","Kinesthetic Teaching"]},{"id":"gaussian-mixture-regression-task-parameterized-gmm","category":"control","sec":3,"tier":3,"sources":[{"title":"Calinon, A Tutorial on Task-Parameterized Movement Learning and Retrieval (Intelligent Service Robotics, 2016)","url":"https://calinon.ch/papers/Calinon-JIST2015.pdf"}],"as_of":"","related_ids":["gaussian-mixture-model","dynamic-movement-primitives","probabilistic-movement-primitives","imitation-learning","demonstration-data","coordinate-frame"],"name":"Gaussian Mixture Regression / Task-Parameterized GMM","alt":"高斯混合回归 / 任务参数化 GMM","abbr":"GMR / TP-GMM","aliases":["GMR","TP-GMM","Task-Parameterized Gaussian Mixture Model"],"one_liner":"Summarizing demonstrated trajectories with a handful of Gaussian distributions, then generating a matching motion for a new object position.","explanation":"These are two classic probabilistic methods for learning from demonstration, systematically laid out in a 2016 tutorial paper by Sylvain Calinon. A Gaussian mixture model (GMM) first fits data points of ‘time plus position’ from several demonstrated trajectories into a set of Gaussian components; GMR then takes a given input (such as time t) and computes its conditional distribution, producing a smooth trajectory along with its variance — a large variance means the demonstrations disagreed a lot at that point, so execution can be looser there. TP-GMM treats coordinate frames such as the start point or the target object as ‘task parameters,’ learning a separate GMM in each frame; for a new scene, these are transformed into the new frame and combined by a product of Gaussians, so just a few demonstrations can adapt to new object positions. It needs little data and is interpretable, but it doesn't capture complex, multimodal motion as well as deep generative methods like diffusion policy.","example":"Demonstrate placing a cup on a tray 5 times: learn one GMM in the ‘cup frame’ and another in the ‘tray frame’; for a new arrangement, multiply the two together to get a trajectory that passes through the new cup position and lands on the new tray position.","related":["Gaussian Mixture Model","Dynamic Movement Primitives","Probabilistic Movement Primitives","Imitation Learning","Demonstration Data","Coordinate Frame"]},{"id":"optimal-control","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Optimal control","url":"https://en.wikipedia.org/wiki/Optimal_control"},{"title":"Underactuated Robotics（Tedrake）: Linear Quadratic Regulators","url":"https://underactuated.mit.edu/lqr.html"},{"title":"Apollo Control 模块说明（README_cn）","url":"https://raw.githubusercontent.com/ApolloAuto/apollo/master/modules/control/control_component/README_cn.md"}],"as_of":"","related_ids":["linear-quadratic-regulator","model-predictive-control","trajectory-optimization","iterative-linear-quadratic-regulator","differential-dynamic-programming","reinforcement-learning"],"name":"Optimal Control","alt":"最优控制","abbr":"","aliases":["Optimal Control Problem","OCP"],"one_liner":"Finding, subject to the system's dynamics, a sequence of control inputs that minimizes total cost.","explanation":"Optimal control is a branch of control theory studying how to choose control inputs over a time span so an objective function is minimized. The standard form is min J = φ(x(T)) + ∫ L(x, u) dt, subject to ẋ = f(x, u): x is the state (e.g., joint angles and velocities), u the control input (e.g., torque), L the running cost (deviation from the goal, energy use), φ the terminal cost, and f the dynamics model. Its theoretical foundations are Pontryagin's maximum principle and Bellman's dynamic programming, both from the 1950s. The special case of linear dynamics with quadratic cost is the LQR, whose solution is linear feedback u = −Kx. Trajectory optimization, iLQR/DDP, and model predictive control (which re-solves a finite-horizon optimal control problem every cycle) in robotics all fall under this umbrella; reinforcement learning can be seen as solving the same class of problem when the model is unknown.","example":"Baidu Apollo's autonomous-driving lateral controller uses LQR to compute steering: it weights ‘deviation from the reference trajectory’ and ‘steering effort’ into a cost, then finds the feedback gain that minimizes total cost.","related":["Linear Quadratic Regulator","Model Predictive Control","Trajectory Optimization","Iterative Linear Quadratic Regulator","Differential Dynamic Programming","Reinforcement Learning"]},{"id":"trajectory-optimization","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Trajectory optimization","url":"https://en.wikipedia.org/wiki/Trajectory_optimization"},{"title":"Russ Tedrake, Underactuated Robotics: Trajectory Optimization","url":"https://underactuated.mit.edu/trajopt.html"},{"title":"Lynch & Park, Modern Robotics (2017 preprint), 10.7 Nonlinear Optimization","url":"https://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["optimal-control","direct-collocation","multiple-shooting","model-predictive-control","iterative-linear-quadratic-regulator","contact-implicit-trajectory-optimization"],"name":"Trajectory Optimization","alt":"轨迹优化","abbr":"TO","aliases":["TO"],"one_liner":"Writing ‘how to move’ as an optimization problem: satisfy the dynamics and constraints while minimizing a cost.","explanation":"Trajectory optimization casts ‘how to move’ as a mathematical optimization problem: the decision variables are the state and control input over a stretch of time, the objective is to minimize a cost (energy, time, deviation from a target), and the constraints include the dynamics equations, torque and joint limits, obstacle avoidance, and start/end conditions. It is essentially open-loop optimal control solved for a single initial state, which is much cheaper than solving for a feedback law over the entire state space. Numerically, direct methods are most common: direct shooting optimizes only the control, with the state obtained by simulating forward; direct collocation treats both state and control as variables, enforcing the dynamics ẋ = f(x, u) only at collocation points; multiple shooting sits between the two. Gradient-based methods generally find only a local optimum, so the initial guess matters a great deal. Re-solving every control cycle and executing only the first step is model predictive control (MPC).","example":"Srinivasan and Ruina's 2006 study in Nature used trajectory optimization to find the most energy-efficient way for a simplified biped model to move, finding that ‘walking’ was cheaper at low speed and ‘running’ cheaper at high speed.","related":["Optimal Control","Direct Collocation","Multiple Shooting","Model Predictive Control","Iterative Linear Quadratic Regulator","Contact-Implicit Trajectory Optimization"]},{"id":"model-predictive-control","category":"control","sec":4,"tier":1,"sources":[{"title":"Wikipedia: Model predictive control（含 MPC vs. LQR）","url":"https://en.wikipedia.org/wiki/Model_predictive_control"},{"title":"Qin & Badgwell, A survey of industrial model predictive control technology (Control Engineering Practice, 2003)","url":"https://doi.org/10.1016/S0967-0661(02)00186-7"},{"title":"Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"},{"title":"qiayuanl/legged_control (GitHub)","url":"https://github.com/qiayuanl/legged_control"}],"as_of":"","related_ids":["convex-mpc","nonlinear-model-predictive-control","sampling-based-mpc","whole-body-control","linear-quadratic-regulator","trajectory-optimization"],"name":"Model Predictive Control","alt":"模型预测控制","abbr":"MPC","aliases":["MPC","Receding Horizon Control"],"one_liner":"Uses a model to predict the near future at every step, optimizes an action sequence, executes only the first step, and repeats.","explanation":"Model predictive control originated in the 1970s in process industries such as oil refining and chemicals, and became widespread in those industries during the 1980s. Each control cycle it does the same thing: using a system model x_{k+1} = f(x_k, u_k) (x is the state, u is the control input) to predict N steps ahead, it solves for the sequence of inputs that minimizes some cost (such as tracking error plus energy use) subject to constraints like joint limits and torque caps — then executes only the first input, and re-solves from scratch next cycle using the newest measurement, which is why it's also called receding horizon control. Compared with PID, which only reacts to an error that has already appeared, it can plan ahead for the future and for constraints; compared with a linear quadratic regulator (LQR), which solves for a fixed feedback gain once offline, it re-optimizes online at every step, so it can directly handle constraints and nonlinear models. The cost is that it must solve an optimization online, which demands compute and depends on model accuracy. In legged robots it's commonly used with a simplified model to plan foot forces, which are then handed off to whole-body control to execute.","example":"MIT's Mini Cheetah uses MPC to optimize foot-contact forces, combined with whole-body impulse control, to run at up to 3.7 m/s; the open-source framework legged_control runs nonlinear MPC on Unitree's A1, with the authors reporting a solve rate near 200 Hz on an 11th-generation Intel NUC mini PC.","related":["Convex MPC","Nonlinear Model Predictive Control","Sampling-based MPC","Whole-Body Control","Linear Quadratic Regulator","Trajectory Optimization"]},{"id":"convex-optimization","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Convex optimization","url":"https://en.wikipedia.org/wiki/Convex_optimization"},{"title":"Boyd & Vandenberghe: Convex Optimization（官方免费电子版页面）","url":"https://web.stanford.edu/~boyd/cvxbook/"},{"title":"Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (IROS 2018)","url":"https://dspace.mit.edu/handle/1721.1/138000"}],"as_of":"","related_ids":["quadratic-programming","convex-mpc","model-predictive-control","trajectory-optimization","sequential-quadratic-programming","osqp"],"name":"Convex Optimization","alt":"凸优化","abbr":"","aliases":["Convex Programming"],"one_liner":"Minimizing a convex function over a convex feasible region — a problem where any local optimum found is also the global one.","explanation":"Convex optimization studies minimizing a convex function over a convex set: min f(x) subject to x ∈ C. f is a convex function (its graph looks like a bowl, and the line segment between any two points on it lies above the curve), and C is a convex set (the line segment between any two points in the set stays inside the set). Its most important property is that any local optimum is automatically the global optimum; categories such as linear programming, quadratic programming, second-order cone programming, and semidefinite programming all have polynomial-time algorithms (such as interior-point methods), so they solve quickly and reliably. Real-time robot control leans on this heavily: convex MPC, the quadratic programs inside whole-body control, and contact-force allocation are all deliberately formulated as convex problems specifically so they can be solved reliably within a millisecond-scale budget. Non-convex problems, such as general trajectory optimization, are commonly approximated by breaking them into a sequence of convex subproblems, an approach called sequential quadratic programming.","example":"MIT Cheetah 3 (2018) simplified the quadruped into a single rigid body and formulated planning of ground-reaction forces over a horizon of up to 0.5 s as a convex quadratic program, solving it in under 1 ms and re-solving repeatedly at 20–30 Hz, producing gaits including trotting, galloping, and pronking.","related":["Quadratic Programming","Convex MPC","Model Predictive Control","Trajectory Optimization","Sequential Quadratic Programming","OSQP"]},{"id":"quadratic-programming","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Quadratic programming","url":"https://en.wikipedia.org/wiki/Quadratic_programming"},{"title":"OSQP Documentation","url":"https://osqp.org/docs/"},{"title":"Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"}],"as_of":"","related_ids":["convex-optimization","whole-body-control","model-predictive-control","hierarchical-quadratic-programming","osqp","friction-cone"],"name":"Quadratic Programming","alt":"二次规划","abbr":"QP","aliases":["QP","Convex QP"],"one_liner":"An optimization problem with a quadratic objective and linear equality or inequality constraints, extremely common in control.","explanation":"Quadratic programming is a class of mathematical optimization problems: min ½xᵀQx + cᵀx, subject to Ax ≤ b. Here x is the variable being solved for (e.g., joint accelerations, torques, or contact forces), Q and c describe the objective (often ‘squared deviation from a desired value’), and A, b encode linear constraints (torque limits, a linearized friction cone, etc.). When Q is positive semi-definite the problem is convex, has a global optimum, and can be solved quickly with interior-point or active-set methods; OSQP and qpOASES are commonly used open-source solvers in robotics. When Q doesn't meet that condition, the problem is generally NP-hard. QP shows up constantly in robot control because ‘minimize squared tracking error subject to physical limits’ is naturally this shape: whole-body control solves a QP every cycle to allocate torques and contact forces, and convex MPC is also often written as a QP.","example":"In MIT Mini Cheetah's controller, an MPC solves a QP at 30 Hz to plan ground reaction forces at each foot, and a 500 Hz whole-body impulse controller solves a small additional QP to refine those forces before converting them into joint torques.","related":["Convex Optimization","Whole-Body Control","Model Predictive Control","Hierarchical Quadratic Programming","OSQP","Friction Cone"]},{"id":"sequential-quadratic-programming","category":"control","sec":4,"tier":3,"sources":[{"title":"Sequential quadratic programming - Wikipedia","url":"https://en.wikipedia.org/wiki/Sequential_quadratic_programming"},{"title":"OCS2 Toolbox 文档（SLQ / iLQR / SQP / IPM 求解器）","url":"https://leggedrobotics.github.io/ocs2/"},{"title":"SciPy minimize(method='SLSQP') 文档","url":"https://docs.scipy.org/doc/scipy/reference/optimize.minimize-slsqp.html"}],"as_of":"","related_ids":["quadratic-programming","trajectory-optimization","nonlinear-model-predictive-control","multiple-shooting","ocs2","acados"],"name":"Sequential Quadratic Programming","alt":"序列二次规划","abbr":"SQP","aliases":["SQP","Lagrange-Newton Method"],"one_liner":"An iterative algorithm that solves nonlinear constrained optimization by repeatedly approximating it as a quadratic program.","explanation":"SQP solves optimization problems where both the objective and the constraints are smooth nonlinear functions. Each iteration, near the current solution, it approximates the Lagrangian (the objective plus the constraints weighted by multipliers) as a quadratic function and linearizes the constraints, producing a quadratic programming (QP) subproblem; solving that subproblem gives a search direction, the solution is updated, and the process repeats until the Karush-Kuhn-Tucker (KKT) optimality conditions are satisfied. This is essentially Newton's method extended to constrained problems — it converges fast, but only guarantees a local optimum, depends on a good initial guess, and needs a line search or trust region to avoid diverging. Robot trajectory optimization and nonlinear MPC rely on it heavily, writing dynamics, joint limits, and obstacle avoidance as constraints; online MPC can warm-start from the previous cycle's solution, and real-time-iteration (RTI) schemes even perform just one SQP round per cycle. The other major family of solvers, running alongside SQP, is interior-point methods such as Ipopt.","example":"ETH Zurich's open-source optimal control library OCS2 provides a multiple-shooting SQP solver built on HPIPM, used for nonlinear MPC on quadrupeds and mobile manipulators. SciPy's minimize(method='SLSQP') is also a form of SQP and can be used directly for small-scale inverse kinematics or parameter fitting.","related":["Quadratic Programming","Trajectory Optimization","Nonlinear Model Predictive Control","Multiple Shooting","OCS2","acados (fast embedded optimal control solver)"]},{"id":"linear-quadratic-regulator","category":"control","sec":4,"tier":3,"sources":[{"title":"Underactuated Robotics (Russ Tedrake), Ch. Linear Quadratic Regulators","url":"https://underactuated.mit.edu/lqr.html"},{"title":"Linear–quadratic regulator - Wikipedia","url":"https://en.wikipedia.org/wiki/Linear%E2%80%93quadratic_regulator"}],"as_of":"","related_ids":["linear-quadratic-gaussian-control","iterative-linear-quadratic-regulator","optimal-control","model-predictive-control","proportional-integral-derivative-control","inverted-pendulum-model"],"name":"Linear Quadratic Regulator","alt":"线性二次调节器","abbr":"LQR","aliases":["LQR"],"one_liner":"The optimal feedback controller for a linear system: trading off state error against control effort, solving for a fixed gain u = −Kx.","explanation":"LQR is the most fundamental optimal-control method, with its theory laid down by Kálmán and others around 1960. It assumes the system is linear: ẋ = Ax + Bu, where x is the state, u the control input, and A, B matrices describing the system; the cost is quadratic, the time integral of xᵀQx + uᵀRu. Q reflects how much you care about state deviation, R how much you care about control effort, and both are tuned by the designer. Solving a Riccati equation gives the gain matrix K, and the optimal control is u = −Kx — at runtime this is just a single matrix multiplication. Compared to PID, LQR handles multiple-input, multiple-output systems naturally and directly trades off ‘error’ against ‘effort’; compared to MPC, it can't handle constraints like torque limits. Robots are nonlinear, so a common approach is to linearize around an equilibrium point and apply LQR there, valid only near that point; linearizing at every point along a trajectory instead gives a time-varying LQR, which is also the basis of iLQR.","example":"Linearizing an inverted pendulum or an Acrobot (a two-link gymnast robot) around its upright equilibrium and computing an LQR gain lets it be balanced steadily upright; the balance controller on a two-wheeled self-balancing vehicle is often built the same way.","related":["Linear Quadratic Gaussian Control","Iterative Linear Quadratic Regulator","Optimal Control","Model Predictive Control","Proportional-Integral-Derivative Control","Inverted Pendulum Model (IPM)"]},{"id":"state-observer","category":"control","sec":4,"tier":3,"sources":[{"title":"State observer - Wikipedia","url":"https://en.wikipedia.org/wiki/State_observer"}],"as_of":"","related_ids":["kalman-filter","disturbance-observer","generalized-momentum-observer","state-estimation","extended-kalman-filter","linear-quadratic-gaussian-control"],"name":"State Observer","alt":"状态观测器","abbr":"","aliases":["Luenberger Observer","State Estimator"],"one_liner":"A model-based estimator that reconstructs a system's unmeasured internal states in real time from its known inputs and outputs.","explanation":"Many controllers need the full state — such as joint velocity or body velocity — but sensors typically measure only part of it. A state observer uses the system model together with the measurable inputs and outputs to reconstruct the unmeasured internal states in real time. It maintains an estimate x̂ that evolves alongside the system model and is corrected by the difference between the measured output y and the predicted output Cx̂: dx̂/dt = A x̂ + B u + L (y − C x̂), where A, B, C are the linear model matrices, u is the control input, and L is the observer gain — a larger L corrects faster but also amplifies noise more. This is the Luenberger observer, proposed by David Luenberger, and it requires the system to be observable, meaning the state can in principle be inferred uniquely from the output; by the separation principle, state feedback and the observer can be designed independently. A Kalman filter can be viewed as an observer whose gain L is chosen optimally from the noise statistics, and specialized variants such as the disturbance observer and momentum observer used in robotics estimate external forces rather than the full state.","example":"A motor has only an encoder measuring position θ. Setting the state to [θ, ω] (with ω the angular velocity) and treating the current-derived torque as the input u, the observer corrects the position estimate while also producing an estimate of ω — one that is less noisy than simply differentiating the measured position to get velocity.","related":["Kalman Filter","Disturbance Observer","Generalized Momentum Observer","State Estimation","Extended Kalman Filter","Linear Quadratic Gaussian Control"]},{"id":"kalman-filter","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Kalman filter","url":"https://en.wikipedia.org/wiki/Kalman_filter"},{"title":"Kalman, A New Approach to Linear Filtering and Prediction Problems, J. Basic Engineering 82 (1960)","url":"https://doi.org/10.1115/1.3662552"},{"title":"MIT Cheetah-Software: PositionVelocityEstimator.h（LinearKFPositionVelocityEstimator）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/common/include/Controllers/PositionVelocityEstimator.h"}],"as_of":"","related_ids":["extended-kalman-filter","unscented-kalman-filter","state-estimation","multi-sensor-fusion","inertial-measurement-unit","leg-odometry"],"name":"Kalman Filter","alt":"卡尔曼滤波","abbr":"KF","aliases":["KF","Linear Kalman Filter"],"one_liner":"Fusing a model's prediction with a noisy measurement, weighted by how much each is trusted, to estimate a system's state.","explanation":"Proposed by Rudolf Kálmán in his 1960 paper ‘A New Approach to Linear Filtering and Prediction Problems,’ and later used for orbit estimation in the Apollo program. Each cycle has two steps: prediction, where a motion model projects the current state and its uncertainty (covariance) forward; and update, where a new measurement z corrects that prediction via x̂ = x̂⁻ + K(z − Hx̂⁻), with x̂⁻ the predicted value, H mapping the state to the measurement it should produce, and K the Kalman gain — the more trustworthy the measurement, the larger K becomes. For a linear model with Gaussian noise of known covariance, this is the optimal estimator. Robot models are usually nonlinear, so variants such as the Extended Kalman Filter (EKF) and Unscented Kalman Filter (UKF) are commonly used instead, fusing IMU, encoder, and vision data to estimate a robot body's pose and velocity.","example":"The open-source control code for MIT's Cheetah 3 and Mini Cheetah uses a linear Kalman filter to estimate body position and velocity: IMU acceleration drives the prediction step, leg kinematics gives each foot's relative position and velocity as the measurement, and trust in each leg is adjusted based on whether it is touching the ground.","related":["Extended Kalman Filter","Unscented Kalman Filter","State Estimation","Multi-Sensor Fusion","Inertial Measurement Unit","Leg Odometry"]},{"id":"extended-kalman-filter","category":"control","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Extended Kalman filter","url":"https://en.wikipedia.org/wiki/Extended_Kalman_filter"},{"title":"robot_localization 文档首页（ekf_localization_node，15 维状态）","url":"https://github.com/cra-ros-pkg/robot_localization/blob/ros2/doc/index.rst"},{"title":"State Estimation for Legged Robots - Consistent Fusion of Leg Kinematics and IMU (Bloesch et al., RSS 2012)","url":"https://www.roboticsproceedings.org/rss08/p03.pdf"}],"as_of":"","related_ids":["kalman-filter","unscented-kalman-filter","error-state-kalman-filter","state-estimation","leg-odometry","multi-sensor-fusion"],"name":"Extended Kalman Filter","alt":"扩展卡尔曼滤波","abbr":"EKF","aliases":["EKF"],"one_liner":"Linearizes a nonlinear system around the current estimate, then applies the ordinary Kalman filter for state estimation.","explanation":"The extended Kalman filter generalizes the Kalman filter to nonlinear systems, and was developed mainly at NASA's Ames Research Center for navigation problems in the 1960s. The standard Kalman filter only applies to linear models; at every step, the EKF uses a Jacobian matrix (the partial derivatives of each output with respect to each state) to take a first-order Taylor expansion of the nonlinear motion and observation models around the current estimate, then runs the usual predict-update cycle: first the motion model projects forward the new state and its uncertainty, then a sensor reading corrects it. It's computationally cheap and its implementation is mature, making it the de facto standard in navigation; but it's only an approximation, and can diverge when nonlinearity is strong, in which case an unscented Kalman filter or an error-state Kalman filter is often used instead. Legged robots commonly use it to fuse IMU data with leg kinematics to estimate the body's pose and velocity.","example":"ROS's robot_localization package provides an ekf_localization_node that can fuse any number of input sources — wheel odometry, IMU, GPS, and so on — into a 15-dimensional state output: 3D position, orientation (roll/pitch/yaw), linear velocity, angular velocity, and linear acceleration.","related":["Kalman Filter","Unscented Kalman Filter","Error-State Kalman Filter","State Estimation","Leg Odometry","Multi-Sensor Fusion"]},{"id":"linear-quadratic-gaussian-control","category":"control","sec":4,"tier":3,"sources":[{"title":"Linear–quadratic–Gaussian control - Wikipedia","url":"https://en.wikipedia.org/wiki/Linear%E2%80%93quadratic%E2%80%93Gaussian_control"}],"as_of":"","related_ids":["linear-quadratic-regulator","kalman-filter","state-estimation","optimal-control","robust-control","state-observer"],"name":"Linear Quadratic Gaussian Control","alt":"线性二次高斯控制","abbr":"LQG","aliases":["LQG"],"one_liner":"When the state is only partially measured and noisy, first estimate it with a Kalman filter, then compute control with LQR.","explanation":"LQG is a classic optimal-control problem: the system is linear, both the process and the measurements carry Gaussian white noise, the state can't be fully measured directly, and the goal is to minimize the expected value of a quadratic cost. The solution splits into two pieces: a Kalman filter estimates the current state x̂ from noisy sensor readings, and LQR then computes the control from that estimate, u = −K·x̂ (K the LQR gain). The two pieces can be designed separately, each optimal on its own, and the combination is still overall optimal — this is the separation principle. The difference from plain LQR is that LQR assumes the full state is known exactly, while LQG faces the real-world case of partial, noisy measurement. It's worth noting that LQR comes with good stability margins by default, while LQG offers no such guarantee — a point Doyle's 1978 paper specifically raised — so robustness needs to be checked separately in engineering practice. LQG is a foundational model for understanding how robots split ‘state estimation’ from ‘control.’","example":"A cart-pole balancing system with only encoders measuring cart position and pole angle, where velocity has to be estimated and the readings are noisy: a Kalman filter estimates all four states — position, velocity, angle, and angular velocity — and multiplying by the LQR gain gives the motor command, together forming an LQG controller.","related":["Linear Quadratic Regulator","Kalman Filter","State Estimation","Optimal Control","Robust Control","State Observer"]},{"id":"iterative-linear-quadratic-regulator","category":"control","sec":4,"tier":3,"sources":[{"title":"Underactuated Robotics (Russ Tedrake), Ch. Trajectory Optimization: Iterative LQR and DDP","url":"https://underactuated.mit.edu/trajopt.html"},{"title":"Differential dynamic programming - Wikipedia","url":"https://en.wikipedia.org/wiki/Differential_dynamic_programming"},{"title":"MuJoCo MPC (MJPC) README","url":"https://github.com/google-deepmind/mujoco_mpc"}],"as_of":"","related_ids":["linear-quadratic-regulator","differential-dynamic-programming","trajectory-optimization","model-predictive-control","nonlinear-model-predictive-control","mujoco-mpc"],"name":"Iterative Linear Quadratic Regulator","alt":"迭代线性二次调节器","abbr":"iLQR","aliases":["iLQR","Iterative LQR"],"one_liner":"A trajectory-optimization algorithm that repeatedly linearizes a nonlinear system along its current trajectory and solves an LQR each round.","explanation":"iLQR is a trajectory-optimization algorithm generally credited to Li and Todorov's 2004 paper. LQR only handles linear systems, but robot dynamics is nonlinear. iLQR's approach: start with an initial control sequence and simulate out a trajectory; at every point along it, linearize the dynamics and take a quadratic approximation of the cost, giving a time-varying LQR problem; run a Riccati recursion backward from the endpoint (a backward pass) to get a correction and a feedback gain at each step, then re-simulate a new trajectory forward from the start (a forward pass); repeat until convergence. It's a simplified version of differential dynamic programming (DDP, proposed by Mayne in 1966): DDP also uses the dynamics' second derivatives, which iLQR drops, and in practice the two converge at similar speed while iLQR is cheaper to compute. Besides a trajectory, the result comes with feedback gains along the way. Combined with MPC, it's re-solved from the current state every control cycle; the noise-aware extension is called iLQG.","example":"DeepMind's open-source MuJoCo MPC (MJPC) has a built-in iLQG planner: every control cycle it re-optimizes a future window of control starting from the current state, used to control a simulated robot model in real time.","related":["Linear Quadratic Regulator","Differential Dynamic Programming","Trajectory Optimization","Model Predictive Control","Nonlinear Model Predictive Control","MuJoCo MPC (MJPC)"]},{"id":"differential-dynamic-programming","category":"control","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Differential dynamic programming","url":"https://en.wikipedia.org/wiki/Differential_dynamic_programming"},{"title":"loco-3d/crocoddyl（solvers based on DDP algorithms）","url":"https://github.com/loco-3d/crocoddyl"}],"as_of":"","related_ids":["iterative-linear-quadratic-regulator","linear-quadratic-regulator","trajectory-optimization","model-predictive-control","optimal-control","crocoddyl"],"name":"Differential Dynamic Programming","alt":"微分动态规划","abbr":"DDP","aliases":["DDP"],"one_liner":"A trajectory-optimization method that repeatedly makes a second-order approximation along the current trajectory and sweeps back and forth to improve the control sequence.","explanation":"DDP is an iterative algorithm for solving nonlinear optimal control, proposed by David Mayne in 1966 and later systematized in a book of the same name with Jacobson. Given an initial control sequence, it repeats two steps: a backward pass, which takes a second-order expansion of the dynamics and cost around the current trajectory and works backward from the endpoint to compute a correction k = −Q_uu⁻¹Q_u and a feedback gain K = −Q_uu⁻¹Q_ux at each step, where Q is the total cost of ‘choosing control u this step and then acting optimally from then on,’ with subscripts denoting derivatives with respect to u or the state x; and a forward pass, which re-simulates a new trajectory using u = ū + αk + K(x − x̄) (ū, x̄ being the old trajectory and α a line-search step size). It converges quadratically near the optimum and produces feedback gains as a byproduct, making it well suited to MPC. Dropping the dynamics' second-derivative terms gives the more commonly used iLQR. DDP belongs to the shooting-method family, with the state obtained by simulation, which handles state constraints less conveniently than collocation methods.","example":"The open-source library Crocoddyl uses DDP and its variant FDDP as its core solvers, leveraging Pinocchio's analytical derivatives to compute optimal trajectories — and the accompanying feedback gains — with contact sequences for legged robots and similar systems.","related":["Iterative Linear Quadratic Regulator","Linear Quadratic Regulator","Trajectory Optimization","Model Predictive Control","Optimal Control","Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)"]},{"id":"direct-collocation","category":"control","sec":4,"tier":3,"sources":[{"title":"Underactuated Robotics (Tedrake), Ch. Trajectory Optimization","url":"https://underactuated.mit.edu/trajopt.html"},{"title":"Drake: DirectCollocation Class Reference","url":"https://drake.mit.edu/doxygen_cxx/classdrake_1_1planning_1_1trajectory__optimization_1_1_direct_collocation.html"},{"title":"Matthew Kelly: Trajectory Optimization tutorials","url":"https://www.matthewpeterkelly.com/tutorials/trajectoryOptimization/index.html"}],"as_of":"","related_ids":["trajectory-optimization","multiple-shooting","sequential-quadratic-programming","contact-implicit-trajectory-optimization","drake","interior-point-optimizer"],"name":"Direct Collocation","alt":"直接配点法","abbr":"","aliases":["Collocation Method"],"one_liner":"Cutting a trajectory into nodes, writing the dynamics as constraints between nodes, and turning the whole thing into one big optimization problem.","explanation":"Direct collocation is one of the most common ‘transcription’ methods in trajectory optimization, proposed by Hargraves and Paris in 1987, and it works by turning a continuous-time optimal control problem into a finite-dimensional nonlinear program (NLP). Time is cut into N segments, with the state x_k and control u_k at every node treated as optimization variables, connected between nodes by polynomials (commonly piecewise-linear control and cubic-spline state); the polynomial's derivative is then required to match the dynamics ẋ = f(x, u) at collocation points (often the segment midpoints), with these equality constraints enforcing physical validity. Shooting methods optimize only the control, with the state obtained by simulating forward, which becomes very sensitive to the initial guess over a long horizon; collocation treats the state as a variable too, so an initial guess can be sketched freely, state constraints are easy to add, and the resulting problem is sparse, suiting solvers like IPOPT or SNOPT. The drawback is more variables, and the trajectory partway through solving may not satisfy the dynamics at all. Multiple shooting sits between the two approaches.","example":"Swinging up a cart-pole: cut the whole time span into nodes, with the cart position, pole angle, velocities, and applied force at each node as variables, constraints requiring adjacent nodes to satisfy the dynamics and the pole to end upright, and the objective minimizing total squared force — solving gives a swing-up trajectory. Matthew Kelly's introductory tutorial in SIAM Review focuses on exactly this method, and Drake includes a ready-made DirectCollocation class.","related":["Trajectory Optimization","Multiple Shooting","Sequential Quadratic Programming","Contact-Implicit Trajectory Optimization","Drake","Interior Point OPTimizer (Ipopt)"]},{"id":"multiple-shooting","category":"control","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Direct multiple shooting method","url":"https://en.wikipedia.org/wiki/Direct_multiple_shooting_method"},{"title":"Perceptive Locomotion through Nonlinear Model Predictive Control (arXiv 2208.08373)","url":"https://arxiv.org/abs/2208.08373"}],"as_of":"","related_ids":["direct-collocation","trajectory-optimization","sequential-quadratic-programming","nonlinear-model-predictive-control","optimal-control","acados"],"name":"Multiple Shooting","alt":"多重打靶法","abbr":"","aliases":["Direct Multiple Shooting"],"one_liner":"Cutting a long trajectory into segments, integrating each one separately, then stitching them together with continuity constraints.","explanation":"Multiple shooting is a transcription method for solving trajectory optimization and optimal control problems — turning a continuous-time problem into a finite-dimensional nonlinear program. Consider single shooting first: only the control is treated as an optimization variable, integrated in one continuous run from the initial state to the endpoint, checking whether it ‘hits’ the target and adjusting; over a long trajectory, or with an unstable system, small changes in the control get amplified, making convergence hard. Multiple shooting instead cuts time into several segments, treating each segment's starting state as an optimization variable too, integrating each segment separately, then adding matching conditions: the state at the end of one segment's integration must equal the state at the start of the next. This spreads the nonlinearity across short segments, making the numerics far more stable, and the segments can even be integrated in parallel. Bock and Plitt applied it to optimal control in 1984. It is one of two mainstream transcription methods alongside direct collocation, with the resulting problem commonly solved by sequential quadratic programming (SQP); NMPC tools such as OCS2 and acados both support it.","example":"ANYmal's perceptive NMPC (Grandia et al., 2022) discretizes a future motion window using multiple shooting, solving with SQP at 100 Hz to generate, in real time, motions that cross gaps and step across stepping stones.","related":["Direct Collocation","Trajectory Optimization","Sequential Quadratic Programming","Nonlinear Model Predictive Control","Optimal Control","acados (fast embedded optimal control solver)"]},{"id":"contact-implicit-trajectory-optimization","category":"control","sec":4,"tier":3,"sources":[{"title":"Posa, Cantu, Tedrake: A Direct Method for Trajectory Optimization of Rigid Bodies Through Contact","url":"https://groups.csail.mit.edu/robotics-center/public_papers/Posa13.pdf"},{"title":"Le Cleac'h et al., Fast Contact-Implicit Model-Predictive Control (arXiv:2107.05616)","url":"https://arxiv.org/abs/2107.05616"}],"as_of":"","related_ids":["trajectory-optimization","linear-complementarity-problem","sequential-quadratic-programming","direct-collocation","multi-contact-planning","gait-planning"],"name":"Contact-Implicit Trajectory Optimization","alt":"接触隐式轨迹优化","abbr":"","aliases":["Contact-Implicit MPC","CI-MPC","Through-Contact Trajectory Optimization"],"one_liner":"Letting the optimizer decide, on its own, when and where to contact the environment, instead of fixing the contact order in advance.","explanation":"Traditional trajectory optimization treats events like ‘foot touches down’ or ‘hand touches object’ as discrete modes, requiring a human to specify the mode order in advance (left foot first, then right), then optimizing within each segment. For tasks like multi-finger manipulation, where contact changes constantly, there's simply no way to write that order down beforehand. MIT's Posa, Cantu, and Tedrake proposed a direct method (WAFR 2012, journal version in IJRR 2014) that makes contact force itself a decision variable, describing contact with complementarity constraints: distance φ ≥ 0, normal force λ ≥ 0, and φ·λ = 0 — meaning ‘no contact, no force; force means touching’ — solved with sequential quadratic programming, so the contact schedule and the trajectory come out together. These constraints originate from the linear complementarity problem used in physics engines. The cost is that the problem is non-convex and hard to converge, often needing relaxation or smoothing; later work developed contact-implicit MPC that can run online.","example":"Posa et al. used the same method to plan a finger rotating a fixed-axis object, simple grasping and manipulation tasks, planar walking for Spring Flamingo, and high-speed bipedal running for FastRunner, all without specifying a contact order in advance; Le Cleac'h et al. (2021) proposed a fast contact-implicit MPC that generated non-periodic motions in real time on a real quadruped robot.","related":["Trajectory Optimization","Linear Complementarity Problem","Sequential Quadratic Programming","Direct Collocation","Multi-contact Planning","Gait Planning"]},{"id":"nonlinear-model-predictive-control","category":"control","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Model predictive control","url":"https://en.wikipedia.org/wiki/Model_predictive_control"},{"title":"GitHub: qiayuanl/legged_control","url":"https://github.com/qiayuanl/legged_control"},{"title":"Perceptive Locomotion through Nonlinear Model Predictive Control (arXiv 2208.08373)","url":"https://arxiv.org/abs/2208.08373"}],"as_of":"","related_ids":["model-predictive-control","convex-mpc","multiple-shooting","sequential-quadratic-programming","whole-body-control","legged-control"],"name":"Nonlinear Model Predictive Control","alt":"非线性模型预测控制","abbr":"NMPC","aliases":["NMPC","Nonlinear MPC"],"one_liner":"MPC that uses a nonlinear dynamics model to predict the future and solves an optimization online every control cycle.","explanation":"Model predictive control (MPC) works by using a system model, every control cycle, to predict a future window of states, solving an optimization for the sequence of control inputs that minimizes cost subject to constraints, then executing only the first one — recomputing with a fresh measurement next cycle, a scheme called receding-horizon control. NMPC is MPC whose model or constraints are nonlinear — for example, working directly with a robot's full rigid-body or centroidal dynamics rather than linearizing first. The benefit is more accurate prediction during large, fast motion, and the ability to directly handle nonlinear constraints like friction cones, foothold regions, and obstacle avoidance; the cost is a non-convex optimization that's computationally heavy and offers no guarantee of finding the global optimum. In practice the problem is usually discretized with multiple shooting or direct collocation and solved with sequential quadratic programming (SQP), using ‘real-time iteration’: only one or two iterations per cycle, without waiting for full convergence, to keep up with the control rate. Common tools include acados, OCS2, and CasADi. Convex MPC instead simplifies the model to something linear, making the problem convex and faster, but narrower in scope.","example":"The open-source framework legged_control uses NMPC (built on OCS2, solved with multiple shooting plus SQP) on a Unitree A1 quadruped to convert a desired torso velocity into a future state trajectory, which is then handed to a whole-body controller to compute joint torques.","related":["Model Predictive Control","Convex MPC","Multiple Shooting","Sequential Quadratic Programming","Whole-Body Control","legged_control (NMPC + WBC framework for legged robots)"]},{"id":"sampling-based-mpc","category":"control","sec":4,"tier":3,"sources":[{"title":"Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo (Howell et al., arXiv 2212.00541)","url":"https://arxiv.org/abs/2212.00541"},{"title":"Information Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving (Williams et al., arXiv 1707.02342)","url":"https://arxiv.org/abs/1707.02342"},{"title":"Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing (DIAL-MPC, arXiv 2409.15610)","url":"https://arxiv.org/abs/2409.15610"}],"as_of":"","related_ids":["model-predictive-control","model-predictive-path-integral-control","cross-entropy-method","dial-mpc","mujoco-mpc","trajectory-optimization"],"name":"Sampling-based MPC","alt":"采样式 MPC","abbr":"","aliases":["Sampling-Based Model Predictive Control"],"one_liner":"MPC that samples many candidate action sequences each cycle, scores them by simulation, and executes the best one's first step.","explanation":"Model predictive control (MPC) optimizes a short sequence of future actions every cycle but executes only the first step before replanning. Traditional solvers rely on gradients or quadratic programming, which requires a differentiable, smooth cost function. Sampling-based MPC instead samples tens to thousands of candidate action sequences by adding noise around the previous optimal sequence, rolls them out in parallel through a dynamics model (often a physics simulator) to score each one, and then picks the best or takes a cost-weighted average to form the new sequence. Representative methods include MPPI (Model Predictive Path Integral control), which weights samples proportional to exp(−cost/λ) with temperature λ; the cross-entropy method (CEM), which keeps the lowest-cost batch and refits the sampling distribution; and Predictive Sampling, the simplest version, released by DeepMind alongside MuJoCo MPC in 2022. These methods need no derivatives, can handle non-smooth dynamics such as contact, and parallelize well on GPUs — though sampling efficiency drops as the action dimension grows.","example":"DIAL-MPC, proposed by CMU and others in 2024, borrows the step-by-step denoising and annealing idea from diffusion models to do force-level sampling optimization directly on full quadruped dynamics without any training. The paper reports tracking error 13.4 times lower than standard MPPI, and demonstrated precise loaded jumps on a real robot.","related":["Model Predictive Control","Model Predictive Path Integral Control","Cross-Entropy Method","DIAL-MPC","MuJoCo MPC (MJPC)","Trajectory Optimization"]},{"id":"model-predictive-path-integral-control","category":"control","sec":4,"tier":3,"sources":[{"title":"Williams et al.: Information Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving (arXiv:1707.02342, T-RO)","url":"https://arxiv.org/abs/1707.02342"},{"title":"Williams, Aldrich, Theodorou: Model Predictive Path Integral Control using Covariance Variable Importance Sampling (arXiv:1509.01149)","url":"https://arxiv.org/abs/1509.01149"},{"title":"Nav2 MPPI Controller README","url":"https://github.com/ros-navigation/navigation2/blob/main/nav2_mppi_controller/README.md"}],"as_of":"2026-09","related_ids":["model-predictive-control","sampling-based-mpc","cross-entropy-method","dial-mpc","ros-2-navigation-stack","optimal-control"],"name":"Model Predictive Path Integral Control","alt":"模型预测路径积分控制","abbr":"MPPI","aliases":["MPPI","Information-Theoretic MPC","IT-MPC"],"one_liner":"A sampling-based MPC that randomly samples large numbers of control sequences each cycle, simulates them, and averages them weighted by cost.","explanation":"MPPI is a sampling-based form of model predictive control, proposed between 2015 and 2017 by Georgia Tech's Grady Williams, Evangelos Theodorou, and colleagues, with theoretical roots in path-integral optimal control and information theory. Like ordinary MPC, it executes only the first step of the optimized result each cycle, then starts over from the new state. What's different is how it solves: centered on the previous cycle's control sequence, it adds random noise and samples hundreds or thousands of candidate control sequences, forward-simulates each with the dynamics model to get a cost S, then computes a weighted average of all of them using exp(−S/λ), where λ is a temperature parameter — smaller values favor the lowest-cost samples more heavily. It needs no gradients or linearization, the cost function need not be smooth (a collision penalty is fine), the dynamics can even be a neural network, and each sample is independent, making it well suited to GPU parallelization. The drawback is needing large numbers of samples, with efficiency dropping as action dimension grows. Compared to the cross-entropy method, which uses only the lowest-cost batch of samples to update, MPPI instead uses every sample, weighted exponentially.","example":"On a 1:5-scale AutoRally off-road car, MPPI simulated thousands of 2–3-second trajectories in parallel on a GPU at a 40–60 Hz control rate, executing aggressive driving on a dirt track. ROS 2's Nav2 navigation framework also includes a built-in MPPI controller, sampling 1,000 trajectories per cycle by default and running above 50 Hz on an ordinary CPU.","related":["Model Predictive Control","Sampling-based MPC","Cross-Entropy Method","DIAL-MPC","ROS 2 Navigation Stack (Nav2)","Optimal Control"]},{"id":"cross-entropy-method","category":"control","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Cross-entropy method","url":"https://en.wikipedia.org/wiki/Cross-entropy_method"},{"title":"Hafner et al., Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv:1811.04551)","url":"https://arxiv.org/abs/1811.04551"},{"title":"Kalashnikov et al., QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (arXiv:1806.10293)","url":"https://arxiv.org/abs/1806.10293"}],"as_of":"","related_ids":["sampling-based-mpc","model-predictive-path-integral-control","model-predictive-control","model-based-reinforcement-learning","planet","qt-opt"],"name":"Cross-Entropy Method","alt":"交叉熵方法","abbr":"CEM","aliases":["CEM","CEM Planning","CEM Optimization"],"one_liner":"A gradient-free optimizer that repeatedly samples a batch, keeps the best, and refits the sampling distribution to them.","explanation":"The cross-entropy method was proposed by Rubinstein in the late 1990s, originally for estimating rare-event probabilities, and later became a general stochastic optimization method. The procedure: sample a batch of candidate solutions from a distribution (usually Gaussian), score each one, keep the best fraction of them (the ‘elite’ samples), refit the distribution's mean and variance from the elites, then sample again; after a few rounds, the distribution concentrates around good solutions. It needs no gradients and is easy to parallelize, making it a good fit when the objective is a simulator or a neural network — a black box. Two common uses in robotics: as a sampling-based MPC, running CEM over a sequence of future actions and executing only the first step before replanning; and finding the action that maximizes a Q-function in a continuous action space. It differs from MPPI in that CEM uses only the elite samples with equal weighting, while MPPI weights and uses all samples according to an exponential function of their cost.","example":"PlaNet plans inside a learned world model using CEM: a 12-step horizon, sampling 1,000 action sequences per round, keeping the best 100, over 10 iterations. QT-Opt uses CEM to find the best grasp action on a Q-function, sampling 64 per round, keeping the best 6, over 2 iterations, used both for computing target values during training and for selecting actions on the real robot.","related":["Sampling-based MPC","Model Predictive Path Integral Control","Model Predictive Control","Model-Based Reinforcement Learning","PlaNet","QT-Opt"]},{"id":"visual-foresight","category":"control","sec":4,"tier":3,"sources":[{"title":"Finn, Levine: Deep Visual Foresight for Planning Robot Motion (arXiv:1610.00696, ICRA 2017)","url":"https://arxiv.org/abs/1610.00696"},{"title":"Ebert et al.: Visual Foresight: Model-Based Deep RL for Vision-Based Robotic Control (arXiv:1812.00568)","url":"https://arxiv.org/abs/1812.00568"},{"title":"DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (arXiv:2411.04983)","url":"https://arxiv.org/abs/2411.04983"}],"as_of":"2025-06","related_ids":["world-model","video-prediction-model","model-predictive-control","cross-entropy-method","dino-wm","v-jepa-2"],"name":"Visual Foresight","alt":"视觉预见 / 基于学习模型的规划","abbr":"","aliases":["Visual MPC","Planning with Learned World Models","Visual Model Predictive Control"],"one_liner":"Learning to predict what the camera view will look like after an action, then picking whichever imagined action reaches the goal best.","explanation":"Visual foresight was proposed by Chelsea Finn and Sergey Levine at UC Berkeley (as ‘deep visual foresight’) at ICRA 2017, and assembled into a complete framework by Frederik Ebert, Finn, and colleagues in 2018. It works in two steps. First, an action-conditioned video prediction model is trained on unlabeled data collected by a robot autonomously pushing objects around: given the current image and a sequence of candidate actions, it predicts the next several frames. Second, the system performs visual model predictive control (MPC): it samples a large number of candidate action sequences — often refined iteratively with the cross-entropy method — lets the model ‘imagine’ the outcome of each, scores them by how close they get to the goal, executes only the first step of the best sequence, and then replans. The goal can be ‘move this pixel here,’ a target image, or a classifier. The approach needs no reward function or human labels and works on objects it has never seen. It is the direct predecessor of today's ‘world model plus planning’ approaches — DINO-WM and V-JEPA 2-AC, for example, move the prediction from pixels into a pretrained feature space but still use MPC to choose actions.","example":"You point at a pixel on some object in the robot's camera view and specify where it should end up. The robot imagines trying a large number of action sequences, picks whichever one's predicted outcome pushes that pixel closest to the target, executes just the first step, and then replans.","related":["World Model","Video Prediction Model","Model Predictive Control","Cross-Entropy Method","DINO-WM","V-JEPA 2"]},{"id":"gait","category":"control","sec":5,"tier":1,"sources":[{"title":"Wikipedia: Gait（duty factor, symmetrical/asymmetrical gaits）","url":"https://en.wikipedia.org/wiki/Gait"},{"title":"Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior (arXiv:2212.03238)","url":"https://arxiv.org/html/2212.03238"}],"as_of":"","related_ids":["trot-gait","pace-gait","bound-gait","gait-cycle-and-duty-factor","gait-planning","legged-locomotion"],"name":"Gait","alt":"步态","abbr":"","aliases":["Gait Pattern","Locomotion Gait"],"one_liner":"The sequence and rhythm in which a legged animal or robot's legs lift and touch down while moving.","explanation":"Gait refers to the timing pattern of when each leg touches down and lifts off during legged locomotion. It's commonly described with two quantities: duty factor, the fraction of a gait cycle a leg spends on the ground (above 50% is generally considered walking, below 50% running), and the phase offsets between legs. Common quadruped gaits include the trot (diagonal leg pairs together), the pace (same-side leg pairs together), the bound (front pair and rear pair each together), the pronk (all four legs together), and the gallop; humanoid robots mainly alternate between two-legged walking and running gaits. Traditional controllers usually fix the gait timing first and then plan footholds and support forces around it; reinforcement-learning locomotion controllers can let a gait emerge naturally from just a velocity command, or take gait frequency and inter-leg phase offsets as explicit command inputs to specify which gait to use.","example":"Walk These Ways specifies a quadruped's gait using three inter-leg phase offsets: (0.5, 0, 0) is trotting, (0, 0.5, 0) is bounding, (0, 0, 0.5) is pacing, and all zeros is pronking — the same policy can switch gaits just by changing this input.","related":["Trot Gait","Pace Gait","Bound Gait","Gait Cycle and Duty Factor","Gait Planning","Legged Locomotion"]},{"id":"trot-gait","category":"control","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Trot","url":"https://en.wikipedia.org/wiki/Trot"},{"title":"Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior (arXiv:2212.03238)","url":"https://arxiv.org/abs/2212.03238"}],"as_of":"","related_ids":["gait","pace-gait","bound-gait","gallop-gait","gait-cycle-and-duty-factor","support-polygon"],"name":"Trot Gait","alt":"对角小跑步态","abbr":"","aliases":["Trotting"],"one_liner":"A quadruped gait where diagonal leg pairs touch down together and the two pairs alternate.","explanation":"The trot originally describes a horse's gait: the left front and right rear leg form one pair, the right front and left rear form the other, and the two pairs alternate touching down — a two-beat rhythm, with a moment where all four hooves leave the ground between beats when a horse trots fast. It's also the most common basic gait for quadruped robots. Compared to other two-beat gaits: pacing moves both legs on the same side together, bounding pairs the two front legs and the two back legs, and pronking lifts and lands all four legs together. During a trot, only two feet are usually on the ground at once, so the support polygon collapses to a diagonal line — the robot can't stand still stably and needs continuous stepping to maintain dynamic balance; the benefit is front-back and left-right symmetry and a wide usable speed range.","example":"MIT's Walk These Ways (2022) used a single policy to switch among multiple gaits on a Unitree Go1, describing each leg's touchdown timing with three phase offsets: (0.5, 0, 0) is a trot, (0, 0, 0) is a synchronized four-leg pronk, (0, 0.5, 0) is a bound, and (0, 0, 0.5) is a pace; a separate stepping-frequency parameter, e.g. 3 Hz, means each foot touches down 3 times per second.","related":["Gait","Pace Gait","Bound Gait","Gallop Gait","Gait Cycle and Duty Factor","Support Polygon"]},{"id":"gait-planning","category":"control","sec":5,"tier":2,"sources":[{"title":"MIT Cheetah-Software ConvexMPCLocomotion.cpp（步态偏移与时长定义）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/user/MIT_Controller/Controllers/convexMPC/ConvexMPCLocomotion.cpp"},{"title":"Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior (arXiv 2212.03238)","url":"https://arxiv.org/abs/2212.03238"},{"title":"Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (IROS 2018)","url":"https://dspace.mit.edu/handle/1721.1/138000"}],"as_of":"","related_ids":["gait","gait-planning","gait-cycle-and-duty-factor","footstep-planning","convex-mpc","central-pattern-generator"],"name":"Gait Planning","alt":"步态规划","abbr":"","aliases":["Gait Scheduling","Gait Generation"],"one_liner":"Deciding when each of a legged robot's feet touches down and lifts off, and the rhythm of its steps.","explanation":"Gait planning decides the contact timing of a legged robot's legs: within one gait cycle, when each leg touches down (stance phase) and when it lifts off (swing phase), what fraction of the cycle each leg spends on the ground (duty factor), and the phase offsets and stepping frequency between legs. Different combinations correspond to different gaits: for a quadruped, diagonal pairs together is trotting, same-side pairs together is pacing, front and rear pairs together is bounding, and all four legs lifting and landing together is pronking. In model-based control, a gait scheduler first lays out this contact timetable, MPC then allocates ground-reaction force to the supporting legs accordingly, and footstep planning decides where the swinging legs should land. Reinforcement-learning locomotion controllers commonly use a phase-clock signal or reward terms to guide the gait, or let a gait emerge on its own. It complements footstep planning: one decides when to step, the other decides where.","example":"MIT Cheetah's open-source code divides one gait cycle into 10 segments: a trot is written as the four legs having phase offsets (0,5,5,0), each supporting for 5 segments — meaning diagonal legs move together, and each leg is on the ground half the time — while a pronk has all four offsets at zero. Walk These Ways (2022) turns these phase offsets into a policy input command, with (0.5,0,0) giving a trot and (0,0.5,0) giving a bound.","related":["Gait","Gait Planning","Gait Cycle and Duty Factor","Footstep Planning","Convex MPC","Central Pattern Generator"]},{"id":"gait-phase","category":"control","sec":5,"tier":3,"sources":[{"title":"Siekmann et al., Sim-to-Real Learning of All Common Bipedal Gaits via Periodic Reward Composition (arXiv 2011.01387, ICRA 2021)","url":"https://arxiv.org/abs/2011.01387"},{"title":"unitree_rl_gym: legged_gym/envs/g1/g1_env.py","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/g1/g1_env.py"}],"as_of":"","related_ids":["gait","stance-phase-swing-phase","gait-cycle-and-duty-factor","gait-planning","trot-gait","gait-phase-variable-clock-input"],"name":"Gait Phase (Phase Clock)","alt":"步态相位（相位时钟）","abbr":"","aliases":["Gait Clock"],"one_liner":"A number that cycles from 0 to 1, indicating where the robot currently is in its gait cycle.","explanation":"Legged walking is periodic motion, with each leg alternating between stance phase (on the ground, bearing weight) and swing phase (lifted, swinging forward). Gait phase φ normalizes one cycle to [0,1): φ = (t mod T)/T, where t is time and T the gait period. Giving each leg its own phase offset defines a gait — for example, left and right legs offset by 0.5 gives alternating stepping in a biped, and diagonal legs sharing the same phase gives a quadruped trot. In model-based control it's produced by the gait scheduler and determines when contact switches; in reinforcement learning, Siekmann et al. (ICRA 2021) used it to build a periodic reward — penalizing foot velocity during stance and foot force during swing — learning walking, running, and hopping gaits on the Cassie biped with no reference motion at all.","example":"unitree_rl_gym's G1 walking environment uses a 0.8-second period with a 0.5 phase offset between legs; a leg with phase under 0.55 is treated as expected to be in stance, and the policy is rewarded when the actual contact state matches.","related":["Gait","Stance Phase / Swing Phase","Gait Cycle and Duty Factor","Gait Planning","Trot Gait","Gait Phase Variable / Clock Input (sin/cos Phase Encoding)"]},{"id":"central-pattern-generator","category":"control","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Central pattern generator","url":"https://en.wikipedia.org/wiki/Central_pattern_generator"},{"title":"Ijspeert A. J. Central pattern generators for locomotion control in animals and robots: a review. Neural Networks, 2008","url":"https://europepmc.org/article/MED/18555958"},{"title":"Bellegarda, Ijspeert. CPG-RL: Learning Central Pattern Generators for Quadruped Locomotion (arXiv:2211.00458)","url":"https://arxiv.org/abs/2211.00458"}],"as_of":"","related_ids":["gait","gait-phase","legged-locomotion","bio-inspired-robot","rl-based-locomotion-control","limit-cycle"],"name":"Central Pattern Generator","alt":"中枢模式发生器","abbr":"CPG","aliases":["CPG","CPG Oscillator"],"one_liner":"A neural circuit or oscillator that produces rhythmic output on its own, with no rhythmic input, used to drive a gait.","explanation":"The central pattern generator was originally a neuroscience concept: neural circuits in the spinal cord and elsewhere can generate rhythmic output on their own, with no rhythmic input, driving walking, swimming, and breathing. Graham Brown's 1911 experiments were the first to show that the spinal cord could independently generate a stepping pattern, with the lamprey and the cat as classic study subjects. Robotics borrows this idea: a set of coupled oscillators generates a periodic signal for each leg or joint, where amplitude sets stride length, frequency sets stepping speed, and the phase difference between oscillators switches the gait. The Ijspeert team's salamander robot, published in Science in 2007, used a spinal CPG model to switch between swimming and walking. CPGs give stable rhythms with few parameters, but handling complex terrain requires added sensory feedback. The ‘gait phase clock’ commonly used in reinforcement-learning locomotion is similar in spirit, and work such as CPG-RL instead has a neural network modulate the oscillators' parameters.","example":"EPFL's CPG-RL (2022) gives each quadruped leg its own oscillator, with a reinforcement-learning policy only outputting each oscillator's amplitude and frequency; deployed on a Unitree A1, it could carry an extra 13.75 kg load — 115% of the robot's own weight.","related":["Gait","Gait Phase (Phase Clock)","Legged Locomotion","Bio-inspired Robot","RL-based Locomotion Control","Limit Cycle"]},{"id":"footstep-planning","category":"control","sec":5,"tier":2,"sources":[{"title":"Footstep Planning on Uneven Terrain with Mixed-Integer Convex Optimization (Deits & Tedrake, 2014)","url":"https://groups.csail.mit.edu/robotics-center/public_papers/Deits14a.pdf"}],"as_of":"","related_ids":["gait-planning","swing-foot-trajectory-planning","raibert-heuristic","perceptive-locomotion","elevation-map","multi-contact-planning"],"name":"Footstep Planning","alt":"落足点规划","abbr":"","aliases":["Foothold Planning"],"one_liner":"Computing where and in what orientation a legged robot should plant each of its next footsteps.","explanation":"Footstep planning finds a sequence of safe foot placements (positions and orientations) for a biped or quadruped, taking it from its current position to a goal, while satisfying constraints such as keeping consecutive steps within the leg's reach and avoiding obstacles or gaps. It's a simplified version of full contact motion planning: it handles the whole-body dynamics only loosely, deciding just where the feet go and leaving exactly how the body moves to a controller further downstream. Approaches fall into roughly two categories. Discrete search builds a set of candidate steps and searches over them on a tree using something like A*. Continuous optimization is the other route — MIT's Deits and Tedrake, for example, decomposed the reachable footholds into a set of convex regions in 2014 and used mixed-integer convex optimization to find the globally optimal footstep sequence. Reinforcement-learning locomotion controllers, by contrast, usually decide footholds implicitly.","example":"Deits and Tedrake's planner guided an Atlas humanoid across a row of stepping stones; removing one stone made it automatically switch to a longer detour route. Short sequences of a few steps solved in under 1 second, while sequences of 10–30 steps took tens of seconds to a few minutes on a laptop.","related":["Gait Planning","Swing Foot Trajectory Planning","Raibert Heuristic","Perceptive Locomotion","Elevation Map","Multi-contact Planning"]},{"id":"raibert-heuristic","category":"control","sec":5,"tier":3,"sources":[{"title":"MIT Leg Laboratory: 3D One-Leg Hopper (1983–1984)","url":"http://www.ai.mit.edu/projects/leglab/robots/3D_hopper/3D_hopper.html"},{"title":"mit-biomimetics/Cheetah-Software: ConvexMPCLocomotion.cpp（footstep placement）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/user/MIT_Controller/Controllers/convexMPC/ConvexMPCLocomotion.cpp"},{"title":"Di Carlo et al., Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (2018)","url":"https://dspace.mit.edu/handle/1721.1/138000"}],"as_of":"","related_ids":["footstep-planning","swing-foot-trajectory-planning","convex-mpc","capture-point","spring-loaded-inverted-pendulum","quadruped-robot"],"name":"Raibert Heuristic","alt":"Raibert 启发式","abbr":"","aliases":["Raibert Foot Placement","Raibert's Foot Placement Formula"],"one_liner":"A foot-placement rule that shifts the landing spot forward or back based on body speed to regulate a legged robot's forward motion.","explanation":"The Raibert heuristic comes from Marc Raibert's work on single-leg hopping robots at Carnegie Mellon University in the early 1980s (his lab later moved to MIT and became the Leg Lab). He split hopping control into three independent parts — hop height, body attitude, and forward speed — with forward speed regulated through foot placement. The landing spot, relative to the hip, is x_f = ẋ·T_s/2 + k·(ẋ − ẋ_d), where ẋ is the current forward speed, T_s is the stance-phase duration, ẋ_d is the desired speed, and k is a feedback gain. The first term is the ‘neutral point’: landing there lets the body pass symmetrically fore and aft over the foot during stance, leaving speed roughly unchanged. The second term is a correction — running faster than desired shifts the foot forward to decelerate, and running slower shifts it back to accelerate. The heuristic needs no full dynamics model, and variants of it are still used today by quadruped controllers such as MIT Cheetah to plan where the swing leg should land, leaving the support-force computation to MPC or whole-body control.","example":"MIT's open-source Cheetah-Software computes swing-leg placement in its convex-MPC gait controller as 'velocity × stance duration × 0.5 + 0.03 × (actual velocity − desired velocity)', plus a turning correction term — a direct variant of the Raibert formula.","related":["Footstep Planning","Swing Foot Trajectory Planning","Convex MPC","Capture Point","Spring-Loaded Inverted Pendulum","Quadruped Robot"]},{"id":"swing-foot-trajectory-planning","category":"control","sec":5,"tier":3,"sources":[{"title":"MIT Cheetah-Software：FootSwingTrajectory.cpp（Bezier 摆动轨迹）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/common/src/Controllers/FootSwingTrajectory.cpp"},{"title":"MIT Cheetah-Software：ConvexMPCLocomotion.cpp（落足点与抬脚高度）","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/user/MIT_Controller/Controllers/convexMPC/ConvexMPCLocomotion.cpp"},{"title":"Humanoid-Gym humanoid_env.py（_reward_feet_clearance 抬脚高度奖励）","url":"https://github.com/roboterax/humanoid-gym/blob/main/humanoid/envs/custom/humanoid_env.py"}],"as_of":"","related_ids":["footstep-planning","raibert-heuristic","stance-phase-swing-phase","gait-planning","bezier-curve-trajectory","whole-body-control"],"name":"Swing Foot Trajectory Planning","alt":"足端轨迹规划","abbr":"","aliases":["Swing Leg Trajectory Planning"],"one_liner":"Planning the path a legged robot's airborne foot follows from lift-off to landing during the swing phase of a gait.","explanation":"Legged robots alternate each leg between stance phase, when the foot is on the ground bearing weight, and swing phase, when the foot is in the air moving toward its next landing spot. Swing foot trajectory planning handles the swing phase: given the lift-off position, the foothold (usually computed by foothold planning or the Raibert heuristic based on body speed), and the swing duration, it generates the foot's position, velocity, and acceleration at every point in between. The requirements are that the foot lifts high enough to clear the ground and step over obstacles like stairs, starts and ends smoothly, and slows to near-zero velocity just before landing to reduce impact — and ideally the landing point can still be updated mid-swing if a new speed command arrives. Cycloids, splines, or Bézier curves are common choices, often designed separately for the horizontal and vertical directions. The resulting foot trajectory is then converted into joint commands through inverse kinematics, Cartesian foot PD control, or whole-body control. Reinforcement-learning-based locomotion controllers generally don't plan this curve explicitly, shaping it indirectly instead through reward terms like foot clearance height and time in the air.","example":"MIT Cheetah's open-source control code connects lift-off to landing with a cubic Bézier curve horizontally, and two Bézier segments (rise and fall) vertically, with lift height set to 6 cm in its convex-MPC gait. The foothold is computed from the hip position plus a velocity-dependent correction, then tracked with Cartesian foot PD control.","related":["Footstep Planning","Raibert Heuristic","Stance Phase / Swing Phase","Gait Planning","Bézier Curve Trajectory","Whole-Body Control"]},{"id":"balance-control","category":"control","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Zero moment point","url":"https://en.wikipedia.org/wiki/Zero_moment_point"},{"title":"Pratt et al., Capture Point: A Step toward Humanoid Push Recovery (Humanoids 2006)","url":"https://doi.org/10.1109/ICHR.2006.321385"}],"as_of":"","related_ids":["zero-moment-point","capture-point","push-recovery","support-polygon","linear-inverted-pendulum-model","ankle-hip-and-stepping-strategies"],"name":"Balance Control","alt":"平衡控制","abbr":"","aliases":["Balance Maintenance","Postural Balance Control"],"one_liner":"Keeps a legged or humanoid robot from falling over while standing, walking, or being pushed.","explanation":"Balance control is the most fundamental capability of a legged robot: keeping the body from tipping over under gravity and ground-reaction forces. Traditional methods rely on simplified models: the zero moment point (ZMP, introduced by Vukobratović in 1968) is the point where the ground-reaction force produces no horizontal moment, and as long as it stays inside the support polygon (the region enclosed by the feet on the ground), the robot won't tip over the edge of its feet — early humanoids such as Honda's ASIMO planned their gaits this way; Pratt and colleagues introduced the capture point in 2006, answering where the foot should land after a push in order to stop. In response to a disturbance, robots can recover, roughly in order of increasing push strength, with ankle adjustments, hip swinging, or stepping. Today's humanoids and quadrupeds mostly use reinforcement learning: the robot is randomly pushed in simulation and the policy learns on its own to stay upright and to step to recover, then transfers to the real machine. In whole-body control, balance is usually the highest-priority task.","example":"A standing humanoid gets pushed from the side by a person: if the push is light, the ankles alone pull the center of mass back; if it's too strong for the ankles to handle, the robot steps sideways, planting its foot near the capture point to regain balance.","related":["Zero Moment Point","Capture Point","Push Recovery","Support Polygon","Linear Inverted Pendulum Model (LIPM)","Ankle, Hip and Stepping Strategies"]},{"id":"push-recovery","category":"control","sec":5,"tier":2,"sources":[{"title":"Pratt et al.: Capture Point: A Step toward Humanoid Push Recovery (Humanoids 2006)","url":"https://doi.org/10.1109/ICHR.2006.321385"},{"title":"Stéphane Caron: Capture point","url":"https://scaron.info/robotics/capture-point.html"},{"title":"legged_gym: legged_robot_config.py（push_robots / push_interval_s / max_push_vel_xy）","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["capture-point","ankle-hip-and-stepping-strategies","balance-control","zero-moment-point","domain-randomization","fall-recovery"],"name":"Push Recovery","alt":"推恢复","abbr":"","aliases":["Disturbance Recovery"],"one_liner":"A robot's ability to adjust its posture or take a step after being pushed or bumped, regaining balance without falling.","explanation":"Push recovery refers to a legged robot — especially bipeds and humanoids — regaining balance after being pushed or bumped by an external force. The classic approach grades the response by disturbance size: a small push is handled by ankle torque shifting the center of pressure (the ‘ankle strategy’), a somewhat larger one by bending the hip and swinging the upper body to generate angular momentum (the ‘hip strategy’), and a large one forces a step (the ‘stepping strategy’). In 2006, Pratt et al. introduced the capture point: the point on the ground where, if the robot steps exactly there, it can come to a complete stop, giving a principled target for where to step. Today's reinforcement-learning locomotion controllers instead learn this skill by randomly shoving the robot in simulation, letting the policy learn to resist pushes on its own — a form of domain randomization. This differs from fall recovery, which is about getting back up after already having fallen.","example":"The open-source framework legged_gym by default randomizes the robot torso's horizontal velocity up to 1 m/s every 15 seconds during training, simulating a sudden shove, so the policy learns to take a stabilizing step after being pushed.","related":["Capture Point","Ankle, Hip and Stepping Strategies","Balance Control","Zero Moment Point","Domain Randomization","Fall Recovery"]},{"id":"ankle-hip-and-stepping-strategies","category":"control","sec":5,"tier":3,"sources":[{"title":"Horak F. B., Nashner L. M. Central programming of postural movements: adaptation to altered support-surface configurations. J Neurophysiol, 1986","url":"https://europepmc.org/article/MED/3734861"},{"title":"Stephens B. Humanoid push recovery. IEEE-RAS Humanoids, 2007","url":"https://doi.org/10.1109/ICHR.2007.4813931"},{"title":"Push Recovery of a Position-Controlled Humanoid Robot Based on Capture Point Feedback Control (arXiv:1710.10598)","url":"https://arxiv.org/abs/1710.10598"}],"as_of":"","related_ids":["push-recovery","balance-control","capture-point","center-of-pressure","centroidal-moment-pivot","support-polygon"],"name":"Ankle, Hip and Stepping Strategies","alt":"踝策略/髋策略/跨步策略","abbr":"","aliases":["Ankle Strategy","Hip Strategy","Stepping Strategy"],"one_liner":"A three-tier response to being pushed while balancing: move the ankle, swing the torso, or take a step.","explanation":"These concepts come from research on human postural control. Horak and Nashner (1986) had subjects stand on a support surface that suddenly translated, and found that when standing normally, people mostly rotate the body about the ankle to recover balance — the ‘ankle strategy’; standing on a surface narrower than the foot, they switch to flexing mainly at the hip — the ‘hip strategy’; and with a large enough disturbance, people step to rebuild their base of support — the ‘stepping strategy.’ Humanoid push-recovery controllers borrow this same tiered scheme: the ankle strategy corresponds to adjusting the center of pressure (CoP) within the footprint; the hip strategy relies on rapidly rotating the upper body to generate angular momentum, equivalent to adjusting the centroidal moment pivot (CMP); and once the capture point (the ground point the robot must step to in order to stop completely) falls outside the support polygon, a step becomes unavoidable. Stephens (2007) derived analytic boundaries, based on these three strategies, for judging whether not stepping will necessarily lead to a fall.","example":"A humanoid robot given a light push while standing can recover with ankle torque alone; a bigger push has it swing its torso or arms forward and back to absorb the impulse; a still bigger one forces a step in the direction of the push.","related":["Push Recovery","Balance Control","Capture Point","Center of Pressure (CoP)","Centroidal Moment Pivot","Support Polygon"]},{"id":"double-support-phase","category":"control","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Gait (human)","url":"https://en.wikipedia.org/wiki/Gait_(human)"},{"title":"Phase-based NMPC for Humanoid Walking Stabilization with Single and Double Support Time Adjustments (arXiv:2506.03856)","url":"https://arxiv.org/abs/2506.03856"},{"title":"Hybrid Zero Dynamics Control for Bipedal Walking with a Non-Instantaneous Double Support Phase (arXiv:2303.05165)","url":"https://arxiv.org/abs/2303.05165"}],"as_of":"","related_ids":["gait","support-polygon","zero-moment-point","stance-phase-swing-phase","flight-phase","gait-cycle-and-duty-factor"],"name":"Double Support Phase","alt":"双支撑期","abbr":"DSP","aliases":["DSP","Double Stance","Double Limb Support"],"one_liner":"The part of walking when both feet are on the ground at once, as the body's support shifts from the rear foot to the front foot.","explanation":"The double support phase is the part of bipedal walking when both feet are on the ground simultaneously, occurring between when the front foot has just landed and the rear foot has yet to leave; it happens twice per gait cycle, with the rest of the time spent in single support, when only one foot is down. Based on Wikipedia's figures for human walking — roughly 60% stance phase and 40% swing phase — the two double-support periods together make up around 20% of the cycle; this period shortens as walking speeds up, and disappears entirely in running, replaced by a flight phase where both feet are off the ground. For a humanoid robot, this period is the most stable: the support polygon covers both feet, the zero moment point (ZMP) must shift from the rear foot to the front foot during it, the two legs and the ground form a closed kinematic chain with more actuation than degrees of freedom (requiring a decision on how much force each foot contributes), and landing impact and contact switching happen right around it. Many simplified models treat double support as instantaneous, while more detailed gait planners explicitly optimize its duration.","example":"A phase-based nonlinear MPC proposed by a Seoul National University team in 2025 jointly optimizes ZMP regulation, foothold placement, single-support duration, and double-support duration in one problem, and forbids updating the foothold during double support, improving a humanoid's walking balance against external pushes.","related":["Gait","Support Polygon","Zero Moment Point","Stance Phase / Swing Phase","Flight Phase (Aerial Phase)","Gait Cycle and Duty Factor"]},{"id":"zmp-preview-control","category":"control","sec":5,"tier":3,"sources":[{"title":"Kajita et al.: Biped walking pattern generation by using preview control of zero-moment point (ICRA 2003)","url":"https://doi.org/10.1109/robot.2003.1241826"},{"title":"Stéphane Caron: Linear inverted pendulum model","url":"https://scaron.info/robotics/linear-inverted-pendulum-model.html"}],"as_of":"","related_ids":["zero-moment-point","cart-table-model","linear-inverted-pendulum-model","gait-planning","model-predictive-control","support-polygon"],"name":"ZMP Preview Control","alt":"ZMP 预观控制","abbr":"","aliases":["Preview Control of ZMP","Kajita Preview Control"],"one_liner":"Generating stable biped gaits by looking ahead at a planned future ZMP trajectory and moving the center of mass to match it in advance.","explanation":"ZMP preview control was proposed by Shuuji Kajita and colleagues at Japan's National Institute of Advanced Industrial Science and Technology (AIST) at ICRA 2003, and is a classic method for generating humanoid walking patterns. The zero moment point (ZMP) is the point where the horizontal component of the ground reaction moment is zero; keeping it inside the support polygon (the foot's contact area) keeps the foot from tipping. The method first lays out a reference ZMP trajectory based on the footholds, turning the problem into: how should the center of mass (CoM) move so the actual ZMP tracks that reference? The paper uses a ‘cart-table model,’ where ZMP and CoM satisfy p = x − (z_c/g)·ẍ — p is the ZMP position, x is the CoM's horizontal position, z_c is a constant CoM height, and g is gravitational acceleration. Because CoM acceleration can only be produced by a deviation between the ZMP and the CoM, the CoM must start moving ahead of time — looking only at the current reference is too late, hence ‘preview.’ The controller takes CoM jerk as its input, minimizes ZMP tracking error, and feeds forward a weighted sum of future reference ZMP values within a preview window; the gains can be computed offline, so the online computation stays light. Later ZMP-based MPC methods add constraints on top of this foundation.","example":"A biped walks four steps: a staircase-shaped ZMP reference is laid out based on the footholds, jumping to the center of the new support foot at each step. The preview controller sees the next step coming and smoothly shifts the CoM toward the next support foot ahead of time, keeping the actual ZMP inside the foot's support area throughout.","related":["Zero Moment Point","Cart-Table Model","Linear Inverted Pendulum Model (LIPM)","Gait Planning","Model Predictive Control","Support Polygon"]},{"id":"human-like-gait","category":"control","sec":5,"tier":3,"sources":[{"title":"Ogura et al.: Human-like Walking with Knee Stretched, Heel-contact and Toe-off Motion by a Humanoid Robot (IROS 2006)","url":"https://gaoyichao.com/Xiaotu/robot_cases/papers/2006%20-%20Human-like%20walking%20with%20knee%20stretched,%20heel-contact%20and%20toe-off%20motion%20by%20a%20humanoid%20robot%20-%20Ogura%20et%20al.pdf"},{"title":"Figure: Natural Humanoid Walk Using Reinforcement Learning (2025-03)","url":"https://www.figure.ai/news/reinforcement-learning-walking"},{"title":"科创板日报：深圳人形机器人行走视频走红 拟人步态震惊英伟达科学家（2025-01）","url":"https://www.cls.cn/detail/1915921"}],"as_of":"2025-03","related_ids":["straight-knee-walking","gait","zero-moment-point","bipedal-locomotion","adversarial-motion-priors","stance-phase-swing-phase"],"name":"Human-like Gait (Straight-knee, Heel-to-toe Walking)","alt":"拟人步态（直膝行走 / 足跟-足尖行走）","abbr":"","aliases":["Human-like Walking","Heel-Strike / Toe-Off Walking"],"one_liner":"Having a humanoid robot walk the way people do — knee straight, heel touching down first, then pushing off with the toe.","explanation":"‘Human-like gait’ is common usage in the humanoid robotics industry for a walking style close to a person's: the stance leg's knee stays essentially straight, the heel lands first, the weight rolls across the foot, and the toe pushes off at the end, with the arms swinging opposite the legs. Early zero-moment-point (ZMP)-based humanoids mostly walked with bent knees, flat feet, and constant hip height, partly because a straight knee sits close to a kinematic singularity, making hip-height control difficult. Waseda University's WABIAN-2R, in 2006, was an early success at straight-knee, heel-to-toe walking, using hip motion to avoid the singularity and adding a passive toe joint to the foot. In recent years this has mostly been achieved with reinforcement learning: a reward term imitating a human walking reference trajectory is added, balanced against velocity tracking and energy cost.","example":"In March 2025, Figure released a reinforcement-learning walking controller for Figure 02: rewarded in GPU simulation for imitating a human walking reference trajectory, it learned heel strike, toe-off, and arm swing synchronized with the legs, and transferred zero-shot to the real robot. China's EngineAI launched its SE01 in October 2024, reportedly marketed with straight-knee gait as a key selling point.","related":["Straight-Knee Walking","Gait","Zero Moment Point","Bipedal Locomotion","Adversarial Motion Priors","Stance Phase / Swing Phase"]},{"id":"hybrid-zero-dynamics","category":"control","sec":5,"tier":3,"sources":[{"title":"Grizzle & Chevallereau: Virtual Constraints and Hybrid Zero Dynamics for Realizing Underactuated Bipedal Locomotion (arXiv:1706.01127)","url":"https://arxiv.org/abs/1706.01127"},{"title":"Gong et al.: Feedback Control of a Cassie Bipedal Robot: Walking, Standing, and Riding a Segway (arXiv:1809.07279)","url":"https://arxiv.org/abs/1809.07279"}],"as_of":"","related_ids":["underactuation","bipedal-locomotion","limit-cycle","zero-moment-point","agility-robotics-cassie","gait-phase"],"name":"Hybrid Zero Dynamics","alt":"混合零动态","abbr":"HZD","aliases":["HZD","Virtual Constraints","HZD Control"],"one_liner":"Using ‘virtual constraints’ to compress bipedal walking into a low-dimensional system, then designing and proving a gait stable.","explanation":"Hybrid zero dynamics was proposed by Westervelt, Grizzle, and Koditschek in 2003 for underactuated bipedal robots — point-foot robots, for instance, can't apply ankle torque to the ground at all. Walking is a ‘hybrid’ system: continuous dynamics while a leg swings, and a discrete jump at footstrike, with impact and a leg switch. The method chooses a phase variable θ that advances monotonically with the gait, and uses feedback to make every actuated joint angle q_a track a specified function h(θ), i.e., enforce y = q_a − h(θ) = 0 — these are the ‘virtual constraints.’ The robot is thereby compressed onto a low-dimensional surface, and whatever underactuated part remains is the zero dynamics. A Poincaré map is then used to check whether the state returns to the same periodic orbit (limit cycle) after each step's impact, which is how gait stability is proven. It doesn't require the foot to stay flat on the ground — the main way it differs from ZMP-based methods.","example":"University of Michigan's Grizzle group used virtual constraints plus a gait library to control Cassie in 2018: about six weeks after receiving the robot, they had it walking on sidewalks, grass, snow, and sand, and even balancing while standing on a Segway.","related":["Underactuation","Bipedal Locomotion","Limit Cycle","Zero Moment Point","Agility Robotics Cassie","Gait Phase (Phase Clock)"]},{"id":"virtual-model-control","category":"control","sec":5,"tier":3,"sources":[{"title":"Pratt, Chew, Torres, Dilworth, Pratt: Virtual Model Control: An Intuitive Approach for Bipedal Locomotion (IJRR 2001)","url":"https://doi.org/10.1177/02783640122067309"}],"as_of":"","related_ids":["jacobian-transpose-method","impedance-control","torque-control","bipedal-locomotion","whole-body-control","raibert-heuristic"],"name":"Virtual Model Control","alt":"虚拟模型控制","abbr":"VMC","aliases":["VMC","Virtual Spring-Damper Control"],"one_liner":"Imagining virtual springs and dampers attached to the robot, then converting the forces they'd produce into joint torques.","explanation":"Virtual model control was proposed by Jerry Pratt and colleagues at the MIT Leg Lab, published in IJRR in 2001, and first used on planar biped walking. The method imagines springs, dampers, and other ‘virtual components’ hung between the robot's body and some reference point — for instance, a spring pulling the torso up toward a target height. It computes a virtual force F from the spring-damper equation, then converts it into support-leg joint torques via the Jacobian transpose: τ = Jᵀ F, where J is the matrix mapping joint velocity to body velocity. It needs no full dynamics model and no inverse kinematics, and its parameters have an intuitive physical meaning — tuning it feels like adjusting how stiff a spring is. The robot in the original paper relied only on foot contact detection and walked over slopes and uneven ground without knowing the grade in advance. It ignores the legs' own inertia, making it a quasi-static approximation whose error grows with more aggressive motion. It's since also commonly used for leg force control on quadrupeds and wheel-legged robots.","example":"A biped in stance phase: a vertical virtual spring-damper on the torso maintains body height, a torsional spring keeps the torso upright, and a horizontal damper controls forward speed. The three virtual forces are converted via τ = Jᵀ F into hip, knee, and ankle torques.","related":["Jacobian Transpose Method","Impedance Control","Torque Control","Bipedal Locomotion","Whole-Body Control","Raibert Heuristic"]},{"id":"convex-mpc","category":"control","sec":5,"tier":3,"sources":[{"title":"Di Carlo et al., Dynamic Locomotion in the MIT Cheetah 3 Through Convex Model-Predictive Control (IROS 2018)","url":"https://dspace.mit.edu/handle/1721.1/138000"},{"title":"Kim et al., Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"}],"as_of":"","related_ids":["model-predictive-control","single-rigid-body-dynamics-model","nonlinear-model-predictive-control","contact-force-optimization","gait-planning","whole-body-control"],"name":"Convex MPC","alt":"凸 MPC","abbr":"","aliases":["Single Rigid-Body MPC"],"one_liner":"Simplifying a legged robot to a single rigid body so MPC becomes a quadratic program that solves quickly to the global optimum.","explanation":"Convex MPC generally refers to the quadruped control method published by MIT Biomimetic Robotics Lab's Di Carlo, Wensing, Kim, et al. at IROS 2018. Model predictive control needs to solve, every cycle, an optimization subject to dynamics constraints; a full-body nonlinear model is too slow. They instead treat the robot as a single rigid body pushed by ground reaction forces (ignoring leg mass) and apply three approximations: small roll and pitch angles, the state staying close to the commanded trajectory (linearizing around commanded yaw and planned foothold positions), and dropping nonlinear terms involving angular velocity. With each foot's contact timing fixed in advance by the gait scheduler, the only decision variables left are the ground reaction forces, and with friction-cone constraints added, the problem becomes a convex quadratic program: it's guaranteed to find the global optimum, and it solves fast. It outputs the reaction forces over a future window, which joint torque control or whole-body control then executes. The approximation breaks down for large rolling motions, where nonlinear MPC is used instead.","example":"On Cheetah 3, a force plan with a prediction horizon up to 0.5 seconds solves in under 1 millisecond, running at 20–30 Hz; the same set of gains produced standing, trotting, flying trot, bounding, pronking, pacing, a three-leg gait, and 3D galloping, reaching forward speeds up to 3 m/s. After porting to Mini Cheetah in 2019, paired with a 500 Hz whole-body impulse controller (WBIC), it reached 3.7 m/s.","related":["Model Predictive Control","Single Rigid Body Dynamics Model","Nonlinear Model Predictive Control","Contact Force Optimization (Force Distribution)","Gait Planning","Whole-Body Control"]},{"id":"contact-force-optimization","category":"control","sec":5,"tier":3,"sources":[{"title":"MIT Cheetah-Software: BalanceController.cpp（接触力 QP，qpOASES）","url":"https://raw.githubusercontent.com/mit-biomimetics/Cheetah-Software/master/user/MIT_Controller/Controllers/BalanceController/BalanceController.cpp"},{"title":"Kim et al., Highly Dynamic Quadruped Locomotion via Whole-Body Impulse Control and Model Predictive Control (arXiv:1909.06586)","url":"https://arxiv.org/abs/1909.06586"}],"as_of":"","related_ids":["quadratic-programming","friction-cone","friction-pyramid","ground-reaction-force","convex-mpc","whole-body-control"],"name":"Contact Force Optimization (Force Distribution)","alt":"接触力优化","abbr":"","aliases":["Force Distribution","Ground Reaction Force Distribution"],"one_liner":"Given the total force and torque the body needs, solving how much force each foot or finger should contribute.","explanation":"Contact force optimization, also called force distribution, applies whenever a robot has force at multiple contact points at once — a quadruped standing on four feet, two hands holding a box, or a multi-fingered grasp. A higher-level controller first computes the total force and torque the body needs (together called a wrench), then solves an optimization problem to split it among the contact points. With four feet on the ground, for instance, each foot contributes 3 force components for 12 unknowns total, but the balance equations give only 6 constraints, so there are infinitely many solutions — an objective is added to pick one: stay close to the desired wrench, and keep the forces themselves small. Constraints are added too: tangential force must not exceed the friction coefficient times normal force (the friction cone, commonly linearized into a friction pyramid), normal force must stay within upper and lower bounds, and any foot in the air must have zero force. It's typically a quadratic program solved once per control cycle, with the resulting forces converted to joint torques via the foot Jacobian transpose. It only looks at the current instant; convex MPC extends the same problem across time, and whole-body control further folds in the full dynamics.","example":"The BalanceController in MIT Cheetah's open-source code: given desired body linear and angular acceleration, it solves for 12 force components across the 4 feet, subject to a linearized friction cone and per-foot upper/lower bounds on normal force (with the bounds pinned to 0 for any foot in the air), solving with qpOASES and warm-starting from the previous solution; the code comments cite Focchi et al.'s 2016 paper on steep-slope walking as the method's basis.","related":["Quadratic Programming","Friction Cone","Friction Pyramid","Ground Reaction Force (GRF)","Convex MPC","Whole-Body Control"]},{"id":"multi-contact-planning","category":"control","sec":5,"tier":3,"sources":[{"title":"Simultaneous Contact Sequence and Patch Planning for Dynamic Locomotion (arXiv 2508.12928)","url":"https://arxiv.org/abs/2508.12928"},{"title":"Online Multi-Contact Receding Horizon Planning via Value Function Approximation (arXiv 2306.04732)","url":"https://arxiv.org/abs/2306.04732"}],"as_of":"","related_ids":["footstep-planning","gait-planning","centroidal-dynamics","friction-cone","contact-implicit-trajectory-optimization","monte-carlo-tree-search"],"name":"Multi-contact Planning","alt":"多接触规划","abbr":"","aliases":["Contact Planning"],"one_liner":"Planning which body part — foot, hand, knee — contacts the environment where and in what order, to make use of that contact for leverage.","explanation":"Multi-contact planning is about deciding, for a legged or humanoid robot, which body part touches the environment, when, and where, in a way that's mechanically feasible for the whole motion. Ordinary bipedal walking alternates just two feet; climbing a steep slope, squeezing through a tight gap, getting over an obstacle, or catching a fall may call for a hand against a wall or a knee on the ground too, with neither the number nor the order of contact points fixed in advance. The difficulty is that it mixes discrete decisions (which limb, what order, which surface to land on) with continuous ones (the body trajectory, keeping contact forces inside their friction cones), making it a mixed discrete-continuous optimization problem. Traditional approaches split it into layers: plan the contact sequence and positions first, then generate the motion with centroidal dynamics or whole-body trajectory optimization; recent work uses Monte Carlo tree search or a learned value function to combine the two layers, run in a receding-horizon, online-replanning fashion. It is in the same family as footstep planning, but without being restricted to a periodic gait.","example":"A Talos humanoid walking up a slope too steep to balance statically on: Wang et al.'s 2023 method uses a learned value function to approximate long-term consequences, planning the next contact and body motion online in a receding horizon.","related":["Footstep Planning","Gait Planning","Centroidal Dynamics","Friction Cone","Contact-Implicit Trajectory Optimization","Monte Carlo Tree Search"]},{"id":"whole-body-control","category":"control","sec":5,"tier":1,"sources":[{"title":"Sentis & Khatib, Synthesis of Whole-Body Behaviors through Hierarchical Control of Behavioral Primitives (IJHR 2005)","url":"https://doi.org/10.1142/S0219843605000594"},{"title":"HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots (arXiv 2410.21229)","url":"https://arxiv.org/abs/2410.21229"},{"title":"A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots (arXiv 2506.20487)","url":"https://arxiv.org/abs/2506.20487"}],"as_of":"2025-03","related_ids":["learning-based-whole-body-control","task-prioritization","hierarchical-quadratic-programming","null-space-control","operational-space-control","hover"],"name":"Whole-Body Control","alt":"全身控制","abbr":"WBC","aliases":["WBC","Whole-Body Controller"],"one_liner":"Computes commands for all of a humanoid's or legged robot's joints together, so several tasks are accomplished at once.","explanation":"Whole-body control means computing commands for every joint of a floating-base robot (one whose body isn't fixed to the ground, such as a humanoid or quadruped) together, rather than controlling the legs and arms separately. The classic framework comes from a 2005 paper by Sentis and Khatib: things like the center of mass, hands, feet, joint limits, and contacts are written as “tasks” or constraints and given priorities — higher-priority tasks (such as staying within joint limits or not falling) must be satisfied, while lower-priority ones use whatever redundant freedom remains to do as well as possible; in practice this is usually formulated as a quadratic program (QP) or a hierarchy of QPs, solved every control cycle. It matters because reaching out with one hand shifts a humanoid's center of mass, so the legs and torso have to compensate at the same time. Learned whole-body control took off around 2024 — HOVER (ICRA 2025), for example, uses a single neural-network policy to unify navigation, mobile manipulation, and tabletop manipulation, among other control modes. A higher-level VLA or a teleoperator supplies the goal, and WBC turns it into whole-body joint commands.","example":"A humanoid bends down to pick up a box with one hand: WBC simultaneously satisfies “hand reaches the box,” “center-of-mass projection stays within the foot-support region,” and “joints stay within limits,” so it automatically bends at the waist and knees while swinging the other arm back for counterweight.","related":["Learning-Based Whole-Body Control","Task Prioritization","Hierarchical Quadratic Programming","Null-Space Control","Operational Space Control","HOVER"]},{"id":"null-space-control","category":"control","sec":5,"tier":3,"sources":[{"title":"franka_ros cartesian_impedance_example_controller.cpp（零空间力矩实现）","url":"https://raw.githubusercontent.com/frankaemika/franka_ros/develop/franka_example_controllers/src/cartesian_impedance_example_controller.cpp"},{"title":"StudyWolf: Robot control part 5 - Controlling in the null space","url":"https://studywolf.wordpress.com/2013/09/17/robot-control-5-controlling-in-the-null-space/"},{"title":"Implicit Null-space Manifold Generation for Redundant Robotic Systems (arXiv 2605.25770)","url":"https://arxiv.org/abs/2605.25770"}],"as_of":"","related_ids":["kinematic-redundancy","null-space","jacobian-pseudoinverse","operational-space-control","task-prioritization","hierarchical-quadratic-programming"],"name":"Null-Space Control","alt":"零空间控制","abbr":"","aliases":["Null-Space Projection","Redundancy Resolution"],"one_liner":"Using spare degrees of freedom to accomplish a secondary goal without disturbing the primary task.","explanation":"When a robot has more joints than the task strictly needs (a 7-DOF arm for a 6D end-effector pose, say), the same end-effector pose corresponds to infinitely many joint configurations — this is called kinematic redundancy. The Jacobian matrix J maps joint velocity q̇ to end-effector velocity; any joint motion satisfying J·q̇ = 0 forms J's null space — the joints move, but the end-effector doesn't, like an elbow tracing a circle while the hand stays put. Null-space control filters a secondary task's command through a projection matrix N = I − J⁺J (J⁺ the pseudoinverse of J) before adding it on top of the primary task, guaranteeing the secondary task never disturbs the primary one. Common secondary objectives include staying away from joint limits, avoiding singularities, keeping the elbow clear of obstacles, and holding a comfortable posture. At the torque level, Khatib's 1987 operational space control uses a ‘dynamically consistent’ pseudoinverse that accounts for the mass matrix — otherwise a secondary-task torque would still cause end-effector acceleration. Projecting several tasks layer by layer gives task priority and hierarchical quadratic programming.","example":"Franka's Cartesian impedance control example: the primary task tracks the end-effector toward a target pose, while a secondary ‘spring-damper pulling back toward the startup joint posture’ torque is projected through (I − Jᵀ(Jᵀ)⁺) and added on top — pushing the elbow away lets it return on its own, without disturbing end-effector tracking.","related":["Kinematic Redundancy","Null Space","Jacobian Pseudoinverse","Operational Space Control","Task Prioritization","Hierarchical Quadratic Programming"]},{"id":"task-prioritization","category":"control","sec":5,"tier":3,"sources":[{"title":"Nakamura, Hanafusa, Yoshikawa: Task-Priority Based Redundancy Control of Robot Manipulators (IJRR 1987)","url":"https://doi.org/10.1177/027836498700600201"},{"title":"Siciliano, Slotine: A general framework for managing multiple tasks in highly redundant robotic systems (ICAR 1991)","url":"https://doi.org/10.1109/icar.1991.240390"}],"as_of":"","related_ids":["null-space","null-space-control","kinematic-redundancy","jacobian-pseudoinverse","hierarchical-quadratic-programming","whole-body-control"],"name":"Task Prioritization","alt":"任务优先级","abbr":"","aliases":["Prioritized Task Control","Task-Priority Control","Multi-Task Priority Control"],"one_liner":"Ranking a robot's simultaneous goals by importance so lower-priority tasks can never disturb higher-priority ones.","explanation":"Redundant robots — those with more joints than a task strictly needs, such as 7-axis arms or humanoids — often have to satisfy several goals at once: reaching a target, staying balanced, avoiding joint limits, keeping a natural posture. Task prioritization ranks these goals into a hierarchy. Yoshihiko Nakamura and colleagues proposed Jacobian-pseudoinverse-based task-priority redundancy control in 1987, and Bruno Siciliano and Jean-Jacques Slotine generalized it to an arbitrary number of tasks in 1991. The core formula is q̇ = J₁⁺ẋ₁ + (I − J₁⁺J₁)q̇₀: J₁ is the primary task's Jacobian (mapping joint velocity to task velocity), J₁⁺ is its pseudoinverse, and the term (I − J₁⁺J₁) projects the secondary task's desired joint velocity q̇₀ into the null space of the primary task, guaranteeing it can't disturb the primary task's result. This differs from a weighted sum of tasks, where the tasks compromise with each other — strict prioritization instead guarantees the higher-priority task always wins. Modern whole-body controllers commonly implement this with hierarchical quadratic programming, which can also handle inequality constraints.","example":"A humanoid reaching for a cup: the first priority keeps its center of mass within the support region so it doesn't fall, the second priority moves the hand to the cup, and the third keeps the joints near a comfortable posture. When reaching for the cup conflicts with balance, the controller sacrifices some reaching accuracy rather than balance.","related":["Null Space","Null-Space Control","Kinematic Redundancy","Jacobian Pseudoinverse","Hierarchical Quadratic Programming","Whole-Body Control"]},{"id":"hierarchical-quadratic-programming","category":"control","sec":5,"tier":3,"sources":[{"title":"Escande, Mansard, Wieber, Hierarchical Quadratic Programming: Fast Online Humanoid-Robot Motion Generation (IJRR 2014)","url":"https://gepettoweb.laas.fr/uploads/Publications/2014_escande_ijrr.pdf"}],"as_of":"","related_ids":["task-prioritization","whole-body-control","quadratic-programming","null-space-control","operational-space-control","inverse-kinematics"],"name":"Hierarchical Quadratic Programming","alt":"分层二次规划","abbr":"HQP","aliases":["HQP","Hierarchical QP"],"one_liner":"Solving multiple control objectives in strict priority order, layer by layer, so lower priorities never interfere with higher ones.","explanation":"Robots often need to satisfy several potentially conflicting objectives at once: don't fall over, don't let feet slip, reach the target with the hand, stay within joint limits. A weighted QP multiplies each by a weight and sums them into one cost, which can only produce a compromise, and an important task can still get sacrificed. HQP instead orders them strictly: solve the highest-priority quadratic program first, treat its set of optimal solutions as a constraint, then solve the next layer within the remaining degrees of freedom (the null space) — so lower layers can never break a higher layer's result. It generalizes the classic null-space projection method, with the difference that every layer here can carry inequality constraints too (joint limits, friction cones, etc.). Escande, Mansard, and Wieber's 2014 IJRR paper gave a fast solver supporting both equality and inequality constraints at every layer, capable of generating whole-body motion for the HRP-2 humanoid at control-loop rates; it's commonly used for inverse kinematics or inverse dynamics solves within whole-body control.","example":"A humanoid reaching for a distant object: layer 1 keeps the support foot planted and the center of mass inside the support polygon; layer 2 gets the hand to the target; layer 3 orients the head toward the object and returns other joints to a default posture. If the reach isn't possible, it's the lower layers that get sacrificed, not balance.","related":["Task Prioritization","Whole-Body Control","Quadratic Programming","Null-Space Control","Operational Space Control","Inverse Kinematics (IK)"]},{"id":"rl-based-locomotion-control","category":"control","sec":5,"tier":2,"sources":[{"title":"Learning agile and dynamic motor skills for legged robots (Hwangbo et al., Science Robotics 2019)","url":"https://arxiv.org/abs/1901.08652"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (Rudin et al.)","url":"https://arxiv.org/abs/2109.11978"},{"title":"unitree_rl_gym：Go2 训练配置","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/go2/go2_config.py"}],"as_of":"2026-09","related_ids":["legged-locomotion","sim-to-real-transfer","domain-randomization","proximal-policy-optimization","stiffness-and-damping-gains","velocity-command-tracking"],"name":"RL-based Locomotion Control","alt":"强化学习运控","abbr":"","aliases":["RL Locomotion"],"one_liner":"Training a neural network with reinforcement learning in simulation to directly control a legged robot's walking.","explanation":"RL-based locomotion control means using reinforcement learning (maximizing reward through trial and error) to train a neural-network controller that lets quadrupeds, humanoids, and other legged robots walk, run, and get up after falling. A typical pipeline runs thousands of robots in parallel inside a GPU simulator such as Isaac Gym or Isaac Lab, training with PPO; the policy outputs joint target angles at roughly 50 Hz, which per-joint PD controllers convert into torques; domain randomization and actuator modeling then help bridge the sim-to-real gap for deployment on hardware. ETH's Hwangbo et al. validated this approach on the ANYmal quadruped in 2019, and Rudin et al. cut flat-ground walking training down to under 4 minutes on a single GPU in 2021. Compared to model-based methods such as MPC plus whole-body control, it doesn't require hand-written gaits or a precise model and tends to be more robust on complex terrain, but it depends heavily on reward design and extensive tuning.","example":"In Unitree's open-source unitree_rl_gym, training Go2 uses a 5-millisecond simulation step with the action updated every 4 steps — so the policy outputs 12 joint target angles every 20 milliseconds — tracked by PD control with Kp=20, Kd=0.5; the trained network is then exported and deployed to the real robot.","related":["Legged Locomotion","Sim-to-Real Transfer","Domain Randomization","Proximal Policy Optimization","Stiffness and Damping Gains","Velocity Command Tracking"]},{"id":"velocity-command-tracking","category":"control","sec":5,"tier":2,"sources":[{"title":"legged_gym: legged_robot.py（_reward_tracking_lin_vel / _resample_commands）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"legged_gym: legged_robot_config.py（commands / rewards）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"},{"title":"Walk These Ways (arXiv:2212.03238)","url":"https://arxiv.org/abs/2212.03238"}],"as_of":"","related_ids":["rl-based-locomotion-control","reward-function","cmd-vel-topic","legged-gym","walk-these-ways","terrain-curriculum"],"name":"Velocity Command Tracking","alt":"速度指令跟踪","abbr":"","aliases":["Velocity Tracking","Command Tracking"],"one_liner":"The standard training task where a legged robot walks according to a given forward, lateral, and turning speed.","explanation":"Velocity command tracking is the most basic task setup in reinforcement-learning locomotion for legged robots: at intervals, the robot is randomly given a target velocity — typically forward body speed vx, lateral speed vy, and yaw rate ωz about the vertical axis — and the policy outputs joint actions so the body's actual velocity matches the command as closely as possible. The reward is often written as exp(−‖v_command − v_actual‖²/σ), approaching 1 as the error shrinks, with σ controlling the tolerance. Once trained, the deployed policy takes velocity commands from a joystick or a higher-level navigation module, which only decides ‘which way, how fast,’ while the policy itself handles gait and balance. It's different from motor-level ‘velocity control,’ which governs a single joint's rotational speed.","example":"legged_gym's default configuration uses a 4-dimensional command (vx, vy, ωz, and heading), resampled every 10 seconds with linear velocity ranging ±1 m/s; the linear-velocity tracking reward has weight 1.0 and the angular-velocity reward 0.5, with σ = 0.25, and commanded lateral velocities under 0.2 m/s are zeroed out.","related":["RL-based Locomotion Control","Reward Function","cmd_vel Topic (geometry_msgs/Twist velocity command)","legged_gym","Walk These Ways","Terrain Curriculum"]},{"id":"gait-phase-variable-clock-input","category":"control","sec":5,"tier":3,"sources":[{"title":"unitree_rl_gym: legged_gym/envs/g1/g1_env.py","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/g1/g1_env.py"},{"title":"Siekmann et al., Periodic Reward Composition (arXiv 2011.01387)","url":"https://arxiv.org/abs/2011.01387"}],"as_of":"","related_ids":["gait-phase","rl-based-locomotion-control","walk-these-ways","observation","stance-phase-swing-phase","positional-encoding"],"name":"Gait Phase Variable / Clock Input (sin/cos Phase Encoding)","alt":"步态相位（相位时钟输入）","abbr":"","aliases":["Clock Input","Phase Observation","sin/cos Phase Encoding"],"one_liner":"Encoding gait phase as a pair of numbers, sin and cos, fed into the policy network so it knows the beat.","explanation":"This is how gait phase gets used concretely in learned locomotion control: when training a reinforcement-learning walking policy, in addition to joint angles, angular velocities, and velocity commands, sin(2πφ) and cos(2πφ) are also fed in as observations, where φ cycles through 0–1. φ isn't fed in directly because it jumps discontinuously from 0.99 back to 0, whereas sin/cos correspond to a point on a circle and stay continuous, letting the network see that these two moments are actually adjacent. With this external beat available, even a memoryless feedforward network can output a stable periodic gait, aligned with the stance/swing schedule used in a periodic reward. Walk These Ways (Margolis and Agrawal) feeds each of the four legs its own sinusoidal timing signal, and uses stepping frequency and inter-leg phase offset as commands, letting it switch online among trotting, pacing, bounding, and other gaits.","example":"unitree_rl_gym's G1 environment appends sin(2πφ) and cos(2πφ) to the end of the observation vector, alongside angular velocity, projected gravity, velocity commands, joint positions and velocities, and the previous action, all fed into the policy network together.","related":["Gait Phase (Phase Clock)","RL-based Locomotion Control","Walk These Ways","Observation","Stance Phase / Swing Phase","Positional Encoding"]},{"id":"learning-based-whole-body-control","category":"control","sec":5,"tier":2,"sources":[{"title":"HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots (arXiv 2410.21229)","url":"https://arxiv.org/abs/2410.21229"},{"title":"HOVER 全文（HTML 版，动作空间与 DAgger 蒸馏细节）","url":"https://arxiv.org/html/2410.21229"},{"title":"A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots (arXiv 2506.20487)","url":"https://arxiv.org/abs/2506.20487"}],"as_of":"2025-11","related_ids":["whole-body-control","hover","motion-tracking","teacher-student-distillation","rl-based-locomotion-control","braincerebellum-architecture"],"name":"Learning-Based Whole-Body Control","alt":"学习型全身控制","abbr":"","aliases":["Neural WBC","Neural Whole-Body Controller","RL Whole-Body Control"],"one_liner":"A neural network, trained with reinforcement learning in simulation, that coordinates a humanoid robot's entire body from one policy.","explanation":"Traditional whole-body control (WBC) solves a prioritized quadratic program from a dynamics model every control cycle to get full-body joint torques, which requires an accurate model and heavy tuning. Learning-based whole-body control instead trains a neural-network policy with reinforcement learning in simulation: it takes in the robot's own state plus a high-level command (walking speed, joint angles, or target positions for the hands and head, sometimes a snippet of human motion) and outputs full-body joint target positions, which per-joint PD controllers then execute; training commonly adds domain randomization to help the policy transfer to the real robot. The motion data used for training is usually retargeted from human motion-capture datasets such as AMASS. Notable examples include ExBody, OmniH2O, HOVER, and SONIC. It often serves as the ‘cerebellum’ in a brain-cerebellum split, receiving high-level commands from a VLA model or a teleoperator.","example":"HOVER (Tairan He et al., ICRA 2025) unified several command modes — root-velocity tracking, local joint-angle tracking, and keypoint-position tracking — into one policy on a 19-degree-of-freedom Unitree H1, using masking, and distilled it from a privileged teacher policy via DAgger; the unified policy beat specialized controllers on at least 7 of 12 metrics under each mode.","related":["Whole-Body Control","HOVER","Motion Tracking","Teacher-Student Distillation","RL-based Locomotion Control","Brain–Cerebellum Architecture"]},{"id":"decoupled-whole-body-control","category":"control","sec":5,"tier":3,"sources":[{"title":"NVlabs/GR00T-WholeBodyControl（Decoupled WBC: RL for lower body, IK for upper body）","url":"https://github.com/NVlabs/GR00T-WholeBodyControl"},{"title":"HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit (arXiv:2502.13013)","url":"https://arxiv.org/abs/2502.13013"}],"as_of":"2026-05","related_ids":["whole-body-control","learning-based-whole-body-control","inverse-kinematics","rl-based-locomotion-control","loco-manipulation","sonic"],"name":"Decoupled Whole-Body Control","alt":"解耦全身控制","abbr":"Decoupled WBC","aliases":["Decoupled WBC","Upper/Lower Body Split Control","Upper-Body IK + Lower-Body RL"],"one_liner":"A humanoid setup where the lower body uses reinforcement learning to walk and balance, and the upper body uses inverse kinematics to control the arms.","explanation":"This is an engineering-pragmatic approach to humanoid whole-body control that splits the robot into two separate controllers. The lower body runs a walking policy trained with reinforcement learning in simulation, taking commands like forward speed, turning, and body height, and handling walking, crouching, and balance — trained with randomized upper-body poses so it learns to treat arm motion as a disturbance it can absorb. The upper body is controlled directly by inverse kinematics (IK, computing joint angles from a desired hand pose) or by a teleoperation device's joint mapping. The benefit is a clean division of labor and accurate arm positioning, convenient for teleoperated data collection and for hooking up a VLA model; the cost is that the two halves aren't coordinated, so motions requiring the whole body to work together — bending down to reach the floor, leaning forward to extend reach, exerting full-body force to move something heavy — don't come out well. NVIDIA's GR00T N1.5 and N1.6 use exactly this controller (RL lower body, IK upper body) on the Unitree G1, though its official VLA pipeline switched starting in 2026 to work with a unified whole-body controller, SONIC, instead.","example":"HOMIE (Shanghai AI Lab et al., 2025): the operator uses foot pedals to send walk/turn/crouch commands to the lower-body RL policy, while the arms are controlled by a joint-matched isomorphic exoskeleton and the hands by motion-capture gloves — letting a single person teleoperate the humanoid to walk and manipulate at the same time while collecting training data.","related":["Whole-Body Control","Learning-Based Whole-Body Control","Inverse Kinematics (IK)","RL-based Locomotion Control","Loco-manipulation","SONIC"]},{"id":"motion-tracking","category":"control","sec":5,"tier":2,"sources":[{"title":"DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills (arXiv:1804.02717)","url":"https://arxiv.org/abs/1804.02717"},{"title":"BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion (arXiv:2508.08241)","url":"https://arxiv.org/abs/2508.08241"},{"title":"GMT: General Motion Tracking for Humanoid Whole-Body Control (arXiv:2506.14770)","url":"https://arxiv.org/abs/2506.14770"}],"as_of":"2025-11","related_ids":["motion-retargeting","deepmimic","beyondmimic","gmt","whole-body-control","motion-capture"],"name":"Motion Tracking","alt":"运动跟踪","abbr":"","aliases":["Motion Imitation","Reference Motion Tracking"],"one_liner":"A whole-body control task where a robot reproduces a reference motion, such as human motion-capture data, in real time.","explanation":"In humanoid robotics, motion tracking means giving a policy a reference motion — usually human motion-capture data retargeted into the robot's joint angles — and having the robot reproduce it frame by frame without falling over or violating its physical limits. The dominant approach follows 2018's DeepMimic: train with reinforcement learning in simulation, with a reward based on how closely each body part matches the reference pose, then transfer the trained policy to the real robot. It is the basic skill behind humanoids ‘learning human motion,’ underlying teleoperation, dancing, martial-arts performance, and many general-purpose whole-body controllers such as GMT and BeyondMimic. Don't confuse it with computer-vision object tracking or with motion capture itself — those are about ‘seeing,’ while motion tracking is about ‘doing.’","example":"BeyondMimic trained a tracking policy on the LAFAN1 motion-capture dataset; a single set of hyperparameters let a humanoid robot perform cartwheels, spinning kicks, and sprints on the real hardware.","related":["Motion Retargeting","DeepMimic","BeyondMimic","GMT","Whole-Body Control","Motion Capture"]},{"id":"fall-mitigation-and-fall-recovery","category":"control","sec":5,"tier":2,"sources":[{"title":"Unified Humanoid Fall-Safety Policy from a Few Demonstrations (arXiv 2511.07407)","url":"https://arxiv.org/abs/2511.07407"},{"title":"Learning Humanoid Standing-up Control across Diverse Postures (HoST, arXiv 2502.08378)","url":"https://arxiv.org/abs/2502.08378"},{"title":"Unified Multi-Contact Fall Mitigation Planning for Humanoids via Contact Transition Tree Optimization (arXiv 1807.08667)","url":"https://arxiv.org/abs/1807.08667"}],"as_of":"2025-11","related_ids":["fall-recovery","host","push-recovery","balance-control","damping-mode","humanoid-robot"],"name":"Fall Mitigation and Fall Recovery","alt":"跌倒保护与摔倒恢复","abbr":"","aliases":["Humanoid Fall Safety","Fall Protection"],"one_liner":"Two related abilities for a humanoid: minimizing damage while it falls, and getting itself back up afterward.","explanation":"A humanoid robot has a high center of mass and a small support base, so being pushed, stepping into a gap, or slipping can all cause a fall, and a single hard fall can damage joints, cameras, and the outer shell. Fall safety is usually split into three stages. Where possible, balance control and stepping (push recovery) are used to avoid falling in the first place. Once a fall becomes unavoidable, fall mitigation kicks in: bending the knees to lower the center of mass, reaching out an arm or stepping to make contact with the ground early, and adjusting posture to spread the impact onto sturdier parts of the body. After landing, fall recovery means getting back up from whatever position the robot ended up in. Earlier work mostly relied on hand-designed motion sequences and trajectory optimization — for example, Wang and Hauser in 2018 used contact-sequence tree search to plan multi-contact protective motions like stepping and bracing with the hands — while the mainstream approach in the last couple of years has been training a policy in simulation with reinforcement learning and transferring it to the real robot, with recent work starting to fold all three stages into a single policy.","example":"HoST (RSS 2025) used reinforcement learning in simulation to train Unitree's G1 to get up from a wide variety of fallen postures, without relying on any preset motion trajectory, and deployed it directly to indoor and outdoor scenes on the real robot; another piece of work from November 2025 combined a small number of human demonstrations with reinforcement learning on the G1 to unify fall prevention, impact mitigation, and getting up into a single policy.","related":["Fall Recovery","HoST (Humanoid Standing-up)","Push Recovery","Balance Control","Damping Mode","Humanoid Robot"]},{"id":"planning-and-control","category":"control","sec":6,"tier":2,"sources":[{"title":"Apollo Planning 模块说明（README_cn）","url":"https://raw.githubusercontent.com/ApolloAuto/apollo/master/modules/planning/planning_component/README_cn.md"},{"title":"Apollo Control 模块说明（README_cn）","url":"https://raw.githubusercontent.com/ApolloAuto/apollo/master/modules/control/control_component/README_cn.md"}],"as_of":"","related_ids":["motion-planning","trajectory-planning","model-predictive-control","whole-body-control","braincerebellum-architecture","autonomous-driving"],"name":"Planning and Control","alt":"规控（规划与控制）","abbr":"PnC","aliases":["PnC","Planning & Control"],"one_liner":"The engineering layer, and job title, covering both deciding how to move and making the actuators follow that decision.","explanation":"‘Planning and Control,’ abbreviated PnC, is common shorthand in Chinese industry for the layer that sits between perception and the actuators. The planning half turns localization, perception results, and the task goal into a path and a timed trajectory; the control half computes motor or actuator commands from the current state so the robot tracks that trajectory. The term is especially common in the self-driving industry: Baidu Apollo splits its system into a planning module (which outputs a trajectory) and a control module (which computes steering, throttle, and braking commands using LQR, PID, or MPC). At robotics companies, a PnC team is typically responsible for motion planning, trajectory optimization, MPC, and whole-body control — roughly the ‘cerebellum’ half of a brain-cerebellum split. End-to-end models now hand part of the planning job to a neural network, but low-level control and safety constraints are still mostly handled by PnC.","example":"In Apollo, the planning module outputs a driving trajectory with speed and acceleration profiles; the control module uses LQR for lateral steering and PID for longitudinal throttle and braking to keep the car on that trajectory.","related":["Motion Planning","Trajectory Planning","Model Predictive Control","Whole-Body Control","Brain–Cerebellum Architecture","Autonomous Driving"]},{"id":"motion-planning","category":"control","sec":6,"tier":1,"sources":[{"title":"Wikipedia: Motion planning","url":"https://en.wikipedia.org/wiki/Motion_planning"},{"title":"MoveIt Docs: Motion Planning","url":"https://moveit.picknik.ai/main/doc/concepts/motion_planning.html"}],"as_of":"","related_ids":["path-planning","trajectory-planning","configuration-space","rapidly-exploring-random-tree","collision-checking","moveit-motion-planning-framework"],"name":"Motion Planning","alt":"运动规划","abbr":"","aliases":["Motion Planner"],"one_liner":"Computing a sequence of poses or a trajectory from start to goal that avoids collisions and satisfies constraints.","explanation":"Motion planning answers the question of how a robot moves from its current state to a goal state: finding a continuous, feasible sequence of configurations that avoids obstacles, avoids self-collision, and stays within joint limits — the classic formulation of this is the 'piano mover's problem.' Planning is usually done in configuration space, the space of every possible combination of the robot's joint values; a six-axis arm's configuration space is 6-dimensional. Common methods include grid search (such as A*), suited to low dimensions; sampling-based planning (such as RRT or PRM), which scatters random points in high-dimensional space and connects them into a path, the mainstream approach for robot arms; artificial potential fields; and trajectory optimization, which treats smoothness and time as costs to minimize. The resulting geometric path is then time-parameterized to respect velocity and acceleration limits, turning it into a trajectory that motion control tracks. End-to-end VLA models output actions directly, skipping explicit planning, but planners remain the workhorse in industrial and navigation settings.","example":"Using MoveIt to have a robot arm retrieve a cup from a cabinet: a sampling-based planner in OMPL finds a collision-free path in joint space around the cabinet door, which is then time-parameterized according to each joint's velocity and acceleration limits before being sent for execution.","related":["Path Planning","Trajectory Planning","Configuration Space (C-Space)","Rapidly-exploring Random Tree","Collision Checking","MoveIt Motion Planning Framework"]},{"id":"path-planning","category":"control","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Motion planning","url":"https://en.wikipedia.org/wiki/Motion_planning"},{"title":"Lynch & Park, Modern Robotics（§9.1 path 与 trajectory 的定义；第 10 章 Motion Planning）","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["motion-planning","trajectory-planning","a-star-search","rapidly-exploring-random-tree","probabilistic-roadmap","configuration-space"],"name":"Path Planning","alt":"路径规划","abbr":"","aliases":["Path Search","Pathfinding","Piano Mover's Problem"],"one_liner":"Finding, without collisions, a sequence of positions or poses to pass through from a start to a goal.","explanation":"Path planning is the problem of computing a collision-free geometric route through an environment given a start and a goal; it is often used interchangeably with ‘motion planning,’ and is also called the ‘piano mover's problem.’ Strictly, a path only describes which configurations (combinations of position and orientation) are visited in sequence, with no notion of time; adding a velocity and acceleration at each moment turns a path into a trajectory, a step called trajectory planning or time parameterization. Common algorithms fall into three families: grid search (A*, Dijkstra), suited to 2D map navigation; sampling-based methods (PRM, RRT), suited to high-dimensional configuration spaces such as robot arms; and artificial potential fields, simple but prone to getting stuck in local minima. Mobile robots typically run a global path planner first, then a local planner that avoids obstacles as it goes.","example":"A mobile robot searches a grid map with A* for a route from its charging dock to the front door that avoids furniture, then hands that jagged path to trajectory planning and the chassis controller for execution.","related":["Motion Planning","Trajectory Planning","A* Search","Rapidly-exploring Random Tree","Probabilistic Roadmap","Configuration Space (C-Space)"]},{"id":"obstacle-avoidance","category":"control","sec":6,"tier":1,"sources":[{"title":"Wikipedia: Obstacle avoidance","url":"https://en.wikipedia.org/wiki/Obstacle_avoidance"},{"title":"Wikipedia: Motion planning","url":"https://en.wikipedia.org/wiki/Motion_planning"}],"as_of":"","related_ids":["motion-planning","collision-checking","costmap","dynamic-window-approach","control-barrier-function","collision-detection"],"name":"Obstacle Avoidance","alt":"避障","abbr":"","aliases":["Collision Avoidance"],"one_liner":"A robot senses an obstacle and adjusts its path or motion ahead of time so it doesn't run into it.","explanation":"Obstacle avoidance means a robot detects an obstacle while moving and steers around it while still reaching its goal. It's a real-time sense-decide-act process: sensors such as lidar, depth cameras, or ultrasonics detect the obstacle, and a planning or control module adjusts the path. Approaches fall into roughly three categories: computing a collision-free route up front during global planning, using algorithms like A* or RRT; reacting locally in real time, such as a mobile robot's costmap plus dynamic window approach, or a robot arm's collision checking or control barrier functions; and learning avoidance behavior directly from sensor data with methods like reinforcement learning. It's different from 'collision detection (robot safety)': obstacle avoidance steers around a collision before it happens, while collision detection notices a collision after it has already occurred and stops or backs away. Moving obstacles, such as walking people, are much harder to handle than static ones, since the robot also has to predict how they'll move.","example":"A vacuuming robot follows its planned cleaning route and encounters a slipper freshly dropped on the floor; the local planner steers around it in real time based on the lidar data and updates the obstacle into its map.","related":["Motion Planning","Collision Checking","Costmap","Dynamic Window Approach","Control Barrier Function","Collision Detection (Robot Safety)"]},{"id":"costmap","category":"control","sec":6,"tier":2,"sources":[{"title":"Nav2 Docs: Environmental Representation","url":"https://docs.nav2.org/rolling/getting_started/navigation_concepts/environmental_representation/"},{"title":"Nav2 Docs: Costmap 2D 配置（含 local_costmap 示例）","url":"https://docs.nav2.org/rolling/configuration_and_development/configuration_guide/core_servers/costmap_2d/"},{"title":"navigation2 源码 cost_values.hpp（代价取值定义）","url":"https://github.com/ros-navigation/navigation2/blob/main/nav2_costmap_2d/include/nav2_costmap_2d/cost_values.hpp"}],"as_of":"2026-09","related_ids":["occupancy-grid-map","global-planning-and-local-planning","ros-2-navigation-stack","path-planning","obstacle-avoidance","dynamic-window-approach"],"name":"Costmap","alt":"代价地图","abbr":"","aliases":["Cost Grid Map","costmap_2d"],"one_liner":"A grid map of the ground where each cell is labeled with a cost of passing through it, used for navigation planning and obstacle avoidance.","explanation":"A costmap is the most common environment representation for mobile-robot navigation, built around ROS's costmap_2d and central to Nav2. It divides the ground into a regular 2D grid, with each cell storing a cost value from 0 to 255: in Nav2, 0 means free, 254 means a lethal obstacle, 253 means the robot's center entering that cell would definitely cause a collision (inflated based on the robot's footprint), and 255 means unknown. The map is built by stacking several layers: a static layer comes from a pre-built map, an obstacle layer and a voxel layer are written in real time from sensor data, and an inflation layer adds decaying cost around obstacles so planned paths automatically keep some distance from them. Global planning searches for a path on a costmap covering the whole map, while local control avoids obstacles on a smaller costmap centered on and moving with the robot.","example":"In a typical Nav2 example configuration, the local costmap is a 3 m × 3 m rolling window centered on the robot at 0.05 m resolution — a 60×60 grid — continuously refreshed as the robot moves.","related":["Occupancy Grid Map","Global Planning and Local Planning","ROS 2 Navigation Stack (Nav2)","Path Planning","Obstacle Avoidance","Dynamic Window Approach"]},{"id":"global-planning-and-local-planning","category":"control","sec":6,"tier":2,"sources":[{"title":"Nav2 Docs: Navigation Servers（planner 与 controller）","url":"https://docs.nav2.org/rolling/getting_started/navigation_concepts/navigation_servers/"},{"title":"Nav2 Docs: Controller Server 配置（controller_frequency 默认 20 Hz）","url":"https://docs.nav2.org/rolling/configuration_and_development/configuration_guide/core_servers/controller_server/"}],"as_of":"2026-09","related_ids":["costmap","ros-2-navigation-stack","a-star-search","dynamic-window-approach","timed-elastic-band","path-planning"],"name":"Global Planning and Local Planning","alt":"全局规划与局部规划","abbr":"","aliases":["Global Planner","Local Planner","Local Controller"],"one_liner":"Navigation split into two layers: global planning finds a route across the whole map, local planning follows it while dodging nearby obstacles.","explanation":"Mobile-robot navigation is usually split into two layers. The global planner (called 'planner' in Nav2) works on a costmap covering the entire map, using an algorithm such as Dijkstra, A*, or hybrid A* to compute a complete path from the current position to the goal, updated at a relatively low frequency. The local planner (called 'controller' in Nav2) only looks at a costmap covering a few meters around the robot and moving with it, computing velocity commands at a higher frequency (20 Hz by default in Nav2) to follow the global path while avoiding transient obstacles like pedestrians or carts, using algorithms such as the dynamic window approach, timed elastic bands, MPPI, or pure pursuit. This split exists because searching the whole map is too slow to do in real time, while looking only locally risks getting stuck in dead ends.","example":"An AMR delivering goods in a warehouse: the global planner first lays out a route down the shelving aisles to bin 3; along the way, someone pushes a cart across the aisle, the local controller spots the new obstacle in its local costmap, slows down and steers around it, and then rejoins the original route.","related":["Costmap","ROS 2 Navigation Stack (Nav2)","A* Search","Dynamic Window Approach","Timed Elastic Band","Path Planning"]},{"id":"dijkstra-s-algorithm","category":"control","sec":6,"tier":2,"sources":[{"title":"Wikipedia: Dijkstra's algorithm","url":"https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm"},{"title":"Nav2 Docs: NavFn Planner（wavefront Dijkstra or A*）","url":"https://docs.nav2.org/rolling/configuration_and_development/configuration_guide/planners_plugins/configuring_navfn/"}],"as_of":"","related_ids":["a-star-search","path-planning","costmap","global-planning-and-local-planning","hybrid-a-star"],"name":"Dijkstra's Algorithm","alt":"Dijkstra 算法","abbr":"","aliases":[],"one_liner":"A classic algorithm that finds the shortest path from a start node to every other node in a graph with non-negative edge weights.","explanation":"Dijkstra's algorithm was conceived by the Dutch computer scientist Edsger Dijkstra in 1956 and published in 1959, for finding the shortest single-source paths on a graph whose edge weights (traversal costs) are all non-negative. It works by tracking the currently known shortest distance to every node, repeatedly picking the not-yet-finalized node with the smallest distance, and using it to update the distances of its neighbors. It's guaranteed to find the optimal solution, and with a binary heap it runs in about O((V+E)logV), where V and E are the numbers of nodes and edges. A* is a generalization of it: A* adds an estimate of the remaining distance to the goal (a heuristic function) and prioritizes searching toward the target, so it expands fewer nodes. In robot navigation, each cell of a grid map is a node, and the value in the costmap serves as the edge weight; Nav2's default NavFn planner can be configured to expand with either Dijkstra or A*.","example":"Three points A, B, C: A→B costs 1, B→C costs 2, and a direct A→C link costs 4. The algorithm finalizes B first (at distance 1), then uses B to update C's distance from 4 down to 3, giving A→B→C as the final shortest path.","related":["A* Search","Path Planning","Costmap","Global Planning and Local Planning","Hybrid A*"]},{"id":"a-star-search","category":"control","sec":6,"tier":2,"sources":[{"title":"Wikipedia: A* search algorithm","url":"https://en.wikipedia.org/wiki/A*_search_algorithm"}],"as_of":"","related_ids":["dijkstra-s-algorithm","hybrid-a-star","path-planning","costmap","global-planning-and-local-planning","rapidly-exploring-random-tree"],"name":"A* Search","alt":"A* 算法","abbr":"A*","aliases":["A*","A-Star"],"one_liner":"A shortest-path search algorithm that expands whichever node has the lowest cost so far plus estimated cost to the goal.","explanation":"A* is a search algorithm for finding the shortest path on a graph or grid map, proposed in 1968 by Hart, Nilsson, and Raphael at the Stanford Research Institute for the Shakey mobile robot project. At each step it expands the node with the smallest f(n) = g(n) + h(n): g(n) is the actual cost already spent getting from the start to node n, and h(n) is an estimate of the cost from n to the goal (the heuristic function, such as straight-line distance). As long as h never overestimates the true remaining cost (called admissible), A* is guaranteed to find the shortest path; when h is always zero, it degenerates into Dijkstra's algorithm, which searches outward evenly in all directions. It's commonly used for global path planning in mobile robots and in game pathfinding; a robot arm's high-dimensional joint space usually calls for sampling-based planners like RRT instead. A variant that accounts for a vehicle's turning constraints is called hybrid A*.","example":"A vacuuming robot navigating from the living room to the bedroom on an occupancy grid with 5 cm cells: each step costs 1, h is taken as the straight-line distance from the current cell to the bedroom, and A* finds the shortest route around occupied cells, which is then handed to a local planner to follow.","related":["Dijkstra's Algorithm","Hybrid A*","Path Planning","Costmap","Global Planning and Local Planning","Rapidly-exploring Random Tree"]},{"id":"hybrid-a-star","category":"control","sec":6,"tier":3,"sources":[{"title":"Dolgov, Thrun, Montemerlo, Diebel: Practical Search Techniques in Path Planning for Autonomous Driving (2008)","url":"https://ai.stanford.edu/~ddolgov/papers/dolgov_gpp_stair08.pdf"},{"title":"Open Robotics Discourse: [Nav2] SmacPlanner (Hybrid-A*, 2D A*) Now Available (2020-10)","url":"https://discourse.openrobotics.org/t/nav2-smacplanner-hybrid-a-2d-a-now-available-reminder-meeting-oct-15-cancelled/16759"}],"as_of":"","related_ids":["a-star-search","nonholonomic-constraint","path-planning","kinodynamic-planning","path-smoothing","autonomous-driving"],"name":"Hybrid A*","alt":"混合 A*","abbr":"","aliases":["Hybrid-A*"],"one_liner":"An A* variant that searches over a grid but tracks a vehicle's continuous position and heading, so the resulting path is actually drivable.","explanation":"Hybrid A* was designed by Dolgov, Thrun, and colleagues for Stanford's self-driving car Junior, for the 2007 DARPA Urban Challenge. Plain A* just hops between grid-cell centers, giving a jagged path a car can't actually follow, since a car can't turn in place (a nonholonomic constraint). Hybrid A* instead searches over position x, y, and heading θ: grid cells are still used to deduplicate nodes, but each node keeps the vehicle's true continuous pose; expansion simulates a short arc using a handful of fixed steering angles (including reverse), each consistent with the vehicle's kinematics. The heuristic takes the larger of two estimates: the shortest distance ignoring obstacles but accounting for turning ability, and the shortest grid distance accounting for obstacles but ignoring heading; the final path is smoothed with numerical optimization. It's commonly used for parking and three-point turns.","example":"Junior used it in the competition to back into parking spots and make U-turns on blocked roads, with the paper reporting a full replan taking roughly 50–300 milliseconds. ROS 2's Nav2 navigation framework also provides a Hybrid-A* implementation in its Smac Planner, aimed at Ackermann-steered, car-like bases.","related":["A* Search","Nonholonomic Constraint","Path Planning","Kinodynamic Planning","Path Smoothing (Shortcutting)","Autonomous Driving"]},{"id":"replanning","category":"control","sec":6,"tier":2,"sources":[{"title":"Nav2 默认行为树 navigate_to_pose_w_replanning_and_recovery.xml（1 Hz 重规划）","url":"https://github.com/ros-navigation/navigation2/blob/main/nav2_bt_navigator/behavior_trees/navigate_to_pose_w_replanning_and_recovery.xml"},{"title":"Inner Monologue: Embodied Reasoning through Planning with Language Models","url":"https://arxiv.org/abs/2207.05608"}],"as_of":"","related_ids":["failure-recovery","model-predictive-control","d-star-d-star-lite","closed-loop-control","behavior-tree","llm-based-task-planning"],"name":"Replanning","alt":"重规划","abbr":"","aliases":["Online Replanning","Dynamic Replanning"],"one_liner":"Recomputing a path or set of steps mid-execution in response to new information, instead of sticking to the original plan.","explanation":"Replanning means a robot recomputes its remaining path, trajectory, or task steps during execution because the environment changed, execution failed, or new information arrived. A plan computed in advance assumes a static world, but in reality new obstacles appear, objects get bumped out of place, and grasps fail — continuing with the original plan in these cases leads to failure. It happens at every level: in navigation, the ROS 2 navigation stack Nav2's default behavior tree recomputes the global path periodically at 1 Hz; in control, model predictive control re-solves a short future window every cycle, which can be seen as very high-frequency replanning; at the task level, Inner Monologue feeds success detection and scene descriptions back to a large language model so it can adjust its next step after a failure.","example":"A mobile robot navigating to the kitchen with Nav2 is on its way when someone moves a chair into the hallway; once the costmap updates, the next replanning cycle produces a new route around the chair, and if no path can be found it triggers recovery behaviors such as clearing the costmap, rotating in place, or backing up.","related":["Failure Recovery","Model Predictive Control","D* / D* Lite","Closed-loop Control","Behavior Tree","LLM-based Task Planning"]},{"id":"d-star-d-star-lite","category":"control","sec":6,"tier":3,"sources":[{"title":"Wikipedia: D*","url":"https://en.wikipedia.org/wiki/D*"},{"title":"Koenig & Likhachev, D* Lite (AAAI 2002)","url":"https://idm-lab.org/bib/abstracts/Koen02e.html"}],"as_of":"","related_ids":["a-star-search","dijkstra-s-algorithm","replanning","path-planning","global-planning-and-local-planning","costmap"],"name":"D* / D* Lite","alt":"D* / D* Lite","abbr":"","aliases":["Dynamic A*","Focused D*","Incremental Replanning"],"one_liner":"A graph-search algorithm that, as the map changes while a robot moves, patches only the affected part to quickly recompute the shortest path.","explanation":"D* was proposed by Anthony Stentz in 1994 — the name comes from ‘Dynamic A*’ — and in 2002 Sven Koenig and Maxim Likhachev built on LPA* to propose the simpler D* Lite, which is more widely used today. It handles navigation when the map isn't fully known in advance: the robot first plans a route from what it knows (treating unexplored areas as passable by default), then discovers new obstacles with its sensors as it drives. Plain A* has to search from scratch every time; the D* family instead searches backward from the goal to the start and keeps the cost values from the previous round, so when the cost of some edges changes, only the affected nodes need patching, making replanning fast. According to Wikipedia, navigation systems based on it were prototyped on the Mars rovers Spirit and Opportunity, and it was also used on the self-driving car that won CMU's entry in the DARPA Urban Challenge. It's a global planner, commonly paired with a local planner such as DWA.","example":"A warehouse mobile robot plans a route through a certain aisle using an old map, but partway through, its lidar discovers the aisle blocked by a pallet. Once the costmap updates, D* Lite only updates the cost of nodes near the blocked cell to produce a detour, without rerunning A* over the entire map.","related":["A* Search","Dijkstra's Algorithm","Replanning","Path Planning","Global Planning and Local Planning","Costmap"]},{"id":"artificial-potential-field","category":"control","sec":6,"tier":2,"sources":[{"title":"Khatib, Real-Time Obstacle Avoidance for Manipulators and Mobile Robots (IJRR 1986)","url":"https://doi.org/10.1177/027836498600500106"},{"title":"Wikipedia: Motion planning（Artificial potential fields 一节）","url":"https://en.wikipedia.org/wiki/Motion_planning"}],"as_of":"","related_ids":["obstacle-avoidance","path-planning","rapidly-exploring-random-tree","a-star-search","riemannian-motion-policies","control-barrier-function"],"name":"Artificial Potential Field","alt":"人工势场法","abbr":"APF","aliases":["APF","Potential Field Method"],"one_liner":"The goal generates an attractive force and obstacles generate repulsive forces; the robot moves along the combined force in real time.","explanation":"The artificial potential field method was introduced by Khatib in papers at ICRA 1985 and IJRR 1986 for real-time obstacle avoidance in robot arms and mobile robots, and was first implemented on a PUMA 560 arm. It treats space as a potential-energy map: potential is lowest at the goal, which produces an attractive force, and it's high near obstacles, which produces a repulsive force (usually only within a limited range). At every control cycle, the robot computes the direction of steepest potential decrease (the negative gradient, i.e. the combined-force direction) and takes a small step that way. Its advantages are very low computational cost, the ability to update in real time from sensor data, and the ability to handle moving obstacles; its main drawback is getting stuck in local minima — when the attractive and repulsive forces happen to cancel out, the robot stalls or oscillates in place — and it doesn't guarantee the shortest path. It's therefore mostly used as a local obstacle-avoidance layer alongside global planners like A* or RRT; later reactive methods such as Riemannian motion policies follow a similar idea.","example":"A mobile robot needs to pass through a narrow doorway, with a repulsive field on each side of the frame and an attractive field from the goal beyond it; if the doorway is narrow enough, the two side repulsive forces combined can cancel out the attraction, and the robot stalls in front of the door — a local minimum.","related":["Obstacle Avoidance","Path Planning","Rapidly-exploring Random Tree","A* Search","Riemannian Motion Policies","Control Barrier Function"]},{"id":"dynamic-window-approach","category":"control","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Dynamic window approach","url":"https://en.wikipedia.org/wiki/Dynamic_window_approach"},{"title":"Nav2 DWB Controller README（successor to DWA in ROS 1）","url":"https://github.com/ros-navigation/navigation2/tree/main/nav2_dwb_controller"}],"as_of":"","related_ids":["global-planning-and-local-planning","timed-elastic-band","artificial-potential-field","costmap","ros-2-navigation-stack","obstacle-avoidance"],"name":"Dynamic Window Approach","alt":"动态窗口法","abbr":"DWA","aliases":["DWA","dwa_local_planner"],"one_liner":"Sampling, simulating, and scoring the velocities a robot could reach next, then picking the best one — a classic local obstacle-avoidance method.","explanation":"The dynamic window approach, proposed by Dieter Fox, Wolfram Burgard, and Sebastian Thrun in 1997, is one of the most classic local obstacle-avoidance algorithms for mobile robots. It searches for a control command directly in velocity space — for a differential-drive base, the candidates are combinations of linear velocity v and angular velocity ω. The ‘dynamic window’ is the small range of velocities actually reachable in the next control cycle given the robot's maximum acceleration and deceleration; velocities that can't brake in time to avoid a collision are then filtered out. Each remaining (v, ω) pair is forward-simulated along a short circular-arc trajectory and scored with G = α·heading toward goal + β·distance from obstacles + γ·speed, and the highest-scoring pair is executed. It accounts for the robot's acceleration limits and is cheap to compute, but since it only looks a short distance ahead, it can get stuck at U-shaped obstacles, so it usually needs a global planner such as A* or D* to give a rough route first. Both ROS 1's dwa_local_planner and Nav2's DWB controller are built on this idea.","example":"A differential-drive robot running in the ROS navigation stack: a global planner gives a route, and each control cycle a local planner samples a batch of (v, ω) pairs, simulates each one's short-term trajectory, discards any that would hit a costmap obstacle, and sends the chassis the pair that stays close to the route while moving fastest.","related":["Global Planning and Local Planning","Timed Elastic Band","Artificial Potential Field","Costmap","ROS 2 Navigation Stack (Nav2)","Obstacle Avoidance"]},{"id":"timed-elastic-band","category":"control","sec":6,"tier":3,"sources":[{"title":"teb_local_planner README（rst-tu-dortmund, GitHub）","url":"https://github.com/rst-tu-dortmund/teb_local_planner"},{"title":"Quinlan, Khatib: Elastic bands: connecting path planning and control (ICRA 1993)","url":"https://doi.org/10.1109/robot.1993.291936"}],"as_of":"","related_ids":["global-planning-and-local-planning","dynamic-window-approach","costmap","trajectory-optimization","ackermann-steering-chassis","ros-2-navigation-stack"],"name":"Timed Elastic Band","alt":"时间弹性带","abbr":"TEB","aliases":["TEB","teb_local_planner"],"one_liner":"A local motion planner that treats a timed trajectory as an elastic band, optimized in real time against obstacles and motion limits.","explanation":"TEB was proposed in 2012 by Christoph Rösmann and colleagues at TU Dortmund, and its open-source implementation, teb_local_planner, is a commonly used local-planner plugin in the ROS navigation stack. It builds on the 1993 ‘elastic band’ idea from Sean Quinlan and Oussama Khatib: a global path is treated as a rubber band that obstacles push away from and that otherwise pulls itself taut. TEB's addition is a time interval ΔT between adjacent poses, which lets it simultaneously optimize total travel time, distance to obstacles, and kinematic and dynamic constraints such as velocity, acceleration, and minimum turning radius. Each objective involves only a handful of neighboring poses, so the problem stays sparse and can be solved in real time with the g2o graph-optimization library. Later versions optimize several topologically distinct routes in parallel — going around an obstacle on the left versus the right — and pick the best one, and they support Ackermann-steered, car-like robots. Unlike the dynamic window approach, which only samples the next instant's velocity, TEB explicitly plans an entire timed trajectory, which serves it better in narrow passages and on car-like chassis.","example":"An Ackermann-steered inspection robot meets an obstacle in a corridor: TEB simultaneously optimizes ‘go around on the left’ and ‘go around on the right’ candidate trajectories, picks whichever takes less time while respecting the minimum turning radius, and re-optimizes against the updated costmap every control cycle.","related":["Global Planning and Local Planning","Dynamic Window Approach","Costmap","Trajectory Optimization","Ackermann Steering Chassis","ROS 2 Navigation Stack (Nav2)"]},{"id":"pure-pursuit","category":"control","sec":6,"tier":3,"sources":[{"title":"R. C. Coulter, Implementation of the Pure Pursuit Path Tracking Algorithm (CMU-RI-TR-92-01, 1992)","url":"https://publications.ri.cmu.edu/storage/publications/pub_files/pub3/coulter_r_craig_1992_1/coulter_r_craig_1992_1.pdf"},{"title":"Nav2 Regulated Pure Pursuit Controller README","url":"https://github.com/ros-navigation/navigation2/tree/main/nav2_regulated_pure_pursuit_controller"}],"as_of":"","related_ids":["trajectory-tracking","path-planning","global-planning-and-local-planning","dynamic-window-approach","timed-elastic-band","ros-2-navigation-stack"],"name":"Pure Pursuit","alt":"纯追踪算法","abbr":"","aliases":["Pure Pursuit Controller","Pure Tracking"],"one_liner":"A path-tracking method that picks a lookahead point ahead on the path and steers along the arc needed to reach it.","explanation":"Pure pursuit is a geometric path-tracking algorithm. Carnegie Mellon first used it in the 1980s on the Terragator mobile robot and later on the NavLab self-driving car; R. Craig Coulter's 1992 technical report gives the standard derivation. The method: find a target point on the path at a fixed 'lookahead distance' L from the vehicle, construct an arc from the vehicle's current position that is tangent to its heading and passes through the target point, and compute its curvature as κ = 2x/L², where x is the target point's lateral offset in the vehicle's own coordinate frame. That curvature is then converted into a steering angle or angular velocity. The only tuning parameter is L: a larger L returns to the path more smoothly with less oscillation but cuts corners; a smaller L tracks more tightly but is prone to oscillating. Pure pursuit only handles geometry, not vehicle dynamics or obstacles, so it is usually paired with a global planner and local obstacle avoidance.","example":"The ROS 2 navigation stack Nav2 ships a Regulated Pure Pursuit controller for differential-drive, Ackermann-steered, and legged robots: it varies the lookahead distance with speed and automatically slows down on sharp turns or near obstacles.","related":["Trajectory Tracking","Path Planning","Global Planning and Local Planning","Dynamic Window Approach","Timed Elastic Band","ROS 2 Navigation Stack (Nav2)"]},{"id":"velocity-obstacles","category":"control","sec":6,"tier":3,"sources":[{"title":"Fiorini, Shiller: Motion Planning in Dynamic Environments Using Velocity Obstacles (IJRR 1998)","url":"https://doi.org/10.1177/027836499801700706"},{"title":"ORCA: Optimal Reciprocal Collision Avoidance（UNC GAMMA 项目页）","url":"https://gamma.cs.unc.edu/ORCA/"},{"title":"RVO2 Library: Reciprocal Collision Avoidance for Real-Time Multi-Agent Simulation","url":"https://gamma.cs.unc.edu/RVO2/"}],"as_of":"","related_ids":["obstacle-avoidance","social-navigation","multi-agent-path-finding","multi-robot-collaboration","dynamic-window-approach","autonomous-mobile-robot"],"name":"Velocity Obstacles","alt":"速度障碍法 / ORCA","abbr":"","aliases":["VO","Reciprocal Velocity Obstacles (RVO)","Optimal Reciprocal Collision Avoidance (ORCA)","RVO2"],"one_liner":"A method that marks the set of velocities in velocity space that would cause a collision, then picks a safe one closest to what's desired.","explanation":"Velocity obstacles (VO) were introduced by Paolo Fiorini and Zvi Shiller in a 1998 IJRR paper: assuming an obstacle keeps its current velocity, the method collects every velocity that would collide with it at some future time into a cone-shaped region in velocity space, and picking any velocity outside that cone avoids collision. When multiple robots do this at once, they tend to dodge in sync and then swing back in sync, oscillating. In 2008, Jur van den Berg and colleagues proposed reciprocal velocity obstacles (RVO), which assume the other side dodges too; van den Berg, Stephen Guy, Ming Lin, and Dinesh Manocha then proposed ORCA (Optimal Reciprocal Collision Avoidance, formally published in 2011), where each pair of agents takes on exactly half the responsibility for avoiding each other. Each neighbor's constraint becomes a half-plane in velocity space, and solving a low-dimensional linear program picks the feasible velocity closest to the one the agent actually wants. ORCA needs no communication and can handle thousands of agents in real time; the open-source RVO2 library that implements it is widely used in crowd simulation, games, and multi-robot navigation. It only accounts for velocity and geometry, not more complex dynamics.","example":"Several autonomous mobile robots (AMRs) in a warehouse meet at an intersection. Each independently computes its move with ORCA, taking on half the avoidance responsibility for each neighbor, so all of them swerve and slow slightly and pass through without needing any central coordination.","related":["Obstacle Avoidance","Social Navigation","Multi-Agent Path Finding","Multi-robot Collaboration","Dynamic Window Approach","Autonomous Mobile Robot"]},{"id":"multi-agent-path-finding","category":"control","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Multi-agent pathfinding","url":"https://en.wikipedia.org/wiki/Multi-agent_pathfinding"},{"title":"Multi-Agent Pathfinding: Definitions, Variants, and Benchmarks (arXiv 1906.08291)","url":"https://arxiv.org/abs/1906.08291"}],"as_of":"","related_ids":["path-planning","multi-robot-collaboration","a-star-search","fleet-management-system","velocity-obstacles","autonomous-mobile-robot"],"name":"Multi-Agent Path Finding","alt":"多智能体路径规划","abbr":"MAPF","aliases":["MAPF","Multi-Robot Path Planning"],"one_liner":"Planning routes for a whole group of robots to their respective goals at the same time, guaranteeing none of them collide.","explanation":"Multi-agent path finding studies the problem: given a map with multiple agents (robots, AGVs, etc.), each with its own start and goal, compute collision-free paths for all of them together. The classic setup discretizes the map into a grid or graph and time into steps, with each agent, each step, allowed to move to an adjacent cell or wait. There are two main kinds of conflict: two agents occupying the same cell at the same time (a vertex conflict), or two agents swapping positions between adjacent cells (an edge conflict). Common objectives are the sum of all agents' arrival times, or the time of the last arrival (makespan). Finding the optimal solution is NP-hard; common algorithms include conflict-based search (CBS, which plans each agent independently first, then adds constraints and replans whenever a conflict is found) and prioritized planning (planning agents in order, with later ones yielding to earlier ones). A 2019 survey by Stern et al. unified the definitions of various variants and gave grid-based benchmarks. It is the core problem behind warehouse robot fleet scheduling, and complements single-robot path planning and local avoidance methods such as ORCA.","example":"In an e-commerce warehouse, large numbers of transport robots carry shelves to picking stations (Amazon's Kiva system is a well-known example of this scenario); every time a new task arrives, the scheduling system has to replan conflict-free paths for these robots.","related":["Path Planning","Multi-robot Collaboration","A* Search","Fleet Management System (e.g. Open-RMF)","Velocity Obstacles","Autonomous Mobile Robot"]},{"id":"frontier-based-exploration","category":"control","sec":6,"tier":3,"sources":[{"title":"Yamauchi, A Frontier-Based Approach for Autonomous Exploration (IEEE CIRA 1997)","url":"https://www.cs.cmu.edu/~motionplanning/papers/sbp_papers/integrated1/yamauchi_frontiers.pdf"}],"as_of":"","related_ids":["active-exploration","occupancy-grid-map","simultaneous-localization-and-mapping","object-goal-navigation","vlfm","next-best-view-planning"],"name":"Frontier-based Exploration","alt":"前沿探索","abbr":"","aliases":["Frontier Exploration"],"one_liner":"A robot repeatedly drives toward the boundary between known open space and unknown territory, mapping as it goes.","explanation":"Proposed by Brian Yamauchi of the U.S. Naval Research Laboratory in 1997 (IEEE CIRA conference), this method answers: in an unfamiliar environment, where should the robot go next to gain the most new information? The approach finds ‘frontiers’ on an occupancy grid map (space divided into cells labeled free, occupied, or unknown) — free cells adjacent to unknown cells — clusters neighboring frontier cells into regions, picks one (the original paper picks the nearest) and navigates there, updates the map with sensor readings on arrival, then finds new frontiers, repeating until no frontiers remain. It's simple, makes no assumptions about wall orientation or room shape, and remains a common baseline for autonomous mapping and object-goal navigation in mobile robotics today; later work has replaced ‘nearest’ with information gain, or had a vision-language model score frontiers (as in VLFM).","example":"A quadruped robot placed inside an office building with no pre-existing map builds an occupancy grid with lidar SLAM while repeatedly driving to the nearest frontier; once every hallway and room has been explored, no frontiers remain and exploration ends.","related":["Active Exploration","Occupancy Grid Map","Simultaneous Localization and Mapping","Object-Goal Navigation","VLFM","Next-Best-View Planning"]},{"id":"coverage-path-planning","category":"control","sec":6,"tier":3,"sources":[{"title":"Mier et al., Fields2Cover: An Open-Source Coverage Path Planning Library for Unmanned Agricultural Vehicles (arXiv:2210.07838)","url":"https://arxiv.org/abs/2210.07838"},{"title":"open-navigation/opennav_coverage（Nav2 全覆盖任务服务器）","url":"https://github.com/open-navigation/opennav_coverage"},{"title":"Choset & Pignon: Coverage Path Planning: The Boustrophedon Cellular Decomposition (1998)","url":"https://publications.ri.cmu.edu/coverage-path-planning-the-boustrophedon-cellular-decomposition/"}],"as_of":"","related_ids":["path-planning","global-planning-and-local-planning","occupancy-grid-map","robot-vacuum-cleaner","ros-2-navigation-stack","multi-agent-path-finding"],"name":"Coverage Path Planning","alt":"覆盖路径规划","abbr":"CPP","aliases":["CPP","Complete Coverage Path Planning"],"one_liner":"Planning a path that sweeps a robot, or its tool, across every reachable part of an area.","explanation":"Ordinary path planning only cares about getting from A to B; coverage path planning requires a robot's working footprint (a suction nozzle, a cutting blade, a spray nozzle) to sweep every obstacle-free part of a target area, while minimizing repeated coverage, turns, and total time. Three classic families of methods exist: cell decomposition, most commonly the boustrophedon (‘ox-turning,’ i.e. back-and-forth like plowing a field) decomposition proposed by Choset and Pignon in 1998, which cuts the region into cells around obstacles and sweeps back and forth within each like plowing; grid-based methods, which divide the map into cells visited one by one; and spiral patterns, among others. Methods are further split into offline and online depending on whether the map is known in advance. Galceran and Carreras's 2013 survey in Robotics and Autonomous Systems gives a systematic overview. Typical applications include robot vacuums, lawn-mowing robots, agricultural machinery, spraying drones, and arm-based spray-painting or sanding of curved surfaces.","example":"The open-source library Fields2Cover, aimed at farm machinery, splits planning into headland generation, swath generation, swath-sequence ordering, and final path generation; Nav2's opennav_coverage, built on it, supports boustrophedon, serpentine, and spiral swath orderings, letting ROS 2 robots run full-coverage tasks directly.","related":["Path Planning","Global Planning and Local Planning","Occupancy Grid Map","Robot Vacuum Cleaner","ROS 2 Navigation Stack (Nav2)","Multi-Agent Path Finding"]},{"id":"collision-checking","category":"control","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Motion planning（collision detection 与采样规划）","url":"https://en.wikipedia.org/wiki/Motion_planning"},{"title":"FCL: The Flexible Collision Library (GitHub)","url":"https://github.com/flexible-collision-library/fcl"},{"title":"MoveIt 2 Docs: Planning Scene tutorial（collision checking）","url":"https://moveit.picknik.ai/main/doc/examples/planning_scene/planning_scene_tutorial.html"}],"as_of":"","related_ids":["self-collision-checking","collision-detection-2","flexible-collision-library","moveit-motion-planning-framework","sampling-based-planning","bounding-volume"],"name":"Collision Checking","alt":"碰撞检查","abbr":"","aliases":["Collision Query"],"one_liner":"During planning, querying whether a given pose or path segment would make the robot hit itself or the environment.","explanation":"Collision checking is a geometric query used in motion planning: given a set of joint angles (a configuration), forward kinematics first computes where every link is located in space, and then the query determines whether those geometric shapes intersect environmental obstacles or other parts of the robot itself; it can also return the closest distance and contact point. Sampling-based planners like RRT and PRM call it for every sampled point and every connecting edge, so it often accounts for a large share of total planning time. To keep it fast, a coarse pass is usually done first with bounding-volume hierarchies, followed by exact checks between convex shapes using an algorithm such as GJK. The open-source library FCL supports collision, distance, and continuous collision-detection queries; MoveIt checks both self-collision and environment collision, using an 'allowed collision matrix' to ignore link pairs that are always in contact by design. Three related but distinct concepts are worth separating: collision checking happens before execution, on the robot model; collision detection inside a physics engine is used for simulating contact; and collision detection for robot safety notices unexpected impacts during real-robot motion.","example":"A robot arm needs to place a cup inside a cabinet; RRT samples a set of joint angles and runs collision checking on it: the forearm's capsule shape intersects the cabinet door's box shape, so that sample is discarded; the straight-line connection between two valid samples also has to be checked point by point at small steps.","related":["Self-Collision Checking","Collision Detection","Flexible Collision Library (FCL)","MoveIt Motion Planning Framework","Sampling-Based Planning","Bounding Volume (AABB / OBB)"]},{"id":"self-collision-checking","category":"control","sec":7,"tier":2,"sources":[{"title":"MoveIt Setup Assistant Tutorial: Generate Self-Collision Matrix","url":"https://moveit.picknik.ai/main/doc/examples/setup_assistant/setup_assistant_tutorial.html"},{"title":"Franka FCI 文档：franka_selfcollision","url":"https://frankarobotics.github.io/docs/doc/franka_ros2_jazzy/franka_selfcollision/doc/index.html"}],"as_of":"2026-09","related_ids":["collision-checking","collision-detection","semantic-robot-description-format","collision-geometry","bimanual-manipulation","moveit-motion-planning-framework"],"name":"Self-Collision Checking","alt":"自碰撞检测","abbr":"","aliases":["Self-Collision"],"one_liner":"Determining whether a robot's own links would collide with each other in a given pose.","explanation":"Self-collision checking is a form of collision checking that only tests whether a robot's own links intersect each other at a given joint configuration — for example, the two arms of a dual-arm robot colliding, an arm hitting the torso, or the two legs touching. It is used alongside checking against environmental obstacles in motion planning, inverse-kinematics solving, and teleoperation, and can also serve as a runtime safety monitor. To save computation, an allowed-collision matrix is usually built first to rule out link pairs that never need checking: the MoveIt Setup Assistant, for instance, randomly samples 10,000 configurations by default and flags adjacent, never-colliding, or always-colliding link pairs into the SRDF. It differs from collision detection for intrinsic safety, which uses torque or momentum observers to detect a collision that has already happened; self-collision checking instead predicts geometrically, ahead of time.","example":"Franka's FR3 Duo dual-arm setup automatically runs a self-collision-monitoring node at startup, keeping a default 0.045-meter safety margin around each link, and publishes an alert on the collision_detected topic — naming the two links involved — whenever the arms get too close.","related":["Collision Checking","Collision Detection (Robot Safety)","Semantic Robot Description Format (SRDF)","Collision Geometry (Collider)","Bimanual Manipulation","MoveIt Motion Planning Framework"]},{"id":"bounding-volume","category":"control","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Bounding volume","url":"https://en.wikipedia.org/wiki/Bounding_volume"},{"title":"Wikipedia: Bounding volume hierarchy","url":"https://en.wikipedia.org/wiki/Bounding_volume_hierarchy"},{"title":"FCL (Flexible Collision Library) README","url":"https://github.com/flexible-collision-library/fcl"}],"as_of":"","related_ids":["collision-checking","broad-phase-narrow-phase-collision-detection","gilbert-johnson-keerthi-algorithm","collision-geometry","convex-decomposition","flexible-collision-library"],"name":"Bounding Volume (AABB / OBB)","alt":"包围盒","abbr":"AABB/OBB","aliases":["Bounding Volume Hierarchy","BVH","Axis-Aligned Bounding Box","Oriented Bounding Box"],"one_liner":"Wrapping a complex object in a simple shape, like a box or sphere, for a quick first pass before exact collision checks.","explanation":"A bounding volume is a simple geometric shape that fully encloses an object, used to speed up collision detection: if two bounding volumes don't overlap, the objects inside them definitely don't collide, so the expensive exact check can be skipped. The most common are AABB (axis-aligned bounding box, with edges parallel to the world axes, so intersection tests are just per-axis min/max comparisons — but it must be recomputed whenever the object rotates, and it can enclose a lot of empty space) and OBB (oriented bounding box, which rotates with the object's own frame for a tighter fit at the cost of a more expensive test); bounding spheres, capsules, k-DOPs, and convex hulls are also used. Organizing large numbers of bounding volumes into a tree gives a bounding volume hierarchy (BVH): if a parent doesn't overlap, none of its children need checking, cutting query cost from linear to logarithmic. In physics engines and motion planning, AABBs are typically used for broad-phase filtering of candidate colliding pairs, with exact narrow-phase algorithms such as GJK used afterward.","example":"The collision library FCL generally recommends a dynamic AABB tree for its broad-phase manager; for triangle-mesh models, it builds an OBBRSS (a combination of OBB and a rectangle-swept sphere) BVH by default for narrow-phase queries.","related":["Collision Checking","Broad-phase / Narrow-phase Collision Detection","Gilbert-Johnson-Keerthi Algorithm","Collision Geometry (Collider)","Convex Decomposition","Flexible Collision Library (FCL)"]},{"id":"gilbert-johnson-keerthi-algorithm","category":"control","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Gilbert–Johnson–Keerthi distance algorithm","url":"https://en.wikipedia.org/wiki/Gilbert%E2%80%93Johnson%E2%80%93Keerthi_distance_algorithm"},{"title":"MuJoCo documentation: Computation (collision detection pipelines)","url":"https://github.com/google-deepmind/mujoco/blob/main/doc/computation/index.rst"}],"as_of":"","related_ids":["collision-detection-2","collision-checking","convex-decomposition","broad-phase-narrow-phase-collision-detection","collision-geometry","flexible-collision-library"],"name":"Gilbert-Johnson-Keerthi Algorithm","alt":"GJK 算法","abbr":"GJK","aliases":["GJK","GJK Distance Algorithm"],"one_liner":"A classic iterative algorithm that tells whether two convex shapes intersect and computes the distance between them.","explanation":"Published by Elmer Gilbert, Daniel Johnson, and S. Sathiya Keerthi in 1988, GJK needs only a ‘support function’ for each convex shape — given a direction, it returns the farthest point on the shape in that direction — so spheres, boxes, capsules, and convex meshes can all be handled uniformly. Its central fact is that two convex bodies intersect if and only if their Minkowski difference (the set formed by subtracting every point of B from every point of A) contains the origin. The algorithm iteratively builds points, line segments, triangles, and tetrahedra (simplices) inside that difference set, progressively closing in on the point nearest the origin, which gives either the separation distance or confirms intersection. It's fast and memory-efficient, making it a common core of narrow-phase collision detection in physics engines — MuJoCo's default collision pipeline is based on GJK plus EPA (the expanding polytope algorithm, used to compute penetration depth). Non-convex objects need convex decomposition first.","example":"When checking clearance during motion planning, computing the distance between an arm link (represented as a convex hull) and a box on a table takes GJK only a few iterations, giving the planner a distance value to judge whether the path leaves enough margin.","related":["Collision Detection","Collision Checking","Convex Decomposition","Broad-phase / Narrow-phase Collision Detection","Collision Geometry (Collider)","Flexible Collision Library (FCL)"]},{"id":"sampling-based-planning","category":"control","sec":7,"tier":2,"sources":[{"title":"LaValle, Planning Algorithms, Chapter 5: Sampling-Based Motion Planning","url":"https://lavalle.pl/planning/ch5.pdf"},{"title":"OMPL: Available Planners","url":"https://ompl.kavrakilab.org/planners.html"}],"as_of":"","related_ids":["motion-planning","rapidly-exploring-random-tree","probabilistic-roadmap","probabilistic-completeness","configuration-space","open-motion-planning-library"],"name":"Sampling-Based Planning","alt":"基于采样的规划","abbr":"","aliases":["Sampling-Based Motion Planning"],"one_liner":"A family of motion-planning methods that randomly sample points in configuration space and connect the collision-free ones into a path.","explanation":"Sampling-based planning is one of the dominant approaches to motion planning (finding a collision-free path from start to goal). Steven LaValle's textbook Planning Algorithms frames the idea as: rather than explicitly building the set of configurations that collide, randomly sample points in configuration space (the space of all possible joint angles) and hand each point and each connecting segment to a collision-checking module as a black box. Representative algorithms include the probabilistic roadmap (PRM), which builds a reusable, multi-query road network; the rapidly-exploring random tree (RRT), which grows a tree from the start for single-query use; and the asymptotically optimal RRT* and PRM*. It excels at high-dimensional problems such as 6–7 degree-of-freedom arms, but only offers probabilistic completeness (the probability of finding a solution approaches 1 as samples grow), and paths are often jagged, needing further smoothing and time parameterization. The open-source library OMPL implements most of these algorithms.","example":"To reach a 7-DOF arm's hand from above a table into a cabinet, RRT-Connect grows one random tree from the start and one from the goal; once the trees connect, a collision-free path results, which path smoothing then cleans up by removing unnecessary turns before execution.","related":["Motion Planning","Rapidly-exploring Random Tree","Probabilistic Roadmap","Probabilistic Completeness","Configuration Space (C-Space)","Open Motion Planning Library (OMPL)"]},{"id":"probabilistic-completeness","category":"control","sec":7,"tier":3,"sources":[{"title":"S. M. LaValle, Planning Algorithms, Chapter 5: Sampling-Based Motion Planning","url":"http://lavalle.pl/planning/ch5.pdf"},{"title":"Karaman & Frazzoli, Sampling-based Algorithms for Optimal Motion Planning (arXiv:1105.1186, IJRR 2011)","url":"https://arxiv.org/abs/1105.1186"}],"as_of":"","related_ids":["sampling-based-planning","rapidly-exploring-random-tree","rapidly-exploring-random-tree","probabilistic-roadmap","configuration-space","motion-planning"],"name":"Probabilistic Completeness","alt":"概率完备性","abbr":"","aliases":["Probabilistically Complete"],"one_liner":"A planner property: whenever a solution exists, the chance of finding a feasible path approaches 1 as sampling increases.","explanation":"In motion planning, an algorithm is ‘complete’ if it finds a solution in finite time whenever one exists, and correctly reports failure when none does — only exact, combinatorial planners achieve this. Sampling-based planners like PRM and RRT can't, so they settle for the weaker guarantee of probabilistic completeness: as Steven LaValle's textbook Planning Algorithms puts it, as long as a solution exists, the probability of finding it converges to 1 as the number of samples grows. Sertac Karaman and Emilio Frazzoli gave a rigorous definition in 2011 and showed that the probability RRT or PRM fail to find an existing solution decays exponentially with the number of samples. It has two practical limits: a feasible path needs some clearance from obstacles, so very narrow passages are almost never sampled; and when no solution exists, the algorithm just keeps running rather than ever declaring failure. It is often mentioned alongside the stronger property of asymptotic optimality, where path cost converges almost surely to the optimum as samples grow — plain RRT has only probabilistic completeness, while RRT* has both.","example":"When MoveIt calls an OMPL planner, you set a planning time limit: a probabilistically complete algorithm is more likely to find a solution the more time it is given. But if the goal pose is completely boxed in by obstacles, it won't report 'no solution' — it will just run until the timeout and return failure.","related":["Sampling-Based Planning","Rapidly-exploring Random Tree","Rapidly-exploring Random Tree","Probabilistic Roadmap","Configuration Space (C-Space)","Motion Planning"]},{"id":"probabilistic-roadmap","category":"control","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Probabilistic roadmap","url":"https://en.wikipedia.org/wiki/Probabilistic_roadmap"},{"title":"Lynch & Park, Modern Robotics（§10.5.2 The PRM Algorithm）","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","sampling-based-planning","probabilistic-completeness","configuration-space","path-planning","a-star-search"],"name":"Probabilistic Roadmap","alt":"概率路线图","abbr":"PRM","aliases":["PRM","PRM Planner"],"one_liner":"A planning algorithm that first scatters random points into space and connects them into a road network, then searches it for a path.","explanation":"The probabilistic roadmap is a sampling-based motion planning algorithm generally credited to a 1996 paper by Lydia Kavraki and colleagues. It works in two phases: a construction phase that randomly samples collision-free points in configuration space (the space of all possible robot poses) and connects nearby points that can be joined by a straight, collision-free line, forming a graph — the ‘roadmap’; and a query phase that connects the start and goal into this graph and searches it with Dijkstra or A*. Its strength is that the roadmap, once built, can be queried repeatedly (multi-query), which suits settings where the environment stays largely fixed and planning happens often; its weakness is that narrow passages are hard to sample into. It has probabilistic completeness: given enough samples, the probability of finding an existing path approaches 1. RRT, by contrast, grows a single tree from the start each time and suits one-off queries better.","example":"For a fixed workstation arm, a PRM roadmap is built once in joint space ahead of time; every time the pick-and-place targets change, only the new start and goal need to be connected into the existing network and searched, without replanning from scratch.","related":["Rapidly-exploring Random Tree","Sampling-Based Planning","Probabilistic Completeness","Configuration Space (C-Space)","Path Planning","A* Search"]},{"id":"rapidly-exploring-random-tree","category":"control","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Rapidly exploring random tree","url":"https://en.wikipedia.org/wiki/Rapidly_exploring_random_tree"},{"title":"Lynch & Park, Modern Robotics（§10.5.1 The RRT Algorithm）","url":"http://hades.mech.northwestern.edu/images/7/7f/MR.pdf"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","rrt-connect","probabilistic-roadmap","sampling-based-planning","path-smoothing","collision-checking"],"name":"Rapidly-exploring Random Tree","alt":"快速扩展随机树","abbr":"RRT","aliases":["RRT","RRT Algorithm"],"one_liner":"A planning algorithm that grows a tree from the start by repeatedly sampling random points and reaching toward them, until a branch reaches the goal.","explanation":"Proposed by Steven LaValle in 1998 and further developed with James Kuffner, the rapidly-exploring random tree is one of the most widely used sampling-based planning algorithms. Each round does four things: sample a random point in configuration space; find the tree's nearest existing node to it; step a small distance from that node toward the sample; and add the new point to the tree if that step is collision-free. Large unexplored regions are more likely to be sampled, so the tree naturally grows fastest toward unexplored space. It needs only collision checking, not an explicit description of free space, which suits arms with six or seven degrees of freedom and systems with dynamics constraints. Basic RRT is probabilistically complete, but its paths are usually jagged and not optimal; common improvements include the bidirectionally growing RRT-Connect and the asymptotically optimal RRT*, and results are typically smoothed afterward.","example":"In a 2D maze, RRT grows random branches from the entrance; once a branch nears the exit, tracing it back to the root gives a jagged path, which shortcutting then straightens and shortens.","related":["Rapidly-exploring Random Tree","RRT-Connect","Probabilistic Roadmap","Sampling-Based Planning","Path Smoothing (Shortcutting)","Collision Checking"]},{"id":"rrt-star","category":"control","sec":7,"tier":3,"sources":[{"title":"Karaman & Frazzoli, Sampling-based Algorithms for Optimal Motion Planning (arXiv:1105.1186, IJRR 2011)","url":"https://arxiv.org/abs/1105.1186"},{"title":"OMPL: ompl::geometric::RRTstar","url":"https://ompl.kavrakilab.org/classompl_1_1geometric_1_1RRTstar.html"},{"title":"MoveIt 2 Documentation: OMPL Planner","url":"https://moveit.picknik.ai/main/doc/examples/ompl_interface/ompl_interface_tutorial.html"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","informed-rrt-star","batch-informed-trees","probabilistic-completeness","probabilistic-roadmap","rrt-connect"],"name":"RRT*","alt":"RRT*","abbr":"RRT*","aliases":["RRT-star","RRTstar","Optimal RRT"],"one_liner":"An extension of RRT that adds parent selection and rewiring so path cost converges to optimal as sampling increases.","explanation":"RRT* was proposed by MIT researchers Sertac Karaman and Emilio Frazzoli, with the systematic treatment appearing in their 2011 IJRR paper. They proved that the path returned by plain RRT converges almost surely to a non-optimal cost, and fixed this with two added steps: when a new node is added, among its neighbors within a certain radius, RRT* picks whichever one minimizes the total cost from the start to the new node as its parent; it then checks those same neighbors and rewires any of them to the new node if that path turns out cheaper. The neighborhood radius shrinks as the node count n grows, following γ(log n / n)^{1/d}, where d is the dimension of the space and γ is a constant tied to its size. As a result, path cost converges almost surely to the optimum as sampling increases — called asymptotic optimality — while the extra computation over plain RRT is only a constant factor. The downside is that convergence can be slow in practice, so implementations are usually just given a fixed time budget and run until it expires. Follow-up methods such as Informed RRT* and BIT* are designed specifically to speed up this convergence.","example":"OMPL's RRTstar planner doesn't return as soon as it finds a first feasible path — it keeps sampling and rewiring, continuously shortening the path within the given time budget. If a cost threshold is set, it can also stop early once the path cost drops below it.","related":["Rapidly-exploring Random Tree","Informed RRT* (Informed Sampling)","Batch Informed Trees","Probabilistic Completeness","Probabilistic Roadmap","RRT-Connect"]},{"id":"informed-rrt-star","category":"control","sec":7,"tier":3,"sources":[{"title":"Gammell, Srinivasa, Barfoot: Informed RRT* (arXiv:1404.2334, IROS 2014)","url":"https://arxiv.org/abs/1404.2334"},{"title":"OMPL: ompl::geometric::InformedRRTstar Class Reference","url":"https://ompl.kavrakilab.org/classompl_1_1geometric_1_1InformedRRTstar.html"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","rapidly-exploring-random-tree","batch-informed-trees","sampling-based-planning","open-motion-planning-library","probabilistic-completeness"],"name":"Informed RRT* (Informed Sampling)","alt":"Informed RRT*","abbr":"","aliases":["Informed Sampling"],"one_liner":"After finding a path once, sampling only inside the ellipsoidal region that could still shorten it, so RRT* converges faster.","explanation":"Informed RRT* was proposed by Gammell, Srinivasa, and Barfoot at IROS 2014. After RRT* finds a feasible path, it keeps sampling and rewiring branches to push the path toward the shortest one, but it still samples uniformly over the whole space, and most samples have no chance of improving the result. Let c_best be the length of the current best path; a point x that could shorten the path must satisfy ‖x−x_s‖ + ‖x−x_g‖ < c_best (x_s, x_g the start and goal) — in 2D this is an ellipse with the start and goal as foci, and a long, thin ellipsoid in higher dimensions. Informed RRT* samples only inside that ellipsoid and prunes nodes outside it; the shorter the path gets, the thinner the ellipsoid, and the more concentrated the search becomes. It keeps RRT*'s asymptotic optimality while converging noticeably faster in high dimensions and large maps, and later algorithms such as BIT* build on the same idea.","example":"The open-source motion-planning library OMPL provides an InformedRRTstar planner, which can be wired into frameworks like MoveIt through OMPL, letting an arm keep shortening its trajectory in any remaining time after first finding a feasible one.","related":["Rapidly-exploring Random Tree","Rapidly-exploring Random Tree","Batch Informed Trees","Sampling-Based Planning","Open Motion Planning Library (OMPL)","Probabilistic Completeness"]},{"id":"batch-informed-trees","category":"control","sec":7,"tier":3,"sources":[{"title":"Gammell, Srinivasa, Barfoot. Batch Informed Trees (BIT*) (arXiv:1405.5848, ICRA 2015)","url":"https://arxiv.org/abs/1405.5848"},{"title":"OMPL: ompl::geometric::BITstar Class Reference","url":"https://ompl.kavrakilab.org/classompl_1_1geometric_1_1BITstar.html"}],"as_of":"","related_ids":["informed-rrt-star","rapidly-exploring-random-tree","a-star-search","sampling-based-planning","probabilistic-completeness","open-motion-planning-library"],"name":"Batch Informed Trees","alt":"BIT*","abbr":"BIT*","aliases":["BIT*","BITstar"],"one_liner":"A planning algorithm that samples points in batches and searches them in a heuristic order to converge on an optimal path.","explanation":"BIT* (Batch Informed Trees) is a sampling-based motion planning algorithm by Gammell, Srinivasa, and Barfoot, published at ICRA 2015 with a complete version in IJRR in 2020. Where RRT*-family algorithms grow a tree one random point at a time, BIT* samples a whole batch at once, treats them as an implicit random geometric graph (edges aren't pre-connected — collision-checked only when needed), and searches it in an A*-like order based on the heuristic ‘how short could a path through this edge possibly be.’ After finding a first solution, it samples the next batch only inside the ellipsoidal region that could still improve the current solution (following Informed RRT*), refining round by round. It can report its best-so-far solution at any time, keeps improving, and is both probabilistically complete and asymptotically optimal; the paper's experiments show it finding better solutions faster than RRT*, Informed RRT*, and FMT*, especially in high dimensions. OMPL already includes BIT*, and later variants include ABIT*, AIT*, and EIT*.","example":"Planning a path for a 7-DOF arm around a shelf's dividers, one can swap the OMPL planner from RRTConnect to BITstar: the former just wants a usable path as fast as possible, while the latter keeps shortening the path within the given time budget.","related":["Informed RRT* (Informed Sampling)","Rapidly-exploring Random Tree","A* Search","Sampling-Based Planning","Probabilistic Completeness","Open Motion Planning Library (OMPL)"]},{"id":"rrt-connect","category":"control","sec":7,"tier":3,"sources":[{"title":"OMPL: ompl::geometric::RRTConnect（Kuffner & LaValle, ICRA 2000, pp. 995–1001）","url":"https://ompl.kavrakilab.org/classompl_1_1geometric_1_1RRTConnect.html"},{"title":"S. M. LaValle, Planning Algorithms, Chapter 5（bidirectional RDT/RRT）","url":"http://lavalle.pl/planning/ch5.pdf"},{"title":"moveit_resources: panda_moveit_config/config/ompl_planning.yaml","url":"https://github.com/moveit/moveit_resources/blob/ros2/panda_moveit_config/config/ompl_planning.yaml"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","rapidly-exploring-random-tree","sampling-based-planning","probabilistic-completeness","open-motion-planning-library","path-smoothing"],"name":"RRT-Connect","alt":"RRT-Connect","abbr":"","aliases":["Bidirectional RRT","RRTConnect"],"one_liner":"A path planner that grows two random trees, one from the start and one from the goal, and greedily connects them.","explanation":"RRT-Connect was introduced by James Kuffner and Steven LaValle at ICRA 2000 as a bidirectional version of the rapidly-exploring random tree (RRT). A plain RRT grows only one tree from the start and occasionally attempts to connect to the goal; RRT-Connect instead grows a tree from both the start and the goal. Each round, one tree extends one step toward a random sample to create a new node, and then the other tree extends step by step toward that new node until it connects or is blocked by an obstacle — the ‘CONNECT’ heuristic. A connection means a path has been found, and the two trees then swap roles for the next round. This greedy connection strategy is usually fast in spaces that aren't too cluttered, which suits arms with six or seven degrees of freedom that just need to answer a single ‘get from A to B’ query. It is probabilistically complete but doesn't aim for the shortest path, so the result is often zigzagging and needs post-processing to smooth. The open-source planning library OMPL implements it as RRTConnect, and MoveIt can call it directly.","example":"MoveIt's official Panda arm configuration lists RRTConnectkConfigDefault among its OMPL planners. After it plans a collision-free path, the path still typically needs smoothing and time parameterization before it is sent to the robot.","related":["Rapidly-exploring Random Tree","Rapidly-exploring Random Tree","Sampling-Based Planning","Probabilistic Completeness","Open Motion Planning Library (OMPL)","Path Smoothing (Shortcutting)"]},{"id":"path-smoothing","category":"control","sec":7,"tier":3,"sources":[{"title":"OMPL: ompl::geometric::PathSimplifier Class Reference","url":"https://ompl.kavrakilab.org/classompl_1_1geometric_1_1PathSimplifier.html"},{"title":"Hauser & Ng-Thow-Hing, Fast Smoothing of Manipulator Trajectories using Optimal Bounded-Acceleration Shortcuts (ICRA 2010)","url":"https://motion.cs.illinois.edu/papers/icra10-smoothing.pdf"}],"as_of":"","related_ids":["rapidly-exploring-random-tree","probabilistic-roadmap","collision-checking","time-parameterization","trajectory-optimization","open-motion-planning-library"],"name":"Path Smoothing (Shortcutting)","alt":"路径平滑","abbr":"","aliases":["Shortcutting","Path Simplification"],"one_liner":"Post-processing a jagged planned path to cut out detours and round off corners, so the robot moves shorter and smoother.","explanation":"Path smoothing is a post-processing step in motion planning. Sampling-based planners such as RRT and PRM produce paths made of many waypoints strung into a jagged line, often taking a roundabout route with back-and-forth wiggle, which would be slow and unnatural to execute as-is. The most common technique is randomized shortcutting: repeatedly pick two random points on the path and try connecting them with a shorter straight line (or a smooth curve respecting velocity and acceleration limits); if that segment is collision-free, it replaces the original path between the two points; repeating this many times keeps shortening the path. Other approaches include removing redundant waypoints and fitting a B-spline; the open-source planning library OMPL's PathSimplifier provides shortcutting, point removal, and B-spline smoothing functions. After smoothing, the path usually still needs time parameterization to assign it speed and timing before it becomes an executable trajectory. It differs from trajectory optimization in that it only makes local improvements to an existing feasible solution — cheap to compute, but with no optimality guarantee.","example":"Hauser and Ng-Thow-Hing (2010) had an arm reach under a table to grab a cup: applying 100 random shortcut attempts, each respecting velocity and acceleration limits, to a sampling planner's jagged path cut the execution time from 9.4 seconds to 4.0 seconds.","related":["Rapidly-exploring Random Tree","Probabilistic Roadmap","Collision Checking","Time Parameterization","Trajectory Optimization","Open Motion Planning Library (OMPL)"]},{"id":"kinodynamic-planning","category":"control","sec":7,"tier":3,"sources":[{"title":"Kinodynamic planning - Wikipedia","url":"https://en.wikipedia.org/wiki/Kinodynamic_planning"},{"title":"Rapidly exploring random tree - Wikipedia（引 LaValle & Kuffner 2001, Randomized Kinodynamic Planning, IJRR）","url":"https://en.wikipedia.org/wiki/Rapidly_exploring_random_tree"}],"as_of":"","related_ids":["motion-planning","rapidly-exploring-random-tree","trajectory-optimization","time-optimal-path-parameterization","nonholonomic-constraint","state-space"],"name":"Kinodynamic Planning","alt":"动力学约束规划","abbr":"","aliases":["Kinodynamic Motion Planning"],"one_liner":"Planning that satisfies obstacle avoidance and dynamics limits like velocity, acceleration, and torque together, producing a directly executable trajectory.","explanation":"Ordinary path planning only handles geometry — finding a route from A to B that avoids obstacles, with speed figured out separately. Kinodynamic planning, a term coined by Donald, Xavier, Canny, and Reif in 1993, requires the result to satisfy both kinematic constraints (obstacle avoidance, joint limits) and dynamics constraints (velocity, acceleration, force, or torque limits) at the same time. It therefore usually searches over a state space that includes both position and velocity, expanding only along motions the system can actually perform — a car, for example, can't move sideways. A representative method is LaValle and Kuffner's 2001 randomized kinodynamic planning: rather than connecting tree nodes with straight lines, it extends the RRT search tree by applying a control input and integrating the dynamics model forward for a short time to get a new state. Another common route plans a purely geometric path first and applies time parameterization afterward — simpler, but it may fail to find a feasible solution for strongly dynamic tasks like running, jumping, or high-speed flight. Trajectory optimization and MPC are also often used to solve this class of problem.","example":"Getting a quadruped robot to jump onto a step requires more than just a geometrically feasible foothold route: the joint torque available at takeoff must be enough, the center of mass must follow a parabola while airborne, and landing speed must stay within limits — all of which have to be folded into the planning together as dynamics constraints.","related":["Motion Planning","Rapidly-exploring Random Tree","Trajectory Optimization","Time-Optimal Path Parameterization","Nonholonomic Constraint","State Space"]},{"id":"euclidean-signed-distance-field","category":"control","sec":7,"tier":3,"sources":[{"title":"Voxblox: Incremental 3D Euclidean Signed Distance Fields for On-Board MAV Planning (arXiv:1611.03631)","url":"https://arxiv.org/abs/1611.03631"},{"title":"EGO-Planner: An ESDF-free Gradient-based Local Planner for Quadrotors (arXiv:2008.08835)","url":"https://arxiv.org/abs/2008.08835"}],"as_of":"","related_ids":["signed-distance-field-function","truncated-signed-distance-function","nvblox","covariant-hamiltonian-optimization-for-motion-planning","collision-checking","occupancy-grid-map"],"name":"Euclidean Signed Distance Field","alt":"欧氏符号距离场","abbr":"ESDF","aliases":["ESDF","ESDF Map"],"one_liner":"A 3D grid where every cell stores the true distance to the nearest obstacle, positive outside and negative inside, with a computable gradient.","explanation":"A Euclidean signed distance field is a 3D map representation: space is divided into voxels (small cubic cells), and each voxel stores its Euclidean distance to the nearest obstacle surface, positive outside obstacles and negative inside. It differs from TSDF (truncated signed distance function), which only holds values near an object's surface, with distance estimated along the camera's line of sight; an ESDF instead gives the true nearest distance everywhere on the map and supports computing a gradient, whose direction points ‘further from the obstacle.’ This makes it well suited to gradient-based trajectory optimization, such as CHOMP and many drone local planners: approximating the robot as a string of small spheres, as long as the distance at each sphere's center exceeds that sphere's radius there's no collision, and the gradient of the collision cost directly pushes the trajectory away from obstacles. ETH's Voxblox (2016) introduced incrementally building an ESDF from a TSDF, and NVIDIA's nvblox moved this kind of mapping onto the GPU. The downside is that maintaining a full distance field is computationally expensive, which is why work like EGO-Planner skips ESDF altogether.","example":"In Voxblox's experiments, a drone builds a TSDF on the fly from onboard sensors while flying, incrementally converts it to an ESDF, and a local trajectory optimizer reads the distance and gradient, completing mapping and replanning in real time on a single onboard CPU core.","related":["Signed Distance Field / Function","Truncated Signed Distance Function","nvblox","Covariant Hamiltonian Optimization for Motion Planning","Collision Checking","Occupancy Grid Map"]},{"id":"covariant-hamiltonian-optimization-for-motion-planning","category":"control","sec":7,"tier":3,"sources":[{"title":"Ratliff, Zucker, Bagnell, Srinivasa: CHOMP: Gradient Optimization Techniques for Efficient Motion Planning (ICRA 2009)","url":"https://publications.ri.cmu.edu/storage/publications/pub_files/2009/5/icra09-chomp.pdf"},{"title":"Zucker et al., CHOMP: Covariant Hamiltonian Optimization for Motion Planning (IJRR 2013)","url":"https://publications.ri.cmu.edu/chomp-covariant-hamiltonian-optimization-for-motion-planning/"},{"title":"MoveIt: CHOMP Planner 教程","url":"https://moveit.picknik.ai/main/doc/how_to_guides/chomp_planner/chomp_planner_tutorial.html"}],"as_of":"","related_ids":["trajectory-optimization","stomp","trajopt","signed-distance-field-function","moveit-motion-planning-framework","path-smoothing"],"name":"Covariant Hamiltonian Optimization for Motion Planning","alt":"CHOMP","abbr":"CHOMP","aliases":["CHOMP","CHOMP Planner"],"one_liner":"A trajectory-optimization planner that uses gradient descent to push an initial trajectory away from obstacles while keeping it smooth.","explanation":"CHOMP was proposed by CMU's Ratliff, Zucker, Bagnell, and Srinivasa at ICRA 2009, with an extended version in IJRR in 2013. It discretizes a trajectory into a sequence of waypoints, with cost split into two terms: a smoothness term (the sum of squared velocity and acceleration computed by finite differences), and an obstacle term (approximating the robot as a string of small spheres and querying a signed distance field of the environment — the closer to an obstacle, the higher the cost). ‘Covariant’ refers to premultiplying the gradient update by the inverse of a smoothness metric matrix, so a single change spreads smoothly across the whole trajectory rather than yanking a single waypoint; ‘Hamiltonian’ in the name refers to using Hamiltonian Monte Carlo plus random perturbation to escape local optima. It doesn't require the initial trajectory to be collision-free — it can converge starting from a straight line that passes right through an obstacle. Its drawback is getting stuck in local optima, occasionally cutting straight through a thin obstacle, so it's often given a collision-free initial guess from a sampling planner such as OMPL first. It belongs, alongside STOMP and TrajOpt, to the family of optimization-based motion planners.","example":"The original paper used it for manipulation planning on a Barrett WAM arm, and also to generate walking trajectories for the LittleDog quadruped robot. MoveIt integrates CHOMP, with smoothness_cost_weight and obstacle_cost_weight as tunable parameters balancing smoothness against obstacle avoidance.","related":["Trajectory Optimization","STOMP","TrajOpt","Signed Distance Field / Function","MoveIt Motion Planning Framework","Path Smoothing (Shortcutting)"]},{"id":"stomp","category":"control","sec":7,"tier":3,"sources":[{"title":"MoveIt 2 文档：STOMP Planner","url":"https://moveit.picknik.ai/main/doc/how_to_guides/stomp_planner/stomp_planner.html"},{"title":"MoveIt 1 教程：STOMP Planner（引用 Kalakrishnan et al. ICRA 2011）","url":"https://moveit.github.io/moveit_tutorials/doc/stomp_planner/stomp_planner_tutorial.html"}],"as_of":"","related_ids":["covariant-hamiltonian-optimization-for-motion-planning","trajopt","trajectory-optimization","motion-planning","moveit-motion-planning-framework","model-predictive-path-integral-control"],"name":"STOMP","alt":"STOMP","abbr":"STOMP","aliases":["Stochastic Trajectory Optimization for Motion Planning"],"one_liner":"A gradient-free trajectory optimizer that samples noisy trajectories around an initial guess and updates it by cost-weighted averaging.","explanation":"STOMP was proposed by Mrinal Kalakrishnan, Sachin Chitta, Evangelos Theodorou, Peter Pastor, and Stefan Schaal at ICRA 2011 as a trajectory optimization method. It starts from an initial trajectory, which may pass straight through obstacles, and each round adds noise around it to generate several candidate trajectories; it computes per-waypoint costs such as collision, constraint violation, and smoothness, then combines the noise into an update weighted so lower-cost trajectories count for more, iterating toward a smooth, collision-free result. STOMP is derived from PI², a path-integral method from reinforcement learning, and it needs no gradient of the cost function, so it can incorporate costs — like torque, energy, or end-effector orientation — that are hard to differentiate. Compared with the gradient-based CHOMP, STOMP's randomness makes it less prone to getting stuck in local optima and needs less tuning; compared with OMPL's sampling-based planners, it is usually slower but produces smoother trajectories that often need no further smoothing afterward.","example":"MoveIt implements STOMP as a planner plugin. Its behavior is set by parameters such as num_rollouts (how many noisy trajectories to generate per round), stddev (the noise magnitude per joint), and cost functions such as CollisionCheck and ObstacleDistanceGradient.","related":["Covariant Hamiltonian Optimization for Motion Planning","TrajOpt","Trajectory Optimization","Motion Planning","MoveIt Motion Planning Framework","Model Predictive Path Integral Control"]},{"id":"trajopt","category":"control","sec":7,"tier":3,"sources":[{"title":"Schulman et al.: Finding Locally Optimal, Collision-Free Trajectories with Sequential Convex Optimization (RSS 2013)","url":"https://www.roboticsproceedings.org/rss09/p31.html"},{"title":"TrajOpt documentation (UC Berkeley RLL)","url":"https://rll.berkeley.edu/trajopt/doc/sphinx_build/html/"},{"title":"tesseract-robotics/trajopt (GitHub)","url":"https://github.com/tesseract-robotics/trajopt"}],"as_of":"","related_ids":["trajectory-optimization","sequential-quadratic-programming","covariant-hamiltonian-optimization-for-motion-planning","stomp","collision-checking","open-motion-planning-library"],"name":"TrajOpt","alt":"TrajOpt","abbr":"","aliases":["Sequential Convex Trajectory Optimization","trajopt_ros","Tesseract TrajOpt"],"one_liner":"A motion-planning method and open-source library that finds locally optimal, collision-free robot trajectories via sequential convex optimization.","explanation":"TrajOpt was proposed by John Schulman, Pieter Abbeel, and colleagues at UC Berkeley at RSS 2013, with an extended version published in IJRR in 2014. It formulates motion planning as trajectory optimization: the decision variables are a sequence of joint waypoints, the objective favors a short, smooth path, and the constraints are joint limits and collision avoidance. The collision constraint is non-convex, so TrajOpt uses sequential convex optimization: each round it linearizes the cost and constraints near the current trajectory into a convex quadratic program, solves it within a trust region, and iterates. Collisions are represented using signed distance (negative when penetrating) plus a hinge penalty, with the penalty coefficient increased in an outer loop if it isn't enough; it also checks the convex hull of the robot's shape between two adjacent timesteps, guaranteeing it never ‘passes through’ a thin obstacle even in continuous time. The paper reports it solving more problems, faster, than OMPL's sampling-based planners and CHOMP. Its drawback is that it only guarantees a local optimum and can fail when the initial trajectory is too poor. The ROS-Industrial project Tesseract now maintains its C++ implementation.","example":"To make a 7-axis arm reach into a bookshelf compartment to grab an item, you give it an initial joint-space linear-interpolation trajectory that passes straight through the shelf; TrajOpt then iteratively pushes the colliding waypoints away from the shelf while keeping the trajectory smooth.","related":["Trajectory Optimization","Sequential Quadratic Programming","Covariant Hamiltonian Optimization for Motion Planning","STOMP","Collision Checking","Open Motion Planning Library (OMPL)"]},{"id":"graphs-of-convex-sets","category":"control","sec":7,"tier":3,"sources":[{"title":"Marcucci et al., Motion Planning around Obstacles with Convex Optimization (arXiv 2205.04422)","url":"https://arxiv.org/abs/2205.04422"},{"title":"Marcucci et al., Shortest Paths in Graphs of Convex Sets (arXiv 2101.11565, SIAM J. Optim. 2024)","url":"https://arxiv.org/abs/2101.11565"},{"title":"Drake: GcsTrajectoryOptimization","url":"https://drake.mit.edu/doxygen_cxx/classdrake_1_1planning_1_1trajectory__optimization_1_1_gcs_trajectory_optimization.html"}],"as_of":"","related_ids":["motion-planning","rapidly-exploring-random-tree","probabilistic-roadmap","convex-optimization","trajectory-optimization","drake"],"name":"Graphs of Convex Sets","alt":"凸集图规划","abbr":"GCS","aliases":["GCS","Graph of Convex Sets","GCS Motion Planning"],"one_liner":"Splitting free space into convex chunks connected into a graph, then using convex optimization to find a globally good, smooth trajectory.","explanation":"Proposed by Tobia Marcucci and colleagues in MIT Russ Tedrake's group: the underlying ‘shortest paths in graphs of convex sets’ theory was published in SIAM Journal on Optimization (2024), with the motion-planning application in Science Robotics (2023). Configuration space's collision-free region is first decomposed into a set of convex regions, each a node in a graph, with edges between regions that overlap; a Bézier curve then represents each trajectory segment within a region, and the problem decides both which regions to pass through and the shape of each curve. This is fundamentally a mixed-integer optimization, but its convex relaxation is tight enough that solving one convex program and rounding is usually enough to get a globally optimal or near-optimal trajectory, with an accompanying optimality bound. Compared to sampling-based planners like RRT or PRM, the resulting trajectory is smoother and higher quality; the cost is needing a convex decomposition up front. Drake already includes an implementation.","example":"Using Drake's GcsTrajectoryOptimization: given several collision-free convex regions for an arm moving among shelves, the solver outputs a smooth trajectory that passes through these regions, respects velocity limits, and takes as little time as possible.","related":["Motion Planning","Rapidly-exploring Random Tree","Probabilistic Roadmap","Convex Optimization","Trajectory Optimization","Drake"]},{"id":"riemannian-motion-policies","category":"control","sec":7,"tier":3,"sources":[{"title":"Ratliff et al., Riemannian Motion Policies (arXiv:1801.02854)","url":"https://arxiv.org/abs/1801.02854"},{"title":"Cheng et al., RMPflow: A Computational Graph for Automatic Motion Policy Generation (arXiv:1811.07049, WAFR 2018)","url":"https://arxiv.org/abs/1811.07049"},{"title":"Isaac Sim Documentation: RMPflow","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/manipulators/concepts/rmpflow.html"}],"as_of":"2026-09","related_ids":["geometric-fabrics","operational-space-control","obstacle-avoidance","artificial-potential-field","nvidia-isaac-sim","curobo"],"name":"Riemannian Motion Policies","alt":"黎曼运动策略","abbr":"RMP","aliases":["RMP","RMPflow","Riemannian Motion Policy"],"one_liner":"A reactive control framework where each sub-goal contributes a desired acceleration and an importance matrix, combined by weighted sum.","explanation":"RMPs were introduced by Nathan Ratliff, Dieter Fox, and colleagues at NVIDIA in 2018. An RMP is a pair (a, M): a is a desired acceleration defined in some task space — such as the end-effector reaching for a goal, a link moving away from an obstacle, or a joint staying away from its limit — and M is a Riemannian metric, roughly ‘how important this sub-goal is in each direction.’ Each sub-policy is defined in its own space, pulled back into joint space through a Jacobian, and combined into a single joint acceleration weighted by M. This lets conflicting requirements be handled in a unified way, with the weights changing as the state changes — the closer the robot gets to an obstacle, the more avoidance dominates. RMPflow, published the same year, organizes this composition into an automated computation graph and provides stability conditions. RMPs are recomputed every control cycle, so they can react to moving targets and obstacles, but they are purely local and still need a global planner for complex scenes. NVIDIA's later Geometric Fabrics work continues this line.","example":"Isaac Sim's Lula motion generation toolkit uses RMPflow as its core, combining goal-reaching, collision-avoidance, joint-limit, and damping RMPs so an arm's end-effector can follow a moving target in real time while avoiding obstacles. As of September 2026, NVIDIA's official documentation recommends new development use the experimental Robot Motion API instead.","related":["Geometric Fabrics","Operational Space Control","Obstacle Avoidance","Artificial Potential Field","NVIDIA Isaac Sim","cuRobo (NVIDIA GPU-accelerated motion planning)"]},{"id":"geometric-fabrics","category":"control","sec":7,"tier":3,"sources":[{"title":"Van Wyk et al., Geometric Fabrics: Generalizing Classical Mechanics to Capture the Physics of Behavior (arXiv 2109.10443)","url":"https://arxiv.org/abs/2109.10443"},{"title":"NVlabs/FABRICS (GitHub)","url":"https://github.com/NVlabs/FABRICS"},{"title":"DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric Fabrics","url":"https://arxiv.org/html/2407.02274v2"}],"as_of":"2026-09","related_ids":["riemannian-motion-policies","dextrah-g","obstacle-avoidance","joint-limits","motion-planning","operational-space-control"],"name":"Geometric Fabrics","alt":"Geometric Fabrics","abbr":"","aliases":["Fabrics"],"one_liner":"An NVIDIA reactive motion-generation framework that synthesizes obstacle avoidance and joint-limit behavior in real time with provable stability.","explanation":"Proposed by Karl Van Wyk, Nathan Ratliff, and colleagues at NVIDIA Research (IEEE RA-L 2022), this is the successor to Riemannian Motion Policies (RMP). Rather than planning a whole trajectory in advance, it computes joint acceleration directly every control cycle: behaviors such as approaching a goal, avoiding obstacles, staying away from joint limits, and holding a posture are each written as a second-order differential equation in their own space, then combined into joint space through Jacobian mappings. Compared to RMP, it uses more general Finsler geometry, allowing behaviors to be shaped more flexibly while still guaranteeing stability. It's often used as a safe action space for reinforcement learning: DextrAH-G has the policy output only low-dimensional targets like a palm pose, leaving collision avoidance and joint constraints to the fabric. NVIDIA has open-sourced a GPU-parallel FABRICS library built on Warp on GitHub.","example":"In DextrAH-G, a Kuka arm plus an Allegro hand with 23 motors total has an RL policy output a palm target and low-dimensional finger commands, while the fabric converts these into joint commands at 60 Hz, automatically avoiding the table and self-collisions.","related":["Riemannian Motion Policies","DextrAH-G","Obstacle Avoidance","Joint Limits","Motion Planning","Operational Space Control"]},{"id":"neural-motion-planning","category":"control","sec":7,"tier":3,"sources":[{"title":"Neural MP: A Generalist Neural Motion Planner (arXiv 2409.05864)","url":"https://arxiv.org/abs/2409.05864"},{"title":"Motion Policy Networks (arXiv 2210.12209)","url":"https://arxiv.org/abs/2210.12209"},{"title":"Motion Planning Networks (arXiv 1806.05767)","url":"https://arxiv.org/abs/1806.05767"}],"as_of":"2024-09","related_ids":["motion-planning","motion-policy-networks","rapidly-exploring-random-tree","imitation-learning","collision-checking","curobo"],"name":"Neural Motion Planning","alt":"神经运动规划","abbr":"","aliases":["Learning-Based Motion Planning"],"one_liner":"Using a neural network trained on huge numbers of planning examples to learn to generate collision-free motion directly.","explanation":"Neural motion planning uses a neural network to perform or assist motion planning: given a scene observation (often a point cloud) and the current and goal configuration, it outputs a collision-free path or the next action. Traditional sampling-based planners (like RRT) and trajectory optimizers restart from scratch on every new problem, sometimes taking minutes on cluttered scenes, and they depend on an accurate geometric model of the scene. Neural approaches instead first use a traditional planner to generate large numbers of expert solutions in simulation, then distill them into a network with imitation learning, so an action comes out with just a forward pass at inference time. Representative work includes 2018's MPNet (which encodes an obstacle point cloud and predicts the next configuration step by step, combinable with RRT* for a success guarantee), NVIDIA and University of Washington's 2022 MπNets (trained on over 500,000 environments and 3 million planning problems, acting directly from a depth camera's point cloud), and CMU's 2024 Neural MP (which procedurally generates huge numbers of scenes to train a general policy, adding a lightweight optimization step at deployment for safety). The difficulty is that the network's raw output carries no collision-free guarantee, so it is usually paired with a collision check or a fallback mechanism.","example":"Neural MP was tested on 64 tasks across 4 real-world environment types, generating arm motion directly from scene point clouds, with success rates 23, 17, and 79 percentage points higher than sampling-based, optimization-based, and learning-based baselines respectively.","related":["Motion Planning","Motion Policy Networks","Rapidly-exploring Random Tree","Imitation Learning","Collision Checking","cuRobo (NVIDIA GPU-accelerated motion planning)"]},{"id":"grasp-planning","category":"control","sec":7,"tier":2,"sources":[{"title":"Data-Driven Grasp Synthesis - A Survey (Bohg et al., IEEE T-RO 2014, arXiv 1309.2660)","url":"https://arxiv.org/abs/1309.2660"},{"title":"Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics (arXiv 1703.09312)","url":"https://arxiv.org/abs/1703.09312"}],"as_of":"","related_ids":["grasping","grasp-pose-detection","pre-grasp-pose","force-closure","grasp-quality-metric","motion-planning"],"name":"Grasp Planning","alt":"抓取规划","abbr":"","aliases":["Grasp Synthesis","Grasp Pose Planning"],"one_liner":"Computing where and how a gripper or dexterous hand should grip an object — position, orientation, and finger configuration.","explanation":"Grasp planning (also called grasp synthesis) answers ‘where and how to grip’: it outputs a 6D grasp pose (position plus orientation) and opening width for a parallel-jaw gripper, or finger joint angles for a dexterous hand. Motion planning then generates a collision-free path, typically moving to a pre-grasp pose before closing the gripper. Early methods were analytic: given a known object model and friction coefficient, they searched for grasps that maximize quality metrics such as force closure (contact forces able to resist an external force from any direction). A 2014 survey by Bohg et al. grouped data-driven methods into three categories based on whether the object had been seen before: known, similar, or novel. Later, deep learning began predicting grasps directly from depth images or point clouds, as in Dex-Net, Contact-GraspNet, and AnyGrasp. End-to-end vision-language-action (VLA) models usually don't treat this as a separate step, but it remains common in industrial bin-picking and modular pipelines.","example":"UC Berkeley's Dex-Net 2.0 trained a grasp-quality network called GQ-CNN on 6.7 million synthetic point clouds paired with analytic grasp metrics; on an ABB YuMi robot it planned a grasp in about 0.8 seconds and reached 93% success on 8 known objects.","related":["Grasping","Grasp Pose Detection","Pre-grasp Pose","Force Closure","Grasp Quality Metric","Motion Planning"]},{"id":"pre-grasp-pose","category":"control","sec":7,"tier":2,"sources":[{"title":"MoveIt Tutorials: Pick and Place","url":"https://moveit.github.io/moveit_tutorials/doc/pick_place/pick_place_tutorial.html"},{"title":"moveit_msgs/Grasp.msg","url":"https://raw.githubusercontent.com/moveit/moveit_msgs/master/msg/Grasp.msg"}],"as_of":"","related_ids":["grasp-planning","grasp-pose-detection","movej-movel","moveit-motion-planning-framework","end-effector-pose","grasping"],"name":"Pre-grasp Pose","alt":"预抓取位姿","abbr":"","aliases":["Approach Pose","Pregrasp Pose"],"one_liner":"An intermediate pose where the gripper pauses near an object, aligned with the grasp direction, just before the actual grasp.","explanation":"Arm grasping usually happens in stages: move to the pre-grasp pose, approach the grasp pose in a straight line along the approach direction, close the gripper, then lift and retreat. The pre-grasp pose is typically obtained by backing off the grasp pose a short distance along the approach direction (usually the direction the gripper faces), with the gripper pre-opened to a suitable width. This separates ‘large-scale travel’ from ‘fine motion near the object’: the first leg can be handled quickly by motion planning, while the second is short and straight, less likely to bump the object or its surroundings. ROS's MoveIt describes this in its Grasp message with pre_grasp_approach (direction and distance) and pre_grasp_posture (hand shape before grasping). In dexterous-hand research, ‘pre-grasp’ also often refers to the finger shape formed just before the hand closes.","example":"In the MoveIt grasping tutorial, a Panda arm first stops about 11.5 cm from the grasp pose (desired_distance of 0.115 m) with the gripper open, approaches and grasps along the x-axis, then retreats about 25 cm upward along z.","related":["Grasp Planning","Grasp Pose Detection","MoveJ / MoveL","MoveIt Motion Planning Framework","End-Effector Pose","Grasping"]},{"id":"finite-state-machine","category":"control","sec":8,"tier":2,"sources":[{"title":"Wikipedia: Finite-state machine","url":"https://en.wikipedia.org/wiki/Finite-state_machine"},{"title":"unitree_sdk2 g1_loco_client.hpp（SetFsmId / Damp / ZeroTorque）","url":"https://github.com/unitreerobotics/unitree_sdk2/blob/main/include/unitree/robot/g1/loco/g1_loco_client.hpp"},{"title":"Behavior Trees in Robotics and AI: An Introduction (arXiv 1709.00084)","url":"https://arxiv.org/abs/1709.00084"}],"as_of":"","related_ids":["behavior-tree","damping-mode","task-planning","hierarchical-control","behaviortree-cpp","unitree-sdk2"],"name":"Finite State Machine","alt":"有限状态机","abbr":"FSM","aliases":["FSM","State Machine"],"one_liner":"A model where a system is always in exactly one of a fixed set of states, switching states according to fixed rules on each event.","explanation":"A finite state machine is a computational model in which a system is, at any moment, in exactly one of a fixed set of states, and moves to a different state according to predetermined rules whenever it receives a particular input or event, with each state corresponding to a fixed set of behaviors. A classic example is a subway turnstile: inserting a coin switches it from 'locked' to 'unlocked,' and someone passing through switches it back to 'locked.' In robotics it's commonly used to manage operating modes and task flow — such as the sequence from power-on zero-torque, to standing, to walking, to damping, or the 'approach → grasp → lift → place' sequence in a pick-and-place routine — with the benefit of clear, easy-to-check logic. Its drawback is that it becomes hard to maintain once the number of states and transitions grows large, which is why complex tasks increasingly use behavior trees instead, since they're more modular and adapt more easily to new situations on the fly.","example":"Unitree's G1 high-level motion interface manages modes directly by numbered states: 0 is zero-torque, 1 is damping, 2 is squatting, 3 is sitting, 4 is standing up, and 500 is start-moving; calling Damp() in the SDK simply switches the state machine to state 1.","related":["Behavior Tree","Damping Mode","Task Planning","Hierarchical Control","BehaviorTree.CPP","Unitree SDK2"]},{"id":"behavior-tree","category":"control","sec":8,"tier":2,"sources":[{"title":"Wikipedia: Behavior tree (artificial intelligence, robotics and control)","url":"https://en.wikipedia.org/wiki/Behavior_tree_(artificial_intelligence,_robotics_and_control)"},{"title":"Colledanchise & Ögren, Behavior Trees in Robotics and AI: An Introduction (arXiv 1709.00084)","url":"https://arxiv.org/abs/1709.00084"},{"title":"ros-navigation/navigation2 README","url":"https://github.com/ros-navigation/navigation2"}],"as_of":"","related_ids":["finite-state-machine","behaviortree-cpp","ros-2-navigation-stack","task-planning","skill-primitive","failure-recovery"],"name":"Behavior Tree","alt":"行为树","abbr":"BT","aliases":["BT","Behaviour Tree"],"one_liner":"A tree-structured way of organizing actions and conditions to decide what a robot should do next.","explanation":"A behavior tree is a way of organizing how a robot or game character switches among multiple tasks. It first became popular in the game industry (Damian Isla's 2005 GDC talk on Halo 2's AI is often credited with popularizing it) before spreading into robotics. A root node periodically sends a 'tick' signal down the tree, and each node returns success, failure, or running. Leaves are action nodes (such as 'grab the cup') and condition nodes (such as 'is the cup in view'); a sequence node runs its children left to right and fails as soon as one fails; a fallback (selector) node tries its children left to right and succeeds as soon as one succeeds, which is well suited to writing 'try A first, and if that fails, do B as a fallback.' Compared with a finite state machine, it doesn't require writing a transition for every pair of states, and adding or removing subtrees is easier. ROS 2's Nav2 navigation framework uses a behavior tree to orchestrate navigation and recovery actions, commonly implemented with the C++ library BehaviorTree.CPP.","example":"A water-delivery robot's root node is a sequence: 'navigate to the table' → 'detect the cup' → a fallback node (try 'grab the cup,' and if that fails, 'reposition and retry the grab') → 'hand it to the user.' If some step returns running, the next tick simply continues that same step.","related":["Finite State Machine","BehaviorTree.CPP","ROS 2 Navigation Stack (Nav2)","Task Planning","Skill Primitive","Failure Recovery"]},{"id":"task-planning","category":"control","sec":8,"tier":2,"sources":[{"title":"Wikipedia: Automated planning and scheduling","url":"https://en.wikipedia.org/wiki/Automated_planning_and_scheduling"},{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)","url":"https://arxiv.org/abs/2204.01691"}],"as_of":"","related_ids":["task-and-motion-planning","planning-domain-definition-language","hierarchical-task-network","symbolic-planning","llm-based-task-planning","saycan"],"name":"Task Planning","alt":"任务规划","abbr":"","aliases":["Automated Planning","AI Planning","High-Level Planning"],"one_liner":"Given the current state and a goal, finding a sequence of high-level action steps that achieves it.","explanation":"Task planning is the application of automated planning from AI to robotics: given an initial state, a goal, and a set of available actions (each with preconditions and effects), find a sequence of actions that, once executed, makes the goal true. Classical approaches describe the problem in STRIPS or PDDL (Planning Domain Definition Language) and hand it to a general-purpose planner to search; hierarchical task networks (HTN) instead break a large task down into subtasks layer by layer. Task planning only decides what to do and in what order — not exactly how the arm should move, which is left to motion planning; combining the two gives task and motion planning. Large language models are now commonly used for task planning — SayCan, for instance, has a language model propose candidate skills, then uses each skill's value function to judge whether it can actually succeed in the current scene, grounding the model's language knowledge in a real robot.","example":"Given the instruction ‘put the coke on the table into the fridge,’ task planning produces: walk to the table → pick up the coke → walk to the fridge → open the fridge door → put in the coke → close the door, with each step then handed to lower-level skills such as navigation and grasping.","related":["Task and Motion Planning","Planning Domain Definition Language","Hierarchical Task Network","Symbolic Planning","LLM-based Task Planning","SayCan"]},{"id":"symbolic-planning","category":"control","sec":8,"tier":3,"sources":[{"title":"Stanford Research Institute Problem Solver (STRIPS) - Wikipedia","url":"https://en.wikipedia.org/wiki/Stanford_Research_Institute_Problem_Solver"},{"title":"Automated planning and scheduling - Wikipedia（经典规划的假设与求解方法）","url":"https://en.wikipedia.org/wiki/Automated_planning_and_scheduling"},{"title":"LLM+P: Empowering Large Language Models with Optimal Planning Proficiency (arXiv 2304.11477)","url":"https://arxiv.org/abs/2304.11477"}],"as_of":"","related_ids":["planning-domain-definition-language","task-planning","task-and-motion-planning","hierarchical-task-network","llm-based-task-planning","pddlstream"],"name":"Symbolic Planning","alt":"符号规划","abbr":"","aliases":["Classical Planning","STRIPS","Automated Planning"],"one_liner":"Planning that describes states and actions as logical symbols, then searches for an action sequence that reaches a goal.","explanation":"Symbolic planning abstracts the world into a set of logical propositions, such as (on cup table) or (handempty), and writes each action as a precondition plus an effect (which facts it adds or deletes). It then searches from the initial state for a sequence of actions that makes all the goal propositions true. The field traces back to STRIPS, developed by Richard Fikes and Nils Nilsson at SRI in 1971 for the Shakey robot; today's common planning description language, PDDL, is built on it, and solving typically uses heuristic-guided state-space search. Classical planning assumes the initial state is known and action outcomes are deterministic, so its results are verifiable and explainable — but turning perceived objects into symbols, and grounding symbolic actions back into continuous motion, is left to task-and-motion planning. The large-language-model era has brought approaches where an LLM translates natural language into PDDL and hands it to a classical planner to solve, such as LLM+P (2023).","example":"A PDDL ‘pick up’ action might be written with parameters ?o (an object) and ?l (a location); precondition (at robot ?l), (on ?o ?l), (handempty); and effect: add (holding ?o), delete (on ?o ?l) and (handempty). From rules like this, a planner can automatically sequence a plan such as ‘walk to the table → pick up the cup → walk to the sink → put it down.’","related":["Planning Domain Definition Language","Task Planning","Task and Motion Planning","Hierarchical Task Network","LLM-based Task Planning","PDDLStream"]},{"id":"planning-domain-definition-language","category":"control","sec":8,"tier":3,"sources":[{"title":"Wikipedia: Planning Domain Definition Language","url":"https://en.wikipedia.org/wiki/Planning_Domain_Definition_Language"},{"title":"LLM+P: Empowering Large Language Models with Optimal Planning Proficiency (arXiv 2304.11477)","url":"https://arxiv.org/abs/2304.11477"},{"title":"PDDLStream: Integrating Symbolic Planners and Blackbox Samplers (arXiv 1802.08705)","url":"https://arxiv.org/abs/1802.08705"}],"as_of":"","related_ids":["task-planning","symbolic-planning","task-and-motion-planning","pddlstream","llm-based-task-planning","hierarchical-task-network"],"name":"Planning Domain Definition Language","alt":"规划领域定义语言","abbr":"PDDL","aliases":["PDDL"],"one_liner":"A standard language for writing down actions' preconditions, effects, and a task's goal, for a general-purpose planner to solve.","explanation":"PDDL is the standard description language of classical AI planning, created in 1998 by Drew McDermott and colleagues for the International Planning Competition (IPC), drawing on earlier planning formalisms such as STRIPS and ADL. It splits a planning problem into two parts: a domain file, which defines predicates (true/false statements describing world state, such as ‘block A is on block B’) and actions (each action's preconditions and the effects of executing it); and a problem file, which gives the specific objects, initial state, and goal. Once written, this is handed to a general-purpose planner to automatically search out a sequence of actions. Later versions added numeric and durative actions (PDDL2.1), continuous processes and events (PDDL+), and preferences and trajectory constraints (PDDL3.0). In robotics it commonly serves as the symbolic interface at the task-planning layer: the task-and-motion-planning framework PDDLStream wires continuous-parameter samplers — for grasp poses, placement positions, and so on — into PDDL as ‘streams’; LLM+P has a large language model translate a natural-language instruction into PDDL, then hands it to a classical planner to compute a correct plan, addressing how easily an LLM goes wrong doing long-horizon planning directly.","example":"‘Put the cup on the table into the cabinet’ can be written in PDDL as: action pick(?obj) has precondition ‘hand empty’ and ‘object reachable,’ with effect ‘holding the object’; further defining open(?door) and place(?obj ?loc), and giving the initial state and the goal ‘cup is in the cabinet,’ a planner outputs ‘open the cabinet door → pick up the cup → place it in the cabinet.’","related":["Task Planning","Symbolic Planning","Task and Motion Planning","PDDLStream","LLM-based Task Planning","Hierarchical Task Network"]},{"id":"hierarchical-task-network","category":"control","sec":8,"tier":3,"sources":[{"title":"Erol, Hendler, Nau: HTN Planning: Complexity and Expressivity (AAAI 1994)","url":"https://www.cs.umd.edu/~nau/papers/erol1994htn.pdf"},{"title":"Nau et al.: SHOP2: An HTN Planning System (JAIR 2003)","url":"https://www.jair.org/index.php/jair/article/view/10362"},{"title":"Höller et al.: HDDL: An Extension to PDDL for Expressing Hierarchical Planning Problems (AAAI 2020)","url":"https://ojs.aaai.org/index.php/AAAI/article/view/6542"}],"as_of":"","related_ids":["task-planning","planning-domain-definition-language","symbolic-planning","llm-based-task-planning","behavior-tree","long-horizon-task"],"name":"Hierarchical Task Network","alt":"分层任务网络","abbr":"HTN","aliases":["HTN","HTN Planning"],"one_liner":"Breaking a large task down, layer by layer, into directly executable actions using human-written ‘decomposition methods.’","explanation":"The hierarchical task network is a form of classical AI planning, with roots in 1970s planners such as NOAH, formally defined by Erol, Hendler, and Nau in 1994. Tasks are split into two kinds: primitive tasks, which can be executed directly, and compound tasks, which need further decomposition. A domain expert writes ‘methods’ for each compound task, specifying under what conditions it breaks into which subtasks and in what order; the planner then repeatedly decomposes the top-level task until only a sequence of primitive actions remains. Compared to PDDL, which just gives a goal and lets the planner search for a combination of actions on its own, HTN leans on human-authored knowledge to search faster and produce more predictable results, at the cost of someone having to write the methods by hand. A representative system is SHOP2. In robotics it's commonly used for high-level decomposition of long-horizon tasks, and often paired with LLM-based task planning and behavior trees.","example":"‘Clear the table’ decomposes into ‘collect the dishes’ and ‘wipe the table’; the method for ‘collect the dishes’ further decomposes into ‘move to the table → identify a bowl → pick up the bowl → put it in the sink,’ repeating while dishes remain — with every bottom-level step corresponding to a skill the robot already has.","related":["Task Planning","Planning Domain Definition Language","Symbolic Planning","LLM-based Task Planning","Behavior Tree","Long-horizon Task"]},{"id":"monte-carlo-tree-search","category":"control","sec":8,"tier":3,"sources":[{"title":"Wikipedia: Monte Carlo tree search","url":"https://en.wikipedia.org/wiki/Monte_Carlo_tree_search"},{"title":"Simultaneous Contact Sequence and Patch Planning for Dynamic Locomotion (arXiv 2508.12928)","url":"https://arxiv.org/abs/2508.12928"}],"as_of":"","related_ids":["task-planning","exploration-vs-exploitation","partially-observable-markov-decision-process","muzero","multi-contact-planning","value-function"],"name":"Monte Carlo Tree Search","alt":"蒙特卡洛树搜索","abbr":"MCTS","aliases":["MCTS"],"one_liner":"A decision algorithm that estimates how good each choice is through large numbers of random simulations, growing a search tree as it goes.","explanation":"Monte Carlo tree search is used to find good decisions in sequential-choice problems. Rémi Coulom coined the name in 2006, and that same year Kocsis and Szepesvári proposed the most commonly used variant, UCT. Each iteration does four steps: selection (descend from the root, picking the currently most-promising branch at each level), expansion (add a new node), simulation (play out randomly, or according to a policy, from the new node to some end state, getting a result), and backpropagation (update the statistics of every node along the path with that result). Selection commonly uses the UCB formula w/n + c·√(ln N / n): w/n is the branch's average score, n how many times it's been tried, N the parent's visit count, and c a knob balancing exploring new branches against exploiting known-good ones. It needs no hand-written position-evaluation function, and can report its current best answer even mid-search. AlphaGo, which beat Lee Sedol in 2016, combined it with neural networks; in robotics it is used for task planning, decision-making under partial observability, and contact-sequence planning.","example":"For a legged robot crossing stepping stones, Dhédin et al. (2025) used MCTS to search, at a discrete level, over which leg, in what order, and onto which foothold region to step, handing each candidate sequence to whole-body trajectory optimization to check feasibility and using the result to update the search tree.","related":["Task Planning","Exploration vs. Exploitation","Partially Observable Markov Decision Process","MuZero","Multi-contact Planning","Value Function"]},{"id":"task-and-motion-planning","category":"control","sec":8,"tier":2,"sources":[{"title":"Integrated Task and Motion Planning (Garrett et al., Annual Review of Control, Robotics, and Autonomous Systems 2021)","url":"https://arxiv.org/abs/2010.01083"},{"title":"PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning","url":"https://arxiv.org/abs/1802.08705"}],"as_of":"","related_ids":["task-planning","motion-planning","planning-domain-definition-language","pddlstream","long-horizon-task","llm-based-task-planning"],"name":"Task and Motion Planning","alt":"任务与运动规划","abbr":"TAMP","aliases":["TAMP","Integrated Task and Motion Planning"],"one_liner":"Jointly solving both which actions to take and exactly how each action should move, at the same time.","explanation":"Task and motion planning (TAMP) decides two kinds of problems together: the discrete task level (which object to move, where to put it, in what order) and the continuous motion level (grasp poses, placement positions, collision-free arm trajectories). A 2021 survey by Garrett, Kaelbling, and Lozano-Pérez et al. notes that TAMP spans discrete task planning, discrete-continuous mathematical programming, and continuous motion planning, and that none of these fields alone solves it well. The reason is that the two levels depend on each other: a plan that is symbolically valid may be physically unreachable or blocked, while pure motion planning can't handle long sequences of decisions. A representative tool, PDDLStream, extends PDDL with ‘streams’ that treat inverse kinematics, grasp sampling, and collision checking as black-box samplers to call during search.","example":"Asking a robot to put a green block into a cabinet whose door is blocked by a red block: pure task planning doesn't know the red block is in the way, and pure motion planning can't find a feasible path; TAMP produces a plan that first moves the red block aside, while simultaneously computing where to move it, how to grasp it, and how the arm should travel.","related":["Task Planning","Motion Planning","Planning Domain Definition Language","PDDLStream","Long-horizon Task","LLM-based Task Planning"]},{"id":"llm-based-task-planning","category":"control","sec":8,"tier":2,"sources":[{"title":"SayCan 项目主页 (Do As I Can, Not As I Say)","url":"https://say-can.github.io/"},{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (arXiv 2204.01691)","url":"https://arxiv.org/abs/2204.01691"},{"title":"LLM+P: Empowering Large Language Models with Optimal Planning Proficiency (arXiv 2304.11477)","url":"https://arxiv.org/abs/2304.11477"}],"as_of":"2025-09","related_ids":["task-planning","saycan","code-as-policies","planning-domain-definition-language","gemini-robotics-er","dual-system-architecture"],"name":"LLM-based Task Planning","alt":"大模型任务规划","abbr":"","aliases":["LLM Planning","VLM Planning"],"one_liner":"Using a large language or vision-language model to break a high-level instruction into a sequence of sub-tasks a robot can execute.","explanation":"Task planning decides what to do first and what comes next. Traditionally this required hand-written formal descriptions in PDDL (Planning Domain Definition Language), which had to be rewritten for every new scenario. Starting around 2022, researchers began having large language models generate steps directly from an instruction: Google's SayCan has an LLM score candidate skills, then multiplies that score by a value function estimating how likely each skill is to succeed right now, to pick the next step; Code as Policies has the model write code that calls perception and control APIs; Inner Monologue feeds execution feedback back into the prompt for closed-loop replanning. LLMs readily produce steps that sound plausible but can't actually be executed, so LLM+P instead has the model translate only the natural-language goal into PDDL, which a classical planner then solves. Today, vision-language models (VLMs) are often used to plan directly from images — Gemini Robotics-ER, for instance — handing off each sub-task to a VLA model to execute.","example":"Given ‘I spilled my coke, can you bring me something to clean it up?,’ SayCan on a mobile manipulator selected, in sequence: find a sponge, pick up the sponge, bring it to you, done; the project page reported an 84% planning success rate and 74% execution success rate across 101 instructions.","related":["Task Planning","SayCan","Code as Policies","Planning Domain Definition Language","Gemini Robotics-ER","Dual-System Architecture (System 1 / System 2)"]},{"id":"shared-autonomy","category":"control","sec":8,"tier":3,"sources":[{"title":"Shared Autonomy via Hindsight Optimization (Javdani, Srinivasa, Bagnell, arXiv 1503.07619)","url":"https://arxiv.org/abs/1503.07619"},{"title":"Shared Autonomy via Deep Reinforcement Learning (Reddy, Dragan, Levine, RSS 2018, arXiv 1802.01744)","url":"https://arxiv.org/abs/1802.01744"}],"as_of":"","related_ids":["teleoperation","human-in-the-loop","human-robot-collaboration","human-robot-interaction","intent-understanding","partially-observable-markov-decision-process"],"name":"Shared Autonomy","alt":"共享自主","abbr":"","aliases":["Shared Control"],"one_liner":"A control scheme where the robot infers a human operator's intent and blends its own autonomous action with the person's input.","explanation":"Shared autonomy sits between pure teleoperation and full autonomy: a person gives rough, noisy commands through a joystick, brain-computer interface, or similar, while the robot infers what the person is trying to achieve and blends that inferred goal with the person's own input — commonly by weighting the mix according to how confident it is about the person's intent, an approach called arbitration or policy blending. It was first used for assistive robotic arms, smart wheelchairs, and teleoperation, aiming to reduce the operator's workload without taking away their control. Siddhartha Javdani and colleagues at CMU modeled it in 2015 as a partially observable Markov decision process (POMDP) with an unknown goal, using inverse optimal control to infer a distribution over goals from the person's history of inputs; users completed tasks faster with less input, though some reported feeling they had lost a sense of control. Reddy, Dragan, and Levine used deep reinforcement learning in 2018 to remove the need for a predefined set of possible goals.","example":"In an experiment by Reddy and colleagues, a person plays the Lunar Lander game while an assistive agent first filters out actions whose Q-value falls below a threshold, then picks whichever of the remaining actions is closest to the human's input — helping the person land more steadily without ever knowing the actual target landing site.","related":["Teleoperation","Human-in-the-Loop","Human-Robot Collaboration","Human-Robot Interaction","Intent Understanding","Partially Observable Markov Decision Process"]},{"id":"emergency-stop","category":"control","sec":9,"tier":1,"sources":[{"title":"Wikipedia (de): Not-Halt（停止类别 0/1/2、ISO 13850、急停按钮形态）","url":"https://de.wikipedia.org/wiki/Not-Halt"},{"title":"Wikipedia: Kill switch（emergency stop, ISO 13850）","url":"https://en.wikipedia.org/wiki/Kill_switch"}],"as_of":"","related_ids":["protective-stop","functional-safety","safe-torque-off","damping-mode","watchdog","collision-detection"],"name":"Emergency Stop","alt":"急停","abbr":"E-stop","aliases":["E-Stop","Stop Category 0/1/2"],"one_liner":"A safety function that stops a machine with a single action in an emergency — typically a red mushroom button on yellow.","explanation":"An emergency stop is a mandatory safety feature on any machine; the international standard ISO 13850 requires that anyone present be able to stop the machine with a single action, without having to think about it. The typical E-stop button is a red mushroom-shaped head on a yellow background; pressing it mechanically latches in place, and someone has to manually reset it before the machine can restart. IEC 60204-1 classifies stopping into three categories: category 0 cuts power to the drives immediately; category 1 brings the machine to a controlled stop first, then cuts power; category 2 brings it to a controlled stop while keeping power on. Only categories 0 or 1 are allowed for an emergency stop. Which one to use depends on whether cutting power immediately would actually be more dangerous: legged and humanoid robots collapse the instant power is cut, so they're often tested on a safety harness, and beyond the E-stop, a softer stop such as a damping mode is also provided. An E-stop is the last line of defense — it can't substitute for other safety measures like collision detection or speed limiting.","example":"A collaborative arm's teach pendant usually has a built-in E-stop button; during real-robot VLA experiments, the operator typically also keeps a separate, standalone E-stop within reach, ready to hit it the moment the policy does something abnormal.","related":["Protective Stop","Functional Safety","Safe Torque Off (STO)","Damping Mode","Watchdog","Collision Detection (Robot Safety)"]},{"id":"protective-stop","category":"control","sec":9,"tier":3,"sources":[{"title":"Universal Robots: Safety FAQ（stop categories, protective stop, monitored standstill）","url":"https://www.universal-robots.com/articles/ur/safety/safety-faq/"},{"title":"OSHA Technical Manual, Section IV Chapter 4: Industrial Robot Systems and Industrial Robot System Safety","url":"https://www.osha.gov/otm/section-4-safety-hazards/chapter-4"}],"as_of":"2026-05","related_ids":["emergency-stop","functional-safety","collision-detection","collision-reaction","power-and-force-limiting","speed-and-separation-monitoring"],"name":"Protective Stop","alt":"保护性停止","abbr":"","aliases":["Safety Stop"],"one_liner":"A safety function that brings a robot to an automatic, controlled halt as soon as it detects an overload or hazard.","explanation":"Protective stop is a basic safety function defined in the ISO 10218 industrial and collaborative robot safety standards: when collision detection fires, force or speed exceeds a configured safety limit, a safety gate opens, or a light curtain is broken, the controller brings the robot to an orderly stop. It differs from an emergency stop (E-stop): an E-stop is a person hitting the red button and is reserved for real emergencies — Universal Robots' E-stop first performs a controlled stop and then cuts power to the drives — whereas a protective stop is usually triggered automatically by the system itself. IEC 60204-1 defines three stop categories: category 0 removes power immediately, category 1 performs a controlled stop before removing power, and category 2 stops and then holds position without removing power, which resumes fastest. A related concept is 'safety-rated monitored stop' in collaborative mode, where the robot halts and waits when a person enters the collaborative workspace and resumes automatically once they leave; per Universal Robots' documentation, ISO 10218-1:2025 has renamed this ‘monitored standstill’.","example":"A UR collaborative arm running into an obstacle will automatically trigger a protective stop the moment any force or torque parameter exceeds its safety settings; an operator must clear the cause and confirm before the robot can resume.","related":["Emergency Stop","Functional Safety","Collision Detection (Robot Safety)","Collision Reaction","Power and Force Limiting","Speed and Separation Monitoring"]},{"id":"damping-mode","category":"control","sec":9,"tier":2,"sources":[{"title":"unitree_rl_gym 真机部署说明（中文 README）","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/deploy/deploy_real/README.zh.md"},{"title":"unitree_rl_gym command_helper.py（create_damping_cmd）","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/deploy/deploy_real/common/command_helper.py"},{"title":"unitree_sdk2 g1_loco_client.hpp（Damp / ZeroTorque 接口）","url":"https://github.com/unitreerobotics/unitree_sdk2/blob/main/include/unitree/robot/g1/loco/g1_loco_client.hpp"}],"as_of":"2026-09","related_ids":["damping","stiffness-and-damping-gains","mit-mode","emergency-stop","protective-stop","unitree-sdk2"],"name":"Damping Mode","alt":"阻尼模式","abbr":"","aliases":["Damp Mode","Protective Damping State"],"one_liner":"A joint stops chasing its target position and only resists motion with damping, letting the robot settle down gently.","explanation":"Damping mode is a protective state for joint motors, provided directly by Unitree's SDK and remote control, among others. A joint motor commonly computes its output torque as τ = kp(q_target − q) + kd(q̇_target − q̇) + τff, where q is joint angle, q̇ is angular velocity, kp and kd are the stiffness and damping gains, and τff is a feedforward torque. Damping mode sets kp, τff, and the target velocity all to zero while keeping only kd, so τ = −kd·q̇: the joint applies no force while at rest, and resists as soon as it starts moving. Compared with zero-torque mode (where kd is also zero and the joint goes completely limp), damping mode lets the robot settle down slowly instead of collapsing abruptly, which is why it's commonly used as the safety fallback when a program exits or something goes wrong.","example":"In Unitree's unitree_rl_gym real-robot deployment workflow, pressing the remote's select button during locomotion control puts the robot into damping mode, where it sinks down and the program exits; in code, the damping command is simply kp = 0 and kd = 8 with zero feedforward torque for every motor, i.e. τ = −8·q̇.","related":["Damping","Stiffness and Damping Gains","MIT Mode","Emergency Stop","Protective Stop","Unitree SDK2"]},{"id":"soft-limits","category":"control","sec":9,"tier":2,"sources":[{"title":"ros2_control joint_limits.hpp：SoftJointLimits（来自 URDF safety_controller）","url":"https://github.com/ros-controls/ros2_control/blob/master/joint_limits/include/joint_limits/joint_limits.hpp"},{"title":"legged_gym legged_robot_config.py：soft_dof_pos_limit","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["joint-limits","torque-limiting","protective-stop","workspace","unified-robot-description-format","functional-safety"],"name":"Soft Limits (Software Joint Limits)","alt":"软限位","abbr":"","aliases":["Software Limits","Virtual Wall","Workspace Limits"],"one_liner":"A software-defined motion boundary set more conservatively than the mechanical limit, stopping the robot before it reaches the hard stop.","explanation":"Soft limits are motion boundaries set by control software, positioned inside the joint's mechanical limits (hard stops, limit blocks) so the robot is stopped by software before it ever reaches the mechanical stop. At the joint level, URDF lets you write soft_lower_limit / soft_upper_limit in a safety_controller tag; ros2_control's documentation explains that once a joint reaches this boundary, a safety controller starts constraining its position, with k_position and k_velocity determining how quickly the allowed velocity and torque get tightened. At the Cartesian level, many arm controllers support defining planes or regions the end-effector may not cross, commonly called a virtual wall. In reinforcement-learning locomotion, legged_gym uses soft_dof_pos_limit to proportionally shrink the URDF's joint range and gives a negative reward for exceeding it, teaching the policy to stay away from the limits.","example":"A joint with a mechanical range of ±170° might have its soft limit set to ±165°; during teleoperation, if the operator pushes the controller all the way, the joint stops near 165° rather than slamming into its mechanical stop.","related":["Joint Limits","Torque Limiting (Saturation)","Protective Stop","Workspace","Unified Robot Description Format","Functional Safety"]},{"id":"torque-limiting","category":"control","sec":9,"tier":2,"sources":[{"title":"legged_gym: legged_robot.py（_compute_torques 中的 torch.clip）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"Isaac Lab Docs: Actuators（ideal PD actuator clipping）","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/actuators.html"},{"title":"libfranka rate_limiting.h（kMaxTorqueRate）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/rate_limiting.h"}],"as_of":"","related_ids":["peak-torque","rated-torque","torque-speed-curve","integral-windup-anti-windup","soft-limits","torque-control"],"name":"Torque Limiting (Saturation)","alt":"力矩限幅","abbr":"","aliases":["Torque Saturation","Output Clipping","Torque Clipping","Actuator Saturation"],"one_liner":"Clipping the controller's computed torque to what the motor can actually deliver, discarding whatever exceeds that.","explanation":"Torque limiting is the final clip in the control chain: no matter how much torque the higher levels compute, it gets clamped to [−τmax, τmax] before reaching the motor. τmax is set by the motor's peak torque, the drive's current limit, and the gearbox's strength, and in practice it also drops as speed increases. It protects the motor from burning out and the gearbox from breaking, and it also stops a simulated robot from applying ‘unlimited’ force. The side effect is that once saturated, the controller thinks it's sending one thing while the actual output is another: a PID's integral term keeps accumulating (integral windup), producing a large overshoot once it comes out of saturation, which requires anti-windup handling; the limits used in simulation training should also match the real hardware, or a policy's learned behavior won't transfer correctly. Besides the magnitude, some systems also cap how fast torque is allowed to change.","example":"legged_gym's _compute_torques function computes torque from the policy's target joint angle via the PD formula, then clips it with torch.clip to ±torque_limits in its final line before passing it to the simulator; libfranka additionally limits the rate of torque change, to roughly 1000 N·m/s per joint.","related":["Peak Torque","Rated Torque","Torque-Speed Curve","Integral Windup / Anti-windup","Soft Limits (Software Joint Limits)","Torque Control"]},{"id":"watchdog","category":"control","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Watchdog timer","url":"https://en.wikipedia.org/wiki/Watchdog_timer"},{"title":"ur_rtde: rtde_control_interface.h（setWatchdog / kickWatchdog）","url":"https://gitlab.com/sdurobotics/ur_rtde/-/raw/master/include/ur_rtde/rtde_control_interface.h"},{"title":"ros2_controllers: diff_drive_controller（cmd_vel_timeout）","url":"https://control.ros.org/rolling/doc/ros2_controllers/diff_drive_controller/doc/userdoc.html"}],"as_of":"","related_ids":["emergency-stop","protective-stop","damping-mode","real-time-control","control-latency","functional-safety"],"name":"Watchdog","alt":"看门狗","abbr":"","aliases":["Watchdog Timer","WDT","Communication Timeout Protection"],"one_liner":"A timer that requires a program to check in periodically and forces a safe state if it stops responding in time.","explanation":"A watchdog was originally a hardware or software timer in embedded systems: while a program runs normally, it must periodically ‘feed the dog’ (kick it, resetting the timer); if the program hangs or crashes and misses a feeding, the timer expires and triggers a corrective action, usually putting the output into a safe state (motors off, high voltage cut) before restarting. In robotics its most common use is as communication-timeout protection: the host computer or policy process must send commands at no less than some minimum frequency, and a timeout triggers braking, a damping mode, or a protective stop — preventing the robot from continuing on its last command if the network drops or the process crashes. For example, the ur_rtde library provides setWatchdog and kickWatchdog interfaces that, by default, require updates of at least 10 Hz or shut down control; ROS 2's differential-drive controller, diff_drive_controller, has a cmd_vel_timeout parameter that by default treats a velocity command as stale if no new one arrives within 0.5 seconds. Deploying a learned policy on a real robot generally requires a watchdog configured at this low level.","example":"A mobile base is teleoperated over Wi-Fi from a laptop, and the signal suddenly drops. The base controller hasn't received a new cmd_vel in 0.5 seconds, so the old command is judged stale and the base stops — rather than continuing to drive forward at its last commanded speed.","related":["Emergency Stop","Protective Stop","Damping Mode","Real-Time Control","Control Latency","Functional Safety"]},{"id":"collision-detection","category":"control","sec":9,"tier":2,"sources":[{"title":"Haddadin, De Luca, Albu-Schäffer, Robot Collisions: A Survey on Detection, Isolation, and Identification (IEEE T-RO 2017)","url":"https://doi.org/10.1109/TRO.2017.2723903"},{"title":"libfranka robot.h（setCollisionBehavior / automaticErrorRecovery）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"}],"as_of":"","related_ids":["generalized-momentum-observer","collision-reaction","power-and-force-limiting","protective-stop","joint-torque-sensor","iso-ts-15066-robots-and-robotic-devices-collaborative-robots"],"name":"Collision Detection (Robot Safety)","alt":"碰撞检测（本体安全）","abbr":"","aliases":["Collision Protection","Sensorless Collision Detection"],"one_liner":"Detecting, in real time while the robot moves, that it has hit a person or object, and immediately stopping or yielding.","explanation":"Collision detection, in this safety sense, is a real-time robot function: the moment the body unexpectedly hits a person or object, the controller must notice as fast as possible and respond, such as stopping, backing off, or switching into a compliant mode. Most collaborative robots don't rely on skin-like sensors, using only proprioception instead: a dynamics model predicts the joint torque normal motion should require, that prediction is compared against the actual torque measured from motor current or joint torque sensors, and a difference beyond some threshold is flagged as a collision. The generalized momentum observer used by De Luca, Haddadin, and colleagues is the classic method; their 2017 survey in IEEE Transactions on Robotics breaks the process into detection, isolation (which link was hit), and identification (how large the impact force was). A manufacturer's 'collision sensitivity' setting is exactly this threshold: set it too low and the robot stops on false alarms; set it too high and a real collision hits harder before it's caught. Franka's libfranka, for example, lets each joint have separate 'contact' and 'collision' thresholds, and the robot stops moving once the collision threshold is exceeded.","example":"A collaborative arm's elbow bumps into a nearby worker's shoulder while carrying something; the joint's measured torque spikes above the model's prediction and crosses the threshold, triggering a stop — the worker feels only a gentle nudge — after which the error has to be cleared before operation can resume.","related":["Generalized Momentum Observer","Collision Reaction","Power and Force Limiting","Protective Stop","Joint Torque Sensor","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots"]},{"id":"generalized-momentum-observer","category":"control","sec":9,"tier":3,"sources":[{"title":"Haddadin, De Luca, Albu-Schäffer, Robot Collisions: A Survey on Detection, Isolation, and Identification (IEEE T-RO 2017)","url":"https://portal.fis.tum.de/en/publications/robot-collisions-a-survey-on-detection-isolation-and-identificati"},{"title":"Collision detection and external force estimation for robot manipulators using a composite momentum observer (AIMS Electronics and Electrical Engineering, 2024)","url":"https://www.aimspress.com/article/doi/10.3934/electreng.2024011?viewType=HTML"}],"as_of":"","related_ids":["collision-detection","collision-reaction","disturbance-observer","sensorless-force-estimation","mass-matrix","friction-compensation"],"name":"Generalized Momentum Observer","alt":"动量观测器","abbr":"","aliases":["Momentum Observer","GMO"],"one_liner":"Estimating the external torque acting on a robot using only joint position, velocity, and motor torque — no force sensor needed.","explanation":"Proposed by Alessandro De Luca and colleagues around 2003 (originally for actuator-fault detection), and used for collision detection on the DLR-III lightweight arm in 2006, this is the classic method letting collaborative robots sense collisions without a force sensor. It tracks the generalized momentum p = M(q)q̇ (M the mass matrix, q̇ joint velocity), comparing the momentum change predicted by the dynamics model against what's actually measured to get a residual r, satisfying ṙ = K(τ_ext − r): r is exactly the estimate of external torque τ_ext after first-order low-pass filtering, with a larger gain K responding faster but more sensitive to noise. It needs no joint acceleration measurement and no inversion of the mass matrix. Once the residual exceeds a threshold, a collision is declared and a collision reaction triggered; because modeling errors like friction leak into the residual too, it's commonly paired with friction compensation.","example":"An arm in motion is blocked by a person's hand; the momentum residual on the first few joints quickly exceeds the threshold, and the controller declares a collision, halting the original trajectory and switching to a compliant or yielding mode.","related":["Collision Detection (Robot Safety)","Collision Reaction","Disturbance Observer","Sensorless Force Estimation","Mass Matrix","Friction Compensation"]},{"id":"collision-reaction","category":"control","sec":9,"tier":3,"sources":[{"title":"Haddadin, De Luca, Albu-Schäffer. Robot Collisions: A Survey on Detection, Isolation, and Identification. IEEE T-RO, 2017","url":"https://doi.org/10.1109/TRO.2017.2723903"},{"title":"De Luca et al. Collision Detection and Safe Reaction with the DLR-III Lightweight Manipulator Arm. IROS 2006","url":"https://doi.org/10.1109/IROS.2006.282053"},{"title":"libfranka robot.h（setCollisionBehavior / automaticErrorRecovery）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot.h"}],"as_of":"","related_ids":["collision-detection","generalized-momentum-observer","protective-stop","power-and-force-limiting","damping-mode","physical-human-robot-interaction"],"name":"Collision Reaction","alt":"碰撞反应","abbr":"","aliases":["Post-Collision Response Strategy","Collision Reflex"],"one_liner":"The strategy a robot follows after detecting a collision — deciding whether to stop, go soft, or yield.","explanation":"Collision reaction is the last step in a robot's collision-handling pipeline. A 2017 survey by Haddadin, De Luca, et al. places it at the end of a chain: first detect whether a collision has happened, then localize which link was hit and estimate the force's magnitude and direction, judge whether it was accidental or intentional contact, and only then decide how to react. Common reactions include: stopping immediately (a protective stop); switching to a torque mode that only compensates gravity, making the arm soft and easy to push away; actively yielding in the direction of the collision force; or switching to impedance or admittance control to continue compliantly. De Luca et al. (2006) demonstrated collision detection on the DLR-III lightweight arm using a generalized-momentum-based method that relies only on the robot's own sensors, and that also gives the direction of the collision for use by different reaction strategies. If a person is pinned between the robot and a table, simply stopping doesn't relieve the pressure, which is why yielding and going soft matter especially in collaborative settings.","example":"Franka arms let you use setCollisionBehavior to set separate ‘contact’ and ‘collision’ thresholds for each joint and each end-effector direction: a force between the two thresholds is just logged as contact, while exceeding the collision threshold stops the robot immediately, requiring a call to automaticErrorRecovery to reset before it can continue.","related":["Collision Detection (Robot Safety)","Generalized Momentum Observer","Protective Stop","Power and Force Limiting","Damping Mode","Physical Human-Robot Interaction"]},{"id":"power-and-force-limiting","category":"control","sec":9,"tier":3,"sources":[{"title":"OSHA Technical Manual, Section IV Chapter 4: Industrial Robot Systems and Industrial Robot System Safety","url":"https://www.osha.gov/otm/section-4-safety-hazards/chapter-4"},{"title":"3D Collision-Force-Map for Safe Human-Robot Collaboration (arXiv:2009.01036, ICRA 2021)","url":"https://arxiv.org/abs/2009.01036"},{"title":"Universal Robots: Safety FAQ","url":"https://www.universal-robots.com/articles/ur/safety/safety-faq/"}],"as_of":"2026-05","related_ids":["collaborative-robot","iso-ts-15066-robots-and-robotic-devices-collaborative-robots","speed-and-separation-monitoring","protective-stop","collision-detection","physical-human-robot-interaction"],"name":"Power and Force Limiting","alt":"功率与力限制","abbr":"PFL","aliases":["PFL","PFL Collaborative Mode"],"one_liner":"A collaborative-safety mode that lets a robot touch a person but caps contact force and pressure below injury thresholds.","explanation":"PFL is one of several collaborative operating modes defined by the ISO 10218 series and ISO/TS 15066 (the others are safety-rated monitored stop, hand guiding, and speed and separation monitoring), and it is the main way collaborative robots work alongside people without safety fencing. Rather than avoiding contact, it limits the consequences of contact: TS 15066 specifies allowable force and pressure by body region, distinguishing transient contact (the person can be pushed away) from quasi-static contact (the body part is pinned), which gets a lower threshold. In practice this relies on a lightweight, slow-moving robot body, or on safety functions such as joint torque sensors that slow or stop the robot when limits are exceeded. TS 15066 gives a simplified formula for allowed speed: v ≤ F_max/√k · √(1/m_R + 1/m_H), where F_max is the maximum allowed force for that body region, k is its equivalent stiffness, and m_R and m_H are the equivalent masses of the robot and the human body part. According to Universal Robots' own materials, the 2025 edition of ISO 10218 has absorbed most of TS 15066's content.","example":"Under TS 15066, a pinched back of the hand (quasi-static contact) allows up to 140 N, while unconstrained transient contact allows up to 280 N. A 2021 ICRA paper used this formula to calculate that, where pinching risk exists, a UR10e end-effector's speed must be capped at about 0.13 m/s.","related":["Collaborative Robot","ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Speed and Separation Monitoring","Protective Stop","Collision Detection (Robot Safety)","Physical Human-Robot Interaction"]},{"id":"speed-and-separation-monitoring","category":"control","sec":9,"tier":3,"sources":[{"title":"Implementing Speed and Separation Monitoring in Collaborative Robot Workcells (Marvel & Norcross, NIST, Robot Comput Integr Manuf 2017)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5117641/"},{"title":"NIST 出版物页面：Implementing Speed and Separation Monitoring in Collaborative Robot Workcells","url":"https://www.nist.gov/publications/implementing-speed-and-separation-monitoring-collaborative-robot-workcells"}],"as_of":"","related_ids":["iso-ts-15066-robots-and-robotic-devices-collaborative-robots","power-and-force-limiting","collaborative-robot","protective-stop","safety-laser-scanner-safety-light-curtain","human-robot-collaboration"],"name":"Speed and Separation Monitoring","alt":"速度与分离监控","abbr":"SSM","aliases":["SSM"],"one_liner":"A collaborative-robot safeguard that tracks human-robot distance in real time and slows or stops the robot once it gets too close.","explanation":"Speed and separation monitoring is one of the collaborative operating modes defined by ISO 10218 and ISO/TS 15066 (the others are safety-rated monitored stop, hand guiding, and power and force limiting). External sensors — safety laser scanners, 3D cameras, and the like — continuously track both the person and the robot, and the system computes a protective separation distance in real time: S = S_h + S_r + S_s + C + Z_S + Z_R. Here S_h is the distance the person covers during the robot's reaction and braking time (the person's approach speed toward the robot is usually taken as 1.6 m/s), S_r and S_s are the distances the robot itself travels during its reaction time and while braking, C is an allowance for how far a person's hand might reach in, and Z_S and Z_R are the position-measurement uncertainties for the person and the robot. If the actual distance drops below S, a safety-rated monitored stop is triggered — the slower the robot moves, the smaller S becomes, letting the person approach closer. The difference from power and force limiting is that PFL tolerates limited contact, while SSM aims to avoid contact altogether.","example":"Suppose the robot's reaction time is 0.1 s and its braking time is 0.3 s, and a person approaches at 1.6 m/s. The S_h term alone is 1.6 × (0.1 + 0.3) = 0.64 m; adding the robot's own travel distance, the hand-reach allowance, and measurement error gives the minimum distance that must be maintained. Once the robot slows down, it brakes faster and travels less, so the person can stand closer.","related":["ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Power and Force Limiting","Collaborative Robot","Protective Stop","Safety Laser Scanner / Safety Light Curtain","Human-Robot Collaboration"]},{"id":"functional-safety","category":"control","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Functional safety","url":"https://en.wikipedia.org/wiki/Functional_safety"},{"title":"Wikipedia: IEC 61508","url":"https://en.wikipedia.org/wiki/IEC_61508"},{"title":"Wikipedia: ISO 13849","url":"https://en.wikipedia.org/wiki/ISO_13849"}],"as_of":"2023","related_ids":["emergency-stop","protective-stop","safe-torque-off","iso-13849-performance-level-safety-integrity-level","iso-10218-1-2-2025-robotics-safety-requirements","embodied-safety"],"name":"Functional Safety","alt":"功能安全","abbr":"","aliases":[],"one_liner":"Relying on automated protective functions to bring equipment to a safe state after a fault, with the reliability of that response quantified.","explanation":"Functional safety is the part of overall safety that depends on automated protective functions acting correctly — for example, emergency stop, safe torque off, and speed monitoring. Its overarching standard is IEC 61508 (first edition 1998–2000, second edition 2010), which expresses reliability with Safety Integrity Levels SIL 1–4 — the higher the level, the lower the allowed probability of dangerous failure; the machinery-specific IEC 62061 and automotive ISO 26262 both derive from it. Robotics more often cites ISO 13849-1 (currently the 2023 fourth edition), expressed with Performance Levels PL a–e and architecture Categories B, 1–4, where Categories 3 and 4 require redundant channels. The process runs hazard analysis → set a target level for each safety function → prove that level is met through redundancy, self-diagnosis, and verification. It's about whether hardware and the control system reliably stop when something goes wrong — a different concern from embodied safety, which is about a large model ‘doing the wrong thing.’","example":"Many collaborative arms design their emergency-stop and protective-stop functions to ISO 13849-1's PL d, Category 3: two independent channels cross-check each other, and either channel failing can still cut motor torque.","related":["Emergency Stop","Protective Stop","Safe Torque Off (STO)","ISO 13849 Performance Level (PL a–e, Category B–4) / Safety Integrity Level (SIL, IEC 62061/61508)","ISO 10218-1/-2:2025 Robotics — Safety Requirements","Embodied Safety"]},{"id":"iso-13849-performance-level-safety-integrity-level","category":"control","sec":9,"tier":3,"sources":[{"title":"ISO 13849-1:2023 Safety of machinery — Safety-related parts of control systems — Part 1","url":"https://www.iso.org/standard/73481.html"},{"title":"Spilma: ISO 13849 and IEC 62061: machinery functional safety","url":"https://www.spilma.com/en/guides/iec-62061-iso-13849-machinery-functional-safety"},{"title":"IBF Solutions: New standards for industrial robots EN ISO 10218-1 and -2","url":"https://www.ibf-solutions.com/en/seminars-and-news/news/new-standards-for-industrial-robots-en-iso-10218-1-and-2"}],"as_of":"2025-02","related_ids":["functional-safety","safe-torque-off","emergency-stop","protective-stop","iso-10218-1-2-2025-robotics-safety-requirements","iso-25785-1"],"name":"ISO 13849 Performance Level (PL a–e, Category B–4) / Safety Integrity Level (SIL, IEC 62061/61508)","alt":"ISO 13849 性能等级 PL（安全完整性等级 SIL）","abbr":"PL / SIL","aliases":["PL","SIL","Performance Level","Safety Integrity Level","ISO 13849-1:2023","IEC 62061","IEC 61508","PL d","SIL 2"],"one_liner":"Rating scales for how reliably a machine's safety functions work: ISO 13849's PL a–e, and the IEC family's SIL.","explanation":"Functional safety asks whether safety functions such as emergency stop, protective stop, and speed limiting can be relied on to actually trigger when needed. ISO 13849-1 (current 2023 edition) measures this with a Performance Level, rising from a to e, corresponding to a per-hour probability of dangerous failure PFHd — PL d, for example, is 10⁻⁷ to 10⁻⁶. A required level PLr is first set from injury severity, exposure frequency, and whether the hazard is avoidable; the achieved PL is then computed from the architecture Category (B, 1, 2, 3, 4 — Categories 3 and 4 require that a single fault not lose the safety function), the mean time to dangerous failure (MTTFd) of components, and diagnostic coverage (DC), and must be no lower than PLr. The other major scale is IEC 61508's SIL 1–4, whose machinery-specific version IEC 62061 uses SIL 1–3; the two roughly correspond by PFHd: PL b/c ≈ SIL 1, PL d ≈ SIL 2, PL e ≈ SIL 3.","example":"The 2011 edition of ISO 10218-1 required all safety-related control functions on industrial robots to uniformly reach PL d, Category 3; the 2025 edition instead assigns a default level per function — a protective stop, for example, defaults to PL d or SIL 2, adjustable if the risk assessment allows.","related":["Functional Safety","Safe Torque Off (STO)","Emergency Stop","Protective Stop","ISO 10218-1/-2:2025 Robotics — Safety Requirements","ISO 25785-1 (Safety Requirements for Industrial Mobile Robots with Actively Controlled Stability — Part 1: Robots)"]},{"id":"iso-10218-1-2-2025-robotics-safety-requirements","category":"control","sec":9,"tier":3,"sources":[{"title":"ISO 10218-1:2025 Robotics — Safety requirements — Part 1: Industrial robots","url":"https://www.iso.org/standard/73933.html"},{"title":"ISO 10218-2:2025 Part 2: Industrial robot applications and robot cells","url":"https://www.iso.org/standard/73934.html"},{"title":"IBF Solutions: New standards for industrial robots EN ISO 10218-1 and -2","url":"https://www.ibf-solutions.com/en/seminars-and-news/news/new-standards-for-industrial-robots-en-iso-10218-1-and-2"}],"as_of":"2025-02","related_ids":["iso-ts-15066-robots-and-robotic-devices-collaborative-robots","power-and-force-limiting","speed-and-separation-monitoring","collaborative-robot","industrial-robot","iso-13849-performance-level-safety-integrity-level"],"name":"ISO 10218-1/-2:2025 Robotics — Safety Requirements","alt":"ISO 10218 工业机器人安全标准（2025 版）","abbr":"","aliases":["ISO 10218:2025","ISO 10218-1:2025","ISO 10218-2:2025","ANSI/A3 R15.06-2025"],"one_liner":"The core international safety standard for industrial robot design and integration, revised with a new edition in 2025.","explanation":"ISO 10218 is developed by ISO/TC 299, the robotics technical committee; Part 1 covers the design of the industrial robot itself, and Part 2 covers integrating a robot into an application and a work cell. The edition published in February 2025 replaces the 2011 version, with major changes: the previously separate collaborative-robot specification ISO/TS 15066 is folded in, with human contact force and pressure limits moved into Part 2; the object of analysis shifts to the ‘robot application,’ including the workpiece, program, and surrounding equipment; functional safety no longer uniformly demands PL d, Category 3, but instead assigns a default level to each safety function; robots are now classified into Type 1 and Type 2 by hazard level; and cybersecurity requirements are added. It doesn't apply to service robots accessible to the public, or to medical or human-carrying applications — those fall under standards such as ISO 13482.","example":"A collaborative arm sharing a workstation with a worker to drive screws must have its integrator run a risk assessment under ISO 10218-2:2025; if power and force limiting is used, the contact force and pressure at every body part the arm might touch must be measured and verified to stay under the standard's limits.","related":["ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","Power and Force Limiting","Speed and Separation Monitoring","Collaborative Robot","Industrial Robot","ISO 13849 Performance Level (PL a–e, Category B–4) / Safety Integrity Level (SIL, IEC 62061/61508)"]},{"id":"iso-ts-15066-robots-and-robotic-devices-collaborative-robots","category":"control","sec":9,"tier":3,"sources":[{"title":"OSHA Technical Manual, Section IV Chapter 4: Industrial Robot Systems and Industrial Robot System Safety","url":"https://www.osha.gov/otm/section-4-safety-hazards/chapter-4"},{"title":"ISO 10218 - Wikipedia","url":"https://en.wikipedia.org/wiki/ISO_10218"}],"as_of":"2025","related_ids":["collaborative-robot","power-and-force-limiting","speed-and-separation-monitoring","iso-10218-1-2-2025-robotics-safety-requirements","human-robot-collaboration","physical-human-robot-interaction"],"name":"ISO/TS 15066 Robots and Robotic Devices — Collaborative Robots","alt":"ISO/TS 15066 协作机器人安全标准","abbr":"ISO/TS 15066","aliases":["ISO/TS 15066","ISO/TS 15066:2016","TS 15066"],"one_liner":"A 2016 ISO technical specification setting force and pressure limits for human contact with collaborative robots.","explanation":"ISO/TS 15066 is a technical specification (TS) published by the International Organization for Standardization in 2016, supplementing the industrial robot safety standard ISO 10218 with guidance specifically for people and robots sharing a workspace without a fence between them. It details four collaborative modes: safety-rated monitored stop (the robot stops and is monitored when a person enters the collaborative zone), hand guiding (a person leads the robot by hand), speed and separation monitoring (the robot slows as a person gets closer and stops if they get too close), and power and force limiting (contact is allowed, but the force of any impact is limited). Its most-cited content is a table of force and pressure limits by body part, covering two scenarios: transient contact (the person can freely pull away after being struck) and quasi-static contact (the person is pinned between the robot and a fixed object), with sensitive areas such as the face, temple, and throat requiring avoidance of contact altogether. The corresponding U.S. technical report is RIA TR R15.606-2016. As reported, the 2025 revision of ISO 10218-2 has folded the collaborative-application requirements directly into its main text.","example":"When a collaborative arm and a worker share an assembly table, integrators typically configure it under ‘power and force limiting,’ then use a dedicated force-measurement device to test the actual collision force and pressure at spots on the worker's body the arm might touch, confirming they stay under the specification's limits before putting it into service.","related":["Collaborative Robot","Power and Force Limiting","Speed and Separation Monitoring","ISO 10218-1/-2:2025 Robotics — Safety Requirements","Human-Robot Collaboration","Physical Human-Robot Interaction"]},{"id":"iso-13482","category":"control","sec":9,"tier":3,"sources":[{"title":"ISO 13482:2014 Robots and robotic devices — Safety requirements for personal care robots","url":"https://www.iso.org/standard/53820.html"},{"title":"ISO/FDIS 13482 Robotics — Safety requirements for service robots","url":"https://www.iso.org/standard/83498.html"},{"title":"CYBERDYNE: HAL received the world-first certificates of ISO 13482 (2014-11)","url":"https://www.cyberdyne.jp/en/news/1441.html"}],"as_of":"2026-09","related_ids":["service-robot","exoskeleton","iso-10218-1-2-2025-robotics-safety-requirements","functional-safety","physical-human-robot-interaction","embodied-safety"],"name":"ISO 13482 (Safety Requirements for Personal Care Robots)","alt":"ISO 13482 个人护理机器人安全标准","abbr":"","aliases":["ISO 13482:2014","ISO/FDIS 13482","Safety Requirements for Service Robots"],"one_liner":"An international safety standard for mobile service, wearable-assist, and person-carrying robots that operate in close proximity to people's daily lives.","explanation":"Published in February 2014, ISO 13482 covers three categories of ground robots that operate close to people outside industrial and medical settings: mobile servant robots (e.g., delivery robots), physical assistant robots (e.g., wearable exoskeletons), and person carrier robots. It specifies inherently safe design, protective measures, and instructions for use, and it permits physical contact between a person and the robot. It doesn't apply to robots traveling faster than 20 km/h, toys, underwater or flying robots, industrial robots (covered by ISO 10218), medical devices, or military and police use; the standard also notes that at the time of publication there was no internationally agreed limit for collision pain or injury. According to the ISO website, a revision, ISO/FDIS 13482, has reached the final draft stage, retitled ‘Safety requirements for service robots’ and expanded in scope to cover both personal and professional/commercial service robots.","example":"In November 2014, Cyberdyne's lumbar-support exoskeletons HAL for Labor Support and HAL for Care Support received certification from Japan's Quality Assurance Organization (JQA) under ISO 13482:2014 — officially described as the world's first such certifications.","related":["Service Robot","Exoskeleton","ISO 10218-1/-2:2025 Robotics — Safety Requirements","Functional Safety","Physical Human-Robot Interaction","Embodied Safety"]},{"id":"iso-25785-1","category":"control","sec":9,"tier":3,"sources":[{"title":"ISO/CD 25785-1 Robotics — Safety requirements for dynamically stable industrial mobile robots — Part 1: Robots","url":"https://www.iso.org/standard/91469.html"},{"title":"Synapticon: ISO 25785-1: Safety Standard for Dynamically Stable Robots","url":"https://www.synapticon.com/en/newslist/iso-25785-sicherheit-dynamisch-stabile-roboter"},{"title":"Provael: ISO 25785-1 crosswalk（CD 于 2026-05-08 登记）","url":"https://www.provael.com/compliance/iso-25785"}],"as_of":"2026-09","related_ids":["humanoid-robot","dynamic-stability","safe-torque-off","fall-mitigation-and-fall-recovery","iso-10218-1-2-2025-robotics-safety-requirements","quadruped-robot"],"name":"ISO 25785-1 (Safety Requirements for Industrial Mobile Robots with Actively Controlled Stability — Part 1: Robots)","alt":"ISO 25785-1 动态稳定移动机器人安全标准","abbr":"","aliases":["ISO/CD 25785-1","Humanoid Robot Safety Standard","Actively Controlled Stability"],"one_liner":"A safety standard, still being drafted, specifically for humanoid, quadruped, and other industrial mobile robots that need active balancing just to stay upright.","explanation":"ISO 25785-1 is being drafted by ISO/TC 299 Working Group 12, initiated in May 2025, with experts from Agility Robotics, Boston Dynamics, and the U.S. Association for Advancing Automation (A3) leading the effort. It targets ‘actively controlled stability’ industrial mobile robots — bipedal humanoids, quadrupeds, and self-balancing wheeled robots (possibly with arms) that would fall over if power or control were lost — restricted to industrial settings the public can't freely enter. Traditional robots, on encountering danger, simply cut motor torque (safe torque off) and stopping is itself safe; an actively balanced robot instead falls over the moment power is cut, so falling and the center of mass leaving the support base have to be treated as hazards in their own right. Part 1 covers the robot itself; Part 2, covering integrated applications, will be developed separately. As of the committee draft (ISO/CD) stage reached in May 2026, it had not yet been published as of September 2026.","example":"A humanoid robot in a warehouse carrying a box detects a fault; if it simply cut joint torque the way a traditional arm would, it would fall over completely. Under this draft standard's approach, it should instead assess which way it's likely to fall and how hard the impact would be, and be designed to execute a controlled stop or crouch, with a possible fall zone marked out in advance.","related":["Humanoid Robot","Dynamic Stability","Safe Torque Off (STO)","Fall Mitigation and Fall Recovery","ISO 10218-1/-2:2025 Robotics — Safety Requirements","Quadruped Robot"]},{"id":"lyapunov-stability","category":"control","sec":9,"tier":3,"sources":[{"title":"Lyapunov stability - Wikipedia","url":"https://en.wikipedia.org/wiki/Lyapunov_stability"}],"as_of":"","related_ids":["control-lyapunov-function","control-barrier-function","passivity-based-control","impedance-control","dynamic-stability","robust-control"],"name":"Lyapunov Stability","alt":"李雅普诺夫稳定性","abbr":"","aliases":["Lyapunov Function","Lyapunov's Second Method","Lyapunov's Direct Method"],"one_liner":"A theory for judging whether a disturbed system returns to equilibrium, typically proven using an ‘energy function’ that never increases.","explanation":"Lyapunov stability comes from the Russian mathematician Lyapunov's 1892 doctoral thesis, and is the foundation for analyzing the stability of nonlinear systems. It comes in grades: a state starting near equilibrium that never wanders far is Lyapunov stable; if it not only stays close but eventually returns to equilibrium, that's asymptotic stability; if the return rate is at least exponential, that's exponential stability. The most commonly used tool is the second method (the direct method): rather than solving the differential equations, find a function V(x) that equals 0 at equilibrium and is positive everywhere else, and that never increases over time along the system's trajectories (dV/dt ≤ 0) — this proves stability; strictly decreasing proves asymptotic stability. V can be thought of as a generalized energy, but it doesn't have to be actual physical energy. The hard part is that there's no general recipe for constructing V. In robotics, stability proofs for PD plus gravity compensation, impedance control, and passivity-based control all rely on it, and both the control Lyapunov function and control barrier function are built on top of it.","example":"A damped pendulum: taking V = kinetic energy + potential energy (zero at the lowest point), one can compute that along the motion dV/dt = −b·θ̇² (b the damping coefficient, θ̇ angular velocity) — energy only decreases, never increases, so the lowest point is stable without ever solving the pendulum's equations of motion; adding LaSalle's invariance principle further proves the pendulum eventually comes to rest there.","related":["Control Lyapunov Function","Control Barrier Function","Passivity-Based Control","Impedance Control","Dynamic Stability","Robust Control"]},{"id":"control-lyapunov-function","category":"control","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Control-Lyapunov function","url":"https://en.wikipedia.org/wiki/Control-Lyapunov_function"},{"title":"Ames, Galloway, Sreenath, Grizzle: Rapidly Exponentially Stabilizing Control Lyapunov Functions and Hybrid Zero Dynamics (IEEE TAC 2014)","url":"https://ieeexplore.ieee.org/document/6709752"},{"title":"Ames et al., Control Barrier Functions: Theory and Applications（CLF-CBF-QP 一节）","url":"https://arxiv.org/abs/1903.11199"}],"as_of":"","related_ids":["lyapunov-stability","control-barrier-function","quadratic-programming","hybrid-zero-dynamics","feedback-linearization","optimal-control"],"name":"Control Lyapunov Function","alt":"控制李雅普诺夫函数","abbr":"CLF","aliases":["CLF","CLF-QP"],"one_liner":"An energy-like function that, as long as some control choice can always make it decrease, proves the system can be steered to its goal.","explanation":"A Lyapunov function V(x) behaves like a system's ‘energy’: zero at the goal, positive everywhere else. A plain Lyapunov function is used to analyze whether an already-designed closed loop is stable; a control Lyapunov function instead applies to a system that still has a free input u and no fixed control law yet: if, at every non-goal state, there exists some u making V̇ < 0, the system can be steered — stabilized — to the goal. This theory was developed by Artstein and Sontag in the 1980s, and Sontag also gave a general formula for constructing a control law directly from a CLF. Robotics commonly uses CLF-QP: each control cycle, solve for a u satisfying V̇ ≤ −λV (exponential convergence at rate λ) while keeping torque as small as possible, which makes it easy to add torque limits, friction cones, and similar constraints at the same time. Ames et al. applied it to bipedal walking; it is also often combined with a control barrier function in the same quadratic program, with the safety constraint kept hard and the convergence constraint relaxed as a soft constraint.","example":"For a first-order system ẋ = u, take V = x²/2, so V̇ = x·u. Requiring V̇ ≤ −V (i.e., λ = 1) at x = 2 becomes 2u ≤ −2, i.e., u ≤ −1. The CLF-QP picks the smallest-magnitude u satisfying that, u = −1, and the state converges to 0 at an exponential rate.","related":["Lyapunov Stability","Control Barrier Function","Quadratic Programming","Hybrid Zero Dynamics","Feedback Linearization","Optimal Control"]},{"id":"control-barrier-function","category":"control","sec":9,"tier":3,"sources":[{"title":"Ames et al., Control Barrier Functions: Theory and Applications (arXiv:1903.11199)","url":"https://arxiv.org/abs/1903.11199"},{"title":"Ames, Xu, Grizzle, Tabuada: Control Barrier Function Based Quadratic Programs for Safety Critical Systems (arXiv:1609.06408)","url":"https://arxiv.org/abs/1609.06408"}],"as_of":"","related_ids":["control-lyapunov-function","safety-filter","quadratic-programming","hamilton-jacobi-reachability-analysis","safe-reinforcement-learning","artificial-potential-field"],"name":"Control Barrier Function","alt":"控制障碍函数","abbr":"CBF","aliases":["CBF","CBF-QP","Barrier Function"],"one_liner":"Writing ‘don't cross this boundary’ as a constraint, nudging the control command only minimally, and only when safety is close to being violated.","explanation":"A control barrier function defines a safe set with a function h(x): h(x) ≥ 0 counts as safe, where x is the system state (e.g., position and velocity). Its core condition is ḣ ≥ −α·h (α a positive constant): the closer to the boundary — the smaller h is — the more the allowed rate of decrease of h is slowed down, and right at the boundary it can no longer decrease at all, which keeps the state inside the safe set forever. Ames, Tabuada, and others combined this with quadratic programming (arXiv 2016, journal version in IEEE TAC): among all control values satisfying that condition, find the one closest to the original desired command u_des — this is the CBF-QP. Because it only adjusts the command when necessary, it's often used as a safety filter wrapped around teleoperation or a learned policy. It's the counterpart of the control Lyapunov function: a CLF guarantees ‘converge to the goal,’ while a CBF guarantees ‘never leave the safe region’; compared to artificial potential fields, it offers a formal safety guarantee.","example":"A 1D cart moving as ẋ = u toward a wall at x = 5 meters: taking h = 5 − x and α = 2, the condition becomes u ≤ 2(5 − x). At x = 4, the speed limit is 2 m/s, so an original command of 1 m/s is unaffected; at x = 4.9, the limit drops to 0.2 m/s, and the CBF-QP clamps the command to 0.2 m/s. The closer the cart gets to the wall, the slower it's allowed to go, so it never hits it.","related":["Control Lyapunov Function","Safety Filter","Quadratic Programming","Hamilton-Jacobi Reachability Analysis","Safe Reinforcement Learning","Artificial Potential Field"]},{"id":"safety-filter","category":"control","sec":9,"tier":3,"sources":[{"title":"The Safety Filter: A Unified View of Safety-Critical Control in Autonomous Systems (Hsu, Hu, Fisac, arXiv 2309.05837)","url":"https://arxiv.org/abs/2309.05837"},{"title":"Control Barrier Functions: Theory and Applications (Ames et al., arXiv 1903.11199)","url":"https://arxiv.org/abs/1903.11199"}],"as_of":"","related_ids":["control-barrier-function","hamilton-jacobi-reachability-analysis","model-predictive-control","safe-reinforcement-learning","embodied-safety","quadratic-programming"],"name":"Safety Filter","alt":"安全滤波器","abbr":"","aliases":["Safety Shield"],"one_liner":"A layer between a policy and the actuators that only minimally overrides an action when it would otherwise be unsafe.","explanation":"A safety filter splits ‘completing the task’ from ‘staying safe’ into two separate layers. A task policy — a hand-written controller, a reinforcement-learning policy, or a VLA model — proposes a nominal action, and the filter checks whether executing it would keep the system inside a defined safe set. If it's safe, the action passes through unchanged; if not, it's replaced with the closest safe alternative. Three implementations are common: control barrier functions (CBFs) solve a small quadratic program every step to find the action closest to the nominal one that keeps a safety function h(x) from dropping below zero; Hamilton-Jacobi (HJ) reachability analysis precomputes which states will lead to a hazard no matter what control is applied afterward; and MPC-based shielding predicts online whether a future trajectory can still return to the safe region. A 2023 survey by Kai-Chieh Hsu, Haimin Hu, and Jaime Fisac unifies these under a ‘monitor and intervene’ framework. Learned policies carry no inherent safety guarantee on their own, so adding this outer layer is what provides a provable constraint.","example":"For a mobile robot avoiding people, define h(x) = ‖p − p_human‖² − d² (p is the robot's position, d is the minimum allowed distance, and h ≥ 0 means safe). Each cycle it solves min‖u − u_nom‖² subject to ḣ ≥ −α·h, where u_nom is the velocity from the navigation policy and α > 0 sets how early it starts slowing as it approaches. Far from people the constraint is inactive and u simply equals u_nom.","related":["Control Barrier Function","Hamilton-Jacobi Reachability Analysis","Model Predictive Control","Safe Reinforcement Learning","Embodied Safety","Quadratic Programming"]},{"id":"hamilton-jacobi-reachability-analysis","category":"control","sec":9,"tier":3,"sources":[{"title":"Bansal, Chen, Herbert, Tomlin, Hamilton-Jacobi Reachability: A Brief Overview and Recent Advances (arXiv 1709.07523)","url":"https://arxiv.org/abs/1709.07523"}],"as_of":"","related_ids":["control-barrier-function","safety-filter","safe-reinforcement-learning","optimal-control","value-function","embodied-safety"],"name":"Hamilton-Jacobi Reachability Analysis","alt":"HJ 可达性分析","abbr":"HJ reachability","aliases":["HJ Reachability"],"one_liner":"Solving a partial differential equation to compute the safe region — the set of states from which danger is guaranteed avoidable.","explanation":"A formal safety-verification method long championed by Claire Tomlin and colleagues, with a 2017 survey by Bansal, Chen, Herbert, and Tomlin. Given the system's dynamics, bounded disturbances, and a ‘failure set’ (states already counted as a collision, say), solving a Hamilton-Jacobi partial differential equation yields a value function V(x) whose sign carves out the backward reachable set: states from which, under worst-case disturbance, no control strategy can avoid entering the failure set; every other state forms the safe set, and the boundary's optimal safe control also falls out of the solution. It supports nonlinear dynamics and gives strict guarantees, but computation grows exponentially with state dimension (the curse of dimensionality), limiting it to low-dimensional models. It's commonly used as a safety filter: a learned policy controls the robot normally, and control switches to the HJ-derived safe action only near the boundary of the safe set. Compared to a control barrier function, it computes the safe set directly, while a CBF usually requires a human to supply a candidate function first.","example":"Two drones flying toward each other: computing the backward reachable set from their relative position and heading, the original controller is left alone while the relative state stays outside that set, and the moment it touches the boundary, the HJ-derived avoidance maneuver takes over immediately.","related":["Control Barrier Function","Safety Filter","Safe Reinforcement Learning","Optimal Control","Value Function","Embodied Safety"]},{"id":"proprioception","category":"perception","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Proprioception","url":"https://en.wikipedia.org/wiki/Proprioception"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164v1"},{"title":"unitree_rl_gym g1_env.py（G1 策略观测定义）","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/legged_gym/envs/g1/g1_env.py"}],"as_of":"","related_ids":["exteroception","rotary-encoder","state-proprioception-encoder","state-estimation","blind-locomotion","inertial-measurement-unit"],"name":"Proprioception","alt":"本体感知","abbr":"","aliases":["Proprioceptive State","Robot Self-State Sensing"],"one_liner":"A robot's sense of its own joint positions, velocities, forces, and posture, as opposed to sensing the outside world.","explanation":"Proprioception was originally a physiology term for the sense of one's own limb position, movement, and effort; the British physiologist Charles Sherrington used the word in his writing in 1906, contrasting it with exteroception, the sensing of the outside world. In robotics it refers to measurement of the robot's own state: encoders give joint angle and velocity, motor current or torque sensors give force, an IMU gives body attitude and angular velocity, and gripper opening width is often counted too. This data is low-dimensional, low-noise, and high-frequency, so nearly every policy uses it — for example, π0's observation consists of multiple camera images, a language instruction, and a proprioceptive state made of joint angles, with the proprioceptive state projected through a linear layer and fed into the model alongside the other tokens. Legged locomotion that relies only on proprioception, with no camera at all, is called blind locomotion.","example":"In Unitree's open-source unitree_rl_gym, the G1 walking policy's observation includes IMU angular velocity, projected gravity, each joint's angle relative to its default pose, and joint velocity — all proprioception — concatenated with a velocity command, the previous action, and the gait phase. Body linear velocity is hard to measure accurately on the real robot, so it appears only in the privileged observation used for simulation training.","related":["Exteroception","Rotary Encoder","State / Proprioception Encoder","State Estimation","Blind Locomotion","Inertial Measurement Unit"]},{"id":"exteroception","category":"perception","sec":0,"tier":3,"sources":[{"title":"Wikipedia: Exteroception","url":"https://en.wikipedia.org/wiki/Exteroception"},{"title":"Learning robust perceptive locomotion for quadrupedal robots in the wild (Science Robotics 2022)","url":"https://arxiv.org/abs/2201.08117"}],"as_of":"","related_ids":["proprioception","multi-sensor-fusion","perceptive-locomotion","blind-locomotion","state-estimation","lidar"],"name":"Exteroception","alt":"外部感知","abbr":"","aliases":["Exteroceptive Sensing","Environment Perception"],"one_liner":"A robot’s sensing of its surrounding environment using sensors like cameras and lidar.","explanation":"Exteroception comes from physiology, referring to an organism’s ability to sense stimuli from outside the body — sight, hearing, skin touch — as opposed to proprioception, which senses the body’s own limb position and motion. Robotics keeps this same distinction: cameras, depth cameras, lidar, ultrasonic sensors, and tactile sensors that measure the environment fall under exteroception, while joint encoders, IMUs, and motor current that measure the robot’s own state fall under proprioception. Exteroception lets a robot see steps, obstacles, and target objects ahead of time, but it is vulnerable to lighting, occlusion, reflections, and noise; proprioception is stable, but it can only sense terrain once contact has already happened. Legged robots therefore often fuse the two: adjusting gait ahead of time when exteroceptive information is reliable, and falling back on proprioception when it isn’t.","example":"ETH Zurich’s legged locomotion work (Science Robotics 2022) end-to-end fuses exteroception and proprioception with an attention-based recurrent encoder, letting the robot adjust its gait before making contact with terrain, and it completed an Alpine hiking trail within the time recommended for human hikers.","related":["Proprioception","Multi-Sensor Fusion","Perceptive Locomotion","Blind Locomotion","State Estimation","LiDAR"]},{"id":"multimodal-perception","category":"perception","sec":0,"tier":2,"sources":[{"title":"See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation (CoRL 2022)","url":"https://arxiv.org/abs/2212.03858"},{"title":"Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks (ICRA 2019)","url":"https://arxiv.org/abs/1810.10191"}],"as_of":"","related_ids":["multi-sensor-fusion","visuo-tactile-fusion","proprioception","tactile-sensor","contact-rich-manipulation","vision-tactile-language-action-model"],"name":"Multimodal Perception","alt":"多模态感知","abbr":"","aliases":["Multi-Sensory Perception"],"one_liner":"Understanding a scene and task using vision together with touch, force, sound, and other senses at once.","explanation":"Multimodal perception means a robot understands a scene and its task progress using not just vision but simultaneously touch, force and torque, sound, proprioception (sensing of its own state, such as joint angle and motor current), and even language instructions. Different modalities are good at different things: a 2022 paper, See, Hear, and Feel, summarizes it as vision seeing the global picture but often getting occluded, sound catching key moments that aren't visible, and touch providing precise local geometry. The difficulty is that modalities differ hugely in frequency, dimensionality, and noise characteristics, requiring a separate encoder for each one before they can be combined through concatenation, attention, or similar mechanisms, and data collection is also more involved. Multiple studies show that contact-rich manipulation tasks like peg insertion and pouring are more stable when touch and force are added rather than relying on vision alone.","example":"See, Hear, and Feel (CoRL 2022) has an arm use a camera, a contact microphone, and a vision-based tactile sensor together to do dense packing and pouring, fusing the three signals with self-attention, and outperforming setups that use only one or two modalities.","related":["Multi-Sensor Fusion","Visuo-Tactile Fusion","Proprioception","Tactile Sensor","Contact-rich Manipulation","Vision-Tactile-Language-Action Model"]},{"id":"computer-vision","category":"perception","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Computer vision","url":"https://en.wikipedia.org/wiki/Computer_vision"},{"title":"Richard Szeliski, Computer Vision: Algorithms and Applications (2nd ed., 2022)","url":"https://szeliski.org/Book/"}],"as_of":"","related_ids":["object-detection","semantic-segmentation","3d-vision","vision-foundation-model","depth-estimation","vision-encoder"],"name":"Computer Vision","alt":"计算机视觉","abbr":"CV","aliases":["CV"],"one_liner":"The field of making computers extract useful information from images and video and understand what a scene contains.","explanation":"Computer vision studies how to make a computer automatically extract, analyze, and understand information from a single image or a video sequence, covering tasks such as classification, object detection, segmentation, tracking, pose estimation, and 3D reconstruction. The field got its start in AI labs in the late 1960s — in 1966, someone reportedly thought hooking a camera up to a computer and having it ‘describe what it sees’ would take an undergraduate just one summer project; the problem has instead occupied researchers for more than half a century since. Since deep learning took off, convolutional networks, vision transformers, and vision foundation models such as CLIP, DINOv2, and SAM have become the mainstream tools. In embodied AI, the camera is a robot's main channel for information about the outside world, and the vision encoders inside VLA models mostly reuse models pretrained by the computer vision community directly.","example":"Before a robot clears a table, it first uses object detection to locate a cup, segmentation to cut out the cup's outline, and then combines that with a depth map to compute the cup's 3D position — all of these steps belong to computer vision.","related":["Object Detection","Semantic Segmentation","3D Vision","Vision Foundation Model","Depth Estimation","Vision Encoder"]},{"id":"machine-vision","category":"perception","sec":0,"tier":2,"sources":[{"title":"Machine vision - Wikipedia","url":"https://en.wikipedia.org/wiki/Machine_vision"}],"as_of":"","related_ids":["computer-vision","3d-vision-guided-robotics","bin-picking","quality-inspection","structured-light","mech-mind-robotics"],"name":"Machine Vision","alt":"机器视觉（工业视觉）","abbr":"","aliases":["Industrial Vision","Industrial Machine Vision"],"one_liner":"Using cameras and software on a factory floor to automatically inspect, measure, identify, and guide robots.","explanation":"Machine vision refers to using cameras, lenses, light sources, processors, and image-processing software on a factory floor to automate inspection, measurement, recognition, and localization — turning an image directly into an actionable conclusion such as ‘pass or fail’ or ‘the part is here.’ Typical applications include appearance defect detection and sorting, dimensional measurement, code reading, and visual guidance that gives a robot arm a workpiece's position and orientation. Its relationship to computer vision is that CV is the academic field studying image understanding, while machine vision is more systems-engineering focused, emphasizing lighting design, cycle time, and long-term stability, often relying on rule-based algorithms with deep learning as a supplement rather than the core. As 3D cameras have become widespread, 3D-vision-guided bin picking and part feeding have become common configurations on industrial robots, and machine vision is also the reference point embodied AI has to reckon with as it moves into factories.","example":"On a production line, a camera photographs each part as it passes, and software judges whether it has scratches and triggers rejection; at the next station, a 3D camera identifies the pose of parts piled randomly in a bin and guides a robot arm to pick them up one by one for loading.","related":["Computer Vision","3D Vision-Guided Robotics","Bin Picking","Quality Inspection","Structured Light","Mech-Mind Robotics"]},{"id":"rgb-camera","category":"perception","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Bayer filter","url":"https://en.wikipedia.org/wiki/Bayer_filter"},{"title":"ALOHA / ACT project page (Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware)","url":"https://tonyzhaozh.github.io/aloha/"},{"title":"arXiv 2304.13705: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["depth-camera","stereo-camera","wrist-camera","monocular-depth-estimation","camera-intrinsics","global-shutter-rolling-shutter"],"name":"RGB Camera","alt":"RGB相机","abbr":"","aliases":["Color Camera","Monocular Camera"],"one_liner":"An ordinary camera that outputs color images, recording red, green, and blue brightness at every pixel.","explanation":"An RGB camera is simply the most common kind of color camera: its image sensor is covered by a color filter array, most commonly the Bayer filter patented by Kodak engineer Bryce Bayer in 1976, in which half the pixels sense green and a quarter each sense red and blue; a demosaicing algorithm then interpolates a full R, G, B value for every pixel. RGB cameras are cheap, high-resolution, and information-rich, making them the main input for most imitation-learning policies and VLA models. Their limitation is that a single RGB (monocular) camera only captures a 2D projection of the 3D world and has no direct sense of how far away things are; getting depth requires a depth camera, a stereo camera, or a monocular depth estimation model. Other factors to weigh when mounting one include field of view, frame rate, and whether it uses a global shutter (the whole frame exposed at once) or a rolling shutter (exposed row by row, which can skew fast-moving subjects).","example":"The original ALOHA dual-arm platform is fitted with 4 ordinary webcams: two mounted on the wrists, one facing forward, and one overhead looking down. ACT, a policy that uses only these color images plus joint angles — no depth — learned to open a semi-transparent condiment cup and insert a battery from about 10 minutes of demonstrations, reaching 80–90% success.","related":["Depth Camera","Stereo Camera","Wrist Camera","Monocular Depth Estimation","Camera Intrinsics","Global Shutter / Rolling Shutter"]},{"id":"wrist-camera","category":"perception","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Visual servoing (eye-in-hand vs. eye-to-hand)","url":"https://en.wikipedia.org/wiki/Visual_servoing"},{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset (project page)","url":"https://droid-dataset.github.io/"},{"title":"ALOHA project page","url":"https://tonyzhaozh.github.io/aloha/"}],"as_of":"","related_ids":["eye-in-hand","third-person-camera","head-camera","hand-eye-calibration","multi-view","occlusion"],"name":"Wrist Camera","alt":"腕部相机","abbr":"","aliases":["Wrist-Mounted Camera","Eye-in-Hand Camera"],"one_liner":"A camera mounted on a robot arm's wrist or gripper that moves with the hand, giving a close-up view of the workspace.","explanation":"A wrist camera is fixed near the gripper on a robot arm's end effector, moving together with the hand. This is the ‘eye-in-hand’ configuration described in visual servoing, where the camera moves with the hand and sees the hand's position relative to the target; the opposite is a fixed, third-person camera in the environment, called ‘eye-to-hand.’ The advantage of a wrist camera is that it sees the final few centimeters of a grasp or insertion most clearly, is less likely to be blocked by the arm itself, and shows the object's position directly in relation to the gripper — so imitation-learning and VLA data collection almost always includes one, feeding into the policy alongside one or two third-person cameras. Its drawback is that its viewpoint changes drastically with the arm's motion, it has no global view, and a depth camera can become inaccurate at very close range. Using its images for geometric calculations first requires hand-eye calibration.","example":"The DROID dataset mounts 3 cameras on each Franka arm: 2 adjustable-position external ZED 2 stereo cameras, plus 1 ZED Mini on the wrist.","related":["Eye-in-Hand","Third-Person Camera","Head Camera","Hand-Eye Calibration","Multi-View","Occlusion"]},{"id":"third-person-camera","category":"perception","sec":0,"tier":2,"sources":[{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset","url":"https://arxiv.org/abs/2403.12945"}],"as_of":"","related_ids":["wrist-camera","head-camera","eye-to-hand","camera-extrinsics","multi-view","droid"],"name":"Third-Person Camera","alt":"第三视角相机","abbr":"","aliases":["Static Camera","External Camera","Exterior Camera","Global Camera","3rd-Person Camera"],"one_liner":"A camera mounted off the robot’s body, fixed in place, viewing the whole workspace from the side.","explanation":"A third-person camera is mounted on a tripod, wall, or table edge and doesn’t move with the robot arm, in contrast to a wrist camera or head camera. It has a wide field of view and captures the robot, objects, and environment together, giving global position information; the tradeoff is that distant details are hard to see and the arm can easily occlude objects, so manipulation datasets usually pair it with a wrist camera. This setup is a form of “eye-to-hand” mounting, and its extrinsics — the camera’s position and orientation relative to the robot — must be calibrated before image coordinates can be converted into the robot’s base frame. If the camera moves even slightly, the scene a policy sees changes, and success rates often drop, which is why generalization across viewpoints is a common evaluation axis. In papers this camera is also called an exterior camera or 3rd-person camera.","example":"Each capture rig in the DROID dataset has two repositionable ZED 2 stereo cameras serving as third-person views, plus a ZED Mini mounted on the wrist; every time the scene changes, the operator repositions the third-person cameras and recalibrates their extrinsics with a checkerboard.","related":["Wrist Camera","Head Camera","Eye-to-Hand","Camera Extrinsics","Multi-View","DROID (Distributed Robot Interaction Dataset)"]},{"id":"head-camera","category":"perception","sec":0,"tier":2,"sources":[{"title":"Mobile ALOHA (arXiv:2401.02117)","url":"https://arxiv.org/html/2401.02117"},{"title":"Open-TeleVision (arXiv:2407.01512)","url":"https://arxiv.org/html/2407.01512"},{"title":"Unitree G1 产品页","url":"https://www.unitree.com/g1"}],"as_of":"","related_ids":["wrist-camera","third-person-camera","egocentric-video","open-television","multi-view","camera-extrinsics"],"name":"Head Camera","alt":"头部相机","abbr":"","aliases":["Head-Mounted Camera","Top Camera"],"one_liner":"A camera mounted on a robot's head or upper body that gives a near-first-person view of the scene.","explanation":"A head camera is mounted on a humanoid's head, or above a mobile manipulation platform, giving a near-first-person view similar to looking out through one's own eyes. It's usually paired with a wrist camera and a fixed third-person camera: the head camera sees the overall scene and both hands, while the wrist camera sees close-up detail near the gripper. Mobile ALOHA, for instance, uses two wrist cameras plus one forward-facing camera mounted up top; Unitree's G1 carries a depth camera and a 3D lidar on its head. When the head can rotate, the camera's extrinsics change with the neck joint and must be recomputed in real time through kinematics. Teleoperation setups also use an ‘active head camera’: Open-TeleVision mounts a ZED Mini stereo camera on a 2-DOF gimbal that follows the operator's head rotation inside an Apple Vision Pro headset and streams back a stereo view. When collecting human data, a camera worn on a person's head or on glasses is likewise called head-mounted, capturing first-person video.","example":"Training a policy for dual-arm clothes-folding feeds in three camera streams: the head camera sees the whole garment laid out flat, while the two wrist cameras check whether the gripper has hold of the garment's corner.","related":["Wrist Camera","Third-Person Camera","Egocentric Video","Open-TeleVision","Multi-View","Camera Extrinsics"]},{"id":"palm-camera","category":"perception","sec":0,"tier":3,"sources":[{"title":"Introducing Figure 03 (Figure AI, 2025-10-09)","url":"https://www.figure.ai/news/introducing-figure-03"},{"title":"DexWild 项目主页","url":"https://dexwild.github.io/"}],"as_of":"2025-10","related_ids":["wrist-camera","head-camera","occlusion","dexwild","figure-03","cross-embodiment-data"],"name":"Palm Camera","alt":"手掌相机","abbr":"","aliases":["Palm-Mounted Camera","In-Hand Palm Camera"],"one_liner":"A camera mounted in a robot’s palm that gives a close-up view of the hand and object during grasping.","explanation":"A palm camera is a small camera embedded in a robot’s palm, or in the palm of a handheld data-collection device, positioned closer to the contact point than a wrist camera is. A head-mounted camera often loses sight of the contact area once the hand reaches into a cabinet or is blocked by the object or the arm itself; a palm camera keeps the object in view through the final few centimeters of a grasp. Figure AI’s Figure 03, released in October 2025, includes a wide-angle, low-latency palm camera in each hand, which the company says provides redundant close-range visual feedback when the main cameras are blocked. Carnegie Mellon’s DexWild (RSS 2025) mounts the same kind of palm camera on both its handheld human data-collection device and the robot hand; because this viewpoint mostly frames the task and environment rather than the hand itself, human and robot data are easier to train on together.","example":"As Figure 03 reaches into a cabinet for a cup, its head camera is blocked by the cabinet door, so it keeps adjusting its finger positions using the palm camera’s view instead.","related":["Wrist Camera","Head Camera","Occlusion","DexWild","Figure 03","Cross-Embodiment Data"]},{"id":"occlusion","category":"perception","sec":0,"tier":2,"sources":[{"title":"Occlusion Handling in Generic Object Detection: A Review (SAMI 2021)","url":"https://arxiv.org/abs/2101.08845"},{"title":"See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation（视觉易受遮挡）","url":"https://arxiv.org/abs/2212.03858"},{"title":"Vision-Based Manipulators Need to Also See from Their Hands (ICLR 2022)","url":"https://arxiv.org/abs/2203.12677"}],"as_of":"","related_ids":["multi-view","wrist-camera","active-perception","point-cloud-completion-shape-completion","visuo-tactile-fusion","object-tracking"],"name":"Occlusion","alt":"遮挡","abbr":"","aliases":["Self-Occlusion","Mutual Occlusion"],"one_liner":"A target being partly or fully blocked from a sensor's view by something else, including the robot's own body.","explanation":"Occlusion means part or all of a target is blocked by something else, so a sensor can't see it fully. Three scenarios are common in robotics: objects blocking each other (cluttered parts), the robot blocking its own view (the arm reaching in front of an object, or a dexterous hand's fingers blocking an object in the palm — called self-occlusion), and, during manipulation, the hand and the object blocking each other. Occlusion causes missing or wrong values in detection, segmentation, pose estimation, and depth, and a blocked object can effectively ‘disappear’ from a policy's input, leading to misjudgment. Because the location, size, and proportion of occlusion all vary unpredictably, it's one of the main reasons detection models still fall short of human performance. Common countermeasures include multiple camera views, a wrist camera, actively moving the viewpoint (active perception), adding touch or sound, or having the model complete or remember the blocked part.","example":"An arm relying only on a top-down camera to grasp a cup has its own arm block the cup right as the end-effector gets close, so the policy loses sight of the cup's position; adding a wrist camera fills in that missing view.","related":["Multi-View","Wrist Camera","Active Perception","Point Cloud Completion / Shape Completion","Visuo-Tactile Fusion","Object Tracking"]},{"id":"multi-view","category":"perception","sec":0,"tier":2,"sources":[{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"},{"title":"Vision-Based Manipulators Need to Also See from Their Hands (ICLR 2022)","url":"https://arxiv.org/abs/2203.12677"}],"as_of":"","related_ids":["wrist-camera","head-camera","third-person-camera","multi-view-stereo","occlusion","camera-extrinsics"],"name":"Multi-View","alt":"多视角","abbr":"","aliases":["Multi-Camera"],"one_liner":"Observing the same scene from multiple cameras or multiple angles at once.","explanation":"Multi-view refers to observing the same scene from two or more positions, whether that's several cameras shooting at once or a single camera moving while it shoots. It mainly shows up in two places: 3D vision, where images from different angles are used for triangulation, multi-view stereo reconstruction, or training a NeRF; and robot learning, where a robot commonly feeds a head camera, a wrist camera (close to the gripper), and a fixed third-person camera into the policy together. The benefit is reduced occlusion and depth information that a single image can't provide; research has also found that adding a wrist view improves training efficiency and generalization to new scenes. The cost is more overhead for calibration, synchronization, and compute, and every camera's intrinsics and extrinsics need to be recorded.","example":"The DROID dataset's collection rig pairs two adjustable-position ZED 2 stereo cameras with one wrist-mounted ZED Mini, collecting 76,000 trajectories in total across 1,417 camera viewpoints.","related":["Wrist Camera","Head Camera","Third-Person Camera","Multi-View Stereo","Occlusion","Camera Extrinsics"]},{"id":"pinhole-camera-model","category":"perception","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Pinhole camera model","url":"https://en.wikipedia.org/wiki/Pinhole_camera_model"},{"title":"Intel RealSense Wiki: Projection in RealSense SDK 2.0","url":"https://github.com/IntelRealSense/librealsense/wiki/Projection-in-RealSense-SDK-2.0"}],"as_of":"","related_ids":["camera-intrinsics","camera-extrinsics","lens-distortion","projection-back-projection","camera-calibration","camera-coordinate-frame"],"name":"Pinhole Camera Model","alt":"针孔相机模型","abbr":"","aliases":["Pinhole Model","Perspective Projection Model"],"one_liner":"A mathematical model treating a camera as an ideal tiny hole, projecting 3D points onto an image plane along straight lines.","explanation":"The pinhole camera model is computer vision's standard model for how a 3D point becomes a pixel: it assumes every ray of light passes through an ideal tiny hole, the optical center, and a point (X, Y, Z) in the camera's coordinate frame lands at u = fx·X/Z + cx, v = fy·Y/Z + cy. Here fx and fy are the focal lengths in pixels, and (cx, cy) is the principal point — together these four numbers are the camera's intrinsics. Dividing by Z produces perspective — near things look big, far things look small. The model ignores lens distortion and defocus blur, so it's only a first-order approximation; in practice, distortion coefficients are calibrated first to undistort the image, and the pinhole model is applied afterward. Turning a depth map into a point cloud, converting a pixel position into a 3D grasp point, and doing hand-eye calibration are all built on this formula.","example":"A camera has fx = fy = 600 pixels and principal point (320, 240). A point 1 meter directly ahead and 0.1 meters to the right lands at column 600 × 0.1/1 + 320 = 380; moving that same point to 2 meters away lands it at column 350, closer to the center of the image.","related":["Camera Intrinsics","Camera Extrinsics","Lens Distortion","Projection / Back-Projection","Camera Calibration","Camera Coordinate Frame"]},{"id":"camera-coordinate-frame","category":"perception","sec":0,"tier":2,"sources":[{"title":"ROS REP 103: Standard Units of Measure and Coordinate Conventions","url":"https://raw.githubusercontent.com/ros-infrastructure/rep/master/rep-0103.rst"},{"title":"MATLAB: What Is Camera Calibration?","url":"https://www.mathworks.com/help/vision/ug/camera-calibration.html"}],"as_of":"","related_ids":["coordinate-frame","camera-extrinsics","projection-back-projection","right-handed-frame-and-axis-conventions","tf-tf2-transform-tree","pinhole-camera-model"],"name":"Camera Coordinate Frame","alt":"相机坐标系","abbr":"","aliases":["Camera Frame","Camera Optical Frame"],"one_liner":"A 3D coordinate frame centered on the camera's optical center, with the z axis along the optical axis, used to locate objects relative to the camera.","explanation":"The camera coordinate frame has its origin at the camera's optical center; by computer vision convention, the z axis points forward along the optical axis, x points to the image's right, and y points down. Point clouds back-projected from a depth map, and tag poses solved by PnP, are both initially expressed in this frame, and must be multiplied by the camera's extrinsics (the camera-to-base or camera-to-world transform) before they can be used in a robot arm's base frame. A common pitfall is mismatched conventions: ROS's REP 103 defines a robot's body frame as x-forward, y-left, z-up, while a camera's ‘optical frame’ (named with an _optical suffix) is z-forward, x-right, y-down — the two differ by a fixed rotation. Graphics libraries like OpenGL use yet another convention, with the camera frame's z pointing backward and y pointing up. Mixing these up scrambles a point cloud's orientation entirely.","example":"RealSense's ROS driver publishes both camera_link (x-forward) and a frame with an _optical suffix. A point cloud message's frame_id is the latter, so treating it as x-forward directly will get the orientation wrong.","related":["Coordinate Frame","Camera Extrinsics","Projection / Back-Projection","Right-Handed Frame & Axis Conventions","TF / tf2 Transform Tree","Pinhole Camera Model"]},{"id":"camera-intrinsics","category":"perception","sec":0,"tier":1,"sources":[{"title":"OpenCV calib3d.hpp：camera intrinsic matrix（fx、fy、cx、cy）","url":"https://raw.githubusercontent.com/opencv/opencv/4.x/modules/calib3d/include/opencv2/calib3d.hpp"},{"title":"RealSense SDK rs_types.h：rs2_intrinsics 结构体","url":"https://github.com/realsenseai/librealsense/blob/master/include/librealsense2/h/rs_types.h"},{"title":"Wikipedia: Camera resectioning","url":"https://en.wikipedia.org/wiki/Camera_resectioning"}],"as_of":"","related_ids":["camera-extrinsics","pinhole-camera-model","lens-distortion","projection-back-projection","camera-calibration","calibration-board"],"name":"Camera Intrinsics","alt":"相机内参","abbr":"","aliases":["Intrinsic Parameters","Intrinsic Matrix","K Matrix","Focal Length and Principal Point"],"one_liner":"The parameters describing how a camera itself forms an image, determining which pixel a 3D point lands on.","explanation":"Camera intrinsics are usually written as a 3×3 matrix K, containing the focal lengths fx and fy in pixels and the principal point cx, cy (where the optical axis meets the image plane, generally near the image center); lens distortion coefficients are typically listed separately. Intrinsics describe how a 3D point in the camera's own coordinate frame projects onto a pixel, and depend only on the camera and lens, not on what's being photographed — so they stay valid as long as the focal length doesn't change. They can be found by calibrating against a checkerboard (Zhang Zhengyou's method), and depth cameras such as the RealSense ship factory-calibrated, with intrinsics readable straight from the SDK. A common pitfall: if an image is resized or cropped before being fed to a model, the intrinsics must be adjusted by the same scale and offset, or any point cloud reprojected from it will come out misaligned.","example":"For a 640×480 image with fx = fy = 600, cx = 320, cy = 240: a pixel at (420, 240) with a measured depth of 1 meter corresponds to a camera-frame point of X = (420 − 320) × 1/600 ≈ 0.17 m, Y = 0, Z = 1 m.","related":["Camera Extrinsics","Pinhole Camera Model","Lens Distortion","Projection / Back-Projection","Camera Calibration","Calibration Board"]},{"id":"camera-extrinsics","category":"perception","sec":0,"tier":1,"sources":[{"title":"OpenCV calib3d.hpp：针孔相机模型与外参 R、t 的定义","url":"https://raw.githubusercontent.com/opencv/opencv/4.x/modules/calib3d/include/opencv2/calib3d.hpp"},{"title":"Wikipedia: Camera resectioning","url":"https://en.wikipedia.org/wiki/Camera_resectioning"},{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"}],"as_of":"","related_ids":["camera-intrinsics","hand-eye-calibration","homogeneous-transformation-matrix","coordinate-transformation","perspective-n-point","camera-calibration"],"name":"Camera Extrinsics","alt":"相机外参","abbr":"","aliases":["Extrinsic Parameters","Camera Pose","Extrinsic Matrix"],"one_liner":"The rotation-plus-translation parameters describing where a camera sits in space and which way it points.","explanation":"Camera extrinsics consist of a 3×3 rotation matrix R and a translation vector t, often combined into a 4×4 homogeneous transformation matrix, describing where the camera sits and which way it points relative to a world frame or a robot's base frame. OpenCV's convention transforms a point in the world frame into the camera frame; many robotics codebases instead store the opposite direction — a ‘camera pose’ that goes from camera to world — so it's worth confirming the direction before reusing someone else's data. Extrinsics are typically found using a calibration board together with PnP, or through hand-eye calibration, and need to be redone if the camera gets bumped. Once known, they let point clouds from multiple cameras be merged into one common frame, or let an object's position as seen by the camera be converted into coordinates a robot arm can act on.","example":"The DROID dataset was collected with two external ZED 2 cameras plus one wrist-mounted ZED Mini, covering 1,417 camera viewpoints in total, and ships each viewpoint's intrinsic and extrinsic calibration alongside the data.","related":["Camera Intrinsics","Hand-Eye Calibration","Homogeneous Transformation Matrix","Coordinate Transformation","Perspective-n-Point","Camera Calibration"]},{"id":"projection-back-projection","category":"perception","sec":0,"tier":2,"sources":[{"title":"Intel RealSense Wiki: Projection in RealSense SDK 2.0","url":"https://github.com/IntelRealSense/librealsense/wiki/Projection-in-RealSense-SDK-2.0"},{"title":"Open3D API: open3d.geometry.PointCloud (create_from_depth_image)","url":"https://www.open3d.org/docs/release/python_api/open3d.geometry.PointCloud.html"}],"as_of":"","related_ids":["pinhole-camera-model","camera-intrinsics","depth-map","point-cloud","depth-to-color-alignment","camera-extrinsics"],"name":"Projection / Back-Projection","alt":"投影与反投影","abbr":"","aliases":["Depth-to-Point-Cloud Conversion","Unprojection","Deprojection"],"one_liner":"Projection maps a 3D point onto a pixel; back-projection uses a pixel plus its depth to recover the 3D point.","explanation":"Projection and back-projection are the two directions of the pinhole camera model. Projection: given a 3D point (X, Y, Z) in the camera's coordinate frame and the intrinsics, compute which pixel (u, v) it lands on. Back-projection: a single pixel only defines a ray, so recovering the 3D point also needs that pixel's depth Z: X = (u − cx)·Z/fx, Y = (v − cy)·Z/fy. Doing this for every pixel of a depth map produces a point cloud, and both Open3D and the RealSense SDK provide ready-made functions for it. A few practical points: raw depth values are usually stored as integers (RealSense D400 cameras default to 1 unit = 1 millimeter) and must be converted to meters; color and depth images from different sensors need to be aligned first; and the result must then be multiplied by the camera's extrinsics to land in the robot's base frame.","example":"A VLM points at a cup's handle at pixel (412, 305) on the color image. Looking up the aligned depth gives 0.52 meters; back-projection with the intrinsics gives the 3D point in the camera frame, and the hand-eye-calibrated extrinsics then convert it into the arm's base frame as the grasp target.","related":["Pinhole Camera Model","Camera Intrinsics","Depth Map","Point Cloud","Depth-to-Color Alignment","Camera Extrinsics"]},{"id":"lens-distortion","category":"perception","sec":0,"tier":2,"sources":[{"title":"OpenCV Tutorial: Camera Calibration","url":"https://raw.githubusercontent.com/opencv/opencv/4.x/doc/py_tutorials/py_calib3d/py_calibration/py_calibration.markdown"},{"title":"Distortion (optics) - Wikipedia","url":"https://en.wikipedia.org/wiki/Distortion_(optics)"}],"as_of":"","related_ids":["camera-calibration","camera-intrinsics","pinhole-camera-model","fisheye-camera","calibration-board","opencv"],"name":"Lens Distortion","alt":"镜头畸变","abbr":"","aliases":["Radial Distortion","Tangential Distortion","Undistortion"],"one_liner":"The way a real lens's image deviates from the ideal pinhole model, bending straight lines into curves.","explanation":"The pinhole camera model assumes a straight line in space still photographs as a straight line, but real lenses deviate from this — that deviation is lens distortion. Radial distortion worsens toward the edges of the image: wide-angle and fisheye lenses commonly bow outward (barrel distortion), while telephoto lenses commonly bow inward (pincushion distortion); tangential distortion comes from the lens not being mounted perfectly parallel to the sensor. It's commonly described with the Brown-Conrady model, which OpenCV represents using three radial coefficients (k1, k2, k3) and two tangential coefficients (p1, p2); these are solved together with focal length and the principal point through checkerboard camera calibration, and an undistort function then straightens the image. Robots need to undistort before computing a 3D position from pixels, performing hand-eye calibration, or running visual SLAM, or the error near the edges of the frame grows significantly.","example":"A wrist camera's wide-angle lens photographs the table's edge as a curved arc. After calibrating the distortion coefficients and calling OpenCV's cv.undistort, the table edge straightens back out, letting edge objects' pixel coordinates convert accurately into arm coordinates.","related":["Camera Calibration","Camera Intrinsics","Pinhole Camera Model","Fisheye Camera","Calibration Board","OpenCV (Open Source Computer Vision Library)"]},{"id":"field-of-view","category":"perception","sec":0,"tier":2,"sources":[{"title":"Field of view - Wikipedia","url":"https://en.wikipedia.org/wiki/Field_of_view"},{"title":"Universal Manipulation Interface (arXiv:2402.10329)","url":"https://arxiv.org/html/2402.10329"}],"as_of":"","related_ids":["pinhole-camera-model","camera-intrinsics","fisheye-camera","lens-distortion","wrist-camera","depth-camera"],"name":"Field of View","alt":"视场角","abbr":"FoV","aliases":["FoV","Angle of View"],"one_liner":"The angular range a camera or sensor can see at once, usually given as horizontal, vertical, and diagonal figures.","explanation":"Field of view is the angular range a camera or sensor covers at any given instant, usually reported separately as horizontal, vertical, and diagonal, in degrees. For an ordinary pinhole camera, it's set by the sensor size s and the focal length f: FoV = 2·arctan(s/2f) — a shorter focal length gives a wider field. Lidar and depth-camera spec sheets likewise list horizontal and vertical FoV separately. In embodied AI, field of view determines how much a camera can actually see: a wrist camera sits close to the object, so too narrow a FoV loses the surrounding context, which is why UMI added a 155° fisheye lens to its wrist-mounted GoPro. Conversely, a wider FoV means more distortion at the edges and fewer pixels devoted to each degree of view. Swapping a camera or lens changes the distribution of the images a policy sees, so an already-trained vision policy often needs extra data to adapt.","example":"A camera with a 36 mm-wide sensor and an 18 mm focal length has a horizontal field of view of 2·arctan(36/36) = 90°; switching to a 36 mm focal length shrinks that to about 53°.","related":["Pinhole Camera Model","Camera Intrinsics","Fisheye Camera","Lens Distortion","Wrist Camera","Depth Camera"]},{"id":"fisheye-camera","category":"perception","sec":0,"tier":2,"sources":[{"title":"Fisheye lens - Wikipedia","url":"https://en.wikipedia.org/wiki/Fisheye_lens"},{"title":"Kalibr: Supported camera and distortion models","url":"https://github.com/ethz-asl/kalibr/wiki/supported-models"},{"title":"Universal Manipulation Interface (arXiv:2402.10329)","url":"https://arxiv.org/html/2402.10329"}],"as_of":"","related_ids":["field-of-view","lens-distortion","camera-calibration","universal-manipulation-interface","wrist-camera","visual-slam"],"name":"Fisheye Camera","alt":"鱼眼相机","abbr":"","aliases":["Wide-Angle Camera","Fisheye Lens"],"one_liner":"A camera with an ultra-wide fisheye lens, often around 180° field of view, whose images bow visibly at the edges.","explanation":"A fisheye camera uses an ultra-wide-angle fisheye lens, typically with a field of view of 100° to 180°, sometimes more. Unlike an ordinary lens, which keeps straight lines straight, it maps the scene through a specific projection (equidistant, equisolid-angle, and others) to compress a very wide view into one image, giving the image a convex, barrel-like distortion; the term ‘fisheye’ was coined by the physicist Robert W. Wood in 1906. The ordinary pinhole model can't describe distortion this severe, so calibration needs specialized models, such as the equidistant, double-sphere, and omnidirectional models supported by the Kalibr toolbox. Two uses are common in robotics: visual SLAM and panoramic sensing use fisheye lenses to widen the field of view and reduce tracking loss, while wrist cameras use them to add close-range environmental context — the UMI paper notes that a fisheye lens preserves resolution at the center of the image while compressing peripheral information into the frame.","example":"UMI's handheld gripper adds a 155° fisheye lens to its wrist-mounted GoPro, so that even when the gripper gets close to an object, enough of the surrounding tabletop is still visible in frame.","related":["Field of View","Lens Distortion","Camera Calibration","Universal Manipulation Interface","Wrist Camera","Visual SLAM"]},{"id":"cmos-image-sensor","category":"perception","sec":0,"tier":3,"sources":[{"title":"Active-pixel sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Active-pixel_sensor"}],"as_of":"","related_ids":["rgb-camera","global-shutter-rolling-shutter","image-signal-processor","event-camera","depth-camera","mipi-csi-2"],"name":"CMOS Image Sensor","alt":"CMOS 图像传感器","abbr":"CIS","aliases":["CIS","Active-Pixel Sensor"],"one_liner":"The chip inside a camera that turns light into a digital image, with amplification built into every pixel — the dominant sensor today.","explanation":"A CMOS image sensor is an active-pixel sensor built with CMOS manufacturing: each pixel has a photodiode that collects light-generated charge, which a handful of transistors then amplify and read out on the spot (commonly a 4-transistor design). In the early 1990s, Mitsubishi Electric and NASA’s Jet Propulsion Laboratory (Eric Fossum and colleagues) each independently built CMOS active-pixel sensors; they’re cheap, low-power, and can integrate analog-to-digital conversion and image processing on the same chip, and sales overtook CCDs in 2007 — today CMOS is overwhelmingly dominant. Sony is the largest manufacturer. For robotics, the sensor determines a camera’s resolution, frame rate, low-light noise, and dynamic range, as well as its shutter type: most CMOS sensors use a rolling shutter that exposes row by row, which skews the image under fast motion, so VIO and high-speed grasping applications often pick global-shutter models instead.","example":"When picking a wrist camera for a robot arm, check the shutter type, pixel size, and frame rate on the spec sheet; for a fast-moving arm, choose a global-shutter CMOS sensor so object edges in the image don’t come out skewed.","related":["RGB Camera","Global Shutter / Rolling Shutter","Image Signal Processor","Event Camera","Depth Camera","MIPI CSI-2"]},{"id":"image-signal-processor","category":"perception","sec":0,"tier":3,"sources":[{"title":"Wikipedia: Image processor","url":"https://en.wikipedia.org/wiki/Image_processor"},{"title":"NVIDIA Jetson Linux Developer Guide: Camera Software Development Solution","url":"https://docs.nvidia.com/jetson/archives/r36.4/DeveloperGuide/SD/CameraDevelopment/CameraSoftwareDevelopmentSolution.html"}],"as_of":"","related_ids":["cmos-image-sensor","mipi-csi-2","rgb-camera","nvidia-jetson","global-shutter-rolling-shutter"],"name":"Image Signal Processor","alt":"ISP（图像信号处理器）","abbr":"ISP","aliases":["ISP","Image Processor"],"one_liner":"The processing unit that turns an image sensor’s raw output into a normal color image.","explanation":"An ISP is a dedicated processor in a camera’s imaging pipeline — sometimes a standalone chip, but more often integrated into a system-on-chip (SoC) like a phone’s or a Jetson’s. A CMOS image sensor outputs raw Bayer-pattern data, where each pixel records only red, green, or blue; the ISP performs demosaicing (interpolating to recover a full-color image), noise reduction, white balance, auto-exposure, gamma and color correction, sharpening, and more, in sequence, outputting an RGB or YUV image. When a robot connects a bare sensor over a MIPI interface, it needs the host chip’s ISP to process and tune the image; USB cameras and some camera modules with a built-in ISP output an already-processed image directly. Auto-exposure and auto-white-balance make the brightness and color of the same scene drift over time, which affects training-data consistency, so they are often locked manually while collecting data.","example":"NVIDIA Jetson developer kits’ OV5693 camera module has no ISP of its own — reading it directly over V4L2 gives raw Bayer data; going through libargus (nvarguscamerasrc) routes it through the Jetson’s built-in ISP, producing a normal color image.","related":["CMOS Image Sensor","MIPI CSI-2","RGB Camera","NVIDIA Jetson","Global Shutter / Rolling Shutter"]},{"id":"global-shutter-rolling-shutter","category":"perception","sec":0,"tier":3,"sources":[{"title":"Wikipedia: Rolling shutter","url":"https://en.wikipedia.org/wiki/Rolling_shutter"}],"as_of":"","related_ids":["cmos-image-sensor","rgb-camera","visual-inertial-odometry","wrist-camera","event-camera","camera-calibration"],"name":"Global Shutter / Rolling Shutter","alt":"全局快门 / 卷帘快门","abbr":"","aliases":["Global Exposure","Rolling Exposure","Jello Effect"],"one_liner":"The two ways a camera can expose an image: all pixels at once, or row by row in sequence.","explanation":"These are the two exposure and readout schemes an image sensor can use. A global shutter exposes all pixels at the same instant; a rolling shutter exposes and reads out row by row from top to bottom, with a time gap between rows. Many CMOS cameras use a rolling shutter because it is structurally simpler, cheaper, and more light-sensitive. But when photographing a fast-moving object, or when the camera itself is shaking, the image skews, warps, or judders — this is the “jello effect”; under a flash, only part of the frame may get lit. For robots, a wrist camera moving with the arm, or a legged robot’s body shaking while walking, means a rolling shutter introduces errors into SLAM, visual-inertial odometry, and camera calibration, which is why these scenarios often use a global-shutter camera instead, or model and compensate for the rolling shutter’s row-by-row time offset in the algorithm.","example":"Photographing a fast-spinning propeller with a rolling-shutter camera makes the blades come out bent or even broken up into strange shapes; a global-shutter camera doesn’t show this.","related":["CMOS Image Sensor","RGB Camera","Visual-Inertial Odometry","Wrist Camera","Event Camera","Camera Calibration"]},{"id":"mipi-csi-2","category":"perception","sec":0,"tier":3,"sources":[{"title":"MIPI CSI-2 Specification - MIPI Alliance","url":"https://www.mipi.org/specifications/csi-2"}],"as_of":"","related_ids":["gigabit-multimedia-serial-link","cmos-image-sensor","image-signal-processor","nvidia-jetson","embedded-system"],"name":"MIPI CSI-2","alt":"MIPI CSI-2 相机接口","abbr":"MIPI CSI-2","aliases":["CSI-2","MIPI Camera Serial Interface 2","CSI Camera Interface"],"one_liner":"A high-speed serial interface standard that carries image data from a camera sensor to a processor.","explanation":"MIPI CSI-2 is a camera serial interface specification defined by the MIPI Alliance (Mobile Industry Processor Interface Alliance) that specifies how an image sensor sends raw pixel data to a processor — an SoC, or system-on-chip — at high speed. Nearly every phone camera uses it, and embedded boards such as NVIDIA’s Jetson and the Raspberry Pi also expose CSI ports. It has low latency and low power draw, and can connect directly into a chip’s onboard ISP (image signal processor). Its cable runs are short, however, so when a robot’s or car’s camera sits far from the main compute unit, the signal is typically extended over longer distances with serializers such as GMSL (Gigabit Multimedia Serial Link).","example":"A CSI camera module is wired to a Jetson Orin dev board with a ribbon cable, feeding images straight into the board’s onboard ISP.","related":["Gigabit Multimedia Serial Link","CMOS Image Sensor","Image Signal Processor","NVIDIA Jetson","Embedded System"]},{"id":"gigabit-multimedia-serial-link","category":"perception","sec":0,"tier":3,"sources":[{"title":"Wikipedia: Gigabit Multimedia Serial Link","url":"https://en.wikipedia.org/wiki/Gigabit_Multimedia_Serial_Link"},{"title":"NVIDIA Jetson 开发者指南：Jetson Virtual Channel with GMSL Camera Framework","url":"https://docs.nvidia.com/jetson/archives/r36.4/DeveloperGuide/SD/CameraDevelopment/JetsonVirtualChannelWithGmslCameraFramework.html"}],"as_of":"2024","related_ids":["mipi-csi-2","nvidia-jetson","multi-sensor-time-synchronization","cmos-image-sensor","autonomous-driving","rgb-camera"],"name":"Gigabit Multimedia Serial Link","alt":"GMSL 相机接口","abbr":"GMSL","aliases":["GMSL","GMSL2","GMSL3"],"one_liner":"An automotive-grade serial link that powers a camera and carries high-speed image data over a single coaxial cable.","explanation":"GMSL is an automotive serial transmission technology introduced by Maxim in 2008, now owned by Analog Devices (ADI) since it acquired Maxim in 2021. A serializer at the camera end packs image data and sends it over a single coaxial cable or shielded twisted pair to a deserializer at the host end, which converts it to MIPI CSI-2 for the processor; the same cable also carries power and bidirectional control signals, over cable runs of up to 15 meters. GMSL2 offers 6 Gb/s of bandwidth, and GMSL3 offers 12 Gb/s. USB cameras have short cable runs, connectors that come loose easily, and poor noise immunity; GMSL comes out of automotive driver-assistance systems, with good electromagnetic interference resistance, and multiple camera streams can be combined into the same deserializer — which is why multi-camera setups in autonomous driving as well as humanoid and mobile robots commonly use it to connect cameras to a controller like a Jetson.","example":"On a Jetson AGX Orin, multiple GMSL cameras connect through a deserializer into the same CSI port, distinguished by virtual channels; NVIDIA’s documentation states the AGX Orin series supports up to 16 virtual channels when using the ISP.","related":["MIPI CSI-2","NVIDIA Jetson","Multi-Sensor Time Synchronization","CMOS Image Sensor","Autonomous Driving","RGB Camera"]},{"id":"event-camera","category":"perception","sec":0,"tier":3,"sources":[{"title":"Event-based Vision: A Survey (Gallego et al., TPAMI 2020)","url":"https://arxiv.org/abs/1904.08405"},{"title":"Wikipedia: Event camera","url":"https://en.wikipedia.org/wiki/Event_camera"},{"title":"Event-based Agile Object Catching with a Quadrupedal Robot (ICRA 2023)","url":"https://arxiv.org/abs/2303.17479"}],"as_of":"","related_ids":["rgb-camera","global-shutter-rolling-shutter","visual-odometry","optical-flow","exteroception","multi-sensor-fusion"],"name":"Event Camera","alt":"事件相机","abbr":"DVS","aliases":["Dynamic Vision Sensor","DVS","Neuromorphic Camera","Silicon Retina"],"one_liner":"A camera where each pixel independently outputs an “event” only when brightness changes, instead of capturing full frames.","explanation":"An event camera is a biologically inspired vision sensor, also called a dynamic vision sensor (DVS) or neuromorphic camera. An ordinary camera outputs a complete image at a fixed frame rate; each pixel in an event camera works independently and asynchronously, firing an event only when its brightness changes past a threshold, with each event carrying the pixel coordinates, a timestamp, and whether it got brighter or darker (polarity). A 2020 survey by Gallego and colleagues in TPAMI summarizes its advantages: microsecond-level time resolution, a dynamic range around 140 dB (versus about 60 dB for ordinary cameras), low power, and almost no motion blur. The tradeoff is that a static scene produces almost no output, and there is no color or absolute brightness information, so it needs dedicated processing algorithms. It suits high-speed motion and scenes with dramatic lighting changes, such as drone obstacle avoidance, high-speed grasping, and visual odometry.","example":"Researchers at the University of Zurich and ETH Zurich (ICRA 2023) mounted an event camera on a quadruped robot to catch thrown objects, catching items flying in from 4 meters away at up to 15 m/s with an 83% success rate, running the algorithm at 100 Hz on a Jetson Orin.","related":["RGB Camera","Global Shutter / Rolling Shutter","Visual Odometry","Optical Flow","Exteroception","Multi-Sensor Fusion"]},{"id":"thermal-camera","category":"perception","sec":0,"tier":3,"sources":[{"title":"Thermographic camera (Wikipedia)","url":"https://en.wikipedia.org/wiki/Thermographic_camera"}],"as_of":"","related_ids":["inspection-robot","rgb-camera","multi-sensor-fusion","exteroception","special-purpose-robot"],"name":"Thermal Camera","alt":"热成像相机（红外热像仪）","abbr":"","aliases":["Infrared Camera","Thermal Imager","Long-Wave Infrared Camera"],"one_liner":"A camera that captures infrared heat radiation instead of visible light, revealing an object’s temperature distribution.","explanation":"A thermal camera captures the infrared radiation objects emit themselves (long-wave infrared, for most consumer and industrial models), converting temperature differences into an image; uncooled models commonly use a microbolometer as the sensor. Because it doesn’t depend on ambient light, it can form an image in total darkness, smoke, or partial fog, and it is very sensitive to warm-blooded people and animals. In robotics it is commonly used for inspection (finding overheating electrical equipment or pipe leaks), search and rescue, and detecting people for security, and some research fuses it with RGB and depth cameras to improve perception at night. A few caveats: ordinary glass is opaque to long-wave infrared, so a thermal camera can’t see through a window; its resolution is usually lower than an RGB camera’s; and the temperature it reads is affected by the object’s emissivity.","example":"A substation inspection robot scans a switchgear cabinet with a thermal camera and finds one connector running noticeably hotter than its surroundings, automatically flagging a possible loose connection.","related":["Inspection Robot","RGB Camera","Multi-Sensor Fusion","Exteroception","Special-purpose Robot"]},{"id":"3d-vision","category":"perception","sec":1,"tier":2,"sources":[{"title":"Wikipedia: 3D reconstruction","url":"https://en.wikipedia.org/wiki/3D_reconstruction"},{"title":"Wikipedia: Computer stereo vision","url":"https://en.wikipedia.org/wiki/Computer_stereo_vision"}],"as_of":"","related_ids":["point-cloud","depth-camera","stereo-camera","6d-object-pose-estimation","feed-forward-3d-reconstruction","3d-vla"],"name":"3D Vision","alt":"3D视觉","abbr":"","aliases":["3D Perception"],"one_liner":"The umbrella term for techniques that recover an object's or scene's 3D geometry from images or sensor data.","explanation":"3D vision is the branch of computer vision that studies 3D geometry, with the goal of recovering depth, point clouds, meshes, and object poses — in short, where things are and what shape they have. It splits into two families by how the data is captured: active methods emit their own signal to measure distance, such as structured light, time-of-flight cameras, and lidar; passive methods use only an ordinary camera, relying on stereo disparity, multi-view structure from motion, or neural-network monocular depth estimation. Common tasks include depth estimation, 3D reconstruction, point cloud segmentation and registration, 6D pose estimation, and 3D object detection. A robot reaching for, avoiding, or placing an object needs precise distance information that a plain 2D image can't provide, which is why grasp planning, 3D Diffusion Policy, and 3D VLA models all depend on 3D vision; feed-forward reconstruction models such as DUSt3R and VGGT have made it substantially easier to use.","example":"Before an arm grasps a cup, a depth camera turns the scene into a point cloud, an algorithm segments out the cup within that point cloud and estimates the position and orientation of its handle, and that result is converted into the 3D coordinates the gripper needs to reach.","related":["Point Cloud","Depth Camera","Stereo Camera","6D Object Pose Estimation","Feed-Forward 3D Reconstruction","3D VLA"]},{"id":"depth-map","category":"perception","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Depth map","url":"https://en.wikipedia.org/wiki/Depth_map"},{"title":"RealSense SDK 文档：Depth from Stereo（视差转深度与深度单位）","url":"https://github.com/realsenseai/librealsense/blob/master/doc/depth-from-stereo.md"}],"as_of":"","related_ids":["depth-camera","point-cloud","projection-back-projection","depth-holes","monocular-depth-estimation","depth-to-color-alignment"],"name":"Depth Map","alt":"深度图","abbr":"","aliases":["Depth Image","Range Image"],"one_liner":"An image whose pixel values store distance to the camera instead of color.","explanation":"A depth map is a single-channel image where each pixel's value represents the distance from the corresponding scene point to the camera — generally measured along the camera's optical axis (the Z axis), not the straight-line distance to the lens. It can come directly from a depth camera's measurement, or be predicted from an ordinary color image by a monocular depth estimation model such as Depth Anything, which sometimes outputs only relative distance and sometimes metric depth in meters. Depth is commonly stored as a 16-bit integer that must be multiplied by a depth unit to get meters. A depth map combined with the camera's intrinsics can be back-projected, pixel by pixel, into a point cloud. Common problems include holes where nothing was measured (value 0), ‘flying pixels’ at object edges, and a mismatch between the depth map's and the color image's viewpoints, which needs to be corrected by alignment before the two are used together.","example":"If the depth unit is 1 millimeter, a pixel value of 850 in the depth map means that point is 0.85 meters from the camera along the optical axis; a value of 0 usually means nothing was measured there.","related":["Depth Camera","Point Cloud","Projection / Back-Projection","Depth Holes","Monocular Depth Estimation","Depth-to-Color Alignment"]},{"id":"point-cloud","category":"perception","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Point cloud","url":"https://en.wikipedia.org/wiki/Point_cloud"},{"title":"PCL 文档：The PCD (Point Cloud Data) file format","url":"https://pointclouds.org/documentation/tutorials/pcd_file_format.html"},{"title":"3D Diffusion Policy 项目页","url":"https://3d-diffusion-policy.github.io/"}],"as_of":"","related_ids":["depth-map","point-cloud-encoder","farthest-point-sampling","iterative-closest-point","pointnet-pointnet-plus-plus","3d-diffusion-policy"],"name":"Point Cloud","alt":"点云","abbr":"","aliases":["3D Point Cloud","PCD"],"one_liner":"Data made of many 3D coordinate points used to represent the shape of an object or scene.","explanation":"A point cloud is a set of discrete points in 3D space, each carrying X, Y, Z coordinates and optionally color, a surface normal, a timestamp, and so on. It typically comes from lidar, from a depth camera (a depth map back-projected using the camera's intrinsics), or from multi-view 3D reconstruction. Point clouds are unordered, sparse, and variable in size, so ordinary image convolutional networks don't directly apply to them, which is why specialized point cloud networks like PointNet exist. Common processing steps include voxel downsampling or farthest point sampling (both reduce the point count), aligning two clouds with ICP (Iterative Closest Point, a registration method), and segmenting out a target object; the open-source PCL library defines the widely used PCD file format. In embodied AI, 3D Diffusion Policy (DP3) uses a single-view point cloud as its policy input, and grasp-detection methods such as AnyGrasp predict grasp poses directly on point clouds.","example":"A 640×480 depth image can be back-projected into more than 300,000 points at most. For 3D policies, the region outside the tabletop is usually cropped out first, and farthest point sampling then reduces the cloud to a few hundred or a few thousand points before it's fed into the encoder.","related":["Depth Map","Point Cloud Encoder","Farthest Point Sampling","Iterative Closest Point","PointNet / PointNet++","3D Diffusion Policy"]},{"id":"depth-camera","category":"perception","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Range imaging（深度相机的几类原理）","url":"https://en.wikipedia.org/wiki/Range_imaging"},{"title":"RealSense 白皮书：Projectors for D400 series（主动双目与红外投射器）","url":"https://dev.realsenseai.com/docs/projectors"},{"title":"RealSense D435i 产品页","url":"https://www.realsenseai.com/stereo-depth-cameras/depth-camera-d435i/"}],"as_of":"","related_ids":["depth-map","active-stereo","structured-light","time-of-flight","point-cloud","realsense-depth-camera"],"name":"Depth Camera","alt":"深度相机","abbr":"RGB-D","aliases":["RGB-D Camera","3D Camera","Range Camera"],"one_liner":"A camera that outputs, for every pixel, both a color and the distance from that point to the camera.","explanation":"A depth camera, also called an RGB-D camera, outputs a depth map (D) alongside an ordinary color image (RGB). Three ranging principles dominate: stereo vision, where two lenses view the same point and triangulate distance from the disparity between them, often paired with an infrared projector casting a random texture onto the scene to help matching on textureless surfaces like white walls (called active stereo); structured light, which projects a known pattern and computes depth from how it deforms; and time-of-flight (ToF), which measures how long light takes to travel out and back. Depth cameras are cheaper and smaller than lidar and more precise at close range, but error grows with distance, and transparent or mirror-like objects are often measured incorrectly. They're commonly mounted on a robot's head or wrist to produce point clouds and support grasping and obstacle avoidance; common brands include RealSense, Orbbec, and ZED.","example":"The RealSense D435i is an active-stereo camera officially rated for an ideal working range of 0.3–3 meters, with error under 2% at 2 meters, and is widely used in tabletop manipulation and mobile-robot research.","related":["Depth Map","Active Stereo","Structured Light","Time of Flight","Point Cloud","RealSense Depth Camera (D435i / D405)"]},{"id":"stereo-camera","category":"perception","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Computer stereo vision","url":"https://en.wikipedia.org/wiki/Computer_stereo_vision"},{"title":"Stereolabs ZED 2i product page","url":"https://www.stereolabs.com/products/zed-2"}],"as_of":"","related_ids":["disparity","stereo-matching","stereo-baseline","active-stereo","depth-camera","foundationstereo"],"name":"Stereo Camera","alt":"双目相机","abbr":"","aliases":["Stereo Vision","Passive Stereo"],"one_liner":"Two side-by-side cameras that shoot the same scene at once, computing distance from how much the two images differ.","explanation":"A stereo camera mounts two cameras side by side at a fixed separation, called the baseline, mimicking human eyes. The same physical point appears at a slightly different horizontal position in the left and right images — the disparity — and after calibration and epipolar rectification, stereo matching finds corresponding points and computes depth as Z = f·B/d, where f is focal length, B is the baseline, and d is the disparity: a larger disparity means closer, and a longer baseline gives more accurate readings at range. Stereo that relies only on ambient light is called passive stereo, and it struggles on textureless surfaces such as white walls or smooth tabletops where matching fails; active stereo cameras such as the RealSense D435 and D455 add an infrared speckle projector to work around this. Recent work also uses deep networks for the matching step, such as FoundationStereo. Stereo cameras commonly serve as a robot's head- or wrist-mounted depth sensor.","example":"A stereo camera with a 700-pixel focal length and a 0.12-meter baseline observes a 42-pixel disparity at some point, giving a depth of Z = 700 × 0.12 ÷ 42 = 2 meters. Stereolabs' ZED 2i has a baseline of 120 mm.","related":["Disparity","Stereo Matching","Stereo Baseline","Active Stereo","Depth Camera","FoundationStereo"]},{"id":"disparity","category":"perception","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Computer stereo vision","url":"https://en.wikipedia.org/wiki/Computer_stereo_vision"}],"as_of":"","related_ids":["stereo-camera","stereo-matching","stereo-baseline","depth-map","epipolar-geometry","foundationstereo"],"name":"Disparity","alt":"视差","abbr":"","aliases":["Disparity Map","Stereo Disparity"],"one_liner":"The horizontal position difference of the same point between the left and right images — bigger for closer objects.","explanation":"Disparity is the fundamental quantity in binocular stereo vision: with two cameras mounted side by side and the images rectified (so the same point falls on the same row in both), the difference d = xL − xR between a point’s x-coordinate in the left image and its x-coordinate in the right image is its disparity. Disparity is inversely proportional to depth: Z = f × B / d, where f is the focal length in pixels and B is the baseline distance between the two cameras. Computing disparity pixel by pixel gives a disparity map, which converts to a depth map with this formula. The process of computing disparity is called stereo matching; classic methods include OpenCV’s StereoBM and SGBM, and recent deep-learning methods include RAFT-Stereo and FoundationStereo. Because of the inverse relationship, distant objects have a disparity of only a few pixels or less, so a small matching error causes a large depth error — stereo ranging accuracy drops off quickly with distance, and lengthening the baseline improves accuracy at range.","example":"With focal length f = 640 pixels and baseline B = 5 centimeters, a point with disparity 16 pixels has depth Z = 640 × 0.05 / 16 = 2 meters; if the disparity is off by 1 pixel to 15, the computed depth becomes about 2.13 meters.","related":["Stereo Camera","Stereo Matching","Stereo Baseline","Depth Map","Epipolar Geometry","FoundationStereo"]},{"id":"stereo-baseline","category":"perception","sec":1,"tier":3,"sources":[{"title":"Depth Map from Stereo Images - OpenCV","url":"https://docs.opencv.org/4.x/dd/d53/tutorial_py_depthmap.html"}],"as_of":"","related_ids":["stereo-camera","disparity","stereo-matching","triangulation","camera-extrinsics","depth-camera"],"name":"Stereo Baseline","alt":"基线","abbr":"","aliases":["Baseline"],"one_liner":"The distance between a stereo camera’s two lens centers, which sets how far and how accurately it can measure depth.","explanation":"The baseline is the distance between the optical centers of a stereo camera’s left and right lenses. The stereo depth formula is depth = focal length × baseline / disparity, where disparity is the horizontal pixel difference between the same point as seen in the left and right images. A longer baseline gives more disparity at a given depth, so far-away measurements are more accurate, but the close-range blind zone grows and the overlap between the left and right images shrinks; a short baseline suits close range instead. So a camera should be chosen by baseline to match its working distance: wrist cameras usually use a short baseline, while mobile robots looking far ahead use a longer one. Depth error grows roughly with the square of distance and is inversely proportional to the baseline. In structured-light and active-stereo cameras, the distance between the projector and the camera plays a similar role.","example":"The ZED 2i has a baseline of about 12 cm, suited to scenes a few meters to over ten meters away; the wrist-mounted RealSense D405 has a very short baseline, suited to close range within tens of centimeters.","related":["Stereo Camera","Disparity","Stereo Matching","Triangulation","Camera Extrinsics","Depth Camera"]},{"id":"stereo-matching","category":"perception","sec":1,"tier":3,"sources":[{"title":"Middlebury Stereo Vision Page","url":"https://vision.middlebury.edu/stereo/"}],"as_of":"","related_ids":["disparity","stereo-baseline","stereo-camera","epipolar-geometry","foundationstereo","depth-estimation"],"name":"Stereo Matching","alt":"立体匹配","abbr":"","aliases":["Stereo Depth Estimation"],"one_liner":"Finding the same point’s position in both the left and right images to compute disparity, then converting that into depth.","explanation":"Stereo matching is the core step of stereo depth sensing: the left and right images are first rectified onto the same horizontal lines (epipolar rectification), and then, for each pixel in the left image, the most similar point along the same row of the right image is found; the horizontal distance between them is the disparity, which is converted to depth via depth = focal length × baseline / disparity. Classic methods include block matching and semi-global matching (SGM); deep-learning methods build a cost volume with a neural network or refine it iteratively, as in RAFT-Stereo, and recent models such as FoundationStereo can generalize zero-shot to new scenes. The main difficulties are texture-less regions, reflective and transparent surfaces, and occlusion boundaries. Stereo matching determines the quality of a stereo camera’s depth map, and it’s often the part that gets replaced or improved when a robot needs to grasp transparent objects.","example":"Feeding the left and right infrared images from a ZED or RealSense camera into FoundationStereo produces a more complete depth map than the camera’s own built-in algorithm.","related":["Disparity","Stereo Baseline","Stereo Camera","Epipolar Geometry","FoundationStereo","Depth Estimation"]},{"id":"active-stereo","category":"perception","sec":1,"tier":3,"sources":[{"title":"Intel RealSense Stereoscopic Depth Cameras (Keselman et al., arXiv:1705.05548)","url":"https://arxiv.org/abs/1705.05548"},{"title":"Orbbec Gemini 335 产品页","url":"https://www.orbbec.com/products/stereo-vision-camera/gemini-335/"}],"as_of":"","related_ids":["stereo-camera","structured-light","speckle-structured-light","depth-camera","stereo-matching","realsense-depth-camera"],"name":"Active Stereo","alt":"主动双目","abbr":"","aliases":["IR-Projected Stereo","Projected Texture Stereo"],"one_liner":"A stereo camera pair with an infrared projector adding a speckle pattern to help match the left and right images and compute depth.","explanation":"Active stereo is a depth-camera design: two infrared cameras triangulate depth from disparity just like ordinary stereo, while an infrared projector simultaneously casts a random dot pattern onto the scene. Passive (ordinary) stereo matches the two images using natural texture, and fails on texture-free surfaces like blank walls or plain tabletops, leaving holes in the depth map; the projected speckle pattern artificially adds texture to these surfaces so matching is no longer ambiguous. Intel RealSense’s R200 and D400 series use this approach; Intel’s technical paper notes that, unlike structured light, the projected pattern doesn’t need to be known in advance — it only needs to be dense and non-repeating along the matching direction. Because the cameras can also exploit natural texture lit by sunlight, this kind of camera still works outdoors. Orbbec’s Gemini 330 series is also marketed as combining active and passive stereo.","example":"A camera like the RealSense D435 photographing a plain white table gets a depth map full of holes with the infrared projector off; turning the projector on covers the table in a speckle pattern, and the holes disappear.","related":["Stereo Camera","Structured Light","Speckle Structured Light","Depth Camera","Stereo Matching","RealSense Depth Camera (D435i / D405)"]},{"id":"realsense-depth-camera","category":"perception","sec":1,"tier":1,"sources":[{"title":"PR Newswire: Cognex to Acquire RealSense（2026-09-22）","url":"https://www.prnewswire.com/news-releases/cognex-to-acquire-realsense-expanding-machine-vision-leadership-into-high-growth-robotic-perception-market-302885738.html"},{"title":"RealSense completes spin-out from Intel, raises $50 million","url":"https://www.realsenseai.com/cn/news-insights/news/realsense-completes-spin-out-from-intel-raises-50-million-to-accelerate-ai-powered-vision-for-robotics-and-biometrics/"},{"title":"RealSense D435i 产品页","url":"https://www.realsenseai.com/stereo-depth-cameras/depth-camera-d435i/"}],"as_of":"2026-09","related_ids":["depth-camera","active-stereo","inertial-measurement-unit","intel-realsense-sdk-2-0","orbbec-gemini-330-series","wrist-camera"],"name":"RealSense Depth Camera (D435i / D405)","alt":"RealSense 深度相机（D435i / D405）","abbr":"","aliases":["Intel RealSense","D435","D435i","D455","D405","D555"],"one_liner":"A widely used stereo depth camera series, formerly an Intel division, common in robotics research and products.","explanation":"RealSense grew out of an Intel 3D-camera project started in 2014; its D400 series uses stereo vision plus an onboard chip to compute depth directly. It spun off from Intel as an independent company on July 11, 2025, completing a $50 million Series A round, and on September 22, 2026, the machine-vision company Cognex announced it would acquire RealSense for roughly $500 million in cash, with the deal expected to close in the fourth quarter of that year. Two models are most common in embodied-AI circles: the D435i is an active-stereo camera with an IR projector and a built-in IMU, with an ideal range of 0.3–3 meters, often mounted on a robot's head or used as a third-person view; the D405 is designed specifically for close range, 7–50 cm, and is often wrist-mounted instead. The open-source RealSense SDK provides drivers, depth filtering, and a ROS 2 interface.","example":"The ALOHA 2 dual-arm platform replaced the consumer webcams used in its first generation with the RealSense D405, citing its wider field of view, depth sensing, global shutter, and smaller size.","related":["Depth Camera","Active Stereo","Inertial Measurement Unit","Intel RealSense SDK 2.0 (librealsense)","Orbbec Gemini 330 Series","Wrist Camera"]},{"id":"orbbec-gemini-330-series","category":"perception","sec":1,"tier":2,"sources":[{"title":"Orbbec Unveils Gemini 330 Series of Stereo Vision 3D Cameras（2024-04-30）","url":"https://www.orbbec.com/news/orbbec-unveils-gemini-330-series-of-stereo-vision-3d-cameras-powered-by-latest-asic-for-outdoor-and-indoor-performance/"},{"title":"Orbbec Gemini 335 产品页","url":"https://www.orbbec.com/products/stereo-vision-camera/gemini-335/"},{"title":"Orbbec Gemini 335L 产品页","url":"https://www.orbbec.com/products/stereo-vision-camera/gemini-335l/"}],"as_of":"2026-03","related_ids":["depth-camera","active-stereo","stereo-camera","realsense-depth-camera","orbbec","inertial-measurement-unit"],"name":"Orbbec Gemini 330 Series","alt":"奥比中光 Gemini 330 系列","abbr":"","aliases":["Gemini 335","Gemini 335L","Gemini 336","Gemini 336L","Gemini 335Lg"],"one_liner":"A stereo depth camera series from Orbbec that works indoors and out, commonly mounted on robots.","explanation":"The Gemini 330 series is a family of stereo depth cameras from Orbbec, released on April 30, 2024, with the first models being the Gemini 335 and 335L, followed later by the 336, 336L, and the GMSL2-interface 335Lg. It uses a hybrid of active and passive stereo: it can do passive matching using ambient light alone, but also has an IR projector to add texture onto surfaces that otherwise lack it. Depth is computed onboard by Orbbec's own MX6800 chip, a single USB cable handles both power and data, it has a built-in IMU, and it supports hardware-triggered synchronization. The 335 has a 50 mm baseline with a best range of 0.26–3 meters; the 335L has a 95 mm baseline with a best range of 0.25–6 meters and is IP65-rated; the 336 adds an infrared-only filter, suited to strong light and reflective scenes. It was integrated into NVIDIA's Isaac Perceptor in June 2024, and according to an Orbbec press release from March 2026, Honor's first humanoid robot uses this camera series.","example":"Orbbec's development kit, paired with NVIDIA Isaac Perceptor, hardware-synchronizes 4 Gemini 335L units to stitch together 360-degree depth perception around an autonomous mobile robot (AMR).","related":["Depth Camera","Active Stereo","Stereo Camera","RealSense Depth Camera (D435i / D405)","Orbbec","Inertial Measurement Unit"]},{"id":"stereolabs-zed","category":"perception","sec":1,"tier":3,"sources":[{"title":"ZED 2i | Stereo Camera | Stereolabs","url":"https://www.stereolabs.com/store/products/zed-2i"},{"title":"ZED X - AI Stereo Camera for Robotics | Stereolabs","url":"https://www.stereolabs.com/products/zed-x"}],"as_of":"2026-09","related_ids":["stereo-camera","stereolabs","stereo-baseline","stereo-matching","gigabit-multimedia-serial-link","realsense-depth-camera"],"name":"Stereolabs ZED","alt":"ZED 双目相机","abbr":"","aliases":["ZED 2i","ZED X","ZED Mini"],"one_liner":"A passive stereo depth camera line from the French company Stereolabs, common in mobile robots and data collection.","explanation":"ZED is Stereolabs’ stereo camera product line, which computes depth through stereo matching using two RGB lenses rather than projecting any active light, so it keeps working outdoors in sunlight and reaches farther range than a structured-light camera. The ZED 2i has a 12 cm baseline, connects over USB, and includes a built-in IMU, barometer, and magnetometer, with IP66 protection; the ZED X targets robotics, using a global shutter and a GMSL2 interface (an automotive-grade camera serial link well suited to connecting with a Jetson), with IP67 protection; the ZED Mini has a shorter baseline suited to close range. The companion ZED SDK provides depth maps, point clouds, visual-inertial positioning, and object detection, with a ROS 2 driver available. In embodied AI, it is commonly used as a head-mounted or third-person camera, and also for teleoperation and mobile-robot mapping.","example":"A wheeled humanoid robot carries a ZED X on its head, connected to a Jetson Orin, which outputs a point cloud to the navigation and grasping modules.","related":["Stereo Camera","Stereolabs","Stereo Baseline","Stereo Matching","Gigabit Multimedia Serial Link","RealSense Depth Camera (D435i / D405)"]},{"id":"luxonis-oak-d","category":"perception","sec":1,"tier":3,"sources":[{"title":"Luxonis Docs: OAK-D","url":"https://docs.luxonis.com/hardware/products/OAK-D"},{"title":"Luxonis Docs: RVC4 平台（OAK 4 系列）","url":"https://docs.luxonis.com/hardware/platform/rvc/rvc4"},{"title":"GitHub: luxonis/depthai-core","url":"https://github.com/luxonis/depthai-core"}],"as_of":"2026-09","related_ids":["depth-camera","stereo-camera","realsense-depth-camera","orbbec-gemini-330-series","stereolabs-zed","on-device-edge-deployment"],"name":"Luxonis OAK-D","alt":"Luxonis OAK-D 相机","abbr":"","aliases":["OAK 4","OpenCV AI Kit","DepthAI Camera"],"one_liner":"A stereo depth camera from Luxonis with a built-in AI chip, letting neural networks run directly on the camera.","explanation":"OAK is Luxonis’s camera product line, with the OAK-D as its most common model: two global-shutter monochrome cameras do stereo depth sensing (a 7.5 cm baseline, with an official ideal range of about 0.8–12 meters), a color camera sits in the middle, a 9-axis IMU is built in, and it connects to a computer over USB. Its distinguishing feature is an onboard compute chip (the RVC2 platform, rated at 4 TOPS, with 1.4 TOPS available for AI), so depth computation and neural networks like object detection can run entirely on the camera, sending only the results to the host — well suited to small robots with limited onboard compute. The newer OAK 4 series switches to a Qualcomm QCS8550 chip (the RVC4 platform, 48 INT8 TOPS) with a 6-core ARM CPU running Linux, letting it operate independently without a host computer. Its companion open-source library is called DepthAI (written in C++, with Python bindings, MIT licensed), with an official ROS driver, depthai-ros, also available.","example":"Mounting an OAK-D on a low-cost mobile robot, the official ROS driver publishes color and depth images to ROS topics while an object-detection network runs on the camera itself, so the host only receives detection results, saving onboard compute.","related":["Depth Camera","Stereo Camera","RealSense Depth Camera (D435i / D405)","Orbbec Gemini 330 Series","Stereolabs ZED","On-Device / Edge Deployment"]},{"id":"structured-light","category":"perception","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Structured-light 3D scanner","url":"https://en.wikipedia.org/wiki/Structured-light_3D_scanner"},{"title":"Wikipedia: Kinect（初代结构光与 Xbox One 版 ToF）","url":"https://en.wikipedia.org/wiki/Kinect"}],"as_of":"","related_ids":["speckle-structured-light","active-stereo","time-of-flight","depth-camera","triangulation","transparent-and-reflective-object-perception"],"name":"Structured Light","alt":"结构光","abbr":"","aliases":["Structured-Light Camera","Coded Structured Light","Structured-Light 3D Scanning"],"one_liner":"Projects a known pattern onto a scene and computes depth from how the pattern gets distorted.","explanation":"Structured light is an active 3D-sensing method: a projector casts a known pattern — stripes, a coded pattern, or an infrared dot array — onto the scene, and a camera views it from a different angle. Surface bumps and dips distort the pattern, and knowing the relative position of the projector and camera lets the system triangulate depth at each point. It doesn’t depend on surface texture, so it works even on a blank white wall, and it’s accurate at close range. Well-known examples include the original 2010 Kinect, which projects a near-infrared dot pattern, and the iPhone’s Face ID, which projects over 30,000 infrared dots. Reflective or transparent surfaces make the pattern disappear, and strong sunlight washes out its contrast, so structured light is mostly used indoors at short-to-medium range. Active stereo also projects a speckle pattern, but computes depth through stereo matching between two cameras — that’s the key difference between the two methods.","example":"The original Kinect projects a near-infrared dot pattern across a living room; its infrared camera captures how the dots shift around people and furniture, and from that shift it computes per-pixel depth to generate a depth map for motion-controlled games.","related":["Speckle Structured Light","Active Stereo","Time of Flight","Depth Camera","Triangulation","Transparent & Reflective Object Perception"]},{"id":"speckle-structured-light","category":"perception","sec":1,"tier":3,"sources":[{"title":"Structured-light 3D scanner - Wikipedia","url":"https://en.wikipedia.org/wiki/Structured-light_3D_scanner"}],"as_of":"","related_ids":["structured-light","active-stereo","depth-camera","triangulation","time-of-flight"],"name":"Speckle Structured Light","alt":"散斑结构光","abbr":"","aliases":["Laser Speckle Pattern","Speckle Projection"],"one_liner":"Projecting a random infrared speckle pattern onto a scene and computing depth from how it deforms or matches.","explanation":"Speckle structured light is a form of structured light: an infrared laser, passed through a diffractive element, projects a pseudo-random pattern of speckles; a camera photographs where the speckles land on the object’s surface, and comparing this against a pre-calibrated reference pattern gives the depth of every point via triangulation. Because each small patch of the speckle pattern is unique, matching can be done from a single frame, which makes it well suited to dynamic scenes. The original Kinect (using PrimeSense’s design) and many phone face-recognition modules use this approach. Active stereo cameras (such as the RealSense D400 series) also project a speckle pattern, but there it serves to “add texture” to an otherwise texture-less surface, with depth still computed through stereo matching. Its drawback is that the speckle pattern is easily washed out in strong outdoor sunlight, and accuracy drops at longer range.","example":"The original Microsoft Kinect outputs a depth map using infrared speckle projection combined with a single infrared camera.","related":["Structured Light","Active Stereo","Depth Camera","Triangulation","Time of Flight"]},{"id":"laser-triangulation","category":"perception","sec":1,"tier":3,"sources":[{"title":"Wikipedia: 3D scanning（Triangulation 一节）","url":"https://en.wikipedia.org/wiki/3D_scanning"},{"title":"KEYENCE LJ-X8000 2D/3D Laser Profiler","url":"https://www.keyence.com/products/measure/laser-2d/lj-x8000/"},{"title":"Micro-Epsilon scanCONTROL Laser Profile Scanners","url":"https://www.micro-epsilon.com/2d-3d-measurement/laser-profile-scanners"}],"as_of":"2026-09","related_ids":["structured-light","triangulation","machine-vision","3d-vision-guided-robotics","point-cloud","time-of-flight"],"name":"Laser Triangulation","alt":"激光三角测量（线激光轮廓扫描）","abbr":"","aliases":["Line-Laser Profiler","Laser Profilometer","Laser Profile Scanner"],"one_liner":"A laser projects a dot or line, a camera views the spot from the side, and triangulation gives distance and profile.","explanation":"Laser triangulation is an active optical ranging method: a laser projects a dot or a line onto an object, and a camera set a known distance away from the laser, viewing from a different angle, captures where the spot lands; the laser, camera, and spot form a triangle, and the spot’s position in the image gives the distance. Swapping the dot for a line turns this into line-laser profile scanning, which captures an entire cross-sectional profile in one shot; moving the object on a conveyor belt, or sweeping the sensor with a robot arm, stitches these profiles into a 3D point cloud. It is highly accurate, down to tens of micrometers, but its range is usually limited to within a few meters, and it struggles with occlusion and highly reflective surfaces. Industrially it is used for dimensional measurement, defect detection, and weld-seam tracking, a common approach in machine vision and 3D vision-guided robotics; structured-light cameras are based on the same triangulation principle.","example":"A welding robot mounts a line-laser profile scanner in front of its torch (such as the Micro-Epsilon scanCONTROL 8x00, with 4,224 points per profile at a 10 kHz profile rate), measuring the weld seam’s cross-section position in real time and guiding the arm to correct its path along the seam.","related":["Structured Light","Triangulation","Machine Vision","3D Vision-Guided Robotics","Point Cloud","Time of Flight"]},{"id":"time-of-flight","category":"perception","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Time-of-flight camera","url":"https://en.wikipedia.org/wiki/Time-of-flight_camera"},{"title":"Azure Kinect DK depth camera（Microsoft Learn）","url":"https://learn.microsoft.com/en-us/previous-versions/azure/kinect-dk/depth-camera"}],"as_of":"","related_ids":["direct-time-of-flight","indirect-time-of-flight","lidar","structured-light","flying-pixels","depth-camera"],"name":"Time of Flight","alt":"飞行时间法","abbr":"ToF","aliases":["ToF","ToF Camera","Time-of-Flight Camera"],"one_liner":"Measures how long light takes to bounce off an object and return, then converts that time into distance.","explanation":"Time of flight measures distance from the round-trip travel time of light: distance equals the speed of light times the round-trip time, divided by two. A ToF camera carries its own near-infrared light source and measures a distance at every pixel, producing a depth map directly. Direct time of flight (dToF) sends short pulses and times them directly — most lidar uses this; indirect time of flight (iToF) sends modulated light and measures the phase shift of the returning signal to compute distance, which is how the Kinect for Xbox One and Azure Kinect work. ToF doesn’t depend on surface texture and the module is compact; common problems include multipath interference (light bouncing repeatedly in corners, inflating the measured distance), flying pixels at object edges, and weak returns under strong sunlight or from dark surfaces. Robots commonly use it for close-range obstacle avoidance and tabletop depth sensing.","example":"Azure Kinect uses amplitude-modulated continuous-wave iToF: it emits modulated near-infrared light and computes depth from the phase of the returning signal; when it images a corner, the light bounces back and forth between the two walls, so those pixels get marked invalid with a depth of zero.","related":["Direct Time of Flight","Indirect Time of Flight","LiDAR","Structured Light","Flying Pixels","Depth Camera"]},{"id":"direct-time-of-flight","category":"perception","sec":1,"tier":3,"sources":[{"title":"Sony Semiconductor Solutions: ToF image sensors (dToF vs iToF)","url":"https://www.sony-semicon.com/en/technology/industry/tof.html"},{"title":"Wikipedia: Time-of-flight camera","url":"https://en.wikipedia.org/wiki/Time-of-flight_camera"},{"title":"Apple Newsroom: Apple unveils new iPad Pro with LiDAR Scanner (2020-03)","url":"https://www.apple.com/newsroom/2020/03/apple-unveils-new-ipad-pro-with-lidar-scanner-and-trackpad-support-in-ipados/"}],"as_of":"","related_ids":["time-of-flight","indirect-time-of-flight","single-photon-avalanche-diode","lidar","depth-camera","vertical-cavity-surface-emitting-laser"],"name":"Direct Time of Flight","alt":"直接飞行时间","abbr":"dToF","aliases":["dToF","Pulsed ToF"],"one_liner":"Emits a light pulse and directly times its echo, computing distance from the round-trip time.","explanation":"Direct time of flight is one variant of time-of-flight (ToF) sensing: the sensor emits an extremely short laser pulse and directly measures the time t it takes for the light to hit an object and bounce back, with distance equal to the speed of light times t, divided by two. For a target 1 meter away, the round trip takes only about 6.7 nanoseconds, so the circuitry needs to resolve extremely short time intervals: modern designs commonly use single-photon avalanche diodes (SPADs, which can trigger a signal from a single photon) as pixels, paired with a time-to-digital converter that times a large number of pulses and bins them into a histogram, taking the peak as the echo time. Compared with indirect time of flight (iToF, which computes distance from the phase difference of modulated light), dToF can measure farther and is more robust to ambient light, working both indoors and outdoors, and it’s the ranging principle behind many lidars; the cost is more complex pixel circuitry, and resolution is usually lower than iToF’s.","example":"Apple’s 2020 iPad Pro LiDAR Scanner measures distances up to 5 meters, indoors or outdoors, and Apple describes it as operating at the photon level and nanosecond speeds; reports indicate it uses a SPAD-based dToF design.","related":["Time of Flight","Indirect Time of Flight","Single-Photon Avalanche Diode","LiDAR","Depth Camera","Vertical-Cavity Surface-Emitting Laser"]},{"id":"indirect-time-of-flight","category":"perception","sec":1,"tier":3,"sources":[{"title":"Wikipedia: Time-of-flight camera","url":"https://en.wikipedia.org/wiki/Time-of-flight_camera"},{"title":"Orbbec Femto Bolt 产品页","url":"https://www.orbbec.com/products/tof-camera/femto-bolt/"}],"as_of":"","related_ids":["time-of-flight","direct-time-of-flight","depth-camera","flying-pixels","orbbec-femto-bolt","azure-kinect-dk"],"name":"Indirect Time of Flight","alt":"间接飞行时间","abbr":"iToF","aliases":["iToF","Phase-Based ToF","Continuous-Wave ToF","AMCW ToF"],"one_liner":"A depth-sensing method that emits modulated light and computes distance from the phase shift of its echo.","explanation":"Indirect time of flight (iToF) is one variant of time-of-flight sensing: the light source emits continuous infrared light modulated at a set frequency, and each pixel in the sensor measures the phase difference between the echo and the emitted signal, converting that phase into a round-trip time and then a distance. Compared with dToF, which times a light pulse’s round trip directly, iToF has a simpler pixel structure, making it easier to achieve a high-resolution depth map with good close-range accuracy, so it is commonly used in depth cameras and phones. It has several limitations: the phase repeats every 2π, so an object beyond the unambiguous range (c/2f, where f is the modulation frequency) gets computed as if it were closer, requiring multiple modulation frequencies combined to resolve the ambiguity; strong sunlight can swamp the modulated signal, so outdoor performance is poor; and multipath interference, where light bounces between multiple surfaces, plus flying pixels where foreground and background mix in the same pixel at an object’s edge, both introduce error.","example":"Both the Microsoft Azure Kinect and the Orbbec Femto Bolt use Microsoft’s iToF depth solution: a 1-megapixel depth sensor, up to 1024×1024 at 15 fps in wide-field-of-view mode, with a working range of roughly 0.25–5.46 m depending on the mode.","related":["Time of Flight","Direct Time of Flight","Depth Camera","Flying Pixels","Orbbec Femto Bolt","Azure Kinect DK"]},{"id":"azure-kinect-dk","category":"perception","sec":1,"tier":3,"sources":[{"title":"Microsoft ending production of Azure Kinect Developer Kit - The Robot Report","url":"https://www.therobotreport.com/microsoft-ending-production-of-azure-kinect-developer-kit/"},{"title":"Femto Bolt Comparison with Azure Kinect DK - Orbbec","url":"https://www.orbbec.com/documentation/comparison-with-azure-kinect-dk/"},{"title":"Azure Kinect - Wikipedia","url":"https://en.wikipedia.org/wiki/Azure_Kinect"}],"as_of":"2026-09","related_ids":["depth-camera","indirect-time-of-flight","orbbec-femto-bolt","human-pose-estimation","microphone-array","inertial-measurement-unit"],"name":"Azure Kinect DK","alt":"Azure Kinect 深度相机","abbr":"","aliases":["Kinect","Microsoft Kinect","Microsoft Azure Kinect DK"],"one_liner":"A developer-oriented RGB-D camera from Microsoft that combines a depth sensor, color camera, microphone array, and IMU in one device.","explanation":"Azure Kinect DK is a developer kit Microsoft released in 2019. It packs a time-of-flight (ToF) depth camera — which measures distance from how long light takes to bounce back — a high-resolution color camera, a 7-microphone array, and an IMU (inertial measurement unit) into a single unit, along with a sensor SDK and a body-tracking SDK. Its depth data is high quality and comes with skeletal tracking, which made it a de facto standard RGB-D sensor in robotics and human-motion research. Microsoft announced in August 2023 that it was discontinuing the product. Orbbec’s Femto Bolt uses the same indirect time-of-flight (iToF) depth technology and is compatible with the Azure Kinect SDK interface; Microsoft has pointed to it as a replacement.","example":"The Azure Kinect body-tracking SDK reads out real-time skeletal joint positions, which a human-robot interaction system uses for pose recognition.","related":["Depth Camera","Indirect Time of Flight","Orbbec Femto Bolt","Human Pose Estimation","Microphone Array","Inertial Measurement Unit"]},{"id":"orbbec-femto-bolt","category":"perception","sec":1,"tier":3,"sources":[{"title":"Orbbec Femto Bolt 产品页","url":"https://www.orbbec.com/products/tof-camera/femto-bolt/"},{"title":"orbbec/OrbbecSDK-K4A-Wrapper (GitHub)","url":"https://github.com/orbbec/OrbbecSDK-K4A-Wrapper"}],"as_of":"2026-09","related_ids":["azure-kinect-dk","indirect-time-of-flight","depth-camera","orbbec","orbbec-gemini-330-series","third-person-camera"],"name":"Orbbec Femto Bolt","alt":"奥比中光 Femto Bolt","abbr":"","aliases":["Femto Bolt"],"one_liner":"An indirect time-of-flight RGB-D camera from Orbbec, positioned as a drop-in replacement for Microsoft’s discontinued Azure Kinect.","explanation":"Femto Bolt is an RGB-D depth camera from Orbbec that uses Microsoft’s indirect time-of-flight (iToF) technology, which computes distance from the phase difference between emitted infrared light and its reflection. Orbbec states that Femto Bolt’s depth modes and performance match the Microsoft Azure Kinect DK, and it ships with a K4A-compatible wrapper SDK, so code written for Azure Kinect needs little to no change. Depth modes include a narrow field of view at 640×576 @ 30fps and a wide field of view up to 1024×1024 @ 15fps, with a range of about 0.25–5.46 meters; color goes up to 4K, and it has a built-in 6-axis IMU. It is powered and connected over USB-C and is intended for indoor use only. Labs commonly mount it as a fixed third-person camera for collecting point clouds.","example":"A lab replaces its old Azure Kinect with a Femto Bolt; after installing the official K4A wrapper, the existing point-cloud capture scripts keep running unchanged.","related":["Azure Kinect DK","Indirect Time of Flight","Depth Camera","Orbbec","Orbbec Gemini 330 Series","Third-Person Camera"]},{"id":"vertical-cavity-surface-emitting-laser","category":"perception","sec":1,"tier":3,"sources":[{"title":"Vertical-cavity surface-emitting laser - Wikipedia","url":"https://en.wikipedia.org/wiki/Vertical-cavity_surface-emitting_laser"}],"as_of":"","related_ids":["structured-light","speckle-structured-light","time-of-flight","lidar","single-photon-avalanche-diode","depth-camera"],"name":"Vertical-Cavity Surface-Emitting Laser","alt":"VCSEL（垂直腔面发射激光器）","abbr":"VCSEL","aliases":["VCSEL","Surface-Emitting Laser"],"one_liner":"A semiconductor laser that emits light perpendicular to the chip surface, a common light source in depth cameras and lidar.","explanation":"The VCSEL is a type of semiconductor laser, proposed in 1977 by Kenichi Iga at the Tokyo Institute of Technology. An ordinary edge-emitting laser emits light from the side of a cleaved chip; a VCSEL instead sandwiches a light-emitting layer between two distributed Bragg reflector mirrors, top and bottom, so light exits perpendicular to the chip surface. This means it can be tested across a whole wafer before dicing, tens of thousands can be produced from one wafer at once, and it’s easy to arrange into a 2D array — giving low cost and good consistency. In embodied-AI-relevant hardware, VCSELs are the core of many active infrared light sources: the structured-light dot projector behind iPhone Face ID, time-of-flight (ToF) depth cameras, and lidar in phones and cars all commonly use VCSELs as their light source.","example":"Inside a structured-light depth camera, a VCSEL array projects thousands of infrared dots onto an object, and the camera computes depth at each point from how far its dot has shifted.","related":["Structured Light","Speckle Structured Light","Time of Flight","LiDAR","Single-Photon Avalanche Diode","Depth Camera"]},{"id":"single-photon-avalanche-diode","category":"perception","sec":1,"tier":3,"sources":[{"title":"Single-photon avalanche diode - Wikipedia","url":"https://en.wikipedia.org/wiki/Single-photon_avalanche_diode"}],"as_of":"","related_ids":["direct-time-of-flight","solid-state-lidar","vertical-cavity-surface-emitting-laser","lidar","adaps-photonics"],"name":"Single-Photon Avalanche Diode","alt":"SPAD（单光子雪崩二极管）","abbr":"SPAD","aliases":["SPAD"],"one_liner":"A photodetector sensitive enough to register a single photon, at the core of dToF lidar and ranging chips.","explanation":"A SPAD is a photodiode operated in Geiger mode (reverse-biased above its breakdown voltage), where a single incoming photon can trigger an avalanche of current, recording the precise moment the photon arrived. It is the key component behind direct time-of-flight (dToF) ranging, which fires a laser pulse and directly measures the round-trip time of its echo: paired with time-to-digital converter circuitry, building a histogram of when large numbers of photons arrive lets the system compute distance. SPAD arrays let lidar be built as chip-based, solid-state devices, lowering cost and size, which is why they show up in automotive lidar, phone ranging modules, and the small dToF sensors used in robotics.","example":"Many solid-state and hybrid solid-state lidars use a SPAD array chip on the receiving side, paired with a VCSEL emitter, to do dToF ranging.","related":["Direct Time of Flight","Solid-State LiDAR","Vertical-Cavity Surface-Emitting Laser","LiDAR","Adaps Photonics"]},{"id":"robosense-ac1","category":"perception","sec":1,"tier":3,"sources":[{"title":"RoboSense AC1 产品页（英文）","url":"https://www.robosense.ai/en/rslidar/AC1"},{"title":"速腾聚创 AC1 产品页（中文）","url":"https://www.robosense.cn/rslidar/AC1"},{"title":"RoboSense-Robotics GitHub（AC 驱动、标定、SLAM 开源仓库）","url":"https://github.com/RoboSense-Robotics"}],"as_of":"2026-09","related_ids":["depth-camera","direct-time-of-flight","single-photon-avalanche-diode","inertial-measurement-unit","multi-sensor-fusion","robosense"],"name":"RoboSense AC1","alt":"速腾聚创 AC1 主动相机","abbr":"","aliases":["RoboSense AC1 Active Camera","Active Camera AC1"],"one_liner":"A robot vision sensor from RoboSense that combines active depth sensing, a color camera, and an IMU in one module.","explanation":"The AC1 is the first product in lidar maker RoboSense’s “Active Camera” line. It packages a fully solid-state active depth-sensing module (a VCSEL laser emitter paired with an SPAD single-photon detector chip), an RGB color camera, and an IMU (inertial measurement unit) into one unit, with synchronization and fusion done in hardware, so it outputs already-aligned depth, image, and pose data. Per the published specs, it ranges up to 70 meters, has roughly 1 cm depth accuracy within 5 meters, a depth field of view of 120°×60°, and works in bright light up to 100 kLux. It targets two problems: the hassle of calibrating and synchronizing a lidar, a camera, and an IMU bolted together separately, and the fact that ordinary depth cameras fail outdoors in strong sunlight. Its companion “AI-Ready” ecosystem open-sources drivers, calibration tools, SLAM, and perception algorithms; the follow-up AC2 model is aimed at close-range manipulation.","example":"An AC1 mounted on an outdoor inspection robot runs the official open-source robosense_ac_slam package, using depth, image, and IMU data together for lidar-inertial-visual odometry and mapping.","related":["Depth Camera","Direct Time of Flight","Single-Photon Avalanche Diode","Inertial Measurement Unit","Multi-Sensor Fusion","RoboSense"]},{"id":"depth-holes","category":"perception","sec":1,"tier":3,"sources":[{"title":"RealSense: Depth Post-Processing for Intel RealSense Depth Camera D400 Series","url":"https://dev.realsenseai.com/docs/depth-post-processing-for-intel-realsense-depth-camera-d400-series/"},{"title":"librealsense: Post-processing filters","url":"https://github.com/realsenseai/librealsense/blob/master/doc/post-processing-filters.md"}],"as_of":"","related_ids":["depth-completion","depth-camera","active-stereo","flying-pixels","occlusion","transparent-and-reflective-object-perception"],"name":"Depth Holes","alt":"深度空洞","abbr":"","aliases":["Missing Depth","Invalid Depth Pixels"],"one_liner":"Pixels where a depth camera can’t measure distance, usually recorded as zero in the depth map.","explanation":"Depth holes are pixels in a depth map with no valid measurement; cameras like the RealSense record them as zero, which shows up as black patches. Taking stereo-based depth cameras as an example, Intel RealSense’s official white paper lists several causes: occlusion (the left and right cameras can’t both see the same point), lack of surface texture, repeating patterns causing multiple matches, over- or under-exposure, and objects closer than the minimum working distance; transparent surfaces, specular reflections, and very distant surfaces also often fail to register. Holes directly affect downstream processing: missing chunks in a point cloud cause grasp detection and obstacle avoidance to misjudge, and feeding a raw zero into a network as if it were a real distance makes the network think the object is pressed right up against the camera. Common fixes include post-processing filters (filling holes from neighboring pixels, or from previous frames’ values), learned depth completion, or masking out missing regions during training.","example":"Photographing a stainless-steel spoon and a glass on a table with a stereo depth camera: the highlight on the spoon and the whole glass often come back as a patch of zeros in the depth map, so these two objects are nearly missing once converted to a point cloud.","related":["Depth Completion","Depth Camera","Active Stereo","Flying Pixels","Occlusion","Transparent & Reflective Object Perception"]},{"id":"flying-pixels","category":"perception","sec":1,"tier":3,"sources":[{"title":"Azure Kinect DK depth camera（Invalidation / Multi-path 一节）","url":"https://learn.microsoft.com/en-us/previous-versions/azure/kinect-dk/depth-camera"},{"title":"Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers","url":"https://arxiv.org/abs/2510.07316"}],"as_of":"","related_ids":["depth-camera","time-of-flight","depth-map","point-cloud","depth-holes","monocular-depth-estimation"],"name":"Flying Pixels","alt":"飞点","abbr":"","aliases":["Flying Points","Mixed Pixels"],"one_liner":"Erroneous depth points at object edges in a depth map, floating between the foreground and background.","explanation":"Flying pixels are a common artifact in depth data at object outlines: edge pixels that should belong to either the foreground or the background instead get a depth value somewhere in between, which, once converted to a point cloud, becomes a scatter of points floating in midair — like a veil trailing off the object’s edge. In ToF (time-of-flight) depth cameras, this happens because a single pixel receives light reflected back from both the foreground and background at once, measuring a mixed signal; Microsoft’s Azure Kinect documentation classifies this kind of edge pixel as a multipath problem and marks it invalid. In stereo matching and depth-estimation networks, regression output gets smoothed across sharp depth discontinuities, which also produces flying pixels. Flying pixels cause errors in grasp poses and collision checking, and a common fix is filtering out suspicious edge points based on depth gradient or neighborhood consistency.","example":"Pixel-Perfect Depth (NeurIPS 2025) points out that generative depth models which compress a depth map into a latent space with a VAE first introduce flying pixels at edges and fine detail, so it instead does diffusion generation directly in pixel space, producing a point cloud with almost no flying pixels.","related":["Depth Camera","Time of Flight","Depth Map","Point Cloud","Depth Holes","Monocular Depth Estimation"]},{"id":"lidar","category":"perception","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Lidar","url":"https://en.wikipedia.org/wiki/Lidar"},{"title":"Livox Mid-360 产品页","url":"https://www.livoxtech.com/mid-360"}],"as_of":"","related_ids":["point-cloud","lidar-channels","solid-state-lidar","lidar-slam","time-of-flight","livox-mid-360"],"name":"LiDAR","alt":"激光雷达","abbr":"LiDAR","aliases":["Light Detection and Ranging","Laser Scanner","3D LiDAR"],"one_liner":"A sensor that fires lasers and times their return to measure distance, scanning out a 3D point cloud of its surroundings.","explanation":"Lidar sends out laser pulses and measures how long the light takes to bounce off an object and return, computing distance from that time; rotating or scanning the beam then covers a wide field of view, producing a 3D point cloud of the surroundings. The first system of this kind was built by Hughes Aircraft Company in 1961, shortly after the laser was invented. Lidar comes in three basic designs: fully mechanical units that spin the whole sensor, semi-solid-state units that scan with a rotating mirror or a MEMS micro-mirror, and solid-state units with no moving parts at all; by beam count, it's either a 2D single-line unit that scans a single plane or a multi-line 3D unit. Compared with a depth camera, lidar measures farther, is more accurate at range, and is less affected by ambient light, so it keeps working reliably outdoors in direct sunlight, though its points are sparser, carry no color, and cost more. On robots it is used mainly for SLAM mapping and localization, navigation and obstacle avoidance, and terrain sensing, and is common on self-driving cars, quadrupeds, and humanoids.","example":"The Livox Mid-360 is a common 3D lidar on mobile robots, with a 360° horizontal and 59° vertical field of view; it can measure a target with 10% reflectivity out to 40 meters, with a minimum range of 0.1 meters.","related":["Point Cloud","LiDAR Channels","Solid-State LiDAR","LiDAR SLAM","Time of Flight","Livox Mid-360"]},{"id":"2d-lidar","category":"perception","sec":1,"tier":2,"sources":[{"title":"SLAMTEC RPLIDAR A1","url":"http://www.slamtec.com/en/lidar/a1"},{"title":"Wikipedia: Lidar","url":"https://en.wikipedia.org/wiki/Lidar"}],"as_of":"","related_ids":["lidar","lidar-slam","adaptive-monte-carlo-localization","occupancy-grid-map","costmap","robot-vacuum-cleaner"],"name":"2D LiDAR","alt":"2D激光雷达","abbr":"","aliases":["Single-Line Laser Scanner","Laser Scanner","Planar LiDAR"],"one_liner":"A spinning laser rangefinder that scans one horizontal plane, measuring the distance to obstacles all around.","explanation":"A 2D lidar, also called a single-line lidar, has just one laser beam for ranging, spun by a motor or a rotating mirror to sweep out a horizontal plane; each revolution outputs a series of angle-and-distance readings that, plotted together, trace an outline of the surroundings. Ranging is done by time-of-flight (d = c·t / 2, where c is the speed of light and t is the round-trip time), though cheaper units use triangulation instead. Being cheap, low in data volume, and stable in precision, 2D lidars are common in robot vacuums, warehouse AMRs, and AGVs, used for 2D lidar SLAM, AMCL localization, and obstacle avoidance; ROS represents their data with the LaserScan message. Their limitation is that they only see the single plane at their mounting height, so obstacles like tabletops or overhead bars that sit outside that plane get missed, requiring a depth camera or a multi-line lidar to fill the gap.","example":"The Slamtec RPLIDAR A1 performs a 360° scan, ranging more than 8,000 times per second at an adjustable 2–10 Hz scan rate using triangulation, and is officially positioned for SLAM mapping, robot navigation, and robot vacuums.","related":["LiDAR","LiDAR SLAM","Adaptive Monte Carlo Localization","Occupancy Grid Map","Costmap","Robot Vacuum Cleaner"]},{"id":"lidar-channels","category":"perception","sec":1,"tier":3,"sources":[{"title":"Ouster OS1 Lidar Sensor","url":"https://ouster.com/products/hardware/os1-lidar-sensor"},{"title":"Hesai JT128 产品页","url":"https://www.hesaitech.com/product/jt128/"},{"title":"Wikipedia: Velodyne Lidar","url":"https://en.wikipedia.org/wiki/Velodyne_Lidar"}],"as_of":"2026-09","related_ids":["lidar","solid-state-lidar","field-of-view","point-cloud","hesai-jt128","livox-mid-360"],"name":"LiDAR Channels","alt":"激光雷达线数","abbr":"","aliases":["Beams","Lines","Channel Count"],"one_liner":"The number of laser channels arranged vertically in a multi-line lidar, which sets how dense the point cloud is top to bottom.","explanation":"Channel count refers to the number of laser transmit-receive channels a multi-line lidar arranges vertically. A traditional mechanically spinning lidar stacks multiple laser channels at different pitch angles into a column and spins the whole assembly horizontally, with each channel tracing out one “scan line” per rotation — so a 16-channel lidar produces 16 rings of points per frame. More channels means higher vertical resolution and more points landing on distant objects, but also higher cost, data volume, and compute requirements. Velodyne’s HDL-64E, released around 2007, spins 64 laser channels, producing about a million points per second; the Ouster OS1 offers up to 128 channels with a 45° vertical field of view. Solid-state or non-repetitive-scan lidars have no fixed rings, so vendors often describe point-cloud density using an “equivalent channel count” — for example, Livox’s Mid-360 is rated at an equivalent of 40 lines.","example":"The Hesai JT128 is a 128-channel lidar built for robots: 360° horizontal by 189° vertical field of view, 0.74° vertical angular resolution, about 1.15 million points per second in single-return mode, and a weight of 265 grams.","related":["LiDAR","Solid-State LiDAR","Field of View","Point Cloud","Hesai JT128","Livox Mid-360"]},{"id":"solid-state-lidar","category":"perception","sec":1,"tier":3,"sources":[{"title":"Lidar - Wikipedia","url":"https://en.wikipedia.org/wiki/Lidar"}],"as_of":"","related_ids":["lidar","single-photon-avalanche-diode","direct-time-of-flight","lidar-channels","field-of-view","frequency-modulated-continuous-wave-lidar"],"name":"Solid-State LiDAR","alt":"固态激光雷达","abbr":"","aliases":["Hybrid Solid-State LiDAR","Mechanical Spinning LiDAR (contrast)"],"one_liner":"Lidar with no bulk rotating part, scanning instead with electronics or a small internal mechanism.","explanation":"A traditional mechanical spinning lidar scans by using a motor to spin an entire row of transmit/receive modules through a full rotation, giving a 360° field of view but at the cost of bulk, high cost, and a lifetime limited by the rotating parts. Solid-state lidar removes any large moving part; common approaches include Flash (illuminating the whole field of view at once) and OPA (optical phased array, which steers the beam electronically by controlling phase). Between the two is what’s called hybrid or semi-solid-state lidar, which keeps only a small moving part such as a rotating mirror, a MEMS micromirror, or a prism. Going solid-state gives a smaller size, easier mass production, and better vibration tolerance, but the field of view is usually narrower than a spinning design. Robots and cars today mostly use hybrid solid-state products, while fully solid-state units are mostly used for close-range blind-spot coverage.","example":"","related":["LiDAR","Single-Photon Avalanche Diode","Direct Time of Flight","LiDAR Channels","Field of View","Frequency-Modulated Continuous-Wave LiDAR"]},{"id":"frequency-modulated-continuous-wave-lidar","category":"perception","sec":1,"tier":3,"sources":[{"title":"Aurora: FMCW Lidar — The Self-Driving Game-Changer","url":"https://aurora.tech/blog/fmcw-lidar-the-self-driving-game-changer"},{"title":"Aeva 官网（FMCW 4D LiDAR）","url":"https://www.aeva.com/"},{"title":"Wikipedia: Continuous-wave radar（FMCW 原理）","url":"https://en.wikipedia.org/wiki/Continuous-wave_radar"}],"as_of":"2026-09","related_ids":["lidar","time-of-flight","solid-state-lidar","millimeter-wave-radar","autonomous-driving","multi-sensor-fusion"],"name":"Frequency-Modulated Continuous-Wave LiDAR","alt":"FMCW 激光雷达（调频连续波激光雷达）","abbr":"FMCW","aliases":["FMCW LiDAR","Coherent LiDAR","4D LiDAR"],"one_liner":"A lidar that continuously emits frequency-swept laser light, measuring both distance and velocity at every point.","explanation":"This is one of the ranging methods used in lidar. The common time-of-flight (ToF) lidar emits short pulses and times the round trip of the echo; FMCW lidar instead continuously emits laser light whose frequency sweeps periodically over time, mixes the returning echo coherently with a local copy of the light, and reads distance from the resulting frequency difference, while the Doppler shift directly gives that point’s velocity along the line of sight — which is why vendors often call it 4D lidar. It only responds to light matching its own frequency and wavelength, so it’s much less affected by sunlight or crosstalk from other lidars; it commonly operates in the 1550 nm band, where eye-safety limits allow higher power. The cost is that it needs coherent transceiver components, making it structurally complex and expensive. It’s currently used mainly in autonomous driving, with vendors like Aeva and Aurora; robots still mostly use ToF lidar.","example":"Every point from an Aeva FMCW lidar carries a velocity value in addition to its 3D coordinates, so a single frame can separate a walking pedestrian from a stationary pole without comparing two consecutive frames.","related":["LiDAR","Time of Flight","Solid-State LiDAR","Millimeter-Wave Radar","Autonomous Driving","Multi-Sensor Fusion"]},{"id":"livox-mid-360","category":"perception","sec":1,"tier":2,"sources":[{"title":"Livox Mid-360 规格参数","url":"https://www.livoxtech.com/mid-360/specs"},{"title":"Livox Mid-360 产品页","url":"https://www.livoxtech.com/mid-360"},{"title":"livox_ros_driver2 README","url":"https://raw.githubusercontent.com/Livox-SDK/livox_ros_driver2/master/README.md"}],"as_of":"2026-09","related_ids":["lidar","lidar-slam","solid-state-lidar","lidar-inertial-odometry","fast-lio2","livox"],"name":"Livox Mid-360","alt":"览沃 Mid-360","abbr":"","aliases":["Mid-360","Mid360"],"one_liner":"A small 360° hybrid solid-state lidar from Livox with a built-in IMU, widely used on mobile robots.","explanation":"The Mid-360 is a small, 360° hybrid solid-state lidar from Livox Technology, aimed at mobile robots. Its official specs are a 360° horizontal and −7° to 52° vertical field of view, a range of 40 meters for a target with 10% reflectivity and a minimum range of 0.1 meters, 200,000 points per second, a weight of 265 grams, and a built-in ICM40609-model IMU (inertial measurement unit). Its angular resolution improves with integration time — the longer it stays pointed somewhere, the denser the point cloud gets. It's officially positioned to replace 2D lidar, RGB-D cameras, and ultrasonic sensors for indoor navigation and obstacle avoidance; its driver, livox_ros_driver2, supports both ROS 1 and ROS 2, and the built-in IMU also makes it convenient to run lidar-inertial odometry directly.","example":"Mounting a Mid-360 on a wheeled or legged robot, publishing its point cloud and IMU topics with livox_ros_driver2, and feeding them into FAST-LIO2 for mapping, the result can then be handed off to the navigation stack for obstacle avoidance and path planning.","related":["LiDAR","LiDAR SLAM","Solid-State LiDAR","LiDAR-Inertial Odometry","FAST-LIO2","Livox"]},{"id":"hesai-jt128","category":"perception","sec":1,"tier":3,"sources":[{"title":"Hesai JT128/64P 产品页","url":"https://www.hesaitech.com/product/jt128/"},{"title":"Hesai Newsroom（含 2025-05 JT 系列产品介绍）","url":"https://www.hesaitech.com/news/"}],"as_of":"2026-09","related_ids":["lidar","lidar-channels","field-of-view","hesai-technology","livox-mid-360","robosense-airy"],"name":"Hesai JT128","alt":"禾赛 JT128","abbr":"","aliases":["Hesai JT Series","JT64P"],"one_liner":"A compact, ultra-wide-field-of-view 128-channel 3D lidar from Hesai Technology aimed at robots.","explanation":"The JT128 is the 128-channel model in Hesai Technology’s JT series of lidars, which also includes the JT64P and JT16. “Channel count” refers to the number of laser channels in the vertical direction — more channels means a denser point cloud. According to Hesai’s official specs, the JT128 has a 360°×189° field of view, covering more than a full hemisphere; it ranges to 40 m at 10% reflectivity, with a maximum of 60 m; it produces about 1.152 million points per second in single-return mode; and it measures 62.5 mm in diameter and 73 mm tall, weighs 265 g, and is rated IPX7. Traditional spinning lidars usually have a vertical field of view of only a few tens of degrees, which tends to leave blind spots when mounted on a robot; the JT series uses an ultra-wide vertical field of view to reduce these blind spots. Hesai’s website lists embodied AI robots, delivery robots, AGVs/AMRs, and cleaning robots among its applications, for mapping, localization, obstacle avoidance, and terrain sensing.","example":"","related":["LiDAR","LiDAR Channels","Field of View","Hesai Technology","Livox Mid-360","RoboSense Airy"]},{"id":"robosense-airy","category":"perception","sec":1,"tier":3,"sources":[{"title":"RoboSense Airy 产品页","url":"https://www.robosense.ai/en/rslidar/Airy"}],"as_of":"2026-09","related_ids":["lidar","lidar-channels","field-of-view","lidar-slam","livox-mid-360","robosense"],"name":"RoboSense Airy","alt":"速腾聚创 Airy","abbr":"","aliases":["RoboSense Airy Hemispherical LiDAR"],"one_liner":"A hemispherical-field-of-view digital lidar from RoboSense for robots, covering 360°×90° with a single unit.","explanation":"Airy is a robot-oriented lidar from RoboSense, described by the company as the first “digital hemispherical lidar.” It is roughly the size of a ping-pong ball (60 mm diameter, 63 mm tall) and weighs under 240 grams; it has a 360° horizontal field of view and a 90° vertical field of view, in 192-line and 96-line versions, with a range of 60 meters (about 30 meters against a 10%-reflectivity target) and roughly 1 cm ranging accuracy. Traditional spinning lidars have a narrow vertical field of view, so when mounted on a robot they often can’t see the ground right beneath or beside it, requiring multiple units to be combined; a hemispherical field of view lets a single unit cover both the surroundings and the nearby ground at once. It uses a chip-based transmit/receive design with digital detection, and targets quadruped robots, humanoids, home and garden robots, AMRs (autonomous mobile robots), and autonomous forklifts, for SLAM mapping, obstacle avoidance, and traversability assessment.","example":"A quadruped robot dog carries one Airy unit on its head, which sees both the obstacles ahead and the steps beneath its feet, cutting down blind spots during obstacle avoidance and mapping.","related":["LiDAR","LiDAR Channels","Field of View","LiDAR SLAM","Livox Mid-360","RoboSense"]},{"id":"unitree-4d-lidar-l1-l2","category":"perception","sec":1,"tier":3,"sources":[{"title":"Unitree 4D LiDAR L2 官网","url":"https://www.unitree.com/L2"},{"title":"Unitree 4D LiDAR L1 官网","url":"https://www.unitree.com/LiDAR"},{"title":"Unitree Go2 官网","url":"https://www.unitree.com/go2"}],"as_of":"2026-09","related_ids":["lidar","livox-mid-360","unitree-go2","lidar-inertial-odometry","unitree-robotics","inertial-measurement-unit"],"name":"Unitree 4D LiDAR L1 / L2","alt":"宇树 4D 激光雷达 L1 / L2","abbr":"","aliases":["Unitree L1","Unitree L2"],"one_liner":"Unitree’s own low-cost, hemispherical-field-of-view lidar, where every point carries a 3D coordinate plus a grayscale value.","explanation":"This is a small lidar (a sensor that measures range with a laser and scans out a 3D point cloud of its surroundings) developed in-house by Unitree Robotics. “4D” refers to each point carrying a 3D position plus one extra dimension of grayscale (reflected intensity). Per the published specs: the L1 has a 360°×90° field of view, outputs about 21,600 points per second, and starts at US$249; the L2 has a 360°×96° field of view, outputs 64,000 points per second, ranges to 30 m (at 90% reflectivity), has a close-range blind zone of only 0.05 m, weighs 230 g, and costs $419. Both have a built-in IMU and use a non-repetitive scan pattern. The price is far below traditional lidar; Unitree’s Go2 quadruped is listed as shipping with the L2, and Unitree also provides an open-source mapping pipeline based on Point-LIO, with the sensor commonly used for SLAM and obstacle avoidance.","example":"On a Unitree Go2, the head-mounted L2’s point cloud plus its built-in IMU are run through Point-LIO to build a 3D indoor map while walking, used for navigation and obstacle avoidance.","related":["LiDAR","Livox Mid-360","Unitree Go2","LiDAR-Inertial Odometry","Unitree Robotics","Inertial Measurement Unit"]},{"id":"millimeter-wave-radar","category":"perception","sec":1,"tier":3,"sources":[{"title":"The fundamentals of millimeter wave radar sensors - Texas Instruments","url":"https://www.ti.com/lit/wp/spyy005a/spyy005a.pdf"}],"as_of":"","related_ids":["lidar","multi-sensor-fusion","autonomous-driving","ultrasonic-sensor","exteroception"],"name":"Millimeter-Wave Radar","alt":"毫米波雷达","abbr":"mmWave Radar","aliases":["mmWave Radar","mmWave","4D Millimeter-Wave Radar"],"one_liner":"A radar that transmits millimeter-wavelength electromagnetic waves to measure a target’s distance, speed, and bearing.","explanation":"Millimeter-wave radar operates in roughly the 30–300 GHz band; automotive systems typically use frequencies near 77 GHz. It transmits a frequency-modulated electromagnetic wave and computes a target’s distance, radial velocity (via the Doppler effect), and bearing angle from the reflected signal. Its advantages are that it works through rain, fog, and darkness, and measures velocity directly; its drawbacks are coarse angular resolution and a sparse point cloud, which make it hard to resolve an object’s shape. “4D” millimeter-wave radar additionally measures elevation angle, giving a denser point cloud. In self-driving cars it is fused with cameras and lidar as part of multi-sensor fusion; in robotics it is used mainly for safety guarding and detecting people nearby.","example":"A car’s forward-facing 77 GHz radar can still measure the distance and relative speed of the vehicle ahead in rain or fog, which is what makes adaptive cruise control work in bad weather.","related":["LiDAR","Multi-Sensor Fusion","Autonomous Driving","Ultrasonic Sensor","Exteroception"]},{"id":"ultrasonic-sensor","category":"perception","sec":1,"tier":3,"sources":[{"title":"Ultrasonic transducer - Wikipedia","url":"https://en.wikipedia.org/wiki/Ultrasonic_transducer"},{"title":"Parking sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Parking_sensor"}],"as_of":"","related_ids":["vision-only-approach","proximity-sensor","time-of-flight","obstacle-avoidance","millimeter-wave-radar","transparent-and-reflective-object-perception"],"name":"Ultrasonic Sensor","alt":"超声波传感器","abbr":"USS","aliases":["USS","Ultrasonic Ranging Sensor"],"one_liner":"A cheap, short-range sensor that measures distance by emitting ultrasound and listening for its echo.","explanation":"An ultrasonic sensor emits a pulse of sound above 20 kHz (inaudible to humans) and measures how long the echo takes to return; the speed of sound times half that time gives the obstacle’s distance. It’s cheap, and because it doesn’t depend on light, an object’s color and reflectivity don’t affect the measurement, so it can catch obstacles like glass doors that cameras and lidar often miss — which is why it’s common in car parking sensors and short-range obstacle avoidance on mobile robots. Its drawbacks are a wide beam that can’t resolve fine directional detail, and a range usually limited to a few meters; sound-absorbing materials or steeply angled surfaces produce a weak echo that can go undetected. In cars this is often called an “ultrasonic radar”; Tesla reportedly removed it from some models starting in 2022, switching to a camera-only approach instead.","example":"While backing up, an ultrasonic sensor on the bumper detects an obstacle getting closer behind the car; the warning beep goes from intermittent to rapid, and finally to a continuous tone.","related":["Vision-Only Approach","Proximity Sensor","Time of Flight","Obstacle Avoidance","Millimeter-Wave Radar","Transparent & Reflective Object Perception"]},{"id":"proximity-sensor","category":"perception","sec":1,"tier":3,"sources":[{"title":"Proximity sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Proximity_sensor"},{"title":"Proximity Perception in Human-Centered Robotics: A Survey on Sensing Systems and Applications (arXiv 2108.07206)","url":"https://arxiv.org/abs/2108.07206"}],"as_of":"","related_ids":["tactile-sensor","electronic-skin","time-of-flight","ultrasonic-sensor","human-robot-collaboration","speed-and-separation-monitoring"],"name":"Proximity Sensor","alt":"接近觉传感器","abbr":"","aliases":["Proximity Sensing","Near-Field Sensing"],"one_liner":"A sensor that detects whether something is nearby, and how close, without needing to touch it.","explanation":"A proximity sensor detects the presence or distance of a nearby object without contact, using principles that include capacitive sensing (a nearby object changes an electric field), inductive sensing (sensitive only to metal), infrared and other optical methods, time-of-flight (measuring how long light takes to travel out and back), and ultrasonic sensing. The sensor that turns off a phone’s screen when it’s held to your ear works this way. In robotics, proximity sensing fills the gap between vision and touch: a camera is often blocked at close range by the arm or the object itself, while touch sensing needs actual contact to produce a signal. There are two common uses: mounting sensors on a robot arm’s housing as a “sensing skin” that slows or steers the arm away when a person gets close, supporting safety in human-robot collaboration; and mounting them on gripper fingertips to fine-tune finger position in the last small gap before contact. A 2021 survey by Navarro and colleagues systematically reviews these systems.","example":"A collaborative robot arm’s housing is ringed with capacitive proximity sensors; when a worker’s hand gets within a set distance, the arm automatically slows down and stops.","related":["Tactile Sensor","Electronic Skin","Time of Flight","Ultrasonic Sensor","Human-Robot Collaboration","Speed and Separation Monitoring"]},{"id":"safety-laser-scanner-safety-light-curtain","category":"perception","sec":1,"tier":3,"sources":[{"title":"Light curtain - Wikipedia","url":"https://en.wikipedia.org/wiki/Light_curtain"},{"title":"SICK Safety laser scanners","url":"https://www.sick.com/us/en/catalog/products/safety-systems-and-solutions/safety-laser-scanners/c/g187234"}],"as_of":"","related_ids":["functional-safety","speed-and-separation-monitoring","protective-stop","2d-lidar","autonomous-mobile-robot","iso-10218-1-2-2025-robotics-safety-requirements"],"name":"Safety Laser Scanner / Safety Light Curtain","alt":"安全激光扫描仪 / 安全光幕","abbr":"","aliases":["Safety Laser Scanner","Safety Light Curtain","Electro-Sensitive Protective Equipment (ESPE)"],"one_liner":"Certified safety devices that detect a person entering a danger zone and stop the machine before contact.","explanation":"Both devices fall under electro-sensitive protective equipment (ESPE): industrial safety sensors that make a machine “stop when it sees a person,” certified to standards such as IEC 61496, whose output is a safety-rated stop signal rather than ordinary data meant for an algorithm to interpret. A safety light curtain consists of a pair of emitter and receiver columns; the emitter sends a row of parallel infrared beams to the receiver, and blocking any single beam triggers a stop, commonly installed at the entrance to a machine tool or robot work cell. A safety laser scanner is a certified 2D laser scanning device that can define a “warning zone” and a “protective zone” within its scan plane: entering the warning zone slows the machine or triggers an alarm, and entering the protective zone stops it. Both let robots avoid needing a full physical fence, and are common hardware for AMR collision avoidance and speed-and-separation monitoring in human-robot collaboration.","example":"A pair of safety light curtains guards the entrance to a robot arm’s loading station; when a worker reaches in to grab material, the beam is blocked and the arm immediately performs a protective stop.","related":["Functional Safety","Speed and Separation Monitoring","Protective Stop","2D LiDAR","Autonomous Mobile Robot","ISO 10218-1/-2:2025 Robotics — Safety Requirements"]},{"id":"vision-only-approach","category":"perception","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Tesla Autopilot hardware（Tesla Vision：2021 年去毫米波雷达、2022 年去超声波雷达）","url":"https://en.wikipedia.org/wiki/Tesla_Autopilot_hardware"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model（局限：仅支持单图输入）","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"2022-10","related_ids":["rgb-camera","lidar","depth-camera","multi-sensor-fusion","visuo-tactile-fusion","autonomous-driving"],"name":"Vision-Only Approach","alt":"纯视觉方案","abbr":"","aliases":["Camera-Only"],"one_liner":"Perceives the environment using only cameras, without lidar, millimeter-wave radar, or other ranging sensors.","explanation":"A vision-only approach means perception relies solely on cameras — usually ordinary RGB cameras — with distance and 3D structure both inferred from images by an algorithm, without lidar, millimeter-wave radar, ultrasonic sensors, or other ranging hardware. The best-known example is Tesla: starting in May 2021, North American–built Model 3/Y dropped millimeter-wave radar under the name “Tesla Vision,” and in October 2022 Tesla announced it was removing ultrasonic sensors too. The advantages are cheap hardware and data that’s easy to scale; the cost is that depth can only be estimated, so errors are more likely in low light, backlight, and with transparent or reflective objects. In embodied AI, “vision-only” also describes a policy that looks only at images, without tactile or force input — for example, OpenVLA takes just a single RGB image and a language instruction. Whether a vision-only approach or multi-sensor fusion is better remains debated.","example":"OpenVLA receives only a single RGB image and one instruction, with no depth, force, or proprioceptive input, and outputs the robot arm’s end-effector action directly.","related":["RGB Camera","LiDAR","Depth Camera","Multi-Sensor Fusion","Visuo-Tactile Fusion","Autonomous Driving"]},{"id":"rotary-encoder","category":"perception","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Rotary encoder","url":"https://en.wikipedia.org/wiki/Rotary_encoder"}],"as_of":"","related_ids":["absolute-encoder","incremental-encoder","magnetic-encoder","dual-encoder","proprioception","joint-actuator-module"],"name":"Rotary Encoder","alt":"编码器","abbr":"","aliases":["Joint Encoder","Angle Encoder"],"one_liner":"A position sensor mounted on a motor or joint that converts shaft angle into an electrical signal.","explanation":"A rotary encoder is an electromechanical device that converts a shaft's angular position or amount of rotation into an analog or digital signal; a robot has at least one on every joint, telling the controller ‘how many degrees have I turned right now.’ There are two output types: incremental encoders only report how much they've turned, usually via two channels A and B offset by 90° in phase to determine direction, and they lose absolute position when power is cut, needing a homing routine afterward; absolute encoders directly report the current angle and keep it through a power loss. The underlying principle can be optical, magnetic, capacitive, or inductive, with resolution given as pulses per revolution or in bits. Joint modules often mount one encoder on the motor side and another at the reducer's output — a dual-encoder setup — so the true output-side angle is measured directly, canceling out the error introduced by gear backlash (the small play between gear teeth). The joint-angle information in a policy's proprioceptive input comes from encoder readings like these.","example":"A 14-bit absolute magnetic encoder divides a full turn into 2^14 = 16,384 parts, giving an angular resolution of about 360° ÷ 16,384 ≈ 0.022°.","related":["Absolute Encoder","Incremental Encoder","Magnetic Encoder","Dual Encoder","Proprioception","Joint Actuator Module"]},{"id":"inertial-measurement-unit","category":"perception","sec":2,"tier":1,"sources":[{"title":"Wikipedia: Inertial measurement unit","url":"https://en.wikipedia.org/wiki/Inertial_measurement_unit"},{"title":"unitree_rl_gym deploy_real.py（IMU 角速度与投影重力作为观测）","url":"https://github.com/unitreerobotics/unitree_rl_gym/blob/main/deploy/deploy_real/deploy_real.py"}],"as_of":"","related_ids":["sensor-drift","projected-gravity","state-estimation","visual-inertial-odometry","proprioception","camera-imu-calibration"],"name":"Inertial Measurement Unit","alt":"惯性测量单元","abbr":"IMU","aliases":["IMU","Six-Axis IMU","Nine-Axis IMU","Inertial Sensor"],"one_liner":"A sensor combining an accelerometer and gyroscope that measures an object's acceleration and rotation rate.","explanation":"An inertial measurement unit (IMU) combines an accelerometer and a gyroscope, which measure linear acceleration along three axes (at rest, this reads gravity, which is how the IMU can tell which way is down) and angular velocity around three axes — together called six-axis; adding a three-axis magnetometer makes it nine-axis. IMUs are found in phones, drones, and VR headsets. For a robot, an IMU provides the body's orientation (especially which way gravity points) and its rotation rate, which is basic input for balance in legged and humanoid robots. Its readings carry bias and noise, so integrating them to get angle or position accumulates drift over time, which is why IMUs are usually fused with a camera, lidar, or leg odometry. Colloquially, some people call an IMU an ‘inertial nav’ system, but strictly speaking an inertial navigation system is the larger system that further integrates the IMU's readings to solve for position and velocity.","example":"In Unitree's open-source unitree_rl_gym, the first 6 dimensions of the G1 walking policy's observation are the body angular velocity measured by the IMU's gyroscope, plus the projected gravity direction computed from the IMU's attitude quaternion.","related":["Sensor Drift","Projected Gravity","State Estimation","Visual-Inertial Odometry","Proprioception","Camera-IMU Calibration"]},{"id":"attitude-estimation","category":"perception","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Attitude and heading reference system","url":"https://en.wikipedia.org/wiki/Attitude_and_heading_reference_system"},{"title":"x-io Technologies: Open source IMU and AHRS algorithms (Madgwick / Fusion)","url":"https://x-io.co.uk/open-source-imu-and-ahrs-algorithms/"}],"as_of":"","related_ids":["inertial-measurement-unit","complementary-filter","kalman-filter","projected-gravity","roll-pitch-yaw","quaternion"],"name":"Attitude Estimation","alt":"姿态解算（AHRS 航姿参考系统）","abbr":"AHRS","aliases":["Attitude and Heading Reference System","AHRS","IMU Attitude Fusion"],"one_liner":"Fuses gyroscope, accelerometer, and (optionally) magnetometer data to compute roll, pitch, and yaw angles in real time.","explanation":"Attitude estimation computes an object’s orientation in space from inertial sensor data, typically expressed as roll, pitch, and yaw angles or as a quaternion. AHRS (attitude and heading reference system) was originally an aviation instrument, built from a three-axis gyroscope, accelerometer, and magnetometer plus an onboard computer that outputs attitude and heading directly; a plain IMU usually only outputs raw angular velocity and acceleration. In principle, integrating the gyroscope gives an angle that’s accurate short-term but drifts; the accelerometer can sense the direction of gravity, which corrects roll and pitch; yaw has no gravity reference, so it needs correction from a magnetometer, vision, or another source. Common fusion algorithms include the Kalman filter, the complementary filter, the Madgwick algorithm (proposed during a 2009 PhD project, now open-sourced under the name Fusion), and the Mahony filter. Control policies for legged and humanoid robots often take the estimated attitude (or projected gravity) and angular velocity as input, so attitude error directly affects balance control.","example":"When a quadruped robot is shoved from the side, IMU attitude estimation gives the real-time change in body roll angle, and the locomotion control policy steps to recover balance based on it.","related":["Inertial Measurement Unit","Complementary Filter","Kalman Filter","Projected Gravity","Roll-Pitch-Yaw (RPY)","Quaternion"]},{"id":"complementary-filter","category":"perception","sec":2,"tier":3,"sources":[{"title":"Complementary Filter - AHRS 文档","url":"https://ahrs.readthedocs.io/en/latest/filters/complementary.html"},{"title":"complementaryFilter - MATLAB Sensor Fusion and Tracking Toolbox","url":"https://www.mathworks.com/help/fusion/ref/complementaryfilter-system-object.html"}],"as_of":"","related_ids":["inertial-measurement-unit","kalman-filter","extended-kalman-filter","attitude-estimation","sensor-drift","state-estimation"],"name":"Complementary Filter","alt":"互补滤波","abbr":"","aliases":["Mahony Complementary Filter"],"one_liner":"A simple filter that blends the gyroscope for short-term accuracy and the accelerometer for long-term stability to estimate attitude angle.","explanation":"The complementary filter is the simplest and most common method for IMU attitude estimation. The gyroscope measures angular velocity, which integrates into an angle that’s smooth and accurate over short periods, but bias makes error accumulate into drift over time (a low-frequency error); the accelerometer can compute pitch and roll directly from the direction of gravity, which doesn’t drift long-term, but it’s noisy and disturbed by vibration and motion acceleration (a high-frequency error), and yaw needs a magnetometer as an additional reference. The complementary filter high-pass-filters the gyroscope-integrated result and low-pass-filters the accelerometer result, then adds them; a common formula is: angle = α × (previous angle + angular velocity × Δt) + (1 − α) × accelerometer angle, with α set close to 1. It’s extremely cheap to compute, which suits microcontrollers; Mahony and colleagues extended it to full 3D rotation (SO(3)) in IEEE TAC in 2008, giving the commonly used Mahony filter. It needs no noise model, but its weight has to be tuned by hand.","example":"A self-balancing robot’s main controller reads the IMU once per control cycle and computes the body’s pitch angle with a complementary filter using α = 0.98, feeding it back into the balance controller.","related":["Inertial Measurement Unit","Kalman Filter","Extended Kalman Filter","Attitude Estimation","Sensor Drift","State Estimation"]},{"id":"sensor-drift","category":"perception","sec":2,"tier":3,"sources":[{"title":"Inertial measurement unit - Wikipedia","url":"https://en.wikipedia.org/wiki/Inertial_measurement_unit"}],"as_of":"","related_ids":["inertial-measurement-unit","kalman-filter","allan-variance","visual-inertial-odometry","six-axis-force-torque-sensor","state-estimation"],"name":"Sensor Drift","alt":"传感器漂移（零漂）","abbr":"","aliases":["Zero Drift","Bias Drift","Thermal Drift","IMU Drift"],"one_liner":"A sensor’s reading slowly shifting over time or temperature even when the true input hasn’t changed.","explanation":"An ideal sensor should read zero with no input, but a real sensor always carries some offset, or bias, and that bias itself changes slowly with time, temperature, and each time the sensor is powered on — this is drift, also called zero drift or thermal drift. The most common case in robotics is IMU (inertial measurement unit) drift: orientation, velocity, and position are all obtained by integrating gyroscope and accelerometer readings, so a tiny bias keeps accumulating — a constant gyroscope bias makes velocity error grow with the square of time and position error grow with its cube, so relying on an IMU alone to estimate position quickly diverges. Six-axis force/torque sensors and joint torque sensors also drift, so they typically need to be re-zeroed before use. Countermeasures include calibration, temperature compensation, characterizing noise with Allan variance, and fusing with cameras, lidar, or GPS, with the bias itself estimated online as part of the state in a Kalman filter.","example":"A humanoid robot standing perfectly still still sees its yaw angle, computed purely by integrating the IMU, slowly drift over time, which needs correcting by fusing in visual-inertial or legged odometry.","related":["Inertial Measurement Unit","Kalman Filter","Allan Variance","Visual-Inertial Odometry","Six-Axis Force/Torque Sensor","State Estimation"]},{"id":"allan-variance","category":"perception","sec":2,"tier":3,"sources":[{"title":"Kalibr Wiki: IMU Noise Model","url":"https://github.com/ethz-asl/kalibr/wiki/IMU-Noise-Model"},{"title":"Wikipedia: Allan variance","url":"https://en.wikipedia.org/wiki/Allan_variance"}],"as_of":"","related_ids":["inertial-measurement-unit","sensor-drift","camera-imu-calibration","kalibr","visual-inertial-odometry","kalman-filter"],"name":"Allan Variance","alt":"Allan 方差（IMU 噪声标定）","abbr":"","aliases":["Allan Deviation","Two-Sample Variance","AVAR"],"one_liner":"A statistical method for how sensor noise changes with averaging time, commonly used to calibrate IMU noise parameters.","explanation":"Allan variance is the “two-sample variance” that David W. Allan proposed in 1966 to measure the frequency stability of atomic clocks and crystal oscillators; it later became the standard method for calibrating noise in gyroscopes and accelerometers, and IEEE Std 952-1997 (the fiber-optic gyroscope test standard) documents how to read noise parameters off it. In practice, an IMU sits still and records a long stretch of data; the data is divided into segments of different averaging times τ, the mean of each segment is computed, and the variance of the differences between neighboring segment means is plotted on a log-log curve. A segment of the curve with slope −1/2 corresponds to white noise (noise density), while a slope of +1/2 corresponds to bias random walk. Kalibr’s IMU noise model documentation recommends recording 15 to 24 hours while stationary, reading white noise at τ = 1 second and the random-walk line at τ = 3 seconds. These readings feed into the configuration of visual-inertial odometry, camera-IMU joint calibration, and Kalman filters; getting them wrong makes the fusion algorithm trust the IMU too much or too little.","example":"Before doing camera-IMU joint calibration, the IMU is left recording overnight while stationary; allan_variance_ros plots the curve, and the gyroscope’s and accelerometer’s noise density and random walk are read off it and written into Kalibr’s IMU configuration file.","related":["Inertial Measurement Unit","Sensor Drift","Camera-IMU Calibration","Kalibr","Visual-Inertial Odometry","Kalman Filter"]},{"id":"six-axis-force-torque-sensor","category":"perception","sec":2,"tier":1,"sources":[{"title":"ATI Industrial Automation: Multi-Axis Force/Torque Sensors","url":"https://www.ati-ia.com/products/ft/sensors.aspx"},{"title":"Robotiq FT 300-S Force Torque Sensor","url":"https://robotiq.com/products/ft-300-force-torque-sensor"},{"title":"Wikipedia: Strain gauge","url":"https://en.wikipedia.org/wiki/Strain_gauge"}],"as_of":"","related_ids":["strain-gauge","force-control","impedance-control","peg-in-hole-insertion","crosstalk","joint-torque-sensor"],"name":"Six-Axis Force/Torque Sensor","alt":"六维力传感器","abbr":"F/T","aliases":["F/T Sensor","Force/Torque Sensor","Multi-Axis Force Sensor"],"one_liner":"A sensor that measures force along three axes and torque around three axes at once, often mounted on an arm's wrist.","explanation":"A six-axis force/torque (F/T) sensor measures six quantities at once: the forces Fx, Fy, Fz along the x, y, and z axes, and the torques Tx, Ty, Tz around those same axes. A common design bonds strain gauges to a metal elastic body; the tiny deformation under load changes electrical resistance, which is converted through a Wheatstone bridge and calibration into the six components, though capacitive designs also exist. On a robot arm it usually sits between the wrist flange and the gripper, telling the controller how hard, and in which direction, the end-effector is being pushed or twisted — the basis for force control, impedance control, peg-in-hole assembly, sanding, kinesthetic teaching, and collision detection. Representative vendors include ATI and Robotiq, and in China, Sunrise Instruments and Kunwei Technology. Some arms skip this sensor entirely and instead estimate end-effector force indirectly from torque sensors on each joint.","example":"The Robotiq FT 300-S mounts between a collaborative arm's wrist and gripper, rated for ±300 N of force and ±30 N·m of torque, outputting data at 100 Hz, and is commonly used for peg-in-hole assembly and surface finishing.","related":["Strain Gauge","Force Control","Impedance Control","Peg-in-Hole Insertion","Crosstalk (Inter-Axis Coupling)","Joint Torque Sensor"]},{"id":"joint-torque-sensor","category":"perception","sec":2,"tier":2,"sources":[{"title":"KUKA LBR iiwa 产品页","url":"https://www.kuka.com/en-us/products/robotics-systems/industrial-robots/lbr-iiwa"},{"title":"libfranka robot_state.h（tau_J: measured link-side joint torque sensor signals）","url":"https://raw.githubusercontent.com/frankaemika/libfranka/master/include/franka/robot_state.h"}],"as_of":"","related_ids":["six-axis-force-torque-sensor","strain-gauge","sensorless-force-estimation","torque-control","impedance-control","collision-detection"],"name":"Joint Torque Sensor","alt":"关节力矩传感器","abbr":"","aliases":["Joint Torque Sensing","One-Axis Torque Sensor"],"one_liner":"A sensor built into a robot joint that directly measures the torque the joint is actually producing.","explanation":"A joint torque sensor sits between a joint's reducer output and the link it drives, measuring torque only around that single joint's rotation axis — hence it's also called a one-axis torque sensor — and is mostly strain-gauge based, converting a tiny deformation of an elastic element into a torque reading. Motor current can also be used to estimate torque, but friction in the reducer makes that estimate inaccurate; measuring directly at the output lets the robot sense even very light external forces. KUKA's LBR iiwa carries one of these sensors on all seven axes, and the manufacturer states that this lets it immediately sense contact and reduce force and speed, allowing it to work alongside people without safety fencing; Franka's control interface likewise reports each joint's measured torque directly as tau_J. This is the basis for collision detection, torque control, impedance control, and kinesthetic teaching, at the cost of a higher price and a more complex joint design.","example":"A person pushes on a moving KUKA iiwa arm by hand. The joint torque sensor's reading exceeds what the dynamics model predicts, and the controller judges that a collision has occurred, immediately stopping or yielding to the push.","related":["Six-Axis Force/Torque Sensor","Strain Gauge","Sensorless Force Estimation","Torque Control","Impedance Control","Collision Detection (Robot Safety)"]},{"id":"strain-gauge","category":"perception","sec":2,"tier":3,"sources":[{"title":"Strain gauge - Wikipedia","url":"https://en.wikipedia.org/wiki/Strain_gauge"}],"as_of":"","related_ids":["six-axis-force-torque-sensor","joint-torque-sensor","foot-force-sensor","sensor-drift","crosstalk"],"name":"Strain Gauge","alt":"应变片","abbr":"","aliases":["Resistive Strain Gauge","Strain-Gauge Force Sensing"],"one_liner":"A small element bonded to metal whose electrical resistance changes as it deforms, the basic building block of force sensors.","explanation":"A strain gauge is a sensing element that converts a tiny deformation into a change in electrical resistance, typically a metal foil grid or semiconductor pattern printed on an insulating backing. Bonded to the surface of an elastic structure, it stretches or compresses slightly as the structure is loaded, changing its resistance, which is then read out with a Wheatstone bridge — a circuit that amplifies a small resistance change into a voltage signal. Six-axis force/torque sensors, joint torque sensors, foot-force sensors, and load cells are mostly built on strain gauges: several are bonded to a specially designed elastic structure, and the forces and torques along each axis are computed from their combined readings. Strain gauges are accurate and low-cost, but sensitive to temperature and subject to zero drift, requiring temperature compensation and periodic recalibration.","example":"Inside a six-axis force/torque sensor mounted on an arm’s wrist, several groups of strain gauges are bonded to a cross-beam elastic structure to measure force and torque along three axes each.","related":["Six-Axis Force/Torque Sensor","Joint Torque Sensor","Foot Force Sensor","Sensor Drift","Crosstalk (Inter-Axis Coupling)"]},{"id":"crosstalk","category":"perception","sec":2,"tier":3,"sources":[{"title":"A Novel 6-axis Force/Torque Sensor Using Inductance Sensors (arXiv 2505.09069)","url":"https://arxiv.org/abs/2505.09069"},{"title":"ATI Industrial Automation: Six-Axis F/T Transducer Installation and Operation Manual","url":"https://www.ati-ia.com/app_content/documents/9620-05-transducer%20section.pdf"}],"as_of":"","related_ids":["six-axis-force-torque-sensor","strain-gauge","wrench","sensor-drift","force-control","joint-torque-sensor"],"name":"Crosstalk (Inter-Axis Coupling)","alt":"维间耦合","abbr":"","aliases":["Cross-Axis Coupling","Inter-Axis Interference"],"one_liner":"When a multi-axis force sensor loaded along only one direction still shows nonzero readings on the other axes.","explanation":"Crosstalk is a metric for evaluating multi-axis force sensors, such as six-axis force sensors: when a load is applied along only one axis — say, a pure vertical push — the other channels that should read zero (Fx, Fy, and the torques) show a reading anyway. This happens because the strain gauges inside the sensor respond to loads from multiple directions at once, compounded by manufacturing and mounting errors in the elastic body. Manufacturers usually calibrate a decoupling matrix that converts the raw signals into three forces and three torques; the crosstalk that remains after calibration is typically expressed as a percentage of full scale (%FS), tested with both single-axis loading and simultaneous multi-axis loading. It directly affects force-control accuracy: in tasks like assembly or polishing, coupling error can make a robot believe it feels a lateral force that isn’t really there, triggering an incorrect compliant motion.","example":"A six-axis force sensor loaded with only 100 N straight down reads 1 N on the Fx channel; if Fx has a 100 N range, this crosstalk figure is 1% FS.","related":["Six-Axis Force/Torque Sensor","Strain Gauge","Wrench","Sensor Drift","Force Control","Joint Torque Sensor"]},{"id":"sensorless-force-estimation","category":"perception","sec":2,"tier":3,"sources":[{"title":"arXiv 2512.13009: K-VARK for Sensorless Force Estimation in Collaborative Robots","url":"https://arxiv.org/abs/2512.13009"},{"title":"arXiv 2609.13779: Force-Aware RL with Hybrid Sensorless Force Estimation for Wheeled-Legged Loco-Manipulation","url":"https://arxiv.org/abs/2609.13779"}],"as_of":"2026-09","related_ids":["generalized-momentum-observer","torque-constant","collision-detection","friction-compensation","proprioceptive-actuator","six-axis-force-torque-sensor"],"name":"Sensorless Force Estimation","alt":"无传感器力估计","abbr":"","aliases":["Current-Based Force Estimation","Motor-Current Force Estimation"],"one_liner":"Estimating the external force on a robot from motor current and a dynamics model, without installing a dedicated force sensor.","explanation":"Sensorless force estimation infers the force a robot experiences during contact with its environment without installing a dedicated six-axis force/torque sensor or joint torque sensor. Instead, it uses motor current — multiplied by the torque constant to approximate the motor’s output torque — together with joint position and velocity and a dynamics model of the robot. The idea is that the model predicts how much torque would be needed with no external force, and the difference between the torque actually measured and that prediction is attributed to the external force. A commonly used tool is the generalized momentum observer, which avoids differentiating acceleration and so is less noisy. The benefit is skipping an expensive, fragile force sensor, letting low-cost arms, quadrupeds, and humanoids do collision detection and rough force control; the difficulty is that gearbox friction, model error, and temperature changes introduce significant error, so precision is generally worse than with a dedicated sensor. Recent work uses learned methods to compensate for the residual torque, such as K-VARK (2025) and a 2026 hybrid estimation method for wheel-legged robots.","example":"An arm with no force sensor compares its motor current against what the dynamics model predicts; a sudden large mismatch in joint torque is taken to mean it has hit a person, and it stops immediately.","related":["Generalized Momentum Observer","Torque Constant (Kt)","Collision Detection (Robot Safety)","Friction Compensation","Proprioceptive Actuator","Six-Axis Force/Torque Sensor"]},{"id":"contact-detection","category":"perception","sec":2,"tier":2,"sources":[{"title":"Multimodal Contact Detection using Auditory and Force Features for Reliable Object Placing in Household Environments (arXiv 2012.01583)","url":"https://arxiv.org/abs/2012.01583"},{"title":"Tactile sensor（Wikipedia）","url":"https://en.wikipedia.org/wiki/Tactile_sensor"}],"as_of":"","related_ids":["tactile-sensor","six-axis-force-torque-sensor","slip-detection","extrinsic-contact-sensing","sensorless-force-estimation","contact-estimation"],"name":"Contact Detection","alt":"接触检测","abbr":"","aliases":["Contact Sensing"],"one_liner":"Determining whether, when, and where a robot has touched an object or its environment.","explanation":"Contact detection determines whether, when, and where a robot part — a finger, gripper, tool, or the body itself — has touched an object or the environment. Signals come mainly from three sources: a change in a tactile sensor's or electronic skin's reading; a sudden jump in the force or torque read by a wrist F/T sensor; or, when no dedicated sensor is available, the residual between motor current or joint torque and what a dynamics model predicts. Vision alone is often unreliable here, since the hand frequently blocks the object right at the moment of contact. Contact detection is the trigger condition for many manipulation routines — close the gripper until contact then stop, lower until the table is touched then release, judge whether an assembly step is complete. Fixed thresholds are prone to misjudging in home settings, so some work instead fuses sound with force signals, or uses a learned detector. A legged robot's judgment of whether a foot has touched the ground is usually given its own separate name, contact estimation.","example":"While lowering an arm slowly to place a cup, the wrist force sensor reads a sudden jump in vertical force past a threshold, which is taken to mean the cup's base has touched the table, and the gripper opens immediately.","related":["Tactile Sensor","Six-Axis Force/Torque Sensor","Slip Detection","Extrinsic Contact Sensing","Sensorless Force Estimation","Contact Estimation"]},{"id":"foot-force-sensor","category":"perception","sec":2,"tier":3,"sources":[{"title":"Unitree Go2 产品页（规格表：足端力传感器仅 EDU 版）","url":"https://www.unitree.com/go2"},{"title":"unitree_legged_sdk comm.h（LowState footForce 字段）","url":"https://github.com/unitreerobotics/unitree_legged_sdk/blob/master/include/unitree_legged_sdk/comm.h"}],"as_of":"2026-09","related_ids":["contact-estimation","six-axis-force-torque-sensor","ground-reaction-force","zero-moment-point","leg-odometry","sensorless-force-estimation"],"name":"Foot Force Sensor","alt":"足底力传感器","abbr":"","aliases":["Foot Contact Sensor"],"one_liner":"A sensor mounted on a legged robot’s foot or ankle that measures the contact force between foot and ground.","explanation":"A foot force sensor measures ground-contact force at a legged robot’s foot; common forms include a single-axis force or pressure sensor at a quadruped’s foot tip, a six-axis force sensor at a humanoid’s ankle, or a pressure array laid across the sole. It answers two questions: is this foot on the ground, and how large is the ground reaction force? The first answer feeds contact estimation and leg odometry (you need to know which foot is bearing weight to infer how the body is moving); the second lets you compute the center of pressure and zero-moment point for judging balance. The foot takes an impact every step, so the sensor wears out and its readings drift, which is why many quadruped products skip it and instead do sensorless force estimation from joint torque or motor current. On the Unitree Go2, for example, only the EDU version has foot force sensors.","example":"Older Unitree Go1-generation quadruped SDKs (unitree_legged_sdk) provide a footForce field in the low-level LowState message with one force reading per foot, which the control program can use to judge each foot’s touchdown moment.","related":["Contact Estimation","Six-Axis Force/Torque Sensor","Ground Reaction Force (GRF)","Zero Moment Point","Leg Odometry","Sensorless Force Estimation"]},{"id":"contact-estimation","category":"perception","sec":2,"tier":3,"sources":[{"title":"Legged Robot State Estimation using Invariant Kalman Filtering and Learned Contact Events (arXiv 2106.15713, CoRL 2021)","url":"https://arxiv.org/abs/2106.15713"}],"as_of":"","related_ids":["leg-odometry","invariant-extended-kalman-filter","state-estimation","foot-force-sensor","learned-state-estimator","stance-phase-swing-phase"],"name":"Contact Estimation","alt":"接触估计","abbr":"","aliases":["Touchdown Detection","Foot Contact Estimation"],"one_liner":"Determines whether each foot of a legged robot is currently on the ground or in the air.","explanation":"Contact estimation is a subproblem of state estimation for legged robots: at every control cycle, it decides whether each leg is touching the ground, usually outputting a contact probability. There are several general approaches: reading foot-force sensors and thresholding them; using joint torques and a dynamics model to infer ground reaction force; probabilistically fusing gait phase and foot height; or, as in Lin and colleagues (CoRL 2021), training a neural network to learn touchdown events directly from proprioceptive data like joint encoders and the IMU (inertial measurement unit), with no dedicated contact sensor needed. Downstream modules all depend on this judgment: leg odometry assumes a foot in contact isn’t moving and uses that to infer body velocity, so a wrong call causes drift; the controller also relies on it to switch between stance and swing phases.","example":"If a quadruped robot trotting over gravel misjudges a slipping foot as “stably in contact,” leg odometry built on an invariant extended Kalman filter will mistake the slip for body motion, and the position estimate will drift as a result.","related":["Leg Odometry","Invariant Extended Kalman Filter","State Estimation","Foot Force Sensor","Learned State Estimator","Stance Phase / Swing Phase"]},{"id":"tactile-sensor","category":"perception","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Tactile sensor","url":"https://en.wikipedia.org/wiki/Tactile_sensor"},{"title":"arXiv 2005.14679: DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor","url":"https://arxiv.org/abs/2005.14679"}],"as_of":"","related_ids":["vision-based-tactile-sensor","electronic-skin","gelsight","digit","slip-detection","visuo-tactile-fusion"],"name":"Tactile Sensor","alt":"触觉传感器","abbr":"","aliases":["Tactile Sensing"],"one_liner":"A sensor that lets a robot feel contact, measuring pressure distribution, contact location, shear force, and slip.","explanation":"A tactile sensor measures the information generated when a robot contacts an object, including pressure distribution, contact location, normal and shear force, vibration, and slip; some also sense texture and temperature. Common principles include piezoresistive, capacitive, piezoelectric, and magnetic sensing, as well as vision-based tactile sensing, where a small camera behind a transparent elastomer photographs how the surface deforms at the point of contact — representative examples are GelSight, which originated at MIT, and DIGIT, a low-cost open-source sensor from Facebook AI and collaborators published in 2020. A camera can't see the small patch where a finger meets an object, and it can't measure weight, softness, or friction — exactly the information needed for tasks like twisting off a bottle cap, inserting a cable, or gently squeezing an egg. Tactile data is commonly fed into a policy as tactile images or arrays of individual sensing elements (taxels), alongside vision.","example":"In the DIGIT paper, researchers equipped a multi-fingered robotic hand with DIGIT sensors and trained a neural-network controller on the tactile signal to roll a glass marble around in the hand's palm.","related":["Vision-Based Tactile Sensor","Electronic Skin","GelSight","DIGIT","Slip Detection","Visuo-Tactile Fusion"]},{"id":"electronic-skin","category":"perception","sec":3,"tier":2,"sources":[{"title":"Electronic skin（Wikipedia）","url":"https://en.wikipedia.org/wiki/Electronic_skin"},{"title":"ReSkin: versatile, replaceable, lasting tactile skins (arXiv 2111.00071)","url":"https://arxiv.org/abs/2111.00071"},{"title":"Meta AI: Teaching robots to perceive, understand, and interact through touch","url":"https://ai.meta.com/blog/teaching-robots-to-perceive-understand-and-interact-through-touch/"}],"as_of":"","related_ids":["tactile-sensor","piezoresistive-tactile-sensing","capacitive-tactile-sensing","taxel","reskin-anyskin","vision-based-tactile-sensor"],"name":"Electronic Skin","alt":"电子皮肤","abbr":"E-skin","aliases":["E-Skin","Robot Skin","Flexible Tactile Sensor"],"one_liner":"A soft, bendable, large-area tactile sensing layer applied to a robot's surface, mimicking human skin.","explanation":"Electronic skin refers to a flexible, stretchable, large-area sensing layer that conforms to curved surfaces, mimicking human or animal skin in sensing pressure, temperature, and strain; some materials can even self-heal. Common sensing principles are piezoresistive, capacitive, piezoelectric, and magnetic, built from large arrays of individual sensing elements (taxels). The field originated in flexible-electronics and materials research and is also used in prosthetics and wearable health monitoring. On a robot, electronic skin can cover the palm, an arm, or even the entire body, sensing contact location, human-robot collisions, and grip pressure distribution — things a fingertip-sized vision-based tactile sensor simply can't do. The difficulties are wear resistance, wiring and cross-talk between elements, and keeping the taxels consistent and calibrated. A representative example is ReSkin (CoRL 2021), from Meta and Carnegie Mellon University, which uses magnetic sensing so the skin's surface layer can be swapped out and replaced.","example":"Applying a layer of piezoresistive electronic skin to a dexterous hand's palm and reading out each taxel's pressure during a grasp produces a pressure-distribution map showing which region of the palm the object presses on, and how hard.","related":["Tactile Sensor","Piezoresistive Tactile Sensing","Capacitive Tactile Sensing","Taxel","ReSkin / AnySkin","Vision-Based Tactile Sensor"]},{"id":"tactile-array","category":"perception","sec":3,"tier":3,"sources":[{"title":"Tactile Sensing—From Humans to Humanoids (Dahiya et al., IEEE T-RO 2010)","url":"https://doi.org/10.1109/TRO.2009.2033627"}],"as_of":"","related_ids":["taxel","tactile-image","tactile-sensor","electronic-skin","piezoresistive-tactile-sensing","capacitive-tactile-sensing"],"name":"Tactile Array","alt":"触觉阵列","abbr":"","aliases":["Tactile Array Sensor","Array-Type Tactile Sensor"],"one_liner":"Many small tactile sensing elements arranged in a grid, measuring the pressure distribution across a contact surface.","explanation":"A tactile array is a sensor made of many tactile elements (taxels, each of which measures pressure at one point) arranged in rows and columns on a flexible circuit or skin. A single force sensor can only tell you the total force applied, whereas an array gives the distribution of pressure across the contact surface: where contact is happening, how large the contact area is, its shape, and where the center of force is. Common sensing principles are piezoresistive, capacitive, piezoelectric, and magnetic, typically read out by scanning row by row and column by column. Tactile arrays are commonly mounted on dexterous-hand fingertips, finger pads, and palms, or built into large-area electronic skin. Design involves trading off element density (spatial resolution), sampling rate, the number of wires, and durability. Compared with a vision-based tactile sensor that watches a gel deform through a camera, an array is thinner and easier to fit onto a curved surface, but usually has lower resolution.","example":"A 4×4 capacitive tactile array is attached to each dexterous-hand fingertip; while grasping an egg, it shows exactly which elements the pressure concentrates on, revealing whether the egg is about to slip.","related":["Taxel","Tactile Image","Tactile Sensor","Electronic Skin","Piezoresistive Tactile Sensing","Capacitive Tactile Sensing"]},{"id":"taxel","category":"perception","sec":3,"tier":3,"sources":[{"title":"Tactile Sensing—From Humans to Humanoids (Dahiya et al., IEEE T-RO 2010)","url":"https://doi.org/10.1109/TRO.2009.2033627"}],"as_of":"","related_ids":["tactile-array","tactile-image","tactile-sensor","electronic-skin","piezoresistive-tactile-sensing","uskin"],"name":"Taxel","alt":"触觉单元（触元）","abbr":"","aliases":["Tactile Pixel","Tactile Element"],"one_liner":"The smallest individual sensing element in a tactile array — the tactile equivalent of an image pixel.","explanation":"“Taxel” is a blend of “tactile” and “pixel,” referring to the smallest individually readable sensing element in a tactile array. Just as an image is made of pixels, a tactile array is made of many taxels, each measuring the pressure over the small patch of surface it covers (some also measure force along three directions). The number and spacing of taxels sets the tactile spatial resolution: closer spacing resolves finer shapes and edges, but more taxels also means more wiring, more readout circuitry, and more data. A tactile sensor is often described by how many taxels it has and how many millimeters apart they are, as a measure of its fineness. Vision-based tactile sensors have no discrete taxels, and are usually described instead by camera pixel count or the number of tracked markers.","example":"A 16×16 fingertip tactile array has 256 taxels, each outputting one pressure value, which together form a 16×16 tactile image.","related":["Tactile Array","Tactile Image","Tactile Sensor","Electronic Skin","Piezoresistive Tactile Sensing","uSkin"]},{"id":"piezoresistive-tactile-sensing","category":"perception","sec":3,"tier":3,"sources":[{"title":"Piezoresistive effect - Wikipedia","url":"https://en.wikipedia.org/wiki/Piezoresistive_effect"},{"title":"Force-sensing resistor - Wikipedia","url":"https://en.wikipedia.org/wiki/Force-sensing_resistor"},{"title":"3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv 2410.24091)","url":"https://arxiv.org/abs/2410.24091"}],"as_of":"","related_ids":["tactile-sensor","tactile-array","piezoelectric-tactile-sensing","capacitive-tactile-sensing","strain-gauge","3d-vitac"],"name":"Piezoresistive Tactile Sensing","alt":"压阻式触觉传感","abbr":"","aliases":["Piezoresistive Sensor","Piezoresistive Tactile Array","Force-Sensing Resistor (FSR)"],"one_liner":"Measuring pressure from a material’s change in electrical resistance under load; cheap and easy to build into large sensor arrays.","explanation":"The piezoresistive effect is a change in a material’s electrical resistivity when it is deformed by force; Lord Kelvin first observed it in metals in 1856, and in 1954 Smith found the effect to be far larger in silicon and germanium. Piezoresistive tactile sensors apply this effect to conductive rubber, conductive polymer film, or silicon strain elements: resistance changes under pressure, so measuring resistance gives pressure, and many such elements arranged in a row-and-column electrode grid can output a full pressure-distribution map. The technology is structurally simple, cheap, thin, and flexible, and can measure static force, making it one of the most common tactile approaches; the force-sensing resistor (FSR) belongs to this family. Its drawbacks are relatively high hysteresis and drift, and lower precision. 3D-ViTac (2024) sandwiches a Velostat piezoresistive film between conductive yarn electrodes to build a 16×16 tactile pad.","example":"3D-ViTac attaches piezoresistive tactile pads to gripper fingers, at a cost of about US$20 per pad including the readout board, giving the policy a pressure value at every contact point.","related":["Tactile Sensor","Tactile Array","Piezoelectric Tactile Sensing","Capacitive Tactile Sensing","Strain Gauge","3D-ViTac"]},{"id":"capacitive-tactile-sensing","category":"perception","sec":3,"tier":3,"sources":[{"title":"A Flexible and Robust Large Scale Capacitive Tactile System for Robots (arXiv 1411.6837)","url":"https://arxiv.org/abs/1411.6837"},{"title":"Tactile sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Tactile_sensor"},{"title":"Capacitive sensing - Wikipedia","url":"https://en.wikipedia.org/wiki/Capacitive_sensing"}],"as_of":"","related_ids":["tactile-sensor","piezoresistive-tactile-sensing","piezoelectric-tactile-sensing","electronic-skin","taxel","proximity-sensor"],"name":"Capacitive Tactile Sensing","alt":"电容式触觉传感","abbr":"","aliases":["Capacitive Sensor","Capacitive Skin"],"one_liner":"A tactile-sensing method that infers contact force from the change in capacitance as pressure alters electrode spacing or a dielectric layer.","explanation":"Capacitive tactile sensing is one of the main tactile-sensor principles, alongside piezoresistive, piezoelectric, magnetic, and vision-based sensing. The basic structure is two layers of electrodes sandwiching a compressible dielectric material: pressure changes the spacing or overlap area between the electrodes, which changes the capacitance, and readout circuitry converts that into pressure; arranging many such cells (taxels) into an array gives a full contact-pressure distribution. The advantages are high sensitivity, the ability to make sensors small, thin, and flexible, ease of covering large surfaces, and the ability to sense a hand approaching before contact; the drawbacks are susceptibility to parasitic capacitance and electromagnetic interference, temperature drift, and hysteresis introduced by the elastic dielectric layer. A well-known example is the large-area capacitive skin the Italian Institute of Technology (IIT) built for the iCub humanoid, with an improved version using a 3D fabric dielectric layer to reduce hysteresis.","example":"A layer of capacitive tactile skin stuck to a collaborative robot arm’s casing sees a sudden capacitance change when a human hand touches or approaches a given region, and the controller immediately slows down or stops in response.","related":["Tactile Sensor","Piezoresistive Tactile Sensing","Piezoelectric Tactile Sensing","Electronic Skin","Taxel","Proximity Sensor"]},{"id":"piezoelectric-tactile-sensing","category":"perception","sec":3,"tier":3,"sources":[{"title":"Piezoelectric sensor - Wikipedia","url":"https://en.wikipedia.org/wiki/Piezoelectric_sensor"},{"title":"Tactile Robotics: An Outlook (arXiv 2508.11261)","url":"https://arxiv.org/abs/2508.11261"},{"title":"A-SLIP: Acoustic Sensing for Continuous In-hand Slip Estimation (arXiv 2604.08528)","url":"https://arxiv.org/abs/2604.08528"}],"as_of":"2026-04","related_ids":["tactile-sensor","piezoresistive-tactile-sensing","capacitive-tactile-sensing","slip-detection","contact-microphone","electronic-skin"],"name":"Piezoelectric Tactile Sensing","alt":"压电式触觉传感","abbr":"","aliases":["Piezoelectric Sensor","Piezoelectric Tactile Sensor","PVDF Tactile Sensing"],"one_liner":"Sensing contact and vibration using the piezoelectric effect, where a material generates an electric charge when it is deformed by force.","explanation":"Piezoelectric tactile sensing relies on the piezoelectric effect: materials such as quartz, lead zirconate titanate (PZT) ceramic, or polyvinylidene fluoride (PVDF) film generate a surface charge when deformed by force, and reading out that electrical signal senses changes in force. It responds quickly and has a wide frequency bandwidth, making it very sensitive to the high-frequency vibration produced by sliding, impact, or rubbing against a textured surface; however, the charge gradually leaks away, so it cannot measure a constant static pressure, and is therefore often paired with a piezoresistive or capacitive sensing element. It is one of the main tactile-sensing principles, alongside piezoresistive, capacitive, magnetic, and optical sensing. Recently, several robotics projects have repurposed piezoelectric microphones as tactile sensors — for example, A-SLIP (2026) places several piezoelectric microphones behind a gripper’s silicone pad and uses a neural network to estimate whether an object is slipping, and in which direction and by how much.","example":"A-SLIP mounts 4 piezoelectric microphones behind the textured silicone pad of a parallel gripper; the vibration from an object beginning to slip is detected, and the controller immediately increases the grip force.","related":["Tactile Sensor","Piezoresistive Tactile Sensing","Capacitive Tactile Sensing","Slip Detection","Contact Microphone","Electronic Skin"]},{"id":"magnetic-tactile-sensing","category":"perception","sec":3,"tier":3,"sources":[{"title":"ReSkin: versatile, replaceable, lasting tactile skins (CoRL 2021)","url":"https://arxiv.org/abs/2111.00071"},{"title":"AnySkin: Plug-and-play Skin Sensing for Robotic Touch","url":"https://arxiv.org/abs/2409.08276"}],"as_of":"2024-09","related_ids":["tactile-sensor","reskin-anyskin","uskin","vision-based-tactile-sensor","electronic-skin","slip-detection"],"name":"Magnetic Tactile Sensing","alt":"磁性触觉传感","abbr":"","aliases":["Hall-Effect Tactile Sensing","Magnetic Tactile Skin"],"one_liner":"Mixes magnetic particles into a soft gel and uses a magnetometer to sense contact from the magnetic-field change deformation causes.","explanation":"Magnetic tactile sensing is one approach to building tactile sensors: magnetized particles or small magnets are embedded in a soft elastomer, with a Hall-effect sensor or magnetometer placed underneath. When the elastomer touches an object it deforms, the embedded magnetic source moves with it, and the magnetic-field reading changes accordingly; calibration or machine learning then converts this reading into contact location, pressure, and shear force (force along the surface). The advantage is that the soft gel layer is separate from the circuit board, so a worn gel pad can simply be swapped out, keeping cost low and shape easy to customize; the drawback is susceptibility to external magnetic fields and nearby motors, with spatial resolution usually lower than a vision-based tactile sensor’s. Landmark examples include ReSkin (CoRL 2021), a collaboration between Carnegie Mellon University and Meta, and Lerrel Pinto’s group’s AnySkin (2024), the latter built around the idea that swapping in a fresh skin lets an already-trained manipulation policy keep working without recalibration.","example":"The AnySkin paper mounts magnetic skin on gripper fingertips for slip detection and policy learning, and shows that after swapping in a different, fresh skin the policy still works directly.","related":["Tactile Sensor","ReSkin / AnySkin","uSkin","Vision-Based Tactile Sensor","Electronic Skin","Slip Detection"]},{"id":"reskin-anyskin","category":"perception","sec":3,"tier":3,"sources":[{"title":"ReSkin: versatile, replaceable, lasting tactile skins (arXiv 2111.00071)","url":"https://arxiv.org/abs/2111.00071"},{"title":"ReSkin 项目主页","url":"https://reskin.dev/"},{"title":"AnySkin: Plug-and-play Skin Sensing for Robotic Touch 项目主页","url":"https://any-skin.github.io/"}],"as_of":"2024-09","related_ids":["tactile-sensor","magnetic-tactile-sensing","electronic-skin","slip-detection","tactile-representation-learning","vision-based-tactile-sensor"],"name":"ReSkin / AnySkin","alt":"ReSkin / AnySkin","abbr":"","aliases":["ReSkin","AnySkin"],"one_liner":"A magnetic tactile skin: a soft pad embedded with magnetic particles sits over a magnetometer that reads contact from field changes.","explanation":"ReSkin was proposed by Bhirangi, Hellebrekers, Majidi, and Gupta at Carnegie Mellon and Meta AI (FAIR), published at CoRL 2021. It mixes magnetized particles into an elastomer; when the elastomer is deformed by pressure, the magnetic field it produces changes, which a magnetometer circuit underneath measures to estimate contact location and force. The soft skin wears out but is separate from the electronics, so it can simply be swapped for a new one, with self-supervised learning used to adapt to the replacement skin. In 2024, Lerrel Pinto’s group at NYU and collaborators introduced AnySkin, which simplifies fabrication and installation — it snaps onto a gripper or dexterous hand like a phone case — and makes the signal more consistent across different individual skins: the same manipulation policy loses only about 13% performance after a skin swap with AnySkin, versus about 43% with ReSkin. AnySkin’s design files are open-sourced, and it is also available through commercial channels.","example":"An AnySkin pad is mounted on the fingertip of a Franka gripper; a policy trained on the tactile signal to plug in a USB cable still mostly works after the pad is swapped for a fresh one.","related":["Tactile Sensor","Magnetic Tactile Sensing","Electronic Skin","Slip Detection","Tactile Representation Learning","Vision-Based Tactile Sensor"]},{"id":"uskin","category":"perception","sec":3,"tier":3,"sources":[{"title":"XELA Robotics 官网","url":"https://www.xelarobotics.com/"},{"title":"Self-supervised perception for tactile skin covered dexterous hands (Sparsh-skin, arXiv:2505.11420)","url":"https://arxiv.org/abs/2505.11420"}],"as_of":"2026-09","related_ids":["tactile-sensor","magnetic-tactile-sensing","electronic-skin","allegro-hand","slip-detection","xela-robotics"],"name":"uSkin","alt":"uSkin","abbr":"","aliases":["XELA uSkin"],"one_liner":"A magnetic tactile skin from XELA Robotics where every sensing point measures both pressure and shear force at once.","explanation":"uSkin is a tactile sensor product line from XELA Robotics of Japan; XELA was spun out of Waseda University in 2018 and is headquartered in Tokyo. uSkin uses magnetic tactile sensing: when its soft outer layer deforms under force, the magnetic field inside it changes, and a magnetic sensor reads out that change in flux, with each taxel (sensing element) reporting both normal pressure and tangential shear force. It comes in versions that attach to a gripper, a dexterous hand, or any arbitrary surface, and in research it is commonly mounted on the fingertips, knuckles, and palm of an Allegro Hand for slip detection and contact-force estimation. Magnetic signals are hard to interpret and troublesome to calibrate, which is why Sparsh-skin, from Akash Sharma and colleagues in 2025, uses self-supervised learning to train a general-purpose representation for this kind of magnetic hand-skin sensor.","example":"With uSkin attached to the fingertips of an Allegro Hand, a sudden change in shear force while grasping a cup signals that the object is about to slip, and the hand immediately grips harder.","related":["Tactile Sensor","Magnetic Tactile Sensing","Electronic Skin","Allegro Hand","Slip Detection","XELA Robotics"]},{"id":"paxini-px-6ax","category":"perception","sec":3,"tier":3,"sources":[{"title":"帕西尼 PX6AX GEN4 产品页","url":"https://www.paxini.com/cn/ax/gen4"},{"title":"帕西尼 PX6AX GEN3 产品页","url":"https://www.paxini.com/cn/ax/gen3"},{"title":"帕西尼感知官网","url":"https://www.paxini.com/cn"}],"as_of":"2026-09","related_ids":["tactile-sensor","paxini-tech","paxini-tora-one","dexterous-hand","normal-force-and-tangential-force","slip-detection"],"name":"PaXini PX-6AX","alt":"帕西尼 PX-6AX 多维触觉传感器","abbr":"","aliases":["PX6AX","PX6AX GEN3","PX6AX GEN4","PaXini PX-6AX Multi-Dimensional Tactile Sensor"],"one_liner":"A multi-dimensional tactile sensor from PaXini that outputs distributed force, net force, and torque at the point of contact.","explanation":"The PX-6AX is a tactile sensor series from PaXini (帕西尼感知), described by the company as an ITPU multi-dimensional tactile sensing unit, now in its fourth generation (GEN4); the transduction principle it uses has not been disclosed in detail. It outputs a 3D array of distributed force, a 3D net force, and 3D torque, and the company states it can resolve more than 15 tactile parameters. Per the published specs: normal-force range 0–25 N, shear (tangential) force ±10 N, spatial resolution 0.1 mm, maximum output rate 1000 Hz, IP68 rated, with chip-level magnetic shielding highlighted; the smallest detectable force is 0.01 N for GEN3 and 0.005 N for GEN4. It comes in sizes ranging from fingertip pads to palm pads, and is used on dexterous-hand fingertips, grippers, and robot skin — PaXini’s own DexH13 dexterous hand and TORA-ONE humanoid both feature this kind of tactile sensing.","example":"A PX-6AX mounted on a dexterous-hand fingertip reads out normal pressure and shear force at the same time while pinching an egg; a sudden jump in shear force signals an incipient slip.","related":["Tactile Sensor","PaXini Tech","PaXini TORA-ONE","Dexterous Hand","Normal Force and Tangential (Shear) Force","Slip Detection"]},{"id":"biotac","category":"perception","sec":3,"tier":3,"sources":[{"title":"Interpreting and Predicting Tactile Signals for the SynTouch BioTac (NVIDIA, arXiv:2101.05452)","url":"https://arxiv.org/abs/2101.05452"},{"title":"ACROSS: A Deformation-Based Cross-Modal Representation for Robotic Tactile Perception (arXiv:2411.08533)","url":"https://arxiv.org/abs/2411.08533"}],"as_of":"2024-11","related_ids":["tactile-sensor","slip-detection","electronic-skin","vision-based-tactile-sensor","gelsight","dexterous-hand"],"name":"BioTac","alt":"BioTac","abbr":"","aliases":["SynTouch BioTac","BioTac SP"],"one_liner":"SynTouch’s biomimetic fingertip tactile sensor, which senses contact force, micro-vibration, and temperature.","explanation":"BioTac is a fingertip-shaped tactile sensor from the US company SynTouch, developed out of biomimetic tactile research by Gerald Loeb’s group (Wettels, Fishel, and others). Its structure is a rigid core wrapped in a rubber skin, with conductive liquid filling the space between; the core’s surface carries 19 sensing electrodes and 4 excitation electrodes, and when something makes contact, the liquid layer deforms and each electrode’s voltage changes accordingly, from which contact location, force magnitude, and direction can be inferred. It can also sense micro-vibration and temperature — the former is useful for recognizing texture and detecting slip, the latter for telling materials apart. BioTac was long a high-end benchmark in robot tactile research, commonly mounted on dexterous-hand fingertips, and it built up public datasets on slip direction and grasp stability. Its drawback is that the relationship between electrode signals and deformation is complex and hard to interpret — an NVIDIA paper in 2021 built a model specifically to address this. According to a 2024 paper, the sensor has since been discontinued, which has made these datasets hard to keep using.","example":"The BioTac SP slip-direction dataset mounts BioTac on a robotic fingertip and records electrode signals during grasping, used to train a model that judges which direction an object is sliding.","related":["Tactile Sensor","Slip Detection","Electronic Skin","Vision-Based Tactile Sensor","GelSight","Dexterous Hand"]},{"id":"vision-based-tactile-sensor","category":"perception","sec":3,"tier":2,"sources":[{"title":"GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force (Sensors 2017)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5751610/"},{"title":"DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor (RA-L 2020)","url":"https://arxiv.org/abs/2005.14679"}],"as_of":"","related_ids":["gelsight","digit","tactile-sensor","photometric-stereo","marker-tracking","tactile-image"],"name":"Vision-Based Tactile Sensor","alt":"视触觉传感器","abbr":"VBTS","aliases":["VBTS","Visuotactile Sensor","Optical Tactile Sensor","GelSight-Style Sensor"],"one_liner":"A tactile sensor that uses a built-in camera to photograph how a soft surface deforms on contact.","explanation":"A vision-based tactile sensor coats the surface of a transparent elastic gel with a reflective membrane, with LEDs and a tiny camera inside; when an object presses on it, the gel surface deforms, and what the camera captures is essentially a “tactile image.” Photometric stereo — reconstructing surface normals from how shading changes under lights from different directions — can recover fine geometry of the contact surface, while the displacement of markers embedded in the gel reveals shear force and slip. The best-known examples are MIT’s GelSight, first prototyped in 2009, and Meta’s DIGIT, an open-source, small, low-cost design released in 2020 that’s compact enough to mount on a dexterous fingertip. The advantage is high resolution with image output that vision networks can process directly; the downsides are that the gel surface wears down, the sensor has some thickness, and force values still need separate calibration.","example":"Researchers mounted two DIGIT sensors on two fingers of an Allegro dexterous hand and used a tactile model-predictive controller to roll a glass marble back and forth between the fingers, predicting contact changes from the tactile images.","related":["GelSight","DIGIT","Tactile Sensor","Photometric Stereo","Marker Tracking","Tactile Image"]},{"id":"photometric-stereo","category":"perception","sec":3,"tier":3,"sources":[{"title":"Photometric stereo - Wikipedia","url":"https://en.wikipedia.org/wiki/Photometric_stereo"},{"title":"GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force (Sensors 2017, PMC)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5751610/"}],"as_of":"","related_ids":["vision-based-tactile-sensor","gelsight","tactile-image","surface-normal-estimation","digit","taxim-an-example-based-simulation-model-for-gelsight-tactile"],"name":"Photometric Stereo","alt":"光度立体","abbr":"","aliases":[],"one_liner":"Keeping the camera fixed and lighting a surface from several directions, then recovering its shape from how the shading changes.","explanation":"Photometric stereo was proposed by Woodham in 1980. The camera stays in a fixed position while the object is lit from at least three known directions; under the assumption of Lambertian reflectance (a surface that scatters light evenly in every direction), each pixel’s brightness depends only on the surface normal and the lighting direction, so solving a small linear equation recovers the normal at that point, and integrating the normals yields a depth map. It can recover very fine surface texture, but performs poorly on metallic, glossy, or transparent surfaces. In embodied AI, its most important use is inside vision-based tactile sensors: GelSight illuminates a soft elastic gel surface with red, green, and blue LEDs from different directions, and a camera behind the gel takes a single image; a calibrated lookup table then converts the colors into surface normals, reconstructing the 3D shape of whatever is touching the gel.","example":"A GelSight fingertip is pressed onto a coin; the height map reconstructed by photometric stereo shows the coin’s embossed relief pattern.","related":["Vision-Based Tactile Sensor","GelSight","Tactile Image","Surface Normal Estimation","DIGIT","Taxim: An Example-based Simulation Model for GelSight Tactile Sensors"]},{"id":"gelsight","category":"perception","sec":3,"tier":2,"sources":[{"title":"Improved GelSight Tactile Sensor for Measuring Geometry and Slip (arXiv:1708.00922)","url":"https://arxiv.org/abs/1708.00922"},{"title":"GelSight 官网","url":"https://www.gelsight.com/"},{"title":"GelSight Mini 产品页","url":"https://www.gelsight.com/gelsightmini/"}],"as_of":"2026-09","related_ids":["vision-based-tactile-sensor","gelsight-mini","digit","photometric-stereo","marker-tracking","tactile-image"],"name":"GelSight","alt":"GelSight","abbr":"","aliases":["GelSight Sensor","GelSight Inc."],"one_liner":"A vision-based tactile sensor that reads touch from a camera watching an elastic gel deform, also the name of the company that makes it.","explanation":"GelSight is a family of vision-based tactile sensors: a piece of elastic gel with a reflective coating on its surface is lit from different directions by internal LEDs, and a camera on the back photographs how the gel deforms under pressure, using photometric stereo — inferring surface orientation from how shading changes under different lighting directions — to reconstruct the 3D shape of the contact surface. The approach originated in Edward Adelson's group at MIT, first appearing in a 2009 CVPR paper by Micah Johnson and Adelson; later versions added marker dots to the gel's surface, tracking their movement to measure shear force and slip. GelSight is also the name of the company that spun out of this research (GelSight Inc., based in Waltham, Massachusetts), whose products include the robotics-focused GelSight Mini (which supports ROS/ROS 2 and PyTouch) and industrial measurement products such as Mobile and Modulus. DIGIT, GelSlim, and 9DTact all follow the same ‘camera watching a gel’ approach.","example":"Mounting a GelSight Mini on a gripper's fingertip to pick up a screw, the tactile image clearly shows the imprint of the screw's threads, letting you judge which way the screw is oriented in the hand, and whether it's slipping.","related":["Vision-Based Tactile Sensor","GelSight Mini","DIGIT","Photometric Stereo","Marker Tracking","Tactile Image"]},{"id":"gelsight-mini","category":"perception","sec":3,"tier":3,"sources":[{"title":"GelSight Mini 官方产品页","url":"https://www.gelsight.com/gelsightmini/"},{"title":"gelsightinc/gsrobotics GitHub（Mini SDK 与 FAQ）","url":"https://github.com/gelsightinc/gsrobotics"}],"as_of":"2026-09","related_ids":["gelsight","vision-based-tactile-sensor","photometric-stereo","marker-tracking","tactile-image","slip-detection"],"name":"GelSight Mini","alt":"GelSight Mini","abbr":"","aliases":[],"one_liner":"GelSight’s commercial compact vision-based tactile sensor, ready to use over USB.","explanation":"GelSight Mini is a compact vision-based tactile sensor from the company GelSight, which grew out of the GelSight technology developed in Edward Adelson’s group at MIT. A camera inside the sensor photographs how a soft silicone pad deforms under pressure, and photometric stereo recovers the contact surface’s 3D shape; when the gel surface carries markers, tracking their displacement can estimate shear force and slip. The company describes it as the first commercial tactile sensor with spatial resolution exceeding human touch. According to the official GitHub, it uses a single USB 3 cable for both power and data, runs at 25 FPS, and ships with Python examples for 3D point cloud reconstruction and marker tracking. It’s compact, priced for researchers and hobbyists, and suits mounting directly on a gripper fingertip — one of the most commonly used vision-based tactile sensors in labs.","example":"Mounting two GelSight Mini sensors on the two fingertips of a parallel gripper lets you watch the contact region and marker displacement in the tactile image after picking up an object, to judge whether it is starting to slip.","related":["GelSight","Vision-Based Tactile Sensor","Photometric Stereo","Marker Tracking","Tactile Image","Slip Detection"]},{"id":"digit","category":"perception","sec":3,"tier":2,"sources":[{"title":"DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor (arXiv 2005.14679)","url":"https://arxiv.org/abs/2005.14679"},{"title":"Meta AI: Teaching robots to perceive, understand, and interact through touch","url":"https://ai.meta.com/blog/teaching-robots-to-perceive-understand-and-interact-through-touch/"}],"as_of":"","related_ids":["vision-based-tactile-sensor","gelsight","digit-360","tacto-a-fast-flexible-and-open-source-simulator-for-high-res","tactile-image","agility-robotics-digit"],"name":"DIGIT","alt":"DIGIT 视触觉传感器","abbr":"","aliases":["Meta DIGIT","DIGIT Tactile Sensor"],"one_liner":"Meta's open-source, fingertip-sized, low-cost vision-based tactile sensor that reads contact from a camera watching a gel deform.","explanation":"DIGIT is a vision-based tactile sensor open-sourced by Facebook AI Research (now Meta FAIR) and published in IEEE Robotics and Automation Letters in 2020. It works like GelSight: an internal camera shoots through a layer of elastic gel, and when an object presses against it, the gel's deformation is captured in an image that records the shape and texture of the contact area, forming a ‘tactile image.’ It's small, cheap, and easy to manufacture, making it well suited to the fingertips of a multi-fingered dexterous hand — the paper demonstrates it being used to roll a glass marble around in a hand. In 2021, Meta announced a partnership with GelSight to mass-produce it, and open-sourced the tactile library PyTouch and the simulator TACTO alongside it; a newer generation is called Digit 360. Note that it shares its name only, with no other connection, with Digit, the humanoid robot made by Agility Robotics.","example":"Mounting one DIGIT on each of a gripper's two fingers: during a grasp, the position of the contact region in the tactile image shows whether the object is held straight, and the gel's shear deformation gives early warning that the object is about to slip.","related":["Vision-Based Tactile Sensor","GelSight","Digit 360","TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors","Tactile Image","Agility Robotics Digit"]},{"id":"digit-360","category":"perception","sec":3,"tier":3,"sources":[{"title":"Meta AI blog: Advancing embodied AI through progress in touch perception, dexterity, and human-robot interaction","url":"https://ai.meta.com/blog/fair-robotics-open-source/"},{"title":"arXiv 2411.02479: Digitizing Touch with an Artificial Multimodal Fingertip","url":"https://arxiv.org/abs/2411.02479"}],"as_of":"2024-11","related_ids":["vision-based-tactile-sensor","digit","gelsight","sparsh","taxel","meta-fundamental-ai-research"],"name":"Digit 360","alt":"Digit 360","abbr":"","aliases":["Meta Digit 360"],"one_liner":"Meta’s humanoid fingertip-shaped multimodal tactile sensor, sensing force, vibration, temperature, and more.","explanation":"Digit 360 is a humanoid fingertip tactile sensor Meta FAIR released in late October 2024, manufactured and sold by the company GelSight. It continues the operating principle of DIGIT-style vision-based tactile sensors — an internal camera photographing the deformation of a soft gel surface — but shaped as a hemispherical fingertip that can sense contact from any direction. The paper reports about 8.3 million taxels (tactile pixels), able to distinguish surface detail down to 7 micrometers, with normal- and shear-force resolution around 1 mN, sensing vibration up to 10 kHz, and also sensing temperature and even smell; the fingertip has a built-in AI accelerator chip that can do fast, reflex-like local processing. Meta simultaneously released the Sparsh tactile encoder and the Digit Plexus platform, which connects multiple types of tactile sensors to the same hand.","example":"Meta’s Digit Plexus platform can mount Digit 360, DIGIT, ReSkin, and other tactile sensors on the same robotic hand, sending all their data back over a single cable, making it convenient to collect tactile data for dexterous manipulation.","related":["Vision-Based Tactile Sensor","DIGIT","GelSight","Sparsh","Taxel","Meta Fundamental AI Research"]},{"id":"gelslim","category":"perception","sec":3,"tier":3,"sources":[{"title":"GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing Finger (arXiv 1803.00628)","url":"https://arxiv.org/abs/1803.00628"},{"title":"GelSlim3.0: High-Resolution Measurement of Shape, Force and Slip in a Compact Tactile-Sensing Finger (arXiv 2103.12269)","url":"https://arxiv.org/abs/2103.12269"}],"as_of":"2021-03","related_ids":["gelsight","vision-based-tactile-sensor","slip-detection","bin-picking","parallel-jaw-gripper","gelsight-mini"],"name":"GelSlim","alt":"GelSlim","abbr":"","aliases":["GelSlim 3.0"],"one_liner":"A slim, finger-shaped vision-based tactile sensor from MIT that can be mounted on a gripper for grasping.","explanation":"GelSlim is a series of finger-shaped vision-based tactile sensors from Alberto Rodriguez’s group (the MCube Lab) at MIT, developed with Edward Adelson and others, part of the GelSight family. The original GelSight has a long optical path and a bulky body, making it hard to fit into cluttered environments; the first GelSlim (Donlon and colleagues, 2018) redesigned the optical path with mirrors and light guides to make the sensor a slim finger, swapped in a more wear-resistant gel and fabric skin, and used calibration to keep imaging stable. GelSlim 3.0 (Taylor, Dong, Rodriguez, 2021) shrank it further, measuring contact shape in real time, estimating a 3D distributed contact-force field, and detecting incipient slip; its fingertip module is swappable, and the design and software are open source.","example":"Mounting GelSlim 3.0 on a small parallel gripper for bin picking uses the tactile image to estimate the object’s position in the hand and detect whether it’s starting to slip.","related":["GelSight","Vision-Based Tactile Sensor","Slip Detection","Bin Picking","Parallel Jaw Gripper","GelSight Mini"]},{"id":"9dtact","category":"perception","sec":3,"tier":3,"sources":[{"title":"9DTact: A Compact Vision-Based Tactile Sensor for Accurate 3D Shape Reconstruction and Generalizable 6D Force Estimation (arXiv 2308.14277)","url":"https://arxiv.org/abs/2308.14277"},{"title":"9DTact project page","url":"https://linchangyi1.github.io/9DTact/"},{"title":"9DTact GitHub repository","url":"https://github.com/linchangyi1/9DTact"}],"as_of":"2024-05","related_ids":["vision-based-tactile-sensor","gelsight","digit","six-axis-force-torque-sensor","photometric-stereo","slip-detection"],"name":"9DTact","alt":"9DTact","abbr":"","aliases":[],"one_liner":"An open-source, compact vision-based tactile sensor from Tsinghua’s Huazhe Xu group that measures contact shape and 6D force.","explanation":"9DTact is a vision-based tactile sensor from Huazhe Xu’s group at Tsinghua University along with the Shanghai Qi Zhi Institute and other collaborators; the paper was published in IEEE RA-L and presented at ICRA 2024. The “9D” in the name refers to 3D shape plus 6D force. It’s a GelSight-style design: an internal camera photographs the back of a soft gel pad, and when an object presses on it, the gel deforms and reflects less light, so the image’s shading pattern reflects contact depth. It exploits the optical properties of a translucent gel to estimate three-axis force and three-axis torque directly from the image with a neural network, without needing markers on the gel surface. It was trained on about 100,000 paired image-force samples from 175 objects and generalizes to unseen objects. It measures about 32.5 × 25.5 × 25.5 millimeters, with an average shape-reconstruction error of about 0.046 millimeters. Both hardware and software are open source, with a build guide included.","example":"Mounted on a gripper fingertip, 9DTact reads the normal and shear forces at the fingertip during a grasp, which can be used to judge whether an object is starting to slip and whether to increase grip force.","related":["Vision-Based Tactile Sensor","GelSight","DIGIT","Six-Axis Force/Torque Sensor","Photometric Stereo","Slip Detection"]},{"id":"daimon-dm-tac-visuotactile-sensor","category":"perception","sec":3,"tier":3,"sources":[{"title":"戴盟机器人：DM-Tac W2 通用视触觉传感器","url":"https://www.dmrobot.com/devices/dm-tac-w2.html"},{"title":"戴盟机器人：DM-Tac F 指尖视触觉传感器","url":"https://www.dmrobot.com/devices/dm-tac-f.html"},{"title":"戴盟机器人：关于我们","url":"https://www.dmrobot.com/about.html"}],"as_of":"2026-09","related_ids":["vision-based-tactile-sensor","daimon-robotics","daimon-infinity","gelsight","vision-tactile-language-action-model","tactile-data"],"name":"Daimon DM-Tac Visuotactile Sensor","alt":"戴盟 DM-Tac 视触觉传感器","abbr":"","aliases":["DM-Tac W2","DM-Tac X","DM-Tac F"],"one_liner":"A line of vision-based tactile sensors from Daimon Robotics that senses force by imaging how a contact surface deforms.","explanation":"DM-Tac is the vision-based tactile sensor product line from Shenzhen-based Daimon Robotics, a company incubated at the Hong Kong University of Science and Technology that began formal operations in 2023. A vision-based tactile sensor places a small camera behind an elastic contact surface and infers contact shape and force from how that surface deforms in the image. According to the company’s website (as of 2026), the line includes the general-purpose DM-Tac W2 (available in two sizes, W2L and W2M), the 28°-pointed DM-Tac X for sharp-edge contact, DM-Tac F for dexterous-hand fingertips (covering the pad, side, and tip of a finger), and DM-Tac G, a two-finger gripper with sensors built in. The W2 and X sense at 384×288 resolution (about 110,000 effective sensing points) at a 120 Hz sampling rate, and the company states they can output contact geometry, a 3D deformation field, 3D distributed force, and a 6D lumped force. Daimon also uses these sensors to collect robot datasets that include touch.","example":"Mounting two DM-Tac W2 sensors on the two sides of a parallel gripper to pick up an egg lets the sensors give real-time contact deformation and distributed force, which the controller uses to adjust grip force and avoid crushing or dropping it.","related":["Vision-Based Tactile Sensor","Daimon Robotics","Daimon-Infinity","GelSight","Vision-Tactile-Language-Action Model","Tactile Data"]},{"id":"tactip","category":"perception","sec":3,"tier":3,"sources":[{"title":"The TacTip Family: Soft Optical Tactile Sensors with 3D-Printed Biomimetic Morphologies (Ward-Cherrier et al., Soft Robotics 2018)","url":"https://doi.org/10.1089/soro.2017.0052"}],"as_of":"","related_ids":["vision-based-tactile-sensor","marker-tracking","gelsight","tactile-sensor","slip-detection","tactile-image"],"name":"TacTip","alt":"TacTip","abbr":"","aliases":["Biomimetic Optical Tactile Fingertip","TacTip Family"],"one_liner":"A biomimetic optical tactile fingertip from the Bristol Robotics Laboratory that uses a camera to watch internal marker pins move.","explanation":"TacTip is a family of optical tactile sensors developed at the Bristol Robotics Laboratory in the UK, with the earliest version published around 2009. It mimics the dermal papillae structure beneath human fingertip skin: the outside is a soft, hemispherical skin, the inside is lined with rows of small pins tipped with white markers, the space between is filled with transparent gel, and a camera sits inside. When the finger touches an object, the skin deforms and moves the pins with it; the camera tracks the displacement of these markers, from which contact location, shape, edges, and shear direction can be inferred. Most of its parts can be 3D printed, keeping cost low, and it has since spawned a “TacTip family” of different sizes and shapes that can be mounted on robotic fingers or grippers. Unlike GelSight, which photographs surface texture under lighting, TacTip follows a marker-tracking approach.","example":"A TacTip mounted on the end of a robot arm is slid along an object’s contour; the displacement of its marker pins reveals the direction of the edge, letting it trace out the object’s outline.","related":["Vision-Based Tactile Sensor","Marker Tracking","GelSight","Tactile Sensor","Slip Detection","Tactile Image"]},{"id":"marker-tracking","category":"perception","sec":3,"tier":3,"sources":[{"title":"GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force (Sensors 2017)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5751610/"}],"as_of":"","related_ids":["vision-based-tactile-sensor","gelsight","slip-detection","photometric-stereo","normal-force-and-tangential-force","tactile-image"],"name":"Marker Tracking","alt":"标记点跟踪","abbr":"","aliases":["Tactile Markers","Marker Displacement","Marker Flow"],"one_liner":"Prints a dot pattern on a vision-based tactile sensor’s soft gel and tracks the dots’ displacement to estimate shear force and slip.","explanation":"Marker tracking is a common signal-processing method for vision-based tactile sensors (which photograph a soft gel’s deformation with a camera, like GelSight): rows of black dots are printed between the elastomer and its reflective coating, and on contact the camera captures how these dots shift laterally across the gel surface, with point-by-point tracking giving a displacement field. The 2017 GelSight paper by Yuan, Dong, and Adelson summarized how to read it: under normal pressure the markers spread outward from the contact center; under shear force they shift together in the shear direction; a rotational pattern indicates an in-plane torque; and displacement magnitude is roughly proportional to force. When an object is about to slip, points near the edge of the contact region start moving first while the center stays still — this non-uniformity can be used to detect slip before it fully happens. Looking only at the gel’s surface bumps (photometric stereo) gives shape alone; markers add tangential-force information. Sensors like TacTip instead use the displacement of internal pins as their main signal.","example":"Gripping a cup with a gripper fitted with GelSight Mini sensors, if the outer-ring markers start moving in the same direction while the center markers stay roughly still, that signals the cup is about to slip, so the gripper can increase its grip force.","related":["Vision-Based Tactile Sensor","GelSight","Slip Detection","Photometric Stereo","Normal Force and Tangential (Shear) Force","Tactile Image"]},{"id":"slip-detection","category":"perception","sec":3,"tier":3,"sources":[{"title":"Slip Detection with Combined Tactile and Visual Information (arXiv)","url":"https://arxiv.org/abs/1802.10153"}],"as_of":"","related_ids":["tactile-sensor","vision-based-tactile-sensor","slip","friction-cone",null,"contact-detection"],"name":"Slip Detection","alt":"滑移检测","abbr":"","aliases":["Slippage Detection"],"one_liner":"Determining whether an object held in a hand is starting to slide, so the grip can be adjusted in time.","explanation":"Slip detection means judging, in real time during grasping and manipulation, whether an object is beginning to slide relative to the fingers — including incipient slip, the very early stage before full slippage. Common signals include a sudden change or vibration in shear (tangential) force on a tactile sensor, changes in the displacement field of surface markers on a vision-based tactile sensor such as GelSight, and the ratio of shear force to normal force on a force sensor approaching the friction coefficient. It addresses the problem that gripping too loosely drops the object while gripping too tightly can crush it: once slip is detected, the controller can increase grip force or adjust its pose. This capability matters for fragile, deformable, or unknown-weight objects, and for in-hand manipulation, and is one of the most direct uses of touch sensing in robotics.","example":"Watching marker displacement on a GelSight contact surface, the system flags the start of slipping once the markers’ overall displacement becomes uneven, and the gripper immediately increases its force.","related":["Tactile Sensor","Vision-Based Tactile Sensor","Slip","Friction Cone","Marker Tracking (Tactile Markers)","Contact Detection"]},{"id":"extrinsic-contact-sensing","category":"perception","sec":3,"tier":3,"sources":[{"title":"Extrinsic Contact Sensing with Relative-Motion Tracking from Distributed Tactile Measurements (ICRA 2021)","url":"https://arxiv.org/abs/2103.08108"},{"title":"Perceiving Extrinsic Contacts from Touch Improves Learning Insertion Policies","url":"https://arxiv.org/abs/2309.16652"}],"as_of":"","related_ids":["contact-detection","tactile-sensor","peg-in-hole-insertion","contact-rich-manipulation","tool-use","extrinsic-dexterity"],"name":"Extrinsic Contact Sensing","alt":"外部接触感知","abbr":"","aliases":["Extrinsic Contact Estimation","Tool-Environment Contact Sensing"],"one_liner":"The perception task of inferring where and how an object held in the hand is making contact with the outside environment.","explanation":"When a robot holds an object to do work, there are two kinds of contact: contact between the fingers and the object is called intrinsic contact, while contact between the held object and the environment — a tabletop, the wall of a hole — is called extrinsic contact. Extrinsic contact happens outside the hand, is often occluded, and can’t be felt by the fingers directly, so it has to be inferred indirectly. In an ICRA 2021 paper, MIT’s Ma, Dong, and Rodriguez used a distributed tactile sensor to track the object’s tiny relative motion inside the hand, and combined this with rigid-body constraints like non-penetration and friction to estimate the location of a point or line contact, arguing that distributed tactile sensing suits this better than looking only at the net force from a six-axis force sensor. Later work also infers contact location from touch, vision, or even sound using neural networks. This is crucial for contact-rich tasks like peg-in-hole assembly and tool use: knowing where contact happened is what tells the robot which way to adjust.","example":"Higuera and colleagues’ NCF-v2 infers the contact between a held object and the environment from gripper tactile sensing; plugged into an insertion policy, it raised the success rate of placing a cup into a saucer by 33% and made execution 1.36 times faster, and raised the success rate of placing a bowl on a rack by 13%.","related":["Contact Detection","Tactile Sensor","Peg-in-Hole Insertion","Contact-rich Manipulation","Tool Use","Extrinsic Dexterity"]},{"id":"tactile-image","category":"perception","sec":3,"tier":3,"sources":[{"title":"GelSight: High-Resolution Robot Tactile Sensors for Estimating Geometry and Force (Yuan et al., Sensors 2017)","url":"https://www.mdpi.com/1424-8220/17/12/2762"}],"as_of":"","related_ids":["tactile-array","vision-based-tactile-sensor","gelsight","visuo-tactile-fusion","tactile-representation-learning","photometric-stereo"],"name":"Tactile Image","alt":"触觉图像","abbr":"","aliases":["Pressure Distribution Map"],"one_liner":"Arranging tactile sensor readings into a 2D image, so ordinary vision models can process them directly.","explanation":"A tactile image organizes a tactile signal into 2D image form. It comes from two sources: a tactile array, where each element’s reading becomes one pixel, giving a low-resolution pressure-distribution map; or a vision-based tactile sensor such as GelSight or DIGIT, whose internal camera directly photographs the soft gel deforming under pressure, which is already a color image that clearly shows the texture and shape of whatever it touched. The benefit of framing it as an image is that off-the-shelf vision models — convolutional networks, ViTs — can be applied directly, and it can be fed into a policy alongside camera images for combined vision-touch fusion. A sequence of tactile images sampled over time can also capture how a contact evolves, useful for detecting slip or estimating shear force.","example":"Pressing a GelSight onto a coin produces a tactile image that clearly shows the coin’s embossed relief, from which photometric stereo recovers the local 3D shape.","related":["Tactile Array","Vision-Based Tactile Sensor","GelSight","Visuo-Tactile Fusion","Tactile Representation Learning","Photometric Stereo"]},{"id":"sparsh","category":"perception","sec":3,"tier":3,"sources":[{"title":"Sparsh: Self-supervised touch representations for vision-based tactile sensing (arXiv)","url":"https://arxiv.org/abs/2410.24090"},{"title":"facebookresearch/sparsh (GitHub)","url":"https://github.com/facebookresearch/sparsh"}],"as_of":"2024-10","related_ids":["tactile-representation-learning","vision-based-tactile-sensor","digit","self-supervised-learning","joint-embedding-predictive-architecture","anytouch"],"name":"Sparsh","alt":"Sparsh","abbr":"","aliases":["Self-Supervised Touch Representations for Vision-Based Tactile Sensing"],"one_liner":"A general-purpose vision-based-touch representation model from Meta, pretrained with self-supervision on a large set of tactile images.","explanation":"Sparsh is a family of tactile encoders from Meta FAIR and collaborators, published at CoRL 2024, built for vision-based tactile sensors — sensors that perceive contact by having a camera image the deformation of a soft gel surface. Previously, each sensor type and each task needed its own labeled training data; Sparsh instead uses self-supervised learning, pretraining on more than 460,000 unlabeled tactile images with methods such as MAE, DINO, and JEPA, to produce a representation that generalizes across sensors including DIGIT, GelSight 2017, and GelSight Mini. The authors also released the TacBench benchmark, covering six categories of tasks from recognizing tactile properties to force estimation, slip detection, and manipulation planning. The paper reports that self-supervised pretraining outperforms task- and sensor-specific end-to-end training by 95.1% on average on TacBench. Code and weights are open-sourced.","example":"A Sparsh encoder attached behind a DIGIT sensor lets a force-estimation or slip-detection head be trained with only a small amount of labeled data.","related":["Tactile Representation Learning","Vision-Based Tactile Sensor","DIGIT","Self-Supervised Learning","Joint-Embedding Predictive Architecture","AnyTouch"]},{"id":"anytouch","category":"perception","sec":3,"tier":3,"sources":[{"title":"AnyTouch (ICLR 2025, arXiv:2502.12191)","url":"https://arxiv.org/abs/2502.12191"},{"title":"AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception (arXiv:2602.09617)","url":"https://arxiv.org/abs/2602.09617"}],"as_of":"2026-02","related_ids":["vision-based-tactile-sensor","tactile-representation-learning","gelsight-mini","digit","sparsh","visuo-tactile-fusion"],"name":"AnyTouch","alt":"AnyTouch","abbr":"","aliases":["AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors","AnyTouch 2"],"one_liner":"A unified tactile-representation model spanning multiple vision-based tactile sensors, from Renmin University’s Di Hu group, covering both static and dynamic touch.","explanation":"AnyTouch is a visuo-tactile representation learning framework proposed by Di Hu’s group at Renmin University of China together with Bin Fang at Beijing University of Posts and Telecommunications and others, published at ICLR 2025. Vision-based tactile sensors use a camera to image the deformation of an elastic gel layer, but different sensor models image very differently, making data and models hard to share across them. The authors first collected the TacQuad dataset: four sensors — GelSight Mini, DIGIT, a custom DuraGel, and Tac3D — touching the same location on the same object, yielding 72,606 aligned frames of contact data, paired with visual images and text descriptions of tactile properties. The model takes both tactile images and tactile video as input, uses masked modeling to learn pixel-level detail, and then learns sensor-agnostic semantic features by aligning with vision and language and by matching across sensors. In February 2026 the team released AnyTouch 2, shifting the focus to dynamic tactile sensing with force information.","example":"In the paper’s real-robot experiment, a robot arm pours 60 grams of small beads out of a cylinder using only tactile feedback, and the error between the poured mass and the target mass is used to evaluate different tactile representations.","related":["Vision-Based Tactile Sensor","Tactile Representation Learning","GelSight Mini","DIGIT","Sparsh","Visuo-Tactile Fusion"]},{"id":"contact-microphone","category":"perception","sec":3,"tier":3,"sources":[{"title":"Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation (arXiv 2405.08576)","url":"https://arxiv.org/abs/2405.08576"},{"title":"ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data (arXiv 2406.19464)","url":"https://arxiv.org/abs/2406.19464"}],"as_of":"","related_ids":["tactile-sensor","multimodal-perception","universal-manipulation-interface","contact-rich-manipulation","visuo-tactile-fusion","microphone-array"],"name":"Contact Microphone","alt":"接触式麦克风（音频触觉）","abbr":"","aliases":["Audio as Tactile Signal","Piezo Contact Microphone"],"one_liner":"A microphone stuck to a gripper or object that picks up contact vibration, used as a cheap tactile sensor.","explanation":"A contact microphone doesn’t pick up sound through the air — it’s attached to a solid surface and picks up vibration directly, most commonly using a piezoelectric element. In robotics it’s used as a cheap, durable substitute for touch: the high-frequency vibrations from scraping, collision, friction, or shaking can reveal whether contact has occurred, what material a surface is, or whether something is inside a container — things a camera often can’t see. Hearing Touch (ICRA 2024) uses a contact microphone as a stand-in for touch and improves a manipulation policy with representations learned from large-scale audio-visual pretraining; ManiWAV (CoRL 2024) embeds a piezoelectric contact microphone in one finger of a UMI handheld gripper, so collected human demonstrations carry both video and sound. Its limitation is that it can’t help when the interaction itself makes no sound.","example":"In ManiWAV’s dice-pouring task, the robot first shakes a cup and listens through the contact microphone for the vibration of dice hitting the cup, judging whether the cup has been fully emptied.","related":["Tactile Sensor","Multimodal Perception","Universal Manipulation Interface","Contact-rich Manipulation","Visuo-Tactile Fusion","Microphone Array"]},{"id":"visuo-tactile-fusion","category":"perception","sec":3,"tier":3,"sources":[{"title":"Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks (arXiv)","url":"https://arxiv.org/abs/1810.10191"},{"title":"3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv)","url":"https://arxiv.org/abs/2410.24091"}],"as_of":"","related_ids":["multimodal-fusion","tactile-sensor","vision-based-tactile-sensor","3d-vitac","contact-rich-manipulation","vision-tactile-language-action-model"],"name":"Visuo-Tactile Fusion","alt":"视触觉融合","abbr":"","aliases":["Visual-Tactile Fusion"],"one_liner":"Combining what a camera sees with what a tactile sensor feels into a single, unified signal.","explanation":"Visuo-tactile fusion means using both vision and touch signals together in perception or a policy, and combining them into a unified representation. Vision is good at seeing the big picture and finding a target, but a finger is often occluded the moment it touches an object, and vision can’t tell how much force is being applied or whether something is slipping; tactile sensing fills in exactly that: contact location, pressure distribution, and slip. Common approaches include encoding the two modalities separately and concatenating their features, using attention for cross-modal fusion, or projecting tactile readings into a 3D point cloud and merging it with the visual point cloud. A Stanford team used self-supervised learning to fuse vision and force for peg-in-hole insertion at ICRA 2019; 3D-ViTac (CoRL 2024) merges a flexible tactile array into a point cloud alongside a diffusion policy to handle fragile objects, clearly outperforming vision alone. Note this is distinct from a “vision-based tactile sensor,” which is a type of tactile sensor that uses a camera to photograph an elastic membrane.","example":"While a robotic hand grasps an egg, the camera finds the egg and guides the hand toward it; once contact is made, the tactile array reports the pressure distribution, and the policy keeps its grip force just below the point of crushing it.","related":["Multimodal Fusion","Tactile Sensor","Vision-Based Tactile Sensor","3D-ViTac","Contact-rich Manipulation","Vision-Tactile-Language-Action Model"]},{"id":"camera-calibration","category":"perception","sec":4,"tier":2,"sources":[{"title":"Zhang, A Flexible New Technique for Camera Calibration (IEEE TPAMI 2000)","url":"https://www.microsoft.com/en-us/research/publication/a-flexible-new-technique-for-camera-calibration/"},{"title":"MATLAB: What Is Camera Calibration?","url":"https://www.mathworks.com/help/vision/ug/camera-calibration.html"}],"as_of":"","related_ids":["camera-intrinsics","camera-extrinsics","lens-distortion","calibration-board","hand-eye-calibration","pinhole-camera-model"],"name":"Camera Calibration","alt":"相机标定","abbr":"","aliases":["Intrinsic Calibration","Zhang's Calibration Method"],"one_liner":"Estimating a camera's intrinsics, distortion coefficients, and extrinsics so pixels and 3D coordinates can convert into each other.","explanation":"Camera calibration is the process of estimating a camera's imaging parameters: its intrinsics (focal length, and the principal point where the optical axis meets the image), its lens distortion coefficients (radial and tangential), and its extrinsics (rotation and translation relative to some reference frame). The most widely used method is Zhang Zhengyou's, published in IEEE TPAMI in 2000: it shoots a planar calibration board from at least two different poses, computes a closed-form solution first, and then refines it with maximum-likelihood optimization. Calibration quality is judged by reprojection error — reprojecting 3D points back onto the image using the estimated parameters and comparing the pixel distance to the points actually detected. Without accurate intrinsics and extrinsics, a robot can't turn a depth map into a point cloud or map vision results into arm coordinates; depth cameras usually ship with factory-calibrated intrinsics, but the relationship between the camera and the arm still has to be calibrated separately.","example":"Processing twenty-some checkerboard photos with OpenCV's calibrateCamera gives fx, fy (focal length in pixels), the principal point cx, cy, and a set of distortion coefficients, along with an average reprojection error — the smaller that number, the better the calibration.","related":["Camera Intrinsics","Camera Extrinsics","Lens Distortion","Calibration Board","Hand-Eye Calibration","Pinhole Camera Model"]},{"id":"calibration-board","category":"perception","sec":4,"tier":2,"sources":[{"title":"OpenCV Tutorial: Detection of ChArUco Boards","url":"https://raw.githubusercontent.com/opencv/opencv/4.x/doc/tutorials/objdetect/charuco_detection/charuco_detection.markdown"},{"title":"MATLAB: What Is Camera Calibration?","url":"https://www.mathworks.com/help/vision/ug/camera-calibration.html"}],"as_of":"","related_ids":["camera-calibration","hand-eye-calibration","aruco-marker","apriltag","camera-intrinsics","lens-distortion"],"name":"Calibration Board","alt":"标定板","abbr":"","aliases":["Checkerboard","ChArUco Board","Calibration Target","Dot-Array Calibration Board"],"one_liner":"A flat board printed with a pattern of precisely known size, used to give a camera calibration accurate reference points.","explanation":"A calibration board is a flat target with a printed pattern of precisely known dimensions, used for camera calibration and hand-eye calibration. The most common type is the black-and-white checkerboard: it detects the internal corners where squares meet, achieving sub-pixel precision, but usually requires the whole board to be visible in frame. A dot-array board uses circle centers as its features instead. A ChArUco board embeds ArUco markers into the checkerboard's white squares, using the markers to identify each corner individually, so it still works when partially occluded or only partly in view, while keeping the checkerboard's corner-detection precision; OpenCV has built-in support for it. In practice, using one well requires an accurately printed size, a flat board (often mounted on aluminum or glass), and shots taken from multiple angles and distances so the detected corners cover the whole frame.","example":"Generate a ChArUco board with OpenCV, print it, and mount it on a flat aluminum plate, measuring the square edge length precisely. Hand-hold the camera and shoot from a dozen or more different angles, then feed the images to the calibration function to solve for intrinsics and distortion.","related":["Camera Calibration","Hand-Eye Calibration","ArUco Marker","AprilTag","Camera Intrinsics","Lens Distortion"]},{"id":"reprojection-error","category":"perception","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Reprojection error","url":"https://en.wikipedia.org/wiki/Reprojection_error"},{"title":"MathWorks: Evaluating the Accuracy of Single Camera Calibration","url":"https://www.mathworks.com/help/vision/ug/evaluating-the-accuracy-of-single-camera-calibration.html"},{"title":"Wikipedia: Bundle adjustment","url":"https://en.wikipedia.org/wiki/Bundle_adjustment"}],"as_of":"","related_ids":["camera-calibration","bundle-adjustment","perspective-n-point","triangulation","pinhole-camera-model","hand-eye-calibration"],"name":"Reprojection Error","alt":"重投影误差","abbr":"","aliases":["Re-Projection Error"],"one_liner":"The pixel distance between a 3D point reprojected onto the image and where it was actually observed.","explanation":"Reprojection error measures how accurate a set of estimated camera parameters, camera pose, and 3D points really are: take an estimated 3D point, reproject it onto the image using the estimated intrinsics and extrinsics, and compare that to the pixel actually observed for the corresponding point — the distance between them, in pixels, is the reprojection error. It's the most common optimization objective and quality indicator in geometric vision: in camera calibration, a smaller average reprojection error on the calibration board's corners is better (a MATLAB documentation example gives 0.19 pixels), and if it's noticeably large, the worst images can be removed and the camera recalibrated; bundle adjustment jointly adjusts all camera poses and 3D points to minimize the total reprojection error; and PnP pose solving, triangulation, and SLAM back-ends are all optimized around it.","example":"Calibrating a wrist camera with a checkerboard, OpenCV's calibrateCamera returns an overall RMS reprojection error; if a few images show noticeably higher error (usually from misdetected corners), removing them and recalibrating is standard practice.","related":["Camera Calibration","Bundle Adjustment","Perspective-n-Point","Triangulation","Pinhole Camera Model","Hand-Eye Calibration"]},{"id":"apriltag","category":"perception","sec":4,"tier":2,"sources":[{"title":"AprilTag（University of Michigan APRIL Lab）","url":"https://april.eecs.umich.edu/software/apriltag"},{"title":"AprilRobotics/apriltag（AprilTag 3）","url":"https://github.com/AprilRobotics/apriltag"}],"as_of":"","related_ids":["aruco-marker","hand-eye-calibration","perspective-n-point","camera-extrinsics","6d-object-pose-estimation","calibration-board"],"name":"AprilTag","alt":"AprilTag","abbr":"","aliases":["Visual Fiducial Marker","Fiducial Marker","AprilTag 3"],"one_liner":"A printable black-and-white square marker that a camera can detect to compute the marker's ID and 6D pose.","explanation":"AprilTag is a visual fiducial marker system proposed by Edwin Olson's team (the APRIL lab) at the University of Michigan, presented at ICRA 2011, and it looks like a simplified QR code. It only encodes a single ID number rather than large amounts of data, but the detector reliably finds its four corners, and combining those with the camera's intrinsics and the tag's known physical size, PnP gives its 6D pose relative to the camera. In robotics it's commonly used for hand-eye calibration, providing ground-truth poses for objects or worktables, and aligning the extrinsics of multiple cameras. The current open-source version, AprilTag 3 (a C implementation under the BSD license), offers several tag families such as tag36h11 and tagStandard41h12, and can also detect ArUco markers.","example":"Pasting tag36h11 markers at the four corners of a worktable lets a fixed third-person camera detect them and solve for the table's pose relative to the camera, which is then used to bring several cameras into one shared coordinate frame.","related":["ArUco Marker","Hand-Eye Calibration","Perspective-n-Point","Camera Extrinsics","6D Object Pose Estimation","Calibration Board"]},{"id":"aruco-marker","category":"perception","sec":4,"tier":3,"sources":[{"title":"OpenCV 教程：Detection of ArUco Markers","url":"https://github.com/opencv/opencv/blob/4.x/doc/tutorials/objdetect/aruco_detection/aruco_detection.markdown"},{"title":"ArUco 项目页（University of Córdoba, AVA group）","url":"https://www.uco.es/investiga/grupos/ava/portfolio/aruco/"}],"as_of":"","related_ids":["apriltag","calibration-board","hand-eye-calibration","perspective-n-point","camera-calibration","opencv"],"name":"ArUco Marker","alt":"ArUco码","abbr":"","aliases":["ArUco","Fiducial Marker"],"one_liner":"A square tag with a black border around a black-and-white grid code; a camera that spots it can compute its own pose relative to the tag.","explanation":"ArUco is a square fiducial marker system and open-source detection library developed by the AVA research group (Rafael Muñoz, Sergio Garrido, and others) at the University of Córdoba in Spain. Each marker has a thick black border around a binary matrix of black-and-white cells encoding an ID: the black border makes it fast to find in an image, and the binary code allows error detection and correction, as well as telling how far the marker has rotated. A usable set of codes is called a dictionary — a 4×4 marker, for instance, encodes 16 bits. The four corners of a single marker are enough to solve for the camera’s 6D pose relative to the marker using PnP. OpenCV’s aruco module is built on this library, and from version 4.7 it moved into the objdetect module. In robotics, ArUco markers are commonly used for hand-eye calibration, ChArUco calibration boards (ArUco markers embedded in a checkerboard), or stuck on objects and worktables to quickly get ground-truth pose. It belongs to the same family of techniques as AprilTag.","example":"When collecting data, sticking an ArUco marker to a table corner lets the wrist camera detect it every frame and convert the camera pose into the table’s coordinate frame, which also serves as a quick check on how accurate the hand-eye calibration is.","related":["AprilTag","Calibration Board","Hand-Eye Calibration","Perspective-n-Point","Camera Calibration","OpenCV (Open Source Computer Vision Library)"]},{"id":"perspective-n-point","category":"perception","sec":4,"tier":3,"sources":[{"title":"Perspective-n-Point - Wikipedia","url":"https://en.wikipedia.org/wiki/Perspective-n-Point"}],"as_of":"","related_ids":["camera-intrinsics","6d-object-pose-estimation","random-sample-consensus","aruco-marker","hand-eye-calibration","opencv"],"name":"Perspective-n-Point","alt":"PnP（透视n点）","abbr":"PnP","aliases":["PnP","PnP Problem","solvePnP"],"one_liner":"Recovering a camera’s pose from a set of known 3D points and where each one lands in the image.","explanation":"PnP is a classic computer vision problem: given the camera’s intrinsic parameters, n 3D points, and their 2D projections in the image, find the camera’s rotation and translation relative to those points — 6 degrees of freedom in total. At minimum, 3 point pairs are needed (P3P), but 3 points can yield up to 4 solutions, so a 4th point is usually needed to disambiguate. EPnP, proposed by Lepetit and colleagues in 2009, solves the problem using 4 virtual control points, with computation that scales linearly in the number of points. Real matches often include errors, so PnP is typically paired with Random Sample Consensus (RANSAC) to reject outliers; OpenCV’s solvePnP and solvePnPRansac are the most commonly used implementations. In robotics, it is often used to recover a pose from the corners of a calibration board or an ArUco marker.","example":"A camera sees an ArUco marker of known side length; feeding the 3D coordinates of its 4 corners and their detected pixel positions into solvePnP gives the marker’s position and orientation relative to the camera.","related":["Camera Intrinsics","6D Object Pose Estimation","Random Sample Consensus","ArUco Marker","Hand-Eye Calibration","OpenCV (Open Source Computer Vision Library)"]},{"id":"hand-eye-calibration","category":"perception","sec":4,"tier":1,"sources":[{"title":"OpenCV calib3d.hpp：calibrateHandEye 与 AX=XB 说明","url":"https://raw.githubusercontent.com/opencv/opencv/4.x/modules/calib3d/include/opencv2/calib3d.hpp"},{"title":"Wikipedia: Hand eye calibration problem","url":"https://en.wikipedia.org/wiki/Hand_eye_calibration_problem"}],"as_of":"","related_ids":["eye-in-hand","eye-to-hand","camera-extrinsics","calibration-board","homogeneous-transformation-matrix","wrist-camera"],"name":"Hand-Eye Calibration","alt":"手眼标定","abbr":"","aliases":["AX=XB","Camera-to-Arm Calibration"],"one_liner":"Finding the fixed coordinate transform that relates a camera to a robot arm.","explanation":"Hand-eye calibration solves for the fixed, unchanging transform between the ‘eye’ (camera) and the ‘hand’ (robot arm). There are two configurations: eye-in-hand, where the camera is mounted on the end-effector and the goal is the camera's pose relative to the end-effector flange; and eye-to-hand, where the camera is fixed nearby and the goal is the camera's pose relative to the robot's base. The method moves the arm through several poses, recording the end-effector's pose and the calibration board's pose as seen by the camera at each one; the problem is written as AX = XB, where X is the transform being solved for. Classic solutions such as Tsai-Lenz (1989) are already implemented in OpenCV's calibrateHandEye. At least two motions with non-parallel rotation axes are required — meaning at least 3 poses — though more are collected in practice. An inaccurate calibration means objects the camera sees get mapped to the wrong arm coordinates, and grasping simply fails.","example":"To calibrate a wrist camera, the arm carries it through a dozen or more angles photographing a fixed ChArUco calibration board. Forward kinematics gives the end-effector's pose at each shot, PnP gives the board's pose relative to the camera, and OpenCV's calibrateHandEye then solves for the transform from the camera to the flange.","related":["Eye-in-Hand","Eye-to-Hand","Camera Extrinsics","Calibration Board","Homogeneous Transformation Matrix","Wrist Camera"]},{"id":"eye-in-hand","category":"perception","sec":4,"tier":2,"sources":[{"title":"MoveIt: Hand-Eye Calibration Tutorial","url":"https://moveit.picknik.ai/main/doc/examples/hand_eye_calibration/hand_eye_calibration_tutorial.html"},{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"}],"as_of":"","related_ids":["eye-to-hand","hand-eye-calibration","wrist-camera","visual-servoing","third-person-camera","forward-kinematics"],"name":"Eye-in-Hand","alt":"眼在手上","abbr":"","aliases":["Camera Mounted on End-Effector"],"one_liner":"A camera-mounting scheme where the camera is fixed to the robot arm's end-effector and moves along with it.","explanation":"Eye-in-hand is a camera-mounting configuration in a hand-eye system: the camera is fixed to the arm's end-effector or gripper and moves along with it; the opposite configuration, with the camera fixed in the environment, is called ‘eye-to-hand.’ Its advantage is a close-up view of the target that's less likely to be blocked by the arm itself, which suits fine alignment and visual servoing. Its drawback is a narrow field of view with no global picture, and every observation has to be converted into the base frame by multiplying the end-effector pose (from forward kinematics) by a one-time-calibrated camera-to-end-effector transform. That transform comes from hand-eye calibration, classically formulated as AX = XB. The wrist camera used in learned policies is this configuration, and its images are often fed in alongside third-person cameras — the DROID platform, for example, mounts a ZED Mini stereo camera on each Franka arm's wrist.","example":"During a grasp, a fixed third-person camera first roughly localizes the object; once the arm gets close, the system switches to the wrist camera's close-up view to fine-tune the gripper's position before closing it.","related":["Eye-to-Hand","Hand-Eye Calibration","Wrist Camera","Visual Servoing","Third-Person Camera","Forward Kinematics (FK)"]},{"id":"eye-to-hand","category":"perception","sec":4,"tier":2,"sources":[{"title":"easy_handeye: eye-in-hand and eye-on-base calibration (GitHub)","url":"https://github.com/IFL-CAMP/easy_handeye"},{"title":"MoveIt 2 Hand-Eye Calibration Tutorial","url":"https://moveit.picknik.ai/main/doc/examples/hand_eye_calibration/hand_eye_calibration_tutorial.html"}],"as_of":"","related_ids":["hand-eye-calibration","eye-in-hand","camera-extrinsics","calibration-board","third-person-camera","aruco-marker"],"name":"Eye-to-Hand","alt":"眼在手外","abbr":"","aliases":["Fixed Camera Calibration","Eye-on-Base","Eye-to-Hand Calibration"],"one_liner":"A camera-mounting scheme where the camera is fixed outside the robot and doesn't move with the arm, plus its matching calibration.","explanation":"Eye-to-hand is one of the two hand-eye calibration configurations: the camera is fixed — on a tripod, table edge, or stand — and doesn't move with the arm, as opposed to ‘eye-in-hand,’ where the camera is mounted on the end-effector. Eye-to-hand requires the fixed transform from the camera to the robot's base (the camera extrinsics); once known, it lets objects seen by the camera be mapped into coordinates the arm can act on. Calibration works by fixing a calibration board or an ArUco marker to the end-effector, moving the arm through several poses while the camera photographs each one, and solving an equation of the form AX = XB, commonly with the Tsai-Lenz algorithm; MoveIt and easy_handeye both offer ready-made tools for this. Its advantage is a stable view that can see the whole worktable; its drawback is that it's easily blocked by the arm itself, and bumping the camera means recalibrating.","example":"A fixed depth camera on a tripod at the table's edge looks down at the workspace. A ChArUco board is clamped to the arm's end-effector and moved through a dozen or more poses, each photographed once, to solve for the camera's pose relative to the base. Afterward, cup coordinates the camera detects can be sent directly to the arm to grasp.","related":["Hand-Eye Calibration","Eye-in-Hand","Camera Extrinsics","Calibration Board","Third-Person Camera","ArUco Marker"]},{"id":"depth-to-color-alignment","category":"perception","sec":4,"tier":3,"sources":[{"title":"librealsense rs-align example (IntelRealSense GitHub)","url":"https://github.com/IntelRealSense/librealsense/tree/master/examples/align"},{"title":"realsense-ros README (align_depth.enable)","url":"https://github.com/IntelRealSense/realsense-ros"}],"as_of":"","related_ids":["depth-camera","camera-intrinsics","camera-extrinsics","projection-back-projection","point-cloud","realsense-depth-camera"],"name":"Depth-to-Color Alignment","alt":"深度与彩色对齐","abbr":"","aliases":["Depth Registration","RGB-D Alignment"],"one_liner":"Reprojects a depth map into the color camera’s viewpoint so the two images’ pixels line up one to one.","explanation":"In an RGB-D camera, the depth sensor and the color lens sit in different physical positions with different intrinsics and viewpoints, so the same pixel in each image doesn’t correspond to the same point if they’re simply overlaid. Alignment works by using the depth camera’s intrinsics to back-project each depth pixel into a 3D point, using the extrinsics between the two sensors to transform that point into the color camera’s coordinate frame, and then projecting it onto the color image plane using the color camera’s intrinsics. The resulting depth map matches the color image’s resolution and viewpoint, so a box detected or a mask segmented on the color image can look up depth directly and produce a colored point cloud. Intel RealSense SDK’s rs2::align and the ROS 2 driver’s align_depth.enable parameter do exactly this. The result is a computed approximation: changing viewpoint requires resampling, and occluded regions produce holes or misalignment, especially visible around object edges.","example":"Launching the RealSense driver in ROS 2 with align_depth.enable turned on publishes an extra topic, /camera/camera/aligned_depth_to_color/image_raw; a grasping program that boxes a cup in the color image can look up depth at the same pixel coordinates in this depth map, then back-project to get the cup’s 3D position.","related":["Depth Camera","Camera Intrinsics","Camera Extrinsics","Projection / Back-Projection","Point Cloud","RealSense Depth Camera (D435i / D405)"]},{"id":"camera-imu-calibration","category":"perception","sec":4,"tier":3,"sources":[{"title":"ethz-asl/kalibr（GitHub）","url":"https://github.com/ethz-asl/kalibr"},{"title":"Online Temporal Calibration for Monocular Visual-Inertial Systems (arXiv 1808.00692, IROS 2018)","url":"https://arxiv.org/abs/1808.00692"}],"as_of":"","related_ids":["visual-inertial-odometry","kalibr","inertial-measurement-unit","camera-extrinsics","multi-sensor-time-synchronization-timestamp-alignment","allan-variance"],"name":"Camera-IMU Calibration","alt":"相机-IMU联合标定","abbr":"","aliases":["Visual-Inertial Calibration","Camera-IMU Extrinsic and Time-Offset Calibration"],"one_liner":"Finds the relative pose and time offset between a camera and an IMU so their data can be aligned and fused.","explanation":"Visual-inertial odometry (VIO, which estimates its own motion using a camera plus an IMU) needs to process both sensors’ data together, which requires knowing their spatial extrinsics (the IMU’s rotation and translation relative to the camera) and their time offset (the fixed delay between the two devices’ timestamps caused by triggering and transmission); solving for these two things is camera-IMU calibration. The most widely used offline tool is ETH’s open-source Kalibr, based on work by Furgale and colleagues at IROS 2013: it represents the motion trajectory with a continuous-time B-spline, has the device waved thoroughly along every axis in front of an Aprilgrid calibration board, and jointly optimizes the extrinsics and time offset; the IMU’s noise parameters are usually measured beforehand with Allan variance and fed in as input. Systems like VINS-Mono can also estimate the time offset online while running (Qin and Shen, IROS 2018). When calibration is inaccurate, VIO drifts noticeably or can even diverge.","example":"To calibrate the RealSense D435i on a handheld data-collection rig: record a data bag of translating and rotating in front of an Aprilgrid, use Kalibr to solve for the camera-to-IMU transform matrix and the time offset (usually on the order of milliseconds), and write it into VINS-Fusion’s configuration file.","related":["Visual-Inertial Odometry","Kalibr","Inertial Measurement Unit","Camera Extrinsics","Multi-sensor Time Synchronization / Timestamp Alignment","Allan Variance"]},{"id":"camera-lidar-extrinsic-calibration","category":"perception","sec":4,"tier":3,"sources":[{"title":"What Is Lidar-Camera Calibration? - MATLAB & Simulink","url":"https://www.mathworks.com/help/lidar/ug/lidar-and-camera-calibration.html"},{"title":"koide3/direct_visual_lidar_calibration（GitHub）","url":"https://github.com/koide3/direct_visual_lidar_calibration"}],"as_of":"","related_ids":["camera-extrinsics","lidar","multi-sensor-fusion","calibration-board","camera-calibration","birds-eye-view"],"name":"Camera-LiDAR Extrinsic Calibration","alt":"相机-激光雷达联合标定","abbr":"","aliases":["LiDAR-Camera Calibration"],"one_liner":"Finds the rotation and translation between a lidar and a camera so a point cloud can be projected accurately onto the image.","explanation":"Robots and self-driving cars often carry both a lidar and a camera: the lidar gives accurate range, the camera gives color and texture. Fusing the two first requires knowing the rigid-body transform — three rotation plus three translation parameters — from the lidar’s coordinate frame to the camera’s, i.e., the extrinsics. Target-based methods use a checkerboard, ChArUco, or Aprilgrid: corners are detected in the image, a matching plane is fit in the point cloud to find corresponding corners, and the transform is then solved. Target-free methods register directly against the structure and texture of the environment — for example, the tool Koide and colleagues at Japan’s AIST open-sourced at ICRA 2023, which also supports non-repetitive-scan lidars like Livox. The result is used to color a point cloud, project it onto an image to assist detection, or feed BEV fusion perception; a common sanity check is whether the edges of the projected point cloud line up with object edges in the image.","example":"A quadruped robot carries a Livox Mid-360 and an RGB camera; after calibration, the point cloud is projected onto the image to check whether points along a step’s edge land on the step’s outline in the image, and the same extrinsics are then used to color the point cloud into a colored map.","related":["Camera Extrinsics","LiDAR","Multi-Sensor Fusion","Calibration Board","Camera Calibration","Bird’s-Eye View"]},{"id":"multi-sensor-time-synchronization","category":"perception","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Precision Time Protocol","url":"https://en.wikipedia.org/wiki/Precision_Time_Protocol"},{"title":"ROS 2 message_filters 文档（ExactTime / ApproximateTime 同步策略）","url":"https://raw.githubusercontent.com/ros2/message_filters/rolling/doc/index.rst"},{"title":"Orbbec Gemini 335L（硬件触发与多机统一硬件时间戳）","url":"https://www.orbbec.com/products/stereo-vision-camera/gemini-335l/"}],"as_of":"","related_ids":["multi-sensor-fusion","precision-time-protocol","camera-imu-calibration","observation-action-pair","visual-inertial-odometry","ros-bag"],"name":"Multi-Sensor Time Synchronization","alt":"多传感器时间同步","abbr":"","aliases":["Timestamp Alignment","Hardware Synchronization","Hard Sync / Soft Sync"],"one_liner":"Making sure data from different sensors used together actually comes from the same moment in time, not slightly different ones.","explanation":"A robot's camera, depth camera, lidar, IMU, and joint encoders each run on their own clock and sample rate. Time synchronization ensures that the different pieces of data used together in a computation actually come from the same instant. It works on two levels: hardware synchronization, where a trigger line makes multiple devices expose at the same moment, or PTP (the IEEE 1588 Precision Time Protocol, which can achieve sub-microsecond accuracy on a local network) lets devices share a single clock; and software alignment, where every piece of data gets a timestamp and is then paired or interpolated by timestamp — ROS's message_filters package, for instance, offers exact-match and approximate-match policies for this. Even a gap of a few tens of milliseconds can misalign a point cloud and an image when the robot is moving quickly, degrading the accuracy of visual-inertial odometry, and it can likewise misalign the observation-action pairs collected for training, making a trained policy's actions lag behind what it sees. It's a prerequisite for multi-sensor fusion, calibration, and data collection generally.","example":"During dual-arm teleoperation data collection, several cameras output frames at 30 Hz while joint states are reported at a higher frequency. When saving the data, each image's timestamp must be matched to the closest joint reading, or the action labels end up offset from the images.","related":["Multi-Sensor Fusion","Precision Time Protocol","Camera-IMU Calibration","Observation-Action Pair","Visual-Inertial Odometry","ROS Bag"]},{"id":"object-detection","category":"perception","sec":5,"tier":1,"sources":[{"title":"Wikipedia: Object detection","url":"https://en.wikipedia.org/wiki/Object_detection"},{"title":"You Only Look Once: Unified, Real-Time Object Detection (arXiv 1506.02640)","url":"https://arxiv.org/abs/1506.02640"},{"title":"Grounding DINO (arXiv 2303.05499)","url":"https://arxiv.org/abs/2303.05499"}],"as_of":"","related_ids":["bounding-box","open-vocabulary-object-detection","yolo","grounding-dino","intersection-over-union","instance-segmentation"],"name":"Object Detection","alt":"目标检测","abbr":"","aliases":["Object Recognition and Localization"],"one_liner":"Finding what objects appear in an image and where, marking each with a bounding box.","explanation":"Object detection is a foundational computer vision task: given an image, it outputs the category, bounding box (the rectangle enclosing the object), and confidence score for each object in it. An early representative was the Viola-Jones face detector, built on hand-crafted features; after deep learning took over, the R-CNN family of two-stage detectors emerged, followed by YOLO in 2015, which treats detection directly as a regression problem over boxes and class probabilities and can run in real time. Evaluation uses intersection over union (IoU, the overlap area between two boxes divided by their combined area) to judge whether a predicted box matches the ground truth, aggregated into mean average precision (mAP, the average of the average precision across all classes). Traditional detectors only recognize the classes they were trained on; open-vocabulary detectors such as Grounding DINO can find objects from an arbitrary text description instead. In robotics, detection is often the first step of a modular pipeline: find the box, then segment it, estimate its pose, and plan a grasp.","example":"A user says ‘hand me the red cup.’ The system feeds ‘red cup’ into Grounding DINO to get a bounding box, passes that to SAM to cut out a mask, combines the mask with the depth map to compute the cup's 3D position, and hands it off to the arm to grasp.","related":["Bounding Box","Open-Vocabulary Object Detection","YOLO","Grounding DINO","Intersection over Union","Instance Segmentation"]},{"id":"bounding-box","category":"perception","sec":5,"tier":2,"sources":[{"title":"Dive into Deep Learning: Object Detection and Bounding Boxes","url":"https://d2l.ai/chapter_computer-vision/bounding-box.html"}],"as_of":"","related_ids":["object-detection","intersection-over-union","non-maximum-suppression","3d-object-detection","grounding-dino","mask"],"name":"Bounding Box","alt":"检测框（边界框）","abbr":"BBox","aliases":["BBox","Detection Box"],"one_liner":"The rectangle an object detector draws around an object, given as a few coordinates marking its location in the image.","explanation":"A bounding box is the most common output of an object detector: a rectangle enclosing an object, usually paired with a class label and a confidence score. Two representations are common — top-left plus bottom-right corners (x1, y1, x2, y2), or center plus width and height (cx, cy, w, h) — and different datasets follow different conventions, so mixing them up is a common bug. How well a predicted box matches the ground truth is judged with intersection over union (IoU), and duplicate boxes are removed with non-maximum suppression. Extended forms include rotated boxes with an added angle and 3D bounding boxes. In embodied AI, a bounding box is often an intermediate result: an open-vocabulary detector finds a box from text first, and it's then handed to SAM for a mask or used to crop a region for grasping. This is a different concept from the ‘bounding box’ used for collision checking in physics engines.","example":"Given the instruction ‘pick up the red cup,’ Grounding DINO returns a box with a confidence around 0.6; downstream modules then only segment and search for grasp points within that box.","related":["Object Detection","Intersection over Union","Non-Maximum Suppression","3D Object Detection","Grounding DINO","Mask"]},{"id":"yolo","category":"perception","sec":5,"tier":2,"sources":[{"title":"You Only Look Once: Unified, Real-Time Object Detection (arXiv 1506.02640)","url":"https://arxiv.org/abs/1506.02640"},{"title":"Ultralytics Docs: Models Supported by Ultralytics","url":"https://docs.ultralytics.com/models/"}],"as_of":"2026-01","related_ids":["object-detection","bounding-box","non-maximum-suppression","yolo-world","mean-average-precision"],"name":"YOLO","alt":"YOLO","abbr":"YOLO","aliases":["You Only Look Once","YOLOv8","YOLO11","YOLO26"],"one_liner":"A family of real-time object detectors that locate every object in an image with a single pass through the network.","explanation":"YOLO was introduced by Joseph Redmon and colleagues in 2015 (published at CVPR 2016), framing object detection as a regression problem: the whole image passes through the network only once, producing the location and class of every detection box simultaneously; the original version ran at 45 frames per second, far faster than two-stage methods that first propose candidate boxes and then classify each one. YOLO has since grown into a large family released in relays by different teams; Ultralytics’ YOLOv5, YOLOv8, and YOLO11 are the most widely used. According to its documentation, YOLO26, released in January 2026, can optionally drop non-maximum suppression (NMS, the post-processing step that removes duplicate boxes) and supports segmentation, pose, and rotated-box tasks. Robots commonly use YOLO for fast 2D detection, then combine it with a depth map to get an object’s 3D position; for finding arbitrary categories by text, there’s the open-vocabulary version, YOLO-World.","example":"Desktop sorting: a wrist-camera image is fed into YOLO to get a bounding box for “cup”; the depth pixels inside the box are back-projected into 3D as the target position for the robot arm’s grasp.","related":["Object Detection","Bounding Box","Non-Maximum Suppression","YOLO-World","Mean Average Precision"]},{"id":"intersection-over-union","category":"perception","sec":5,"tier":2,"sources":[{"title":"Jaccard index - Wikipedia","url":"https://en.wikipedia.org/wiki/Jaccard_index"},{"title":"COCO Detection Evaluation","url":"https://raw.githubusercontent.com/cocodataset/cocodataset.github.io/master/dataset/detection-eval.htm"}],"as_of":"","related_ids":["object-detection","bounding-box","mean-average-precision","non-maximum-suppression","instance-segmentation","precision-recall"],"name":"Intersection over Union","alt":"交并比","abbr":"IoU","aliases":["IoU","Jaccard Index"],"one_liner":"The overlap area of two regions divided by their combined area, used to judge how accurate a predicted box or mask is.","explanation":"Intersection over union measures how much two regions overlap: intersection area divided by union area, with 1 meaning perfect overlap and 0 meaning no overlap at all — mathematically, this is the Jaccard index, proposed by Paul Jaccard in 1901. Vision tasks use it to judge whether a detected box or mask is accurate: a prediction counts as correct only if its IoU with the ground truth exceeds some threshold. PASCAL VOC uses a threshold of 0.5; COCO is stricter, computing average precision (AP) at thresholds from 0.50 to 0.95 in steps of 0.05 and then averaging them — the commonly cited AP50 and AP75 are the results at the 0.5 and 0.75 thresholds specifically. Segmentation tasks use the same formula, just applied to mask pixels instead of a box.","example":"A ground-truth box and a predicted box each have an area of 100 and overlap by 60. Their union is then 140, giving an IoU of about 0.43 — below 0.5, so by the VOC standard this detection wouldn't count as correct.","related":["Object Detection","Bounding Box","Mean Average Precision","Non-Maximum Suppression","Instance Segmentation","Precision / Recall"]},{"id":"non-maximum-suppression","category":"perception","sec":5,"tier":3,"sources":[{"title":"torchvision.ops.nms - PyTorch Documentation","url":"https://pytorch.org/vision/stable/generated/torchvision.ops.nms.html"}],"as_of":"","related_ids":["object-detection","bounding-box","intersection-over-union","yolo","detr","grasp-pose-detection"],"name":"Non-Maximum Suppression","alt":"非极大值抑制","abbr":"NMS","aliases":["NMS","NMS Post-Processing"],"one_liner":"A post-processing step that keeps only the highest-scoring box among a cluster of overlapping detection boxes.","explanation":"Object detectors often output several overlapping candidate boxes for the same object. Non-maximum suppression (NMS) sorts these by confidence score, keeps the highest-scoring box, and removes every other box whose IoU (intersection over union — the overlap area divided by the union area) with it exceeds a threshold; this repeats until no boxes remain unprocessed. It is the standard post-processing step in detectors such as YOLO and Faster R-CNN, and is also used in grasp-pose detection to remove duplicate grasp candidates. Setting the threshold too high leaves duplicate boxes behind; setting it too low can wrongly delete real objects that happen to sit close together. End-to-end detectors such as DETR are designed not to need NMS at all.","example":"A detector outputs five overlapping boxes for the same cup; NMS keeps the highest-scoring one and discards every other box whose IoU with it exceeds 0.5.","related":["Object Detection","Bounding Box","Intersection over Union","YOLO","DETR","Grasp Pose Detection"]},{"id":"detr","category":"perception","sec":5,"tier":3,"sources":[{"title":"arXiv 2005.12872: End-to-End Object Detection with Transformers","url":"https://arxiv.org/abs/2005.12872"},{"title":"ECCV 2020 paper page: End-to-End Object Detection with Transformers","url":"https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/832_ECCV_2020_paper.php"},{"title":"facebookresearch/detr (GitHub)","url":"https://github.com/facebookresearch/detr"}],"as_of":"","related_ids":["object-detection","transformer","bounding-box","non-maximum-suppression","grounding-dino","learnable-query"],"name":"DETR","alt":"DETR","abbr":"","aliases":["DEtection TRansformer","End-to-End Object Detection with Transformers"],"one_liner":"An end-to-end detector that frames object detection directly as “set prediction” using a Transformer.","explanation":"DETR is an object detection model proposed by Nicolas Carion and colleagues at Facebook AI (now Meta), published at ECCV 2020. Earlier detectors like Faster R-CNN first lay down a huge number of anchor boxes (preset candidate boxes) and finally use non-maximum suppression to delete duplicates. DETR first extracts image features with a CNN, feeds them into a Transformer encoder-decoder, and has the decoder use a fixed number of learnable “object queries” to each output one box and class; during training, bipartite matching pairs predictions with ground truth one to one, so anchor boxes and NMS are no longer needed. It matches a well-tuned Faster R-CNN’s accuracy on COCO, but converges more slowly during training. Later work — Deformable DETR, Grounding DINO, DINO-X — continued down this path; in embodied AI, ACT decodes actions with learnable queries too, with code adapted from DETR.","example":"The original DETR (ResNet-50 backbone), trained for 500 epochs on the COCO 2017 validation set, reaches 42.0 AP, on par with a same-backbone Faster R-CNN while using about half the computation.","related":["Object Detection","Transformer","Bounding Box","Non-Maximum Suppression","Grounding DINO","Learnable Query"]},{"id":"precision-recall","category":"perception","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Precision and recall","url":"https://en.wikipedia.org/wiki/Precision_and_recall"},{"title":"scikit-learn: Precision, recall and F-measures","url":"https://scikit-learn.org/stable/modules/model_evaluation.html"}],"as_of":"","related_ids":["intersection-over-union","mean-average-precision","object-detection","non-maximum-suppression","success-detector"],"name":"Precision / Recall","alt":"精确率 / 召回率","abbr":"","aliases":["F1 Score","F1"],"one_liner":"Precision measures how many reported results were correct; recall measures how many of the true targets were found.","explanation":"Precision and recall are a pair of basic metrics for detection, classification, and retrieval tasks. Results are split into true positives (TP, reported and correct), false positives (FP, reported but wrong — a false alarm), and false negatives (FN, should have been reported but wasn't — a miss). Precision is TP/(TP+FP), and recall is TP/(TP+FN). The two usually trade off against each other: lowering the confidence threshold finds more targets but also raises the false-alarm rate. F1 is their harmonic mean, 2PR/(P+R), capturing both in a single number. Object detection first uses IoU to decide whether a predicted box counts as a hit, and then computes these metrics; sweeping the threshold traces out a precision-recall curve, whose area underneath is average precision (AP), which averaged across classes gives mAP. Robots also use these metrics to evaluate grasp detectors, contact detectors, and success detectors.","example":"A cup detector reports 100 boxes on a test set, 80 of which correctly match real cups, out of 120 actual cups in the set: precision is 80%, recall is about 67%, and F1 is about 0.73.","related":["Intersection over Union","Mean Average Precision","Object Detection","Non-Maximum Suppression","Success Detector"]},{"id":"mean-average-precision","category":"perception","sec":5,"tier":3,"sources":[{"title":"COCO Detection Evaluation（官方评测说明）","url":"https://raw.githubusercontent.com/cocodataset/cocodataset.github.io/master/dataset/detection-eval.htm"},{"title":"Ultralytics: Performance Metrics Deep Dive","url":"https://docs.ultralytics.com/guides/yolo-performance-metrics/"}],"as_of":"","related_ids":["intersection-over-union","precision-recall","object-detection","coco-lvis","non-maximum-suppression","mask-r-cnn"],"name":"Mean Average Precision","alt":"平均精度均值","abbr":"mAP","aliases":["mAP","AP","mAP50","mAP50-95"],"one_liner":"The most common accuracy metric for object detection and segmentation: average precision (AP) per class, averaged across classes.","explanation":"Mean average precision (mAP) is the most common combined metric for object detection and instance segmentation. For a given class, all predictions are sorted by confidence, and a prediction only counts as correct if its intersection-over-union (IoU — the overlap area between predicted and ground-truth box, divided by their union area) with a ground-truth box exceeds a threshold; plotting this gives a precision-recall curve, and the area under that curve is the AP for that class, and averaging over all classes gives mAP. The number varies a lot depending on the threshold: the PASCAL VOC era commonly used IoU = 0.5 (mAP50), while COCO instead computes it at ten thresholds from 0.5 to 0.95 in steps of 0.05 and averages those (mAP50-95), which is a much stricter test of localization accuracy. Note that the “AP” reported in the COCO paper and its leaderboard has already been averaged over classes — the official documentation states explicitly that it makes no distinction between AP and mAP. When reading a robot perception paper, first check which threshold convention and which dataset the reported number uses.","example":"The Mask R-CNN paper reports a mask AP of 35.7 for the ResNet-101-FPN version on COCO — here, “AP” is the mAP averaged over 10 IoU thresholds and every class.","related":["Intersection over Union","Precision / Recall","Object Detection","COCO / LVIS","Non-Maximum Suppression","Mask R-CNN"]},{"id":"semantic-segmentation","category":"perception","sec":5,"tier":2,"sources":[{"title":"Fully Convolutional Networks for Semantic Segmentation (Long, Shelhamer, Darrell, CVPR 2015)","url":"https://arxiv.org/abs/1411.4038"},{"title":"Wikipedia: Image segmentation","url":"https://en.wikipedia.org/wiki/Image_segmentation"}],"as_of":"","related_ids":["instance-segmentation","panoptic-segmentation","open-vocabulary-segmentation","mask","semantic-map","segment-anything-model"],"name":"Semantic Segmentation","alt":"语义分割","abbr":"","aliases":["Semantic Seg","Pixel-Level Classification"],"one_liner":"Labels every pixel in an image with a category, like “table,” “cup,” or “floor.”","explanation":"Semantic segmentation is a core computer vision task: deciding which category every pixel in an image belongs to, and producing a category map the same size as the original image. It only cares about class, not individual identity — two cups both get labeled “cup” without being told apart; separating individual objects is instance segmentation, and combining the two is panoptic segmentation. In 2015, Long and colleagues introduced the fully convolutional network (FCN), which let a network predict pixel labels end to end, and this became the standard approach. Robots use semantic segmentation to find the region of a graspable object, identify walkable floor, or project labels onto a 3D point cloud to build a semantic map. Early models could only recognize the fixed set of categories they were trained on; today, open-vocabulary segmentation is common, letting a model segment any category described in text.","example":"Given a photo of a kitchen, the table pixels are labeled “table,” both cups are labeled “cup,” and everything else is “background”; a robot then pulls out the point cloud corresponding to the “cup” region to plan a grasp.","related":["Instance Segmentation","Panoptic Segmentation","Open-Vocabulary Segmentation","Mask","Semantic Map","Segment Anything Model"]},{"id":"instance-segmentation","category":"perception","sec":5,"tier":2,"sources":[{"title":"Mask R-CNN (arXiv:1703.06870)","url":"https://arxiv.org/abs/1703.06870"},{"title":"Panoptic Segmentation (arXiv:1801.00868)","url":"https://arxiv.org/abs/1801.00868"}],"as_of":"","related_ids":["semantic-segmentation","panoptic-segmentation","mask","object-detection","mask-r-cnn","segment-anything-model"],"name":"Instance Segmentation","alt":"实例分割","abbr":"","aliases":["Instance-Level Segmentation"],"one_liner":"Cutting out every object in an image pixel by pixel, keeping separate objects of the same class distinct from each other.","explanation":"Instance segmentation has to answer two questions at once — ‘what is this’ and ‘which one is this’ — generating a separate pixel-level mask and class label for every countable object (‘things,’ in the terminology of the field) in an image. This differs from semantic segmentation, which only assigns a class to each pixel: three cups on a table would merge into a single ‘cup’ region under semantic segmentation, whereas instance segmentation splits them into cup 1, cup 2, and cup 3. The representative method is Mask R-CNN, from Kaiming He and colleagues in 2017, which adds a parallel mask-prediction branch on top of an object-detection box. Panoptic segmentation, proposed in 2018, further folds in uncountable background regions (‘stuff’) such as sky and ground. Robot grasping usually uses instance segmentation first to cut out the target object, then crops the corresponding point cloud to compute a grasp pose.","example":"Two identical apples sit side by side on a table. Semantic segmentation only gives one merged ‘apple’ region, while instance segmentation outputs two separate masks, letting the robot grasp specifically the one on the left as instructed.","related":["Semantic Segmentation","Panoptic Segmentation","Mask","Object Detection","Mask R-CNN","Segment Anything Model"]},{"id":"panoptic-segmentation","category":"perception","sec":5,"tier":3,"sources":[{"title":"Panoptic Segmentation (arXiv 1801.00868)","url":"https://arxiv.org/abs/1801.00868"}],"as_of":"","related_ids":["semantic-segmentation","instance-segmentation","mask","scene-understanding","semantic-map","mask-r-cnn"],"name":"Panoptic Segmentation","alt":"全景分割","abbr":"","aliases":[],"one_liner":"A segmentation task that labels every pixel with a category while also telling individual object instances apart.","explanation":"Panoptic segmentation was proposed by Alexander Kirillov, Kaiming He, and colleagues in 2018 (published at CVPR 2019). It merges two separate tasks: semantic segmentation, which labels every pixel with a category but does not distinguish separate instances of the same category, and instance segmentation, which distinguishes individual object instances but ignores amorphous background categories such as sky or ground (which the paper calls “stuff”). Panoptic segmentation requires every pixel to get a category label, with an additional instance number for anything belonging to a countable object category (“things”), producing one complete, non-overlapping parse of the scene, evaluated with a newly proposed Panoptic Quality (PQ) metric. For robots, this simultaneously answers “where is the floor, where is the countertop” and “which cup is this, the first or the second,” and it is commonly used for scene understanding and semantic mapping.","example":"In a kitchen image, the floor, walls, and countertop are each labeled as one solid region by category, while three bowls on the counter are separately labeled bowl 1, bowl 2, and bowl 3.","related":["Semantic Segmentation","Instance Segmentation","Mask","Scene Understanding","Semantic Map","Mask R-CNN"]},{"id":"mask","category":"perception","sec":5,"tier":2,"sources":[{"title":"Segment Anything (arXiv:2304.02643)","url":"https://arxiv.org/abs/2304.02643"},{"title":"COCO Data Format（segmentation: polygon / RLE）","url":"https://raw.githubusercontent.com/cocodataset/cocodataset.github.io/master/dataset/format-data.htm"}],"as_of":"","related_ids":["instance-segmentation","semantic-segmentation","segment-anything-model","grounded-sam","attention-mask","point-cloud-segmentation"],"name":"Mask","alt":"掩码","abbr":"","aliases":["Segmentation Mask","Binary Mask"],"one_liner":"A pixel-by-pixel map, the same size as the image, marking which pixels belong to a given object.","explanation":"A mask is an image the same size as the original, where each pixel takes a value of 0 or 1 (or a class ID), marking which pixels belong to a given object. The output of a segmentation model is essentially a mask: semantic segmentation produces one per class, and instance segmentation produces one per object. A mask is more precise than a bounding box, which often mixes in background, since a mask hugs the object's actual outline. For storage, the COCO dataset uses polygon vertices for a single object and run-length encoding (RLE) compression for groups of objects. Meta's SAM was trained on more than a billion annotated masks across 11 million images, and can produce a mask for essentially any object from just a click or a box. Robots commonly use a mask to pull just the target object out of a depth map or point cloud before computing a grasp. Note that the ‘attention mask’ used in transformers is an unrelated concept — don't confuse the two.","example":"Given the instruction ‘grasp the red cup,’ Grounded-SAM first produces the cup's mask, overlays it on the aligned depth map, back-projects only the cup's pixels into a point cloud, and hands that to the grasp-detection network.","related":["Instance Segmentation","Semantic Segmentation","Segment Anything Model","Grounded SAM","Attention Mask","Point Cloud Segmentation"]},{"id":"mask-r-cnn","category":"perception","sec":5,"tier":3,"sources":[{"title":"arXiv: Mask R-CNN","url":"https://arxiv.org/abs/1703.06870"},{"title":"GitHub: facebookresearch/Detectron（注明 Marr Prize at ICCV 2017）","url":"https://github.com/facebookresearch/Detectron"}],"as_of":"","related_ids":["instance-segmentation","object-detection","mask","segment-anything-model","grounded-sam","mean-average-precision"],"name":"Mask R-CNN","alt":"Mask R-CNN","abbr":"","aliases":[],"one_liner":"An instance segmentation model from 2017, proposed by Kaiming He and colleagues, that outputs both detection boxes and a per-object mask.","explanation":"Mask R-CNN was proposed in 2017 by Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick at Facebook AI Research, winning the Marr Prize (best paper) at ICCV 2017. Building on the two-stage detector Faster R-CNN, it adds a branch to each candidate region that predicts a pixel-level mask in parallel, so one network does object detection and instance segmentation (separating out which pixels belong to each object) at the same time; an additional branch can also do human keypoint detection. The paper introduces RoIAlign, which replaces RoIPool’s rounding with bilinear interpolation, removing the misalignment between features and pixels, which noticeably improves mask accuracy. The ResNet-101-FPN version reaches a mask AP of 35.7 on COCO at about 5 frames per second. In robotics, it was long the standard baseline for segmenting objects in grasping and picking pipelines, though it is now often replaced by or paired with open-vocabulary methods like the SAM family and Grounded-SAM.","example":"In a bin-picking pipeline, a Mask R-CNN fine-tuned on photos of the company’s own parts first segments a mask for every part in the bin, and the point cloud corresponding to each mask is then passed to the grasp-pose-detection module.","related":["Instance Segmentation","Object Detection","Mask","Segment Anything Model","Grounded SAM","Mean Average Precision"]},{"id":"segment-anything-model","category":"perception","sec":5,"tier":2,"sources":[{"title":"Segment Anything (arXiv 2304.02643)","url":"https://arxiv.org/abs/2304.02643"},{"title":"GitHub: facebookresearch/segment-anything","url":"https://github.com/facebookresearch/segment-anything"},{"title":"Meta AI Blog: Segment Anything Model 3","url":"https://ai.meta.com/blog/segment-anything-model-3/"}],"as_of":"2025-11","related_ids":["sam-2","sam-3","grounded-sam","instance-segmentation","mask","foundation-model"],"name":"Segment Anything Model","alt":"分割一切模型","abbr":"SAM","aliases":["SAM","Segment Anything"],"one_liner":"Meta's 2023 promptable segmentation foundation model: give it a point or box and it returns that object's mask.","explanation":"The Segment Anything Model (SAM) is an image segmentation foundation model released by Meta AI in April 2023, introducing the task of ‘promptable segmentation’: a user gives a point, a box, or a rough mask as a prompt, and the model outputs a mask for the corresponding object — even for objects it has never seen — and can also automatically segment an entire image on its own. A large ViT image encoder runs once per image, and a lightweight decoder then answers multiple prompts quickly. Its training data, SA-1B, contains 11 million images and more than 1.1 billion masks, and the code and weights are open-sourced under Apache 2.0. The original SAM gives only a mask with no class label, so it's often chained with Grounding DINO into a pipeline called Grounded-SAM for text-based segmentation. It was later followed by the video-focused SAM 2 (2024) and SAM 3 (2025), which supports text-phrase prompts.","example":"Clicking on a screwdriver in a workbench image, SAM returns its pixel mask; back-projecting the depth pixels inside that mask gives a point cloud containing only the screwdriver, ready for grasp detection.","related":["SAM 2","SAM 3","Grounded SAM","Instance Segmentation","Mask","Foundation Model"]},{"id":"keypoint-detection","category":"perception","sec":5,"tier":2,"sources":[{"title":"COCO Data Format（keypoints 标注格式）","url":"https://raw.githubusercontent.com/cocodataset/cocodataset.github.io/master/dataset/format-data.htm"},{"title":"kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation (arXiv:1903.06684)","url":"https://arxiv.org/abs/1903.06684"}],"as_of":"","related_ids":["human-pose-estimation","semantic-keypoints","feature-points","6d-object-pose-estimation","rekep","tracking-any-point"],"name":"Keypoint Detection","alt":"关键点检测","abbr":"","aliases":["Keypoints","Keypoint Localization"],"one_liner":"Locating a handful of pre-defined, meaningful points in an image, like a wrist, a cup's handle, or a box's corner.","explanation":"Keypoint detection locates a small number of points with predefined meaning in an image, such as a wrist, a cup's handle, or a box's corner. It's more precise than a bounding box and lighter-weight than a segmentation mask; it also differs from feature points like SIFT or ORB, which only need to be locally recognizable and carry no fixed meaning. The most common case is human body keypoints: COCO labels each point as ‘not labeled,’ ‘occluded,’ or ‘visible,’ and scores predictions with OKS similarity. In robot manipulation, MIT's kPAM (2019) represents a whole category of objects with a few 3D semantic keypoints, so swapping in a differently shaped cup still works under the same rule — handling shape variation within a category better than estimating a single 6D pose would.","example":"To hang various mugs on a mug rack, the robot first detects each mug's base, rim, and handle as 3D keypoints, then plans motion so the handle lines up with the hook — the same rule works even though the mugs differ in size and shape.","related":["Human Pose Estimation","Semantic Keypoints","Feature Points","6D Object Pose Estimation","ReKep","Tracking Any Point"]},{"id":"object-tracking","category":"perception","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Video tracking","url":"https://en.wikipedia.org/wiki/Video_tracking"},{"title":"Simple Online and Realtime Tracking (SORT, ICIP 2016)","url":"https://arxiv.org/abs/1602.00763"}],"as_of":"","related_ids":["object-detection","kalman-filter","video-object-segmentation","tracking-any-point","embodied-visual-tracking","pose-tracking"],"name":"Object Tracking","alt":"目标跟踪","abbr":"","aliases":["Multi-Object Tracking","MOT","Single-Object Tracking","Visual Tracking"],"one_liner":"Continuously finding the same object across video frames and keeping its identity consistent over time.","explanation":"Object tracking locates the same object frame by frame through a video, deciding whether ‘the object in this frame is the same one as in the last frame.’ Single-object tracking follows one object given its box in the first frame; multi-object tracking (MOT) tracks many objects at once while keeping each one's own identity. A common recipe is ‘detect then associate’: run a detector every frame, predict each object's motion with a Kalman filter, and match the new detections to existing tracks with the Hungarian algorithm — SORT, from 2016, does exactly this. The difficulty lies in occlusion, changes in appearance, fast motion, and similar-looking objects getting confused with one another. Robots use tracking to keep a lock on an object they intend to grasp, to follow a pedestrian, or to track the object a hand is manipulating in a human video; video segmentation models such as SAM 2 can also track objects in the form of masks.","example":"On a sorting line, packages keep moving on the conveyor belt. The system detects each package every frame and uses tracking to maintain its ID, so the arm can grab the correct one at the right moment.","related":["Object Detection","Kalman Filter","Video Object Segmentation","Tracking Any Point","Embodied Visual Tracking","Pose Tracking"]},{"id":"video-object-segmentation","category":"perception","sec":5,"tier":3,"sources":[{"title":"DAVIS: Densely Annotated VIdeo Segmentation","url":"https://davischallenge.org/"},{"title":"SAM 2: Segment Anything in Images and Videos (arXiv:2408.00714)","url":"https://arxiv.org/abs/2408.00714"}],"as_of":"2024-10","related_ids":["instance-segmentation","object-tracking","sam-2","segment-anything-model","mask","auto-labeling"],"name":"Video Object Segmentation","alt":"视频目标分割","abbr":"VOS","aliases":["VOS"],"one_liner":"Continuously cutting out a pixel mask for a specified object in every frame of a video.","explanation":"Given a video, video object segmentation (VOS) outputs a pixel-level mask of the target object (marking which pixels belong to it) in every frame, and keeps recognizing it as the same object even as it moves, deforms, or is occluded. It’s grouped by how much guidance is given: semi-supervised VOS is given the target’s mask on the first frame and propagates it forward; unsupervised VOS is given no hint and must find the main object automatically; interactive VOS lets a user click partway through to correct the mask. DAVIS is a classic benchmark. In 2024, Meta’s SAM 2 unified image and video segmentation using a transformer with memory, letting a single click keep tracking an object through an entire video. In robotics, VOS is commonly used to keep a mask on a manipulation target, to automatically label data, or to give a policy an object-centric input.","example":"Clicking once on a cup on the table in the first frame, SAM 2 outputs a mask for the cup in every frame of the whole grasping sequence, keeping track of it even when the gripper partially covers it.","related":["Instance Segmentation","Object Tracking","SAM 2","Segment Anything Model","Mask","Auto-labeling"]},{"id":"sam-2","category":"perception","sec":5,"tier":2,"sources":[{"title":"SAM 2: Segment Anything in Images and Videos (arXiv 2408.00714)","url":"https://arxiv.org/abs/2408.00714"},{"title":"Meta AI Blog: Introducing SAM 2","url":"https://ai.meta.com/blog/segment-anything-2/"},{"title":"GitHub: IDEA-Research/Grounded-SAM-2","url":"https://github.com/IDEA-Research/Grounded-SAM-2"}],"as_of":"2024-10","related_ids":["segment-anything-model","sam-3","video-object-segmentation","object-tracking","grounded-sam","mask"],"name":"SAM 2","alt":"SAM 2（视频分割一切）","abbr":"","aliases":["Segment Anything 2","SAM 2.1"],"one_liner":"Meta's image-and-video segmentation model: mark a target once and it keeps segmenting and tracking it through the whole video.","explanation":"SAM 2 is Meta FAIR's second-generation ‘segment anything’ model, released in July 2024, extending SAM from single images to video: given a point, box, or mask specifying a target on one frame, the model continues to output that target's mask on all the frames that follow. Its core is a streaming memory: while processing frame by frame, it stores information about the target from previous frames in a memory bank that the current frame can reference, plus an occlusion head that judges whether the target is currently visible. It was released alongside the SA-V dataset (roughly 51,000 videos and more than 600,000 spatio-temporal masks). The paper reports that video segmentation needs three times fewer user interactions than before, and that image segmentation is more accurate and six times faster than the original SAM. Code and weights are open-sourced under Apache 2.0. Robots often combine it with Grounding DINO: find the object by text first, then track its mask throughout.","example":"In the Grounded-SAM-2 pipeline, Grounding DINO first detects a box for ‘red cup’ in the first frame, and SAM 2 then tracks the cup's mask throughout the manipulation video, for use in data labeling or feeding a target mask to the policy.","related":["Segment Anything Model","SAM 3","Video Object Segmentation","Object Tracking","Grounded SAM","Mask"]},{"id":"optical-flow","category":"perception","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Optical flow","url":"https://en.wikipedia.org/wiki/Optical_flow"},{"title":"RAFT: Recurrent All-Pairs Field Transforms for Optical Flow","url":"https://arxiv.org/abs/2003.12039"}],"as_of":"","related_ids":["raft","scene-flow","tracking-any-point","visual-odometry","event-camera","feature-matching"],"name":"Optical Flow","alt":"光流","abbr":"","aliases":["Dense Optical Flow","Sparse Optical Flow"],"one_liner":"How far, and in which direction, every pixel appears to move between two consecutive video frames.","explanation":"Optical flow describes the apparent motion of the brightness pattern in an image when the camera and the scene move relative to each other, and the result is usually a 2D displacement field the same size as the image. Classic methods rely on a ‘brightness constancy’ assumption — the same point keeps the same brightness across adjacent frames — but one equation can't solve for two unknowns (the aperture problem), so the 1981 Lucas-Kanade method assumes a small local window moves consistently, while Horn-Schunck instead adds a global smoothness constraint. The deep-learning-era representative is RAFT (ECCV 2020), which estimates dense optical flow using correlation volumes between all pairs of pixels plus recurrent iteration. Flow computed at only a few feature points is called sparse; flow computed at every pixel is called dense. Robots use optical flow for visual odometry, obstacle avoidance, and motion segmentation, and it's also used as an intermediate action representation: some manipulation policies first predict how points on an object will move, then convert that prediction into a robot action.","example":"As a drone flies forward, objects ahead appear to expand outward in the frame, with closer objects showing larger optical flow — this can be used to judge whether a collision is imminent.","related":["RAFT","Scene Flow","Tracking Any Point","Visual Odometry","Event Camera","Feature Matching"]},{"id":"raft","category":"perception","sec":5,"tier":3,"sources":[{"title":"RAFT: Recurrent All-Pairs Field Transforms for Optical Flow (arXiv 2003.12039)","url":"https://arxiv.org/abs/2003.12039"},{"title":"princeton-vl/RAFT (GitHub)","url":"https://github.com/princeton-vl/RAFT"}],"as_of":"","related_ids":["optical-flow","scene-flow","tracking-any-point","stereo-matching","droid-slam","recurrent-neural-network"],"name":"RAFT","alt":"RAFT 光流","abbr":"RAFT","aliases":["Recurrent All-Pairs Field Transforms","RAFT Optical Flow"],"one_liner":"A classic optical-flow network that computes correlation between every pair of pixels, then refines the flow through repeated recurrent updates.","explanation":"RAFT was proposed by Zachary Teed and Jia Deng at Princeton in 2020, published at ECCV 2020, where it reportedly won the Best Paper Award. Optical flow describes how far every pixel moves between two adjacent frames. RAFT first computes feature correlation between every pair of pixels across the two frames, producing a 4D correlation volume, and then uses a module based on a GRU (a type of recurrent neural network unit) to repeatedly refine the flow at a single high resolution, replacing the traditional coarse-to-fine pyramid approach. According to the paper, it reduced error by 16% on KITTI and 30% on Sintel compared with prior work, and it remains a widely used baseline today. The same team’s RAFT-Stereo and DROID-SLAM both reuse this iterative-update idea. In robotics research, optical flow is often used to estimate object motion, or to infer motion from video that has no action labels.","example":"Given two frames of a robot arm before and after pushing a block, RAFT outputs each pixel’s displacement, revealing which way the block and the arm each moved.","related":["Optical Flow","Scene Flow","Tracking Any Point","Stereo Matching","DROID-SLAM","Recurrent Neural Network"]},{"id":"scene-flow","category":"perception","sec":5,"tier":3,"sources":[{"title":"Three-Dimensional Scene Flow (Vedula et al., CMU RI)","url":"https://publications.ri.cmu.edu/three-dimensional-scene-flow/"},{"title":"arXiv 1806.01411: FlowNet3D","url":"https://arxiv.org/abs/1806.01411"},{"title":"arXiv 1612.02590: Scene Flow Estimation: A Survey","url":"https://arxiv.org/abs/1612.02590"}],"as_of":"","related_ids":["optical-flow","tracking-any-point","point-cloud","intermediate-representation","raft","4d-reconstruction"],"name":"Scene Flow","alt":"场景流","abbr":"","aliases":["3D Optical Flow","3D Flow"],"one_liner":"The 3D motion vector of every point in a scene between two adjacent frames.","explanation":"Scene flow is the 3D version of optical flow: optical flow describes the 2D displacement of every pixel in an image between two frames, while scene flow describes the 3D displacement of every point in the real world. The concept was introduced by Vedula, Kanade, and colleagues at CMU in 1999, with a journal version published in TPAMI in 2005. Early methods solved for it from multiple viewpoints or stereo images; once depth cameras and lidar became common, networks that learn directly on two point clouds emerged, such as FlowNet3D in 2018. Scene flow tells a system which things are moving, in what direction, and how fast, and is used for motion segmentation and dynamic-object tracking in self-driving cars. In robot manipulation, predicting the future 3D motion trajectory of points on an object — often called 3D flow — is also used as an embodiment-agnostic intermediate representation, for learning a skill from human video and then transferring it to a robot.","example":"General Flow (2024) trains a model on human RGB-D video to predict the future 3D trajectory of points on an object given a language instruction, then converts that into robot actions for zero-shot skill transfer.","related":["Optical Flow","Tracking Any Point","Point Cloud","Intermediate Representation","RAFT","4D Reconstruction"]},{"id":"tracking-any-point","category":"perception","sec":5,"tier":3,"sources":[{"title":"TAP-Vid: A Benchmark for Tracking Any Point in a Video (arXiv 2211.03726)","url":"https://arxiv.org/abs/2211.03726"},{"title":"CoTracker: It is Better to Track Together (arXiv 2307.07635)","url":"https://arxiv.org/abs/2307.07635"}],"as_of":"","related_ids":["tapir","cotracker","optical-flow","scene-flow","atm","3d-point-tracking"],"name":"Tracking Any Point","alt":"任意点跟踪","abbr":"TAP","aliases":["TAP","Point Tracking","Long-Term Point Tracking"],"one_liner":"Given any point in a video, outputting its position in every later frame, and whether it becomes occluded.","explanation":"Tracking Any Point (TAP) is a vision task formally introduced by Google DeepMind in 2022 alongside the TAP-Vid benchmark: a user marks an arbitrary point on some frame of a video — on an object’s surface, on a piece of cloth, or in the background — and the model must output that point’s pixel position in every other frame, and whether it is occluded. It differs from optical flow, which only computes motion between two adjacent frames and drifts when accumulated over a long sequence, and can’t handle a point disappearing behind an occluder and reappearing; it differs from object tracking in that it tracks a point rather than a whole object’s bounding box. Representative models include TAPIR and CoTracker. In robotics, point trajectories serve as an embodiment-agnostic intermediate representation: the motion of points on an object can be extracted from human video and used to guide policy learning (as in ATM) or to drive visual servoing.","example":"Clicking a few points on a cup’s handle in a video of a person pouring water, and tracking their trajectories over time, gives a policy-learning target for “how the cup should move.”","related":["TAPIR","CoTracker","Optical Flow","Scene Flow","ATM","3D Point Tracking"]},{"id":"tapir","category":"perception","sec":5,"tier":3,"sources":[{"title":"TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement (arXiv 2306.08637)","url":"https://arxiv.org/abs/2306.08637"}],"as_of":"2023-06","related_ids":["tracking-any-point","cotracker","optical-flow","atm","keypoint-detection","occlusion"],"name":"TAPIR","alt":"TAPIR","abbr":"","aliases":["Tracking Any Point with per-frame Initialization and temporal Refinement","TAPNet"],"one_liner":"A DeepMind point-tracking model that first finds a rough match per frame, then refines the trajectory over time.","explanation":"TAPIR is a point-tracking model proposed in 2023 (ICCV 2023) by Google DeepMind and Oxford’s VGG group, for the “Tracking Any Point” (TAP) task: given any point on any frame of a video, output that point’s position in every other frame, and whether it is occluded. The method has two stages: a matching stage that independently finds, on every frame, the candidate location most similar to the query point, used as an initialization; and a refinement stage that uses local correlation to repeatedly update the whole trajectory and the query feature over time. The paper reports a clear improvement over prior methods on the TAP-Vid benchmark. TAPNet is the earlier baseline the same team introduced in the TAP-Vid paper, and its code repository still carries the name “tapnet.” In robotics, TAPIR has been used to track keypoints on objects or on a gripper — for instance, DeepMind’s RoboTAP uses point trajectories for few-shot imitation.","example":"Clicking on one corner of a towel in a video of a robot arm folding it, TAPIR outputs that corner’s pixel coordinates in every later frame, picking it back up even after the gripper briefly hides it.","related":["Tracking Any Point","CoTracker","Optical Flow","ATM","Keypoint Detection","Occlusion"]},{"id":"cotracker","category":"perception","sec":5,"tier":3,"sources":[{"title":"CoTracker: It is Better to Track Together (arXiv 2307.07635)","url":"https://arxiv.org/abs/2307.07635"},{"title":"CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos (arXiv 2410.11831)","url":"https://arxiv.org/abs/2410.11831"},{"title":"facebookresearch/co-tracker (GitHub)","url":"https://github.com/facebookresearch/co-tracker"}],"as_of":"2025-01","related_ids":["tracking-any-point","tapir","optical-flow","atm","3d-point-tracking","occlusion"],"name":"CoTracker","alt":"CoTracker","abbr":"","aliases":["CoTracker: It is Better to Track Together","CoTracker3","CoTracker2"],"one_liner":"Meta’s open-source video point-tracking model that jointly tracks large numbers of pixels, including ones that get occluded.","explanation":"CoTracker was proposed by Karaev and colleagues at Meta AI and Oxford’s VGG group, posted to arXiv in July 2023 and published at ECCV 2024. It belongs to the tracking-any-point family: given any pixel in a video, it outputs that point’s position and visibility in every subsequent frame. Most earlier methods tracked each point independently; CoTracker uses a Transformer to track a large batch of points jointly, exploiting the correlations between points to be more robust to occlusion and points leaving the frame, and it processes video in a sliding short time window so it can run online. The October 2024 CoTracker3 simplifies the architecture and trains using pseudo-labels generated on unlabeled real videos by existing models, needing roughly 1,000 times less training data than earlier methods. In robot learning, it’s commonly used to automatically label point trajectories on demonstration videos.","example":"ATM (Any-Point Trajectory Modeling) uses CoTracker to generate point trajectories on demonstration videos as ground truth to train a trajectory-prediction model, then uses the predicted future trajectories to guide a robot policy.","related":["Tracking Any Point","TAPIR","Optical Flow","ATM","3D Point Tracking","Occlusion"]},{"id":"3d-point-tracking","category":"perception","sec":5,"tier":3,"sources":[{"title":"SpatialTrackerV2: 3D Point Tracking Made Easy (arXiv 2507.12462)","url":"https://arxiv.org/abs/2507.12462"},{"title":"TAPVid-3D: A Benchmark for Tracking Any Point in 3D (arXiv 2407.05921)","url":"https://arxiv.org/abs/2407.05921"},{"title":"General Flow as Foundation Affordance for Scalable Robot Learning (arXiv 2401.11439)","url":"https://arxiv.org/abs/2401.11439"}],"as_of":"2025-10","related_ids":["tracking-any-point","cotracker","scene-flow","monocular-depth-estimation","4d-reconstruction"],"name":"3D Point Tracking","alt":"3D 点跟踪","abbr":"","aliases":["TAP-3D","Tracking Any Point in 3D","SpatialTrackerV2"],"one_liner":"Continuously tracks arbitrary pixels through a video and outputs their motion trajectory in 3D space.","explanation":"3D point tracking is the three-dimensional version of tracking any point — following any chosen pixel through a video: given a video and a set of query points, it outputs each point’s 3D coordinates and whether it’s occluded at every frame. 2D tracking can’t tell whether the object is moving or the camera is, and it carries no depth; a 3D trajectory directly describes how an object moves and rotates in space. SpatialTracker (CVPR 2024) uses monocular depth estimation to lift pixels into 3D before tracking them; the follow-up SpatialTrackerV2 (ICCV 2025) folds point tracking, monocular depth, and camera pose estimation into one feed-forward model that works from monocular video alone. The TAPVid-3D benchmark is used to evaluate this task. In robot learning, 3D point trajectories can serve as an intermediate representation for turning human videos into robot actions — for example, General Flow predicts the future 3D trajectories of points on an object, guided by language instructions, to direct manipulation.","example":"Filming someone pulling open a drawer and running SpatialTrackerV2 to track points on the handle gives a 3D trajectory moving outward in a straight line; from that, the drawer’s sliding direction can be inferred, and a robot can pull along the same direction.","related":["Tracking Any Point","CoTracker","Scene Flow","Monocular Depth Estimation","4D Reconstruction"]},{"id":"open-vocabulary-object-detection","category":"perception","sec":5,"tier":2,"sources":[{"title":"Open-Vocabulary Object Detection Using Captions (CVPR 2021)","url":"https://arxiv.org/abs/2011.10678"},{"title":"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection","url":"https://arxiv.org/abs/2303.05499"}],"as_of":"","related_ids":["open-vocabulary","grounding-dino","owl-vit-owlv2","yolo-world","grounded-sam","object-detection"],"name":"Open-Vocabulary Object Detection","alt":"开放词汇检测","abbr":"OVD","aliases":["OVD","Open-Set Detection","Text-Guided Detection"],"one_liner":"Detecting whatever object a piece of text names, instead of being limited to a fixed list of trained classes.","explanation":"Traditional detectors only recognize the fixed set of classes they saw box-labeled at training time (COCO's 80, for instance). Open-vocabulary detection instead lets a detector accept an arbitrary text description and box classes it never saw with box labels during training. The setting was formally proposed by Alireza Zareian and colleagues at CVPR 2021: box annotations from a small set of base classes are combined with vision-language alignment learned from large-scale image-text pairs to detect new classes. OWL-ViT, Grounding DINO, and YOLO-World all followed this path afterward, with Grounding DINO reaching 52.5 AP in zero-shot detection without any COCO training. This is very practical for robots: if a user says ‘bring me the blue mug,’ the system can box it directly from that sentence. Strictly speaking, ‘open-set detection’ originally meant recognizing unseen objects as simply ‘unknown,’ but the term is now often used interchangeably with open-vocabulary detection.","example":"In the Grounded-SAM pipeline, Grounding DINO first boxes the target using the text ‘banana,’ then the box is handed to SAM to get a pixel-level mask, and the robot computes a grasp point from that.","related":["Open-vocabulary","Grounding DINO","OWL-ViT / OWLv2","YOLO-World","Grounded SAM","Object Detection"]},{"id":"grounding-dino","category":"perception","sec":5,"tier":2,"sources":[{"title":"Grounding DINO (arXiv:2303.05499)","url":"https://arxiv.org/abs/2303.05499"},{"title":"IDEA-Research/GroundingDINO (GitHub)","url":"https://github.com/IDEA-Research/GroundingDINO"},{"title":"IDEA-Research/Grounding-DINO-1.5-API (GitHub)","url":"https://github.com/IDEA-Research/Grounding-DINO-1.5-API"}],"as_of":"2024-07","related_ids":["open-vocabulary-object-detection","grounded-sam","segment-anything-model","owl-vit-owlv2","yolo-world","dino-x"],"name":"Grounding DINO","alt":"Grounding DINO","abbr":"","aliases":["GroundingDINO"],"one_liner":"An open-vocabulary detector that finds and boxes whatever object a text description names, given an image and text.","explanation":"Grounding DINO is an open-set object detection model released in March 2023 by IDEA Research together with Tsinghua University and other collaborators, with the paper later accepted at ECCV 2024. Traditional detectors only recognize the few dozen fixed classes they were trained on; Grounding DINO instead combines the transformer-based detector DINO with a text encoder, performing multi-layer fusion between image and text features, so a user can input a class name or a short description (such as ‘red cup’) and get back matching detection boxes along with the matched words. The paper reports 52.5 AP in zero-shot detection on COCO without using any COCO training data. The code is open-sourced under Apache 2.0 and has been integrated into Hugging Face Transformers. It's often chained with the segmentation model SAM into a pipeline called Grounded-SAM — box by text first, then cut a mask — a common front end for robots that need to ‘find things by instruction.’ Later versions, Grounding DINO 1.5 and 1.6, are only available through an API.","example":"A user says ‘put the banana in the bowl.’ The system runs Grounding DINO with the prompt ‘banana. bowl.’ to box both objects, uses SAM to get their masks, and combines that with the depth map to compute the grasp point and the placement point.","related":["Open-Vocabulary Object Detection","Grounded SAM","Segment Anything Model","OWL-ViT / OWLv2","YOLO-World","DINO-X"]},{"id":"owl-vit-owlv2","category":"perception","sec":5,"tier":3,"sources":[{"title":"Simple Open-Vocabulary Object Detection with Vision Transformers (arXiv 2205.06230)","url":"https://arxiv.org/abs/2205.06230"},{"title":"Scaling Open-Vocabulary Object Detection (arXiv 2306.09683)","url":"https://arxiv.org/abs/2306.09683"},{"title":"Hugging Face Transformers: OWLv2","url":"https://huggingface.co/docs/transformers/model_doc/owlv2"}],"as_of":"","related_ids":["open-vocabulary-object-detection","clip","grounding-dino","yolo-world","vision-transformer","ok-robot"],"name":"OWL-ViT / OWLv2","alt":"OWL-ViT / OWLv2","abbr":"","aliases":["Open-World Localization Vision Transformer","OWL-ST"],"one_liner":"A Google open-vocabulary object detection model that can find objects from a text description or an example image.","explanation":"OWL-ViT was proposed by Matthias Minderer and colleagues at Google, published at ECCV 2022. The approach first pretrains a vision transformer with CLIP-style image-text contrastive learning, then fine-tunes it end to end into a detector: each image patch outputs a box and a feature vector, which is compared against text features by similarity, allowing detection of categories never seen during training; it also supports one-shot detection from a single example image. OWLv2, from 2023, scales up the training data through self-training: an existing detector automatically generates pseudo-box labels on web image-text pairs, producing over 1 billion examples, which raised average precision on LVIS rare categories from 31.2% to 44.6%. Both models are available in Hugging Face Transformers and are commonly used in robotic systems to locate objects from language instructions.","example":"Given a desktop image and the text “a red mug”, OWLv2 returns a detection box and confidence score for the mug, and the robot estimates a grasp position from the depth inside that box.","related":["Open-Vocabulary Object Detection","CLIP","Grounding DINO","YOLO-World","Vision Transformer","OK-Robot"]},{"id":"yolo-world","category":"perception","sec":5,"tier":3,"sources":[{"title":"YOLO-World: Real-Time Open-Vocabulary Object Detection (arXiv)","url":"https://arxiv.org/abs/2401.17270"},{"title":"AILab-CVC/YOLO-World (GitHub)","url":"https://github.com/AILab-CVC/YOLO-World"}],"as_of":"2025-02","related_ids":["open-vocabulary-object-detection","yolo","grounding-dino","owl-vit-owlv2","object-detection","clip"],"name":"YOLO-World","alt":"YOLO-World","abbr":"","aliases":["Real-Time Open-Vocabulary Object Detection"],"one_liner":"An open-vocabulary detector that detects objects in real time just from a typed category name.","explanation":"YOLO-World is an open-vocabulary object detector proposed in 2024 by Tencent AI Lab, ARC Lab, and Huazhong University of Science and Technology, published at CVPR 2024. Traditional YOLO can only detect the fixed set of categories it was trained on; YOLO-World attaches a text encoder to YOLO, uses a network called RepVL-PAN to let image features and text features interact with each other, and pretrains at large scale with a region-text contrastive loss, so a user can detect any category just by typing its name. It uses a “prompt-then-detect” strategy: the user’s vocabulary is pre-encoded and re-parameterized into the network, so the text encoder doesn’t need to run again at inference time, keeping speed close to ordinary YOLO. The paper reports 35.4 AP zero-shot on LVIS and 52 FPS on a V100. In robotics it is commonly used to find a target in real time from a language instruction, handing the result off to a grasping or navigation module.","example":"Given the instruction “bring me the red mug,” a program sets “red mug” as YOLO-World’s vocabulary, draws a box around the mug in real time from the wrist camera feed, and hands it to a grasp-pose detection module.","related":["Open-Vocabulary Object Detection","YOLO","Grounding DINO","OWL-ViT / OWLv2","Object Detection","CLIP"]},{"id":"dino-x","category":"perception","sec":5,"tier":3,"sources":[{"title":"arXiv 2411.14347: DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding","url":"https://arxiv.org/abs/2411.14347"},{"title":"IDEA-Research/DINO-X-API (GitHub)","url":"https://github.com/IDEA-Research/DINO-X-API"}],"as_of":"2025-07","related_ids":["grounding-dino","open-vocabulary-object-detection","detr","object-detection","segment-anything-model","grounded-sam"],"name":"DINO-X","alt":"DINO-X（开放世界检测）","abbr":"","aliases":["DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding","DINO-X Pro","DINO-X Edge"],"one_liner":"IDEA Research’s open-world detection model that can box objects from text, examples, or no prompt at all.","explanation":"DINO-X is an object-centric vision model released by IDEA Research (the Guangdong-Hong Kong-Macao Greater Bay Area Institute of Digital Economy) in November 2024, with an architecture built on Grounding DINO 1.5’s Transformer encoder-decoder. It supports text prompts, visual prompts (giving example boxes), and customized prompts, and its “universal object prompt” mode lets the model box every object in an image with no prompt at all. It was trained on Grounding-100M, a set of over 100 million grounding samples the team curated. Beyond the detection head, it also carries segmentation, keypoint, and object-captioning heads, so it can output boxes, masks, poses, and text descriptions all at once. The paper reports DINO-X Pro reaching 56.0 AP on zero-shot COCO detection; there’s also a DINO-X Edge version for edge devices. It’s mainly offered through an API, and in robotics it can be used to find a target from a single instruction, then hand off to segmentation and grasping modules.","example":"Sending a tabletop photo and the prompt “cup” to the DINO-X API returns a detection box and confidence score for every cup; feeding those boxes to SAM 2 gives pixel-level masks, which combined with depth can compute each cup’s position for grasping.","related":["Grounding DINO","Open-Vocabulary Object Detection","DETR","Object Detection","Segment Anything Model","Grounded SAM"]},{"id":"open-vocabulary-segmentation","category":"perception","sec":5,"tier":3,"sources":[{"title":"Language-driven Semantic Segmentation (LSeg, arXiv 2201.03546)","url":"https://arxiv.org/abs/2201.03546"},{"title":"Towards Open Vocabulary Learning: A Survey (arXiv 2306.15880)","url":"https://arxiv.org/abs/2306.15880"}],"as_of":"","related_ids":["open-vocabulary-object-detection","semantic-segmentation","clip","grounded-sam","sam-3","open-vocabulary"],"name":"Open-Vocabulary Segmentation","alt":"开放词汇分割","abbr":"","aliases":["Open-Set Segmentation","Open-Vocabulary Semantic Segmentation"],"one_liner":"Segmenting exactly the pixels that match any text description of a category, even one the model never saw during training.","explanation":"Traditional segmentation models can only separate out the few dozen to few hundred categories fixed at training time. Open-vocabulary segmentation instead lets a category be specified at test time with arbitrary text — for example, “the blue dish sponge” — and the model outputs a pixel mask for it. It is built on CLIP-style image-text contrastive pretraining: features for each pixel or region of the image and features for the text are placed in the same embedding space, and similarity determines the match. Representative work includes LSeg (ICLR 2022); later work combined open-vocabulary detectors with SAM in Grounded-SAM, and SAM 3 accepts noun-phrase prompts directly. For robots, this means finding a particular item on a table no longer requires labeling new data and retraining for every new object.","example":"A user says “put the charging cable away”; the robot feeds the words “charging cable” into a segmentation model as text, gets back a pixel mask for the cable, and combines it with a depth map to compute a grasp point.","related":["Open-Vocabulary Object Detection","Semantic Segmentation","CLIP","Grounded SAM","SAM 3","Open-vocabulary"]},{"id":"grounded-sam","category":"perception","sec":5,"tier":3,"sources":[{"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks (arXiv 2401.14159)","url":"https://arxiv.org/abs/2401.14159"},{"title":"IDEA-Research/Grounded-Segment-Anything GitHub","url":"https://github.com/IDEA-Research/Grounded-Segment-Anything"}],"as_of":"2024-01","related_ids":["grounding-dino","segment-anything-model","open-vocabulary-segmentation","open-vocabulary-object-detection","sam-2","auto-labeling"],"name":"Grounded SAM","alt":"Grounded-SAM","abbr":"","aliases":["Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","Grounded-Segment-Anything"],"one_liner":"An open-source pipeline that boxes and finely segments the object a text phrase refers to in an image.","explanation":"Grounded SAM is an open-world vision pipeline from IDEA Research, with a technical report released in January 2024. Its core idea is chaining two models together: Grounding DINO finds an object from a text description and outputs a detection box, and SAM (the Segment Anything Model) takes that box and outputs a pixel-level mask. This means any text — say, “red cup” — can produce a mask for the matching object, with no retraining needed for a new category; it can also connect to RAM and BLIP for automatic labeling, and to Stable Diffusion for image editing. The report states it reaches 48.7 mAP on the SegInW zero-shot segmentation benchmark. In robotics, it’s commonly used to specify a target by language and extract that object’s point cloud for grasping or pose estimation, and also for automatic data labeling; the later Grounded SAM 2 connects to SAM 2 and can track objects across video.","example":"Given the instruction “put the banana in the bowl,” a robot first uses Grounded-SAM to segment masks for “banana” and “bowl” separately, then combines these with a depth map to compute both objects’ 3D positions, handing this to the grasping and motion-planning modules.","related":["Grounding DINO","Segment Anything Model","Open-Vocabulary Segmentation","Open-Vocabulary Object Detection","SAM 2","Auto-labeling"]},{"id":"sam-3","category":"perception","sec":5,"tier":2,"sources":[{"title":"SAM 3: Segment Anything with Concepts (arXiv 2511.16719)","url":"https://arxiv.org/abs/2511.16719"},{"title":"Meta AI Blog: Segment Anything Model 3","url":"https://ai.meta.com/blog/segment-anything-model-3/"},{"title":"GitHub: facebookresearch/sam3","url":"https://github.com/facebookresearch/sam3"}],"as_of":"2026-03","related_ids":["segment-anything-model","sam-2","open-vocabulary-segmentation","sam-3d","grounding-dino","instance-segmentation"],"name":"SAM 3","alt":"SAM 3（可提示概念分割）","abbr":"","aliases":["Promptable Concept Segmentation","SAM 3.1"],"one_liner":"Meta's third-generation segmentation model that finds every instance matching a short phrase or an example image.","explanation":"SAM 3 is Meta's third-generation segmentation model, released in November 2025, introducing the task of ‘promptable concept segmentation’: given a short noun phrase (such as ‘yellow school bus’), an example image patch, or both, the model finds every instance of that concept in an image or video, outputs a mask and identity for each one, and tracks them continuously. Before this, SAM and SAM 2 could only segment one user-selected target at a time and didn't accept text prompts. SAM 3 has about 848 million parameters and shares a vision encoder between a DETR-style detector and a memory-based video tracker, with a separate ‘presence head’ added to split judging ‘is it there’ from ‘where is it.’ It was released alongside the SA-Co benchmark, covering roughly 270,000 concepts. SAM 3.1, released in March 2026, sped up multi-object video tracking.","example":"Given a kitchen video and the text ‘cup,’ SAM 3 segments and continuously tracks every cup in frame, assigning each its own ID. Taking one cup's mask and combining it with the depth map then gives a point cloud for grasp-pose estimation.","related":["Segment Anything Model","SAM 2","Open-Vocabulary Segmentation","SAM 3D","Grounding DINO","Instance Segmentation"]},{"id":"visual-grounding","category":"perception","sec":5,"tier":2,"sources":[{"title":"Modeling Context in Referring Expressions (RefCOCO / RefCOCO+ / RefCOCOg, ECCV 2016)","url":"https://arxiv.org/abs/1608.00272"},{"title":"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection","url":"https://arxiv.org/abs/2303.05499"}],"as_of":"","related_ids":["referring-expression-segmentation","3d-visual-grounding","grounding-dino","open-vocabulary-object-detection","pointing","vision-language-model"],"name":"Visual Grounding","alt":"视觉定位（Grounding）","abbr":"","aliases":["Referring Expression Comprehension","REC","Phrase Grounding"],"one_liner":"Finds the location in an image that a phrase or sentence refers to.","explanation":"Visual grounding links language to image regions: given an image and a piece of text, it outputs the box, mask, or point of the object referred to. Finding a region for a phrase is phrase grounding; finding the single region that a uniquely identifying description like “the red cup on the left” points to is referring expression comprehension (REC), commonly evaluated on the RefCOCO family of datasets. It demands more than plain detection — understanding attributes and spatial relationships too. In 2023, IDEA Research’s Grounding DINO could locate objects from arbitrary text, and today most vision-language models can output boxes or point coordinates directly. When a robot follows an instruction to pick something up, grounding is usually the first step. Note that the Chinese term for this, 视觉定位, is also commonly used for a robot figuring out its own location from a camera (visual localization) — a different, unrelated task despite the similar name.","example":"Given the instruction “hand me the red cup on the left,” the model first draws a box around that one cup in the camera image, ignoring the blue cup next to it, then projects the depth points inside the box into 3D space for the grasping module to use.","related":["Referring Expression Segmentation","3D Visual Grounding","Grounding DINO","Open-Vocabulary Object Detection","Pointing","Vision-Language Model"]},{"id":"referring-expression-segmentation","category":"perception","sec":5,"tier":3,"sources":[{"title":"Segmentation from Natural Language Expressions (arXiv 1603.06180)","url":"https://arxiv.org/abs/1603.06180"}],"as_of":"","related_ids":["visual-grounding","open-vocabulary-segmentation","instance-segmentation","mask","grounded-sam","language-grounding"],"name":"Referring Expression Segmentation","alt":"指代表达分割","abbr":"RES","aliases":["RES","Referring Image Segmentation"],"one_liner":"Given a sentence describing one specific object, precisely segmenting out exactly that object in the image.","explanation":"Referring expression segmentation takes an image plus a natural-language description (such as “the two people sitting on the bench on the right”) and outputs a pixel-level mask of the object the sentence refers to. It is more fine-grained than open-vocabulary segmentation: the latter segments every object of a named category, while referring segmentation must pick out one specific instance based on color, position, relationships, and other descriptive cues. Hu et al. proposed an early end-to-end method in 2016, encoding the sentence with an LSTM and fusing it with a convolutional network’s feature map to predict a mask pixel by pixel. Common benchmarks include RefCOCO, RefCOCO+, and G-Ref. Today the task is often handled by multimodal large models or grounding models paired with a SAM-style segmentation model. When a robot hears “hand me that red cup on the left,” it must first use this to find the pixel region of the target, then combine that with depth to get a 3D position for grasping.","example":"Given the instruction “pick up the blue block on the far left of the table,” the model outputs a mask for exactly that block in the wrist camera’s image.","related":["Visual Grounding","Open-Vocabulary Segmentation","Instance Segmentation","Mask","Grounded SAM","Language Grounding"]},{"id":"depth-estimation","category":"perception","sec":6,"tier":2,"sources":[{"title":"Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer (MiDaS, arXiv 1907.01341)","url":"https://arxiv.org/abs/1907.01341"},{"title":"Depth Anything V2 (arXiv 2406.09414)","url":"https://arxiv.org/abs/2406.09414"}],"as_of":"","related_ids":["monocular-depth-estimation","stereo-matching","depth-camera","metric-depth-relative-depth","depth-completion","depth-anything"],"name":"Depth Estimation","alt":"深度估计","abbr":"","aliases":["Depth Prediction"],"one_liner":"Inferring how far every pixel in an image is from the camera, producing a depth map.","explanation":"Depth estimation is the task of inferring, from an image, the distance from the camera to each point in the scene, usually producing a depth map the same size as the image. Roughly three approaches exist: binocular stereo matching, which computes distance by triangulation from the disparity between two cameras; active ranging, such as structured light, time-of-flight depth cameras, and lidar; and monocular depth estimation, where a neural network infers depth from a single image. Monocular estimation has an inherent scale ambiguity, so early models mostly gave only relative depth; MiDaS (TPAMI 2020) improved cross-scene generalization by training on a mix of multiple datasets, and more recent models such as Depth Pro and Metric3D can output metric depth in meters directly. For robots, depth is the foundation for turning pixels into point clouds and for grasping and obstacle avoidance; learned methods can also patch the holes that depth cameras leave on transparent or reflective objects.","example":"Photographing a glass cup with a depth camera leaves a large chunk of the cup's body with missing depth. Using a learned model to estimate depth from the RGB image instead fills in that hole, producing a complete point cloud usable for grasp detection.","related":["Monocular Depth Estimation","Stereo Matching","Depth Camera","Metric Depth / Relative Depth","Depth Completion","Depth Anything"]},{"id":"monocular-depth-estimation","category":"perception","sec":6,"tier":2,"sources":[{"title":"Depth Map Prediction from a Single Image using a Multi-Scale Deep Network (arXiv:1406.2283)","url":"https://arxiv.org/abs/1406.2283"},{"title":"Depth Anything V2 (arXiv:2406.09414)","url":"https://arxiv.org/abs/2406.09414"}],"as_of":"","related_ids":["depth-estimation","metric-depth-relative-depth","depth-anything","depth-pro","depth-map","stereo-matching"],"name":"Monocular Depth Estimation","alt":"单目深度估计","abbr":"MDE","aliases":["MDE","Single-Image Depth Estimation"],"one_liner":"Predicting how far away every pixel is from just one ordinary color photo.","explanation":"Monocular depth estimation predicts the depth of every pixel from just a single ordinary RGB image, taken with one camera rather than a stereo pair, producing a depth map. It's inherently ambiguous: the same photo could show a small nearby object or a large distant one, and scale is the hardest thing to pin down, so a model can only rely on learned cues such as typical object size, perspective, and occlusion. Researchers at NYU, David Eigen and colleagues, were among the first to use a deep network to predict depth coarse-to-fine in 2014; Depth Anything V2 (NeurIPS 2024) trains on synthetic data plus large-scale pseudo-labeled real images. Output is either relative depth (only the ordering of distances is known) or metric depth (given in meters). For robots, it can supply 3D information when no depth camera is available, and it's also commonly used to recover scene geometry from ordinary internet video.","example":"An arm has only a wrist-mounted RGB camera. Using the metric-depth version of Depth Anything V2 estimates a depth map from each frame, which is then back-projected into a point cloud using the camera's intrinsics for grasp planning.","related":["Depth Estimation","Metric Depth / Relative Depth","Depth Anything","Depth Pro","Depth Map","Stereo Matching"]},{"id":"metric-depth-relative-depth","category":"perception","sec":6,"tier":3,"sources":[{"title":"ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth","url":"https://arxiv.org/abs/2302.12288"},{"title":"MiDaS: Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer","url":"https://arxiv.org/abs/1907.01341"}],"as_of":"","related_ids":["monocular-depth-estimation","depth-estimation","metric3d","depth-anything","depth-pro","depth-camera"],"name":"Metric Depth / Relative Depth","alt":"度量深度 / 相对深度","abbr":"","aliases":["Absolute Depth","Scale Ambiguity","Affine-Invariant Depth"],"one_liner":"Metric depth gives true distance in meters; relative depth only gives near/far ordering, off by an unknown scale and offset.","explanation":"Depth estimation output comes in two kinds. Metric depth (also called absolute depth) gives every pixel’s true distance from the camera, in meters — this is what a depth camera or lidar directly measures; relative depth only tells you what’s nearer and what’s farther, with values differing from true distance by an unknown scale and offset. A monocular image inherently has scale ambiguity: the same photo could be a nearby scale model or a distant real house, and different cameras have different focal lengths, so training on a mix of these without accounting for it creates contradictions. This is why methods like MiDaS use a loss insensitive to scale and offset, training on large mixed datasets to predict relative depth — generalizing well, but without metric scale — while ZoeDepth, Metric3D, and Depth Pro instead find ways to output metric depth directly. Robot grasping and obstacle avoidance need to know exactly how far an object is, so they need metric depth, or need a few real range measurements to align relative depth to true scale.","example":"Given the same tabletop photo, a relative-depth model can only tell you the cup is closer than the wall behind it; a metric-depth model gives roughly how many meters the cup is from the camera, which a robot arm needs in order to plan a grasp.","related":["Monocular Depth Estimation","Depth Estimation","Metric3D","Depth Anything","Depth Pro","Depth Camera"]},{"id":"depth-anything","category":"perception","sec":6,"tier":2,"sources":[{"title":"Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data (arXiv 2401.10891)","url":"https://arxiv.org/abs/2401.10891"},{"title":"Depth Anything V2 (arXiv 2406.09414)","url":"https://arxiv.org/abs/2406.09414"},{"title":"DepthAnything/Depth-Anything-V2（GitHub）","url":"https://github.com/DepthAnything/Depth-Anything-V2"}],"as_of":"2025-01","related_ids":["monocular-depth-estimation","metric-depth-relative-depth","depth-estimation","depth-anything-3","prompt-depth-anything","vision-foundation-model"],"name":"Depth Anything","alt":"Depth Anything","abbr":"","aliases":["Depth Anything V1","Depth Anything V2"],"one_liner":"A monocular depth estimation foundation model from HKU and TikTok that produces a depth map from a single photo.","explanation":"Depth Anything is a series of monocular depth estimation models from research teams at the University of Hong Kong and TikTok. V1 (CVPR 2024) was trained on about 1.5 million labeled images plus about 62 million unlabeled images via pseudo-labeling, giving it strong generalization, and it outputs relative depth. V2 (NeurIPS 2024) instead trains a large teacher model on synthetic data and then trains a student model on large-scale pseudo-labeled real images, producing sharper detail; it comes in several sizes from about 25 million to 1.3 billion parameters, plus a metric-depth fine-tuned version. The Small version is licensed under Apache 2.0, while the larger versions are restricted to non-commercial use. In robotics it's often used to supplement depth cameras, providing depth priors for 3D perception, navigation, and data generation. A successor, Depth Anything 3, has since been released.","example":"Running Depth Anything V2-Small on a wrist camera's RGB frame gives per-pixel relative depth; fitting a scale and offset to a small number of points actually measured by a depth camera then recovers metric depth in meters.","related":["Monocular Depth Estimation","Metric Depth / Relative Depth","Depth Estimation","Depth Anything 3","Prompt Depth Anything","Vision Foundation Model"]},{"id":"depth-pro","category":"perception","sec":6,"tier":3,"sources":[{"title":"Depth Pro: Sharp Monocular Metric Depth in Less Than a Second (arXiv 2410.02073)","url":"https://arxiv.org/abs/2410.02073"},{"title":"apple/ml-depth-pro (GitHub)","url":"https://github.com/apple/ml-depth-pro"}],"as_of":"2025-04","related_ids":["monocular-depth-estimation","metric-depth-relative-depth","depth-anything","metric3d","moge","camera-intrinsics"],"name":"Depth Pro","alt":"Depth Pro","abbr":"","aliases":["Apple Depth Pro","Depth Pro: Sharp Monocular Metric Depth in Less Than a Second"],"one_liner":"Apple’s open-source monocular metric-depth model that outputs a sharp depth map in meters from a single image.","explanation":"Depth Pro was proposed by Bochkovskii, Koltun, and colleagues at Apple; the paper, code, and weights were released in October 2024, and it was published at ICLR 2025. It does zero-shot monocular metric depth estimation: given only a single image, with no metadata like camera intrinsics required, it outputs absolute depth in meters while also estimating the camera’s focal length. It’s built for sharp edges and fine detail, and the paper reports generating a 2.25-megapixel depth map in 0.3 seconds on an ordinary GPU. Technically, it uses an efficient multi-scale vision Transformer for dense prediction, trains on a mix of real and synthetic data, and introduces a dedicated metric for measuring the accuracy of depth boundaries. Relative depth only tells you what’s nearer or farther; metric depth carries true scale, which is what lets it be back-projected directly into a point cloud a robot can use.","example":"Given a kitchen photo downloaded from the web with no capture parameters, Depth Pro outputs how many meters each pixel is from the camera along with an estimated focal length, letting a true-scale point cloud be recovered with a pinhole camera model.","related":["Monocular Depth Estimation","Metric Depth / Relative Depth","Depth Anything","Metric3D","MoGe","Camera Intrinsics"]},{"id":"metric3d","category":"perception","sec":6,"tier":3,"sources":[{"title":"arXiv: Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image","url":"https://arxiv.org/abs/2307.10984"},{"title":"arXiv: Metric3Dv2: A Versatile Monocular Geometric Foundation Model","url":"https://arxiv.org/abs/2404.15506"},{"title":"GitHub: YvanYin/Metric3D","url":"https://github.com/YvanYin/Metric3D"}],"as_of":"2025-01","related_ids":["monocular-depth-estimation","metric-depth-relative-depth","camera-intrinsics","depth-anything","depth-pro","moge"],"name":"Metric3D","alt":"Metric3D","abbr":"","aliases":["Metric3D v2","Metric3Dv2"],"one_liner":"A model that estimates metric depth from a single image, resolving scale ambiguity across cameras with a canonical camera-space transform.","explanation":"Metric3D was proposed by Wei Yin, Chunhua Shen, and colleagues, published at ICCV 2023, aiming to estimate depth at true scale from a single photo in a zero-shot setting. The difficulty is that different cameras have different focal lengths, so objects of the same size appear at different apparent distances on the image, and training on data mixed together from many cameras creates conflicting scale signals. Its solution is a “canonical camera-space transform”: during training, every sample is converted according to its focal length into a shared virtual canonical camera, and at inference time the result is converted back using the real camera’s intrinsics. This lets it train on over 8 million images from more than a thousand different cameras, giving metric depth even for unseen cameras, and it achieved the best results at the time on seven zero-shot benchmarks. The 2024 Metric3D v2 (TPAMI) switched to a DINOv2 backbone, expanded training data to over 16 million images, and added surface normal estimation. The code is open source under a BSD license.","example":"Monocular SLAM using just an ordinary camera has no way to know the scene’s true scale; the Metric3D repository recommends an open-source implementation that feeds its predicted depth into DROID-SLAM to get a mapping result at metric scale.","related":["Monocular Depth Estimation","Metric Depth / Relative Depth","Camera Intrinsics","Depth Anything","Depth Pro","MoGe"]},{"id":"moge","category":"perception","sec":6,"tier":3,"sources":[{"title":"microsoft/MoGe - GitHub","url":"https://github.com/microsoft/moge"},{"title":"CVPR 2025 Oral: MoGe","url":"https://cvpr.thecvf.com/virtual/2025/oral/35291"}],"as_of":"2025-06","related_ids":["monocular-depth-estimation","pointmap","metric-depth-relative-depth","depth-anything","dust3r","depth-pro"],"name":"MoGe","alt":"MoGe","abbr":"","aliases":["Monocular Geometry Estimation","MoGe-2"],"one_liner":"A Microsoft Research model that estimates full 3D geometry from a single image, predicting a 3D point for every pixel.","explanation":"MoGe is a monocular geometry estimation model from Microsoft Research, presented as an oral paper at CVPR 2025. From a single ordinary photo, it directly predicts a point map — a 3D coordinate for every pixel — along with a depth map and the camera’s field of view. The first version predicts an affine-invariant point map, correct only up to an unknown scale and translation, trained with a specially designed global-alignment loss plus a multi-scale local-geometry loss. MoGe-2, released in June 2025, predicts point maps at true metric (real-world) scale and adds surface-normal estimation. Robots can use it to recover a scene’s geometry from a single camera image, without needing multiple viewpoints or stereo.","example":"Given a photo of a kitchen found online, MoGe outputs a 3D coordinate for every pixel plus the camera’s field of view, which can be converted directly into a point cloud.","related":["Monocular Depth Estimation","Pointmap","Metric Depth / Relative Depth","Depth Anything","DUSt3R","Depth Pro"]},{"id":"foundationstereo","category":"perception","sec":6,"tier":3,"sources":[{"title":"FoundationStereo: Zero-Shot Stereo Matching (arXiv 2501.09898)","url":"https://arxiv.org/abs/2501.09898"},{"title":"NVlabs/FoundationStereo GitHub","url":"https://github.com/NVlabs/FoundationStereo"}],"as_of":"2025-12","related_ids":["stereo-matching","disparity","stereo-camera","depth-estimation","foundationpose","depth-anything"],"name":"FoundationStereo","alt":"FoundationStereo","abbr":"","aliases":["FoundationStereo: Zero-Shot Stereo Matching","Fast-FoundationStereo"],"one_liner":"NVIDIA’s stereo-matching foundation model that produces depth in a new scene with no fine-tuning needed.","explanation":"FoundationStereo is a stereo-matching model proposed by Bowen Wen and colleagues at NVIDIA, published at CVPR 2025. Given left and right camera images, it outputs a dense disparity map, which is then converted to depth or a point cloud using the baseline and focal length. Earlier stereo-matching networks usually needed re-fine-tuning to work in a new scene; FoundationStereo is trained on about 1 million pairs of photorealistic synthetic stereo images (the FSD dataset) and uses side-tuning to bring in monocular-depth priors from a vision foundation model, narrowing the sim-to-real gap and achieving zero-shot generalization — NVIDIA said it ranked first on the Middlebury and ETH3D leaderboards at release. In robotics it’s commonly used to get more complete depth from stereo images for grasping and pose estimation; a real-time version, Fast-FoundationStereo, followed in December 2025.","example":"Photographing a tabletop with a stereo camera and feeding the left and right images into FoundationStereo produces a depth map and point cloud, which is then passed to FoundationPose to estimate a target object’s 6D pose.","related":["Stereo Matching","Disparity","Stereo Camera","Depth Estimation","FoundationPose","Depth Anything"]},{"id":"prompt-depth-anything","category":"perception","sec":6,"tier":3,"sources":[{"title":"Prompt Depth Anything 项目主页","url":"https://promptda.github.io/"},{"title":"DepthAnything/PromptDA (GitHub)","url":"https://github.com/DepthAnything/PromptDA"},{"title":"Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation (arXiv 2412.14015)","url":"https://arxiv.org/abs/2412.14015"}],"as_of":"2025-06","related_ids":["depth-anything","monocular-depth-estimation","metric-depth-relative-depth","depth-completion","lidar","transparent-and-reflective-object-perception"],"name":"Prompt Depth Anything","alt":"Prompt Depth Anything","abbr":"","aliases":["PromptDA","Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation"],"one_liner":"Using a cheap lidar’s sparse depth readings as a “prompt” so a large depth model outputs accurate 4K metric depth.","explanation":"Prompt Depth Anything was proposed by teams from Zhejiang University and ByteDance Seed, among others, published at CVPR 2025, with code and models released under the Apache-2.0 license. Monocular depth models such as Depth Anything capture fine edges well but don’t know the real-world scale; the low-cost lidar sensors built into devices like the iPhone give true distances but at very low resolution — 24×24 in the paper’s example. Prompt Depth Anything treats the lidar depth as a prompt and fuses it into Depth Anything at multiple scales inside the decoder, producing metric (real-world-scale) depth at up to 4K resolution. The paper includes grasping experiments on a Unitree H1: a policy trained only on diffuse (non-reflective, non-transparent) objects was able to grasp transparent and reflective objects once given this depth, outperforming versions that used only RGB or only lidar.","example":"An iPhone records RGB and lidar data at the same time; Prompt Depth Anything turns them into high-resolution depth, which is then converted into a point cloud for a grasping policy to use.","related":["Depth Anything","Monocular Depth Estimation","Metric Depth / Relative Depth","Depth Completion","LiDAR","Transparent & Reflective Object Perception"]},{"id":"depth-completion","category":"perception","sec":6,"tier":3,"sources":[{"title":"Deep Depth Completion of a Single RGB-D Image (arXiv 1803.09326, CVPR 2018)","url":"https://arxiv.org/abs/1803.09326"},{"title":"ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation (arXiv 1910.02550)","url":"https://arxiv.org/abs/1910.02550"}],"as_of":"","related_ids":["depth-holes","depth-camera","transparent-and-reflective-object-perception","monocular-depth-estimation","prompt-depth-anything","lingbot-depth"],"name":"Depth Completion","alt":"深度补全","abbr":"","aliases":["Depth Inpainting","Depth Map Completion"],"one_liner":"Fills in missing or sparse pixels in a depth map to produce complete, dense depth.","explanation":"Depth completion takes an incomplete depth map — one with holes, or sparse like lidar output — usually paired with a matching color image, and predicts dense depth with a value at every pixel. The RGB-D depth cameras commonly used on robots often fail to measure depth on transparent, reflective, overly bright, or distant surfaces, yet grasping and obstacle avoidance both depend on complete geometry, so completion is a common preprocessing step. The simplest approach interpolates from neighboring pixels to fill holes; Zhang and Funkhouser at CVPR 2018 proposed first predicting surface normals and occlusion boundaries from the color image, then jointly solving for complete depth together with the raw depth; ClearGrasp (2019) specifically completes depth for transparent objects to support grasping. Recent work also completes depth by combining a monocular depth foundation model with sparse depth input.","example":"A gripper needs to grasp a glass cup on a table, but the depth camera returns almost entirely invalid values over the cup’s region; a method like ClearGrasp first fills in the cup surface’s depth, which is then passed to the grasp-detection network.","related":["Depth Holes","Depth Camera","Transparent & Reflective Object Perception","Monocular Depth Estimation","Prompt Depth Anything","LingBot-Depth"]},{"id":"lingbot-depth","category":"perception","sec":6,"tier":3,"sources":[{"title":"Masked Depth Modeling for Spatial Perception (arXiv 2601.17895)","url":"https://arxiv.org/abs/2601.17895"},{"title":"GitHub: Robbyant/lingbot-depth","url":"https://github.com/Robbyant/lingbot-depth"},{"title":"Robbyant 官网：LingBot-Depth","url":"https://technology.robbyant.com/lingbot-depth"}],"as_of":"2026-09","related_ids":["depth-completion","depth-camera","transparent-and-reflective-object-perception","depth-holes","masked-autoencoder","robbyant"],"name":"LingBot-Depth","alt":"蚂蚁灵波 LingBot-Depth","abbr":"","aliases":["Masked Depth Modeling for Spatial Perception","Masked Depth Modeling"],"one_liner":"An open-source depth-completion model from Ant Group’s Robbyant that uses a color image to repair a depth camera’s holes and noise.","explanation":"LingBot-Depth is a depth model released and open-sourced in January 2026 by Robbyant, the embodied-AI company under Ant Group; the paper is titled “Masked Depth Modeling for Spatial Perception,” and the GitHub page states it has been accepted at ECCV 2026. RGB-D depth cameras often can’t measure depth on transparent or reflective surfaces, leaving holes and noise. LingBot-Depth treats these missing regions as a natural “mask”: given a color image, raw depth, and camera intrinsics, it uses a ViT-Large backbone to fuse the two modalities and outputs a completed metric depth map and a point cloud in the camera’s coordinate frame. It was trained on about 3 million paired RGB-D samples (about 2 million real, 1 million synthetic). The team states its depth-completion error is 40–50% lower than the best prior methods. Code, weights, and the dataset are all open source.","example":"In the official grasping experiments, grasping hard-to-sense objects using depth repaired by LingBot-Depth: success rate for a transparent storage box rose from 0% to 50%, for a glass cup from 60% to 80%, and for a steel cup from 65% to 85%.","related":["Depth Completion","Depth Camera","Transparent & Reflective Object Perception","Depth Holes","Masked Autoencoder","Robbyant"]},{"id":"camera-depth-models","category":"perception","sec":6,"tier":3,"sources":[{"title":"Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots (arXiv 2509.02530)","url":"https://arxiv.org/abs/2509.02530"},{"title":"Manipulation as in Simulation 项目页","url":"https://manipulation-as-in-simulation.github.io/"}],"as_of":"2025-09","related_ids":["depth-completion","sim-to-real-transfer","sim-to-real-gap","depth-camera","transparent-and-reflective-object-perception","lingbot-depth"],"name":"Camera Depth Models","alt":"CDM 相机深度模型","abbr":"CDM","aliases":["CDM","Manipulation as in Simulation"],"one_liner":"A depth-repair model from ByteDance Seed that cleans up noisy depth-camera output into near-simulation-quality accurate depth.","explanation":"CDM comes from the September 2025 paper “Manipulation as in Simulation,” released by ByteDance Seed together with Shanghai Jiao Tong University, Zhejiang University, and Tsinghua University. Ordinary depth cameras produce noisy, incomplete depth on reflective, transparent, thin, and edge regions, which makes it hard for policies trained on depth or point clouds to transfer directly from simulation to a real robot. CDM is a software plug-in that sits right after the camera: it takes an RGB image and raw depth as input and outputs denoised, metric depth (real-world scale, in meters). Its training data comes from the authors’ “neural data engine,” which generates paired data at scale by simulating the depth noise patterns of specific camera models; models are provided per camera model, covering several RealSense models, the Azure Kinect, and the ZED 2i, and both the models and the ByteCameraDepth dataset are open source. The paper shows that a manipulation policy trained only on clean simulated depth, once paired with CDM, can handle articulated, reflective, and thin objects on a real robot with no added noise and no fine-tuning.","example":"A “put the bowl in the microwave” policy is trained in simulation using only clean depth images; at deployment, the RealSense’s raw depth is first passed through the matching CDM model, and the cleaned-up depth is then fed to the policy, with no real-robot data needed for fine-tuning.","related":["Depth Completion","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Depth Camera","Transparent & Reflective Object Perception","LingBot-Depth"]},{"id":"transparent-and-reflective-object-perception","category":"perception","sec":6,"tier":3,"sources":[{"title":"ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation (arXiv:1910.02550)","url":"https://arxiv.org/abs/1910.02550"}],"as_of":"","related_ids":["depth-completion","depth-holes","depth-camera","surface-normal-estimation","grasp-pose-detection","active-stereo"],"name":"Transparent & Reflective Object Perception","alt":"透明/反光物体感知","abbr":"","aliases":["Transparent Object Depth Estimation","Highly Reflective Object Perception","Transparent Object Perception"],"one_liner":"Getting a robot to correctly see glass, metal, and other objects that ordinary depth cameras get wrong.","explanation":"This refers to detecting, segmenting, and estimating the 3D shape of objects such as glass, clear plastic, and mirror-like metal. Common depth cameras measure range using reflected infrared light, but that light passes straight through transparent objects or bounces away off a mirror-like surface, so the resulting depth map ends up with holes (pixels with no reading) or, worse, reports the distance to whatever is behind the object, causing a robot that grasps based on that depth to miss entirely. A typical approach uses a neural network to predict a mask of the transparent region, surface normals, and occlusion boundaries from the color image, and then performs depth completion (filling in the missing depth); ClearGrasp, from Sajjan, Andy Zeng, Shuran Song, and colleagues in 2019, is representative work in this area. Glasses and bottles are everywhere in household, lab, and retail settings, so this is a problem any real grasping system has to confront.","example":"ClearGrasp estimates surface normals, a mask, and occlusion boundaries for a transparent object from a single RGB-D image, corrects the depth, and hands it to a grasping algorithm, improving a robot arm’s success rate at grasping transparent objects.","related":["Depth Completion","Depth Holes","Depth Camera","Surface Normal Estimation","Grasp Pose Detection","Active Stereo"]},{"id":"transparent-and-reflective-object-perception-datasets-and-be","category":"perception","sec":6,"tier":3,"sources":[{"title":"ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation (arXiv)","url":"https://arxiv.org/abs/1910.02550"},{"title":"TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline (arXiv)","url":"https://arxiv.org/abs/2202.08471"},{"title":"TransCG 数据集主页 (GraspNet)","url":"https://graspnet.net/transcg"}],"as_of":"2022","related_ids":["transparent-and-reflective-object-perception","depth-completion","depth-holes","depth-camera","grasp-pose-detection","realsense-depth-camera"],"name":"Transparent & Reflective Object Perception Datasets and Benchmarks","alt":"透明/反光物体感知（数据集基准）","abbr":"","aliases":["Transparent Object Depth Completion Datasets","ClearGrasp","TransCG"],"one_liner":"Depth and segmentation datasets collected specifically for glass, metal, and other objects that depth cameras struggle to measure.","explanation":"This category of dataset targets transparent or highly reflective objects such as drinking glasses, plastic bottles, and stainless-steel tableware. Ordinary depth cameras depend on light reflecting back normally, and fail on these materials with depth holes, or by mistaking the surface behind the object (such as a tabletop seen through a glass) for the object’s own surface, causing grasp failures. Representative work includes ClearGrasp, released in 2019, which provides more than 50,000 synthetic RGB-D images plus 286 real images with ground truth, and repairs depth by predicting surface normals, a transparent-region mask, and occlusion boundaries; and TransCG (RA-L 2022), from Cewu Lu’s group at Shanghai Jiao Tong University, which used two RealSense cameras to capture 57,715 real RGB-D images across 130 scenes covering 60 transparent objects, with ground-truth depth, normals, and masks included. These datasets are mainly used to train and evaluate depth-completion networks, whose repaired depth is then handed to a grasping algorithm.","example":"In a kitchen scene captured by a wrist-mounted RealSense, the region covering a drinking glass is almost entirely depth holes; only after a depth-completion network trained on TransCG fills it in can the grasp-detection module produce a usable grasp pose.","related":["Transparent & Reflective Object Perception","Depth Completion","Depth Holes","Depth Camera","Grasp Pose Detection","RealSense Depth Camera (D435i / D405)"]},{"id":"feature-points","category":"perception","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Scale-invariant feature transform","url":"https://en.wikipedia.org/wiki/Scale-invariant_feature_transform"},{"title":"Wikipedia: Oriented FAST and rotated BRIEF","url":"https://en.wikipedia.org/wiki/Oriented_FAST_and_rotated_BRIEF"},{"title":"ORB-SLAM: a Versatile and Accurate Monocular SLAM System","url":"https://arxiv.org/abs/1502.00956"}],"as_of":"","related_ids":["feature-matching","visual-slam","orb-slam3","superpoint-superglue-lightglue","keypoint-detection","structure-from-motion"],"name":"Feature Points","alt":"特征点","abbr":"","aliases":["SIFT","ORB","Local Features","Keypoints"],"one_liner":"Corners, blobs, and other points in an image that can be reliably found again, together with a vector describing their surroundings.","explanation":"Feature points are points in an image with distinctive texture that can be found again reliably even under a change of viewpoint or lighting — corners and blobs are typical examples. A feature point has two parts: a location (sometimes with a scale and orientation too), and a descriptor — a vector summarizing the appearance of the small patch around it, used to compare against other images. SIFT was proposed by David Lowe in 1999, with the full paper published in 2004; it’s robust to scaling, rotation, and lighting changes, with a 128-dimensional floating-point descriptor, and its patent expired in 2020. ORB was proposed by Rublee and colleagues in 2011, combining FAST corner detection with an improved binary BRIEF descriptor, making it much faster than SIFT and well suited to real-time systems. Feature points underlie visual SLAM, structure from motion, and image stitching, and the deep-learning era has also produced learned feature points like SuperPoint.","example":"ORB-SLAM (IEEE T-RO 2015) uses the same set of ORB features across all four of its stages: tracking, mapping, relocalization, and loop closure.","related":["Feature Matching","Visual SLAM","ORB-SLAM3","SuperPoint / SuperGlue / LightGlue","Keypoint Detection","Structure from Motion"]},{"id":"feature-matching","category":"perception","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Scale-invariant feature transform（含 Lowe 比值检验）","url":"https://en.wikipedia.org/wiki/Scale-invariant_feature_transform"},{"title":"LightGlue: Local Feature Matching at Light Speed","url":"https://arxiv.org/abs/2306.13643"}],"as_of":"","related_ids":["feature-points","random-sample-consensus","superpoint-superglue-lightglue","structure-from-motion","visual-slam","epipolar-geometry"],"name":"Feature Matching","alt":"特征匹配","abbr":"","aliases":["Keypoint Matching","Correspondence Matching"],"one_liner":"Finds pairs of feature points across two images that correspond to the same physical point.","explanation":"Feature matching finds pairs of features between two images (or two point clouds) that correspond to the same physical point. The classic pipeline: extract feature points and descriptors on each image (such as SIFT or ORB); find nearest neighbors by descriptor distance, using Euclidean distance for floating-point descriptors and Hamming distance for binary descriptors like ORB’s; then apply the ratio test proposed by David Lowe to discard ambiguous matches — if the distance ratio between the nearest and second-nearest neighbor is too close to 1, the match is dropped; finally, fit a geometric model with RANSAC to reject outliers that don’t fit. Recently, SuperGlue and LightGlue use attention networks to learn matching directly, and are more reliable under large viewpoint and lighting changes. Matching results underlie camera pose estimation, triangulation, visual SLAM, and structure from motion.","example":"ETH Zurich’s LightGlue (ICCV 2023) improves on SuperGlue by adaptively stopping computation early based on how easy or hard an image pair is, matching noticeably faster when two images overlap a lot and look similar.","related":["Feature Points","Random Sample Consensus","SuperPoint / SuperGlue / LightGlue","Structure from Motion","Visual SLAM","Epipolar Geometry"]},{"id":"superpoint-superglue-lightglue","category":"perception","sec":6,"tier":3,"sources":[{"title":"LightGlue: Local Feature Matching at Light Speed (arXiv 2306.13643)","url":"https://arxiv.org/abs/2306.13643"},{"title":"SuperPoint: Self-Supervised Interest Point Detection and Description (arXiv 1712.07629)","url":"https://arxiv.org/abs/1712.07629"},{"title":"SuperGlue: Learning Feature Matching with Graph Neural Networks (arXiv 1911.11763)","url":"https://arxiv.org/abs/1911.11763"}],"as_of":"2023-06","related_ids":["feature-points","feature-matching","structure-from-motion","visual-slam","relocalization","random-sample-consensus"],"name":"SuperPoint / SuperGlue / LightGlue","alt":"SuperPoint / LightGlue","abbr":"","aliases":["Learned Feature Matching"],"one_liner":"A line of neural-network models for detecting and matching feature points across images, replacing hand-crafted methods like SIFT.","explanation":"This is a lineage of learned feature-matching techniques. SuperPoint, proposed by Magic Leap in 2018, uses a single network to jointly output keypoint locations and descriptors (vectors describing what the area around each point looks like) in an image, replacing hand-designed features such as SIFT and ORB. SuperGlue (2020) uses a graph neural network with attention to match keypoints between two images, and can also tell which points have no correspondence in the other image. LightGlue (2023, ETH Zurich) is an improved version of SuperGlue that the paper describes as using less memory and compute, being more accurate, easier to train, and able to adaptively stop inference early depending on how hard the match is. These models are commonly plugged into the front end of SfM, visual SLAM, and relocalization pipelines, to keep finding reliable correspondences even under large changes in lighting or viewpoint.","example":"Swapping the default SIFT for SuperPoint keypoints and LightGlue matching inside an hloc or COLMAP pipeline finds more correct matches between photos taken in daylight and at night.","related":["Feature Points","Feature Matching","Structure from Motion","Visual SLAM","Relocalization","Random Sample Consensus"]},{"id":"dense-correspondence","category":"perception","sec":6,"tier":3,"sources":[{"title":"Dense Object Nets: Learning Dense Visual Object Descriptors By and For Robotic Manipulation (arXiv 1806.08756)","url":"https://arxiv.org/abs/1806.08756"},{"title":"Emergent Correspondence from Image Diffusion (DIFT, arXiv 2306.03881)","url":"https://arxiv.org/abs/2306.03881"}],"as_of":"","related_ids":["feature-matching","semantic-keypoints","optical-flow","dinov2","keypoint-detection","neural-descriptor-fields"],"name":"Dense Correspondence","alt":"稠密对应","abbr":"","aliases":["Semantic Correspondence","Dense Matching"],"one_liner":"Finds, for every pixel in one image, the matching location in another image.","explanation":"Dense correspondence means finding a match for every pixel (or nearly every pixel) between two images, as opposed to sparse matching, which matches only a small number of feature points. It comes in two flavors: geometric correspondence finds where the same physical point appears from a different viewpoint or at a different time — optical flow between neighboring video frames is one example; semantic correspondence finds parts with the same meaning across different objects, like the handles of two differently shaped cups. Recent work often uses nearest-neighbor matching on features from pretrained vision models directly, such as DINO-family features, or the features DIFT extracts from a diffusion model at NeurIPS 2023, with no dedicated training needed. In robotics, this lets a grasp point or keypoint from a demonstration transfer to a new object; MIT’s 2018 Dense Object Nets used self-supervised dense descriptors to transfer grasps between objects of the same category.","example":"In a demonstration, a person grasps the handle of a red cup; swapping in a never-before-seen blue mug, dense correspondence using DINOv2 features finds the matching handle pixels on the blue mug, and combined with depth this gives the grasp location.","related":["Feature Matching","Semantic Keypoints","Optical Flow","DINOv2","Keypoint Detection","Neural Descriptor Fields"]},{"id":"random-sample-consensus","category":"perception","sec":6,"tier":3,"sources":[{"title":"Random sample consensus - Wikipedia","url":"https://en.wikipedia.org/wiki/Random_sample_consensus"}],"as_of":"","related_ids":["feature-matching","homography","perspective-n-point","point-cloud-registration","epipolar-geometry","point-cloud-segmentation"],"name":"Random Sample Consensus","alt":"随机采样一致性","abbr":"RANSAC","aliases":["RANSAC","RANSAC Algorithm"],"one_liner":"Repeatedly fitting a model to small random subsets of data and keeping the one most data points agree with, to reject outliers.","explanation":"Random Sample Consensus (RANSAC) was proposed by Fischler and Bolles at SRI in 1981. Real-world data is often mixed with a large number of wrong points (outliers), and fitting directly with least squares gets skewed by them. RANSAC instead randomly draws the minimum number of samples needed to fit the model (for example, 3 points to fit a plane) and computes a candidate model; counts how many data points fall within an error threshold of it (inliers); repeats this many times and keeps the model with the most inliers; and finally refines that model once more using all of its inliers. It is a foundational tool in computer vision, commonly used to estimate a homography or fundamental matrix after feature matching, to solve for camera pose together with PnP, and to do coarse alignment before point cloud registration. In a robot tabletop scene, it is often used first to fit and remove the tabletop plane, leaving only the points belonging to objects on the table.","example":"Running RANSAC plane-fitting on an RGB-D point cloud finds and removes the tabletop; clustering what remains then yields the individual objects sitting on it.","related":["Feature Matching","Homography","Perspective-n-Point","Point Cloud Registration","Epipolar Geometry","Point Cloud Segmentation"]},{"id":"epipolar-geometry","category":"perception","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Epipolar geometry","url":"https://en.wikipedia.org/wiki/Epipolar_geometry"},{"title":"Wikipedia: Essential matrix","url":"https://en.wikipedia.org/wiki/Essential_matrix"}],"as_of":"","related_ids":["triangulation","stereo-matching","homography","camera-calibration","structure-from-motion","random-sample-consensus"],"name":"Epipolar Geometry","alt":"对极几何","abbr":"","aliases":["Essential Matrix","Fundamental Matrix","Epipolar Constraint"],"one_liner":"The geometric relationship that forces the matching point of a scene point, seen by two cameras, to lie on a specific line.","explanation":"Epipolar geometry describes the geometric constraint between two viewpoints. A 3D point X together with the two cameras’ optical centers forms a plane (the epipolar plane); its intersection with each image plane is an epipolar line, and the projection of one camera’s optical center onto the other image is an epipole. The core result: the match for a point in the left image must lie on a specific epipolar line in the right image, so a search for correspondences only needs to scan along that line — reducing a 2D search to 1D. Mathematically this is expressed by the 3×3 fundamental matrix F, with corresponding points x and x′ satisfying x′ᵀFx = 0; when both cameras’ intrinsics K and K′ are known, this can be written as the essential matrix E = K′ᵀFK, which encodes the rotation and the direction-only (unscaled) translation between the two cameras, introduced to computer vision by Longuet-Higgins in 1981. F or E is usually estimated from matched points with an algorithm like the eight-point algorithm combined with RANSAC, and relative pose is then decomposed from it. It underlies stereo rectification, stereo matching, triangulation, structure from motion, and visual SLAM initialization.","example":"Stereo rectification for a stereo camera pair is exactly the process of warping both images so all epipolar lines become horizontal and aligned row by row, letting stereo matching search left-right along a single row to solve for disparity.","related":["Triangulation","Stereo Matching","Homography","Camera Calibration","Structure from Motion","Random Sample Consensus"]},{"id":"homography","category":"perception","sec":6,"tier":3,"sources":[{"title":"Wikipedia: Homography (computer vision)","url":"https://en.wikipedia.org/wiki/Homography_(computer_vision)"},{"title":"Zhang: A Flexible New Technique for Camera Calibration（IEEE TPAMI 2000）","url":"https://www.microsoft.com/en-us/research/publication/a-flexible-new-technique-for-camera-calibration/"}],"as_of":"","related_ids":["pinhole-camera-model","camera-calibration","epipolar-geometry","feature-matching","random-sample-consensus","birds-eye-view"],"name":"Homography","alt":"单应性矩阵","abbr":"","aliases":["Homography Matrix","H Matrix","Homographic Transform"],"one_liner":"A 3×3 matrix describing how pixels on the same plane correspond between two images.","explanation":"A homography H is a 3×3 matrix; because uniform scaling doesn’t change the mapping, it has only 8 degrees of freedom. Under the pinhole camera model, when the same physical plane is photographed from two viewpoints, corresponding pixels in the two images satisfy x′ ∝ Hx (with x in homogeneous coordinates); when a camera only rotates about its optical center with no translation, any two images of an arbitrary scene also satisfy a homography relationship. Solving for H needs at least 4 point correspondences, commonly using the direct linear transform (DLT), with RANSAC used to reject mismatches. Homographies are used for image stitching, perspective correction, and turning an image of the ground or a tabletop into a top-down view; they are also used in camera calibration — Zhang Zhengyou’s calibration method first solves for the homography from a checkerboard plane to each image, then decomposes the camera intrinsics from those. When photographing a general 3D scene with a camera that also translates, the fundamental or essential matrix from epipolar geometry is needed instead.","example":"In a tabletop manipulation experiment, sticking markers at the table’s four corners and measuring their coordinates on the table lets you solve for the homography from image to tabletop plane, converting a detected object’s pixel position directly into x, y coordinates on the table.","related":["Pinhole Camera Model","Camera Calibration","Epipolar Geometry","Feature Matching","Random Sample Consensus","Bird’s-Eye View"]},{"id":"triangulation","category":"perception","sec":6,"tier":3,"sources":[{"title":"Triangulation (computer vision) - Wikipedia","url":"https://en.wikipedia.org/wiki/Triangulation_(computer_vision)"}],"as_of":"","related_ids":["epipolar-geometry","camera-intrinsics","camera-extrinsics","stereo-camera","structure-from-motion","reprojection-error"],"name":"Triangulation","alt":"三角化","abbr":"","aliases":[],"one_liner":"Recovering a point’s 3D coordinates from two cameras’ known positions and where that point lands in each image.","explanation":"Triangulation is a basic geometric operation in computer vision: a 3D point projects to one pixel in each of two or more images, and given each camera’s projection matrix (determined by its intrinsics and extrinsics), a ray is drawn from each camera’s optical center through the corresponding image point; the intersection of these rays is the point’s 3D location. In practice, lens distortion and feature-localization error mean the rays usually don’t intersect exactly, so the best estimate is instead found with the midpoint method, a linear solve (DLT), or by minimizing reprojection error — the gap between the observed point and the 3D point projected back into the image. Stereo depth sensing, structure from motion (SfM), and visual SLAM mapping all rely on triangulation to turn 2D matched points into 3D points.","example":"A stereo camera matches the same corner of a cup’s rim in its left and right images; given both lenses’ intrinsics and relative pose, triangulation computes how far that corner is from the camera.","related":["Epipolar Geometry","Camera Intrinsics","Camera Extrinsics","Stereo Camera","Structure from Motion","Reprojection Error"]},{"id":"bundle-adjustment","category":"perception","sec":6,"tier":3,"sources":[{"title":"Bundle adjustment - Wikipedia","url":"https://en.wikipedia.org/wiki/Bundle_adjustment"}],"as_of":"","related_ids":["structure-from-motion","reprojection-error","simultaneous-localization-and-mapping","slam-front-end-back-end","factor-graph-optimization","ceres-solver"],"name":"Bundle Adjustment","alt":"光束法平差","abbr":"BA","aliases":["BA"],"one_liner":"Jointly fine-tunes all camera poses and 3D point coordinates to minimize reprojection error.","explanation":"Bundle adjustment traces back to photogrammetry in the 1950s; “bundle” refers to the bundle of light rays traveling from each 3D point to a camera’s optical center. Given feature points matched across multiple images, it treats every camera’s pose — sometimes intrinsics and distortion too — and every 3D point’s coordinates as unknowns, and minimizes reprojection error: the sum of squared distances between where each 3D point projects under the current parameters and where it was actually observed in the image. When image noise is zero-mean Gaussian, this is equivalent to maximum likelihood estimation. It is typically solved with nonlinear least-squares methods like Levenberg–Marquardt, exploiting the problem’s sparse structure for speed. Bundle adjustment is the core step of structure from motion (SfM) and the back end of visual SLAM — COLMAP reconstruction and ORB-SLAM’s local and global optimization are both running BA — with common solver libraries including Ceres, g2o, and GTSAM.","example":"Walking around a table taking dozens of photos with a phone: SfM first roughly estimates each photo’s camera pose and triangulates a sparse point cloud, then runs one pass of BA to jointly refine everything, dropping the reprojection error and making both the point cloud and camera trajectory more accurate.","related":["Structure from Motion","Reprojection Error","Simultaneous Localization and Mapping","SLAM Front-end / Back-end","Factor Graph Optimization","Ceres Solver"]},{"id":"structure-from-motion","category":"perception","sec":6,"tier":3,"sources":[{"title":"Structure-from-Motion Revisited (Schönberger & Frahm, CVPR 2016)","url":"https://openaccess.thecvf.com/content_cvpr_2016/html/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.html"},{"title":"COLMAP 官方文档","url":"https://colmap.github.io/"}],"as_of":"","related_ids":["bundle-adjustment","feature-matching","triangulation","colmap","multi-view-stereo","feed-forward-3d-reconstruction"],"name":"Structure from Motion","alt":"运动恢复结构","abbr":"SfM","aliases":["SfM"],"one_liner":"Computing both camera poses and a scene’s 3D points at once from a set of photos taken from different angles.","explanation":"Structure from motion is a classic 3D reconstruction method: given photos of the same scene taken from multiple viewpoints (which can be in any order), it outputs the camera pose (position and orientation) for every photo along with a sparse 3D point cloud. The standard pipeline first extracts feature points and matches them across images, uses epipolar geometry to reject bad matches, then adds images one at a time while triangulating 3D points, and finally runs bundle adjustment — jointly fine-tuning all poses and 3D points to minimize reprojection error — as a global refinement. It differs from SLAM in that it usually runs offline and doesn’t require the images to be in chronological order. The open-source tool COLMAP is the most commonly used implementation; the first step in building a NeRF, a 3D Gaussian Splat, or bringing a real scene into simulation (real-to-sim) is often to compute camera poses with SfM. Recently, feed-forward models such as DUSt3R and VGGT have tried to output both pose and geometry directly, in one shot.","example":"Fifty phone photos taken while circling a mug on a table are fed into COLMAP, producing each photo’s camera pose and a sparse point cloud of the mug, which is then used to train a 3D Gaussian Splat.","related":["Bundle Adjustment","Feature Matching","Triangulation","COLMAP","Multi-View Stereo","Feed-Forward 3D Reconstruction"]},{"id":"multi-view-stereo","category":"perception","sec":6,"tier":3,"sources":[{"title":"COLMAP Tutorial","url":"https://colmap.github.io/tutorial.html"},{"title":"MVSNet: Depth Inference for Unstructured Multi-view Stereo (arXiv)","url":"https://arxiv.org/abs/1804.02505"}],"as_of":"","related_ids":["structure-from-motion","colmap","triangulation","bundle-adjustment","neural-radiance-fields","3d-gaussian-splatting"],"name":"Multi-View Stereo","alt":"多视图立体","abbr":"MVS","aliases":["MVS","Multi-View 3D Reconstruction"],"one_liner":"A method that recovers a scene’s dense 3D structure from multiple photos whose camera positions are already known.","explanation":"Multi-view stereo (MVS) is a classic 3D reconstruction technique. Given each photo’s camera intrinsics and extrinsics — usually solved first by structure from motion (SfM) — it finds correspondences for the same physical point across multiple images, then triangulates to compute depth, producing a dense depth map or point cloud that can be turned into a mesh. SfM alone yields only a sparse set of feature points; MVS is what fills that in to a dense model. COLMAP is a widely used implementation of the classic pipeline; since 2018, deep-learning methods such as MVSNet have used neural networks to regress depth directly. MVS is a basic building block for digital twins, scanning real objects into simulation assets, and novel view synthesis.","example":"COLMAP first runs structure from motion on a set of photos taken around an object to recover the camera poses, then runs multi-view stereo to produce a dense point cloud.","related":["Structure from Motion","COLMAP","Triangulation","Bundle Adjustment","Neural Radiance Fields","3D Gaussian Splatting"]},{"id":"feed-forward-3d-reconstruction","category":"perception","sec":6,"tier":3,"sources":[{"title":"DUSt3R: Geometric 3D Vision Made Easy","url":"https://arxiv.org/abs/2312.14132"},{"title":"VGGT: Visual Geometry Grounded Transformer","url":"https://arxiv.org/abs/2503.11651"},{"title":"MapAnything: Universal Feed-Forward Metric 3D Reconstruction","url":"https://arxiv.org/abs/2509.13414"}],"as_of":"2025-09","related_ids":["dust3r","vggt","pi3","mapanything","pointmap","structure-from-motion"],"name":"Feed-Forward 3D Reconstruction","alt":"前馈式三维重建","abbr":"","aliases":["3D Reconstruction Foundation Model"],"one_liner":"A 3D reconstruction approach that outputs camera parameters, depth, and a point cloud directly from images in a single network forward pass.","explanation":"Traditional 3D reconstruction follows a structure from motion (SfM) plus multi-view stereo (MVS) pipeline: match feature points, estimate camera poses, then iteratively refine everything with bundle adjustment — a process with many steps that takes a long time and easily fails with little texture or few viewpoints. Feed-forward 3D reconstruction instead uses a large network (usually a Transformer) trained on large-scale 3D data; given one or more images, a single forward pass directly outputs a point map (a 3D coordinate for every pixel), depth, and camera intrinsics and extrinsics. Landmark examples include Naver’s DUSt3R (CVPR 2024, needing no prior camera calibration), Meta and Oxford’s VGGT (CVPR 2025 best paper), π³, and MapAnything. For robotics, this lets scene geometry be recovered quickly from ordinary RGB images, useful for mapping, pose estimation, or feeding 3D input to a policy.","example":"VGGT takes anywhere from one to several hundred photos of the same scene and, according to the paper, can directly predict camera parameters, depth maps, point maps, and 3D point trajectories in under a second, with no post-hoc optimization like bundle adjustment needed.","related":["DUSt3R","VGGT","π³ (Pi3)","MapAnything","Pointmap","Structure from Motion"]},{"id":"pointmap","category":"perception","sec":6,"tier":3,"sources":[{"title":"DUSt3R: Geometric 3D Vision Made Easy (arXiv 2312.14132)","url":"https://arxiv.org/abs/2312.14132"}],"as_of":"","related_ids":["feed-forward-3d-reconstruction","dust3r","vggt","depth-map","point-cloud","camera-intrinsics"],"name":"Pointmap","alt":"点图","abbr":"","aliases":["Point Map","Per-Pixel 3D Point Map"],"one_liner":"An array the same size as an image, where each pixel stores a 3D coordinate instead of a color.","explanation":"A pointmap arranges the 3D coordinate (x, y, z) corresponding to every pixel of an image into an H×W×3 array, giving a one-to-one correspondence between pixels and 3D points. DUSt3R, proposed in 2023 by Naver Labs Europe and others, made this its core output: given two images, the network directly regresses two pointmaps, both expressed in the camera coordinate frame of the first image, so it needs no prior knowledge of camera intrinsics (such as focal length) or camera poses. Compared with a depth map, a pointmap implicitly encodes depth, camera parameters, and pixel correspondence between the two images all at once, and the focal length, relative pose, and matching points can all be recovered from it. Later feed-forward 3D reconstruction models such as MASt3R, VGGT, and π³ adopted or remained compatible with this representation, letting robots quickly recover a scene’s 3D structure from ordinary RGB images.","example":"Two phone photos of a tabletop are fed into DUSt3R, producing two pointmaps that, combined, form a dense point cloud of the tabletop scene.","related":["Feed-Forward 3D Reconstruction","DUSt3R","VGGT","Depth Map","Point Cloud","Camera Intrinsics"]},{"id":"dust3r","category":"perception","sec":6,"tier":3,"sources":[{"title":"arXiv 2312.14132: DUSt3R: Geometric 3D Vision Made Easy","url":"https://arxiv.org/abs/2312.14132"},{"title":"naver/dust3r (GitHub)","url":"https://github.com/naver/dust3r"}],"as_of":"2024-06","related_ids":["pointmap","mast3r","vggt","feed-forward-3d-reconstruction","structure-from-motion","multi-view-stereo"],"name":"DUSt3R","alt":"DUSt3R","abbr":"","aliases":["DUSt3R: Geometric 3D Vision Made Easy","Dense and Unconstrained Stereo 3D Reconstruction"],"one_liner":"A feed-forward reconstruction model that regresses a per-pixel 3D point map directly from two images, with no camera parameters needed.","explanation":"DUSt3R is a 3D reconstruction model proposed by Shuzhe Wang and colleagues at Naver Labs Europe and Aalto University, published at CVPR 2024. The traditional pipeline — structure from motion (SfM) plus multi-view stereo (MVS) — first estimates camera intrinsics and extrinsics, then triangulates matched points, a process with many steps that’s easy to break. DUSt3R flips this around: it feeds two images into a Transformer encoder-decoder and directly regresses a 3D coordinate for every pixel, called a point map, with both images’ points expressed in the first image’s camera coordinate frame and each carrying a confidence score; depth, pixel correspondence, relative pose, and focal length can all be read off the point map. With more than two images, a global alignment step then unifies all pairwise results into one coordinate system. DUSt3R is the landmark work in “feed-forward 3D reconstruction,” and later work like MASt3R, VGGT, and π³ all build on this approach. The code is released under a CC BY-NC-SA 4.0 non-commercial license.","example":"Snapping two casual phone photos of a tabletop with no camera calibration, DUSt3R outputs the corresponding 3D point cloud for both images along with the two cameras’ relative pose, which can be used to quickly reconstruct a robot’s workspace.","related":["Pointmap","MASt3R","VGGT","Feed-Forward 3D Reconstruction","Structure from Motion","Multi-View Stereo"]},{"id":"mast3r","category":"perception","sec":6,"tier":3,"sources":[{"title":"arXiv: Grounding Image Matching in 3D with MASt3R","url":"https://arxiv.org/abs/2406.09756"},{"title":"GitHub: naver/mast3r","url":"https://github.com/naver/mast3r"}],"as_of":"","related_ids":["dust3r","mast3r-slam","feature-matching","pointmap","feed-forward-3d-reconstruction","vggt"],"name":"MASt3R","alt":"MASt3R","abbr":"","aliases":["MASt3R: Grounding Image Matching in 3D"],"one_liner":"Naver’s model that adds a matching head onto DUSt3R, outputting both a 3D point map and dense matching features.","explanation":"MASt3R was proposed by Vincent Leroy, Yohann Cabon, and Jérôme Revaud at Naver in 2024, published at ECCV 2024. Its starting premise is that image matching — finding the pixels in two images that correspond to the same physical point — is fundamentally a 3D problem, tightly bound up with camera pose and scene geometry. It builds on DUSt3R, which takes two images and directly regresses a 3D coordinate for every pixel (a point map), robust to large viewpoint changes but with limited matching accuracy — and adds a new head that outputs dense local features, trained with an additional matching loss, paired with a fast mutual-nearest-neighbor matching algorithm that speeds up matching by several orders of magnitude. It substantially leads on several matching benchmarks, with a 30-point absolute improvement in VCRE AUC on the Map-free localization dataset. Both MASt3R-SfM and MASt3R-SLAM are built around it. The code is released under a CC BY-NC-SA 4.0 license, non-commercial use only.","example":"Given two tabletop photos of a robot’s workspace taken from very different viewpoints, MASt3R can directly output the pixel correspondences between the two images along with each one’s 3D point map, from which the relative pose between the two viewpoints can be computed.","related":["DUSt3R","MASt3R-SLAM","Feature Matching","Pointmap","Feed-Forward 3D Reconstruction","VGGT"]},{"id":"vggt","category":"perception","sec":6,"tier":3,"sources":[{"title":"VGGT: Visual Geometry Grounded Transformer (arXiv)","url":"https://arxiv.org/abs/2503.11651"},{"title":"facebookresearch/vggt (GitHub)","url":"https://github.com/facebookresearch/vggt"}],"as_of":"2025-07","related_ids":["feed-forward-3d-reconstruction","dust3r","pi3","structure-from-motion","pointmap","bundle-adjustment"],"name":"VGGT","alt":"VGGT","abbr":"VGGT","aliases":["Visual Geometry Grounded Transformer","VGGT-1B"],"one_liner":"A 3D vision model that computes camera parameters, depth, and a point cloud from multiple images in a single forward pass.","explanation":"VGGT is a feed-forward 3D reconstruction model from Oxford’s Visual Geometry Group (VGG) and Meta AI, with about 1 billion parameters, and won the CVPR 2025 Best Paper Award. Given anywhere from one image to a few hundred, it runs a single forward pass to output every image’s camera parameters, depth map, pointmap (the 3D coordinate corresponding to each pixel), and 3D point tracks — usually in under a second, with no need for bundle adjustment (the repeated optimization of cameras and 3D points that traditional reconstruction relies on). Work that previously required a multi-stage pipeline such as COLMAP is compressed into a single network. In embodied AI, it is commonly used to quickly get scene geometry from multi-view images. The original weights are non-commercial only; in July 2025, a separate set of commercially usable weights and open training code were released.","example":"Twenty phone photos taken circling a tabletop are fed into VGGT, producing each photo’s camera pose and a dense point cloud of the whole table in about a second.","related":["Feed-Forward 3D Reconstruction","DUSt3R","π³ (Pi3)","Structure from Motion","Pointmap","Bundle Adjustment"]},{"id":"pi3","category":"perception","sec":6,"tier":3,"sources":[{"title":"π³: Permutation-Equivariant Visual Geometry Learning (arXiv)","url":"https://arxiv.org/abs/2507.13347"},{"title":"yyfz/Pi3 (GitHub)","url":"https://github.com/yyfz/Pi3"}],"as_of":"2025-12","related_ids":["vggt","dust3r","feed-forward-3d-reconstruction","pointmap","camera-extrinsics","shanghai-artificial-intelligence-laboratory"],"name":"π³ (Pi3)","alt":"π³（Pi3）","abbr":"","aliases":["Pi3","Pi3X","Permutation-Equivariant Visual Geometry Learning"],"one_liner":"A feed-forward 3D reconstruction model with no reference viewpoint, giving the same result no matter what order the images come in.","explanation":"π³ is a feed-forward 3D reconstruction model released in July 2025 by the Shanghai AI Laboratory, Zhejiang University, and other institutions, with its repository marked ICLR 2026. Earlier methods such as DUSt3R and VGGT need one image designated as a reference viewpoint, with every result expressed in that image’s coordinate frame, so a poorly chosen reference image degrades the reconstruction. π³ instead uses a fully permutation-equivariant architecture (shuffle the input order and the output just follows the same reordering, unchanged in content), so it needs no reference frame at all, directly predicting an affine-invariant camera pose and a scale-invariant local pointmap for every image. The paper reports state-of-the-art results at the time on camera pose estimation, monocular and video depth estimation, and dense pointmap reconstruction. An upgraded version, Pi3X, released in December 2025, can take known poses, intrinsics, or depth as input and produce output at approximately true scale. The code is BSD-licensed; the weights are non-commercial only.","example":"The same set of 10 indoor photos is fed into π³ in different orders; the resulting point cloud and relative camera poses stay essentially the same, whereas a reference-frame-based method’s result changes depending on which image is chosen first.","related":["VGGT","DUSt3R","Feed-Forward 3D Reconstruction","Pointmap","Camera Extrinsics","Shanghai Artificial Intelligence Laboratory"]},{"id":"cut3r","category":"perception","sec":6,"tier":3,"sources":[{"title":"Continuous 3D Perception Model with Persistent State (arXiv 2501.12387)","url":"https://arxiv.org/abs/2501.12387"},{"title":"CUT3R project page","url":"https://cut3r.github.io/"}],"as_of":"2025-06","related_ids":["dust3r","pointmap","feed-forward-3d-reconstruction","4d-reconstruction","vggt","mast3r-slam"],"name":"CUT3R","alt":"CUT3R","abbr":"","aliases":["Continuous 3D Perception Model with Persistent State","Continuous Updating Transformer for 3D Reconstruction"],"one_liner":"A 3D reconstruction model with a persistent memory state that reads images one at a time and updates the whole scene online.","explanation":"CUT3R stands for Continuous Updating Transformer for 3D Reconstruction, proposed by researchers at UC Berkeley and Google DeepMind, and presented as an oral paper at CVPR 2025. It follows the same idea as DUSt3R-style models — regressing point maps directly from images, with one 3D point per pixel — but adds a continuously updated state: as each new image arrives, the model first uses it to update the state, then outputs that image’s point map in a shared coordinate system, at real-world scale; the reconstruction gradually gets more complete as more input arrives, with no per-video optimization needed. It can process either a video stream or an unordered set of photos, supports dynamic scenes with moving objects, and can even infer unseen regions from virtual viewpoints that were never actually photographed.","example":"Feeding a robot head camera’s video into CUT3R frame by frame produces, for each new frame, a point map aligned into the same coordinate system; accumulated together, this becomes a room point cloud that keeps getting more complete.","related":["DUSt3R","Pointmap","Feed-Forward 3D Reconstruction","4D Reconstruction","VGGT","MASt3R-SLAM"]},{"id":"mapanything","category":"perception","sec":6,"tier":3,"sources":[{"title":"arXiv: MapAnything: Universal Feed-Forward Metric 3D Reconstruction","url":"https://arxiv.org/abs/2509.13414"},{"title":"MapAnything 项目主页","url":"https://map-anything.github.io/"}],"as_of":"2026-01","related_ids":["feed-forward-3d-reconstruction","dust3r","vggt","mast3r","structure-from-motion","metric-depth-relative-depth"],"name":"MapAnything","alt":"MapAnything","abbr":"","aliases":["MapAnything: Universal Feed-Forward Metric 3D Reconstruction","Map Anything"],"one_liner":"Meta and CMU’s general-purpose feed-forward 3D reconstruction model that directly outputs true-scale scene structure.","explanation":"MapAnything was released by Meta Reality Labs and Carnegie Mellon University in September 2025, a Transformer-based feed-forward 3D reconstruction model: given one or more images, optionally supplemented with camera intrinsics, poses, depth, or partial reconstruction results, a single forward pass outputs each image’s depth map, local ray map, camera pose, and a unified metric scale factor — together forming a 3D scene at true, meter-based scale. Tasks that used to each need their own algorithm — uncalibrated structure from motion, multi-view stereo, monocular depth estimation, camera localization, depth completion — are all covered by this one model. It belongs to the same feed-forward reconstruction lineage as DUSt3R and VGGT, with code and weights open-sourced (including an Apache-licensed version). For robotics, it can turn a handful of photos into a metric point cloud that is directly usable for planning.","example":"Taking a few phone photos around a tabletop with no camera parameters provided, MapAnything can output a metric-scale point cloud and each photo’s camera pose; if camera intrinsics are known, they can also be supplied as a constraint.","related":["Feed-Forward 3D Reconstruction","DUSt3R","VGGT","MASt3R","Structure from Motion","Metric Depth / Relative Depth"]},{"id":"depth-anything-3","category":"perception","sec":6,"tier":3,"sources":[{"title":"Depth Anything 3: Recovering the Visual Space from Any Views (arXiv 2511.10647)","url":"https://arxiv.org/abs/2511.10647"},{"title":"ByteDance-Seed/Depth-Anything-3 (GitHub)","url":"https://github.com/ByteDance-Seed/Depth-Anything-3"}],"as_of":"2025-12","related_ids":["depth-anything","vggt","feed-forward-3d-reconstruction","monocular-depth-estimation","3d-gaussian-splatting","dinov2"],"name":"Depth Anything 3","alt":"Depth Anything 3","abbr":"DA3","aliases":["DA3","Depth Anything 3: Recovering the Visual Space from Any Views"],"one_liner":"ByteDance Seed’s geometry model that recovers depth and camera poses from any number of images, with poses optional.","explanation":"Depth Anything 3 was released by ByteDance’s Seed team in November 2025, the third generation of the Depth Anything series. The first two generations only did depth estimation from a single image; DA3 extends this to any number of input images, with camera poses either known or unknown, outputting spatially consistent depth and camera parameters, which can be further turned into a point cloud or 3D Gaussians. The design is deliberately simple: the backbone is just an ordinary Transformer (the original DINO encoder), and the training objective is unified into a single “depth plus ray” prediction. The paper reports, on its own visual geometry benchmark, that camera pose accuracy is on average 44.3% higher than the previous best method, VGGT, and geometric accuracy is 25.1% higher. Models range in size from Small (0.08B) to Giant (1.15B), with additional metric-depth and monocular-specific versions available.","example":"Walking around a table taking 5 photos on a phone with no camera parameters provided, DA3 gives every image’s depth and camera pose in a single forward pass and fuses them into one tabletop point cloud.","related":["Depth Anything","VGGT","Feed-Forward 3D Reconstruction","Monocular Depth Estimation","3D Gaussian Splatting","DINOv2"]},{"id":"voxel","category":"perception","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Voxel","url":"https://en.wikipedia.org/wiki/Voxel"},{"title":"Open3D Tutorial: Point cloud (Voxel downsampling)","url":"https://www.open3d.org/docs/release/tutorial/geometry/pointcloud.html"},{"title":"PerAct: Perceiver-Actor project page","url":"https://peract.github.io/"}],"as_of":"","related_ids":["point-cloud","occupancy-grid-map","octomap","truncated-signed-distance-function","point-cloud-filtering-and-voxel-downsampling","peract"],"name":"Voxel","alt":"体素","abbr":"","aliases":["Voxel Grid","Voxel Downsampling","Volumetric Pixel"],"one_liner":"A small cube in 3D space — the volumetric equivalent of a pixel.","explanation":"A voxel (a blend of “volume” and “pixel”) is one small cube produced by slicing 3D space into a regular grid; each cell can store whether it’s occupied, a color, a distance to the nearest surface, or a feature vector, just as a pixel does in a 2D image. Compared with a point cloud — a scattered set of 3D points — voxels are arranged regularly, so they can be processed directly with 3D convolutions and make it easy to check whether a given location is blocked; the tradeoff is that memory use grows fast with resolution, since halving the cell size multiplies the cell count by eight. Common robotics uses include occupancy grid maps and octree maps, which use voxels to mark where a robot can go; the truncated signed distance function (TSDF), which stores a distance value in each voxel for 3D reconstruction; and voxel downsampling of point clouds, which merges all points inside one cell into a single average point. The manipulation policy PerAct divides the workspace into a 100×100×100 voxel grid and directly predicts which cell the gripper should move to next.","example":"In Open3D, a single line — pcd.voxel_down_sample(voxel_size=0.05) — downsamples a dense point cloud to one point per 5-centimeter cell (with coordinates in meters), sharply cutting the point count before it’s sent on for registration or grasp detection.","related":["Point Cloud","Occupancy Grid Map","OctoMap","Truncated Signed Distance Function","Point Cloud Filtering & Voxel Downsampling","PerAct"]},{"id":"point-cloud-filtering-and-voxel-downsampling","category":"perception","sec":7,"tier":2,"sources":[{"title":"Open3D: Point cloud outlier removal","url":"https://www.open3d.org/docs/release/tutorial/geometry/pointcloud_outlier_removal.html"},{"title":"PCL Tutorial: Downsampling a PointCloud using a VoxelGrid filter","url":"https://pcl.readthedocs.io/projects/tutorials/en/latest/voxel_grid.html"},{"title":"3D Diffusion Policy (arXiv 2403.03954)","url":"https://arxiv.org/html/2403.03954"}],"as_of":"","related_ids":["point-cloud","voxel","farthest-point-sampling","point-cloud-registration","flying-pixels","3d-diffusion-policy"],"name":"Point Cloud Filtering & Voxel Downsampling","alt":"点云滤波与降采样（体素降采样 / 离群点去除）","abbr":"","aliases":["Voxel Downsampling","Voxel Grid Filter","Outlier Removal"],"one_liner":"Preprocessing steps that crop irrelevant regions, remove noisy points, and shrink a raw point cloud down to a usable size.","explanation":"Raw point clouds straight from a depth camera or lidar often have hundreds of thousands of points, mixed with stray outlier points from measurement error, and with irrelevant regions like the tabletop or the floor. Feeding this directly into downstream algorithms is both slow and unstable, so it's preprocessed first. Common steps include: cropping by a bounding box to keep only the work region; voxel downsampling, which divides space into fixed-size small cubes (voxels) and merges the points within each one into a single point (PCL takes the centroid, Open3D also averages color and normal), drastically reducing the point count and making density more uniform; and outlier removal, where a statistical method removes points whose average distance to their k nearest neighbors deviates too far from the global mean, and a radius-based method removes points with too few neighbors within a given radius. Both PCL and Open3D provide ready-made functions for these steps.","example":"PCL's tutorial downsamples one scan with a 1 cm voxel, reducing the point count from 460,400 to 41,049. 3D Diffusion Policy (DP3) instead first crops out the tabletop and background with a bounding box, then keeps only 512 or 1,024 points via farthest point sampling.","related":["Point Cloud","Voxel","Farthest Point Sampling","Point Cloud Registration","Flying Pixels","3D Diffusion Policy"]},{"id":"farthest-point-sampling","category":"perception","sec":7,"tier":3,"sources":[{"title":"PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space","url":"https://arxiv.org/abs/1706.02413"},{"title":"3D Diffusion Policy (DP3)","url":"https://arxiv.org/html/2403.03954"}],"as_of":"","related_ids":["point-cloud","pointnet-pointnet-plus-plus","point-cloud-filtering-and-voxel-downsampling","3d-diffusion-policy","point-cloud-encoder","depth-camera"],"name":"Farthest Point Sampling","alt":"最远点采样","abbr":"FPS","aliases":["FPS","Furthest Point Sampling"],"one_liner":"Repeatedly picks the point farthest from the already-selected set, downsampling a point cloud evenly to a fixed number of points.","explanation":"Farthest point sampling is a point-cloud downsampling algorithm: starting from one chosen point, it repeatedly picks the remaining point whose distance to the nearest already-selected point is the largest, until the desired number of points is reached. Compared with random sampling, it covers an entire object or scene more evenly with the same number of points, and is less likely to lose structure in sparse regions; the drawback is that it requires point-by-point iteration, making it slower than random sampling or voxel downsampling on large point clouds, and its result depends on the starting point. PointNet++ (NeurIPS 2017) uses it to pick the center point of each local region at every layer, and many later point-cloud networks have followed this practice. When a robot policy takes a point cloud as input, the table and background are usually cropped out first, and then farthest point sampling brings the count down to a few hundred to a few thousand points to control compute.","example":"The 3D Diffusion Policy DP3 (RSS 2024) first crops the point cloud produced by a depth camera, then uses farthest point sampling to bring it down to 512 or 1,024 points, which the authors say is enough for both simulated and real-robot tasks.","related":["Point Cloud","PointNet / PointNet++","Point Cloud Filtering & Voxel Downsampling","3D Diffusion Policy","Point Cloud Encoder","Depth Camera"]},{"id":"surface-normal-estimation","category":"perception","sec":7,"tier":3,"sources":[{"title":"PCL Tutorial: Estimating Surface Normals in a PointCloud","url":"https://pcl.readthedocs.io/projects/tutorials/en/latest/normal_estimation.html"},{"title":"Open3D Point Cloud Tutorial (Vertex normal estimation)","url":"https://www.open3d.org/docs/release/tutorial/geometry/pointcloud.html"}],"as_of":"","related_ids":["point-cloud","point-cloud-registration","fast-point-feature-histograms","grasp-pose-detection","vacuum-suction-cup","point-cloud-library"],"name":"Surface Normal Estimation","alt":"法向量估计","abbr":"","aliases":["Point Cloud Normal Estimation","Normal Estimation"],"one_liner":"Computing which way each point on an object’s surface faces — the vector perpendicular to the surface there.","explanation":"A surface normal is the unit vector perpendicular to an object’s surface at a given point, indicating which way that patch of surface faces. A point cloud is just a set of coordinates with no built-in notion of a “surface,” so normals have to be estimated: the most common approach takes each point’s k nearest neighbors and runs principal component analysis (PCA), where the direction of least variance is taken as the normal; because this leaves the front-versus-back direction ambiguous, normals are then reoriented consistently to face the camera or point outward. Neural networks that predict a normal for every pixel directly from a single RGB image also exist. Normals are used heavily in robotics: suction grasping needs the cup to sit flush and perpendicular to the surface, two-finger grasps often approach along the normal direction, and point-to-plane ICP registration, FPFH point cloud features, and Poisson surface reconstruction all depend on them. Normals visibly jitter when the point cloud is noisy or the neighborhood size is poorly chosen.","example":"Calling estimate_normals on a point cloud in Open3D, then orient_normals_towards_camera_location to make them consistent, gives an approach direction for every candidate suction-grasp point.","related":["Point Cloud","Point Cloud Registration","Fast Point Feature Histograms","Grasp Pose Detection","Vacuum Suction Cup","Point Cloud Library (PCL)"]},{"id":"fast-point-feature-histograms","category":"perception","sec":7,"tier":3,"sources":[{"title":"PCL Tutorials: Fast Point Feature Histograms (FPFH) descriptors","url":"https://pcl.readthedocs.io/projects/tutorials/en/latest/fpfh_estimation.html"},{"title":"Open3D: Global registration","url":"https://www.open3d.org/docs/release/tutorial/pipelines/global_registration.html"}],"as_of":"","related_ids":["point-cloud-registration","surface-normal-estimation","iterative-closest-point","random-sample-consensus","point-cloud-library","open3d"],"name":"Fast Point Feature Histograms","alt":"FPFH 点云特征","abbr":"FPFH","aliases":["FPFH","FPFH Descriptor"],"one_liner":"A hand-crafted feature that describes each point’s local shape using a histogram of normal-vector geometric relationships in its neighborhood.","explanation":"FPFH was proposed by Radu Rusu, Nico Blodow, and Michael Beetz at ICRA 2009, as a faster version of PFH (Point Feature Histograms). The process estimates a normal vector for every point, then computes several angular relationships between it and its neighboring points, summarized into a histogram as the descriptor. PFH computes relationships between every pair of points in a neighborhood, with complexity O(nk²); FPFH only computes relationships between the center point and each neighbor, then combines neighbors’ results weighted by distance, bringing complexity down to O(nk) with little loss of discriminative power. The default implementations in both PCL and Open3D are 33-dimensional. It requires no training, and its most typical use is global point cloud registration: FPFH first finds correspondences between two point clouds, RANSAC solves for a rough pose, and ICP then refines it.","example":"Open3D’s global registration tutorial first voxel-downsamples two point clouds and estimates normals, computes 33-dimensional FPFH features, uses RANSAC to solve for an initial pose, and then refines it with ICP.","related":["Point Cloud Registration","Surface Normal Estimation","Iterative Closest Point","Random Sample Consensus","Point Cloud Library (PCL)","Open3D"]},{"id":"point-cloud-registration","category":"perception","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Point-set registration","url":"https://en.wikipedia.org/wiki/Point-set_registration"},{"title":"Open3D: ICP registration","url":"https://www.open3d.org/docs/release/tutorial/pipelines/icp_registration.html"}],"as_of":"","related_ids":["iterative-closest-point","fast-point-feature-histograms","random-sample-consensus","normal-distributions-transform","6d-object-pose-estimation","simultaneous-localization-and-mapping"],"name":"Point Cloud Registration","alt":"点云配准","abbr":"","aliases":["Point Cloud Alignment","Scan Matching","Point Set Registration"],"one_liner":"Finding the rotation and translation that lines up two point clouds within one shared coordinate frame.","explanation":"Point cloud registration solves for a spatial transform that aligns a source point cloud onto a target one; rigid registration allows only rotation and translation, while non-rigid registration also permits deformation. It's needed whenever a scene is scanned from different viewpoints or times and the results need to be stitched into a full model, when estimating how far a robot has moved in lidar odometry (often called scan matching in that context), or when aligning an object's model onto an observed point cloud to recover its pose. A common pipeline does coarse global registration first, using local features such as FPFH to find correspondences plus RANSAC to reject bad matches, and then fine registration with iterative closest point (ICP), which repeatedly finds the nearest points and solves for the optimal transform — though it depends on a reasonably good initial guess or it gets stuck in a local optimum. Open3D's ICP implementation also reports an overlap ratio (fitness) and inlier RMSE to judge the result's quality.","example":"A wrist camera photographs the same cup from two angles. FPFH plus RANSAC first coarsely aligns the two point clouds, then ICP refines the alignment to stitch together a more complete shape of the cup.","related":["Iterative Closest Point","Fast Point Feature Histograms","Random Sample Consensus","Normal Distributions Transform","6D Object Pose Estimation","Simultaneous Localization and Mapping"]},{"id":"iterative-closest-point","category":"perception","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Iterative closest point","url":"https://en.wikipedia.org/wiki/Iterative_closest_point"},{"title":"Open3D Tutorial: ICP registration","url":"https://www.open3d.org/docs/release/tutorial/pipelines/icp_registration.html"}],"as_of":"","related_ids":["point-cloud-registration","point-cloud","fast-point-feature-histograms","normal-distributions-transform","6d-object-pose-estimation","loam"],"name":"Iterative Closest Point","alt":"迭代最近点","abbr":"ICP","aliases":["ICP","Point-to-Point ICP","Point-to-Plane ICP"],"one_liner":"A classic registration algorithm that repeatedly finds nearest-point pairs and solves for rotation and translation to align two point clouds.","explanation":"Iterative closest point is the classic algorithm for point cloud registration (aligning two point clouds into the same coordinate frame), independently proposed by Chen and Medioni (1991) and Besl and McKay (1992). It loops through four steps: for every point in the source cloud, find its nearest point in the target cloud as a correspondence; solve for the rotation and translation that minimizes the sum of squared distances between these corresponding points; transform the source cloud accordingly; and repeat until the error stops decreasing. Common variants are point-to-point and the normal-vector-based point-to-plane, with the latter usually converging faster. ICP only performs local fine alignment and needs a roughly correct initial pose, or it easily gets stuck in a local optimum — so it is usually preceded by a global coarse registration step. It is widely used in lidar odometry, stitching together multiple scans, and refining object pose, with ready-made implementations in both PCL and Open3D.","example":"Refining an object pose in Open3D: first do global coarse registration with FPFH features plus RANSAC, then feed that result as an initial guess to point-to-plane ICP, which outputs a 4×4 transform matrix along with a fitness score measuring overlap and an inlier_rmse measuring correspondence error.","related":["Point Cloud Registration","Point Cloud","Fast Point Feature Histograms","Normal Distributions Transform","6D Object Pose Estimation","LOAM (LiDAR Odometry and Mapping)"]},{"id":"normal-distributions-transform","category":"perception","sec":7,"tier":3,"sources":[{"title":"How to use Normal Distributions Transform - PCL Tutorials","url":"https://pointclouds.org/documentation/tutorials/normal_distributions_transform.html"}],"as_of":"","related_ids":["point-cloud-registration","iterative-closest-point","lidar-slam","relocalization","point-cloud-library","lidar"],"name":"Normal Distributions Transform","alt":"NDT 配准（正态分布变换）","abbr":"NDT","aliases":["NDT","NDT Scan Matching","NDT Point Cloud Registration"],"one_liner":"A point-cloud registration method that divides space into a grid, models each cell as a normal distribution, and aligns scans against it.","explanation":"The Normal Distributions Transform (NDT) is a point-cloud registration method, first proposed by Biber and Straßer in 2003 for matching 2D laser scans and later extended to 3D. It divides a reference point cloud into a voxel grid, represents the points in each voxel as a normal distribution defined by a mean and covariance, and then optimizes the pose of the incoming point cloud so that its transformed points fall in the highest-probability regions of these distributions. Unlike Iterative Closest Point (ICP), which searches for point-to-point correspondences, NDT needs no nearest-neighbor search, which makes it more robust to a poor initial guess and to noise. It is widely used to localize a lidar sensor within a pre-built map, and is implemented in both the Point Cloud Library (PCL) and Autoware.","example":"A self-driving car registers its current lidar scan against a pre-built point-cloud map using NDT to find its position within the map.","related":["Point Cloud Registration","Iterative Closest Point","LiDAR SLAM","Relocalization","Point Cloud Library (PCL)","LiDAR"]},{"id":"chamfer-distance","category":"perception","sec":7,"tier":3,"sources":[{"title":"A Point Set Generation Network for 3D Object Reconstruction from a Single Image (arXiv 1612.00603)","url":"https://arxiv.org/abs/1612.00603"},{"title":"pytorch3d.loss.chamfer_distance 文档","url":"https://pytorch3d.readthedocs.io/en/latest/modules/loss.html"}],"as_of":"","related_ids":["point-cloud","point-cloud-completion-shape-completion","single-image-3d-reconstruction","loss-function","pointnet-pointnet-plus-plus","feed-forward-3d-reconstruction"],"name":"Chamfer Distance","alt":"倒角距离","abbr":"CD","aliases":["CD"],"one_liner":"Finds each point’s nearest neighbor in the other point cloud and averages the distances, to measure how close two shapes are.","explanation":"Chamfer distance measures how similar two point sets are: for each point in set A, it finds the nearest point in B and averages these distances (usually squared Euclidean distance); it then does the same from B to A and adds the two terms together. It doesn’t require the two sets to have the same number of points, or a one-to-one correspondence between points — it only needs nearest-neighbor search, and it’s differentiable almost everywhere, which makes it suitable as a loss function. The name comes from chamfer matching in early image matching work. Fan, Su, and Guibas used it alongside earth mover’s distance (EMD) as a loss for generating point clouds from a single image at CVPR 2017, after which it became a standard loss and evaluation metric for 3D reconstruction, point cloud completion, and shape generation; libraries like PyTorch3D provide it directly. Its drawback is insensitivity to point density, so predicted points can clump together, which is why papers often report it alongside EMD and F-score.","example":"A point cloud completion network takes in a point cloud of a chair captured from only one side and outputs a full set of points for the whole chair; during training, the chamfer distance between the output and the true complete point cloud is used as the loss — the smaller the distance, the better the completion.","related":["Point Cloud","Point Cloud Completion / Shape Completion","Single-Image 3D Reconstruction","Loss Function","PointNet / PointNet++","Feed-Forward 3D Reconstruction"]},{"id":"pointnet-pointnet-plus-plus","category":"perception","sec":7,"tier":2,"sources":[{"title":"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv 1612.00593)","url":"https://arxiv.org/abs/1612.00593"},{"title":"PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space (arXiv 1706.02413)","url":"https://arxiv.org/abs/1706.02413"},{"title":"3D Diffusion Policy (arXiv 2403.03954)","url":"https://arxiv.org/html/2403.03954"}],"as_of":"","related_ids":["point-cloud","point-cloud-encoder","farthest-point-sampling","point-cloud-segmentation","point-transformer-v3","3d-diffusion-policy"],"name":"PointNet / PointNet++","alt":"PointNet","abbr":"","aliases":["PointNet++"],"one_liner":"A pioneering network that classifies and segments point clouds directly, with PointNet++ its hierarchical upgrade.","explanation":"PointNet was proposed by Charles R. Qi, Leonidas Guibas, and colleagues at Stanford, published at CVPR 2017. Before it, point clouds were usually converted to voxels or multi-view images before processing, which lost detail. PointNet instead extracts a feature for each point with a shared multi-layer perceptron (MLP), then aggregates these into a global feature with max pooling; because max pooling is order-independent, shuffling the points doesn't change the result (permutation invariance), and the network can directly perform classification, part segmentation, and scene semantic segmentation. Its shortcoming is that it doesn't model local neighborhoods. PointNet++ (NeurIPS 2017) fixes this: it first picks center points with farthest point sampling, recursively applies PointNet within their neighborhoods, and extracts multi-scale local features layer by layer, while also adapting to uneven point density. Both remain common baselines for point cloud encoders today.","example":"3D Diffusion Policy (DP3)'s point cloud encoder follows the ‘per-point MLP plus max pooling’ idea, using only three MLP layers; the paper found that more complex encoders like PointNet and PointNet++ actually underperformed this small encoder on its manipulation tasks.","related":["Point Cloud","Point Cloud Encoder","Farthest Point Sampling","Point Cloud Segmentation","Point Transformer V3","3D Diffusion Policy"]},{"id":"point-transformer-v3","category":"perception","sec":7,"tier":3,"sources":[{"title":"Point Transformer V3: Simpler, Faster, Stronger (arXiv 2312.10035)","url":"https://arxiv.org/abs/2312.10035"},{"title":"Pointcept/PointTransformerV3 (GitHub)","url":"https://github.com/Pointcept/PointTransformerV3"}],"as_of":"2024-06","related_ids":["point-cloud-encoder","point-cloud-segmentation","pointnet-pointnet-plus-plus","transformer","backbone-network","self-attention"],"name":"Point Transformer V3","alt":"Point Transformer V3","abbr":"PTv3","aliases":["PTv3","Point Transformer V3: Simpler, Faster, Stronger"],"one_liner":"A transformer backbone for point clouds that swaps costly neighbor search for a fast serialization trick, letting it see farther and run faster.","explanation":"Point Transformer V3 (PTv3) was proposed by Xiaoyang Wu, Hengshuang Zhao, and colleagues from institutions including the University of Hong Kong and the Shanghai AI Laboratory; it was an oral paper at CVPR 2024, with code integrated into the open-source point cloud framework Pointcept. Earlier point cloud transformers had to search for each point’s neighbors and compute complex relative position encodings, which was slow and memory-hungry. PTv3 instead first arranges the unordered points into a 1D sequence using a space-filling curve (“serialization”), then computes attention within chunks of that sequence. According to the paper, it is about 3x faster and uses about 10x less memory than PTv2, expands the receptive field from 16 points to 1024 points, and achieved state-of-the-art results on more than 20 indoor and outdoor tasks at the time. It is commonly used as a point cloud encoder, and the self-supervised pretraining model Sonata also uses it as a backbone.","example":"An indoor point cloud from an RGB-D camera is fed into PTv3, which outputs a semantic category for every point — wall, floor, chair, and so on.","related":["Point Cloud Encoder","Point Cloud Segmentation","PointNet / PointNet++","Transformer","Backbone Network","Self-Attention"]},{"id":"point-cloud-segmentation","category":"perception","sec":7,"tier":3,"sources":[{"title":"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv 1612.00593)","url":"https://arxiv.org/abs/1612.00593"},{"title":"Point Transformer V3: Simpler, Faster, Stronger (arXiv 2312.10035)","url":"https://arxiv.org/abs/2312.10035"}],"as_of":"","related_ids":["point-cloud","semantic-segmentation","instance-segmentation","pointnet-pointnet-plus-plus","point-transformer-v3","grasp-pose-detection"],"name":"Point Cloud Segmentation","alt":"点云分割","abbr":"","aliases":["3D Semantic Segmentation","3D Instance Segmentation","Point Cloud Semantic Segmentation"],"one_liner":"Labeling every point in a point cloud with which object category, or which individual object, it belongs to.","explanation":"Point cloud segmentation is the 3D counterpart of 2D image segmentation: semantic segmentation assigns each point a category (table, cup, floor); instance segmentation further distinguishes separate individuals within the same category (cup 1, cup 2); and part segmentation breaks a single object down into parts such as a handle or a lid. PointNet, from Stanford in 2017, was the first network to operate directly on unordered point sets for classification and segmentation, and it was followed by backbones such as PointNet++ and the Point Transformer series. Common benchmarks include the indoor datasets ScanNet and S3DIS, and the outdoor dataset SemanticKITTI. A robot arm typically segments out the target object’s points before estimating its pose or detecting a grasp; a mobile robot uses it to separate the ground, obstacles, and traversable area.","example":"In a point cloud of a tabletop scene, the points belonging to “cup” are segmented out on their own and fed into a grasp-detection network.","related":["Point Cloud","Semantic Segmentation","Instance Segmentation","PointNet / PointNet++","Point Transformer V3","Grasp Pose Detection"]},{"id":"3d-object-detection","category":"perception","sec":7,"tier":3,"sources":[{"title":"KITTI 3D Object Detection Evaluation 2017","url":"https://www.cvlibs.net/datasets/kitti/eval_object.php?obj_benchmark=3d"},{"title":"Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild (arXiv 2207.10660)","url":"https://arxiv.org/abs/2207.10660"}],"as_of":"","related_ids":["object-detection","point-cloud","lidar","birds-eye-view","intersection-over-union","6d-object-pose-estimation"],"name":"3D Object Detection","alt":"3D目标检测","abbr":"","aliases":["3D Bounding Box Detection","3D Detection"],"one_liner":"Finds objects in 3D space and outputs their category along with an oriented 3D box.","explanation":"3D object detection extends 2D object detection into three dimensions: the input can be a lidar point cloud, an RGB-D image, or even a single color image, and the output is each object’s category plus a 3D bounding box, usually described by a center coordinate, length, width, height, and orientation angle. A 2D box only says which pixels an object occupies; robots that need to grasp, avoid, or navigate around objects need to know their position, size, and orientation in real space. Autonomous driving has been the main driver of this field — the classic KITTI benchmark has 7,481 training images with matching point clouds, evaluating cars, pedestrians, and cyclists, where a car’s 3D box needs an intersection-over-union (IoU) of 0.7 to count as correct. For indoor and general scenes there’s the Omni3D benchmark (234,000 images, 98 categories) with its companion Cube R-CNN model. 3D object detection is often contrasted with 6D pose estimation, which is more fine-grained and gives an object’s full 3D rotation.","example":"A warehouse robot runs 3D detection on a lidar point cloud to get each crate’s center position, size, and orientation, then plans where to slide a fork or place a grasp accordingly.","related":["Object Detection","Point Cloud","LiDAR","Bird’s-Eye View","Intersection over Union","6D Object Pose Estimation"]},{"id":"point-cloud-completion-shape-completion","category":"perception","sec":7,"tier":3,"sources":[{"title":"PCN: Point Completion Network (arXiv 1808.00671)","url":"https://arxiv.org/abs/1808.00671"},{"title":"Shape Completion Enabled Robotic Grasping (arXiv 1609.08546)","url":"https://arxiv.org/abs/1609.08546"}],"as_of":"","related_ids":["point-cloud","occlusion","single-image-3d-reconstruction","chamfer-distance","grasp-planning","point-cloud-encoder"],"name":"Point Cloud Completion / Shape Completion","alt":"点云补全 / 形状补全","abbr":"","aliases":["Point Cloud Completion","Shape Completion","3D Shape Completion"],"one_liner":"Given a partial point cloud, inferring and filling in the parts of an object that were occluded or never scanned.","explanation":"A camera or lidar only sees the side of an object facing it, so the resulting point cloud — a set of points with 3D coordinates — is always incomplete. Point cloud completion takes such a partial point cloud as input and produces the full shape; when the output is a point cloud this is usually called point cloud completion, and when the output is a mesh or voxel grid it is often called shape completion. PCN, from Carnegie Mellon in 2018, was an early deep network that completed shapes by operating directly on point sets, first generating a coarse shape and then progressively densifying and refining it. For robots, completion lets a system estimate the geometry of an object’s hidden back side, making grasp planning more reliable: Varley et al. used a 3D convolutional network in 2016 to complete a single-view point cloud and plan grasps from it, verified on a real robot. Evaluation commonly uses Chamfer distance, the average nearest-point distance between two point sets.","example":"A wrist camera sees only the front of a mug; a completion network infers the shape of its back side and handle, and the grasp planner uses that to choose where to grip.","related":["Point Cloud","Occlusion","Single-Image 3D Reconstruction","Chamfer Distance","Grasp Planning","Point Cloud Encoder"]},{"id":"triangle-mesh","category":"perception","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Polygon mesh","url":"https://en.wikipedia.org/wiki/Polygon_mesh"},{"title":"MuJoCo XML Reference: asset/mesh（「collision detection works with the convex hull of the mesh」）","url":"https://raw.githubusercontent.com/google-deepmind/mujoco/main/doc/XMLreference.rst"}],"as_of":"","related_ids":["point-cloud","voxel","mesh-file","collision-geometry","convex-decomposition","unified-robot-description-format"],"name":"Triangle Mesh","alt":"网格（三角网格）","abbr":"","aliases":["Mesh","Polygon Mesh"],"one_liner":"Represents a 3D object’s surface using vertices connected into triangular faces.","explanation":"A triangle mesh is the most common surface representation in 3D graphics: a set of vertex coordinates plus an index of which three vertices form each triangle, together tracing out an object’s surface, optionally with normals, colors, and texture coordinates. GPUs render in units of triangles, so most 3D assets are stored as meshes, in formats like OBJ, STL, PLY, DAE, and glTF. Compared with a point cloud, a mesh has faces and connectivity, so it can be rendered and used for collision directly. Robot description files in URDF and MJCF use meshes to describe a link’s visual appearance and collision shape; but engines like MuJoCo only use a mesh’s convex hull for collision, so concave objects like cups and bowls need convex decomposition first. The outputs of 3D reconstruction and of the SMPL human body model are also meshes.","example":"Importing a mug’s OBJ model directly into MuJoCo for collision treats it as a convex hull, which effectively “seals” the cup’s opening so a small ball can’t be dropped inside; only after decomposing it into several convex pieces with CoACD can objects actually be placed inside the cup.","related":["Point Cloud","Voxel","Mesh File (STL / OBJ / DAE / glTF)","Collision Geometry (Collider)","Convex Decomposition","Unified Robot Description Format"]},{"id":"truncated-signed-distance-function","category":"perception","sec":7,"tier":3,"sources":[{"title":"Open3D: RGBD integration (TSDF volume)","url":"https://www.open3d.org/docs/release/tutorial/pipelines/rgbd_integration.html"}],"as_of":"","related_ids":["signed-distance-field-function","voxel","triangle-mesh","euclidean-signed-distance-field","nvblox","depth-map"],"name":"Truncated Signed Distance Function","alt":"截断符号距离函数","abbr":"TSDF","aliases":["TSDF","TSDF Fusion","Truncated Signed Distance Field"],"one_liner":"Storing the signed distance to the nearest surface in a voxel grid, used to fuse multiple depth frames into a 3D model.","explanation":"TSDF is a 3D scene representation: space is divided into small cubes (voxels), and each voxel stores its distance to the nearest object surface — positive in front of the surface, negative behind it — with values far from the surface truncated to a fixed limit, keeping only information near the surface itself. Curless and Levoy proposed fusing multiple depth maps with this kind of volumetric method in 1996, and KinectFusion made it real-time on a GPU in 2011. As each new depth frame arrives, its observations are weighted and averaged into the voxel grid according to the camera pose, and noise gets smoothed out as more frames accumulate; finally, the Marching Cubes algorithm extracts the zero-distance surface as a triangle mesh. TSDF is widely used for robot mapping, obstacle avoidance, and grasping, with implementations in Open3D and NVIDIA’s nvblox.","example":"A handheld RGB-D camera is swept around a table; Open3D fuses each frame’s depth into a TSDF volume according to its pose, and finally exports a mesh model of the whole tabletop and the objects on it.","related":["Signed Distance Field / Function","Voxel","Triangle Mesh","Euclidean Signed Distance Field","nvblox","Depth Map"]},{"id":"implicit-vs-explicit-3d-representation","category":"perception","sec":7,"tier":3,"sources":[{"title":"arXiv 2003.08934: NeRF（ECCV 2020）","url":"https://arxiv.org/abs/2003.08934"},{"title":"arXiv 1901.05103: DeepSDF","url":"https://arxiv.org/abs/1901.05103"},{"title":"3D Gaussian Splatting for Real-Time Radiance Field Rendering（SIGGRAPH 2023）","url":"https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/"}],"as_of":"","related_ids":["neural-radiance-fields","3d-gaussian-splatting","signed-distance-field-function","truncated-signed-distance-function","point-cloud","voxel"],"name":"Implicit vs. Explicit 3D Representation","alt":"隐式表示 / 显式表示","abbr":"","aliases":["Implicit Representation","Explicit Representation","Neural Implicit Representation"],"one_liner":"Whether a 3D scene is stored directly as geometric elements, or as a function you query by coordinate.","explanation":"An explicit representation stores geometric elements directly: a point cloud stores points, a mesh stores vertices and triangular faces, a voxel grid stores a value per cell, and 3D Gaussian splatting (3DGS) stores a large number of colored, translucent 3D Gaussian ellipsoids. An implicit representation instead writes the scene as a function: input a spatial coordinate, output that point’s property, with the surface hidden inside some level set of the function. A signed distance function (SDF) outputs the signed distance to the surface, with the surface located where the value is 0; DeepSDF and neural radiance fields (NeRF) use a neural network to fit this kind of function. Implicit representations are continuous and compact, with resolution not limited by a grid, but extracting a surface or rendering requires repeatedly querying the network, which is slow; explicit representations render and edit fast, with accuracy limited by discretization. Robot planning and physics simulation still usually need a mesh or voxels, so an implicit reconstruction is often extracted into a mesh afterward.","example":"NeRF uses a fully connected network to map a 5D coordinate (3D position plus viewing direction) to density and color — a typical implicit representation; 3DGS directly stores a batch of 3D Gaussians and renders them by rasterization — an explicit representation, which is why it can render in real time.","related":["Neural Radiance Fields","3D Gaussian Splatting","Signed Distance Field / Function","Truncated Signed Distance Function","Point Cloud","Voxel"]},{"id":"neural-radiance-fields","category":"perception","sec":7,"tier":2,"sources":[{"title":"NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (ECCV 2020)","url":"https://arxiv.org/abs/2003.08934"},{"title":"Dex-NeRF (CoRL 2021, arXiv 2110.14217)","url":"https://arxiv.org/abs/2110.14217"},{"title":"LERF: Language Embedded Radiance Fields (ICCV 2023)","url":"https://www.lerf.io/"}],"as_of":"","related_ids":["3d-gaussian-splatting","novel-view-synthesis","distilled-feature-fields","lerf","f3rm","multilayer-perceptron"],"name":"Neural Radiance Fields","alt":"神经辐射场","abbr":"NeRF","aliases":["NeRF","Radiance Field"],"one_liner":"A neural network that memorizes the color and density of every point in a scene, letting it render any new viewpoint.","explanation":"Neural radiance fields were proposed by Ben Mildenhall and colleagues at ECCV 2020. A NeRF represents an entire scene with a multi-layer perceptron: given a spatial position (x, y, z) and a viewing direction, it outputs the volume density and color at that point, and volume rendering along each camera ray then synthesizes an image. Training needs only photos with known camera poses, since the rendering process is differentiable — comparing the rendered image against a real photo and minimizing the error is enough to optimize the network. NeRF was originally used for novel view synthesis and later applied to 3D reconstruction and robotics: Dex-NeRF uses it to recover the geometry of transparent objects that depth cameras can't measure accurately, for grasping; LERF and F3RM embed CLIP-like semantic features into a radiance field, making it possible to find objects in a 3D scene using language. Its drawback is that it needs to be optimized separately for every scene and renders slowly; 3D Gaussian splatting, from 2023, renders much faster, and many projects have since switched to it.","example":"Dex-NeRF (CoRL 2021) photographed transparent glassware from multiple angles to train a NeRF, rendered depth from it, and handed that to Dex-Net for grasp planning, reaching 90 to 100 percent grasp success on an ABB YuMi.","related":["3D Gaussian Splatting","Novel View Synthesis","Distilled Feature Fields","LERF","F3RM","Multilayer Perceptron"]},{"id":"3d-gaussian-splatting","category":"perception","sec":7,"tier":2,"sources":[{"title":"3D Gaussian Splatting for Real-Time Radiance Field Rendering (project page, Inria)","url":"https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/"},{"title":"arXiv 2308.04079: 3D Gaussian Splatting for Real-Time Radiance Field Rendering","url":"https://arxiv.org/abs/2308.04079"}],"as_of":"","related_ids":["neural-radiance-fields","novel-view-synthesis","structure-from-motion","gaussian-splatting-based-simulation","gaussian-splatting-slam","real-to-sim"],"name":"3D Gaussian Splatting","alt":"3D高斯泼溅","abbr":"3DGS","aliases":["3DGS","Gaussian Splatting","GS"],"one_liner":"Representing a scene as many colored 3D Gaussian ellipsoids that can be rendered from any new viewpoint in real time.","explanation":"3D Gaussian splatting was proposed by Bernhard Kerbl and colleagues at Inria and other institutions, published at SIGGRAPH 2023. It represents a scene as a large number of 3D Gaussian distributions — imaginable as semi-transparent colored ellipsoids — each with its own position, shape and orientation, opacity, and color that can vary with viewing angle. The method first runs structure-from-motion on a set of multi-view photos to recover camera poses and a sparse point cloud for initialization, then uses differentiable rasterization to repeatedly compare a rendered image against the real photos and optimize the Gaussians accordingly. The paper reports real-time novel-view synthesis at 1080p, no less than 30 frames per second, much faster than neural radiance fields (NeRF). In embodied AI it's often used to bring real scenes into simulation, to perform Gaussian-splatting SLAM, or to render novel views for data augmentation.","example":"Walking around a table taking photos, then using COLMAP to recover the camera poses and a sparse point cloud and training 3DGS on them, lets you view that table from any angle in real time on a computer.","related":["Neural Radiance Fields","Novel View Synthesis","Structure from Motion","Gaussian Splatting-based Simulation","Gaussian Splatting SLAM","Real-to-Sim"]},{"id":"novel-view-synthesis","category":"perception","sec":7,"tier":3,"sources":[{"title":"NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (arXiv)","url":"https://arxiv.org/abs/2003.08934"}],"as_of":"","related_ids":["neural-radiance-fields","3d-gaussian-splatting","multi-view-stereo","real-to-sim","gaussian-splatting-based-simulation","video-generation-model"],"name":"Novel View Synthesis","alt":"新视角合成","abbr":"NVS","aliases":["NVS","New View Synthesis"],"one_liner":"Generating an image of a scene from a camera viewpoint that was never actually photographed, using a handful of existing photos.","explanation":"Novel view synthesis takes several images of a scene, usually with known camera poses, and renders what the scene would look like from a new camera position. Classic approaches reconstruct geometry with multi-view stereo and then apply texture; Neural Radiance Fields (NeRF), introduced in 2020, instead represent the scene with a neural network and render it through volume rendering, which substantially improved quality. 3D Gaussian Splatting, from 2023, represents the scene explicitly as a large number of 3D Gaussians and can render in real time. More recently, video generation models have also been used to generate novel views directly. In embodied AI, novel view synthesis is used to bring real-world scenes into simulation (real-to-sim) and to augment training data for policies with extra viewpoints.","example":"Dozens of phone photos taken while walking around an object are used to train a 3D Gaussian Splatting model, which can then render the object from any new angle.","related":["Neural Radiance Fields","3D Gaussian Splatting","Multi-View Stereo","Real-to-Sim","Gaussian Splatting-based Simulation","Video Generation Model"]},{"id":"4d-reconstruction","category":"perception","sec":7,"tier":3,"sources":[{"title":"4D Gaussian Splatting for Real-Time Dynamic Scene Rendering (arXiv 2310.08528)","url":"https://arxiv.org/abs/2310.08528"},{"title":"MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion (arXiv 2410.03825)","url":"https://arxiv.org/abs/2410.03825"},{"title":"Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis (arXiv 2308.09713)","url":"https://arxiv.org/abs/2308.09713"}],"as_of":"","related_ids":["3d-gaussian-splatting","neural-radiance-fields","feed-forward-3d-reconstruction","dust3r","3d-point-tracking","4d-world-model"],"name":"4D Reconstruction","alt":"4D重建","abbr":"","aliases":["Dynamic Scene Reconstruction","4D Gaussian Splatting"],"one_liner":"Reconstructs a 3D scene that moves — the geometry plus how it changes over time.","explanation":"4D reconstruction recovers a 3D scene that changes over time from video, monocular or multi-view — the fourth dimension is time: it needs to capture not just what the scene looks like, but the shape and position of people, hands, and objects at every moment. Static-reconstruction methods like NeRF and 3D Gaussian splatting assume the scene doesn’t move, and produce ghosting artifacts on moving objects. One family of approaches optimizes per scene: Dynamic 3D Gaussians lets Gaussian points move and rotate over time, and 4D Gaussian Splatting (CVPR 2024) uses a deformation field to predict each Gaussian’s displacement at every moment, rendering 800×800 frames at 82 FPS on an RTX 3090. Another family is feed-forward — MonST3R (ICLR 2025) extends the static reconstruction model DUSt3R to dynamic video, outputting point maps directly frame by frame. In embodied AI, 4D reconstruction is used to recover the 3D motion of hand-object interaction from human videos and to build replayable digital twins.","example":"Filming someone pouring water with a phone and running 4D reconstruction recovers the 3D shape of the hand, cup, and pitcher at every frame, which can be replayed from any viewpoint or mined for the cup’s motion trajectory to have a robot imitate.","related":["3D Gaussian Splatting","Neural Radiance Fields","Feed-Forward 3D Reconstruction","DUSt3R","3D Point Tracking","4D World Model"]},{"id":"single-image-3d-reconstruction","category":"perception","sec":7,"tier":3,"sources":[{"title":"Zero-1-to-3: Zero-shot One Image to 3D Object (arXiv)","url":"https://arxiv.org/abs/2303.11328"}],"as_of":"","related_ids":["hunyuan3d","sam-3d","simulation-assets","real-to-sim","point-cloud-completion-shape-completion","diffusion-model"],"name":"Single-Image 3D Reconstruction","alt":"单图生成3D","abbr":"","aliases":["Image-to-3D","Single-View Reconstruction"],"one_liner":"Given just one photo, having a model fill in an object’s complete 3D shape and texture.","explanation":"Single-image 3D reconstruction takes one RGB photo as input and outputs a 3D mesh, point cloud, or Gaussian representation of the object or scene. A single photo only shows the front; the back and any occluded parts have to be “guessed” using priors the model has learned from large amounts of 3D data, so the dominant recent approach is to use a diffusion model — a model that generates data by progressively removing noise — to first generate several new viewpoints and then reconstruct from them, or to generate the 3D representation directly. For embodied AI, this lets an object in a real photo be quickly turned into a simulation asset (a 3D model usable inside a simulator), useful for digital twins, real-to-sim, and synthetic data; it can also help a robot estimate the unseen back side of an object to assist grasp planning.","example":"A photo of a mug on a table is turned into a textured mesh with Hunyuan3D, then imported into Isaac Sim as an object for grasping training.","related":["Hunyuan3D","SAM 3D","Simulation Assets","Real-to-Sim","Point Cloud Completion / Shape Completion","Diffusion Model"]},{"id":"hunyuan3d","category":"perception","sec":7,"tier":3,"sources":[{"title":"GitHub: Tencent-Hunyuan/Hunyuan3D-2（含版本时间线）","url":"https://github.com/Tencent-Hunyuan/Hunyuan3D-2"},{"title":"GitHub: Tencent-Hunyuan/Hunyuan3D-2.1","url":"https://github.com/Tencent-Hunyuan/Hunyuan3D-2.1"},{"title":"arXiv 2506.16504: Hunyuan3D 2.5","url":"https://arxiv.org/abs/2506.16504"}],"as_of":"2025-06","related_ids":["single-image-3d-reconstruction","simulation-assets","triangle-mesh","physically-based-rendering","convex-decomposition","hunyuanworld"],"name":"Hunyuan3D","alt":"混元3D（Hunyuan3D）","abbr":"","aliases":["Hunyuan3D 2.0","Hunyuan3D-2.1","Hunyuan3D 2.5","Tencent Hunyuan3D"],"one_liner":"Tencent’s open-source image-to-3D asset model series, generating a textured 3D mesh from a picture.","explanation":"Hunyuan3D is a series of 3D-asset generation models from Tencent’s Hunyuan team. Version 2.0 was open-sourced in January 2025, working in two steps: a flow-matching-based diffusion Transformer (Hunyuan3D-DiT) generates geometry from an input image, then Hunyuan3D-Paint generates a texture for the mesh. Version 2.1, open-sourced in June 2025, upgraded the texture to PBR materials (physically based rendering materials, including metalness and roughness); the shape model has 3.3B parameters and the texture model has 2B parameters, and the training code was released too. A technical report for version 2.5, published the same month, scaled the shape model up to as much as 10B parameters. In embodied AI, models like this can quickly turn a single photo into an object mesh to expand a simulation asset library — though the generated result still needs mass, friction, and other physical properties added, and a collision shape worked out, before it can go into physics simulation.","example":"Feeding a photo of a mug to Hunyuan3D-2.1 produces a GLB mesh with a PBR texture; after converting the format, generating a collision shape with convex decomposition, and setting the mass, it can be imported into Isaac Sim as a graspable object.","related":["Single-Image 3D Reconstruction","Simulation Assets","Triangle Mesh","Physically Based Rendering","Convex Decomposition","HunyuanWorld (Tencent)"]},{"id":"sam-3d","category":"perception","sec":7,"tier":3,"sources":[{"title":"Meta AI Blog: Introducing SAM 3D","url":"https://ai.meta.com/blog/sam-3d/"},{"title":"facebookresearch/sam-3d-objects","url":"https://github.com/facebookresearch/sam-3d-objects"},{"title":"facebookresearch/sam-3d-body","url":"https://github.com/facebookresearch/sam-3d-body"}],"as_of":"2026-06","related_ids":["segment-anything-model","sam-3","single-image-3d-reconstruction","human-mesh-recovery","simulation-assets","smpl"],"name":"SAM 3D","alt":"SAM 3D","abbr":"","aliases":["SAM 3D Objects","SAM 3D Body"],"one_liner":"Meta’s open single-image 3D reconstruction models, released as separate versions for objects and for human bodies.","explanation":"SAM 3D is a pair of single-image 3D reconstruction models Meta released on November 19, 2025, part of the “Segment Anything” (SAM) family. SAM 3D Objects takes an image and a mask of a target object and outputs that object’s complete 3D shape, texture, and position and orientation in the scene, handling real photos with occlusion and oblique viewing angles. SAM 3D Body estimates 3D human pose and body shape from a single image, outputting a mesh parameterized by Meta’s proposed MHR (Momentum Human Rig) human body model, and can be prompted with 2D keypoints or a mask. Weights, inference code, and evaluation sets have been released under the SAM License. For embodied AI, it lets a photographed object be quickly turned into a simulation asset, and lets human pose be extracted from images and video of people.","example":"A photo of a cup on a table is first segmented with SAM 3, then handed to SAM 3D Objects to generate a textured 3D mesh, which is imported into a simulator as a grasping-training asset.","related":["Segment Anything Model","SAM 3","Single-Image 3D Reconstruction","Human Mesh Recovery","Simulation Assets","SMPL"]},{"id":"6d-object-pose-estimation","category":"perception","sec":8,"tier":2,"sources":[{"title":"BOP: Benchmark for 6D Object Pose Estimation","url":"https://bop.felk.cvut.cz/home/"},{"title":"arXiv 2312.08344: FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects","url":"https://arxiv.org/abs/2312.08344"}],"as_of":"","related_ids":["pose","foundationpose","bop","category-level-pose-estimation","pose-tracking","average-distance-of-model-points"],"name":"6D Object Pose Estimation","alt":"6D位姿估计","abbr":"6DoF Pose","aliases":["6DoF Pose Estimation","Object Pose Estimation"],"one_liner":"Computing an object's 3D position and 3D orientation relative to the camera — six degrees of freedom in total.","explanation":"6D pose estimation takes an RGB or RGB-D image and solves for a rigid object's 3 translations and 3 rotations relative to the camera — usually output as a rotation matrix plus a translation vector. Methods are grouped by how much is known about the object beforehand: instance-level (the specific object and its CAD model have been seen before), category-level (only that it belongs to a class such as ‘cup’), and methods for novel objects, which further split into model-based (given a CAD model) and model-free (given only reference images). The BOP benchmark has organized evaluations for this task since 2018. NVIDIA's FoundationPose, from 2023, handles both familiar and novel objects within one framework and can also track a pose continuously over time. Grasping, assembly, and AR all rely on it: knowing where an object is and which way it faces is what lets a system compute where a gripper needs to go.","example":"A drill with a known CAD model sits on a table. A pose estimation model computes its position and orientation in the camera's frame from an RGB-D image, and combining that with the hand-eye calibration result gives the target pose the arm needs to grasp the handle.","related":["Pose","FoundationPose","BOP (Benchmark for 6D Object Pose Estimation)","Category-Level Pose Estimation","Pose Tracking","Average Distance of Model Points"]},{"id":"foundationpose","category":"perception","sec":8,"tier":2,"sources":[{"title":"FoundationPose (arXiv:2312.08344)","url":"https://arxiv.org/abs/2312.08344"},{"title":"NVlabs/FoundationPose (GitHub)","url":"https://github.com/NVlabs/FoundationPose"}],"as_of":"2024-03","related_ids":["6d-object-pose-estimation","pose-tracking","bop","megapose","sam-6d","nvidia"],"name":"FoundationPose","alt":"FoundationPose","abbr":"","aliases":["NVIDIA FoundationPose"],"one_liner":"NVIDIA's general-purpose 6D pose model that estimates and tracks the pose of new objects without retraining.","explanation":"FoundationPose is a 6D object pose (3D position plus 3D orientation) estimation and tracking model from Bowen Wen and colleagues at NVIDIA, published at CVPR 2024 as a Highlight paper. It targets objects the model has never seen during training, and needs no fine-tuning at test time: given either a CAD model or about 16 reference photos of the object, along with an RGB-D image and a detected region for the object, it outputs the pose. Its method scatters a large number of initial pose hypotheses uniformly around the object, compares a rendered image against the real one to refine each hypothesis, and then uses a ranking network to pick the best one; subsequent frames only need refinement, letting it track at roughly 32 Hz. Its training data is large-scale synthetic data generated with the help of large language models. As of March 2024, it ranked first on the BOP leaderboard for model-based pose estimation of novel objects, and an Isaac ROS version is also available.","example":"An arm needs to insert a workpiece it has never seen in training into a fixture: the workpiece's CAD model is scanned first, a segmentation model boxes the workpiece in the first frame, FoundationPose estimates its 6D pose and tracks it continuously, and the planner computes the grasp and insertion trajectory from that.","related":["6D Object Pose Estimation","Pose Tracking","BOP (Benchmark for 6D Object Pose Estimation)","MegaPose","SAM-6D","NVIDIA"]},{"id":"megapose","category":"perception","sec":8,"tier":3,"sources":[{"title":"arXiv: MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare","url":"https://arxiv.org/abs/2212.06870"}],"as_of":"","related_ids":["6d-object-pose-estimation","foundationpose","sam-6d","bop","ycb-object-and-model-set","pose-tracking"],"name":"MegaPose","alt":"MegaPose","abbr":"","aliases":["MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare"],"one_liner":"A method that estimates 6D pose for a new object given only its CAD model, with no retraining needed.","explanation":"MegaPose was proposed by Yann Labbé, Dieter Fox, Josef Sivic, and colleagues at Inria, NVIDIA, and other institutions, published at CoRL 2022. 6D pose estimation computes an object’s 3D position and orientation in the camera’s coordinate frame; earlier methods mostly needed to be trained separately for each object, so a new part meant starting over. MegaPose takes a “render and compare” approach: given the region of the image containing the object and its CAD model, it first roughly estimates a pose, renders a synthetic image of the model at that pose, compares it against the real image, and has a network predict how to correct the pose, repeating this iteratively. It is trained on large-scale, photorealistic synthetic data covering thousands of different objects; the paper argues that this object diversity is the key to generalization, letting the method work directly on hundreds of new objects with no retraining, achieving competitive results on benchmarks like YCB-Video and BOP. NVIDIA’s later FoundationPose used a similar render-and-compare refinement step and compared itself against MegaPose.","example":"A factory receives a new type of part; given only its CAD file, a detector first boxes the part, then MegaPose estimates its 6D pose for a robot arm to grasp, with no need to collect and train on data specific to this part.","related":["6D Object Pose Estimation","FoundationPose","SAM-6D","BOP (Benchmark for 6D Object Pose Estimation)","YCB Object and Model Set","Pose Tracking"]},{"id":"sam-6d","category":"perception","sec":8,"tier":3,"sources":[{"title":"arXiv 2311.15707: SAM-6D","url":"https://arxiv.org/abs/2311.15707"},{"title":"JiehongLin/SAM-6D GitHub","url":"https://github.com/JiehongLin/SAM-6D"}],"as_of":"2024-06","related_ids":["6d-object-pose-estimation","segment-anything-model","foundationpose","megapose","bop","point-cloud-registration"],"name":"SAM-6D","alt":"SAM-6D","abbr":"","aliases":["Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation"],"one_liner":"A method that uses SAM’s segmentation to estimate the 6D pose of objects it was never trained on, given only a CAD model.","explanation":"SAM-6D is a zero-shot 6D object pose estimation method proposed in 2023 by Jiehong Lin, Kui Jia, and colleagues, from institutions including the Chinese University of Hong Kong, Shenzhen, and South China University of Technology, published at CVPR 2024. 6D pose means an object’s 3D position plus its 3D orientation, essential for both grasping and assembly; “zero-shot” means that for a new object never seen during training, the method can estimate its pose given only a CAD model, with no retraining required. The method works in two steps: first, SAM generates all candidate regions, which are scored and filtered by semantics, appearance, and geometry to find the target object; then pose estimation is framed as partial-to-partial point matching between the object’s model point cloud and the observed point cloud, solved in a coarse-to-fine two-stage process. Inputs are an RGB-D image, camera intrinsics, and a CAD model; the paper reports results exceeding prior methods across all 7 core datasets of the BOP benchmark.","example":"Given a CAD model of a part, SAM-6D locates that part in an RGB-D image of a cluttered parts bin and outputs its 6D pose for a robot arm to grasp.","related":["6D Object Pose Estimation","Segment Anything Model","FoundationPose","MegaPose","BOP (Benchmark for 6D Object Pose Estimation)","Point Cloud Registration"]},{"id":"pose-tracking","category":"perception","sec":8,"tier":3,"sources":[{"title":"FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects (arXiv 2312.08344)","url":"https://arxiv.org/abs/2312.08344"},{"title":"FoundationPose 项目主页 (NVIDIA)","url":"https://nvlabs.github.io/FoundationPose/"}],"as_of":"2024-06","related_ids":["6d-object-pose-estimation","foundationpose","visual-servoing","in-hand-manipulation","occlusion","pose"],"name":"Pose Tracking","alt":"位姿跟踪","abbr":"","aliases":["6D Pose Tracking","Object Pose Tracking"],"one_liner":"Continuously estimating an object’s 3D position and orientation across a video’s successive frames.","explanation":"Pose means an object’s 3D position plus 3D orientation, 6 degrees of freedom in total, which is why it is also called 6D pose. 6D pose estimation usually computes the pose from scratch in a single frame; pose tracking instead uses the previous frame’s result and makes only a small correction in the new frame, which makes it faster and smoother, and suitable for real-time closed-loop control. NVIDIA’s FoundationPose (a CVPR 2024 Highlight paper) unifies estimation and tracking in one framework: given an RGB-D image, it needs only a CAD model or a handful of reference images for objects it has never seen, and refines the pose iteratively using a “render-and-compare” approach. Robots rely on it when grasping moving objects, adjusting an object in-hand, or doing visual servoing; if the object is occluded or moves too fast, tracking can be lost and a fresh global estimate is needed. Estimating a camera’s own pose in SLAM is also sometimes called tracking, but that is a different target.","example":"While a robot unscrews a bottle cap, FoundationPose tracks the bottle’s 6D pose every frame, and the controller adjusts the gripper position accordingly.","related":["6D Object Pose Estimation","FoundationPose","Visual Servoing","In-hand Manipulation","Occlusion","Pose"]},{"id":"category-level-pose-estimation","category":"perception","sec":8,"tier":3,"sources":[{"title":"Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation (arXiv 1901.02970, CVPR 2019)","url":"https://arxiv.org/abs/1901.02970"}],"as_of":"","related_ids":["6d-object-pose-estimation","normalized-object-coordinate-space","foundationpose","pose-tracking","object-generalization","bop"],"name":"Category-Level Pose Estimation","alt":"类别级位姿估计","abbr":"","aliases":["Category-Level 6D Pose","Category-Level 6D Object Pose and Size Estimation"],"one_liner":"Estimates 6D pose and size for an unseen object in a known category, without needing that exact object’s CAD model.","explanation":"Instance-level 6D pose estimation requires an accurate 3D model of each specific object beforehand, so it can only handle the exact objects seen during training. Category-level pose estimation relaxes this to: knowing only which category an object belongs to — cup, bowl, laptop, and so on — and estimating position, orientation, and 3D size for a new, unseen instance of that category. The difficulty is that objects in the same category can vary a lot in shape. He Wang and colleagues in Leonidas Guibas’s group at Stanford introduced NOCS (Normalized Object Coordinate Space) at CVPR 2019, which has a network map every pixel to a standardized coordinate shared across the category, then combines this with a depth map to solve for pose and size; they also released the CAMERA (synthetic) and REAL275 (real) datasets that are widely used since. This suits household tasks like “pick up any mug”; zero-shot methods like FoundationPose take a different approach, requiring a model or reference image of the new object instead.","example":"A robot sees a mug on the table it has never seen before; a network maps the mug’s pixels to the “mug” category’s standard coordinates, aligns this with the depth point cloud, and computes the mug’s position, orientation, and size, from which it plans a pose for grasping the handle.","related":["6D Object Pose Estimation","Normalized Object Coordinate Space","FoundationPose","Pose Tracking","Object Generalization","BOP (Benchmark for 6D Object Pose Estimation)"]},{"id":"normalized-object-coordinate-space","category":"perception","sec":8,"tier":3,"sources":[{"title":"Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation (arXiv)","url":"https://arxiv.org/abs/1901.02970"}],"as_of":"","related_ids":["category-level-pose-estimation","6d-object-pose-estimation","instance-segmentation","mask-r-cnn","depth-camera"],"name":"Normalized Object Coordinate Space","alt":"NOCS 归一化物体坐标空间","abbr":"NOCS","aliases":["NOCS","NOCS Map"],"one_liner":"A shared, standardized coordinate frame for all objects in a category, used to estimate the pose and size of objects never seen before.","explanation":"NOCS was proposed by He Wang and colleagues at Stanford in a CVPR 2019 paper, for category-level pose estimation — estimating the pose of a specific object instance the model has never seen, within a category it was trained on. The idea is to align and rescale every object in a category into a shared canonical space normalized to a unit cube. Built on top of Mask R-CNN, the network predicts, for every pixel, that pixel’s coordinate within this canonical space (the NOCS map); this is then aligned with the depth map using a similarity transform to recover the object’s 6D pose and 3D size in one step. It requires no CAD model of the specific object, which makes it a representative approach for category-level pose estimation.","example":"A mug never seen during training is placed on a table; the model segments it and predicts its NOCS map, then combines that with the depth map to compute the mug’s position, orientation, and size.","related":["Category-Level Pose Estimation","6D Object Pose Estimation","Instance Segmentation","Mask R-CNN","Depth Camera"]},{"id":"average-distance-of-model-points","category":"perception","sec":8,"tier":3,"sources":[{"title":"PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes (arXiv:1711.00199)","url":"https://arxiv.org/abs/1711.00199"},{"title":"BOP Challenge 2019：pose-error functions","url":"https://bop.felk.cvut.cz/challenges/bop-challenge-2019/"}],"as_of":"","related_ids":["6d-object-pose-estimation","bop","foundationpose","ycb-object-and-model-set","chamfer-distance","pose"],"name":"Average Distance of Model Points","alt":"ADD / ADD-S 位姿误差指标","abbr":"ADD / ADD-S","aliases":["ADD","ADD-S","ADI"],"one_liner":"Poses an object’s 3D model under both the true and predicted pose and averages the point-to-point distance, to score 6D pose estimation.","explanation":"ADD is the most common error metric for 6D object pose estimation, introduced by Hinterstoisser and colleagues at ACCV 2012 alongside the LINEMOD dataset. It works by taking points on an object’s 3D model, transforming them by the ground-truth pose and by the predicted pose separately, and averaging the distance between corresponding points; a common criterion counts a prediction correct if this average distance is under 10% of the model’s diameter, and reports the resulting accuracy. For symmetric objects (bowls, cans) that look the same after a rotation, matching points one to one unfairly penalizes reasonable predictions, so ADD-S is used instead: for each point, it finds the nearest point in the other point set before averaging — the BOP benchmark calls this ADI. PoseCNN (2018) swept this threshold from 0 to 10 centimeters on YCB-Video and reported the area under the accuracy-threshold curve (AUC), a format that later became common. The BOP benchmark notes that ADI can assign a low error to poses that are visibly misaligned, so it switched to three other metrics instead: VSD, MSSD, and MSPD.","example":"For a pot 20 centimeters in diameter, the 10%-of-diameter rule sets the threshold at 2 centimeters: a predicted pose whose model points deviate 1.5 centimeters on average from the ground truth counts as correct, while a 3-centimeter deviation counts as wrong.","related":["6D Object Pose Estimation","BOP (Benchmark for 6D Object Pose Estimation)","FoundationPose","YCB Object and Model Set","Chamfer Distance","Pose"]},{"id":"bop","category":"perception","sec":8,"tier":3,"sources":[{"title":"BOP: Benchmark for 6D Object Pose Estimation（官网）","url":"https://bop.felk.cvut.cz/home/"},{"title":"BOP: Benchmark for 6D Object Pose Estimation (arXiv 1808.08319, ECCV 2018)","url":"https://arxiv.org/abs/1808.08319"},{"title":"BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation (arXiv 2504.02812)","url":"https://arxiv.org/abs/2504.02812"}],"as_of":"2025-11","related_ids":["6d-object-pose-estimation","foundationpose","megapose","sam-6d","ycb-object-and-model-set","average-distance-of-model-points"],"name":"BOP (Benchmark for 6D Object Pose Estimation)","alt":"BOP 位姿估计基准","abbr":"BOP","aliases":["BOP Challenge"],"one_liner":"A public benchmark and yearly challenge for 6D object pose estimation, maintained by the Czech Technical University.","explanation":"BOP was introduced by Tomáš Hodaň and colleagues at ECCV 2018 and is maintained by the Czech Technical University in Prague. It unifies multiple object pose datasets — LM-O, T-LESS, YCB-V, and others — into a common format, provides a shared error function that handles pose ambiguity for symmetric objects, and runs an online evaluation system; it has held challenges in 2019, 2020, and 2022 through 2025, paired with workshops like R6D. The task has expanded from “seen objects” to “unseen objects”: the 2023 edition required methods to handle new objects quickly from just a CAD model, 2024 added a model-free task using only a reference video plus the BOP-H3 dataset, and 2025 added the BOP-Industrial dataset for industrial scenes. The 2024 report notes that 2D detection of unseen objects still lags detection of seen objects by about 35%, the main bottleneck. Methods including FoundationPose, MegaPose, and SAM-6D are all compared here.","example":"A newly proposed zero-shot pose estimation method is run on the BOP-Classic-Core dataset, and its results are uploaded to the BOP online evaluation system for an AR score, which is then compared against methods like FoundationPose on the leaderboard.","related":["6D Object Pose Estimation","FoundationPose","MegaPose","SAM-6D","YCB Object and Model Set","Average Distance of Model Points"]},{"id":"3d-vision-guided-robotics","category":"perception","sec":8,"tier":3,"sources":[{"title":"Wikipedia: Machine vision","url":"https://en.wikipedia.org/wiki/Machine_vision"},{"title":"Wikipedia: Bin picking","url":"https://en.wikipedia.org/wiki/Bin_picking"},{"title":"Mech-Mind Robotics 官网（3D 相机与视觉引导应用）","url":"https://www.mech-mind.com/"}],"as_of":"","related_ids":["bin-picking","machine-vision","structured-light","hand-eye-calibration","6d-object-pose-estimation","mech-mind-robotics"],"name":"3D Vision-Guided Robotics","alt":"3D 视觉引导","abbr":"","aliases":["3D Vision-Guided Robot","Vision-Guided Robotics","VGR"],"one_liner":"Uses a 3D camera to find a workpiece’s position and orientation and guides an industrial robot to pick or process it.","explanation":"3D vision-guided robotics is the industrial-automation term for this setup: a 3D camera — commonly structured light, stereo, or laser triangulation — is mounted at a robot workstation, captures a point cloud of the workpiece, and recognition plus pose estimation gives each workpiece’s 3D position and orientation; hand-eye calibration then converts this into the robot’s coordinate frame to guide the arm to pick, place, assemble, or polish. Traditional industrial robots rely on teach-and-playback, which requires the workpiece to be in exactly the same spot every time; adding 3D vision lets the robot handle parts that sit in random positions or messy stacks. Typical applications include bin picking (picking loose parts from a bin), mixed-case depalletizing, machine tending, and locating parts for assembly; Chinese vendors in this space include Mech-Mind and Percipio. It can be seen as the mature, factory-grade engineering version of the “perceive, then grasp” pipeline, though the object types and tasks are usually fixed in advance.","example":"At a parts factory, a 3D camera photographs metal parts piled randomly in a bin; software matches the point cloud against each part’s CAD model to compute its pose and a graspable point, guiding the robot arm to pick up parts one by one and place them on a machine tool — this is bin picking for machine tending.","related":["Bin Picking","Machine Vision","Structured Light","Hand-Eye Calibration","6D Object Pose Estimation","Mech-Mind Robotics"]},{"id":"grasp-pose-detection","category":"perception","sec":8,"tier":2,"sources":[{"title":"Grasp Pose Detection in Point Clouds (arXiv:1706.09911)","url":"https://arxiv.org/abs/1706.09911"},{"title":"GraspNet-1Billion 官网","url":"https://graspnet.net/"}],"as_of":"","related_ids":["grasping","6d-object-pose-estimation","anygrasp","contact-graspnet","graspnet-1billion","bin-picking"],"name":"Grasp Pose Detection","alt":"抓取位姿检测","abbr":"","aliases":["Grasp Detection","6-DoF Grasp Detection"],"one_liner":"Computing directly from an image or point cloud where and in what orientation a gripper should grasp an object.","explanation":"Grasp pose detection is the perception step of robotic grasping: given an RGB-D image or point cloud, it outputs a batch of feasible grasp poses along with scores. A planar grasp is commonly represented as an oriented rectangle in the image, giving the gripper's position, angle, and opening width; a 6-DOF grasp instead gives the gripper's 3D position, approach direction, rotation, and opening width, letting it grasp from the side or at an angle from above. Unlike 6D pose estimation, it needs no object model and doesn't need to recognize what the object even is, which suits grasping unfamiliar objects in clutter. In 2017, Andreas ten Pas and colleagues' GPD treated the problem in a detection-like way: sample candidates first, then classify each one. GraspNet-1Billion provides a benchmark with 88 objects, 190 scenes, and over 1.1 billion grasp annotations; AnyGrasp, Contact-GraspNet, and GraspGen are commonly used models, with the detected grasp then handed to motion planning to execute.","example":"A bin holds various parts piled together. A depth camera shoots a single frame as a point cloud, AnyGrasp outputs dozens of scored two-finger gripper poses, and the system picks the highest-scoring one that won't hit the bin wall, handing it to motion planning to execute.","related":["Grasping","6D Object Pose Estimation","AnyGrasp","Contact-GraspNet","GraspNet-1Billion","Bin Picking"]},{"id":"anygrasp","category":"perception","sec":8,"tier":2,"sources":[{"title":"AnyGrasp (arXiv 2212.08333, IEEE T-RO)","url":"https://arxiv.org/abs/2212.08333"},{"title":"AnyGrasp 项目页（SJTU MVIG）","url":"https://graspnet.net/anygrasp.html"},{"title":"graspnet/anygrasp_sdk","url":"https://github.com/graspnet/anygrasp_sdk"}],"as_of":"2026-07","related_ids":["grasp-pose-detection","bin-picking","graspnet-1billion","contact-graspnet","ok-robot","point-cloud"],"name":"AnyGrasp","alt":"AnyGrasp","abbr":"","aliases":["AnyGrasp SDK"],"one_liner":"A general-purpose grasp-detection model from Shanghai Jiao Tong University that outputs many usable grasp poses directly from a point cloud.","explanation":"AnyGrasp is a grasp-perception system from Cewu Lu's team (MVIG) at Shanghai Jiao Tong University, with the paper published in IEEE Transactions on Robotics (2023). Given a point cloud from a depth camera, it outputs a dense set of 7-degree-of-freedom grasp poses across the scene — gripper position, orientation, and opening width — each with a score. It's fairly robust to depth noise and can match the same grasp across consecutive frames, letting it track a moving object. The paper reports a single-arm system completing more than 900 grasps per hour. AnyGrasp is often used as a ready-made grasping module plugged into a larger system — OK-Robot, for instance, uses it to decide how to grasp. An SDK is released on GitHub, though the core library requires a machine-specific license.","example":"Shooting a single RGB-D frame of a cluttered tabletop, AnyGrasp returns hundreds of candidate grasps; the target object's segmentation mask filters out any that don't land on the target, and the highest-scoring remaining grasp is sent to the arm to execute.","related":["Grasp Pose Detection","Bin Picking","GraspNet-1Billion","Contact-GraspNet","OK-Robot","Point Cloud"]},{"id":"contact-graspnet","category":"perception","sec":8,"tier":3,"sources":[{"title":"Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes (arXiv 2103.14127)","url":"https://arxiv.org/abs/2103.14127"},{"title":"NVlabs/contact_graspnet (GitHub)","url":"https://github.com/NVlabs/contact_graspnet"}],"as_of":"2021-03","related_ids":["grasp-pose-detection","anygrasp","acronym-a-large-scale-grasp-dataset-based-on-simulation","graspnet-1billion","point-cloud","parallel-jaw-gripper"],"name":"Contact-GraspNet","alt":"Contact-GraspNet","abbr":"","aliases":["Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes"],"one_liner":"NVIDIA’s grasping network that generates 6-DoF grasps directly from a depth point cloud in cluttered scenes.","explanation":"Contact-GraspNet was proposed by Sundermeyer, Mousavian, Triebel, and Fox at NVIDIA, published at ICRA 2021. It targets two-finger parallel-jaw grippers: given a depth image plus camera intrinsics (or a point cloud directly), with an optional object segmentation mask, it outputs a batch of 6-DoF grasp poses with confidence scores end to end. Its key design is treating every observed 3D point as a possible fingertip contact point, which leaves only the gripper’s 3D orientation and opening width — 4 degrees of freedom — left to predict, greatly easing the learning problem. The model doesn’t distinguish between object categories, and it was trained on 17 million simulated grasps (grasp annotations from the ACRONYM dataset); the paper reports over 90% success grasping unseen objects in cluttered real-robot scenes. It’s commonly used as the grasping module right after segmentation or open-vocabulary detection.","example":"A table is cluttered with cups, boxes, and toys; a segmentation model first cuts out the target cup, and the depth image plus the cup’s mask are fed into Contact-GraspNet, which picks the highest-confidence grasp among the candidates that land on the cup for motion planning to execute.","related":["Grasp Pose Detection","AnyGrasp","ACRONYM: A Large-Scale Grasp Dataset Based on Simulation","GraspNet-1Billion","Point Cloud","Parallel Jaw Gripper"]},{"id":"graspgen","category":"perception","sec":8,"tier":3,"sources":[{"title":"GraspGen (arXiv 2507.13097)","url":"https://arxiv.org/abs/2507.13097"},{"title":"GraspGen 项目主页","url":"https://graspgen.github.io/"}],"as_of":"2025-07","related_ids":["grasp-pose-detection","diffusion-model","contact-graspnet","anygrasp","grasp-planning","curobo"],"name":"GraspGen","alt":"GraspGen","abbr":"","aliases":["GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training"],"one_liner":"NVIDIA’s framework that generates 6-DoF grasp poses with a diffusion model, then scores and filters them.","explanation":"GraspGen is a 6-DoF grasp-generation framework NVIDIA released publicly in July 2025 (Murali, Fox, Eppner, and colleagues). Given an object’s point cloud, it first uses a diffusion Transformer to generate a large batch of candidate grasp poses (the gripper’s 3D position and orientation), then uses a discriminator to score each grasp and filter out the poor ones; the discriminator is trained directly on grasps sampled from the generator, which the paper calls on-generator training. It was released alongside a dataset of over 53 million simulated grasps, covering the Franka gripper, the Robotiq 2F-140, suction cups, and other end effectors, aimed at the problem that learned grasping models often don’t transfer well to a new gripper or a new real scene. The paper reports state-of-the-art results on the FetchBench simulation benchmark, and it also works on noisy real-world point clouds.","example":"A segmentation model first cuts the target object out of a depth point cloud; the object’s point cloud is fed into GraspGen to get a batch of scored grasp poses, and the highest-scoring reachable one is passed to cuRobo to plan an arm trajectory to execute.","related":["Grasp Pose Detection","Diffusion Model","Contact-GraspNet","AnyGrasp","Grasp Planning","cuRobo (NVIDIA GPU-accelerated motion planning)"]},{"id":"semantic-keypoints","category":"perception","sec":8,"tier":3,"sources":[{"title":"arXiv 1903.06684: kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation","url":"https://arxiv.org/abs/1903.06684"}],"as_of":"","related_ids":["keypoint-detection","rekep","omnimanip","6d-object-pose-estimation","category-level-pose-estimation","affordance"],"name":"Semantic Keypoints","alt":"语义关键点","abbr":"","aliases":["Task Keypoints","Semantic 3D Keypoints"],"one_liner":"Representing an object with a handful of meaningful points on it — a mug’s handle, a kettle’s spout — to make planning manipulation easier.","explanation":"Semantic keypoints are a small number of 3D points on an object that carry clear meaning — for example, the center of a mug’s handle, the center of its base, or the heel of a shoe. Unlike 6D pose, this does not require an exact CAD template for every object; objects within the same category whose shapes vary a lot can still have corresponding points found on them, which makes the representation well suited to category-level generalization. kPAM, from Russ Tedrake’s group at MIT in 2019, applied this to robot manipulation: it first detects semantic 3D keypoints, then expresses the task as geometric constraints on those points (such as “hang the mug’s handle on the hook” or “press the mug’s base flat against the table”), and solves an optimization for the arm’s target pose, letting it manipulate new objects it has never seen. Recent work such as ReKep and OmniManip has vision-language models propose keypoints and constraints directly from an image, which is why the term “task keypoints” is also common.","example":"In kPAM, detecting only a few keypoints — the handle and the base — on mugs of many different, unseen shapes is enough to plan the motion for hanging each one on a mug rack.","related":["Keypoint Detection","ReKep","OmniManip","6D Object Pose Estimation","Category-Level Pose Estimation","Affordance"]},{"id":"affordance-detection","category":"perception","sec":8,"tier":3,"sources":[{"title":"AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection (arXiv:1709.07326)","url":"https://arxiv.org/abs/1709.07326"},{"title":"Learning Affordance Grounding from Exocentric Images (CVPR 2022, arXiv:2203.09905)","url":"https://arxiv.org/abs/2203.09905"}],"as_of":"","related_ids":["affordance","grasp-pose-detection","task-oriented-grasping","robopoint","vrb","intermediate-representation"],"name":"Affordance Detection","alt":"可供性检测","abbr":"","aliases":["Affordance Grounding","Affordance Map","Affordance Prediction"],"one_liner":"Finds where on an object you can grip, press, or pour from an image or point cloud, and marks it to a specific region.","explanation":"Affordance refers to the actions an object offers an agent — a handle affords gripping, a button affords pressing. Affordance detection pins this possibility down to specific pixels or points: given an RGB image, depth image, or point cloud, it outputs an action label or heatmap (an affordance map) for each region. An early landmark, AffordanceNet (Do et al., ICRA 2018), detects objects and simultaneously assigns each pixel of the object its most likely affordance label. “Affordance grounding” puts more emphasis on finding a region for a given action word — for instance, Luo and colleagues’ AGD20K dataset (CVPR 2022, over 20,000 images across 36 affordance classes) learns from third-person human-object interaction images and transfers this knowledge to object-only images. This fills the gap between “recognizing an object” and “knowing how to operate it,” and the result is often used as an intermediate representation for grasp planning or a policy. Recent work also uses vision-language models to predict actionable points directly, such as RoboPoint.","example":"Given the instruction “pour a cup of water,” affordance detection on an image of a kettle labels the handle a “grip” region and the spout a “pour” region, so the grasping module only samples grasp poses on the handle.","related":["Affordance","Grasp Pose Detection","Task-Oriented Grasping","RoboPoint","VRB","Intermediate Representation"]},{"id":"articulation-estimation","category":"perception","sec":8,"tier":3,"sources":[{"title":"Category-Level Articulated Object Pose Estimation (ANCSH, arXiv:1912.11913)","url":"https://arxiv.org/abs/1912.11913"},{"title":"Ditto: Building Digital Twins of Articulated Objects from Interaction (arXiv:2202.08227)","url":"https://arxiv.org/abs/2202.08227"}],"as_of":"","related_ids":["articulated-object","articulated-object-manipulation","partnet-mobility","interactive-perception","digital-twin","6d-object-pose-estimation"],"name":"Articulation Estimation","alt":"铰接结构估计","abbr":"","aliases":["Articulated Object Pose Estimation","Joint Parameter Estimation"],"one_liner":"Infers an object’s parts, how they’re joined, which axis each part moves around, and how far it’s currently open.","explanation":"Articulated objects — cabinet doors, drawers, laptops, scissors — are made of multiple rigid parts connected by joints. Articulation estimation infers, from an image or point cloud: how the object divides into parts, whether each joint is revolute or prismatic, the position and direction of each joint axis, and the current joint state (say, how many degrees a door is open). A robot needs this information before opening a door or pulling a drawer — pushing in the wrong direction can jam or damage the object. Notable work includes ANCSH, by Xiaolong Li, He Wang, Shuran Song, and colleagues, which estimates part poses, joint parameters, and joint states for unseen objects of a known category from a single depth point cloud; and Ditto (Zhenyu Jiang, Yuke Zhu, and colleagues, CVPR 2022), which uses two observations, before and after an interaction, to reconstruct part geometry and estimate a joint model whose result can be dropped straight into physics simulation — effectively building a digital twin of the articulated object.","example":"Facing an unfamiliar cabinet, a robot first gives the door a push, compares the point clouds before and after, and estimates that the hinge is a vertical axis at the left edge of the cabinet — then pulls the door open along that arc.","related":["Articulated Object","Articulated Object Manipulation","PartNet-Mobility","Interactive Perception","Digital Twin","6D Object Pose Estimation"]},{"id":"human-pose-estimation","category":"perception","sec":9,"tier":2,"sources":[{"title":"Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields (OpenPose, arXiv:1611.08050)","url":"https://arxiv.org/abs/1611.08050"},{"title":"COCO Keypoint Evaluation (OKS)","url":"https://raw.githubusercontent.com/cocodataset/cocodataset.github.io/master/dataset/keypoints-eval.htm"},{"title":"MediaPipe Pose Landmarker 官方文档","url":"https://developers.google.com/edge/mediapipe/solutions/vision/pose_landmarker"}],"as_of":"","related_ids":["keypoint-detection","markerless-motion-capture","motion-retargeting","smpl","hand-pose-estimation","mediapipe"],"name":"Human Pose Estimation","alt":"人体姿态估计","abbr":"HPE","aliases":["HPE","Human Keypoint Detection"],"one_liner":"Finding the positions of a person's body joints from an image or video and connecting them into a skeleton.","explanation":"Human pose estimation takes an image or video as input and outputs the positions of a person's joints — shoulders, elbows, wrists, hips, knees, ankles, and so on — as 2D pixel coordinates or 3D positions, which together form a skeleton. The COCO dataset annotates 17 keypoints per person and scores predictions with object keypoint similarity (OKS), which normalizes by body scale; Carnegie Mellon's OpenPose (CVPR 2017) first finds all the joints in an image and then assigns them to individual people, running in real time even with multiple people in frame; Google's MediaPipe can output 33 body points. For embodied AI, this is the first step in turning human motion into robot data: teleoperation reads the operator's pose in real time, or motion is extracted from human video, and motion retargeting then maps it onto the robot's joints.","example":"Stanford's HumanPlus uses just one RGB camera to estimate the operator's body and hand pose in real time, letting a custom 33-DOF humanoid mimic it synchronously — used both for teleoperation and to collect demonstration data.","related":["Keypoint Detection","Markerless Motion Capture","Motion Retargeting","SMPL","Hand Pose Estimation","MediaPipe"]},{"id":"hand-pose-estimation","category":"perception","sec":9,"tier":2,"sources":[{"title":"HaMeR: Reconstructing Hands in 3D with Transformers (arXiv:2312.05251)","url":"https://arxiv.org/abs/2312.05251"},{"title":"MANO 官网","url":"https://mano.is.tue.mpg.de/"},{"title":"Open-TeleVision (arXiv:2407.01512)","url":"https://arxiv.org/html/2407.01512"}],"as_of":"","related_ids":["mano","hamer","mediapipe","motion-retargeting","hand-object-interaction","human-pose-estimation"],"name":"Hand Pose Estimation","alt":"手部姿态估计","abbr":"","aliases":["Hand Tracking","Hand Mesh Recovery"],"one_liner":"Estimating the positions of a human hand's joints and how the fingers are bent, from an image or sensor data.","explanation":"Hand pose estimation infers the positions of a human hand's joints and its finger configuration from an image, depth data, or headset sensor data. Two outputs are common: 21 2D or 3D keypoints (the wrist plus 4 points per finger), which is what MediaPipe outputs; or a parametric hand mesh, most commonly the MANO model (proposed in 2017, with 778 vertices controlled by pose and shape parameters), with HaMeR using a large vision transformer to regress a MANO mesh directly from a single image. The difficulty is that fingers are thin and prone to self-occlusion, and are further blocked by whatever object is being held. For embodied AI, this is the first step in turning human hand motion into robot motion: teleoperation gets real-time keypoints from a headset's hand tracking, which motion retargeting then converts to drive a dexterous hand; learning from human video likewise requires the hand's trajectory to be estimated first, as a pseudo-action label.","example":"Open-TeleVision uses an Apple Vision Pro to get the operator's hand keypoints in real time, which dex-retargeting then optimizes into dexterous-hand joint angles; when the operator makes a fist, the Unitree H1's hand follows suit.","related":["MANO","HaMeR","MediaPipe","Motion Retargeting","Hand-Object Interaction","Human Pose Estimation"]},{"id":"mediapipe","category":"perception","sec":9,"tier":2,"sources":[{"title":"MediaPipe Hand Landmarker guide","url":"https://developers.google.com/edge/mediapipe/solutions/vision/hand_landmarker"},{"title":"MediaPipe Pose Landmarker guide","url":"https://developers.google.com/edge/mediapipe/solutions/vision/pose_landmarker"},{"title":"google-ai-edge/mediapipe (GitHub)","url":"https://github.com/google-ai-edge/mediapipe"}],"as_of":"2026-09","related_ids":["hand-pose-estimation","human-pose-estimation","keypoint-detection","motion-retargeting","hamer","dex-retargeting"],"name":"MediaPipe","alt":"MediaPipe（手部/人体关键点）","abbr":"","aliases":["Google MediaPipe","MediaPipe Hands","MediaPipe Pose","Hand Landmarker","Pose Landmarker","BlazePose"],"one_liner":"Google's open-source on-device toolkit that extracts hand and body keypoints in real time from an ordinary camera feed.","explanation":"MediaPipe is Google's open-source (Apache 2.0), cross-platform, on-device machine learning framework, and MediaPipe Tasks provides ready-made vision tasks within it. Two are most relevant to embodied AI: Hand Landmarker first detects the palm, then regresses 21 hand keypoints within the cropped region, outputting image coordinates, metric world coordinates in meters, and left/right handedness; Pose Landmarker, built on BlazePose, outputs 33 body keypoints. Both run in real time on a phone's CPU and support mobile, web, and Python. Because it needs nothing more than an ordinary RGB webcam, MediaPipe is commonly used for low-cost gesture-based teleoperation, retargeting human hand motion onto a dexterous hand, or labeling keypoints on human videos; its 3D accuracy is limited under occlusion and complex gestures, so 3D hand-reconstruction models like HaMeR are used instead when higher accuracy is needed. The older Solutions interface stopped being supported in March 2023.","example":"An example program from dex-retargeting uses MediaPipe to detect human hand keypoints in real time from a laptop webcam, then retargets them into joint angles for an Allegro dexterous hand, making the simulated robotic hand mimic the person's hand movements.","related":["Hand Pose Estimation","Human Pose Estimation","Keypoint Detection","Motion Retargeting","HaMeR","dex-retargeting"]},{"id":"markerless-motion-capture","category":"perception","sec":9,"tier":2,"sources":[{"title":"OpenCap: Human movement dynamics from smartphone videos (PLOS Computational Biology, 2023)","url":"https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1011462"},{"title":"GVHMR: World-Grounded Human Motion Recovery via Gravity-View Coordinates (arXiv:2409.06662)","url":"https://arxiv.org/abs/2409.06662"}],"as_of":"","related_ids":["motion-capture","optical-motion-capture","human-pose-estimation","human-mesh-recovery","smpl","motion-retargeting"],"name":"Markerless Motion Capture","alt":"无标记动捕","abbr":"","aliases":["Video-Based Motion Capture"],"one_liner":"Recovering 3D human motion straight from ordinary video, with no reflective markers worn on the body.","explanation":"Traditional optical motion capture sticks reflective markers on the body and tracks them with multiple infrared cameras, giving high precision but requiring expensive equipment that only works in a dedicated studio space. Markerless motion capture removes the markers, instead relying on human pose estimation to find joints from one or a few ordinary camera videos, then fitting the result to a 3D skeleton or a parametric model such as SMPL. Stanford's OpenCap, from 2023, films with two or more iPhones and reports an average joint-angle error of about 4.5°; Zhejiang University's GVHMR (SIGGRAPH Asia 2024) can recover human motion in world-frame coordinates from a single monocular video alone. Markerless capture still lags behind optical capture on occlusion, fast motion, and finger detail. For embodied AI, it makes it possible to turn large amounts of human video into motion references for humanoid robots.","example":"Filming a dance with a phone, then recovering it into a world-frame SMPL motion sequence with GVHMR and retargeting it onto a humanoid's joints, gives a reference motion for reinforcement-learning motion tracking.","related":["Motion Capture","Optical Motion Capture","Human Pose Estimation","Human Mesh Recovery","SMPL","Motion Retargeting"]},{"id":"smpl","category":"perception","sec":9,"tier":2,"sources":[{"title":"SMPL 官方项目页（Max Planck Institute for Intelligent Systems）","url":"https://smpl.is.tue.mpg.de/"},{"title":"Expressive Body Capture: 3D Hands, Face, and Body from a Single Image（SMPL-X，CVPR 2019）","url":"https://arxiv.org/abs/1904.05866"},{"title":"Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation（H2O）","url":"https://arxiv.org/abs/2403.04436"}],"as_of":"","related_ids":["mano","amass","motion-retargeting","human-mesh-recovery","human-pose-estimation","h2o"],"name":"SMPL","alt":"SMPL 人体模型","abbr":"SMPL","aliases":["Skinned Multi-Person Linear Model","SMPL-X","SMPL+H"],"one_liner":"A parametric 3D human body model that generates a full body mesh from a small set of shape and pose parameters.","explanation":"SMPL was published in 2015 by Michael Black’s group at the Max Planck Institute for Intelligent Systems in Germany, learned from thousands of body scans. Given shape parameters β (usually 10 numbers controlling height and build) and pose parameters θ (rotations for 23 joints plus global orientation), it generates a triangle mesh with 6,890 vertices; linear blend skinning plus corrective deformations keep the mesh looking natural as joints bend. SMPL-X, released in 2019, folds in the MANO hand model and the FLAME face model for 10,475 vertices total. SMPL has become the common format for human motion data: AMASS unifies 15 motion-capture datasets into SMPL parameters, and humanoid robots doing motion retargeting usually start from SMPL motion clips. The model is free for research use only; commercial use requires a license.","example":"H2O first optimizes SMPL’s shape parameters so the body skeleton proportions match Unitree’s H1 humanoid, then retargets about 10,000 SMPL motion clips from AMASS onto the H1 to train a whole-body motion-tracking policy.","related":["MANO","AMASS (Archive of Motion Capture as Surface Shapes)","Motion Retargeting","Human Mesh Recovery","Human Pose Estimation","H2O"]},{"id":"human-mesh-recovery","category":"perception","sec":9,"tier":3,"sources":[{"title":"arXiv 1712.06584: End-to-end Recovery of Human Shape and Pose（HMR, CVPR 2018）","url":"https://arxiv.org/abs/1712.06584"},{"title":"4D-Humans / HMR 2.0 项目主页（ICCV 2023）","url":"https://shubham-goel.github.io/4dhumans/"}],"as_of":"","related_ids":["smpl","gvhmr","hamer","human-pose-estimation","markerless-motion-capture","motion-retargeting"],"name":"Human Mesh Recovery","alt":"人体网格恢复","abbr":"HMR","aliases":["HMR","3D Human Pose and Shape Estimation","Human Mesh Reconstruction"],"one_liner":"The task of estimating a person’s complete 3D body mesh — pose plus body shape — from an image or video.","explanation":"Human mesh recovery estimates a person’s complete 3D body surface from an RGB image or video, rather than just a dozen or so joint points. The common approach regresses the parameters of a parametric body model like SMPL: pose parameters describe each joint’s rotation, shape parameters describe height and build, and plugging these into the model produces a human mesh. The name comes from Kanazawa, Black, Malik, and colleagues’ CVPR 2018 paper HMR, which regresses SMPL parameters directly from image pixels and uses an adversarial discriminator to constrain the result to look like a real human body. Later work, HMR 2.0 / 4D-Humans, switched to a ViT backbone and added cross-frame tracking, and methods like GVHMR further recover the person’s motion trajectory in world coordinates; HaMeR is the corresponding method for hands. In embodied AI, this is a key step for capturing full-body motion from human video and retargeting it onto a humanoid robot.","example":"4D-Humans can reconstruct every person’s SMPL mesh frame by frame from a multi-person video and track each person’s identity across frames, with the result handed to a motion-retargeting tool to convert into joint trajectories for a humanoid robot.","related":["SMPL","GVHMR","HaMeR","Human Pose Estimation","Markerless Motion Capture","Motion Retargeting"]},{"id":"gvhmr","category":"perception","sec":9,"tier":3,"sources":[{"title":"arXiv 2409.06662: World-Grounded Human Motion Recovery via Gravity-View Coordinates","url":"https://arxiv.org/abs/2409.06662"},{"title":"GVHMR 项目主页（ZJU3DV）","url":"https://zju3dv.github.io/gvhmr/"},{"title":"GMR: General Motion Retargeting（GitHub README）","url":"https://github.com/YanjieZe/GMR"}],"as_of":"2024-12","related_ids":["human-mesh-recovery","smpl","markerless-motion-capture-2","general-motion-retargeting","motion-retargeting","visual-odometry"],"name":"GVHMR","alt":"GVHMR","abbr":"","aliases":["GVHMR: World-Grounded Human Motion Recovery via Gravity-View Coordinates"],"one_liner":"A method that recovers a person’s 3D motion in world coordinates from monocular video.","explanation":"GVHMR is a method from Zhejiang University’s ZJU3DV team, published at SIGGRAPH Asia 2024: given a monocular video, it outputs SMPL-X human body parameters (a parametric human body model) and the person’s motion trajectory in world coordinates. The difficulty is that the camera itself is also moving — estimating pose only in the camera’s coordinate frame can’t tell which way the person is walking or whether their body is upright. GVHMR defines a “gravity-view” coordinate frame for every frame: one axis aligned with gravity, another referencing the camera’s line of sight, predicts body orientation within this frame, and then converts back to world coordinates using the camera’s relative rotation from visual odometry or a gyroscope. It predicts frame by frame in parallel, avoiding the error accumulation that autoregressive methods like WHAM suffer on long videos. In embodied AI, it is commonly used to extract human motion from internet videos, which is then retargeted onto a humanoid robot.","example":"The GMR general-purpose motion retargeting tool supports first extracting human motion from a monocular video with GVHMR, then retargeting it onto a humanoid robot.","related":["Human Mesh Recovery","SMPL","Markerless (Video-Based) Motion Capture","General Motion Retargeting","Motion Retargeting","Visual Odometry"]},{"id":"hamer","category":"perception","sec":9,"tier":3,"sources":[{"title":"arXiv 2312.05251: Reconstructing Hands in 3D with Transformers","url":"https://arxiv.org/abs/2312.05251"},{"title":"HaMeR 项目主页","url":"https://geopavlakos.github.io/hamer/"},{"title":"OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation","url":"https://arxiv.org/html/2410.11792"}],"as_of":"2024-06","related_ids":["mano","hand-pose-estimation","wilor","human-mesh-recovery","okami","human-video-data"],"name":"HaMeR","alt":"HaMeR","abbr":"","aliases":["Hand Mesh Recovery","Reconstructing Hands in 3D with Transformers"],"one_liner":"A Transformer model that reconstructs a 3D hand mesh from a single RGB image.","explanation":"HaMeR was proposed by researchers at UC Berkeley, the University of Michigan, and NYU, published at CVPR 2024. It uses a ViT-H vision Transformer as its backbone, regressing MANO hand model parameters (MANO describes a hand’s 3D mesh with a small number of pose and shape parameters) plus camera parameters from the hand region of an image. The authors combined 10 datasets with 2D or 3D hand annotations into about 2.7 million training samples, and also annotated the HInt evaluation set from videos like Ego4D, specifically to test hands in real-world footage. It is a single-frame method, but its results are reasonably smooth when applied to video too. In embodied AI, it is a commonly used tool for extracting finger poses from human video, with results retargeted to a dexterous hand or gripper as an action source for imitation learning.","example":"When OKAMI teaches a humanoid robot to manipulate objects from a human demonstration video, it reconstructs body motion with SLAHMR while estimating each hand’s pose separately with HaMeR.","related":["MANO","Hand Pose Estimation","WiLoR","Human Mesh Recovery","OKAMI","Human Video Data"]},{"id":"wilor","category":"perception","sec":9,"tier":3,"sources":[{"title":"WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild (arXiv)","url":"https://arxiv.org/abs/2409.12259"},{"title":"WiLoR 项目主页","url":"https://rolpotamias.github.io/WiLoR/"}],"as_of":"2025-03","related_ids":["mano","hand-pose-estimation","hamer","motion-retargeting","human-video-data","egocentric-video"],"name":"WiLoR","alt":"WiLoR","abbr":"","aliases":["End-to-End 3D Hand Localization and Reconstruction in-the-wild"],"one_liner":"Quickly finding every hand in an image and reconstructing each one as a 3D hand mesh.","explanation":"WiLoR is a 3D hand reconstruction method from a team at Imperial College London and Shanghai Jiao Tong University, published at CVPR 2025. It works in two steps: a real-time, fully convolutional network first detects every hand in the image, and then a Vision Transformer-based reconstruction network regresses, coarse to fine, the MANO parameters (a parametric hand model that describes hand shape and pose with a small number of parameters) and camera parameters for each hand, producing a 3D hand mesh. The authors also curated WHIM, a dataset of more than 2 million in-the-wild hand images. With no temporal module at all, it produces reasonably smooth hand tracking frame by frame from monocular video; code, models, and data are all open-sourced. In embodied AI, it can be used to extract 3D hand poses from human video and then retarget them onto a dexterous hand as imitation-learning data.","example":"A first-person video of a person folding laundry is fed into WiLoR frame by frame, producing MANO poses and 3D fingertip positions for both hands, which are then retargeted into joint targets for a dexterous hand.","related":["MANO","Hand Pose Estimation","HaMeR","Motion Retargeting","Human Video Data","Egocentric Video"]},{"id":"hand-object-interaction","category":"perception","sec":9,"tier":2,"sources":[{"title":"HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction (arXiv:2203.01577)","url":"https://arxiv.org/abs/2203.01577"},{"title":"DexYCB: A Benchmark for Capturing Hand Grasping of Objects (arXiv:2104.04631)","url":"https://arxiv.org/abs/2104.04631"}],"as_of":"","related_ids":["hand-pose-estimation","6d-object-pose-estimation","dexycb","hoi4d","arctic-a-dataset-for-dexterous-bimanual-hand-object-manipula","human-video-data"],"name":"Hand-Object Interaction","alt":"手物交互","abbr":"","aliases":["HOI (Hand-Object)"],"one_liner":"The study of how human hands contact, grasp, and manipulate objects, including estimating their joint 3D pose and contact.","explanation":"Hand-object interaction is a research direction in vision and robotics focused specifically on how human hands contact, grasp, and manipulate objects — narrower than general human-object interaction, which covers the whole body. Typical tasks include jointly estimating hand pose (usually a MANO mesh) and the object's 6D pose, inferring the contact region and grasp type, recognizing what action the hand is performing, and generating plausible grasps. Representative datasets include DexYCB (CVPR 2021, which annotates MANO hand pose and object 6D pose while a hand grasps YCB objects), HOI4D (CVPR 2022, with 2.4 million frames of egocentric RGB-D video across 16 object categories and 800 instances), ARCTIC, and OakInk. For robots, human hands are a ready-made source of dexterous manipulation demonstrations: extracting hand and object trajectories from human video can be converted into dexterous-hand actions or grasp targets, and DexYCB specifically targets scenarios where an object is handed from a human to a robot.","example":"From a first-person cooking video, HaMeR reconstructs the hand's 3D pose while FoundationPose tracks the spatula's pose, giving the position of the hand gripping the handle and the stir-fry trajectory — which is then retargeted into demonstration data for a dexterous hand.","related":["Hand Pose Estimation","6D Object Pose Estimation","DexYCB","HOI4D","ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation","Human Video Data"]},{"id":"gesture-recognition","category":"perception","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Gesture recognition","url":"https://en.wikipedia.org/wiki/Gesture_recognition"},{"title":"MediaPipe Gesture Recognizer 官方文档","url":"https://developers.google.com/edge/mediapipe/solutions/vision/gesture_recognizer"}],"as_of":"","related_ids":["hand-pose-estimation","human-robot-interaction","keypoint-detection","mediapipe","data-glove","teleoperation"],"name":"Gesture Recognition","alt":"手势识别","abbr":"","aliases":["Hand Gesture Recognition"],"one_liner":"Lets a machine recognize a gesture a person makes from images or sensor signals and understand its meaning.","explanation":"Gesture recognition is a task in computer vision and human-robot interaction: recognizing a person’s hand gesture — a fist, a thumbs-up, pointing somewhere — from an ordinary camera, a depth camera, a data glove, or EMG signals. A static gesture looks only at the hand’s shape at one moment; a dynamic gesture needs to look at a motion trajectory over time, like waving. The common approach today is to detect hand keypoints first, then classify based on those keypoints. It differs from hand pose estimation: the latter recovers the complete 3D hand shape, while gesture recognition only outputs a category. In embodied AI, gesture recognition is used for human-robot interaction — for example, using a gesture to make a robot stop, follow, or fetch a pointed-at object; teleoperation systems also map specific gestures to control commands, such as opening and closing a gripper or switching modes.","example":"Google’s MediaPipe Gesture Recognizer can recognize a closed fist, an open palm, a pointing index finger, thumbs up, thumbs down, a victory sign, and more by default, while also outputting hand keypoints, which can be used to implement “open palm makes the robot stop.”","related":["Hand Pose Estimation","Human-Robot Interaction","Keypoint Detection","MediaPipe","Data Glove","Teleoperation"]},{"id":"eye-tracking-gaze-estimation","category":"perception","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Eye tracking","url":"https://en.wikipedia.org/wiki/Eye_tracking"},{"title":"Gaze-based dual resolution deep imitation learning for high-precision dexterous robot manipulation (RA-L 2021)","url":"https://arxiv.org/abs/2102.01295"},{"title":"Project Aria Gen 1 hardware specifications","url":"https://facebookresearch.github.io/projectaria_tools/docs/tech_spec/hardware_spec"}],"as_of":"","related_ids":["egocentric-video","project-aria-glasses","human-robot-interaction","intent-understanding","imitation-learning","active-perception"],"name":"Eye Tracking / Gaze Estimation","alt":"眼动追踪 / 注视估计","abbr":"","aliases":["Gaze Tracking"],"one_liner":"Measures the movement of a person’s eyes and estimates where they are looking right now.","explanation":"Eye tracking measures the eyeball’s movement relative to the head, while gaze estimation goes a step further and computes where the line of sight lands — where the person is actually looking. The dominant method shines infrared light on the eye and uses a camera to capture both the pupil center and the corneal reflection at the same time, computing gaze direction from their relative position; another approach uses a neural network to regress gaze directly from an ordinary image. In embodied AI it has two main uses: recording a demonstrator’s gaze point while collecting demonstration data, as an extra signal for “where in the frame to pay attention,” helping a policy concentrate computation on the key region; and judging a person’s intent and the object of their attention during human-robot interaction. Head-worn devices like Meta’s Project Aria glasses have a built-in eye-tracking camera, so gaze information can be recorded alongside first-person data.","example":"Kim and colleagues (RA-L 2021) recorded an operator’s gaze point during teleoperation; the resulting trained policy used high-resolution image detail near the gaze point and lower resolution at the periphery, and completed a high-precision task like robotic needle threading.","related":["Egocentric Video","Project Aria Glasses","Human-Robot Interaction","Intent Understanding","Imitation Learning","Active Perception"]},{"id":"microphone-array","category":"perception","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Microphone array","url":"https://en.wikipedia.org/wiki/Microphone_array"},{"title":"Unitree G1 产品页（规格表：4 Microphone Array）","url":"https://www.unitree.com/g1"}],"as_of":"2026-09","related_ids":["automatic-speech-recognition","audio-visual-navigation","multimodal-perception","human-robot-interaction","multi-sensor-fusion","contact-microphone"],"name":"Microphone Array","alt":"麦克风阵列","abbr":"","aliases":["Sound Source Localization","Beamforming"],"one_liner":"Multiple microphones arranged in a fixed geometry working together to pick up sound directionally and judge where it comes from.","explanation":"A microphone array is a set of microphones arranged in a known geometry — a line, a ring, and so on — that capture and jointly process signals together. Sound reaches each microphone at a slightly different time, which can be used to estimate the direction a sound came from (sound source localization, also called direction-of-arrival or DOA estimation); delaying and weighting the different channels before combining them can boost sound from one particular direction while suppressing noise from others — this is called beamforming. A single microphone can do neither of these. Microphone arrays are widely used in phones, smart speakers, hearing aids, and as the front end for speech recognition. For robots, a microphone array lets it pick out a command clearly in a noisy environment, judge where a speaker is and turn to face them, and it is also the hardware foundation for “listen to find the target” research like audio-visual navigation. Many humanoid robots treat it as standard equipment — for instance, Unitree’s G1 spec sheet lists a 4-microphone array.","example":"A user calls the robot’s name from across the living room; the robot estimates the sound’s direction with its microphone array, turns to face the user, then uses beamforming to boost speech from that direction before sending it to speech recognition.","related":["Automatic Speech Recognition","Audio-Visual Navigation","Multimodal Perception","Human-Robot Interaction","Multi-Sensor Fusion","Contact Microphone"]},{"id":"automatic-speech-recognition","category":"perception","sec":9,"tier":3,"sources":[{"title":"Robust Speech Recognition via Large-Scale Weak Supervision (Whisper, arXiv:2212.04356)","url":"https://arxiv.org/abs/2212.04356"},{"title":"Wikipedia: Speech recognition","url":"https://en.wikipedia.org/wiki/Speech_recognition"}],"as_of":"","related_ids":["microphone-array","large-language-model","instruction-following","human-robot-interaction","native-multimodal","vision-language-action-model"],"name":"Automatic Speech Recognition","alt":"语音识别","abbr":"ASR","aliases":["ASR","Speech-to-Text","STT"],"one_liner":"Automatically converts spoken words into text — the first step in a robot understanding a spoken command.","explanation":"Automatic speech recognition converts a speech signal into text, also called speech-to-text (STT). The standard evaluation metric is word error rate (WER): the number of substitutions plus deletions plus insertions, divided by the number of reference words; for Chinese, this is usually computed per character and called character error rate (CER). The leading recent model is OpenAI’s Whisper, released in 2022, trained on 680,000 hours of weakly supervised, multilingual audio; it approaches the performance of supervised methods on several standard benchmarks without fine-tuning, and both the model and the inference code are open source. In embodied AI systems, speech recognition is usually the entry point of the interaction pipeline: a microphone array picks up audio, noise reduction cleans it up, ASR converts it to text, that text goes to a large language model or a VLA to generate a plan and actions, and finally text-to-speech (TTS) is used to reply. Some natively multimodal models take audio directly as input without converting to text first. On robots, ASR also has to cope with motor noise and far-field pickup.","example":"A user says “hand me the red cup on the table”; an ASR model like Whisper converts it to text first, which is then passed to a VLA model as a language instruction to execute.","related":["Microphone Array","Large Language Model","Instruction Following","Human-Robot Interaction","Native Multimodal","Vision-Language-Action Model"]},{"id":"state-estimation","category":"perception","sec":10,"tier":2,"sources":[{"title":"Timothy Barfoot 主页：State Estimation for Robotics（第二版 2024，含免费 PDF 与中译本信息）","url":"https://asrl.utias.utoronto.ca/~tdb/"},{"title":"Contact-Aided Invariant Extended Kalman Filtering for Legged Robot State Estimation (RSS 2018)","url":"https://arxiv.org/abs/1805.10410"}],"as_of":"","related_ids":["kalman-filter","invariant-extended-kalman-filter","leg-odometry","inertial-measurement-unit","proprioception","learned-state-estimator"],"name":"State Estimation","alt":"状态估计","abbr":"","aliases":["Robot State Estimation"],"one_liner":"Estimates a robot’s current position, orientation, and velocity from noisy sensor readings.","explanation":"State estimation infers quantities that can’t be measured directly — such as body pose, velocity, and sensor bias — from sensor measurements and a motion model. Every sensor has noise and drift, so readings must be fused according to how much each can be trusted, typically using the Kalman filter family, particle filters, or factor-graph optimization. Take a legged robot as an example: the IMU measures angular velocity and acceleration, joint encoders combined with leg kinematics give each foot’s position relative to the body, and assuming the stance foot doesn’t slip lets the system estimate body velocity. A 2018 University of Michigan paper validated this approach with a contact-aided invariant EKF on the Cassie biped. The body linear velocity that reinforcement-learning locomotion policies need comes from a state estimator; some work instead trains a neural network to learn that estimate directly.","example":"While a quadruped robot walks blind, IMU readings, joint angles, and foot contact states are fed into an extended Kalman filter every control cycle, which outputs body orientation and linear velocity for the locomotion policy to use.","related":["Kalman Filter","Invariant Extended Kalman Filter","Leg Odometry","Inertial Measurement Unit","Proprioception","Learned State Estimator"]},{"id":"multi-sensor-fusion","category":"perception","sec":10,"tier":2,"sources":[{"title":"Wikipedia: Sensor fusion","url":"https://en.wikipedia.org/wiki/Sensor_fusion"},{"title":"ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM (arXiv 2007.11898)","url":"https://arxiv.org/abs/2007.11898"}],"as_of":"","related_ids":["multimodal-perception","kalman-filter","extended-kalman-filter","tightly-coupled-vs-loosely-coupled-fusion","visual-inertial-odometry","multi-sensor-time-synchronization"],"name":"Multi-Sensor Fusion","alt":"多传感器融合","abbr":"","aliases":["Sensor Fusion","Early Fusion / Late Fusion"],"one_liner":"Combining data from a camera, lidar, IMU, and other sensors to get an estimate more accurate and stable than any one alone.","explanation":"Multi-sensor fusion combines data from a camera, lidar, an IMU (which measures acceleration and angular velocity), joint encoders, tactile sensors, and more to jointly estimate the same quantity, reducing the uncertainty below what any single sensor could give alone. Fusion is commonly grouped by the stage at which it happens: data-level (or early fusion, merging raw data directly), feature-level (extracting features from each sensor separately and merging those), and decision-level (or late fusion, where each sensor gives its own result and the results are voted or weighted together). The classic approach uses probabilistic estimators from the Kalman-filter family, though neural networks that learn how to fuse the data directly are now also common. Fusion matters because every sensor has blind spots: cameras struggle in the dark and with glare, IMUs drift, and lidar carries no color. Robot state estimation, SLAM, and self-driving perception all depend on it, with the prerequisite that the sensors have already been calibrated and time-synchronized.","example":"A quadruped estimating its own velocity: the IMU gives high-frequency acceleration and angular velocity, joint encoders combined with foot-contact detection give a leg-odometry estimate, and the two are merged with an extended Kalman filter to reduce the drift either source alone would have.","related":["Multimodal Perception","Kalman Filter","Extended Kalman Filter","Tightly-Coupled vs. Loosely-Coupled Fusion","Visual-Inertial Odometry","Multi-Sensor Time Synchronization"]},{"id":"wheel-odometry","category":"perception","sec":10,"tier":2,"sources":[{"title":"Wikipedia: Odometry","url":"https://en.wikipedia.org/wiki/Odometry"},{"title":"ros2_controllers: diff_drive_controller user documentation","url":"https://control.ros.org/rolling/doc/ros2_controllers/diff_drive_controller/doc/userdoc.html"}],"as_of":"","related_ids":["rotary-encoder","differential-drive-kinematics","visual-odometry","leg-odometry","extended-kalman-filter","slip"],"name":"Wheel Odometry","alt":"轮式里程计","abbr":"","aliases":["Wheel Encoder Odometry"],"one_liner":"Counts how much the wheels have turned to estimate how far a robot has traveled and how much it has turned.","explanation":"Wheel odometry is the most basic method of relative localization: an encoder on each drive wheel reads how far it has rotated, which multiplied by the wheel radius gives distance traveled; for a differential-drive base, dividing the difference between the left and right wheels’ distances by the wheelbase gives the body’s turning angle, and accumulating this frame by frame gives the robot’s pose relative to its starting point. It’s cheap, runs at high frequency, is unaffected by lighting, and comes built into almost every wheeled base. The drawback is that error only grows: wheel slip, uneven ground, and inaccurate wheel-radius calibration all push the estimate further off over time — drift — so wheel odometry is usually fused with an IMU, lidar, or visual odometry, for example with an extended Kalman filter, and then corrected further by a SLAM or localization algorithm. ROS 2’s diff_drive_controller computes and publishes the odom topic from exactly this kind of left/right wheel feedback.","example":"A differential-drive robot with a 0.4-meter wheelbase has its left wheel travel 1.0 meter and its right wheel travel 1.2 meters over some interval: the body center moves forward about 1.1 meters while turning left by (1.2 − 1.0)/0.4 = 0.5 radians, about 29°.","related":["Rotary Encoder","Differential Drive Kinematics","Visual Odometry","Leg Odometry","Extended Kalman Filter","Slip"]},{"id":"leg-odometry","category":"perception","sec":10,"tier":3,"sources":[{"title":"State Estimation for Legged Robots - Consistent Fusion of Leg Kinematics and IMU (Bloesch et al., RSS 2012)","url":"https://www.roboticsproceedings.org/rss08/p03.html"},{"title":"Contact-Aided Invariant Extended Kalman Filtering for Robot State Estimation (arXiv 1904.09251)","url":"https://arxiv.org/abs/1904.09251"},{"title":"MIT Cheetah-Software: PositionVelocityEstimator.h","url":"https://github.com/mit-biomimetics/Cheetah-Software/blob/master/common/include/Controllers/PositionVelocityEstimator.h"}],"as_of":"","related_ids":["state-estimation","proprioception","contact-estimation","extended-kalman-filter","invariant-extended-kalman-filter","wheel-odometry"],"name":"Leg Odometry","alt":"腿式里程计","abbr":"","aliases":["Proprioceptive Odometry","Kinematic Odometry"],"one_liner":"Uses joint encoders, the IMU, and contact detection to work out how far and in what direction a legged robot has moved.","explanation":"Leg odometry is a method for a legged robot to estimate its body position and velocity using only its own sensors (proprioception). The basic idea: while a supporting leg’s foot is planted on the ground and not moving, joint encoder readings and forward kinematics give that foot’s position relative to the body, which can be used to infer how the body itself is moving; this is then fused with the IMU’s acceleration and angular velocity, usually with a Kalman-filter-family method. ETH’s Bloesch and colleagues proposed at RSS 2012 putting the foothold position itself into the extended Kalman filter’s state, requiring no assumptions about the terrain, and validated this on a quadruped; the later invariant extended Kalman filter converges more reliably on the Cassie biped. It is unaffected by lighting or texture and runs at high frequency, making it a foundational input for motion control; but foot slip or a wrong contact judgment introduces error, and it drifts over long durations, so it is often corrected with visual or lidar odometry.","example":"MIT Cheetah 3 and Mini Cheetah’s open-source LinearKFPositionVelocityEstimator uses a linear Kalman filter to estimate body position and velocity, with the IMU as the prediction step and foot position and velocity computed from leg kinematics as the measurement.","related":["State Estimation","Proprioception","Contact Estimation","Extended Kalman Filter","Invariant Extended Kalman Filter","Wheel Odometry"]},{"id":"learned-state-estimator","category":"perception","sec":10,"tier":3,"sources":[{"title":"Concurrent Training of a Control Policy and a State Estimator for Dynamic and Robust Legged Locomotion (Ji et al., RA-L 2022, arXiv 2202.05481)","url":"https://arxiv.org/abs/2202.05481"}],"as_of":"","related_ids":["state-estimation","rl-based-locomotion-control","privileged-information","leg-odometry","dreamwaq","him"],"name":"Learned State Estimator","alt":"学习型状态估计器","abbr":"","aliases":["Concurrent Estimator Network","Estimator Network"],"one_liner":"A neural network that estimates quantities like body velocity from joint and IMU data, often trained jointly with a locomotion policy.","explanation":"A learned state estimator uses a neural network to replace or supplement a Kalman filter, regressing hard-to-measure states directly from proprioceptive data — a history of joint angles, joint velocities, and IMU readings. A landmark example is a 2022 RA-L paper from Hwangbo’s group at KAIST in Korea: it trains a control policy and an estimator network simultaneously in simulation, the policy with PPO reinforcement learning and the estimator with supervised learning against simulation ground truth, estimating body linear velocity, foot height, and contact probability, which then feed into the policy as input. This needs no predefined gait and no foot-contact sensor, and the estimator sees plenty of randomized slippery and rough terrain during training. In the paper, the quadruped robot reached a top speed of 3.75 m/s on flat ground and 3.54 m/s on a wet surface with a friction coefficient of 0.22. Many later legged reinforcement-learning locomotion projects have adopted a similar estimator network.","example":"In Ji and colleagues’ real-robot system, the estimator network first estimates body linear velocity, foot height, and contact probability from a history of proprioception, and the policy network then outputs desired joint positions based on this, completing high-speed walking over slopes, wet boards, and bumpy ground.","related":["State Estimation","RL-based Locomotion Control","Privileged Information","Leg Odometry","DreamWaQ","HIM"]},{"id":"visual-odometry","category":"perception","sec":10,"tier":2,"sources":[{"title":"Scaramuzza & Fraundorfer, Visual Odometry Part I: The First 30 Years and Fundamentals (IEEE RAM 2011)","url":"https://rpg.ifi.uzh.ch/docs/VO_Part_I_Scaramuzza.pdf"},{"title":"Wikipedia: Visual odometry","url":"https://en.wikipedia.org/wiki/Visual_odometry"}],"as_of":"","related_ids":["visual-inertial-odometry","visual-slam","simultaneous-localization-and-mapping","loop-closure-detection","feature-points","wheel-odometry"],"name":"Visual Odometry","alt":"视觉里程计","abbr":"VO","aliases":["VO"],"one_liner":"Compares consecutive camera frames to estimate, frame by frame, how far and how much a camera has turned.","explanation":"Visual odometry uses a sequence of images from one or more cameras to estimate the camera’s relative motion frame by frame, then accumulates those estimates into a trajectory. The name was coined by Nistér and colleagues in 2004, borrowed from wheel odometry, which accumulates wheel rotations to estimate displacement; unlike wheel odometry, it isn’t thrown off by wheel slip, and NASA used it on both of its Mars rovers. There are two main approaches: feature-based methods extract and match features across frames, while direct methods solve for motion from pixel brightness error. A single, monocular camera can only recover a trajectory with unknown scale; stereo cameras or adding an IMU (making it visual-inertial odometry) give true scale. Visual odometry only tracks motion locally between neighboring frames, so error accumulates into drift over time; adding loop closure and global optimization turns it into visual SLAM.","example":"As a quadruped robot walks down a corridor, its front-facing stereo camera matches feature points frame by frame, estimating how far forward and how many degrees it turned relative to the previous frame and accumulating this into a trajectory; after going all the way around a large loop back to the start, the estimated endpoint usually doesn’t line up with the actual starting point — that mismatch is drift.","related":["Visual-Inertial Odometry","Visual SLAM","Simultaneous Localization and Mapping","Loop Closure Detection","Feature Points","Wheel Odometry"]},{"id":"simultaneous-localization-and-mapping","category":"perception","sec":10,"tier":1,"sources":[{"title":"Wikipedia: Simultaneous localization and mapping","url":"https://en.wikipedia.org/wiki/Simultaneous_localization_and_mapping"},{"title":"Cartographer documentation","url":"https://google-cartographer.readthedocs.io/en/latest/"},{"title":"Durrant-Whyte & Bailey, Simultaneous Localisation and Mapping (SLAM): Part I (IEEE RAM 2006)，含 SLAM 缩写 1995 年在 ISRR 提出的说明","url":"https://people.eecs.berkeley.edu/~pabbeel/cs287-fa09/readings/Durrant-Whyte_Bailey_SLAM-tutorial-I.pdf"}],"as_of":"","related_ids":["visual-slam","lidar-slam","loop-closure-detection","visual-inertial-odometry","occupancy-grid-map","adaptive-monte-carlo-localization"],"name":"Simultaneous Localization and Mapping","alt":"同步定位与建图","abbr":"SLAM","aliases":["SLAM"],"one_liner":"A robot building a map of an unfamiliar environment while simultaneously figuring out where it is on that map.","explanation":"Simultaneous localization and mapping refers to a robot, on entering an unfamiliar environment, building a map of it while also estimating its own pose within that map at the same time. The difficulty is that the two problems depend on each other: localization needs a map, but mapping needs to know where you are. The field traces back to Randall Smith and Peter Cheeseman's 1986 work on spatial uncertainty and to Hugh Durrant-Whyte's group's research in the early 1990s; the acronym SLAM itself first appeared in a survey paper at the 1995 International Symposium on Robotics Research (ISRR). Classic solutions include the extended Kalman filter, particle filters, and graph optimization; approaches are also grouped by sensor, into lidar SLAM, visual SLAM, and so on, with open-source implementations including Google's Cartographer and ORB-SLAM3. It underlies navigation in robot vacuums, warehouse AMRs, and AR headsets.","example":"A robot vacuum entering a new home for the first time uses its lidar to scan out a floor plan while moving, simultaneously computing its own coordinates on that map in real time; subsequent cleaning runs then plan their routes using this map.","related":["Visual SLAM","LiDAR SLAM","Loop Closure Detection","Visual-Inertial Odometry","Occupancy Grid Map","Adaptive Monte Carlo Localization"]},{"id":"slam-front-end-back-end","category":"perception","sec":10,"tier":3,"sources":[{"title":"Past, Present, and Future of Simultaneous Localization And Mapping (Cadena et al., arXiv)","url":"https://arxiv.org/abs/1606.05830"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","visual-odometry","factor-graph-optimization","bundle-adjustment","loop-closure-detection","orb-slam3"],"name":"SLAM Front-end / Back-end","alt":"SLAM 前端 / 后端","abbr":"","aliases":["Front-end Odometry","Back-end Optimization"],"one_liner":"The two-part division of labor in a SLAM system: the front end estimates motion from sensor data, the back end globally corrects error.","explanation":"Simultaneous Localization and Mapping (SLAM) is usually split into two parts. The front end works directly on sensor data: it extracts and matches feature points, tracks adjacent frames, estimates the relative motion of the camera or robot, and performs data association (deciding which observations correspond to the same map point) plus loop-closure candidate detection. The back end takes the constraints the front end produces and performs a global estimate using factor graph optimization, bundle adjustment, or filtering methods such as the Kalman filter, spreading out accumulated drift and keeping the map consistent. This division was systematically laid out in a 2016 SLAM survey by Cadena and colleagues. When reading code for systems like ORB-SLAM3, VINS, or FAST-LIO, sorting out which part is the front end and which is the back end makes the rest much easier to follow.","example":"In ORB-SLAM3, extracting ORB features for frame-to-frame tracking belongs to the front end, while local and global bundle adjustment belong to the back end.","related":["Simultaneous Localization and Mapping","Visual Odometry","Factor Graph Optimization","Bundle Adjustment","Loop Closure Detection","ORB-SLAM3"]},{"id":"loop-closure-detection","category":"perception","sec":10,"tier":3,"sources":[{"title":"Wikipedia: Simultaneous localization and mapping（Loop closure 一节）","url":"https://en.wikipedia.org/wiki/Simultaneous_localization_and_mapping"},{"title":"GitHub: dorian3d/DBoW2","url":"https://github.com/dorian3d/DBoW2"},{"title":"GitHub: gisbi-kim/scancontext (Scan Context, IROS 2018)","url":"https://github.com/gisbi-kim/scancontext"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","visual-place-recognition","relocalization","factor-graph-optimization","iterative-closest-point","slam-front-end-back-end"],"name":"Loop Closure Detection","alt":"回环检测","abbr":"","aliases":["Loop Closure","Place Recognition"],"one_liner":"Lets a robot recognize that it has returned to a place it has already been, used to remove SLAM’s accumulated drift.","explanation":"Loop closure detection is a step in SLAM (simultaneous localization and mapping): deciding whether the scene the robot currently sees is a place it has already visited. Every step of odometry carries a small error, and the farther the robot travels the more this drift accumulates, so by the time it loops back to the starting point, the map no longer lines up. Once a return to a familiar place is confirmed, a constraint is added between the two poses and handed to the back-end pose-graph or factor-graph optimization, pulling the whole trajectory and map back into consistency. Visual SLAM commonly compares image features using a bag-of-words model, such as the DBoW2 library used in the ORB-SLAM series; lidar SLAM commonly uses a global point-cloud descriptor like Scan Context to find candidates, then aligns them precisely with ICP. A false detection is costly — one wrong loop closure can warp the entire map — so a geometric consistency check is usually added as well.","example":"DBoW2 converts an image’s ORB or BRIEF features into a bag-of-words vector; according to its README, processing BRIEF features takes about 3 milliseconds per image, and it can search a database of tens of thousands of images for likely loop-closure candidates.","related":["Simultaneous Localization and Mapping","Visual Place Recognition","Relocalization","Factor Graph Optimization","Iterative Closest Point","SLAM Front-end / Back-end"]},{"id":"visual-place-recognition","category":"perception","sec":10,"tier":3,"sources":[{"title":"Where is your place, Visual Place Recognition? (Garg et al., IJCAI 2021)","url":"https://arxiv.org/abs/2103.06443"},{"title":"NetVLAD: CNN architecture for weakly supervised place recognition (CVPR 2016)","url":"https://arxiv.org/abs/1511.07247"}],"as_of":"","related_ids":["loop-closure-detection","relocalization","visual-slam","topological-map","feature-matching","navigation"],"name":"Visual Place Recognition","alt":"视觉位置识别","abbr":"VPR","aliases":["VPR","Place Recognition"],"one_liner":"Looking at an image and determining which place it is, and whether the system has been there before.","explanation":"Visual place recognition (VPR) takes a newly captured image and finds which place it was taken, by searching a pre-built database of images tagged with location — fundamentally an image retrieval problem. The difficulty is that the same place can look very different depending on lighting, season, time of day, and viewing angle. A common approach aggregates an image’s local features into a single global descriptor vector for similarity comparison; NetVLAD (CVPR 2016) uses a trainable aggregation layer, weakly supervised with Google Street View images from different years, and is a widely used baseline. In robotics, VPR mainly serves loop closure detection in SLAM and relocalization after tracking is lost (re-determining where the robot is on the map), and is also used in navigation based on topological maps.","example":"A robot vacuum is picked up and set down in a different room; it compares the current view against the keyframes stored during mapping one by one, finds the closest match, and figures out which room it’s in.","related":["Loop Closure Detection","Relocalization","Visual SLAM","Topological Map","Feature Matching","Navigation"]},{"id":"relocalization","category":"perception","sec":10,"tier":3,"sources":[{"title":"ORB-SLAM: a Versatile and Accurate Monocular SLAM System (arXiv 1502.00956)","url":"https://arxiv.org/abs/1502.00956"},{"title":"PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization (arXiv 1505.07427)","url":"https://arxiv.org/abs/1505.07427"},{"title":"UZ-SLAMLab/ORB_SLAM3 (GitHub)","url":"https://github.com/UZ-SLAMLab/ORB_SLAM3"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","loop-closure-detection","visual-place-recognition","perspective-n-point","random-sample-consensus","adaptive-monte-carlo-localization"],"name":"Relocalization","alt":"重定位","abbr":"","aliases":["Camera Relocalization","Re-Localization"],"one_liner":"Figuring out where a robot is and which way it’s facing on an existing map, after tracking is lost or it restarts.","explanation":"Relocalization is a step within SLAM (Simultaneous Localization and Mapping) and visual localization: when tracking is lost due to fast motion, occlusion, or a lighting change, or when a robot restarts or is picked up and moved elsewhere (the so-called “kidnapped robot” problem), the system must recover its pose on an already-built map using only its current observation. A classic approach, used in ORB-SLAM, retrieves the keyframe most similar to the current image with a bag-of-words model (DBoW2), performs feature matching, and then solves for the camera pose using PnP combined with RANSAC. Among learned approaches, PoseNet, from Cambridge in 2015, was the first to regress a 6-degree-of-freedom camera pose directly from a single image using a convolutional network. Relocalization uses techniques similar to loop closure detection, but with a different purpose: loop closure corrects accumulated drift, while relocalization recovers a lost position.","example":"A robot vacuum is picked up and set down in a different room; it compares the current view against its stored map, relocalizes, and resumes cleaning.","related":["Simultaneous Localization and Mapping","Loop Closure Detection","Visual Place Recognition","Perspective-n-Point","Random Sample Consensus","Adaptive Monte Carlo Localization"]},{"id":"visual-slam","category":"perception","sec":10,"tier":2,"sources":[{"title":"MathWorks: What Is SLAM (Simultaneous Localization and Mapping)?","url":"https://www.mathworks.com/discovery/slam.html"},{"title":"ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM (arXiv 2007.11898)","url":"https://arxiv.org/abs/2007.11898"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","visual-odometry","visual-inertial-odometry","loop-closure-detection","orb-slam3","absolute-trajectory-error-relative-pose-error"],"name":"Visual SLAM","alt":"视觉SLAM","abbr":"vSLAM","aliases":["vSLAM","VSLAM"],"one_liner":"Uses only cameras to simultaneously estimate where it is and build a map of the surroundings.","explanation":"Visual SLAM is the branch of simultaneous localization and mapping (SLAM) that uses a camera as the main sensor — monocular, stereo, or RGB-D — often fused with an IMU (inertial measurement unit). Systems are usually split into a front end and a back end: the front end extracts feature points from images and matches them across frames to estimate camera motion (this part is also called visual odometry); the back end uses graph optimization or bundle adjustment to reduce accumulated error, and loop closure — recognizing that the camera has returned to a place it has already been — removes drift. Cameras are cheap and information-rich, but a monocular camera can’t recover absolute scale, and weak texture, reflections, and fast motion make tracking easy to lose. Well-known open-source systems include ORB-SLAM3, which supports monocular, stereo, RGB-D, and visual-inertial modes, and the learning-based DROID-SLAM; visual SLAM is widely used for localization and navigation in mobile robots, AR glasses, and drones.","example":"The ORB-SLAM3 paper reports that, running on the EuRoC drone dataset with stereo cameras plus an IMU, the average error of the estimated trajectory is about 3.6 centimeters; the input is an image sequence, and the output is the camera pose for every frame plus a sparse map point cloud.","related":["Simultaneous Localization and Mapping","Visual Odometry","Visual-Inertial Odometry","Loop Closure Detection","ORB-SLAM3","Absolute Trajectory Error / Relative Pose Error"]},{"id":"orb-slam3","category":"perception","sec":10,"tier":2,"sources":[{"title":"ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM (IEEE T-RO 2021)","url":"https://arxiv.org/abs/2007.11898"},{"title":"GitHub: UZ-SLAMLab/ORB_SLAM3","url":"https://github.com/UZ-SLAMLab/ORB_SLAM3"}],"as_of":"2021-12","related_ids":["simultaneous-localization-and-mapping","visual-slam","visual-inertial-odometry","feature-points","loop-closure-detection","bundle-adjustment"],"name":"ORB-SLAM3","alt":"ORB-SLAM3","abbr":"","aliases":["ORB_SLAM3","ORB-SLAM Family"],"one_liner":"A classic open-source visual SLAM system from the University of Zaragoza, supporting monocular, stereo, RGB-D, and IMU input.","explanation":"ORB-SLAM3 is an open-source SLAM (simultaneous localization and mapping) library from Juan D. Tardós and José M. M. Montiel's team at the University of Zaragoza in Spain, with the paper published in IEEE Transactions on Robotics in 2021. It uses ORB feature points (a corner feature that's fast to compute) for tracking and mapping, supports monocular, stereo, and RGB-D cameras with either pinhole or fisheye lenses, and can be tightly coupled with an IMU. Its multi-map system, called Atlas, lets it start a new map whenever tracking is lost and automatically merges maps when the robot returns to a previously mapped area. The paper reports an average error of 3.6 cm for stereo-plus-IMU on the EuRoC drone dataset. The code is open-sourced under GPLv3 and is often used as a visual SLAM baseline and teaching tool; it builds a sparse feature-point map that isn't directly usable for obstacle avoidance, and like other feature-based methods it can still lose tracking in low-texture scenes or under drastic lighting changes.","example":"Running ORB-SLAM3's stereo-plus-IMU or RGB-D mode on a mobile robot equipped with an IMU-carrying depth camera (such as a RealSense D435i) gives a real-time camera trajectory, used as the pose source for navigation or data collection.","related":["Simultaneous Localization and Mapping","Visual SLAM","Visual-Inertial Odometry","Feature Points","Loop Closure Detection","Bundle Adjustment"]},{"id":"lidar-slam","category":"perception","sec":10,"tier":2,"sources":[{"title":"LOAM: Lidar Odometry and Mapping in Real-time (RSS 2014)","url":"https://www.roboticsproceedings.org/rss10/p07.html"},{"title":"FAST-LIO2: Fast Direct LiDAR-inertial Odometry (arXiv:2107.06829)","url":"https://arxiv.org/abs/2107.06829"},{"title":"Cartographer 官方文档","url":"https://google-cartographer.readthedocs.io/en/latest/"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","lidar","lidar-inertial-odometry","fast-lio2","visual-slam","loop-closure-detection"],"name":"LiDAR SLAM","alt":"激光SLAM","abbr":"","aliases":["Lidar-Based SLAM"],"one_liner":"Using a lidar's scanned point clouds to localize a robot and build a map of the environment at the same time.","explanation":"Lidar SLAM is simultaneous localization and mapping that uses lidar as its primary sensor: as the robot moves, it aligns each new frame's point cloud with the existing map to work out its own position, while stitching the new point cloud into the map. Lidar measures distance directly and isn't affected by darkness, making it more stable than visual SLAM in large scenes and low-light places, though it lacks color and texture. 2D lidar SLAM is common in vacuums and warehouse AGVs; Google's open-source Cartographer supports both 2D and 3D. In 3D, Carnegie Mellon's 2014 LOAM splits the problem into a high-frequency odometry step and a low-frequency mapping step, and the MARS Lab at the University of Hong Kong's FAST-LIO2 further fuses in an IMU, running at up to 100 Hz and working well with narrow-field-of-view solid-state lidars too.","example":"Mounting a 3D lidar on a quadruped's back and driving it remotely around a campus loop, FAST-LIO2 builds a point-cloud map; afterward, during autonomous navigation, matching the live point cloud against this map tells the robot where it is.","related":["Simultaneous Localization and Mapping","LiDAR","LiDAR-Inertial Odometry","FAST-LIO2","Visual SLAM","Loop Closure Detection"]},{"id":"loam","category":"perception","sec":10,"tier":3,"sources":[{"title":"LOAM: Lidar Odometry and Mapping in Real-time (Zhang & Singh, RSS 2014)","url":"https://www.roboticsproceedings.org/rss10/p07.html"},{"title":"GitHub: RobustFieldAutonomyLab/LeGO-LOAM","url":"https://github.com/RobustFieldAutonomyLab/LeGO-LOAM"}],"as_of":"","related_ids":["lidar-slam","lio-sam","lidar-inertial-odometry","iterative-closest-point","loop-closure-detection","fast-lio2"],"name":"LOAM (LiDAR Odometry and Mapping)","alt":"LOAM 系列激光里程计","abbr":"","aliases":["LeGO-LOAM","A-LOAM"],"one_liner":"A classic method that splits lidar SLAM into a high-frequency odometry step and a low-frequency mapping step, plus its lightweight variants.","explanation":"LOAM is a lidar odometry and mapping method proposed by Ji Zhang and Sanjiv Singh at Carnegie Mellon University at RSS 2014. Its core idea is splitting the problem into two parts running at frequencies about an order of magnitude apart: odometry estimates the lidar’s motion roughly, at high frequency, while mapping registers the point cloud finely into the map at low frequency. It picks out edge points and planar points from each scan based on local curvature and matches only these features, keeping computation light; it achieves low drift with no IMU needed, reaching accuracy close to offline batch methods on the KITTI odometry benchmark. LeGO-LOAM, published by Tixiao Shan and Brendan Englot at IROS in 2018, lightened this for ground vehicles: it segments out ground points before extracting features, solves for the 6-DoF pose in two steps, and adds ICP-based loop closure. The later LIO-SAM then tightly coupled in the IMU.","example":"LeGO-LOAM’s original configuration targets the Clearpath Jackal ground robot: a horizontally mounted Velodyne VLP-16 lidar plus an optional IMU, outputting 6-DoF pose in real time; the README also warns that its simple ICP loop closure often fails when odometry drift gets too large.","related":["LiDAR SLAM","LIO-SAM","LiDAR-Inertial Odometry","Iterative Closest Point","Loop Closure Detection","FAST-LIO2"]},{"id":"occupancy-grid-map","category":"perception","sec":10,"tier":2,"sources":[{"title":"Wikipedia: Occupancy grid mapping","url":"https://en.wikipedia.org/wiki/Occupancy_grid_mapping"},{"title":"ROS 2 nav_msgs/OccupancyGrid 消息定义","url":"https://raw.githubusercontent.com/ros2/common_interfaces/rolling/nav_msgs/msg/OccupancyGrid.msg"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","costmap","octomap","occupancy-network","path-planning","lidar-slam"],"name":"Occupancy Grid Map","alt":"占据栅格地图","abbr":"","aliases":["Occupancy Grid","Grid Map"],"one_liner":"A map that divides the environment into cells, each storing the probability that it's blocked by an obstacle.","explanation":"The occupancy grid map was proposed by Hans Moravec and Alberto Elfes in 1985, and is one of the most common map forms in mobile robotics. It divides a plane (or a volume) into uniform cells, each storing the probability of being occupied, usually classified for use into free, occupied, and unknown. Each lidar or depth-camera frame updates the cells one by one through a binary Bayes filter, commonly implemented using log-odds to make the updates easy to accumulate. The classic algorithm assumes the robot's pose is already known, so it's commonly paired with SLAM. Path planning and obstacle avoidance can compute directly on top of it: ROS's nav_msgs/OccupancyGrid message is exactly this kind of map, with −1 marking unknown cells. A 3D version exists as octree-based maps such as OctoMap, and self-driving perception's occupancy networks instead predict 3D occupancy directly with a neural network.","example":"A robot vacuum builds a planar grid map with its lidar while moving: walls and furniture get marked occupied, floor it has already crossed gets marked free, and unexplored areas stay unknown; it then plans its cleaning route over the free cells.","related":["Simultaneous Localization and Mapping","Costmap","OctoMap","Occupancy Network","Path Planning","LiDAR SLAM"]},{"id":"particle-filter","category":"perception","sec":10,"tier":3,"sources":[{"title":"Particle filter - Wikipedia","url":"https://en.wikipedia.org/wiki/Particle_filter"},{"title":"Monte Carlo localization - Wikipedia","url":"https://en.wikipedia.org/wiki/Monte_Carlo_localization"},{"title":"Nav2 nav2_amcl README","url":"https://github.com/ros-navigation/navigation2/blob/main/nav2_amcl/README.md"}],"as_of":"","related_ids":["kalman-filter","extended-kalman-filter","adaptive-monte-carlo-localization","state-estimation","simultaneous-localization-and-mapping","importance-sampling"],"name":"Particle Filter","alt":"粒子滤波","abbr":"PF","aliases":["PF","Sequential Monte Carlo"],"one_liner":"A recursive estimation method that approximates a probability distribution over states using a large set of weighted random samples.","explanation":"The particle filter is also called Sequential Monte Carlo; its classic starting point is the 1993 bootstrap filter proposed by Gordon and colleagues. It represents uncertainty using many “particles” — each one a hypothesis about the state, such as a robot’s pose — together with a weight for each particle, and repeats three steps at every timestep: move the particles forward using a motion model and add noise (prediction), re-score the particles against a sensor observation (update), and discard low-weight particles while duplicating high-weight ones (resampling). A Kalman filter can only represent a single-peaked Gaussian distribution, but a particle filter can handle nonlinear, non-Gaussian, and multi-modal situations — for instance, a robot that is unsure which of two similar-looking corridors it is in. The most well-known robotics application is Monte Carlo Localization, proposed in 1999, and its adaptive version, AMCL.","example":"When a mobile robot starts up, it scatters particles across the whole map; after a few steps, particles that agree with the lidar scan survive, and the particle cloud gradually shrinks around the robot’s true location.","related":["Kalman Filter","Extended Kalman Filter","Adaptive Monte Carlo Localization","State Estimation","Simultaneous Localization and Mapping","Importance Sampling"]},{"id":"adaptive-monte-carlo-localization","category":"perception","sec":10,"tier":2,"sources":[{"title":"ROS amcl package.xml (ros-planning/navigation, noetic-devel)","url":"https://raw.githubusercontent.com/ros-planning/navigation/noetic-devel/amcl/package.xml"},{"title":"Nav2 nav2_amcl README","url":"https://raw.githubusercontent.com/ros-navigation/navigation2/main/nav2_amcl/README.md"},{"title":"Wikipedia: Monte Carlo localization","url":"https://en.wikipedia.org/wiki/Monte_Carlo_localization"}],"as_of":"","related_ids":["particle-filter","simultaneous-localization-and-mapping","2d-lidar","wheel-odometry","occupancy-grid-map","ros-2-navigation-stack"],"name":"Adaptive Monte Carlo Localization","alt":"AMCL 自适应蒙特卡洛定位","abbr":"AMCL","aliases":["AMCL","KLD-Sampling Monte Carlo Localization","nav2_amcl"],"one_liner":"A ROS localization module that estimates a robot's position on a known map using particle filtering and laser scans.","explanation":"AMCL is the 2D localization package in the ROS navigation stack, implementing Dieter Fox's adaptive (KLD-sampling) Monte Carlo localization; the ROS 1 version is credited to Brian Gerkey, and Nav2 ported it over largely unchanged as nav2_amcl. Monte Carlo localization itself was proposed by Frank Dellaert, Dieter Fox, Wolfram Burgard, and Sebastian Thrun in 1999, and represents the robot's possible positions with a set of particles: as the robot moves, the particles are pushed forward according to odometry plus noise, and each new laser scan gives higher weight to particles that agree with the map; resampling by weight then lets the particles gradually converge on the true position. ‘Adaptive’ refers to the particle count automatically growing or shrinking with the level of uncertainty. AMCL only handles localization — the map itself must already have been built with SLAM.","example":"After a warehouse AMR powers on, the operator roughly marks its initial position in RViz using ‘2D Pose Estimate.’ As the robot starts moving, AMCL's particle cloud gradually shrinks from a broad spread into a tight cluster, and the position locks in.","related":["Particle Filter","Simultaneous Localization and Mapping","2D LiDAR","Wheel Odometry","Occupancy Grid Map","ROS 2 Navigation Stack (Nav2)"]},{"id":"gnss-real-time-kinematic-positioning","category":"perception","sec":10,"tier":3,"sources":[{"title":"Wikipedia: Real-time kinematic positioning","url":"https://en.wikipedia.org/wiki/Real-time_kinematic_positioning"}],"as_of":"","related_ids":["inertial-measurement-unit","multi-sensor-fusion","state-estimation","lidar-slam","inspection-robot","autonomous-driving"],"name":"GNSS / Real-Time Kinematic Positioning","alt":"GNSS / RTK 定位","abbr":"GNSS / RTK","aliases":["RTK","Carrier-Phase Differential GPS","Satellite Positioning"],"one_liner":"Combines satellite navigation with base-station differential correction to achieve centimeter-level positioning outdoors.","explanation":"GNSS is the umbrella term for global navigation satellite systems, including GPS, BeiDou, GLONASS, and Galileo; used alone, positioning error is typically at the meter level. RTK (real-time kinematic positioning) sets up a base station at a known location, streams its observed satellite carrier-phase data to a mobile receiver in real time, differences the two to cancel out most shared errors, and then resolves the carrier-phase integer ambiguity, achieving centimeter-level accuracy; network RTK uses multiple base stations to widen coverage. Its limitation is that it needs a clear view of the sky — it fails indoors, between buildings, and under tree cover due to blockage and multipath reflection, so real systems are usually fused with an IMU, wheel odometry, or lidar SLAM. In embodied AI, this is mainly used on outdoor robots: inspection quadrupeds, agricultural and lawn-mowing robots, drones, and self-driving vehicles.","example":"An outdoor inspection quadruped robot carries an RTK antenna on its back, fusing it with IMU and lidar for localization, and autonomously patrols a campus following preset latitude-longitude waypoints.","related":["Inertial Measurement Unit","Multi-Sensor Fusion","State Estimation","LiDAR SLAM","Inspection Robot","Autonomous Driving"]},{"id":"ultra-wideband-positioning","category":"perception","sec":10,"tier":3,"sources":[{"title":"Ultra-wideband - Wikipedia","url":"https://en.wikipedia.org/wiki/Ultra-wideband"}],"as_of":"","related_ids":["multi-sensor-fusion","state-estimation","wheel-odometry","gnss-real-time-kinematic-positioning","autonomous-mobile-robot","kalman-filter"],"name":"Ultra-Wideband Positioning","alt":"UWB 定位（超宽带）","abbr":"UWB","aliases":["UWB","UWB Ranging"],"one_liner":"Using extremely short radio pulses to measure signal flight time, for centimeter-level indoor ranging and positioning.","explanation":"Ultra-wideband (UWB) is a low-power radio technology that occupies a very wide frequency band, defined by the US FCC as a bandwidth exceeding 500 MHz or 20% of the center frequency, whichever is smaller. Its pulses are extremely narrow, which allows precise measurement of signal flight time and makes it well suited to ranging: common approaches are two-way ranging (TWR) or time difference of arrival (TDoA), where a tagged device measures its distance to several fixed anchors, and geometry is used to solve for its coordinates, reaching centimeter-level accuracy, though this degrades with occlusion and multipath reflection. GPS doesn’t work indoors, so UWB is commonly used for warehouse robot localization, follow-me robots, and relative localization between multiple robots; phones from the iPhone 11 onward, and Apple’s AirTag, also have UWB built in.","example":"UWB anchors are mounted in the four corners of a warehouse, and a transport robot carries a UWB tag; it measures its distance to each anchor in real time, solves for its position in the warehouse, and fuses that with wheel odometry.","related":["Multi-Sensor Fusion","State Estimation","Wheel Odometry","GNSS / Real-Time Kinematic Positioning","Autonomous Mobile Robot","Kalman Filter"]},{"id":"tightly-coupled-vs-loosely-coupled-fusion","category":"perception","sec":10,"tier":3,"sources":[{"title":"VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator (arXiv 1708.03852)","url":"https://arxiv.org/abs/1708.03852"},{"title":"LIO-SAM: Tightly-coupled Lidar Inertial Odometry via Smoothing and Mapping (arXiv 2007.00258)","url":"https://arxiv.org/abs/2007.00258"}],"as_of":"","related_ids":["multi-sensor-fusion","visual-inertial-odometry","lidar-inertial-odometry","imu-preintegration","vins-mono-vins-fusion","extended-kalman-filter"],"name":"Tightly-Coupled vs. Loosely-Coupled Fusion","alt":"紧耦合 / 松耦合融合","abbr":"","aliases":["Tightly-Coupled Fusion","Loosely-Coupled Fusion"],"one_liner":"Two architectures for multi-sensor fusion: combining raw measurements together, versus computing each sensor’s own result and merging afterward.","explanation":"This distinction comes up constantly in multi-sensor fusion (such as camera+IMU or lidar+IMU). In loosely-coupled fusion, each sensor first computes its own result independently — for example, visual odometry computes a pose on its own, which is then weighted and merged with the pose from integrating the IMU inside a filter; this is simple to implement and its modules are swappable, but if one sensor fails on its own (say, tracking is lost from too little texture), its bad result feeds straight into the fusion. Tightly-coupled fusion instead puts raw measurements — such as the reprojection error of image feature points and IMU pre-integration terms — directly into one shared state estimator for joint optimization or filtering; this uses the information more fully and is usually more accurate and robust, at the cost of a more complex implementation and heavier computation. Systems such as VINS-Mono, OKVIS, FAST-LIO2, and LIO-SAM are all tightly-coupled designs.","example":"A drone flying past a blank white wall has almost no visual features: a loosely-coupled system’s visual pose estimate can go badly wrong, while a tightly-coupled system can still maintain a good estimate using the few available feature points plus the IMU constraint.","related":["Multi-Sensor Fusion","Visual-Inertial Odometry","LiDAR-Inertial Odometry","IMU Preintegration","VINS-Mono / VINS-Fusion","Extended Kalman Filter"]},{"id":"error-state-kalman-filter","category":"perception","sec":10,"tier":3,"sources":[{"title":"Joan Solà: Quaternion kinematics for the error-state Kalman filter (arXiv 1711.02508)","url":"https://arxiv.org/abs/1711.02508"},{"title":"PX4 Docs: Using PX4's Navigation Filter (EKF2)","url":"https://docs.px4.io/main/en/advanced_config/tuning_the_ecl_ekf"}],"as_of":"","related_ids":["kalman-filter","extended-kalman-filter","inertial-measurement-unit","imu-preintegration","quaternion","multi-sensor-fusion"],"name":"Error-State Kalman Filter","alt":"误差状态卡尔曼滤波","abbr":"ESKF","aliases":["ESKF","ES-EKF"],"one_liner":"A Kalman filter formulation that estimates the difference between a nominal state and the true state, rather than the state itself.","explanation":"The error-state Kalman filter is a way of writing the extended Kalman filter (EKF, which applies a Kalman filter after locally linearizing a nonlinear system), commonly used to fuse an IMU (inertial measurement unit) with a camera, lidar, or GPS. It splits the state into two parts: a “nominal state” integrated at high frequency from IMU readings, and a small “error state.” The filter only estimates the error; after each measurement update, the error is folded back into the nominal state and reset to zero. The benefit is that the error stays small, so linearization is more accurate; attitude error can be represented with a 3-dimensional rotation vector, avoiding the covariance singularity that comes from a quaternion using 4 numbers to represent 3 degrees of freedom. Joan Solà’s 2017 lecture notes are the most frequently cited reference for the derivation.","example":"PX4 flight controller’s EKF2 estimator fuses IMU, GPS, magnetometer, and other data; the official documentation states it uses an “error-state” formulation so that rotational uncertainty can be represented as a 3D vector.","related":["Kalman Filter","Extended Kalman Filter","Inertial Measurement Unit","IMU Preintegration","Quaternion","Multi-Sensor Fusion"]},{"id":"unscented-kalman-filter","category":"perception","sec":10,"tier":3,"sources":[{"title":"Unscented transform - Wikipedia","url":"https://en.wikipedia.org/wiki/Unscented_transform"},{"title":"robot_localization: State Estimation Nodes","url":"https://raw.githubusercontent.com/cra-ros-pkg/robot_localization/ros2/doc/state_estimation_nodes.rst"}],"as_of":"","related_ids":["kalman-filter","extended-kalman-filter","particle-filter","state-estimation","multi-sensor-fusion","error-state-kalman-filter"],"name":"Unscented Kalman Filter","alt":"无迹卡尔曼滤波","abbr":"UKF","aliases":["UKF","Sigma-Point Kalman Filter"],"one_liner":"A Kalman filter for nonlinear systems that propagates a small set of sample points instead of taking derivatives.","explanation":"The Unscented Kalman Filter is one nonlinear variant of the Kalman filter, proposed by Julier and Uhlmann in the mid-to-late 1990s. The Extended Kalman Filter (EKF) handles nonlinearity by linearizing around the current estimate using a Jacobian matrix, which introduces large errors when the nonlinearity is strong and requires deriving that Jacobian by hand. The UKF takes a different approach: it picks a small, deterministic set of sample points (sigma points) based on the current mean and covariance, passes each one directly through the true motion or observation model, and recomputes the mean and covariance from the transformed points — a step called the unscented transform. It needs no Jacobian, and is usually more accurate than the EKF under strong nonlinearity, at the cost of more computation. In robotics it is used to fuse IMU, wheel speed, and GPS for state estimation; ROS’s robot_localization package provides both EKF and UKF nodes.","example":"Using robot_localization’s ukf_localization_node to fuse wheel odometry with IMU data gives a mobile base a smooth, continuous pose estimate.","related":["Kalman Filter","Extended Kalman Filter","Particle Filter","State Estimation","Multi-Sensor Fusion","Error-State Kalman Filter"]},{"id":"invariant-extended-kalman-filter","category":"perception","sec":10,"tier":3,"sources":[{"title":"Contact-Aided Invariant Extended Kalman Filtering for Robot State Estimation (Hartley et al., arXiv 1904.09251)","url":"https://arxiv.org/abs/1904.09251"},{"title":"The Invariant Extended Kalman Filter as a Stable Observer (Barrau & Bonnabel, arXiv 1410.1465)","url":"https://arxiv.org/abs/1410.1465"},{"title":"GitHub: RossHartley/invariant-ekf","url":"https://github.com/RossHartley/invariant-ekf"}],"as_of":"","related_ids":["extended-kalman-filter","error-state-kalman-filter","lie-group","leg-odometry","state-estimation","inertial-measurement-unit"],"name":"Invariant Extended Kalman Filter","alt":"不变扩展卡尔曼滤波","abbr":"InEKF","aliases":["InEKF","IEKF","Contact-Aided InEKF"],"one_liner":"An extended Kalman filter that defines error on a Lie group, converging more reliably, commonly used for legged-robot state estimation.","explanation":"The invariant extended Kalman filter was systematically developed by French researchers Barrau and Bonnabel around 2014, as a variant of the extended Kalman filter (EKF). An ordinary EKF linearizes around the current estimate, so if the estimate drifts far off, the linearization becomes inaccurate and the filter can diverge. InEKF instead places states like attitude, velocity, and position on a Lie group (a mathematical structure describing rotation and translation) and defines the error using the group’s own operation; for a broad class of systems, the error’s evolution no longer depends on the current estimate, giving a larger and more stable region of convergence. A 2019 paper by Hartley, Grizzle, and colleagues at the University of Michigan applied this to fusing an IMU with leg kinematics and contact information, outperforming a quaternion-based EKF on the Cassie series of biped robots, and released a C++ implementation.","example":"The University of Michigan’s open-source invariant-ekf library (C++, depends on Eigen, BSD-3 licensed) uses the IMU as its motion model, incorporates joint kinematics and foot-contact measurements, and estimates the body’s 3D pose, velocity, and IMU bias, usable on biped or quadruped robots.","related":["Extended Kalman Filter","Error-State Kalman Filter","Lie Group","Leg Odometry","State Estimation","Inertial Measurement Unit"]},{"id":"factor-graph-optimization","category":"perception","sec":10,"tier":3,"sources":[{"title":"GTSAM: Factor Graphs and GTSAM tutorial","url":"https://gtsam.org/tutorials/intro.html"},{"title":"Dellaert & Kaess: Factor Graphs for Robot Perception (Foundations and Trends in Robotics, 2017)","url":"https://www.cs.cmu.edu/~kaess/pub/Dellaert17fnt.html"},{"title":"LIO-SAM: Tightly-coupled Lidar Inertial Odometry via Smoothing and Mapping","url":"https://arxiv.org/abs/2007.00258"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","gtsam","bundle-adjustment","imu-preintegration","loop-closure-detection","multi-sensor-fusion"],"name":"Factor Graph Optimization","alt":"因子图优化","abbr":"","aliases":["Factor Graph","Graph Optimization"],"one_liner":"Draws every measurement constraint as a “variable-factor” graph, then finds the most likely state by least squares.","explanation":"A factor graph is a bipartite graph: one type of node holds the variables being estimated, such as a robot’s pose at each moment or a landmark’s 3D coordinates; the other type of node holds factors, each representing a probabilistic constraint from a measurement or a prior — for example, the IMU gives the relative motion between two frames, or a camera gives an observation of a landmark. Finding the maximum a posteriori estimate — the most likely state given all observations — is equivalent to minimizing the sum of squared errors over all factors, which can be solved with nonlinear least-squares methods like Gauss-Newton or Levenberg-Marquardt, sped up by exploiting the graph’s sparsity. A 2017 survey by Frank Dellaert and Michael Kaess systematically lays out this framework, and Georgia Tech’s open-source GTSAM is a commonly used implementation. It is the dominant formulation for the back end of modern SLAM and multi-sensor fusion systems — adding a new sensor just means adding a new type of factor.","example":"LIO-SAM (IROS 2020) adds lidar odometry, IMU preintegration, and loop closure all as factors into the same factor graph, jointly optimizing the robot’s trajectory and using the optimization result to estimate IMU bias.","related":["Simultaneous Localization and Mapping","GTSAM (Georgia Tech Smoothing and Mapping)","Bundle Adjustment","IMU Preintegration","Loop Closure Detection","Multi-Sensor Fusion"]},{"id":"imu-preintegration","category":"perception","sec":10,"tier":3,"sources":[{"title":"arXiv 1512.02363: On-Manifold Preintegration for Real-Time Visual-Inertial Odometry（IEEE T-RO 2016）","url":"https://arxiv.org/abs/1512.02363"},{"title":"arXiv 1708.03852: VINS-Mono","url":"https://arxiv.org/abs/1708.03852"}],"as_of":"","related_ids":["inertial-measurement-unit","visual-inertial-odometry","factor-graph-optimization","vins-mono-vins-fusion","tightly-coupled-vs-loosely-coupled-fusion","lidar-inertial-odometry"],"name":"IMU Preintegration","alt":"IMU 预积分","abbr":"","aliases":["IMU Preintegration Factor","Inertial Preintegration"],"one_liner":"A technique that pre-integrates a large batch of IMU readings between two frames into a single relative-motion constraint.","explanation":"An IMU (inertial measurement unit) outputs angular velocity and acceleration at several hundred hertz, while camera or lidar keyframes come at only ten to a few tens of hertz. An optimization-based visual-inertial odometry system needs to add an IMU constraint between neighboring keyframes, but directly integrating depends on the pose and velocity at the starting point — every time optimization changes that starting point, the integration would need to be redone, which is slow. Preintegration instead integrates this stretch of IMU data, expressed in the body frame at the starting point, into relative rotation, velocity, and position increments that don’t depend on the global pose, so it only needs to be computed once; when the estimated bias (the IMU’s systematic offset) changes, a first-order approximation corrects it without recomputing from scratch. Forster and colleagues gave a complete derivation on the rotation manifold SO(3) in IEEE T-RO in 2016, and this approach has become a standard component in visual/lidar-inertial odometry systems like VINS-Mono and LIO-SAM, and in factor-graph libraries like GTSAM.","example":"VINS-Mono, in its tightly coupled nonlinear optimization, preintegrates the IMU data between neighboring keyframes into a single constraint, solved jointly with visual feature observations to get the camera trajectory.","related":["Inertial Measurement Unit","Visual-Inertial Odometry","Factor Graph Optimization","VINS-Mono / VINS-Fusion","Tightly-Coupled vs. Loosely-Coupled Fusion","LiDAR-Inertial Odometry"]},{"id":"visual-inertial-odometry","category":"perception","sec":10,"tier":3,"sources":[{"title":"Visual odometry (Wikipedia)","url":"https://en.wikipedia.org/wiki/Visual_odometry"},{"title":"VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator (arXiv)","url":"https://arxiv.org/abs/1708.03852"}],"as_of":"","related_ids":["visual-odometry","inertial-measurement-unit","vins-mono-vins-fusion","imu-preintegration","tightly-coupled-vs-loosely-coupled-fusion","visual-slam"],"name":"Visual-Inertial Odometry","alt":"视觉惯性里程计","abbr":"VIO","aliases":["VIO","Visual-Inertial Navigation"],"one_liner":"Fusing a camera and an IMU to continuously estimate a device’s own motion trajectory.","explanation":"Visual-inertial odometry (VIO) fuses camera images with IMU (inertial measurement unit) readings to continuously estimate a device’s own position and orientation; using only a camera is called visual odometry (VO). The two sensors complement each other: the IMU samples fast and keeps working even during violent motion, but drifts quickly once its readings are integrated; the camera constrains that drift, but is thrown off by motion blur and texture-less surfaces, and a monocular camera alone doesn’t know the true scale — the gravity and acceleration the IMU measures can restore that scale. Implementations split into filtering-based and optimization-based approaches (such as VINS-Mono), and separately into loosely-coupled and tightly-coupled fusion. VIO only performs local estimation, so error still accumulates over a long run; adding loop closure detection and global optimization turns it into visual-inertial SLAM. It is commonly used to localize drones, AR devices, and legged robots in environments without GPS.","example":"A drone flying inside an indoor warehouse has no GPS signal, so it relies on its onboard camera and IMU running VIO to know where it has flown in real time.","related":["Visual Odometry","Inertial Measurement Unit","VINS-Mono / VINS-Fusion","IMU Preintegration","Tightly-Coupled vs. Loosely-Coupled Fusion","Visual SLAM"]},{"id":"vins-mono-vins-fusion","category":"perception","sec":10,"tier":3,"sources":[{"title":"VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator (arXiv)","url":"https://arxiv.org/abs/1708.03852"},{"title":"HKUST-Aerial-Robotics/VINS-Mono (GitHub)","url":"https://github.com/HKUST-Aerial-Robotics/VINS-Mono"},{"title":"HKUST-Aerial-Robotics/VINS-Fusion (GitHub)","url":"https://github.com/HKUST-Aerial-Robotics/VINS-Fusion"}],"as_of":"2019-01","related_ids":["visual-inertial-odometry","visual-slam","inertial-measurement-unit","imu-preintegration","loop-closure-detection","camera-imu-calibration"],"name":"VINS-Mono / VINS-Fusion","alt":"VINS-Mono / VINS-Fusion","abbr":"","aliases":["VINS","Visual-Inertial Navigation System"],"one_liner":"An open-source visual-inertial localization system from HKUST that fuses a camera and an IMU to estimate pose in real time.","explanation":"VINS-Mono is a monocular visual-inertial state estimator open-sourced by Shaojie Shen’s team (the Aerial Robotics Group) at the Hong Kong University of Science and Technology, published in IEEE T-RO. Using just one camera plus one IMU (inertial measurement unit, which measures acceleration and angular velocity), it computes a device’s position and orientation in real time. It works by tightly coupling pre-integrated IMU data with image feature points inside a nonlinear optimization, and includes automatic initialization, online calibration of the camera-IMU extrinsics, loop closure detection, and 4-degree-of-freedom pose graph optimization. VINS-Fusion, released in 2019, is an extended version supporting monocular+IMU, stereo, and stereo+IMU configurations, and also demonstrates fusion with GPS. Both are built on ROS and open-sourced under GPLv3, and are commonly used as a baseline localization solution for drones and mobile robots.","example":"A quadruped robot dog is fitted with a stereo camera that has a built-in IMU; running VINS-Fusion lets it output its body trajectory in real time indoors, with no GPS available.","related":["Visual-Inertial Odometry","Visual SLAM","Inertial Measurement Unit","IMU Preintegration","Loop Closure Detection","Camera-IMU Calibration"]},{"id":"lidar-inertial-odometry","category":"perception","sec":10,"tier":3,"sources":[{"title":"FAST-LIO2: Fast Direct LiDAR-inertial Odometry (arXiv 2107.06829)","url":"https://arxiv.org/abs/2107.06829"},{"title":"LIO-SAM: Tightly-coupled Lidar Inertial Odometry via Smoothing and Mapping (arXiv 2007.00258)","url":"https://arxiv.org/abs/2007.00258"}],"as_of":"","related_ids":["lidar-slam","fast-lio2","lio-sam","imu-preintegration","tightly-coupled-vs-loosely-coupled-fusion","visual-inertial-odometry"],"name":"LiDAR-Inertial Odometry","alt":"激光惯性里程计","abbr":"LIO","aliases":["LIO","LiDAR-IMU Odometry"],"one_liner":"Fuses lidar point clouds with IMU readings to estimate a robot’s pose in real time while building a point-cloud map.","explanation":"LiDAR-inertial odometry fuses a lidar with an IMU (inertial measurement unit) to estimate its own motion, forming the core of lidar SLAM. The two are complementary: a single point-cloud scan takes a stretch of time to capture (commonly 10 Hz, or 0.1 second), and the robot’s motion during that time distorts the point cloud — the high-frequency IMU can estimate this motion to correct the distortion (deskewing) and provide an initial guess for registration; registering the point cloud against the map, in turn, corrects the drift that accumulates from integrating the IMU. Systems are categorized as loosely or tightly coupled by how they fuse the two, with tightly coupled now the mainstream approach. Landmark systems include Tixiao Shan and colleagues’ LIO-SAM at IROS 2020 (factor-graph optimization) and the University of Hong Kong MARS Lab’s FAST-LIO / FAST-LIO2 (iterated Kalman filter), the latter skipping feature extraction to register raw points directly, reaching up to 100 Hz and also working with solid-state lidars.","example":"Walking a loop through a building holding a lidar with a built-in IMU, FAST-LIO2 can output the trajectory and point-cloud map in real time; in the paper’s tests, pose estimation stayed reliable even while the sensor was spun rapidly at angular rates up to about 1000°/s.","related":["LiDAR SLAM","FAST-LIO2","LIO-SAM","IMU Preintegration","Tightly-Coupled vs. Loosely-Coupled Fusion","Visual-Inertial Odometry"]},{"id":"fast-lio2","category":"perception","sec":10,"tier":2,"sources":[{"title":"FAST-LIO2: Fast Direct LiDAR-inertial Odometry (arXiv:2107.06829)","url":"https://arxiv.org/abs/2107.06829"},{"title":"hku-mars/FAST_LIO (GitHub)","url":"https://github.com/hku-mars/FAST_LIO"}],"as_of":"2021-07","related_ids":["lidar-inertial-odometry","lidar","inertial-measurement-unit","lidar-slam","lio-sam","solid-state-lidar"],"name":"FAST-LIO2","alt":"FAST-LIO / FAST-LIO2","abbr":"","aliases":["FAST-LIO","Fast LiDAR-Inertial Odometry"],"one_liner":"An open-source lidar-plus-IMU odometry system from HKU that's fast and works with many lidar types.","explanation":"FAST-LIO is an open-source lidar-inertial odometry system from the MARS Lab at the University of Hong Kong (Fu Zhang's group): it fuses lidar point clouds with an IMU (which measures acceleration and angular velocity) to estimate the robot's own pose in real time while simultaneously building a point-cloud map. The first version used a tightly coupled iterated extended Kalman filter. FAST-LIO2, released in July 2021, made two key changes: it no longer hand-extracts edge and plane features but registers raw points directly against the map, so it works with both spinning lidars (Velodyne, Ouster) and solid-state ones (Livox); and it maintains its map with a custom incremental k-d tree called ikd-Tree, with the paper reporting rates up to 100 Hz in large scenes. It's a common open-source baseline for localization and mapping on quadrupeds, drones, and humanoids, is licensed under GPL-2.0, and runs on ARM boards.","example":"Mounting a Livox solid-state lidar on a quadruped's back and running FAST-LIO2 gives its self-pose and a point-cloud map in real time, which is then handed to a navigation module for path planning.","related":["LiDAR-Inertial Odometry","LiDAR","Inertial Measurement Unit","LiDAR SLAM","LIO-SAM","Solid-State LiDAR"]},{"id":"lio-sam","category":"perception","sec":10,"tier":3,"sources":[{"title":"LIO-SAM: Tightly-coupled Lidar Inertial Odometry via Smoothing and Mapping (IROS 2020, arXiv 2007.00258)","url":"https://arxiv.org/abs/2007.00258"},{"title":"GitHub: TixiaoShan/LIO-SAM","url":"https://github.com/TixiaoShan/LIO-SAM"}],"as_of":"","related_ids":["lidar-inertial-odometry","factor-graph-optimization","imu-preintegration","gtsam","loop-closure-detection","loam"],"name":"LIO-SAM","alt":"LIO-SAM","abbr":"","aliases":["LiDAR Inertial Odometry via Smoothing and Mapping"],"one_liner":"An open-source tightly coupled lidar-inertial odometry and mapping system built on factor-graph optimization.","explanation":"LIO-SAM is a lidar-inertial SLAM system by Tixiao Shan, Brendan Englot, Daniela Rus, and colleagues, published at IROS 2020; Shan and Englot had previously built LeGO-LOAM. It formulates localization and mapping as a factor graph (a graph model that treats each measurement as a constraint and jointly optimizes a sequence of poses), supporting four types of factors: IMU preintegration, lidar odometry, GPS, and loop closure. IMU preintegration both deskews the point cloud and provides an initial guess for registration. To stay real-time, it only selects keyframes and registers each new keyframe against a local map made from a fixed number of nearby historical keyframes, rather than matching against the global map. Its back end is built on the GTSAM library. The code is open source under a BSD-3 license, with a ROS 1 version and a ROS 2 branch.","example":"The README notes that LIO-SAM only works with a 9-axis IMU that can output roll, pitch, and yaw, recommended at 200 Hz or higher; when using an Ouster lidar, its built-in 6-axis IMU is not sufficient, so a separate external 9-axis IMU is needed.","related":["LiDAR-Inertial Odometry","Factor Graph Optimization","IMU Preintegration","GTSAM (Georgia Tech Smoothing and Mapping)","Loop Closure Detection","LOAM (LiDAR Odometry and Mapping)"]},{"id":"rtab-map","category":"perception","sec":10,"tier":3,"sources":[{"title":"RTAB-Map 官方主页（IntRoLab）","url":"https://introlab.github.io/rtabmap/"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","visual-slam","loop-closure-detection","occupancy-grid-map","common-ros-slam-packages","depth-camera"],"name":"RTAB-Map","alt":"RTAB-Map","abbr":"","aliases":["Real-Time Appearance-Based Mapping","rtabmap","rtabmap_ros"],"one_liner":"An open-source graph-based SLAM library that supports mapping and localization with RGB-D, stereo, or lidar sensors.","explanation":"RTAB-Map (Real-Time Appearance-Based Mapping) is an open-source SLAM (Simultaneous Localization and Mapping) library developed by Mathieu Labbé and François Michaud at IntRoLab, Université de Sherbrooke, in Canada. It organizes the places a robot has visited into a pose graph, and uses a bag-of-words model — which quantizes image features into “visual words” for comparison — to perform loop closure detection: recognizing “I’ve been here before” adds a constraint to the graph, which is then cleaned up with graph optimization to remove accumulated error. Its memory-management mechanism keeps only a subset of locations active for real-time detection and optimization, so it stays real-time even in large environments over long runs. It supports RGB-D cameras, stereo cameras, 3D lidar, and 2D lidar, and connects to ROS through rtabmap_ros, making it one of the most commonly used off-the-shelf solutions for mobile robot mapping and navigation.","example":"In ROS 2, an RGB-D camera runs rtabmap_ros while a robot is pushed around a lab; the result is a 3D point-cloud map and a 2D occupancy grid, handed off to Nav2 for navigation.","related":["Simultaneous Localization and Mapping","Visual SLAM","Loop Closure Detection","Occupancy Grid Map","Common ROS SLAM Packages (GMapping / SLAM Toolbox / RTAB-Map)","Depth Camera"]},{"id":"droid-slam","category":"perception","sec":10,"tier":3,"sources":[{"title":"arXiv 2108.10869: DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras","url":"https://arxiv.org/abs/2108.10869"},{"title":"princeton-vl/DROID-SLAM (GitHub)","url":"https://github.com/princeton-vl/DROID-SLAM"}],"as_of":"","related_ids":["visual-slam","bundle-adjustment","raft","orb-slam3","mast3r-slam","visual-odometry"],"name":"DROID-SLAM","alt":"DROID-SLAM","abbr":"","aliases":["DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras"],"one_liner":"A deep-learning visual SLAM system from Princeton that iteratively estimates camera pose and dense depth.","explanation":"DROID-SLAM is a visual SLAM (simultaneous localization and mapping) system from Princeton University’s Zachary Teed and Jia Deng, published at NeurIPS 2021. It borrows the architecture of the same group’s RAFT optical flow model, computing dense pixel correspondences between correlated frames, and uses a recurrent network to repeatedly update the camera poses and the depth of every pixel, with a differentiable dense bundle adjustment (BA) layer embedded in the middle — jointly optimizing camera poses and 3D structure — putting geometric constraints directly inside the network. It is trained purely on monocular video from the synthetic TartanAir dataset, but at test time it can also take stereo or RGB-D input, and it clearly outperforms earlier methods in accuracy on TartanAir, EuRoC, TUM-RGBD, and ETH3D, with far fewer catastrophic failures. The cost is dependence on a GPU — inference needs at least 11 GB of VRAM. Later deep-SLAM work like MASt3R-SLAM often uses it as a comparison baseline.","example":"Given a video of someone walking around a room with a handheld camera, DROID-SLAM outputs the camera pose and dense depth for every frame, which can be stitched into a point cloud of the room, or used to recover the camera trajectory from a human demonstration video.","related":["Visual SLAM","Bundle Adjustment","RAFT","ORB-SLAM3","MASt3R-SLAM","Visual Odometry"]},{"id":"mast3r-slam","category":"perception","sec":10,"tier":3,"sources":[{"title":"arXiv: MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors","url":"https://arxiv.org/abs/2412.12392"},{"title":"GitHub: rmurai0610/MASt3R-SLAM","url":"https://github.com/rmurai0610/MASt3R-SLAM"}],"as_of":"2025-06","related_ids":["mast3r","dust3r","simultaneous-localization-and-mapping","visual-slam","droid-slam","pointmap"],"name":"MASt3R-SLAM","alt":"MASt3R-SLAM","abbr":"","aliases":["MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors"],"one_liner":"A real-time monocular dense SLAM system built on MASt3R as a prior, able to map and localize from ordinary uncalibrated video.","explanation":"MASt3R-SLAM was proposed by Andrew Davison’s group at Imperial College London (Riku Murai, Eric Dexheimer, and colleagues), published at CVPR 2025 as a Highlight paper. It is a real-time monocular dense SLAM (simultaneous localization and mapping) system designed from the ground up around MASt3R, a two-view 3D reconstruction and matching model: for each new frame, it uses MASt3R to predict the point map and dense correspondences between that frame and the keyframes, then performs camera tracking, local fusion, loop closure, and a second-order global optimization, yielding a globally consistent camera trajectory and dense 3D geometry at about 15 frames per second. Traditional monocular SLAM usually needs the camera intrinsics calibrated beforehand; this system only assumes a single optical center and doesn’t depend on a fixed parametric camera model, so it can run on uncalibrated video too — and with known calibration, a small modification lets it reach the state of the art at the time. The open-source code supports live RealSense input, MP4 video, and folders of images.","example":"Recording a video indoors with a phone, with no camera calibration done, MASt3R-SLAM can estimate the camera trajectory and reconstruct a dense point cloud of the room.","related":["MASt3R","DUSt3R","Simultaneous Localization and Mapping","Visual SLAM","DROID-SLAM","Pointmap"]},{"id":"gaussian-splatting-slam","category":"perception","sec":10,"tier":3,"sources":[{"title":"SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM (arXiv 2312.02126)","url":"https://arxiv.org/abs/2312.02126"},{"title":"Gaussian Splatting SLAM / MonoGS (arXiv 2312.06741)","url":"https://arxiv.org/abs/2312.06741"}],"as_of":"2024-06","related_ids":["3d-gaussian-splatting","simultaneous-localization-and-mapping","visual-slam","neural-radiance-fields","novel-view-synthesis","real-to-sim"],"name":"Gaussian Splatting SLAM","alt":"高斯泼溅 SLAM","abbr":"","aliases":["3DGS SLAM","SplaTAM","MonoGS"],"one_liner":"A family of SLAM methods that use 3D Gaussian splatting as the map, building a photorealistically renderable scene while localizing.","explanation":"This is a family of SLAM methods that use 3D Gaussian splatting — representing a scene with a large number of colored, translucent 3D Gaussian ellipsoids that render very fast — as the map representation, appearing in a cluster starting in late 2023. Landmark examples include SplaTAM (Carnegie Mellon University and others, CVPR 2024, taking RGB-D video as input) and MonoGS (Imperial College London, Andrew Davison’s group, CVPR 2024, the first monocular version, running in real time at about 3 frames per second). During tracking, the current map is rendered into an image and compared against the real observation, optimizing the camera pose in reverse; during mapping, Gaussians are added, removed, and adjusted. Compared with a point-cloud or TSDF map, this approach gets a dense map that can be rendered from any new viewpoint at the same time as localization, which suits reality-to-simulation, digital twins, and navigation mapping for robots.","example":"SplaTAM scans a room once with a handheld RGB-D camera, estimating the camera trajectory while building a Gaussian map at the same time; afterward, the room can be rendered from a viewpoint that was never actually photographed.","related":["3D Gaussian Splatting","Simultaneous Localization and Mapping","Visual SLAM","Neural Radiance Fields","Novel View Synthesis","Real-to-Sim"]},{"id":"absolute-trajectory-error-relative-pose-error","category":"perception","sec":10,"tier":3,"sources":[{"title":"TUM RGB-D Dataset: Useful tools (ATE / RPE evaluation)","url":"https://cvg.cit.tum.de/data/datasets/rgbd-dataset/tools"},{"title":"evo: Python package for the evaluation of odometry and SLAM","url":"https://github.com/MichaelGrupp/evo"}],"as_of":"","related_ids":["visual-slam","visual-odometry","ground-truth","trajectory","loop-closure-detection"],"name":"Absolute Trajectory Error / Relative Pose Error","alt":"ATE / RPE（绝对 / 相对轨迹误差）","abbr":"ATE / RPE","aliases":["ATE","RPE","APE","Absolute Pose Error"],"one_liner":"Two standard metrics for how far a SLAM or odometry system’s estimated trajectory deviates from ground truth.","explanation":"ATE and RPE are the two most common metrics for evaluating localization algorithms like visual SLAM and visual odometry; the TUM RGB-D benchmark from the Technical University of Munich gave the widely used definitions and evaluation scripts. ATE (absolute trajectory error) first matches the estimated trajectory to the ground-truth trajectory by timestamp, then aligns the two as a whole with a single rigid-body transform (monocular SLAM also needs an extra scale estimate), and computes the root-mean-square error (RMSE, in meters) of the position differences at each moment, reflecting the trajectory’s overall global consistency. RPE, strictly “relative pose error,” compares relative motion over a fixed time or distance interval — for example, translational error in m/s and rotational error in deg/s — reflecting local drift, which suits evaluating odometry that has no loop closure. The popular open-source tool evo supports TUM, KITTI, EuRoC, and other formats, where the metric corresponding to ATE is called APE.","example":"Comparing a SLAM system’s output trajectory to ground truth with evo_ape: an ATE RMSE of 0.02 means the whole trajectory deviates by about 2 centimeters on average after alignment; running evo_rpe with a 1-second interval then shows how much drift accumulates every second.","related":["Visual SLAM","Visual Odometry","Ground Truth","Trajectory","Loop Closure Detection"]},{"id":"scene-understanding","category":"perception","sec":11,"tier":2,"sources":[{"title":"ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes (arXiv 1702.04405)","url":"https://arxiv.org/abs/1702.04405"},{"title":"ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning (arXiv 2309.16650)","url":"https://arxiv.org/abs/2309.16650"}],"as_of":"","related_ids":["3d-scene-graph","semantic-map","3d-vision","semantic-segmentation","conceptgraphs","scannet"],"name":"Scene Understanding","alt":"场景理解","abbr":"","aliases":["3D Scene Understanding"],"one_liner":"Working out what's in an environment, where it is, and how things relate to each other from images, depth, or point clouds.","explanation":"Scene understanding is a goal rather than a single algorithm: figuring out, from an image, depth, or point cloud, what's in an environment, where it is, and how things relate to each other. That covers geometry (where objects are, how big they are, where's walkable), semantics (what each thing is), and relations (a cup is on the table, a drawer can be pulled open). It's assembled from object detection, semantic and instance segmentation, depth estimation, 3D reconstruction, pose estimation, and affordance detection. Because robots need to navigate and manipulate in 3D space, embodied AI cares especially about 3D scene understanding: ScanNet (2017) provides 1,513 indoor scenes and about 2.5 million semantically annotated RGB-D frames, and ConceptGraphs (2023) fuses results from 2D foundation models across multiple views into open-vocabulary 3D scene graphs that a large language model can use for instruction-based planning.","example":"Entering a kitchen, a household robot first builds a 3D scene graph: a fridge, a dining table, two bowls on the table, with the bowls sitting on the table. Given the instruction ‘put the bowls in the sink,’ it looks up the bowls' and sink's 3D positions from this graph and plans accordingly.","related":["3D Scene Graph","Semantic Map","3D Vision","Semantic Segmentation","ConceptGraphs","ScanNet"]},{"id":"octomap","category":"perception","sec":11,"tier":3,"sources":[{"title":"OctoMap 官方主页","url":"https://octomap.github.io/"},{"title":"MoveIt Perception Pipeline Tutorial","url":"https://moveit.picknik.ai/main/doc/examples/perception_pipeline/perception_pipeline_tutorial.html"}],"as_of":"","related_ids":["occupancy-grid-map","voxel","point-cloud","truncated-signed-distance-function","moveit-motion-planning-framework","collision-checking"],"name":"OctoMap","alt":"八叉树地图","abbr":"","aliases":["Octree Map","Octree Occupancy Map"],"one_liner":"A 3D probabilistic occupancy map stored hierarchically as an octree, commonly used for robot obstacle avoidance and planning.","explanation":"OctoMap is an open-source C++ library developed by Armin Hornung, Kai M. Wurm, and colleagues at the University of Freiburg, described in a 2013 paper in Autonomous Robots. It recursively splits space into eight sub-cubes (an octree), and each cell stores a probability of being occupied, distinguishing three states — occupied, free, and unknown — with sensor noise and environment changes corrected gradually through probabilistic updates. Compared with dividing all of space into equal-sized voxels, it subdivides only where there is something to represent, which saves a great deal of memory, and it can be queried at a coarse or fine resolution as needed. Point clouds from a depth camera or lidar can be written directly into it for collision checking; MoveIt’s perception pipeline uses OctoMap to represent obstacles around the robot.","example":"A depth camera mounted beside a robot arm continuously feeds its point cloud into OctoMap through MoveIt, so the arm automatically plans around clutter newly placed on the table.","related":["Occupancy Grid Map","Voxel","Point Cloud","Truncated Signed Distance Function","MoveIt Motion Planning Framework","Collision Checking"]},{"id":"nvblox","category":"perception","sec":11,"tier":3,"sources":[{"title":"nvblox: GPU-Accelerated Incremental Signed Distance Field Mapping (arXiv)","url":"https://arxiv.org/abs/2311.00626"},{"title":"nvblox Documentation","url":"https://nvidia-isaac.github.io/nvblox/v0.0.9/index.html"}],"as_of":"2024-05","related_ids":["truncated-signed-distance-function","euclidean-signed-distance-field","nvidia-isaac-ros","ros-2-navigation-stack","costmap","obstacle-avoidance"],"name":"nvblox","alt":"nvblox","abbr":"","aliases":["NVIDIA nvblox","Isaac ROS nvblox","isaac_ros_nvblox"],"one_liner":"An open-source, GPU-accelerated 3D mapping library from NVIDIA that builds real-time distance-field maps for obstacle avoidance.","explanation":"nvblox is an open-source voxel mapping library from NVIDIA, described in a paper at ICRA 2024. Running on the GPU, it incrementally fuses data from an RGB-D camera or lidar into a TSDF (truncated signed distance function, which records how far each voxel is from the nearest surface), and computes an ESDF (Euclidean signed distance field) in real time for collision checking during path planning. The paper reports speedups of up to 177x for surface reconstruction and up to 31x for distance-field computation. nvblox provides a ROS 2 interface as part of Isaac ROS, and can output costmaps for the Nav2 navigation stack to consume.","example":"A mobile robot builds a map on the GPU with an RGB-D camera and nvblox while driving, feeding the resulting costmap to Nav2 for obstacle avoidance.","related":["Truncated Signed Distance Function","Euclidean Signed Distance Field","NVIDIA Isaac ROS","ROS 2 Navigation Stack (Nav2)","Costmap","Obstacle Avoidance"]},{"id":"occupancy-network","category":"perception","sec":11,"tier":3,"sources":[{"title":"Occupancy Networks: Learning 3D Reconstruction in Function Space (arXiv)","url":"https://arxiv.org/abs/1812.03828"}],"as_of":"","related_ids":["occupancy-grid-map","implicit-vs-explicit-3d-representation","birds-eye-view","autonomous-driving","vision-only-approach","tesla-ai-day"],"name":"Occupancy Network","alt":"占用网络","abbr":"","aliases":["Occupancy Networks","3D Occupancy Prediction","Occupancy Grid Prediction"],"one_liner":"A neural network that predicts, for every location in 3D space, whether it is occupied by an object.","explanation":"“Occupancy network” originally referred to work by Mescheder et al. at CVPR 2019: a network that takes an image or point cloud as input and, for any queried 3D point, outputs the probability that the point lies inside an object — an implicit representation used to reconstruct 3D shapes at a resolution not limited by a voxel grid. A second, unrelated sense of the term comes from self-driving cars: in 2022 Tesla publicly described an occupancy network that uses multiple onboard cameras to predict whether each voxel of space around the vehicle is occupied, and “3D occupancy prediction” has since become a popular task in autonomous-driving research. Its advantage is that it does not depend on a predefined list of object categories, so it can represent oddly shaped obstacles; it is closely related to occupancy grid maps and bird’s-eye-view perception.","example":"A camera-only self-driving system predicts whether each voxel of space around the car is occupied using multiple cameras, letting it avoid oddly shaped obstacles that aren’t in its list of detectable categories.","related":["Occupancy Grid Map","Implicit vs. Explicit 3D Representation","Bird’s-Eye View","Autonomous Driving","Vision-Only Approach","Tesla AI Day"]},{"id":"birds-eye-view","category":"perception","sec":11,"tier":3,"sources":[{"title":"BEVFormer (arXiv 2203.17270, ECCV 2022)","url":"https://arxiv.org/abs/2203.17270"},{"title":"Lift, Splat, Shoot (arXiv 2008.05711, ECCV 2020)","url":"https://arxiv.org/abs/2008.05711"}],"as_of":"","related_ids":["autonomous-driving","3d-object-detection","occupancy-network","elevation-map","multi-sensor-fusion","semantic-map"],"name":"Bird’s-Eye View","alt":"鸟瞰图","abbr":"BEV","aliases":["BEV","BEV Perception","BEV Representation"],"one_liner":"A 2D top-down grid that fuses information from multiple sensors onto a single ground plane.","explanation":"A bird’s-eye view (BEV) is a 2D grid representation laid over the ground plane, viewed from directly above, where each cell stores features or semantics for that location. Autonomous driving was the first field to use it at scale: Lift-Splat-Shoot (NVIDIA, ECCV 2020) predicts a depth distribution for every pixel and then “splats” image features onto a BEV grid; BEVFormer (ECCV 2022, Shanghai AI Lab and others) uses Transformer cross-attention to query BEV features from multiple camera streams and fuses in past frames. The benefit is that multiple cameras and lidar can all be aligned into one coordinate frame, so detection, segmentation, and path planning can happen directly on this map. The occupancy grids, elevation maps, and semantic maps used in robot navigation are also essentially BEV representations; because BEV compresses away height information, tasks that depend more on height, like tabletop manipulation, usually switch to point clouds or voxels instead.","example":"A self-driving car converts images from its six surround-view cameras into a single BEV feature map centered on itself, covering tens of meters ahead, behind, and to each side; it draws boxes around vehicles and lane lines directly on this map before handing it to the planning module.","related":["Autonomous Driving","3D Object Detection","Occupancy Network","Elevation Map","Multi-Sensor Fusion","Semantic Map"]},{"id":"elevation-map","category":"perception","sec":11,"tier":3,"sources":[{"title":"ANYbotics/elevation_mapping (GitHub)","url":"https://github.com/ANYbotics/elevation_mapping"},{"title":"leggedrobotics/elevation_mapping_cupy (GitHub)","url":"https://github.com/leggedrobotics/elevation_mapping_cupy"}],"as_of":"","related_ids":["height-scan","traversability-estimation","perceptive-locomotion","occupancy-grid-map","rough-terrain-locomotion","anybotics-anymal"],"name":"Elevation Map","alt":"高程图","abbr":"","aliases":["2.5D Elevation Map","Elevation Mapping","Height Map"],"one_liner":"A 2.5D terrain map that divides the ground into a grid and stores one height value per cell, commonly used by legged robots.","explanation":"An elevation map is a 2.5D map: the ground around a robot is divided into a regular grid on the horizontal plane, and each cell stores just one height value (usually with a variance representing uncertainty too), rather than a full 3D voxel grid. The data comes from range sensors like depth cameras and lidar, continuously fused and updated together with the robot’s pose estimate. elevation_mapping, developed starting in 2014 by Péter Fankhauser and colleagues at ETH Zurich, is a commonly used open-source implementation that builds the map robot-centered and explicitly accounts for pose-drift uncertainty; it is no longer maintained. The 2022 elevation_mapping_cupy moved the computation to the GPU and added traversability, semantic, and other layers. An elevation map is more compact and faster to query than a point cloud, and perceptive locomotion for quadruped and humanoid robots commonly reads terrain height in a patch around the feet from it as policy input. Its limitation is that each cell holds only one height, so it can’t represent overhangs like the underside of a table or a bridge.","example":"When training perceptive legged locomotion in Isaac Lab, a grid of terrain heights around the robot (a height scan) is commonly used as an observation; once deployed on the real robot, these heights are looked up from an elevation map built in real time.","related":["Height Scan","Traversability Estimation","Perceptive Locomotion","Occupancy Grid Map","Rough-terrain Locomotion","ANYbotics ANYmal"]},{"id":"height-scan","category":"perception","sec":11,"tier":3,"sources":[{"title":"Isaac Lab velocity_env_cfg.py（height_scanner 与 height_scan 观测定义）","url":"https://raw.githubusercontent.com/isaac-sim/IsaacLab/main/source/isaaclab_tasks/isaaclab_tasks/manager_based/locomotion/velocity/velocity_env_cfg.py"},{"title":"Learning robust perceptive locomotion for quadrupedal robots in the wild（Science Robotics 2022）","url":"https://arxiv.org/html/2201.08117"},{"title":"legged_gym legged_robot_config.py（measured_points 配置）","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["elevation-map","perceptive-locomotion","blind-locomotion","teacher-student-distillation","nvidia-isaac-lab","legged-gym"],"name":"Height Scan","alt":"高度扫描","abbr":"","aliases":["height_scan","Terrain Height Map","Height Samples"],"one_liner":"Samples ground height around a robot in a fixed pattern, used as terrain input for a locomotion policy.","explanation":"A height scan is a commonly used terrain observation in reinforcement-learning locomotion for legged robots: centered on the body, or around each foot, it samples ground height at a number of points in a fixed pattern, then stacks these into a vector fed to the policy network. In simulation, this is usually done by casting rays straight down from above the robot and reading off ground-truth height directly; a real robot has no such bird’s-eye view, so it first has to build an elevation map (a 2.5D grid of ground height) with a depth camera or lidar, then sample from that map. A real elevation map has noise and occlusion, so a common approach trains a teacher policy on noise-free height first, then distills it into a student policy that reads the real, noisy elevation map. A policy that relies only on proprioception, with no height scan, is called blind locomotion.","example":"Isaac Lab’s velocity-tracking task mounts a ray-casting sensor on the body, measuring ground height on a 1.6 m × 1.0 m grid at 0.1 m spacing; Miki and colleagues’ 2022 ANYmal work instead samples 52 points in 5 concentric rings around each foot, for 208 dimensions across all four legs.","related":["Elevation Map","Perceptive Locomotion","Blind Locomotion","Teacher-Student Distillation","NVIDIA Isaac Lab","legged_gym"]},{"id":"traversability-estimation","category":"perception","sec":11,"tier":3,"sources":[{"title":"Fast Traversability Estimation for Wild Visual Navigation (RSS 2023, arXiv:2305.08510)","url":"https://arxiv.org/abs/2305.08510"}],"as_of":"","related_ids":["elevation-map","costmap","perceptive-locomotion","rough-terrain-locomotion","navigation","semantic-segmentation"],"name":"Traversability Estimation","alt":"可通行性估计","abbr":"","aliases":["Traversable-Area Analysis","Terrain Traversability"],"one_liner":"Judging which parts of the terrain ahead can be crossed, which can’t, and how costly each path would be.","explanation":"This is a question wheeled and legged mobile robots must answer before navigating: can each patch of ground ahead be crossed safely, and how difficult would it be? Traditional methods compute slope, roughness, and step height from an elevation map (a 2D map recording the ground height at every grid cell); this purely geometric approach mistakes tall grass or branches for obstacles, which is why later work added semantic information (grass, mud, water) and self-supervised methods that learn from the robot’s own walking experience. The result is usually output as a traversability map or costmap, handed to a path planner to choose a route. Wild Visual Navigation, from ETH Zurich and Oxford in 2023, let the quadruped ANYmal learn to tell walkable vegetation apart from real obstacles after less than 5 minutes of on-site training.","example":"Using Wild Visual Navigation, the ANYmal quadruped, after being briefly led by a person, can push through tall grass and go around tree trunks on its own, completing a 1.4 km trail navigation.","related":["Elevation Map","Costmap","Perceptive Locomotion","Rough-terrain Locomotion","Navigation","Semantic Segmentation"]},{"id":"topological-map","category":"perception","sec":11,"tier":3,"sources":[{"title":"Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs (arXiv 2407.07775)","url":"https://arxiv.org/abs/2407.07775"},{"title":"ViNT: A Foundation Model for Visual Navigation (arXiv 2306.14846)","url":"https://arxiv.org/abs/2306.14846"}],"as_of":"","related_ids":["navigation","occupancy-grid-map","semantic-map","a-star-search","vint","mobility-vla"],"name":"Topological Map","alt":"拓扑地图","abbr":"","aliases":["Topological Graph"],"one_liner":"A map made of place nodes and connecting edges, recording only what leads where, not exact coordinates.","explanation":"A topological map represents an environment as a graph: nodes are representative places (a particular room, an intersection, or a keyframe image), and edges mean that two places can be reached directly from each other. Unlike a metric map such as an occupancy grid, it does not require an exact metric coordinate for every location, so it takes far less storage, can scale to very large environments, and is closer to how a person describes a route (“go out the door, turn left down the hallway, then into the kitchen”). During navigation, an algorithm such as A* first searches the graph for a sequence of nodes, and a local controller or policy then moves between adjacent nodes. Recent learned navigation systems make heavy use of this idea: ViNT and NoMaD use images as nodes for long-range navigation, while Mobility VLA builds a topological map from a video of a person walking around giving a tour, then uses a vision-language model to interpret an instruction and find the target node.","example":"A robot is first walked around an office building while recording video, saving one image as a node every few seconds; later, told “go to the printer,” it finds the matching node in the graph and plans a path of nodes to reach it.","related":["Navigation","Occupancy Grid Map","Semantic Map","A* Search","ViNT","Mobility VLA"]},{"id":"semantic-map","category":"perception","sec":11,"tier":3,"sources":[{"title":"VLMaps 项目主页","url":"https://vlmaps.github.io/"},{"title":"arXiv 2210.05714: Visual Language Maps for Robot Navigation","url":"https://arxiv.org/abs/2210.05714"}],"as_of":"","related_ids":["semantic-slam","vlmaps","conceptgraphs","3d-scene-graph","object-goal-navigation","occupancy-grid-map"],"name":"Semantic Map","alt":"语义地图","abbr":"","aliases":["Open-Vocabulary Semantic Map","Language Map"],"one_liner":"A map that labels what each place is, on top of its geometry, so it can be queried by object name or natural language.","explanation":"An ordinary SLAM map only records where obstacles are and where surfaces sit; a semantic map goes further, attaching category labels or features to the points, cells, or objects in the map, so it can answer questions like “where is the refrigerator” or “where is the kitchen.” Early approaches labeled the map using detection or segmentation networks with a fixed set of categories; in recent years, open-vocabulary semantic maps have become popular, storing features from vision-language models such as CLIP or LSeg directly in the map so it can be queried with arbitrary natural language. A representative example is VLMaps (2022, from the University of Freiburg, Google, and others), which extracts pixel-level vision-language features from RGB-D video, back-projects them into 3D using depth, and stores them in a top-down grid map, letting it follow navigation instructions with spatial relationships such as “go between the sofa and the TV.” Semantic maps can also be organized as 3D scene graphs, and are a common intermediate representation for object-goal navigation, vision-and-language navigation, and mobile manipulation.","example":"VLMaps lets a LoCoBot and a drone share the same language map, navigating from instructions such as “move three meters to the right of the chair.”","related":["Semantic SLAM","VLMaps","ConceptGraphs","3D Scene Graph","Object-Goal Navigation","Occupancy Grid Map"]},{"id":"semantic-slam","category":"perception","sec":11,"tier":3,"sources":[{"title":"arXiv 1910.02490: Kimera","url":"https://arxiv.org/abs/1910.02490"},{"title":"arXiv 2505.12384: Is Semantic SLAM Ready for Embedded Systems?","url":"https://arxiv.org/abs/2505.12384"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","semantic-map","visual-slam","3d-scene-graph","gaussian-splatting-slam","conceptgraphs"],"name":"Semantic SLAM","alt":"语义SLAM","abbr":"","aliases":["Metric-Semantic SLAM"],"one_liner":"SLAM that recognizes object categories while localizing and mapping, producing a map with semantic labels attached.","explanation":"Semantic SLAM is a family of methods that adds semantic information to traditional SLAM (Simultaneous Localization and Mapping): while estimating its own pose and building a geometric map, the system also uses a detection or segmentation network to recognize walls, tables, chairs, and other objects, and writes those categories into the map. Semantics can also help SLAM itself — for instance, by filtering out moving objects such as people and cars to reduce their interference with pose estimation, or by using objects as landmarks for loop closure detection. Representative work includes SLAM++ (2013), which builds maps at the level of individual objects, and Kimera, open-sourced by MIT’s SPARK Lab in 2019, which uses a camera and IMU to generate a semantically labeled 3D mesh in real time on a CPU. Recent work has also explored semantic SLAM based on NeRF and 3D Gaussian Splatting, though the computational cost is still high for embedded platforms. Its output is, in effect, a semantic map.","example":"Kimera walks a stereo camera and IMU around an indoor space and outputs, in real time, a 3D mesh map labeled with semantics such as “wall,” “floor,” and “chair.”","related":["Simultaneous Localization and Mapping","Semantic Map","Visual SLAM","3D Scene Graph","Gaussian Splatting SLAM","ConceptGraphs"]},{"id":"3d-scene-graph","category":"perception","sec":11,"tier":3,"sources":[{"title":"3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera (arXiv 1910.02527)","url":"https://arxiv.org/abs/1910.02527"},{"title":"Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization (arXiv 2201.13360)","url":"https://arxiv.org/abs/2201.13360"},{"title":"ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning (arXiv 2309.16650)","url":"https://arxiv.org/abs/2309.16650"}],"as_of":"","related_ids":["conceptgraphs","sayplan","semantic-map","scene-understanding","llm-based-task-planning"],"name":"3D Scene Graph","alt":"3D场景图","abbr":"","aliases":["3DSG"],"one_liner":"Organizes a 3D scene’s floors, rooms, objects, and their relationships into a layered graph.","explanation":"A 3D scene graph is a structured scene representation: nodes are entities in the scene — buildings, floors, rooms, objects, camera positions, and so on — and edges are the relationships between them, like “inside,” “on top of,” or “next to,” with each node carrying attributes such as category, size, and 3D position. Armeni and colleagues introduced this structure at ICCV 2019, with building, room, object, and camera layers; MIT’s Hydra (RSS 2022) lets a robot build a hierarchical scene graph in real time while exploring. Compared with a dense point cloud or mesh, a scene graph is compact and semantically clear — it can directly answer “which room is the fridge in?” — and it’s easy to convert to text for a large language model to use in task planning. ConceptGraphs builds open-vocabulary scene graphs using 2D foundation models, while SayPlan has a large model search and plan long-horizon, multi-room tasks directly over a scene graph.","example":"A robot given the instruction “bring the apple from the kitchen table to the living room” first follows kitchen → table → apple through the scene graph to find the target node and the living-room node, then calls its navigation, grasping, and placing skills in sequence.","related":["ConceptGraphs","SayPlan","Semantic Map","Scene Understanding","LLM-based Task Planning"]},{"id":"conceptgraphs","category":"perception","sec":11,"tier":3,"sources":[{"title":"ConceptGraphs (arXiv 2309.16650)","url":"https://arxiv.org/abs/2309.16650"},{"title":"ConceptGraphs 项目页","url":"https://concept-graphs.github.io/"}],"as_of":"2024-05","related_ids":["3d-scene-graph","open-vocabulary","semantic-map","clip","llm-based-task-planning","sayplan"],"name":"ConceptGraphs","alt":"ConceptGraphs","abbr":"","aliases":["ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning"],"one_liner":"Fuses multi-view images into an open-vocabulary 3D object scene graph for large-model reasoning and robot planning.","explanation":"ConceptGraphs was proposed by a team from MIT, Université de Montréal, the University of Toronto, and others (lead author Qiao Gu), published at ICRA 2024. It uses a general-purpose segmentation model frame by frame to cut out object regions, extracts CLIP features for each region, projects them into 3D points using depth, and then merges regions belonging to the same object across viewpoints to obtain a set of 3D objects with semantic features attached. It then uses LLaVA to generate text descriptions for the objects and GPT-4 to infer spatial relationships between objects as edges, forming a 3D scene graph — nodes are objects, edges are relationships. The whole pipeline needs no 3D annotation and no model fine-tuning. The resulting graph can be converted to text for a large model to answer queries like “find a basketball”; the paper demonstrates object retrieval, relocalization, and navigation tasks that judge which obstacles can be pushed aside, on mobile robots like the Jackal.","example":"A robot first drives a loop through a house to build the scene graph; when the user says “find something to prop up my phone,” a large model picks “book” from the object descriptions in the graph, and the robot navigates to that node’s 3D location.","related":["3D Scene Graph","Open-vocabulary","Semantic Map","CLIP","LLM-based Task Planning","SayPlan"]},{"id":"distilled-feature-fields","category":"perception","sec":11,"tier":3,"sources":[{"title":"arXiv 2205.15585: Decomposing NeRF for Editing via Feature Field Distillation","url":"https://arxiv.org/abs/2205.15585"},{"title":"arXiv 2308.07931: Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation","url":"https://arxiv.org/abs/2308.07931"},{"title":"F3RM project page","url":"https://f3rm.github.io/"}],"as_of":"","related_ids":["f3rm","lerf","neural-radiance-fields","clip","knowledge-distillation","3d-gaussian-splatting"],"name":"Distilled Feature Fields","alt":"蒸馏特征场","abbr":"DFF","aliases":["DFF","Feature Fields","Language-Embedded Fields"],"one_liner":"Distills features from 2D models like CLIP and DINO into a 3D scene, so every point in space carries semantics.","explanation":"A distilled feature field adds a learned feature vector at every point in space to a neural radiance field (NeRF, a neural representation that reconstructs a 3D scene from multi-view photos) or a Gaussian splat, on top of color and density; the training objective is to make this feature, once rendered from any viewpoint, match the image features extracted by a 2D foundation model like CLIP or DINO — effectively distilling the 2D model’s knowledge into 3D. The name comes from a 2022 NeurIPS paper by Kobayashi, Sitzmann, and colleagues at Preferred Networks and MIT, originally used to select and edit objects in a NeRF by text or by clicking. It combines accurate geometry with the ability to locate objects in 3D space using natural language. In robotics, MIT’s F3RM uses CLIP-distilled feature fields for few-shot, language-guided 6-DoF grasping and placing, winning the CoRL 2023 best paper award; LERF takes a similar approach.","example":"F3RM first takes a set of multi-view photos of a tabletop and reconstructs a scene carrying CLIP features; with just two demonstrations for a task like “grasp the cup by the rim,” it can then grasp objects of shapes and categories it has never seen, based on a text instruction.","related":["F3RM","LERF","Neural Radiance Fields","CLIP","Knowledge Distillation","3D Gaussian Splatting"]},{"id":"lerf","category":"perception","sec":11,"tier":3,"sources":[{"title":"LERF: Language Embedded Radiance Fields (ICCV 2023 项目主页)","url":"https://www.lerf.io/"},{"title":"Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping (LERF-TOGO, arXiv 2309.07970)","url":"https://arxiv.org/abs/2309.07970"}],"as_of":"","related_ids":["neural-radiance-fields","clip","distilled-feature-fields","f3rm","open-vocabulary","3d-visual-grounding"],"name":"LERF","alt":"LERF","abbr":"LERF","aliases":["Language Embedded Radiance Fields","LERF-TOGO"],"one_liner":"Embeds CLIP language features into a NeRF, so objects in a 3D scene can be found with natural language.","explanation":"LERF was proposed by Justin Kerr, Chung Min Kim, Angjoo Kanazawa, and colleagues at UC Berkeley, published at ICCV 2023 as an oral presentation. On top of a neural radiance field (NeRF, which uses a network to represent color and density at every point in a scene), it additionally learns a multi-scale language field: a point in space maps to a CLIP vector at each of several scales, supervised during training by CLIP features extracted from an image pyramid across multiple viewpoints, and regularized with DINO features to sharpen object boundaries. Once built, feeding in any text produces a relevance heatmap rendered in 3D space, with no detection box, segmentation mask, or model fine-tuning needed. It is an early landmark for bringing a vision-language model’s open-vocabulary ability into 3D scenes, with code integrated into Nerfstudio.","example":"LERF-TOGO (2023) first uses LERF to localize an object part like “the handle of the cup” in a scene, then ranks the candidate grasps from an off-the-shelf grasp planner; across 31 real objects, it picked the correct part 81% of the time, with a 69% grasp success rate.","related":["Neural Radiance Fields","CLIP","Distilled Feature Fields","F3RM","Open-vocabulary","3D Visual Grounding"]},{"id":"3d-visual-grounding","category":"perception","sec":11,"tier":3,"sources":[{"title":"ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language (ECCV 2020)","url":"https://daveredrum.github.io/ScanRefer/"},{"title":"ReferIt3D: Neural Listeners for Fine-Grained 3D Object Identification in Real-World Scenes (ECCV 2020)","url":"https://referit3d.github.io/"}],"as_of":"","related_ids":["visual-grounding","scannet","3d-object-detection","3d-scene-graph","open-vocabulary-object-detection","language-grounding"],"name":"3D Visual Grounding","alt":"3D视觉定位","abbr":"","aliases":["3D Grounding","3D Referring Expression Grounding"],"one_liner":"Finds the object in a 3D scene that a sentence describes.","explanation":"3D visual grounding brings 2D visual grounding — using text to draw a box around a target in an image — into 3D scenes: given a scene’s point cloud or multi-view images plus a description such as “the trash can next to the chair by the window,” it outputs the target object’s 3D bounding box. The Chinese term for this, 视觉定位, is also commonly used for a camera estimating its own position, but here it means grounding. The main difficulty is that a scene often has multiple objects of the same category, so the model must understand spatial relationships like “on the left,” “the biggest one,” or “next to the door” to tell them apart. ScanRefer and ReferIt3D, both published at ECCV 2020, built benchmarks on ScanNet indoor scans: ScanRefer has about 52,000 descriptions covering 11,000 objects, while ReferIt3D provides two sets, Nr3D (natural descriptions) and Sr3D (spatial-relation descriptions). For robots, this is the key step that grounds a language instruction in an actual object.","example":"A user says “hand me the blue cup on the far right of the desk”; the robot runs 3D visual grounding on the reconstructed room point cloud to get that cup’s 3D box, then plans a grasp.","related":["Visual Grounding","ScanNet","3D Object Detection","3D Scene Graph","Open-Vocabulary Object Detection","Language Grounding"]},{"id":"visual-question-answering","category":"perception","sec":11,"tier":2,"sources":[{"title":"VQA: Visual Question Answering (ICCV 2015)","url":"https://arxiv.org/abs/1505.00468"},{"title":"VQA 官网（VQA v2 数据集与挑战赛）","url":"https://visualqa.org/"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization（协同训练用到 VQAv2 等网络数据）","url":"https://arxiv.org/abs/2504.16054"}],"as_of":"2025-04","related_ids":["vision-language-model","multimodal-large-language-model","embodied-question-answering","co-training","erqa","vsi-bench"],"name":"Visual Question Answering","alt":"视觉问答","abbr":"VQA","aliases":["VQA","Image QA"],"one_liner":"Given an image and a natural-language question about it, the model answers in words.","explanation":"Visual question answering takes an image and a question about it, such as “how many cups are on the table?”, and outputs a natural-language answer, requiring the model to understand the image, the language, and relevant common sense together. The task was introduced by Antol, Agrawal, and colleagues at ICCV 2015; the VQA dataset has about 250,000 images and 760,000 questions. The 2017 VQA v2 rebalanced the dataset to reduce cases where a model could guess the answer from the question alone, without looking at the image. Today most vision-language models are trained and evaluated with question-answering formats, and embodied and spatial-reasoning benchmarks like ERQA and VSI-Bench are also framed as question answering; VQA data is often mixed into VLA training as a co-training task — for example, π0.5 uses web data including VQAv2 to preserve its image-understanding ability. Unlike embodied question answering, VQA only looks at a given image and never requires the robot to move around to find the answer.","example":"Given a photo of a kitchen counter, ask “is the cup to the left of the sink empty?” and the model answers “yes, it’s empty”; an embodied-reasoning benchmark would instead ask something like “which object should the robot’s gripper move above first?”","related":["Vision-Language Model","Multimodal Large Language Model","Embodied Question Answering","Co-training","ERQA","VSI-Bench"]},{"id":"pointing","category":"perception","sec":11,"tier":2,"sources":[{"title":"Molmo and PixMo (arXiv 2409.17146)","url":"https://arxiv.org/abs/2409.17146"},{"title":"RoboPoint (arXiv 2406.10721)","url":"https://arxiv.org/abs/2406.10721"},{"title":"Gemini API Docs: Gemini Robotics-ER overview","url":"https://ai.google.dev/gemini-api/docs/robotics-overview"}],"as_of":"2026-09","related_ids":["molmo","robopoint","gemini-robotics-er","visual-prompting","affordance-detection","projection-back-projection"],"name":"Pointing","alt":"指向（点预测）","abbr":"","aliases":["2D Point Prediction","Point Prediction"],"one_liner":"A vision-language model answering by marking exact pixel coordinates on the image, following a text instruction.","explanation":"Pointing lets a vision-language model (VLM) answer by literally marking a spot on the image: given an image and an instruction, it outputs one or more 2D pixel coordinates. Ai2's Molmo (2024) specifically collected the PixMo-Points dataset for this and can point at and count objects; RoboPoint (2024) fine-tunes a VLM on automatically synthesized data to predict keypoints for where to grasp or where to place something; Google DeepMind's Gemini Robotics-ER outputs coordinates in [y, x] format, normalized to a 0–1000 range. A point is more fine-grained than a box and more directly usable downstream than plain text: combined with back-projection from a depth map, it gives a 3D target position that can be handed directly to grasping, motion planning, or a VLA model.","example":"Asking Gemini Robotics-ER to ‘point to all the bananas in the image’ returns a list of entries like point: [376, 508], label: small banana, with each point marking one banana's location.","related":["Molmo (Ai2)","RoboPoint","Gemini Robotics-ER","Visual Prompting","Affordance Detection","Projection / Back-Projection"]},{"id":"visual-prompting","category":"perception","sec":11,"tier":3,"sources":[{"title":"Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V (arXiv)","url":"https://arxiv.org/abs/2310.11441"},{"title":"MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv)","url":"https://arxiv.org/abs/2403.03174"}],"as_of":"","related_ids":["visual-prompting-2","moka","pivot","visual-grounding","segment-anything-model","vision-language-model"],"name":"Visual Prompting","alt":"视觉提示","abbr":"","aliases":["Set-of-Mark","SoM","Marker-Based Visual Prompting"],"one_liner":"Drawing boxes or numbers directly on an image so a multimodal large model can answer by referring to those marks.","explanation":"Visual prompting means overlaying boxes, points, arrows, or numbers directly on an input image, without changing any model weights, to guide a multimodal large model to attend to and refer to specific regions. The representative method is Set-of-Mark (SoM), proposed by a Microsoft team in 2023: a segmentation model such as SAM or SEEM first cuts the image into regions, each region is labeled with a number, mask, or box, and GPT-4V is then asked to answer using those numbers; the paper reports that this beats fully fine-tuned specialist models zero-shot on the RefCOCOg referring task. It addresses the fact that large models can describe a scene in words but struggle to state a precise pixel location in text. In robotics, work such as MOKA and PIVOT uses this kind of marking to let a vision-language model pick out a grasp point or a direction to move, which is then handed to low-level control to execute.","example":"Objects in a tabletop photo are labeled 1 through 8; asked “which object could be used to scoop soup,” the model answers “number 5,” and the program converts region 5 into 3D coordinates for the robot arm.","related":["Visual Prompting (Set-of-Mark)","MOKA","PIVOT","Visual Grounding","Segment Anything Model","Vision-Language Model"]},{"id":"vsi-bench","category":"perception","sec":11,"tier":3,"sources":[{"title":"Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces (arXiv)","url":"https://arxiv.org/abs/2412.14171"},{"title":"Thinking in Space 项目主页","url":"https://vision-x-nyu.github.io/thinking-in-space.github.io/"},{"title":"vision-x-nyu/thinking-in-space (GitHub)","url":"https://github.com/vision-x-nyu/thinking-in-space"}],"as_of":"2025-06","related_ids":["spatial-intelligence","spatial-reasoning","multimodal-large-language-model","benchmark","erqa","3d-vision"],"name":"VSI-Bench","alt":"VSI-Bench 空间智能基准","abbr":"VSI-Bench","aliases":["Thinking in Space: Visual-Spatial Intelligence Benchmark","Thinking in Space"],"one_liner":"A benchmark that uses indoor videos to test how well multimodal large models understand space.","explanation":"VSI-Bench comes from the paper Thinking in Space, released in December 2024 by Saining Xie’s group at NYU together with Yale and Stanford (including Fei-Fei Li), and selected as an oral presentation at CVPR 2025. It draws 288 first-person videos from three indoor scanning datasets — ScanNet, ScanNet++, and ARKitScenes — and builds more than 5,000 question-answer pairs covering 8 task types: object counting, relative distance, relative direction, object size, absolute distance, room size, order of appearance, and route planning. The authors tested 15 multimodal large models that support video input, and all scored well below human performance (humans average about 79%), mostly struggling with spatial reasoning; language-prompting techniques such as chain-of-thought actually hurt scores, while having the model first draw a “cognitive map” improved distance judgments.","example":"Shown a video panning around a living room, the model is asked “how many meters apart are the sofa and the TV” or “standing in front of the fridge facing the sink, is the stove on your left or your right,” and is scored by numeric error or multiple-choice accuracy.","related":["Spatial Intelligence","Spatial Reasoning","Multimodal Large Language Model","Benchmark","ERQA","3D Vision"]},{"id":"active-perception","category":"perception","sec":11,"tier":3,"sources":[{"title":"Revisiting Active Perception (Bajcsy, Aloimonos, Tsotsos, arXiv:1603.02729)","url":"https://arxiv.org/abs/1603.02729"}],"as_of":"","related_ids":["interactive-perception","next-best-view-planning","active-exploration","perception-action-loop","embodied-perception","head-camera"],"name":"Active Perception","alt":"主动感知","abbr":"","aliases":["Active Vision","Active Sensing"],"one_liner":"A robot deliberately moves its sensors or body to see or feel a target better, rather than just passively observing.","explanation":"Active perception means an agent doesn’t just passively receive sensor data, but actively decides where, how, and when to look based on the task and its current judgment. The idea was proposed by Ruzena Bajcsy in 1985 and published in Proceedings of the IEEE in 1988; that same year Aloimonos and colleagues independently proposed “active vision.” A 2016 survey by Bajcsy, Aloimonos, and Tsotsos defined it this way: an agent knows why it needs to perceive something, chooses what to perceive, and decides how, when, and where to perceive it. It addresses the problem of a single viewpoint not giving enough information — occlusion, a target outside the field of view, or poor lighting. Common forms in embodied AI include turning the head for a different view, adjusting a wrist camera, or reaching out to touch something to confirm its material; it is closely related to next-best-view planning, interactive perception, and active exploration.","example":"A cup is half-hidden behind a cardboard box; a humanoid robot turns its head and bends down for another shot from a different angle, confirms which way the handle is facing, and only then reaches to grasp it.","related":["Interactive Perception","Next-Best-View Planning","Active Exploration","Perception-Action Loop","Embodied Perception","Head Camera"]},{"id":"next-best-view-planning","category":"perception","sec":11,"tier":3,"sources":[{"title":"The determination of next best views (Connolly, ICRA 1985) - Semantic Scholar","url":"https://www.semanticscholar.org/paper/The-determination-of-next-best-views-Connolly/3d18fbad2b81c108955e8c293a51fe985b0e127e"}],"as_of":"","related_ids":["active-perception","active-exploration","occupancy-grid-map","octomap","occlusion","wrist-camera"],"name":"Next-Best-View Planning","alt":"下一最佳视角","abbr":"NBV","aliases":["NBV","NBV Planning","Best Next View"],"one_liner":"Deciding where a sensor should look next, given what it has already seen, to gain the most useful information.","explanation":"Next-best-view (NBV) planning addresses this problem: given a partial set of observations so far, choose the next camera or scanner pose that will add the most new information for the least cost. It traces back to a 1985 ICRA paper by Connolly, who evaluated candidate viewpoints using a partial octree model. A common modern approach scores each candidate viewpoint by how many unknown voxels it would newly reveal in an occupancy map, combined with the cost of moving there. It is used for 3D-scanning objects, exploring a scene, and grasping occluded objects, and is one of the central problems in active perception.","example":"A robot arm with a wrist-mounted camera scans an unknown object, moving at each step to the pose that reveals the most previously unobserved area, until the 3D model is complete.","related":["Active Perception","Active Exploration","Occupancy Grid Map","OctoMap","Occlusion","Wrist Camera"]},{"id":"interactive-perception","category":"perception","sec":11,"tier":3,"sources":[{"title":"arXiv 1604.03670: Interactive Perception: Leveraging Action in Perception and Perception in Action（IEEE T-RO 2017）","url":"https://arxiv.org/abs/1604.03670"}],"as_of":"","related_ids":["active-perception","articulation-estimation","affordance","perception-action-loop","non-prehensile-manipulation","embodied-perception"],"name":"Interactive Perception","alt":"交互式感知","abbr":"","aliases":[],"one_liner":"A robot actively pushes, pokes, or pulls objects, using the resulting changes to understand its environment.","explanation":"Interactive perception means a robot changes the scene through physical contact — pushing, poking, pulling, picking something up to look at it — and uses the sensory changes caused by that action to understand the environment. A 2017 survey by Bohg, Hausman, Brock, Kragic, and colleagues in IEEE T-RO systematically laid out this direction, with two core ideas: interaction produces signals that passive observation can’t get, and knowing what action you just took lets you better predict and interpret the signal that follows. It differs from active perception, which usually just moves the camera or changes viewpoint without changing the environment; interactive perception actually changes the environment by acting on it. Typical uses include segmenting objects that are pressed together, estimating the joint axis of a drawer or cabinet door, and judging physical properties like an object’s mass and friction.","example":"Several blocks sit pressed together, and looking at the image alone can’t tell their boundaries apart; the robot gives them a push, and the pixels that move together belong to the same block, letting it segment them. Tracking a handle’s motion while pulling open a drawer can likewise estimate the drawer’s sliding direction.","related":["Active Perception","Articulation Estimation","Affordance","Perception-Action Loop","Non-prehensile Manipulation","Embodied Perception"]},{"id":"python-and-c-plus-plus","category":"software","sec":0,"tier":1,"sources":[{"title":"Python official site","url":"https://www.python.org/"},{"title":"ROS 2 Client Libraries (rclcpp / rclpy)","url":"https://docs.ros.org/en/jazzy/Concepts/Basic/About-Client-Libraries.html"}],"as_of":"","related_ids":["pytorch","robot-operating-system-2","software-development-kit","torchscript-libtorch","real-time-control"],"name":"Python and C++","alt":"Python 与 C++（具身常用编程语言）","abbr":"","aliases":["Python","C++"],"one_liner":"The two languages embodied AI relies on most: Python for models and experiments, C++ for real-time control.","explanation":"Embodied AI development runs on these two languages almost everywhere, with a fairly clear division of labor. Python is quick to write and has a rich ecosystem: deep-learning frameworks (PyTorch, JAX), simulation training (Isaac Lab, MuJoCo's Python bindings), data processing, and VLA model training are nearly all done in it. C++ runs efficiently with predictable latency, which suits the parts that need hard real-time guarantees: low-level motor control, whole-body control, motion planning, sensor drivers, high-performance ROS nodes, and most robot manufacturers' SDKs. A common pattern is to train a policy in Python, then export the model to be loaded and run in C++, or to have a Python layer on top call down into a C++ library. Newcomers usually get comfortable with Python first, then pick up C++ as their direction demands it.","example":"A quadruped walking policy is trained in Python inside Isaac Lab, exported to ONNX (a general-purpose format for exchanging neural network models), and then run at a fixed rate inside a C++ control program that sends out motor commands.","related":["PyTorch","Robot Operating System 2","Software Development Kit","TorchScript (torch.jit) / LibTorch","Real-Time Control"]},{"id":"ubuntu-linux","category":"software","sec":0,"tier":2,"sources":[{"title":"About Ubuntu","url":"https://ubuntu.com/about"},{"title":"ROS 2 Jazzy Installation","url":"https://docs.ros.org/en/jazzy/Installation.html"}],"as_of":"2024-05","related_ids":["ros-distribution","robot-operating-system-2","nvidia-jetpack-sdk","windows-subsystem-for-linux-2","docker","preempt-rt"],"name":"Ubuntu Linux","alt":"Ubuntu","abbr":"","aliases":[],"one_liner":"The Linux distribution maintained by Canonical that serves as the default operating system for robotics development.","explanation":"Ubuntu is an open-source Linux distribution maintained by the UK company Canonical, first released in 2004, with a new Long-Term Support (LTS) version every two years in April. It is the de facto standard operating system for robotics development: every ROS 2 distribution is tied to a specific Ubuntu LTS (Humble to 22.04, Jazzy to 24.04), NVIDIA Jetson's JetPack system is also Ubuntu-based, and Isaac Sim and most GPU training servers list it as their primary supported platform. The most common pitfall for newcomers setting up an environment is a mismatch between the installed Ubuntu version and the ROS version they're trying to use.","example":"To install ROS 2 Humble, first install Ubuntu 22.04; Windows users can practice by first setting up an Ubuntu installation through WSL2.","related":["ROS Distribution","Robot Operating System 2","NVIDIA JetPack SDK","Windows Subsystem for Linux 2 (WSL2)","Docker","PREEMPT_RT"]},{"id":"windows-subsystem-for-linux-2","category":"software","sec":0,"tier":3,"sources":[{"title":"What is Windows Subsystem for Linux (Microsoft Learn)","url":"https://learn.microsoft.com/en-us/windows/wsl/about"},{"title":"Comparing WSL Versions (Microsoft Learn)","url":"https://learn.microsoft.com/en-us/windows/wsl/compare-versions"},{"title":"CUDA on WSL User Guide (NVIDIA)","url":"https://docs.nvidia.com/cuda/wsl-user-guide/index.html"}],"as_of":"","related_ids":["ubuntu-linux","robot-operating-system-2","docker","cuda","conda","secure-shell"],"name":"Windows Subsystem for Linux 2 (WSL2)","alt":"WSL2（Windows 下的 Linux 子系统）","abbr":"WSL2","aliases":["WSL 2","WSL"],"one_liner":"Microsoft's official way to run a full Linux environment inside Windows, without dual-booting.","explanation":"WSL2 is the second version of the Linux subsystem Microsoft provides for Windows 10/11, running a real Linux kernel inside a lightweight virtual machine, which gives it better system-call compatibility and faster file and build performance than the first version. It solves the problem of having only a Windows laptop when almost all robotics software runs on Ubuntu: once installed, you can use apt, Conda, and Docker directly, train models on an NVIDIA GPU through CUDA on WSL, and even display graphical tools like RViz. Its downsides are that USB device access, real-time performance, and network configuration aren't as smooth as native Linux, so connecting to a real robot for debugging often needs extra workarounds.","example":"Run wsl --install on a Windows laptop to set up Ubuntu 22.04, then install ROS 2 Humble inside it and try the classic turtlesim tutorial.","related":["Ubuntu Linux","Robot Operating System 2","Docker","CUDA","Conda","Secure Shell (SSH)"]},{"id":"github","category":"software","sec":0,"tier":1,"sources":[{"title":"GitHub Docs: What is GitHub?（Git 与 GitHub 的区别）","url":"https://docs.github.com/en/get-started/start-your-journey/what-is-github"}],"as_of":"","related_ids":["open-source-license","lerobot","arxiv-preprint","open-weight-model","hugging-face"],"name":"GitHub","alt":"GitHub（代码托管平台）","abbr":"","aliases":["GitHub Repository"],"one_liner":"The world's largest code-hosting platform, where nearly all paper code and open-source robotics projects live.","explanation":"GitHub is a code-hosting platform, acquired by Microsoft in 2018. It isn't the same thing as Git: Git is a version-control tool that records every change to code locally, while GitHub puts Git repositories in the cloud and adds collaboration features on top. Each project is a “repository,” containing code, documentation (a README), an open-source license, and issue discussions (Issues). Most paper code, model-training frameworks, simulation environments, and robot SDKs in embodied AI are open-sourced on GitHub, so after reading a paper, checking its official repository is the first step toward reproducing it. The typical workflow is to run git clone to pull a repository locally, install dependencies and download weights as the README describes, and run the example scripts; when something breaks, searching the Issues tab first often turns up someone who has already hit the same problem.","example":"Hugging Face's LeRobot code lives at github.com/huggingface/lerobot; after git clone, installing it is just a matter of following the README.","related":["Open-Source License (Apache 2.0 / MIT / Non-commercial)","LeRobot","arXiv Preprint","Open-weight Model","Hugging Face"]},{"id":"conda","category":"software","sec":0,"tier":2,"sources":[{"title":"conda documentation","url":"https://docs.conda.io/"}],"as_of":"","related_ids":["uv","docker","pytorch","cuda","package-mirror-sources-in-china"],"name":"Conda","alt":"Conda","abbr":"","aliases":["Anaconda","Miniconda","Miniforge","conda environment"],"one_liner":"A widely used Python package and virtual-environment manager that isolates dependencies between projects.","explanation":"Conda is an open-source package- and environment-management tool, originally released by the Anaconda company. It can create a separate environment for each project, each with its own versions of Python and libraries that don’t interfere with one another; beyond Python packages, it can also install non-Python dependencies such as the CUDA toolkit and compilers. Anaconda is the full distribution that bundles a large set of scientific-computing packages; Miniconda is a stripped-down version with just conda itself; Miniforge is the community edition, defaulting to downloads from the conda-forge channel. Dependency versions in embodied AI projects often conflict with each other (a simulator, for instance, might require a specific Python version), so most open-source install instructions start with creating a conda environment.","example":"conda create -n lerobot python=3.10, then conda activate lerobot, then pip install -e . as the project’s instructions describe.","related":["uv","Docker","PyTorch","CUDA","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)"]},{"id":"uv","category":"software","sec":0,"tier":3,"sources":[{"title":"uv documentation","url":"https://docs.astral.sh/uv/"},{"title":"astral-sh/uv GitHub","url":"https://github.com/astral-sh/uv"}],"as_of":"","related_ids":["conda","docker","openpi","package-mirror-sources-in-china","lerobot"],"name":"uv","alt":"uv","abbr":"","aliases":["Astral uv"],"one_liner":"An extremely fast Python package and virtual-environment manager written in Rust.","explanation":"uv is a Python package manager open-sourced in 2024 by Astral (also the maker of the Ruff linter), implemented in Rust. It folds the jobs of pip (installing packages), virtualenv (creating virtual environments), pip-tools (locking versions), and pyenv (managing Python versions) into a single command, and resolves dependencies and installs packages far faster than pip. A project declares its dependencies in pyproject.toml, and uv generates a uv.lock file that pins exact versions, so the same environment can be reproduced on a different machine. Open-source embodied-AI code tends to have many dependencies and frequent version conflicts, and more and more projects are switching to uv to manage their environments. A mirror can be configured for faster access from mainland China.","example":"Following the openpi repository's instructions, clone the code and run uv sync once to install every dependency, then use uv run to launch the training script.","related":["Conda","Docker","openpi (Physical Intelligence)","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)","LeRobot"]},{"id":"package-mirror-sources-in-china","category":"software","sec":0,"tier":2,"sources":[{"title":"清华大学开源软件镜像站 PyPI 帮助","url":"https://mirrors.tuna.tsinghua.edu.cn/help/pypi/"},{"title":"阿里云开源镜像站","url":"https://developer.aliyun.com/mirror/"}],"as_of":"","related_ids":["conda","docker","uv","hugging-face-mirror","modelscope","ubuntu-linux"],"name":"Package Mirror Sources in China (pip / conda / apt / Docker mirrors)","alt":"国内镜像源（换源：清华 TUNA / 阿里云等）","abbr":"","aliases":["mirror source","Tsinghua mirror","Aliyun mirror","switching sources"],"one_liner":"Redirecting pip, conda, apt, and similar tools to a Chinese mirror instead of the slow or blocked default servers.","explanation":"Tools like pip, conda, apt, and Docker default to downloading packages from servers overseas, which are often slow or time out entirely from within China. Universities and cloud providers maintain synced copies — such as the mirror hosted by Tsinghua University's TUNA student association, plus mirrors run by Alibaba Cloud, USTC, and Shanghai Jiao Tong University — and pointing a tool's configured download URL at one of these is what's called “switching sources” (换源). Embodied-AI environments have a lot of dependencies (PyTorch, ROS, simulators, assorted Python packages), so switching sources is nearly always the first step in setting one up. A few caveats: mirror syncing lags behind, so the very latest release may not be available yet; multiple Chinese Docker Hub mirrors reportedly stopped operating starting in 2024, so pulling images may require finding another working source; and Hugging Face models have their own separate options, the Hugging Face mirror and ModelScope.","example":"Running pip config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simple switches pip's default source to the Tsinghua mirror.","related":["Conda","Docker","uv","Hugging Face Mirror (hf-mirror.com)","ModelScope","Ubuntu Linux"]},{"id":"cmake","category":"software","sec":0,"tier":2,"sources":[{"title":"CMake 官网","url":"https://cmake.org/"}],"as_of":"","related_ids":["colcon","ament","catkin","cross-compilation","python-and-c-plus-plus"],"name":"CMake","alt":"CMake","abbr":"","aliases":["CMakeLists.txt"],"one_liner":"A cross-platform build-configuration tool for C/C++ projects that spells out exactly how to compile them.","explanation":"CMake is an open-source build-system generator led by Kitware. C/C++ code has to be compiled and linked to become an executable program, and once a project has many files, something needs to tell the compiler what to build first and which libraries it depends on. CMake reads a project’s CMakeLists.txt and generates the build files for a given platform (such as a Makefile or Ninja file), which then invoke the actual compiler. A large amount of low-level robotics code is C++; ROS 1’s catkin and ROS 2’s ament_cmake are both built on top of CMake, and libraries such as Pinocchio and OpenCV also need it when installed from source.","example":"The common three-step source install for a library: mkdir build && cd build, cmake .., make -j8.","related":["colcon","ament","catkin","Cross-compilation","Python and C++"]},{"id":"docker","category":"software","sec":0,"tier":2,"sources":[{"title":"Docker Docs: What is Docker?","url":"https://docs.docker.com/get-started/docker-overview/"}],"as_of":"","related_ids":["conda","ubuntu-linux","cuda","robot-operating-system-2","inference-deployment","autodl"],"name":"Docker","alt":"Docker","abbr":"","aliases":["Container","Image","Dockerfile"],"one_liner":"Packaging a program together with its runtime environment into a container that runs the same way on any machine.","explanation":"Docker is a container platform from Docker, Inc., open-sourced in 2013. It packages a program together with its system libraries, dependencies, and configuration into an “image”; running that image produces a “container,” which shares the host machine’s kernel and is therefore much lighter weight than a virtual machine. An image is described by a Dockerfile script, and can be pushed to a registry for others to pull. Robotics development frequently runs into problems like a ROS version being tied to a specific Ubuntu version, or a simulator with complicated dependencies; Docker lets different environments run side by side on the same machine, and paired with the NVIDIA Container Toolkit, a container can use the GPU too. Many simulators and deployment toolchains ship an official Docker image.","example":"On an Ubuntu 24.04 host, pulling the official ROS 2 Humble image and building and running a driver package that needs a 22.04 environment inside the container.","related":["Conda","Ubuntu Linux","CUDA","Robot Operating System 2","Inference Deployment","AutoDL"]},{"id":"cuda","category":"software","sec":0,"tier":1,"sources":[{"title":"NVIDIA CUDA Toolkit","url":"https://developer.nvidia.com/cuda-toolkit"},{"title":"CUDA Programming Guide（NVIDIA 官方文档）","url":"https://docs.nvidia.com/cuda/cuda-programming-guide/"},{"title":"PyTorch Forums: Cuda versioning and pytorch compatibility（PyTorch 二进制包自带 CUDA 运行库，只需驱动支持）","url":"https://discuss.pytorch.org/t/cuda-versioning-and-pytorch-compatibility/189777"}],"as_of":"","related_ids":["cuda-deep-neural-network-library","pytorch","nvidia-tensorrt","gpu-memory","nvidia"],"name":"CUDA","alt":"CUDA","abbr":"CUDA","aliases":["Compute Unified Device Architecture","CUDA Toolkit"],"one_liner":"NVIDIA's general-purpose computing platform for GPUs, which is what lets deep learning run on a graphics card.","explanation":"CUDA is NVIDIA's parallel computing platform and programming model, including a compiler, runtime libraries, and math libraries, that lets developers call an NVIDIA GPU directly for general-purpose computation from languages such as C/C++. Deep learning's huge number of matrix operations get their speed from GPU parallelism, and frameworks such as PyTorch call into CUDA and its companion libraries (such as cuDNN, which implements neural-network operators) whenever they run on an NVIDIA card. Newcomers usually meet it first while setting up their environment: the GPU build of PyTorch installed via pip bundles its own matching CUDA runtime libraries, so it works as long as the graphics driver is new enough to support that CUDA version — too old a driver either errors out or falls back to CPU only. Only when compiling a custom CUDA extension yourself does the locally installed CUDA Toolkit version need to match PyTorch's. GPU parallelism in simulation (such as Isaac Lab) and deployment acceleration (such as TensorRT) are both built on top of CUDA too.","example":"Running torch.cuda.is_available() in Python returns True only when PyTorch can actually use the GPU.","related":["CUDA Deep Neural Network Library (cuDNN)","PyTorch","NVIDIA TensorRT","GPU Memory (VRAM)","NVIDIA"]},{"id":"secure-shell","category":"software","sec":0,"tier":2,"sources":[{"title":"OpenSSH","url":"https://www.openssh.com/"},{"title":"RFC 4251: The Secure Shell (SSH) Protocol Architecture","url":"https://www.rfc-editor.org/rfc/rfc4251"}],"as_of":"","related_ids":["ubuntu-linux","autodl","nvidia-jetson","docker","host-computer"],"name":"Secure Shell (SSH)","alt":"SSH 远程登录","abbr":"SSH","aliases":["ssh","OpenSSH"],"one_liner":"An encrypted network protocol for logging into and controlling another computer's command line remotely.","explanation":"SSH (Secure Shell) is an encrypted remote-login protocol, proposed in 1995 by Finland's Tatu Ylönen and later standardized by the IETF; the most widely used implementation is the open-source OpenSSH. It lets you securely log into another machine's terminal from your own and run commands there, as well as transfer files (via scp or rsync) and forward ports. It's used constantly in embodied-AI development: logging into a robot's onboard Jetson to edit code, logging into a cloud GPU server to run training, and remote development in VS Code or Cursor is also built on SSH. A common setup is key-based login combined with aliases in ~/.ssh/config.","example":"Once the robot and a laptop are on the same Wi-Fi network, running ssh unitree@192.168.123.164 logs into the onboard computer to start the policy-inference program.","related":["Ubuntu Linux","AutoDL","NVIDIA Jetson","Docker","Host Computer"]},{"id":"robot-operating-system","category":"software","sec":1,"tier":1,"sources":[{"title":"ROS official site","url":"https://www.ros.org/"},{"title":"ROS Wiki: Introduction","url":"https://wiki.ros.org/ROS/Introduction"}],"as_of":"2025-05","related_ids":["robot-operating-system-2","node","topic","ros-distribution","rviz-rviz2","moveit-motion-planning-framework"],"name":"Robot Operating System","alt":"机器人操作系统","abbr":"ROS","aliases":["ROS","ROS 1"],"one_liner":"The most widely used open-source middleware framework for robot development — not actually an operating system.","explanation":"Despite its name, ROS isn’t really an operating system; it’s a set of open-source middleware and tools that runs on top of Linux (usually Ubuntu). It originated at Stanford and was later driven forward by Willow Garage. ROS breaks robot software into nodes that communicate through mechanisms like topics and services, and it ships with a large collection of ready-made packages: sensor drivers, coordinate transforms (TF), visualization (RViz), data recording (rosbag), navigation, motion planning, and more. Its value is that people don’t have to reinvent these pieces — a great many robots in both academia and industry expose a ROS interface. Noetic, the final release of ROS 1, reached end of maintenance in 2025; new projects should use ROS 2, though a large amount of existing code and many tutorials are still ROS 1.","example":"roslaunch starts the camera driver, the arm driver, and the MoveIt planning node all at once, and dragging a target pose around in RViz then makes the arm move to it.","related":["Robot Operating System 2","Node (ROS)","Topic (ROS)","ROS Distribution","RViz / RViz2","MoveIt Motion Planning Framework"]},{"id":"robot-operating-system-2","category":"software","sec":1,"tier":1,"sources":[{"title":"ROS 2 Documentation","url":"https://docs.ros.org/en/jazzy/index.html"},{"title":"ROS 2 Distributions（各发行版发布与停更时间）","url":"https://docs.ros.org/en/jazzy/Releases.html"},{"title":"Lyrical Luth Supported Platforms（平台支持等级与默认中间件）","url":"https://docs.ros.org/en/lyrical/Releases/lyrical/supported-platforms.html"}],"as_of":"2026-09","related_ids":["robot-operating-system","data-distribution-service","ros-2-quality-of-service","node","topic","colcon"],"name":"Robot Operating System 2","alt":"ROS 2","abbr":"ROS 2","aliases":["ROS 2","ROS2"],"one_liner":"The ground-up rewrite of ROS that defaults to DDS for its underlying communication; the standard choice for new projects.","explanation":"ROS 2 is a redesign of ROS 1, led by Open Robotics. Its biggest change is switching the default communication layer to DDS (Data Distribution Service), an industrial middleware standard, which removes the need for a central master node (roscore) and lets nodes discover each other automatically; it also adds Quality of Service (QoS) settings and better real-time and security support. Ubuntu and Windows are officially first-tier supported; macOS has to be built from source. It keeps the same core concepts — nodes, topics, services, actions — but its interfaces and build tools (colcon, ament) differ from ROS 1, so code isn't directly portable between them. The long-term support (LTS) releases still maintained are Humble, Jazzy, and Lyrical, released in May 2026 and supported through 2031. Nav2 navigation, MoveIt 2, ros2_control, and many humanoid and quadruped robot SDKs have all migrated to ROS 2.","example":"After installing ROS 2 on Ubuntu, running ros2 topic list lists every topic currently active in the system.","related":["Robot Operating System","Data Distribution Service","ROS 2 Quality of Service (QoS)","Node (ROS)","Topic (ROS)","colcon"]},{"id":"ros-distribution","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 Documentation: Distributions","url":"https://docs.ros.org/en/rolling/Releases.html"},{"title":"ROS Wiki: Distributions","url":"https://wiki.ros.org/Distributions"}],"as_of":"2026-09","related_ids":["robot-operating-system-2","robot-operating-system","ubuntu-linux","fishros-one-click-installer","ros1-bridge"],"name":"ROS Distribution","alt":"ROS 发行版","abbr":"","aliases":["Noetic","Humble","Jazzy","Kilted","Lyrical Luth","ROS release"],"one_liner":"A versioned, fixed bundle of ROS core libraries and packages, with each release tied to a specific Ubuntu version.","explanation":"A ROS distribution is a set of ROS core libraries and packages that have been tested together and pinned to fixed versions, each with its own codename, released in alphabetical order. ROS 2 ships a new distribution every May; even-numbered years get a Long-Term Support (LTS) release, supported for about five years, each tied to the Ubuntu LTS current at the time: Humble pairs with 22.04, Jazzy with 24.04, Kilted Kaiju is a non-LTS release, and 2026's Lyrical Luth is the newest LTS; there's also a continuously updated Rolling development distribution. Noetic, the final ROS 1 release, reached end of maintenance in May 2025. Picking a distribution means making sure the OS, the ROS release, and any third-party packages all line up — a mismatch is a frequent source of dependency failures — and a new project usually goes with whichever LTS is current.","example":"To install ROS 2 Humble, first install Ubuntu 22.04, then activate the environment with source /opt/ros/humble/setup.bash.","related":["Robot Operating System 2","Robot Operating System","Ubuntu Linux","FishROS One-Click Installer","ros1_bridge"]},{"id":"fishros-one-click-installer","category":"software","sec":1,"tier":3,"sources":[{"title":"fishros/install (GitHub)","url":"https://github.com/fishros/install"},{"title":"鱼香ROS 社区","url":"https://fishros.org.cn/forum/"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","ros-distribution","rosdep","ubuntu-linux","package-mirror-sources-in-china"],"name":"FishROS One-Click Installer","alt":"鱼香 ROS 一键安装","abbr":"","aliases":["fishros","Xiaoyu one-click install"],"one_liner":"A one-line install script from the Chinese ROS community FishROS that sets up ROS and its environment.","explanation":"The FishROS one-click installer is an open-source install script maintained by “FishROS” (鱼香ROS), a Chinese-language ROS community whose author goes by “Xiaoyu” (小鱼). Newcomers installing ROS on Ubuntu commonly get stuck on steps like blocked software sources, expired keys, or a failed rosdep initialization (rosdep being the tool that auto-installs dependencies); this script bundles switching to domestic mirror sources, installing any ROS 1 or ROS 2 distribution, installing the domestic rosdepc, configuring Docker, and more, into an interactive menu — just pick a number and follow the prompts. It's one of the most common shortcuts for getting started with ROS within China, though it's still worth reading through the official installation steps afterward, so there's somewhere to start troubleshooting if something goes wrong later.","example":"On a fresh Ubuntu 22.04 install, run wget http://fishros.com/install -O fishros && . fishros, then pick “one-click install ROS” from the menu and select the Humble desktop version.","related":["Robot Operating System","Robot Operating System 2","ROS Distribution","rosdep","Ubuntu Linux","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)"]},{"id":"node","category":"software","sec":1,"tier":1,"sources":[{"title":"ROS 2 Documentation: Understanding nodes","url":"https://docs.ros.org/en/jazzy/Tutorials/Beginner-CLI-Tools/Understanding-ROS2-Nodes/Understanding-ROS2-Nodes.html"},{"title":"ROS Wiki: Nodes","url":"https://wiki.ros.org/Nodes"}],"as_of":"","related_ids":["topic","service","action","robot-operating-system-2","robot-operating-system","ros-client-library"],"name":"Node (ROS)","alt":"节点","abbr":"","aliases":["ROS Node"],"one_liner":"A single-purpose program unit in ROS; a whole system is made of many nodes talking to each other.","explanation":"A node is the most basic unit of execution in ROS (Robot Operating System); typically, one node is responsible for one job — reading a camera, running object detection, planning a path, driving a motor. Nodes don’t call each other’s functions directly; instead they exchange messages through topics (a continuous publish/subscribe stream), services (a single request and response), and actions (a longer-running task with progress feedback). Splitting the system up this way means modules can be developed, swapped, and debugged independently; one node crashing doesn’t necessarily bring the whole system down, and different nodes can easily be spread across multiple computers. In ROS 2, nodes are written using rclpy (Python) or rclcpp (C++), and ros2 node list is the usual command to see which nodes are currently running.","example":"A camera-driver node publishes images; a detection node subscribes to the images and publishes an object’s position; a control node uses that to send the arm to grasp it.","related":["Topic (ROS)","Service (ROS)","Action (ROS Action)","Robot Operating System 2","Robot Operating System","ROS Client Library (rclcpp / rclpy)"]},{"id":"topic","category":"software","sec":1,"tier":1,"sources":[{"title":"ROS 2 Documentation: Understanding topics","url":"https://docs.ros.org/en/jazzy/Tutorials/Beginner-CLI-Tools/Understanding-ROS2-Topics/Understanding-ROS2-Topics.html"},{"title":"ROS Wiki: Topics","url":"https://wiki.ros.org/Topics"}],"as_of":"","related_ids":["node","service","action","message","publish-subscribe","cmd-vel-topic"],"name":"Topic (ROS)","alt":"话题","abbr":"","aliases":["ROS Topic"],"one_liner":"A named channel for continuously streaming data between ROS nodes, working on a publish/subscribe model.","explanation":"A topic is the most commonly used communication mechanism in ROS. Every topic has a name and a fixed message type; publisher nodes continuously send messages to the topic, and subscriber nodes receive them, with neither side needing to know who the other is — a single topic can have multiple publishers and multiple subscribers. This suits a continuous stream of data, such as images, lidar point clouds, joint states, or velocity commands. By contrast, a service is a one-shot request-response call (a client sends a request, the server processes it and replies with a result), suited to a query or a single triggered action; an action is for a longer-running task that needs progress feedback partway through, or the ability to be cancelled, such as navigating to a point. Together these three form the basis of ROS communication, and a newcomer should get comfortable with topics first.","example":"A base-driver node subscribes to the /cmd_vel topic (a Twist-type velocity command); a keyboard teleop node publishes velocities to that same topic, and the base moves accordingly.","related":["Node (ROS)","Service (ROS)","Action (ROS Action)","Message (msg / srv / action interface)","Publish-Subscribe","cmd_vel Topic (geometry_msgs/Twist velocity command)"]},{"id":"publish-subscribe","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 Documentation: Understanding topics","url":"https://docs.ros.org/en/humble/Tutorials/Beginner-CLI-Tools/Understanding-ROS2-Topics/Understanding-ROS2-Topics.html"},{"title":"Publish–subscribe pattern - Wikipedia","url":"https://en.wikipedia.org/wiki/Publish%E2%80%93subscribe_pattern"}],"as_of":"","related_ids":["topic","node","service","data-distribution-service","ros-2-quality-of-service","middleware"],"name":"Publish-Subscribe","alt":"发布/订阅","abbr":"Pub/Sub","aliases":["pub/sub","publisher/subscriber"],"one_liner":"Senders publish messages on a topic and receivers subscribe to it, with neither side needing to know the other.","explanation":"Publish-subscribe is a communication pattern: publishers send messages to a named “topic,” and every receiver subscribed to that topic gets them, without the publisher needing to know who's listening or subscribers needing to know who's sending. That keeps modules loosely coupled — adding a new subscriber, say for recording or visualization, doesn't require touching existing code. ROS / ROS 2's topic mechanism is a textbook example: a camera driver node publishes an image topic, and a perception node and RViz can both subscribe to it at once. It suits continuously flowing data such as sensor readings and state; a one-off request-response interaction should use a service instead, and a long-running task should use an action. ROS 2 implements this over DDS underneath, with QoS settings for tuning reliability and buffer depth.","example":"A lidar node publishes the /scan topic, and both a mapping node and a rosbag recording session subscribe to it at the same time.","related":["Topic (ROS)","Node (ROS)","Service (ROS)","Data Distribution Service","ROS 2 Quality of Service (QoS)","Middleware"]},{"id":"message","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 文档：About interfaces","url":"https://docs.ros.org/en/rolling/Concepts/Basic/About-Interfaces.html"}],"as_of":"","related_ids":["topic","service","action","publish-subscribe","robot-operating-system-2","cmd-vel-topic"],"name":"Message (msg / srv / action interface)","alt":"消息","abbr":"msg","aliases":["msg","ROS message","interface type","interface"],"one_liner":"The agreed-upon data format that ROS nodes use to exchange information.","explanation":"Before ROS nodes can exchange data, they need to agree on what that data looks like — that agreement is a message, collectively called an interface in ROS 2. There are three kinds of interface files: .msg defines the data structure carried on a topic; .srv defines a service's request and response, separated by three dashes; and .action defines a goal, result, and intermediate feedback for an action. These files only declare field types and names; the build tools generate the corresponding C++ and Python code automatically. Commonly used standard message packages include std_msgs, geometry_msgs (e.g. the velocity command Twist, or the pose type Pose), and sensor_msgs (e.g. the image type Image, joint state JointState, or point cloud PointCloud2). Reusing a standard message instead of defining a custom one keeps a node easy to connect to existing tools.","example":"Running ros2 interface show geometry_msgs/msg/Twist shows it's made of a linear and an angular three-dimensional vector, which is what base velocity commands use.","related":["Topic (ROS)","Service (ROS)","Action (ROS Action)","Publish-Subscribe","Robot Operating System 2","cmd_vel Topic (geometry_msgs/Twist velocity command)"]},{"id":"cmd-vel-topic","category":"software","sec":1,"tier":3,"sources":[{"title":"geometry_msgs/msg/Twist.msg (ros2/common_interfaces)","url":"https://github.com/ros2/common_interfaces/blob/rolling/geometry_msgs/msg/Twist.msg"},{"title":"Nav2 Documentation","url":"https://docs.nav2.org/"}],"as_of":"","related_ids":["topic","message","robot-operating-system","robot-operating-system-2","ros-2-navigation-stack","differential-drive-kinematics"],"name":"cmd_vel Topic (geometry_msgs/Twist velocity command)","alt":"cmd_vel 速度指令话题（Twist 消息）","abbr":"","aliases":["/cmd_vel"],"one_liner":"ROS's de facto standard topic for sending a mobile robot linear and angular velocity commands.","explanation":"cmd_vel is a topic name that has become a convention in ROS / ROS 2 (a topic being a publish-subscribe data channel), carrying messages of type geometry_msgs/Twist, which holds a three-component linear velocity and a three-component angular velocity. Navigation stacks, teleoperation, and keyboard-control programs all publish to /cmd_vel, and a base-driver node subscribes and converts it into left/right wheel speeds or a steering angle. Its purpose is to decouple high-level decision-making from the low-level base: differential-drive, mecanum, and Ackermann bases can all expose this same interface externally. Nav2 publishes this topic by default too, and a timestamped variant, TwistStamped, has become more common in recent years.","example":"Publishing commands to /cmd_vel with teleop_twist_keyboard drives a TurtleBot forward at 0.2 m/s while turning in place at 0.5 rad/s.","related":["Topic (ROS)","Message (msg / srv / action interface)","Robot Operating System","Robot Operating System 2","ROS 2 Navigation Stack (Nav2)","Differential Drive Kinematics"]},{"id":"service","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 Docs: About services","url":"https://docs.ros.org/en/rolling/Concepts/Basic/About-Services.html"},{"title":"ROS 2 Tutorial: Understanding services","url":"https://docs.ros.org/en/humble/Tutorials/Beginner-CLI-Tools/Understanding-ROS2-Services/Understanding-ROS2-Services.html"}],"as_of":"","related_ids":["topic","action","node","message","publish-subscribe","robot-operating-system-2"],"name":"Service (ROS)","alt":"服务","abbr":"","aliases":["ROS service"],"one_liner":"ROS's request-response communication pattern: a client sends a request and the server sends back a result.","explanation":"A service is one of ROS's three basic communication patterns, alongside topics and actions. One node acts as the server, offering a service; other nodes act as clients, sending a request and waiting for a response, with the request and response formats declared in a .srv file. The difference from a topic is that a topic is a continuous, one-way stream of data suited to sensor readings, while a service is a single call-and-response suited to quick operations like “reset,” “open the gripper,” or “query status.” A task that takes a long time, or needs feedback partway through or the ability to be cancelled, should use an action instead.","example":"Running ros2 service call /reset_world std_srvs/srv/Empty from the command line resets a simulation; or a /gripper/open service can be written so that calling it once opens the gripper.","related":["Topic (ROS)","Action (ROS Action)","Node (ROS)","Message (msg / srv / action interface)","Publish-Subscribe","Robot Operating System 2"]},{"id":"action","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 Documentation: Understanding actions","url":"https://docs.ros.org/en/rolling/Tutorials/Beginner-CLI-Tools/Understanding-ROS2-Actions/Understanding-ROS2-Actions.html"}],"as_of":"","related_ids":["topic","service","node","message","ros-2-navigation-stack","moveit-motion-planning-framework"],"name":"Action (ROS Action)","alt":"动作","abbr":"","aliases":["ROS Action","actionlib"],"one_liner":"A ROS communication style for tasks that take a while, with progress feedback partway through and the option to cancel.","explanation":"Action is one of ROS’s three main communication mechanisms, alongside topics (continuously broadcast data) and services (a single request and response). It’s designed for tasks that take some time to finish: a client sends a goal, the server sends back feedback continuously while it executes, and returns a result when done, with the option to cancel partway through. In ROS 1 it’s implemented by the actionlib library; in ROS 2 it’s a built-in mechanism, built underneath from a combination of services and topics. Tasks such as navigating to a point or an arm executing a trajectory are usually exposed as an action interface.","example":"Nav2’s NavigateToPose action: send a goal point, and the robot reports its remaining distance while it walks, returning success or failure once it arrives.","related":["Topic (ROS)","Service (ROS)","Node (ROS)","Message (msg / srv / action interface)","ROS 2 Navigation Stack (Nav2)","MoveIt Motion Planning Framework"]},{"id":"ros-parameter-server","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS Wiki: Parameter Server","url":"http://wiki.ros.org/Parameter%20Server"},{"title":"ROS 2 文档：About parameters","url":"https://docs.ros.org/en/rolling/Concepts/Basic/About-Parameters.html"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","ros-master","node","launch-file","topic"],"name":"ROS Parameter Server","alt":"参数服务器","abbr":"","aliases":["Parameter Server","rosparam"],"one_liner":"A shared key-value store in ROS 1 that holds global configuration values any node can read or write.","explanation":"The Parameter Server is a mechanism in ROS 1: the ROS master (roscore) maintains a globally shared key-value dictionary, and nodes read and write entries in it over the network — things like control gains, a camera's topic name, or a robot's physical dimensions. Parameters are typically loaded in bulk from launch files or YAML files, and can also be viewed or changed from the command line with rosparam. It's meant for static configuration that rarely changes, not high-frequency data. ROS 2 removed the global parameter server; instead, each node declares and holds its own parameters, accessed via ros2 param, with callbacks that can fire when a parameter changes. Note that distributed machine learning has an unrelated concept with the same name, referring to servers that store model parameters during training — that's a different thing.","example":"A launch file uses rosparam to load pid.yaml, so the control node reads gains such as /arm_controller/kp on startup.","related":["Robot Operating System","Robot Operating System 2","ROS Master (roscore)","Node (ROS)","Launch File","Topic (ROS)"]},{"id":"launch-file","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 文档：Launch 教程","url":"https://docs.ros.org/en/rolling/Tutorials/Intermediate/Launch/Launch-Main.html"},{"title":"ROS Wiki: roslaunch","url":"http://wiki.ros.org/roslaunch"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","node","package","ros-parameter-server","topic"],"name":"Launch File","alt":"启动文件","abbr":"","aliases":["roslaunch","ros2 launch",".launch file"],"one_liner":"A ROS script that starts multiple nodes at once and configures their parameters.","explanation":"A robot system typically has to run a dozen or more nodes at the same time — camera drivers, state estimation, planning, control — and starting each one by hand in its own terminal is tedious and error-prone. A launch file writes down, in one place, which nodes to start, what parameters to pass them, how to remap their topics, and which namespace to put them in, so a single command brings the whole system up. ROS 1 uses XML-format .launch files with the roslaunch command; ROS 2 supports Python (.launch.py), XML, and YAML, launched with ros2 launch, and the Python form — which allows conditionals and loops — is the most common. Launch files are usually kept in a package's launch directory.","example":"A single command, ros2 launch realsense2_camera rs_launch.py, starts the RealSense camera driver and publishes color and depth image topics.","related":["Robot Operating System","Robot Operating System 2","Node (ROS)","Package (ROS)","ROS Parameter Server","Topic (ROS)"]},{"id":"ros-client-library","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS 2 Documentation: Client libraries","url":"https://docs.ros.org/en/rolling/Concepts/Basic/About-Client-Libraries.html"}],"as_of":"","related_ids":["robot-operating-system-2","robot-operating-system","node","ros-middleware-interface","ros-2-executor-and-callback-groups","python-and-c-plus-plus"],"name":"ROS Client Library (rclcpp / rclpy)","alt":"ROS 客户端库","abbr":"RCL","aliases":["rcl","rclcpp","rclpy","rospy","roscpp"],"one_liner":"The programming API used to write ROS nodes — rclcpp for C++, rclpy for Python.","explanation":"A client library is the API developers call directly when writing ROS programs — creating nodes, publishing and subscribing to topics, calling services, reading parameters, and so on. In the ROS 1 era, the main client libraries were roscpp (C++) and rospy (Python). ROS 2 switched to a layered design: at the bottom is rcl, a general-purpose library written in C, on top of which each language gets its own wrapper — the official ones are rclcpp for C++ and rclpy for Python, with community-maintained versions for Rust, Java, and others. This means the communication logic only has to be written once, and behavior stays consistent across languages. In general, control nodes with strict real-time requirements tend to use rclcpp, while rapid prototyping and hooking up deep learning models tend to use rclpy.","example":"In Python, import rclpy, subclass rclpy.node.Node to write a node, subscribe to a camera topic, and feed the images into a VLA model for inference.","related":["Robot Operating System 2","Robot Operating System","Node (ROS)","ROS Middleware Interface (RMW)","ROS 2 Executor & Callback Groups","Python and C++"]},{"id":"package","category":"software","sec":1,"tier":2,"sources":[{"title":"ROS 2 Documentation: Creating a package","url":"https://docs.ros.org/en/humble/Tutorials/Beginner-Client-Libraries/Creating-Your-First-ROS2-Package.html"}],"as_of":"","related_ids":["ros-workspace","colcon","rosdep","ament","node","launch-file"],"name":"Package (ROS)","alt":"功能包","abbr":"","aliases":["ROS package","package"],"one_liner":"The basic unit for organizing code in ROS — one package groups a set of related nodes and configuration.","explanation":"A package is the basic unit for organizing and distributing code in ROS / ROS 2. A package typically contains source code, launch files, message definitions, and configuration parameters, along with a package.xml manifest describing its name, version, and dependencies; a C++ package also has a CMakeLists.txt, and a Python package has a setup.py. In ROS 2, ros2 pkg create scaffolds a new package, and colcon builds an entire workspace at once. Splitting code into packages lets others reuse just your camera driver or navigation module on its own, and lets rosdep automatically install dependencies based on package.xml. For newcomers reading an open-source robotics project, looking at what packages exist and what each one is responsible for is the fastest way in.","example":"Running ros2 pkg create my_robot_bringup --build-type ament_python generates the skeleton of a Python package.","related":["ROS Workspace","colcon","rosdep","ament","Node (ROS)","Launch File"]},{"id":"ros-workspace","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS 2 Docs: Creating a workspace","url":"https://docs.ros.org/en/jazzy/Tutorials/Beginner-Client-Libraries/Creating-A-Workspace/Creating-A-Workspace.html"},{"title":"ROS Wiki: catkin workspaces","url":"http://wiki.ros.org/catkin/workspaces"}],"as_of":"","related_ids":["package","colcon","catkin","ament","rosdep","robot-operating-system-2"],"name":"ROS Workspace","alt":"ROS 工作空间","abbr":"","aliases":["catkin_ws","ros2_ws","colcon workspace"],"one_liner":"A directory convention for holding, building, and installing your own ROS packages.","explanation":"A workspace is ROS's directory convention for organizing code: a top-level src folder holds the source of packages (a package being ROS's smallest unit of code), and building it produces build (intermediate files), install or devel (the built output), and log directories. ROS 1 builds with catkin, and the workspace is conventionally named catkin_ws; ROS 2 builds with colcon, and it's usually called ros2_ws. After building, you have to source the corresponding setup script before the terminal can find the packages inside. Workspaces can be overlaid: you source the system-installed ROS (the underlay) first, then your own workspace on top, and if a package exists in both, the top one wins. The single most common beginner mistake is forgetting to source, or sourcing things in the wrong order.","example":"mkdir -p ~/ros2_ws/src, clone your code into src, run colcon build from ros2_ws, then source install/setup.bash — after that, ros2 run can find your own node.","related":["Package (ROS)","colcon","catkin","ament","rosdep","Robot Operating System 2"]},{"id":"colcon","category":"software","sec":1,"tier":2,"sources":[{"title":"colcon documentation","url":"https://colcon.readthedocs.io/"},{"title":"ROS 2 Documentation: Using colcon to build packages","url":"https://docs.ros.org/en/rolling/Tutorials/Beginner-Client-Libraries/Colcon-Tutorial.html"}],"as_of":"","related_ids":["ros-workspace","package","ament","catkin","cmake","rosdep"],"name":"colcon","alt":"colcon","abbr":"","aliases":["collective construction","colcon build"],"one_liner":"The recommended build tool for ROS 2, compiling an entire workspace with a single command.","explanation":"colcon is the officially recommended command-line build tool for ROS 2; its name comes from “collective construction.” A ROS workspace usually contains many packages with dependencies on each other, and colcon automatically works out the build order and calls each package’s own build system in turn or in parallel — CMake/ament_cmake for C++ packages, setuptools for Python packages — installing the results into an install directory. It replaced catkin_make and catkin_tools from the ROS 1 era. After building, you need to source install/setup.bash before the terminal can find the newly built packages.","example":"Running colcon build --symlink-install in ~/ros2_ws, then source install/setup.bash, followed by ros2 run to launch a node you wrote yourself.","related":["ROS Workspace","Package (ROS)","ament","catkin","CMake","rosdep"]},{"id":"ament","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS 2 Docs: About the build system","url":"https://docs.ros.org/en/rolling/Concepts/Advanced/About-Build-System.html"},{"title":"ROS 2 Design: ament","url":"https://design.ros2.org/articles/ament.html"}],"as_of":"","related_ids":["robot-operating-system-2","colcon","catkin","package","ros-workspace","cmake"],"name":"ament","alt":"ament（ROS 2 构建系统）","abbr":"","aliases":["ament_cmake","ament_python"],"one_liner":"ROS 2's package build system, with ament_cmake for C++ packages and ament_python for Python packages.","explanation":"ament is the build system ROS 2 adopted, succeeding catkin from ROS 1. It specifies how a ROS 2 package declares dependencies, gets compiled, gets installed, and gets found by other packages. The two most common build types are ament_cmake (CMake-based, for C++ packages, though it can also install Python scripts) and ament_python (based on Python's setuptools, for pure-Python packages), specified via build_type in package.xml. Day to day, developers don't call ament directly; instead they use the colcon build tool to compile an entire workspace at once, and colcon invokes the matching ament process for each package based on its type. Newcomers creating a package in ROS 2 just pick ament_cmake or ament_python with ros2 pkg create.","example":"Running ros2 pkg create my_controller --build-type ament_python generates a Python package with a setup.py, and colcon build then compiles the whole workspace.","related":["Robot Operating System 2","colcon","catkin","Package (ROS)","ROS Workspace","CMake"]},{"id":"catkin","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS Wiki: catkin","url":"http://wiki.ros.org/catkin"},{"title":"catkin_tools documentation","url":"https://catkin-tools.readthedocs.io/"}],"as_of":"2025-05","related_ids":["robot-operating-system","ament","colcon","ros-workspace","package","cmake"],"name":"catkin","alt":"catkin（ROS 1 构建系统）","abbr":"","aliases":["catkin_make","catkin build"],"one_liner":"ROS 1's official build system, using CMake to compile and manage packages in a workspace.","explanation":"catkin is ROS 1's official build system, extended from CMake; it compiles the packages under a workspace's src directory in dependency order and produces the build and devel directories along with environment scripts (source devel/setup.bash) that let the system find those packages. Two commands are commonly used: catkin_make, which comes with ROS and compiles all packages as a single CMake project, and catkin build, from catkin_tools, which builds each package in isolation and makes errors easier to trace back to their source. By ROS 2, catkin had been replaced by ament and colcon. Noetic, ROS 1's final distribution, reached end of maintenance in 2025, but a large amount of legacy code and tutorial material still uses catkin, so it's still something you run into when reading older projects.","example":"Run mkdir -p ~/catkin_ws/src, put packages into src, run catkin_make from within catkin_ws, then source devel/setup.bash.","related":["Robot Operating System","ament","colcon","ROS Workspace","Package (ROS)","CMake"]},{"id":"rosdep","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS 2 Docs: Managing Dependencies with rosdep","url":"https://docs.ros.org/en/jazzy/Tutorials/Intermediate/Rosdep.html"},{"title":"ROS Wiki: rosdep","url":"http://wiki.ros.org/rosdep"}],"as_of":"","related_ids":["ros-workspace","package","colcon","fishros-one-click-installer","package-mirror-sources-in-china","ubuntu-linux"],"name":"rosdep","alt":"rosdep","abbr":"","aliases":["rosdepc"],"one_liner":"A ROS command-line tool that installs a package's system dependencies automatically from its declarations.","explanation":"rosdep reads the dependencies declared in each package's package.xml across a workspace, looks them up against the mapping rules in the rosdistro repository to translate an abstract dependency name into the actual package name on the current system — say, a specific apt development library or a pip package — and then calls the matching package manager to install it. The first time it's used, you have to run sudo rosdep init and rosdep update, both of which need to reach rule files hosted on GitHub; since that often fails from mainland China's network, the Yuxiang ROS (鱼香ROS) community built rosdepc, a drop-in replacement that goes through a domestic mirror instead. Running rosdep before building is the standard first step after cloning someone else's ROS repository.","example":"After cloning someone else's ROS 2 repository into src, run rosdep install --from-paths src --ignore-src -r -y to pull in all dependencies, then colcon build.","related":["ROS Workspace","Package (ROS)","colcon","FishROS One-Click Installer","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)","Ubuntu Linux"]},{"id":"ros-master","category":"software","sec":1,"tier":3,"sources":[{"title":"ROS Wiki: Master","url":"http://wiki.ros.org/Master"},{"title":"ROS Wiki: roscore","url":"http://wiki.ros.org/roscore"}],"as_of":"","related_ids":["robot-operating-system","node","topic","ros-parameter-server","robot-operating-system-2","data-distribution-service"],"name":"ROS Master (roscore)","alt":"ROS 主节点","abbr":"","aliases":["roscore","rosmaster"],"one_liner":"The central process in ROS 1 that registers nodes and helps them find each other.","explanation":"The ROS Master is the central coordinating process in ROS 1 (the first generation of the Robot Operating System), typically started with the roscore command, which also brings up the Parameter Server and the logging node rosout alongside it. When a node — a process that does one job — starts up, it registers with the Master which topics (data channels) it publishes or subscribes to, and which services it offers. The Master's only job is to tell publishers and subscribers each other's addresses; after that, data flows directly between nodes, peer to peer, without passing through the Master. Its weakness is being a single point of failure: if the Master dies, no new node can join, and communicating across machines requires setting ROS_MASTER_URI on each one. ROS 2 removed the Master entirely, letting the underlying DDS layer discover nodes automatically instead.","example":"To run a robot in ROS 1: open one terminal and run roscore, then use rosrun to start the camera driver and mapping node; a second computer must point ROS_MASTER_URI at this machine to see its topics.","related":["Robot Operating System","Node (ROS)","Topic (ROS)","ROS Parameter Server","Robot Operating System 2","Data Distribution Service"]},{"id":"ros1-bridge","category":"software","sec":1,"tier":3,"sources":[{"title":"ros2/ros1_bridge (GitHub)","url":"https://github.com/ros2/ros1_bridge"}],"as_of":"2025-05","related_ids":["robot-operating-system","robot-operating-system-2","ros-master","topic","service","message"],"name":"ros1_bridge","alt":"ros1_bridge","abbr":"","aliases":["dynamic_bridge"],"one_liner":"A bridge program that lets ROS 1 and ROS 2 nodes exchange topics and services with each other.","explanation":"ros1_bridge is an officially maintained ROS 2 package that, when running, connects to both a ROS 1 Master and the ROS 2 network at once, converting message types between the two and forwarding them, covering both topics and services. It's mainly used during migration: an old driver still lives in ROS 1 while new algorithms are already written in ROS 2. Because it has to generate conversion code for every message type, it usually needs to be built from source in an environment with both generations of ROS installed, and any custom messages have to be built for both sides. ROS 1's final release, Noetic, reached end of life in May 2025, so the bridge is increasingly treated as a stopgap rather than a long-term solution.","example":"After running ros2 run ros1_bridge dynamic_bridge, the /scan topic published by a ROS 1 lidar driver becomes visible to Nav2 running on ROS 2.","related":["Robot Operating System","Robot Operating System 2","ROS Master (roscore)","Topic (ROS)","Service (ROS)","Message (msg / srv / action interface)"]},{"id":"guyuehome","category":"software","sec":1,"tier":3,"sources":[{"title":"古月居官网","url":"https://www.guyuehome.com/"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","fishros-one-click-installer","ros-2-navigation-stack","moveit-motion-planning-framework"],"name":"Guyuehome (Chinese ROS community)","alt":"古月居（ROS 中文社区）","abbr":"","aliases":["Guyuehome ROS community"],"one_liner":"An early, sizable Chinese-language learning community and tutorial site for ROS robotics.","explanation":"Guyuehome (古月居) is a Chinese-language robotics community founded by Hu Chunxu, who goes by “Guyue” (古月) online and is also the author of the book ROS Robot Development in Practice. The site centers on ROS (Robot Operating System, a robot software framework), and gathers blog posts, Q&A, and video courses ranging from beginner to advanced (some paid), covering ROS 1/ROS 2, SLAM, navigation, robot arms, and mobile robots. For Chinese-speaking readers, it's a common place to look up ROS errors or find introductory tutorials, and together with tools like the FishROS one-click installer, it forms part of the domestic Chinese ROS learning ecosystem.","example":"A newcomer works through Guyuehome's ROS 2 introductory course, starting by writing their first publisher and subscriber node.","related":["Robot Operating System","Robot Operating System 2","FishROS One-Click Installer","ROS 2 Navigation Stack (Nav2)","MoveIt Motion Planning Framework"]},{"id":"unified-robot-description-format","category":"software","sec":2,"tier":1,"sources":[{"title":"ROS Wiki: urdf","url":"https://wiki.ros.org/urdf"},{"title":"ROS 2 Documentation: URDF tutorials","url":"https://docs.ros.org/en/jazzy/Tutorials/Intermediate/URDF/URDF-Main.html"},{"title":"Gazebo Classic tutorial: URDF in Gazebo（URDF 只能描述单个机器人、不能表达闭链等局限）","url":"https://classic.gazebosim.org/tutorials?tut=ros_urdf&cat=connect_ros"}],"as_of":"","related_ids":["xacro","mjcf","link","revolute-joint","kinematic-tree","mesh-file"],"name":"Unified Robot Description Format","alt":"统一机器人描述格式","abbr":"URDF","aliases":["URDF","URDF File"],"one_liner":"An XML file format that describes a robot's links, joints, geometry, and mass, used as a standard across ROS.","explanation":"URDF is the XML format used across the ROS ecosystem to describe a robot model. In the file, each link element describes one rigid body's visual mesh, collision shape, and mass/inertia, and each joint element describes how two links connect — including the joint type (revolute, prismatic, fixed, and so on), its axis, and its range of motion — together forming a kinematic tree. With a URDF, RViz can display the robot, kinematics and dynamics libraries can compute forward and inverse kinematics from it, and simulators such as Isaac Sim, MuJoCo, and PyBullet can import it directly. Its limitations are that it can only describe a single robot, can only express a tree structure, and handles closed kinematic chains poorly; writing it by hand is also long and repetitive, which is why the Xacro macro system is commonly used to simplify it. MuJoCo has its own separate format, MJCF, which can encode more simulation parameters.","example":"Robot manufacturers typically ship a URDF and mesh files with the robot; importing that URDF into Isaac Lab is enough to start training a policy for that robot in simulation.","related":["Xacro (XML Macros)","MJCF (MuJoCo XML Format)","Link","Revolute Joint","Kinematic Tree","Mesh File (STL / OBJ / DAE / glTF)"]},{"id":"xacro","category":"software","sec":2,"tier":2,"sources":[{"title":"ROS Wiki: xacro","url":"http://wiki.ros.org/xacro"},{"title":"ros/xacro (GitHub)","url":"https://github.com/ros/xacro"}],"as_of":"","related_ids":["unified-robot-description-format","semantic-robot-description-format","mjcf","robot-state-publisher","robot-operating-system-2","mesh-file"],"name":"Xacro (XML Macros)","alt":"Xacro（XML 宏）","abbr":"Xacro","aliases":[".xacro file"],"one_liner":"An XML macro language that adds variables, macros, and includes to URDF, cutting down on repetitive robot descriptions.","explanation":"Xacro is an XML macro language and companion tool package in the ROS ecosystem, built specifically for writing URDF (the Unified Robot Description Format, which uses XML to describe a robot's links, joints, mass, and shape). Plain URDF has no variables or functions, so a symmetric dual-arm robot ends up needing nearly identical links and joints copied twice, and changing one dimension means editing many places. Xacro adds properties (constants), macros (parameterized templates), basic math, conditional expansion, and an include mechanism for combining multiple files. A finished .xacro file is expanded into plain URDF at launch time by the xacro command, then handed off to robot_state_publisher, RViz, Gazebo, or MoveIt. Most vendor-provided ROS robot description packages ship in .xacro form.","example":"Write an arm macro parameterized by prefix and mounting position for a dual-arm robot, then call it twice with left_ and right_ to generate all the links and joints for both arms.","related":["Unified Robot Description Format","Semantic Robot Description Format (SRDF)","MJCF (MuJoCo XML Format)","robot_state_publisher","Robot Operating System 2","Mesh File (STL / OBJ / DAE / glTF)"]},{"id":"semantic-robot-description-format","category":"software","sec":2,"tier":3,"sources":[{"title":"ROS Wiki: srdf","url":"http://wiki.ros.org/srdf"},{"title":"MoveIt Docs: URDF and SRDF","url":"https://moveit.picknik.ai/main/doc/examples/urdf_srdf/urdf_srdf_tutorial.html"}],"as_of":"","related_ids":["unified-robot-description-format","moveit-motion-planning-framework","xacro","self-collision-checking","motion-planning","mjcf"],"name":"Semantic Robot Description Format (SRDF)","alt":"语义机器人描述格式","abbr":"SRDF","aliases":["SRDF",".srdf file"],"one_liner":"An XML file that adds planning groups and collision exceptions on top of a URDF.","explanation":"SRDF is an XML format used by MoveIt (ROS's motion-planning framework) alongside URDF (the Unified Robot Description Format, which describes links, joints, and geometry). URDF only describes what the robot looks like; SRDF adds the semantic information motion planning needs: planning groups (for example, which joints make up the “arm” group versus the “gripper” group), end effectors, named poses (such as home or ready), virtual joints (which fix the robot to the world frame, or to a mobile base), passive joints, and pairs of links for which collision checking should be disabled. Links that are adjacent or can never physically touch don't need self-collision checks, which noticeably speeds up planning. It's usually generated automatically with the MoveIt Setup Assistant and then fine-tuned by hand.","example":"When generating a MoveIt configuration for a Franka arm, the SRDF defines a panda_arm group containing its 7 joints and a named pose called ready, and lists the adjacent links panda_link1 and panda_link2 as exempt from collision checking.","related":["Unified Robot Description Format","MoveIt Motion Planning Framework","Xacro (XML Macros)","Self-Collision Checking","Motion Planning","MJCF (MuJoCo XML Format)"]},{"id":"mjcf","category":"software","sec":2,"tier":2,"sources":[{"title":"MuJoCo 文档：XML Reference","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html"},{"title":"MuJoCo Menagerie GitHub","url":"https://github.com/google-deepmind/mujoco_menagerie"}],"as_of":"","related_ids":["mujoco","unified-robot-description-format","mujoco-menagerie","mesh-file","simulation-description-format","universal-scene-description"],"name":"MJCF (MuJoCo XML Format)","alt":"MJCF","abbr":"MJCF","aliases":["MuJoCo XML"],"one_liner":"The XML model format the MuJoCo simulator uses to describe robots and scenes.","explanation":"MJCF is MuJoCo's native model format, describing a simulated scene in XML: rigid bodies are nested into a tree by parent-child relationships, and each body carries its joints, geoms (a sphere, box, or referenced mesh file), and mass and inertia; the format can also specify actuators, sensors, tendons, contact parameters, and the simulation time step. Compared with URDF, it can express more simulation-specific detail and is written more compactly. MuJoCo can also read URDF, but files are usually converted to MJCF and then given extra parameters. MuJoCo Menagerie, maintained by DeepMind, collects a large number of well-tuned robot MJCF models, and tools like MJX and MuJoCo Playground use the format directly.","example":"The unitree_g1 directory in MuJoCo Menagerie provides an MJCF file for the Unitree G1, which can be loaded directly into simulation for reinforcement-learning locomotion training.","related":["MuJoCo (Multi-Joint dynamics with Contact)","Unified Robot Description Format","MuJoCo Menagerie","Mesh File (STL / OBJ / DAE / glTF)","Simulation Description Format (SDFormat)","Universal Scene Description (OpenUSD)"]},{"id":"mujoco-menagerie","category":"software","sec":2,"tier":3,"sources":[{"title":"mujoco_menagerie GitHub","url":"https://github.com/google-deepmind/mujoco_menagerie"}],"as_of":"","related_ids":[null,null,"mujoco-playground",null,null,null],"name":"MuJoCo Menagerie","alt":"MuJoCo Menagerie","abbr":"","aliases":[],"one_liner":"Google DeepMind's curated collection of high-quality MuJoCo robot models.","explanation":"MuJoCo Menagerie is an open-source model collection maintained by Google DeepMind, gathering MJCF models (MuJoCo's robot description files) for a large number of common robots — arms like the Franka and UR5e, quadrupeds and humanoids like the Unitree Go2 and G1, the ALOHA dual-arm platform, and dexterous hands like Shadow's. Every model has had its mass, inertia, joint limits, actuators, and collision geometry cleaned up, along with notes on its source and license. Converting a robot from URDF yourself often runs into issues like wrong collision geometry or missing parameters, and using Menagerie directly skips that work — projects like MuJoCo Playground and mink use it as their default source of models.","example":"Download the unitree_g1 directory from Menagerie and load scene.xml directly in the MuJoCo viewer to look at the Unitree G1.","related":["MuJoCo (Multi-Joint dynamics with Contact)","MJCF (MuJoCo XML Format)","MuJoCo Playground","mink (MuJoCo inverse kinematics)","Simulation Assets","Google DeepMind"]},{"id":"simulation-description-format","category":"software","sec":2,"tier":3,"sources":[{"title":"SDFormat 官网","url":"http://sdformat.org/"}],"as_of":"","related_ids":["unified-robot-description-format","gazebo","mjcf","universal-scene-description","xacro"],"name":"Simulation Description Format (SDFormat)","alt":"SDFormat 仿真描述格式","abbr":"SDF","aliases":["SDF","SDFormat"],"one_liner":"The XML format used by Gazebo that can describe a robot and an entire simulated world.","explanation":"SDFormat is an XML-based description format, originally designed for the Gazebo simulator and now maintained by Open Robotics. Unlike URDF (the Unified Robot Description Format), which describes only a single robot, SDFormat can describe an entire “world” in one file: multiple models, lights, the ground plane, sensors, and physics parameters, and it also supports closed kinematic chains, which URDF can only represent as a tree. It's commonly used to write scenes for Gazebo simulation; a common workflow is to start from a URDF and automatically convert it to SDF for loading. Note that this SDF is unrelated to the “signed distance field,” which uses the same abbreviation.","example":"Build a warehouse scene in Gazebo: use a single .sdf file to place shelving, the floor, light sources, and a mobile robot fitted with a lidar.","related":["Unified Robot Description Format","Gazebo","MJCF (MuJoCo XML Format)","Universal Scene Description (OpenUSD)","Xacro (XML Macros)"]},{"id":"universal-scene-description","category":"software","sec":2,"tier":2,"sources":[{"title":"OpenUSD 官网","url":"https://openusd.org/"},{"title":"Alliance for OpenUSD","url":"https://aousd.org/"}],"as_of":"","related_ids":["nvidia-isaac-sim","nvidia-omniverse","unified-robot-description-format","mjcf","simready-assets","simulation-assets"],"name":"Universal Scene Description (OpenUSD)","alt":"通用场景描述","abbr":"USD","aliases":["OpenUSD","USD"],"one_liner":"Pixar's open-source 3D scene-description format, now the standard scene format for NVIDIA's simulation platforms.","explanation":"USD is a 3D scene-description framework originally developed by Pixar Animation Studios for film production and open-sourced in 2016, now referred to as OpenUSD. It stores not just model meshes but also materials, lighting, hierarchy, and physical properties, and supports layering multiple files on top of one another, which makes collaborative editing of a shared scene easier. In 2023, Pixar, Apple, Adobe, Autodesk, NVIDIA, and others formed the Alliance for OpenUSD (AOUSD) to push standardization forward. NVIDIA Omniverse and Isaac Sim both use USD as their core format; a robot's URDF gets converted to USD after import, and simulation assets and digital-twin scenes circulate largely as .usd files.","example":"Import a robot's URDF into Isaac Sim, save it as robot.usd, then combine it with a kitchen scene's USD file to build a simulation environment.","related":["NVIDIA Isaac Sim","NVIDIA Omniverse","Unified Robot Description Format","MJCF (MuJoCo XML Format)","SimReady Assets","Simulation Assets"]},{"id":"mesh-file","category":"software","sec":2,"tier":2,"sources":[{"title":"Khronos glTF","url":"https://www.khronos.org/gltf/"},{"title":"ROS Wiki: URDF link（mesh 引用）","url":"http://wiki.ros.org/urdf/XML/link"}],"as_of":"","related_ids":["unified-robot-description-format","mjcf","collision-geometry","convex-decomposition","triangle-mesh","solidworks-to-urdf-exporter"],"name":"Mesh File (STL / OBJ / DAE / glTF)","alt":"网格文件","abbr":"","aliases":["STL","OBJ","DAE","glTF","model file"],"one_liner":"A 3D file that describes an object's shape as a collection of triangles — the visual “shell” of a robot model.","explanation":"A mesh file breaks a part or object's surface into many triangles and records their vertex coordinates and faces. Common formats include: STL, which stores geometry only, with no color or defined units; OBJ, which stores geometry plus texture coordinates, with materials in a companion .mtl file; DAE (COLLADA), an XML format that can carry materials and color and is common in ROS; and glTF, a newer format from the Khronos Group that supports physically based materials and suits web and rendering use cases. Robot description files like URDF and MJCF don't store shape directly — they reference these mesh files separately for the visual appearance and for the collision geometry used in physics checks. Collision meshes usually need to be simplified or convex-decomposed, or simulation becomes slow and unstable; another common pitfall is mixing up whether an STL file's units are millimeters or meters, which can make a model a thousand times too big.","example":"In a URDF exported from SolidWorks, each link references a meshes/xxx.STL file for both its visual appearance and its collision shape.","related":["Unified Robot Description Format","MJCF (MuJoCo XML Format)","Collision Geometry (Collider)","Convex Decomposition","Triangle Mesh","SolidWorks to URDF Exporter"]},{"id":"cad-software","category":"software","sec":2,"tier":2,"sources":[{"title":"Onshape 官网","url":"https://www.onshape.com/"},{"title":"SolidWorks 官网","url":"https://www.solidworks.com/"}],"as_of":"","related_ids":["unified-robot-description-format","solidworks-to-urdf-exporter","onshape-to-robot","mesh-file","mjcf","3d-printing"],"name":"CAD Software (SolidWorks / Onshape / Fusion 360)","alt":"CAD 建模软件（SolidWorks / Onshape / Fusion 360）","abbr":"CAD","aliases":["Computer-Aided Design","SolidWorks","Onshape","Fusion 360","STEP File"],"one_liner":"Software for drawing 3D models of mechanical parts and assemblies, the starting point for designing a robot’s structure.","explanation":"CAD (computer-aided design) software is used to draw 3D models of parts and assemblies. Three tools are common in robotics: Dassault Systèmes’ SolidWorks (desktop software, the most widely used in industry), PTC’s Onshape (browser-based, cloud-hosted CAD, convenient for collaboration), and Autodesk’s Fusion 360. Models are usually exchanged between different programs using the STEP format, a general-purpose, vendor-neutral 3D file format. For embodied AI, a CAD model is where simulation starts: a plugin exports the assembly into URDF or MJCF (robot model files describing links, joints, and mass), which, paired with mesh files, can be loaded straight into MuJoCo or Isaac Sim.","example":"onshape-to-robot exports an open-source arm assembly drawn in Onshape directly into a URDF, which is then loaded into a simulator.","related":["Unified Robot Description Format","SolidWorks to URDF Exporter","onshape-to-robot","Mesh File (STL / OBJ / DAE / glTF)","MJCF (MuJoCo XML Format)","3D Printing (FDM / Resin SLA)"]},{"id":"solidworks-to-urdf-exporter","category":"software","sec":2,"tier":3,"sources":[{"title":"sw_urdf_exporter - ROS Wiki","url":"http://wiki.ros.org/sw_urdf_exporter"}],"as_of":"","related_ids":["unified-robot-description-format","cad-software","onshape-to-robot","mesh-file","inertial-parameters"],"name":"SolidWorks to URDF Exporter","alt":"SolidWorks 转 URDF 插件","abbr":"sw2urdf","aliases":["sw2urdf","sw_urdf_exporter"],"one_liner":"A plugin that exports a SolidWorks assembly directly into a URDF robot model.","explanation":"This is a SolidWorks plugin maintained by the ROS community. Mechanical engineers typically design a robot in SolidWorks (a 3D CAD modeling program), while simulators and ROS need a model description in URDF format. The plugin lets users specify links, joint types, rotation axes, and coordinate frames directly in the assembly, then exports a ROS package containing the URDF file, an STL mesh for each link, and a launch file, with mass and inertia computed automatically from the CAD model's mass properties. The exported result usually still needs a manual check of joint directions, limits, and inertia before it's used in a simulator such as Isaac Sim or MuJoCo.","example":"Design a six-axis arm in SolidWorks, set a reference axis for each joint, export it with the plugin, and check the joint rotation directions in RViz.","related":["Unified Robot Description Format","CAD Software (SolidWorks / Onshape / Fusion 360)","onshape-to-robot","Mesh File (STL / OBJ / DAE / glTF)","Inertial Parameters"]},{"id":"onshape-to-robot","category":"software","sec":2,"tier":3,"sources":[{"title":"Rhoban/onshape-to-robot (GitHub)","url":"https://github.com/Rhoban/onshape-to-robot"}],"as_of":"","related_ids":["unified-robot-description-format","mjcf","solidworks-to-urdf-exporter","cad-software","mesh-file","simulation-description-format"],"name":"onshape-to-robot","alt":"onshape-to-robot","abbr":"","aliases":[],"one_liner":"Open-source tool that exports robot assemblies built in Onshape directly into URDF or MJCF files.","explanation":"onshape-to-robot is an open-source Python tool from the Rhoban team at the University of Bordeaux, France. Robot structures are often designed in CAD software, but simulators and ROS need robot description files such as URDF, SDF, or MJCF, and manually transcribing links, joints, mass, and inertia from a CAD model is slow and error-prone. This tool reads an assembly through the API of Onshape, a cloud-based CAD platform, identifies joints from a naming convention, and automatically generates the description files and meshes, including mass and inertia data. Open-hardware projects commonly use it to move a design straight into simulators like MuJoCo or PyBullet.","example":"Name the joint mates in Onshape something like dof_knee, run onshape-to-robot, and get an MJCF file plus STL meshes ready to load directly into MuJoCo.","related":["Unified Robot Description Format","MJCF (MuJoCo XML Format)","SolidWorks to URDF Exporter","CAD Software (SolidWorks / Onshape / Fusion 360)","Mesh File (STL / OBJ / DAE / glTF)","Simulation Description Format (SDFormat)"]},{"id":"blender","category":"software","sec":2,"tier":3,"sources":[{"title":"Blender 官网","url":"https://www.blender.org/"}],"as_of":"","related_ids":["mesh-file","simulation-assets","synthetic-data","unified-robot-description-format","photorealistic-rendering","path-tracing"],"name":"Blender","alt":"Blender","abbr":"","aliases":[],"one_liner":"Free, open-source 3D modeling and rendering software, often used to process robot models and generate synthetic data.","explanation":"Blender is a free, open-source 3D creation suite maintained by the Blender Foundation, covering modeling, materials, animation, and rendering (with the built-in Cycles path tracer and the real-time EEVEE renderer), plus a full Python API for scripting batch operations. It isn't robotics software per se, but it gets heavy use in embodied AI: cleaning up and simplifying mesh files for a robot or object before exporting them as STL, OBJ, or glTF for a URDF or simulator; building simulation scenes and assets; and scripting batch rendering of annotated images for synthetic data — the German Aerospace Center's open-source BlenderProc, for example, is built on Blender. Plugins also exist for editing and exporting URDF from inside Blender.","example":"Open a high-poly gripper mesh exported from CAD in Blender, reduce its polygon count, and export it as OBJ for use as the visual model in a URDF.","related":["Mesh File (STL / OBJ / DAE / glTF)","Simulation Assets","Synthetic Data","Unified Robot Description Format","Photorealistic Rendering","Path Tracing"]},{"id":"coacd","category":"software","sec":2,"tier":3,"sources":[{"title":"CoACD - GitHub","url":"https://github.com/SarahWeiii/CoACD"}],"as_of":"","related_ids":["convex-decomposition","collision-geometry","simulation-assets","mesh-file","unified-robot-description-format","mjcf"],"name":"CoACD (Collision-Aware Approximate Convex Decomposition)","alt":"CoACD 凸分解工具","abbr":"","aliases":["CoACD"],"one_liner":"A tool that automatically slices a mesh into a set of convex chunks for use as a physics engine's collision shape.","explanation":"CoACD is an open-source approximate convex-decomposition tool from Hao Su's group at UC San Diego (Wei et al., SIGGRAPH 2022). Physics engines compute collisions fastest against convex shapes, but simulation assets are usually arbitrarily shaped triangle meshes, which are slow and unstable to use directly; the older, commonly used V-HACD tends to fill in important concave structures — a cup's rim, the gap of a drawer — making grasping and insertion tasks impossible to simulate correctly. CoACD explicitly accounts for collision-aware concavity during decomposition and uses a tree search to choose where to cut, preserving those structures. In practice, an OBJ or STL is run through the tool once, and the resulting set of convex pieces is written into a URDF's or MJCF's collision fields.","example":"Before importing a cup or piece of furniture's mesh into MuJoCo or Isaac Sim, run it through CoACD first to generate convex collision geometry, so the gripper doesn't fail to fit into the cup's opening.","related":["Convex Decomposition","Collision Geometry (Collider)","Simulation Assets","Mesh File (STL / OBJ / DAE / glTF)","Unified Robot Description Format","MJCF (MuJoCo XML Format)"]},{"id":"tf-tf2-transform-tree","category":"software","sec":2,"tier":2,"sources":[{"title":"ROS 2 Docs: About tf2","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Tf2.html"},{"title":"ROS Wiki: tf2","url":"http://wiki.ros.org/tf2"}],"as_of":"","related_ids":["coordinate-transformation","coordinate-frame","robot-state-publisher","rep-105","rviz-rviz2","homogeneous-transformation-matrix"],"name":"TF / tf2 Transform Tree","alt":"TF 坐标树","abbr":"TF","aliases":["tf2","TF tree","transform tree"],"one_liner":"ROS's tree structure for tracking and querying the live transform between every pair of coordinate frames.","explanation":"TF is ROS's coordinate-transform library; the current version is tf2, originally developed by Tully Foote and others at Willow Garage. A robot carries many coordinate frames — the map, the base, each link, cameras, grippers — and their relative poses change over time. TF organizes these frames into a tree, where every edge is a timestamped transform, published by the relevant node to the /tf (dynamic) and /tf_static (fixed) topics. Any program can then query the transform between any two frames at a given moment — for example, converting an object's coordinates as seen by a camera into the arm's base frame so it can be grasped.","example":"Running ros2 run tf2_ros tf2_echo base_link camera_link shows the camera's pose relative to the base; robot_state_publisher automatically publishes the whole arm's TF tree from the URDF and joint angles.","related":["Coordinate Transformation","Coordinate Frame","robot_state_publisher","REP 105","RViz / RViz2","Homogeneous Transformation Matrix"]},{"id":"robot-state-publisher","category":"software","sec":2,"tier":3,"sources":[{"title":"robot_state_publisher on GitHub","url":"https://github.com/ros/robot_state_publisher"},{"title":"ROS Wiki: robot_state_publisher","url":"http://wiki.ros.org/robot_state_publisher"}],"as_of":"","related_ids":["tf-tf2-transform-tree","unified-robot-description-format","rviz-rviz2","forward-kinematics","robot-operating-system-2","topic"],"name":"robot_state_publisher","alt":"robot_state_publisher","abbr":"","aliases":["joint_state_publisher"],"one_liner":"ROS package that reads a URDF and joint angles, then publishes each link's coordinate transform.","explanation":"robot_state_publisher is a fundamental package in both ROS and ROS 2. It reads a robot's URDF model (via the robot_description parameter), subscribes to joint angles on the /joint_states topic, runs forward kinematics, and publishes the pose of each link relative to its parent onto the TF transform tree. With it in place, RViz can draw the robot's current pose, and other nodes can query things like “where is the gripper relative to the base.” Its usual companion, joint_state_publisher, is responsible for publishing the joint angles themselves: without a real robot connected, its GUI version lets you drag sliders to set joint angles by hand, which is convenient for checking a model.","example":"After writing a robot arm's URDF, launch robot_state_publisher, joint_state_publisher_gui, and RViz together, and dragging the sliders will visibly move each joint of the model.","related":["TF / tf2 Transform Tree","Unified Robot Description Format","RViz / RViz2","Forward Kinematics (FK)","Robot Operating System 2","Topic (ROS)"]},{"id":"rep-105","category":"software","sec":2,"tier":3,"sources":[{"title":"REP 105 -- Coordinate Frames for Mobile Platforms","url":"https://www.ros.org/reps/rep-0105.html"}],"as_of":"","related_ids":["tf-tf2-transform-tree","coordinate-frame","ros-2-navigation-stack","wheel-odometry","adaptive-monte-carlo-localization","right-handed-frame-and-axis-conventions"],"name":"REP 105","alt":"REP 105 坐标系约定（map / odom / base_link）","abbr":"","aliases":["REP 105: Coordinate Frames for Mobile Platforms","map–odom–base_link"],"one_liner":"The ROS convention defining the map, odom, and base_link coordinate frames for mobile robots.","explanation":"REP 105 is a ROS Enhancement Proposal — a community specification — that fixes the names and meanings of the coordinate frames commonly used by mobile robots. base_link is fixed to the robot's own body; odom is the odometry frame, in which the robot's pose is continuous and smooth but slowly drifts over time; map is the global map frame, where the pose doesn't drift but can jump when corrected by localization. The three are chained into a TF tree as map → odom → base_link: the localization module is responsible only for publishing map → odom, while odometry publishes odom → base_link. This shared convention is what lets independently written navigation, SLAM, and localization packages plug together — Nav2 and others all follow it.","example":"A robot publishes odom → base_link from wheel odometry, while an AMCL localization node publishes map → odom based on laser-scan matching; combined, the two give the robot's position on the map.","related":["TF / tf2 Transform Tree","Coordinate Frame","ROS 2 Navigation Stack (Nav2)","Wheel Odometry","Adaptive Monte Carlo Localization","Right-Handed Frame & Axis Conventions"]},{"id":"rviz-rviz2","category":"software","sec":2,"tier":2,"sources":[{"title":"ros2/rviz (GitHub)","url":"https://github.com/ros2/rviz"},{"title":"ROS Wiki: rviz","url":"http://wiki.ros.org/rviz"}],"as_of":"","related_ids":["robot-operating-system-2","tf-tf2-transform-tree","unified-robot-description-format","moveit-motion-planning-framework","rqt","foxglove-studio"],"name":"RViz / RViz2","alt":"RViz","abbr":"","aliases":["RViz2","rviz"],"one_liner":"ROS's built-in 3D visualization tool for viewing a robot model, coordinate frames, and sensor data.","explanation":"RViz is ROS's official 3D visualization tool; the ROS 2 version is called RViz2. It subscribes to ROS topics and draws the robot model (from a URDF), TF coordinate frames, lidar point clouds, camera images, planned paths, and custom markers all in the same 3D scene. It doesn't do any simulation itself — its only job is to show what the robot currently believes the world looks like — which makes it the standard tool for debugging sensor calibration, coordinate transforms, and navigation planning. Both MoveIt's and Nav2's interactive interfaces are provided as RViz plugins.","example":"Add a RobotModel and a PointCloud2 display in RViz2 to check whether the depth camera's point cloud lines up with the arm model, as a way of verifying hand-eye calibration.","related":["Robot Operating System 2","TF / tf2 Transform Tree","Unified Robot Description Format","MoveIt Motion Planning Framework","rqt","Foxglove Studio"]},{"id":"middleware","category":"software","sec":3,"tier":2,"sources":[{"title":"ROS 2 文档：Different ROS 2 middleware vendors","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Different-Middleware-Vendors.html"}],"as_of":"","related_ids":["ros-middleware-interface","data-distribution-service","eprosima-fast-dds","eclipse-cyclone-dds","eclipse-zenoh","ros-2-quality-of-service"],"name":"Middleware","alt":"中间件","abbr":"","aliases":["communication middleware","robot middleware"],"one_liner":"The software layer sitting between the operating system and applications that handles communication between programs.","explanation":"Middleware is a layer of software between the operating system and applications that takes care of generic problems on the developer's behalf: how programs find each other, how data gets packaged, how it moves between processes or machines, and how packet loss and latency are handled. Robot software is made up of many independent processes, sensor data volumes are large, and latency matters, so communication middleware is critical. ROS 2 does not implement transport itself; instead it plugs into an underlying implementation through the ROS middleware interface (RMW), defaulting to a DDS (Data Distribution Service) implementation such as Fast DDS or Cyclone DDS, though it can also be switched to Zenoh. Frameworks like ROS, LCM, and dora-rs are sometimes referred to collectively as robot middleware as well.","example":"Setting the environment variable RMW_IMPLEMENTATION=rmw_cyclonedds_cpp switches ROS 2's underlying communication from the default Fast DDS to Cyclone DDS.","related":["ROS Middleware Interface (RMW)","Data Distribution Service","eProsima Fast DDS","Eclipse Cyclone DDS","Eclipse Zenoh","ROS 2 Quality of Service (QoS)"]},{"id":"data-distribution-service","category":"software","sec":3,"tier":2,"sources":[{"title":"OMG: Data Distribution Service specification","url":"https://www.omg.org/spec/DDS/"},{"title":"ROS 2 Documentation: Different ROS 2 middleware vendors","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Different-Middleware-Vendors.html"}],"as_of":"","related_ids":["robot-operating-system-2","ros-middleware-interface","eprosima-fast-dds","eclipse-cyclone-dds","ros-2-quality-of-service","eclipse-zenoh"],"name":"Data Distribution Service","alt":"数据分发服务","abbr":"DDS","aliases":["DDS"],"one_liner":"A publish/subscribe communication standard from the OMG, and the default underlying transport for ROS 2.","explanation":"DDS is a real-time publish/subscribe communication middleware standard defined by the Object Management Group (OMG), originally used in distributed systems such as defense and aerospace that demand high real-time performance and reliability. It has no central node — participants automatically discover each other over the network — and its QoS (Quality of Service) policies configure things like reliable versus best-effort delivery and how much message history to keep. ROS 2 dropped ROS 1’s home-grown communication layer in favor of connecting to DDS through the RMW (ROS middleware interface), with common implementations including eProsima Fast DDS, Eclipse Cyclone DDS, and RTI Connext. When multi-machine communication breaks or a topic isn’t receiving anything, the DDS configuration is often where to start looking.","example":"Setting the environment variable RMW_IMPLEMENTATION=rmw_cyclonedds_cpp switches ROS 2’s underlying communication from the default Fast DDS to Cyclone DDS.","related":["Robot Operating System 2","ROS Middleware Interface (RMW)","eProsima Fast DDS","Eclipse Cyclone DDS","ROS 2 Quality of Service (QoS)","Eclipse Zenoh"]},{"id":"ros-2-quality-of-service","category":"software","sec":3,"tier":3,"sources":[{"title":"Quality of Service settings (ROS 2 Documentation)","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Quality-of-Service-Settings.html"}],"as_of":"","related_ids":["robot-operating-system-2","data-distribution-service","publish-subscribe","topic","middleware","ros-middleware-interface"],"name":"ROS 2 Quality of Service (QoS)","alt":"服务质量","abbr":"QoS","aliases":["QoS","QoS policies"],"one_liner":"The ROS 2 settings that control how a message is delivered — reliably or not, how many kept.","explanation":"Quality of Service (QoS) is originally a general networking concept. ROS 2 is built on DDS (Data Distribution Service) middleware, and applies QoS settings to every publisher and subscriber individually. Common policies include: reliability (guaranteed delivery with retransmission, versus best-effort), durability (whether a subscriber that joins late still receives earlier messages), history (keep the last N messages, or all of them), queue depth, plus deadline and liveliness settings. Choosing them well trades off dropped data against latency — for example, for high-rate sensor data, it's often better to drop a frame than to let messages queue up. A common beginner mistake is mismatched QoS between publisher and subscriber: the topic name is correct, but no messages ever arrive.","example":"A camera driver publishes images as best-effort; if the subscriber is set to reliable mode, the two are incompatible, and the subscribing node receives no frames at all.","related":["Robot Operating System 2","Data Distribution Service","Publish-Subscribe","Topic (ROS)","Middleware","ROS Middleware Interface (RMW)"]},{"id":"ros-2-domain-id","category":"software","sec":3,"tier":3,"sources":[{"title":"ROS 2 Documentation: The ROS_DOMAIN_ID","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Domain-ID.html"}],"as_of":"","related_ids":["robot-operating-system-2","data-distribution-service","middleware","node","topic","ros-2-quality-of-service"],"name":"ROS 2 Domain ID","alt":"ROS_DOMAIN_ID（域 ID）","abbr":"","aliases":["ROS_DOMAIN_ID"],"one_liner":"A ROS 2 environment variable that decides which nodes can discover and talk to each other.","explanation":"ROS 2 communicates over DDS middleware under the hood, and nodes on the same local network discover each other automatically. ROS_DOMAIN_ID is an environment variable used to group nodes: only nodes with the same domain ID can see one another, and the default value is 0. It comes from DDS's own concept of a “domain,” and different domain IDs map to different network ports. The official documentation recommends staying within the 0–101 range to avoid clashing with system ports. In a lab where multiple people and multiple robots share one network, forgetting to set separate domain IDs means everyone ends up receiving each other's topic data and interfering with one another — a common beginner pitfall.","example":"Two students each controlling a robot on the same Wi-Fi run export ROS_DOMAIN_ID=11 and export ROS_DOMAIN_ID=12 respectively in their terminals, so their topics no longer cross over.","related":["Robot Operating System 2","Data Distribution Service","Middleware","Node (ROS)","Topic (ROS)","ROS 2 Quality of Service (QoS)"]},{"id":"ros-middleware-interface","category":"software","sec":3,"tier":3,"sources":[{"title":"ROS 2 Docs: About different ROS 2 middleware vendors","url":"https://docs.ros.org/en/jazzy/Concepts/Intermediate/About-Different-Middleware-Vendors.html"},{"title":"ROS 2 Design: ROS 2 middleware interface","url":"https://design.ros2.org/articles/ros_middleware_interface.html"}],"as_of":"","related_ids":["robot-operating-system-2","middleware","data-distribution-service","eprosima-fast-dds","eclipse-cyclone-dds","eclipse-zenoh"],"name":"ROS Middleware Interface (RMW)","alt":"ROS 中间件接口","abbr":"RMW","aliases":["rmw","RMW implementation"],"one_liner":"The ROS 2 abstraction layer that separates the upper-level API from the actual communication backend.","explanation":"RMW is a C interface layer in ROS 2, sitting between the client libraries (rclcpp, rclpy, via rcl) and the actual communication middleware. ROS 2 communicates over DDS (Data Distribution Service, a publish/subscribe communication standard) by default, and since several vendors implement DDS, RMW lets each one plug in as a swappable backend — for example rmw_fastrtps_cpp (Fast DDS) or rmw_cyclonedds_cpp (Cyclone DDS); newer distributions have also added rmw_zenoh_cpp, based on Zenoh. User code doesn't need to change at all — switching backends is just a matter of setting the RMW_IMPLEMENTATION environment variable. When multi-machine discovery fails or large data transfers stutter, trying a different RMW implementation is a common troubleshooting step; nodes within the same system should generally all use the same one.","example":"Run export RMW_IMPLEMENTATION=rmw_cyclonedds_cpp before ros2 run, and the node switches to sending and receiving messages over Cyclone DDS.","related":["Robot Operating System 2","Middleware","Data Distribution Service","eProsima Fast DDS","Eclipse Cyclone DDS","Eclipse Zenoh"]},{"id":"eprosima-fast-dds","category":"software","sec":3,"tier":3,"sources":[{"title":"Fast DDS 官方文档","url":"https://fast-dds.docs.eprosima.com/"},{"title":"Fast DDS GitHub","url":"https://github.com/eProsima/Fast-DDS"}],"as_of":"","related_ids":["data-distribution-service","eclipse-cyclone-dds","ros-middleware-interface","robot-operating-system-2","ros-2-quality-of-service","ros-distribution"],"name":"eProsima Fast DDS","alt":"Fast DDS","abbr":"","aliases":["Fast-RTPS","Fast RTPS"],"one_liner":"An open-source DDS implementation from the Spanish company eProsima, the default middleware in several ROS 2 releases.","explanation":"Fast DDS is an open-source DDS (Data Distribution Service) implementation developed by the Spanish company eProsima, written in C++, formerly named Fast RTPS (RTPS being the wire protocol underlying DDS). ROS 2 connects to a specific DDS implementation through its RMW interface, with the package for this one being rmw_fastrtps_cpp; mainstream ROS 2 distributions such as Humble and Jazzy use Fast DDS by default. It supports shared-memory transport and a Discovery Server mode (which cuts down on node-discovery traffic in large networks), among other features. Newcomers who run into topics not being received, or cross-machine communication not working, often end up needing to understand Fast DDS's XML configuration and discovery mechanism.","example":"With many robots online at once, discovery traffic gets too heavy; switching to Fast DDS's Discovery Server mode puts one server node in charge of node discovery for everyone.","related":["Data Distribution Service","Eclipse Cyclone DDS","ROS Middleware Interface (RMW)","Robot Operating System 2","ROS 2 Quality of Service (QoS)","ROS Distribution"]},{"id":"eclipse-cyclone-dds","category":"software","sec":3,"tier":3,"sources":[{"title":"Eclipse Cyclone DDS 官网","url":"https://cyclonedds.io/"},{"title":"Cyclone DDS GitHub","url":"https://github.com/eclipse-cyclonedds/cyclonedds"}],"as_of":"","related_ids":["data-distribution-service","eprosima-fast-dds","ros-middleware-interface","eclipse-iceoryx","unitree-sdk2","robot-operating-system-2"],"name":"Eclipse Cyclone DDS","alt":"Cyclone DDS","abbr":"","aliases":["CycloneDDS"],"one_liner":"The Eclipse Foundation's open-source DDS implementation, one of the optional communication middlewares underlying ROS 2.","explanation":"Cyclone DDS is an open-source DDS (Data Distribution Service — an industrial publish-subscribe communication standard) implementation hosted by the Eclipse Foundation, written in C and focused on being lightweight and low-latency. ROS 2 doesn't send and receive data directly; instead it goes through the RMW (ROS middleware interface) to call into a specific DDS implementation, and Cyclone DDS is one of the officially supported options, corresponding to rmw_cyclonedds_cpp. It can use iceoryx to achieve zero-copy transport over shared memory. Quite a few robot vendors' SDKs are built directly on it — Unitree SDK2's communication layer, for instance, uses Cyclone DDS. Switching to Cyclone DDS is a common troubleshooting step when the default middleware shows slow discovery or dropped packets.","example":"Setting the environment variable RMW_IMPLEMENTATION=rmw_cyclonedds_cpp switches ROS 2 nodes over to communicating through Cyclone DDS.","related":["Data Distribution Service","eProsima Fast DDS","ROS Middleware Interface (RMW)","Eclipse iceoryx","Unitree SDK2","Robot Operating System 2"]},{"id":"eclipse-zenoh","category":"software","sec":3,"tier":3,"sources":[{"title":"Eclipse Zenoh 官网","url":"https://zenoh.io/"},{"title":"rmw_zenoh GitHub","url":"https://github.com/ros2/rmw_zenoh"}],"as_of":"","related_ids":["robot-operating-system-2","ros-middleware-interface","data-distribution-service","eclipse-cyclone-dds","publish-subscribe","dataflow-oriented-robotic-architecture"],"name":"Eclipse Zenoh","alt":"Zenoh","abbr":"","aliases":["rmw_zenoh"],"one_liner":"An open-source communication protocol unifying publish-subscribe, storage, and query, usable in ROS 2 as a DDS replacement.","explanation":"Zenoh is an open-source communication protocol and implementation hosted by the Eclipse Foundation, developed mainly by the company ZettaScale, with its core written in Rust. It unifies publish-subscribe, data storage, and querying under one interface, designed to work across everything from microcontrollers to the cloud and to span routers and wide-area networks. DDS often runs into heavy node-discovery traffic and configuration complexity over Wi-Fi, in multi-robot setups, or across network segments; Zenoh eases these problems with a router-based approach (zenohd). ROS 2 officially provides rmw_zenoh, letting Zenoh serve as the underlying middleware in place of DDS, and a zenoh-bridge can also carry existing DDS traffic across.","example":"Start a Zenoh router in ROS 2, then set RMW_IMPLEMENTATION=rmw_zenoh_cpp so the robot and a remote workstation on a different network segment can exchange topics with each other.","related":["Robot Operating System 2","ROS Middleware Interface (RMW)","Data Distribution Service","Eclipse Cyclone DDS","Publish-Subscribe","Dataflow-Oriented Robotic Architecture (dora)"]},{"id":"zero-copy","category":"software","sec":3,"tier":3,"sources":[{"title":"Eclipse iceoryx","url":"https://iceoryx.io/"},{"title":"Configure Zero Copy Loaned Messages (ROS 2 Documentation)","url":"https://docs.ros.org/en/rolling/How-To-Guides/Configure-ZeroCopy-loaned-messages.html"}],"as_of":"","related_ids":["eclipse-iceoryx","composable-nodes-and-intra-process-communication","robot-operating-system-2","data-distribution-service","middleware","control-latency"],"name":"Zero-Copy (Shared-Memory IPC)","alt":"零拷贝","abbr":"","aliases":["shared-memory communication","zero-copy communication"],"one_liner":"Passing large data between processes by sharing a memory location instead of copying the content.","explanation":"Zero-copy refers to inter-process communication (IPC) in which the sender writes data directly into a block of shared memory that both sides can access, and the receiver gets only a reference to that memory rather than a fresh copy of it. On a robot, camera images, depth maps, and point clouds routinely run to tens or hundreds of megabytes per second; if every subscriber has to go through serialization plus a copy, CPU usage and latency both climb noticeably. Eclipse iceoryx is middleware built specifically for this; ROS 2 achieves zero-copy through loaned messages combined with iceoryx or Fast DDS's shared-memory transport, and within a single process, composable nodes can also use intra-process communication. The limitation is that it generally only works on a single machine, and works best when messages are a fixed size.","example":"Enable Fast DDS shared-memory transport between a camera driver and a detection node in ROS 2, so 1080p images are no longer copied frame by frame, cutting end-to-end latency.","related":["Eclipse iceoryx","Composable Nodes (Components) & Intra-process Communication","Robot Operating System 2","Data Distribution Service","Middleware","Control Latency"]},{"id":"eclipse-iceoryx","category":"software","sec":3,"tier":3,"sources":[{"title":"Eclipse iceoryx 官网","url":"https://iceoryx.io/"},{"title":"iceoryx GitHub","url":"https://github.com/eclipse-iceoryx/iceoryx"}],"as_of":"","related_ids":["zero-copy","eclipse-cyclone-dds","middleware","robot-operating-system-2","dataflow-oriented-robotic-architecture"],"name":"Eclipse iceoryx","alt":"iceoryx","abbr":"","aliases":["iceoryx2"],"one_liner":"A shared-memory, zero-copy inter-process communication library often used by DDS and ROS 2 to move large data.","explanation":"iceoryx is an open-source inter-process communication (IPC) middleware hosted by the Eclipse Foundation, originally developed by a Bosch team for automotive software. Ordinary communication copies data from the sender to the receiver, and for large messages like images or point clouds that copy alone can take a while; iceoryx instead lets the sender write data straight into shared memory, with the receiver getting only a reference to it — “zero-copy” — so transfer latency is essentially independent of message size. It only handles processes on the same machine; communication across machines still needs a network protocol like DDS. Cyclone DDS can use it as its local shared-memory transport layer. A Rust rewrite, iceoryx2, has since been released as well.","example":"Configuring Cyclone DDS on a robot's main controller to enable iceoryx shared memory means high-resolution images sent from a camera node to a perception node no longer get copied.","related":["Zero-Copy (Shared-Memory IPC)","Eclipse Cyclone DDS","Middleware","Robot Operating System 2","Dataflow-Oriented Robotic Architecture (dora)"]},{"id":"composable-nodes-and-intra-process-communication","category":"software","sec":3,"tier":3,"sources":[{"title":"Intra-process Communications in ROS 2 (design.ros2.org)","url":"https://design.ros2.org/articles/intraprocess_communications.html"},{"title":"About Composition - ROS 2 Documentation","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Composition.html"}],"as_of":"","related_ids":["node","robot-operating-system-2","middleware","data-distribution-service","zero-copy","ros-2-executor-and-callback-groups"],"name":"Composable Nodes (Components) & Intra-process Communication","alt":"可组合节点 / 进程内通信","abbr":"","aliases":["ROS 2 components","Component"],"one_liner":"Loading several ROS 2 nodes into a single process, so they pass data as raw pointers instead of copying it.","explanation":"This is a ROS 2 mechanism: a node is compiled into a shared library, and at runtime a container process loads several of them on demand — nodes built this way are called composable nodes. Publish-subscribe within the same process can then use intra-process communication, passing messages as smart pointers instead of paying for serialization and a memory copy; communication across processes still has to go through the middleware (DDS) and the network stack. This matters a great deal for high-bandwidth data — if several camera image streams and point clouds get copied at every hop, CPU load and latency both suffer. The practice is to register a node as a component when writing it, then use a launch file to load a camera driver, an image-processing node, and an inference node into the same container; Isaac ROS's vision pipelines are organized this way by default.","example":"Load a camera driver, an undistortion node, and an object-detection node into one component container, so images flow between them within the process with zero copying.","related":["Node (ROS)","Robot Operating System 2","Middleware","Data Distribution Service","Zero-Copy (Shared-Memory IPC)","ROS 2 Executor & Callback Groups"]},{"id":"ros-2-executor-and-callback-groups","category":"software","sec":3,"tier":3,"sources":[{"title":"ROS 2 Documentation: Executors","url":"https://docs.ros.org/en/rolling/Concepts/Intermediate/About-Executors.html"},{"title":"ROS 2 Documentation: Using Callback Groups","url":"https://docs.ros.org/en/rolling/How-To-Guides/Using-callback-groups.html"}],"as_of":"","related_ids":["robot-operating-system-2","node","ros-client-library","composable-nodes-and-intra-process-communication","real-time-control","service"],"name":"ROS 2 Executor & Callback Groups","alt":"ROS 2 执行器与回调组","abbr":"","aliases":["Executor","Callback Group","MultiThreadedExecutor"],"one_liner":"The ROS 2 machinery that decides which threads run callbacks, and whether they can run concurrently.","explanation":"In a ROS 2 node, events from subscriptions, timers, and services all trigger callback functions, and the Executor is responsible for scheduling them: a single-threaded executor runs only one at a time, while a multi-threaded executor can run several in parallel. Callback groups further control which callbacks are allowed to run concurrently: callbacks in a mutually exclusive group never run at the same time, while those in a reentrant group can. This addresses real-time and deadlock concerns — for example, if a callback synchronously calls a service and waits for the result, and the service's response callback is blocked behind it on the same thread, the node can deadlock. Putting them in different callback groups and running a multi-threaded executor avoids this.","example":"Put a control node's 100 Hz timer callback and its slower camera-image-processing callback in different callback groups and run them with a MultiThreadedExecutor, so the control loop isn't dragged down by image processing.","related":["Robot Operating System 2","Node (ROS)","ROS Client Library (rclcpp / rclpy)","Composable Nodes (Components) & Intra-process Communication","Real-Time Control","Service (ROS)"]},{"id":"lifecycle-node","category":"software","sec":3,"tier":3,"sources":[{"title":"ROS 2 Design: Managed nodes","url":"https://design.ros2.org/articles/node_lifecycle.html"}],"as_of":"","related_ids":[null,null,null,null,null,"ros2-control"],"name":"Lifecycle Node (Managed Node)","alt":"生命周期节点","abbr":"","aliases":["managed node","LifecycleNode"],"one_liner":"A ROS 2 node type with a standard state machine, letting an external manager configure, start, pause, and shut it down in order.","explanation":"A lifecycle node is a node type ROS 2 introduced. An ordinary node starts working the moment it launches, and there's no guarantee about which of several nodes becomes ready first; a lifecycle node instead has a built-in standard state machine, with main states unconfigured, inactive, active, and finalized, moving between them through transitions such as configure, activate, deactivate, cleanup, and shutdown — each transition calling a callback the developer has written (for example, reading parameters and allocating resources during “configure,” or starting to publish messages during “activate”). An external manager can then bring up a group of nodes in a defined order, and shut them all down uniformly if something goes wrong, making system startup and fault handling controllable. Nav2's navigation stack uses a lifecycle manager to start up its various server nodes this way.","example":"Nav2's lifecycle_manager configures and activates the map server and localization node in sequence, then, once they're confirmed ready, activates the planner and controller — so the controller never starts outputting velocity commands before a map exists.","related":["ROS 2","Node (ROS)","ROS 2 Navigation Stack (Nav2)","Finite State Machine","Launch File","ros2_control"]},{"id":"lightweight-communications-and-marshalling","category":"software","sec":3,"tier":3,"sources":[{"title":"LCM 官方文档","url":"https://lcm-proj.github.io/lcm/"},{"title":"lcm-proj/lcm (GitHub)","url":"https://github.com/lcm-proj/lcm"}],"as_of":"","related_ids":[null,null,null,"drake","zeromq","mit-mini-cheetah"],"name":"Lightweight Communications and Marshalling (LCM)","alt":"LCM","abbr":"LCM","aliases":["lcm-proj"],"one_liner":"A lightweight publish-subscribe messaging library over UDP multicast, low-latency and common in real-time robot control.","explanation":"LCM is a messaging library developed and open-sourced by an MIT team during their participation in the DARPA Urban Challenge (a self-driving-car competition). Developers define message structures in .lcm files, and the tool automatically generates the send/receive code for C, C++, Python, Java, and other languages; messages travel over UDP multicast in a publish-subscribe pattern, need no central node, and have few dependencies and low latency. It also comes with tools like lcm-spy (for viewing messages live) and lcm-logger (for recording and replay). Compared with ROS, LCM handles communication only and nothing else, which is why it's often chosen for control loops with tight real-time requirements — the MIT Cheetah series of quadrupeds, the Drake toolbox, and some robot vendors' early SDKs among them.","example":"In a Drake simulation, the controller and a visualizer exchange robot state over an LCM channel, and a developer uses lcm-spy to watch each channel's message rate and content live.","related":["Publish-Subscribe","Middleware","Robot Operating System","Drake","ZeroMQ","MIT Mini Cheetah"]},{"id":"zeromq","category":"software","sec":3,"tier":3,"sources":[{"title":"ZeroMQ 官网","url":"https://zeromq.org/"},{"title":"ZeroMQ Guide (zguide)","url":"https://zguide.zeromq.org/"},{"title":"ZeroMQ - Wikipedia","url":"https://en.wikipedia.org/wiki/ZeroMQ"}],"as_of":"","related_ids":["policy-server","publish-subscribe","grpc-remote-procedure-calls","lightweight-communications-and-marshalling","protocol-buffers","middleware"],"name":"ZeroMQ","alt":"ZeroMQ","abbr":"ZMQ","aliases":["ØMQ","0MQ","ZMQ"],"one_liner":"A lightweight, open-source messaging library that lets programs exchange data in just a few lines.","explanation":"ZeroMQ is an open-source asynchronous messaging library started by Pieter Hintjens, Martin Sustrik, and others at the company iMatix; its core is libzmq, written in C++, with pyzmq as the Python binding. It needs no separate message broker, and packages common communication patterns into a handful of socket types: request-reply (REQ/REP), publish-subscribe (PUB/SUB), push-pull (PUSH/PULL), and others, usable across processes or across machines. In embodied AI it's commonly used to build a lightweight policy server: a GPU machine runs the model, and the robot sends over observations and gets back actions, without pulling in the whole ROS stack. It only handles moving bytes around — the data format itself still needs a serialization scheme such as msgpack or Protobuf.","example":"The robot side uses a REQ socket to send camera images and joint state to a remote GPU server, which uses a REP socket to return a chunk of actions.","related":["Policy Server (Remote Inference)","Publish-Subscribe","gRPC Remote Procedure Calls","Lightweight Communications and Marshalling (LCM)","Protocol Buffers","Middleware"]},{"id":"protocol-buffers","category":"software","sec":3,"tier":3,"sources":[{"title":"Protocol Buffers 官方文档","url":"https://protobuf.dev/"}],"as_of":"","related_ids":["grpc-remote-procedure-calls","message","mcap","middleware","zeromq","lightweight-communications-and-marshalling"],"name":"Protocol Buffers","alt":"Protobuf","abbr":"Protobuf","aliases":["Protobuf","proto"],"one_liner":"Google's open-source, language-neutral format for serializing structured data compactly and quickly.","explanation":"Protobuf (Protocol Buffers) is a data serialization mechanism developed by Google and open-sourced in 2008. Developers first define a message's structure — field names, types, and numeric tags — in a .proto file, then use the protoc compiler to generate reading and writing code for languages such as C++, Python, and Go; the data itself is transmitted and stored in a compact binary format. Compared with JSON, it uses less bandwidth, parses faster, and lets old and new versions of a message stay compatible after fields are added. It's the default interface description and encoding format for gRPC, and is also supported by robotics tools such as the MCAP logging format and Foxglove, making it common for communication between a robot and the cloud or an inference server.","example":"When deploying a policy as a remote inference service, define the “image + joint state → action chunk” request and response format in a .proto file, then call it over gRPC.","related":["gRPC Remote Procedure Calls","Message (msg / srv / action interface)","MCAP","Middleware","ZeroMQ","Lightweight Communications and Marshalling (LCM)"]},{"id":"grpc-remote-procedure-calls","category":"software","sec":3,"tier":3,"sources":[{"title":"Introduction to gRPC","url":"https://grpc.io/docs/what-is-grpc/introduction/"}],"as_of":"","related_ids":["protocol-buffers","policy-server","websocket","zeromq","middleware"],"name":"gRPC Remote Procedure Calls","alt":"gRPC","abbr":"gRPC","aliases":["gRPC"],"one_liner":"Google's open-source RPC framework, letting programs on different machines call each other like local functions.","explanation":"gRPC is Google's open-source RPC (remote procedure call) framework, now hosted by the Cloud Native Computing Foundation (CNCF). It defines interfaces and messages with Protobuf (a binary serialization format), runs over HTTP/2, supports bidirectional streaming, and can automatically generate client and server code in C++, Python, Go, and many other languages. In embodied AI, it's commonly used to connect “the GPU server running the model” to “the robot itself”: the robot sends images and joint state to a policy server and gets an action back, which saves the trouble of hand-rolling a socket protocol and makes cross-language use easy. It isn't a real-time control bus, though — millisecond-scale low-level motor loops are still left to something like EtherCAT or CAN.","example":"Boston Dynamics' official Spot SDK exposes a gRPC-based interface, letting Python code send commands and read state remotely.","related":["Protocol Buffers","Policy Server (Remote Inference)","WebSocket","ZeroMQ","Middleware"]},{"id":"websocket","category":"software","sec":3,"tier":3,"sources":[{"title":"RFC 6455: The WebSocket Protocol","url":"https://datatracker.ietf.org/doc/html/rfc6455"},{"title":"The WebSocket API — MDN","url":"https://developer.mozilla.org/en-US/docs/Web/API/WebSockets_API"}],"as_of":"","related_ids":["policy-server","rosbridge","foxglove-studio","openpi","web-real-time-communication","grpc-remote-procedure-calls"],"name":"WebSocket","alt":"WebSocket","abbr":"","aliases":["WS"],"one_liner":"A protocol for sending messages both ways, in real time, over a single TCP connection.","explanation":"WebSocket is a standard protocol published by the IETF in 2011 (RFC 6455). Ordinary HTTP is “the client asks once, the server answers once”; the server can't push data on its own. WebSocket starts with a single HTTP request that upgrades the connection, after which both sides can send each other messages at any time over the same TCP connection, with low overhead and low latency, and ready-made libraries exist for browsers and nearly every language. It's used widely in robotics software: rosbridge uses it to expose ROS topics to web pages, Foxglove's live connection is also built on it, and some VLA projects use it to move observations and actions between a GPU server and a robot. Because it's built on TCP, a dropped packet triggers a retransmission wait; high-bitrate video streaming usually switches to WebRTC instead.","example":"openpi starts a WebSocket policy server on a GPU server; on every step, the robot sends over an image and its joint state, and gets back an action chunk to execute.","related":["Policy Server (Remote Inference)","rosbridge","Foxglove Studio","openpi (Physical Intelligence)","Web Real-Time Communication (WebRTC)","gRPC Remote Procedure Calls"]},{"id":"web-real-time-communication","category":"software","sec":3,"tier":3,"sources":[{"title":"WebRTC 官网","url":"https://webrtc.org/"},{"title":"WebRTC: Real-Time Communication in Browsers (W3C)","url":"https://www.w3.org/TR/webrtc/"}],"as_of":"","related_ids":["websocket","teleoperation","control-latency","vuer","ffmpeg"],"name":"Web Real-Time Communication (WebRTC)","alt":"WebRTC 实时音视频传输","abbr":"WebRTC","aliases":["WebRTC"],"one_liner":"A browser-native standard for low-latency, peer-to-peer audio, video, and data transfer.","explanation":"WebRTC is an open standard jointly developed by the W3C (web APIs) and the IETF (network protocols), driven primarily by Google, and it became an official W3C recommendation in 2021. It lets browsers and applications send audio, video, and arbitrary data directly to each other, peer to peer, without routing through a server, with built-in codecs, jitter buffering, and congestion control; end-to-end latency can typically be kept under a few hundred milliseconds. The two ends are often on different local networks, so an ICE process working with STUN/TURN servers is needed to get through NAT (the address translation done by routers). In embodied AI, it's commonly used to stream a robot's camera feed live to a remote operator, and some robot makers also use it to connect a phone app to the robot.","example":"During remote teleoperation, the robot encodes its head-camera feed and pushes it over WebRTC to the operator's browser or VR headset, with control commands sent back over WebRTC's data channel.","related":["WebSocket","Teleoperation","Control Latency","Vuer","FFmpeg"]},{"id":"rosbridge","category":"software","sec":3,"tier":3,"sources":[{"title":"RobotWebTools/rosbridge_suite (GitHub)","url":"https://github.com/RobotWebTools/rosbridge_suite"},{"title":"ROS Wiki: rosbridge_suite","url":"http://wiki.ros.org/rosbridge_suite"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","websocket","topic","foxglove-studio",null],"name":"rosbridge","alt":"rosbridge","abbr":"","aliases":["rosbridge_suite","rosbridge_server"],"one_liner":"A WebSocket-and-JSON bridge that exposes ROS to web pages and other non-ROS programs.","explanation":"rosbridge (the package set is called rosbridge_suite) defines a JSON-based protocol and provides a WebSocket server, rosbridge_server. Non-ROS programs — a browser page, Unity, a phone app — can subscribe to and publish topics, call services, and read or write parameters simply by opening a WebSocket connection and exchanging JSON, without needing ROS installed locally at all. Web front-ends usually pair it with the roslibjs library. It works with both ROS 1 and ROS 2. The tradeoff is that JSON serialization is relatively heavy, making it a poor fit for high-rate point clouds or images; it's better suited to lightweight use cases like monitoring dashboards or remote teleoperation.","example":"Launch the rosbridge_websocket launch file on the robot, and a web page using roslibjs can connect to ws://<robot-IP>:9090 to display a battery-level topic and publish cmd_vel to drive the base.","related":["Robot Operating System","Robot Operating System 2","WebSocket","Topic (ROS)","Foxglove Studio","cmd_vel velocity command topic"]},{"id":"model-context-protocol","category":"software","sec":3,"tier":3,"sources":[{"title":"Model Context Protocol 官网","url":"https://modelcontextprotocol.io/"},{"title":"ros-mcp-server GitHub","url":"https://github.com/robotmcp/ros-mcp-server"}],"as_of":"2026-09","related_ids":[null,"rosbridge",null,null,null],"name":"Model Context Protocol (ROS MCP Server)","alt":"MCP / ROS-MCP-Server","abbr":"MCP","aliases":["model context protocol","ros-mcp-server"],"one_liner":"A bridge that lets large models call ROS robot interfaces through a single, unified protocol.","explanation":"MCP (Model Context Protocol) is an open standard Anthropic released in November 2024, defining a common format for how an LLM application calls external tools and data sources. ROS-MCP-Server is a community open-source project that exposes ROS as an MCP tool: it connects to a robot through rosbridge (ROS's WebSocket interface), letting models such as Claude or GPT list topics, read sensor data, publish velocity commands, or call services — all without modifying the robot's existing code, and it supports both ROS 1 and ROS 2. It's well suited to quickly prototyping natural-language robot control, but having a large model issue commands directly lacks real-time guarantees and safety assurances, so it can't substitute for low-level control.","example":"Telling a large model “move the cart forward one meter” makes it publish a velocity message to the cmd_vel topic via ROS-MCP-Server.","related":["Large Language Model (LLM)","rosbridge","ROS 2","LLM-based Task Planning","cmd_vel Topic (geometry_msgs/Twist velocity command)"]},{"id":"dataflow-oriented-robotic-architecture","category":"software","sec":3,"tier":3,"sources":[{"title":"dora-rs GitHub","url":"https://github.com/dora-rs/dora"},{"title":"dora-rs 官网","url":"https://dora-rs.ai/"}],"as_of":"","related_ids":["robot-operating-system-2","middleware","zero-copy","eclipse-iceoryx","publish-subscribe","eclipse-zenoh"],"name":"Dataflow-Oriented Robotic Architecture (dora)","alt":"dora-rs","abbr":"dora","aliases":["dora"],"one_liner":"A Rust-based robotics dataflow middleware positioned as a lower-latency alternative to ROS.","explanation":"dora-rs is an open-source robotics middleware framework with its core written in Rust. It splits a robot program into a number of nodes and describes who sends data to whom as a dataflow graph; nodes pass data to each other through shared memory and Apache Arrow (a columnar in-memory data format), avoiding copies wherever possible and thereby cutting transmission latency for large data such as camera images and point clouds. Nodes can be written in Python, Rust, or C/C++, which makes it fairly friendly toward Python nodes running AI models. It targets the same class of problems as ROS 2 — inter-process communication and system orchestration — and provides a bridge for interoperating with ROS 2, and is often benchmarked against ROS 2 on latency.","example":"A YAML file declares a three-stage dataflow — “camera node → VLA inference node → arm control node” — and dora starts each node and passes images and actions between them.","related":["Robot Operating System 2","Middleware","Zero-Copy (Shared-Memory IPC)","Eclipse iceoryx","Publish-Subscribe","Eclipse Zenoh"]},{"id":"aimrt","category":"software","sec":3,"tier":3,"sources":[{"title":"AimRT/AimRT (GitHub)","url":"https://github.com/AimRT/AimRT"}],"as_of":"2024","related_ids":["robot-operating-system-2","middleware","dataflow-oriented-robotic-architecture","agibot","grpc-remote-procedure-calls","data-distribution-service"],"name":"AimRT (AgiBot robot runtime framework)","alt":"AimRT（智元机器人运行时框架）","abbr":"","aliases":["AgiBot AimRT"],"one_liner":"AgiBot's open-source, modern C++ robot runtime framework, interoperable with ROS 2 and other communication backends.","explanation":"AimRT is a robot runtime framework open-sourced by AgiBot (智元机器人) in 2024, written in modern C++, occupying a similar niche to ROS 2: it splits robot software into modules and handles loading, scheduling, inter-module communication, and logging configuration for them. Its distinguishing feature is a lightweight core, with communication backends and other capabilities plugged in as modules — it can use ROS 2, HTTP, or gRPC as its transport, so it can run standalone or interoperate with existing ROS 2 nodes, and it supports deployment all the way from the edge to the cloud. It aims to address some of the performance, resource-usage, and engineering-deployment pain points found in ROS 2. It's a useful reference point for developers using AgiBot hardware, or anyone curious about Chinese-built robot middleware.","example":"","related":["Robot Operating System 2","Middleware","Dataflow-Oriented Robotic Architecture (dora)","AgiBot","gRPC Remote Procedure Calls","Data Distribution Service"]},{"id":"software-development-kit","category":"software","sec":4,"tier":1,"sources":[{"title":"unitree_sdk2 (GitHub)","url":"https://github.com/unitreerobotics/unitree_sdk2"},{"title":"Franka Robotics 文档：libfranka client library","url":"https://frankarobotics.github.io/docs/doc/libfranka/docs/index.html"}],"as_of":"","related_ids":["secondary-development","unitree-sdk2","libfranka-franka-control-interface","driver","robot-operating-system-2"],"name":"Software Development Kit","alt":"软件开发工具包","abbr":"SDK","aliases":["SDK"],"one_liner":"The libraries, interfaces, and examples a manufacturer provides so developers can control a device from code.","explanation":"An SDK is the full package of development materials a software or hardware maker gives to developers, typically including libraries, API documentation, sample code, drivers, and debugging tools. In robotics, once you’ve bought a robot arm, quadruped, or humanoid, reading its joint states, sending position or torque commands, and pulling images from its camera is usually all done through the manufacturer’s SDK — what’s commonly called building on top of the platform. Most SDKs offer C++ and Python interfaces, and many also ship a ROS / ROS 2 wrapper. How open an SDK is directly determines what you can actually do with the robot — for instance, whether it exposes low-level joint torque control, or only lets you call high-level walking and gesture commands.","example":"The Unitree G1 reads out every joint angle and sends motor commands through unitree_sdk2, which is how researchers deploy their own trained locomotion policies on it.","related":["Secondary Development (custom development on SDK)","Unitree SDK2","libfranka / Franka Control Interface (FCI)","Driver (Device Driver)","Robot Operating System 2"]},{"id":"secondary-development","category":"software","sec":4,"tier":2,"sources":[{"title":"Software development kit - Wikipedia","url":"https://en.wikipedia.org/wiki/Software_development_kit"},{"title":"Unitree 开发者文档","url":"https://support.unitree.com/home/zh/developer"}],"as_of":"","related_ids":["software-development-kit","unitree-sdk2","edu-edition","driver","robot-operating-system-2","hardware-abstraction-layer"],"name":"Secondary Development (custom development on SDK)","alt":"二次开发","abbr":"","aliases":["open interface","SDK development"],"one_liner":"Writing your own programs against a vendor's SDK or interface to customize a robot, without touching its firmware.","explanation":"“Secondary development” (二次开发) is a common phrase in Chinese engineering circles for building custom functionality on top of a vendor's published SDK, API, or ROS interface, without modifying the vendor's underlying firmware. For embodied-AI research, whether a purchased robot supports secondary development matters a great deal: only if you can read joint states and camera data, and send joint or velocity commands, can you deploy a policy you trained yourself. Because of this, many manufacturers sell a standard version alongside an EDU (research/education) version with a more open low-level interface — the two differ in both how much access they give and in price.","example":"Buy the EDU edition of the Unitree G1, use the Unitree SDK to read joint angles and send torque commands, and deploy a reinforcement-learning walking policy trained in simulation onto the real robot.","related":["Software Development Kit","Unitree SDK2","EDU Edition","Driver (Device Driver)","Robot Operating System 2","Hardware Abstraction Layer (HAL)"]},{"id":"unitree-sdk2","category":"software","sec":4,"tier":2,"sources":[{"title":"unitreerobotics/unitree_sdk2 (GitHub)","url":"https://github.com/unitreerobotics/unitree_sdk2"},{"title":"unitreerobotics/unitree_ros2 (GitHub)","url":"https://github.com/unitreerobotics/unitree_ros2"}],"as_of":"","related_ids":["unitree-robotics","software-development-kit","secondary-development","eclipse-cyclone-dds","robot-operating-system-2","unitree-g1"],"name":"Unitree SDK2","alt":"宇树 SDK","abbr":"","aliases":["unitree_sdk2","unitree_sdk2_python","unitree_ros2"],"one_liner":"Unitree's official open-source development kit for reading sensors and commanding joints and motion on its robots.","explanation":"Unitree SDK2 refers to unitree_sdk2 (C++) and its Python counterpart unitree_sdk2_python, open-sourced by Unitree Robotics on GitHub for the Go2, B2, H1, and G1 platforms. It's built on DDS communication and offers two levels of interface: a high-level motion interface that lets you send velocity and posture commands directly to make the robot walk, and a low-level interface that sends target position, velocity, stiffness, damping, and feedforward torque to individual joints while reading back IMU and joint state — this is the level used when deploying a reinforcement-learning policy. Because ROS 2 is also built on DDS, the companion unitree_ros2 package lets ROS 2 programs communicate with the robot directly.","example":"Using unitree_sdk2_python to read a G1's joint states in a loop at roughly 500 Hz, then sending the policy network's output joint-angle targets with kp and kd gains attached, completes a sim-to-real deployment.","related":["Unitree Robotics","Software Development Kit","Secondary Development (custom development on SDK)","Eclipse Cyclone DDS","Robot Operating System 2","Unitree G1"]},{"id":"libfranka-franka-control-interface","category":"software","sec":4,"tier":3,"sources":[{"title":"Franka Control Interface 文档","url":"https://frankaemika.github.io/docs/"},{"title":"frankaemika/libfranka (GitHub)","url":"https://github.com/frankaemika/libfranka"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"libfranka / Franka Control Interface (FCI)","alt":"libfranka","abbr":"FCI","aliases":["franka_ros","franka_ros2","Franka Control Interface"],"one_liner":"Franka's official low-level C++ control library, exchanging commands with the robot in real time at 1 kHz through the FCI.","explanation":"libfranka is the official open-source C++ control library for Franka arms (Panda / FR3), connecting over Ethernet to the robot's Franka Control Interface (FCI). In a callback function, the user reads the robot's state once every millisecond — joint angles, velocities, estimated external forces — and returns a command each time; available control modes include joint torque, joint position/velocity, and Cartesian pose/velocity, which means the computer running it needs a real-time kernel (such as PREEMPT_RT) to guarantee it never misses a cycle. franka_ros and franka_ros2 are its ROS wrappers. Many robot-learning frameworks — Polymetis and Deoxys among them — call into it under the hood to implement impedance control and policy execution.","example":"On a machine running a PREEMPT_RT kernel, write a 1 kHz torque callback with libfranka to implement Cartesian impedance control, so a VLA policy's output end-effector target pose gets tracked compliantly.","related":["Franka Emika Panda / Franka Research 3","Torque Control","Cartesian Impedance Control","PREEMPT_RT Real-Time Linux Patch","RTDE (Real-Time Data Exchange, Universal Robots)","Software Development Kit (SDK)"]},{"id":"rtde","category":"software","sec":4,"tier":3,"sources":[{"title":"Universal Robots: Real-Time Data Exchange (RTDE) Guide","url":"https://www.universal-robots.com/articles/ur/interface-communication/real-time-data-exchange-rtde-guide/"},{"title":"ur_rtde documentation (SDU Robotics)","url":"https://sdurobotics.gitlab.io/ur_rtde/"}],"as_of":"","related_ids":["universal-robots-ur5e","universal-robots","streaming-servo-control","software-development-kit","libfranka-franka-control-interface","host-computer"],"name":"RTDE","alt":"RTDE","abbr":"RTDE","aliases":["Real-Time Data Exchange","ur_rtde"],"one_liner":"Universal Robots' real-time interface that lets external programs read arm state and send commands.","explanation":"RTDE is a communication interface built into Universal Robots (UR) controllers. An external computer connects over TCP/IP and can synchronously read state — joint angles, end-effector pose, torques — at a fixed rate, and also write to input registers; e-Series controllers support up to 500 Hz. It solves the problem of a host computer needing to exchange data with the arm at high, stable rates, and is more structured than the older ports 30001–30003. The most widely used community tool is ur_rtde, an open-source C++/Python library from the University of Southern Denmark, which wraps commands like moveL and servoJ directly; many real-robot VLA deployments use it to control UR arms.","example":"Use ur_rtde's RTDEReceiveInterface to read a UR5e's current joint angles at 500 Hz, then use RTDEControlInterface's servoJ to stream a policy's target joint angles to the arm.","related":["Universal Robots UR5e","Universal Robots","Streaming Servo Control","Software Development Kit","libfranka / Franka Control Interface (FCI)","Host Computer"]},{"id":"intel-realsense-sdk-2-0","category":"software","sec":4,"tier":3,"sources":[{"title":"IntelRealSense/librealsense (GitHub)","url":"https://github.com/IntelRealSense/librealsense"},{"title":"IntelRealSense/realsense-ros (GitHub)","url":"https://github.com/IntelRealSense/realsense-ros"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"Intel RealSense SDK 2.0 (librealsense)","alt":"RealSense SDK","abbr":"","aliases":["librealsense","realsense-ros","pyrealsense2"],"one_liner":"The official open-source driver and development library for RealSense depth cameras, handling capture, alignment, and point clouds.","explanation":"RealSense SDK 2.0 is the open-source library librealsense, the official software development kit for RealSense depth cameras (such as the D435i and D405), originally released by Intel, supporting Linux, Windows, and macOS. It connects to the camera and reads its color, depth, infrared, and IMU (inertial measurement unit) data streams, and provides post-processing filters for aligning depth to color, generating point clouds, and hole filling, along with the realsense-viewer visualization and debugging tool. It's mainly written in C++, with a Python interface, pyrealsense2, and a ROS wrapper, realsense-ros. Embodied-AI labs use RealSense heavily as both wrist cameras and third-person cameras, so this library shows up constantly in both data collection and policy deployment.","example":"Open a D435i with pyrealsense2, align the depth image to the color image, back-project it into a point cloud, and feed that to a 3D diffusion policy as input.","related":["RealSense Depth Camera (D435i / D405)","Depth Camera (RGB-D Camera)","Depth-to-Color Alignment (Depth Registration)","Point Cloud","Camera Intrinsics","Driver (device driver)"]},{"id":"driver","category":"software","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Device driver","url":"https://en.wikipedia.org/wiki/Device_driver"},{"title":"IntelRealSense/realsense-ros (GitHub)","url":"https://github.com/IntelRealSense/realsense-ros"}],"as_of":"","related_ids":["firmware","software-development-kit","hardware-abstraction-layer","ros2-control","intel-realsense-sdk-2-0","node"],"name":"Driver (Device Driver)","alt":"驱动","abbr":"","aliases":["ROS Driver","Device Driver"],"one_liner":"The layer of software that lets an operating system or higher-level program read and write to a piece of hardware.","explanation":"A driver sits between hardware and higher-level software, translating requests like “read a frame of video” or “turn the motor to this angle” into low-level instructions the hardware understands, and organizing whatever the hardware returns into a format a program can use. The term has two related senses: an operating-system-level device driver, such as a graphics driver or a USB-serial driver; and a “ROS driver” in robot software — a ROS node, written by the manufacturer or the community, that publishes data from a camera, lidar, or motor onto standard topics. When bringing new hardware into a system, checking whether a driver is already available is usually the first step.","example":"After installing the realsense-ros driver package, launching it makes a RealSense camera’s color and depth images available on topics such as /camera/color/image_raw.","related":["Firmware","Software Development Kit","Hardware Abstraction Layer (HAL)","ros2_control","Intel RealSense SDK 2.0 (librealsense)","Node (ROS)"]},{"id":"hardware-abstraction-layer","category":"software","sec":4,"tier":3,"sources":[{"title":"Hardware abstraction - Wikipedia","url":"https://en.wikipedia.org/wiki/Hardware_abstraction"},{"title":"ros2_control 官方文档","url":"https://control.ros.org/"}],"as_of":"","related_ids":["ros2-control","driver","software-development-kit","secondary-development","one-brain-multiple-robots"],"name":"Hardware Abstraction Layer (HAL)","alt":"硬件抽象层","abbr":"HAL","aliases":[],"one_liner":"A layer of code that hides specific hardware differences and exposes a uniform interface to the software above it.","explanation":"A hardware abstraction layer is a common design pattern in operating systems and software engineering: a uniform interface is placed between specific hardware — motor drivers, cameras, sensors — and the higher-level program, so the upper layer only calls standard functions like “read joint angle” or “send torque command” without caring whether the bus underneath is CAN or EtherCAT, or which vendor's motor it is. That way, switching hardware only requires changing or adding this one layer, leaving the control algorithms and policy code untouched. In robotics, ROS 2's ros2_control implements exactly this through “hardware interface” plugins; embodied models aiming for “one brain, many bodies” — the same policy running on different robot bodies — also can't do without a clean HAL.","example":"Writing a hardware interface plugin for a particular arm model in ros2_control lets MoveIt and controllers above it drive that arm directly.","related":["ros2_control","Driver (Device Driver)","Software Development Kit","Secondary Development (custom development on SDK)","One Brain, Multiple Robots"]},{"id":"ros2-control","category":"software","sec":4,"tier":2,"sources":[{"title":"ros2_control 官方文档","url":"https://control.ros.org/"},{"title":"ros-controls/ros2_control (GitHub)","url":"https://github.com/ros-controls/ros2_control"}],"as_of":"","related_ids":["robot-operating-system-2","hardware-abstraction-layer","moveit-motion-planning-framework","unified-robot-description-format","real-time-control","proportional-derivative-control"],"name":"ros2_control","alt":"ros2_control","abbr":"","aliases":["ros_control"],"one_liner":"The standard ROS 2 control framework that connects control algorithms to real motor hardware.","explanation":"ros2_control is the official control framework in the ROS 2 ecosystem, maintained by the ros-controls community as the successor to ROS 1's ros_control. It splits a system into three layers: the hardware interface (written once per robot, unifying a motor's position, velocity, and torque reads and writes into a standard interface), the controller manager (which runs a fixed-frequency control loop), and controllers themselves (such as a joint trajectory controller or a differential-drive controller). The benefit is that control algorithms are decoupled from specific hardware — switching to a new robot only requires rewriting the hardware interface, while controllers can be reused directly. Trajectories planned by MoveIt are typically handed off to it for execution, and the Gazebo simulator has a matching plugin.","example":"Write a hardware_interface plugin for a custom robot arm, declare its joint interfaces in the URDF, then load a joint_trajectory_controller so trajectories planned by MoveIt can drive the real motors.","related":["Robot Operating System 2","Hardware Abstraction Layer (HAL)","MoveIt Motion Planning Framework","Unified Robot Description Format","Real-Time Control","Proportional-Derivative Control"]},{"id":"embedded-software-development","category":"software","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Embedded software","url":"https://en.wikipedia.org/wiki/Embedded_software"}],"as_of":"","related_ids":["microcontroller-unit","real-time-operating-system","firmware","cross-compilation","stmicroelectronics-stm32-mcu-family","controller-area-network"],"name":"Embedded Software Development","alt":"嵌入式软件开发","abbr":"","aliases":["Embedded Development","Microcontroller Programming"],"one_liner":"Writing the program that runs on the chip inside a device such as a motor driver board or a sensor.","explanation":"Embedded software runs on the chip built into a device — usually a microcontroller (MCU, such as an STM32) or an SoC (system-on-chip) — which has limited resources and needs to run reliably for long stretches. Development is usually done in C/C++, either operating hardware registers directly (bare metal) or running on top of a real-time operating system such as FreeRTOS; the code is cross-compiled on a computer and then flashed onto the chip, with a JTAG/SWD debugger used to step through it. Joint-motor control, sensor sampling, battery management, and low-level bus communication (CAN, EtherCAT) in a robot all belong to this layer, and it determines whether control is real-time and reliable — it’s also an important job category at embodied AI companies, distinct from algorithms work.","example":"A current-loop program written for an STM32 reads the encoder at 20 kHz, computes the PWM duty cycle, and reports the joint angle back to the host computer over a CAN bus.","related":["Microcontroller Unit (MCU)","Real-Time Operating System (RTOS)","Firmware","Cross-compilation","STMicroelectronics STM32 MCU Family","Controller Area Network (CAN)"]},{"id":"cross-compilation","category":"software","sec":4,"tier":3,"sources":[{"title":"Cross compiler - Wikipedia","url":"https://en.wikipedia.org/wiki/Cross_compiler"},{"title":"cmake-toolchains(7) - CMake Documentation","url":"https://cmake.org/cmake/help/latest/manual/cmake-toolchains.7.html"}],"as_of":"","related_ids":["embedded-software-development","cmake","nvidia-jetson","docker","on-device-edge-deployment","firmware"],"name":"Cross-compilation","alt":"交叉编译","abbr":"","aliases":["cross-building"],"one_liner":"Compiling on an x86 computer a program meant to run on a robot's ARM-based onboard controller.","explanation":"Cross-compilation means building, on one machine (the host), an executable meant to run on a different CPU architecture or operating system. In robotics this is a very practical need: the onboard controller is often an ARM board like a Jetson, Raspberry Pi, or Rockchip device, with limited compute and memory, and compiling a large project directly on it — a ROS 2 workspace, OpenCV, an inference engine — can take forever or fail to complete at all, while the development machine is x86. The approach is to install a toolchain for the target architecture (e.g. aarch64-linux-gnu-gcc) plus the target system's headers and libraries, and configure a CMake toolchain file accordingly; Docker combined with QEMU emulation of the target architecture is also common, as a way to sidestep environment mismatches.","example":"Compile ROS 2 nodes and an inference program on an x86 workstation using an aarch64 toolchain, then copy them onto a Jetson Orin to run directly.","related":["Embedded Software Development","CMake","NVIDIA Jetson","Docker","On-Device / Edge Deployment","Firmware"]},{"id":"firmware","category":"software","sec":4,"tier":3,"sources":[{"title":"Firmware - Wikipedia","url":"https://en.wikipedia.org/wiki/Firmware"}],"as_of":"","related_ids":["driver","over-the-air-update","microcontroller-unit","embedded-system","servo-drive","embedded-software-development"],"name":"Firmware","alt":"固件","abbr":"","aliases":[],"one_liner":"The program burned into a hardware chip that directly controls a device's lowest-level behavior.","explanation":"Firmware is a program stored in a device's non-volatile memory (such as flash), usually running on a microcontroller (MCU), handling the work closest to the hardware. On a robot, joint drivers, dexterous hands, IMUs, cameras, and the battery management system each have their own firmware — a driver's firmware, for instance, implements the current loop, the velocity loop, and over-current and over-temperature protection, while a camera's firmware handles exposure and depth computation. A higher-level ROS node or SDK is really just talking to this firmware over a bus. A firmware version mismatch is a common cause of communication errors or missing functionality, so checking and updating firmware is often the first step when setting up new hardware or debugging a problem; a whole robot can also often be updated remotely via OTA.","example":"When a RealSense camera's depth output looks wrong, first check its firmware version with the official tool and update to the version the SDK recommends before testing further.","related":["Driver (Device Driver)","Over-the-Air (OTA) Update","Microcontroller Unit (MCU)","Embedded System","Servo Drive (Motor Driver)","Embedded Software Development"]},{"id":"over-the-air-update","category":"software","sec":4,"tier":2,"sources":[{"title":"Over-the-air update - Wikipedia","url":"https://en.wikipedia.org/wiki/Over-the-air_update"}],"as_of":"","related_ids":["firmware","driver","mass-production","scaled-deployment","on-device-edge-deployment"],"name":"Over-the-Air (OTA) Update","alt":"OTA 升级","abbr":"OTA","aliases":["OTA","over-the-air update"],"one_liner":"Updating a device's firmware, system, or models remotely over a network, without opening it up or plugging in a cable.","explanation":"An OTA update means a device downloads and installs new software wirelessly, over Wi-Fi or cellular, first popularized in phones and cars. For a robot, what gets updated can be low-level firmware (the code in driver boards or joint motors), the higher-level operating system, or even a motion-control policy or large model's weights. It solves a specific problem: once machines are sold and scattered across many sites, how do you keep fixing bugs and adding features without shipping them back to the factory or sending an engineer out? Building OTA properly means handling rollback after a power loss, version verification, and staged rollout — get it wrong, and a failed update can leave a robot unable to boot. At the mass-production and scaled-deployment stage, OTA capability is often treated as a sign that a product has matured.","example":"A humanoid robot maker pushes a new motion-control policy to robots already delivered to customers; the user taps update in an app, the robot downloads it, and gains the new behavior after a restart.","related":["Firmware","Driver (Device Driver)","Mass Production","Scaled Deployment","On-Device / Edge Deployment"]},{"id":"real-time-operating-system","category":"software","sec":4,"tier":2,"sources":[{"title":"Real-time operating system - Wikipedia","url":"https://en.wikipedia.org/wiki/Real-time_operating_system"},{"title":"FreeRTOS","url":"https://www.freertos.org/"}],"as_of":"","related_ids":["freertos","rt-thread","preempt-rt","xenomai","hard-real-time","jitter"],"name":"Real-Time Operating System (RTOS)","alt":"实时操作系统","abbr":"RTOS","aliases":["RTOS","real-time system"],"one_liner":"An operating system that guarantees tasks finish within a set deadline, used for timing-sensitive control.","explanation":"What defines a real-time operating system is not raw speed but punctuality: through mechanisms such as priority-based preemptive scheduling, it guarantees critical tasks execute within a fixed deadline, with small and predictable timing jitter. A regular desktop Linux is optimized for throughput, and an occasional few milliseconds of delay doesn't matter; but a robot's joint current loop or force-control loop needs to run at a fixed period — 1 kHz or higher — reliably, and missing a cycle can cause jitter or instability. There are two common approaches: running a lightweight RTOS such as FreeRTOS, RT-Thread, or Zephyr on a microcontroller, or using a main controller running Linux patched with PREEMPT_RT, or Xenomai, to get real-time behavior. Higher-level AI policies typically run on a non-real-time system, with low-level control kept on the real-time system, the two communicating over a bus or shared memory.","example":"The MCU on a joint driver board runs FreeRTOS to execute the current loop at a fixed period, while the main controller runs a PREEMPT_RT Linux kernel to drive a 1 kHz EtherCAT master.","related":["FreeRTOS","RT-Thread","PREEMPT_RT","Xenomai","Hard Real-Time","Jitter (Timing Jitter)"]},{"id":"hard-real-time","category":"software","sec":4,"tier":3,"sources":[{"title":"Real-time computing - Wikipedia","url":"https://en.wikipedia.org/wiki/Real-time_computing"}],"as_of":"","related_ids":["real-time-operating-system","preempt-rt","xenomai","ethercat","control-frequency","jitter"],"name":"Hard Real-Time","alt":"硬实时","abbr":"","aliases":[],"one_liner":"A guarantee that every single execution finishes before its deadline — missing even once counts as a system failure.","explanation":"Hard real-time is the strictest tier of real-time computing: a task must complete within its specified deadline every single time, and missing it even once is treated as an error, potentially leading to an accident. The looser tier, soft real-time, tolerates the occasional miss with just some degradation in quality — a video dropping a frame now and then, for instance. A robot's low-level motor current loop, joint torque control, and bus communication typically demand hard real-time, often with periods of 1 kHz or higher, because computing a cycle late can cause the motor to jitter or lose control entirely. Achieving hard real-time generally relies on a real-time operating system (RTOS), Linux patched with PREEMPT_RT, or Xenomai, combined with a deterministic bus such as EtherCAT; large-model inference doesn't meet hard-real-time requirements, so it's usually kept at a higher level, with real-time control reserved for a dedicated low-level controller.","example":"A humanoid robot's joint drivers run a torque loop at 1 kHz, which must finish and send its output every 1 millisecond without fail — not something that can be run on an ordinary Python process.","related":["Real-Time Operating System (RTOS)","PREEMPT_RT","Xenomai","EtherCAT (Ethernet for Control Automation Technology)","Control Frequency","Jitter (Timing Jitter)"]},{"id":"freertos","category":"software","sec":4,"tier":3,"sources":[{"title":"FreeRTOS 官网","url":"https://www.freertos.org/"},{"title":"FreeRTOS/FreeRTOS-Kernel (GitHub)","url":"https://github.com/FreeRTOS/FreeRTOS-Kernel"}],"as_of":"","related_ids":["real-time-operating-system","microcontroller-unit","stmicroelectronics-stm32-mcu-family","micro-ros","embedded-system","rt-thread"],"name":"FreeRTOS","alt":"FreeRTOS","abbr":"","aliases":[],"one_liner":"A small open-source real-time operating system kernel that runs on microcontrollers.","explanation":"FreeRTOS is an open-source real-time operating system (RTOS — an OS that guarantees tasks respond within a set deadline) kernel first released by Richard Barry around 2003, taken over by Amazon Web Services (AWS) in 2017, and released under the MIT license. It's extremely small, able to run on microcontrollers with only tens of kilobytes of memory, and provides basic facilities like task scheduling, queues, semaphores, and timers, letting a developer split motor control, sensor reading, and communication into several tasks that run by priority. Lower-level boards on a robot — motor driver boards, IMU boards, dexterous-hand controller boards — commonly run it; Espressif's official ESP32 development framework is also built on FreeRTOS, and micro-ROS supports it as well.","example":"On an STM32 joint driver board, FreeRTOS sets up three tasks: a 1 kHz current-loop task, a CAN-communication task for receiving commands, and a low-priority temperature-monitoring task.","related":["Real-Time Operating System (RTOS)","Microcontroller Unit (MCU)","STMicroelectronics STM32 MCU Family","micro-ROS","Embedded System","RT-Thread"]},{"id":"rt-thread","category":"software","sec":4,"tier":3,"sources":[{"title":"RT-Thread 官网","url":"https://www.rt-thread.org/"},{"title":"RT-Thread/rt-thread (GitHub)","url":"https://github.com/RT-Thread/rt-thread"}],"as_of":"","related_ids":["real-time-operating-system","freertos","microcontroller-unit","stmicroelectronics-stm32-mcu-family","embedded-system","lower-level-controller"],"name":"RT-Thread","alt":"RT-Thread","abbr":"","aliases":["RT-Thread Nano","RT-Thread Smart"],"one_liner":"An open-source embedded real-time operating system developed by a team in China.","explanation":"RT-Thread is an open-source embedded real-time operating system (RTOS — a system that guarantees a task gets a response within a set time) led and maintained by the Shanghai company RT-Thread (Ruixinde Electronics). The project started in 2006 and is now licensed under Apache 2.0. It comes in a standard edition with a driver framework, file system, network stack, and package ecosystem; a Nano edition for microcontrollers with very limited resources; and a Smart edition for processors with a memory management unit, supporting architectures including ARM Cortex-M and RISC-V. In robotics, it commonly runs on microcontrollers such as joint drive boards, sensor boards, and chassis controllers, handling hard-real-time jobs like motor control and bus communication, working alongside a host computer running Linux and ROS; its niche is similar to FreeRTOS's.","example":"Use an STM32 as a wheeled base's control board: run threads on RT-Thread for 1 kHz motor control and CAN communication, then send odometry over a serial link to a host computer running ROS 2.","related":["Real-Time Operating System (RTOS)","FreeRTOS","Microcontroller Unit (MCU)","STMicroelectronics STM32 MCU Family","Embedded System","Lower-Level Controller"]},{"id":"micro-ros","category":"software","sec":4,"tier":3,"sources":[{"title":"micro-ROS 官网","url":"https://micro.ros.org/"}],"as_of":"","related_ids":[null,null,"freertos",null,null],"name":"micro-ROS","alt":"micro-ROS","abbr":"","aliases":[],"one_liner":"A framework that brings ROS 2 onto microcontrollers.","explanation":"micro-ROS is an open-source framework that lets even microcontrollers (chips with only tens to hundreds of kilobytes of memory) join a ROS 2 system, originating from an EU-funded project with contributions from eProsima, Bosch, and others, now maintained by an official ROS working group. A microcontroller can't run a full DDS communication middleware, so micro-ROS substitutes the lightweight Micro XRCE-DDS instead, running on real-time operating systems such as FreeRTOS or Zephyr, and connecting over a serial port or UDP to an Agent running on a host computer, which forwards its traffic into standard ROS 2 topics. That lets a motor driver board or sensor board publish and subscribe to topics directly, without the need to hand-write a serial protocol.","example":"Running micro-ROS on an ESP32 publishes IMU data as a ROS 2 topic, which a node on a computer can subscribe to directly.","related":["ROS 2","Microcontroller Unit (MCU)","FreeRTOS","Data Distribution Service (DDS)","Lower-level Controller (Slave Computer)"]},{"id":"preempt-rt","category":"software","sec":4,"tier":3,"sources":[{"title":"Real-Time Linux (Linux Foundation Wiki)","url":"https://wiki.linuxfoundation.org/realtime/start"}],"as_of":"2024-11","related_ids":["real-time-operating-system","hard-real-time","real-time-control","jitter","ethercat-master","xenomai"],"name":"PREEMPT_RT","alt":"实时内核补丁","abbr":"PREEMPT_RT","aliases":["PREEMPT_RT real-time Linux patch","RT-Linux","real-time Linux"],"one_liner":"A real-time patch set that makes the Linux kernel almost fully preemptible, cutting worst-case latency.","explanation":"PREEMPT_RT is a real-time patch set maintained for years by the Linux community. It converts most of the kernel's non-interruptible code paths into preemptible ones, turns interrupt handlers into threads, and replaces spinlocks with sleepable mutexes, pushing the worst-case response latency for high-priority tasks down to roughly tens of microseconds. It was merged into the mainline kernel in 2024 (version 6.12); before that, it had to be patched in and compiled by hand. Robot control needs to send commands on time, every cycle — a stock kernel's occasional multi-millisecond stall can make a motor jitter or trip a safety cutoff — so tools like ros2_control, EtherCAT masters, and Franka's libfranka typically require or recommend a real-time kernel.","example":"Before controlling a Franka arm at 1 kHz with libfranka, install a PREEMPT_RT kernel on Ubuntu first, or communication timeout errors are likely.","related":["Real-Time Operating System (RTOS)","Hard Real-Time","Real-Time Control","Jitter (Timing Jitter)","EtherCAT Master (SOEM / IgH)","Xenomai"]},{"id":"xenomai","category":"software","sec":4,"tier":3,"sources":[{"title":"Xenomai 官网","url":"https://xenomai.org/"},{"title":"Xenomai - Wikipedia","url":"https://en.wikipedia.org/wiki/Xenomai"}],"as_of":"","related_ids":["real-time-operating-system","preempt-rt","hard-real-time","ethercat-master","jitter","real-time-control"],"name":"Xenomai","alt":"Xenomai","abbr":"","aliases":[],"one_liner":"An open-source framework that adds hard real-time capability to Linux, common in low-level robot control.","explanation":"Xenomai is an open-source real-time extension framework for Linux, originally started by Philippe Gerum. Its classic approach is a “dual kernel” design: a real-time kernel (Cobalt) runs alongside ordinary Linux, with real-time tasks scheduled by Cobalt at top priority while Linux runs only when the system is otherwise idle, pushing the timing jitter of a control loop (the variation in exactly when each cycle actually executes) down to the microsecond range. Robot joint control often needs a loop running at 1 kHz or higher that can never run late, which stock Linux cannot reliably guarantee, so Xenomai is commonly paired with an EtherCAT master on an industrial control computer. The other common route is patching Linux with PREEMPT_RT, which is simpler to set up but usually can't match the dual-kernel approach's worst-case latency.","example":"On an industrial PC, run the IgH EtherCAT master under Xenomai, sending position commands to each joint driver on a robot arm at a 1 kHz cycle.","related":["Real-Time Operating System (RTOS)","PREEMPT_RT","Hard Real-Time","EtherCAT Master (SOEM / IgH)","Jitter (Timing Jitter)","Real-Time Control"]},{"id":"ethercat-master","category":"software","sec":4,"tier":3,"sources":[{"title":"SOEM GitHub","url":"https://github.com/OpenEtherCATsociety/SOEM"},{"title":"IgH EtherCAT Master（EtherLab）","url":"https://gitlab.com/etherlab.org/ethercat"}],"as_of":"","related_ids":["ethercat","preempt-rt","cyclic-synchronous-position-velocity-torque-modes","servo-drive","hard-real-time","canopen-cia-402-drive-profile"],"name":"EtherCAT Master (SOEM / IgH)","alt":"EtherCAT 主站","abbr":"","aliases":["SOEM","IgH EtherCAT Master"],"one_liner":"The end of an EtherCAT bus that issues commands and gathers data, usually running on the main control computer.","explanation":"EtherCAT is a real-time industrial bus built on Ethernet, with one master and a number of slaves on the network: the slaves are devices such as joint drivers and sensors, and the master sends out a data frame every fixed cycle, reading and writing each slave's data in turn. Commercial masters include Beckhoff's TwinCAT; the two most common open-source options are SOEM (Simple Open EtherCAT Master), a user-space C library that's easy to get started with, and IgH EtherCAT Master (the EtherLab project), which runs as a Linux kernel module and offers better real-time performance. Humanoid and legged robots commonly run IgH or SOEM on Linux patched with PREEMPT_RT, sending commands to each joint at roughly 1 kHz.","example":"On the main controller, use SOEM to scan 12 joint drivers on the bus, switch them into CSP mode, and send a target position while reading back encoder values every 1 ms.","related":["EtherCAT (Ethernet for Control Automation Technology)","PREEMPT_RT","Cyclic Synchronous Position / Velocity / Torque Modes (CiA 402)","Servo Drive (Motor Driver)","Hard Real-Time","CANopen / CiA 402 Drive Profile"]},{"id":"socketcan","category":"software","sec":4,"tier":3,"sources":[{"title":"SocketCAN - Controller Area Network (Linux kernel docs)","url":"https://docs.kernel.org/networking/can.html"}],"as_of":"","related_ids":["controller-area-network","canopen-cia-402-drive-profile","servo-drive","host-computer","ubuntu-linux"],"name":"SocketCAN","alt":"SocketCAN","abbr":"","aliases":["can-utils"],"one_liner":"The Linux kernel's built-in CAN bus driver, which makes a CAN port behave like a network interface.","explanation":"SocketCAN is the CAN bus driver and protocol stack built into the Linux kernel, originally contributed by Volkswagen's research division. It abstracts a CAN interface into a network interface just like an Ethernet card (for example, can0), so programs can send and receive CAN frames using the ordinary socket API, without depending on the proprietary driver of any particular USB-CAN adapter. Many robot joint motors, grippers, and arms communicate over CAN bus (a serial bus known for strong noise resistance); a common setup is to plug a USB-CAN device into the host computer, bring up the interface at the right bit rate with the ip command, and then debug with candump and cansend from can-utils — a routine step when building desktop arms and motor drives.","example":"Run ip link set can0 up type can bitrate 1000000 to bring up the interface, then use candump can0 to view the feedback frames sent back by a motor.","related":["Controller Area Network (CAN)","CANopen / CiA 402 Drive Profile","Servo Drive (Motor Driver)","Host Computer","Ubuntu Linux"]},{"id":"precision-time-protocol","category":"software","sec":4,"tier":3,"sources":[{"title":"Precision Time Protocol (Wikipedia)","url":"https://en.wikipedia.org/wiki/Precision_Time_Protocol"},{"title":"The Linux PTP Project","url":"https://linuxptp.sourceforge.net/"}],"as_of":"","related_ids":["multi-sensor-time-synchronization-timestamp-alignment","multi-sensor-fusion","lidar","ethercat","preempt-rt"],"name":"Precision Time Protocol","alt":"精确时间协议","abbr":"PTP","aliases":["PTP","IEEE 1588","gPTP"],"one_liner":"A synchronization protocol that aligns multiple devices' clocks over Ethernet to sub-microsecond precision.","explanation":"The Precision Time Protocol (PTP) is defined by the IEEE 1588 standard, used to synchronize clocks across devices on a local network. The network elects one device as the master clock; the others exchange timestamped messages with it to measure link delay and clock offset, then correct themselves accordingly. When network cards support hardware timestamping, accuracy can reach the sub-microsecond range, far tighter than common NTP synchronization. gPTP, defined by IEEE 802.1AS, is a trimmed-down profile of PTP used mainly in automotive and other time-sensitive networks. On a robot, the camera, lidar, IMU, and main computer each have their own clock, and without synchronization, multi-sensor fusion and data logging end up misaligned in time — which is why PTP is often used to give everything a common time base. On Linux, linuxptp (ptp4l, phc2sys) is the common implementation.","example":"On a data-collection rig, the main computer acts as the PTP master clock, with the lidar and industrial cameras as slaves, so every point-cloud frame and image frame carries a directly comparable timestamp.","related":["Multi-sensor Time Synchronization / Timestamp Alignment","Multi-Sensor Fusion","LiDAR","EtherCAT (Ethernet for Control Automation Technology)","PREEMPT_RT"]},{"id":"matlab-simulink","category":"software","sec":5,"tier":2,"sources":[{"title":"MATLAB 产品页","url":"https://www.mathworks.com/products/matlab.html"},{"title":"Simulink 产品页","url":"https://www.mathworks.com/products/simulink.html"}],"as_of":"","related_ids":["proportional-integral-derivative-control","field-oriented-control","system-identification","embedded-software-development","model-predictive-control","python-and-c-plus-plus"],"name":"MATLAB / Simulink","alt":"MATLAB / Simulink","abbr":"","aliases":["MATLAB","Simulink"],"one_liner":"MathWorks' commercial numerical computing and graphical modeling-and-simulation software, widely used in controls.","explanation":"MATLAB is commercial software from MathWorks that serves as both a programming language and a computing environment, strong at matrix operations, signal processing, and plotting. Simulink is its companion graphical tool, where dynamic systems are built by dragging blocks and wiring them together and then simulated; models can also be auto-converted into embedded C code and flashed onto a controller. In robotics, the pair is mainly used for control algorithm design: modeling a motor's current loop and field-oriented control, tuning PID gains, system identification (inferring a system's model from data), and arm kinematics, aided by toolboxes like Robotics System Toolbox and ROS Toolbox. Many controls and motor courses, and engineering teams generally, rely on it; researchers working on learning-based methods tend to use Python instead.","example":"Build a cascaded current-loop, velocity-loop, and position-loop model for a joint motor in Simulink, tune the PID gains, then generate C code and flash it to the driver board.","related":["Proportional-Integral-Derivative Control","Field-Oriented Control","System Identification","Embedded Software Development","Model Predictive Control","Python and C++"]},{"id":"robotics-toolbox-for-python","category":"software","sec":5,"tier":3,"sources":[{"title":"Robotics Toolbox for Python on GitHub","url":"https://github.com/petercorke/robotics-toolbox-python"}],"as_of":"","related_ids":["forward-kinematics","inverse-kinematics","denavit-hartenberg-parameters","jacobian-matrix","pinocchio","franka-emika-panda-franka-research-3"],"name":"Robotics Toolbox for Python","alt":"Robotics Toolbox for Python（Peter Corke 机器人工具箱）","abbr":"","aliases":["roboticstoolbox-python","RTB"],"one_liner":"Peter Corke's Python robotics toolbox, a popular way to learn kinematics and dynamics hands-on.","explanation":"Robotics Toolbox for Python is an open-source Python library from Professor Peter Corke's group at Queensland University of Technology in Australia. It's the successor to the MATLAB Robotics Toolbox he maintained for many years, and is also the companion code for his textbook Robotics, Vision and Control. It includes built-in models for common arms such as the Panda and UR5, lets users build models from DH parameters or a URDF, and provides functions for forward and inverse kinematics, the Jacobian matrix, dynamics, and trajectory generation, plus a built-in visualizer that can animate an arm's motion. For a beginner, using it to work through a textbook's formulas is much faster than writing the code from scratch.","example":"Load a Franka model with rtb.models.Panda(), call fkine to compute the end-effector pose, then use ikine_LM to solve for the joint angles at a target pose.","related":["Forward Kinematics (FK)","Inverse Kinematics (IK)","Denavit-Hartenberg (DH) Parameters","Jacobian Matrix","Pinocchio","Franka Emika Panda / Franka Research 3"]},{"id":"eigen","category":"software","sec":5,"tier":3,"sources":[{"title":"Eigen 官网","url":"https://eigen.tuxfamily.org/"},{"title":"Eigen GitLab","url":"https://gitlab.com/libeigen/eigen"}],"as_of":"","related_ids":["pinocchio","rotation-matrix","quaternion","jacobian-matrix","python-and-c-plus-plus"],"name":"Eigen (C++ linear algebra library)","alt":"Eigen","abbr":"","aliases":[],"one_liner":"The most widely used open-source C++ linear-algebra library, behind most matrix, vector, and rotation math in robotics code.","explanation":"Eigen is an open-source, header-only C++ template library for linear algebra — no separate compilation or linking needed, just include it. It provides matrix and vector operations, matrix decompositions, linear-system solving, and geometry modules for quaternions, rotation matrices, and affine transforms. Robotics C++ code is constantly computing coordinate transforms, Jacobians, and dynamics equations, which has made Eigen a de facto standard: ROS, Pinocchio, MoveIt, and OCS2 all depend on it. It's roughly the C++ world's equivalent of NumPy, though details like memory alignment and quaternion component ordering are common early pitfalls.","example":"In a controller, represent the end-effector's orientation with Eigen::Quaterniond, and store a Jacobian in an Eigen::MatrixXd, taking its pseudoinverse to compute joint velocities.","related":["Pinocchio","Rotation Matrix","Quaternion","Jacobian Matrix","Python and C++"]},{"id":"pinocchio","category":"software","sec":5,"tier":2,"sources":[{"title":"stack-of-tasks/pinocchio - GitHub","url":"https://github.com/stack-of-tasks/pinocchio"}],"as_of":"","related_ids":["rigid-body-dynamics","recursive-newton-euler-algorithm","articulated-body-algorithm","crocoddyl","tsid","pink"],"name":"Pinocchio","alt":"Pinocchio","abbr":"","aliases":["pin"],"one_liner":"An open-source rigid-body dynamics library for fast robot kinematics, dynamics, and their derivatives.","explanation":"Pinocchio is an open-source C++ library, with a Python interface, developed mainly by the LAAS-CNRS and INRIA teams in France. It loads a robot model such as a URDF and efficiently computes forward and inverse kinematics, Jacobians, the recursive Newton-Euler algorithm (for inverse dynamics), the articulated-body algorithm (for forward dynamics), the mass matrix, and more — and it can also produce analytical derivatives of these quantities with respect to joint state. Those derivatives matter a lot for trajectory optimization and model predictive control, which is why libraries such as Crocoddyl, TSID, and Pink are all built on top of it. It's close to a standard tool for humanoid and quadruped whole-body control and MPC research; researchers working on learning-based methods also often use it for forward kinematics or IK-based retargeting.","example":"Load a G1 humanoid's URDF with Pinocchio and call forwardKinematics to get the wrist end-effector's pose in the world frame.","related":["Rigid-Body Dynamics","Recursive Newton-Euler Algorithm","Articulated Body Algorithm","Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)","TSID (Task Space Inverse Dynamics)","Pink"]},{"id":"rbdl","category":"software","sec":5,"tier":3,"sources":[{"title":"RBDL on GitHub","url":"https://github.com/rbdl/rbdl"}],"as_of":"","related_ids":["rigid-body-dynamics","recursive-newton-euler-algorithm","articulated-body-algorithm","pinocchio","orocos-kdl","unified-robot-description-format"],"name":"RBDL","alt":"RBDL","abbr":"RBDL","aliases":["Rigid Body Dynamics Library"],"one_liner":"Open-source C++ library that efficiently computes a robot's forward and inverse dynamics.","explanation":"RBDL is an open-source C++ library developed by Martin Felis while at Heidelberg University in Germany. It implements the classic rigid-body dynamics algorithms from Featherstone's textbook, including the Recursive Newton-Euler Algorithm (computes inverse dynamics — torques from motion), the Articulated Body Algorithm (computes forward dynamics — accelerations from torques), and the Composite Rigid Body Algorithm (computes the mass matrix). It can also load a model from a URDF or Lua file, and provides Python bindings. Deriving the dynamics equations for a multi-joint robot by hand is tedious, and libraries like this turn it into a few lines of function calls; RBDL is commonly used in legged-robot control and motion-optimization research, and its role overlaps closely with Pinocchio.","example":"Given a humanoid's URDF plus its current joint angles, velocities, and target accelerations, calling RBDL's InverseDynamics function computes the torque required at each joint.","related":["Rigid-Body Dynamics","Recursive Newton-Euler Algorithm","Articulated Body Algorithm","Pinocchio","Orocos KDL","Unified Robot Description Format"]},{"id":"orocos-kdl","category":"software","sec":5,"tier":3,"sources":[{"title":"Orocos KDL 官方页面","url":"https://www.orocos.org/kdl.html"},{"title":"orocos/orocos_kinematics_dynamics (GitHub)","url":"https://github.com/orocos/orocos_kinematics_dynamics"}],"as_of":"","related_ids":["forward-kinematics","inverse-kinematics","jacobian-matrix","trac-ik","moveit-motion-planning-framework","robot-state-publisher"],"name":"Orocos KDL","alt":"KDL","abbr":"KDL","aliases":["KDL","Orocos Kinematics and Dynamics Library"],"one_liner":"A C++ kinematics and dynamics library from the Orocos project, long a default choice for kinematics in ROS.","explanation":"KDL is part of Orocos, an open-source robot control software project originally developed at institutions including KU Leuven in Belgium. It's implemented in C++, with Python bindings also available. KDL models a robot as a kinematic chain of links connected by joints, and provides solvers for forward kinematics, inverse kinematics, the Jacobian matrix, and inverse dynamics. It's used extremely widely across the ROS ecosystem: kdl_parser converts a URDF into a KDL model, robot_state_publisher uses it to compute the pose of every link, and MoveIt's default kinematics plugin is also built on it. Its numerical inverse-kinematics solver tends to fail near joint limits, which is why it's often replaced by TRAC-IK; newer projects are also increasingly switching to Pinocchio.","example":"After a robot arm publishes its joint angles in ROS, robot_state_publisher calls KDL to compute the pose of the end effector and every link, then publishes them to the TF transform tree.","related":["Forward Kinematics (FK)","Inverse Kinematics (IK)","Jacobian Matrix","TRAC-IK","MoveIt Motion Planning Framework","robot_state_publisher"]},{"id":"trac-ik","category":"software","sec":5,"tier":3,"sources":[{"title":"TRAC-IK — TRACLabs","url":"https://traclabs.com/projects/trac-ik/"},{"title":"trac_ik — ROS Wiki","url":"https://wiki.ros.org/trac_ik"}],"as_of":"","related_ids":["inverse-kinematics","numerical-inverse-kinematics","orocos-kdl","ikfast","moveit-motion-planning-framework","joint-limits"],"name":"TRAC-IK","alt":"TRAC-IK","abbr":"","aliases":["trac_ik"],"one_liner":"An open-source numerical inverse-kinematics solver that's more reliable and faster than KDL.","explanation":"TRAC-IK is an open-source inverse kinematics library — solving for joint angles from a target end-effector pose — released by Patrick Beeson and Barrett Ames of TRACLabs at the 2015 IEEE Humanoids conference. The KDL solver commonly used in ROS iterates with Newton's method, which tends to get stuck and fail near joint limits. TRAC-IK instead runs two methods in parallel: a KDL-style Jacobian iteration with random restarts, and a version that formulates IK as nonlinear optimization solved with SQP (Sequential Quadratic Programming); whichever finishes first is used, giving a clear boost to both success rate and speed. It provides a C++ interface and a MoveIt plugin, and works with any serial robot arm.","example":"Switch the solver in MoveIt's kinematics.yaml from KDL to trac_ik_kinematics_plugin, and a 7-axis arm sees fewer IK failures when working near its joint limits.","related":["Inverse Kinematics (IK)","Numerical Inverse Kinematics","Orocos KDL","IKFast","MoveIt Motion Planning Framework","Joint Limits"]},{"id":"ikfast","category":"software","sec":5,"tier":3,"sources":[{"title":"OpenRAVE ikfast 文档","url":"http://openrave.org/docs/latest_stable/openravepy/ikfast/"},{"title":"MoveIt IKFast Kinematics Solver 教程","url":"https://moveit.picknik.ai/main/doc/examples/ikfast/ikfast_tutorial.html"}],"as_of":"","related_ids":[null,null,null,"trac-ik","moveit-motion-planning-framework",null],"name":"IKFast","alt":"IKFast","abbr":"","aliases":["OpenRAVE IKFast"],"one_liner":"OpenRAVE's analytical IK generator, auto-producing extremely fast C++ inverse-kinematics code for a specific arm.","explanation":"IKFast is an inverse-kinematics compiler Rosen Diankov developed within the OpenRAVE robot-planning framework. Inverse kinematics means working backward from a desired end-effector pose to the joint angles that achieve it. Numerical methods rely on iterative approximation, which is slow and can fail to converge; IKFast instead performs a symbolic derivation specific to one arm's kinematic structure and generates an analytical-solution C++ file offline, which at runtime computes every solution directly, usually in microseconds. It works best for 6-degree-of-freedom arms; a 7-DOF redundant arm needs one “free joint” designated and swept over a discrete set of values. MoveIt provides a workflow for wrapping IKFast-generated code as a kinematics plugin, commonly used to replace the default numerical solver.","example":"Generate an analytical inverse-kinematics solution for a UR5 arm with IKFast and package it as a MoveIt plugin — each IK solve drops from milliseconds to microseconds, and returns all 8 solutions at once.","related":["Inverse Kinematics (IK)","Analytical Inverse Kinematics","Numerical Inverse Kinematics","TRAC-IK","MoveIt Motion Planning Framework","Kinematic Redundancy"]},{"id":"pink","category":"software","sec":5,"tier":3,"sources":[{"title":"stephane-caron/pink (GitHub)","url":"https://github.com/stephane-caron/pink"}],"as_of":"","related_ids":["pinocchio","mink","differential-kinematics","inverse-kinematics","quadratic-programming","whole-body-inverse-kinematics"],"name":"Pink","alt":"Pink（基于 Pinocchio 的微分逆运动学库）","abbr":"","aliases":["pink (Python package)"],"one_liner":"A Python differential inverse-kinematics library built on Pinocchio, solving joint velocities from weighted tasks.","explanation":"Pink is an open-source Python library by Stéphane Caron that uses Pinocchio (a rigid-body dynamics library) under the hood to compute kinematics and Jacobians. It takes a differential inverse-kinematics approach: goals like “move the end effector to this pose,” “hold this orientation,” or “stay within joint limits” are written as weighted tasks and constraints, and on every control cycle Pink solves a quadratic program to get joint velocities, which are then integrated to update the joint angles. The advantage is that multiple objectives can be balanced at once, and redundant degrees of freedom and joint limits are handled naturally. It's commonly used for teleoperation, motion retargeting, and whole-body inverse kinematics on humanoids and robot arms; the equivalent library in the MuJoCo ecosystem is mink.","example":"When teleoperating a humanoid in VR, use Pink to track both wrist poses while constraining the torso orientation, solving for all arm joint angles in real time.","related":["Pinocchio","mink (MuJoCo inverse kinematics)","Differential Kinematics","Inverse Kinematics (IK)","Quadratic Programming","Whole-Body Inverse Kinematics"]},{"id":"mink","category":"software","sec":5,"tier":3,"sources":[{"title":"mink GitHub","url":"https://github.com/kevinzakka/mink"}],"as_of":"","related_ids":[null,null,null,null,"mujoco-menagerie",null],"name":"mink (MuJoCo inverse kinematics)","alt":"mink","abbr":"","aliases":[],"one_liner":"An open-source Python library for differential inverse kinematics, built on MuJoCo.","explanation":"mink is an open-source Python library by Kevin Zakka that performs inverse kinematics (working backward from a desired end-effector position to joint angles) using a MuJoCo model. It uses differential inverse kinematics: at each step, requirements such as “get the end effector close to the target,” “hold the desired orientation,” “stay within joint limits,” and “avoid collisions” are written as a quadratic program, solved for joint velocities, which are then integrated forward — a design that draws on the Pinocchio-based Pink library. Because it reads an MJCF model directly, it works out of the box with robots from MuJoCo Menagerie, and it's commonly used for teleoperation retargeting, generating demonstration trajectories, and whole-body IK for humanoids.","example":"Use mink to set an end-effector target pose for the Franka model in Menagerie and iteratively solve for a sequence of joint angles.","related":["Inverse Kinematics (IK)","Differential Kinematics","MuJoCo (Multi-Joint dynamics with Contact)","Pink (Python inverse kinematics based on Pinocchio)","MuJoCo Menagerie","Quadratic Programming"]},{"id":"pyroki","category":"software","sec":5,"tier":3,"sources":[{"title":"chungmin99/pyroki (GitHub)","url":"https://github.com/chungmin99/pyroki"},{"title":"PyRoki: A Modular Toolkit for Robot Kinematic Optimization (arXiv)","url":"https://arxiv.org/abs/2505.03728"}],"as_of":"2025-05","related_ids":["inverse-kinematics","trajectory-optimization","motion-retargeting","jax","viser","pink"],"name":"PyRoki","alt":"PyRoki","abbr":"PyRoki","aliases":["Python Robot Kinematics"],"one_liner":"UC Berkeley's open-source JAX toolkit for solving robot kinematics optimization problems.","explanation":"PyRoki is a Python toolkit open-sourced in 2025 by a team at UC Berkeley, built on JAX (Google's numerical computing library that supports automatic differentiation and compiling to GPU). It formulates problems like inverse kinematics, trajectory optimization, and motion retargeting uniformly as nonlinear optimization over “cost terms plus constraints”: users combine modules for collision avoidance, joint limits, end-effector pose, and so on as needed, and the solver runs on CPU or GPU with support for batched computation. The authors report that it's faster than some existing tools for inverse kinematics. It's often used alongside Viser, a visualization tool from the same group, and is well suited to research prototypes such as retargeting human hand or body motion onto a robot, or generating training data.","example":"Take wrist trajectories estimated from a video of a human demonstration, and use PyRoki to optimize them into joint trajectories for a bimanual robot while avoiding self-collision.","related":["Inverse Kinematics (IK)","Trajectory Optimization","Motion Retargeting","JAX","Viser","Pink"]},{"id":"open-motion-planning-library","category":"software","sec":5,"tier":2,"sources":[{"title":"The Open Motion Planning Library","url":"https://ompl.kavrakilab.org/"}],"as_of":"","related_ids":["moveit-motion-planning-framework","motion-planning","sampling-based-planning","rapidly-exploring-random-tree","probabilistic-roadmap","flexible-collision-library"],"name":"Open Motion Planning Library (OMPL)","alt":"OMPL","abbr":"OMPL","aliases":["OMPL"],"one_liner":"An open-source motion planning library from Rice University that bundles dozens of sampling-based planning algorithms.","explanation":"OMPL is an open-source C++ motion planning library, with Python bindings, developed by the Kavraki Lab at Rice University. It implements dozens of sampling-based planning algorithms, including RRT, RRT-Connect, RRT*, PRM, and BIT*. Its only job is to find a collision-free path through configuration space — it doesn't handle collision checking or robot modeling itself, relying instead on an externally supplied function that answers whether a given state is valid. That decoupling is exactly why MoveIt uses it as its default planning backend, so a lot of collision-avoidance planning for robot arms ultimately runs on OMPL underneath. It's also a convenient way to compare different planning algorithms when learning about motion planning.","example":"When MoveIt plans a path for a UR5e to avoid a cup on a table, it calls OMPL's RRTConnect by default.","related":["MoveIt Motion Planning Framework","Motion Planning","Sampling-Based Planning","Rapidly-exploring Random Tree","Probabilistic Roadmap","Flexible Collision Library (FCL)"]},{"id":"flexible-collision-library","category":"software","sec":5,"tier":3,"sources":[{"title":"flexible-collision-library/fcl (GitHub)","url":"https://github.com/flexible-collision-library/fcl"},{"title":"coal-library/coal (GitHub)","url":"https://github.com/coal-library/coal"}],"as_of":"","related_ids":["collision-checking","bounding-volume","gilbert-johnson-keerthi-algorithm","moveit-motion-planning-framework","pinocchio","open-motion-planning-library"],"name":"Flexible Collision Library (FCL)","alt":"FCL","abbr":"FCL","aliases":["hpp-fcl","Coal"],"one_liner":"The most widely used open-source C++ collision-checking and distance-computation library in robot motion planning.","explanation":"FCL is an open-source C++ library presented by Jia Pan, Sachin Chitta, and Dinesh Manocha at ICRA 2012, used to test whether two geometric shapes collide and to compute the closest distance and penetration depth between them, supporting spheres, boxes, cylinders, convex shapes, triangle meshes, and octree maps. Motion planning has to run collision checks against thousands of candidate poses, so speed and robustness directly determine how fast planning runs; FCL speeds this up with a bounding-volume hierarchy (a coarse rejection pass followed by finer computation). MoveIt uses it as its default collision checker. A fork maintained by the LAAS/INRIA team in France, hpp-fcl, made performance improvements and was later renamed Coal, and now serves as Pinocchio's collision backend.","example":"When MoveIt plans a path for an arm to go around a water cup on a table, every candidate set of joint angles is checked by FCL for intersection between the arm mesh and the cup or the table.","related":["Collision Checking","Bounding Volume (AABB / OBB)","Gilbert-Johnson-Keerthi Algorithm","MoveIt Motion Planning Framework","Pinocchio","Open Motion Planning Library (OMPL)"]},{"id":"moveit-motion-planning-framework","category":"software","sec":5,"tier":2,"sources":[{"title":"MoveIt 官网","url":"https://moveit.ai/"},{"title":"MoveIt 2 文档","url":"https://moveit.picknik.ai/main/index.html"}],"as_of":"","related_ids":["open-motion-planning-library","motion-planning","inverse-kinematics","semantic-robot-description-format","rviz-rviz2","curobo"],"name":"MoveIt Motion Planning Framework","alt":"MoveIt","abbr":"","aliases":["MoveIt 2","MoveIt2"],"one_liner":"The most widely used open-source motion planning framework for robot arms, built on ROS.","explanation":"MoveIt is an open-source, ROS-based motion planning framework for robot arms. It originated at Willow Garage and is now led by PickNik Robotics together with the community; the ROS 2 version is called MoveIt 2. It packages together a set of commonly needed arm capabilities: inverse kinematics (computing joint angles from a desired end-effector pose), calling planners such as OMPL to find a collision-free path, using a collision-checking library to detect self-collision and collisions with the environment, and adding velocity and timing information to a path before handing it to a controller for execution. Its Setup Assistant can generate an SRDF and configuration package from a URDF, and an RViz plugin lets users drag the end effector and plan interactively. Traditional teaching and pick-and-place pipelines rely on it heavily; learned policies that output actions directly often bypass it.","example":"Import a Franka arm's URDF with the MoveIt Setup Assistant, generate a configuration package, then drag the end-effector marker in RViz and click Plan & Execute to move the arm around obstacles to a target pose.","related":["Open Motion Planning Library (OMPL)","Motion Planning","Inverse Kinematics (IK)","Semantic Robot Description Format (SRDF)","RViz / RViz2","cuRobo (NVIDIA GPU-accelerated motion planning)"]},{"id":"topp-ra","category":"software","sec":5,"tier":3,"sources":[{"title":"A New Approach to Time-Optimal Path Parameterization based on Reachability Analysis (arXiv)","url":"https://arxiv.org/abs/1707.07239"},{"title":"toppra GitHub","url":"https://github.com/hungpham2511/toppra"}],"as_of":"","related_ids":["time-optimal-path-parameterization","time-parameterization","trajectory-planning","ruckig","drake","moveit-motion-planning-framework"],"name":"TOPP-RA","alt":"TOPP-RA","abbr":"TOPP-RA","aliases":["Time-Optimal Path Parameterization based on Reachability Analysis","toppra"],"one_liner":"Open-source library that takes a path and computes the fastest way to traverse it within limits.","explanation":"TOPP-RA is a time-optimal path parameterization method proposed by Hung Pham and Quang-Cuong Pham, published in IEEE T-RO (2018) and open-sourced as the Python/C++ library toppra. The problem it solves: a motion planner typically outputs only a sequence of geometric waypoints, with no information about how fast to move at each moment. TOPP-RA discretizes the path into segments and uses reachability analysis (computing, segment by segment, the range of speeds reachable at the next step) to break the problem into a series of small linear programs, producing the shortest-time speed profile subject to constraints on joint velocity, acceleration, and torque. It's more numerically stable and less failure-prone than earlier TOPP methods based on numerical integration. Libraries such as Drake include it built in, and it's commonly chained after path planning to produce an executable trajectory.","example":"After RRT plans a collision-free path for a robot arm, use toppra to redistribute timing according to each joint's velocity and acceleration limits, producing a trajectory ready to send straight to the controller.","related":["Time-Optimal Path Parameterization","Time Parameterization","Trajectory Planning","Ruckig","Drake","MoveIt Motion Planning Framework"]},{"id":"ruckig","category":"software","sec":5,"tier":3,"sources":[{"title":"pantor/ruckig (GitHub)","url":"https://github.com/pantor/ruckig"},{"title":"Jerk-limited Real-time Trajectory Generation with Arbitrary Target States (arXiv)","url":"https://arxiv.org/abs/2105.04830"}],"as_of":"","related_ids":["jerk","s-curve-velocity-profile","time-parameterization","time-optimal-path-parameterization","topp-ra","moveit-motion-planning-framework"],"name":"Ruckig","alt":"Ruckig","abbr":"","aliases":["Ruckig Pro"],"one_liner":"Open-source library that generates jerk-limited, time-optimal trajectories in real time, every control cycle.","explanation":"Ruckig is an open-source C++ library (with a Python interface) by Lars Berscheid, with its paper published at RSS 2021. Given the current state (position, velocity, acceleration), a target state, and each joint's limits on velocity, acceleration, and jerk (the rate of change of acceleration), it computes a time-optimal trajectory in microseconds — fast enough to re-plan on every single control cycle. That means when the target suddenly changes, the robot can smoothly transition from whatever motion state it's currently in, which is the “online” part of its name. MoveIt 2 uses it for trajectory smoothing and limiting in real-time servoing. The core is open source, with a commercial Ruckig Pro edition adding features such as intermediate waypoints.","example":"During visual servoing, the target position changes every frame; the controller calls Ruckig every 1 ms to re-plan from the current velocity and acceleration to the new target, so the joints don't jerk when the target jumps.","related":["Jerk","S-Curve Velocity Profile","Time Parameterization","Time-Optimal Path Parameterization","TOPP-RA","MoveIt Motion Planning Framework"]},{"id":"curobo","category":"software","sec":5,"tier":3,"sources":[{"title":"cuRobo Documentation","url":"https://curobo.org/"},{"title":"NVlabs/curobo - GitHub","url":"https://github.com/NVlabs/curobo"}],"as_of":"","related_ids":["motion-planning","inverse-kinematics","trajectory-optimization","nvidia-isaac-ros-cumotion","signed-distance-field-function","nvidia-isaac-ros"],"name":"cuRobo (NVIDIA GPU-accelerated motion planning)","alt":"cuRobo","abbr":"","aliases":["CuRobo"],"one_liner":"NVIDIA's GPU-parallel motion-planning library that produces collision-free arm trajectories in milliseconds.","explanation":"cuRobo is NVIDIA's open-source motion-generation library for robot arms, writing inverse kinematics, collision checking (against meshes, point clouds, or signed distance fields), and trajectory optimization all as parallel GPU kernels, optimizing thousands of candidate trajectories at once — which lets it produce smooth, collision-free trajectories that respect joint limits on the order of milliseconds. It targets the problem that traditional sampling-based planners (the RRT family) are too slow and produce jittery results for arm grasping, and it's also often used as the low-level executor underneath a large-model system: the upper layers just supply a target pose, and cuRobo figures out how to get there. cuMotion in Isaac ROS is a wrapper around it.","example":"Once an upstream system provides a target end-effector pose, cuRobo solves for a joint trajectory that avoids the table and any boxes in tens of milliseconds.","related":["Motion Planning","Inverse Kinematics (IK)","Trajectory Optimization","NVIDIA Isaac ROS cuMotion","Signed Distance Field / Function","NVIDIA Isaac ROS"]},{"id":"nvidia-isaac-ros-cumotion","category":"software","sec":5,"tier":3,"sources":[{"title":"Isaac ROS cuMotion 文档","url":"https://nvidia-isaac-ros.github.io/repositories_and_packages/isaac_ros_cumotion/index.html"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"NVIDIA Isaac ROS cuMotion","alt":"cuMotion","abbr":"","aliases":["isaac_ros_cumotion"],"one_liner":"NVIDIA's GPU-based ROS 2 motion-planning package for robot arms, built on cuRobo.","explanation":"cuMotion is the robot-arm motion-planning package within NVIDIA's Isaac ROS, with its core algorithm coming from cuRobo (NVIDIA's GPU-parallel motion-planning library). It plugs into MoveIt 2 as a plugin, replacing or supplementing traditional sampling-based planners, optimizing large numbers of candidate trajectories in parallel on the GPU to produce smooth, obstacle-avoiding joint trajectories in a short amount of time. It can also avoid obstacles using the environment map built by nvblox, and it provides a way to segment the robot's own arm out of a depth image so the robot doesn't mistake itself for an obstacle.","example":"Select the cuMotion planner in MoveIt 2's RViz interface to have the arm plan a path around obstacles on the table to reach a grasp.","related":["cuRobo (NVIDIA GPU-accelerated motion planning)","MoveIt Motion Planning Framework","NVIDIA Isaac ROS","Motion Planning","Obstacle Avoidance","NVIDIA nvblox (GPU TSDF/ESDF Mapping)"]},{"id":"pddlstream","category":"software","sec":5,"tier":3,"sources":[{"title":"caelan/pddlstream (GitHub)","url":"https://github.com/caelan/pddlstream"},{"title":"PDDLStream: Integrating Symbolic Planners and Blackbox Samplers via Optimistic Adaptive Planning (arXiv)","url":"https://arxiv.org/abs/1802.08705"}],"as_of":"","related_ids":["task-and-motion-planning","planning-domain-definition-language","symbolic-planning","motion-planning","inverse-kinematics"],"name":"PDDLStream","alt":"PDDLStream","abbr":"","aliases":[],"one_liner":"A framework that combines a symbolic planner with continuous samplers to do task and motion planning.","explanation":"PDDLStream is a planning framework proposed by Caelan Garrett, Tomás Lozano-Pérez, and Leslie Kaelbling at MIT, with the paper published at ICAPS 2020 and the code open-sourced. It extends PDDL (Planning Domain Definition Language, a standard way of describing planning problems) with “streams”: conditional samplers that generate continuous values on demand, such as grasp poses, placement locations, inverse-kinematics solutions, or collision-free paths. The planner first assumes these values exist and optimistically searches out a symbolic plan, then calls the samplers to verify and fill in the actual values, backtracking and re-searching if that fails. It addresses a gap between pure symbolic planners, which can't handle continuous geometry, and pure motion planners, which don't understand task ordering — making it a common baseline in task and motion planning (TAMP).","example":"Have a robot arm retrieve a cup blocked by a box: PDDLStream plans “move the box aside, then grasp the cup,” and uses its streams to sample a grasp pose and collision-free trajectory for each step.","related":["Task and Motion Planning","Planning Domain Definition Language","Symbolic Planning","Motion Planning","Inverse Kinematics (IK)"]},{"id":"drake","category":"software","sec":5,"tier":2,"sources":[{"title":"Drake 官网","url":"https://drake.mit.edu/"}],"as_of":"","related_ids":["mujoco","pinocchio","trajectory-optimization","multibody-dynamics","toyota-research-institute","inverse-kinematics"],"name":"Drake","alt":"Drake","abbr":"","aliases":["pydrake"],"one_liner":"A model-based robotics simulation and optimization toolbox led by MIT and the Toyota Research Institute.","explanation":"Drake is an open-source C++ toolbox started by Russ Tedrake’s group at MIT, with deep involvement from the Toyota Research Institute (TRI), and a Python interface called pydrake. It positions itself around “model-based design and verification”: it includes multibody dynamics computation, physics simulation with contact, and a set of mathematical optimization interfaces that let you write trajectory optimization, inverse kinematics, or controller-design problems directly and hand them to a range of solvers. Unlike simulators built mainly for large-scale parallel training, Drake puts more emphasis on physical and numerical rigor. The companion code for Tedrake’s public courses Underactuated Robotics and Robotic Manipulation is built on Drake.","example":"Loading an arm’s URDF with pydrake and writing an inverse-kinematics optimization problem to solve for the joint angles that bring the end effector to a target pose.","related":["MuJoCo (Multi-Joint dynamics with Contact)","Pinocchio","Trajectory Optimization","Multibody Dynamics","Toyota Research Institute","Inverse Kinematics (IK)"]},{"id":"casadi","category":"software","sec":5,"tier":3,"sources":[{"title":"CasADi 官网","url":"https://web.casadi.org/"},{"title":"casadi/casadi (GitHub)","url":"https://github.com/casadi/casadi"}],"as_of":"","related_ids":["acados","interior-point-optimizer","trajectory-optimization","model-predictive-control","direct-collocation","multiple-shooting"],"name":"CasADi","alt":"CasADi","abbr":"","aliases":[],"one_liner":"An open-source symbolic toolkit for numerical optimization and automatic differentiation, widely used for MPC and trajectory-optimization modeling.","explanation":"CasADi is an open-source tool for nonlinear optimization and algorithmic differentiation (automatic differentiation), developed by Joel Andersson, Joris Gillis, Moritz Diehl, and others at KU Leuven in Belgium, with support for Python, MATLAB, and C++. Users first write out system dynamics, cost functions, and constraints as symbolic expressions; CasADi then automatically computes gradients, Jacobians, and Hessians, and calls a solver such as IPOPT, qpOASES, or OSQP to solve the resulting problem — it can also generate C code for the whole problem. It solves the pain of deriving derivatives by hand when formulating optimal-control problems, which is slow and error-prone, and has become a standard modeling layer for building model predictive control, trajectory optimization, and parameter identification; acados also uses it to describe models.","example":"Use CasADi's Opti interface to write a cart-pole swing-up problem — with state and control as decision variables discretized via direct multiple shooting — and call IPOPT to solve for the optimal control sequence.","related":["acados (fast embedded optimal control solver)","Interior Point OPTimizer (Ipopt)","Trajectory Optimization","Model Predictive Control","Direct Collocation","Multiple Shooting"]},{"id":"interior-point-optimizer","category":"software","sec":5,"tier":3,"sources":[{"title":"coin-or/Ipopt (GitHub)","url":"https://github.com/coin-or/Ipopt"},{"title":"Ipopt 官方文档","url":"https://coin-or.github.io/Ipopt/"}],"as_of":"","related_ids":[null,null,"casadi",null,null,"acados"],"name":"Interior Point OPTimizer (Ipopt)","alt":"Ipopt","abbr":"Ipopt","aliases":["IPOPT"],"one_liner":"COIN-OR's open-source large-scale nonlinear solver, based on the interior-point method, commonly used for trajectory optimization.","explanation":"Ipopt is a nonlinear programming solver maintained by the COIN-OR open-source community, with its core algorithm proposed by Andreas Wächter and Lorenz Biegler, using the interior-point method (stepping toward the optimum from inside the feasible region, guided by a “barrier function”) to solve large-scale problems with both equality and inequality constraints. Trajectory optimization and nonlinear model predictive control in robotics repeatedly need to solve exactly this type of problem: the variables are the states and controls over a time horizon, and the constraints are the dynamics equations, joint limits, friction cones, and so on. Ipopt usually isn't called directly on its own but rather through a modeling tool such as CasADi, Drake, or Pyomo, and it needs to be paired with a linear-system solver such as MUMPS or HSL.","example":"Write out a direct-collocation formulation of a quadruped's jump in CasADi, then call nlpsol('solver', 'ipopt', nlp) to solve for a takeoff trajectory that satisfies the dynamics and friction constraints.","related":["Trajectory Optimization","Nonlinear Model Predictive Control","CasADi","Direct Collocation","Sequential Quadratic Programming","acados (fast embedded optimal control solver)"]},{"id":"osqp","category":"software","sec":5,"tier":3,"sources":[{"title":"OSQP 官网","url":"https://osqp.org/"},{"title":"OSQP: an operator splitting solver for quadratic programs (arXiv)","url":"https://arxiv.org/abs/1711.08013"}],"as_of":"","related_ids":["quadratic-programming","model-predictive-control","convex-mpc","whole-body-control","qpoases","convex-optimization"],"name":"OSQP","alt":"OSQP","abbr":"OSQP","aliases":["Operator Splitting Quadratic Program solver"],"one_liner":"Open-source quadratic-program solver based on operator splitting, widely used in MPC and whole-body control.","explanation":"OSQP is an open-source solver for quadratic programs (QPs) — optimization problems with a quadratic objective and linear inequality constraints. It was developed by researchers at Oxford and Stanford (Stellato, Banjac, Goulart, Bemporad, Boyd, and others), with the paper published in 2020. In robotics, model predictive control, whole-body control, and contact-force allocation are all commonly formulated as QPs that must be solved within a single control cycle, often just a few milliseconds. OSQP uses ADMM (Alternating Direction Method of Multipliers), an algorithm that splits a large problem into simpler subproblems solved in alternation. It's implemented in pure C with no external dependencies, supports warm-starting (using the previous time step's solution as the initial guess), and can generate embedded C code, which makes it well suited to real-time control. It also has Python, MATLAB, and C++ interfaces.","example":"A quadruped robot's convex MPC controller formulates the foot-force optimization for the next ten steps as a QP on every control cycle, and solves it with warm-started OSQP in a few milliseconds.","related":["Quadratic Programming","Model Predictive Control","Convex MPC","Whole-Body Control","qpOASES","Convex Optimization"]},{"id":"qpoases","category":"software","sec":5,"tier":3,"sources":[{"title":"coin-or/qpOASES (GitHub)","url":"https://github.com/coin-or/qpOASES"}],"as_of":"","related_ids":["quadratic-programming","model-predictive-control","convex-mpc","osqp","acados","casadi"],"name":"qpOASES","alt":"qpOASES","abbr":"","aliases":[],"one_liner":"Open-source quadratic-program solver based on an online active-set method, commonly used in MPC.","explanation":"qpOASES is an open-source C++ quadratic programming (QP) solver developed by Hans Joachim Ferreau and colleagues at KU Leuven in Belgium, now hosted on COIN-OR. It uses a parametric online active-set method: since two consecutive problems to be solved are usually only slightly different, it can reuse information about which constraints were active last time as a “warm start,” which makes it especially well suited to model predictive control (MPC), where a QP must be solved on every control cycle. It works best on small-to-medium, fairly dense problems, provides MATLAB and Python interfaces, and is integrated into optimization frameworks such as CasADi and acados. It shows up frequently in convex-MPC implementations for legged robots.","example":"A quadruped robot's convex MPC needs to solve for foot contact forces over the next several steps on every cycle, which can be handed to qpOASES for a warm-started solve.","related":["Quadratic Programming","Model Predictive Control","Convex MPC","OSQP","acados (fast embedded optimal control solver)","CasADi"]},{"id":"acados","category":"software","sec":5,"tier":3,"sources":[{"title":"acados documentation","url":"https://docs.acados.org/"},{"title":"acados/acados (GitHub)","url":"https://github.com/acados/acados"}],"as_of":"","related_ids":["model-predictive-control","nonlinear-model-predictive-control","casadi","sequential-quadratic-programming","ocs2","crocoddyl"],"name":"acados (fast embedded optimal control solver)","alt":"acados","abbr":"","aliases":[],"one_liner":"An open-source solver for real-time optimal control and MPC that can generate embeddable C code.","explanation":"acados is an open-source suite of software for optimal control and model predictive control (MPC — solving an optimization problem over a short future horizon at every control step), developed mainly by Moritz Diehl's group at the University of Freiburg and collaborators, and can be seen as a successor to the earlier ACADO toolkit. Its core is written in C and relies on BLASFEO (a small-matrix linear-algebra library) and HPIPM (a structured quadratic-programming solver) underneath, offering algorithms such as SQP (sequential quadratic programming) and real-time iteration (RTI), aimed at solving nonlinear MPC within a millisecond-scale control period. Users typically write system dynamics and cost functions in Python or MATLAB using CasADi, have acados generate C code from them, and then deploy that code to an industrial PC or embedded board. It's common in MPC research for legged robots, drones, and robot arms.","example":"Write a quadrotor's dynamics in CasADi in Python, generate a solver with acados, and run nonlinear MPC for trajectory tracking on an onboard computer at roughly 100 Hz.","related":["Model Predictive Control","Nonlinear Model Predictive Control","CasADi","Sequential Quadratic Programming","OCS2","Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)"]},{"id":"tsid","category":"software","sec":5,"tier":3,"sources":[{"title":"stack-of-tasks/tsid (GitHub)","url":"https://github.com/stack-of-tasks/tsid"}],"as_of":"","related_ids":["whole-body-control","inverse-dynamics","quadratic-programming","pinocchio","task-space-control","crocoddyl"],"name":"TSID (Task Space Inverse Dynamics)","alt":"TSID","abbr":"TSID","aliases":["Task Space Inverse Dynamics library"],"one_liner":"A C++ library that solves whole-body control as a task-space inverse-dynamics quadratic program.","explanation":"TSID is an open-source C++ library (with a Python interface) developed by Andrea Del Prete and colleagues at LAAS-CNRS in France, built on the Pinocchio dynamics library. It formulates whole-body control as a quadratic program (QP): given several tasks — such as maintaining the center of mass, tracking a hand's position, or holding an orientation — plus constraints like keeping contact forces inside the friction cone and staying under joint torque limits, it solves, on every control cycle, for the joint accelerations, torques, and contact forces that satisfy the constraints while getting as close as possible to each task's goal. It's commonly used for model-based whole-body control on humanoid and quadruped robots, and is also popular as teaching code for learning how whole-body control (WBC) works.","example":"Give a humanoid robot three tasks — keep both feet in contact, track a reference center-of-mass trajectory, and move the right hand to a target point — and use TSID to solve for joint torques every 1 millisecond.","related":["Whole-Body Control","Inverse Dynamics","Quadratic Programming","Pinocchio","Task-Space Control","Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)"]},{"id":"crocoddyl","category":"software","sec":5,"tier":3,"sources":[{"title":"loco-3d/crocoddyl - GitHub","url":"https://github.com/loco-3d/crocoddyl"}],"as_of":"","related_ids":["differential-dynamic-programming","optimal-control","trajectory-optimization","model-predictive-control","pinocchio","ocs2"],"name":"Crocoddyl (Contact RObot COntrol by Differential DYnamic programming Library)","alt":"Crocoddyl","abbr":"","aliases":[],"one_liner":"A contact-aware optimal control library that solves robot trajectories using differential dynamic programming.","explanation":"The name is an acronym for Contact RObot COntrol by Differential DYnamic programming Library, open-sourced by researchers at the University of Edinburgh and LAAS-CNRS (Mastalli and others). It formulates optimal-control problems for multi-body dynamics with contact in a unified way, then iterates on solvers such as DDP and FDDP to produce joint torques and a full trajectory — fast enough to run as online model predictive control. Legged robots can't avoid this class of tool: jumping or climbing a step requires simultaneously deciding foothold locations, contact forces, and whole-body motion, which inverse kinematics alone can't produce. It uses Pinocchio for the underlying dynamics, provides C++ and Python interfaces, and is often compared with OCS2 and acados.","example":"To plan a jump for a quadruped robot, define the contact sequence and cost function in Crocoddyl and solve for the torque trajectory with FDDP.","related":["Differential Dynamic Programming","Optimal Control","Trajectory Optimization","Model Predictive Control","Pinocchio","OCS2"]},{"id":"ocs2","category":"software","sec":5,"tier":3,"sources":[{"title":"leggedrobotics/ocs2 (GitHub)","url":"https://github.com/leggedrobotics/ocs2"}],"as_of":"","related_ids":["nonlinear-model-predictive-control","legged-control","eth-zurich-robotic-systems-lab","iterative-linear-quadratic-regulator","pinocchio","multiple-shooting"],"name":"OCS2","alt":"OCS2","abbr":"OCS2","aliases":["Optimal Control for Switched Systems"],"one_liner":"Open-source C++ optimal-control toolbox from ETH Zurich, commonly used for MPC on legged robots.","explanation":"OCS2 is an open-source C++ toolbox developed by the Robotic Systems Lab (RSL) at ETH Zurich for solving optimal control problems in “switched systems” — systems whose dynamics change as they switch between modes. A legged robot is one example: each leg alternates between stance and swing, and the dynamics change with the contact state. OCS2 provides algorithms including SLQ, iLQR, and multiple-shooting SQP, generates derivatives via automatic differentiation, and uses Pinocchio for rigid-body dynamics, which lets it run nonlinear model predictive control (MPC) in real time. It ships with examples for quadrupeds and mobile manipulators, plus a ROS interface, and underlies many legged-robot MPC projects, including legged_control.","example":"legged_control uses OCS2 to run nonlinear MPC for a quadruped robot, computing the center-of-mass trajectory and foot forces for the near future, which are then handed to a whole-body controller for execution.","related":["Nonlinear Model Predictive Control","legged_control (NMPC + WBC framework for legged robots)","ETH Zurich Robotic Systems Lab","Iterative Linear Quadratic Regulator","Pinocchio","Multiple Shooting"]},{"id":"legged-control","category":"software","sec":5,"tier":3,"sources":[{"title":"qiayuanl/legged_control (GitHub)","url":"https://github.com/qiayuanl/legged_control"}],"as_of":"","related_ids":[null,null,null,null,null,null],"name":"legged_control (NMPC + WBC framework for legged robots)","alt":"legged_control","abbr":"","aliases":["qiayuanl/legged_control"],"one_liner":"An open-source legged-robot control framework combining nonlinear MPC planning with whole-body control tracking, built on OCS2 and ROS.","explanation":"legged_control is an open-source, model-based control framework for quadruped robots by Qiayuan Liao, built on ROS and ros_control. The upper level uses the OCS2 library to run nonlinear model predictive control (NMPC — solving online for an optimal trajectory over a short future horizon at every control step); the lower level uses whole-body control (WBC) based on hierarchical quadratic programming to convert that plan into individual joint torques, together with a state estimator that fuses leg odometry and IMU data. It provides Gazebo simulation and real-robot examples on platforms such as the Unitree A1, making it a common reference implementation to get running and modify when learning model-based legged locomotion control, and it's also frequently used as a baseline to compare against reinforcement-learning-based locomotion control.","example":"Launch a Unitree A1 model in Gazebo and use legged_control's NMPC+WBC controller to have the robot switch into a diagonal trot and follow velocity commands from a game controller.","related":["Nonlinear Model Predictive Control","Whole-Body Control","OCS2 (Optimal Control for Switched Systems)","Hierarchical Quadratic Programming","Legged Locomotion","Model-Based Control"]},{"id":"mujoco-mpc","category":"software","sec":5,"tier":3,"sources":[{"title":"mujoco_mpc GitHub","url":"https://github.com/google-deepmind/mujoco_mpc"},{"title":"Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo (arXiv)","url":"https://arxiv.org/abs/2212.00541"}],"as_of":"","related_ids":[null,null,null,null,null],"name":"MuJoCo MPC (MJPC)","alt":"MuJoCo MPC","abbr":"MJPC","aliases":["MJPC"],"one_liner":"Google DeepMind's open-source tool for interactive, real-time model predictive control inside MuJoCo.","explanation":"MuJoCo MPC (MJPC) is an interactive tool Google DeepMind open-sourced in 2022 that uses the MuJoCo simulator to run model predictive control — at every instant, simulating a short window into the future, optimizing an action sequence, executing only the first step, and repeating. It has built-in planners including iLQG, gradient descent, and Predictive Sampling; a user adjusts cost-function weights in a graphical interface and sees the robot's behavior change live. It doesn't train a neural network at all, which makes it well suited to quickly generating motion for quadrupeds, humanoids, or dexterous hands, and it's also commonly used as a model-based control baseline to compare against reinforcement learning.","example":"Pick a quadruped task in the MJPC interface, drag a target point, and the robot walks over to it via real-time online planning.","related":["Model Predictive Control (MPC)","Sampling-based MPC","Iterative Linear Quadratic Regulator (iLQR)","MuJoCo (Multi-Joint dynamics with Contact)","DIAL-MPC (Diffusion-Inspired Annealing for Legged MPC)"]},{"id":"opencv","category":"software","sec":6,"tier":2,"sources":[{"title":"OpenCV - Open Computer Vision Library","url":"https://opencv.org/"}],"as_of":"","related_ids":["camera-calibration","aruco-marker","perspective-n-point","feature-points","point-cloud-library","open3d"],"name":"OpenCV (Open Source Computer Vision Library)","alt":"OpenCV","abbr":"OpenCV","aliases":["cv2"],"one_liner":"The most widely used open-source computer vision library, for image, camera, and geometry processing.","explanation":"OpenCV is an open-source computer vision library originally started by Intel, with its core written in C++ and bindings for Python and other languages; in Python it's used via import cv2. It covers image reading and resizing, color-space conversion, filtering, feature-point extraction and matching, camera calibration, distortion correction, PnP pose estimation, and ArUco marker detection, among other classic vision tasks. Deep learning has taken over most recognition tasks, but OpenCV still gets used constantly in robotics: reading camera frames, preprocessing, calibrating camera intrinsics and extrinsics, and visualizing results — it's hard to avoid.","example":"Use cv2.calibrateCamera with a checkerboard pattern to calibrate a wrist camera's intrinsics, then cv2.undistort to remove lens distortion.","related":["Camera Calibration","ArUco Marker","Perspective-n-Point","Feature Points","Point Cloud Library (PCL)","Open3D"]},{"id":"open3d","category":"software","sec":6,"tier":2,"sources":[{"title":"Open3D: A Modern Library for 3D Data Processing","url":"https://www.open3d.org/"},{"title":"Open3D paper (arXiv:1801.09847)","url":"https://arxiv.org/abs/1801.09847"}],"as_of":"","related_ids":["point-cloud","point-cloud-library","iterative-closest-point","truncated-signed-distance-function","opencv","3d-diffusion-policy"],"name":"Open3D","alt":"Open3D","abbr":"","aliases":[],"one_liner":"An open-source library for processing point clouds, meshes, and 3D reconstruction.","explanation":"Open3D is an open-source 3D data-processing library initiated by researchers at Intel Labs, with a C++ core and an easy-to-use Python interface. It supports reading and writing point clouds and meshes, voxel downsampling, normal estimation, ICP point-cloud registration, RGB-D fusion reconstruction (TSDF), and interactive visualization. Compared with PCL (Point Cloud Library), which is more full-featured but heavier, Open3D is quicker to pick up in Python, so it's commonly used in robotics research: turning depth-camera data into point clouds, cropping out a tabletop region, preparing input for a 3D policy, or sanity-checking a calibration result.","example":"Use Open3D to fuse a RealSense camera's color and depth images into a point cloud, voxel-downsample it, and feed it as input to a 3D diffusion policy.","related":["Point Cloud","Point Cloud Library (PCL)","Iterative Closest Point","Truncated Signed Distance Function","OpenCV (Open Source Computer Vision Library)","3D Diffusion Policy"]},{"id":"point-cloud-library","category":"software","sec":6,"tier":2,"sources":[{"title":"Point Cloud Library (PCL)","url":"https://pointclouds.org/"},{"title":"PointCloudLibrary/pcl - GitHub","url":"https://github.com/PointCloudLibrary/pcl"}],"as_of":"","related_ids":["point-cloud","open3d","iterative-closest-point","point-cloud-registration","point-cloud-segmentation","opencv"],"name":"Point Cloud Library (PCL)","alt":"点云库","abbr":"PCL","aliases":["PCL"],"one_liner":"An open-source C++ library for processing 3D point clouds — filtering, registration, and segmentation included.","explanation":"PCL is an open-source C++ library for point-cloud processing, originally started at Willow Garage and formally introduced in a 2011 ICRA paper, released under the BSD license. It packages the common steps of point-cloud processing into modules: voxel downsampling and outlier removal (filtering); computing features like surface normals and FPFH; ICP and NDT registration; RANSAC plane fitting and clustering-based segmentation; and surface reconstruction. Depth cameras and lidar both output point clouds, so PCL has long been a foundational tool for 3D perception in robotics, and ROS interconverts with it through the perception_pcl package. If most of the work is happening in Python, Open3D is a lighter-weight alternative.","example":"Use PCL's RANSAC to fit and remove the tabletop plane, then run Euclidean clustering on the remaining points to get a separate point cloud for each object on the table.","related":["Point Cloud","Open3D","Iterative Closest Point","Point Cloud Registration","Point Cloud Segmentation","OpenCV (Open Source Computer Vision Library)"]},{"id":"kalibr","category":"software","sec":6,"tier":3,"sources":[{"title":"ethz-asl/kalibr (GitHub)","url":"https://github.com/ethz-asl/kalibr"}],"as_of":"","related_ids":[null,null,"apriltag",null,null,"vins-mono-vins-fusion"],"name":"Kalibr","alt":"Kalibr","abbr":"","aliases":["ethz-asl/kalibr"],"one_liner":"ETH Zurich's open-source calibration toolbox specializing in multi-camera and camera-IMU intrinsic, extrinsic, and time calibration.","explanation":"Kalibr is an open-source sensor-calibration toolbox from ETH Zurich's Autonomous Systems Lab (ASL), running on ROS. It can calibrate intrinsics and extrinsics across multiple cameras, the spatial extrinsics and time offset between a camera and an IMU, multiple IMUs together, and rolling-shutter cameras, commonly using an AprilGrid (a grid made of AprilTag markers) as the calibration target. Algorithms like visual-inertial odometry and SLAM need to know the relative pose and timing offset between the camera and IMU precisely — get those parameters wrong and drift follows directly — which makes Kalibr one of the most commonly used tools for this job.","example":"Hold a RealSense D435i with an IMU in front of an AprilGrid and wave it around while recording a rosbag, then run kalibr_calibrate_imu_camera to get the camera-to-IMU transform and their time offset, which then get plugged into VINS-Fusion's configuration.","related":["Camera-IMU Calibration","Camera Calibration","AprilTag","Visual-Inertial Odometry","Camera Extrinsics","VINS-Mono / VINS-Fusion"]},{"id":"colmap","category":"software","sec":6,"tier":3,"sources":[{"title":"COLMAP Documentation","url":"https://colmap.github.io/"},{"title":"colmap/colmap - GitHub","url":"https://github.com/colmap/colmap"}],"as_of":"","related_ids":["structure-from-motion","multi-view-stereo","3d-gaussian-splatting","neural-radiance-fields","bundle-adjustment","camera-extrinsics"],"name":"COLMAP","alt":"COLMAP","abbr":"","aliases":[],"one_liner":"An open-source 3D reconstruction pipeline combining structure-from-motion and multi-view stereo.","explanation":"COLMAP is a general-purpose 3D reconstruction pipeline created by Johannes Schönberger and others. It first runs structure-from-motion (SfM — recovering camera poses and a sparse point cloud simultaneously from multiple photos), then multi-view stereo (MVS — densifying that into a full point cloud and mesh). It matters because it has become close to the de facto standard for “getting camera poses from photos”: training data for neural radiance fields and 3D Gaussian splatting is almost always pose-labeled using COLMAP first. It ships with both a command-line and a graphical interface, and is also commonly called from scripts; in robotics, it's often used to turn a video taken circling an object or scene into a usable 3D asset.","example":"Before training a 3D Gaussian splatting model, run a set of images taken circling a tabletop through COLMAP to get camera intrinsics, extrinsics, and a sparse point cloud.","related":["Structure from Motion","Multi-View Stereo","3D Gaussian Splatting","Neural Radiance Fields","Bundle Adjustment","Camera Extrinsics"]},{"id":"nerfstudio-gsplat","category":"software","sec":6,"tier":3,"sources":[{"title":"Nerfstudio 官网","url":"https://docs.nerf.studio/"},{"title":"gsplat GitHub","url":"https://github.com/nerfstudio-project/gsplat"}],"as_of":"","related_ids":[null,null,null,null,"colmap",null],"name":"Nerfstudio / gsplat","alt":"Nerfstudio / gsplat（NeRF 与高斯泼溅开源工具库）","abbr":"","aliases":[],"one_liner":"Open-source toolkits for reconstructing 3D scenes from photos, covering both NeRF and Gaussian splatting.","explanation":"Nerfstudio is an open-source neural radiance field (NeRF — representing a 3D scene with a neural network) framework from a UC Berkeley team, presented at SIGGRAPH 2023, providing data processing, training, and a web viewer, with methods like Nerfacto built in. gsplat is a CUDA-accelerated 3D Gaussian splatting (representing a scene as a large number of colored 3D Gaussian points) rendering and training library from the same team, used by projects such as Nerfstudio's own Splatfacto. In robotics, they're commonly used to scan a real scene into a renderable 3D asset for real-to-sim transfer, Gaussian-splatting-based simulation, and data augmentation.","example":"Film a video circling a tabletop with a phone, process and train it through Nerfstudio, and get a Gaussian-splatting scene that can be rendered from novel viewpoints.","related":["Neural Radiance Fields (NeRF)","3D Gaussian Splatting","Real-to-Sim","Gaussian-Splatting-based Simulation","COLMAP","Novel View Synthesis"]},{"id":"ceres-solver","category":"software","sec":6,"tier":3,"sources":[{"title":"Ceres Solver 官网","url":"http://ceres-solver.org/"},{"title":"ceres-solver/ceres-solver (GitHub)","url":"https://github.com/ceres-solver/ceres-solver"}],"as_of":"","related_ids":["bundle-adjustment","simultaneous-localization-and-mapping","g2o","gtsam","factor-graph-optimization","camera-calibration"],"name":"Ceres Solver","alt":"Ceres Solver","abbr":"","aliases":["Ceres"],"one_liner":"Google's open-source C++ library for nonlinear least-squares optimization, a common backend for SLAM and calibration.","explanation":"Ceres Solver is an open-source C++ optimization library from Google, mainly used to solve nonlinear least-squares problems (minimizing the sum of squares of a set of error terms), and it can also handle general unconstrained optimization. Users only need to specify how each error term (residual) is computed; Ceres can differentiate automatically, and it provides algorithms such as Levenberg-Marquardt and Dogleg along with sparse linear solvers that handle problems with thousands of variables. Many problems in robotics and 3D vision can be cast in this form — bundle adjustment, SLAM pose-graph optimization, camera-IMU calibration, and hand-eye calibration among them. Open-source systems such as VINS-Mono and Google's Cartographer use it as their optimization backend. It occupies roughly the same niche as g2o and GTSAM, and the three are often treated as interchangeable tools.","example":"During camera calibration, write the reprojection error of each detected corner as a residual block, and let Ceres jointly optimize the intrinsics and each image's extrinsics.","related":["Bundle Adjustment","Simultaneous Localization and Mapping","g2o (General Graph Optimization)","GTSAM (Georgia Tech Smoothing and Mapping)","Factor Graph Optimization","Camera Calibration"]},{"id":"g2o","category":"software","sec":6,"tier":3,"sources":[{"title":"RainerKuemmerle/g2o (GitHub)","url":"https://github.com/RainerKuemmerle/g2o"}],"as_of":"","related_ids":["factor-graph-optimization","bundle-adjustment",null,"gtsam","ceres-solver","orb-slam3"],"name":"g2o (General Graph Optimization)","alt":"g2o 图优化库","abbr":"g2o","aliases":[],"one_liner":"A C++ library for solving graph-structured optimization problems like SLAM pose graphs and bundle adjustment.","explanation":"g2o is an open-source C++ framework presented at ICRA 2011 by Rainer Kümmerle, Giorgio Grisetti, Kurt Konolige, Wolfram Burgard, and others, built specifically to solve nonlinear least-squares problems that can be drawn as a graph: nodes are the quantities being estimated (such as camera poses or map points), and edges are the observation constraints between them. SLAM's (simultaneous localization and mapping) back-end optimization and bundle adjustment both fall into this category; g2o exploits the sparse structure of these problems to solve them efficiently and lets users define custom node and edge types. The ORB-SLAM series uses g2o for its back end. GTSAM and Ceres Solver are tools in the same category.","example":"After ORB-SLAM3 detects a loop closure, it uses g2o to optimize the whole pose graph, spreading the accumulated drift back out across the entire trajectory.","related":["Factor Graph Optimization","Bundle Adjustment","Simultaneous Localization and Mapping (SLAM)","GTSAM (Georgia Tech Smoothing and Mapping)","Ceres Solver","ORB-SLAM3"]},{"id":"gtsam","category":"software","sec":6,"tier":3,"sources":[{"title":"GTSAM 官网","url":"https://gtsam.org/"},{"title":"borglab/gtsam (GitHub)","url":"https://github.com/borglab/gtsam"}],"as_of":"","related_ids":["factor-graph-optimization",null,"imu-preintegration","g2o","ceres-solver","lio-sam"],"name":"GTSAM (Georgia Tech Smoothing and Mapping)","alt":"GTSAM","abbr":"GTSAM","aliases":[],"one_liner":"A factor-graph optimization library from Georgia Tech, a common backend for SLAM and sensor fusion.","explanation":"GTSAM is an open-source C++ library (BSD-licensed, with Python and MATLAB bindings) developed by Frank Dellaert's team at Georgia Tech. It represents an estimation problem as a factor graph: variables are the unknowns being estimated, such as poses and velocities, and factors are the constraints coming from sensor measurements, with nonlinear optimization then finding the most likely solution. Its distinguishing feature is the iSAM2 incremental solver, which, when a new measurement arrives, updates only the parts of the solution it affects, making it well suited to real-time SLAM; it also has built-in support for common factors like IMU preintegration. Lidar-inertial SLAM systems such as LIO-SAM use it as their back end. It's in the same category as g2o and Ceres Solver, though GTSAM leans more toward probabilistic modeling.","example":"LIO-SAM feeds lidar odometry, IMU preintegration, GPS, and loop-closure constraints into GTSAM as factors, using iSAM2 to optimize the robot's trajectory in real time.","related":["Factor Graph Optimization","Simultaneous Localization and Mapping (SLAM)","IMU Preintegration","g2o (General Graph Optimization)","Ceres Solver","LIO-SAM"]},{"id":"common-ros-slam-packages","category":"software","sec":6,"tier":3,"sources":[{"title":"slam_toolbox - GitHub","url":"https://github.com/SteveMacenski/slam_toolbox"},{"title":"gmapping - ROS Wiki","url":"https://wiki.ros.org/gmapping"},{"title":"RTAB-Map","url":"http://introlab.github.io/rtabmap/"}],"as_of":"","related_ids":["simultaneous-localization-and-mapping","lidar-slam","occupancy-grid-map","ros-2-navigation-stack","rtab-map","package"],"name":"Common ROS SLAM Packages (GMapping / SLAM Toolbox / RTAB-Map)","alt":"ROS 常用 SLAM 建图包（GMapping / SLAM Toolbox / RTAB-Map）","abbr":"","aliases":[],"one_liner":"Three ready-to-run ROS packages for mapping and localization that just need a config file, not custom code.","explanation":"GMapping is the oldest 2D lidar SLAM package (particle-filter-based, originating from OpenSLAM), and was the default choice for teaching robots in the ROS 1 era. SLAM Toolbox, developed by Steve Macenski, is the main option for 2D lidar mapping in ROS 2 today; it's graph-optimization-based and supports continuing to map on top of an existing map, as well as large environments. RTAB-Map, developed by Mathieu Labbé, targets RGB-D and stereo cameras, with built-in appearance-based loop-closure detection, and can output a dense 3D map. Their value is in turning SLAM into something you can get running just by changing a configuration file, producing a navigable occupancy grid map that plugs directly into Nav2; the tradeoff is that accuracy and robustness fall short of an algorithm specially tuned for a given environment.","example":"Run SLAM Toolbox on a mobile base with a 2D lidar to build a grid map of a building floor, then hand it to Nav2 for navigation.","related":["Simultaneous Localization and Mapping","LiDAR SLAM","Occupancy Grid Map","ROS 2 Navigation Stack (Nav2)","RTAB-Map","Package (ROS)"]},{"id":"google-cartographer","category":"software","sec":6,"tier":3,"sources":[{"title":"cartographer-project/cartographer (GitHub)","url":"https://github.com/cartographer-project/cartographer"},{"title":"Cartographer 文档","url":"https://google-cartographer.readthedocs.io/"}],"as_of":"","related_ids":["lidar-slam",null,"occupancy-grid-map","loop-closure-detection","2d-lidar","common-ros-slam-packages"],"name":"Google Cartographer","alt":"Cartographer","abbr":"","aliases":["cartographer_ros"],"one_liner":"Google's open-source real-time lidar SLAM system, supporting both 2D and 3D mapping.","explanation":"Cartographer is a real-time SLAM (simultaneous localization and mapping) system Google open-sourced in 2016, described in the accompanying paper by Hess et al., “Real-Time Loop Closure in 2D LIDAR SLAM.” It uses lidar (optionally combined with an IMU and odometry) to build local submaps, localizes through scan matching, and uses branch-and-bound search for loop-closure detection to eliminate the drift that accumulates over long runs, supporting both 2D and 3D. It connects to ROS through cartographer_ros, and was once one of the most common lidar mapping choices for cleaning robots, service robots, and ROS teaching. According to community feedback, the project has seen fewer updates in recent years, and ROS 2 users have often moved to SLAM Toolbox instead.","example":"Drive a TurtleBot fitted with a 2D lidar around an office once, run Cartographer to produce an occupancy grid map, and hand that map to Nav2 for navigation.","related":["LiDAR SLAM","Simultaneous Localization and Mapping (SLAM)","Occupancy Grid Map","Loop Closure Detection","2D LiDAR","Common ROS SLAM Packages (GMapping / SLAM Toolbox / RTAB-Map)"]},{"id":"ros-2-navigation-stack","category":"software","sec":6,"tier":2,"sources":[{"title":"Nav2 Documentation","url":"https://docs.nav2.org/"},{"title":"ros-navigation/navigation2 - GitHub","url":"https://github.com/ros-navigation/navigation2"}],"as_of":"","related_ids":["navigation","costmap","adaptive-monte-carlo-localization","behavior-tree","global-planning-and-local-planning","robot-operating-system-2"],"name":"ROS 2 Navigation Stack (Nav2)","alt":"Nav2","abbr":"Nav2","aliases":["Navigation2","ROS navigation stack","move_base"],"one_liner":"ROS 2's official navigation framework, letting a mobile robot autonomously drive itself from point A to point B.","explanation":"Nav2 is ROS 2's navigation framework, succeeding the move_base-centered navigation stack from ROS 1, led primarily by Steve Macenski and others. Given a map and a goal point, it handles localization (commonly AMCL), maintains a costmap, does global path planning (e.g. NavFn or the Smac Planner), tracks a local trajectory (e.g. DWB, MPPI, or Regulated Pure Pursuit), and orchestrates the whole process with a behavior tree — automatically backing up, rotating, or performing other recovery behaviors when the robot gets stuck. Every stage is a swappable plugin. Indoor navigation for wheeled bases and quadrupeds most often starts here, typically paired with a SLAM package to build a map first, then Nav2 to run point-to-point navigation.","example":"Clicking a “2D Goal Pose” in RViz gives Nav2 a target; it plans a path and drives a TurtleBot around obstacles to reach it.","related":["Navigation","Costmap","Adaptive Monte Carlo Localization","Behavior Tree","Global Planning and Local Planning","Robot Operating System 2"]},{"id":"behaviortree-cpp","category":"software","sec":6,"tier":3,"sources":[{"title":"BehaviorTree.CPP 官方文档","url":"https://www.behaviortree.dev/"},{"title":"BehaviorTree/BehaviorTree.CPP (GitHub)","url":"https://github.com/BehaviorTree/BehaviorTree.CPP"}],"as_of":"","related_ids":["behavior-tree","finite-state-machine","ros-2-navigation-stack","groot2","task-planning","robot-operating-system-2"],"name":"BehaviorTree.CPP","alt":"BehaviorTree.CPP（C++ 行为树库）","abbr":"BT.CPP","aliases":["BT.CPP","Groot behavior-tree editor"],"one_liner":"The most widely used C++ behavior-tree library in robotics, describing task logic in XML.","explanation":"BehaviorTree.CPP is an open-source C++ behavior-tree library, primarily authored by Davide Faconti. A behavior tree is a tree structure for organizing a robot's task logic: leaf nodes execute actions or check conditions, while internal control nodes — sequence, selector, parallel, and so on — decide what to try next and what to fall back on after a failure. Compared with a finite state machine, a behavior tree is easier to break apart, reuse, and extend. BT.CPP lets developers write each action node in C++ and then assemble them into a tree using an XML file, loaded and run without recompiling — so the task flow can be adjusted without touching code — and its companion graphical editor, Groot / Groot2, lets you build the tree by dragging nodes and monitor its execution state live. ROS 2's navigation framework, Nav2, uses it to orchestrate navigation behavior.","example":"Nav2's default navigation behavior tree: compute a path, then follow it; if that fails partway through, it tries recovery actions in sequence — clearing the costmap, rotating in place, backing up.","related":["Behavior Tree","Finite State Machine","ROS 2 Navigation Stack (Nav2)","Groot2 (BehaviorTree.CPP IDE)","Task Planning","Robot Operating System 2"]},{"id":"groot2","category":"software","sec":6,"tier":3,"sources":[{"title":"Groot2 - BehaviorTree.CPP 官方文档","url":"https://www.behaviortree.dev/groot/"},{"title":"BehaviorTree.CPP GitHub","url":"https://github.com/BehaviorTree/BehaviorTree.CPP"}],"as_of":"","related_ids":["behavior-tree","behaviortree-cpp","ros-2-navigation-stack","finite-state-machine","task-planning"],"name":"Groot2 (BehaviorTree.CPP IDE)","alt":"Groot2（行为树可视化编辑器，非英伟达 GR00T）","abbr":"","aliases":["Groot"],"one_liner":"A graphical editor and debugger for the behavior trees used by BehaviorTree.CPP.","explanation":"Groot2 is a desktop tool built by the same team behind BehaviorTree.CPP; its predecessor, Groot, was open source. A behavior tree writes a robot's task logic as a tree of sequence, selector, condition, and action nodes, and BehaviorTree.CPP describes that tree in an XML file. Writing that XML by hand is error-prone and makes it hard to see where execution currently is, so Groot2 lets you build the tree by dragging nodes and exporting XML, and it can also connect to a running program to watch each node's status — success, failure, or running — live, and replay logs. The name looks similar to NVIDIA's humanoid model GR00T, but the two have nothing to do with each other. It's also commonly used to inspect and debug the behavior trees in ROS 2's Nav2 navigation stack.","example":"In Nav2, open the navigation behavior tree's XML in Groot2 to see which recovery node keeps failing when the robot gets stuck.","related":["Behavior Tree","BehaviorTree.CPP","ROS 2 Navigation Stack (Nav2)","Finite State Machine","Task Planning"]},{"id":"autoware","category":"software","sec":6,"tier":3,"sources":[{"title":"Autoware Foundation","url":"https://autoware.org/"},{"title":"autowarefoundation/autoware (GitHub)","url":"https://github.com/autowarefoundation/autoware"}],"as_of":"","related_ids":["robot-operating-system-2","autonomous-driving","ros-2-navigation-stack","simultaneous-localization-and-mapping","carla","model-predictive-control"],"name":"Autoware (open-source autonomous driving stack on ROS 2)","alt":"Autoware（开源自动驾驶软件栈）","abbr":"","aliases":["Autoware Universe","Autoware Core"],"one_liner":"A full open-source autonomous-driving software stack built on ROS 2, maintained by the Autoware Foundation.","explanation":"Autoware is a full open-source autonomous-driving software stack, originally started around 2015 on ROS 1 by Shinpei Kato's team at Nagoya University in Japan (later driven primarily by the company Tier IV), now maintained by the Autoware Foundation, with the current version built on ROS 2. It splits self-driving into modules — localization, perception, planning, control, mapping, and the vehicle interface — with each module made up of a set of ROS 2 nodes that developers can swap out individually. Its significance is in providing a complete, runnable reference implementation that universities and companies can build research on, or use as a base for campus shuttles and small vehicles. For embodied AI more broadly, it's a good example to study for how a large ROS 2 system is organized, how sensors get fused, and how planning connects to control.","example":"","related":["Robot Operating System 2","Autonomous Driving","ROS 2 Navigation Stack (Nav2)","Simultaneous Localization and Mapping","CARLA","Model Predictive Control"]},{"id":"px4-ardupilot","category":"software","sec":6,"tier":3,"sources":[{"title":"PX4 Autopilot 官网","url":"https://px4.io/"},{"title":"ArduPilot 官网","url":"https://ardupilot.org/"}],"as_of":"","related_ids":["unmanned-aerial-vehicle","aerial-vision-and-language-navigation","robot-operating-system-2","hardware-in-the-loop-software-in-the-loop-simulation","gazebo"],"name":"PX4 / ArduPilot","alt":"PX4 / ArduPilot","abbr":"","aliases":["PX4","ArduPilot","open-source flight stacks"],"one_liner":"The two leading open-source autopilot software stacks, used on drones, fixed-wing aircraft, and uncrewed ground or water vehicles.","explanation":"PX4 and ArduPilot are the two major open-source autopilot software stacks. PX4 originated at ETH Zurich and is now hosted by the Dronecode Foundation under the Linux Foundation, under a BSD license; ArduPilot grew out of the Arduino hobbyist community and uses a GPLv3 license. Both run on a flight-controller board and handle attitude estimation, attitude and position control, waypoint missions, and failsafe behavior; both support multirotors, fixed-wing aircraft, VTOL, ground vehicles, and boats, communicate with a ground station or onboard computer over the MAVLink protocol, and offer software-in-the-loop simulation. For aerial-robotics or drone-navigation research, they're typically used as the low-level controller, with higher-level algorithms running in ROS 2 on an onboard computer.","example":"When researching vision-language navigation for drones, the policy on the onboard computer outputs velocity commands, which are sent to PX4 over ROS 2, and PX4 handles the low-level attitude and motor control.","related":["Unmanned Aerial Vehicle (UAV)","Aerial Vision-and-Language Navigation","Robot Operating System 2","Hardware-in-the-Loop / Software-in-the-Loop Simulation","Gazebo"]},{"id":"fleet-management-system","category":"software","sec":6,"tier":3,"sources":[{"title":"Open-RMF 官网","url":"https://www.open-rmf.org/"},{"title":"open-rmf/rmf (GitHub)","url":"https://github.com/open-rmf/rmf"}],"as_of":"","related_ids":["multi-robot-collaboration","autonomous-mobile-robot","automated-guided-vehicle","multi-agent-path-finding","robot-operating-system-2","ros-2-navigation-stack"],"name":"Fleet Management System (e.g. Open-RMF)","alt":"多机调度系统","abbr":"FMS","aliases":["robot fleet scheduling system","Open-RMF"],"one_liner":"Upper-level software that assigns tasks, plans routes, and prevents conflicts across a whole fleet of robots.","explanation":"A fleet management system oversees a group of robots: it receives tasks and decides which robot to send, plans each one's route, prevents collisions and deadlocks at hallways and intersections, and coordinates shared infrastructure such as elevators, automatic doors, and charging stations. Warehouse AGV/AMR (automated guided vehicle / autonomous mobile robot) vendors typically each have their own proprietary scheduling system, which makes it hard to manage a mixed fleet of different brands together. Open-RMF (Robotics Middleware Framework), led by Open Robotics, is an open-source, ROS 2–based framework that provides traffic scheduling, task allocation, facility interfaces, and adapters for different robot brands, letting heterogeneous robots operate together in settings like hospitals and office buildings. It sits one level above single-robot navigation systems such as Nav2.","example":"A hospital's medicine-delivery robots and cleaning robots come from different manufacturers; Open-RMF coordinates them so they share elevators and don't collide in the same hallway.","related":["Multi-robot Collaboration","Autonomous Mobile Robot","Automated Guided Vehicle","Multi-Agent Path Finding","Robot Operating System 2","ROS 2 Navigation Stack (Nav2)"]},{"id":"ros-bag","category":"software","sec":7,"tier":2,"sources":[{"title":"ROS 2 Documentation: Recording and playing back data","url":"https://docs.ros.org/en/humble/Tutorials/Beginner-CLI-Tools/Recording-And-Playing-Back-Data/Recording-And-Playing-Back-Data.html"},{"title":"ros2/rosbag2 - GitHub","url":"https://github.com/ros2/rosbag2"}],"as_of":"","related_ids":["mcap","topic","trajectory-episode-replay","multi-sensor-time-synchronization-timestamp-alignment","lerobotdataset","foxglove-studio"],"name":"ROS Bag","alt":"rosbag","abbr":"","aliases":["ros2 bag","rosbag2","bag file"],"one_liner":"ROS's data-recording tool, which logs topic messages with timestamps so they can be replayed later.","explanation":"rosbag is ROS's built-in tool for recording and replaying data. It subscribes to specified topics and writes each message to a file along with its timestamp; on playback, it republishes the messages in their original time order, so downstream nodes behave as if the real robot were running live. ROS 1 uses the .bag format; the ROS 2 counterpart is rosbag2 (the ros2 bag command), which used SQLite storage by default early on and switched to the MCAP format by default starting with the Iron release. It's used to reproduce bugs, debug algorithms offline, and share data, and it's also the raw capture format behind many robot datasets — for imitation learning, the images and joint-state topics in a bag are commonly time-aligned and converted into training formats such as LeRobot or HDF5.","example":"ros2 bag record -a records every topic, ros2 bag info shows the duration and message counts, and ros2 bag play replays it.","related":["MCAP","Topic (ROS)","Trajectory / Episode Replay","Multi-sensor Time Synchronization / Timestamp Alignment","LeRobotDataset","Foxglove Studio"]},{"id":"ffmpeg","category":"software","sec":7,"tier":3,"sources":[{"title":"FFmpeg 官网","url":"https://ffmpeg.org/"},{"title":"LeRobot GitHub","url":"https://github.com/huggingface/lerobot"}],"as_of":"","related_ids":["lerobotdataset","data-cleaning","multi-sensor-time-synchronization-timestamp-alignment","demonstration-data"],"name":"FFmpeg","alt":"FFmpeg（音视频编解码工具）","abbr":"","aliases":[],"one_liner":"An open-source audio and video processing toolkit, used for transcoding, cutting, frame extraction, and compression.","explanation":"FFmpeg is a set of open-source audio and video processing tools and libraries, including command-line programs like ffmpeg and ffprobe as well as codec libraries such as libavcodec, and it supports nearly every common audio and video format. In embodied AI, robot datasets need to store huge amounts of camera footage, and storing every frame as a separate image file takes up far too much space, so footage is usually encoded into video files like MP4 and decoded back into frames at training time. The LeRobot dataset format stores each camera's footage as video, with encoding and decoding relying on FFmpeg or libraries built on it. It's also used for extracting frames for annotation, aligning sample rates, and compressing demo videos.","example":"Running ffmpeg -i episode_0.mp4 -vf fps=10 frames/%05d.png extracts a 30-frames-per-second demonstration video into images at 10 frames per second.","related":["LeRobotDataset","Data Cleaning","Multi-sensor Time Synchronization / Timestamp Alignment","Demonstration Data"]},{"id":"rqt","category":"software","sec":7,"tier":3,"sources":[{"title":"ROS 2 Docs: Overview and usage of RQt","url":"https://docs.ros.org/en/jazzy/Concepts/Intermediate/About-RQt.html"},{"title":"ROS Wiki: rqt","url":"http://wiki.ros.org/rqt"}],"as_of":"","related_ids":["rviz-rviz2","plotjuggler","foxglove-studio","node","topic","qt"],"name":"rqt","alt":"rqt","abbr":"","aliases":["rqt_graph","rqt_plot","rqt_console","rqt_image_view"],"one_liner":"ROS's built-in, plugin-based set of graphical debugging tools.","explanation":"rqt is a Qt-based GUI framework for ROS that packages various debugging widgets as plugins, which can be docked together into one window or launched individually. Common plugins include: rqt_graph, which draws the connections between nodes and topics; rqt_plot, which plots numeric values live; rqt_console, for viewing logs; rqt_image_view, for viewing camera feeds; rqt_reconfigure, for changing parameters on the fly; and rqt_topic, for inspecting topic content and rate. Both ROS 1 and ROS 2 have it. It's well suited to quickly checking whether data is actually being published and who is connected to whom; for 3D visualization, people use RViz instead, and for detailed curve analysis, they often switch to PlotJuggler or Foxglove.","example":"When a node isn't receiving data, running rqt_graph reveals that the subscriber's topic name has an extra namespace prefix, so it was never actually connected to the publisher.","related":["RViz / RViz2","PlotJuggler","Foxglove Studio","Node (ROS)","Topic (ROS)","Qt"]},{"id":"plotjuggler","category":"software","sec":7,"tier":3,"sources":[{"title":"facontidavide/PlotJuggler (GitHub)","url":"https://github.com/facontidavide/PlotJuggler"},{"title":"PlotJuggler 官网","url":"https://plotjuggler.io/"}],"as_of":"","related_ids":["ros-bag","robot-operating-system-2","foxglove-studio","rerun","rviz-rviz2","topic"],"name":"PlotJuggler","alt":"PlotJuggler","abbr":"","aliases":[],"one_liner":"Open-source time-series plotting tool commonly used to inspect ROS topics and log data.","explanation":"PlotJuggler is an open-source desktop tool by Davide Faconti, built on Qt, dedicated to plotting time-series data as curves. It can read CSV files, ROS/ROS 2 rosbag logs, PX4's ULog format, and other logs, and can also subscribe live to ROS topics, MQTT, ZeroMQ, and other data streams through plugins. Fields can be plotted just by dragging them in, and it supports multiple linked, synchronized-zoom windows and custom formulas — for example, differentiating a velocity signal or computing the difference between two signals. When debugging a robot, it's most commonly used to check whether the actual joint position is keeping up with the commanded one, whether the IMU is noisy, or whether the control loop is running at a stable rate.","example":"While tuning a legged robot, overlay the desired joint angle and the encoder reading on the same plot to spot at a glance which joint is lagging behind its command.","related":["ROS Bag","Robot Operating System 2","Foxglove Studio","Rerun","RViz / RViz2","Topic (ROS)"]},{"id":"foxglove-studio","category":"software","sec":7,"tier":3,"sources":[{"title":"Foxglove 官网","url":"https://foxglove.dev/"},{"title":"Foxglove Docs","url":"https://docs.foxglove.dev/"}],"as_of":"2024","related_ids":["rviz-rviz2","ros-bag","mcap","rerun","plotjuggler","robot-operating-system-2"],"name":"Foxglove Studio","alt":"Foxglove","abbr":"","aliases":["Foxglove"],"one_liner":"A robot data visualization tool that replays rosbags and shows curves, point clouds, and images in one view.","explanation":"Foxglove is a robot data visualization and debugging tool from the company Foxglove, available as both a web app and a desktop app. It can connect to a running ROS 1 or ROS 2 system, and can also open recorded rosbag and MCAP files (MCAP being a robot logging format that Foxglove itself helped launch), showing camera footage, 3D point clouds, the TF coordinate tree, joint curves, and logs all in one interface at once. Compared with RViz, it doesn't require ROS to be installed locally, and its layouts can be saved and shared, which makes it convenient for a team reviewing data together. It started out as the open-source Foxglove Studio; reportedly, starting in 2024 it became closed-source and was simply renamed Foxglove, with usage limits on the free tier — Rerun and PlotJuggler are open-source alternatives.","example":"After a real-robot policy run fails, drag the recorded MCAP file into Foxglove and compare the wrist-camera footage against the joint-torque curves to pinpoint the moment of failure.","related":["RViz / RViz2","ROS Bag","MCAP","Rerun","PlotJuggler","Robot Operating System 2"]},{"id":"rerun","category":"software","sec":7,"tier":2,"sources":[{"title":"Rerun documentation","url":"https://rerun.io/docs"},{"title":"rerun-io/rerun - GitHub","url":"https://github.com/rerun-io/rerun"}],"as_of":"","related_ids":["rviz-rviz2","foxglove-studio","plotjuggler","lerobot","viser"],"name":"Rerun","alt":"Rerun","abbr":"","aliases":["rerun.io","Rerun Viewer"],"one_liner":"An open-source visualization tool for multimodal, time-series robot data — images, point clouds, and curves played back in sync.","explanation":"Rerun is an open-source visualization SDK and viewer developed by the Swedish company Rerun, with bindings for Python, Rust, and C++. Developers call rr.log in their code to record images, depth maps, point clouds, 3D poses, scalar curves, text, and more, each tagged with a timestamp; the viewer lays all of it out on one shared timeline that can be scrubbed, replayed, and compared. It doesn't depend on ROS — installing a single pip package is enough to use it — which makes it convenient for debugging perception algorithms, inspecting a dataset, or watching a policy's inference process. Hugging Face's LeRobot uses it to visualize the camera footage and joint trajectories in its datasets, and many embodied-AI projects use it as a quicker alternative to RViz.","example":"Opening an episode with LeRobot's dataset-visualization script shows the wrist camera footage and each joint's angle curve side by side in Rerun.","related":["RViz / RViz2","Foxglove Studio","PlotJuggler","LeRobot","Viser"]},{"id":"meshcat","category":"software","sec":7,"tier":3,"sources":[{"title":"meshcat GitHub","url":"https://github.com/meshcat-dev/meshcat"},{"title":"meshcat-python GitHub","url":"https://github.com/meshcat-dev/meshcat-python"}],"as_of":"","related_ids":["drake","pinocchio","viser",null,null],"name":"Meshcat","alt":"Meshcat","abbr":"","aliases":["meshcat-python"],"one_liner":"A lightweight visualization tool that displays robots and 3D scenes in a web browser.","explanation":"Meshcat is an open-source 3D visualization tool Robin Deits developed during his PhD at MIT, built on WebGL (the browser's 3D graphics interface) and three.js. A backend program sends meshes, point clouds, coordinate frames, and poses over WebSocket to a web page, which can then be opened in a browser to view and rotate — no desktop graphics environment needed, which makes it well suited to remote servers and Jupyter notebooks. Drake has it built in, and libraries such as Pinocchio provide a Meshcat interface too; it's commonly used to check inverse-kinematics results, replay trajectories, or inspect collision geometry.","example":"In Jupyter, load a URDF with Pinocchio, call MeshcatVisualizer, and watch the arm's pose update in the browser as joint angles change.","related":["Drake","Pinocchio","Viser","RViz / RViz2","Unified Robot Description Format (URDF)"]},{"id":"viser","category":"software","sec":7,"tier":3,"sources":[{"title":"Viser documentation","url":"https://viser.studio/"},{"title":"nerfstudio-project/viser GitHub","url":"https://github.com/nerfstudio-project/viser"}],"as_of":"","related_ids":["rerun","meshcat","nerfstudio-gsplat","pyroki","unified-robot-description-format","vuer"],"name":"Viser","alt":"Viser","abbr":"","aliases":[],"one_liner":"A Python library for showing 3D scenes and interactive controls right in the browser.","explanation":"Viser is an open-source Python 3D visualization library from the nerfstudio team at UC Berkeley. Starting a server from Python opens a 3D view in the browser, into which you can add point clouds, meshes, camera frustums, and coordinate axes, as well as sliders, buttons, and other controls whose callbacks are received back in Python. Because it works over the web, when the program runs on a remote GPU server, forwarding the port to a local browser is enough to view it — no graphical desktop environment needed. It comes with a built-in tool for loading a URDF (a robot's model-description file), and projects such as the nerfstudio viewer and PyRoki use it for visualization and debugging.","example":"While training a reconstruction model on a server, use Viser to show the current point cloud and the robot's URDF live, with a slider to drag a joint angle by hand and check that the kinematics look right.","related":["Rerun","Meshcat","Nerfstudio / gsplat","PyRoki","Unified Robot Description Format","Vuer"]},{"id":"vuer","category":"software","sec":7,"tier":3,"sources":[{"title":"vuer-ai/vuer GitHub","url":"https://github.com/vuer-ai/vuer"},{"title":"Vuer documentation","url":"https://docs.vuer.ai/"}],"as_of":"","related_ids":["open-television","vr-teleoperation","apple-vision-pro","unitree-xr-teleoperate","viser","web-real-time-communication"],"name":"Vuer","alt":"Vuer（网页 3D / XR 可视化与遥操作工具）","abbr":"","aliases":[],"one_liner":"A Python-driven 3D visualization framework that runs inside a VR headset's browser.","explanation":"Vuer is an open-source Python 3D visualization framework by Ge Yang, built on web technology and supporting WebXR (the browser standard for accessing VR/AR devices). After starting a server on the Python side, headsets such as the Apple Vision Pro or Meta Quest can enter the 3D scene just by opening a page in their browser, with no native app needed. It can stream the rendered view to the headset while simultaneously streaming head- and hand-keypoint poses back to Python in real time — exactly the two-way channel teleoperation needs, which is why projects such as Open-TeleVision and Unitree's xr_teleoperate use it as their VR teleoperation front end.","example":"Open-TeleVision uses Vuer to display a robot's stereo head-camera feed inside the Vision Pro, while reading the operator's hand keypoints and retargeting them into arm and dexterous-hand motions for a humanoid robot.","related":["Open-TeleVision","VR Teleoperation","Apple Vision Pro","Unitree xr_teleoperate","Viser","Web Real-Time Communication (WebRTC)"]},{"id":"qt","category":"software","sec":7,"tier":3,"sources":[{"title":"Qt 官网","url":"https://www.qt.io/"},{"title":"Qt Documentation","url":"https://doc.qt.io/"}],"as_of":"","related_ids":["host-computer","rviz-rviz2","rqt","plotjuggler","python-and-c-plus-plus"],"name":"Qt","alt":"Qt（上位机图形界面开发框架）","abbr":"","aliases":["PyQt","PySide"],"one_liner":"A cross-platform C++ GUI framework commonly used for robot operator-side software interfaces.","explanation":"Qt is a cross-platform C++ application development framework, originally released by the Norwegian company Trolltech in the 1990s and now maintained by The Qt Company under a dual open-source (GPL/LGPL) and commercial license. It offers both a traditional widget-based interface (Qt Widgets) and a declarative one (QML / Qt Quick), and the same codebase can be built for Windows, Linux, macOS, and embedded devices; Python developers can use it through PyQt or the official PySide bindings. Operator-side robotics software — tuning panels, data-collection clients, teaching interfaces — is frequently written in Qt, and ROS's own RViz, rqt, and PlotJuggler are all built on it as well.","example":"Write an operator interface for a custom robot arm: use Qt to build a UI that shows each joint's angle and current draw, with enable, home, and e-stop buttons.","related":["Host Computer","RViz / RViz2","rqt","PlotJuggler","Python and C++"]},{"id":"pytorch","category":"software","sec":8,"tier":1,"sources":[{"title":"PyTorch official site","url":"https://pytorch.org/"},{"title":"PyTorch Foundation","url":"https://pytorch.org/foundation"}],"as_of":"","related_ids":["cuda","jax","lerobot","checkpoint","python-and-c-plus-plus","hugging-face-transformers"],"name":"PyTorch","alt":"PyTorch","abbr":"","aliases":["torch"],"one_liner":"Today’s dominant deep-learning framework, used to train most embodied AI models.","explanation":"PyTorch is an open-source deep-learning framework originally developed by Meta’s (then Facebook’s) AI research team, now hosted by the PyTorch Foundation under the Linux Foundation. It’s built around tensor (multi-dimensional array) operations and automatic differentiation, with code that reads close to ordinary Python and is easy to debug, which is why it dominates in academia. Most embodied AI models — Diffusion Policy, ACT, OpenVLA, the PyTorch version of π0, LeRobot, and more — are built on PyTorch, which calls into an NVIDIA GPU through CUDA for acceleration. The core things a newcomer needs to learn are tensor operations, building a network with nn.Module, loading data with DataLoader, writing an optimizer training loop, and saving and loading checkpoints.","example":"After import torch, model.to('cuda') moves the network onto the GPU for training, and torch.save writes out a checkpoint once training finishes.","related":["CUDA","JAX","LeRobot","Checkpoint","Python and C++","Hugging Face Transformers"]},{"id":"jax","category":"software","sec":8,"tier":2,"sources":[{"title":"jax-ml/jax GitHub","url":"https://github.com/jax-ml/jax"},{"title":"Flax 文档","url":"https://flax.readthedocs.io/"}],"as_of":"","related_ids":["pytorch","tensorflow","mujoco-xla","brax","openpi","octo"],"name":"JAX","alt":"JAX","abbr":"","aliases":["Flax"],"one_liner":"Google's NumPy-like framework for numerical computing and deep learning, with automatic differentiation and compilation.","explanation":"JAX is Google's open-source Python library for numerical computing, with an interface close to NumPy plus a set of composable function transforms: grad for automatic differentiation, jit for compiling code into fast GPU/TPU programs via the XLA compiler, and vmap for automatic batching. Flax is a neural-network library built on top of JAX that defines layers and manages parameters; the two are usually used together. In embodied AI, JAX shows up heavily in large-scale parallel simulation and in Google-adjacent work: MJX and Brax use it to run physics simulation on GPUs, and Octo and Physical Intelligence's openpi were originally implemented in JAX. Compared with PyTorch, it leans more functional in style and has a somewhat steeper learning curve.","example":"The π0 training code in the openpi repository was originally written in JAX and Flax.","related":["PyTorch","TensorFlow","MuJoCo XLA","Brax","openpi (Physical Intelligence)","Octo"]},{"id":"tensorflow","category":"software","sec":8,"tier":3,"sources":[{"title":"TensorFlow 官网","url":"https://www.tensorflow.org/"},{"title":"TensorFlow Datasets","url":"https://www.tensorflow.org/datasets"}],"as_of":"","related_ids":["pytorch","jax","tensorflow-datasets","rlds","tfrecord","open-x-embodiment"],"name":"TensorFlow","alt":"TensorFlow","abbr":"TF","aliases":["TF"],"one_liner":"Google's open-source deep learning framework; robotics datasets often use its data format.","explanation":"TensorFlow is the deep learning framework Google Brain open-sourced in 2015, with version 2.x adopting Keras as its high-level interface. Academic model training has largely shifted to PyTorch and JAX since then, but TensorFlow still shows up regularly in robotics: Google's RT-1 and other work is built on it, and datasets such as Open X-Embodiment are stored in the RLDS format as TFRecord files, read through TensorFlow Datasets (TFDS); the data-loading pipelines for Octo and OpenVLA also depend on tf.data. So even when training a VLA in PyTorch, TensorFlow usually still needs to be installed just to read the data.","example":"Use tfds.builder to load the BridgeData V2 subset from Open X-Embodiment, then convert it to PyTorch tensors to feed into fine-tuning OpenVLA.","related":["PyTorch","JAX","TensorFlow Datasets (TFDS)","RLDS (Reinforcement Learning Datasets)","TFRecord","Open X-Embodiment"]},{"id":"mindspore","category":"software","sec":8,"tier":3,"sources":[{"title":"昇思 MindSpore 官网","url":"https://www.mindspore.cn/"},{"title":"MindSpore GitHub","url":"https://github.com/mindspore-ai/mindspore"}],"as_of":"","related_ids":[null,"pytorch",null,null,null],"name":"MindSpore","alt":"昇思 MindSpore","abbr":"","aliases":["Shengsi"],"one_liner":"Huawei's open-source deep-learning framework, built mainly to pair with Ascend chips.","explanation":"MindSpore (Chinese name 昇思) is a deep-learning framework Huawei open-sourced in 2020, occupying the same role as PyTorch or TensorFlow — defining networks, and running training and inference. It supports CPU and GPU, but is optimized primarily for Huawei's Ascend NPUs, calling into the hardware underneath through Ascend CANN (Huawei's chip operator and compiler software stack). It's something you're likely to encounter when training or deploying models in a domestic Chinese compute environment; running a VLA or similar model on an Ascend server or Ascend edge chip for embodied AI generally means either porting PyTorch code over to MindSpore or using PyTorch's Ascend adaptation plugin.","example":"","related":["Compute Architecture for Neural Networks (Huawei Ascend, CANN)","PyTorch","Huawei","Inference Deployment","Huawei Cloud CloudRobo Embodied AI Platform"]},{"id":"nvidia-warp","category":"software","sec":8,"tier":3,"sources":[{"title":"NVIDIA/warp (GitHub)","url":"https://github.com/NVIDIA/warp"}],"as_of":"2025","related_ids":["newton-physics-engine","mujoco-warp","differentiable-simulation","cuda","gpu-accelerated-parallel-simulation","operator-kernel"],"name":"NVIDIA Warp","alt":"Warp","abbr":"","aliases":["Warp"],"one_liner":"NVIDIA's Python framework that compiles ordinary Python functions into GPU code for writing simulations.","explanation":"Warp is an open-source Python framework from NVIDIA for writing high-performance simulation and graphics code. Developers write ordinary Python functions and mark them with the @wp.kernel decorator; Warp then just-in-time compiles them into CUDA or CPU code that runs in parallel, so there is no need to write C++/CUDA by hand. It also supports automatic differentiation, so a simulation written in Warp can be differentiated with respect to its inputs — an approach known as differentiable simulation. Warp includes built-in tools for geometry, collision, and particles, and can exchange data directly with PyTorch and JAX. NVIDIA's Newton physics engine and MuJoCo Warp are both built on top of it.","example":"Write a cloth simulation in a few dozen lines of Warp code, advance thousands of parallel environments at once on the GPU, and backpropagate gradients into PyTorch to optimize a grasping trajectory.","related":["Newton Physics Engine","MuJoCo Warp","Differentiable Simulation","CUDA","GPU-Accelerated Parallel Simulation","Operator / Kernel"]},{"id":"taichi-programming-language-quadrants","category":"software","sec":8,"tier":3,"sources":[{"title":"taichi-dev/taichi (GitHub)","url":"https://github.com/taichi-dev/taichi"},{"title":"Genesis-Embodied-AI/Genesis (GitHub)","url":"https://github.com/Genesis-Embodied-AI/Genesis"}],"as_of":"2026-09","related_ids":["genesis","material-point-method","differentiable-simulation","nvidia-warp","gpu-accelerated-parallel-simulation"],"name":"Taichi Programming Language / Quadrants","alt":"Taichi（太极）/ Quadrants","abbr":"","aliases":["Taichi","Quadrants (Genesis fork)"],"one_liner":"A high-performance parallel programming language embedded in Python; the engine underneath the Genesis simulator.","explanation":"Taichi is a high-performance, parallel programming language embedded in Python, proposed by Yuanming Hu during his PhD at MIT (SIGGRAPH Asia 2019). Users write computational kernels using Python syntax, and Taichi just-in-time compiles them to run in parallel on CUDA, Vulkan, Metal, or CPU — a particularly good fit for physics simulations like the material point method, fluids, and soft bodies. The Genesis simulator's physics computation is built on Taichi; the Genesis team reportedly later maintained their own fork of Taichi, named Quadrants, so they could keep improving and adapting it for their needs. Students in embodied AI mostly encounter it when reading Genesis's source code or writing custom simulations.","example":"Decorate a Python function with @ti.kernel to write a million-particle material-point-method simulation that runs directly on the GPU.","related":["Genesis","Material Point Method","Differentiable Simulation","NVIDIA Warp","GPU-Accelerated Parallel Simulation"]},{"id":"hugging-face-transformers","category":"software","sec":8,"tier":2,"sources":[{"title":"huggingface/transformers GitHub","url":"https://github.com/huggingface/transformers"},{"title":"Transformers 文档","url":"https://huggingface.co/docs/transformers/index"}],"as_of":"","related_ids":["hugging-face","transformer","vision-language-model","openvla","hugging-face-diffusers","hugging-face-mirror"],"name":"Hugging Face Transformers","alt":"Transformers 库","abbr":"","aliases":["transformers","🤗 Transformers"],"one_liner":"Hugging Face's open-source Python library for loading pretrained models with a few lines of code.","explanation":"Transformers is Hugging Face's open-source library, bundling the architecture definitions and loading code for a huge number of pretrained models, spanning language models, vision models, and vision-language models (VLMs — models that can look at an image and describe it in words). Its core interfaces are AutoModel, AutoProcessor, and from_pretrained: give it a model name and it automatically downloads the weights, configuration, and tokenizer (the tool that splits text into tokens). In embodied AI, many VLA backbones — PaliGemma, Qwen-VL, Llama, and others — are loaded directly through it, and models like OpenVLA also publish their weights in a Transformers-compatible format. That makes it a foundational library that anyone reading code or fine-tuning models in this space has to know.","example":"OpenVLA's official example loads the 7B weights with AutoModelForVision2Seq.from_pretrained, then calls predict_action to output robot arm actions.","related":["Hugging Face","Transformer","Vision-Language Model","OpenVLA","Hugging Face Diffusers","Hugging Face Mirror (hf-mirror.com)"]},{"id":"hugging-face-diffusers","category":"software","sec":8,"tier":3,"sources":[{"title":"Diffusers 官方文档","url":"https://huggingface.co/docs/diffusers/index"},{"title":"Diffusion Policy GitHub","url":"https://github.com/real-stanford/diffusion_policy"}],"as_of":"","related_ids":["diffusion-model","diffusion-policy","denoising-diffusion-probabilistic-model","denoising-diffusion-implicit-model","noise-schedule","hugging-face"],"name":"Hugging Face Diffusers","alt":"Diffusers 库","abbr":"","aliases":["diffusers"],"one_liner":"Hugging Face's open-source diffusion-model toolkit, with ready-made schedulers, networks, and generation pipelines.","explanation":"Diffusers is Hugging Face's open-source library for diffusion models, built on PyTorch. It splits a diffusion model into three pieces: a scheduler (which decides the noising and denoising steps, such as DDPM or DDIM), a model (such as a U-Net or a diffusion transformer), and a pipeline that strings the two together for inference — and it can load many publicly available image and video generation models directly. In robotics, people often only borrow its scheduler: Diffusion Policy swaps “denoising an image” for “denoising a sequence of actions,” and the sampling steps are the same as in image diffusion, so there's no need to rewrite them from scratch. It's also commonly used to load a pretrained video model and fine-tune it when building a world model or doing video generation.","example":"The official Diffusion Policy code calls Diffusers' DDPMScheduler and DDIMScheduler directly to train and sample actions.","related":["Diffusion Model","Diffusion Policy","Denoising Diffusion Probabilistic Model","Denoising Diffusion Implicit Model","Noise Schedule","Hugging Face"]},{"id":"safetensors","category":"software","sec":8,"tier":2,"sources":[{"title":"huggingface/safetensors (GitHub)","url":"https://github.com/huggingface/safetensors"},{"title":"Safetensors 文档","url":"https://huggingface.co/docs/safetensors"}],"as_of":"","related_ids":["checkpoint","pytorch","hugging-face","hugging-face-transformers","open-weight-model"],"name":"safetensors","alt":"safetensors 权重格式","abbr":"","aliases":[".safetensors"],"one_liner":"A safe, fast-loading file format from Hugging Face for storing model weights.","explanation":"safetensors is a tensor-storage format developed and open-sourced by Hugging Face. A file consists of a JSON header — recording each tensor's name, shape, data type, and byte offset — followed by contiguous raw data. It was designed to fix a security problem with PyTorch's usual .pt/.bin files, which are based on Python's pickle format and can execute arbitrary code when loaded; safetensors also supports memory-mapping and reading individual tensors on demand, making large models load faster. Most open weights on Hugging Face today, including many VLA model checkpoints, are released as .safetensors, and libraries such as transformers and LeRobot can read them directly.","example":"A VLA model's download includes a model.safetensors file; calling safetensors.torch.load_file reads out the parameter dictionary directly, with no risk of the file hiding malicious code.","related":["Checkpoint","PyTorch","Hugging Face","Hugging Face Transformers","Open-weight Model"]},{"id":"open-weight-model","category":"software","sec":8,"tier":2,"sources":[{"title":"The Open Source AI Definition - Open Source Initiative","url":"https://opensource.org/ai/open-source-ai-definition"}],"as_of":"","related_ids":["open-source-license","openpi","pre-training","fine-tuning","llama","openvla"],"name":"Open-weight Model","alt":"开放权重","abbr":"","aliases":["open weights"],"one_liner":"A model whose trained parameters are published for download, even if its training data and code aren't.","explanation":"An open-weight model is one whose developer publishes the trained parameter files publicly, so anyone can download them, run inference locally, and fine-tune them. This differs from open source in the strict sense: the Open Source Initiative's (OSI) definition of open-source AI also requires information about training data and complete training code, whereas many open-weight models only provide the weights and inference code, and their licenses may restrict commercial use or cap the number of users. For embodied-AI researchers, open weights mean an existing VLA can be fine-tuned on their own robot data instead of pretraining from scratch — this is a big part of why the open-source ecosystem has grown so quickly over the past couple of years.","example":"Physical Intelligence released π0's weights through openpi, letting researchers download them and fine-tune on their own robot arm data.","related":["Open-Source License (Apache 2.0 / MIT / Non-commercial)","openpi (Physical Intelligence)","Pre-training","Fine-tuning","Llama","OpenVLA"]},{"id":"open-source-license","category":"software","sec":8,"tier":2,"sources":[{"title":"Licenses - Open Source Initiative","url":"https://opensource.org/licenses"},{"title":"Choose an open source license","url":"https://choosealicense.com/"}],"as_of":"","related_ids":["open-weight-model","github","openpi","lerobot","open-source-hardware"],"name":"Open-Source License (Apache 2.0 / MIT / Non-commercial)","alt":"开源许可证","abbr":"","aliases":["open-source license","License"],"one_liner":"The legal terms attached to public code or models that spell out who can use, modify, or sell them, and how.","explanation":"An open-source license is the set of terms attached to code or a model that determines what users are allowed to do and what obligations they take on. MIT is the most permissive — keep the copyright notice and you can use and sell it freely; Apache 2.0 also permits commercial use, adds an explicit patent grant, and requires noting any modifications; GPL requires derivative works to also be open-sourced; and CC BY-NC or various vendors' own “non-commercial” licenses restrict use to research only. In embodied AI, a lot of model code is released under Apache 2.0, but the weights or datasets that go with it may carry separate restrictions — before building a product on top of something, check the licenses for code, weights, and data separately.","example":"The openpi repository's code is released under Apache 2.0, while some of the datasets it uses are restricted to non-commercial research.","related":["Open-weight Model","GitHub","openpi (Physical Intelligence)","LeRobot","Open-Source Hardware (OSHW)"]},{"id":"hugging-face-mirror","category":"software","sec":8,"tier":2,"sources":[{"title":"HF-Mirror","url":"https://hf-mirror.com/"},{"title":"Hugging Face Hub 环境变量文档（HF_ENDPOINT）","url":"https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables"}],"as_of":"","related_ids":["hugging-face","hugging-face-transformers","modelscope","package-mirror-sources-in-china","lerobot","autodl"],"name":"Hugging Face Mirror (hf-mirror.com)","alt":"HF 镜像站","abbr":"","aliases":["hf-mirror","HF_ENDPOINT"],"one_liner":"An unofficial mirror site that makes it fast to download Hugging Face models and datasets from inside China.","explanation":"hf-mirror.com is a community-run mirror of Hugging Face, not an official Hugging Face service. Direct connections to huggingface.co from within China are often slow or fail outright, yet almost all VLA models, vision encoders, and LeRobot datasets are hosted on Hugging Face, so the mirror has become the standard entry point for reproducing papers there. It is simple to use: set the environment variable HF_ENDPOINT to the mirror's address, and downloads from transformers, huggingface-cli, LeRobot, and similar tools automatically route through it instead of the original site. Gated models, which require accepting a license before download, still need permission requested on the official Hugging Face site plus an access token — the mirror does not bypass that step.","example":"After running export HF_ENDPOINT=https://hf-mirror.com, you can use huggingface-cli download to fetch the openvla/openvla-7b weights.","related":["Hugging Face","Hugging Face Transformers","ModelScope","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)","LeRobot","AutoDL"]},{"id":"modelscope","category":"software","sec":8,"tier":2,"sources":[{"title":"ModelScope 魔搭社区","url":"https://modelscope.cn/"},{"title":"modelscope/modelscope GitHub","url":"https://github.com/modelscope/modelscope"}],"as_of":"","related_ids":["hugging-face","hugging-face-mirror","qwen-vl","llama-factory-ms-swift","hugging-face-transformers"],"name":"ModelScope","alt":"魔搭社区","abbr":"","aliases":["Moda"],"one_liner":"Alibaba's domestic open-source model and dataset community, roughly a Chinese counterpart to Hugging Face.","explanation":"ModelScope (魔搭社区) is an open-source model platform launched by Alibaba in 2022 that hosts model weights, datasets, and online demos. It is reliably accessible from within China, and Chinese-made models such as the Qwen series are usually released here in sync with Hugging Face. It provides a modelscope Python library and command-line tool for downloading models directly, along with companion fine-tuning frameworks such as ms-swift. For students in China working on embodied AI, ModelScope and the Hugging Face mirror are the two go-to routes for downloading VLM backbones or robotics datasets.","example":"Running modelscope download --model Qwen/Qwen2.5-VL-7B-Instruct fetches the Qwen2.5-VL weights over a domestic Chinese connection, for use as a VLA's vision-language backbone.","related":["Hugging Face","Hugging Face Mirror (hf-mirror.com)","Qwen-VL","LLaMA-Factory / ms-swift (ModelScope SWIFT)","Hugging Face Transformers"]},{"id":"llama-factory-ms-swift","category":"software","sec":8,"tier":3,"sources":[{"title":"LLaMA-Factory GitHub","url":"https://github.com/hiyouga/LLaMA-Factory"},{"title":"ms-swift GitHub","url":"https://github.com/modelscope/ms-swift"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"LLaMA-Factory / ms-swift (ModelScope SWIFT)","alt":"LLaMA-Factory / ms-swift（大模型与 VLM 微调框架）","abbr":"","aliases":["LlamaFactory","SWIFT","ms-swift"],"one_liner":"Two popular open-source fine-tuning frameworks for LLMs and VLMs, usable by just editing a config file.","explanation":"LLaMA-Factory is an open-source fine-tuning framework maintained by the developer hiyouga (Yaowei Zheng and others), with an accompanying paper presented as a system demonstration at ACL 2024; ms-swift is the equivalent framework from Alibaba's ModelScope team. Both package full-parameter fine-tuning, parameter-efficient methods like LoRA/QLoRA, and preference-alignment and reinforcement-learning methods like DPO and GRPO into a configuration file or command-line interface, and both support multimodal models such as Qwen-VL and InternVL. In embodied AI, they're commonly used to supervised-fine-tune a VLM — for example, training a model for embodied reasoning, spatial question-answering, or task planning — which can then serve as a VLA's backbone or its high-level planner.","example":"Write a YAML file for LLaMA-Factory and use your own annotated robot-scene question-answering data to LoRA-fine-tune Qwen2.5-VL.","related":["Fine-tuning","Low-Rank Adaptation (LoRA)","Supervised Fine-Tuning","Qwen-VL series (Qwen2.5-VL / Qwen3-VL)","ModelScope","Hugging Face Transformers"]},{"id":"lerobot","category":"software","sec":8,"tier":1,"sources":[{"title":"huggingface/lerobot (GitHub)","url":"https://github.com/huggingface/lerobot"},{"title":"LeRobot documentation","url":"https://huggingface.co/docs/lerobot/index"}],"as_of":"","related_ids":["hugging-face","lerobotdataset","so-100-so-101-arm","smolvla","action-chunking-with-transformers","imitation-learning"],"name":"LeRobot","alt":"LeRobot","abbr":"","aliases":["Hugging Face LeRobot","lerobot"],"one_liner":"An open-source robot learning toolkit from Hugging Face, covering the whole pipeline from data collection to training to deployment.","explanation":"LeRobot is a PyTorch robot-learning library open-sourced by Hugging Face in 2024, aimed at lowering the barrier to learning on real robots. It provides a unified dataset format (LeRobotDataset, hosted on the Hugging Face Hub), ships built-in implementations of policies such as ACT, Diffusion Policy, π0, and SmolVLA, and includes scripts for teleoperated data collection, training, evaluation, and running on a real robot. It pairs with low-cost open-source arms such as the SO-100/SO-101, letting a student complete the whole pipeline — collect demonstration data, train with imitation learning, deploy on the real robot — with only a few hundred to a couple thousand dollars of hardware, which has made it one of the most common starting points for getting into embodied AI.","example":"Two SO-101 arms are used in a leader-follower teleoperation setup to record a few dozen grasping demonstrations, which are then used to train an ACT policy in LeRobot and run it on the real arm.","related":["Hugging Face","LeRobotDataset","SO-100 / SO-101 Arm","SmolVLA","Action Chunking with Transformers","Imitation Learning"]},{"id":"lerobot-envhub","category":"software","sec":8,"tier":3,"sources":[{"title":"LeRobot 文档 EnvHub","url":"https://huggingface.co/docs/lerobot/envhub"},{"title":"huggingface/lerobot (GitHub)","url":"https://github.com/huggingface/lerobot"}],"as_of":"2026-09","related_ids":[null,"hugging-face",null,null,null],"name":"LeRobot EnvHub","alt":"EnvHub","abbr":"","aliases":[],"one_liner":"LeRobot's mechanism for sharing simulation environments, publishing environment code to the Hugging Face Hub for one-line loading.","explanation":"EnvHub is a simulation-environment-sharing feature Hugging Face built into its LeRobot framework. Previously, using someone else's simulated task meant cloning their repository, installing its dependencies, and rewriting code to match their interface; EnvHub lets an author publish environment code as a Hub repository, and a user in LeRobot then loads it by repository name and gets back an environment conforming to the Gymnasium interface (reset/step), ready to use for training or evaluating a policy. Because loading it executes code from that repository, it requires explicitly allowing remote code execution, so it should only be loaded from a trusted source. It lives on the same Hub as LeRobot's datasets and models, making it convenient to publish data, a policy, and an evaluation environment together in one place.","example":"","related":["LeRobot (Hugging Face)","Hugging Face","Gymnasium (formerly OpenAI Gym)","Simulation-based Evaluation","LeRobotDataset (LeRobot dataset format)"]},{"id":"openpi","category":"software","sec":8,"tier":2,"sources":[{"title":"Physical-Intelligence/openpi - GitHub","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"2025-09","related_ids":["pi0","pi0-fast","pi0-5","physical-intelligence","lerobot","open-weight-model"],"name":"openpi (Physical Intelligence)","alt":"openpi","abbr":"","aliases":[],"one_liner":"Physical Intelligence's open-source code and weights repository for its π-series VLA models.","explanation":"openpi is the open-source code repository Physical Intelligence released on GitHub in 2025, providing pretrained weights along with inference and fine-tuning code for π0, π0-FAST, and the later π0.5, released under the Apache 2.0 license. The repository also includes fine-tuning examples on platforms such as DROID, ALOHA, and LIBERO, plus an inference server that lets a policy be called as a remote service. It lets teams without large-scale robot data of their own fine-tune an existing, capable VLA on their own data instead of starting from scratch, making it one of the most common starting points for VLA work today in both academia and startups.","example":"Plug your own LeRobot-format data into openpi, fine-tune the π0.5 base model, and use its policy server to drive a real robot.","related":["π0","π0-FAST","π0.5","Physical Intelligence","LeRobot","Open-weight Model"]},{"id":"dexbotic","category":"software","sec":8,"tier":3,"sources":[{"title":"Dexbotic GitHub","url":"https://github.com/Dexmal/dexbotic"}],"as_of":"2026-09","related_ids":["dexmal","vision-language-action-model","lerobot","openpi","starvla","fine-tuning"],"name":"Dexbotic (Dexmal VLA toolbox)","alt":"Dexbotic","abbr":"","aliases":[],"one_liner":"An open-source VLA training-and-deployment toolbox from Dexmal that brings several mainstream VLA methods into one codebase.","explanation":"Dexbotic is an open-source, PyTorch-based toolbox for VLA (vision-language-action model) development, released by Dexmal (原力灵机). Different VLA papers each tend to come with their own repository, data format, and training scripts, which makes reproducing and comparing them a lot of work; Dexbotic folds several mainstream VLA methods into a single unified framework, with a common data interface, pretrained weights, and shared training and evaluation pipelines, so researchers can swap models more easily for experiments, or fine-tune on their own robot data and deploy. It sits in the same category as LeRobot, openpi, and starVLA — a “VLA code base” — and exactly which models it supports is best checked against the official repository.","example":"Convert your own dual-arm data into Dexbotic's data format, fine-tune one of its built-in VLA configurations, then deploy to a real robot using the inference script the repository provides.","related":["Dexmal","Vision-Language-Action Model","LeRobot","openpi (Physical Intelligence)","starVLA","Fine-tuning"]},{"id":"starvla","category":"software","sec":8,"tier":3,"sources":[{"title":"starVLA/starVLA (GitHub)","url":"https://github.com/starVLA/starVLA"}],"as_of":"2026-09","related_ids":["vision-language-action-model","action-head","qwen-vl","libero-benchmark","dexbotic","openpi"],"name":"starVLA","alt":"starVLA","abbr":"","aliases":[],"one_liner":"An open-source codebase for assembling and training VLA models like building blocks.","explanation":"starVLA is an open-source codebase for developing vision-language-action (VLA) models, self-described as “LEGO-style”: it splits a vision-language-model backbone and various types of action head into swappable modules, so researchers can combine autoregressive discrete-action, continuous-regression, or flow-matching action heads within the same codebase and compare them under one unified training and evaluation pipeline. According to its project page, it uses the Qwen-VL model family as its main backbone and plugs into common benchmarks such as LIBERO and SimplerEnv. It suits newcomers who want to quickly reproduce and compare different VLA design choices.","example":"Keep the Qwen-VL backbone fixed and swap only the action head, from discrete tokens to a flow-matching head, then compare the two on LIBERO success rate.","related":["Vision-Language-Action Model","Action Head","Qwen-VL","LIBERO Benchmark","Dexbotic (Dexmal VLA toolbox)","openpi (Physical Intelligence)"]},{"id":"gemini-robotics-sdk","category":"software","sec":8,"tier":3,"sources":[{"title":"Gemini Robotics On-Device brings AI to local robotic devices (Google DeepMind)","url":"https://deepmind.google/discover/blog/gemini-robotics-on-device-brings-ai-to-local-robotic-devices/"}],"as_of":"2025-06","related_ids":["gemini-robotics-on-device","gemini-robotics","software-development-kit","mujoco","fine-tuning","google-deepmind"],"name":"Gemini Robotics SDK","alt":"Gemini Robotics SDK","abbr":"","aliases":[],"one_liner":"Google DeepMind's development kit for evaluating and fine-tuning its Gemini Robotics models.","explanation":"The Gemini Robotics SDK is a software development kit Google DeepMind released in June 2025 alongside Gemini Robotics On-Device (a VLA model that can run locally on a robot), initially made available to select developers through a trusted-tester program. It lets developers evaluate the model on their own tasks, test it inside MuJoCo physics simulation, and adapt it to a new task or a new robot using a small number of demonstrations (the official materials mention roughly 50–100). This means Gemini Robotics is moving beyond papers and demos toward being something outside teams can build on, though access is still controlled rather than freely downloadable.","example":"","related":["Gemini Robotics On-Device","Gemini Robotics","Software Development Kit","MuJoCo (Multi-Joint dynamics with Contact)","Fine-tuning","Google DeepMind"]},{"id":"gymnasium","category":"software","sec":8,"tier":2,"sources":[{"title":"Gymnasium documentation","url":"https://gymnasium.farama.org/"}],"as_of":"","related_ids":["environment","reinforcement-learning","termination-vs-truncation","stable-baselines3","gym-gymnasium-mujoco-tasks","mujoco"],"name":"Gymnasium","alt":"Gymnasium","abbr":"Gym","aliases":["OpenAI Gym","gym"],"one_liner":"The standard interface library for reinforcement learning environments, and the successor to OpenAI Gym.","explanation":"Gymnasium is a reinforcement-learning environment library maintained by the Farama Foundation; it grew out of Gym, which OpenAI released in 2016, and which Farama took over and renamed once OpenAI stopped maintaining it. Its most important contribution is a unified interface: env.reset() starts an episode and returns the initial observation, and env.step(action) executes an action and returns the new observation, the reward, terminated (the task ended naturally), truncated (cut off by something like a step limit), and extra info. It ships with classic tasks such as CartPole and MuJoCo continuous-control tasks built in. Most reinforcement-learning algorithm libraries are written against this interface, and many robot simulation environments provide a compatible wrapper, so an algorithm written once can be pointed at a new environment directly.","example":"env = gymnasium.make(“HalfCheetah-v5”), then calling env.step() in a loop to collect data, which is handed to Stable-Baselines3’s PPO for training.","related":["Environment (Env; reset/step interface)","Reinforcement Learning","Termination vs. Truncation","Stable-Baselines3","Gym/Gymnasium MuJoCo Tasks","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"stable-baselines3","category":"software","sec":8,"tier":3,"sources":[{"title":"Stable-Baselines3 Docs","url":"https://stable-baselines3.readthedocs.io/"},{"title":"DLR-RM/stable-baselines3 (GitHub)","url":"https://github.com/DLR-RM/stable-baselines3"}],"as_of":"","related_ids":["reinforcement-learning","proximal-policy-optimization","soft-actor-critic","gymnasium","cleanrl","skrl"],"name":"Stable-Baselines3","alt":"Stable-Baselines3","abbr":"SB3","aliases":["SB3"],"one_liner":"A PyTorch-based library of classic reinforcement learning algorithms, with a simple API and reliable implementations.","explanation":"Stable-Baselines3 is an open-source, PyTorch-based reinforcement learning library maintained by Antonin Raffin and colleagues at the German Aerospace Center's Robotics Institute (DLR-RM). It's the successor to Stable Baselines, itself derived from OpenAI Baselines, with the associated paper published in JMLR in 2021. It provides reliable implementations of algorithms including PPO, SAC, TD3, DQN, and A2C — training a policy on a Gymnasium environment can take just a few lines of code — with thorough documentation and tests, and it's often used as a starting point or a reference baseline. It mainly targets a single environment or a small number of parallel ones; for large-scale GPU-parallel simulation training, people more often reach for rsl_rl or rl_games instead.","example":"model = PPO(“MlpPolicy”, env); model.learn(100000) — a few lines of code are enough to train a policy on Gymnasium's cart-pole task.","related":["Reinforcement Learning","Proximal Policy Optimization","Soft Actor-Critic","Gymnasium","CleanRL","skrl"]},{"id":"cleanrl","category":"software","sec":8,"tier":3,"sources":[{"title":"vwxyzjn/cleanrl (GitHub)","url":"https://github.com/vwxyzjn/cleanrl"},{"title":"CleanRL (JMLR 2022)","url":"https://www.jmlr.org/papers/v23/21-1342.html"}],"as_of":"","related_ids":["reinforcement-learning","proximal-policy-optimization","stable-baselines3","gymnasium","rsl-rl","weights-and-biases"],"name":"CleanRL","alt":"CleanRL","abbr":"","aliases":[],"one_liner":"An open-source deep reinforcement-learning library that implements each algorithm as a single, readable file.","explanation":"CleanRL is an open-source deep reinforcement-learning library, primarily authored by Shengyi Huang and others, with an accompanying paper published in JMLR in 2022. Its distinguishing feature is “single-file implementation”: algorithms such as PPO, DQN, SAC, TD3, and DDPG are each written in one standalone Python file, with everything from environment creation to the network and training loop on the same page, with none of the layered abstraction found in more modular libraries. The benefit is that newcomers can read every detail of an algorithm start to finish, and researchers can copy and modify a file directly to run an experiment; it also has built-in logging to Weights & Biases and TensorBoard. Compared with a modular library like Stable-Baselines3, it's better suited to learning and research prototyping than to being called as a general-purpose toolkit.","example":"Running python cleanrl/ppo_continuous_action.py --env-id HalfCheetah-v4 trains PPO on a MuJoCo environment, with the curves viewable in TensorBoard.","related":["Reinforcement Learning","Proximal Policy Optimization","Stable-Baselines3","Gymnasium","rsl_rl","Weights & Biases (W&B)"]},{"id":"rsl-rl","category":"software","sec":8,"tier":3,"sources":[{"title":"leggedrobotics/rsl_rl (GitHub)","url":"https://github.com/leggedrobotics/rsl_rl"}],"as_of":"","related_ids":["legged-gym","nvidia-isaac-lab","proximal-policy-optimization","massively-parallel-reinforcement-learning","eth-zurich-robotic-systems-lab","rl-games"],"name":"rsl_rl","alt":"rsl_rl","abbr":"","aliases":["rsl-rl","RSL RL"],"one_liner":"ETH Zurich's open-source reinforcement learning library built for GPU-parallel simulation.","explanation":"rsl_rl is a lightweight, open-source PyTorch reinforcement learning library from the Robotic Systems Lab (RSL) at ETH Zurich, originally built to train quadruped locomotion in Isaac Gym alongside legged_gym. At its core is a PPO (Proximal Policy Optimization) implementation optimized for thousands of parallel environments, with all sampling and update data kept on the GPU to avoid costly CPU round-trips. It's now one of the RL libraries Isaac Lab officially supports out of the box, and it's also used by legged- and humanoid-locomotion projects such as unitree_rl_gym; later versions have added recurrent-network policies, symmetry augmentation, and teacher-student distillation. Reproducing a legged-locomotion paper today can hardly avoid it.","example":"Use rsl_rl's training script in Isaac Lab to train a velocity-tracking policy for the Unitree Go2, sampling thousands of environments in parallel on the GPU before updating with PPO.","related":["legged_gym","NVIDIA Isaac Lab","Proximal Policy Optimization","Massively Parallel Reinforcement Learning","ETH Zurich Robotic Systems Lab","rl_games"]},{"id":"rl-games","category":"software","sec":8,"tier":3,"sources":[{"title":"rl_games on GitHub","url":"https://github.com/Denys88/rl_games"}],"as_of":"","related_ids":["proximal-policy-optimization","massively-parallel-reinforcement-learning","nvidia-isaac-lab","isaac-gym","rsl-rl","skrl"],"name":"rl_games","alt":"rl_games","abbr":"","aliases":[],"one_liner":"A reinforcement learning training library optimized for massively parallel GPU simulation, common in the Isaac stack.","explanation":"rl_games is an open-source reinforcement learning library by Denys Makoviichuk and collaborators, implementing mainly PPO (Proximal Policy Optimization) and SAC. Its defining feature is that training data can stay entirely on the GPU throughout, which suits simulators like Isaac Gym that run thousands of parallel environments at once. NVIDIA's IsaacGymEnvs uses it as the default training backend, and Isaac Lab lists it as one of its supported training libraries; the open-source code behind many dexterous-hand and legged-locomotion papers is built on it. Compared with rsl_rl, it has more features but more configuration surface, and hyperparameters are usually written in YAML files.","example":"Run rl_games' training script inside Isaac Lab with a task name specified, and it will train a robot hand to reorient a cube with PPO across thousands of parallel environments.","related":["Proximal Policy Optimization","Massively Parallel Reinforcement Learning","NVIDIA Isaac Lab","Isaac Gym","rsl_rl","skrl"]},{"id":"skrl","category":"software","sec":8,"tier":3,"sources":[{"title":"skrl documentation","url":"https://skrl.readthedocs.io/"},{"title":"Toni-SM/skrl (GitHub)","url":"https://github.com/Toni-SM/skrl"}],"as_of":"","related_ids":["nvidia-isaac-lab","rsl-rl","rl-games","stable-baselines3","proximal-policy-optimization","gymnasium"],"name":"skrl","alt":"skrl","abbr":"","aliases":[],"one_liner":"A modular Python reinforcement learning library that plugs directly into Isaac Lab.","explanation":"skrl is an open-source Python reinforcement learning library started by Antonio Serrano-Muñoz, supporting both PyTorch and JAX backends. It splits agents, memory (experience storage), models, and trainers into separate, swappable modules, and includes common algorithms such as PPO (Proximal Policy Optimization), SAC, and TD3 out of the box. Its defining feature is native support for Gymnasium as well as NVIDIA's GPU-parallel environments, Isaac Gym and Isaac Lab; it's one of the few reinforcement learning libraries officially supported by Isaac Lab, alongside rsl_rl, rl_games, and Stable-Baselines3.","example":"When training quadruped locomotion in Isaac Lab, swap the training script's library argument from rsl_rl to skrl, and run the same task with its PPO implementation instead.","related":["NVIDIA Isaac Lab","rsl_rl","rl_games","Stable-Baselines3","Proximal Policy Optimization","Gymnasium"]},{"id":"humanoidverse","category":"software","sec":8,"tier":3,"sources":[{"title":"LeCAR-Lab/HumanoidVerse GitHub","url":"https://github.com/LeCAR-Lab/HumanoidVerse"}],"as_of":"2025","related_ids":["asap","sim-to-sim-transfer","sim-to-real-transfer","isaac-gym","genesis","holosoma"],"name":"HumanoidVerse","alt":"HumanoidVerse（多仿真器人形 sim2real 框架）","abbr":"","aliases":[],"one_liner":"An open-source, multi-simulator humanoid robot learning framework from CMU's LeCAR Lab.","explanation":"HumanoidVerse is an open-source humanoid-robot learning framework from Carnegie Mellon University's LeCAR Lab (Guanya Shi's group). It decouples the simulator, the task, and the robot from one another: the same task and training code can switch between different simulation backends such as IsaacGym, Isaac Sim, and Genesis, and can also switch between different humanoid robot bodies. That makes “sim-to-sim” validation convenient — training in one simulator and testing in another — catching a policy that has overfit to one particular physics engine before it ever gets transferred to a real robot. The same lab's ASAP (Aligning Simulation and Real Physics for learning Agile whole-body Skills) is built on top of it.","example":"Train a Unitree G1's motion-tracking policy in IsaacGym using HumanoidVerse, then switch to Genesis to run a sim-to-sim check.","related":["ASAP","Sim-to-Sim Transfer","Sim-to-Real Transfer","Isaac Gym","Genesis","Holosoma (Amazon FAR humanoid RL framework)"]},{"id":"holosoma","category":"software","sec":8,"tier":3,"sources":[{"title":"amazon-far/holosoma GitHub","url":"https://github.com/amazon-far/holosoma"}],"as_of":"2025","related_ids":["amazon-frontier-ai-and-robotics","humanoidverse","rl-based-locomotion-control","sim-to-real-transfer","motion-tracking","unitree-g1"],"name":"Holosoma (Amazon FAR humanoid RL framework)","alt":"Holosoma（亚马逊人形 RL 训练部署框架）","abbr":"","aliases":[],"one_liner":"An open-source humanoid reinforcement-learning framework for training and deployment, from Amazon's FAR team.","explanation":"Holosoma is a humanoid robotics framework open-sourced on GitHub by Amazon's Frontier AI & Robotics (FAR) team. According to its repository, it connects “training a policy in simulation with reinforcement learning” to “deploying that policy on a real robot,” covering both velocity-command-based walking and whole-body motion tracking, able to train across multiple simulation backends and supporting humanoid platforms such as the Unitree G1. A common pain point in humanoid RL projects is that training code, the simulator, and real-robot deployment each end up as a separate, one-off setup, so switching robots or simulators means rebuilding almost everything; the value of a framework like this is in offering a ready-made pipeline. It sits in the same category as HumanoidVerse and legged_gym.","example":"","related":["Amazon Frontier AI & Robotics","HumanoidVerse","RL-based Locomotion Control","Sim-to-Real Transfer","Motion Tracking","Unitree G1"]},{"id":"gr00t-wholebodycontrol","category":"software","sec":8,"tier":3,"sources":[{"title":"NVlabs/GR00T-WholeBodyControl (GitHub)","url":"https://github.com/NVlabs/GR00T-WholeBodyControl"}],"as_of":"2025-12","related_ids":["whole-body-control","decoupled-whole-body-control","nvidia-isaac-gr00t-n1","sonic","unitree-g1","nvidia-generalist-embodied-agent-research-lab"],"name":"GR00T-WholeBodyControl","alt":"GR00T 全身控制代码库","abbr":"GR00T WBC","aliases":["GR00T WBC","GR00T whole-body control platform"],"one_liner":"NVIDIA's open-source codebase for training, evaluating, and deploying whole-body controllers for humanoid robots.","explanation":"GR00T-WholeBodyControl is a codebase NVIDIA open-sourced on GitHub (under NVlabs), providing model weights for humanoid whole-body controllers along with training, evaluation, and deployment scripts. Whole-body control means coordinating the legs, torso, and arms together, so a robot keeps its balance while walking and manipulating objects at the same time. According to the repository, it includes the decoupled whole-body controller used by GR00T N1.5 and N1.6 on the Unitree G1 (a reinforcement-learning policy handling lower-body walking and balance, with inverse kinematics tracking hand targets for the upper body), as well as controllers like SONIC that are based on large-scale motion tracking. Within the GR00T system, it sits at the “cerebellum” layer: the upper-level VLA supplies a goal, and this layer converts it into joint commands.","example":"GR00T N1.6 outputs a walking-velocity command and target poses for both hands on a Unitree G1, and the decoupled controller in this codebase converts them into whole-body joint commands.","related":["Whole-Body Control","Decoupled Whole-Body Control","NVIDIA Isaac GR00T N1","SONIC","Unitree G1","NVIDIA Generalist Embodied Agent Research Lab"]},{"id":"autodl","category":"software","sec":9,"tier":2,"sources":[{"title":"AutoDL 官网","url":"https://www.autodl.com/"},{"title":"AutoDL 帮助文档","url":"https://www.autodl.com/docs/"}],"as_of":"","related_ids":["secure-shell","docker","conda","package-mirror-sources-in-china","common-training-and-inference-gpus","hugging-face-mirror"],"name":"AutoDL","alt":"AutoDL（GPU 算力租用平台）","abbr":"","aliases":["AutoDL GPU Cloud"],"one_liner":"A Chinese pay-by-the-hour GPU cloud platform, popular with students for running experiments.","explanation":"AutoDL is a Chinese GPU-rental cloud platform: users pick a GPU model and a base image on its website, spin up a GPU-backed container instance, and are billed by usage time, accessing it through JupyterLab or SSH. The platform ships pre-built images with common environments such as PyTorch and CUDA already installed, plus a data disk, file storage, and network acceleration aimed at academic sites. Students without a lab server, or without a strong enough local GPU, often use it to train and fine-tune models and run simulations. Remember to shut the instance down when you’re done, since it keeps billing otherwise; data should be kept in a persistent directory, since it’s lost once the instance is released.","example":"Renting an RTX 4090 and booting the official PyTorch image, then SSHing in to clone the LeRobot code and fine-tune an ACT policy.","related":["Secure Shell (SSH)","Docker","Conda","Package Mirror Sources in China (pip / conda / apt / Docker mirrors)","Common Training and Inference GPUs (RTX 4090 / A100 / H100 / B200)","Hugging Face Mirror (hf-mirror.com)"]},{"id":"tensorboard","category":"software","sec":9,"tier":2,"sources":[{"title":"TensorBoard 官方页面","url":"https://www.tensorflow.org/tensorboard"},{"title":"tensorflow/tensorboard (GitHub)","url":"https://github.com/tensorflow/tensorboard"}],"as_of":"","related_ids":["weights-and-biases","swanlab","loss-function","rsl-rl","tensorflow","pytorch"],"name":"TensorBoard","alt":"TensorBoard","abbr":"","aliases":["tensorboard"],"one_liner":"A local, browser-based tool for viewing training curves, images, and other logged data.","explanation":"TensorBoard is an open-source training-visualization tool released by Google alongside TensorFlow, though PyTorch can also write logs to it directly through torch.utils.tensorboard. During training, a program writes scalars like loss, reward, and learning rate, along with images and histograms, to a log directory; running tensorboard --logdir then shows live curves in a browser and lets you compare multiple runs. It runs entirely locally and needs no account, which is why many robot reinforcement-learning frameworks — such as rsl_rl and Isaac Lab's training scripts — log to it by default; teams that need cloud collaboration and experiment management often switch to Weights & Biases or SwanLab instead.","example":"While training a quadruped walking policy in Isaac Lab, open TensorBoard to watch whether the average reward and episode-length curves are trending up, as a sign that training is converging.","related":["Weights & Biases (W&B)","SwanLab","Loss Function","rsl_rl","TensorFlow","PyTorch"]},{"id":"weights-and-biases","category":"software","sec":9,"tier":2,"sources":[{"title":"Weights & Biases 官网","url":"https://wandb.ai/site"},{"title":"W&B 文档","url":"https://docs.wandb.ai/"}],"as_of":"2025-05","related_ids":["tensorboard","swanlab","hyperparameter","checkpoint","lerobot"],"name":"Weights & Biases (W&B)","alt":"Weights & Biases","abbr":"W&B","aliases":["wandb"],"one_liner":"A cloud platform for logging and comparing deep-learning training runs online.","explanation":"Weights & Biases is a machine-learning experiment-management platform built by the company of the same name, acquired by the cloud-compute company CoreWeave in 2025. Adding a few lines — wandb.init and wandb.log — to training code streams loss curves, hyperparameters, GPU utilization, and sample videos to a web dashboard in real time, viewable from a phone or browser, with tools for comparing dozens of runs side by side and sharing them with collaborators. Mainstream open-source embodied-AI frameworks such as LeRobot and openpi have built-in wandb logging. When Chinese network access to it is unreliable, people commonly substitute a local TensorBoard setup or the domestic tool SwanLab.","example":"Adding --wandb.enable=true when training a diffusion policy with LeRobot makes the loss curve and evaluation-episode videos viewable on the wandb website.","related":["TensorBoard","SwanLab","Hyperparameter","Checkpoint","LeRobot"]},{"id":"swanlab","category":"software","sec":9,"tier":3,"sources":[{"title":"SwanHubX/SwanLab (GitHub)","url":"https://github.com/SwanHubX/SwanLab"}],"as_of":"","related_ids":["weights-and-biases","tensorboard","hyperparameter","llama-factory-ms-swift"],"name":"SwanLab","alt":"SwanLab（国产训练实验追踪工具）","abbr":"","aliases":[],"one_liner":"An open-source Chinese experiment-tracking and visualization tool for model training, similar to W&B.","explanation":"SwanLab is an open-source experiment-tracking tool built by a team in China, functionally similar to Weights & Biases: add a few lines to training code, and it logs loss curves, hyperparameters, images, and hardware usage to a web dashboard, making it easy to compare multiple runs and share results with a team. It offers both a cloud-hosted version and a self-hostable one, is directly accessible from within mainland China, and integrates with common frameworks such as PyTorch, Hugging Face Transformers, and LLaMA-Factory. It's a common substitute when training models in China, where reaching W&B's servers can be inconvenient.","example":"Call swanlab.init in a training script, then swanlab.log at every step to record the loss, and compare the curves of three different learning rates on the web dashboard.","related":["Weights & Biases (W&B)","TensorBoard","Hyperparameter","LLaMA-Factory / ms-swift (ModelScope SWIFT)"]},{"id":"hydra","category":"software","sec":9,"tier":3,"sources":[{"title":"Hydra 官网","url":"https://hydra.cc/"},{"title":"facebookresearch/hydra (GitHub)","url":"https://github.com/facebookresearch/hydra"}],"as_of":"","related_ids":[null,null,null,"nvidia-isaac-lab","weights-and-biases"],"name":"Hydra (Meta configuration framework)","alt":"Hydra 配置框架","abbr":"","aliases":["facebookresearch/hydra"],"one_liner":"Meta's open-source Python configuration framework for composing YAML configs and overriding parameters from the command line.","explanation":"Hydra is an open-source Python configuration framework from Meta (originally Facebook AI Research), built on OmegaConf for reading and writing YAML underneath. It splits an experiment's configuration into several small, composable files — one each for the model, the dataset, the task — assembled as needed at runtime, and lets any field be overridden directly from the command line; its multirun feature can also sweep a whole set of hyperparameters (hyperparameters being values set by hand before training, such as the learning rate) in one go. Robot-learning code tends to have a lot of parameters and a lot of experiments, and Hydra means “change one parameter and rerun” doesn't require touching code — it also automatically creates an output directory for each run and saves the full configuration used at the time, which makes results reproducible. Codebases such as Diffusion Policy and Isaac Lab both use it to manage training configuration.","example":"Running python train.py --config-name=train_diffusion_unet_image_workspace task=pusht_image training.seed=42 in the Diffusion Policy repository swaps the task and random seed without touching any code.","related":["Hyperparameter","Random Seed & Reproducibility","Diffusion Policy","NVIDIA Isaac Lab","Weights & Biases (W&B)"]},{"id":"numerical-precision-formats","category":"software","sec":9,"tier":2,"sources":[{"title":"Wikipedia: bfloat16 floating-point format","url":"https://en.wikipedia.org/wiki/Bfloat16_floating-point_format"},{"title":"NVIDIA Transformer Engine: Using FP8","url":"https://docs.nvidia.com/deeplearning/transformer-engine/user-guide/examples/fp8_primer.html"}],"as_of":"","related_ids":["mixed-precision-training","post-training-quantization","quantization-aware-training","gpu-memory","on-device-edge-deployment","nvidia-tensorrt"],"name":"Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)","alt":"数值精度格式（FP32 / FP16 / BF16 / FP8 / INT8 / INT4）","abbr":"","aliases":["FP32","FP16","BF16","FP8","INT8","INT4","half precision","single precision"],"one_liner":"The bit width and format used to store a model's numbers, which sets its memory use, speed, and accuracy.","explanation":"Every parameter and intermediate result in a neural network is a number, and how many bits are used to store it directly affects memory, speed, and precision. FP32 is 32-bit single precision, the most stable but the most memory-hungry; FP16 half precision uses only 16 bits with a small numeric range, so training with it easily overflows and needs loss scaling; BF16 also uses 16 bits but keeps the same exponent range as FP32 — wide range, lower precision — and is the default choice for large-model training today; FP8 uses only 8 bits and needs GPUs from the H100 generation or later; INT8 and INT4 are integer formats mainly used to compress models for inference. As a rough rule, memory equals parameter count times bytes per number (4 for FP32, 2 for BF16, 1 for INT8, 0.5 for INT4), so a 7B-parameter model stored in BF16 needs roughly 14 GB, or about 3.5 GB quantized to INT4 — which determines whether a VLA can fit on a robot's onboard compute.","example":"A 7B model like OpenVLA takes roughly 15 GB of GPU memory loaded in BF16; quantizing it to 4 bits lets it run on smaller GPUs, though action accuracy may drop.","related":["Mixed-Precision Training","Post-Training Quantization","Quantization-Aware Training","GPU Memory (VRAM)","On-Device / Edge Deployment","NVIDIA TensorRT"]},{"id":"distributed-data-parallel","category":"software","sec":9,"tier":3,"sources":[{"title":"PyTorch DDP 文档","url":"https://docs.pytorch.org/docs/stable/generated/torch.nn.parallel.DistributedDataParallel.html"}],"as_of":"","related_ids":["fully-sharded-data-parallel","deepspeed","distributed-training","batch-size","pytorch","gradient-accumulation"],"name":"Distributed Data Parallel (DDP)","alt":"分布式数据并行","abbr":"DDP","aliases":["DistributedDataParallel"],"one_liner":"A multi-GPU training method where each GPU holds a full copy of the model and processes a slice of the data, then syncs gradients.","explanation":"Distributed data parallel is the most common way to train across multiple GPUs, implemented in PyTorch as torch.nn.parallel.DistributedDataParallel. The approach: every GPU (every process) holds a complete copy of the model; a batch of data is split up and divided among the GPUs, each running its own forward and backward pass, then gradients are synchronized through all-reduce (a form of collective communication where every GPU's gradients are averaged together), after which each GPU updates its parameters, keeping every copy identical. It lets training speed scale close to linearly with the number of GPUs, but requires the full model to fit on a single GPU; when the model is too large for that, a parameter-sharding approach like FSDP or DeepSpeed's ZeRO is needed instead. It's typically launched with torchrun.","example":"Running torchrun --nproc_per_node=8 train.py trains a diffusion policy across 8 GPUs on one machine, with a per-GPU batch size of 32 and an effective total batch size of 256.","related":["Fully Sharded Data Parallel (FSDP)","DeepSpeed","Distributed Training","Batch Size","PyTorch","Gradient Accumulation"]},{"id":"fully-sharded-data-parallel","category":"software","sec":9,"tier":3,"sources":[{"title":"PyTorch FSDP 文档","url":"https://pytorch.org/docs/stable/fsdp.html"},{"title":"PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel","url":"https://arxiv.org/abs/2304.11277"}],"as_of":"","related_ids":["distributed-data-parallel","deepspeed","distributed-training","pytorch","full-fine-tuning","mixed-precision-training"],"name":"Fully Sharded Data Parallel (FSDP)","alt":"全分片数据并行","abbr":"FSDP","aliases":["FSDP2","fully_shard"],"one_liner":"A data-parallel training method that splits a model's parameters, gradients, and optimizer state across multiple GPUs.","explanation":"FSDP is PyTorch's built-in distributed training approach, based on the idea behind Microsoft DeepSpeed's ZeRO-3. In ordinary distributed data parallel (DDP), every GPU stores a full copy of the model parameters, gradients, and optimizer state, so once a model gets large enough it simply doesn't fit on one GPU. FSDP shards all three across the available GPUs; before each layer's forward and backward pass, the full parameters for that layer are temporarily gathered from the other GPUs and released again right after, so each GPU only needs to hold a fraction of the total at any moment, making it possible to train models far larger than a single GPU's capacity — at the cost of more inter-GPU communication. Full-parameter fine-tuning of a multi-billion-parameter VLA or video world model commonly uses FSDP or DeepSpeed; newer PyTorch versions recommend the per-layer FSDP2 (fully_shard) interface.","example":"Full-parameter fine-tuning a roughly 3-billion-parameter VLA on 8 GPUs with 80GB each runs out of memory under DDP, but switching to FSDP makes it fit.","related":["Distributed Data Parallel (DDP)","DeepSpeed","Distributed Training","PyTorch","Full Fine-Tuning","Mixed-Precision Training"]},{"id":"deepspeed","category":"software","sec":9,"tier":3,"sources":[{"title":"DeepSpeed 官网","url":"https://www.deepspeed.ai/"},{"title":"DeepSpeed GitHub","url":"https://github.com/deepspeedai/DeepSpeed"}],"as_of":"","related_ids":["fully-sharded-data-parallel","distributed-data-parallel","mixed-precision-training","distributed-training","hugging-face-accelerate","pytorch"],"name":"DeepSpeed","alt":"DeepSpeed","abbr":"","aliases":[],"one_liner":"Microsoft's open-source library for accelerating large-model distributed training, best known for its ZeRO memory optimizer.","explanation":"DeepSpeed is an open-source deep-learning training and inference optimization library from Microsoft, built on top of PyTorch. It's best known for ZeRO (Zero Redundancy Optimizer): instead of every GPU storing a full copy of the optimizer state, gradients, and even the model parameters, ZeRO shards them across multiple GPUs, making it possible to train larger models within limited memory. It also offers mixed precision, gradient accumulation, CPU/NVMe offloading, and pipeline parallelism. When training a VLA or large language model with parameters in the billions, a single GPU often can't hold it all, and DeepSpeed — or PyTorch's own FSDP, which follows a similar idea — is a common solution. Frameworks such as Hugging Face Accelerate can call into it directly.","example":"When fine-tuning a 7B-parameter VLA, turning on DeepSpeed's ZeRO-2 or ZeRO-3 configuration in the training script lets 8 GPUs split the optimizer state and gradients between them.","related":["Fully Sharded Data Parallel (FSDP)","Distributed Data Parallel (DDP)","Mixed-Precision Training","Distributed Training","Hugging Face Accelerate","PyTorch"]},{"id":"hugging-face-accelerate","category":"software","sec":9,"tier":3,"sources":[{"title":"Accelerate 官方文档","url":"https://huggingface.co/docs/accelerate/index"}],"as_of":"","related_ids":["distributed-data-parallel","fully-sharded-data-parallel","deepspeed","mixed-precision-training","pytorch","hugging-face"],"name":"Hugging Face Accelerate","alt":"Accelerate 库","abbr":"","aliases":["accelerate"],"one_liner":"Hugging Face's training-helper library that lets the same PyTorch code run across multiple GPUs with minimal changes.","explanation":"Accelerate is an open-source PyTorch helper library from Hugging Face. A training script written for a single GPU only needs a few added lines — creating an Accelerator and wrapping the model, optimizer, and data loader with it — to run unchanged across a single GPU, multiple GPUs, multiple machines, or with mixed precision, and it can switch to a distributed backend such as DeepSpeed or FSDP (Fully Sharded Data Parallel) without the developer having to write process initialization or gradient synchronization by hand. The companion commands accelerate config and accelerate launch generate a configuration and then launch training. Fine-tuning a large model like a VLA often needs multiple GPUs, and a lot of open-source robot-learning code uses Accelerate to manage that distributed training.","example":"Running accelerate launch starts a VLA fine-tuning script across 8 GPUs, with bf16 mixed precision turned on.","related":["Distributed Data Parallel (DDP)","Fully Sharded Data Parallel (FSDP)","DeepSpeed","Mixed-Precision Training","PyTorch","Hugging Face"]},{"id":"slurm-workload-manager","category":"software","sec":9,"tier":3,"sources":[{"title":"Slurm Workload Manager - Documentation","url":"https://slurm.schedmd.com/documentation.html"}],"as_of":"","related_ids":["distributed-training","distributed-data-parallel","ray","docker","secure-shell"],"name":"Slurm Workload Manager","alt":"Slurm 集群调度","abbr":"","aliases":["Slurm","SLURM"],"one_liner":"The scheduling system that queues jobs and allocates GPUs on a shared compute cluster.","explanation":"Slurm is an open-source Linux cluster job-scheduling system, originally from Lawrence Livermore National Laboratory in the US and now maintained primarily by the company SchedMD; it's used by a large share of supercomputing centers, universities, and companies' GPU clusters. When many people share a pool of machines, instead of logging into a specific machine and running a job directly, everyone submits jobs to Slurm, which queues and allocates nodes and GPUs by resource availability and priority. Common commands include sbatch (submit a script), srun (run directly), squeue (check the queue), and scancel (cancel a job). Training large models like VLAs, or doing multi-node, multi-GPU distributed training, can hardly avoid it.","example":"Write a train.sh that declares a need for 2 nodes with 8 GPUs each, then submit it with sbatch train.sh and check whether it's been scheduled with squeue.","related":["Distributed Training","Distributed Data Parallel (DDP)","Ray","Docker","Secure Shell (SSH)"]},{"id":"ray","category":"software","sec":9,"tier":3,"sources":[{"title":"Ray 官网","url":"https://www.ray.io/"},{"title":"Ray: A Distributed Framework for Emerging AI Applications (arXiv)","url":"https://arxiv.org/abs/1712.05889"}],"as_of":"","related_ids":["verl","distributed-training","actor-learner-architecture","slurm-workload-manager","massively-parallel-reinforcement-learning"],"name":"Ray","alt":"Ray 分布式框架","abbr":"","aliases":[],"one_liner":"Open-source distributed computing framework for scaling Python programs across many machines and GPUs.","explanation":"Ray originated at UC Berkeley's RISELab, with its paper published at OSDI in 2018; it's now developed primarily by Anyscale and remains free and open source. It lets an ordinary Python function or class become a parallel unit — a stateless task or a stateful actor — just by adding a decorator, letting a cluster run it in parallel while Ray handles scheduling, data transfer, and fault tolerance. Built on top of this core are libraries such as Ray Train (distributed training), Ray Tune (hyperparameter search), RLlib (reinforcement learning), Ray Data, and Ray Serve. Reinforcement learning needs to schedule large numbers of simulation, inference, and training processes at once, and large-model RL frameworks such as veRL use Ray to orchestrate these components.","example":"When fine-tuning a VLA with reinforcement learning, use Ray actors to run simulation rollout, policy inference, and parameter updates on separate GPUs, all coordinated by one driver process.","related":["veRL (Volcano Engine Reinforcement Learning)","Distributed Training","Actor-Learner Architecture (Distributed RL)","Slurm Workload Manager","Massively Parallel Reinforcement Learning"]},{"id":"verl","category":"software","sec":9,"tier":3,"sources":[{"title":"volcengine/verl GitHub","url":"https://github.com/volcengine/verl"},{"title":"HybridFlow: A Flexible and Efficient RLHF Framework (arXiv)","url":"https://arxiv.org/abs/2409.19256"}],"as_of":"","related_ids":["reinforcement-fine-tuning","group-relative-policy-optimization","proximal-policy-optimization","vllm","simplevla-rl","ray"],"name":"veRL (Volcano Engine Reinforcement Learning)","alt":"veRL","abbr":"","aliases":["verl","HybridFlow"],"one_liner":"ByteDance's open-source large-model reinforcement learning framework, also used for VLA RL fine-tuning.","explanation":"veRL is a reinforcement learning training library open-sourced by ByteDance's Seed team (released under the Volcano Engine name), and is the open-source implementation of the HybridFlow paper (EuroSys 2025). Doing RL on a large model means repeatedly alternating between two things: using the current model to generate a batch of samples (rollout), then using those samples to update the parameters — and the two steps want different kinds of parallelism. veRL hands generation to inference engines such as vLLM or SGLang, hands training to FSDP or Megatron, and handles the weight synchronization and scheduling between the two, with PPO and GRPO built in. It originally targeted language models, and was later adopted by the embodied-AI community for online reinforcement-learning fine-tuning of VLAs.","example":"SimpleVLA-RL is built on veRL: it runs parallel rollouts in simulators such as LIBERO, using task success or failure as the reward to fine-tune OpenVLA-OFT with GRPO-style reinforcement learning.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Group Relative Policy Optimization","Proximal Policy Optimization","vLLM","SimpleVLA-RL","Ray"]},{"id":"rlinf","category":"software","sec":9,"tier":3,"sources":[{"title":"RLinf on GitHub","url":"https://github.com/RLinf/RLinf"}],"as_of":"2025-09","related_ids":["reinforcement-fine-tuning","vision-language-action-model","verl","simplevla-rl","maniskill","libero-benchmark"],"name":"RLinf","alt":"RLinf","abbr":"","aliases":[],"one_liner":"Open-source, large-scale reinforcement learning training framework for embodied and agentic foundation models.","explanation":"RLinf is a reinforcement learning infrastructure open-sourced in 2025, jointly developed by institutions including Tsinghua University, aimed at making large-model RL training both fast and flexible. It schedules simulation, inference (rollout generation), and training onto the same pool of GPUs, switching between them or running them in a pipeline as needed to keep GPU utilization high. On the embodied side, it supports online reinforcement-learning fine-tuning of vision-language-action models such as OpenVLA, OpenVLA-OFT, and π0 inside simulators like ManiSkill and LIBERO; it can also be used for reasoning-oriented RL on large language models. It suits researchers who want to apply RL post-training to a VLA without building their own distributed system from scratch.","example":"Use RLinf's provided configuration to fine-tune OpenVLA-OFT with PPO or GRPO reinforcement learning inside the ManiSkill simulator, improving pick-and-place success rate.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Vision-Language-Action Model","veRL (Volcano Engine Reinforcement Learning)","SimpleVLA-RL","ManiSkill","LIBERO Benchmark"]},{"id":"inference-deployment","category":"software","sec":10,"tier":1,"sources":[{"title":"NVIDIA TensorRT","url":"https://developer.nvidia.com/tensorrt"},{"title":"openpi (Physical Intelligence) - remote inference","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["inference-latency","on-device-edge-deployment","policy-server","nvidia-tensorrt","post-training-quantization","inference"],"name":"Inference Deployment","alt":"推理部署","abbr":"","aliases":["Model Deployment","Deployment"],"one_liner":"Getting a trained model running on real hardware or a server so it outputs actions in real time.","explanation":"Inference deployment means moving a trained model from its training environment onto the hardware where it will actually be used; “inference” here means the model’s forward computation, not logical reasoning. In embodied AI, this means the model has to be wired up to real-time inputs — camera feeds, joint sensors — and output actions to the controller at a fixed rate. The hard parts are latency and compute: a robot’s onboard computer (such as a Jetson) has limited compute, so a large model often needs to be quantized, pruned, or accelerated with an engine such as TensorRT, or else run on a remote GPU server that sends actions back to the robot over the network (a policy server). How well the deployment is engineered directly affects whether the robot’s motion is smooth and whether it can correct itself in a closed loop.","example":"A fine-tuned π0 model runs as a policy server on a GPU workstation; the robot periodically sends images and joint states over the network and receives back a chunk of actions.","related":["Inference Latency","On-Device / Edge Deployment","Policy Server (Remote Inference)","NVIDIA TensorRT","Post-Training Quantization","Inference"]},{"id":"policy-server","category":"software","sec":10,"tier":2,"sources":[{"title":"Physical-Intelligence/openpi - GitHub","url":"https://github.com/Physical-Intelligence/openpi"},{"title":"LeRobot docs: Asynchronous Inference","url":"https://huggingface.co/docs/lerobot/async"}],"as_of":"","related_ids":["inference-deployment","asynchronous-inference","inference-latency","action-chunking","openpi","cloud-edge-device-collaboration"],"name":"Policy Server (Remote Inference)","alt":"策略服务器","abbr":"","aliases":["remote inference","server-client deployment","inference server"],"one_liner":"Running a large policy on a GPU server and having the robot send it observations and get actions back.","explanation":"A policy server is a deployment pattern: a large model policy, such as a VLA, runs on a GPU-equipped workstation or in the cloud, while a lightweight client on the robot sends camera images, joint states, and instructions over the network; the server runs inference and returns a chunk of actions, which the client passes to the low-level controller for execution. It solves the problem of a robot's onboard compute not being enough to hold a model with billions of parameters, and it also makes it easy for multiple robots to share one model, or to swap models without touching the robot-side code. The tradeoff is network latency and jitter, so it's usually paired with action chunking and asynchronous inference, letting the robot execute its current action chunk while requesting the next one. openpi serves its policy over WebSocket, and LeRobot also has a gRPC-based asynchronous inference option.","example":"Start openpi's serve_policy script on an RTX 4090 workstation to load π0; the client on the robot arm sends an observation each cycle and gets back a chunk of actions.","related":["Inference Deployment","Asynchronous Inference","Inference Latency","Action Chunking","openpi (Physical Intelligence)","Cloud-Edge-Device Collaboration"]},{"id":"on-device-edge-deployment","category":"software","sec":10,"tier":2,"sources":[{"title":"Edge computing - Wikipedia","url":"https://en.wikipedia.org/wiki/Edge_computing"}],"as_of":"","related_ids":["on-device-model","inference-latency","cloud-edge-device-collaboration","policy-server","post-training-quantization","nvidia-jetson"],"name":"On-Device / Edge Deployment","alt":"端侧部署","abbr":"","aliases":["edge deployment","local deployment","onboard deployment"],"one_liner":"Running model inference directly on a robot's own onboard computer, instead of calling a remote server.","explanation":"On-device deployment means model inference runs on the robot itself or on a device right next to it, typically an embedded board such as Jetson or RK3588. The alternative is streaming images to a cloud or lab GPU server and streaming actions back. Running on-device keeps latency low, works without a network connection, and keeps data from leaving the device — at the cost of limited compute, memory, and power, which often means large models need to be quantized, pruned, or distilled and then compiled with tools like TensorRT or RKNN before they can hit an adequate frame rate. Real systems often combine both approaches: low-level locomotion control on-device, large-model inference in the cloud or on an edge server, an arrangement known as cloud-edge-device collaboration.","example":"Gemini Robotics On-Device is a version of the VLA specifically optimized to run locally on a robot.","related":["On-device Model","Inference Latency","Cloud-Edge-Device Collaboration","Policy Server (Remote Inference)","Post-Training Quantization","NVIDIA Jetson"]},{"id":"cloud-edge-device-collaboration","category":"software","sec":10,"tier":3,"sources":[{"title":"Edge computing - Wikipedia","url":"https://en.wikipedia.org/wiki/Edge_computing"},{"title":"openpi (Physical Intelligence) - GitHub","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["on-device-edge-deployment","policy-server","inference-latency","asynchronous-inference","on-device-model","onboard-compute-platform"],"name":"Cloud-Edge-Device Collaboration","alt":"云边端协同","abbr":"","aliases":["cloud-edge collaboration"],"one_liner":"Splitting a robot's compute across the cloud, an edge server, and the robot itself, with each layer handling a different job.","explanation":"This is an approach to splitting compute across layers: the cloud handles training and any large-model inference that needs a lot of GPU memory, an edge server (a local machine in the same building or facility) handles low-latency vision and planning, and the robot's own onboard controller runs only the control loops that absolutely must be real time. This split exists because a large embodied-AI model's compute needs conflict with a robot body's constraints on power, heat, and cost: running everything onboard is often too slow, but running everything in the cloud means a network hiccup can cause a loss of control. A common pattern is a remote policy server for inference, with locomotion control on the robot itself as a fallback — directly related to on-device deployment, inference latency, and asynchronous inference.","example":"openpi, the open-source implementation of π0, provides both a policy server and a client: the policy runs on a GPU-equipped machine, while the robot side only sends observations and receives actions.","related":["On-Device / Edge Deployment","Policy Server (Remote Inference)","Inference Latency","Asynchronous Inference","On-device Model","Onboard Compute Platform"]},{"id":"operator-kernel","category":"software","sec":10,"tier":3,"sources":[{"title":"PyTorch 文档：Custom C++ and CUDA Operators","url":"https://docs.pytorch.org/tutorials/advanced/cpp_custom_ops.html"}],"as_of":"","related_ids":["cuda","flashattention","nvidia-tensorrt","torch-compile","compute-architecture-for-neural-networks","cuda-graphs"],"name":"Operator / Kernel","alt":"算子","abbr":"","aliases":["kernel","CUDA kernel"],"one_liner":"The basic computational building block of a neural network — like matmul or softmax — and its hardware implementation.","explanation":"An operator is the smallest unit of computation in a deep learning framework — for example, matrix multiplication, convolution, normalization, or softmax. A model's forward pass is a computation graph made of operators wired together. A “kernel” is the actual implementation code for a given operator on a specific piece of hardware, such as a CUDA kernel that runs on a GPU. How fast a model runs at inference depends heavily on how well these kernels are written, and on whether several small operators can be “fused” into a single kernel to cut down on memory reads and writes. This fusing, and picking the fastest kernel, is largely what TensorRT and torch.compile do; when engineers talk about “operator support” while porting a model to a Chinese domestic chip, they mean whether that hardware has a matching implementation for each operator the model uses.","example":"FlashAttention fuses several steps inside attention — matrix multiply, softmax, then multiplying by the value matrix — into a single CUDA kernel, substantially cutting memory access and speeding up inference on long sequences.","related":["CUDA","FlashAttention","NVIDIA TensorRT","torch.compile","Compute Architecture for Neural Networks (Huawei Ascend, CANN)","CUDA Graphs"]},{"id":"cuda-deep-neural-network-library","category":"software","sec":10,"tier":3,"sources":[{"title":"NVIDIA cuDNN","url":"https://developer.nvidia.com/cudnn"},{"title":"NVIDIA cuDNN Documentation","url":"https://docs.nvidia.com/deeplearning/cudnn/"}],"as_of":"","related_ids":["cuda","pytorch","nvidia-tensorrt","operator-kernel","gpu-memory","mixed-precision-training"],"name":"CUDA Deep Neural Network Library (cuDNN)","alt":"cuDNN","abbr":"cuDNN","aliases":["cuDNN"],"one_liner":"NVIDIA's GPU-accelerated library of the operators most commonly used in deep learning.","explanation":"cuDNN is NVIDIA's acceleration library for deep-learning primitives, providing highly optimized implementations, across GPU generations, of common operators (an operator being a basic unit of computation in a neural network) such as convolution, pooling, normalization, and attention. Frameworks like PyTorch and TensorFlow don't write these low-level kernels themselves — they call into cuDNN — so whether it's installed correctly, and whether its version matches the CUDA toolkit and driver, directly determines whether training runs at all and how fast it goes. It isn't a standalone tool but a library installed alongside the CUDA ecosystem; in practice, the most common issues are mismatched environment versions, and the first few batches running slower while cuDNN auto-tunes to pick the fastest convolution algorithm.","example":"Turning on torch.backends.cudnn.benchmark in PyTorch lets cuDNN search for the fastest convolution implementation for a fixed input size.","related":["CUDA","PyTorch","NVIDIA TensorRT","Operator / Kernel","GPU Memory (VRAM)","Mixed-Precision Training"]},{"id":"flashattention","category":"software","sec":10,"tier":3,"sources":[{"title":"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness","url":"https://arxiv.org/abs/2205.14135"},{"title":"Dao-AILab/flash-attention (GitHub)","url":"https://github.com/Dao-AILab/flash-attention"}],"as_of":"","related_ids":["attention-mechanism","transformer","operator-kernel","cuda","inference-latency","mixed-precision-training"],"name":"FlashAttention","alt":"FlashAttention","abbr":"","aliases":["FlashAttention-2","FlashAttention-3"],"one_liner":"A GPU implementation of attention that reorders computation around memory hierarchy, giving identical results faster and with less memory.","explanation":"FlashAttention was introduced by Tri Dao and others in 2022 as a GPU implementation of the attention mechanism that produces results numerically identical to standard attention — it's exact, not an approximation. A standard implementation has to write the full N×N attention matrix to GPU memory and read it back, and for long sequences that memory traffic becomes the bottleneck. FlashAttention instead splits Q, K, and V into small blocks, computes them inside the GPU's fast on-chip cache, and writes back only the final result, never materializing the full attention matrix — so memory use grows linearly with sequence length, and speed improves noticeably too. It was followed by FlashAttention-2 and, for the Hopper architecture, FlashAttention-3. Long-sequence models like VLAs and video world models rely on it heavily for both training and inference, and PyTorch's scaled_dot_product_attention has a similar backend built in.","example":"Setting attn_implementation=“flash_attention_2” in Hugging Face Transformers when training a VLA allows a larger batch size or more image tokens within the same amount of GPU memory.","related":["Attention Mechanism","Transformer","Operator / Kernel","CUDA","Inference Latency","Mixed-Precision Training"]},{"id":"torch-compile","category":"software","sec":10,"tier":3,"sources":[{"title":"torch.compile — PyTorch documentation","url":"https://pytorch.org/docs/stable/generated/torch.compile.html"},{"title":"PyTorch 2.x overview","url":"https://pytorch.org/get-started/pytorch-2-x/"}],"as_of":"","related_ids":["pytorch","inference-latency","cuda-graphs","operator-kernel","nvidia-tensorrt","inference-deployment"],"name":"torch.compile","alt":"torch.compile","abbr":"","aliases":["PyTorch 2 compilation"],"one_liner":"The one-line model-compilation acceleration API that PyTorch has offered since version 2.0.","explanation":"torch.compile is a compilation feature introduced in PyTorch 2.0 (2023). Normally, PyTorch executes one operator at a time, eagerly — every step carries Python dispatch overhead, and there's no way to optimize across operators. torch.compile uses TorchDynamo to capture the computation graph inside a piece of Python code at runtime, then hands it to the default backend, TorchInductor, which fuses operators together and generates Triton or C++ kernels (a kernel being the low-level function that actually executes on the GPU or CPU). Using it usually just means wrapping the model in one call; the first call takes time to compile, and training or inference is faster after that. In embodied AI it's commonly used to cut per-step inference latency for large models like VLAs, and can be paired with CUDA Graphs to further reduce kernel-launch overhead.","example":"When deploying a diffusion policy, write policy = torch.compile(policy, mode=“max-autotune”); after a few warm-up calls, the time spent per denoising step drops.","related":["PyTorch","Inference Latency","CUDA Graphs","Operator / Kernel","NVIDIA TensorRT","Inference Deployment"]},{"id":"cuda-graphs","category":"software","sec":10,"tier":3,"sources":[{"title":"Getting Started with CUDA Graphs (NVIDIA Technical Blog)","url":"https://developer.nvidia.com/blog/cuda-graphs/"},{"title":"CUDA semantics - PyTorch Documentation","url":"https://pytorch.org/docs/stable/notes/cuda.html"}],"as_of":"","related_ids":["cuda","inference-latency","torch-compile","nvidia-tensorrt","operator-kernel","policy-inference-frequency"],"name":"CUDA Graphs","alt":"CUDA Graph","abbr":"","aliases":["CUDA graph"],"one_liner":"Recording a sequence of GPU operations as a single graph, so it can be replayed in one submission with less launch overhead.","explanation":"CUDA Graphs is an execution model in CUDA: a sequence of kernel launches, memory copies, and their dependencies within a computation are first recorded as a graph, and afterward only that graph needs to be submitted each time, letting the driver schedule the whole thing in one go. It addresses CPU-side launch overhead — when a model has many small operators, the CPU time spent launching each kernel individually can exceed the time the GPU actually spends computing, becoming the bottleneck. This is especially useful for robot-policy inference: vision-language-action models and diffusion policies repeatedly run a forward pass with a fixed structure within tens of milliseconds, with unchanging tensor shapes, which is exactly the situation graph capture is suited for. PyTorch exposes this capability through graph capture and torch.compile's low-overhead mode.","example":"Capture a diffusion action head's multi-step denoising loop as a CUDA Graph, reducing kernel-launch overhead within a single frame of inference.","related":["CUDA","Inference Latency","torch.compile","NVIDIA TensorRT","Operator / Kernel","Policy Inference Frequency"]},{"id":"torchscript-libtorch","category":"software","sec":10,"tier":3,"sources":[{"title":"TorchScript — PyTorch documentation","url":"https://pytorch.org/docs/stable/jit.html"},{"title":"Loading a TorchScript Model in C++","url":"https://pytorch.org/tutorials/advanced/cpp_export.html"}],"as_of":"","related_ids":["pytorch","open-neural-network-exchange","inference-deployment","legged-gym","on-device-edge-deployment","python-and-c-plus-plus"],"name":"TorchScript (torch.jit) / LibTorch","alt":"TorchScript / LibTorch（策略导出与 C++ 部署）","abbr":"","aliases":["torch.jit","JIT export","LibTorch"],"one_liner":"Exports a PyTorch model to a file that runs without Python, then loads it from C++.","explanation":"TorchScript is PyTorch's built-in model serialization format: torch.jit.trace (records the computation by running one example input through the model) or torch.jit.script (parses the code directly) converts a model into a .pt file that stores both its structure and its weights. LibTorch is the C++ distribution of PyTorch, letting a C++ program load that file with torch::jit::load and run inference. Robot control software is mostly written in C++ and needs a stable real-time loop, and this path lets a trained reinforcement learning policy run on the robot without depending on a Python environment. PyTorch officially now considers TorchScript to be in maintenance mode and recommends newer options such as torch.export or ONNX for new projects, but a large amount of open-source locomotion-control code still uses it.","example":"legged_gym's play script can export a trained walking policy as a JIT-format policy.pt; the deployment code loads it inside a C++ control loop, feeding in observations and reading out joint targets every control cycle.","related":["PyTorch","Open Neural Network Exchange (ONNX)","Inference Deployment","legged_gym","On-Device / Edge Deployment","Python and C++"]},{"id":"open-neural-network-exchange","category":"software","sec":10,"tier":2,"sources":[{"title":"ONNX | Home","url":"https://onnx.ai/"}],"as_of":"","related_ids":["onnx-runtime","nvidia-tensorrt","intel-openvino","rockchip-rknn-toolkit","inference-deployment","pytorch"],"name":"Open Neural Network Exchange (ONNX)","alt":"ONNX","abbr":"ONNX","aliases":["ONNX"],"one_liner":"An open, general-purpose file format for neural networks that makes it easy to move models between frameworks and inference engines.","explanation":"ONNX is an open model-representation format launched in 2017 by Microsoft and Facebook, now hosted by LF AI & Data under the Linux Foundation. It stores a network's architecture and weights as a unified computation graph, and often serves as the bridge when a model is trained in PyTorch but deployed with TensorRT, ONNX Runtime, OpenVINO, or RKNN. The most common headache during deployment is an unsupported operator or a mismatched export version (opset), which usually requires rewriting part of the model or swapping operators. For newcomers, it's enough to think of it as the “universal file format” of the model world.","example":"Use torch.onnx.export to export a trained walking policy as policy.onnx, then hand it to ONNX Runtime for inference on the robot.","related":["ONNX Runtime","NVIDIA TensorRT","Intel OpenVINO","Rockchip RKNN-Toolkit","Inference Deployment","PyTorch"]},{"id":"onnx-runtime","category":"software","sec":10,"tier":3,"sources":[{"title":"ONNX Runtime 官网","url":"https://onnxruntime.ai/"},{"title":"microsoft/onnxruntime (GitHub)","url":"https://github.com/microsoft/onnxruntime"}],"as_of":"","related_ids":["open-neural-network-exchange","nvidia-tensorrt","intel-openvino","inference-deployment","on-device-edge-deployment","rl-based-locomotion-control"],"name":"ONNX Runtime","alt":"ONNX Runtime","abbr":"ORT","aliases":["ORT","onnxruntime"],"one_liner":"Microsoft's open-source, cross-platform engine for running models saved in the ONNX format.","explanation":"ONNX Runtime is an open-source inference engine from Microsoft that executes models in the ONNX format, a general-purpose file format for exchanging neural networks between frameworks. After a model is trained in a framework like PyTorch, it can be exported to ONNX and then run with ONNX Runtime on Windows, Linux, phones, and embedded devices, without needing the original training framework installed. It connects to different hardware backends through “execution providers” — CPU, CUDA, TensorRT, OpenVINO, CoreML, and others — so the same model can use the right hardware acceleration just by changing a setting. In robotics, control policies trained with reinforcement learning are often deployed as ONNX files and run in real time on an onboard computer using ONNX Runtime.","example":"Export a quadruped walking policy trained in Isaac Lab to policy.onnx, then run it at 50 Hz on the robot dog's onboard CPU with ONNX Runtime.","related":["Open Neural Network Exchange (ONNX)","NVIDIA TensorRT","Intel OpenVINO","Inference Deployment","On-Device / Edge Deployment","RL-based Locomotion Control"]},{"id":"nvidia-tensorrt","category":"software","sec":10,"tier":2,"sources":[{"title":"NVIDIA TensorRT","url":"https://developer.nvidia.com/tensorrt"}],"as_of":"2025-09","related_ids":["open-neural-network-exchange","nvidia-tensorrt-llm","inference-latency","post-training-quantization","nvidia-jetpack-sdk","inference-deployment"],"name":"NVIDIA TensorRT","alt":"TensorRT","abbr":"TRT","aliases":["TRT"],"one_liner":"NVIDIA's inference-acceleration library that compiles a trained model into a faster GPU runtime engine.","explanation":"TensorRT is NVIDIA's deep-learning inference optimizer and runtime. It reads in a model in a format such as ONNX, performs layer fusion, picks the fastest available GPU kernels, and lowers numerical precision (FP16, INT8, FP8, and so on) to produce an inference “engine” file compiled for a specific GPU. A robot policy has to produce an action within tens of milliseconds, and running it directly in PyTorch is often too slow, so compiling with TensorRT typically cuts latency noticeably — making it a standard step for both Jetson and server-side deployment. Note that an engine is tied to a specific GPU model and TensorRT version, so switching hardware requires recompiling; large language models have their own dedicated variant, TensorRT-LLM.","example":"Export a trained diffusion policy to ONNX, then compile it into an FP16 engine with trtexec to cut per-step inference latency on a Jetson Orin.","related":["Open Neural Network Exchange (ONNX)","NVIDIA TensorRT-LLM","Inference Latency","Post-Training Quantization","NVIDIA JetPack SDK","Inference Deployment"]},{"id":"nvidia-jetpack-sdk","category":"software","sec":10,"tier":2,"sources":[{"title":"NVIDIA JetPack SDK","url":"https://developer.nvidia.com/embedded/jetpack"}],"as_of":"2025-09","related_ids":["nvidia-jetson","nvidia-jetson-orin","nvidia-jetson-thor","nvidia-tensorrt","cuda","on-device-edge-deployment"],"name":"NVIDIA JetPack SDK","alt":"JetPack","abbr":"","aliases":["JetPack SDK"],"one_liner":"NVIDIA's official operating system and development kit for its Jetson embedded boards.","explanation":"JetPack is NVIDIA's software development kit for the Jetson family of embedded computing boards, bundling Jetson Linux (an Ubuntu-based OS and drivers, also called L4T) along with acceleration libraries such as CUDA, cuDNN, and TensorRT. Flashing a Jetson essentially means installing a particular version of JetPack, and that version determines which CUDA and TensorRT releases are available on the board, which in turn determines whether — and how fast — a given model can run. Different hardware generations map to different major versions: Orin boards commonly run JetPack 6, while Jetson Thor pairs with JetPack 7. Checking JetPack version compatibility against a target framework or model is usually the first step before on-device deployment.","example":"Before deploying a VLA on a Jetson Orin, flash it with JetPack 6 using SDK Manager, then install matching versions of PyTorch and TensorRT.","related":["NVIDIA Jetson","NVIDIA Jetson Orin","NVIDIA Jetson Thor","NVIDIA TensorRT","CUDA","On-Device / Edge Deployment"]},{"id":"nvidia-tensorrt-edge-llm","category":"software","sec":10,"tier":3,"sources":[{"title":"NVIDIA/TensorRT-Edge-LLM (GitHub)","url":"https://github.com/NVIDIA/TensorRT-Edge-LLM"}],"as_of":"2026-01","related_ids":["nvidia-tensorrt","nvidia-tensorrt-llm","nvidia-jetson-thor","on-device-edge-deployment","on-device-model","post-training-quantization"],"name":"NVIDIA TensorRT Edge-LLM","alt":"TensorRT Edge-LLM","abbr":"","aliases":["Edge-LLM"],"one_liner":"NVIDIA's on-device inference framework for running large language and vision-language models on automotive and robotics chips.","explanation":"TensorRT Edge-LLM is an open-source C++ inference framework from NVIDIA, built to run large language models and vision-language models on embedded platforms such as Jetson Thor and DRIVE AGX Thor. NVIDIA's data-center library for this, TensorRT-LLM, depends on a Python environment and targets high throughput across many GPUs; automotive and robotics deployments instead need low per-request latency, a small memory footprint, and lightweight deployment without Python, which is what Edge-LLM is built for. It reportedly supports low-precision quantization formats such as FP8 and NVFP4, along with speculative decoding (predicting several tokens ahead to speed up generation). For embodied AI, it is one option for running a VLA's or VLM's “brain” directly on the robot itself, rather than on a remote server.","example":"Quantize a multi-billion-parameter VLM and deploy it on a Jetson Thor with TensorRT Edge-LLM, so the robot can do scene understanding and task planning locally.","related":["NVIDIA TensorRT","NVIDIA TensorRT-LLM","NVIDIA Jetson Thor","On-Device / Edge Deployment","On-device Model","Post-Training Quantization"]},{"id":"nvidia-tensorrt-llm","category":"software","sec":10,"tier":3,"sources":[{"title":"NVIDIA/TensorRT-LLM (GitHub)","url":"https://github.com/NVIDIA/TensorRT-LLM"}],"as_of":"2025","related_ids":["nvidia-tensorrt","vllm","key-value-cache","speculative-decoding","post-training-quantization","inference-deployment"],"name":"NVIDIA TensorRT-LLM","alt":"TensorRT-LLM","abbr":"","aliases":["TRT-LLM"],"one_liner":"NVIDIA's open-source library that speeds up large language model inference on NVIDIA GPUs.","explanation":"TensorRT-LLM is an open-source large language model inference library that NVIDIA released in 2023, built on top of TensorRT with a Python interface. It packages together the optimizations commonly used for LLM inference: paged management of the KV cache (the stored attention keys and values from previous steps, kept so they don't need recomputing), dynamic batching, quantization formats like FP8 and INT4, speculative decoding, multi-GPU tensor parallelism, and hand-written high-performance kernels for operations such as attention. Users load model weights in a format such as Hugging Face's, and get back an inference service that runs much faster than native PyTorch. In embodied AI, it is often used to accelerate the language-model backbone of a VLA, or a cloud-hosted “brain” service that a robot calls over the network.","example":"Deploy a Llama model at FP8 precision on an H100 GPU using TensorRT-LLM, as a cloud service for robot task planning.","related":["NVIDIA TensorRT","vLLM","Key-Value Cache","Speculative Decoding","Post-Training Quantization","Inference Deployment"]},{"id":"vllm","category":"software","sec":10,"tier":3,"sources":[{"title":"vllm-project/vllm GitHub","url":"https://github.com/vllm-project/vllm"},{"title":"Efficient Memory Management for Large Language Model Serving with PagedAttention (arXiv)","url":"https://arxiv.org/abs/2309.06180"}],"as_of":"","related_ids":["key-value-cache","large-language-model","vision-language-model","nvidia-tensorrt-llm","verl","inference-deployment"],"name":"vLLM","alt":"vLLM","abbr":"","aliases":[],"one_liner":"An open-source, high-throughput inference and serving engine for large language models.","explanation":"vLLM originated at UC Berkeley's Sky Computing Lab; its core technique, PagedAttention, was published at SOSP 2023. During LLM inference, every request has to keep a KV cache (the attention keys and values for content already generated), and its length is unpredictable; the traditional approach reserves memory for the maximum possible length up front, which wastes a great deal of it. PagedAttention borrows the idea of paging from operating systems, allocating the KV cache in small blocks on demand, and combines this with continuous batching (new requests can be inserted into an already-running batch at any time), so a single GPU can serve far more concurrent requests. It supports mainstream open-source LLMs and VLMs and can start an OpenAI-compatible API with one command. In embodied AI it's often used to serve the large model that does task planning, and is also used by frameworks such as veRL as the generation engine for reinforcement learning.","example":"Run vllm serve on a server to load a Qwen2.5-VL model; the robot sends an image and an instruction over HTTP and gets back a list of decomposed subtasks.","related":["Key-Value Cache","Large Language Model","Vision-Language Model","NVIDIA TensorRT-LLM","veRL (Volcano Engine Reinforcement Learning)","Inference Deployment"]},{"id":"intel-openvino","category":"software","sec":10,"tier":3,"sources":[{"title":"OpenVINO 官方文档","url":"https://docs.openvino.ai/"},{"title":"openvinotoolkit/openvino (GitHub)","url":"https://github.com/openvinotoolkit/openvino"}],"as_of":"","related_ids":[null,null,"nvidia-tensorrt","open-neural-network-exchange",null,null],"name":"Intel OpenVINO","alt":"OpenVINO","abbr":"","aliases":["OpenVINO Toolkit"],"one_liner":"Intel's open-source inference-optimization and deployment toolkit for speeding up models on Intel CPUs, integrated graphics, and NPUs.","explanation":"OpenVINO (Open Visual Inference and Neural network Optimization) is Intel's open-source deep-learning inference toolkit. It converts a model from PyTorch, ONNX, TensorFlow, or another format into its own intermediate representation, then applies operator fusion and low-precision quantization (such as INT8) tuned specifically for Intel CPUs, integrated graphics, and NPUs (neural processing units) before running it. It occupies a similar niche to NVIDIA's TensorRT, just aimed at Intel hardware instead. Many robots run on an industrial PC or mini PC with no discrete GPU, and deploying vision detection or segmentation models with OpenVINO in that situation can raise frame rate and cut latency without adding a GPU.","example":"Export a YOLO detection model to OpenVINO format and run real-time tabletop object detection on the CPU of an Intel Core mini PC mounted on the robot.","related":["Inference Deployment","On-Device / Edge Deployment","NVIDIA TensorRT","Open Neural Network Exchange (ONNX)","Post-Training Quantization (PTQ)","Industrial PC"]},{"id":"rockchip-rknn-toolkit","category":"software","sec":10,"tier":3,"sources":[{"title":"airockchip/rknn-toolkit2 on GitHub","url":"https://github.com/airockchip/rknn-toolkit2"}],"as_of":"","related_ids":["rockchip-rk3588","on-device-edge-deployment","post-training-quantization","open-neural-network-exchange","neural-processing-unit","inference-deployment"],"name":"Rockchip RKNN-Toolkit","alt":"RKNN-Toolkit","abbr":"RKNN","aliases":["RKNN","RKNN-Toolkit2","RKNN-Toolkit-Lite2"],"one_liner":"Rockchip's model-conversion toolchain that turns neural networks into a format its NPUs can run.","explanation":"RKNN-Toolkit is the development toolchain Rockchip provides for the NPU (neural processing unit) built into its own chips. It converts models trained in frameworks such as PyTorch, ONNX, or TensorFlow into the RKNN format, optionally applying INT8 quantization along the way, so they can then run on the NPU of chips such as the RK3588. The current generation of chips uses RKNN-Toolkit2; on-device Python inference uses RKNN-Toolkit-Lite2; a C interface is provided by the RKNPU runtime. Many low-cost robots and development boards use the RK3588 as their main computer, and running detection, segmentation, or small policy networks on it generally goes through this toolchain.","example":"Export a trained YOLO model to ONNX, quantize and convert it to a .rknn file with RKNN-Toolkit2, then deploy it on an RK3588 development board for real-time object detection.","related":["Rockchip RK3588","On-Device / Edge Deployment","Post-Training Quantization","Open Neural Network Exchange (ONNX)","Neural Processing Unit (NPU)","Inference Deployment"]},{"id":"horizon-robotics-openexplorer","category":"software","sec":10,"tier":3,"sources":[{"title":"地瓜机器人开发者社区","url":"https://developer.d-robotics.cc/"}],"as_of":"","related_ids":["horizon-robotics","d-robotics-rdk-s100","post-training-quantization","neural-processing-unit","on-device-edge-deployment","open-neural-network-exchange"],"name":"Horizon Robotics OpenExplorer","alt":"地平线天工开物","abbr":"","aliases":["OpenExplorer"],"one_liner":"Horizon Robotics' toolchain for converting, quantizing, and deploying AI models on its own BPU chips.","explanation":"OpenExplorer (天工开物) is an AI development platform and toolchain from Horizon Robotics, built to serve its chips that carry a BPU (Horizon's own neural-network processor). A trained PyTorch or ONNX model can't run directly on this kind of chip — it first needs to be quantized (converting floating-point weights to a lower-precision format like INT8), compiled into chip instructions, and then loaded through a runtime library — and OpenExplorer provides the tools for these steps, along with example models and performance-profiling tools. Embodied-AI developers deploying vision or policy models on D-Robotics' RDK-series development boards commonly end up using this toolchain.","example":"Export a YOLO detection model to ONNX, post-training-quantize and compile it with OpenExplorer's tools, then deploy it to an RDK development board for real-time detection.","related":["Horizon Robotics","D-Robotics RDK S100","Post-Training Quantization","Neural Processing Unit (NPU)","On-Device / Edge Deployment","Open Neural Network Exchange (ONNX)"]},{"id":"compute-architecture-for-neural-networks","category":"software","sec":10,"tier":3,"sources":[{"title":"昇腾 CANN 异构计算架构（华为昇腾社区）","url":"https://www.hiascend.com/software/cann"}],"as_of":"","related_ids":["operator-kernel","mindspore","nvidia-tensorrt","inference-deployment","on-device-edge-deployment","huawei"],"name":"Compute Architecture for Neural Networks (Huawei Ascend, CANN)","alt":"昇腾 CANN","abbr":"CANN","aliases":["CANN"],"one_liner":"The low-level compute architecture and operator compilation toolchain for Huawei's Ascend NPUs.","explanation":"CANN, short for Compute Architecture for Neural Networks, is Huawei's heterogeneous computing software stack for its Ascend AI processors, occupying a role similar to CUDA plus cuDNN on the NVIDIA side: it interfaces with Ascend chips underneath and supports frameworks like MindSpore and PyTorch above. It includes the AscendCL programming interface, a graph compilation and optimization engine, accelerated operator libraries, and Ascend C for writing custom operators. For embodied AI, it matters as the deployment foundation for the domestic Chinese compute route: to run a model on an Ascend board or server, it has to be converted, compiled, and tuned through CANN, occupying a role analogous to TensorRT for NVIDIA or RKNN-Toolkit for Rockchip.","example":"After exporting a vision model trained in PyTorch, convert it into an Ascend offline model with the CANN toolchain to run inference on an Ascend board.","related":["Operator / Kernel","MindSpore","NVIDIA TensorRT","Inference Deployment","On-Device / Edge Deployment","Huawei"]},{"id":"nvidia-isaac-platform","category":"software","sec":11,"tier":2,"sources":[{"title":"NVIDIA Isaac - AI Robot Development Platform","url":"https://developer.nvidia.com/isaac"}],"as_of":"2025-12","related_ids":["nvidia-isaac-sim","nvidia-isaac-lab","nvidia-isaac-ros","nvidia-isaac-gr00t-n1","nvidia-omniverse","nvidia-three-computer-solution"],"name":"NVIDIA Isaac Platform","alt":"NVIDIA Isaac","abbr":"","aliases":["Isaac Platform"],"one_liner":"NVIDIA's umbrella brand for its full stack of robotics software, models, and tools.","explanation":"NVIDIA Isaac is NVIDIA's platform brand for robotics development — not a single piece of software but a family of products: Isaac Sim (a simulator), Isaac Lab (a robot learning framework built on Isaac Sim), Isaac ROS (GPU-accelerated perception and planning packages that run on ROS 2), and Isaac GR00T (foundation models and tools for humanoid robots), among others. It ties together the pipeline of training in simulation, generating synthetic data, and deploying to Jetson boards using NVIDIA's own hardware and software, forming the software layer of NVIDIA's “three computers” narrative. When newcomers see a name with the Isaac prefix, the first thing to figure out is whether it's a simulator, a training framework, a ROS package, or a model.","example":"Train a quadruped locomotion policy in parallel on GPUs with Isaac Lab, then deploy it on a Jetson Orin using Isaac ROS's perception packages.","related":["NVIDIA Isaac Sim","NVIDIA Isaac Lab","NVIDIA Isaac ROS","NVIDIA Isaac GR00T N1","NVIDIA Omniverse","NVIDIA Three-Computer Solution"]},{"id":"nvidia-omniverse","category":"software","sec":11,"tier":2,"sources":[{"title":"NVIDIA Omniverse","url":"https://www.nvidia.com/en-us/omniverse/"}],"as_of":"2025-09","related_ids":["nvidia-isaac-sim","universal-scene-description","physx","digital-twin","nvidia-omniverse-replicator","nvidia-isaac-platform"],"name":"NVIDIA Omniverse","alt":"Omniverse","abbr":"","aliases":[],"one_liner":"NVIDIA's 3D simulation and collaborative development platform, built around OpenUSD.","explanation":"Omniverse is NVIDIA's platform for building 3D applications: it represents scenes using OpenUSD (Universal Scene Description), renders them with real-time ray tracing via RTX, and plugs in the PhysX physics engine. It isn't built specifically for robotics — factory digital twins, film, and autonomous-driving simulation all use it too — but the piece embodied-AI practitioners run into most is what's built on top of it: Isaac Sim is constructed on Omniverse, and its Replicator tool is used to generate large volumes of annotated synthetic data. Understanding Omniverse mainly means understanding why Isaac Sim's scene files are USD, and why its rendering depends on an RTX-capable GPU.","example":"Use Omniverse Replicator to randomize lighting and object placement in a warehouse scene, generating large batches of training images with segmentation labels.","related":["NVIDIA Isaac Sim","Universal Scene Description (OpenUSD)","PhysX","Digital Twin","NVIDIA Omniverse Replicator","NVIDIA Isaac Platform"]},{"id":"nvidia-isaac-ros","category":"software","sec":11,"tier":3,"sources":[{"title":"Isaac ROS 文档","url":"https://nvidia-isaac-ros.github.io/"},{"title":"NVIDIA-ISAAC-ROS GitHub","url":"https://github.com/NVIDIA-ISAAC-ROS"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"NVIDIA Isaac ROS","alt":"Isaac ROS","abbr":"","aliases":[],"one_liner":"A collection of GPU-accelerated ROS 2 packages from NVIDIA.","explanation":"Isaac ROS is an open-source collection of ROS 2 software packages from NVIDIA, aimed at GPU-equipped platforms such as Jetson. It rewrites commonly used modules as GPU-accelerated versions — visual SLAM (cuVSLAM), 3D mapping (nvblox), stereo depth, AprilTag recognition, object pose estimation (FoundationPose), and motion planning (cuMotion) among them — and uses a mechanism called NITROS to keep large data like images in GPU memory as it passes between nodes, cutting down on copies. Developers can drop these in as replacements for the corresponding nodes in an existing ROS 2 system to speed up perception and planning, and it's also an important part of the deployment side of NVIDIA's Isaac platform.","example":"Run Isaac ROS Visual SLAM on a Jetson Orin to estimate a mobile robot's pose in real time using a stereo camera.","related":["ROS 2","NVIDIA Isaac Platform","NVIDIA Jetson","NVIDIA nvblox (GPU TSDF/ESDF Mapping)","NVIDIA Isaac ROS cuMotion","Zero-Copy (Shared-Memory IPC)"]},{"id":"nvidia-osmo","category":"software","sec":11,"tier":3,"sources":[{"title":"NVIDIA OSMO 官方页面","url":"https://developer.nvidia.com/osmo"}],"as_of":"2025","related_ids":["nvidia-three-computer-solution","nvidia-isaac-sim","nvidia-isaac-lab","nvidia-cosmos","synthetic-data","hardware-in-the-loop-software-in-the-loop-simulation"],"name":"NVIDIA OSMO","alt":"NVIDIA OSMO","abbr":"","aliases":["OSMO"],"one_liner":"NVIDIA's cloud orchestration platform that schedules robotics and physical-AI training, simulation, and data generation across different compute.","explanation":"OSMO is a cloud-native orchestration platform from NVIDIA for robotics and physical-AI development. A robotics project usually chains together many steps — synthetic data generation, model training, simulation-based evaluation, and software-in-the-loop/hardware-in-the-loop testing — often running on different machines: cloud GPU clusters, local workstations, or edge devices like Jetson. OSMO lets developers describe this pipeline as a configuration file, and the platform handles allocating compute, queuing jobs, and tracking data, removing the need to manually move data and coordinate jobs across machines by hand. It's one piece of NVIDIA's “three computers” strategy, the layer that connects training and simulation, and is commonly used alongside Isaac Sim, Isaac Lab, and Cosmos.","example":"Turn “Isaac Sim generates synthetic data → train a grasping policy → Isaac Lab runs batch evaluation” into a single OSMO workflow, submit it once, and the platform automatically distributes the jobs across different GPU nodes.","related":["NVIDIA Three-Computer Solution","NVIDIA Isaac Sim","NVIDIA Isaac Lab","NVIDIA Cosmos","Synthetic Data","Hardware-in-the-Loop / Software-in-the-Loop Simulation"]},{"id":"ros-industrial","category":"software","sec":11,"tier":3,"sources":[{"title":"ROS-Industrial 官网","url":"https://rosindustrial.org/"},{"title":"ROS-Industrial GitHub","url":"https://github.com/ros-industrial"}],"as_of":"","related_ids":["robot-operating-system","robot-operating-system-2","moveit-motion-planning-framework","industrial-robot","unified-robot-description-format","system-integrator"],"name":"ROS-Industrial","alt":"ROS-Industrial","abbr":"ROS-I","aliases":["ROS-I"],"one_liner":"An open-source project and consortium that brings ROS to industrial robots and manufacturing.","explanation":"ROS-Industrial is an open-source project that extends ROS into industrial automation. It was launched around 2012, led by the Southwest Research Institute (SwRI) in the US, and is now maintained by three regional consortia covering North America, Europe, and the Asia-Pacific, with members including equipment vendors, manufacturers, and research institutions. It provides ROS drivers and description packages for industrial arms from various brands, integration with MoveIt (the ROS motion-planning framework), and tooling for “scan–plan–execute” applications such as sanding and spray painting. Its significance is letting perception and planning algorithms already developed in academia connect to factory-floor hardware, cutting down the work of writing a separate integration for every controller brand.","example":"Use an ABB or FANUC driver package provided by ROS-Industrial to send a trajectory planned by MoveIt to a real industrial robot arm for execution.","related":["Robot Operating System","Robot Operating System 2","MoveIt Motion Planning Framework","Industrial Robot","Unified Robot Description Format","System Integrator"]},{"id":"togetheros-bot","category":"software","sec":11,"tier":3,"sources":[{"title":"D-Robotics 开发者社区","url":"https://developer.d-robotics.cc/"},{"title":"D-Robotics GitHub","url":"https://github.com/D-Robotics"}],"as_of":"","related_ids":["robot-operating-system-2","d-robotics","d-robotics-rdk-s100","horizon-robotics","zero-copy","robot-operating-system"],"name":"TogetheROS.Bot (D-Robotics)","alt":"TogetheROS.Bot","abbr":"TROS","aliases":["TROS","D-Robotics TROS","TROS.B"],"one_liner":"D-Robotics' robot software stack, built on top of and compatible with ROS 2.","explanation":"TogetheROS.Bot is the robot operating stack that D-Robotics — the robotics business spun off from Horizon Robotics — provides for its RDK line of development boards; it was originally released under the Horizon Robotics name. It's an extension of ROS 2, with an interface compatible with it, so existing ROS 2 nodes can run on it directly. What it adds is mainly capability tied to its own chips: inference nodes that call the onboard BPU (Horizon's neural network accelerator), camera and image-processing nodes, communication optimizations that cut down large-image-data copying, and a set of ready-made perception algorithm examples. For developers building a mobile robot or arm on an RDK board, it removes the work of adapting to the hardware accelerator themselves.","example":"After installing TROS on an RDK development board, use the official object-detection example node to read a USB camera feed, run the detection model on the BPU, and publish the results as a ROS 2 topic for other nodes.","related":["Robot Operating System 2","D-Robotics","D-Robotics RDK S100","Horizon Robotics","Zero-Copy (Shared-Memory IPC)","Robot Operating System"]},{"id":"m-robots-os","category":"software","sec":11,"tier":3,"sources":[{"title":"深开鸿官网","url":"https://www.kaihong.com/"},{"title":"OpenHarmony 开源鸿蒙","url":"https://www.openharmony.cn/"}],"as_of":"2026-09","related_ids":[null,null,null,null,null],"name":"M-Robots OS (OpenHarmony-based Robot OS)","alt":"M-Robots OS","abbr":"","aliases":["Kaihong M-Robots OS"],"one_liner":"A Chinese robot operating system built on open-source HarmonyOS, focused on multi-robot collaboration.","explanation":"M-Robots OS is a robot operating system reportedly led by Shenzhen KaiHong Digital Industry Development Co. (深开鸿), built on top of OpenHarmony, the open-source HarmonyOS project incubated by the OpenAtom Foundation. The problem it's trying to solve is that robots and devices from different manufacturers and with different chips each run their own separate system, making it hard for them to connect and collaborate. M-Robots OS draws on HarmonyOS capabilities like its distributed soft bus to let heterogeneous robots discover each other, communicate, and divide up work. It occupies the same layer of the software stack as ROS 2, but its ecosystem and focus lean toward domestic Chinese technology and multi-device collaboration, and it's currently being promoted mainly through universities and industry alliances.","example":"","related":["Robot Operating System","ROS 2","Multi-Robot Collaboration","Middleware","Huawei"]},{"id":"link-u-os","category":"software","sec":11,"tier":3,"sources":[{"title":"智元机器人官网","url":"https://www.agibot.com/"},{"title":"AimRT (GitHub)","url":"https://github.com/AimRT/AimRT"}],"as_of":"2025-07","related_ids":[null,null,null,null,null,null],"name":"Link-U OS (AgiBot Lingqu OS)","alt":"灵渠 OS（智元具身操作系统）","abbr":"","aliases":["AgiBot Lingqu OS","Lingqu OS"],"one_liner":"AgiBot's operating system for embodied robots, with plans to open-source it in stages.","explanation":"Link-U OS (灵渠 OS) is an operating system for embodied robots from AgiBot (智元机器人), reportedly unveiled along with an open-source roadmap at the 2025 World Artificial Intelligence Conference. According to public descriptions, it follows a layered design: the lower layer handles hardware abstraction and real-time communication (related to AimRT, AgiBot's previously open-sourced robot runtime framework), while the upper layer connects to embodied large models, skills, and applications, aiming to let different robot bodies and different models run on one shared software foundation, cutting down the cost of custom development and porting between robot bodies. It occupies a similar space to ROS 2 and RoboOS; the exact scope and timeline of what gets open-sourced is best confirmed against AgiBot's official announcements.","example":"","related":["AgiBot","AimRT (AgiBot robot runtime framework)","Robot Operating System","ROS 2","RoboOS (BAAI embodied OS)","Hardware Abstraction Layer (HAL)"]},{"id":"agibot-genie-studio","category":"software","sec":11,"tier":3,"sources":[{"title":"AgiBot 智元机器人官网","url":"https://www.agibot.com/"}],"as_of":"2025","related_ids":["agibot","agibot-go-1","genie-sim",null,"software-development-kit","inference-deployment"],"name":"AgiBot Genie Studio","alt":"Genie Studio（智元具身开发平台）","abbr":"","aliases":["AgiBot Genie Studio"],"one_liner":"AgiBot's one-stop embodied-AI development platform, covering data collection, training, evaluation, and deployment.","explanation":"Genie Studio is a developer-facing embodied-AI platform from AgiBot (智元机器人). According to the company's own description, its goal is to string together, on one platform, the several steps needed to build a robot skill: teleoperation-based data collection and management, data annotation, training and fine-tuning on top of AgiBot's own foundation models (such as GO-1), simulation-based evaluation, and deploying the trained model onto AgiBot's hardware. It addresses the problem that the embodied-AI development pipeline is long and its tooling scattered — getting from collected data to a working real robot often means stitching together a lot of custom scripts. For a newcomer, it can be understood as a vendor-supplied “embodied-AI IDE plus cloud training service,” aimed mainly at customers and developers using AgiBot hardware. Its exact feature set and how open it is are best confirmed against AgiBot's own announcements.","example":"","related":["AgiBot","AgiBot GO-1","Genie Sim","AimRT (AgiBot robot runtime framework)","Software Development Kit","Inference Deployment"]},{"id":"om1","category":"software","sec":11,"tier":3,"sources":[{"title":"OpenMind/OM1 (GitHub)","url":"https://github.com/OpenMind/OM1"}],"as_of":"2025","related_ids":["robot-operating-system","robot-operating-system-2","large-language-model","eclipse-zenoh","unitree-go2","link-u-os"],"name":"OM1","alt":"OM1（OpenMind 机器人 OS）","abbr":"","aliases":["OpenMind OM1","OpenMind Modular AI Runtime"],"one_liner":"OpenMind's open-source, modular AI runtime that connects large language models to different robot bodies.","explanation":"OM1 is an open-source, modular AI runtime from OpenMind, a U.S. startup. It's often called a “robot operating system,” but it doesn't replace low-level communication layers such as ROS — instead it runs on top of them: it turns sensor input like camera and microphone feeds into text descriptions, hands that to a large language model or vision-language model for decision-making, and converts the model's output back into robot actions or speech. Each piece plugs in as a module, so the underlying model or hardware can be swapped out. According to OpenMind, OM1 supports various robot platforms, including humanoids and quadrupeds such as Unitree's G1 and Go2, and can connect to lower-level layers through protocols like ROS 2 and Zenoh. It reflects one approach to building an LLM-driven interaction layer for robots that works across different hardware.","example":"Run OM1 on a Unitree Go2 with a multimodal large language model plugged in, so the robot dog can understand spoken commands and describe what it sees.","related":["Robot Operating System","Robot Operating System 2","Large Language Model","Eclipse Zenoh","Unitree Go2","Link-U OS (AgiBot Lingqu OS)"]},{"id":"roboos","category":"software","sec":11,"tier":3,"sources":[{"title":"RoboOS on GitHub (FlagOpen)","url":"https://github.com/FlagOpen/RoboOS"}],"as_of":"2025-07","related_ids":["robobrain","beijing-academy-of-artificial-intelligence","braincerebellum-architecture","one-brain-multiple-robots","multi-robot-collaboration","model-context-protocol"],"name":"RoboOS","alt":"RoboOS","abbr":"","aliases":["BAAI RoboOS"],"one_liner":"BAAI's open-source cross-embodiment, multi-robot coordination framework built on a cloud brain, local cerebellum design.","explanation":"RoboOS is an embodied-AI system framework open-sourced in 2025 by the Beijing Academy of Artificial Intelligence (BAAI). Despite the name “OS,” it isn't a low-level operating system like Linux — it's a coordination layer that runs on top of robots. It uses a brain–cerebellum split: a cloud-based “brain,” built on the embodied foundation model RoboBrain, understands instructions, breaks tasks into steps, and assigns them to different robots; each robot's local “cerebellum” calls skills such as grasping or navigation to execute its part; a shared memory in between keeps scene and state information synchronized across robots. The goal is to let robots of different form factors connect to the same brain and cooperate on a task. RoboOS 2.0 reportedly followed later the same year, adding a skill marketplace and support for the MCP protocol.","example":"A user says “hand me the water on the table”: RoboOS's brain splits the task into navigation and grasping, and assigns the two steps to a wheeled base and a dual-arm robot respectively.","related":["RoboBrain","Beijing Academy of Artificial Intelligence","Brain–Cerebellum Architecture","One Brain, Multiple Robots","Multi-robot Collaboration","Model Context Protocol (ROS MCP Server)"]},{"id":"huisi-kaiwu","category":"software","sec":11,"tier":3,"sources":[{"title":"北京人形机器人创新中心官网","url":"https://www.x-humanoid.com/"}],"as_of":"2025","related_ids":["beijing-humanoid-robot-innovation-center","tiangong","pelican-vl","xr-1","one-brain-multiple-robots","huawei-cloud-cloudrobo-embodied-ai-platform"],"name":"Huisi Kaiwu (X-Humanoid general embodied AI platform)","alt":"慧思开物","abbr":"","aliases":[],"one_liner":"A general-purpose embodied-AI software platform from the Beijing Humanoid Robot Innovation Center.","explanation":"Huisi Kaiwu is a general-purpose embodied-AI platform from the Beijing Humanoid Robot Innovation Center (X-Humanoid), positioned as a software system that can be installed across different robots. According to its public materials, it brings together large-model-driven task understanding and planning, a skill library, and low-level motion control, aiming for “one brain, many skills; one brain, many bodies” — the same software able to drive different robot bodies and carry out a variety of tasks, cutting down on the work of building everything from scratch for every scenario. It belongs to the same technical family as the center's humanoid robot Tien Kung and its embodied models Pelican-VL and XR-1.","example":"","related":["Beijing Humanoid Robot Innovation Center","Tiangong","Pelican-VL","XR-1","One Brain, Multiple Robots","Huawei Cloud CloudRobo Embodied AI Platform"]},{"id":"huawei-cloud-cloudrobo-embodied-ai-platform","category":"software","sec":11,"tier":3,"sources":[{"title":"华为云官网","url":"https://www.huaweicloud.com/"}],"as_of":"2025-06","related_ids":["huawei","cloud-edge-device-collaboration","embodied-foundation-model","tencent-tairos-embodied-ai-open-platform","huisi-kaiwu"],"name":"Huawei Cloud CloudRobo Embodied AI Platform","alt":"华为云 CloudRobo","abbr":"","aliases":["CloudRobo"],"one_liner":"Huawei Cloud's embodied-AI cloud platform, providing robots with large models and services hosted in the cloud.","explanation":"CloudRobo is a cloud platform for embodied AI from Huawei Cloud, reportedly unveiled at Huawei's 2025 developer conference. The idea is to move part of a robot's “brain” into the cloud: the platform offers embodied-AI-related models built on Huawei's Pangu large model (reportedly spanning data generation, task planning, and action execution), along with services for data synthesis and simulation-based training, with robots connecting to the cloud through a connection protocol Huawei has proposed. For robot manufacturers, a platform like this lowers the barrier to building their own large model and compute stack; the tradeoff is dependence on network and cloud service availability, and any control loop with tight real-time requirements still has to be handled on-device.","example":"","related":["Huawei","Cloud-Edge-Device Collaboration","Embodied Foundation Model","Tencent Tairos Embodied AI Open Platform","Huisi Kaiwu (X-Humanoid general embodied AI platform)"]},{"id":"tencent-tairos-embodied-ai-open-platform","category":"software","sec":11,"tier":3,"sources":[{"title":"腾讯 Robotics X 实验室","url":"https://robotics.tencent.com/"}],"as_of":"2026-09","related_ids":["tencent-robotics-x-lab","tencent-hy-embodied","robot-brain-company","hardware-software-decoupling","huawei-cloud-cloudrobo-embodied-ai-platform"],"name":"Tencent Tairos Embodied AI Open Platform","alt":"腾讯 Tairos","abbr":"","aliases":["Tairos"],"one_liner":"Tencent's open platform for embodied AI, offering foundation models and developer tools for robots.","explanation":"Tairos is an embodied-AI open platform launched by Tencent, led by its Robotics X lab, with a Chinese nickname (钛螺丝) that translates roughly as “titanium screw”; it reportedly launched in 2025. Tencent's positioning is not to build complete robots itself, but to supply “brain”-layer capabilities to robot makers and developers: perception, planning, and action models, along with development tools, simulation, and data services, which manufacturers can call on a modular, as-needed basis and install on their own robot bodies. It represents one path for a large Chinese internet company entering embodied AI: build the platform and the models, and partner with hardware makers.","example":"","related":["Tencent Robotics X Lab","Tencent HY-Embodied (Hunyuan Embodied)","Robot-Brain (Model-Only) Company","Hardware-Software Decoupling","Huawei Cloud CloudRobo Embodied AI Platform"]},{"id":"simulator","category":"sim","sec":0,"tier":1,"sources":[{"title":"MuJoCo Documentation: Overview","url":"https://mujoco.readthedocs.io/en/stable/overview.html"},{"title":"NVIDIA Isaac Sim Documentation","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/index.html"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"2026-09","related_ids":["physics-engine","rendering","sim-to-real-transfer","mujoco","nvidia-isaac-sim","simulation-based-evaluation"],"name":"Simulator","alt":"仿真器","abbr":"","aliases":["Simulation"],"one_liner":"Software that models in a computer how robots, objects, and sensors move and interact.","explanation":"A simulator is software that recreates a robot and its environment inside a computer. It typically includes a physics engine, which advances the state forward each step by computing forces, collisions, and contact; a renderer, which generates camera images, depth, and other sensor data; scenes and assets, meaning the robot model and various objects; and an interface for algorithms to call, such as resetting the environment, executing one action, and returning observations and reward. In embodied AI, simulators are used to train reinforcement-learning policies, generate synthetic data, and run simulation-based evaluation, and their appeal is being cheap, safe, parallelizable, and reproducible. Common ones include MuJoCo (maintained by Google DeepMind, open-sourced in 2022), NVIDIA's Isaac Sim/Isaac Lab, PyBullet, Genesis, and SAPIEN. The core challenge is always the gap between it and reality.","example":"The ACT paper built two simulated bimanual tasks in MuJoCo, “transfer cube” and “bimanual insertion,” so others could reproduce the experiments without buying ALOHA hardware.","related":["Physics Engine","Rendering","Sim-to-Real Transfer","MuJoCo (Multi-Joint dynamics with Contact)","NVIDIA Isaac Sim","Simulation-Based Evaluation"]},{"id":"physics-engine","category":"sim","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Physics engine","url":"https://en.wikipedia.org/wiki/Physics_engine"},{"title":"google-deepmind/mujoco (GitHub)","url":"https://github.com/google-deepmind/mujoco"}],"as_of":"","related_ids":["simulator","mujoco","physx","contact-model","constraint-solver","simulation-timestep"],"name":"Physics Engine","alt":"物理引擎","abbr":"","aliases":[],"one_liner":"The software core that computes forces, collisions, and motion step by step according to physical laws.","explanation":"A physics engine is software that approximately simulates a physical system, first used heavily in games and film. On every simulation step, it first runs collision detection to find which objects are touching, then a constraint solver computes contact forces, friction, and joint forces, and finally an integrator updates positions and velocities. Engines can be grouped by their goal into real-time engines, which favor speed, and high-accuracy engines, which favor precision. Common ones in robotics include MuJoCo, PhysX, Bullet (PyBullet), ODE, and the newer Newton. It's worth distinguishing from a simulator: a physics engine only handles how things move, while a simulator such as Isaac Sim or Gazebo wraps it with rendering, sensors, scene editing, and a programming interface. The contact model and step size it uses directly determine how accurate and stable the simulation is, and are one of the main sources of the sim-to-real gap.","example":"The same robot-arm grasping task, run in MuJoCo versus PhysX, can produce different contact forces and success rates — which is why papers usually state which engine they used.","related":["Simulator","MuJoCo (Multi-Joint dynamics with Contact)","PhysX","Contact Model","Constraint Solver","Simulation Timestep"]},{"id":"environment","category":"sim","sec":0,"tier":1,"sources":[{"title":"Gymnasium Documentation: Env API","url":"https://gymnasium.farama.org/api/env/"}],"as_of":"","related_ids":["agentenvironment-interaction","episode","observation","action-space","termination-vs-truncation","gymnasium"],"name":"Environment (Env; reset/step interface)","alt":"环境（Env）与 reset / step 接口","abbr":"Env","aliases":["Gym interface","Gymnasium API","Env"],"one_liner":"The standard reinforcement-learning interface: reset starts an episode, step executes one action.","explanation":"In reinforcement learning, the environment is the world an agent interacts with. Originating with OpenAI Gym and now maintained by the Farama Foundation as Gymnasium, it's standardized into a programming interface: reset() starts a new episode and returns the initial observation plus an info dictionary; step(action) executes one action and returns the new observation, the reward, terminated (the episode ended in a meaningful sense, such as success or falling over), truncated (the episode was cut off externally, such as by a timeout), and info. The environment also declares its action_space and observation_space, fixing the format of actions and observations. Frameworks such as Isaac Lab follow a similar convention, which is what lets the same training algorithm be pointed at different environments. The older Gym API had a single done flag; it was split into two because, on a timeout truncation, the value estimate still needs to bootstrap from the next state, and conflating the two cases leads to incorrect updates.","example":"A typical loop: obs, info = env.reset(), then repeatedly obs, reward, terminated, truncated, info = env.step(policy(obs)), resetting whenever either termination flag comes back true.","related":["Agent–Environment Interaction","Episode","Observation","Action Space","Termination vs. Truncation","Gymnasium"]},{"id":"rollout","category":"sim","sec":0,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"RoboChallenge 论文 HTML 全文 (arXiv 2510.17950)","url":"https://arxiv.org/html/2510.17950"}],"as_of":"","related_ids":["trajectory","episode","success-rate","closed-loop-evaluation","learning-in-imagination","on-policy"],"name":"Rollout","alt":"推演","abbr":"","aliases":[],"one_liner":"Running a policy in an environment from start to finish to produce one complete interaction trajectory.","explanation":"A rollout means placing a policy in an environment, whether simulated or real, and running the “observe, act, environment changes” loop from the initial state until the episode succeeds, fails, or times out, producing one complete trajectory. OpenAI's Spinning Up tutorial notes that a trajectory is also commonly called an episode or a rollout. Common usages: in reinforcement learning, rollouts are used to collect the data that updates a policy; in evaluation, “50 rollouts per task” means running 50 episodes and computing the success rate from them; and in model-based reinforcement learning, taking several imagined steps forward inside a world model is also called a rollout. Reinforcement learning for large language models has carried the same term over to mean sampling several responses for the same prompt.","example":"The RoboChallenge Table30 benchmark runs 10 rollouts per task, then scores each model by success rate and progress score.","related":["Trajectory","Episode","Success Rate","Closed-Loop Evaluation","Learning in Imagination","On-Policy"]},{"id":"termination-vs-truncation","category":"sim","sec":0,"tier":3,"sources":[{"title":"Gymnasium: Handling Time Limits","url":"https://gymnasium.farama.org/tutorials/gymnasium_basics/handling_time_limits/"},{"title":"Time Limits in Reinforcement Learning (arXiv 1712.00378, ICML 2018)","url":"https://arxiv.org/abs/1712.00378"},{"title":"rsl_rl PPO 实现（time_outs 自举）","url":"https://raw.githubusercontent.com/leggedrobotics/rsl_rl/master/rsl_rl/algorithms/ppo.py"}],"as_of":"","related_ids":["episode","early-termination","value-function","bootstrapping","environment","gymnasium"],"name":"Termination vs. Truncation","alt":"终止与截断","abbr":"","aliases":["terminated / truncated","Time-limit Truncation"],"one_liner":"An episode can end two ways: termination means the task itself is over; truncation means it was cut off by a time limit.","explanation":"This is a pair of concepts from the reinforcement-learning environment interface. Termination means the agent has entered a terminal state of the Markov decision process — the task is complete, or the robot has fallen over. Truncation means the episode was cut off for a reason outside the task itself, most commonly a maximum step count set for training. The two affect training differently: after termination there is no future reward, so the value target is just the immediate reward; a truncated state is not actually an endpoint, so its target should still bootstrap using the value estimate of the next state (substituting an estimate for return that hasn't actually been observed yet). Pardo and colleagues' ICML 2018 paper specifically analyzed the problems caused by treating a time limit as termination. Starting with version 0.26, Gymnasium's step() returns terminated and truncated separately, replacing the old single done flag; Isaac Lab similarly uses a time_out flag to distinguish the two.","example":"In legged_gym, an episode terminates when contact force on designated body parts exceeds a threshold; it truncates when the step count exceeds the limit, which is passed to rsl_rl through a time_outs field in extras — PPO adds the discounted value estimate back into the reward at truncated steps, effectively continuing to bootstrap.","related":["Episode","Early Termination","Value Function","Bootstrapping (in Reinforcement Learning)","Environment (Env; reset/step interface)","Gymnasium"]},{"id":"simulation-timestep","category":"sim","sec":0,"tier":2,"sources":[{"title":"MuJoCo Documentation: XML Reference (option timestep)","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html#option-timestep"},{"title":"Isaac Lab source: velocity_env_cfg.py (locomotion velocity task)","url":"https://github.com/isaac-sim/IsaacLab/blob/main/source/isaaclab_tasks/isaaclab_tasks/manager_based/locomotion/velocity/velocity_env_cfg.py"},{"title":"Isaac Lab source: simulation_cfg.py (SimulationCfg)","url":"https://github.com/isaac-sim/IsaacLab/blob/main/source/isaaclab/isaaclab/sim/simulation_cfg.py"}],"as_of":"2026-09","related_ids":["control-decimation","substeps","numerical-integrator","simulation-instability","control-frequency","real-time-factor"],"name":"Simulation Timestep","alt":"仿真步长","abbr":"dt","aliases":["dt","Physics Timestep","Physics dt","Sim dt"],"one_liner":"The amount of simulated time a physics engine advances per step, which trades off accuracy, stability, and speed.","explanation":"The simulation timestep (often written dt) is how much time the physics engine's integrator advances on each step — 0.002 seconds, for instance, means simulating 1 real second takes 500 steps. A smaller timestep computes contact and fast motion more accurately and is less prone to blowing up, but it also needs more steps to cover the same duration, making it slower. MuJoCo defaults to a 0.002-second timestep, and its documentation calls it the single most important parameter for the speed-accuracy trade-off; Isaac Lab's SimulationCfg defaults to 1/60 second, though specific tasks often set it smaller. The simulation timestep is usually shorter than the policy's control period: for every action the policy outputs, the simulation keeps running for several steps, and that multiple is called decimation; some engines further split a single step into several substeps.","example":"Isaac Lab's velocity-tracking task for legged robots sets sim.dt to 0.005 seconds (200 Hz) with a decimation of 4, so the policy outputs actions at 50 Hz.","related":["Control Decimation","Substeps","Numerical Integrator (Semi-implicit Euler / RK4)","Simulation Instability","Control Frequency","Real-Time Factor"]},{"id":"control-decimation","category":"sim","sec":0,"tier":2,"sources":[{"title":"Isaac Lab API: isaaclab.envs（ManagerBasedEnvCfg.decimation）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.envs.html"},{"title":"legged_gym: legged_robot_config.py","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["simulation-timestep","control-frequency","policy-inference-frequency","proportional-derivative-control","substeps","nvidia-isaac-lab"],"name":"Control Decimation","alt":"控制降频","abbr":"","aliases":["decimation","physics-to-control frequency ratio"],"one_liner":"The number of physics-engine steps that run for every single action a policy outputs.","explanation":"A physics engine needs a very small timestep (usually a few milliseconds) to stay numerically stable, but a policy network doesn't need to, and often can't, output actions that fast. Control decimation is the setting that lets a policy output one action while the simulator advances that action (or the PD targets computed from it) forward for N consecutive physics steps; N is called the decimation. Both Isaac Lab and legged_gym use this parameter, and the duration of one environment step equals the physics timestep times the decimation. Choosing it involves a tradeoff: too large, and the policy reacts too slowly; too small, and training compute is wasted, and it may no longer match the real robot's control frequency, widening the sim-to-real gap. When deployed to a real robot, the policy generally runs at that same frequency, while the low-level motor PD control runs at a much higher frequency underneath it.","example":"legged_gym defaults to a 0.005-second physics timestep with decimation = 4, so the policy outputs an action every 0.02 seconds — 50 Hz control — while the joint PD controller executes within 200 Hz physics steps.","related":["Simulation Timestep","Control Frequency","Policy Inference Frequency","Proportional-Derivative Control","Substeps","NVIDIA Isaac Lab"]},{"id":"substeps","category":"sim","sec":0,"tier":3,"sources":[{"title":"Genesis SimOptions 源码（dt / substeps 说明）","url":"https://raw.githubusercontent.com/Genesis-Embodied-AI/Genesis/main/genesis/options/solvers.py"},{"title":"legged_gym 默认配置 legged_robot_config.py","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"},{"title":"Isaac Gym Python API: SimParams（dt / substeps）","url":"https://docs.robotsfan.com/isaacgym/api/python/struct_py.html"}],"as_of":"","related_ids":["simulation-timestep","control-decimation","numerical-integrator","simulation-instability","solver-iteration-count","physics-engine"],"name":"Substeps","alt":"子步","abbr":"","aliases":["Substepping","Physics Substeps"],"one_liner":"Splitting one simulation step into several smaller integration steps, trading extra computation for more stable, accurate physics.","explanation":"Substeps are a setting inside a physics simulator: each time the simulator advances by one step, the engine internally divides that time span into a number of smaller pieces and does integration and constraint solving (the step that computes contact forces and joint constraints) once per piece, in sequence. Isaac Gym's SimParams has both a dt and a substeps parameter; Genesis's SimOptions also has substeps, and its documentation says it means how many times the solver integrates within each scene.step() call. More substeps means a smaller effective timestep, which makes high-speed collisions and stiff contacts less likely to tunnel through geometry or blow up numerically, but the computation cost also grows proportionally. It is not the same thing as control decimation: decimation means the policy only issues a new action every few simulation steps, while substeps subdivide inside a single simulation step. Substeps, together with control decimation and the simulation timestep, jointly determine the physics frequency and the control frequency.","example":"legged_gym's default configuration uses a simulation timestep dt=0.005 seconds, substeps=1, and decimation=4 — physics at 200 Hz and the policy at 50 Hz; changing substeps to 2 makes the engine integrate internally every 2.5 milliseconds while the policy's frequency stays the same.","related":["Simulation Timestep","Control Decimation","Numerical Integrator (Semi-implicit Euler / RK4)","Simulation Instability","Solver Iteration Count (Position / Velocity Iterations)","Physics Engine"]},{"id":"real-time-factor","category":"sim","sec":0,"tier":3,"sources":[{"title":"Gazebo Classic 教程：Physics Parameters（real_time_factor）","url":"https://classic.gazebosim.org/tutorials?tut=physics_params"},{"title":"Gazebo 文档：Understanding the GUI（RTF 显示）","url":"https://gazebosim.org/docs/latest/gui/"},{"title":"Learning agile and dynamic motor skills for legged robots (arXiv 1901.08652)","url":"https://arxiv.org/abs/1901.08652"}],"as_of":"","related_ids":["simulator","simulation-timestep","simulation-throughput","gpu-accelerated-parallel-simulation","hardware-in-the-loop-software-in-the-loop-simulation","gazebo"],"name":"Real-Time Factor","alt":"实时因子","abbr":"RTF","aliases":["RTF","Simulation Speed Multiplier"],"one_liner":"The ratio of simulated time to real elapsed time, measuring whether a simulation runs faster or slower than reality.","explanation":"Real-Time Factor = time advanced inside the simulation ÷ time elapsed in the real world. RTF = 1 means the simulation runs at the same speed as reality; above 1 is faster than real time, so RTF = 10 means 10 seconds of simulation happen in 1 real second; below 1 is slower, common when a scene is complex, has many sensors, or has heavy rendering. Gazebo's interface shows RTF live at the bottom of the screen; in Gazebo Classic, the step size max_step_size multiplied by the update rate real_time_update_rate gives RTF's upper bound, and setting the update rate to 0 runs the simulation as fast as possible. RTF matters greatly for reinforcement learning, which needs huge amounts of interaction — the higher the RTF, the more experience can be collected per wall-clock hour; GPU-parallel simulators instead usually report throughput as steps per second summed across all environments. Conversely, hardware- and software-in-the-loop testing must stay synchronized with a real controller, so RTF is locked to 1. Note that RTF in speech recognition is defined the opposite way (processing time ÷ audio duration, smaller is faster).","example":"In ETH's 2019 ANYmal locomotion work, the simulator advanced nearly 500,000 timesteps per second, roughly a thousand times faster than real time (RTF ≈ 1000), letting the whole training finish on one computer in under eleven hours.","related":["Simulator","Simulation Timestep","Simulation Throughput (Steps / Frames Per Second)","GPU-Accelerated Parallel Simulation","Hardware-in-the-Loop / Software-in-the-Loop Simulation","Gazebo"]},{"id":"vectorized-environments","category":"sim","sec":0,"tier":2,"sources":[{"title":"Gymnasium Documentation: Vector environments","url":"https://gymnasium.farama.org/api/vector/"},{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (arXiv 2108.10470)","url":"https://arxiv.org/abs/2108.10470"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"2026-09","related_ids":["gpu-accelerated-parallel-simulation","massively-parallel-reinforcement-learning","nvidia-isaac-lab","isaac-gym","simulation-throughput","proximal-policy-optimization"],"name":"Vectorized Environments","alt":"并行环境","abbr":"","aliases":["num_envs","Vector Env","Parallel Environments"],"one_liner":"Running many independent copies of an environment at once to collect interaction data in batches.","explanation":"Vectorized environments are a way to speed up sampling in reinforcement learning: the same environment is duplicated N times (N is the num_envs value in code), and each step sends in N actions together and gets back N observations, rewards, and done flags — exactly forming one batch to feed a neural network. Gymnasium provides a sequential SyncVectorEnv and a multi-process AsyncVectorEnv, and automatically resets any copy whose episode ends without waiting for the others. GPU simulators such as Isaac Gym, Isaac Lab, and MJX go a step further, computing thousands of environments at once on a single GPU and handing the data to the training code directly as PyTorch tensors; the Isaac Gym paper claims speedups of two to three orders of magnitude over CPU simulation. Rudin and colleagues used this in 2021 to get an ANYmal quadruped walking on flat ground in under 4 minutes.","example":"Isaac Lab's velocity-tracking task for legged robots defaults to num_envs=4096, simulating 4,096 robot copies at once on a single GPU while training a walking policy with PPO.","related":["GPU-Accelerated Parallel Simulation","Massively Parallel Reinforcement Learning","NVIDIA Isaac Lab","Isaac Gym","Simulation Throughput (Steps / Frames Per Second)","Proximal Policy Optimization"]},{"id":"gpu-accelerated-parallel-simulation","category":"sim","sec":0,"tier":2,"sources":[{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (arXiv 2108.10470)","url":"https://arxiv.org/abs/2108.10470"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["vectorized-environments","isaac-gym","nvidia-isaac-lab","massively-parallel-reinforcement-learning","mujoco-xla","legged-gym"],"name":"GPU-Accelerated Parallel Simulation","alt":"GPU 并行仿真","abbr":"","aliases":["Massively Parallel Simulation","GPU-Accelerated Simulation"],"one_liner":"Running thousands of simulated environments at once on a single GPU, shrinking robot training from days to minutes.","explanation":"GPU-accelerated parallel simulation moves the physics simulation itself onto the GPU, advancing thousands of mutually independent environments at the same time and handing the simulation state directly to a framework like PyTorch as tensors for training — skipping the back-and-forth data copying between CPU and GPU that older pipelines needed. NVIDIA's Isaac Gym brought this approach into the mainstream in 2021; its paper claimed speedups of two to three orders of magnitude over the traditional “CPU simulation plus GPU training” pipeline. The most direct beneficiary is reinforcement learning: once sampling becomes cheap, policies for legged locomotion or dexterous manipulation can be trained on a single GPU in tens of minutes. Commonly used platforms today include Isaac Lab, MuJoCo MJX and MuJoCo Warp, Genesis, ManiSkill3, and Brax. The trade-off is that scenes must be written in a batched form, and complex contact or high-quality rendering noticeably slows things down.","example":"ETH Zürich's Rudin and colleagues simulated thousands of ANYmal quadruped robots in parallel on a single GPU, training a flat-ground walking policy in under 4 minutes and a rough-terrain policy in about 20 minutes, then transferred both successfully to the real robot.","related":["Vectorized Environments","Isaac Gym","NVIDIA Isaac Lab","Massively Parallel Reinforcement Learning","MuJoCo XLA","legged_gym"]},{"id":"simulation-throughput","category":"sim","sec":0,"tier":3,"sources":[{"title":"arXiv 2108.10470 - Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning","url":"https://arxiv.org/abs/2108.10470"},{"title":"arXiv 2109.11978 - Learning to Walk in Minutes Using Massively Parallel Deep RL","url":"https://arxiv.org/abs/2109.11978"},{"title":"arXiv 2410.00425 - ManiSkill3","url":"https://arxiv.org/abs/2410.00425"}],"as_of":"","related_ids":["gpu-accelerated-parallel-simulation","vectorized-environments","real-time-factor","massively-parallel-reinforcement-learning","batched-rendering","sample-efficiency"],"name":"Simulation Throughput (Steps / Frames Per Second)","alt":"仿真吞吐量","abbr":"","aliases":["Simulation FPS","Steps Per Second"],"one_liner":"How many environment steps or frames a simulator produces per second, which sets how fast data can be generated and training can run.","explanation":"Simulation throughput refers to how much interaction data a simulator produces per unit time, usually expressed as steps per second or frames per second (FPS); in GPU-parallel simulation, it is typically reported as the summed steps across thousands of parallel environments. Reinforcement learning often needs hundreds of millions of steps of interaction, so throughput directly determines whether a training run takes days or minutes. NVIDIA's 2021 Isaac Gym put both physics simulation and the neural network on the GPU with no data crossing to the CPU, reporting speedups of 2 to 3 orders of magnitude over the older “CPU simulation plus GPU training” approach. The main factors affecting throughput are the number of parallel environments, the complexity of the model and its contacts, the simulation timestep and number of substeps, solver iteration count, and whether camera images are rendered (rendering is usually much slower). It differs from real-time factor, which measures how much faster a single environment runs than real time. When comparing numbers across papers, check whether they use the same task, the same hardware, and rendering or not.","example":"ETH's legged_gym simulates thousands of ANYmal quadrupeds on a single GPU at once, finishing a flat-ground walking policy in under 4 minutes and a rough-terrain one in about 20; ManiSkill3 reports over 30,000 frames per second with rendering enabled on its benchmark environments.","related":["GPU-Accelerated Parallel Simulation","Vectorized Environments","Real-Time Factor","Massively Parallel Reinforcement Learning","Batched Rendering","Sample Efficiency"]},{"id":"headless-mode","category":"sim","sec":0,"tier":3,"sources":[{"title":"Isaac Lab API: isaaclab.app (AppLauncher)","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.app.html"},{"title":"google-deepmind/dm_control README: Rendering","url":"https://github.com/google-deepmind/dm_control"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","nvidia-isaac-sim","mujoco","batched-rendering","gpu-accelerated-parallel-simulation","simulation-throughput"],"name":"Headless Mode","alt":"无头模式","abbr":"","aliases":["--headless","Headless Rendering"],"one_liner":"Running a simulator with no graphical window open, commonly used for batch training and data collection on a server.","explanation":"Headless means running a program without a graphical user interface (GUI). In embodied AI, this mainly means running simulators such as Isaac Sim/Isaac Lab or MuJoCo on a GPU server with no monitor attached: no window pops up and nothing is displayed in real time, which saves overhead and speeds up training and data generation. Isaac Lab scripts just need a --headless flag added, which its documentation describes as launching “without a graphical interface”; if the environment has cameras that need to output images, --enable_cameras also needs to be added to enable off-screen rendering. Tools in the MuJoCo family use the environment variable MUJOCO_GL=egl for windowless rendering on the GPU, while osmesa does pure CPU software rendering. Note that headless doesn't mean no rendering at all — it just means nothing is displayed; to view the output remotely, Isaac Lab's livestream feature or video recording can be used instead.","example":"Running python scripts/reinforcement_learning/rsl_rl/train.py --task=Isaac-Cartpole-v0 --headless on a server trains a cart-pole policy with no window opened.","related":["NVIDIA Isaac Lab","NVIDIA Isaac Sim","MuJoCo (Multi-Joint dynamics with Contact)","Batched Rendering","GPU-Accelerated Parallel Simulation","Simulation Throughput (Steps / Frames Per Second)"]},{"id":"simulation-fidelity","category":"sim","sec":0,"tier":2,"sources":[{"title":"Choi et al., On the use of simulation in robotics: Opportunities, challenges, and suggestions for moving forward (PNAS 2021)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7817170/"},{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"}],"as_of":"","related_ids":["sim-to-real-gap","physics-engine","rendering","sensor-simulation","digital-twin","domain-randomization"],"name":"Simulation Fidelity","alt":"仿真保真度","abbr":"","aliases":["Physical Fidelity","Visual Fidelity"],"one_liner":"How closely a simulation matches the real world, in both the accuracy of its physics and the realism of its appearance.","explanation":"Simulation fidelity is usually split into two parts: physical fidelity, meaning how accurately dynamics, contact, friction, deformable objects, and motor characteristics are computed; and visual or sensor fidelity, meaning how closely rendered images, depth, and lidar readings match a real sensor. Higher fidelity means a smaller sim-to-real gap, but it also costs more compute. A 2021 PNAS review by Choi and colleagues points out that game engines aim for “looking plausible” rather than being precise, and that practical use often requires compromises such as reducing solver iterations or substituting a rigid ground for a deformable one. In practice this is usually paired with domain randomization and system identification; work such as SIMPLER has also shown that policy evaluation doesn't necessarily require a fully realistic digital twin.","example":"SIMPLER didn't reconstruct a complete digital twin — it only composited a real background onto a green screen, aligned textures, and identified control parameters, which was enough to make its simulated evaluation scores correlate strongly with real-robot results.","related":["Sim-to-Real Gap (Reality Gap)","Physics Engine","Rendering","Sensor Simulation","Digital Twin","Domain Randomization"]},{"id":"rigid-body-simulation","category":"sim","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Rigid body dynamics","url":"https://en.wikipedia.org/wiki/Rigid_body_dynamics"},{"title":"MuJoCo Documentation: Overview","url":"https://mujoco.readthedocs.io/en/stable/overview.html"}],"as_of":"","related_ids":["physics-engine","rigid-body-dynamics","articulated-body-simulation","contact-model","deformable-body-simulation","mujoco"],"name":"Rigid-Body Simulation","alt":"刚体仿真","abbr":"","aliases":["Rigid-Body Dynamics Simulation"],"one_liner":"Physics simulation that assumes objects never deform, computing only their translation, rotation, and collisions with each other.","explanation":"Rigid-body simulation is the most basic and most widely used part of a physics engine. It assumes an object's shape never changes under force, so each object's state can be described with just its position, orientation, velocity, and angular velocity. Simulation advances in fixed timesteps: each step first runs collision detection to find contact points, then a contact model and constraint solver compute contact forces, friction, and joint constraint forces, and finally an integrator updates the state using the Newton-Euler equations. Robots are usually modeled as several rigid links connected by joints (an articulated body), which is why tasks like quadruped locomotion or an arm stacking blocks are all handled within rigid-body simulation. MuJoCo, PhysX, and Bullet all build rigid-body simulation as their core. It cannot handle cloth, rope, liquids, or soft objects, which instead require soft-body, cloth, or fluid simulation; even contact and friction are only approximations, making them one of the main sources of the sim-to-real gap.","example":"When a robot arm pushes a block off a table, both the block and the table are treated as rigid bodies that never deform; each step, the engine detects collisions between the block and the table edge or the ground, computes contact and friction forces, and integrates the block's next position and rotation.","related":["Physics Engine","Rigid-Body Dynamics","Articulated-Body Simulation (Articulation)","Contact Model","Deformable-Body Simulation","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"articulated-body-simulation","category":"sim","sec":1,"tier":2,"sources":[{"title":"NVIDIA PhysX 5 Documentation: Articulations","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/Articulations.html"},{"title":"MuJoCo Documentation: Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Wikipedia: Featherstone's algorithm","url":"https://en.wikipedia.org/wiki/Featherstone%27s_algorithm"}],"as_of":"","related_ids":["rigid-body-simulation","articulated-body-algorithm","generalized-coordinates","generalized-coordinates-vs-cartesian-coordinates","physx","mujoco"],"name":"Articulated-Body Simulation (Articulation)","alt":"关节体仿真","abbr":"","aliases":["Multi-rigid-body simulation","Articulation"],"one_liner":"Simulating a robot as rigid bodies connected by joints, computing how it moves and what forces it feels.","explanation":"A robot is built from links (rigid parts) connected into a tree by joints, and articulated-body simulation is the technique for computing how such a system moves. There are two general approaches: treat every link as an independent rigid body and hold them together with constraints (maximal coordinates), where joints can slowly drift apart or separate; or describe the whole robot using only the base pose plus each joint angle (reduced, or generalized, coordinates), where joints can never come apart by construction — Featherstone's 1987 book laid out the efficient algorithms for this approach. Both PhysX's Articulation feature and MuJoCo take the reduced-coordinate route; PhysX's documentation states that this gives zero joint error and lets it handle much larger mass ratios, with computation scaling with the number of degrees of freedom rather than the number of links. The tradeoff is that it only supports tree structures — a closed loop, such as a parallel ankle mechanism, needs extra constraints added on top.","example":"Importing a Franka arm's URDF into Isaac Sim produces a fixed-base Articulation: 7 revolute joints plus 2 prismatic finger joints, so the simulation state is just the positions and velocities of these 9 joints.","related":["Rigid-Body Simulation","Articulated Body Algorithm","Generalized Coordinates","Generalized (Reduced) Coordinates vs. Cartesian (Maximal) Coordinates","PhysX","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"generalized-coordinates-vs-cartesian-coordinates","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo Documentation: Overview","url":"https://mujoco.readthedocs.io/en/stable/overview.html"},{"title":"PhysX 5 Documentation: Articulations","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/Articulations.html"}],"as_of":"","related_ids":["generalized-coordinates","articulated-body-simulation","multibody-dynamics","constraint-solver","mujoco","physx"],"name":"Generalized (Reduced) Coordinates vs. Cartesian (Maximal) Coordinates","alt":"关节坐标仿真 vs 笛卡尔坐标仿真（约化坐标 / 最大坐标）","abbr":"","aliases":["Reduced-coordinate Articulation","Maximal Coordinates","Generalized-coordinate Simulation"],"one_liner":"Two ways a physics engine can represent a multi-link robot: by joint angles alone, or as separate rigid bodies tied together by constraints.","explanation":"There are two modeling approaches for simulating a robot made of multiple links and joints. Generalized (or reduced) coordinates describe the state using only each joint's angle or displacement; joint constraints hold automatically and can never be pulled apart, and the number of variables equals the number of degrees of freedom, which suits arms and legged robots well — the “articulation” feature in both MuJoCo and PhysX works this way. The cost is a more complex algorithm, which usually requires the link structure to be a tree, with closed loops needing extra constraints. Cartesian (or maximal) coordinates instead give every rigid body its full 6 degrees of freedom and connect them with constraints; this is simpler to implement and is what game physics engines commonly do, but because the constraints are enforced numerically, joints can drift or jitter when the structure is complex or link masses vary wildly. Knowing this distinction matters when choosing a simulator or debugging simulation instability.","example":"The same robot arm can be built in PhysX either as rigid bodies plus joint constraints (maximal coordinates) or as an articulation (reduced coordinates); the official documentation notes that the latter has joint error at zero by design and can tolerate a much larger ratio between link masses, which is why it's recommended for robots.","related":["Generalized Coordinates","Articulated-Body Simulation (Articulation)","Multibody Dynamics","Constraint Solver","MuJoCo (Multi-Joint dynamics with Contact)","PhysX"]},{"id":"collision-detection-2","category":"sim","sec":1,"tier":2,"sources":[{"title":"MuJoCo Documentation: Computation - Collision detection","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"NVIDIA PhysX 5 Documentation: Rigid Body Collision","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyCollision.html"}],"as_of":"","related_ids":["collision-geometry","broad-phase-narrow-phase-collision-detection","gilbert-johnson-keerthi-algorithm","bounding-volume","collision-filtering","continuous-collision-detection"],"name":"Collision Detection","alt":"碰撞检测（物理引擎）","abbr":"","aliases":["broad phase / narrow phase"],"one_liner":"The step where a physics engine works out which objects are touching, where, and how deeply.","explanation":"Collision detection is a step a physics engine runs on every simulation tick, producing a list of contacts (contact points, normal directions, and penetration depth) that's then handed to the solver to compute contact forces. With n objects in a scene, there are n(n−1)/2 possible pairs, and checking every pair exactly is too slow, so engines generally split the work into two phases: broad phase, which uses cheap methods like bounding-box sorting (such as sweep-and-prune) to quickly rule out pairs that obviously can't be touching; and narrow phase, which does exact geometric computation on the remaining pairs, commonly using the GJK/EPA algorithm for convex shapes. MuJoCo adds a middle layer between the two, based on a bounding-volume hierarchy, and uses contype/conaffinity bitmasks to skip pairs that don't need checking at all. Note that this is a different “collision detection” from the kind used in robot safety systems to detect an actual physical impact.","example":"In MuJoCo, all the geometries of a robot, a table, and a block first go through sweep-and-prune to filter down to pairs that might touch, then a bounding-sphere test filters further; finally, basic shapes like spheres and boxes are handled with closed-form formulas, while pairs involving a mesh (first converted to a convex hull) go through GJK/EPA to compute contact points and penetration depth.","related":["Collision Geometry (Collider)","Broad-phase / Narrow-phase Collision Detection","Gilbert-Johnson-Keerthi Algorithm","Bounding Volume (AABB / OBB)","Collision Filtering","Continuous Collision Detection"]},{"id":"collision-geometry","category":"sim","sec":1,"tier":2,"sources":[{"title":"Isaac Sim Documentation: Simulation Fundamentals（Collision approximations）","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/physics/simulation_fundamentals.html"},{"title":"MuJoCo Documentation: Computation - Collision detection","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"}],"as_of":"","related_ids":["convex-decomposition","convex-decomposition","collision-detection-2","unified-robot-description-format","signed-distance-field-function","interpenetration"],"name":"Collision Geometry (Collider)","alt":"碰撞体","abbr":"","aliases":["collision mesh","collision shape","Collider"],"one_liner":"The simplified shape a physics engine uses for collision checks, usually different from what's rendered on screen.","explanation":"A simulated object usually has two separate geometries: a visual mesh, used for rendering and meant to look good, with a high polygon count; and a collision geometry, handed to the physics engine for collision detection, meant to be computed quickly and stably. Collision geometry is usually a basic shape — sphere, capsule, box — or a convex hull (the smallest convex shape that encloses the object), because these have efficient collision algorithms. MuJoCo automatically converts a non-convex mesh supplied by the user into a convex hull for collision purposes; Isaac Sim also defaults to convex hulls, with convex decomposition (splitting an object into several convex pieces) or an SDF (signed distance field) mesh available when a closer fit is needed. Too coarse a collision shape, and a mug gets treated as solid, with nothing able to grip its handle; too fine, and the simulation slows down. This is exactly the distinction between <visual> and <collision> in a URDF, specified separately for every link.","example":"If a coffee mug uses a single convex hull as its collision geometry, the opening gets sealed shut and a spoon can't be placed inside it in simulation; decomposing it into several convex pieces with a tool like CoACD lets the wall, base, and handle collide separately.","related":["Convex Decomposition","Convex Decomposition","Collision Detection","Unified Robot Description Format","Signed Distance Field / Function","Interpenetration"]},{"id":"convex-decomposition","category":"sim","sec":1,"tier":3,"sources":[{"title":"Approximate Convex Decomposition for 3D Meshes with Collision-Aware Concavity and Tree Search (CoACD, SIGGRAPH 2022)","url":"https://arxiv.org/abs/2205.02961"},{"title":"V-HACD GitHub 仓库（已归档，指向 CoACD）","url":"https://github.com/kmammou/v-hacd"},{"title":"MuJoCo 文档 Modeling（mesh geom 碰撞时被凸化）","url":"https://mujoco.readthedocs.io/en/stable/modeling.html"}],"as_of":"2026-09","related_ids":["collision-geometry","collision-detection-2","convex-decomposition","simulation-assets","mujoco","interpenetration"],"name":"Convex Decomposition","alt":"凸分解","abbr":"","aliases":["Approximate Convex Decomposition","ACD","CoACD","V-HACD"],"one_liner":"Cutting a concave mesh into several approximately convex pieces so it can serve as simulation collision geometry.","explanation":"When a physics engine does collision detection, convex shapes (where the line between any two points stays inside the shape) are fast and numerically stable to work with, which is why engines like MuJoCo automatically replace a mesh's collision geometry with its convex hull (the smallest convex shape that contains it). For concave objects like cups, bowls, and drawers, taking the convex hull directly “seals up” the opening, so nothing can be placed inside. Convex decomposition instead cuts a concave mesh into several approximately convex pieces, takes the convex hull of each piece, and assembles them to approximate the original shape. Two tools are commonly used: V-HACD (voxelized hierarchical approximate convex decomposition by Karim Mammou, now unmaintained) and CoACD (from Hao Su's group at UC San Diego, SIGGRAPH 2022, which uses “collision-aware concavity” plus tree search to choose cutting planes, preserving more detail, and installs via pip). When preparing simulation assets, object meshes are usually run through convex decomposition before entering the simulator.","example":"A coffee mug mesh reduced to a single convex hull has its opening sealed shut, so a spoon can't be placed inside it in simulation; decomposed with CoACD into several convex pieces instead, the wall and handle each become their own piece and the interior space is preserved.","related":["Collision Geometry (Collider)","Collision Detection","Convex Decomposition","Simulation Assets","MuJoCo (Multi-Joint dynamics with Contact)","Interpenetration"]},{"id":"signed-distance-field-function","category":"sim","sec":1,"tier":3,"sources":[{"title":"Wikipedia - Signed distance function","url":"https://en.wikipedia.org/wiki/Signed_distance_function"}],"as_of":"","related_ids":["truncated-signed-distance-function","euclidean-signed-distance-field","collision-detection-2","collision-checking","covariant-hamiltonian-optimization-for-motion-planning","implicit-vs-explicit-3d-representation"],"name":"Signed Distance Field / Function","alt":"符号距离场","abbr":"SDF","aliases":["SDF","Signed Distance Function"],"one_liner":"A function over space giving each point's distance to the nearest object surface, signed to distinguish inside from outside.","explanation":"A signed distance field is a function defined over space: given any point, it returns the distance to the nearest object surface, with a sign distinguishing whether the point is inside or outside the object (robotics and graphics usually take positive outside and negative inside, though some mathematical literature uses the opposite convention); the object's surface is exactly where the function equals 0. Its gradient has magnitude 1 everywhere, and points in the direction that moves away from the surface fastest. These two properties make it useful throughout embodied AI: a simulator can query it directly for penetration depth and contact normals during collision detection; motion planning uses it to compute how far a robot is from an obstacle and which way to steer to avoid it; and 3D reconstruction and neural implicit representations often represent an object's surface as an SDF. Note that the same abbreviation SDFormat refers to Gazebo's scene-description file format, an unrelated thing.","example":"When planning an arm's trajectory, first build a signed distance field for the table and obstacles, then query the SDF value at sampled points on each link; if it's below a safety margin, push the trajectory away along the gradient direction — this is how optimization-based planners like CHOMP avoid obstacles.","related":["Truncated Signed Distance Function","Euclidean Signed Distance Field","Collision Detection","Collision Checking","Covariant Hamiltonian Optimization for Motion Planning","Implicit vs. Explicit 3D Representation"]},{"id":"broad-phase-narrow-phase-collision-detection","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo Documentation: Computation - Collision detection","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html#collision"},{"title":"NVIDIA PhysX 5 SDK Documentation: Rigid Body Collision","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyCollision.html"},{"title":"Wikipedia: Collision detection","url":"https://en.wikipedia.org/wiki/Collision_detection"}],"as_of":"","related_ids":["collision-detection-2","bounding-volume","gilbert-johnson-keerthi-algorithm","collision-filtering","collision-geometry","physics-engine"],"name":"Broad-phase / Narrow-phase Collision Detection","alt":"宽相 / 窄相碰撞检测","abbr":"","aliases":["Broad Phase / Narrow Phase","Coarse / Fine Collision Detection","Near-phase"],"one_liner":"Collision detection done in two steps: first a coarse bounding-box pass to find likely pairs, then exact contact computation.","explanation":"A physics engine has to figure out which objects are touching at every step. With n geometries, there are n(n-1)/2 possible pairs, and checking every pair exactly would be too slow, so the work is split into phases. The broad phase uses simple bounding volumes (such as axis-aligned bounding boxes, or AABBs) for a conservative test that quickly rules out pairs that clearly don't overlap, reporting only “possible collisions”; a common algorithm is sweep-and-prune, which sorts objects along an axis and looks for overlapping intervals. The narrow phase then runs an exact, geometry-specific algorithm (such as GJK/EPA for convex shapes) on the remaining candidate pairs to determine whether they actually touch, along with the contact point, normal, and penetration depth, which is handed to the constraint solver. MuJoCo inserts a mid-phase between the two, further filtering with a bounding-volume hierarchy; PhysX offers several broad-phase algorithms, including SAP, MBP, and a GPU variant. Collision filtering is usually inserted between the broad and narrow phases.","example":"MuJoCo's broad phase uses a modified sweep-and-prune, sorting along the dominant eigenvector of the covariance matrix of all geometry centers; in the narrow phase, a non-convex mesh gets replaced with its convex hull before being tested.","related":["Collision Detection","Bounding Volume (AABB / OBB)","Gilbert-Johnson-Keerthi Algorithm","Collision Filtering","Collision Geometry (Collider)","Physics Engine"]},{"id":"collision-filtering","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo Documentation: Computation - Collision detection (Selection)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html#collision"},{"title":"NVIDIA PhysX 5 SDK Documentation: Rigid Body Collision - Collision Filtering","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyCollision.html"},{"title":"Isaac Lab API: isaaclab.scene (InteractiveSceneCfg.filter_collisions)","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.scene.html"}],"as_of":"2026-09","related_ids":["broad-phase-narrow-phase-collision-detection","collision-detection-2","self-collision-checking","collision-geometry","mjcf","vectorized-environments"],"name":"Collision Filtering","alt":"碰撞过滤（碰撞组）","abbr":"","aliases":["Collision Groups","contype/conaffinity","Collision Mask","Collision Layers"],"one_liner":"Declaring in advance which object pairs should never count as colliding, saving computation and avoiding contact that shouldn't exist.","explanation":"Not every pair of objects in a simulation needs its collisions computed: a robot's adjacent links are naturally touching right at a joint, and thousands of parallel training environments running at once shouldn't collide with each other either. Collision filtering discards these pairs after the broad phase but before the narrow phase. A common technique is a bitmask attached to each geometry. In MuJoCo, every geom has two integers, contype and conaffinity, and a pair is only checked if one's contype shares at least one set bit with the other's conaffinity — a mechanism borrowed from the ODE engine, with both defaulting to 1; MuJoCo also skips collisions within the same rigid body and between parent and child bodies by default, and pairs can be explicitly excluded with exclude. PhysX implements similar grouping through a filter shader. Getting this wrong leads to problems like a gripper failing to hold an object or a foot passing through the ground.","example":"Isaac Lab's InteractiveSceneCfg defaults to filter_collisions=True, so cloned environments never collide with each other, while all of them can still collide with global objects such as the ground.","related":["Broad-phase / Narrow-phase Collision Detection","Collision Detection","Self-Collision Checking","Collision Geometry (Collider)","MJCF (MuJoCo XML Format)","Vectorized Environments"]},{"id":"continuous-collision-detection","category":"sim","sec":1,"tier":3,"sources":[{"title":"NVIDIA PhysX 5 SDK Documentation: Advanced Collision Detection - Continuous Collision Detection","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/AdvancedCollisionDetection.html"},{"title":"Wikipedia: Collision detection","url":"https://en.wikipedia.org/wiki/Collision_detection"},{"title":"MuJoCo Documentation: Computation - Convex collisions","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html#convex-collisions"}],"as_of":"","related_ids":["collision-detection-2","interpenetration","simulation-timestep","substeps","broad-phase-narrow-phase-collision-detection","physx"],"name":"Continuous Collision Detection","alt":"连续碰撞检测（隧穿问题）","abbr":"CCD","aliases":["CCD","Tunneling"],"one_liner":"Finding collisions along an object's full motion path within a step, preventing fast-moving objects from passing straight through obstacles.","explanation":"Ordinary discrete collision detection only checks whether objects overlap at the end of each timestep. When an object is very fast or very thin, it can jump from one side of an obstacle to the other within a single step, with neither the before nor after check finding any overlap — so the collision is missed entirely. This is called tunneling. Continuous collision detection instead finds the earliest time of contact along the path an object sweeps out during that step, handling the collision before the object passes through — at the cost of more computation, which grows more noticeable as speed and object density increase. PhysX requires CCD to be enabled separately at the scene, object-pair, and rigid-body levels, and also offers a cheaper speculative CCD (which enlarges the contact distance based on velocity). Another common fix is shrinking the simulation timestep or adding more substeps. Note that in MuJoCo's documentation, “CCD” instead refers to convex collision detection, not continuous collision detection.","example":"A thin rod swung quickly, or a small ball thrown fast, can pass straight through a tabletop in simulation when the timestep is too large; enabling CCD or shrinking the timestep fixes the behavior.","related":["Collision Detection","Interpenetration","Simulation Timestep","Substeps","Broad-phase / Narrow-phase Collision Detection","PhysX"]},{"id":"interpenetration","category":"sim","sec":1,"tier":2,"sources":[{"title":"MuJoCo Documentation: Computation (Soft contacts)","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"NVIDIA PhysX 5 Docs: Advanced Collision Detection","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/AdvancedCollisionDetection.html"}],"as_of":"","related_ids":["collision-geometry","soft-contact-model","continuous-collision-detection","simulation-timestep","convex-decomposition","solver-iteration-count"],"name":"Interpenetration","alt":"穿模","abbr":"","aliases":["Penetration","Clipping"],"one_liner":"A simulation or animation glitch where two objects that should block each other pass through one another instead.","explanation":"Interpenetration (often called clipping) originated as gaming and animation jargon for models overlapping incorrectly; in robot simulation it means two collision bodies overlap — a finger poking through a tabletop, or a foot sinking into the ground. Physics engines tolerate a small amount of penetration by design: MuJoCo uses a soft-contact model that generates contact force based on penetration depth, and PhysX also treats penetration as a negative distance and progressively increases contact force to push objects apart. The problem is penetration that goes too deep. Common causes include too large a simulation timestep, too few solver iterations, collision geometry that is too thin or too coarse, or objects moving too fast (passing all the way through is called tunneling). Deep interpenetration distorts contact forces and grasp outcomes, and a policy can even learn to exploit it as a “cheat” that fails once deployed on a real robot. Fixes include shrinking the timestep, adding substeps and solver iterations, using convex decomposition for collision geometry, and enabling continuous collision detection. Feet sinking into the ground in motion-capture data is also called interpenetration.","example":"When training a grasping policy in simulation, if a cup's collision geometry is a thin hollow mesh, a finger may pass straight through the cup wall; the simulator still counts this as a successful grasp, but the real robot cannot actually pick the cup up. Switching to convex-decomposed collision geometry and increasing solver iterations usually helps.","related":["Collision Geometry (Collider)","Soft Contact Model","Continuous Collision Detection","Simulation Timestep","Convex Decomposition","Solver Iteration Count (Position / Velocity Iterations)"]},{"id":"contact-model","category":"sim","sec":1,"tier":2,"sources":[{"title":"MuJoCo Documentation: Computation - Soft contact model","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Isaac Sim Documentation: Simulation Fundamentals","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/physics/simulation_fundamentals.html"}],"as_of":"","related_ids":["soft-contact-model","linear-complementarity-problem","friction-cone","constraint-solver","collision-detection-2","coulomb-friction"],"name":"Contact Model","alt":"接触模型","abbr":"","aliases":["contact dynamics"],"one_liner":"The physics engine's rule for how much force and friction two touching objects generate.","explanation":"Collision detection only answers whether and where two objects are touching; a contact model answers how much force results from that contact: a normal force that keeps objects from interpenetrating, and a tangential friction force that resists sliding, capped by a friction cone whose size is the friction coefficient times the normal force. The classical approach is a hard-contact model, which writes “no force without contact, and no interpenetration once there is contact” as a complementarity condition, reducing to a Linear Complementarity Problem (LCP) that becomes NP-hard once friction is included, so different engines rely on various approximations. MuJoCo instead uses soft contact: it allows a small amount of interpenetration and defines the contact force as the unique solution to a convex optimization problem, with parameters like solref and solimp used to tune how soft or stiff it is. How accurate the contact model is directly affects whether contact-heavy tasks — grasping, peg insertion, legged locomotion — can transfer successfully from simulation to a real robot.","example":"In MuJoCo, setting a geom's condim to 3 computes only the normal force and in-plane sliding friction; setting it to 4 adds torsional friction, which the official documentation says helps simulate soft fingertips for steadier grasping; setting it to 6 further adds rolling friction.","related":["Soft Contact Model","Linear Complementarity Problem","Friction Cone","Constraint Solver","Collision Detection","Coulomb Friction"]},{"id":"soft-contact-model","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo 文档 - Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"MuJoCo 文档 - Modeling（solref / solimp）","url":"https://mujoco.readthedocs.io/en/stable/modeling.html"}],"as_of":"","related_ids":["contact-model","linear-complementarity-problem","mujoco","interpenetration","projected-gauss-seidel","simulation-timestep"],"name":"Soft Contact Model","alt":"软接触","abbr":"","aliases":["Soft Contact","Compliant Constraint"],"one_liner":"A contact model that allows a small amount of interpenetration and generates contact force like a spring-damper.","explanation":"Soft contact is a family of contact-modeling approaches used in physics engines. The classical hard-contact approach requires that bodies never interpenetrate at all, describing this with a complementarity condition (contact force can only exist when bodies are not separating, and separation implies zero contact force), which is difficult to solve and prone to numerical instability. Soft contact relaxes this requirement: it allows a small amount of penetration and treats contact as a spring-and-damper constraint, where force grows with penetration depth. MuJoCo is the leading example — it drops the strict complementarity condition from the linear complementarity problem and instead formulates contact as a convex optimization problem, tuned with two parameter groups: solref (a time constant and damping ratio, defaulting to 0.02 and 1) and solimp (constraint impedance, which controls how soft the contact is). The benefits are stable solving, smooth results, and a unique inverse-dynamics solution; the cost is that poorly chosen parameters cause visible interpenetration or bounciness, so sim-to-real transfer work often has to tune these parameters to match real materials.","example":"In MuJoCo, increasing the time constant of solref for a gripper fingertip's contact makes the contact softer and allows deeper penetration, smoothing out how contact force changes during a grasp; setting the time constant smaller than the simulation timestep tends to make contact unstable.","related":["Contact Model","Linear Complementarity Problem","MuJoCo (Multi-Joint dynamics with Contact)","Interpenetration","Projected Gauss-Seidel","Simulation Timestep"]},{"id":"constraint-solver","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo Documentation: Computation - Constraint solver","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html#constraint-solver"},{"title":"NVIDIA PhysX 5 SDK Documentation: Rigid Body Dynamics - Constraint Solver","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyDynamics.html"}],"as_of":"2026-09","related_ids":["solver-iteration-count","projected-gauss-seidel","linear-complementarity-problem","contact-model","soft-contact-model","physics-engine"],"name":"Constraint Solver","alt":"约束求解器","abbr":"","aliases":["Contact Solver"],"one_liner":"The physics-engine module that computes contact and joint constraint forces, keeping objects from interpenetrating and joints from coming apart.","explanation":"Collision detection only tells the engine where objects are touching; how much supporting and frictional force each contact needs, and how much force each joint needs to stay connected, is computed by the constraint solver. These forces must simultaneously satisfy conditions like no interpenetration and friction staying within the friction cone, which is fundamentally a complementarity problem or an optimization problem, usually solved iteratively. Projected Gauss-Seidel (PGS) updates one constraint component at a time and sweeps through them repeatedly; PhysX added Temporal Gauss-Seidel (TGS) starting with version 5.1, which splits a step into several substeps to solve, and its documentation generally recommends TGS. MuJoCo defines constraint forces as the unique global solution to a convex optimization problem, offering PGS, Conjugate Gradient (CG), and Newton's method (the default) as solvers. Too few iterations leads to interpenetration, joint drift, and grasps slipping, which is why iteration count is a key parameter to tune.","example":"PhysX rigid bodies default to 4 position iterations and 1 velocity iteration; scenes with heavy contact and many joints, such as a dexterous hand grasping something, often need the position-iteration count turned up.","related":["Solver Iteration Count (Position / Velocity Iterations)","Projected Gauss-Seidel","Linear Complementarity Problem","Contact Model","Soft Contact Model","Physics Engine"]},{"id":"linear-complementarity-problem","category":"sim","sec":1,"tier":3,"sources":[{"title":"Linear complementarity problem - Wikipedia","url":"https://en.wikipedia.org/wiki/Linear_complementarity_problem"},{"title":"MuJoCo Documentation: Computation","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"ODE Manual","url":"http://ode.org/wiki/index.php/Manual"}],"as_of":"","related_ids":["contact-model","constraint-solver","projected-gauss-seidel","coulomb-friction","soft-contact-model","physics-engine"],"name":"Linear Complementarity Problem","alt":"线性互补问题","abbr":"LCP","aliases":["LCP"],"one_liner":"A math problem asking for two sets of non-negative variables where one being positive forces the other to be zero — what contact-force solving reduces to.","explanation":"The linear complementarity problem was formulated by Cottle and Dantzig in 1968: given a matrix M and a vector q, find non-negative vectors z and w such that w = Mz + q, with the added condition that for each corresponding pair of components, at least one of z or w must be zero (the complementarity condition). In simulation, this describes contact exactly: two objects either have a gap between them, with zero contact force, or are pressed together, with positive contact force, and the two situations can never hold at once — a physics engine has to solve for contact forces satisfying this kind of condition at every step. ODE's contact and friction model is based on Dantzig's LCP solver, and game engines commonly use iterative methods like Projected Gauss-Seidel to approximate a solution. MuJoCo, by contrast, notes that LCPs with friction are NP-hard, and instead uses a convex soft-contact model that relaxes the complementarity condition.","example":"A box resting still on a table has zero normal gap and a positive supporting force; once the box is lifted off the table, the gap becomes positive and the supporting force drops to zero. Solving for exactly this pair of complementary conditions is what the solver does at every step.","related":["Contact Model","Constraint Solver","Projected Gauss-Seidel","Coulomb Friction","Soft Contact Model","Physics Engine"]},{"id":"projected-gauss-seidel","category":"sim","sec":1,"tier":3,"sources":[{"title":"PhysX 5 文档：Rigid Body Dynamics（Solver Type: PGS / TGS）","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyDynamics.html"},{"title":"Gazebo Classic 教程：Physics Parameters（quick 求解器为 PGS）","url":"https://classic.gazebosim.org/tutorials?tut=physics_params"},{"title":"Isaac Lab API：isaaclab.sim（PhysxCfg.solver_type 默认 TGS）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.sim.html"}],"as_of":"2026-09","related_ids":["constraint-solver","linear-complementarity-problem","solver-iteration-count","physx","mujoco","contact-model"],"name":"Projected Gauss-Seidel","alt":"投影高斯-赛德尔求解器","abbr":"PGS","aliases":["PGS","Projective Gauss-Seidel","TGS","Temporal Gauss-Seidel"],"one_liner":"A classic iterative method that solves contact and joint forces one constraint at a time, clamping each to a valid range.","explanation":"At every step, a physics engine must compute the contact and joint forces (or impulses) that keep bodies from interpenetrating and keep friction within bounds — a system of equations with inequality conditions, commonly written as a Linear Complementarity Problem (LCP). PGS is the classic way to solve it: process one constraint at a time, assume the forces from all other constraints stay fixed, compute how much force this one constraint needs, then “project” (clamp) the result into a valid range — for example, normal force can't be negative, and friction can't exceed the friction cone. Sweeping through all constraints once counts as one iteration; more iterations means more accuracy but more time. ODE's quick solver, PhysX, and MuJoCo all offer PGS. Starting with PhysX 5.1, NVIDIA added TGS (Temporal Gauss-Seidel), which splits a timestep into several substeps, solves constraints once per substep, and integrates immediately — more accurate for large mass ratios and driven joints, and now the officially recommended default; Isaac Lab also defaults to TGS. The “position/velocity iteration count” in a simulator's settings refers to how many rounds this kind of solver runs.","example":"In Isaac Lab, a dexterous hand's block jittering between the fingers or slowly slipping is often debugged by confirming the solver is set to TGS and raising the position-iteration count.","related":["Constraint Solver","Linear Complementarity Problem","Solver Iteration Count (Position / Velocity Iterations)","PhysX","MuJoCo (Multi-Joint dynamics with Contact)","Contact Model"]},{"id":"solver-iteration-count","category":"sim","sec":1,"tier":3,"sources":[{"title":"NVIDIA PhysX 5 文档 - Rigid Body Dynamics","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/RigidBodyDynamics.html"},{"title":"Isaac Lab API - isaaclab.sim（PhysxCfg）","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.sim.html"},{"title":"Isaac Lab 源码 - franka.py","url":"https://github.com/isaac-sim/IsaacLab/blob/main/source/isaaclab_assets/isaaclab_assets/robots/franka.py"}],"as_of":"2026-09","related_ids":["constraint-solver","projected-gauss-seidel","simulation-timestep","substeps","simulation-instability","physx"],"name":"Solver Iteration Count (Position / Velocity Iterations)","alt":"求解器迭代次数","abbr":"","aliases":["Position Iterations","Velocity Iterations","solver_position_iteration_count"],"one_liner":"How many rounds a physics engine spends per step correcting contact and joint constraints; more rounds means more accuracy but less speed.","explanation":"Every time a physics engine advances one simulation step, it has to satisfy all contact and joint constraints at once, usually approximated round by round with an iterative method like Gauss-Seidel; the number of rounds is the solver iteration count. In NVIDIA PhysX (the physics core behind Isaac Sim / Isaac Lab), iterations split into position iterations, which correct penetration and joint misalignment, and velocity iterations, which correct velocity error, defaulting to 4 position iterations and 1 velocity iteration. NVIDIA's documentation says more iterations give more accurate results, but generally only bodies with many joints or low tolerance for joint error need it raised noticeably. PhysX also offers both PGS and TGS solvers, with TGS converging better at a slightly higher per-iteration cost. Too few iterations commonly shows up as loose or jittery joints and interpenetrating bodies, while raising the count slows down simulation throughput. Together with the simulation timestep and substep count, this is one of the parameters most often adjusted when tuning simulation stability.","example":"Isaac Lab's built-in Franka arm configuration sets solver_position_iteration_count to 8 and solver_velocity_iteration_count to 0, more position iterations than PhysX's default of 4; each object can set its own values, and the scene simulates at the maximum across all objects, clamped between 1 and 255.","related":["Constraint Solver","Projected Gauss-Seidel","Simulation Timestep","Substeps","Simulation Instability","PhysX"]},{"id":"numerical-integrator","category":"sim","sec":1,"tier":3,"sources":[{"title":"MuJoCo Documentation: Computation - Numerical Integration","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html"},{"title":"Gaffer On Games: Integration Basics","url":"https://gafferongames.com/post/integration_basics/"}],"as_of":"2026-09","related_ids":["physics-engine","simulation-timestep","substeps","simulation-instability","mujoco","constraint-solver"],"name":"Numerical Integrator (Semi-implicit Euler / RK4)","alt":"积分器","abbr":"","aliases":["Semi-implicit Euler","Symplectic Euler","RK4","Fourth-order Runge-Kutta","Implicit Integration"],"one_liner":"The numerical method a physics engine uses each step to advance velocity and position forward in time given the current forces.","explanation":"At every simulation step, a physics engine first computes acceleration from the forces acting on a body, then hands off to an integrator to advance velocity and position to the next instant. The most naive method, explicit Euler, updates position using the old velocity; in spring-like systems this steadily manufactures energy out of nowhere until the simulation diverges. Semi-implicit (or symplectic) Euler just swaps the order — update velocity first, then use the new velocity to update position — which is still only first-order accurate but far more stable, and is what most game physics engines use. RK4 (fourth-order Runge-Kutta) evaluates the derivative four times per step for higher accuracy, at roughly four times the cost. Implicit integration folds velocity-dependent forces like damping into the equations being solved, letting it tolerate much larger timesteps. MuJoCo defaults to semi-implicit Euler with implicit joint damping, though its documentation recommends switching most models to implicitfast. The integrator, together with the simulation timestep, determines whether a simulation stays stable or “blows up.”","example":"Setting integrator to implicitfast in a MuJoCo model's option element, while keeping the default 0.002-second timestep, is usually more stable than the default Euler integrator at a similar computational cost.","related":["Physics Engine","Simulation Timestep","Substeps","Simulation Instability","MuJoCo (Multi-Joint dynamics with Contact)","Constraint Solver"]},{"id":"simulation-instability","category":"sim","sec":1,"tier":2,"sources":[{"title":"MuJoCo Documentation: Computation - Numerical integration","url":"https://mujoco.readthedocs.io/en/stable/computation/index.html#numerical-integration"},{"title":"MuJoCo Documentation: XML Reference (option/flag autoreset)","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html#option-flag-autoreset"}],"as_of":"","related_ids":["simulation-timestep","numerical-integrator","constraint-solver","interpenetration","solver-iteration-count","mujoco"],"name":"Simulation Instability","alt":"仿真爆炸","abbr":"","aliases":["Blow-up","Numerical Instability","NaN Explosion","Simulation Blow-up"],"one_liner":"A physics simulation diverging numerically — the robot twitches violently, objects fly apart, and the state turns into NaN or huge numbers.","explanation":"Simulation instability (often called blow-up) is a common failure mode in physics simulation: the velocities and accelerations the integrator computes each step keep growing until they become NaN (not a number) or astronomically large, which shows up on screen as a robot shaking violently, parts flying off, or an object that penetrated another getting launched away. Common causes include too large a simulation timestep, joint PD gains (the stiffness and damping of position control) set too stiff, unreasonable mass or inertia values, objects that start out interpenetrating, or too few constraint-solver iterations. MuJoCo's documentation explicitly warns that too large a timestep makes simulation unstable; the engine checks every step whether acceleration is NaN or exceeds a limit, and if so, issues a warning and resets the simulation automatically by default. In massively parallel reinforcement learning, the NaN values from just one environment can contaminate an entire batch of training data, so problem environments usually need to be detected and reset individually.","example":"In MuJoCo, increasing the timestep while also setting joint stiffness very high can make a robot arm shake violently within a few steps and its state turn into NaN; MuJoCo reports a bad-acceleration (BADQACC) warning and automatically resets the simulation.","related":["Simulation Timestep","Numerical Integrator (Semi-implicit Euler / RK4)","Constraint Solver","Interpenetration","Solver Iteration Count (Position / Velocity Iterations)","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"deformable-body-simulation","category":"sim","sec":2,"tier":2,"sources":[{"title":"NVIDIA PhysX 5 Documentation: Soft Bodies","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/SoftBodies.html"},{"title":"SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation (arXiv 2011.07215)","url":"https://arxiv.org/abs/2011.07215"},{"title":"Genesis GitHub 仓库","url":"https://github.com/Genesis-Embodied-AI/Genesis"}],"as_of":"","related_ids":["finite-element-method","material-point-method","position-based-dynamics","cloth-simulation","deformable-object-manipulation","softgym-benchmarking-deep-reinforcement-learning-for-deforma"],"name":"Deformable-Body Simulation","alt":"软体仿真","abbr":"","aliases":["soft-body simulation","flexible-body simulation"],"one_liner":"Physics simulation of objects that change shape under force, like cloth, rope, dough, or liquid.","explanation":"Rigid-body simulation assumes an object's shape never changes, so a handful of numbers is enough to describe its pose; deformable-body simulation has to handle objects that change shape under force, with state made up of the positions of hundreds or thousands of nodes or particles, which is far more computationally and numerically demanding. Common methods include the finite element method (FEM, which cuts an object into a tetrahedral mesh and computes deformation from material parameters like Young's modulus and Poisson's ratio, well suited to elastic bodies), position-based dynamics (PBD, fast and stable, commonly used for cloth and rope), and the material point method (MPM, suited to large-deformation materials such as dough and sand). PhysX runs soft bodies with FEM on the GPU; Genesis integrates FEM, MPM, and PBD/SPH solvers together. This underlies research on manipulating flexible objects, such as folding laundry or kneading dough, though the gap to real materials is usually larger than for rigid-body simulation.","example":"SoftGym (CoRL 2020), built on the particle simulator NVIDIA FleX, provides tasks such as flattening cloth, folding cloth, straightening rope, and pouring water, used to test how well reinforcement learning algorithms handle deformable objects.","related":["Finite Element Method","Material Point Method","Position-Based Dynamics","Cloth Simulation","Deformable Object Manipulation","SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation"]},{"id":"cloth-simulation","category":"sim","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Cloth modeling","url":"https://en.wikipedia.org/wiki/Cloth_modeling"},{"title":"SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation (arXiv 2011.07215)","url":"https://arxiv.org/abs/2011.07215"},{"title":"MuJoCo Documentation: Modeling - Deformable objects","url":"https://mujoco.readthedocs.io/en/stable/modeling.html#deformable-objects"}],"as_of":"","related_ids":["deformable-body-simulation","garment-manipulation","deformable-object-manipulation","position-based-dynamics","finite-element-method","softgym-benchmarking-deep-reinforcement-learning-for-deforma"],"name":"Cloth Simulation","alt":"布料仿真","abbr":"","aliases":["Fabric Simulation"],"one_liner":"Computer simulation of how cloth stretches, bends, wrinkles, and collides — essential for research on manipulating clothing.","explanation":"Cloth simulation is a category of soft-body simulation that computer graphics has studied since the 1980s. Early geometric methods (such as Weil 1986) approximated wrinkle shapes using catenary curves and ignored dynamics entirely; physics-based methods model cloth as a mesh of point masses connected by springs, accounting for stretch, shear, bending, and gravity; more refined approaches use energy models or the finite element method. The difficulty is that cloth is very thin and extremely prone to self-collision and interpenetration, while being stiff along its stretch direction, so even a moderately large timestep can trigger numerical instability. Implementations commonly used in robotics research include the particle-based NVIDIA FleX and MuJoCo's flex deformable bodies (which can represent rope, cloth, and volumetric soft bodies). Real cloth's physical parameters are hard to measure accurately, so the sim-to-real gap for cloth tasks is usually larger than for rigid-body tasks.","example":"SoftGym (CoRL 2020), built on NVIDIA FleX, provides tasks such as flattening a cloth, folding it in half, and letting it settle flat on the ground, used to test reinforcement learning on deformable-object manipulation.","related":["Deformable-Body Simulation","Garment Manipulation","Deformable Object Manipulation","Position-Based Dynamics","Finite Element Method","SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation"]},{"id":"fluid-simulation","category":"sim","sec":2,"tier":3,"sources":[{"title":"Fluid animation - Wikipedia","url":"https://en.wikipedia.org/wiki/Fluid_simulation"},{"title":"SoftGym project page","url":"https://sites.google.com/view/softgym"},{"title":"Genesis Documentation: What is Genesis","url":"https://genesis-world.readthedocs.io/en/latest/user_guide/overview/what_is_genesis.html"}],"as_of":"","related_ids":["smoothed-particle-hydrodynamics","material-point-method","position-based-dynamics","deformable-body-simulation","softgym-benchmarking-deep-reinforcement-learning-for-deforma","genesis"],"name":"Fluid Simulation","alt":"流体仿真","abbr":"","aliases":["Liquid Simulation"],"one_liner":"Numerically simulating how fluids like water or smoke flow and experience force inside a computer.","explanation":"Fluid simulation means numerically approximating the equations that describe fluid motion (usually the Navier-Stokes equations) to compute how a liquid's or gas's velocity and pressure change over time. The branch chasing scientific precision is called computational fluid dynamics (CFD), used in engineering design for aircraft and piping; graphics and robot simulation instead prioritize speed and visual plausibility. Common approaches fall into three categories: grid-based methods (Eulerian, recording flow velocity on a fixed grid), particle-based methods (Lagrangian, such as smoothed-particle hydrodynamics, or SPH, which treats a fluid as a large swarm of particles), and hybrids of the two such as FLIP and the material point method. In embodied AI, tasks like pouring water, carrying a bowl of soup, or cleaning up a spill all need it; because of the large number of particles involved and the need to couple with the robot arm and container, the computational cost is far higher than rigid-body simulation, so a trade-off between accuracy and speed is usually necessary.","example":"The SoftGym benchmark's PourWater task (pouring all the water into a target cup) and TransportWater task (carrying a cup of water to a target location without spilling); Genesis simulates liquids with an SPH solver.","related":["Smoothed Particle Hydrodynamics","Material Point Method","Position-Based Dynamics","Deformable-Body Simulation","SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation","Genesis"]},{"id":"finite-element-method","category":"sim","sec":2,"tier":3,"sources":[{"title":"Finite element method - Wikipedia","url":"https://en.wikipedia.org/wiki/Finite_element_method"},{"title":"PhysX 5 Documentation: Soft Bodies","url":"https://nvidia-omniverse.github.io/PhysX/physx/5.4.1/docs/SoftBodies.html"},{"title":"Genesis Documentation: What is Genesis","url":"https://genesis-world.readthedocs.io/en/latest/user_guide/overview/what_is_genesis.html"}],"as_of":"","related_ids":["deformable-body-simulation","material-point-method","position-based-dynamics","deformable-object-manipulation","physics-engine","genesis"],"name":"Finite Element Method","alt":"有限元法","abbr":"FEM","aliases":["FEM","Finite Element Analysis","FEA"],"one_liner":"A numerical method that cuts an object into many small elements, solves each approximately, and assembles them into the whole.","explanation":"The finite element method is a general numerical technique for solving partial differential equations: a continuous object or region is divided into a large number of simple small pieces (elements, commonly triangles or tetrahedra), each approximated with a simple function, then assembled into an overall system of equations to solve. It originated in structural-mechanics research in the 1940s–50s, with foundational work by Courant and others, and later became the standard tool for engineering analysis of structural strength, heat transfer, fluids, and electromagnetics. In embodied AI, FEM is mainly used for soft-body simulation: rigid-body simulation treats an object as one undeformable whole, but tasks like grasping a sponge, squeezing a soft package, or designing a soft gripper require computing how the object deforms and how much internal force it experiences, which needs FEM. The cost is heavy computation — a finer mesh is more accurate but also slower, forcing a trade-off between accuracy and speed.","example":"PhysX 5 (the physics backend behind Isaac Sim) simulates soft bodies with FEM using two tetrahedral meshes: a coarser one for computing deformation and a surface-fitted one for collision, and it only runs on the GPU; Genesis also has a built-in FEM solver.","related":["Deformable-Body Simulation","Material Point Method","Position-Based Dynamics","Deformable Object Manipulation","Physics Engine","Genesis"]},{"id":"position-based-dynamics","category":"sim","sec":2,"tier":3,"sources":[{"title":"Position Based Dynamics (Müller et al., VRIPHYS 2006)","url":"https://matthias-research.github.io/pages/publications/posBasedDyn.pdf"},{"title":"XPBD: Position-Based Simulation of Compliant Constrained Dynamics (Macklin et al., MIG 2016)","url":"https://matthias-research.github.io/pages/publications/XPBD.pdf"},{"title":"Newton Physics 文档：Solvers（SolverXPBD）","url":"https://newton-physics.github.io/newton/latest/api/newton_solvers.html"}],"as_of":"2026-09","related_ids":["cloth-simulation","deformable-body-simulation","constraint-solver","physics-engine","newton-physics-engine","projected-gauss-seidel"],"name":"Position-Based Dynamics","alt":"基于位置的动力学","abbr":"PBD","aliases":["PBD","XPBD","Extended Position-Based Dynamics"],"one_liner":"A simulation method that satisfies constraints by directly correcting object positions, fast and stable, popular for cloth and soft bodies.","explanation":"Position-Based Dynamics was proposed by Matthias Müller and colleagues (then at the physics-engine company AGEIA) at the VRIPHYS workshop in 2006. Conventional simulation first computes forces, then integrates acceleration into velocity and position; with a large timestep this easily overshoots or even diverges. PBD skips the force-and-velocity layer entirely: it predicts each particle's new position from inertia, then checks each constraint in turn — such as “the distance between two points must stay fixed” or “this point cannot pass through the ground” — and directly projects any violating point back to a valid position; after a few iterations, velocity is derived from the position change. It is stable, controllable, and handles collisions simply, and was first used for real-time cloth in games. Its drawback is that material stiffness drifts with the number of iterations and the timestep; in 2016, NVIDIA's Macklin and colleagues introduced XPBD, which adds compliance (the inverse of stiffness) and Lagrange multipliers so stiffness no longer depends on either, while also allowing constraint forces to be estimated. NVIDIA's Newton physics engine includes an XPBD solver that can simulate both rigid and soft bodies.","example":"To simulate a tablecloth, discretize the cloth into a grid of particles, add a “fixed distance” constraint between neighboring particles, and each step pull any over-stretched edge back — after a few iterations, the cloth naturally drapes instead of stretching indefinitely.","related":["Cloth Simulation","Deformable-Body Simulation","Constraint Solver","Physics Engine","Newton Physics Engine","Projected Gauss-Seidel"]},{"id":"material-point-method","category":"sim","sec":2,"tier":3,"sources":[{"title":"Material point method - Wikipedia","url":"https://en.wikipedia.org/wiki/Material_point_method"},{"title":"PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable Physics (arXiv 2104.03311)","url":"https://arxiv.org/abs/2104.03311"},{"title":"Genesis 文档：What is Genesis","url":"https://genesis-world.readthedocs.io/en/latest/user_guide/overview/what_is_genesis.html"}],"as_of":"","related_ids":["deformable-body-simulation","finite-element-method","smoothed-particle-hydrodynamics","differentiable-simulation","genesis","deformable-object-manipulation"],"name":"Material Point Method","alt":"物质点法","abbr":"MPM","aliases":["MPM"],"one_liner":"A continuum simulation method where particles carry the material's state while forces are computed on a background grid.","explanation":"The material point method is a hybrid Eulerian-Lagrangian numerical method: an object is discretized into a large number of “material points,” with each particle carrying mass, velocity, stress, and the rest of its full state; at each step, this information is mapped onto a fixed background grid to solve the momentum equation, and the results are mapped back onto the particles. It traces back to the particle-in-cell (PIC) method Harlow introduced in 1957, and was developed into its current form starting in 1993 by Sulsky and colleagues. Compared with the finite element method, it avoids repeatedly re-meshing, making it well suited to materials that undergo large deformation, fracture, or flow, such as snow, sand, mud, and plasticine — Disney's Frozen used it to simulate snow. The cost is heavy memory and computation. In embodied AI it's used for deformable-object manipulation simulation, and simulators such as Genesis have a built-in MPM solver.","example":"The PlasticineLab benchmark uses a differentiable MLS-MPM (Moving Least Squares Material Point Method) to simulate plasticine, letting an agent learn to pinch and mold it into a target shape.","related":["Deformable-Body Simulation","Finite Element Method","Smoothed Particle Hydrodynamics","Differentiable Simulation","Genesis","Deformable Object Manipulation"]},{"id":"smoothed-particle-hydrodynamics","category":"sim","sec":2,"tier":3,"sources":[{"title":"Wikipedia - Smoothed-particle hydrodynamics","url":"https://en.wikipedia.org/wiki/Smoothed-particle_hydrodynamics"},{"title":"GitHub - Genesis-Embodied-AI/Genesis","url":"https://github.com/Genesis-Embodied-AI/Genesis"}],"as_of":"","related_ids":["fluid-simulation","position-based-dynamics","material-point-method","finite-element-method","genesis","deformable-body-simulation"],"name":"Smoothed Particle Hydrodynamics","alt":"光滑粒子流体动力学","abbr":"SPH","aliases":["SPH"],"one_liner":"A meshless simulation method that breaks a fluid into particles and computes motion by weighting over nearby particles.","explanation":"Smoothed Particle Hydrodynamics was proposed by Gingold, Monaghan, and Lucy in 1977, originally to model galaxy and star formation in astrophysics. Rather than dividing space into a grid, it treats a fluid as a swarm of particles that move with the material (a Lagrangian method, meaning it follows material points rather than fixed locations): each particle's density, pressure, and other physical quantities are computed as a weighted sum over neighboring particles within a “smoothing length,” using a kernel function, and the resulting forces then drive the particle forward. Because it needs no mesh, it is naturally suited to free surfaces, splashing, pouring, and other flows with drastic shape changes, and it is also widely used for fluid effects in computer graphics and games. In embodied AI it is used when a robot needs to learn to pour water or scoop liquid — the Genesis simulator, for example, has a built-in SPH solver. The cost is that computation grows heavy as particle counts rise, and precisely maintaining a liquid's incompressibility is not easy.","example":"In Genesis, represent the water in a cup as SPH particles and train a robot arm to pour it into another cup, penalizing the policy for how many particles spill out.","related":["Fluid Simulation","Position-Based Dynamics","Material Point Method","Finite Element Method","Genesis","Deformable-Body Simulation"]},{"id":"discrete-element-method","category":"sim","sec":2,"tier":3,"sources":[{"title":"Discrete element method - Wikipedia","url":"https://en.wikipedia.org/wiki/Discrete_element_method"}],"as_of":"","related_ids":["material-point-method","finite-element-method","deformable-body-simulation","physics-engine","smoothed-particle-hydrodynamics","contact-model"],"name":"Discrete Element Method","alt":"离散元法","abbr":"DEM","aliases":["DEM","Distinct Element Method"],"one_liner":"A simulation method that treats a material as a large number of independent particles and computes each one's forces and motion.","explanation":"The discrete element method is a class of numerical simulation methods that treat a material as a large number of independent small particles, computing each particle's motion individually. Peter Cundall introduced an early version in 1971, formally published with Strack in 1979. At each timestep, the method first finds which particles are in contact with each other, computes collision and friction forces from a contact model, adds gravity, and then numerically integrates to update every particle's position, velocity, and rotation. Unlike the finite element method, which treats a material as a continuum, DEM is naturally suited to granular materials such as sand, gravel, grain, and powder, and is widely used in mining, pharmaceuticals, and agriculture. Its drawback is heavy computation, which slows further as particle count grows; large-scale use now generally relies on GPU parallelism to handle millions of particles. In embodied AI, it can be used to simulate a robot walking on sand or gravel, or digging and scooping granular material; the material point method (MPM) can handle similar problems too.","example":"DEM is used to simulate a patch of sand and let a quadruped robot's foot step into it, computing the force on each grain of sand as it gets pushed aside, to study how the robot sinks and slips on the beach.","related":["Material Point Method","Finite Element Method","Deformable-Body Simulation","Physics Engine","Smoothed Particle Hydrodynamics","Contact Model"]},{"id":"incremental-potential-contact","category":"sim","sec":2,"tier":3,"sources":[{"title":"Incremental Potential Contact project page","url":"https://ipc-sim.github.io/"},{"title":"ipc-sim (GitHub organization)","url":"https://github.com/ipc-sim"}],"as_of":"","related_ids":["contact-model","interpenetration","continuous-collision-detection","deformable-body-simulation","finite-element-method","cloth-simulation"],"name":"Incremental Potential Contact","alt":"增量势接触","abbr":"IPC","aliases":["IPC","Barrier-based Contact"],"one_liner":"A contact-simulation algorithm that uses barrier potentials to guarantee objects never interpenetrate, well suited to large soft-body deformation.","explanation":"IPC was proposed by Minchen Li and colleagues at the University of Pennsylvania, Adobe Research, and NYU, published at SIGGRAPH 2020 (ACM TOG). Traditional physics engines often allow objects to interpenetrate slightly before pushing them apart, which makes interpenetration or numerical blow-up likely once the timestep grows large. IPC instead formulates every implicit timestep as an optimization problem, adding a barrier potential that grows sharply as an object's distance to another approaches zero, combined with a line search that incorporates continuous collision detection, guaranteeing the entire trajectory has no intersections and no mesh elements flip — regardless of material, timestep size, or amount of deformation — and it also supports friction. The cost is heavy computation, running far slower than rigid-body engines like MuJoCo. Later extensions include C-IPC for cloth and rods, a rigid-body version called rigid-ipc, and the IPC Toolkit, which can be embedded in other simulators.","example":"In the paper's demonstrations, IPC handled scenes with up to about 498,000 contacts and 2.3 million tetrahedra within a single timestep, remaining interpenetration-free across timesteps ranging from 2×10⁻⁵ seconds to 2 seconds.","related":["Contact Model","Interpenetration","Continuous Collision Detection","Deformable-Body Simulation","Finite Element Method","Cloth Simulation"]},{"id":"differentiable-simulation","category":"sim","sec":2,"tier":2,"sources":[{"title":"DiffTaichi: Differentiable Programming for Physical Simulation (arXiv 1910.00935)","url":"https://arxiv.org/abs/1910.00935"},{"title":"A Review of Differentiable Simulators (arXiv 2407.05560)","url":"https://arxiv.org/abs/2407.05560"},{"title":"Do Differentiable Simulators Give Better Policy Gradients? (arXiv 2202.00817)","url":"https://arxiv.org/abs/2202.00817"}],"as_of":"","related_ids":["physics-engine","mujoco-xla","brax","nvidia-warp","trajectory-optimization","system-identification"],"name":"Differentiable Simulation","alt":"可微仿真","abbr":"","aliases":["differentiable physics","differentiable physics engine"],"one_liner":"A simulator whose output can be differentiated, so gradient descent can optimize actions or physical parameters directly.","explanation":"An ordinary simulator only computes forward: given this action, what happens next. A differentiable simulator can also use automatic differentiation to compute, backward, the gradient of the result with respect to the action, the initial state, or physical parameters such as mass and friction. That makes it possible to directly optimize a control sequence, a policy network, or an estimate of physical parameters using gradient descent, the same way a neural network is trained, instead of estimating a gradient indirectly through large amounts of trial and error the way reinforcement learning does. Notable examples include DiffTaichi (ICLR 2020), Google's Brax, and MJX, the JAX version of MuJoCo (its Warp backend does not support automatic differentiation). The hard part is contact: collisions introduce sudden changes in the dynamics, so gradients can be huge, noisy, or zero; a 2022 ICML study by Suh and colleagues found that stiffness and discontinuity undermine the usefulness of this kind of first-order gradient.","example":"The DiffTaichi paper implements 10 differentiable simulators, and using them to optimize neural-network controllers typically converges within a few dozen iterations.","related":["Physics Engine","MuJoCo XLA","Brax","NVIDIA Warp","Trajectory Optimization","System Identification"]},{"id":"rendering","category":"sim","sec":3,"tier":2,"sources":[{"title":"Wikipedia: Rendering (computer graphics)","url":"https://en.wikipedia.org/wiki/Rendering_(computer_graphics)"},{"title":"ManiSkill3: GPU Parallelized Robotics Simulation and Rendering (arXiv 2410.00425)","url":"https://arxiv.org/abs/2410.00425"}],"as_of":"","related_ids":["rasterization","ray-tracing","path-tracing","photorealistic-rendering","batched-rendering","headless-mode"],"name":"Rendering","alt":"渲染","abbr":"","aliases":["Rendering Engine","Renderer"],"one_liner":"The process of computing a camera image from a 3D scene; it produces all of a simulator's camera-based sensor data.","explanation":"Rendering is a computer-graphics term for computing a 2D image from 3D models, materials, lighting, and camera parameters. The main methods are rasterization (projecting triangles onto the screen and shading pixel by pixel, which is fast and used by games and most real-time simulators) and ray tracing or path tracing (simulating how light actually travels, which is more realistic but slower). In robot simulation, the physics engine handles how objects move while the renderer handles “taking the picture”: producing RGB images, depth maps, and segmentation masks as camera observations for training and evaluating vision policies, as well as for visualizations meant for humans to watch. Rendering quality determines the size of the visual sim-to-real gap, while rendering speed is often the bottleneck for vision-based reinforcement learning, which is why ManiSkill3, Isaac Lab, and others all implement batched GPU rendering. Running a simulation without opening a graphical window is called headless mode.","example":"ManiSkill3 runs both physics simulation and rendering in parallel on the GPU; its paper reports that simulation with rendering reaches over 30,000 frames per second on its benchmark environments, 10 to 1000 times faster than other platforms and using 2 to 3 times less GPU memory.","related":["Rasterization","Ray Tracing","Path Tracing","Photorealistic Rendering","Batched Rendering","Headless Mode"]},{"id":"rasterization","category":"sim","sec":3,"tier":3,"sources":[{"title":"NVIDIA Blog: What's the Difference Between Ray Tracing and Rasterization?","url":"https://blogs.nvidia.com/blog/whats-difference-between-ray-tracing-rasterization/"},{"title":"Wikipedia: Rasterisation","url":"https://en.wikipedia.org/wiki/Rasterisation"},{"title":"ManiSkill 文档：Sensors / Cameras（shader packs）","url":"https://maniskill.readthedocs.io/en/latest/user_guide/concepts/sensors.html"}],"as_of":"","related_ids":["rendering","ray-tracing","batched-rendering","photorealistic-rendering","sim-to-real-gap","sensor-simulation"],"name":"Rasterization","alt":"光栅化","abbr":"","aliases":["Rasterized Rendering","Rasterisation"],"one_liner":"A rendering method that projects 3D triangles onto the screen and shades them pixel by pixel; fast, and standard for real time.","explanation":"Rendering (turning a 3D scene into a 2D image) has two fundamental approaches, the other being ray tracing. 3D objects are usually built from large numbers of triangles; rasterization projects each triangle's vertices onto the screen, figures out which pixels it covers, then shades those pixels based on texture and lighting, using a depth buffer (z-buffer, which records each pixel's closest distance from the camera) to decide what occludes what. GPUs have dedicated hardware pipelines for this process, making it extremely fast, which is why games and most real-time 3D engines rely on it. The trade-off is that shadows, reflections, refraction, and indirect lighting all need approximation tricks, so it looks less realistic than ray tracing. In embodied simulation, MuJoCo's built-in OpenGL renderer and ManiSkill's non-ray-traced shaders are both rasterization, well suited to quickly rendering images for large numbers of parallel environments when training visual policies; ray tracing is used instead when photorealism is needed to narrow the visual sim-to-real gap.","example":"When training a vision-based policy to grasp a block, rasterization can render wrist-camera images for hundreds of parallel simulated environments at once fast enough for training, though a glass cup's refraction and metal's reflections will look noticeably unrealistic.","related":["Rendering","Ray Tracing","Batched Rendering","Photorealistic Rendering","Sim-to-Real Gap (Reality Gap)","Sensor Simulation"]},{"id":"ray-tracing","category":"sim","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Ray tracing (graphics)","url":"https://en.wikipedia.org/wiki/Ray_tracing_(graphics)"},{"title":"Omniverse 文档：RTX Renderer（Real-Time 2.0 / Interactive Path Tracing）","url":"https://docs.omniverse.nvidia.com/materials-and-rendering/latest/rtx-renderer.html"},{"title":"NVIDIA Blog: What's the Difference Between Ray Tracing and Rasterization?","url":"https://blogs.nvidia.com/blog/whats-difference-between-ray-tracing-rasterization/"}],"as_of":"2026-09","related_ids":["rendering","rasterization","path-tracing","photorealistic-rendering","nvidia-isaac-sim","sensor-simulation"],"name":"Ray Tracing","alt":"光线追踪","abbr":"","aliases":["Ray-Traced Rendering","RTX Rendering"],"one_liner":"A rendering method that traces rays backward from the camera through reflections and refractions; highly realistic but slow.","explanation":"Ray tracing is one of the two fundamental rendering approaches. It casts a ray from the camera through each pixel, finds which object it hits first, and continues tracing reflection, refraction, and shadow rays toward light sources, which naturally produces effects like mirror reflections, transparent refraction, and soft shadows. Path tracing is its more advanced form, letting rays bounce randomly many times to approximate global illumination — the most realistic and slowest option. Because it is computationally expensive, ray tracing was long confined to offline rendering such as film; real-time ray tracing only became practical after NVIDIA launched GeForce RTX graphics cards with dedicated RT cores (hardware that accelerates ray-surface intersection) in 2018, which is where the term “RTX rendering” comes from. The Omniverse RTX renderer used by Isaac Sim is based on path tracing, and sensor simulation for things like RTX lidar depends on it too. In embodied AI, ray tracing generates photorealistic synthetic data and simulates transparent and reflective objects to narrow the visual sim-to-real gap, at the cost of being much slower than rasterization.","example":"Rendering a glass filled with water in path-tracing mode inside Isaac Sim produces near-camera-realistic refraction and highlights, useful for training a grasping model to recognize transparent objects.","related":["Rendering","Rasterization","Path Tracing","Photorealistic Rendering","NVIDIA Isaac Sim","Sensor Simulation"]},{"id":"path-tracing","category":"sim","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Path tracing","url":"https://en.wikipedia.org/wiki/Path_tracing"},{"title":"Omniverse Docs: RTX Interactive (Path Tracing) Mode","url":"https://docs.omniverse.nvidia.com/materials-and-rendering/latest/rtx-renderer_pt.html"},{"title":"Physically Based Rendering: From Theory to Implementation, 4th ed. (online)","url":"https://www.pbr-book.org/4ed/contents"}],"as_of":"2026-09","related_ids":["ray-tracing","rasterization","physically-based-rendering","photorealistic-rendering","rendering","nvidia-isaac-sim"],"name":"Path Tracing","alt":"路径追踪","abbr":"","aliases":["Path-Traced Rendering","Monte Carlo Ray Tracing"],"one_liner":"A rendering method that simulates light bouncing many times through a scene using randomly sampled rays, producing physically realistic images.","explanation":"Path tracing was introduced by Jim Kajiya in 1986, in the same paper that gave the rendering equation (the integral equation describing how light propagates through a scene) and approximated it with Monte Carlo integration (averaging over random samples). It works by casting a ray from the camera through each pixel, and whenever it hits a surface, randomly choosing a new direction to keep bouncing based on the material, accumulating lighting along the way — which naturally produces effects like indirect illumination, soft shadows, and caustics. The cost is that images are noisy when too few samples are used, requiring either very heavy sampling or a neural denoiser to clean up. It belongs to the ray-tracing family and is much slower than rasterization, traditionally reserved for offline rendering such as film. In robot simulation, the Omniverse RTX renderer used by Isaac Sim offers a path-tracing mode that produces more photorealistic synthetic images, at the cost of being slower than its real-time mode.","example":"Omniverse's RTX Interactive (Path Tracing) mode defaults to sampling 1 ray per pixel per frame, accumulating up to 512 samples, with the OptiX denoiser enabled by default.","related":["Ray Tracing","Rasterization","Physically Based Rendering","Photorealistic Rendering","Rendering","NVIDIA Isaac Sim"]},{"id":"physically-based-rendering","category":"sim","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Physically based rendering","url":"https://en.wikipedia.org/wiki/Physically_based_rendering"},{"title":"LearnOpenGL: PBR Theory","url":"https://learnopengl.com/PBR/Theory"},{"title":"Physically Based Rendering: From Theory to Implementation, 4th ed. (online)","url":"https://www.pbr-book.org/4ed/contents"}],"as_of":"","related_ids":["path-tracing","ray-tracing","rendering","photorealistic-rendering","simready-assets","sim-to-real-gap"],"name":"Physically Based Rendering","alt":"基于物理的渲染","abbr":"PBR","aliases":["PBR","PBR Materials","Metallic-Roughness Workflow"],"one_liner":"A rendering approach that models light and materials according to real optics, so objects look believable under any lighting.","explanation":"PBR refers to rendering methods that describe light sources and surface materials according to real-world optical principles. The term emerged in the 1990s, spread through Pharr and colleagues' textbook of the same name, and was later pushed into real-time rendering by Disney and Epic Games; it is now the standard approach in Unreal, Unity, and Blender. It generally satisfies three conditions: surface roughness is described with a microfacet model; energy is conserved, so a surface never reflects more light than it received; and reflectance follows a physically valid BRDF (the function describing how a surface reflects light). Materials are typically described with maps for base color (albedo), normal, metallic, and roughness. When a simulator uses PBR materials, the reflections and highlights in camera views come much closer to reality, which is one precondition for narrowing the visual sim-to-real gap.","example":"For the same stainless-steel pot, setting metallic to 1 and lowering roughness renders sharp highlights and clear environment reflections; raising roughness instead gives it a brushed, matte look.","related":["Path Tracing","Ray Tracing","Rendering","Photorealistic Rendering","SimReady Assets","Sim-to-Real Gap (Reality Gap)"]},{"id":"photorealistic-rendering","category":"sim","sec":3,"tier":2,"sources":[{"title":"Wikipedia: Rendering (computer graphics)","url":"https://en.wikipedia.org/wiki/Rendering_(computer_graphics)"},{"title":"NVIDIA Omniverse: RTX Renderer","url":"https://docs.omniverse.nvidia.com/materials-and-rendering/latest/rtx-renderer.html"},{"title":"GraspVLA (arXiv 2505.03233)","url":"https://arxiv.org/abs/2505.03233"}],"as_of":"2025-05","related_ids":["rendering","path-tracing","physically-based-rendering","sim-to-real-gap","gaussian-splatting-based-simulation","synthetic-data"],"name":"Photorealistic Rendering","alt":"照片级真实感渲染","abbr":"","aliases":["Photo-realistic Rendering"],"one_liner":"Computing an image by simulating real optical behavior, so a simulated picture looks like a photo from a real camera.","explanation":"Photorealistic rendering simulates how light travels, reflects, and refracts through a scene so the resulting image approaches a real photograph, commonly using path tracing (a Monte Carlo method that randomly samples many light paths to compute global illumination) combined with physically based materials. In embodied AI it is mainly used to narrow the visual sim-to-real gap: the more a vision policy's simulated images resemble a real camera's output, the less likely it is to fail after deployment simply because the imagery looks different; it also underlies synthetic training-data generation and camera-sensor simulation. The cost is speed — it is computationally expensive, so massively parallel training often falls back to faster rasterization or lower image quality. NVIDIA's Omniverse / Isaac Sim RTX renderer offers both a real-time mode and a path-traced mode; a separate approach reconstructs a scene from real photos with 3D Gaussian Splatting to get photorealistic imagery directly.","example":"Galbot's GraspVLA used ray-traced rendering in Isaac Sim to generate SynGrasp-1B, a synthetic grasping dataset of about 1 billion frames, while randomizing point lights, directional lights, ambient light, and roughly 2,000 tabletop, floor, and wall textures — a run that took about 10 days on 160 RTX 4090 GPUs.","related":["Rendering","Path Tracing","Physically Based Rendering","Sim-to-Real Gap (Reality Gap)","Gaussian Splatting-based Simulation","Synthetic Data"]},{"id":"batched-rendering","category":"sim","sec":3,"tier":3,"sources":[{"title":"Isaac Lab Documentation: Camera (Tiled Rendering)","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/sensors/camera.html"},{"title":"ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (arXiv 2410.00425)","url":"https://arxiv.org/abs/2410.00425"},{"title":"MuJoCo Playground (arXiv 2502.08844)","url":"https://arxiv.org/abs/2502.08844"}],"as_of":"2026-09","related_ids":["gpu-accelerated-parallel-simulation","vectorized-environments","rendering","nvidia-isaac-lab","maniskill","massively-parallel-reinforcement-learning"],"name":"Batched Rendering","alt":"批量渲染","abbr":"","aliases":["Tiled Rendering","Tiled Camera","Batch Renderer"],"one_liner":"Rendering camera images for hundreds or thousands of parallel environments in one pass, feeding vision policies at high speed.","explanation":"GPU-parallel simulation can run thousands of environments at once, but if each environment's camera were rendered and copied separately, image bandwidth would become the bottleneck: Isaac Lab's documentation estimates that one 800×600 floating-point image is nearly 2 MB, so 60 frames per second alone needs 120 MB/s, before multiplying by the number of cameras and environments. Batched rendering merges the processing of the same camera across all environments. Isaac Lab's TiledCamera (supported from Isaac Sim 4.2.0) is a good example: all the cloned cameras share a single render product, and each environment's image is tiled into one large combined image that can be handed to the training code in a single sync. ManiSkill3's SAPIEN-based parallel rendering and MuJoCo Playground's built-in batched renderer offer similar capability. This is what lets pixel-based vision-motor policies be trained with massively parallel reinforcement learning too.","example":"In Isaac Lab, setting a TiledCameraCfg's camera path to /World/envs/env_.*/Camera retrieves an 80×80 RGB image for every environment in one call; the official guidance suggests around 512 cameras on an RTX 4090-class GPU.","related":["GPU-Accelerated Parallel Simulation","Vectorized Environments","Rendering","NVIDIA Isaac Lab","ManiSkill","Massively Parallel Reinforcement Learning"]},{"id":"sensor-simulation","category":"sim","sec":3,"tier":2,"sources":[{"title":"Isaac Sim 文档：Sensors","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/sensors/index.html"},{"title":"Choi et al., On the use of simulation in robotics (PNAS 2021)","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7817170/"}],"as_of":"","related_ids":["rendering","ray-tracing","tactile-simulation","sim-to-real-gap","simulation-fidelity","visual-randomization"],"name":"Sensor Simulation","alt":"传感器仿真","abbr":"","aliases":["Camera/Lidar/IMU Simulation","Sensor Modeling"],"one_liner":"Generating simulated readings from cameras, lidar, IMUs, force sensors, and other sensors inside a simulator.","explanation":"Sensor simulation means computing what each sensor “should read” based on the state of a virtual scene. Camera images and depth maps come from rendering; lidar and radar cast virtual rays or use ray tracing to compute distance; IMU (inertial measurement unit) readings come from a rigid body's acceleration and angular velocity; and contact forces and joint torques come straight from the physics solver. A good implementation also adds the flaws real sensors have, such as noise, distortion, latency, and missing depth values. NVIDIA Isaac Sim groups sensors into camera-based sensors, RTX ray-traced sensors (lidar, radar, acoustic), and physics-based sensors (contact, IMU, and so on). These are exactly the readings a policy sees, so if they diverge too much from the real sensor, they become a source of sim-to-real gap.","example":"Isaac Sim provides ray-traced RTX lidar and radar simulation; robosuite v1.5 has also added sensor models.","related":["Rendering","Ray Tracing","Tactile Simulation","Sim-to-Real Gap (Reality Gap)","Simulation Fidelity","Visual Randomization"]},{"id":"tactile-simulation","category":"sim","sec":3,"tier":3,"sources":[{"title":"TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors (arXiv 2012.08456)","url":"https://arxiv.org/abs/2012.08456"},{"title":"Taxim: An Example-based Simulation Model for GelSight Tactile Sensors (arXiv 2109.04027)","url":"https://arxiv.org/abs/2109.04027"},{"title":"TacSL: A Library for Visuotactile Sensor Simulation and Learning (arXiv 2408.06506)","url":"https://arxiv.org/abs/2408.06506"}],"as_of":"","related_ids":["vision-based-tactile-sensor","tacto-a-fast-flexible-and-open-source-simulator-for-high-res","taxim-an-example-based-simulation-model-for-gelsight-tactile","tacsl-a-library-for-visuotactile-sensor-simulation-and-learn","sensor-simulation","tactile-image"],"name":"Tactile Simulation","alt":"触觉仿真","abbr":"","aliases":["Tactile Sensor Simulation"],"one_liner":"Simulating what a tactile sensor would read inside a physics simulator, so tactile-equipped policies can be trained in sim.","explanation":"Tactile simulation means generating the output a tactile sensor should produce inside a physics simulator — for instance, the tactile image from a visuotactile sensor (a sensor with a built-in camera that photographs the deformation of an elastic gel pad, such as GelSight or DIGIT), the displacement of markers embedded in the gel, or the pressure map from a piezoresistive array. The hard part is handling contact, soft-material deformation, lighting, and image formation all at once — the result has to look real and still be fast to compute. Common approaches fall into three groups: rendering an image directly from contact geometry, as in Meta's open-source TACTO; calibrating a lookup table from a small number of real samples, as in Carnegie Mellon's Taxim; and computing contact force fields and images in parallel on the GPU, as in NVIDIA's TacSL. It lets sim-to-real transfer extend to contact-rich tasks like peg insertion and screwing, though a gap between simulated and real sensor readings remains, usually addressed with domain randomization or a small amount of real-robot data.","example":"The TACTO paper used simulation to generate tactile data for 1 million grasps, training a model to predict whether a given grasp would be stable.","related":["Vision-Based Tactile Sensor","TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors","Taxim: An Example-based Simulation Model for GelSight Tactile Sensors","TacSL: A Library for Visuotactile Sensor Simulation and Learning","Sensor Simulation","Tactile Image"]},{"id":"tacto-a-fast-flexible-and-open-source-simulator-for-high-res","category":"sim","sec":3,"tier":3,"sources":[{"title":"TACTO (arXiv 2012.08456)","url":"https://arxiv.org/abs/2012.08456"},{"title":"facebookresearch/tacto GitHub 仓库","url":"https://github.com/facebookresearch/tacto"}],"as_of":"2022-02","related_ids":["tactile-simulation","digit","pybullet","taxim-an-example-based-simulation-model-for-gelsight-tactile","vision-based-tactile-sensor"],"name":"TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors","alt":"TACTO","abbr":"","aliases":["TACTO Simulator"],"one_liner":"Meta's open-source visuotactile simulator, using PyBullet to render tactile images for sensors like DIGIT.","explanation":"TACTO was developed by Shaoxiong Wang, Mike Lambeta, Po-Wei Chou, and Roberto Calandra, with code released in Meta's facebookresearch repository; the paper appeared in IEEE RA-L and was presented at ICRA 2022. It uses the PyRender renderer to generate hundreds of high-resolution tactile images per second from the contact geometry between an object and an elastic gel pad, and provides an interface that connects to the PyBullet physics engine, shipping with built-in models and configurations for the DIGIT and OmniTact sensors. The authors note explicitly that TACTO does not itself model the physical accuracy of contact dynamics (deformation, friction) — that is left to existing physics engines. It is open source under the MIT license and installs directly with pip.","example":"The paper used TACTO to simulate 1 million grasps for training a grasp-stability predictor, and also demonstrated a task using touch to control a rolling marble, along with preliminary sim-to-real transfer.","related":["Tactile Simulation","DIGIT","PyBullet","Taxim: An Example-based Simulation Model for GelSight Tactile Sensors","Vision-Based Tactile Sensor"]},{"id":"taxim-an-example-based-simulation-model-for-gelsight-tactile","category":"sim","sec":3,"tier":3,"sources":[{"title":"Taxim: An Example-based Simulation Model for GelSight Tactile Sensors (arXiv 2109.04027)","url":"https://arxiv.org/abs/2109.04027"},{"title":"Robo-Touch/Taxim GitHub 仓库","url":"https://github.com/Robo-Touch/Taxim"}],"as_of":"2021-12","related_ids":["tactile-simulation","gelsight","tacto-a-fast-flexible-and-open-source-simulator-for-high-res","marker-tracking","photometric-stereo"],"name":"Taxim: An Example-based Simulation Model for GelSight Tactile Sensors","alt":"Taxim","abbr":"","aliases":["Taxim Simulation Model"],"one_liner":"An example-based simulation model for GelSight visuotactile sensors, calibrated from a small amount of real data.","explanation":"Taxim was proposed in 2021 by Zilin Si and Wenzhen Yuan at Carnegie Mellon University's RoboTouch lab, built specifically to simulate GelSight-type visuotactile sensors. “Example-based” means it does not model the optical path from scratch; instead it calibrates a polynomial lookup table from real samples captured on the actual sensor, mapping gel-pad deformation geometry directly to camera pixel brightness, while the motion of markers printed on the gel is computed separately by superimposing elastic-deformation theory. Calibration needs fewer than 100 real data points, so it is easy to port to different GelSight models. The paper reports lower per-pixel intensity error than prior methods, and it runs on a CPU. It is often compared with TACTO: TACTO renders images with a general-purpose renderer, while Taxim calibrates its output from measured data.","example":"Give Taxim an object's point cloud and indentation depth, and it returns the corresponding GelSight tactile image; given loads along the x, y, and z directions, it can also return the resulting displacement field of the gel's surface markers.","related":["Tactile Simulation","GelSight","TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors","Marker Tracking","Photometric Stereo"]},{"id":"tacsl-a-library-for-visuotactile-sensor-simulation-and-learn","category":"sim","sec":3,"tier":3,"sources":[{"title":"TacSL: A Library for Visuotactile Sensor Simulation and Learning (arXiv 2408.06506)","url":"https://arxiv.org/abs/2408.06506"},{"title":"TacSL 项目主页","url":"https://iakinola23.github.io/tacsl/"},{"title":"Isaac Lab 文档：Visuo-Tactile Sensor","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/sensors/visuo_tactile_sensor.html"}],"as_of":"2026-09","related_ids":["tactile-simulation","vision-based-tactile-sensor","gelsight","nvidia-isaac-lab","asymmetric-actor-critic","sim-to-real-transfer"],"name":"TacSL: A Library for Visuotactile Sensor Simulation and Learning","alt":"TacSL 视触觉仿真库","abbr":"","aliases":["TacSL"],"one_liner":"NVIDIA's open-source GPU library for simulating vision-based tactile sensors and training policies that use them.","explanation":"TacSL was proposed by NVIDIA's Iretiayo Akinola, Yashraj Narang, and colleagues, posted to arXiv in August 2024, and later published in IEEE Transactions on Robotics. It uses the GPU inside NVIDIA's Isaac simulators to simultaneously generate visuotactile images (the deformation image an internal camera would capture on a sensor's gel pad, as with GelSight-style sensors) and contact force distributions, which the paper reports as more than 200 times faster than the best prior method — necessary for training tactile-input policies at large parallel scale. The library also includes contact-rich training environments, such as peg-in-hole insertion, and an asymmetric actor-critic distillation (AACD) algorithm for learning tactile policies and transferring them to real robots. The code first shipped inside the IsaacGymEnvs repository, and today Isaac Lab's visuotactile sensor module is implemented on top of TacSL.","example":"Isaac Lab's visuotactile sensors offer GelSight R1.5 and GelSight Mini configurations, computing penalty-based normal and shear forces via signed-distance-field queries while also outputting an RGB tactile image, usable directly as an observation for reinforcement or imitation learning.","related":["Tactile Simulation","Vision-Based Tactile Sensor","GelSight","NVIDIA Isaac Lab","Asymmetric Actor-Critic","Sim-to-Real Transfer"]},{"id":"mujoco","category":"sim","sec":4,"tier":1,"sources":[{"title":"google-deepmind/mujoco (GitHub)","url":"https://github.com/google-deepmind/mujoco"},{"title":"Google DeepMind: Opening up a physics simulator for robotics (2021-10)","url":"https://deepmind.google/blog/opening-up-a-physics-simulator-for-robotics/"}],"as_of":"2026-09","related_ids":["physics-engine","mujoco-xla","mjcf","mujoco-menagerie","robosuite","deepmind-control-suite"],"name":"MuJoCo (Multi-Joint dynamics with Contact)","alt":"MuJoCo","abbr":"","aliases":["Mujoco"],"one_liner":"An open-source physics engine known for accurate contact simulation, maintained by Google DeepMind.","explanation":"MuJoCo stands for Multi-Joint dynamics with Contact. It was published by Todorov and colleagues at IROS in 2012, and originally required a paid license. DeepMind acquired it in October 2021 and made it free, then open-sourced it under Apache 2.0 starting in 2022; it's now maintained by Google DeepMind. It solves for contact forces using convex optimization, giving a unique solution and cleanly defined inverse dynamics, which suits it well to control and reinforcement learning research. Robot models are described in MJCF, an XML format. Around it has grown an ecosystem including MJX, a JAX version that runs in parallel on GPU/TPU, and the MuJoCo Menagerie model library; robosuite, LIBERO, and the DeepMind Control Suite are all built on top of it.","example":"The LIBERO benchmark is built on robosuite, with the robot arm's and objects' physics computed by MuJoCo underneath.","related":["Physics Engine","MuJoCo XLA","MJCF (MuJoCo XML Format)","MuJoCo Menagerie","robosuite","DeepMind Control Suite"]},{"id":"mujoco-xla","category":"sim","sec":4,"tier":2,"sources":[{"title":"MuJoCo XLA (MJX) documentation","url":"https://mujoco.readthedocs.io/en/stable/mjx.html"},{"title":"mujoco-mjx (PyPI)","url":"https://pypi.org/project/mujoco-mjx/"}],"as_of":"2026-09","related_ids":["mujoco","jax","mujoco-warp","brax","mujoco-playground","differentiable-simulation"],"name":"MuJoCo XLA","alt":"MJX","abbr":"MJX","aliases":["MJX","MuJoCo MJX","mujoco-mjx","MJX-JAX"],"one_liner":"A JAX implementation of MuJoCo for batched, parallel simulation on GPU/TPU that also supports differentiation.","explanation":"MJX (MuJoCo XLA) is a JAX interface to MuJoCo that Google DeepMind released alongside MuJoCo 3.0 in October 2023: MuJoCo's physics are reimplemented in JAX and, once compiled through the XLA compiler, can run on NVIDIA and AMD GPUs, Apple silicon, and Google TPUs. Combined with JAX's vmap, thousands of scenes can be batched together in one computation, which suits large-scale reinforcement learning, though running a single scene alone can be roughly 10 times slower than plain MuJoCo. The pure-JAX version supports automatic differentiation, enabling differentiable simulation. MJX now has two backends: MJX-JAX, and MJX-Warp, which calls into MuJoCo Warp; the latter is faster on NVIDIA GPUs and contact-heavy scenes but cannot be differentiated. Both the Brax training library and MuJoCo Playground are built on top of it.","example":"A typical three-step workflow: mjx.put_model loads a model onto the accelerator, mjx.make_data creates the state, and jax.vmap(mjx.step) advances thousands of robots one step at once, feeding into Brax's PPO implementation to train a walking policy.","related":["MuJoCo (Multi-Joint dynamics with Contact)","JAX","MuJoCo Warp","Brax","MuJoCo Playground","Differentiable Simulation"]},{"id":"brax","category":"sim","sec":4,"tier":3,"sources":[{"title":"Brax - A Differentiable Physics Engine for Large Scale Rigid Body Simulation (arXiv 2106.13281)","url":"https://arxiv.org/abs/2106.13281"},{"title":"google/brax (GitHub)","url":"https://github.com/google/brax"},{"title":"google-deepmind/mujoco_playground (GitHub)","url":"https://github.com/google-deepmind/mujoco_playground"}],"as_of":"2026-09","related_ids":["mujoco-xla","mujoco-playground","jax","differentiable-simulation","gpu-accelerated-parallel-simulation","massively-parallel-reinforcement-learning"],"name":"Brax","alt":"Brax","abbr":"","aliases":[],"one_liner":"A differentiable rigid-body physics engine from Google written in JAX, bundled with a reinforcement-learning training library.","explanation":"Brax was open-sourced by a Google team in 2021 (lead paper author C. Daniel Freeman), written in JAX (a numerical computing library that supports automatic differentiation and compiles to run on GPU/TPU). Its selling point is that both the physics simulation and the learning algorithm compile to run on the same accelerator, eliminating the back-and-forth data transfer between CPU and GPU, and it can train a usable policy on MuJoCo-like Gym tasks in a few minutes; because the simulation is differentiable, it can also compute gradients through the simulation directly to optimize a policy. Brax later offered four physics backends: MJX (MuJoCo's JAX reimplementation), generalized coordinates, position-based dynamics, and a spring model. According to its repository, as of version 0.13.0 only brax/training (its PPO, SAC, and similar training code) is still actively maintained; for physics simulation it now recommends switching to MJX or MuJoCo Warp, and MuJoCo Playground for environments.","example":"MuJoCo Playground's repository notes that to reproduce results exactly as reported in its paper, users can run Brax's training scripts directly, while the environments themselves are implemented with MJX or MuJoCo Warp.","related":["MuJoCo XLA","MuJoCo Playground","JAX","Differentiable Simulation","GPU-Accelerated Parallel Simulation","Massively Parallel Reinforcement Learning"]},{"id":"mujoco-playground","category":"sim","sec":4,"tier":2,"sources":[{"title":"MuJoCo Playground (arXiv 2502.08844)","url":"https://arxiv.org/abs/2502.08844"},{"title":"google-deepmind/mujoco_playground (GitHub)","url":"https://github.com/google-deepmind/mujoco_playground"},{"title":"MuJoCo Playground 项目主页","url":"https://playground.mujoco.org/"}],"as_of":"2026-09","related_ids":["mujoco-xla","mujoco-warp","mujoco","deepmind-control-suite","sim-to-real-transfer","mjlab"],"name":"MuJoCo Playground","alt":"MuJoCo Playground","abbr":"","aliases":["Playground"],"one_liner":"A collection of GPU-accelerated robot learning environments led by DeepMind, built for fast single-GPU training and zero-shot transfer to real robots.","explanation":"MuJoCo Playground is a collection of robot learning environments open-sourced in February 2025 under the Apache 2.0 license, led by Google DeepMind together with UC Berkeley and other teams. Physics runs on MJX, can now also be switched to the MuJoCo Warp backend, and it ships with a batched renderer for training policies that take images as input. It bundles together the classic tasks from the DeepMind Control Suite along with legged-locomotion and manipulation tasks, covering robots such as Unitree's Go1, G1, and H1, the Booster T1, the Berkeley Humanoid, Spot, Franka, ALOHA, and the LEAP dexterous hand. The paper reports training policies in minutes on a single GPU and demonstrates zero-shot transfer to real robots from both state and pixel inputs, cutting out a large amount of the usual repetitive work of building environments and tuning rewards. It installs with pip install playground.","example":"After installing it, a user can load the G1 or Go1 walking environment, train thousands of simulated robots in parallel on one GPU with PPO, and then deploy the trained policy straight to the real robot for testing.","related":["MuJoCo XLA","MuJoCo Warp","MuJoCo (Multi-Joint dynamics with Contact)","DeepMind Control Suite","Sim-to-Real Transfer","mjlab"]},{"id":"mujoco-warp","category":"sim","sec":4,"tier":2,"sources":[{"title":"google-deepmind/mujoco_warp (GitHub)","url":"https://github.com/google-deepmind/mujoco_warp"},{"title":"MuJoCo Warp documentation","url":"https://mujoco.readthedocs.io/en/latest/mjwarp/index.html"},{"title":"mujoco-warp (PyPI)","url":"https://pypi.org/project/mujoco-warp/"}],"as_of":"2026-09","related_ids":["mujoco","mujoco-xla","newton-physics-engine","nvidia-warp","mjlab","gpu-accelerated-parallel-simulation"],"name":"MuJoCo Warp","alt":"MuJoCo Warp","abbr":"MJWarp","aliases":["MJWarp","mujoco-warp","MJX-Warp"],"one_liner":"A GPU rewrite of MuJoCo built by DeepMind and NVIDIA on the Warp framework, aimed at massive batched parallelism.","explanation":"MuJoCo Warp is a GPU version of MuJoCo developed jointly by Google DeepMind and NVIDIA, reimplemented using NVIDIA Warp (a framework for writing GPU kernels in Python), and it also serves as the core solver of the Newton physics engine. It is built for throughput, advancing hundreds or thousands of simulated “worlds” at once; its developers say it scales better than MJX on scenes with many geometries, high degrees of freedom, and complex contact, and that it integrates more easily with PyTorch. The trade-off is that its per-step latency can be slower than plain MuJoCo, so real-time control still uses the original; it also doesn't support automatic differentiation, so differentiable simulation still requires MJX's JAX implementation. It additionally supports batched ray-traced rendering. Since January 2026 it has shipped on PyPI as mujoco-warp, with version numbers kept in sync with MuJoCo, and it can be called from mjlab, MJX, MuJoCo Playground, and Newton.","example":"A typical workflow uses mjw.put_model to load a model onto the GPU, mjw.make_data(mjm, nworld=100) to create 100 worlds, and a single mjw.step call to advance all 100 simulations one step at once, with a CUDA Graph capturing the loop for further speedup.","related":["MuJoCo (Multi-Joint dynamics with Contact)","MuJoCo XLA","Newton Physics Engine","NVIDIA Warp","mjlab","GPU-Accelerated Parallel Simulation"]},{"id":"physx","category":"sim","sec":4,"tier":2,"sources":[{"title":"NVIDIA-Omniverse/PhysX (GitHub)","url":"https://github.com/NVIDIA-Omniverse/PhysX"},{"title":"NVIDIA PhysX SDK","url":"https://developer.nvidia.com/physx-sdk"},{"title":"PhysX (Wikipedia)","url":"https://en.wikipedia.org/wiki/PhysX"}],"as_of":"2026-09","related_ids":["nvidia-isaac-sim","nvidia-isaac-lab","isaac-gym","nvidia-omniverse","articulated-body-simulation","newton-physics-engine"],"name":"PhysX","alt":"PhysX","abbr":"","aliases":["NVIDIA PhysX","PhysX 5","PhysX SDK","NovodeX"],"one_liner":"NVIDIA's open-source real-time physics engine, and the physics core behind Isaac Sim and Isaac Lab.","explanation":"PhysX is NVIDIA's real-time physics engine SDK. It traces back to NovodeX, founded by ETH researchers and commercialized in 2002; Ageia acquired and renamed it in 2004; and after NVIDIA acquired Ageia in 2008, it was rebuilt to use CUDA GPU acceleration, becoming dominant physics middleware in the game industry for years afterward. It was open-sourced under the BSD-3 license in December 2018, and the open-source version was upgraded to PhysX 5 in November 2022. PhysX 5 supports GPU rigid bodies, reduced-coordinate articulations (describing a robot by its root pose plus joint angles, so joints never “fall apart” — well suited to arms and legged robots), finite-element soft bodies, position-based-dynamics fluids and cloth, and signed-distance-field collision. It is the physics core behind Omniverse, Isaac Sim, Isaac Lab, and the earlier Isaac Gym; tuning solver iteration counts, friction, and similar parameters on those platforms is really tuning PhysX itself.","example":"When training a robot arm to perform peg insertion in Isaac Lab, joint torques, contact forces, and friction are all computed by PhysX on the GPU; when interpenetration shows up, a common fix is to increase PhysX's solver position-iteration count.","related":["NVIDIA Isaac Sim","NVIDIA Isaac Lab","Isaac Gym","NVIDIA Omniverse","Articulated-Body Simulation (Articulation)","Newton Physics Engine"]},{"id":"isaac-gym","category":"sim","sec":4,"tier":2,"sources":[{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (arXiv 2108.10470)","url":"https://arxiv.org/abs/2108.10470"},{"title":"NVIDIA Isaac Gym","url":"https://developer.nvidia.com/isaac-gym"},{"title":"isaac-sim/IsaacGymEnvs (GitHub)","url":"https://github.com/isaac-sim/IsaacGymEnvs"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","physx","legged-gym","massively-parallel-reinforcement-learning","nvidia-isaac-sim","orbit-isaacgymenvs-omniisaacgymenvs"],"name":"Isaac Gym","alt":"Isaac Gym","abbr":"","aliases":["NVIDIA Isaac Gym","IsaacGym","IsaacGymEnvs","Isaac Gym Preview 4"],"one_liner":"NVIDIA's GPU reinforcement-learning simulator launched in 2021; now discontinued.","explanation":"Isaac Gym is a robot reinforcement-learning simulation platform NVIDIA released in 2021 (its technical report appeared in the NeurIPS 2021 Datasets and Benchmarks track), with physics computed by PhysX on the GPU. Its key feature is a tensor interface: physics results are handed to the neural network directly as PyTorch tensors without passing through the CPU, letting a single GPU run thousands of environments at once; NVIDIA claimed speedups of two to three orders of magnitude over a traditional “CPU simulation plus GPU training” pipeline. A wave of legged-locomotion and dexterous-hand work, including legged_gym, was built on it, and its companion repository IsaacGymEnvs shipped example environments such as Ant and ShadowHand. It was only ever released as a Preview build, with Preview 4 the final version; it is now discontinued, and NVIDIA recommends migrating to Isaac Lab, though many open-source projects still depend on it.","example":"Unitree's unitree_rl_gym and ETH's legged_gym still require downloading and installing Isaac Gym Preview 3 or 4 before they can train walking policies for quadruped or humanoid robots.","related":["NVIDIA Isaac Lab","PhysX","legged_gym","Massively Parallel Reinforcement Learning","NVIDIA Isaac Sim","Orbit / IsaacGymEnvs / OmniIsaacGymEnvs (predecessors of Isaac Lab)"]},{"id":"nvidia-isaac-sim","category":"sim","sec":4,"tier":1,"sources":[{"title":"NVIDIA Isaac Sim","url":"https://developer.nvidia.com/isaac/sim"},{"title":"isaac-sim/IsaacSim releases (GitHub)","url":"https://github.com/isaac-sim/IsaacSim/releases"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","nvidia-omniverse","physx","universal-scene-description","synthetic-data","digital-twin"],"name":"NVIDIA Isaac Sim","alt":"Isaac Sim","abbr":"","aliases":[],"one_liner":"NVIDIA's Omniverse-based, open-source robot simulator built for high physical and visual fidelity.","explanation":"Isaac Sim is NVIDIA's reference robot simulation application, built on the Omniverse platform: scenes are described in OpenUSD, physics runs on PhysX, and rendering runs on RTX. It can import robots and scenes from CAD, URDF, or MJCF, simulate sensors such as cameras and lidar, batch-generate synthetic data by randomizing lighting, color, and position, and supports software-in-the-loop/hardware-in-the-loop testing along with a ROS 2 bridge. It's been open source on GitHub under Apache 2.0 since 2025. Its role is to be the high-fidelity simulator, while large-scale policy training is usually handed off to Isaac Lab, which is built on top of it. As of September 2026, the latest stable release is 6.1, with 7.0 available as an alpha preview.","example":"Build a warehouse scene in Isaac Sim, randomize the crate textures and lighting, and batch-render labeled images to train a detection model.","related":["NVIDIA Isaac Lab","NVIDIA Omniverse","PhysX","Universal Scene Description (OpenUSD)","Synthetic Data","Digital Twin"]},{"id":"nvidia-isaac-lab","category":"sim","sec":4,"tier":1,"sources":[{"title":"Isaac Lab Documentation","url":"https://isaac-sim.github.io/IsaacLab/main/index.html"},{"title":"Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning (arXiv 2511.04831)","url":"https://arxiv.org/abs/2511.04831"},{"title":"isaac-sim/IsaacLab releases (GitHub)","url":"https://github.com/isaac-sim/IsaacLab/releases"},{"title":"NVIDIA Developer: Isaac Gym - Preview Release (now deprecated, no longer supported)","url":"https://developer.nvidia.com/isaac-gym"}],"as_of":"2026-09","related_ids":["nvidia-isaac-sim","isaac-gym","orbit-isaacgymenvs-omniisaacgymenvs","gpu-accelerated-parallel-simulation","newton-physics-engine","rl-based-locomotion-control"],"name":"NVIDIA Isaac Lab","alt":"Isaac Lab","abbr":"","aliases":["Orbit","IsaacLab"],"one_liner":"NVIDIA's open-source, GPU-parallel robot learning framework for training policies at scale in simulation.","explanation":"Isaac Lab is an open-source robot learning framework (BSD-3 licensed) led by NVIDIA, evolved from the Orbit framework published in 2023, and NVIDIA's officially recommended replacement for Isaac Gym, which is no longer supported. It's built on top of Isaac Sim, using PhysX for physics and RTX for rendering, and can run large numbers of parallel environments simultaneously on the GPU; it comes with actuator models, a variety of sensors, domain randomization, and data-collection tools, and supports reinforcement learning, imitation learning, and motion planning. Many humanoid and quadruped reinforcement-learning locomotion policies are trained here. As of September 2026, version 3.0 is in early access and has moved to a multi-physics-backend architecture that can plug in the Newton physics engine, with some workflows no longer requiring Isaac Sim to be installed at all.","example":"To train a walking policy for the Unitree G1, open a few thousand parallel environments in Isaac Lab with domain randomization applied, then deploy the trained policy to the real robot.","related":["NVIDIA Isaac Sim","Isaac Gym","Orbit / IsaacGymEnvs / OmniIsaacGymEnvs (predecessors of Isaac Lab)","GPU-Accelerated Parallel Simulation","Newton Physics Engine","RL-based Locomotion Control"]},{"id":"orbit-isaacgymenvs-omniisaacgymenvs","category":"sim","sec":4,"tier":3,"sources":[{"title":"Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments (RA-L 2023)","url":"https://arxiv.org/abs/2301.04195"},{"title":"GitHub: isaac-sim/IsaacGymEnvs (archived)","url":"https://github.com/isaac-sim/IsaacGymEnvs"},{"title":"Isaac Lab Docs: Migrating from Orbit","url":"https://isaac-sim.github.io/IsaacLab/main/source/migration/migrating_from_orbit.html"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","isaac-gym","nvidia-isaac-sim","legged-gym","gpu-accelerated-parallel-simulation","massively-parallel-reinforcement-learning"],"name":"Orbit / IsaacGymEnvs / OmniIsaacGymEnvs (predecessors of Isaac Lab)","alt":"Orbit / IsaacGymEnvs（Isaac Lab 前身）","abbr":"","aliases":["Isaac Orbit","OmniIsaacGymEnvs","OIGE"],"one_liner":"Three now-discontinued NVIDIA robot-learning frameworks that were later unified into Isaac Lab.","explanation":"These are the three robot reinforcement-learning frameworks that existed in NVIDIA's ecosystem before Isaac Lab. IsaacGymEnvs (from 2021) was a library of example environments for the Isaac Gym preview release, including tasks like Ant, ShadowHand, and ANYmal, and it demonstrated training with thousands of parallel environments on a single GPU. OmniIsaacGymEnvs ported that style of task onto Isaac Sim; version 4.0.0 was its last release, and NVIDIA's documentation says it has been folded into Isaac Lab. Orbit was proposed by Mayank Mittal and colleagues (RA-L 2023), built on Isaac Sim, and supported 16 robot platforms with more than 20 benchmark tasks — it became the codebase Isaac Lab is built on. Today the first two repositories are archived, and Isaac Lab's documentation provides migration guides from all three; anyone reading older paper code will still run into them regularly.","example":"Old code that imports from omni.isaac.orbit needs to be changed to import from isaaclab when migrating to Isaac Lab — that's the first step in the official Orbit migration guide.","related":["NVIDIA Isaac Lab","Isaac Gym","NVIDIA Isaac Sim","legged_gym","GPU-Accelerated Parallel Simulation","Massively Parallel Reinforcement Learning"]},{"id":"leisaac","category":"sim","sec":4,"tier":3,"sources":[{"title":"LightwheelAI/leisaac (GitHub)","url":"https://github.com/LightwheelAI/leisaac"},{"title":"Hugging Face LeRobot Docs: LeIsaac × LeRobot EnvHub","url":"https://huggingface.co/docs/lerobot/envhub_leisaac"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","lerobot","lerobot-envhub","so-100-so-101-arm","lightwheel","nvidia-isaac-gr00t-n1"],"name":"LeIsaac","alt":"LeIsaac","abbr":"","aliases":["leisaac"],"one_liner":"Lightwheel's open-source Isaac Lab teleoperation data-collection framework, plugging the SO-101 arm into the LeRobot pipeline.","explanation":"LeIsaac is open-sourced by Lightwheel (光轮智能) under the Apache 2.0 license. It places a simulated SO-101 robot arm (single- or dual-arm) inside NVIDIA Isaac Lab, letting a user teleoperate the simulated follower arm with a real SO-101 leader arm, with keyboard and gamepad also supported; besides manual teleoperation, it can also generate demonstration trajectories automatically with a state-machine script. The collected HDF5 data can be converted to LeRobot's dataset format, used to fine-tune policies such as GR00T N1.5/N1.6 and deploy them to a real robot. It's also the official simulated imitation-learning environment for LeRobot's EnvHub, loadable in one line of code for tasks such as picking oranges, lifting blocks, tidying a toy table, and bimanual cloth folding, and it can also run in the cloud when no local GPU is available.","example":"Using LeRobot's make_env to load the so101_pick_orange task from LightwheelAI/leisaac_env: three oranges are placed onto a plate, and the arm then returns to its initial pose.","related":["NVIDIA Isaac Lab","LeRobot","LeRobot EnvHub","SO-100 / SO-101 Arm","Lightwheel","NVIDIA Isaac GR00T N1"]},{"id":"newton-physics-engine","category":"sim","sec":4,"tier":2,"sources":[{"title":"newton-physics/newton (GitHub)","url":"https://github.com/newton-physics/newton"},{"title":"Linux Foundation: Contribution of Newton by Disney Research, Google DeepMind, and NVIDIA","url":"https://www.linuxfoundation.org/press/linux-foundation-announces-contribution-of-newton-by-disney-research-google-deepmind-and-nvidia-to-accelerate-open-robot-learning"},{"title":"Newton solvers API","url":"https://newton-physics.github.io/newton/latest/api/newton_solvers.html"}],"as_of":"2026-09","related_ids":["mujoco-warp","nvidia-warp","nvidia-isaac-lab","physx","universal-scene-description","differentiable-simulation"],"name":"Newton Physics Engine","alt":"Newton 物理引擎","abbr":"","aliases":["Newton","newton-physics"],"one_liner":"An open-source GPU physics engine for robotics jointly launched by NVIDIA, DeepMind, and Disney.","explanation":"Newton is an open-source physics engine jointly launched by NVIDIA, Google DeepMind, and Disney Research, built on NVIDIA Warp and OpenUSD (Universal Scene Description) and released under the Apache 2.0 license; the Linux Foundation announced it would take over stewardship on September 29, 2025, to keep it vendor-neutral. Its design is one framework with multiple solvers: MuJoCo Warp is the primary rigid-body solver, alongside Featherstone, XPBD, VBD (for cloth and particles), implicit MPM (for granular materials like sand), Style3D cloth, and Kamino for handling closed-loop mechanisms. It supports massive GPU parallelism and differentiability, is compatible with Isaac Sim and Isaac Lab, and NVIDIA positions it as the next-generation robot simulation engine succeeding PhysX. Version 1.0 shipped in 2026, and it had reached 1.6 by September.","example":"Within a single Newton setup, a humanoid robot's walking might use SolverMuJoCo, a towel's cloth might use SolverVBD or SolverStyle3D, and sand particles might use SolverImplicitMPM.","related":["MuJoCo Warp","NVIDIA Warp","NVIDIA Isaac Lab","PhysX","Universal Scene Description (OpenUSD)","Differentiable Simulation"]},{"id":"mjlab","category":"sim","sec":4,"tier":3,"sources":[{"title":"mjlab: A Lightweight Framework for GPU-Accelerated Robot Learning (arXiv 2601.22074)","url":"https://arxiv.org/abs/2601.22074"},{"title":"mjlab GitHub 仓库","url":"https://github.com/mujocolab/mjlab"}],"as_of":"2026-09","related_ids":["mujoco-warp","nvidia-isaac-lab","mujoco-playground","gpu-accelerated-parallel-simulation","rl-based-locomotion-control","unitree-g1"],"name":"mjlab","alt":"mjlab","abbr":"","aliases":[],"one_liner":"A lightweight GPU robot-learning framework combining an Isaac Lab-style interface with MuJoCo Warp physics.","explanation":"mjlab is an open-source framework developed by Kevin Zakka, Koushil Sreenath, Pieter Abbeel, and colleagues, with its paper released in January 2026. It reuses Isaac Lab's “manager-based” interface, where observations, rewards, randomization events, and so on are written as composable modules to assemble an environment; its physics backend is instead MuJoCo Warp, the GPU version of MuJoCo, which can simulate thousands of environments in parallel. Its advantages are that it installs with a single command, has few dependencies, and gives direct access to MuJoCo's native data structures. It ships with three reference task categories — velocity tracking, motion imitation, and manipulation — commonly used for humanoid and quadruped reinforcement-learning locomotion control. Training requires an NVIDIA GPU; macOS can only be used for evaluation.","example":"Running uv run train Mjlab-Velocity-Flat-Unitree-G1 --env.scene.num-envs 4096 trains a Unitree G1 to walk according to velocity commands across 4,096 parallel environments.","related":["MuJoCo Warp","NVIDIA Isaac Lab","MuJoCo Playground","GPU-Accelerated Parallel Simulation","RL-based Locomotion Control","Unitree G1"]},{"id":"genesis","category":"sim","sec":4,"tier":2,"sources":[{"title":"Genesis-Embodied-AI/Genesis (GitHub)","url":"https://github.com/Genesis-Embodied-AI/Genesis"},{"title":"Genesis World Docs: What is Genesis","url":"https://genesis-world.readthedocs.io/en/latest/user_guide/overview/what_is_genesis.html"}],"as_of":"2026-09","related_ids":["gpu-accelerated-parallel-simulation","physics-engine","material-point-method","nvidia-isaac-lab","mujoco","generative-simulation"],"name":"Genesis","alt":"Genesis","abbr":"","aliases":["Genesis World","Genesis Simulator","Genesis-Embodied-AI"],"one_liner":"A multi-physics GPU simulation platform open-sourced in late 2024, built for speed and unified rigid-body, soft-body, and fluid simulation.","explanation":"Genesis (now branded Genesis World) is a physics simulation platform for robotics and embodied AI, open-sourced as an academic project in December 2024 under the Apache 2.0 license and now developed with backing from Genesis AI; it has over 30,000 GitHub stars. It unifies several solvers — rigid body, finite element method (FEM, for soft-body deformation), material point method (MPM, for granular materials like sand), and PBD/SPH particle methods (for cloth and fluids) — inside a single scene, and ships with a photorealistic renderer called Nyx. It is written in Python and compiled to CUDA, Metal, and other backends by Quadrants, a compiler forked from Taichi. The developers claim throughput 10 to 80 times that of GPU simulators such as Isaac Gym/Sim/Lab and MuJoCo MJX. It is commonly used for massively parallel reinforcement learning, such as training locomotion controllers for quadrupeds and humanoids.","example":"The official examples include a script that runs thousands of Unitree Go2 environments in parallel with Genesis and trains a walking policy with PPO, completing training on a single GPU.","related":["GPU-Accelerated Parallel Simulation","Physics Engine","Material Point Method","NVIDIA Isaac Lab","MuJoCo (Multi-Joint dynamics with Contact)","Generative Simulation"]},{"id":"pybullet","category":"sim","sec":4,"tier":2,"sources":[{"title":"bulletphysics/bullet3 (GitHub)","url":"https://github.com/bulletphysics/bullet3"},{"title":"pybullet on PyPI","url":"https://pypi.org/project/pybullet/"},{"title":"Sim-to-Real: Learning Agile Locomotion For Quadruped Robots (arXiv 1804.10332)","url":"https://arxiv.org/abs/1804.10332"}],"as_of":"2025-01","related_ids":["physics-engine","mujoco","isaac-gym","sim-to-real-transfer","unified-robot-description-format","rigid-body-simulation"],"name":"PyBullet","alt":"PyBullet","abbr":"","aliases":["Bullet","Bullet Physics","Bullet Physics SDK"],"one_liner":"The Python interface to the Bullet physics engine, and one of the most widely used simulators in early robot reinforcement learning.","explanation":"Bullet is an open-source physics engine (written in C++, zlib license) led by Erwin Coumans, supporting collision detection along with rigid-body and soft-body (cloth, rope) dynamics; it has been widely used in games and film visual effects, and Coumans received a Sci-Tech Academy Award for it. PyBullet is its Python wrapper, installable with a single pip command, and it can load URDF, SDF, and MJCF robot models directly, with built-in forward/inverse kinematics, inverse dynamics, and basic rendering. Before GPU-parallel simulation became common, a large amount of robot reinforcement-learning and sim-to-real work was done in PyBullet. It runs mainly on the CPU, so its parallel scale falls well short of Isaac Lab, MJX, and similar platforms, and new large-scale training has largely moved to those instead — though PyBullet remains common for teaching and lightweight experiments. The latest version on PyPI, 3.2.7, was released in January 2025.","example":"Google's 2018 paper “Sim-to-Real: Learning Agile Locomotion for Quadruped Robots” trained trotting and galloping gaits for the Minitaur quadruped in PyBullet, then deployed the policy directly to the real robot after system identification, actuator modeling, and randomization.","related":["Physics Engine","MuJoCo (Multi-Joint dynamics with Contact)","Isaac Gym","Sim-to-Real Transfer","Unified Robot Description Format","Rigid-Body Simulation"]},{"id":"raisim","category":"sim","sec":4,"tier":3,"sources":[{"title":"RaiSim 文档：License（v2.7.0）","url":"https://raisim.com/sections/License.html"},{"title":"Per-Contact Iteration Method for Solving Contact Dynamics (IEEE RA-L 2018)","url":"https://ieeexplore.ieee.org/document/8255551"},{"title":"raisimTech/raisim2Lib GitHub","url":"https://github.com/raisimTech/raisim2Lib"}],"as_of":"2026-09","related_ids":["physics-engine","mujoco","isaac-gym","rl-based-locomotion-control","anybotics-anymal","eth-zurich-robotic-systems-lab"],"name":"RaiSim","alt":"RaiSim","abbr":"","aliases":["RaiSim2","raisimLib"],"one_liner":"A multi-body physics engine for robotics and reinforcement learning, known for fast and accurate contact solving.","explanation":"RaiSim was developed by Jemin Hwangbo and colleagues, and is now maintained by RaiSim Tech. Its core comes from a 2018 IEEE RA-L paper by Hwangbo, Lee, and Hutter describing a “per-contact iteration” method that iterates over contact points one at a time and uses bisection to solve for contact force, which the paper reports as roughly twice as fast as two existing methods. ETH Zurich's 2019 paper on training the ANYmal quadruped with reinforcement learning used this same contact solver in its custom simulator. RaiSim focuses on simulating rigid and articulated bodies fast and accurately on the CPU, offering a C++ interface, Python bindings called raisimPy, and reinforcement-learning examples in raisimGymTorch; later versions added soft bodies, granular media, tendons, and camera/depth sensor rendering. The older raisimLib repository was archived in April 2026 and replaced by the binary-distributed RaiSim2, which reached v2.7.0 in September 2026. It is not open source: an academic license is free, while a commercial license costs $1,500 per person per year.","example":"Using raisimGymTorch, run many quadruped-robot environments in parallel on one computer to train a reinforcement-learning policy for ANYmal to walk over height-map terrain.","related":["Physics Engine","MuJoCo (Multi-Joint dynamics with Contact)","Isaac Gym","RL-based Locomotion Control","ANYbotics ANYmal","ETH Zurich Robotic Systems Lab"]},{"id":"sapien","category":"sim","sec":4,"tier":3,"sources":[{"title":"arXiv 2003.08515 - SAPIEN: A SimulAted Part-based Interactive ENvironment","url":"https://arxiv.org/abs/2003.08515"},{"title":"GitHub - haosulab/SAPIEN","url":"https://github.com/haosulab/SAPIEN"}],"as_of":"2026-09","related_ids":["maniskill","partnet-mobility","articulated-object","articulated-object-manipulation","physx","simulator"],"name":"SAPIEN (SimulAted Part-based Interactive ENvironment)","alt":"SAPIEN","abbr":"","aliases":[],"one_liner":"A UC San Diego robot simulation platform specialized for interacting with articulated objects like drawers and cabinet doors.","explanation":"SAPIEN, short for SimulAted Part-based Interactive ENvironment, is a robot simulation environment proposed in 2020 by Hao Su's group at UC San Diego together with collaborators including Stanford, with the paper appearing at CVPR 2020. Its physics runs on NVIDIA PhysX, its rendering is based on Vulkan, and later versions added ray-traced rendering; starting with version 3.0 it supports PhysX 5's GPU-parallel simulation. Its distinguishing feature is the PartNet-Mobility dataset released alongside the paper: a large collection of articulated objects with joint annotations — drawers that pull open, cabinet doors that swing open, faucets, and the like — well suited to research on part-level perception and interaction, such as a robot opening a door or pulling out a drawer. The widely used ManiSkill family of manipulation benchmarks is built on top of SAPIEN.","example":"Load a cabinet model from PartNet-Mobility into SAPIEN and have a robot arm learn to locate the handle, grasp it, and pull the cabinet door open.","related":["ManiSkill","PartNet-Mobility","Articulated Object","Articulated Object Manipulation","PhysX","Simulator"]},{"id":"maniskill","category":"sim","sec":4,"tier":2,"sources":[{"title":"haosulab/ManiSkill (GitHub)","url":"https://github.com/haosulab/ManiSkill"},{"title":"ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI (arXiv 2410.00425)","url":"https://arxiv.org/abs/2410.00425"},{"title":"mani-skill (PyPI)","url":"https://pypi.org/project/mani-skill/"}],"as_of":"2026-04","related_ids":["sapien","gpu-accelerated-parallel-simulation","simulation-based-evaluation","simplerenv","digital-twin","benchmark"],"name":"ManiSkill","alt":"ManiSkill","abbr":"","aliases":["ManiSkill3","ManiSkill2","ManiSkill 1"],"one_liner":"An open-source robot manipulation simulation framework and benchmark from a UC San Diego team, built for GPU parallelism.","explanation":"ManiSkill is a robot manipulation simulation framework and benchmark developed by Hao Su's group at UC San Diego together with Hillbot and others, built on the same team's SAPIEN engine and now in its third generation. The latest version, ManiSkill3 (RSS 2025), runs both physics simulation and camera rendering in parallel on the GPU; its paper claims 10 to 1000 times the speed of other platforms and 2 to 3 times less GPU memory use, generating over 30,000 segmented RGB-D frames per second on a single RTX 4090. It covers 12 task categories, including tabletop manipulation, mobile manipulation, humanoids, dexterous hands, and soft bodies (soft-body tasks don't support batched parallelism), plus digital-twin environments built to match real-world setups for quickly evaluating real-robot policies in simulation. Its interface follows Gymnasium, and it exited beta with version 3.0.1 in April 2026.","example":"ManiSkill3's built-in PickCube-v1 task, where an arm has to pick up a cube, can be run as a thousand parallel environments at once to train a camera-conditioned grasping policy with PPO.","related":["SAPIEN (SimulAted Part-based Interactive ENvironment)","GPU-Accelerated Parallel Simulation","Simulation-Based Evaluation","SimplerEnv","Digital Twin","Benchmark"]},{"id":"robosuite","category":"sim","sec":4,"tier":2,"sources":[{"title":"robosuite 官网","url":"https://robosuite.ai/"},{"title":"robosuite 文档：Environments","url":"https://robosuite.ai/docs/modules/environments.html"},{"title":"robosuite: A Modular Simulation Framework and Benchmark for Robot Learning (arXiv 2009.12293)","url":"https://arxiv.org/abs/2009.12293"}],"as_of":"2024-10","related_ids":["mujoco","robomimic","robocasa","mimicgen","simulator","spacemouse-teleoperation"],"name":"robosuite","alt":"robosuite","abbr":"","aliases":["robosuite v1.5"],"one_liner":"A modular, MuJoCo-based robot-learning simulation framework and benchmark, maintained under the ARISE initiative.","explanation":"robosuite is an open-source simulation framework developed by Stanford and UT Austin researchers under the ARISE initiative, built on MuJoCo, with its paper published in 2020. It breaks robots, grippers, controllers, scenes, and tasks into composable modules, offering both joint-space and Cartesian-space (directly controlling end-effector pose) controllers. It ships with built-in single-arm tasks such as lifting a cube, stacking blocks, opening a door, and wiping a table, plus three bimanual tasks, and it supports teleoperation for collecting demonstrations. Version 1.5, released in October 2024, added humanoid and other embodiments, whole-body controllers, photorealistic rendering, and sensor models. Both robomimic and RoboCasa are built on top of it, making it a common foundation for imitation-learning and reinforcement-learning research.","example":"robomimic's standard datasets, such as Lift, Can, and Square, were all collected inside robosuite environments using a Franka Panda arm.","related":["MuJoCo (Multi-Joint dynamics with Contact)","robomimic","RoboCasa","MimicGen","Simulator","SpaceMouse Teleoperation"]},{"id":"robomimic","category":"sim","sec":4,"tier":2,"sources":[{"title":"robomimic 官网","url":"https://robomimic.github.io/"},{"title":"robomimic v0.1 数据集说明","url":"https://robomimic.github.io/docs/datasets/robomimic_v0.1.html"},{"title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (arXiv 2108.03298)","url":"https://arxiv.org/abs/2108.03298"}],"as_of":"","related_ids":["robosuite","behavior-cloning","offline-reinforcement-learning","demonstration-data","diffusion-policy","roboturk"],"name":"robomimic","alt":"RoboMimic","abbr":"","aliases":[],"one_liner":"A learning-from-demonstration framework from Stanford and UT Austin researchers, with standard datasets and offline learning algorithms.","explanation":"robomimic is a learning-from-demonstration framework developed under the ARISE initiative by researchers at Stanford and UT Austin, with its paper published at CoRL 2021. It provides data for five manipulation tasks in the robosuite simulator — Lift, Can, Square, Transport, and Tool Hang — in three variants: PH is 200 demonstrations from one skilled operator, MH is 300 demonstrations from 6 operators of varying skill, and MG is data generated by a reinforcement-learning agent. The framework includes built-in behavioral cloning, BC-RNN (behavioral cloning that can look at history), and several offline reinforcement-learning algorithms, with diffusion policies added in later versions. The paper found that models which look at history perform better, and that data quality matters a great deal; this dataset later became a standard simulation benchmark for imitation-learning papers.","example":"The Diffusion Policy paper compared success rates against BC-RNN and other methods on robomimic's Lift, Can, Square, Transport, and Tool Hang tasks.","related":["robosuite","Behavior Cloning","Offline Reinforcement Learning","Demonstration Data","Diffusion Policy","RoboTurk"]},{"id":"gazebo","category":"sim","sec":4,"tier":2,"sources":[{"title":"Gazebo (simulator) - Wikipedia","url":"https://en.wikipedia.org/wiki/Gazebo_(simulator)"},{"title":"Gazebo Docs: Getting Started","url":"https://gazebosim.org/docs/latest/getstarted/"}],"as_of":"2026-09","related_ids":["robot-operating-system","robot-operating-system-2","simulator","unified-robot-description-format","simulation-description-format","ros-2-navigation-stack"],"name":"Gazebo","alt":"Gazebo","abbr":"","aliases":["Gazebo Classic","Gazebo Sim","Ignition Gazebo"],"one_liner":"The most widely used open-source robot simulator in the ROS ecosystem, maintained by Open Robotics.","explanation":"Gazebo is an open-source 3D robot simulator that began in 2002 as part of the Player project, became independent in 2011 with support from Willow Garage, and has been maintained since 2012 by the Open Source Robotics Foundation (OSRF, renamed Open Robotics in 2018) — the same organization behind ROS. It can simulate robots, cameras, lidar, IMUs, and other sensors and environments, and is commonly used to test navigation, mapping, and control code before deploying it on a real robot, as well as in competitions such as the DARPA Robotics Challenge. A rewrite that began in 2017 was branded Ignition for a while, then renamed back to Gazebo in 2022 over trademark issues, with the older codebase relabeled Gazebo Classic; Gazebo Classic stopped being maintained in January 2025. As of 2026 the current long-term-support release is Gazebo Jetty. Gazebo is oriented toward software integration testing rather than massively parallel GPU training, so reinforcement-learning work tends to use Isaac Lab, MuJoCo, or similar tools instead.","example":"When building a mobile robot with ROS 2, developers commonly load the robot's URDF model and a warehouse scene into Gazebo first, get Nav2 navigation and SLAM mapping working there, and then deploy the same set of ROS nodes to the real robot.","related":["Robot Operating System","Robot Operating System 2","Simulator","Unified Robot Description Format","Simulation Description Format (SDFormat)","ROS 2 Navigation Stack (Nav2)"]},{"id":"webots","category":"sim","sec":4,"tier":3,"sources":[{"title":"Cyberbotics 官网：Webots","url":"https://cyberbotics.com/"},{"title":"Webots GitHub 仓库","url":"https://github.com/cyberbotics/webots"},{"title":"Webots Releases","url":"https://github.com/cyberbotics/webots/releases"}],"as_of":"2026-09","related_ids":["simulator","physics-engine","gazebo","coppeliasim","robot-operating-system","robot-operating-system-2"],"name":"Webots","alt":"Webots","abbr":"","aliases":[],"one_liner":"An open-source desktop robot simulator maintained by Switzerland's Cyberbotics, beginner-friendly and popular for teaching.","explanation":"Webots is a cross-platform, open-source desktop robot simulation program that originated in 1996 at EPFL (the Swiss Federal Institute of Technology in Lausanne), has been maintained since 1998 by its spin-off company Cyberbotics, and became open source (Apache 2.0 license) in December 2018. It ships with a graphical interface, a physics engine adapted from ODE, and a large library of ready-made robot and sensor models; controllers can be written in C/C++, Java, Python, or MATLAB, and it can also connect to ROS / ROS 2. Compared with simulators like Isaac Lab and MuJoCo, which are built for massively parallel reinforcement learning, Webots leans toward single-machine, interactive simulation, and is more commonly seen in teaching, competitions, and validating mobile-robot algorithms. As of September 2026, the latest stable release is R2025a, from February 2025.","example":"The official first tutorial: place a two-wheeled e-puck robot in a walled arena and write a short controller program to make it move — it can typically be finished in about 30 minutes.","related":["Simulator","Physics Engine","Gazebo","CoppeliaSim","Robot Operating System","Robot Operating System 2"]},{"id":"coppeliasim","category":"sim","sec":4,"tier":3,"sources":[{"title":"Coppelia Robotics 官网","url":"https://www.coppeliarobotics.com/"},{"title":"CoppeliaSim - Wikipedia","url":"https://en.wikipedia.org/wiki/CoppeliaSim"},{"title":"RLBench GitHub（基于 CoppeliaSim 与 PyRep）","url":"https://github.com/stepjam/RLBench"}],"as_of":"2026-09","related_ids":["simulator","physics-engine","rlbench","mujoco","gazebo","webots"],"name":"CoppeliaSim","alt":"CoppeliaSim","abbr":"","aliases":["V-REP","Virtual Robot Experimentation Platform"],"one_liner":"A general-purpose robot simulator from Coppelia Robotics in Switzerland, the successor to V-REP.","explanation":"CoppeliaSim is robot simulation software maintained by Coppelia Robotics, based in Zurich, Switzerland; it was formerly called V-REP, originally developed within Toshiba's R&D division, and remains fully compatible with V-REP after the rename. Its defining feature is breadth: it can switch between five physics engines — MuJoCo, Bullet, ODE, Newton, and Vortex — and has built-in forward/inverse kinematics, collision and distance computation, and OMPL-based motion planning, controllable from Python, Lua, C/C++, and other languages through embedded scripts, plugins, or a remote API. It comes in an educational edition (restricted to students and teachers, non-commercial) and a commercial edition. In embodied learning, it is most commonly used indirectly through the PyRep interface, and the classic manipulation benchmark RLBench is built on top of it. It is primarily single-machine CPU simulation, so reinforcement-learning training that needs thousands of parallel environments generally switches to GPU simulators like Isaac Lab instead.","example":"RLBench's tabletop manipulation tasks, such as opening a drawer or stacking blocks, all run inside CoppeliaSim 4.1, with researchers using PyRep's Python interface to control a simulated Franka arm for data collection and policy testing.","related":["Simulator","Physics Engine","RLBench","MuJoCo (Multi-Joint dynamics with Contact)","Gazebo","Webots"]},{"id":"game-engine","category":"sim","sec":4,"tier":3,"sources":[{"title":"Game engine - Wikipedia","url":"https://en.wikipedia.org/wiki/Game_engine"},{"title":"CARLA Simulator","url":"https://carla.org/"},{"title":"AI2-THOR","url":"https://ai2thor.allenai.org/"}],"as_of":"2026-09","related_ids":["simulator","physics-engine","rendering","photorealistic-rendering","carla","ai2-thor"],"name":"Game Engine","alt":"游戏引擎","abbr":"","aliases":["Unity","Unreal Engine"],"one_liner":"Software framework for building video games, combining rendering, physics, and animation, and often repurposed for simulation.","explanation":"A game engine is a software framework built for developing video games, typically bundling modules for 3D rendering, a physics engine, animation, audio, scripting, and scene management; leading commercial products are Unity and Unreal Engine, with Godot as an open-source option. Because they can render realistic images in real time on consumer graphics cards and make scene editing easy, they're also used for architectural visualization, training simulations, and research. Many embodied-AI simulation platforms are built directly on game engines, gaining realistic visuals and a rich ecosystem of assets and tools; the downside is that game physics is tuned to look plausible rather than to be precise, commonly connecting joints with Cartesian coordinates plus numerical constraints, which can let joint constraints break down in complex structures. That's why fine-grained manipulation and motion-control research more often uses robotics-focused simulators such as MuJoCo and Isaac Sim.","example":"The self-driving simulator CARLA is built on Unreal Engine (upgraded to UE 5.5 starting with version 0.10.0), while the household embodied environment AI2-THOR is built on Unity.","related":["Simulator","Physics Engine","Rendering","Photorealistic Rendering","CARLA","AI2-THOR"]},{"id":"carla","category":"sim","sec":4,"tier":3,"sources":[{"title":"CARLA: An Open Urban Driving Simulator (arXiv 1711.03938)","url":"https://arxiv.org/abs/1711.03938"},{"title":"CARLA Simulator 官网","url":"https://carla.org/"},{"title":"carla-simulator/carla (GitHub)","url":"https://github.com/carla-simulator/carla"}],"as_of":"2026-09","related_ids":["autonomous-driving","simulator","game-engine","closed-loop-evaluation","end-to-end","lidar"],"name":"CARLA","alt":"CARLA","abbr":"","aliases":["CARLA Simulator"],"one_liner":"An open-source, Unreal Engine-based self-driving simulator commonly used for closed-loop testing of driving policies.","explanation":"CARLA was released by Alexey Dosovitskiy, German Ros, Vladlen Koltun, and colleagues at the first CoRL in 2017. It provides open-source digital assets for city road layouts, buildings, and vehicles, and lets users flexibly configure cameras, lidar, depth, and GPS sensors, along with weather, lighting, and other traffic participants. The original paper used it to compare three driving approaches: a traditional modular pipeline, an end-to-end imitation-learning model, and an end-to-end reinforcement-learning model. It has since become a standard platform for closed-loop self-driving evaluation, with an accompanying autonomous-driving Leaderboard. According to its website, version 0.10.0 (December 2024) upgraded to Unreal Engine 5.5, while the Unreal Engine 4-based 0.9.x line is still maintained, with 0.9.16 released in September 2025; the code is MIT-licensed and the assets are CC-BY licensed.","example":"The CARLA autonomous-driving Leaderboard has competing driving systems drive in closed loop along fixed routes and traffic scenarios, automatically scoring them on route completion and rule violations.","related":["Autonomous Driving","Simulator","Game Engine","Closed-Loop Evaluation","End-to-End","LiDAR"]},{"id":"roboverse-towards-a-unified-platform-dataset-and-benchmark-f","category":"sim","sec":4,"tier":3,"sources":[{"title":"arXiv 2504.18904 - RoboVerse","url":"https://arxiv.org/abs/2504.18904"},{"title":"GitHub - RoboVerseOrg/RoboVerse","url":"https://github.com/RoboVerseOrg/RoboVerse"},{"title":"RoboVerse 文档站","url":"https://roboverse.wiki/"}],"as_of":"2026-09","related_ids":["simulator","sim-to-sim-transfer","nvidia-isaac-lab","mujoco","sapien","benchmark"],"name":"RoboVerse (MetaSim): Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning","alt":"RoboVerse","abbr":"","aliases":["MetaSim"],"one_liner":"A robot-learning platform unifying many simulators behind one interface, with a bundled synthetic dataset and benchmark.","explanation":"RoboVerse is an open-source project released in April 2025, with authors including Berkeley's Pieter Abbeel and Jitendra Malik among researchers from multiple institutions; the paper was accepted at RSS 2025, and the code is Apache 2.0 licensed. Its foundation is called MetaSim: a simulator-agnostic middleware layer that unifies backends such as Isaac Lab, Isaac Gym, MuJoCo, SAPIEN, PyBullet, Genesis, and CoppeliaSim behind the same configuration format and API, handling environment launch, asset loading, and physics stepping. This lets the same task, robots, and assets switch between different simulators, easing the problem that each simulator's formats are mutually incompatible and benchmarks are hard to compare directly. On top of this, RoboVerse pulls together tasks and data from existing benchmarks such as LIBERO, ManiSkill, RLBench, Meta-World, and robosuite, and provides a synthetic dataset, a unified evaluation protocol, and training pipelines for imitation learning, reinforcement learning, and VLA models.","example":"Write a grasping task's configuration once in MetaSim, then just by switching the simulation-backend parameter, run that same task with the same robots and assets in both MuJoCo and Isaac Lab, and compare how the policy performs across the two.","related":["Simulator","Sim-to-Sim Transfer","NVIDIA Isaac Lab","MuJoCo (Multi-Joint dynamics with Contact)","SAPIEN (SimulAted Part-based Interactive ENvironment)","Benchmark"]},{"id":"simulation-assets","category":"sim","sec":5,"tier":2,"sources":[{"title":"NVIDIA SimReady 文档","url":"https://docs.omniverse.nvidia.com/simready/latest/index.html"},{"title":"RoboTwin 官网（RoboTwin-OD）","url":"https://robotwin-platform.github.io/"}],"as_of":"","related_ids":["simready-assets","convex-decomposition","articulated-object","objaverse","universal-scene-description","unified-robot-description-format"],"name":"Simulation Assets","alt":"仿真资产","abbr":"","aliases":["3D Assets","Scene Assets","Sim Assets"],"one_liner":"The robot, object, and scene models used in simulation, which need to look right and also carry physical properties.","explanation":"Simulation assets are the raw material for building a simulated environment: robot models (usually described with URDF or MJCF for links and joints), manipulable objects, furniture, and entire room scenes. Unlike a game model, a simulation asset needs more than an appearance mesh and material — it also needs collision geometry (complex meshes are usually broken into convex pieces via convex decomposition), mass, friction coefficients, articulated joints (such as a drawer slide or a door hinge), and semantic labels. NVIDIA's OpenUSD-based SimReady assets specifically emphasize carrying physical properties, semantic annotations, and sensor properties. The quantity and diversity of assets determines how many kinds of training scenes can be produced, which directly affects a policy's generalization; common sources include Objaverse, PartNet-Mobility, and generative 3D models.","example":"RoboTwin 2.0's object library, RoboTwin-OD, contains 731 objects: 534 custom-built meshes, 153 taken from Objaverse, and 44 articulated objects from PartNet-Mobility.","related":["SimReady Assets","Convex Decomposition","Articulated Object","Objaverse","Universal Scene Description (OpenUSD)","Unified Robot Description Format"]},{"id":"articulated-object","category":"sim","sec":5,"tier":2,"sources":[{"title":"SAPIEN: A SimulAted Part-based Interactive ENvironment (PartNet-Mobility)","url":"https://arxiv.org/abs/2003.08515"},{"title":"simpler-env/SimplerEnv (GitHub)","url":"https://github.com/simpler-env/SimplerEnv"}],"as_of":"","related_ids":["articulated-object-manipulation","partnet-mobility","revolute-joint","prismatic-joint","articulation-estimation","unified-robot-description-format"],"name":"Articulated Object","alt":"铰接物体","abbr":"","aliases":["jointed object"],"one_liner":"An object made of parts connected by joints that can rotate or slide relative to each other.","explanation":"An articulated object is made of two or more rigid parts connected by joints that let them move relative to each other — a cabinet or refrigerator door swinging on a hinge (a revolute joint), a drawer sliding out on a rail (a prismatic joint), and also things like a laptop, scissors, or a faucet. Much of what a household robot needs to handle falls into this category. Manipulating them is harder than grasping a single rigid body: the robot has to recognize which part can move, where its joint axis is, and how far it can travel, then apply force along the constrained trajectory — pull it off-axis and it jams. In simulation, articulated objects are usually described as a tree of links and joints using formats such as URDF; a common asset library is PartNet-Mobility, released by the SAPIEN team in 2020, with 46 categories and 2,346 object models annotated with joints.","example":"In SimplerEnv's Google Robot “open/close drawer” task, the drawer is connected to the cabinet body by a prismatic joint, and the robot has to pull it out or push it back in along the rail direction.","related":["Articulated Object Manipulation","PartNet-Mobility","Revolute Joint","Prismatic Joint","Articulation Estimation","Unified Robot Description Format"]},{"id":"simready-assets","category":"sim","sec":5,"tier":3,"sources":[{"title":"NVIDIA Omniverse - SimReady 文档","url":"https://docs.omniverse.nvidia.com/simready/latest/index.html"}],"as_of":"2026-09","related_ids":["simulation-assets","universal-scene-description","nvidia-isaac-sim","nvidia-omniverse","digital-twin","synthetic-data"],"name":"SimReady Assets","alt":"SimReady 资产","abbr":"","aliases":["SimReady"],"one_liner":"An NVIDIA 3D-asset specification: models ship with physics, materials, and semantic labels already attached, ready to drop into simulation.","explanation":"SimReady is a set of 3D asset conventions NVIDIA promotes within its Omniverse ecosystem. An ordinary 3D model usually has only geometry and textures, so dropping it into a simulator still requires manually adding mass, collision shapes, friction, and other parameters. SimReady assets instead require the model to be carried as OpenUSD (Universal Scene Description, a 3D scene file format), packaged together with physical properties based on USD Physics, physically based materials, semantic labels for machine learning use, and non-visual attributes relevant to sensor simulation. An asset built this way can be dragged into a simulator like Isaac Sim and immediately take part in collisions and grasping, and the synthetic images it renders already come with annotations. NVIDIA positions it as a general standard for digital twins of factories, warehouses, and data centers, and for training robotics and autonomous-driving systems.","example":"Building a warehouse scene in Isaac Sim, drag in SimReady cardboard boxes and shelving — they already carry mass, collision shapes, and semantic categories, so they can immediately be used to train a picking policy and to batch-generate synthetic images with segmentation labels.","related":["Simulation Assets","Universal Scene Description (OpenUSD)","NVIDIA Isaac Sim","NVIDIA Omniverse","Digital Twin","Synthetic Data"]},{"id":"heightfield-terrain","category":"sim","sec":5,"tier":3,"sources":[{"title":"MuJoCo XML Reference: asset/hfield","url":"https://mujoco.readthedocs.io/en/stable/XMLreference.html#asset-hfield"},{"title":"Isaac Lab API: isaaclab.terrains","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.terrains.html"}],"as_of":"","related_ids":["terrain-curriculum","rough-terrain-locomotion","elevation-map","height-scan","legged-gym","mujoco"],"name":"Heightfield Terrain","alt":"高度场地形","abbr":"","aliases":["Height Map Terrain","hfield"],"one_liner":"Representing uneven terrain with a 2D grid that stores the ground height at each point — the standard approach for legged-robot training.","explanation":"A heightfield is a 2D elevation matrix: the ground is divided into a grid, and each cell stores just one height value, much like a grayscale image where lighter means higher. Simulators use it as a cheap way to represent ramps, stairs, rubble, and other uneven terrain. MuJoCo's hfield asset can be loaded from a PNG or binary file, or generated at runtime from sensor data, and is treated as a set of triangular prisms during collision detection; Isaac Lab instead converts a heightfield into a triangle mesh, and comes with built-in terrains such as random rough ground, pyramid slopes, stairs, discrete obstacles, waves, and stepping stones, each with a difficulty parameter, making it easy to train from easy to hard using a terrain curriculum. The limitation is that a heightfield can only have one height per horizontal position, so it can't represent overhangs or caves.","example":"Isaac Lab's pyramid stairs terrain has taller steps as its difficulty parameter increases; when training quadruped locomotion, it's often mixed with randomly rough terrain and difficulty is ramped up gradually via a terrain curriculum.","related":["Terrain Curriculum","Rough-terrain Locomotion","Elevation Map","Height Scan","legged_gym","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"ai2-thor","category":"sim","sec":5,"tier":2,"sources":[{"title":"AI2-THOR 官网","url":"https://ai2thor.allenai.org/"},{"title":"AI2-THOR: An Interactive 3D Environment for Visual AI (arXiv 1712.05474)","url":"https://arxiv.org/abs/1712.05474"}],"as_of":"2026-09","related_ids":["procthor","alfred","habitat","object-goal-navigation","embodied-question-answering","allen-institute-for-ai"],"name":"AI2-THOR","alt":"AI2-THOR","abbr":"","aliases":["ai2thor","iTHOR"],"one_liner":"The Allen Institute for AI's interactive indoor 3D simulator, widely used for navigation and household tasks.","explanation":"AI2-THOR is an open-source indoor simulation platform from the Allen Institute for AI (Ai2), with its paper published in 2017, built on the Unity engine. It offers near-photorealistic kitchen, bedroom, bathroom, and living-room scenes that an agent can walk through and interact with — opening a cabinet door, picking up and putting down objects, turning appliances on and off — with objects' open/closed and hot/cold states, among others, changing accordingly. It emphasizes high-level semantic interaction over fine-grained contact physics, so it's mostly used for visual navigation, embodied question answering, instruction following, and task planning. It includes several sub-frameworks: iTHOR (120 rooms with over 2,000 object types), RoboTHOR (apartment scenes with a matching physical test space), ManipulaTHOR (adds robot-arm manipulation), and ProcTHOR (procedurally generates large numbers of houses).","example":"In an iTHOR kitchen scene, an agent can execute a sequence like “walk to the fridge,” “open the fridge,” and “pick up the apple,” with the fridge door's open/closed state changing accordingly.","related":["ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","ALFRED","Habitat","Object-Goal Navigation","Embodied Question Answering","Allen Institute for AI"]},{"id":"virtualhome-simulating-household-activities-via-programs","category":"sim","sec":5,"tier":3,"sources":[{"title":"VirtualHome: Simulating Household Activities via Programs (arXiv 1806.07011, CVPR 2018)","url":"https://arxiv.org/abs/1806.07011"},{"title":"VirtualHome GitHub 仓库","url":"https://github.com/xavierpuigf/virtualhome"}],"as_of":"2026-09","related_ids":["household-tasks","llm-based-task-planning","long-horizon-task","ai2-thor","alfred","simulator"],"name":"VirtualHome: Simulating Household Activities via Programs","alt":"VirtualHome 家庭活动仿真","abbr":"","aliases":["VirtualHome"],"one_liner":"A Unity-based simulation platform where virtual humans perform household chores driven by step-by-step 'programs.'","explanation":"VirtualHome is a household-activity simulation platform proposed by MIT and the University of Toronto at CVPR 2018, built on the Unity3D game engine and controlled through a Python API. It writes chores like “making breakfast” or “watching TV” as “programs” — sequences of atomic actions such as walking to something, picking it up, opening it, and switching it on or off — which a virtual human then executes step by step inside a simulated apartment, rendering an annotated video as it goes. It is concerned with high-level task planning rather than simulating a robot arm's contact mechanics, so it is often used to test how well large language models decompose tasks. The current version, 2.3, supports procedural scene generation, and it has also spun off a human-robot collaboration benchmark called Watch-And-Help.","example":"For instance, “watch TV” can be written as the program “walk to the living room → pick up the remote → turn on the TV → sit on the sofa”; the virtual human executes it step by step, and the platform outputs the corresponding video and annotations.","related":["Household Tasks","LLM-based Task Planning","Long-horizon Task","AI2-THOR","ALFRED","Simulator"]},{"id":"habitat","category":"sim","sec":5,"tier":2,"sources":[{"title":"AI Habitat 官网","url":"https://aihabitat.org/"},{"title":"Habitat: A Platform for Embodied AI Research (arXiv 1904.01201)","url":"https://arxiv.org/abs/1904.01201"},{"title":"Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots (arXiv 2310.13724)","url":"https://arxiv.org/abs/2310.13724"}],"as_of":"2023-10","related_ids":["point-goal-navigation","object-goal-navigation","habitat-matterport-3d-dataset","social-navigation","rearrangement","matterport3d"],"name":"Habitat","alt":"Habitat","abbr":"","aliases":["Habitat-Sim","Habitat-Lab","Habitat 3.0","AI Habitat"],"one_liner":"Meta's open-source embodied-AI simulation platform, used mainly for indoor navigation, object rearrangement, and human-robot collaboration tasks.","explanation":"Habitat is an open-source embodied-AI simulation platform developed by Meta (formerly Facebook AI Research) together with several universities, first described in a paper at ICCV 2019. It has two layers: Habitat-Sim is a high-speed 3D simulator that can load real indoor scans from datasets such as Matterport3D, HM3D, Gibson, and Replica, and can render over 10,000 frames per second using multiple processes on a single GPU; Habitat-Lab is the upper-level library used to define tasks, configure agents, and handle training and evaluation. Habitat 2.0 (2021) added physical interaction and household rearrangement tasks, and Habitat 3.0 (2023) added humanoid avatars and human-in-the-loop control, supporting social navigation and human-robot collaborative rearrangement. It is one of the main platforms for visual navigation research, and hosts an annual Habitat Challenge that requires participants to submit code for evaluation in scenes it has never seen.","example":"In a point-goal navigation task, an agent is placed in a previously unseen apartment scan from HM3D and told only the target's coordinates relative to its starting point; it must reach the target using nothing but its camera feed and depth map.","related":["Point-Goal Navigation","Object-Goal Navigation","Habitat-Matterport 3D Dataset","Social Navigation","Rearrangement","Matterport3D"]},{"id":"igibson","category":"sim","sec":5,"tier":3,"sources":[{"title":"iGibson project page (Stanford SVL)","url":"https://svl.stanford.edu/igibson/"},{"title":"iGibson 1.0 (arXiv 2012.02924)","url":"https://arxiv.org/abs/2012.02924"},{"title":"BEHAVIOR (Stanford) — OmniGibson on Isaac Sim","url":"https://behavior.stanford.edu/"}],"as_of":"2026-09","related_ids":["omnigibson","behavior-1k","habitat","ai2-thor","mobile-manipulation","pybullet"],"name":"iGibson","alt":"iGibson","abbr":"","aliases":["Interactive Gibson","iGibson 1.0","iGibson 2.0"],"one_liner":"Stanford's interactive indoor simulation platform, used for navigation and mobile-manipulation research in household settings.","explanation":"iGibson was developed by Stanford's Vision and Learning Lab (Fei-Fei Li, Silvio Savarese, and colleagues), as an interactive version of the earlier Gibson environment. Version 1.0 (IROS 2021) provides 15 fully interactive scenes reconstructed from real homes, totaling 108 rooms, where furniture, cabinet doors, drawers, and other rigid and articulated objects can all be pushed or opened by the robot; it can output RGB, depth, segmentation, and lidar sensor signals, with physics based on the Bullet engine, and it comes with its own motion planner. Version 2.0 (CoRL 2021) added object states such as temperature, wetness, cleanliness, and being sliced, and supports collecting human demonstrations with VR, used for the BEHAVIOR household-task benchmark. Stanford's BEHAVIOR-1K later switched to OmniGibson, built on NVIDIA Isaac Sim, and most new projects have since moved to that instead.","example":"In an iGibson household scene, a mobile robot is made to navigate to the kitchen and open a cabinet door, using the built-in motion planner to generate a collision-free arm trajectory.","related":["OmniGibson","BEHAVIOR-1K (BEHAVIOR Challenge)","Habitat","AI2-THOR","Mobile Manipulation","PyBullet"]},{"id":"omnigibson","category":"sim","sec":5,"tier":2,"sources":[{"title":"BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation (arXiv 2403.09227)","url":"https://arxiv.org/abs/2403.09227"},{"title":"BEHAVIOR 项目主页","url":"https://behavior.stanford.edu/"},{"title":"StanfordVL/BEHAVIOR-1K (GitHub)","url":"https://github.com/StanfordVL/BEHAVIOR-1K"}],"as_of":"2026-09","related_ids":["behavior-1k","igibson","nvidia-isaac-sim","nvidia-omniverse","household-tasks","long-horizon-task"],"name":"OmniGibson","alt":"OmniGibson","abbr":"","aliases":["BEHAVIOR-1K Simulator"],"one_liner":"Stanford's Omniverse-based household simulator, the underlying engine behind the BEHAVIOR-1K benchmark.","explanation":"OmniGibson is a household-environment simulation platform developed by Stanford's Vision and Learning Lab (SVL, including Fei-Fei Li), built on NVIDIA Omniverse and PhysX 5. It is the underlying simulator for the BEHAVIOR-1K household-task benchmark, replacing the earlier iGibson 2.0. Its physics coverage is broad: rigid bodies, deformable objects, cloth, fluids, and granular materials, plus thermal effects such as fire, steam, and smoke, and it tracks extended object states such as temperature, being cooked or burnt, being soaked, on/off, being cut, and being dirty. The paper notes that more than half of BEHAVIOR-1K's activities cannot be simulated at all without fluids, soft bodies, and cloth. Scenes are rendered with ray tracing or path tracing. BEHAVIOR-1K contains 1,000 everyday activities, 50 interactive scenes, and over 10,000 objects, and it is the platform behind the annual BEHAVIOR Challenge.","example":"In a cooking task where the robot has to put ingredients in a pot and heat them, OmniGibson continuously tracks the ingredients' temperature, marks their state as “cooked” once it crosses a threshold, and judges task completion based on states like this.","related":["BEHAVIOR-1K (BEHAVIOR Challenge)","iGibson","NVIDIA Isaac Sim","NVIDIA Omniverse","Household Tasks","Long-horizon Task"]},{"id":"internutopia","category":"sim","sec":5,"tier":3,"sources":[{"title":"InternRobotics/InternUtopia (GitHub)","url":"https://github.com/InternRobotics/InternUtopia"},{"title":"GRUtopia: Dream General Robots in a City at Scale (arXiv 2407.10943)","url":"https://arxiv.org/abs/2407.10943"},{"title":"百度百科：浦源·桃源","url":"https://baike.baidu.com/item/%E6%B5%A6%E6%BA%90%C2%B7%E6%A1%83%E6%BA%90"}],"as_of":"2025-07","related_ids":["nvidia-isaac-sim","shanghai-artificial-intelligence-laboratory","social-navigation","mobile-manipulation","object-goal-navigation","internvla"],"name":"InternUtopia","alt":"InternUtopia","abbr":"","aliases":["GRUtopia"],"one_liner":"Shanghai AI Lab's city-scale embodied-AI simulation platform built on Isaac Sim, formerly named GRUtopia.","explanation":"InternUtopia was developed by Shanghai AI Lab, previously named GRUtopia (Chinese name 浦源·桃源); it was released with a public paper at the World Artificial Intelligence Conference in July 2024, shipped version 2.0 in February 2025, and was renamed to InternUtopia with version 2.2.0 in July of that year. It's built on top of NVIDIA Isaac Sim and has three parts: GRScenes, a scene library of about 100,000 finely annotated interactive scenes across 89 categories, extending from homes to service settings like supermarkets and hospitals; GRResidents, virtual residents (NPCs) driven by a large language model that can converse with a robot and generate tasks; and the GRBench benchmark, focused on legged robots, with three task categories — object navigation, social navigation, and mobile manipulation. The platform also supports VR teleoperation and hand motion-capture control, and is used together with companion projects InternNav and InternManip.","example":"In GRBench's social-navigation task, a robot can ask a GRResidents-driven virtual resident for information about a target object before walking to its location.","related":["NVIDIA Isaac Sim","Shanghai Artificial Intelligence Laboratory","Social Navigation","Mobile Manipulation","Object-Goal Navigation","InternVLA (Shanghai AI Laboratory)"]},{"id":"procedural-generation","category":"sim","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Procedural generation","url":"https://en.wikipedia.org/wiki/Procedural_generation"},{"title":"ProcTHOR: Large-Scale Embodied AI Using Procedural Generation (arXiv 2206.06994)","url":"https://arxiv.org/abs/2206.06994"},{"title":"Infinigen","url":"https://infinigen.org/"}],"as_of":"","related_ids":["procthor","infinigen","domain-randomization","generative-simulation","scene-generalization","terrain-curriculum"],"name":"Procedural Generation","alt":"程序化生成","abbr":"","aliases":["Procedural Scene Generation","Procedural Content Generation","PCG"],"one_liner":"Automatically producing large numbers of scenes, objects, or terrain from rules plus randomness, instead of building each one by hand.","explanation":"Procedural generation originated in the game industry: content is produced automatically by an algorithm plus randomness, the way Minecraft generates an entire map from a single random seed. In embodied AI, it is used to mass-produce simulated training environments: rules are written first (how to partition a floor plan, where furniture goes, which materials and textures to use, how to light the scene), and then thousands of distinct scenes are randomly sampled from those rules. Building one interactive 3D house by hand is slow, and too few scenes make a policy prone to overfitting, so procedural generation trades rule-writing effort for quantity and diversity, making it a standard tool for improving generalization across scenes. Notable projects include the Allen Institute for AI's ProcTHOR and Princeton's Infinigen; the randomly generated stairs and slopes used in quadruped reinforcement learning also fall into this category. Its counterpart is generative simulation, which instead uses a large model to build scenes from text descriptions.","example":"ProcTHOR can randomly generate floor plans of 1 to 10 rooms, sampling and arranging from 108 categories, 1,633 interactive objects, and 3,278 materials; the paper pretrained an embodied agent on 10,000 generated houses, and without any fine-tuning on downstream data it often beat the previous best methods.","related":["ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","Infinigen","Domain Randomization","Generative Simulation","Scene Generalization","Terrain Curriculum"]},{"id":"procthor","category":"sim","sec":5,"tier":3,"sources":[{"title":"ProcTHOR: Large-Scale Embodied AI Using Procedural Generation (arXiv 2206.06994)","url":"https://arxiv.org/abs/2206.06994"},{"title":"ProcTHOR 项目主页","url":"https://procthor.allenai.org/"},{"title":"NeurIPS 2022 Awards","url":"https://neurips.cc/virtual/2022/awards_detail"}],"as_of":"2024-06","related_ids":["ai2-thor","procedural-generation","object-goal-navigation","rearrangement","poliformer","holodeck"],"name":"ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","alt":"ProcTHOR","abbr":"","aliases":["ProcTHOR-10K"],"one_liner":"Allen Institute for AI's framework for procedurally generating large numbers of interactive 3D houses for embodied-AI training.","explanation":"ProcTHOR was released by the Allen Institute for AI (Ai2) in June 2022 and won an Outstanding Paper award at NeurIPS 2022; it is built on top of the AI2-THOR simulator. Before it, embodied agents could only train in a few dozen to a few hundred hand-built scenes, which made them prone to memorizing the map and failing in a new house. ProcTHOR instead uses procedural generation (building content automatically from rules plus random sampling) to mass-produce houses: it first decides room types and counts, then generates the floor plan and places furniture and small objects, randomizing materials and lighting as it goes, drawing from a library of 108 categories and 1,633 interactive objects. The team publicly released ProcTHOR-10K, a set of 10,000 houses. Agents pretrained purely on RGB images inside these houses achieved then-state-of-the-art results on 6 benchmarks — object navigation, object rearrangement, and arm-pointing navigation among them — and performed strongly zero-shot (with no fine-tuning) on some tasks. Later navigation models such as PoliFormer were also trained with reinforcement learning inside its generated houses.","example":"Specify “two bedrooms plus a kitchen and living room,” and ProcTHOR automatically generates the floor plan, furnishes it with a bed, a sofa, and a refrigerator, and randomizes the flooring material and lighting — producing a new house an agent can enter, open the fridge in, and search for an apple.","related":["AI2-THOR","Procedural Generation","Object-Goal Navigation","Rearrangement","PoliFormer","Holodeck"]},{"id":"infinigen","category":"sim","sec":5,"tier":3,"sources":[{"title":"Infinigen official site","url":"https://infinigen.org/"},{"title":"Infinite Photorealistic Worlds using Procedural Generation (arXiv 2306.09310)","url":"https://arxiv.org/abs/2306.09310"},{"title":"princeton-vl/infinigen (GitHub)","url":"https://github.com/princeton-vl/infinigen"}],"as_of":"2025-05","related_ids":["procedural-generation","synthetic-data","simulation-assets","articulated-object","blender","generative-simulation"],"name":"Infinigen","alt":"Infinigen 程序化世界生成","abbr":"","aliases":["Infinite Photorealistic Worlds using Procedural Generation","Infinigen Indoors","Infinigen-Articulated","Infinigen-Sim"],"one_liner":"Princeton's open-source procedural 3D world generator, where every asset is generated from scratch by randomized mathematical rules.","explanation":"Infinigen was developed by Princeton's Vision and Learning Lab (Jia Deng's group), built on Blender, and open-sourced under the BSD license. The first version (CVPR 2023) generates natural outdoor scenes — terrain, plants, animals, and phenomena like fire, clouds, rain, and snow — with every asset's shape and texture generated from scratch by randomized mathematical rules, using no external assets at all, which allows infinite variation and automatic output of ground-truth depth, optical flow, segmentation, and normals. Infinigen Indoors (CVPR 2024) extended this indoors, using a constraint description language and a solver to place furniture and appliances, and can export to real-time simulators such as Omniverse for training embodied agents. Infinigen-Articulated (also called Infinigen-Sim, 2025) further procedurally generates articulated, simulation-ready objects, exportable as URDF, MJCF, or USD for use in MuJoCo and Isaac.","example":"Running the same indoor-generation rules with a different random seed produces a new room with a different layout, furniture shapes, and materials, complete with pixel-accurate ground-truth depth and segmentation, ready to use as visual training data.","related":["Procedural Generation","Synthetic Data","Simulation Assets","Articulated Object","Blender","Generative Simulation"]},{"id":"generative-simulation","category":"sim","sec":5,"tier":2,"sources":[{"title":"Towards Generalist Robots: A Promising Paradigm via Generative Simulation (arXiv 2305.10455)","url":"https://arxiv.org/abs/2305.10455"},{"title":"RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation (arXiv 2311.01455)","url":"https://arxiv.org/abs/2311.01455"}],"as_of":"","related_ids":["procedural-generation","synthetic-data","simulation-data","robogen","holodeck","genesis"],"name":"Generative Simulation","alt":"生成式仿真","abbr":"","aliases":["Automatic Simulation-Scene Generation","Automatic Simulation-Task Generation"],"one_liner":"Using foundation models to automatically generate simulation tasks, scenes, and training supervision, producing robot training data at scale.","explanation":"Generative simulation is an approach proposed by Zhou Xian and colleagues in a 2023 paper: rather than having a large model output robot actions directly, use foundation models — language models, image and 3D generative models — to fully automatically generate diverse tasks, scenes, and training supervision (such as reward functions and demonstration trajectories), and learn skills at scale inside simulation. It targets the problem that manually building scenes and designing tasks and rewards is slow and limits data diversity. A typical pipeline: a large model first proposes a task, then retrieves or generates 3D assets to build the scene, breaks the task into sub-steps, and uses reinforcement learning or motion planning to generate data and train a policy. Notable projects include RoboGen, GenSim, and Holodeck. It differs from procedural generation (building randomized scenes from hand-written rules) in that it is driven by foundation models rather than fixed rules.","example":"RoboGen's loop works like this: a large model proposes a task and skill to learn, selects relevant object assets and arranges them into a plausible simulated scene, breaks the task into sub-tasks, then chooses its own learning method, generates training supervision, and trains a policy — all without manual task design.","related":["Procedural Generation","Synthetic Data","Simulation Data","RoboGen","Holodeck","Genesis"]},{"id":"holodeck","category":"sim","sec":5,"tier":3,"sources":[{"title":"Holodeck: Language Guided Generation of 3D Embodied AI Environments (arXiv 2312.09067)","url":"https://arxiv.org/abs/2312.09067"},{"title":"allenai/Holodeck (GitHub)","url":"https://github.com/allenai/Holodeck"},{"title":"Holodeck project page","url":"https://yueyang.ai/holodeck/"}],"as_of":"2024-06","related_ids":["ai2-thor","procthor","objaverse","generative-simulation","procedural-generation","object-goal-navigation"],"name":"Holodeck","alt":"Holodeck 语言生成三维环境","abbr":"","aliases":["Language Guided Generation of 3D Embodied AI Environments"],"one_liner":"Given a one-sentence description, automatically generates an interactive 3D indoor scene using GPT-4 and Objaverse assets.","explanation":"Holodeck is a collaboration between the University of Pennsylvania, Stanford, the University of Washington, and the Allen Institute for AI (Ai2), published at CVPR 2024. Embodied agents need large numbers of diverse training scenes, but building 3D rooms by hand is expensive. Holodeck lets a user type a description like “apartment for a researcher with a cat,” and GPT-4 supplies the common-sense knowledge: room layout, wall and floor materials, doors and windows, which objects belong in the scene, and spatial-relationship constraints between objects (such as a chair being next to a table); an optimization algorithm then solves for object placement, with objects retrieved from the Objaverse 3D asset library. The generated scene is loaded and used through AI2-THOR. In the paper's experiments, an object-goal navigation agent trained on Holodeck scenes generalized better zero-shot to new scene types such as music rooms and daycares.","example":"Given the input “apartment for a researcher with a cat,” Holodeck automatically plans the rooms, chooses wall and floor materials, picks furniture from Objaverse, and arranges it according to the constraints; the resulting scene can be loaded directly in AI2-THOR.","related":["AI2-THOR","ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","Objaverse","Generative Simulation","Procedural Generation","Object-Goal Navigation"]},{"id":"molmospaces","category":"sim","sec":5,"tier":3,"sources":[{"title":"MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation (arXiv 2602.11337)","url":"https://arxiv.org/abs/2602.11337"},{"title":"Ai2 Blog: MolmoSpaces","url":"https://allenai.org/blog/molmospaces"},{"title":"MolmoBot: Large-Scale Simulation Enables Zero-Shot Manipulation (arXiv 2603.16861)","url":"https://arxiv.org/abs/2603.16861"}],"as_of":"2026-03","related_ids":["allen-institute-for-ai","simulation-based-evaluation","sim-to-real-correlation","procthor","objaverse","mobile-manipulation"],"name":"MolmoSpaces","alt":"MolmoSpaces","abbr":"","aliases":["MolmoSpaces-Bench"],"one_liner":"The Allen Institute for AI's open ecosystem of large-scale indoor simulation scenes and robot evaluation tools.","explanation":"MolmoSpaces is an open ecosystem released by the Allen Institute for AI (Ai2) in February 2026, for generating data and training and evaluating robot policies at scale. It includes over 230,000 indoor scenes (drawn from iTHOR, ProcTHOR, Holodeck, and others) and 130,000 annotated object models, of which 48,000 are manipulable objects paired with 42 million stable grasp poses. Scenes aren't tied to any one simulator, supporting MuJoCo, ManiSkill, and Isaac. Its companion benchmark, MolmoSpaces-Bench, contains 8 tasks covering tabletop and mobile manipulation, navigation, and cross-room long-horizon tasks; the paper reports that its simulated scores correlate strongly with real-robot scores (R = 0.96), and finds that policies are quite sensitive to instruction wording, initial joint positions, and camera occlusion.","example":"Ai2's MolmoBot procedurally generates 1.8 million simulated trajectories in MolmoSpaces and deploys directly to a Franka FR3 with no real-robot data at all, reaching a real-robot tabletop pick-and-place success rate of 79.2%, compared with 39.2% for π0.5.","related":["Allen Institute for AI","Simulation-Based Evaluation","Sim-to-Real Correlation","ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","Objaverse","Mobile Manipulation"]},{"id":"genie-sim","category":"sim","sec":5,"tier":2,"sources":[{"title":"AgibotTech/genie_sim (GitHub)","url":"https://github.com/AgibotTech/genie_sim"},{"title":"Genie Sim 3.0: A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot (arXiv 2601.02078)","url":"https://arxiv.org/abs/2601.02078"}],"as_of":"2026-08","related_ids":["agibot","nvidia-isaac-sim","simulation-based-evaluation","synthetic-data","gaussian-splatting-based-simulation","generative-simulation"],"name":"Genie Sim","alt":"智元 Genie Sim","abbr":"","aliases":["AgiBot Genie Sim","Genie Sim Benchmark","Genie Sim 3.0"],"one_liner":"AgiBot's open-source embodied-simulation platform, combining scene generation, synthetic data, and an evaluation benchmark.","explanation":"Genie Sim is an open-source robot manipulation simulation platform from AgiBot, built on NVIDIA Isaac Sim, with code released under the MPL 2.0 license. Version 2.x shipped in 2025, and Genie Sim 3.0 shipped with a public technical report in January 2026, later updated to 3.2. It bundles several pieces together: a generator that turns natural-language descriptions into simulated scenes using a large model; 3D Gaussian Splatting to reconstruct simulation assets from real-world scenes; an evaluation benchmark covering over 200 tasks and more than 100,000 scenes that is scored automatically by a vision-language model; and an open-sourced synthetic dataset of over 10,000 hours. The goal is to let simulated data substitute for part of the real-robot data needed to train policies, and AgiBot has reported results on zero-shot sim-to-real transfer.","example":"A developer writes a description such as “arrange several rows of beverage bottles in front of a supermarket shelf,” and Genie Sim's generator retrieves assets and outputs a .usda scene file; AgiBot's Genie G2 robot model is then used in that scene to collect synthetic data or run evaluations.","related":["AgiBot","NVIDIA Isaac Sim","Simulation-Based Evaluation","Synthetic Data","Gaussian Splatting-based Simulation","Generative Simulation"]},{"id":"sim-to-real-gap","category":"sim","sec":6,"tier":1,"sources":[{"title":"Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey (arXiv 2009.13303)","url":"https://arxiv.org/abs/2009.13303"},{"title":"Lilian Weng: Domain Randomization for Sim2Real Transfer","url":"https://lilianweng.github.io/posts/2019-05-05-domain-randomization/"}],"as_of":"","related_ids":["sim-to-real-transfer","domain-randomization","system-identification","domain-adaptation","actuator-modeling","real-to-sim"],"name":"Sim-to-Real Gap (Reality Gap)","alt":"虚实差距","abbr":"","aliases":["Reality Gap","Sim2Real Gap"],"one_liner":"The mismatch between simulation and reality that makes a policy trained in sim perform worse on a real robot.","explanation":"The sim-to-real gap refers to the mismatch between a simulated environment and the real world, which causes a policy trained in simulation to perform worse, or fail outright, once it's transferred to a real robot. The gap has several sources: inaccurate physical parameters (friction, mass, damping, motor characteristics), simplifications in the physics modeling itself (such as soft contact or flexible objects), visual differences (rendered lighting and textures don't match a real camera), and factors that simulation doesn't model at all, such as sensor noise and control latency. Because simulated training is cheap, safe, and can run at massive parallel scale, closing this gap is the central problem in sim-to-real transfer. Common countermeasures include system identification (calibrating simulation parameters to match reality), domain randomization (training the policy to handle a wide range of parameters), domain adaptation (shifting simulated data's distribution to look more like real data), and fine-tuning with a small amount of real-robot data.","example":"A quadruped locomotion policy trained in simulation shakes in place or even falls over on the real robot because the real motors respond more slowly than the simulated ones did — a textbook case of the sim-to-real gap.","related":["Sim-to-Real Transfer","Domain Randomization","System Identification","Domain Adaptation","Actuator Modeling (Actuator Network)","Real-to-Sim"]},{"id":"sim-to-real-transfer","category":"sim","sec":6,"tier":1,"sources":[{"title":"Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World (Tobin et al., 2017)","url":"https://arxiv.org/abs/1703.06907"},{"title":"Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: a Survey (Zhao et al., 2020)","url":"https://arxiv.org/abs/2009.13303"},{"title":"Learning agile and dynamic motor skills for legged robots (Hwangbo et al., 2019)","url":"https://arxiv.org/abs/1901.08652"}],"as_of":"","related_ids":["sim-to-real-gap","domain-randomization","system-identification","actuator-modeling","real-to-sim","teacher-student-distillation"],"name":"Sim-to-Real Transfer","alt":"仿真到现实迁移","abbr":"Sim2Real","aliases":["Sim2Real"],"one_liner":"Training a robot policy in simulation, then deploying it to work on a real robot.","explanation":"Sim-to-real transfer means training a robot policy in a simulator first, then deploying it on a real robot. Simulation can run thousands of environments at once, and a failure costs nothing, so reinforcement-learning policies for quadruped and humanoid walking and dexterous manipulation are mostly trained in simulation first. The difficulty is the sim-to-real gap: friction, mass, motor response, and visuals in simulation never fully match reality, and a policy can fail once it reaches a real robot. Common countermeasures include domain randomization (randomizing physical and visual parameters during training so the real world is just one more case the policy has seen), system identification (calibrating simulation parameters from real-robot data), actuator modeling (using real-robot data to model the simulated motors so they respond like the real ones), and teacher-student distillation (first training a teacher policy that can see the simulator's full state, then having a student policy that only uses sensors available on the real robot imitate it). In 2017, Tobin and colleagues trained an object detector using only simulated images with randomized textures and achieved about 1.5 cm of localization error in the real world — an early landmark result.","example":"ETH Zurich's ANYmal quadruped: a walking policy is trained with reinforcement learning in simulation and combined with an actuator network learned from real-robot data, then deployed directly to the real robot, where it runs faster than earlier methods and can even get back up after falling.","related":["Sim-to-Real Gap (Reality Gap)","Domain Randomization","System Identification","Actuator Modeling (Actuator Network)","Real-to-Sim","Teacher-Student Distillation"]},{"id":"real-to-sim","category":"sim","sec":6,"tier":2,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"},{"title":"RialTo: Real-to-Sim-to-Real Approach for Robust Manipulation (arXiv 2403.03949)","url":"https://arxiv.org/abs/2403.03949"}],"as_of":"","related_ids":["real-to-sim-to-real","digital-twin","sim-to-real-transfer","sim-to-real-gap","simplerenv","gaussian-splatting-based-simulation"],"name":"Real-to-Sim","alt":"现实到仿真","abbr":"Real2Sim","aliases":["Real2Sim","Real-to-Sim Reconstruction"],"one_liner":"Recreating a real scene, its objects, and the robot inside a simulator, so the simulation matches reality as closely as possible.","explanation":"Real-to-sim is the reverse step of sim-to-real: information is captured from the real world and used to build a matching version inside a simulator. It involves geometry and appearance reconstruction (scanning a room and its objects with a phone, NeRF, or 3D Gaussian Splatting), identifying physical parameters (mass, friction, joint damping, and so on), and aligning the robot's controller; the resulting simulated version is often called a digital twin. It serves two main purposes. One is evaluation: a real test setup is recreated in simulation, and the simulated score is used to predict real-robot performance, saving a large amount of real-robot testing. The other is training: data is generated or reinforcement learning is run inside the recreated scene, and the result is transferred back to the real robot — this full loop is called real-to-sim-to-real. The difficulty is that any error in the reconstructed visuals or physics becomes a new sim-to-real gap of its own.","example":"SIMPLER (SimplerEnv) builds matching simulated environments for real experiments with the Google Robot and the WidowX arm, focusing on narrowing both control and visual gaps; the paper shows that policy performance inside it correlates strongly with real-robot results and can even reproduce how sensitive a policy is to various distribution shifts.","related":["Real-to-Sim-to-Real","Digital Twin","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","SimplerEnv","Gaussian Splatting-based Simulation"]},{"id":"real-to-sim-to-real","category":"sim","sec":6,"tier":2,"sources":[{"title":"Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation (RialTo, arXiv 2403.03949)","url":"https://arxiv.org/abs/2403.03949"},{"title":"RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator (arXiv 2411.11839)","url":"https://arxiv.org/abs/2411.11839"}],"as_of":"","related_ids":["real-to-sim","sim-to-real-transfer","digital-twin","gaussian-splatting-based-simulation","reinforcement-fine-tuning","sim-to-real-gap"],"name":"Real-to-Sim-to-Real","alt":"真-仿-真闭环","abbr":"Real2Sim2Real","aliases":["Real2Sim2Real"],"one_liner":"Recreating a real scene in simulation, training or generating data there, and then deploying the resulting policy back to the real robot.","explanation":"Real-to-sim-to-real chains real-to-sim and sim-to-real into a single pipeline: a real environment is scanned and rebuilt as a digital twin inside a simulator; reinforcement learning, large-scale synthetic-data generation, or safe evaluation is then carried out in that recreated scene; and the resulting policy is finally transferred back to the real robot. Compared with plain sim-to-real, the training scene is a direct copy of the deployment scene, so the sim-to-real gap is smaller; compared with using only real-robot data, simulation allows cheap trial and error and deliberately introduced perturbations, which improves robustness. A common recent approach reconstructs photorealistic imagery with 3D Gaussian Splatting and pairs it with a physics engine to handle interaction. The limitation is that every new scene requires reconstructing it again from scratch, and physical parameters such as friction and softness remain hard to identify accurately.","example":"MIT's RialTo builds a scene's digital twin in about 25 minutes using scanning tools like Polycam, brings roughly 15 real demonstrations into simulation to fine-tune with reinforcement learning, and then distills the result back into a vision-based real-robot policy; the paper reports over 67% higher robustness than pure imitation learning across 8 tasks such as stacking plates and placing books on a shelf.","related":["Real-to-Sim","Sim-to-Real Transfer","Digital Twin","Gaussian Splatting-based Simulation","Reinforcement Fine-Tuning (RL Fine-Tuning)","Sim-to-Real Gap (Reality Gap)"]},{"id":"domain-randomization","category":"sim","sec":6,"tier":1,"sources":[{"title":"Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World (arXiv 1703.06907)","url":"https://arxiv.org/abs/1703.06907"},{"title":"Lilian Weng: Domain Randomization for Sim2Real Transfer","url":"https://lilianweng.github.io/posts/2019-05-05-domain-randomization/"}],"as_of":"","related_ids":["sim-to-real-gap","sim-to-real-transfer","dynamics-randomization","visual-randomization","automatic-domain-randomization","system-identification"],"name":"Domain Randomization","alt":"域随机化","abbr":"DR","aliases":["DR"],"one_liner":"Randomizing appearance and physics in simulation during training so the real world looks like just another variation.","explanation":"Domain randomization was proposed by Josh Tobin and colleagues at OpenAI in 2017 (IROS 2017). Instead of trying to make simulation look exactly like reality, it heavily randomizes rendering parameters — textures, lighting, camera settings — and dynamics parameters such as mass, friction, and damping (the latter often called dynamics randomization) during training, so the model sees enough variation that the real world just looks like one more instance of it. The original paper trained an object-localization network purely on randomized simulated RGB images, and after transferring to a real robot, achieved about 1.5 cm of localization accuracy. It's one of the most common techniques for sim-to-real transfer, and frameworks such as Isaac Lab include it out of the box. The difficulty is that the randomization range is usually tuned by hand — too wide, and the policy becomes overly conservative; automatic domain randomization lets the range expand on its own as training progresses.","example":"When training a quadruped locomotion policy, randomize ground friction, body mass, and motor delay independently for every parallel environment, so the resulting policy is less likely to fall over on a real robot.","related":["Sim-to-Real Gap (Reality Gap)","Sim-to-Real Transfer","Dynamics Randomization","Visual Randomization","Automatic Domain Randomization","System Identification"]},{"id":"dynamics-randomization","category":"sim","sec":6,"tier":2,"sources":[{"title":"Sim-to-Real Transfer of Robotic Control with Dynamics Randomization (arXiv 1710.06537)","url":"https://arxiv.org/abs/1710.06537"},{"title":"legged_gym: legged_robot_config.py（domain_rand）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["domain-randomization","sim-to-real-transfer","sim-to-real-gap","visual-randomization","automatic-domain-randomization","system-identification"],"name":"Dynamics Randomization","alt":"动力学随机化","abbr":"","aliases":["physics parameter randomization","friction/mass randomization"],"one_liner":"Randomizing physical parameters like mass and friction during training so a policy can handle a real robot.","explanation":"Dynamics randomization is a form of domain randomization that specifically randomizes physical parameters rather than visual appearance. A simulation's mass, friction, damping, motor gains, and delay can never fully match a real robot, and training under one fixed set of these values leaves a policy prone to overfitting to the simulator. The approach is to resample these parameters from a chosen range at the start of every episode, forcing the policy to succeed under a wide variety of physical conditions; if the real robot's true parameters fall within that range, the policy is more likely to work on it directly. A landmark example is Peng and colleagues' 2018 paper, which randomized 95 parameters for a puck-pushing task and deployed a policy trained purely in simulation directly onto a Fetch robot arm, with performance close to simulation. Today, frameworks such as Isaac Lab and legged_gym routinely randomize ground friction and payload mass, and apply random pushes, when training legged robots.","example":"legged_gym by default samples ground friction uniformly between 0.5 and 1.25, and randomly shoves the robot every 15 seconds, forcing the policy to learn to keep walking across different surfaces and under unexpected disturbances.","related":["Domain Randomization","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Visual Randomization","Automatic Domain Randomization","System Identification"]},{"id":"visual-randomization","category":"sim","sec":6,"tier":2,"sources":[{"title":"Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World (arXiv 1703.06907)","url":"https://arxiv.org/abs/1703.06907"},{"title":"Isaac Lab source: envs/mdp/events.py (randomize_visual_texture_material / randomize_visual_color)","url":"https://github.com/isaac-sim/IsaacLab/blob/main/source/isaaclab/isaaclab/envs/mdp/events.py"}],"as_of":"","related_ids":["domain-randomization","dynamics-randomization","sim-to-real-transfer","sim-to-real-gap","visual-generalization","variant-aggregation"],"name":"Visual Randomization","alt":"视觉随机化","abbr":"","aliases":["Texture Randomization","Lighting Randomization","Visual Domain Randomization"],"one_liner":"Randomly varying a simulation's textures, colors, lighting, and camera during training so a vision policy isn't picky about how things look.","explanation":"Visual randomization is the part of domain randomization aimed specifically at what the policy sees: object and background textures and colors, light position and intensity, and camera pose are all randomly varied in simulation, with distractor objects added on top, so the model sees enough visual diversity to treat the real world as just one more variation. Tobin and colleagues (OpenAI, Berkeley) did this systematically in 2017, training an object detector using only randomly rendered simulated images and reaching about 1.5 cm localization accuracy in real scenes, which they then used for grasping. It mainly narrows the visual portion of the sim-to-real gap; randomizing physical parameters like mass and friction instead is called dynamics randomization. Isaac Lab has built-in event terms for random textures and random colors that can run automatically at the start of every episode.","example":"When training a vision policy to grasp a cube, each episode might randomly swap the tabletop for wood grain, marble, or a solid color, randomly adjust the light's direction and brightness, and slightly shift the camera position.","related":["Domain Randomization","Dynamics Randomization","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Visual Generalization","Variant Aggregation (SimplerEnv)"]},{"id":"nvidia-omniverse-replicator","category":"sim","sec":6,"tier":3,"sources":[{"title":"Omniverse Extensions Docs: Replicator","url":"https://docs.omniverse.nvidia.com/extensions/latest/ext_replicator.html"},{"title":"Isaac Sim Docs: Synthetic Data Generation (Replicator tutorials)","url":"https://docs.isaacsim.omniverse.nvidia.com/latest/replicator_tutorials/index.html"},{"title":"GitHub: NVIDIA-AI-IOT/sdg_pallet_model","url":"https://github.com/NVIDIA-AI-IOT/sdg_pallet_model"}],"as_of":"2026-09","related_ids":["synthetic-data","domain-randomization","nvidia-isaac-sim","nvidia-omniverse","visual-randomization","photorealistic-rendering"],"name":"NVIDIA Omniverse Replicator","alt":"Omniverse Replicator 合成数据生成","abbr":"","aliases":["Replicator","omni.replicator"],"one_liner":"NVIDIA Omniverse's synthetic-data framework that randomizes simulated scenes and automatically outputs labeled training data.","explanation":"Replicator is the framework inside NVIDIA's Omniverse platform for building synthetic data generation (SDG) pipelines, integrated into Isaac Sim as the omni.replicator extension. Collecting and labeling real images is expensive, whereas in simulation the position, category, and depth of every object are already known exactly, so labels come for free and are perfectly accurate. It has three main pieces: randomizers, which sample assets, materials, lighting, and camera poses following the domain-randomization approach; annotators, which output ground truth such as 2D/3D bounding boxes, semantic and instance segmentation, and depth; and writers, which save the results in whatever format a given model needs. In robotics it is commonly used to train perception models such as object detectors and pose estimators, and the Isaac Sim documentation also gives examples for navigation and manipulation scenes.","example":"NVIDIA-AI-IOT's open-source pallet-detection model (sdg_pallet_model) was trained entirely on synthetic data generated with Replicator, and can be deployed to Jetson hardware with TensorRT.","related":["Synthetic Data","Domain Randomization","NVIDIA Isaac Sim","NVIDIA Omniverse","Visual Randomization","Photorealistic Rendering"]},{"id":"automatic-domain-randomization","category":"sim","sec":6,"tier":3,"sources":[{"title":"Solving Rubik's Cube with a Robot Hand (OpenAI, arXiv 1910.07113)","url":"https://arxiv.org/abs/1910.07113"}],"as_of":"2019-10","related_ids":["domain-randomization","dynamics-randomization","sim-to-real-transfer","curriculum-learning","dactyl","shadow-dexterous-hand"],"name":"Automatic Domain Randomization","alt":"自动域随机化","abbr":"ADR","aliases":["ADR"],"one_liner":"A domain-randomization method, proposed by OpenAI, that automatically widens its randomization range as the policy gets better.","explanation":"ADR is an algorithm OpenAI proposed in its 2019 work on solving a Rubik's Cube with a robot hand. Ordinary domain randomization (randomly varying parameters like friction, mass, and appearance in simulation so a policy transfers to a real robot) requires a person to manually set the randomization range for every parameter, which becomes hard to tune as the number of parameters grows. ADR turns this into an automatic curriculum: the initial distribution is concentrated entirely on a single environment set to calibrated real-robot values; during training, one parameter is randomly picked and pinned to the current edge of its range for evaluation, and its range is widened if performance is above an upper threshold or narrowed if it is below a lower threshold. Difficulty then grows in step with the policy's ability, and the final range can end up far wider than anything a person would set by hand. The paper used ADR to jointly train a control policy and a vision-based pose-estimation network, both trained entirely in simulation and then transferred to a real Shadow dexterous hand.","example":"The paper's best policy achieved about a 60% success rate on real-robot cube configurations requiring 15 moves to solve, and about 20% on the hardest configurations requiring 26 moves; the solution sequence itself came from the classical Kociemba solver, while the neural network handled the hand's manipulation.","related":["Domain Randomization","Dynamics Randomization","Sim-to-Real Transfer","Curriculum Learning","Dactyl","Shadow Dexterous Hand"]},{"id":"actuator-modeling","category":"sim","sec":6,"tier":2,"sources":[{"title":"Learning agile and dynamic motor skills for legged robots (Hwangbo et al., Science Robotics 2019)","url":"https://arxiv.org/abs/1901.08652"},{"title":"Isaac Lab Docs: Actuators","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/core-concepts/actuators.html"},{"title":"Isaac Lab API: isaaclab.actuators","url":"https://isaac-sim.github.io/IsaacLab/main/source/api/lab/isaaclab.actuators.html"}],"as_of":"2026-09","related_ids":["sim-to-real-transfer","sim-to-real-gap","system-identification","proportional-derivative-control","series-elastic-actuator","nvidia-isaac-lab"],"name":"Actuator Modeling (Actuator Network)","alt":"执行器建模","abbr":"","aliases":["Actuator Network","Actuator Net","implicit actuator","explicit actuator"],"one_liner":"Simulating how a real motor's torque output actually responds to a command, to narrow the sim-to-real gap.","explanation":"Actuator modeling means describing, inside a simulation, how a real motor and gearbox actually behave: given a target joint position, how much torque it can really deliver, and with how much delay and saturation. A simulator's default ideal PD control responds faster and more cleanly than a real motor does, and a policy trained against that ideal often jitters or lags once it reaches a real robot — a significant source of the sim-to-real gap. There are two general approaches: analytical motor models that add torque limits, speed saturation, or command delay; and actuator networks, which originated in ETH Zurich and Intel's 2019 ANYmal work, where a small neural network is trained on less than 4 minutes of real-robot data to predict torque from a short history of joint position error and velocity, and is then plugged into the simulator for policy training. Isaac Lab calls the case where the physics engine computes PD control internally an “implicit actuator,” and calls this kind of custom model an “explicit actuator.”","example":"Isaac Lab provides actuator classes such as DCMotor (with speed-torque saturation), DelayedPDActuator (with command delay), and the neural-network-based ActuatorNetMLP and ActuatorNetLSTM, chosen to match the real motor's characteristics when training a quadruped policy.","related":["Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","System Identification","Proportional-Derivative Control","Series Elastic Actuator (SEA)","NVIDIA Isaac Lab"]},{"id":"factory-industreal","category":"sim","sec":6,"tier":3,"sources":[{"title":"Factory: Fast Contact for Robotic Assembly (arXiv 2205.03532)","url":"https://arxiv.org/abs/2205.03532"},{"title":"IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality (arXiv 2305.17110)","url":"https://arxiv.org/abs/2305.17110"},{"title":"Isaac Lab Available Environments","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/environments.html"}],"as_of":"2026-09","related_ids":["contact-rich-manipulation","peg-in-hole-insertion","sim-to-real-transfer","nvidia-isaac-lab","physx","signed-distance-field-function"],"name":"Factory / IndustReal","alt":"Factory / IndustReal 接触丰富装配仿真","abbr":"","aliases":["Fast Contact for Robotic Assembly","IndustRealKit","IndustRealLib"],"one_liner":"NVIDIA's contact-rich assembly simulation and transfer pipeline: simulate tasks like nut-threading fast, then transfer the policy to a real robot.","explanation":"Factory is a contact-rich simulation method and learning toolkit NVIDIA published at RSS 2022: it uses a signed distance field (SDF, which records how far every point is from an object's surface) for collision detection, combined with contact-point reduction and a Gauss-Seidel solver, to simulate 1,000 nut-and-bolt assemblies in real time on a single GPU, and ships 60 part models, 3 assembly environments, and 7 controllers, integrated into both PhysX and Isaac Gym. IndustReal (RSS 2023) builds on this, training reinforcement-learning policies for peg and gear assembly with techniques such as an SDF-based reward, a sampling-based curriculum, and a policy-level action integrator, reaching 83–99% success over 600 real-robot trials, and open-sourcing a 3D-printable parts kit along with deployment code. Together they show that high-precision assembly can also follow the “train in simulation, transfer to the real robot” path.","example":"The Isaac-Factory-PegInsert-Direct-v0 (peg insertion), GearMesh (gear meshing), and NutThread (nut threading) environments in Isaac Lab all come from Factory and can be trained directly with PPO.","related":["Contact-rich Manipulation","Peg-in-Hole Insertion","Sim-to-Real Transfer","NVIDIA Isaac Lab","PhysX","Signed Distance Field / Function"]},{"id":"sim-to-sim-transfer","category":"sim","sec":6,"tier":2,"sources":[{"title":"unitreerobotics/unitree_rl_gym GitHub","url":"https://github.com/unitreerobotics/unitree_rl_gym"},{"title":"roboterax/humanoid-gym GitHub","url":"https://github.com/roboterax/humanoid-gym"},{"title":"Humanoid-Gym: Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer (arXiv 2404.05695)","url":"https://arxiv.org/abs/2404.05695"}],"as_of":"","related_ids":["sim-to-real-transfer","sim-to-real-gap","isaac-gym","mujoco","humanoid-gym","unitree-rl-gym"],"name":"Sim-to-Sim Transfer","alt":"仿真到仿真迁移","abbr":"Sim2Sim","aliases":["Sim2Sim","Cross-Simulator Validation"],"one_liner":"Running a policy trained in one simulator inside a different simulator, to expose problems before they reach the real robot.","explanation":"Sim2Sim is common in reinforcement-learning pipelines for legged and humanoid robots: a policy is trained at scale in a GPU-parallel simulator such as Isaac Gym or Isaac Lab, then dropped unchanged into a different simulator, such as MuJoCo, which has a different contact model and solver. If the policy becomes unstable once the simulator changes, that means it was exploiting quirks specific to the original simulator, and deploying it straight to the real robot would be risky. Unitree's unitree_rl_gym documents its pipeline as “train → replay → Sim2Sim → Sim2Real”; RobotEra's Humanoid-Gym, open-sourced in 2024, likewise provides an Isaac Gym to MuJoCo validation pipeline. It is cheap and carries no risk of damaging a real robot, but passing it still does not guarantee success on the real robot.","example":"After training a walking policy for the Unitree G1 in Isaac Gym with unitree_rl_gym, developers first run its Sim2Sim script to check the policy in MuJoCo before deploying it to the real robot.","related":["Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Isaac Gym","MuJoCo (Multi-Joint dynamics with Contact)","Humanoid-Gym","unitree_rl_gym (Unitree RL Gym)"]},{"id":"hardware-in-the-loop-software-in-the-loop-simulation","category":"sim","sec":6,"tier":3,"sources":[{"title":"MathWorks: What Is Hardware-in-the-Loop (HIL)?","url":"https://www.mathworks.com/discovery/hardware-in-the-loop-hil.html"},{"title":"PX4 User Guide: Simulation (SITL / HITL)","url":"https://docs.px4.io/main/en/simulation/"},{"title":"unitreerobotics/unitree_mujoco (GitHub)","url":"https://github.com/unitreerobotics/unitree_mujoco"}],"as_of":"","related_ids":["sim-to-real-transfer","sim-to-sim-transfer","digital-twin","real-time-control","px4-ardupilot","unitree-sdk2"],"name":"Hardware-in-the-Loop / Software-in-the-Loop Simulation","alt":"硬件在环 / 软件在环仿真","abbr":"HIL / SIL","aliases":["HIL / SIL","HITL / SITL","Hardware-in-the-Loop Testing"],"one_liner":"Connecting a real controller or real control code to a simulated plant for testing — a standard validation step before deploying to a real robot.","explanation":"These are two testing methods borrowed from automotive and aerospace embedded engineering. Software-in-the-loop (SIL): the control program runs directly on a computer and interfaces with a simulated robot and sensors. Hardware-in-the-loop (HIL): the control program runs on the actual controller board, connected through real I/O signals or a communication bus to a real-time simulated model, so the controller believes it's controlling the real machine. MathWorks places these alongside model-in-the-loop (MIL) and processor-in-the-loop (PIL) as a chain of tests that progressively approach the real hardware. The benefit is that dangerous conditions and edge cases can be tested repeatedly in simulation with no risk of damaging the real machine. A common robotics practice is to let the same control code connect to either simulation or the real robot by changing only the communication configuration, as with PX4 drones' SITL/HITL or Unitree's unitree_mujoco.","example":"unitree_mujoco uses the same DDS communication as the real robot: a control program written with Unitree's SDK can be debugged in simulation using the local loopback network interface and domain ID 1, and switching to the network interface connected to the real robot with domain ID 0 drives the real robot directly.","related":["Sim-to-Real Transfer","Sim-to-Sim Transfer","Digital Twin","Real-Time Control","PX4 / ArduPilot","Unitree SDK2"]},{"id":"digital-twin","category":"sim","sec":6,"tier":1,"sources":[{"title":"Wikipedia: Digital twin","url":"https://en.wikipedia.org/wiki/Digital_twin"},{"title":"NVIDIA Glossary: What Is a Digital Twin?","url":"https://www.nvidia.com/en-us/glossary/digital-twin/"},{"title":"Automated Creation of Digital Cousins for Robust Policy Learning (ACDC, CoRL 2024, arXiv 2410.07408)","url":"https://arxiv.org/abs/2410.07408"}],"as_of":"","related_ids":["real-to-sim","digital-cousin","sim-to-real-gap","nvidia-isaac-sim","nvidia-omniverse","simulation-fidelity"],"name":"Digital Twin","alt":"数字孪生","abbr":"","aliases":[],"one_liner":"A virtual copy of a real object, robot, or factory that's kept in sync using real data.","explanation":"A digital twin is a computer model that corresponds one-to-one to a physical product, system, or process. What sets it apart from an ordinary simulation is that it continuously receives data from the real object to keep its own state synchronized, and in the strict sense it can also feed back and influence the real system. NASA used ground-based simulators paired with the actual spacecraft for troubleshooting as far back as the Apollo era, and formally defined the concept in 2010; NVIDIA positions it as a flagship use case for Omniverse, applied to factory planning and robot testing. In embodied AI, the term is often used more loosely, to mean a simulated environment that faithfully reproduces a real scene, used to train and evaluate a policy in simulation before it ever touches a real robot. The closer the replica, the smaller the sim-to-real gap, but the higher the modeling cost; an approach that only aims for similarity rather than an exact match is called a “digital cousin.”","example":"According to NVIDIA, Foxconn used a digital twin to optimize a production line's layout and equipment placement in a virtual factory before rolling the changes out to the physical one.","related":["Real-to-Sim","Digital Cousin","Sim-to-Real Gap (Reality Gap)","NVIDIA Isaac Sim","NVIDIA Omniverse","Simulation Fidelity"]},{"id":"digital-cousin","category":"sim","sec":6,"tier":3,"sources":[{"title":"Automated Creation of Digital Cousins for Robust Policy Learning (arXiv 2410.07408)","url":"https://arxiv.org/abs/2410.07408"},{"title":"Digital Cousins 项目主页","url":"https://digital-cousins.github.io/"}],"as_of":"2024-10","related_ids":["digital-twin","real-to-sim-to-real","sim-to-real-transfer","domain-randomization","simulation-assets","omnigibson"],"name":"Digital Cousin","alt":"数字表亲","abbr":"","aliases":["ACDC","Automated Creation of Digital Cousins"],"one_liner":"A simulated scene that only needs to be geometrically and semantically similar to reality, not an exact one-to-one replica.","explanation":"Digital cousin was proposed by Fei-Fei Li and Jiajun Wu's group at Stanford in their CoRL 2024 paper “Automated Creation of Digital Cousins for Robust Policy Learning,” as a counterpart to the “digital twin” (a virtual copy that precisely replicates one real scene down to the last detail). A digital cousin doesn't replicate any specific real scene; it only needs to be geometrically and semantically similar — for a kitchen cabinet, for instance, it's enough for the shape, size, and opening mechanism to be close, built directly by picking similar objects from an existing asset library. This costs less than precise modeling, and generating several cousin scenes at once brings built-in diversity, making a trained policy less prone to overfitting to one specific scene. The paper's ACDC pipeline builds a simulated environment automatically from a single real RGB photo, through three steps: extracting object information, matching similar assets, and generating an interactive scene. In its reported zero-shot sim-to-real transfer, policies trained on digital cousins reached about 90% success versus about 25% for digital twins.","example":"Given one photo of a kitchen cabinet at home, ACDC automatically picks several cabinets from its asset library with similar shapes and opening mechanisms, generates several cousin scenes, trains a cabinet-opening policy in simulation, and deploys it directly to the real cabinet.","related":["Digital Twin","Real-to-Sim-to-Real","Sim-to-Real Transfer","Domain Randomization","Simulation Assets","OmniGibson"]},{"id":"gaussian-splatting-based-simulation","category":"sim","sec":6,"tier":2,"sources":[{"title":"SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting (arXiv 2409.10161)","url":"https://arxiv.org/abs/2409.10161"},{"title":"RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator (arXiv 2411.11839)","url":"https://arxiv.org/abs/2411.11839"},{"title":"3D Gaussian Splatting for Real-Time Radiance Field Rendering (arXiv 2308.04079)","url":"https://arxiv.org/abs/2308.04079"}],"as_of":"2026-01","related_ids":["3d-gaussian-splatting","real-to-sim-to-real","sim-to-real-gap","digital-twin","robogsim-a-real2sim2real-robotic-gaussian-splatting-simulato","real2render2real"],"name":"Gaussian Splatting-based Simulation","alt":"高斯泼溅仿真","abbr":"","aliases":["3DGS Simulation","GS Simulation","Gaussian Splatting Simulator"],"one_liner":"Rendering simulated scenes with 3D Gaussian Splatting reconstructions of real places, paired with a physics engine for motion.","explanation":"This approach plugs 3D Gaussian Splatting (3DGS) — introduced by Kerbl and colleagues in 2023, which represents a scene as a large number of colored 3D Gaussian blobs and can render photorealistic images in real time — into robot simulation. Typically, a phone or camera scans a real scene to reconstruct it as 3DGS, and the objects and robot inside it are then aligned with the collision bodies used by the physics engine: the physics engine computes motion, while 3DGS handles rendering the camera view. The goal is to close the visual sim-to-real gap caused by traditional simulators' obviously synthetic-looking graphics, making it easier for policies trained on simulated vision to transfer zero-shot to a real robot, and enabling reproducible closed-loop evaluation. Notable projects include SplatSim, RoboGSim, and Real2Render2Real; AgiBot's Genie Sim 3.0 also added 3DGS scene reconstruction. Gaussian points carry no physical properties of their own, so collision geometry and a dynamics model still have to be supplied separately.","example":"SplatSim replaced a simulator's usual mesh rendering with Gaussian-splatting rendering; an RGB manipulation policy trained only on simulated data then deployed zero-shot to a real robot at an average success rate of 86.25%, versus 97.5% for a policy trained on real-robot data.","related":["3D Gaussian Splatting","Real-to-Sim-to-Real","Sim-to-Real Gap (Reality Gap)","Digital Twin","RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator","Real2Render2Real"]},{"id":"robogsim-a-real2sim2real-robotic-gaussian-splatting-simulato","category":"sim","sec":6,"tier":3,"sources":[{"title":"arXiv 2411.11839 - RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator","url":"https://arxiv.org/abs/2411.11839"},{"title":"RoboGSim 项目主页","url":"https://robogsim.github.io/"}],"as_of":"2025-08","related_ids":["gaussian-splatting-based-simulation","real-to-sim-to-real","3d-gaussian-splatting","digital-twin","sim-to-real-transfer","nvidia-isaac-sim"],"name":"RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator","alt":"RoboGSim","abbr":"","aliases":[],"one_liner":"A robot simulator that reconstructs real scenes with 3D Gaussian Splatting and couples them to a physics engine.","explanation":"RoboGSim is a paper and system released in November 2024, with authors from the Harbin Institute of Technology (Shenzhen), the Institute of Computing Technology at the Chinese Academy of Sciences, Megvii, and Zhejiang University. It combines 3D Gaussian Splatting (a 3D representation that reconstructs a scene as a large number of colored ellipsoids and renders results close to photographs) with the Isaac Sim physics engine, and is made up of four parts: a Gaussian reconstructor, a digital-twin builder, a scene composer, and an interactive engine. It first reconstructs a real tabletop scene into simulation, then inside simulation swaps camera viewpoints, objects, trajectories, and backgrounds to synthesize new data, and finally deploys back to a real robot — that is, Real2Sim2Real. The paper reports that a policy trained purely on data synthesized this way can run zero-shot on a real robot, matching the performance of a policy trained on real-robot data, and doing better still under new viewpoints and new scenes. It can also serve as a closed-loop evaluation environment for testing VLA models.","example":"First film a lab bench with a camera, reconstruct the tabletop and robot arm inside RoboGSim, then swap in new camera viewpoints to batch-generate pick-and-place trajectories, and use that synthetic data to train a policy that deploys directly to the real arm.","related":["Gaussian Splatting-based Simulation","Real-to-Sim-to-Real","3D Gaussian Splatting","Digital Twin","Sim-to-Real Transfer","NVIDIA Isaac Sim"]},{"id":"discoverse","category":"sim","sec":6,"tier":3,"sources":[{"title":"DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments (arXiv 2507.21981)","url":"https://arxiv.org/abs/2507.21981"},{"title":"DISCOVERSE 项目主页","url":"https://air-discoverse.github.io/"},{"title":"DISCOVERSE GitHub","url":"https://github.com/TATP-233/DISCOVERSE"}],"as_of":"2025-07","related_ids":["gaussian-splatting-based-simulation","3d-gaussian-splatting","mujoco","real-to-sim-to-real","sim-to-real-gap","photorealistic-rendering"],"name":"DISCOVERSE","alt":"DISCOVERSE","abbr":"","aliases":["Efficient Robot Simulation in Complex High-Fidelity Environments"],"one_liner":"An open-source real-to-sim-to-real simulation framework combining 3D Gaussian Splatting rendering with MuJoCo physics.","explanation":"DISCOVERSE is an open-source robot simulation framework released jointly by Tsinghua University, Zhejiang University, and other institutions together with DISCOVER Robotics and D-Robotics (地瓜机器人); its paper was accepted at IROS 2025, and the code is MIT-licensed. It uses 3D Gaussian Splatting (3DGS, a method that reconstructs and quickly renders a real scene using a large number of colored 3D Gaussian points) for visuals and MuJoCo for physics: a real scene is first reconstructed with photorealistic appearance and then paired with collision and physics models, producing a simulated environment that “looks like the real world,” narrowing the visual sim-to-real gap. It supports parallel simulation of multiple cameras and sensors, is compatible with existing 3D assets, robot models, and ROS plugins, and has been adapted to embodiments including Airbot Play, AgileX PiPER, UR5e, Franka Panda, and the LEAP Hand. Its project page reports rendering 5 channels of 640×480 RGB-D cameras at up to 650 FPS (about 240 FPS on a laptop); the paper's imitation-learning experiments show better zero-shot sim-to-real transfer than other simulators.","example":"A lab tabletop is scanned and reconstructed as a 3DGS scene inside DISCOVERSE; demonstrations are collected in simulation to train a grasping policy, which is then deployed directly to the same real tabletop with no real-robot fine-tuning.","related":["Gaussian Splatting-based Simulation","3D Gaussian Splatting","MuJoCo (Multi-Joint dynamics with Contact)","Real-to-Sim-to-Real","Sim-to-Real Gap (Reality Gap)","Photorealistic Rendering"]},{"id":"gs-playground","category":"sim","sec":6,"tier":3,"sources":[{"title":"GS-Playground (arXiv 2604.25459)","url":"https://arxiv.org/abs/2604.25459"},{"title":"GS-Playground project page","url":"https://gsplayground.github.io"}],"as_of":"2026-09","related_ids":["gaussian-splatting-based-simulation","3d-gaussian-splatting","real-to-sim","digital-twin","batched-rendering","discoverse"],"name":"GS-Playground","alt":"GS-Playground","abbr":"","aliases":["A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning"],"one_liner":"A high-throughput simulator from Tsinghua and others that uses batched 3D Gaussian Splatting to render photorealistic images quickly.","explanation":"GS-Playground was proposed by Tsinghua University together with Motphys, Dexmal, and more than a dozen other organizations; its paper was released in April 2026, its project page shows acceptance at RSS 2026, and its code is open-sourced on GitHub. It addresses a tension in vision-based robot learning: traditional simulators don't look realistic, while photorealistic rendering is too slow to run at the scale needed for visual reinforcement learning. Its approach combines a self-developed parallel physics engine, MotrixSim (compatible with the MJCF format), with batched 3D Gaussian Splatting (representing a scene with a large number of colored ellipsoids, which renders quickly), reaching a total rendering throughput of about 10,000 frames per second across 2,048 parallel environments at 640×480 resolution, and it can automatically generate a simulatable digital twin from a single RGB photo. The paper validates the approach on quadruped and humanoid locomotion, visual navigation, and robot-arm grasping, and deploys it to real robots including the Unitree Go2 and G1.","example":"The project page states that generating a scene asset ready to drop into simulation from one RGB photo takes under 5 minutes, and that humanoid physics simulation runs 32 times faster than MuJoCo.","related":["Gaussian Splatting-based Simulation","3D Gaussian Splatting","Real-to-Sim","Digital Twin","Batched Rendering","DISCOVERSE"]},{"id":"success-rate","category":"sim","sec":7,"tier":1,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)","url":"https://arxiv.org/abs/2304.13705"},{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)","url":"https://arxiv.org/abs/2405.05941"}],"as_of":"","related_ids":["progress-score","average-length","episode","simulation-based-evaluation","real-world-evaluation","statistical-rigor-in-policy-evaluation"],"name":"Success Rate","alt":"成功率","abbr":"SR","aliases":["SR","Task Success Rate"],"one_liner":"The fraction of attempts at the same task that a policy completes successfully — the most common robotics metric.","explanation":"Success rate is the most commonly used evaluation metric in robot learning: a policy attempts the same task N times (each attempt called an episode, usually with a randomized initial object position), the number of successes is counted against a predefined criterion (such as “the block ends up in the target zone”), and that count is divided by N. It's intuitive and easy to compare across methods, but it carries limited information — it only looks at the final outcome, and doesn't distinguish a near-miss from a total failure — so it's often reported alongside subtask success rate, a progress score, or average completed length. Success rate is also very sensitive to the number of trials: one extra success out of 25 trials shifts the number by 4 percentage points, so rigorous papers report the number of trials and random seeds used, and ideally a confidence interval. Simulation can run hundreds or thousands of episodes at once; real-robot testing is usually limited to a few dozen.","example":"The ACT paper averages success rate over 3 random seeds with 50 trials each for simulated tasks; for real-robot tasks it runs 25 trials each, with “open a sealed bag” reaching 88% success.","related":["Progress Score","Average Length (CALVIN)","Episode","Simulation-Based Evaluation","Real-World Evaluation","Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)"]},{"id":"real-world-evaluation","category":"sim","sec":7,"tier":1,"sources":[{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"},{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"}],"as_of":"2025-06","related_ids":["simulation-based-evaluation","evaluation-protocol","success-rate","double-blind-pairwise-comparison","roboarena","sim-to-real-gap"],"name":"Real-World Evaluation","alt":"真机评测","abbr":"","aliases":["real-robot evaluation","real-robot testing"],"one_liner":"Running a policy on an actual robot repeatedly and recording success rate and other outcomes.","explanation":"Real-world evaluation means running a policy on an actual robot in an actual scene, recording whether each attempt succeeded, how far it got, and whether a human had to step in. No matter how good a simulation is, a sim-to-real gap remains, so this is the ultimate test of an embodied model. The difficulty is cost and reproducibility: a person has to place objects and reset the scene every time, and results are judged by hand; a slight difference in lighting, placement, or the robot's own state can shift the score, making it hard to compare numbers across different labs. For this reason, papers need to spell out their evaluation protocol clearly — the tasks, number of trials, initial conditions, and success criteria. To make results more trustworthy and more scalable, approaches such as RoboArena's distributed, double-blind pairwise comparisons have emerged, alongside using simulated evaluation (like SimplerEnv) or a world model to predict real-robot performance instead.","example":"RoboArena ran more than 600 double-blind, pairwise real-robot comparisons of 7 general-purpose policies across DROID robot setups at 7 universities, then aggregated the results into a ranking.","related":["Simulation-Based Evaluation","Evaluation Protocol","Success Rate","Double-blind Pairwise Comparison","RoboArena","Sim-to-Real Gap (Reality Gap)"]},{"id":"simulation-based-evaluation","category":"sim","sec":7,"tier":1,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)","url":"https://arxiv.org/abs/2405.05941"},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning","url":"https://arxiv.org/abs/2306.03310"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"2025-02","related_ids":["real-world-evaluation","benchmark","success-rate","simplerenv","libero-benchmark","sim-to-real-correlation"],"name":"Simulation-Based Evaluation","alt":"仿真评测","abbr":"","aliases":["Sim Evaluation"],"one_liner":"Running a policy through many tasks in a simulator and tallying success rate, instead of or alongside real-robot testing.","explanation":"Simulation-based evaluation means placing a trained robot policy in a simulator and running it through tasks under fixed initial conditions and judging rules, tallying success rate and other metrics. Real-robot evaluation requires a person to place objects and reset the scene by hand, so running dozens or hundreds of trials per policy is slow, and setups differ enough between labs that results are hard to reproduce; simulated evaluation, by contrast, can run automatically, at scale, and repeatably, and is commonly used to compare algorithms, run ablation studies, and pick checkpoints. Common benchmarks include LIBERO, CALVIN, SimplerEnv, and RoboTwin. The main concern is that a simulated score doesn't always reflect real-robot performance — visuals, physics, and the controller all differ from reality, and quite a few benchmarks are already close to saturated — so researchers check the correlation between simulated and real-robot scores, and use real-robot evaluation or world-model-based evaluation as a supplement.","example":"The OpenVLA-OFT paper runs simulation-based evaluation on LIBERO's four task suites, raising OpenVLA's average success rate from 76.5% to 97.1%.","related":["Real-World Evaluation","Benchmark","Success Rate","SimplerEnv","LIBERO Benchmark","Sim-to-Real Correlation"]},{"id":"closed-loop-evaluation","category":"sim","sec":7,"tier":2,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/html/2405.05941"},{"title":"Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? (arXiv 2312.03031)","url":"https://arxiv.org/abs/2312.03031"}],"as_of":"","related_ids":["open-loop-evaluation","simulation-based-evaluation","real-world-evaluation","rollout","compounding-error","simplerenv"],"name":"Closed-Loop Evaluation","alt":"闭环评测","abbr":"","aliases":["online evaluation"],"one_liner":"Letting a policy actually control the robot, act, observe, and act again, scored by whether the task finishes.","explanation":"Closed-loop evaluation means actually running a policy inside a simulated or real environment: at every step, it computes an action from the latest observation, the action changes the environment, and the environment returns a new observation, repeating until the task succeeds, fails, or times out, with success rate and similar metrics tallied at the end. This contrasts with open-loop evaluation, which only compares the model's predicted actions against human demonstrations on offline data (for example, using mean squared error) without the model's own actions ever affecting what it sees next. The problem is that imitation learning's small errors can drive the robot into states never seen in the demonstrations, and these errors compound — something an offline error metric can't reveal. The SimplerEnv paper found in practice that validation-set mean squared error doesn't predict a policy's real-robot performance well, and a CVPR 2024 study in autonomous driving similarly found that open-loop planning metrics can be misleading. VLA papers therefore generally treat closed-loop success rate, in simulation or on a real robot, as the metric that counts.","example":"Evaluating a VLA on LIBERO: for each task, run several episodes from different initial states with the policy controlling the simulated arm in real time, then report success rate — that's closed-loop evaluation; computing only its predicted actions' mean squared error against a held-out test set would be open-loop evaluation instead.","related":["Open-loop Evaluation","Simulation-Based Evaluation","Real-World Evaluation","Rollout","Compounding Error","SimplerEnv"]},{"id":"open-loop-evaluation","category":"sim","sec":7,"tier":2,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"},{"title":"Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? (arXiv 2312.03031)","url":"https://arxiv.org/abs/2312.03031"}],"as_of":"","related_ids":["closed-loop-evaluation","open-loop-control","simulation-based-evaluation","real-world-evaluation","simplerenv","sim-to-real-correlation"],"name":"Open-loop Evaluation","alt":"开环评测","abbr":"","aliases":["Offline Metric Evaluation","Offline Evaluation"],"one_liner":"Comparing a model's predicted actions against recorded demonstration actions on offline data, without letting the policy actually control anything.","explanation":"Open-loop evaluation does not let a policy actually drive a robot. Instead, recorded observations from an offline dataset are fed in frame by frame, and the model's predicted actions are compared against the demonstrated actions, typically using mean squared error (MSE) or L2 error. It is cheap, fast, reproducible, and needs neither a robot nor a simulator. The problem is that the policy's output never affects the next frame, so it cannot capture error accumulation or how well the policy recovers from a mistake; and because the same situation often has more than one correct way to act (action multimodality), a valid but different action still gets penalized simply for not matching the one recorded demonstration. The SIMPLER paper (2024) compared 6 checkpoints and found the Pearson correlation between validation-set MSE and real-robot success rate was only 0.308, versus 0.924 for closed-loop simulated evaluation. Self-driving research has reached a similar conclusion: a CVPR 2024 paper found that on nuScenes' open-loop planning metric, using nothing but the ego vehicle's own state, such as its speed, already scores competitively. As a result, open-loop metrics are mostly used for early debugging and screening.","example":"On a validation set from DROID or self-collected data, each frame's image and instruction is fed into a VLA model, and the mean squared error between its output action and the demonstrated action is computed as a rough check for whether training has gone off track.","related":["Closed-Loop Evaluation","Open-loop Control","Simulation-Based Evaluation","Real-World Evaluation","SimplerEnv","Sim-to-Real Correlation"]},{"id":"off-policy-evaluation","category":"sim","sec":7,"tier":3,"sources":[{"title":"Benchmarks for Deep Off-Policy Evaluation (DOPE, ICLR 2021)","url":"https://arxiv.org/abs/2103.16596"},{"title":"Off-Policy Evaluation via Off-Policy Classification (NeurIPS 2019)","url":"https://arxiv.org/abs/1906.01624"}],"as_of":"","related_ids":["offline-reinforcement-learning","off-policy","q-function","real-world-evaluation","world-model-based-policy-evaluation","importance-sampling"],"name":"Off-Policy Evaluation","alt":"离线策略评估","abbr":"OPE","aliases":["OPE","Off-Policy Policy Evaluation"],"one_liner":"Estimating how much return a new policy would actually get once deployed, using only data collected by other policies.","explanation":"Off-policy evaluation is a class of problem in reinforcement learning: data was collected by some behavior policy, and the goal is to estimate the expected return of a different target policy without letting it interact with the environment. Testing a policy on a real robot requires a human minder and wears down hardware, so if bad policies can be screened out using existing logs first, it saves a great deal of evaluation cost, and it also helps offline reinforcement learning pick checkpoints and hyperparameters. Common methods include importance sampling (reweighting old data by the ratio of the probability that each policy would pick the same action), fitted Q evaluation (FQE, which fits the target policy's Q-function from data), doubly robust estimation, and model-based approaches. Fu and colleagues' 2021 DOPE benchmark judges methods not just by value-estimation error but also by rank correlation and regret@k, since in practice getting the ranking right usually matters more than getting the exact number right.","example":"Google's Irpan and colleagues (NeurIPS 2019) reframed OPE as a classification problem for image-based robotic grasping, and using only offline data were able to reliably predict the relative performance of several policies on a real robot, including in sim-to-real transfer settings.","related":["Offline Reinforcement Learning","Off-Policy","Q-Function","Real-World Evaluation","World-Model-based Policy Evaluation","Importance Sampling"]},{"id":"benchmark","category":"sim","sec":7,"tier":1,"sources":[{"title":"Wikipedia: Benchmark (computing)","url":"https://en.wikipedia.org/wiki/Benchmark_(computing)"},{"title":"LIBERO (NeurIPS 2023 Datasets and Benchmarks Track)","url":"https://proceedings.neurips.cc/paper_files/paper/2023/hash/8c3c666820ea055a77726d66fc7d447f-Abstract-Datasets_and_Benchmarks.html"}],"as_of":"","related_ids":["baseline","evaluation-protocol","success-rate","libero-benchmark","benchmark-saturation","leaderboard-chasing"],"name":"Benchmark","alt":"基准测试","abbr":"","aliases":["leaderboard"],"one_liner":"A fixed set of tasks, data, and scoring rules that let different methods be compared under the same conditions.","explanation":"The term “benchmark” originates in computing, referring to a standardized set of tests used to measure the relative performance of something. In embodied AI, a benchmark typically bundles a fixed set of tasks, a simulator or real-world setup, demonstration data (if any), an evaluation protocol, and a metric — usually success rate. Common simulated benchmarks include LIBERO, CALVIN, SimplerEnv, and RoboTwin, alongside real-robot evaluation networks such as RoboArena. Its value is reproducibility and the ability to compare methods head to head; the risk is that methods can be tuned specifically to score well on the leaderboard (“benchmark hacking”) without that reflecting real-world usefulness, and once leading methods get close to a perfect score, the benchmark stops being able to distinguish between them — what's called benchmark saturation.","example":"VLA papers commonly report average success rate across LIBERO's four task suites, placing their numbers in the same table as baselines such as OpenVLA for comparison.","related":["Baseline","Evaluation Protocol","Success Rate","LIBERO Benchmark","Benchmark Saturation","Leaderboard Chasing"]},{"id":"evaluation-protocol","category":"sim","sec":7,"tier":3,"sources":[{"title":"Robot Learning as an Empirical Science: Best Practices for Policy Evaluation (arXiv 2409.09491)","url":"https://arxiv.org/abs/2409.09491"},{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"}],"as_of":"","related_ids":["benchmark","real-world-evaluation","simulation-based-evaluation","success-rate","statistical-rigor-in-policy-evaluation","double-blind-pairwise-comparison"],"name":"Evaluation Protocol","alt":"评测协议","abbr":"","aliases":["Testing Protocol"],"one_liner":"The set of rules specifying under what conditions a policy is tested, how many trials, and what counts as success.","explanation":"An evaluation protocol is a set of rules that spells out exactly how a policy is tested: which scenes and objects it's tested on, how the initial state is arranged, how many trials each condition gets, the maximum duration of a trial, what counts as success, whether human intervention is allowed mid-trial, and what metrics summarize the results. Robot evaluation results are very sensitive to these details — the same policy can get a very different success rate depending on object placement or how loosely success is defined — and if a paper doesn't spell out its protocol, readers can't tell whether the results are reproducible or comparable across methods. A 2024 paper by Kress-Gazit and colleagues, “Robot Learning as an Empirical Science,” recommends clearly reporting experimental conditions and success criteria, supplementing success rate with other metrics, doing statistical analysis, and qualitatively describing failure modes. Simulation benchmarks (such as LIBERO and CALVIN) usually come with a fixed protocol built in; real-robot evaluation more often relies on practices like a fixed table of initial positions, alternating between methods during testing, and double-blind pairwise comparison to keep things fair.","example":"A protocol might read: test each task 20 times, drawing the object's initial position in turn from 20 pre-marked points, cap each trial at 60 seconds, count success only when the object is fully inside the box and the gripper has released, and allow no manual resets mid-trial.","related":["Benchmark","Real-World Evaluation","Simulation-Based Evaluation","Success Rate","Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)","Double-blind Pairwise Comparison"]},{"id":"baseline","category":"sim","sec":7,"tier":1,"sources":[{"title":"Google Machine Learning Glossary: baseline","url":"https://developers.google.com/machine-learning/glossary#baseline"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"","related_ids":["ablation-study","benchmark","state-of-the-art","success-rate","evaluation-protocol"],"name":"Baseline","alt":"基线方法","abbr":"","aliases":[],"one_liner":"An existing or simple method used as a point of comparison to show how much a new method improves.","explanation":"A baseline is the reference point used when evaluating a new method. Google's machine learning glossary defines it as a reference model used to compare against a (usually more complex) model. A baseline can be something simple, like plain behavior cloning or a random policy, or it can be a strong, widely recognized method, like OpenVLA. Reporting “80% success rate” on its own says little; the improvement only means something once it's measured against a baseline under the same benchmark and evaluation protocol. When reading a paper, check whether the baseline is strong enough and whether it was reproduced faithfully to the original authors' settings — a weak baseline makes an improvement look bigger than it is. An ablation study can also be seen as comparing a method against a version of itself with one component removed, which then serves as the baseline.","example":"The OpenVLA-OFT paper uses the original OpenVLA as its baseline, raising average success rate on the four LIBERO task suites from 76.5% to 97.1%.","related":["Ablation Study","Benchmark","State of the Art (SOTA)","Success Rate","Evaluation Protocol"]},{"id":"state-of-the-art","category":"sim","sec":7,"tier":1,"sources":[{"title":"Wikipedia: State of the art","url":"https://en.wikipedia.org/wiki/State_of_the_art"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)","url":"https://arxiv.org/abs/2502.19645"},{"title":"paperswithcode-data issue #116: paperswithcode.com now redirects to huggingface (2025-08)","url":"https://github.com/paperswithcode/paperswithcode-data/issues/116"}],"as_of":"2026-09","related_ids":["benchmark","benchmark-saturation","leaderboard-chasing","success-rate","ablation-study"],"name":"State of the Art (SOTA)","alt":"最先进水平","abbr":"SOTA","aliases":["SOTA","SotA","state-of-the-art"],"one_liner":"The best publicly reported result on a given task or benchmark, commonly abbreviated SOTA.","explanation":"“State of the art” originally refers to the highest level of technical achievement a field has reached at a given time; the English phrase already appears in engineering literature from the early 20th century. In machine learning papers, SOTA refers specifically to the best current result on a public benchmark under the same evaluation protocol — for example, “achieves SOTA on LIBERO.” It's a convenient yardstick for comparing methods, but three caveats apply: it only holds for a specific benchmark and setting, and may not carry over to different data or a different robot; embodied AI's simulated benchmarks saturate easily, so a lead of a fraction of a percentage point can be within statistical noise; and real-robot results are hard to compare across labs because of differences in setup and objects. The website Papers with Code used to aggregate leaderboards for many tasks; that domain now redirects to Hugging Face's papers page.","example":"In 2025, OpenVLA-OFT reached 97.1% average success rate across LIBERO's four task suites, which the paper describes as a new SOTA on that benchmark.","related":["Benchmark","Benchmark Saturation","Leaderboard Chasing","Success Rate","Ablation Study"]},{"id":"ablation-study","category":"sim","sec":7,"tier":2,"sources":[{"title":"Wikipedia: Ablation (artificial intelligence)","url":"https://en.wikipedia.org/wiki/Ablation_(artificial_intelligence)"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["baseline","state-of-the-art","success-rate","action-chunking","temporal-ensembling"],"name":"Ablation Study","alt":"消融实验","abbr":"","aliases":["ablation"],"one_liner":"Removing or swapping one component of a method to see how performance changes, testing whether it matters.","explanation":"An ablation study starts from a complete method and removes or replaces one component at a time — a module, a loss term, an input modality, a data source, or a training trick — then retrains and re-evaluates under the same settings and compares the change in performance, to judge how much each design choice actually contributes. The term is borrowed from biology, where removing part of a tissue is used to study its function; according to Wikipedia, the AI pioneer Allen Newell used it in a 1974 speech-recognition tutorial. It's a paper's main evidence for explaining why it was designed a certain way: if removing a module barely changes performance, that suggests it isn't essential. Other variables need to be held fixed during an ablation, and enough trials need to be run for the comparison, or the difference observed may just be noise.","example":"The ACT paper's ablations separately remove action chunking, temporal ensembling, and CVAE training, finding that when training on human demonstration data, removing the CVAE clearly hurts success rate.","related":["Baseline","State of the Art (SOTA)","Success Rate","Action Chunking","Temporal Ensembling"]},{"id":"statistical-rigor-in-policy-evaluation","category":"sim","sec":7,"tier":3,"sources":[{"title":"Deep Reinforcement Learning at the Edge of the Statistical Precipice (arXiv 2108.13264, NeurIPS 2021)","url":"https://arxiv.org/abs/2108.13264"},{"title":"Robot Learning as an Empirical Science: Best Practices for Policy Evaluation (arXiv 2409.09491)","url":"https://arxiv.org/abs/2409.09491"},{"title":"A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation (TRI, arXiv 2507.05331)","url":"https://arxiv.org/html/2507.05331v1"}],"as_of":"2025-07","related_ids":["success-rate","evaluation-protocol","random-seed-and-reproducibility","double-blind-pairwise-comparison","real-world-evaluation","rollout"],"name":"Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)","alt":"评测统计显著性（置信区间 / 序贯检验 / 多随机种子）","abbr":"","aliases":["Statistically Rigorous Evaluation"],"one_liner":"Using statistics to tell whether a gap in success rate between two policies is real or just noise.","explanation":"Robot-policy evaluation often runs only a few dozen trials, so the success rate itself carries a wide margin of error: 14 successes out of 20 trials gives roughly a 48%–85% 95% confidence interval by the commonly used Wilson method. Statistically rigorous evaluation means reporting a confidence interval or a Bayesian posterior instead of a single percentage; reinforcement-learning results also need to be repeated across multiple random seeds, since the same algorithm can produce very different outcomes with a different seed. Agarwal and colleagues, at NeurIPS 2021, recommended reporting interval estimates and summarizing multi-task results with interquartile means, and open-sourced the rliable library. Sequential testing runs trials while checking as it goes, stopping early once the gap is already clear enough, saving expensive real-robot trials. The Toyota Research Institute also wrote in 2024 arguing that robot learning should be evaluated to the standards of experimental science: state the experimental conditions clearly, use multiple metrics together, and run proper statistical analysis.","example":"In Toyota Research Institute's 2025 large behavior model paper, each real-robot task and policy pair was run 50 times, and each simulated task 200 times, with the evaluator not told which policy was being tested; results were shown as Bayesian posteriors under a Beta prior, pairwise policy comparisons used sequential hypothesis testing, and Bonferroni correction controlled for multiple comparisons.","related":["Success Rate","Evaluation Protocol","Random Seed and Reproducibility","Double-blind Pairwise Comparison","Real-World Evaluation","Rollout"]},{"id":"generalization-robustness-evaluation","category":"sim","sec":7,"tier":2,"sources":[{"title":"THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (arXiv 2402.08191)","url":"https://arxiv.org/abs/2402.08191"},{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models (arXiv 2510.13626)","url":"https://arxiv.org/abs/2510.13626"}],"as_of":"2025-12","related_ids":["generalization","robustness","out-of-distribution","distractor-objects","libero-plus","the-colosseum-a-benchmark-for-evaluating-generalization-for"],"name":"Generalization / Robustness Evaluation","alt":"泛化与鲁棒性评测","abbr":"","aliases":["Perturbation Test","OOD Evaluation","Out-of-Distribution Evaluation","Robustness Benchmark"],"one_liner":"Deliberately changing lighting, positions, objects, or instructions to test how much skill a policy retains outside its training conditions.","explanation":"Rather than only reporting a policy's success rate under conditions that match its training distribution, this kind of evaluation systematically applies perturbations — swapping an object's color or shape, adding distractor objects, changing lighting and background, moving the camera viewpoint, changing the robot's starting pose, rewording the language instruction, adding sensor noise — and measures how far the success rate drops. The question it answers is whether a high score reflects real task competence or just memorization of the training scenes. A common approach is to vary one dimension of perturbation at a time to locate specific weaknesses, though multiple perturbations are also stacked together. Notable benchmarks include The Colosseum, LIBERO-Plus, and LIBERO-PRO; the variant aggregations in SimplerEnv follow the same idea. This kind of evaluation routinely shows that models scoring near-perfectly on standard benchmarks lose a large chunk of their success rate under even mild perturbation.","example":"LIBERO-Plus perturbs LIBERO along seven dimensions — object layout, camera viewpoint, robot initial state, language instructions, lighting, background texture, and sensor noise — and finds that some VLA models' success rates fall from 95% to below 30%, often because the model simply ignores the language instruction.","related":["Generalization","Robustness","Out-of-Distribution","Distractor Objects","LIBERO-Plus","The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation"]},{"id":"progress-score","category":"sim","sec":7,"tier":2,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"HomeRobot: Open-Vocabulary Mobile Manipulation (arXiv 2306.11565)","url":"https://arxiv.org/abs/2306.11565"},{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"}],"as_of":"","related_ids":["success-rate","evaluation-protocol","long-horizon-task","real-world-evaluation","roboarena","average-length"],"name":"Progress Score","alt":"进度分数","abbr":"","aliases":["Task Progress","Partial Success","Partial Credit"],"one_liner":"A metric that scores how many steps of a task a robot completed, rather than just recording binary success or failure.","explanation":"Progress score is a class of evaluation metric: a task is broken into stages or sub-goals ahead of time, and a robot earns credit for each stage it completes, usually normalized to a 0–1 or 0–100 scale. It complements binary success rate. On long-horizon, difficult tasks, every policy's success rate can sit near 0, which makes it impossible to tell them apart; scoring by progress reveals which policy got further and where exactly it got stuck. The downside is that the scoring rubric is defined by whichever researchers wrote it, so scores are not directly comparable across different papers. Physical Intelligence's π0 paper designed a separate scoring rubric for each task; HomeRobot OVMM uses partial success, awarding 1 point per completed stage, as one basis for its ranking; and RoboArena has evaluators assign a 0–100 progress score to each trial.","example":"In the π0 paper's laundry-folding task, taking a garment out of the basket, laying it flat, folding it, and stacking it neatly each earn 1 point; completing all of them is a perfect score, while only laying it flat earns half credit.","related":["Success Rate","Evaluation Protocol","Long-horizon Task","Real-World Evaluation","RoboArena","Average Length (CALVIN)"]},{"id":"intervention-rate","category":"sim","sec":7,"tier":2,"sources":[{"title":"HG-DAgger: Interactive Imitation Learning with Human Experts (arXiv 1810.02890)","url":"https://arxiv.org/abs/1810.02890"},{"title":"Sirius: Robot Learning on the Job (project page)","url":"https://ut-austin-rpl.github.io/sirius/"}],"as_of":"","related_ids":["mean-time-between-interventions","human-in-the-loop","human-gated-dagger","human-intervention-data","remote-teleoperation-takeover","levels-of-autonomy"],"name":"Intervention Rate","alt":"干预率","abbr":"","aliases":["Takeover Rate","Disengagement Rate"],"one_liner":"How often a human has to take over or correct a policy while it runs autonomously; lower means more independent operation.","explanation":"Intervention rate measures how autonomous a system really is: the fraction of the time a human has to take over or correct a robot or self-driving system while it operates, usually counted per episode, per hour, or per kilometer, and sometimes reported instead as the average time between interventions (mean time between interventions). The concept comes from the “takeover” or “disengagement” metric used in self-driving cars, and it matters just as much for deployed robots: success rate alone only tells you the outcome, while intervention rate reveals how much human labor was needed to get there, and directly determines how many robots one person can supervise at once. Human-in-the-loop methods such as HG-DAgger and Sirius have a person take over when the robot makes a mistake or a situation gets risky, and feed those intervention episodes back into training as correction data, so the intervention rate should fall as more deployment rounds accumulate. Because different teams may define what counts as one intervention differently, reports need to state that definition, or the numbers are not comparable across teams.","example":"Sirius lets a human remotely take over whenever the robot makes a mistake and adds that intervention data back into training; as more rounds of deployment accumulate, the amount of human intervention needed drops sharply, and by the third round the robot operates autonomously most of the time.","related":["Mean Time Between Interventions","Human-in-the-Loop","Human-Gated DAgger","Human Intervention Data","Remote Teleoperation Takeover","Levels of Autonomy"]},{"id":"mean-time-between-interventions","category":"sim","sec":7,"tier":3,"sources":[{"title":"Habilis-β: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model (arXiv 2602.18813)","url":"https://arxiv.org/abs/2602.18813"},{"title":"Mean time between failures - Wikipedia","url":"https://en.wikipedia.org/wiki/Mean_time_between_failures"}],"as_of":"2026-02","related_ids":["intervention-rate","mean-time-between-failures","success-rate","real-world-evaluation","fully-autonomous","remote-teleoperation-takeover"],"name":"Mean Time Between Interventions","alt":"平均干预间隔","abbr":"MTBI","aliases":["MTBI"],"one_liner":"How long, on average, a robot can keep working continuously before a human has to step in.","explanation":"MTBI measures a robot's reliability over long continuous operation: total operating time divided by the number of human interventions during that time, with a larger number being better. An intervention generally includes a manual emergency stop, a manual reset, a human repositioning an object for the robot, or a human taking over after a task times out. The idea is borrowed from mean time between failures (MTBF) in industrial reliability. The single-trial success rate commonly reported in academic papers is measured after a careful manual reset each time, which hides problems like state gradually drifting or occasional getting stuck during long runs, and also blends speed and accuracy into one number. Deployment-focused teams have therefore started reporting MTBI under a continuous-operation protocol with no resets, measuring it alongside tasks completed per hour to capture both efficiency and reliability.","example":"The Habilis-β paper runs a 1-hour continuous-operation evaluation: in a real humanoid-robot logistics workflow, its MTBI is 137.4 seconds, compared with 46.1 seconds for π0.5.","related":["Intervention Rate","Mean Time Between Failures","Success Rate","Real-World Evaluation","Fully Autonomous","Remote Teleoperation Takeover"]},{"id":"sim-to-real-correlation","category":"sim","sec":7,"tier":2,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"},{"title":"SIMPLER 项目页","url":"https://simpler-env.github.io/"}],"as_of":"2024-05","related_ids":["simulation-based-evaluation","real-world-evaluation","mean-maximum-rank-violation","simplerenv","visual-matching","sim-to-real-gap"],"name":"Sim-to-Real Correlation","alt":"仿真-真机相关性","abbr":"","aliases":["Pearson r","Sim-Real Alignment"],"one_liner":"A measure of whether simulated evaluation scores actually reflect real-robot performance, usually reported as a Pearson correlation coefficient r.","explanation":"Before using simulation for evaluation, it needs to be confirmed that a policy which scores well in simulation is also strong on the real robot. A common approach picks a set of policies, measures their success rate in both simulation and on the real robot, and computes the Pearson correlation coefficient r between the two sets of numbers (ranging from -1 to 1, with values closer to 1 meaning better agreement). The 2024 SIMPLER (SimplerEnv) paper ran this comparison for policies such as RT-1, RT-1-X, and Octo on the Google Robot and WidowX platforms, and pointed out that r only captures linear fit and can miss cases where the ranking itself gets flipped, so it also proposed Maximum Mean Rank Violation (MMRV, where lower is better). When correlation is high, a simulation benchmark can substitute for expensive real-robot testing when picking a model; when it is low, gains seen in simulation may just reflect overfitting to the simulator.","example":"SIMPLER narrows the gap through visual matching (compositing a real background onto a green screen, aligning textures) and system identification of control parameters, demonstrating a strong correlation between simulated and real-robot scores across about 1,500 evaluation episodes.","related":["Simulation-Based Evaluation","Real-World Evaluation","Mean Maximum Rank Violation","SimplerEnv","Visual Matching (SimplerEnv)","Sim-to-Real Gap (Reality Gap)"]},{"id":"mean-maximum-rank-violation","category":"sim","sec":7,"tier":3,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"}],"as_of":"2024-05","related_ids":["simplerenv","sim-to-real-correlation","visual-matching","variant-aggregation","simulation-based-evaluation","real-world-evaluation"],"name":"Mean Maximum Rank Violation","alt":"平均最大排名违背","abbr":"MMRV","aliases":["MMRV"],"one_liner":"A metric for whether a simulated evaluation's ranking of policies matches the real-robot ranking; lower is better.","explanation":"MMRV was proposed in the 2024 SIMPLER (SimplerEnv) paper. The authors argue that a simulated evaluation doesn't need to reproduce the absolute value of real-robot success rate — what matters is getting the relative ranking of policies right. It's computed as follows: for any pair of policies, if their ordering in simulation is reversed compared to the real robot, that counts as one “rank violation,” with a magnitude equal to the difference between their real-robot success rates; each policy takes its single worst violation, and these are averaged across all policies, giving a value between 0 and 1. This way, two policies that were already close in real-robot performance getting swapped only counts as a small error, while swapping two policies with a large real-robot gap counts as a large error. It's usually reported alongside the Pearson correlation coefficient, which only captures linear relationships and can be thrown off by real-robot evaluation noise when policies are close in skill.","example":"The SIMPLER paper ranks 6 Google Robot policies: ranking by validation-set action MSE gives an average MMRV of 0.375, while ranking with SIMPLER's “visual matching” simulated evaluation brings it down to 0.056.","related":["SimplerEnv","Sim-to-Real Correlation","Visual Matching (SimplerEnv)","Variant Aggregation (SimplerEnv)","Simulation-Based Evaluation","Real-World Evaluation"]},{"id":"elo-rating","category":"sim","sec":7,"tier":2,"sources":[{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"},{"title":"Elo rating system - Wikipedia","url":"https://en.wikipedia.org/wiki/Elo_rating_system"}],"as_of":"2025-11","related_ids":["double-blind-pairwise-comparison","roboarena","real-world-evaluation","success-rate","benchmark"],"name":"Elo Rating","alt":"Elo 评分","abbr":"","aliases":["Bradley-Terry Model","BT Model","Elo Rating System"],"one_liner":"A scoring method that estimates each player's or policy's relative strength from a series of head-to-head wins and losses.","explanation":"Elo rating was designed by physics professor and chess master Arpad Elo to rank chess players, and the US Chess Federation adopted it in 1960. Each player gets a single number; the gap between two ratings predicts the expected win probability, and after each game a rating is nudged by (actual result minus expected result) times a constant K. Mathematically, Elo is a special case of the Bradley-Terry model, which writes the probability that A beats B as a sigmoid function of the difference in their abilities. Embodied AI borrows this to solve a real problem: success rates reported by different labs on different tasks cannot be compared directly. Instead, two policies are compared head-to-head on the same task, and aggregating many such comparisons produces a leaderboard. RoboArena found that plain Elo rankings get distorted when tasks vary widely in difficulty, so it uses a Bradley-Terry variant that adds a task-difficulty parameter.","example":"RoboArena has evaluators at different sites blindly compare two policies on tasks of their own choosing and judge which one did better, then aggregates over 600 real-robot comparisons into a leaderboard of seven general-purpose policies using a modified Bradley-Terry model.","related":["Double-blind Pairwise Comparison","RoboArena","Real-World Evaluation","Success Rate","Benchmark"]},{"id":"double-blind-pairwise-comparison","category":"sim","sec":7,"tier":3,"sources":[{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"}],"as_of":"2025-06","related_ids":["elo-rating","roboarena","real-world-evaluation","evaluation-protocol","droid","benchmark"],"name":"Double-blind Pairwise Comparison","alt":"双盲成对比较","abbr":"","aliases":["A/B Evaluation","Blind A/B Evaluation","Pairwise Preference Evaluation"],"one_liner":"An evaluation method where a judge, without knowing which model is which, decides only which of two policies performed better.","explanation":"Double-blind pairwise comparison is an evaluation method: two models or policies each run once under the same conditions, the evaluator doesn't know beforehand which is which, and simply judges which one did better; a large number of these “who won” records are then aggregated into a ranking. Chatbot Arena, in the large-language-model world, ranks models this way through anonymous head-to-head comparisons. A leading example in embodied AI is RoboArena (2025): 7 universities ran over 600 real-robot pairwise evaluations on the DROID platform to compare 7 generalist policies; evaluators could choose their own tasks and scenes, but always compared two policies blind. This avoids having to standardize scenes and success criteria in advance, and also reduces evaluators favoring their own institution's model; the win/loss records from many evaluation sites are then converted into scores using an Elo or Bradley-Terry model (a statistical model that estimates each competitor's ability score from win/loss outcomes). The paper argues that this kind of distributed evaluation produces more accurate and more scalable rankings than centralized evaluation.","example":"An evaluator gives the instruction “fold the towel” in their own lab, the system sends two anonymous policies, A and B, to perform it in turn, and the evaluator judges B to be better; that preference is recorded toward the leaderboard.","related":["Elo Rating","RoboArena","Real-World Evaluation","Evaluation Protocol","DROID (Distributed Robot Interaction Dataset)","Benchmark"]},{"id":"benchmark-saturation","category":"sim","sec":7,"tier":3,"sources":[{"title":"LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (arXiv 2510.03827)","url":"https://arxiv.org/abs/2510.03827"},{"title":"Dynabench: Rethinking Benchmarking in NLP (arXiv 2104.14337)","url":"https://arxiv.org/abs/2104.14337"},{"title":"Stanford HAI: The 2025 AI Index Report","url":"https://hai.stanford.edu/ai-index/2025-ai-index-report"}],"as_of":"2025-10","related_ids":["benchmark","libero-benchmark","libero-pro","libero-plus","leaderboard-chasing","generalization-robustness-evaluation"],"name":"Benchmark Saturation","alt":"基准饱和","abbr":"","aliases":["Leaderboard Saturation"],"one_liner":"When leading models' scores on a benchmark all cluster near the maximum, so it can no longer distinguish good methods from bad ones.","explanation":"A benchmark — a shared test with fixed tasks and an evaluation protocol — tends to saturate the longer it's used: scores from different groups converge toward the ceiling, with gaps shrinking to a percentage point or two, sometimes within the range of random noise. This can happen because methods genuinely improved, but it can equally happen because everyone has repeatedly tuned against the same test set, or because the test scenes are too similar to the training data, letting a model score well through memorization. The Dynabench paper (2021) in NLP made exactly this point: models quickly achieve excellent benchmark scores yet fail on simple adversarial examples. A typical case in embodied AI is LIBERO, where multiple VLA models now score above 90% success under the standard setup. Once a benchmark saturates, the community usually releases a harder or perturbed successor, or shifts to real-robot evaluation and generalization/robustness evaluation. A small lead on a saturated benchmark carries limited weight.","example":"LIBERO-PRO (2025) shows that a model scoring above 90% success on standard LIBERO drops to 0.0% success once objects are swapped, initial states changed, instructions reworded, or the environment changed — indicating that the high score largely reflected memorization of training trajectories and scene layouts.","related":["Benchmark","LIBERO Benchmark","LIBERO-PRO","LIBERO-Plus","Leaderboard Chasing","Generalization / Robustness Evaluation"]},{"id":"libero-benchmark","category":"sim","sec":8,"tier":1,"sources":[{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning (arXiv 2306.03310)","url":"https://arxiv.org/abs/2306.03310"},{"title":"LIBERO GitHub repository","url":"https://github.com/Lifelong-Robot-Learning/LIBERO"},{"title":"OpenVLA-OFT (arXiv 2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"2026-09","related_ids":["benchmark","benchmark-saturation","libero-plus","libero-pro","robosuite","openvla-oft"],"name":"LIBERO Benchmark","alt":"LIBERO","abbr":"","aliases":["Benchmarking Knowledge Transfer for Lifelong Robot Learning","LIBERO-Spatial","LIBERO-Object","LIBERO-Goal","LIBERO-Long (LIBERO-10)","LIBERO-90","LIBERO-100"],"one_liner":"A simulated benchmark of 130 tabletop manipulation tasks, one of the most commonly reported in VLA papers.","explanation":"LIBERO was proposed by Bo Liu, Yuke Zhu, Peter Stone, and colleagues, published at the NeurIPS 2023 Datasets and Benchmarks track, and originally designed to study knowledge transfer in lifelong learning. It's built on robosuite (which runs on MuJoCo underneath), and procedurally generates 130 tasks with language instructions, split into four suites: Spatial (varying object placement), Object (varying which object), and Goal (varying the goal), each with 10 tasks; LIBERO-100 splits further into LIBERO-90, used for pretraining, and 10 long-horizon tasks called LIBERO-Long (also known as LIBERO-10). Each task comes with 50 human-teleoperated demonstrations. It later became the most commonly used leaderboard for VLA fine-tuning, and leading methods now exceed 97% average success rate, which has led to harder variants with added perturbations, LIBERO-Plus and LIBERO-PRO.","example":"OpenVLA-OFT reaches 97.1% average success rate across LIBERO's four suites, compared with 76.5% for the original OpenVLA.","related":["Benchmark","Benchmark Saturation","LIBERO-Plus","LIBERO-PRO","robosuite","OpenVLA-OFT"]},{"id":"libero-plus","category":"sim","sec":8,"tier":3,"sources":[{"title":"LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models (arXiv 2510.13626)","url":"https://arxiv.org/abs/2510.13626"},{"title":"LIBERO-plus GitHub 仓库","url":"https://github.com/sylvestf/LIBERO-plus"}],"as_of":"2025-10","related_ids":["libero-benchmark","libero-pro","generalization-robustness-evaluation","robustness","vision-language-action-model","openvla-oft"],"name":"LIBERO-Plus","alt":"LIBERO-Plus","abbr":"","aliases":["In-depth Robustness Analysis of Vision-Language-Action Models"],"one_liner":"A benchmark that adds 7 kinds of perturbations to LIBERO specifically to test the robustness of VLA models.","explanation":"LIBERO-Plus is a manipulation evaluation benchmark released in October 2025 by teams from Fudan University, the National University of Singapore, Tongji University, and others. It adds 7 categories of controlled perturbation to the popular LIBERO simulation benchmark — object placement, camera viewpoint, robot initial pose, language instructions, lighting, background texture, and sensor noise — for a total of 10,030 task variants, used to test whether a VLA (vision-language-action) model's high score reflects real skill or just memorization of the training scenes. The results show that a small change in viewpoint or initial pose can drop success rate from 95% to below 30%; models turn out to be relatively insensitive to language perturbation, often because they don't really look at the instruction at all. The repository is a drop-in replacement for the original LIBERO, and perturbed training data has also been open-sourced.","example":"OpenVLA scores 76.5% success on the original LIBERO but falls to 1.1% once the camera viewpoint is changed; replacing OpenVLA-OFT's instructions with blank text barely moves its success rate on the object task group.","related":["LIBERO Benchmark","LIBERO-PRO","Generalization / Robustness Evaluation","Robustness","Vision-Language-Action Model","OpenVLA-OFT"]},{"id":"libero-pro","category":"sim","sec":8,"tier":3,"sources":[{"title":"LIBERO-PRO: Towards Robust and Fair Evaluation of VLA Models Beyond Memorization (arXiv 2510.03827)","url":"https://arxiv.org/abs/2510.03827"},{"title":"LIBERO-PRO GitHub 仓库","url":"https://github.com/Zxy-MLlab/LIBERO-PRO"}],"as_of":"2025-10","related_ids":["libero-benchmark","libero-plus","benchmark-saturation","generalization-robustness-evaluation","out-of-distribution","leaderboard-chasing"],"name":"LIBERO-PRO","alt":"LIBERO-PRO","abbr":"","aliases":["Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization"],"one_liner":"Adds object, position, instruction, and environment perturbations to LIBERO to test whether a VLA model is just memorizing.","explanation":"LIBERO-PRO is an extended evaluation benchmark released in October 2025 by teams from Huazhong University of Science and Technology, Harvard, MIT, Lehigh University, and others. The authors point out that LIBERO's training and test scenes are nearly identical, so a model can score well by memorizing action sequences and tabletop layouts, inflating scores and making fair comparison difficult. LIBERO-PRO adds reasonable perturbations across four dimensions — the manipulated object, initial state, task instruction, and environment — which can be freely combined. Results show that models scoring above 90% success under the standard setup can drop to 0% under the generalization setup: a model keeps reaching for the target object even after it's swapped for an unrelated one, and its actions barely change even when the instruction is scrambled or replaced with gibberish. It appeared around the same time as LIBERO-Plus, and both point to LIBERO having become saturated as a benchmark.","example":"The project homepage reports OpenVLA scoring 0.98 success on the original tasks, dropping to 0.00 after object-position perturbation; π0.5 drops from 0.97 to 0.38.","related":["LIBERO Benchmark","LIBERO-Plus","Benchmark Saturation","Generalization / Robustness Evaluation","Out-of-Distribution","Leaderboard Chasing"]},{"id":"simplerenv","category":"sim","sec":8,"tier":1,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"},{"title":"SIMPLER 项目主页","url":"https://simpler-env.github.io/"},{"title":"simpler-env/SimplerEnv (GitHub)","url":"https://github.com/simpler-env/SimplerEnv"}],"as_of":"2026-09","related_ids":["simulation-based-evaluation","visual-matching","variant-aggregation","mean-maximum-rank-violation","sim-to-real-correlation","maniskill"],"name":"SimplerEnv","alt":"SimplerEnv","abbr":"","aliases":["SIMPLER: Simulated Manipulation Policy Evaluation for Real Robot Setups","SIMPLER"],"one_liner":"A simulated evaluation suite that replicates common real-robot manipulation setups for cheap policy scoring.","explanation":"SimplerEnv (paper name: SIMPLER) was proposed in 2024 by researchers at UC San Diego, Stanford, UC Berkeley, and Google DeepMind, published at CoRL 2024. It replicates two widely used real-robot setups inside the SAPIEN/ManiSkill simulator: the Google Robot platform used to collect Google's RT-1 dataset, and the WidowX arm used with the Bridge dataset — so a policy trained on real-robot data can be scored without ever touching the real robot. To narrow the sim-to-real gap, it calibrates controller parameters and offers two evaluation modes: “visual matching” (compositing simulated objects onto a real background image) and “variant aggregation” (building several simulated variants with different backgrounds, lighting, and tabletop textures, then averaging the results). The authors verified that simulated and real-robot rankings of methods correlate closely, and many VLA papers since have reported success rate on it.","example":"To evaluate a WidowX policy, run tasks like “put the spoon on the towel,” “put the carrot on the plate,” and “put the eggplant in the basket” directly in SimplerEnv and tally success rate, without arranging objects by hand on a real robot each time.","related":["Simulation-Based Evaluation","Visual Matching (SimplerEnv)","Variant Aggregation (SimplerEnv)","Mean Maximum Rank Violation","Sim-to-Real Correlation","ManiSkill"]},{"id":"visual-matching","category":"sim","sec":8,"tier":3,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (arXiv 2405.05941)","url":"https://arxiv.org/abs/2405.05941"},{"title":"SIMPLER 项目主页","url":"https://simpler-env.github.io/"},{"title":"SimplerEnv GitHub 仓库","url":"https://github.com/simpler-env/SimplerEnv"}],"as_of":"2024-05","related_ids":["simplerenv","variant-aggregation","sim-to-real-correlation","mean-maximum-rank-violation","sim-to-real-gap","real-to-sim"],"name":"Visual Matching (SimplerEnv)","alt":"视觉匹配","abbr":"VM","aliases":["VM","SimplerEnv Visual Matching","SIMPLER Visual Matching"],"one_liner":"A SimplerEnv evaluation setup that makes the simulated image look as close as possible to real-robot camera footage.","explanation":"Visual Matching is one of two evaluation setups from SimplerEnv (SIMPLER, released in May 2024 by UC San Diego, Stanford, Berkeley, and Google DeepMind), used to evaluate in simulation manipulation policies that were trained on real-robot data. Such policies often fail as soon as they enter simulation, simply because the images look different. Visual Matching uses “green-screening” to composite the simulated objects and robot arm onto a real background photo, then projects real textures onto the simulated models and adjusts the arm's color to match real footage, bringing the image closer to what a real robot camera sees. The other setup, Variant Aggregation, instead generates multiple variants by changing background, lighting, and distractors and averages across them. Both are checked against real-robot results using the Pearson correlation coefficient and Mean Maximum Rank Violation (MMRV).","example":"VLA papers evaluating on SimplerEnv's Google-robot tasks — picking up a Coke can, moving near an object, opening and closing a drawer — typically report success rate in two separate columns, “Visual Matching” and “Variant Aggregation.”","related":["SimplerEnv","Variant Aggregation (SimplerEnv)","Sim-to-Real Correlation","Mean Maximum Rank Violation","Sim-to-Real Gap (Reality Gap)","Real-to-Sim"]},{"id":"variant-aggregation","category":"sim","sec":8,"tier":3,"sources":[{"title":"Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER, arXiv 2405.05941)","url":"https://arxiv.org/html/2405.05941"},{"title":"simpler-env/SimplerEnv GitHub 仓库","url":"https://github.com/simpler-env/SimplerEnv"}],"as_of":"2024-05","related_ids":["simplerenv","visual-matching","mean-maximum-rank-violation","sim-to-real-correlation","visual-randomization","simulation-based-evaluation"],"name":"Variant Aggregation (SimplerEnv)","alt":"变体聚合","abbr":"VA","aliases":["VA"],"one_liner":"One of SimplerEnv's two evaluation modes: testing a policy across many visually randomized scene variants and averaging the results.","explanation":"Variant Aggregation is one of two “real-to-sim” evaluation approaches proposed by SimplerEnv (SIMPLER), from a paper by Xuanlin Li and colleagues at CoRL 2024. Simulated images and real-robot images always differ somewhat; the other approach, Visual Matching, tries to make the simulated image look as close to reality as possible. Variant Aggregation instead does the opposite: it applies heavy visual randomization to a scene, generating multiple environment variants along axes such as background, lighting, distractor objects, table texture, and camera pose, measures success rate separately in each, then averages them to estimate how a policy performs overall under visual variation. The paper uses Mean Maximum Rank Violation (MMRV) and the Pearson correlation coefficient to measure whether simulated results agree with real-robot results; in the paper's experiments, Visual Matching agreed with real robots better overall. VLA papers reporting SimplerEnv results on Google-robot tasks commonly list both a VM and a VA column.","example":"For the “pick up the Coke can” task, run the same policy through several variants — a different background, different lighting, added distractors, a different table texture, a different camera angle — and average the success rates across them to get its VA score.","related":["SimplerEnv","Visual Matching (SimplerEnv)","Mean Maximum Rank Violation","Sim-to-Real Correlation","Visual Randomization","Simulation-Based Evaluation"]},{"id":"calvin-benchmark","category":"sim","sec":8,"tier":2,"sources":[{"title":"CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv 2112.03227)","url":"https://arxiv.org/abs/2112.03227"},{"title":"CALVIN GitHub 仓库","url":"https://github.com/mees/calvin"}],"as_of":"","related_ids":["average-length","language-conditioned-policy","long-horizon-task","libero-benchmark","pybullet","play-data"],"name":"CALVIN Benchmark","alt":"CALVIN","abbr":"","aliases":["CALVIN ABC→D","Composing Actions from Language and Vision"],"one_liner":"A tabletop manipulation benchmark testing whether a robot can follow 5 language instructions in a row.","explanation":"CALVIN (Composing Actions from Language and Vision) is an open-source simulated benchmark released by Oier Mees, Wolfram Burgard, and colleagues at the University of Freiburg in Germany, published in RA-L 2022, where it won that journal's best-paper award that year. The setup is a table with a 7-DOF Franka arm, plus a drawer, a sliding door, a button, a switch, and three colored blocks, simulated in PyBullet; there are four environments, A, B, C, and D, structurally identical but differing in texture and part placement. It provides about 24 hours of teleoperated “play” data (only 1% of which is labeled with language), defines 34 task types, and its main evaluation requires executing 5 language instructions in a row, scored with Average Length. The most commonly used ABC→D setting trains on three environments and tests on the fourth, unseen one, to measure generalization.","example":"One test chain from the paper: “open the drawer” → “push the block into the drawer” → “take the block back out of the drawer” → “stack the blocks” → “close the drawer,” with the robot only advancing to the next instruction once it completes the current one.","related":["Average Length (CALVIN)","Language-conditioned Policy","Long-horizon Task","LIBERO Benchmark","PyBullet","Play Data"]},{"id":"average-length","category":"sim","sec":8,"tier":2,"sources":[{"title":"CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks (arXiv 2112.03227)","url":"https://arxiv.org/abs/2112.03227"},{"title":"CALVIN 官方评测脚本 evaluate_policy.py","url":"https://github.com/mees/calvin/blob/main/calvin_models/calvin_agent/evaluation/evaluate_policy.py"},{"title":"CALVIN 评测工具 utils.py（avg_seq_len 计算）","url":"https://github.com/mees/calvin/blob/main/calvin_models/calvin_agent/evaluation/utils.py"}],"as_of":"","related_ids":["calvin-benchmark","success-rate","long-horizon-task","language-conditioned-policy","closed-loop-evaluation","progress-score"],"name":"Average Length (CALVIN)","alt":"平均完成长度","abbr":"Avg. Len","aliases":["Avg. Len","Average Successful Sequence Length"],"one_liner":"In the CALVIN benchmark, how many consecutive instructions a policy completes on average, out of 5.","explanation":"Average Length is the core metric of CALVIN's long-horizon evaluation. During evaluation, a policy attempts 1,000 instruction chains in sequence, each made of 5 consecutive language instructions (for example, open a drawer, then push a block into it), with each subtask capped at 360 steps; the chain ends the moment any subtask fails. The number of subtasks completed in a row (0 through 5) is recorded for every chain, and Avg. Len is the average across all chains. It's equal to the sum of the success rates for completing at least 1, at least 2, … up to all 5 tasks in a row, so it captures both single-step competence and whether errors compound over a long horizon. Papers typically report this number under the ABC→D setting (train on environments A, B, and C, test on the unseen environment D), to compare generalization and long-horizon ability.","example":"The CALVIN paper's baseline, MCIL, gets success rates of 48.9%, 12.9%, 2.6%, 0.5%, and 0.08% for completing 1 through 5 tasks in a row under the D→D setting — summing to about 0.65, or an average of well under one completed task.","related":["CALVIN Benchmark","Success Rate","Long-horizon Task","Language-conditioned Policy","Closed-Loop Evaluation","Progress Score"]},{"id":"rlbench","category":"sim","sec":8,"tier":2,"sources":[{"title":"RLBench: The Robot Learning Benchmark & Learning Environment (arXiv 1909.12271)","url":"https://arxiv.org/abs/1909.12271"},{"title":"stepjam/RLBench (GitHub)","url":"https://github.com/stepjam/RLBench"},{"title":"Perceiver-Actor (arXiv 2209.05451)","url":"https://arxiv.org/abs/2209.05451"}],"as_of":"","related_ids":["peract","rvt-2","3d-diffuser-actor","coppeliasim","keyframe-action-prediction","the-colosseum-a-benchmark-for-evaluating-generalization-for"],"name":"RLBench","alt":"RLBench","abbr":"","aliases":["RLBench-18 (PerAct subset)"],"one_liner":"A simulation benchmark of 100 robot-arm manipulation tasks from Imperial College London, built on CoppeliaSim.","explanation":"RLBench was released in 2019 (IEEE RA-L 2020) by Stephen James, Andrew Davison, and colleagues at Imperial College London's Dyson Robotics Lab. It contains 100 hand-designed manipulation tasks, ranging from simple reaching to multi-step tasks like opening a door or an oven, using a Franka Panda arm by default and running on the CoppeliaSim simulator. Each task provides multi-camera RGB, depth, segmentation-mask, and proprioceptive observations, can automatically generate any number of demonstrations via motion planning, and many tasks also have variants that change color, position, and similar attributes. It was originally aimed at reinforcement learning, imitation learning, and few-shot learning, and later became one of the main simulation benchmarks for 3D manipulation policies: the 18-task subset selected by PerAct has since been reused by a large body of work including RVT and 3D Diffuser Actor.","example":"PerAct trained a single multi-task Transformer on 18 RLBench tasks (249 variants), voxelizing multi-view RGB-D observations to predict the end effector's next keyframe pose; later work such as RVT-2 and 3D Diffuser Actor compares success rates on the same set of tasks.","related":["PerAct","RVT-2","3D Diffuser Actor","CoppeliaSim","Keyframe Action Prediction","The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation"]},{"id":"gembench","category":"sim","sec":8,"tier":3,"sources":[{"title":"Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy (arXiv 2410.01345)","url":"https://arxiv.org/abs/2410.01345"},{"title":"GemBench project page","url":"https://www.di.ens.fr/willow/research/gembench/"}],"as_of":"2025-05","related_ids":["rlbench","generalization","compositional-generalization","long-horizon-task","language-conditioned-policy","benchmark"],"name":"GemBench","alt":"GemBench","abbr":"","aliases":["GEMBench"],"one_liner":"An RLBench-based simulation benchmark that tests a language-conditioned manipulation policy's generalization across four difficulty levels.","explanation":"GemBench was proposed by Ricardo Garcia, Shizhe Chen, and Cordelia Schmid at Inria and ENS Paris, published at ICRA 2025. Built on the RLBench simulator, it defines 7 action primitives — press, grasp, push, turn, close, open, and place/stack — trains on 16 tasks (31 variants), and tests on 44 tasks (92 variants), with generalization difficulty split into four levels: new object placements, new rigid objects, new articulated objects, and new long-horizon tasks. The same paper also proposes 3D-LOTUS (a point-cloud-based language-conditioned policy) and 3D-LOTUS++ (which adds a large language model for task planning and a vision-language model for object localization). Results show that pure imitation-learning policies score near-perfectly on familiar tasks but drop off sharply when faced with unlearned combinations of skills.","example":"At level 1 (only object placement changes), 3D-LOTUS scores 94.3% success; at level 4, which requires combining learned actions into new long-horizon tasks, it scores only 0.3%, while 3D-LOTUS++ with added LLM planning reaches 17.4%.","related":["RLBench","Generalization","Compositional Generalization","Long-horizon Task","Language-conditioned Policy","Benchmark"]},{"id":"the-colosseum-a-benchmark-for-evaluating-generalization-for","category":"sim","sec":8,"tier":3,"sources":[{"title":"THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation (arXiv 2402.08191)","url":"https://arxiv.org/abs/2402.08191"},{"title":"The Colosseum 项目主页","url":"https://robot-colosseum.github.io/"},{"title":"robot-colosseum GitHub 仓库","url":"https://github.com/robot-colosseum/robot-colosseum"}],"as_of":"2024-05","related_ids":["rlbench","generalization-robustness-evaluation","distractor-objects","domain-randomization","peract","visual-generalization"],"name":"The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation","alt":"The Colosseum","abbr":"","aliases":["Colosseum"],"one_liner":"A benchmark applying 14 kinds of systematic environmental perturbations to RLBench tasks to test manipulation policies' generalization.","explanation":"The Colosseum was proposed by Wilbert Pumacay, Jiafei Duan, Dieter Fox, and colleagues, and published at RSS 2024. Built on the PyRep simulation framework, it selects 20 of RLBench's 100 tasks, and each task can be perturbed along 14 dimensions: the color, texture, and size of both the manipulated object and static objects, light color, table color and texture, background texture, number of distractor objects, camera pose, and object friction and mass. The authors used it to test methods including PerAct, RVT, R3M, MVP, and VoxPoser, and found that a single perturbation alone drops success rate by 30%–50%, and applying several perturbations together drops it by more than 75%, with distractor count, target-object color, and lighting having the biggest effects. In a real-robot replication experiment, the R² between simulation and real-robot results was 0.614.","example":"For the same task, run separate trials changing only the table texture, only adding distractors, and only changing the camera pose, then compare each success rate against the unperturbed baseline to see which kind of change the policy is most sensitive to.","related":["RLBench","Generalization / Robustness Evaluation","Distractor Objects","Domain Randomization","PerAct","Visual Generalization"]},{"id":"push-t","category":"sim","sec":8,"tier":2,"sources":[{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"huggingface/gym-pusht (GitHub)","url":"https://github.com/huggingface/gym-pusht"}],"as_of":"","related_ids":["diffusion-policy","action-multimodality","imitation-learning","lerobot","non-prehensile-manipulation","aloha-sim"],"name":"Push-T","alt":"Push-T","abbr":"","aliases":["PushT","gym-pusht"],"one_liner":"A 2D task where a round pusher must push a T-shaped block to a target pose, commonly used to test imitation-learning policies.","explanation":"Push-T is a 2D tabletop pushing task: an agent controls a round pusher that can only move the T-shaped block on the table through point contact, and must push it into a fixed target position and orientation. It first appeared in Google's 2021 Implicit Behavioral Cloning (IBC) work, and became a standard imitation-learning benchmark after Diffusion Policy adapted it in 2023; Hugging Face later packaged it as the gym-pusht environment. The task looks simple, but the block's pose has to be adjusted bit by bit through contact, and the same situation often has several equally valid human strategies, such as pushing from the left or from the right (action multimodality), which makes it well suited for testing whether a policy can represent a multimodal action distribution. The metric is the overlap ratio between the T block and the target region; in gym-pusht, 95% overlap counts as success. Observations can be either keypoints or 96×96 images.","example":"The Diffusion Policy paper compared its method against IBC, LSTM-GMM, and others on simulated Push-T, and also built a real-robot version: trained with a UR5 arm and 136 human demonstrations, requiring more precise multi-stage pushing.","related":["Diffusion Policy","Action Multimodality","Imitation Learning","LeRobot","Non-prehensile Manipulation","ALOHA Sim (Transfer Cube / Insertion)"]},{"id":"aloha-sim","category":"sim","sec":8,"tier":2,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)","url":"https://arxiv.org/abs/2304.13705"},{"title":"huggingface/gym-aloha (GitHub)","url":"https://github.com/huggingface/gym-aloha"},{"title":"tonyzhaozh/act (GitHub)","url":"https://github.com/tonyzhaozh/act"}],"as_of":"2026-09","related_ids":["action-chunking-with-transformers","aloha","bimanual-manipulation","lerobot","mujoco","imitation-learning"],"name":"ALOHA Sim (Transfer Cube / Insertion)","alt":"ALOHA 仿真任务","abbr":"","aliases":["gym-aloha","Transfer Cube","Insertion","AlohaTransferCube-v0","AlohaInsertion-v0"],"one_liner":"The two bimanual simulated tasks — cube transfer and peg insertion — built in MuJoCo for the ACT paper.","explanation":"The ALOHA simulated tasks are two bimanual fine-manipulation tasks built in MuJoCo by Tony Zhao and colleagues for the 2023 ACT (Action Chunking Transformer) paper, meant to let others reproduce the results easily. Transfer Cube requires the right arm to pick up a red block on the table and hand it off into the left arm's gripper, with roughly a 1 cm clearance; Insertion requires the two arms to each pick up a socket and a peg and mate them in mid-air, with roughly a 5 mm clearance. Each task ships with 50 scripted demonstrations and 50 human-teleoperated ones. Hugging Face packaged it as the gym-aloha environment (a 14-dimensional action space: 6 joints plus 1 gripper per arm), and provides the dataset through LeRobot, making it a common starting benchmark for learning imitation learning.","example":"The ACT project's own documentation notes that, trained on 50 scripted demonstrations, success rate should land around 90% for cube transfer and around 50% for insertion.","related":["Action Chunking with Transformers","ALOHA","Bimanual Manipulation","LeRobot","MuJoCo (Multi-Joint dynamics with Contact)","Imitation Learning"]},{"id":"robotwin","category":"sim","sec":8,"tier":2,"sources":[{"title":"RoboTwin 官网","url":"https://robotwin-platform.github.io/"},{"title":"RoboTwin 2.0 (arXiv 2506.18088)","url":"https://arxiv.org/abs/2506.18088"}],"as_of":"2025-06","related_ids":["bimanual-manipulation","domain-randomization","synthetic-data","sapien","benchmark","cross-embodiment"],"name":"RoboTwin","alt":"RoboTwin","abbr":"","aliases":["RoboTwin 2.0"],"one_liner":"A bimanual-manipulation simulation data generator and evaluation benchmark from the University of Hong Kong, Shanghai AI Lab, and others.","explanation":"RoboTwin is a simulation data-generation and evaluation platform for bimanual manipulation, released by more than a dozen institutions including the University of Hong Kong's MMLab, Shanghai AI Lab, Tsinghua, and Shanghai Jiao Tong University, built on the SAPIEN simulator; version 1.0 received a CVPR 2025 Highlight. Version 2.0, released in June 2025, has a multimodal large model automatically write expert manipulation code, which is verified in simulation and then used to generate trajectories, covering 50 bimanual tasks, 5 robot embodiments, and 731 objects, with over 100,000 trajectories pre-collected. It applies domain randomization (randomly varying environmental conditions during training) across five dimensions — clutter, background texture, lighting, table height, and language instructions — to improve real-robot robustness, and it also serves as a common bimanual evaluation benchmark for VLA models.","example":"The RoboTwin 2.0 paper reports that a VLA model trained on its heavily randomized data achieved a 367% relative improvement in previously unseen real-world scenes, and the CVPR 2025 MEIS workshop ran a challenge built on it.","related":["Bimanual Manipulation","Domain Randomization","Synthetic Data","SAPIEN (SimulAted Part-based Interactive ENvironment)","Benchmark","Cross-Embodiment"]},{"id":"meta-world","category":"sim","sec":8,"tier":2,"sources":[{"title":"Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning (arXiv 1910.10897)","url":"https://arxiv.org/abs/1910.10897"},{"title":"Farama-Foundation/Metaworld (GitHub)","url":"https://github.com/Farama-Foundation/Metaworld"}],"as_of":"2026-09","related_ids":["meta-reinforcement-learning","multi-task-learning","mujoco","benchmark","success-rate","gymnasium"],"name":"Meta-World","alt":"Meta-World","abbr":"","aliases":["MetaWorld","Meta-World+","MT10","MT50","ML10","ML45"],"one_liner":"A simulation benchmark of 50 robot-arm manipulation tasks, used to test multi-task and meta-reinforcement learning.","explanation":"Meta-World is an open-source simulation benchmark proposed by Tianhe Yu, Chelsea Finn, Sergey Levine, and colleagues at CoRL 2019: a Sawyer robot arm performs 50 distinct manipulation tasks in MuJoCo, such as opening a drawer, pressing a button, pushing a block, or opening a window. It supports two evaluation setups: multi-task learning (MT1/MT10/MT50, learning 1, 10, or 50 tasks at once) and meta-learning (ML1/ML10/ML45, training on one set of tasks and then testing how fast the model adapts to new ones; ML45 uses 45 training tasks plus 5 held-out test tasks). The original paper found that existing algorithms already struggled to learn just 10 tasks at once. The project is now maintained by the Farama Foundation, has switched to the Gymnasium interface, and released Meta-World+ (NeurIPS 2025), which standardized version details; it remains a common benchmark for multi-task learning, meta-reinforcement learning, and imitation learning.","example":"In the ML45 setup, an agent is meta-trained on 45 tasks and then faces 5 tasks it has never practiced, such as a novel way of opening a door; it is scored on how quickly it can learn the new task from a small amount of interaction, measured by success rate.","related":["Meta Reinforcement Learning","Multi-Task Learning","MuJoCo (Multi-Joint dynamics with Contact)","Benchmark","Success Rate","Gymnasium"]},{"id":"franka-kitchen","category":"sim","sec":8,"tier":3,"sources":[{"title":"Franka Kitchen - Gymnasium-Robotics Documentation","url":"https://robotics.farama.org/envs/franka_kitchen/franka_kitchen/"},{"title":"Relay Policy Learning (arXiv 1910.11956)","url":"https://arxiv.org/abs/1910.11956"},{"title":"Minari: D4RL Kitchen datasets","url":"https://minari.farama.org/datasets/D4RL/kitchen/"}],"as_of":"2026-09","related_ids":["d4rl","mujoco","long-horizon-task","offline-reinforcement-learning","sparse-reward","franka-emika-panda-franka-research-3"],"name":"Franka Kitchen","alt":"Franka Kitchen","abbr":"","aliases":["FrankaKitchen","D4RL Kitchen"],"one_liner":"A MuJoCo kitchen scene where a Franka arm completes sub-tasks in sequence, such as opening a microwave and moving a kettle.","explanation":"Franka Kitchen comes from the 2019 CoRL paper Relay Policy Learning (Gupta, Levine, Hausman, and colleagues), which built a kitchen in MuJoCo containing a 9-degree-of-freedom Franka arm (7 arm joints plus 2 fingers) that can operate a microwave door, a kettle, a light switch, a sliding cabinet door, a hinged cabinet door, and stove knobs. Each episode has to complete several specified sub-tasks, earning 1 point per completion, making it a sparse-reward, long-horizon, multi-task setting. It was later incorporated into the D4RL offline reinforcement-learning benchmark, providing complete, partial, and mixed demonstration data, and is now maintained by the Farama Foundation's Gymnasium-Robotics; it's commonly used to test whether offline reinforcement learning, imitation learning, and hierarchical policies can chain multiple skills together.","example":"In D4RL's kitchen-complete data, every demonstration completes the same 4 sub-tasks in order — opening the microwave, moving the kettle, flipping the light switch, and pushing open the sliding cabinet door; the mixed data instead completes these 4 sub-tasks out of order and incompletely, testing whether an algorithm can stitch fragments together.","related":["D4RL","MuJoCo (Multi-Joint dynamics with Contact)","Long-horizon Task","Offline Reinforcement Learning","Sparse Reward","Franka Emika Panda / Franka Research 3"]},{"id":"adroit","category":"sim","sec":8,"tier":3,"sources":[{"title":"Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations (arXiv 1709.10087)","url":"https://arxiv.org/abs/1709.10087"},{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning (arXiv 2004.07219)","url":"https://arxiv.org/abs/2004.07219"},{"title":"Gymnasium-Robotics: Adroit Hand","url":"https://robotics.farama.org/envs/adroit_hand/"}],"as_of":"","related_ids":["dexterous-manipulation","in-hand-manipulation","shadow-dexterous-hand","d4rl","mujoco","sparse-reward"],"name":"Adroit","alt":"Adroit 灵巧手任务","abbr":"","aliases":["Adroit Hand","ADROIT","DAPG Tasks"],"one_liner":"A benchmark for controlling a 24-DoF simulated five-fingered hand in MuJoCo to open doors, hammer nails, and more, across four tasks.","explanation":"The Adroit tasks come from a 2018 paper by Aravind Rajeswaran, Vikash Kumar, Sergey Levine, and colleagues (University of Washington, OpenAI, Berkeley), simulating a 24-degree-of-freedom ADROIT anthropomorphic hand in MuJoCo (D4RL refers to it as a simulated Shadow Hand) across four tasks: moving a ball to a target location, spinning a pen in-hand to a target orientation, hammering a nail, and pulling open a latched door. The authors collected 25 human demonstrations per task using a VR data glove, and proposed DAPG: behavioral cloning on the demonstrations first, followed by reinforcement-learning fine-tuning with a policy-gradient method that incorporates the demonstrations. D4RL later incorporated these tasks into its offline reinforcement-learning benchmark, providing human, cloned, and expert data for each, and Gymnasium-Robotics also maintains these environments. They are commonly used to test high-dimensional dexterous manipulation, sparse rewards, and demonstration-assisted learning.","example":"D4RL's pen-human dataset consists of the 25 human demonstrations for the pen-spinning task, and offline reinforcement-learning papers commonly use it to compare algorithms.","related":["Dexterous Manipulation","In-hand Manipulation","Shadow Dexterous Hand","D4RL","MuJoCo (Multi-Joint dynamics with Contact)","Sparse Reward"]},{"id":"bi-dexhands","category":"sim","sec":8,"tier":3,"sources":[{"title":"Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning (arXiv 2206.08686)","url":"https://arxiv.org/abs/2206.08686"},{"title":"PKU-MARL/DexterousHands (GitHub)","url":"https://github.com/PKU-MARL/DexterousHands"}],"as_of":"2022-10","related_ids":["dexterous-manipulation","bimanual-manipulation","shadow-dexterous-hand","isaac-gym","multi-agent-reinforcement-learning","proximal-policy-optimization"],"name":"Bi-DexHands","alt":"Bi-DexHands","abbr":"","aliases":["DexterousHands"],"one_liner":"A reinforcement-learning benchmark from a Peking University team, using two Shadow dexterous hands in Isaac Gym.","explanation":"Bi-DexHands comes from Peking University's PKU-MARL team; its paper, “Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning,” appeared in the NeurIPS 2022 Datasets and Benchmarks track, and its code repository is named DexterousHands. It places two Shadow dexterous hands in Isaac Gym (NVIDIA's GPU-parallel simulator), with dozens of bimanual tasks and thousands of target objects, and the tasks were designed to correspond to different levels of human motor skill drawn from the cognitive-science literature. The paper reports over 30,000 frames per second on a single RTX 3090. The benchmark covers single-agent, multi-agent (treating each hand as a separate agent), offline, multi-task, and meta-reinforcement learning. It concludes that PPO-family algorithms can master simple tasks at roughly the level of a 48-month-old child, that multi-agent methods help on tasks requiring close two-hand coordination, but that existing algorithms mostly fail under multi-task and few-shot settings.","example":"In the paper, PPO-family algorithms learn tasks like catching a thrown object or opening a bottle; tasks requiring tight two-hand coordination, such as lifting a pot or stacking blocks, are learned more easily with multi-agent reinforcement learning.","related":["Dexterous Manipulation","Bimanual Manipulation","Shadow Dexterous Hand","Isaac Gym","Multi-Agent Reinforcement Learning","Proximal Policy Optimization"]},{"id":"softgym-benchmarking-deep-reinforcement-learning-for-deforma","category":"sim","sec":8,"tier":3,"sources":[{"title":"arXiv 2011.07215 - SoftGym","url":"https://arxiv.org/abs/2011.07215"},{"title":"GitHub - Xingyu-Lin/softgym","url":"https://github.com/Xingyu-Lin/softgym"}],"as_of":"","related_ids":["deformable-object-manipulation","garment-manipulation","cloth-simulation","fluid-simulation","position-based-dynamics","benchmark"],"name":"SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation","alt":"SoftGym","abbr":"","aliases":[],"one_liner":"A deformable-object manipulation benchmark built on NVIDIA FleX, covering cloth, rope, and liquid tasks.","explanation":"SoftGym is an open-source benchmark from Xingyu Lin and colleagues in David Held's group at Carnegie Mellon University, published at CoRL 2020, built specifically to evaluate reinforcement learning on deformable-object manipulation. Earlier RL benchmarks were mostly rigid-body or low-dimensional state tasks, whereas cloth, rope, and liquids have extremely high-dimensional state that is only partially observable — a completely different level of difficulty. SoftGym is built on NVIDIA's FleX particle physics engine (accessed through the Python interface PyFleX) and provides a standard OpenAI Gym interface. Tasks come in two tiers: medium difficulty includes transporting water, pouring water, straightening a rope, spreading cloth flat, and folding cloth; hard includes precise water pouring, folding a wrinkled cloth, dropping and then folding cloth, and shaping a rope into a specified form. The paper's experiments show existing algorithms still struggle noticeably on these tasks, especially when only given image observations. Because it depends on older CUDA and system versions, the authors recommend installing it via Docker.","example":"In SoftGym's FoldCloth task, an agent sees only an overhead camera image and controls two grasp points to fold a flat piece of cloth in half, rewarded by how well the two halves align.","related":["Deformable Object Manipulation","Garment Manipulation","Cloth Simulation","Fluid Simulation","Position-Based Dynamics","Benchmark"]},{"id":"vima-bench","category":"sim","sec":8,"tier":3,"sources":[{"title":"VIMA: General Robot Manipulation with Multimodal Prompts (arXiv 2210.03094)","url":"https://arxiv.org/html/2210.03094"},{"title":"vimalabs/VIMABench GitHub 仓库","url":"https://github.com/vimalabs/VIMABench"}],"as_of":"2023-05","related_ids":["vima","benchmark","compositional-generalization","zero-shot","tabletop-manipulation","transporter-networks"],"name":"VIMA-Bench","alt":"VIMA-Bench","abbr":"","aliases":["VIMABench"],"one_liner":"A tabletop manipulation benchmark driven by interleaved text-and-image prompts, with four levels of generalization testing.","explanation":"VIMA-Bench was introduced alongside the VIMA model, first-authored by Yunfan Jiang with collaborators including Fei-Fei Li and Anima Anandkumar, published at ICML 2023. Built on the Ravens simulator (a PyBullet-based tabletop manipulation environment), it extends to 17 meta-tasks whose instructions are written as multimodal prompts interleaving text and images, spanning categories such as simple object manipulation, visual goal reaching, novel-concept understanding, one-shot video imitation, visual-constraint satisfaction, and visual reasoning, and it provides 650,000 successful demonstration trajectories. Evaluation has four levels: L1 only randomizes object placement; L2 recombines already-seen objects and textures in new ways; L3 introduces new objects and textures; L4 is entirely new tasks — 4 of the 17 tasks are held out specifically to test zero-shot generalization. The action space is one grasp pose plus one placement pose.","example":"The VIMA paper reports that under the hardest zero-shot generalization setting, with the same amount of training data VIMA's task success rate was up to 2.9 times that of other approaches, and still 2.7 times higher with 10 times less training data.","related":["VIMA","Benchmark","Compositional Generalization","Zero-shot","Tabletop Manipulation","Transporter Networks"]},{"id":"genmanip","category":"sim","sec":8,"tier":3,"sources":[{"title":"GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation (arXiv 2506.10966)","url":"https://arxiv.org/abs/2506.10966"},{"title":"GenManip Suite project page","url":"https://genmanip.com/"}],"as_of":"2025-06","related_ids":["nvidia-isaac-sim","instruction-following","generative-simulation","synthetic-data","copa","shanghai-artificial-intelligence-laboratory"],"name":"GenManip","alt":"GenManip","abbr":"","aliases":["LLM-driven Simulation for Generalizable Instruction-Following Manipulation","GenManip-Bench","GenManip Suite"],"one_liner":"A manipulation simulation and evaluation platform from Shanghai AI Lab, built on Isaac Sim, that uses a large model to auto-generate tasks.","explanation":"GenManip was proposed by Shanghai AI Lab together with Zhejiang University, Xi'an Jiaotong University, Nanjing University, and others, published at CVPR 2025. It's a tabletop-manipulation simulation platform built on NVIDIA Isaac Sim, focused on whether a policy can understand a wide range of language instructions. It uses a large language model to generate task-oriented scene graphs (describing which objects are in the scene and what the target relationships are), paired with 10,000 annotated 3D object assets, to automatically synthesize a large amount of diverse tasks and demonstration data. Its evaluation component, GenManip-Bench, contains 200 hand-refined scenes and tests four kinds of generalization: spatial relationships, appearance understanding, common-sense reasoning, and long-horizon tasks. The paper compares two approaches: a modular system that uses foundation models for perception and planning, and an end-to-end policy trained with behavioral cloning, finding that the former generalizes better zero-shot while the latter improves as more data is added. The code is open-sourced on GitHub.","example":"On GenManip-Bench, the best-performing modular system, CoPa (paired with GPT-4.5), reaches an overall success rate of 23.0%; on long-horizon tasks, the evaluated models average only 9.07%.","related":["NVIDIA Isaac Sim","Instruction Following","Generative Simulation","Synthetic Data","CoPa","Shanghai Artificial Intelligence Laboratory"]},{"id":"arnold","category":"sim","sec":8,"tier":3,"sources":[{"title":"ARNOLD: A Benchmark for Language-Grounded Task Learning With Continuous States in Realistic 3D Scenes (arXiv 2304.04321)","url":"https://arxiv.org/abs/2304.04321"},{"title":"ARNOLD project page","url":"https://arnold-benchmark.github.io/"}],"as_of":"","related_ids":["language-conditioned-policy","articulated-object-manipulation","nvidia-isaac-sim","physx","generalization-robustness-evaluation","benchmark"],"name":"ARNOLD","alt":"ARNOLD","abbr":"","aliases":["A Benchmark for Language-Grounded Task Learning with Continuous States in Realistic 3D Scenes"],"one_liner":"A benchmark in Isaac Sim testing whether a robot can manipulate objects to a specified continuous state based on language.","explanation":"ARNOLD was proposed by the Beijing Institute for General Artificial Intelligence (BIGAI) together with UCLA, Peking University, Tsinghua, and Columbia, published at ICCV 2023. Earlier language-conditioned manipulation benchmarks mostly treated the goal as a binary state like open/closed; ARNOLD instead requires reaching a continuous target value, such as pulling a drawer open to a specified degree or pouring out a specified proportion of water, counting success only when the object's state stays within a tolerance range around the target. It is built on NVIDIA Isaac Sim and PhysX 5.0, and includes 8 tasks (picking up an object, adjusting orientation, opening/closing a drawer, opening/closing a cabinet door, pouring water, transferring water), 40 object types, 20 scenes, and 10,000 expert demonstrations, with language instructions generated from templates. Besides in-distribution testing, evaluation separately reports generalization splits for novel objects, novel scenes, and novel target states; the authors found that the language-conditioned policies available at the time performed noticeably worse on these splits.","example":"Given an instruction to open a cabinet door halfway, the robot must not only open the door but also stop its opening angle within an allowed range around the target value to count as successful.","related":["Language-conditioned Policy","Articulated Object Manipulation","NVIDIA Isaac Sim","PhysX","Generalization / Robustness Evaluation","Benchmark"]},{"id":"vlabench-a-large-scale-benchmark-for-language-conditioned-ro","category":"sim","sec":8,"tier":3,"sources":[{"title":"VLABench (arXiv 2412.18194)","url":"https://arxiv.org/abs/2412.18194"},{"title":"VLABench GitHub 仓库（OpenMOSS）","url":"https://github.com/OpenMOSS/VLABench"}],"as_of":"2025-06","related_ids":["vision-language-action-model","instruction-following","long-horizon-task","libero-benchmark","mujoco","benchmark"],"name":"VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks","alt":"VLABench","abbr":"","aliases":[],"one_liner":"A Fudan University language-conditioned manipulation benchmark focused on common sense, implicit intent, and long-horizon multi-step reasoning.","explanation":"VLABench is an open-source benchmark released by Fudan University's OpenMOSS team in December 2024 and accepted by ICCV in 2025, evaluating the ability to manipulate a robot arm from natural-language instructions. Built on MuJoCo and dm_control, it uses a 7-DoF Franka arm by default, across 100 task categories (60 atomic, 40 compositional) with more than 2,000 object assets. Compared with earlier benchmarks whose instructions are mostly fixed templates, its tasks require common sense and world knowledge, instructions carry implicit intent, and long-horizon tasks require multi-step reasoning; both VLA policies and VLM-driven workflows can be evaluated on it. The project also provides automatically generated training data and six evaluation tracks.","example":"Its instructions don't always state directly what to pick up — they carry implicit intent, so the model first has to infer the target object before acting; the paper's results show that even the strongest pretrained VLAs and VLM-based pipelines at the time struggled with these tasks.","related":["Vision-Language-Action Model","Instruction Following","Long-horizon Task","LIBERO Benchmark","MuJoCo (Multi-Joint dynamics with Contact)","Benchmark"]},{"id":"robocerebra","category":"sim","sec":8,"tier":3,"sources":[{"title":"RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation (arXiv 2506.06677)","url":"https://arxiv.org/abs/2506.06677"}],"as_of":"2025-10","related_ids":["long-horizon-task","dual-system-architecture","embodied-memory","libero-benchmark","openvla","benchmark"],"name":"RoboCerebra (A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation)","alt":"RoboCerebra","abbr":"","aliases":[],"one_liner":"A large simulation benchmark testing planning, reflection, and memory in long-horizon robot manipulation.","explanation":"RoboCerebra was released in June 2025 by researchers at Beihang University, the National University of Singapore, Shanghai Jiao Tong University, and others, and was accepted to NeurIPS 2025. It focuses on “System 2”-style slow, deliberate reasoning: GPT is first used to generate long household tasks and break them into subtask sequences, which a human then carries out step by step inside simulation, yielding 100 task variants and 1,000 human-executed trajectories averaging about 2,972 simulation steps each — roughly 6 times longer, the paper says, than existing long-horizon manipulation datasets — with subtask time spans annotated. Test tasks fall into six categories: ideal conditions, random disturbance, observation inconsistency, memory exploration, memory execution, and mixed, with scenes changing mid-execution. A companion hierarchical framework uses a vision-language model (VLM) for high-level planning that writes to memory, and OpenVLA for low-level action execution; the paper uses this setup to compare how GPT-4o, Qwen2.5-VL, and other VLMs perform as planners on planning, reflection (judging whether a subtask is done), and memory.","example":"Given the instruction “prepare a drink, then tidy the table,” the high-level model must first break it into steps like fetching a cup, pouring the drink, and putting things away; if an object is randomly moved mid-execution, it also has to judge whether the current subtask is still complete and replan.","related":["Long-horizon Task","Dual-System Architecture (System 1 / System 2)","Embodied Memory","LIBERO Benchmark","OpenVLA","Benchmark"]},{"id":"mikasa-robo","category":"sim","sec":8,"tier":3,"sources":[{"title":"Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning (arXiv 2502.10550)","url":"https://arxiv.org/abs/2502.10550"},{"title":"MIKASA-Robo GitHub 仓库","url":"https://github.com/CognitiveAISystems/MIKASA-Robo"},{"title":"MIKASA-Robo-VLA Documentation","url":"https://mikasarobo.github.io/"}],"as_of":"2026-09","related_ids":["embodied-memory","memory-augmented-vla","partially-observable-markov-decision-process","maniskill","memory-augmented-vla","benchmark"],"name":"MIKASA-Robo","alt":"MIKASA-Robo","abbr":"","aliases":["MIKASA-Robo-VLA"],"one_liner":"A tabletop-manipulation benchmark specifically testing robot memory, with tasks requiring recall of information that's occluded or gone.","explanation":"MIKASA-Robo is a memory-intensive manipulation benchmark proposed by Cherepanov, Panov, and colleagues in February 2025, part of MIKASA (a suite for evaluating memory-intensive skills), with its paper published at ICLR 2026. Many real-world tasks are partially observable: an object gets blocked from view, critical information appears only briefly, and a policy that only looks at the current frame has no way to succeed. It's built on ManiSkill3, with the first version containing 32 tasks. It later expanded into MIKASA-Robo-VLA, aimed at VLA models: the task count grew to 90, covering 10 categories of memory, each task paired with a language instruction, with 22,500 trajectories released on Hugging Face, usable to test memory-equipped models such as MemoryVLA.","example":"In the ShellGameTouch task, the robot can see the red ball under one of three positions for the first 5 steps; all three are then covered with cups, and the robot must touch the cup hiding the ball.","related":["Embodied Memory","Memory-Augmented VLA","Partially Observable Markov Decision Process","ManiSkill","Memory-Augmented VLA","Benchmark"]},{"id":"vla-arena-an-open-source-framework-for-benchmarking-vision-l","category":"sim","sec":8,"tier":3,"sources":[{"title":"VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models (arXiv 2512.22539)","url":"https://arxiv.org/abs/2512.22539"},{"title":"VLA-Arena 项目主页","url":"https://vla-arena.github.io"}],"as_of":"2026-08","related_ids":["vision-language-action-model","libero-benchmark","robosuite","generalization-robustness-evaluation","embodied-safety","benchmark"],"name":"VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models","alt":"VLA-Arena","abbr":"","aliases":["VLA-Arena Benchmark"],"one_liner":"Peking University's open-source VLA evaluation framework, grading capability boundaries across task, language, and vision axes.","explanation":"VLA-Arena is an open-source VLA (vision-language-action model) evaluation framework released by Peking University's PKU-Alignment team, posted to arXiv in December 2025 and later accepted at ICML 2026, with simulation built on LIBERO and robosuite. It breaks difficulty into three independent axes: task structure, language instructions, and visual observations. There are 11 task suites totaling 170 tasks, split into four categories — safety, distractors, extrapolation, and long-horizon — with each suite graded L0–L2, where fine-tuning is only allowed on L0; language (W0–W4) and vision (V0–V4) perturbations can then be layered onto any task. Using this setup, the authors find that current VLAs tend to memorize their training tasks, understand vision only shallowly, and often disregard safety constraints.","example":"The paper found that a model ranking near the top on L0 could be overtaken by other models once moved to L1 or L2; the authors treat this kind of rank reversal as evidence that the three difficulty levels each carry independent information.","related":["Vision-Language-Action Model","LIBERO Benchmark","robosuite","Generalization / Robustness Evaluation","Embodied Safety","Benchmark"]},{"id":"nvidia-isaac-lab-arena","category":"sim","sec":8,"tier":3,"sources":[{"title":"NVIDIA Technical Blog: Simplify Generalist Robot Policy Evaluation in Simulation with NVIDIA Isaac Lab-Arena","url":"https://developer.nvidia.com/blog/simplify-generalist-robot-policy-evaluation-in-simulation-with-nvidia-isaac-lab-arena/"},{"title":"GitHub: isaac-sim/IsaacLab-Arena","url":"https://github.com/isaac-sim/IsaacLab-Arena"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","simulation-based-evaluation","benchmark","lerobot-envhub","robofinals","robocasa"],"name":"NVIDIA Isaac Lab-Arena","alt":"Isaac Lab-Arena","abbr":"","aliases":["Isaac Lab Arena","IsaacLab-Arena"],"one_liner":"An open-source NVIDIA extension to Isaac Lab for composing simulation benchmarks and evaluating robot policies in parallel.","explanation":"Isaac Lab-Arena is an open-source framework built on top of Isaac Lab, developed by NVIDIA together with the simulation-data company Lightwheel and formally introduced on NVIDIA's blog in January 2026; it is still in alpha. Large-scale evaluation used to mean hand-writing a scene and a script for every single task. Isaac Lab-Arena instead splits an environment into three freely combinable pieces: the scene (which objects are placed where), the embodiment (the robot along with its sensors and controllers), and the task (what needs to be accomplished). It supports perturbation testing of lighting, cameras, and object properties, tracks progress through subtasks such as grasping and placing, and can run thousands of environments simultaneously on a GPU. It plugs into LeRobot's EnvHub, so it can evaluate models such as GR00T and π0.5, and it already ships adapted versions of RoboCasa, LIBERO, and RoboTwin 2.0.","example":"In a comparison NVIDIA published on its blog, running the same batch of evaluations in parallel with Isaac Lab-Arena took 0.76 hours, versus 34.9 hours running them one at a time.","related":["NVIDIA Isaac Lab","Simulation-Based Evaluation","Benchmark","LeRobot EnvHub","RoboFinals (Lightwheel industrial-grade simulation evaluation platform)","RoboCasa"]},{"id":"robofinals","category":"sim","sec":8,"tier":3,"sources":[{"title":"Lightwheel Unveils RoboFinals（光轮官网，2025-12-04）","url":"https://lightwheel.ai/robofinals"},{"title":"RoboFinals-100: An Industrial Benchmark for Embodied AI（光轮官网）","url":"https://lightwheel.ai/media/robofinals-industrial-benchmark"},{"title":"新华网：光轮智能发布十万小时全模态人类行为开源数据集（2026-08-21）","url":"http://www.xinhuanet.com/sci-tech/20260821/57f5b57922604293a05e67b02510c9ec/c.html"}],"as_of":"2026-08","related_ids":["lightwheel","nvidia-isaac-lab-arena","simulation-based-evaluation","benchmark","simready-assets","vision-language-action-model"],"name":"RoboFinals (Lightwheel industrial-grade simulation evaluation platform)","alt":"光轮 RoboFinals 工业级仿真评测平台","abbr":"","aliases":["RoboFinals-100"],"one_liner":"An industrial-grade simulation evaluation platform from Lightwheel, built specifically to test VLA and other robot foundation models.","explanation":"RoboFinals was released on December 4, 2025 by Lightwheel, a simulation-data and evaluation company; it is positioned as a deliberately difficult, industrial-grade simulation evaluation platform built specifically to test frontier VLA (vision-language-action) models and other robot foundation models. Its core benchmark, RoboFinals-100, contains 100 tasks spanning household, factory, and retail scenes, involving rigid bodies, articulated objects like cabinet doors, and deformables such as cables, cloth, and liquids, and supports desktop arms, mobile manipulation, and whole-body manipulation robots. It is built on Isaac Lab-Arena, an evaluation framework Lightwheel co-developed with NVIDIA, with a choice of Newton, PhysX, MuJoCo, or Genesis as the physics backend so a policy's robustness across simulators can be checked, and it uses Real2Sim calibration to narrow the gap with real robots. Lightwheel says Alibaba's Qwen team helped build scenes and evaluation standards, and that teams including Fourier also use it. In August 2026, Lightwheel followed up with RoboFinals pilot-base simulation evaluation infrastructure.","example":"A team connects its own VLA to RoboFinals and, across a large batch of parallel simulated environments, runs tasks like home tidying, factory assembly, and supermarket restocking, comparing success rates by scene category to find weak spots.","related":["Lightwheel","NVIDIA Isaac Lab-Arena","Simulation-Based Evaluation","Benchmark","SimReady Assets","Vision-Language-Action Model"]},{"id":"robocasa","category":"sim","sec":9,"tier":2,"sources":[{"title":"RoboCasa 官网（RoboCasa365）","url":"https://robocasa.ai/"},{"title":"RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots (arXiv 2406.02523)","url":"https://arxiv.org/abs/2406.02523"},{"title":"robocasa/robocasa GitHub README","url":"https://github.com/robocasa/robocasa"}],"as_of":"2026-05","related_ids":["robosuite","mimicgen","simulation-assets","benchmark","household-tasks","nvidia-isaac-gr00t-n1"],"name":"RoboCasa","alt":"RoboCasa","abbr":"","aliases":["RoboCasa365"],"one_liner":"A large-scale kitchen-household simulation framework and benchmark from UT Austin; its newest version is called RoboCasa365.","explanation":"RoboCasa is a household-scene simulation framework from Yuke Zhu's group at UT Austin, built on top of robosuite and MuJoCo. Its first version, published at RSS 2024, included 120 kitchen scenes and 100 tasks, and used MimicGen (a tool that automatically expands a small number of human demonstrations into many more trajectories) to augment its data. RoboCasa365, released in February 2026 (ICLR 2026), expanded this to over 2,500 kitchen scenes, more than 3,200 objects, and 365 everyday tasks, with over 600 hours of human demonstrations and over 1,600 hours of automatically generated data, plus a multi-task leaderboard. It addresses the problem that real-robot household data is expensive and hard to reproduce, and is commonly used to compare generalist policies.","example":"RoboCasa365's official leaderboard compares the multi-task performance of policies such as Diffusion Policy, the π series, and GR00T on the same set of kitchen tasks.","related":["robosuite","MimicGen","Simulation Assets","Benchmark","Household Tasks","NVIDIA Isaac GR00T N1"]},{"id":"behavior-1k","category":"sim","sec":9,"tier":2,"sources":[{"title":"BEHAVIOR 官网","url":"https://behavior.stanford.edu/"},{"title":"BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation (arXiv 2403.09227)","url":"https://arxiv.org/abs/2403.09227"},{"title":"2026 BEHAVIOR Challenge","url":"https://behavior.stanford.edu/challenge/index.html"}],"as_of":"2026-09","related_ids":["omnigibson","nvidia-isaac-sim","household-tasks","long-horizon-task","benchmark","mobile-manipulation"],"name":"BEHAVIOR-1K (BEHAVIOR Challenge)","alt":"BEHAVIOR-1K","abbr":"","aliases":["BEHAVIOR","BEHAVIOR Challenge"],"one_liner":"Stanford's simulated benchmark of 1,000 everyday household tasks, built on the OmniGibson simulator.","explanation":"BEHAVIOR-1K is an embodied-AI benchmark from the Stanford Vision and Learning Lab, guided by Fei-Fei Li, Jiajun Wu, and others, with the first version published at CoRL 2022. Its tasks come from a survey asking what people would want a robot to help with around the house, distilled into 1,000 everyday activities spread across 50 interactive scenes (homes, gardens, restaurants, offices, and more), with more than 9,000 objects annotated with physical and semantic properties; task goals are defined in BDDL, a language that describes goal states with logical predicates. Its companion simulator, OmniGibson, is built on NVIDIA Isaac Sim and supports rigid bodies, soft bodies, and liquids. The team has run the BEHAVIOR Challenge since 2025; the second edition, in 2026, requires completing 100 full household tasks, provides 20,000 teleoperated demonstrations totaling 1,950 hours, and ranks entries by an average task success score that awards partial credit.","example":"The 2026 BEHAVIOR Challenge provides π0.5 and GR00T N1.7 as baselines; entrant policies control the robot using RGB, depth, and proprioceptive input to do household chores, with submissions due October 16, 2026, and a total prize pool of $11,000.","related":["OmniGibson","NVIDIA Isaac Sim","Household Tasks","Long-horizon Task","Benchmark","Mobile Manipulation"]},{"id":"open-vocabulary-mobile-manipulation","category":"sim","sec":9,"tier":2,"sources":[{"title":"HomeRobot: Open-Vocabulary Mobile Manipulation (arXiv 2306.11565)","url":"https://arxiv.org/abs/2306.11565"},{"title":"NeurIPS 2023 HomeRobot OVMM Challenge","url":"https://ovmm.github.io/"}],"as_of":"2023-12","related_ids":["mobile-manipulation","open-vocabulary","habitat","hello-robot-stretch","rearrangement","progress-score"],"name":"Open-Vocabulary Mobile Manipulation","alt":"开放词汇移动操作基准","abbr":"OVMM","aliases":["OVMM","HomeRobot","HomeRobot OVMM","OVMM Challenge"],"one_liner":"A benchmark that has a robot find any named object in an unfamiliar house and place it on a specified piece of furniture.","explanation":"HomeRobot OVMM is a benchmark released in 2023 by Meta FAIR, Georgia Tech, Carnegie Mellon, and Simon Fraser University, and it also served as a NeurIPS 2023 competition. Tasks take the form “move object X from furniture A to furniture B”: the object is specified in text and may belong to a category never seen during training (open vocabulary), and the robot must, in an unfamiliar house, find the object, pick it up, locate the target furniture, and place the object there. The simulated portion is built on Habitat and synthetic HSSD scenes, covering 60 multi-room homes with 2,535 objects across 129 categories; the real-robot portion uses the low-cost Hello Robot Stretch, paired with the open-source HomeRobot software stack. It evaluates navigation, perception, and grasping together as one pipeline, and the paper's baseline achieved a real-robot success rate of about 20%.","example":"Given the instruction “move the toy elephant from the chair to the table,” where the toy elephant belongs to a category never seen during training, the robot earns 1 point for each stage it completes — finding the object, picking it up, finding the table, and placing it down — and only counts as fully successful once all four stages are done.","related":["Mobile Manipulation","Open-vocabulary","Habitat","Hello Robot Stretch","Rearrangement","Progress Score"]},{"id":"alfred","category":"sim","sec":9,"tier":3,"sources":[{"title":"ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks (arXiv 1912.01734)","url":"https://arxiv.org/abs/1912.01734"},{"title":"ALFRED project site","url":"https://askforalfred.com/"}],"as_of":"","related_ids":["instruction-following","long-horizon-task","ai2-thor","alfworld","household-tasks","vision-and-language-navigation"],"name":"ALFRED","alt":"ALFRED","abbr":"","aliases":["Action Learning From Realistic Environments and Directives"],"one_liner":"A benchmark that has an agent complete multi-step household chores in a simulated home by following natural-language instructions.","explanation":"ALFRED was proposed by Mohit Shridhar and colleagues (University of Washington, Carnegie Mellon, the Allen Institute for AI, NVIDIA), published at CVPR 2020, and built on the AI2-THOR simulator. It has 120 indoor scenes (30 each of kitchens, bathrooms, bedrooms, and living rooms), 8,055 expert demonstrations, and 25,743 English instructions, with each demonstration averaging about 50 steps; tasks fall into 7 categories, such as pick-and-place, stack-and-place, heat, cool, clean-and-place, and examining an object under a light. The agent sees only a first-person view and the instructions, and outputs discrete navigation and interaction actions, providing a pixel mask for the target object whenever it interacts. Metrics are task success rate and goal-condition success rate, plus versions weighted by path length. The original paper's baseline scored under 1% success on unseen scenes, versus about 91% for humans. Both ALFWorld and TEACh are built on top of it.","example":"Given the high-level goal “rinse off the mug and put it in the coffee maker,” paired with step-by-step instructions like “walk to the coffee maker on the right,” the agent must navigate, pick up the mug, rinse it at the sink, and then place it in the coffee maker, in sequence.","related":["Instruction Following","Long-horizon Task","AI2-THOR","ALFWorld","Household Tasks","Vision-and-Language Navigation"]},{"id":"alfworld","category":"sim","sec":9,"tier":3,"sources":[{"title":"ALFWorld: Aligning Text and Embodied Environments for Interactive Learning (arXiv 2010.03768)","url":"https://arxiv.org/abs/2010.03768"},{"title":"ALFWorld project site","url":"https://alfworld.github.io/"},{"title":"ReAct: Synergizing Reasoning and Acting in Language Models (arXiv 2210.03629)","url":"https://arxiv.org/abs/2210.03629"}],"as_of":"","related_ids":["alfred","ai2-thor","llm-based-task-planning","planning-domain-definition-language","long-horizon-task","embodiedbench"],"name":"ALFWorld","alt":"ALFWorld","abbr":"","aliases":["Aligning Text and Embodied Environments for Interactive Learning"],"one_liner":"A text-game version of the ALFRED household tasks, letting an agent learn in text first and then act on the actual visuals.","explanation":"ALFWorld was proposed by Shridhar, Xingdi Yuan, and colleagues (University of Washington, Microsoft Research, Carnegie Mellon), published at ICLR 2021. It describes each ALFRED scene in PDDL (Planning Domain Definition Language) and uses Microsoft's TextWorld engine to generate an equivalent text-adventure game: the agent reads a text description of which furniture and items are in the room and acts through text commands like go to desk 1. Tasks fall into 6 categories, with 3,553 games used for training. The authors' BUTLER agent first learns a high-level policy in the text environment, then adds visual recognition (Mask R-CNN) and a low-level control module to transfer it to execution on AI2-THOR's actual visuals, training about 7 times faster than learning directly from visuals alone while also generalizing better. After large language models took off, agent frameworks such as ReAct adopted it as a standard test environment.","example":"For the task “examine the alarm clock under the desk lamp,” the text version is solved by entering go to desk 1, take alarmclock 2 from desk 1, and use desklamp 1 in sequence.","related":["ALFRED","AI2-THOR","LLM-based Task Planning","Planning Domain Definition Language","Long-horizon Task","EmbodiedBench"]},{"id":"room-to-room","category":"sim","sec":9,"tier":2,"sources":[{"title":"Room-to-Room (R2R) 官网","url":"https://bringmeaspoon.org/"},{"title":"VLN-CE 项目页","url":"https://jacobkrantz.github.io/vlnce/"},{"title":"jacobkrantz/VLN-CE GitHub（含 RxR-Habitat）","url":"https://github.com/jacobkrantz/VLN-CE"}],"as_of":"","related_ids":["vision-and-language-navigation","habitat","matterport3d","success-weighted-by-path-length","normalized-dynamic-time-warping","navila"],"name":"Room-to-Room","alt":"R2R / VLN-CE 视觉语言导航基准","abbr":"R2R / VLN-CE","aliases":["R2R","VLN-CE","RxR","R2R-CE","RxR-CE","VLN in Continuous Environments"],"one_liner":"The most widely used vision-language navigation benchmark: follow a human-written route description to reach a destination in an unfamiliar house.","explanation":"R2R (Room-to-Room) was released in 2018 by teams from the Australian National University, the University of Adelaide, and others, alongside the Matterport3D simulator. It is built on scans of 90 real buildings and contains about 22,000 human-written route instructions, with the agent restricted to jumping between nodes of a pre-built navigation graph. In 2020, Oregon State University, Georgia Tech, and Facebook AI proposed VLN-CE, which moved R2R into Habitat's continuous environments: the agent instead has to walk using low-level actions such as “move forward 0.25 meters,” “turn left 15 degrees,” and “stop,” which is much closer to a real robot and causes scores to drop noticeably. Google's RxR (2020) contains 126,000 instructions in English, Hindi, and Telugu. Common metrics include success rate, SPL (success weighted by path length), and nDTW (normalized dynamic time warping).","example":"The official VLN-CE cross-modal attention baseline scores 0.28 success rate and 0.25 SPL on the R2R continuous-environment test set, showing that the continuous-environment version is far harder than the navigation-graph version.","related":["Vision-and-Language Navigation","Habitat","Matterport3D","Success weighted by Path Length","normalized Dynamic Time Warping","NaVILA"]},{"id":"navigation-error-oracle-success-rate-trajectory-length","category":"sim","sec":9,"tier":3,"sources":[{"title":"Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments (R2R, CVPR 2018)","url":"https://arxiv.org/abs/1711.07280"},{"title":"Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (VLN-CE)","url":"https://arxiv.org/abs/2004.02857"}],"as_of":"","related_ids":["vision-and-language-navigation","room-to-room","success-rate","success-weighted-by-path-length","normalized-dynamic-time-warping"],"name":"Navigation Error / Oracle Success Rate / Trajectory Length","alt":"导航误差 / Oracle 成功率 / 轨迹长度","abbr":"NE / OSR / TL","aliases":["NE","OSR","OS","TL","Oracle Success"],"one_liner":"The three standard Vision-and-Language Navigation metrics: how close the agent stops to the goal, whether it ever passed it, and total distance traveled.","explanation":"These three metrics are a standard fixture in Vision-and-Language Navigation (VLN — following a natural-language instruction to reach a goal inside a house) results tables, used since Anderson and colleagues introduced the R2R benchmark at CVPR 2018. Navigation Error (NE) is the shortest-path distance, in meters, from where the agent finally stops to the goal; lower is better, and R2R counts an episode successful if NE is under 3 meters. Oracle Success Rate (OSR, sometimes written OS) instead asks whether the episode would count as successful if the agent had stopped at the point on its own trajectory closest to the goal — this separates “walked past the goal but didn't stop there” from “never got close at all.” Trajectory Length (TL) is the total distance actually walked; an unusually long TL signals wandering or backtracking. The three are normally reported together with Success Rate, SPL, and nDTW.","example":"In the original R2R paper, on a test set of previously unseen buildings, human annotators reached an 86.4% success rate while the authors' sequence-to-sequence baseline achieved only 20.4%.","related":["Vision-and-Language Navigation","Room-to-Room","Success Rate","Success weighted by Path Length","normalized Dynamic Time Warping"]},{"id":"success-weighted-by-path-length","category":"sim","sec":9,"tier":2,"sources":[{"title":"On Evaluation of Embodied Navigation Agents (arXiv 1807.06757)","url":"https://arxiv.org/abs/1807.06757"}],"as_of":"","related_ids":["success-rate","point-goal-navigation","object-goal-navigation","navigation","normalized-dynamic-time-warping","navigation-error-oracle-success-rate-trajectory-length"],"name":"Success weighted by Path Length","alt":"路径长度加权成功率","abbr":"SPL","aliases":["SPL","Success weighted by normalized inverse Path Length"],"one_liner":"A navigation metric that credits both reaching the goal and taking a path close to the shortest possible route.","explanation":"SPL was proposed by Peter Anderson, Jitendra Malik, and 9 other researchers in their 2018 paper “On Evaluation of Embodied Navigation Agents,” which recommended it as the primary metric for embodied navigation. It is computed as follows: for each test episode i, S_i is 1 if the episode succeeded and 0 otherwise; l_i is the length of the shortest path from the start to the goal; p_i is the length of the path the agent actually took; and SPL is the average, over all episodes, of S_i · l_i / max(p_i, l_i). “Success” requires the agent to actively output a stop action while close enough to the goal — the paper recommends a default threshold of twice the agent's body width — and distance is measured along the shortest path around obstacles rather than as a straight line. SPL penalizes taking a roundabout route, preventing an agent that wanders randomly and happens to reach the goal from getting a perfect score under success rate alone. Tasks such as point-goal and object-goal navigation typically report both success rate and SPL together.","example":"The paper gives worked examples: if half the episodes succeed and each follows the shortest path, SPL is 0.5; if all episodes succeed but each path is twice the shortest-path length, SPL is also 0.5; and if half succeed with paths twice as long, SPL is 0.25.","related":["Success Rate","Point-Goal Navigation","Object-Goal Navigation","Navigation","normalized Dynamic Time Warping","Navigation Error / Oracle Success Rate / Trajectory Length"]},{"id":"normalized-dynamic-time-warping","category":"sim","sec":9,"tier":3,"sources":[{"title":"General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping (arXiv 1907.05446)","url":"https://arxiv.org/abs/1907.05446"},{"title":"Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (VLN-CE)","url":"https://arxiv.org/abs/2004.02857"}],"as_of":"","related_ids":["navigation-error-oracle-success-rate-trajectory-length","success-weighted-by-path-length","vision-and-language-navigation","room-to-room","success-rate"],"name":"normalized Dynamic Time Warping","alt":"归一化动态时间规整","abbr":"nDTW","aliases":["nDTW","normalized DTW"],"one_liner":"A 0-to-1 navigation metric for how closely an agent's path matches the shape and order of a reference path.","explanation":"nDTW was proposed by Ilharco, Baldridge, and colleagues in 2019 to evaluate agents that navigate by following language instructions. Success Rate only checks the endpoint, and SPL only checks whether the path avoided needless detours — neither cares whether the agent actually followed the route the instruction described. nDTW borrows Dynamic Time Warping (DTW) from time-series analysis, which aligns two sequences of different lengths in order and sums the distance between aligned points, to compute the minimum cumulative distance between the agent's path and the reference path. That distance is then passed through an exponential function to normalize it to a 0–1 score, where higher is better: it penalizes deviation smoothly and stays sensitive to the order in which places are visited. The authors' human evaluation found it tracks human rankings better than other metrics. The same paper also proposes SDTW, which scores 0 on failed episodes and nDTW on successful ones. It is commonly reported alongside NE, SR, and SPL on R2R, R4R, and VLN-CE.","example":"An instruction says to enter the kitchen before going to the bedroom, but the agent walks straight to the bedroom — it reaches the right endpoint, so Success Rate still counts it as a success, but because the route skipped the kitchen, nDTW would score it noticeably lower.","related":["Navigation Error / Oracle Success Rate / Trajectory Length","Success weighted by Path Length","Vision-and-Language Navigation","Room-to-Room","Success Rate"]},{"id":"openeqa","category":"sim","sec":9,"tier":3,"sources":[{"title":"Meta AI Blog: OpenEQA: From word models to world models","url":"https://ai.meta.com/blog/openeqa-embodied-question-answering-robotics-ar-glasses/"},{"title":"OpenEQA project page","url":"https://open-eqa.github.io/"},{"title":"GitHub: facebookresearch/open-eqa (data)","url":"https://github.com/facebookresearch/open-eqa/tree/main/data"}],"as_of":"2024-04","related_ids":["embodied-question-answering","visual-question-answering","open-vocabulary","embodied-memory","habitat-matterport-3d-dataset","scannet"],"name":"OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark)","alt":"OpenEQA 开放词汇具身问答基准","abbr":"","aliases":["OpenEQA","EM-EQA","A-EQA"],"one_liner":"A Meta benchmark testing whether an agent can answer natural-language questions using what it has observed of a real environment.","explanation":"OpenEQA was released by Meta FAIR in April 2024, with the paper appearing at CVPR 2024. Embodied Question Answering (EQA) requires an agent to understand an environment it currently occupies or has previously seen, and answer questions about it in natural language — for example, “where did I leave my keys?” It contains more than 1,600 question-answer pairs, hand-written rather than generated from templates, drawn from videos and scans of over 180 real environments sourced from HM3D and ScanNet. It defines two settings: episodic memory (EM-EQA), where the agent answers from a previously recorded history of observations, modeling something like smart glasses; and active exploration (A-EQA), where the robot has to move around itself to gather the information it needs. Because answers are open-vocabulary free text, the authors score them automatically with an LLM-based metric called LLM-Match, which they show correlates highly with human judgment.","example":"In results Meta published, GPT-4V scored 48.5% accuracy versus 85.9% for humans; on questions that required spatial understanding, models that could see images performed barely better than models that could only read text.","related":["Embodied Question Answering","Visual Question Answering","Open-vocabulary","Embodied Memory","Habitat-Matterport 3D Dataset","ScanNet"]},{"id":"erqa","category":"sim","sec":9,"tier":3,"sources":[{"title":"embodiedreasoning/ERQA GitHub","url":"https://github.com/embodiedreasoning/ERQA"},{"title":"Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)","url":"https://arxiv.org/abs/2503.20020"}],"as_of":"2025-03","related_ids":["embodied-reasoning","gemini-robotics-er","visual-question-answering","multimodal-large-language-model","embodied-arena","vsi-bench"],"name":"ERQA","alt":"ERQA 具身推理问答基准","abbr":"ERQA","aliases":["Embodied Reasoning Question Answer Benchmark"],"one_liner":"A 400-question multiple-choice benchmark from Google DeepMind, pairing images and text to test embodied reasoning.","explanation":"ERQA is a benchmark Google DeepMind open-sourced in March 2025 alongside Gemini Robotics, used to measure a multimodal model's embodied reasoning: its ability to understand real physical scenes and make judgments for robot action. It has 400 multiple-choice questions (options A–D) that mix images and text, with scenes mostly drawn from real robot-relevant environments; question types include spatial reasoning, trajectory reasoning, action reasoning, state estimation, pointing, multi-view reasoning, and task reasoning, and 28% of questions require looking at multiple images. Because it's purely multiple-choice, it needs no robot or simulator, and any vision-language model can run it directly. In the Gemini Robotics technical report, Gemini 2.0 Pro Experimental scored 48.3%, GPT-4o scored 47.0%, and Claude 3.5 Sonnet scored 35.5%, showing that models at the time were still far from reliable embodied reasoning. Evaluation platforms such as Embodied Arena have also incorporated it.","example":"A typical question shows a photo of a robot arm holding a cup of water and asks, “to pour the water into the bowl next to it, which direction should the gripper rotate next?” with four options, A through D, to choose from.","related":["Embodied Reasoning","Gemini Robotics-ER","Visual Question Answering","Multimodal Large Language Model","Embodied Arena","VSI-Bench"]},{"id":"embodiedbench","category":"sim","sec":9,"tier":3,"sources":[{"title":"EmbodiedBench (arXiv 2502.09560)","url":"https://arxiv.org/abs/2502.09560"},{"title":"EmbodiedBench 项目主页","url":"https://embodiedbench.github.io/"}],"as_of":"2025-02","related_ids":["benchmark","multimodal-large-language-model","alfred","habitat","llm-based-task-planning","embodied-arena"],"name":"EmbodiedBench","alt":"EmbodiedBench","abbr":"","aliases":["Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents","EB-ALFRED","EB-Habitat","EB-Navigation","EB-Manipulation"],"one_liner":"A benchmark with 4 environments testing how well multimodal large models perform as the “brain” of an embodied agent.","explanation":"EmbodiedBench was proposed by Rui Yang, Tong Zhang, and colleagues at UIUC and other institutions, published at ICML 2025, specifically to evaluate how well multimodal large language models (MLLMs, models that can both look at images and read text) perform as the “brain” of an embodied agent. It has 4 environments totaling 1,128 test tasks, split into two levels: EB-ALFRED and EB-Habitat test high-level task decomposition and planning (for example, “put the book on the table,” where the model outputs a sequence of high-level skills); EB-Navigation and EB-Manipulation test low-level action planning (directly outputting control values such as translation and rotation), which demands precise perception and spatial reasoning. Tasks are also grouped by 6 capabilities: basic tasks, common-sense reasoning, complex-instruction understanding, spatial awareness, visual perception, and long-horizon planning. The authors tested 24 closed- and open-source models and found they were good at high-level tasks but struggled with low-level manipulation; the paper's abstract reports the best model, GPT-4o, averaging only 28.9%.","example":"In EB-Manipulation, a model sees a tabletop image and the instruction “stack the red block on the blue block” and must directly output the arm end effector's position, orientation, and gripper state, rather than calling a ready-made “grasp” skill.","related":["Benchmark","Multimodal Large Language Model","ALFRED","Habitat","LLM-based Task Planning","Embodied Arena"]},{"id":"embodied-arena","category":"sim","sec":9,"tier":3,"sources":[{"title":"Embodied Arena (arXiv 2509.15273)","url":"https://arxiv.org/abs/2509.15273"},{"title":"Embodied Arena 论文 HTML 版（作者单位与基准列表）","url":"https://arxiv.org/html/2509.15273v1"}],"as_of":"2025-09","related_ids":["benchmark","embodiedbench","erqa","openeqa","vsi-bench","embodied-question-answering"],"name":"Embodied Arena","alt":"Embodied Arena 具身评测竞技场","abbr":"","aliases":["A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI"],"one_liner":"An evaluation platform that plugs in benchmarks for embodied question answering, navigation, and task planning under one live leaderboard.","explanation":"Embodied Arena is an embodied-AI evaluation platform released jointly by over a dozen institutions in China and abroad, including Tianjin University and Huawei's Noah's Ark Lab, with its paper posted to arXiv in September 2025. It addresses the problem that embodied AI has many benchmarks but each operates in isolation, making models hard to compare directly and leaving unclear what capabilities embodied AI actually requires. The platform first defines a capability taxonomy with three layers — perception, reasoning, and task execution — broken into 7 core capabilities and 25 finer-grained dimensions. It then plugs 22 existing benchmarks into a unified evaluation backend, covering 2D/3D embodied question answering (such as OpenEQA, VSI-Bench, and ERQA), navigation (such as R2R-CE and HM3D), and task planning (such as EB-ALFRED and EB-Habitat), evaluating over 30 models from more than 20 institutions; it also uses an LLM-driven pipeline to automatically generate new evaluation data, keeping the question pool continuously updated. Results are published as three live leaderboards, viewable either by benchmark or by capability dimension.","example":"To check a multimodal model's spatial-reasoning ability, a user can simply look at that dimension's aggregate score in Embodied Arena's capability view, instead of running VSI-Bench, ERQA, and other benchmarks separately.","related":["Benchmark","EmbodiedBench","ERQA","OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark)","VSI-Bench","Embodied Question Answering"]},{"id":"arcade-learning-environment-atari-100k","category":"sim","sec":10,"tier":2,"sources":[{"title":"The Arcade Learning Environment: An Evaluation Platform for General Agents (Bellemare et al.)","url":"https://arxiv.org/abs/1207.4708"},{"title":"Arcade Learning Environment 文档（Farama Foundation）","url":"https://ale.farama.org/"},{"title":"Model-Based Reinforcement Learning for Atari (SimPLe, Kaiser et al., 2019)","url":"https://arxiv.org/abs/1903.00374"}],"as_of":"","related_ids":["reinforcement-learning","deep-q-network","sample-efficiency","world-model","dreamerv3","diamond"],"name":"Arcade Learning Environment (ALE) / Atari 100k","alt":"Atari 游戏基准（街机学习环境 / Atari 100k）","abbr":"ALE","aliases":["Arcade Learning Environment","Atari 100k","Atari 2600 benchmark"],"one_liner":"The standard platform for testing reinforcement learning on Atari 2600 games; Atari 100k is its low-sample variant.","explanation":"The Arcade Learning Environment (ALE) was proposed by Bellemare and colleagues, with the paper published in JAIR in 2013. Built on the Stella emulator, it wraps over a hundred Atari 2600 games behind a uniform interface: the agent sees screen pixels, chooses joystick actions, and receives the game score as reward. After DQN reached human-level performance on it in 2015, it became the most widely used benchmark in deep reinforcement learning, and is now maintained by the Farama Foundation. Atari 100k comes from the 2019 SimPLe paper: across 26 games, the agent is allowed only about 100,000 steps of interaction (roughly two hours of gameplay), specifically to test sample efficiency, and world-model methods such as DreamerV3 and DIAMOND are commonly compared on it. A great deal of the reinforcement-learning and world-model technology used in embodied AI today was first validated here.","example":"Under the Atari 100k setting, an agent playing Breakout gets only about two hours' worth of interaction data to learn from, and its score is then compared, normalized, against human performance.","related":["Reinforcement Learning","Deep Q-Network","Sample Efficiency","World Model","DreamerV3","DIAMOND"]},{"id":"minecraft-environments","category":"sim","sec":10,"tier":3,"sources":[{"title":"MineDojo 官网","url":"https://minedojo.org/"},{"title":"MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge (arXiv 2206.08853)","url":"https://arxiv.org/abs/2206.08853"},{"title":"MineRL: A Large-Scale Dataset of Minecraft Demonstrations (arXiv 1907.13440)","url":"https://arxiv.org/abs/1907.13440"}],"as_of":"2022-11","related_ids":["vpt","voyager","dreamerv3","open-world","long-horizon-task","reward-function"],"name":"Minecraft Environments","alt":"Minecraft 环境（MineDojo / MineRL）","abbr":"","aliases":["MineDojo","MineRL"],"one_liner":"Agent training and evaluation platforms built on the game Minecraft, used for open-world research.","explanation":"Minecraft is an open-world sandbox game with long task chains and open-ended goals that often require crafting items through several steps, which is why it's commonly used as a virtual testbed for agents. MineRL, released in 2019 by a team including Carnegie Mellon University, provides Gym-style environments and over 60 million human demonstration state-action pairs, and has run competitions at NeurIPS, with its classic task being mining a diamond from scratch. MineDojo, released in 2022 by teams from NVIDIA, Caltech, Stanford, and others, contains thousands of open-ended tasks described in language along with a knowledge base built from 730,000 YouTube videos, wiki pages, and Reddit posts, using a video-language model called MineCLIP as a reward function; it won a NeurIPS 2022 Outstanding Paper Award.","example":"Voyager has GPT-4 automatically write code and build up a skill library inside the MineDojo environment, continuously unlocking new tools and items.","related":["VPT","Voyager","DreamerV3","Open-world","Long-horizon Task","Reward Function"]},{"id":"gym-gymnasium-mujoco-tasks","category":"sim","sec":10,"tier":3,"sources":[{"title":"Gymnasium Documentation: MuJoCo environments","url":"https://gymnasium.farama.org/environments/mujoco/"},{"title":"Gymnasium Documentation: Half Cheetah","url":"https://gymnasium.farama.org/environments/mujoco/half_cheetah/"}],"as_of":"2026-09","related_ids":["mujoco","gymnasium","d4rl","soft-actor-critic","proximal-policy-optimization","deepmind-control-suite"],"name":"Gym/Gymnasium MuJoCo Tasks","alt":"Gym MuJoCo 连续控制任务","abbr":"","aliases":["Gymnasium MuJoCo Environments","MuJoCo Locomotion Tasks","HalfCheetah / Hopper / Walker2d"],"one_liner":"A classic set of MuJoCo-based continuous-control environments in Gymnasium, a standard test bed for reinforcement-learning papers.","explanation":"This is a set of 11 MuJoCo physics-simulation environments, originally released with OpenAI Gym and now maintained by the Farama Foundation's Gymnasium; the most commonly used are HalfCheetah (a 2D running cheetah), Hopper (a single-leg hopper), Walker2d (a 2D biped walker), Ant (a quadruped), and Humanoid. Actions are continuous joint torques, the reward is roughly “move forward as fast as possible, minus energy spent on actions,” and each episode runs at most 1,000 steps. Reinforcement-learning algorithms such as SAC, TD3, and PPO all report scores on them, and several of the D4RL offline datasets are also built on top of them. They test only locomotion control, with no vision or real-robot detail involved, making them well suited for learning the ropes and comparing algorithms. Version v5 is currently recommended (requires mujoco≥2.3.3), and scores from different versions cannot be compared directly.","example":"HalfCheetah-v5: the action is 6 joint torques (a 6-dimensional vector), the observation is 17-dimensional joint positions and velocities, there's no fall-based termination, the episode truncates at 1,000 steps, and the score equals a forward-velocity reward minus a control cost.","related":["MuJoCo (Multi-Joint dynamics with Contact)","Gymnasium","D4RL","Soft Actor-Critic","Proximal Policy Optimization","DeepMind Control Suite"]},{"id":"deepmind-control-suite","category":"sim","sec":10,"tier":3,"sources":[{"title":"DeepMind Control Suite (arXiv 1801.00690)","url":"https://arxiv.org/abs/1801.00690"},{"title":"google-deepmind/dm_control GitHub","url":"https://github.com/google-deepmind/dm_control"}],"as_of":"","related_ids":["reinforcement-learning","mujoco","benchmark","dreamerv3","td-mpc2","sample-efficiency"],"name":"DeepMind Control Suite","alt":"DeepMind 控制套件","abbr":"DMC","aliases":["DMC","dm_control","DM Control"],"one_liner":"DeepMind's standard set of MuJoCo-based continuous-control reinforcement-learning tasks.","explanation":"The DeepMind Control Suite was released in 2018 by Yuval Tassa and colleagues at DeepMind, as part of the open-source dm_control library. It uses the MuJoCo physics engine to build a set of continuous-control tasks across domains such as the cart-pole, cheetah, planar biped walker, humanoid, quadruped, and finger, with tasks of varying difficulty within each domain. Every task shares a uniform structure and an interpretable reward: each step's reward falls between 0 and 1, and every episode runs a fixed 1,000 steps, so a perfect score is 1,000. Input can be either joint states or raw camera pixels alone, which is why it became a standard benchmark for sample efficiency and for “learning control from pixels,” with algorithms such as Dreamer, DrQ, and TD-MPC all reporting results on it. dm_control also ships with tools for editing MJCF models, building environments from composable parts, and multi-agent soccer.","example":"In the walker-walk task, a planar biped model learns to walk forward using only camera images as input, and different algorithms are compared on how much return (up to a ceiling of 1,000) they achieve within the same number of interaction steps.","related":["Reinforcement Learning","MuJoCo (Multi-Joint dynamics with Contact)","Benchmark","DreamerV3","TD-MPC2","Sample Efficiency"]},{"id":"d4rl","category":"sim","sec":10,"tier":3,"sources":[{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning (arXiv 2004.07219)","url":"https://arxiv.org/abs/2004.07219"},{"title":"Farama-Foundation/D4RL GitHub（已弃用，迁往 Minari）","url":"https://github.com/Farama-Foundation/D4RL"}],"as_of":"2026-09","related_ids":["offline-reinforcement-learning","benchmark","gym-gymnasium-mujoco-tasks","adroit","franka-kitchen","conservative-q-learning"],"name":"D4RL","alt":"D4RL 离线强化学习基准","abbr":"D4RL","aliases":["Datasets for Deep Data-Driven Reinforcement Learning"],"one_liner":"The most widely used standard datasets and benchmark for offline reinforcement learning.","explanation":"D4RL was released in 2020 by Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine at UC Berkeley and Google Brain. At the time, offline reinforcement learning (training a policy purely from pre-collected data, with no further interaction with the environment during training) lacked dedicated test data, making it hard to compare results across papers. D4RL designs its datasets around how the data was collected: trajectories from hand-coded controllers, human demonstrations, multi-task data, and data mixed from several different policies, covering Maze2D and AntMaze maze navigation, Gym-MuJoCo locomotion control, Adroit dexterous-hand tasks, Franka Kitchen manipulation, and even Flow traffic and CARLA driving. It provides normalized scores to make comparison across tasks easier. Offline RL algorithm papers such as CQL and IQL commonly report D4RL scores. The original repository is no longer maintained: its environments moved to Gymnasium-Robotics and its data moved to Minari.","example":"halfcheetah-medium-v2 consists of running data for the HalfCheetah collected from a policy trained to medium skill level; an offline algorithm can only learn from this fixed batch of data, and results are reported as a normalized score via get_normalized_score.","related":["Offline Reinforcement Learning","Benchmark","Gym/Gymnasium MuJoCo Tasks","Adroit","Franka Kitchen","Conservative Q-Learning"]},{"id":"ogbench-benchmarking-offline-goal-conditioned-rl","category":"sim","sec":10,"tier":3,"sources":[{"title":"OGBench: Benchmarking Offline Goal-Conditioned RL (ICLR 2025)","url":"https://arxiv.org/abs/2410.20092"},{"title":"OGBench project page","url":"https://seohong.me/projects/ogbench/"},{"title":"GitHub: seohongpark/ogbench","url":"https://github.com/seohongpark/ogbench"}],"as_of":"","related_ids":["goal-conditioned-reinforcement-learning","offline-reinforcement-learning","d4rl","benchmark","implicit-q-learning","mujoco"],"name":"OGBench: Benchmarking Offline Goal-Conditioned RL","alt":"OGBench 离线目标条件强化学习基准","abbr":"OGBench","aliases":["OGBench"],"one_liner":"A benchmark built specifically to test offline goal-conditioned reinforcement learning algorithms, with 8 environment types and 85 datasets.","explanation":"OGBench was proposed by Seohong Park, Sergey Levine, and colleagues, and published at ICLR 2025. Goal-conditioned reinforcement learning trains a policy to reach any specified goal state; offline means training uses only a fixed dataset, with no further interaction with the environment. The authors argued that on older benchmarks, different algorithms scored too close together to tell apart, so they designed 8 categories of environments — including maze navigation, block manipulation, and puzzles — with 85 datasets that separately test trajectory stitching (combining pieces of incomplete trajectories into a new route), long-horizon reasoning, image observations, and environment stochasticity, plus another 410 standard offline-RL tasks. It is built on MuJoCo and the Gymnasium interface, installs with pip, and ships reference implementations of six algorithms including GCBC, GCIQL, and HIQL.","example":"In the antmaze-large-navigate-v0 dataset, a four-legged ant robot must walk from a starting point to a goal location specified at evaluation time inside a large maze; most environments provide 5 evaluation goals each.","related":["Goal-Conditioned Reinforcement Learning","Offline Reinforcement Learning","D4RL","Benchmark","Implicit Q-Learning","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"legged-gym","category":"sim","sec":10,"tier":2,"sources":[{"title":"leggedrobotics/legged_gym (GitHub)","url":"https://github.com/leggedrobotics/legged_gym"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/abs/2109.11978"},{"title":"unitreerobotics/unitree_rl_gym (GitHub)","url":"https://github.com/unitreerobotics/unitree_rl_gym"}],"as_of":"2026-09","related_ids":["isaac-gym","rsl-rl","terrain-curriculum","massively-parallel-reinforcement-learning","unitree-rl-gym","nvidia-isaac-lab"],"name":"legged_gym","alt":"legged_gym","abbr":"","aliases":["Isaac Gym Environments for Legged Robots","Legged Gym"],"one_liner":"ETH Zürich's open-source reinforcement-learning training environment for legged robots, built on Isaac Gym.","explanation":"legged_gym is an open-source reinforcement-learning codebase for legged robots from ETH Zürich's Robotic Systems Lab (RSL, Nikita Rudin and colleagues), accompanying their CoRL 2021 paper “Learning to Walk in Minutes…”. It runs thousands of simulated robots on a single GPU via Isaac Gym and trains locomotion with PPO (Proximal Policy Optimization); in the paper, an ANYmal robot learned to walk on flat ground in under 4 minutes and on rough terrain in about 20 minutes. The codebase includes the pieces needed for a policy to survive the transfer to a real robot: actuator networks, friction and mass randomization, observation noise, random pushes, and a terrain curriculum that automatically raises difficulty as the policy improves; it ships with the ANYmal B/C, A1, and Cassie robots built in. In January 2024 the authors announced a move to Isaac Lab and said the original repository would only get limited maintenance, but many projects, including Unitree's unitree_rl_gym, still use it as their backbone.","example":"Unitree's open-source unitree_rl_gym follows the same structure as legged_gym: running python legged_gym/scripts/train.py --task=g1 trains a G1 humanoid to walk inside Isaac Gym, after which the policy is validated with sim-to-sim testing in MuJoCo before being deployed to the real robot.","related":["Isaac Gym","rsl_rl","Terrain Curriculum","Massively Parallel Reinforcement Learning","unitree_rl_gym (Unitree RL Gym)","NVIDIA Isaac Lab"]},{"id":"unitree-rl-gym","category":"sim","sec":10,"tier":3,"sources":[{"title":"unitreerobotics/unitree_rl_gym GitHub 仓库","url":"https://github.com/unitreerobotics/unitree_rl_gym"},{"title":"unitreerobotics/unitree_rl_lab GitHub 仓库","url":"https://github.com/unitreerobotics/unitree_rl_lab"}],"as_of":"2026-09","related_ids":["legged-gym","isaac-gym","rsl-rl","sim-to-sim-transfer","unitree-g1","mujoco"],"name":"unitree_rl_gym (Unitree RL Gym)","alt":"unitree_rl_gym","abbr":"","aliases":["Unitree RL Gym"],"one_liner":"Unitree's official open-source reinforcement-learning example framework, covering everything from simulated training to real-robot deployment.","explanation":"unitree_rl_gym is a reinforcement-learning example repository Unitree Robotics open-sourced on GitHub, built on ETH Zurich's legged_gym (a legged-robot training framework on top of Isaac Gym) and the rsl_rl algorithm library, supporting the Go2 quadruped as well as the H1, H1_2, and G1 humanoids. It splits the workflow into four steps: train in Isaac Gym (Train), replay and check the result in simulation (Play), move the policy into MuJoCo for sim-to-sim validation (Sim2Sim), and finally deploy to the real robot (Sim2Real). The repository can export either MLP or LSTM policy networks, and provides a C++ deployment example for the G1. Unitree separately maintains a unitree_rl_lab repository built on Isaac Lab (2.3 and above), which is a different project.","example":"A common path for newcomers: run the repo's train.py with --task=g1 to train a G1 walking policy, use play.py to replay and export the network, load it in MuJoCo to confirm it doesn't fall, then deploy it to the real robot.","related":["legged_gym","Isaac Gym","rsl_rl","Sim-to-Sim Transfer","Unitree G1","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"robot-lab","category":"sim","sec":10,"tier":3,"sources":[{"title":"GitHub - fan-ziqi/robot_lab","url":"https://github.com/fan-ziqi/robot_lab"},{"title":"robot_lab Releases","url":"https://github.com/fan-ziqi/robot_lab/releases"}],"as_of":"2026-09","related_ids":["nvidia-isaac-lab","legged-gym","unitree-rl-gym","rsl-rl","rl-based-locomotion-control","massively-parallel-reinforcement-learning"],"name":"robot_lab (RL extension library based on Isaac Lab)","alt":"robot_lab（Isaac Lab 强化学习扩展库）","abbr":"","aliases":[],"one_liner":"An open-source library extending Isaac Lab with ready-made reinforcement-learning training environments for many legged and humanoid robots.","explanation":"robot_lab is an open-source project (Apache 2.0 license) maintained by developer Ziqi Fan on GitHub, positioned as an extension library for Isaac Lab (NVIDIA's Isaac Sim-based robot-learning framework): rather than modifying Isaac Lab itself, it keeps robot models and training tasks in a separate directory so they can be developed independently. It ships with more than 25 robots, covering quadrupeds such as ANYmal D, Unitree Go2/A1/B2, and DEEP Robotics Lite3; wheeled-leg robots such as Go2W and B2W; and humanoids such as Unitree G1/H1, Fourier GR1, and Booster T1. Training mainly uses RSL-RL (ETH's open-source PPO implementation), with experimental extras such as AMP motion imitation and BeyondMimic motion tracking. As of September 2026, the latest release is v2.3.2, matching Isaac Lab 2.3. It suits newcomers who want to train locomotion policies for an off-the-shelf robot directly on Isaac Lab without writing an environment from scratch.","example":"After installing Isaac Lab, pick robot_lab's velocity-tracking task for the Unitree Go2, train a walking policy with RSL-RL across thousands of parallel environments, then export the policy for sim-to-sim validation or deployment to the real robot.","related":["NVIDIA Isaac Lab","legged_gym","unitree_rl_gym (Unitree RL Gym)","rsl_rl","RL-based Locomotion Control","Massively Parallel Reinforcement Learning"]},{"id":"humanoid-gym","category":"sim","sec":10,"tier":3,"sources":[{"title":"Humanoid-Gym (arXiv 2404.05695)","url":"https://arxiv.org/abs/2404.05695"},{"title":"roboterax/humanoid-gym (GitHub)","url":"https://github.com/roboterax/humanoid-gym"}],"as_of":"2024-04","related_ids":["legged-gym","isaac-gym","sim-to-sim-transfer","sim-to-real-transfer","robotera","rl-based-locomotion-control"],"name":"Humanoid-Gym","alt":"Humanoid-Gym","abbr":"","aliases":["Reinforcement Learning for Humanoid Robot with Zero-Shot Sim2Real Transfer"],"one_liner":"RobotEra's open-source reinforcement-learning framework for humanoid walking, built around zero-shot sim-to-real transfer.","explanation":"Humanoid-Gym is a reinforcement-learning framework open-sourced in 2024 by the RobotEra (星动纪元) team; the paper's authors are Xinyang Gu, Yen-Jen Wang, and Jianyu Chen (Tsinghua IIIS), and it appeared at the ICRA 2024 workshop on agile robotics. It trains humanoid walking policies (with PPO) at massive scale on NVIDIA Isaac Gym, with code borrowed from legged_gym and rsl_rl. Its distinguishing feature is an Isaac Gym-to-MuJoCo sim2sim pipeline: retesting a trained policy in a different physics engine can expose behavior that only worked because it exploited quirks of the single original simulator, improving the odds of success on the real robot. The authors validated zero-shot transfer to the real world on two humanoid robots, the 1.2-meter XBot-S and the 1.65-meter XBot-L. It's also commonly used as a starting point for getting other humanoid robots walking with reinforcement learning.","example":"A walking policy for XBot-L is first trained with PPO in Isaac Gym, then dropped into MuJoCo to run sim2sim checks for continued stability, and deployed to the real robot once it passes.","related":["legged_gym","Isaac Gym","Sim-to-Sim Transfer","Sim-to-Real Transfer","RobotEra","RL-based Locomotion Control"]},{"id":"protomotions","category":"sim","sec":10,"tier":3,"sources":[{"title":"NVlabs/ProtoMotions GitHub（ProtoMotions3）","url":"https://github.com/NVlabs/ProtoMotions"},{"title":"ProtoMotions 文档","url":"https://protomotions.github.io/"}],"as_of":"2026-09","related_ids":["motion-tracking","maskedmimic","perpetual-humanoid-control","amass","motion-retargeting","nvidia-isaac-lab"],"name":"ProtoMotions (NVIDIA GPU-accelerated humanoid simulation & learning framework)","alt":"ProtoMotions","abbr":"","aliases":["ProtoMotions3"],"one_liner":"NVIDIA's open-source framework for training digital humans and humanoid robots to move using GPU physics simulation.","explanation":"ProtoMotions is an open-source framework from NVIDIA Research (NVlabs); its GitHub repository was created in September 2024 and it is now on its third version, ProtoMotions3, released under the Apache 2.0 license, with authors including Chen Tessler and Xue Bin Peng. It modularizes the whole pipeline of “training a humanoid to move inside physics simulation”: the backend can be swapped between Isaac Gym, Isaac Lab, Newton, and MuJoCo, and it supports the SMPL human body model as well as robots like Unitree's H1_2 and G1. It comes with built-in motion retargeting (transferring human motion capture onto a robot's skeleton), motion tracking, MaskedMimic generative control, and terrain locomotion tasks. The README reports that training on the AMASS dataset (over 40 hours of human motion) takes about 12 hours on 4 A100 GPUs; a general-purpose tracking policy trained on roughly 142,000 BONES-SEED motions can be deployed zero-shot to a real G1 robot just by exporting an ONNX model. It also supports one-command simulator swapping for sim-to-sim testing.","example":"Retarget a human running-and-jumping motion from AMASS onto a Unitree G1 in one step, train a tracking policy in Isaac Gym, validate it in MuJoCo, then export an ONNX model and deploy it to the real robot.","related":["Motion Tracking","MaskedMimic","Perpetual Humanoid Control","AMASS (Archive of Motion Capture as Surface Shapes)","Motion Retargeting","NVIDIA Isaac Lab"]},{"id":"humanoidbench","category":"sim","sec":10,"tier":3,"sources":[{"title":"HumanoidBench (arXiv 2403.10506)","url":"https://arxiv.org/abs/2403.10506"},{"title":"HumanoidBench project page","url":"https://humanoid-bench.github.io/"},{"title":"Robotics: Science and Systems XX (2024) proceedings","url":"https://www.roboticsproceedings.org/rss20/p061.html"}],"as_of":"2024-07","related_ids":["humanoid-robot","unitree-h1","shadow-dexterous-hand","mujoco","whole-body-control","td-mpc2"],"name":"HumanoidBench","alt":"HumanoidBench","abbr":"","aliases":["Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation","humanoid-bench"],"one_liner":"A simulated humanoid benchmark from Berkeley and others, covering 27 whole-body locomotion and manipulation tasks.","explanation":"HumanoidBench was proposed by researchers at UC Berkeley and Yonsei University (Carmelo Sferrazza, Pieter Abbeel, and colleagues), published at RSS 2024. Humanoid robot hardware is expensive and fragile, which makes algorithm research hard to run at scale. The benchmark builds a Unitree H1 fitted with two Shadow dexterous hands in MuJoCo, with 27 whole-body control tasks: 12 locomotion tasks (walk, run, hurdle, crawl, climb stairs, and more) and 15 manipulation tasks (open a cabinet, open a door, push, shoot a basketball, insert, and more). The codebase also provides variants such as an H1 without hands and a Unitree G1, along with DreamerV3, TD-MPC2, SAC, and PPO baselines. The paper found that the strongest reinforcement-learning algorithms at the time performed poorly on most tasks, while methods that first learn low-level skills like walking and reaching, then add hierarchical control on top, performed better.","example":"The same two-handed H1 has to complete locomotion tasks such as walk, hurdle, and stair, as well as manipulation tasks such as cabinet, door, basketball, and insert, all trained and scored with the same algorithm.","related":["Humanoid Robot","Unitree H1","Shadow Dexterous Hand","MuJoCo (Multi-Joint dynamics with Contact)","Whole-Body Control","TD-MPC2"]},{"id":"locomujoco","category":"sim","sec":10,"tier":3,"sources":[{"title":"LocoMuJoCo: A Comprehensive Imitation Learning Benchmark for Locomotion (arXiv 2311.02496)","url":"https://arxiv.org/abs/2311.02496"},{"title":"loco-mujoco GitHub 仓库","url":"https://github.com/robfiras/loco-mujoco"}],"as_of":"2026-09","related_ids":["mujoco","mujoco-xla","imitation-learning","motion-retargeting","amass","adversarial-motion-priors"],"name":"LocoMuJoCo","alt":"LocoMuJoCo","abbr":"","aliases":["Imitation Learning Benchmark for Whole-Body Locomotion"],"one_liner":"A MuJoCo-based imitation-learning benchmark for whole-body locomotion, with humanoid, quadruped, and human musculoskeletal models.","explanation":"LocoMuJoCo is an open-source benchmark released in 2023 by Jan Peters's group at TU Darmstadt in Germany (Firas Al-Hafez and colleagues), specifically for evaluating imitation-learning algorithms aimed at locomotion. Earlier locomotion benchmarks were mostly simplified toy tasks; this one provides quadruped, humanoid, and human musculoskeletal models, paired with real motion-capture data, expert data, and suboptimal data. Later versions switched to JAX, adding MJX and MuJoCo Warp for parallel simulation; it currently contains 12 humanoid and 4 quadruped environments, retargeting over 22,000 motion-capture clips from AMASS, LAFAN1, and other sources onto various humanoids, along with baselines including PPO, GAIL, AMP, and DeepMimic, plus a domain-randomization interface.","example":"Using ImitationFactory to create a Unitree H1 environment, loading dance and walking clips from LAFAN1, and then training a policy to imitate those motions with the built-in AMP algorithm.","related":["MuJoCo (Multi-Joint dynamics with Contact)","MuJoCo XLA","Imitation Learning","Motion Retargeting","AMASS (Archive of Motion Capture as Surface Shapes)","Adversarial Motion Priors"]},{"id":"mean-per-joint-position-error","category":"sim","sec":10,"tier":3,"sources":[{"title":"Perpetual Humanoid Control for Real-time Simulated Avatars (PHC, arXiv 2305.06456)","url":"https://arxiv.org/abs/2305.06456"},{"title":"A simple yet effective baseline for 3d human pose estimation (arXiv 1705.03098)","url":"https://arxiv.org/abs/1705.03098"},{"title":"Human3.6M Dataset","url":"http://vision.imar.ro/human3.6m/description.php"}],"as_of":"","related_ids":["motion-tracking","human-pose-estimation","motion-retargeting","perpetual-humanoid-control","asap","amass"],"name":"Mean Per-Joint Position Error","alt":"平均关节位置误差","abbr":"MPJPE","aliases":["MPJPE"],"one_liner":"The average distance between predicted and ground-truth joint positions, measuring how accurate a pose estimate or motion tracking is.","explanation":"MPJPE was originally the standard metric for 3D human pose estimation, widely used on datasets such as Human3.6M: for every frame, the Euclidean distance between each joint's predicted 3D position and its ground truth is computed, then averaged across all joints and frames, usually in millimeters, with lower being better. There are three common variants: root-relative MPJPE, which first aligns the root joint (the pelvis) and so measures only local pose; global MPJPE (g-MPJPE), computed with no alignment at all, which also captures overall positional drift; and PA-MPJPE, computed after a rigid-body (Procrustes) alignment. In humanoid robotics, motion-tracking work such as PHC and ASAP uses it to measure how precisely a robot reproduces a reference motion in simulation or on the real robot, usually reported together with success rate, velocity error, and acceleration error.","example":"PHC reports both root-relative MPJPE and global MPJPE (in mm) on the AMASS test set, and counts tracking as failed whenever the average joint deviation from the reference motion exceeds 0.5 meters.","related":["Motion Tracking","Human Pose Estimation","Motion Retargeting","Perpetual Humanoid Control","ASAP","AMASS (Archive of Motion Capture as Surface Shapes)"]},{"id":"furniturebench","category":"sim","sec":11,"tier":3,"sources":[{"title":"FurnitureBench project page","url":"https://clvrai.github.io/furniture-bench/"},{"title":"FurnitureBench (arXiv 2305.12821)","url":"https://arxiv.org/abs/2305.12821"}],"as_of":"2023-07","related_ids":["robotic-assembly","long-horizon-task","imitation-learning","offline-reinforcement-learning","factory-industreal","isaac-gym"],"name":"FurnitureBench","alt":"FurnitureBench","abbr":"","aliases":["Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation","FurnitureSim"],"one_liner":"A reproducible real-world furniture-assembly benchmark from KAIST and others, testing long-horizon, precise robot manipulation.","explanation":"FurnitureBench was published at RSS 2023 by KAIST's CLVR Lab in Korea and UC Berkeley, as a reproducible real-world furniture-assembly benchmark. It uses a Franka Panda arm and Intel RealSense cameras, with the task of assembling 8 kinds of furniture — a lamp, a square table, a desk, a drawer, a cabinet, a round table, a stool, and a chair — requiring multi-step coordination of grasping, insertion, and tightening, making it a long-horizon, contact-rich manipulation task. To let different labs set up a consistent environment, the parts are 3D-printed, the hardware uses common off-the-shelf products, and detailed build instructions are included; the authors also provide about 219.6 hours and 5,100 teleoperated demonstrations, plus FurnitureSim, a simulated version built on Isaac Gym and Factory. Evaluation of offline reinforcement-learning and imitation-learning algorithms shows there's still a lot of room for improvement on tasks like these.","example":"Screwing one table leg into a tabletop requires roughly 540 degrees of rotation; limited by wrist rotation range, the robot has to perform at least 5 “rotate 90 degrees, release, and re-grasp” cycles, which is exactly what makes this benchmark hard.","related":["Robotic Assembly","Long-horizon Task","Imitation Learning","Offline Reinforcement Learning","Factory / IndustReal","Isaac Gym"]},{"id":"autoeval","category":"sim","sec":11,"tier":3,"sources":[{"title":"AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World (arXiv 2503.24278)","url":"https://arxiv.org/abs/2503.24278"},{"title":"AutoEval project page","url":"https://auto-eval.github.io/"},{"title":"GitHub: zhouzypaul/auto_eval","url":"https://github.com/zhouzypaul/auto_eval"}],"as_of":"2025-09","related_ids":["real-world-evaluation","success-detector","bridgedata-v2","openvla","roboarena","sim-to-real-correlation"],"name":"AutoEval","alt":"AutoEval","abbr":"","aliases":["Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World"],"one_liner":"Berkeley's round-the-clock, unattended real-robot evaluation system, which judges success automatically and resets the scene itself.","explanation":"AutoEval was proposed by Zhiyuan Zhou, Pranav Atreya, and colleagues in Sergey Levine's group at Berkeley (one author is also affiliated with NVIDIA), posted to arXiv in March 2025, with its code repository noting publication at CoRL 2025. It automates the two most labor-intensive parts of real-robot evaluation: a fine-tuned PaliGemma vision-language model serves as a success detector, answering questions like “is the drawer open?”; and OpenVLA fine-tuned on a few dozen teleoperated demonstrations (or a recorded trajectory played back) serves as a reset policy that restores the scene to its original state. Users submit their own policy server through a webpage, much like submitting a job to a compute cluster, and the system queues it up to run on a WidowX arm in Bridge-style scenes. The paper reports an average Pearson correlation of 0.942 with human evaluation, about 850 episodes run per station in 24 hours with only 3 human interventions needed, a reduction in human time of over 99%, and it has opened public scenes to the community.","example":"A researcher deploys their VLA policy as a publicly reachable server, submits an evaluation for the task “put the eggplant in the basket” on the AutoEval webpage, and the system automatically runs several dozen episodes before returning a success-rate report.","related":["Real-World Evaluation","Success Detector","BridgeData V2","OpenVLA","RoboArena","Sim-to-Real Correlation"]},{"id":"roboarena","category":"sim","sec":11,"tier":2,"sources":[{"title":"RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies (arXiv 2506.18123)","url":"https://arxiv.org/abs/2506.18123"},{"title":"RoboArena project page","url":"https://robo-arena.github.io/"}],"as_of":"2025-11","related_ids":["double-blind-pairwise-comparison","elo-rating","real-world-evaluation","droid","progress-score","policy-server"],"name":"RoboArena","alt":"RoboArena","abbr":"","aliases":["Robo Arena","RoboArena Distributed Real-Robot Evaluation"],"one_liner":"A generalist-robot-policy evaluation platform where multiple labs run blind head-to-head real-robot comparisons that get aggregated into a ranking.","explanation":"RoboArena is a real-robot evaluation framework released in June 2025 by Karl Pertsch, Chelsea Finn, Sergey Levine, and colleagues from 7 academic institutions. It borrows the “arena” approach used to evaluate large language models: an evaluator freely sets up a scene and gives a task in their own lab, the system sends two anonymous policies to each run once from the same starting condition, and the evaluator provides a preference, a 0–100 progress score, and a written rationale; a Bradley-Terry model extended to account for task difficulty then aggregates a large number of these pairwise comparisons into a ranking. Policies connect as remote inference servers, so the evaluating site only needs a robot and a network connection. The first round standardized on the DROID platform (a Franka arm), and the paper reports that its ranking was closer to the reference ranking obtained from exhaustive testing than traditional single-lab, fixed-task evaluation was.","example":"The first evaluation round compared π0-flow-DROID, π0-FAST-DROID, and 5 other DROID policies built on PaliGemma with different action representations, completing over 600 blind pairwise comparisons across the 7 participating institutions.","related":["Double-blind Pairwise Comparison","Elo Rating","Real-World Evaluation","DROID (Distributed Robot Interaction Dataset)","Progress Score","Policy Server (Remote Inference)"]},{"id":"robochallenge","category":"sim","sec":11,"tier":2,"sources":[{"title":"RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies (arXiv 2510.17950)","url":"https://arxiv.org/abs/2510.17950"},{"title":"RoboChallenge 论文 HTML 全文","url":"https://arxiv.org/html/2510.17950"}],"as_of":"2025-10","related_ids":["real-world-evaluation","success-rate","progress-score","evaluation-protocol","dexmal","roboarena"],"name":"RoboChallenge","alt":"RoboChallenge","abbr":"","aliases":["Table30"],"one_liner":"A real-robot online evaluation platform from Dexmal and Hugging Face; its first benchmark is called Table30.","explanation":"RoboChallenge is a real-robot evaluation system launched in October 2025 by Dexmal (原力灵机) together with Hugging Face. Participants don't have to hand over model weights; instead, they run their model on their own machine and remotely control a real robot on the platform through an API. Its first benchmark, Table30, contains 30 tabletop manipulation tasks that probe difficulties such as precise 3D localization, occlusion, and dependence on earlier steps, using four robot types: UR5, Franka Panda, a Cobot Magic ALOHA dual-arm setup, and ARX-5. Each task is rolled out 10 times, reporting success rate and a stage-based progress score. It targets the problem that simulation scores don't equal real-robot capability, and that different labs' real-robot setups aren't otherwise comparable.","example":"The RoboChallenge paper used Table30 to compare π0, π0.5, CogACT, and variants of the OpenVLA series, under both single-task training and generalist training setups.","related":["Real-World Evaluation","Success Rate","Progress Score","Evaluation Protocol","Dexmal","RoboArena"]},{"id":"robodojo","category":"sim","sec":11,"tier":3,"sources":[{"title":"RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies (arXiv 2607.04434)","url":"https://arxiv.org/abs/2607.04434"},{"title":"RoboDojo-Benchmark/RoboDojo GitHub","url":"https://github.com/RoboDojo-Benchmark/RoboDojo"}],"as_of":"2026-09","related_ids":["robotwin","benchmark","simulation-based-evaluation","real-world-evaluation","robochallenge","pi0-5"],"name":"RoboDojo (A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies)","alt":"RoboDojo","abbr":"","aliases":["RoboDojo-RealEval"],"one_liner":"A unified sim-and-real benchmark for general-purpose manipulation policies, with 42 simulated tasks and 18 real-robot tasks.","explanation":"RoboDojo was released and open-sourced in July 2026 by the University of Hong Kong's MMLab together with 18 institutions including Berkeley, Tsinghua, Peking University, and Stanford, with Tianxing Chen as first author. It targets the problem that existing benchmark tasks tend to be too easy, get their scores saturated quickly, and lack real-robot validation. The simulation half is built on Isaac Sim / Isaac Lab, with 42 tasks on an ARX X5 dual-arm platform covering generalization, memory, precision, long-horizon execution, and open-ended instructions, with different tasks able to run in parallel across heterogeneous setups; the real-robot half sets 6 tasks each on three platforms — ARX X5, AgileX Piper, and Piper X — using standardized hardware and a remote cloud evaluation system called RoboDojo-RealEval to ensure reproducibility. Individual policies all connect through a shared interface called XPolicyLab; the first batch evaluated 30 policies, where the best in simulation, Hy-Embodied-0.5-VLA, averaged only an 8.80% success rate, and the best on real robots, π0.5, reached 12.8%, against human-expert scores of 76.03% and 100% respectively. The leaderboard is maintained and continuously updated by a nonprofit organization.","example":"After connecting a self-trained VLA through the XPolicyLab interface, one command runs it through all 42 tasks in Isaac Sim and produces a results table in the same format as the leaderboard.","related":["RoboTwin","Benchmark","Simulation-Based Evaluation","Real-World Evaluation","RoboChallenge","π0.5"]},{"id":"robotarena-infinity","category":"sim","sec":11,"tier":3,"sources":[{"title":"arXiv 2510.23571 - RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation","url":"https://arxiv.org/abs/2510.23571"},{"title":"RobotArena ∞ 论文 HTML 版","url":"https://arxiv.org/html/2510.23571"}],"as_of":"2026-03","related_ids":["simulation-based-evaluation","real-to-sim","double-blind-pairwise-comparison","elo-rating","digital-twin","roboarena"],"name":"RobotArena Infinity (RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation)","alt":"RobotArena ∞","abbr":"","aliases":["RobotArena Infinity","RobotArena ∞"],"one_liner":"A benchmark that auto-converts real robot videos into simulated scenes, then ranks VLA models by scoring and human voting.","explanation":"RobotArena ∞ is a robot-policy evaluation framework released in October 2025 by Katerina Fragkiadaki's group at Carnegie Mellon University. Real-robot evaluation is labor-intensive, slow, unsafe, and hard to reproduce, so this framework instead takes real manipulation videos from public datasets such as BridgeData V2, DROID, and RH20T and, using object segmentation, single-image-to-3D generation, and background inpainting models together with camera calibration and system identification, automatically converts them into digital-twin scenes inside the Genesis simulator, where different VLA models can then execute. Scoring runs two ways: a vision-language model gives a per-frame task-progress score, and crowd reviewers watch two execution videos double-blind and vote for the better one, with results aggregated into an Elo-style ranking using a Bradley-Terry model. It also systematically swaps backgrounds, changes colors, and moves objects to test robustness. The paper evaluated 6 policies — Octo, CogACT, π0, and X-VLA among them — and collected more than 8,500 human preference comparisons.","example":"Convert a real tabletop-manipulation video from DROID into a Genesis scene, let π0 and Octo each execute the same instruction inside it, have crowd reviewers watch both replays double-blind and vote for the better one to feed the leaderboard, then swap the background and rerun to see if the ranking changes.","related":["Simulation-Based Evaluation","Real-to-Sim","Double-blind Pairwise Comparison","Elo Rating","Digital Twin","RoboArena"]},{"id":"neural-simulator","category":"sim","sec":11,"tier":2,"sources":[{"title":"Learning to Simulate Complex Physics with Graph Networks (arXiv 2002.09405)","url":"https://arxiv.org/abs/2002.09405"},{"title":"Learning Interactive Real-World Simulators (UniSim, arXiv 2310.06114)","url":"https://arxiv.org/abs/2310.06114"},{"title":"1X World Model","url":"https://www.1x.tech/discover/1x-world-model"}],"as_of":"","related_ids":["world-model","unisim","1x-world-model","world-model-based-policy-evaluation","physics-engine","learning-in-imagination"],"name":"Neural Simulator","alt":"神经模拟器","abbr":"","aliases":["Learned Simulator","Neural Simulation","World-Model Simulator"],"one_liner":"A simulator learned from data with a neural network, predicting what happens next instead of relying on hand-written physics equations.","explanation":"A neural simulator is a simulator that learns from data, rather than from hand-written physics equations, how a scene changes given the current state and an action. One line of work learns physical quantities directly, such as DeepMind's Graph Network Simulator (GNS, ICML 2020), which represents fluids, rigid bodies, and deformable materials as particles and predicts their motion. A second line learns directly on pixels, which is essentially an action-conditioned video world model: UniSim (2023) learned an interactive simulator from internet and robot data, and policies trained inside it transferred zero-shot to a real robot; in September 2024, 1X trained a world model on thousands of hours of EVE robot data to evaluate policies. Neural simulators can absorb real-world complexity that is hard to hand-model, such as cloth, but objects can deform or vanish inconsistently, physics isn't guaranteed to be conserved, and errors accumulate over long rollouts. They are commonly used for policy evaluation, data generation, and learning in imagination.","example":"1X's world model takes a robot's starting frame plus a sequence of actions and generates the resulting video, which is then used to compare different policy versions on how likely each is to complete tasks like folding laundry or opening a door.","related":["World Model","UniSim","1X World Model","World-Model-based Policy Evaluation","Physics Engine","Learning in Imagination"]},{"id":"world-model-based-policy-evaluation","category":"sim","sec":11,"tier":2,"sources":[{"title":"Evaluating Gemini Robotics Policies in a Veo World Simulator (arXiv 2512.10675)","url":"https://arxiv.org/abs/2512.10675"},{"title":"WorldEval: World Model as Real-World Robot Policies Evaluator (arXiv 2505.19017)","url":"https://arxiv.org/abs/2505.19017"}],"as_of":"2026-07","related_ids":["world-model","veo-world-simulator","worldeval-world-model-as-real-world-robot-policies-evaluator","ctrl-world","sim-to-real-correlation","real-world-evaluation"],"name":"World-Model-based Policy Evaluation","alt":"世界模型评测","abbr":"","aliases":["World Model as Policy Evaluator","World-Model Evaluation"],"one_liner":"Using an interactive video world model in place of a real robot, running the policy in closed loop to estimate its performance.","explanation":"World-model-based evaluation connects a robot policy to an action-conditioned video generation model in a closed loop: the policy looks at the generated frame and outputs an action, the world model predicts the next frame based on that action, and the cycle repeats, with a human or a vision-language model finally judging whether the task was completed. It targets the problem that real-robot evaluation is slow, expensive, and hard to reproduce, while traditional simulators are laborious to set up and still fall short on visual and physical realism. Work such as WorldEval, WorldGym, and Ctrl-World appeared starting in 2025; Google DeepMind used its Veo video model to evaluate Gemini Robotics policies and cross-checked the results against over 1,600 real-robot evaluations, showing it can predict the relative ranking of different policies and can also be used for out-of-distribution generalization and safety red-teaming. New work has continued to appear into 2026, and the approach is currently used mainly for ranking checkpoints and screening for risk.","example":"Google DeepMind swapped in new objects, new backgrounds, and distractors inside the Veo world simulator to compare 8 Gemini Robotics policy checkpoints across 5 bimanual tasks, then checked the ranking against real-robot results.","related":["World Model","Veo World Simulator","WorldEval: World Model as Real-World Robot Policies Evaluator","Ctrl-World","Sim-to-Real Correlation","Real-World Evaluation"]},{"id":"worldeval-world-model-as-real-world-robot-policies-evaluator","category":"sim","sec":11,"tier":3,"sources":[{"title":"WorldEval: World Model as Real-World Robot Policies Evaluator (arXiv 2505.19017)","url":"https://arxiv.org/abs/2505.19017"},{"title":"WorldEval 项目主页","url":"https://worldeval.github.io"},{"title":"dWorldEval (arXiv 2604.22152)","url":"https://arxiv.org/abs/2604.22152"}],"as_of":"2026-04","related_ids":["world-model-based-policy-evaluation","world-model","video-generation-model","latent-action","sim-to-real-correlation","ctrl-world"],"name":"WorldEval: World Model as Real-World Robot Policies Evaluator","alt":"WorldEval","abbr":"","aliases":["dWorldEval"],"one_liner":"Using a video world model instead of a real robot to run rollouts and score and rank robot policies.","explanation":"WorldEval is a robot-policy evaluation method proposed in May 2025 by researchers at Midea Group and East China Normal University, which substitutes a world model for the real robot during evaluation. Testing policies one by one on real hardware is slow and hard to reproduce. WorldEval instead runs a policy in closed loop inside a world model: the policy looks at generated frames and outputs actions, and the world model generates the next stretch of video based on those actions. To make the video strictly follow the actions, the authors propose Policy2Vec, which conditions a video-generation model on latent actions so the frames it generates track the actions given. Experiments show its rankings correlate strongly with real-robot results, and it can also distinguish between checkpoints of the same policy and flag dangerous actions. A May 2026 follow-up, dWorldEval, switched to a discrete diffusion world model and added a “progress token” that automatically judges whether a task is complete.","example":"Before deploying to a real robot, run several candidate policies — or several checkpoints of the same policy — through WorldEval to rank them first, and only take the top-ranked ones to real-robot validation.","related":["World-Model-based Policy Evaluation","World Model","Video Generation Model","Latent Action","Sim-to-Real Correlation","Ctrl-World"]},{"id":"frechet-inception-distance","category":"sim","sec":11,"tier":2,"sources":[{"title":"GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (arXiv 1706.08500)","url":"https://arxiv.org/abs/1706.08500"},{"title":"Fréchet inception distance - Wikipedia","url":"https://en.wikipedia.org/wiki/Fr%C3%A9chet_inception_distance"}],"as_of":"","related_ids":["frechet-video-distance","peak-signal-to-noise-ratio-structural-similarity-index-learn","video-generation-model","world-model-based-policy-evaluation","generative-adversarial-network","generative-model"],"name":"Fréchet Inception Distance","alt":"弗雷歇初始距离","abbr":"FID","aliases":["FID","FID Score"],"one_liner":"A metric for how close a batch of generated images is to real images in overall distribution; lower is better.","explanation":"FID was introduced by Heusel and colleagues in a NeurIPS 2017 paper to evaluate the image quality of GANs and other generative models. It works by feeding a batch of real images and a batch of generated images through a pretrained Inception v3 network (an image classifier), taking the 2048-dimensional features from its final pooling layer, fitting each set to a multivariate Gaussian distribution, and computing the Fréchet distance between the two Gaussians. Because it compares overall distributions rather than pixel by pixel, it captures both image quality and diversity at once; lower scores are better, and 0 means the two sets are statistically identical. Its limitations are that it depends on Inception's learned features, doesn't always track human judgment, and is biased when the sample size is small. The video version is called FVD. In embodied AI, FID, FVD, and PSNR are often reported together when evaluating the generated frames of world models and video-generation models.","example":"To evaluate a robot video world model, researchers take a few thousand frames of its generated manipulation footage and a few thousand real frames from the same scene, run both through Inception v3, and compute FID; a lower score means the generated footage looks more like real data overall.","related":["Fréchet Video Distance","Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch Similarity","Video Generation Model","World-Model-based Policy Evaluation","Generative Adversarial Network","Generative Model"]},{"id":"frechet-video-distance","category":"sim","sec":11,"tier":3,"sources":[{"title":"Towards Accurate Generative Models of Video: A New Metric & Challenges (arXiv 1812.01717)","url":"https://arxiv.org/abs/1812.01717"}],"as_of":"","related_ids":["frechet-inception-distance","video-generation-model","world-model","ewmbench","vbench-comprehensive-benchmark-suite-for-video-generative-mo","peak-signal-to-noise-ratio-structural-similarity-index-learn"],"name":"Fréchet Video Distance","alt":"弗雷歇视频距离","abbr":"FVD","aliases":["FVD"],"one_liner":"A metric for how close a batch of generated videos is to real videos overall; lower is better.","explanation":"FVD was proposed in late 2018 by researchers at Johannes Kepler University, IDSIA, and Google Brain, as the video counterpart to the image metric FID (Fréchet Inception Distance). It works by feeding both real and generated videos through an I3D network (a 3D convolutional network that captures both spatial appearance and motion over time) pretrained on the Kinetics action-recognition dataset to extract features, fitting each set to a multivariate Gaussian distribution, and computing the Fréchet distance between the two distributions. A lower value means the generated videos are closer to the real ones in both image quality and temporal coherence; the authors' human evaluation showed it correlates reasonably well with subjective human judgment. It's one of the most commonly used metrics in video generation and world model papers, but because it compares overall distributions, it doesn't check whether the physics or robot actions in any individual video are correct, which is why embodied settings often pair it with dedicated benchmarks such as EWMBench.","example":"To evaluate a robot video world model, a batch of real manipulation videos and videos the model generates from the same starting frames are both run through I3D to extract features; a smaller resulting FVD means the generated videos look more like real ones overall.","related":["Fréchet Inception Distance","Video Generation Model","World Model","EWMBench","VBench: Comprehensive Benchmark Suite for Video Generative Models","Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch Similarity"]},{"id":"peak-signal-to-noise-ratio-structural-similarity-index-learn","category":"sim","sec":11,"tier":3,"sources":[{"title":"The Unreasonable Effectiveness of Deep Features as a Perceptual Metric (LPIPS, CVPR 2018)","url":"https://arxiv.org/abs/1801.03924"},{"title":"Wikipedia: Structural similarity index measure","url":"https://en.wikipedia.org/wiki/Structural_similarity_index_measure"},{"title":"Wikipedia: Peak signal-to-noise ratio","url":"https://en.wikipedia.org/wiki/Peak_signal-to-noise_ratio"}],"as_of":"","related_ids":["novel-view-synthesis","3d-gaussian-splatting","world-model","frechet-video-distance","frechet-inception-distance","neural-radiance-fields"],"name":"Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch Similarity","alt":"PSNR / SSIM / LPIPS 图像相似度指标","abbr":"PSNR / SSIM / LPIPS","aliases":["PSNR","SSIM","LPIPS","Structural Similarity Index Measure"],"one_liner":"Three widely used metrics for how similar a generated image is to a reference image, each closer to human perception than the last.","explanation":"All three compare a generated or reconstructed image against a reference image, one pair at a time. PSNR (Peak Signal-to-Noise Ratio) converts mean squared error into decibels; typical values for 8-bit images fall between 30 and 50 dB, with higher being better, but it only looks at per-pixel differences and correlates poorly with what humans actually perceive. SSIM (Structural Similarity Index), published by Wang, Bovik, and colleagues in 2004, compares local image patches along luminance, contrast, and structure, with a maximum score of 1, higher is better. LPIPS, proposed by Zhang and colleagues at CVPR 2018, feeds both images through a pretrained deep network and compares their intermediate-layer features; lower is better, and it aligns with human perceptual judgment more closely than the other two. In embodied AI, all three are commonly used to evaluate novel-view synthesis, Gaussian-splatting reconstructions, and world-model predictions, but a high score does not guarantee the result is physically plausible.","example":"The original 3D Gaussian Splatting paper reports PSNR 27.21, SSIM 0.815, and LPIPS 0.214 on the Mip-NeRF360 dataset, matching the image quality of the then-best Mip-NeRF360 method.","related":["Novel View Synthesis","3D Gaussian Splatting","World Model","Fréchet Video Distance","Fréchet Inception Distance","Neural Radiance Fields"]},{"id":"vbench-comprehensive-benchmark-suite-for-video-generative-mo","category":"sim","sec":11,"tier":3,"sources":[{"title":"VBench: Comprehensive Benchmark Suite for Video Generative Models (arXiv 2311.17982)","url":"https://arxiv.org/abs/2311.17982"},{"title":"Vchitect/VBench GitHub 仓库","url":"https://github.com/Vchitect/VBench"}],"as_of":"2025-03","related_ids":["video-generation-model","world-model","frechet-video-distance","physics-iq","worldscore-a-unified-evaluation-benchmark-for-world-generati","ewmbench"],"name":"VBench: Comprehensive Benchmark Suite for Video Generative Models","alt":"VBench 视频生成评测基准","abbr":"","aliases":["VBench","VBench++","VBench-2.0"],"one_liner":"An open-source benchmark that scores video-generation quality separately across 16 dimensions instead of giving one overall number.","explanation":"VBench was proposed by a team from Nanyang Technological University and the Shanghai AI Laboratory, among others (first author Ziqi Huang), released in November 2023 and selected as a CVPR 2024 Highlight. Rather than a single vague score, it breaks text-to-video quality into 16 dimensions — subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, object class, and spatial relationships among them — each with dedicated prompts and an automated evaluation method, checked against human preference annotations to see whether the scoring agrees with human judgment. The follow-up VBench++ extended it to image-to-video, long-video, and trustworthiness evaluation; VBench-2.0, from March 2025, shifted toward “intrinsic faithfulness” — commonsense reasoning, physical realism, and human motion. In embodied AI, evaluating video world models often borrows some of its metrics.","example":"When evaluating a model meant to generate robot training videos, VBench's subject-consistency and motion-smoothness dimensions can check whether the robot arm deforms or jitters partway through the clip; checking whether the motion obeys physics still requires a dedicated benchmark such as Physics-IQ.","related":["Video Generation Model","World Model","Fréchet Video Distance","Physics-IQ (Do generative video models understand physical principles?)","WorldScore: A Unified Evaluation Benchmark for World Generation","EWMBench"]},{"id":"worldscore-a-unified-evaluation-benchmark-for-world-generati","category":"sim","sec":11,"tier":3,"sources":[{"title":"WorldScore (arXiv 2504.00983, ICCV 2025)","url":"https://arxiv.org/abs/2504.00983"},{"title":"WorldScore 项目主页","url":"https://haoyi-duan.github.io/WorldScore/"}],"as_of":"2025-11","related_ids":["world-model","video-generation-model","4d-world-model","marble","worldarena-a-unified-benchmark-for-evaluating-perception-and","vbench-comprehensive-benchmark-suite-for-video-generative-mo"],"name":"WorldScore: A Unified Evaluation Benchmark for World Generation","alt":"WorldScore 世界生成评测","abbr":"","aliases":["WorldScore"],"one_liner":"Stanford's unified world-generation benchmark that lets 3D and 4D scene-generation models and video models be compared on the same scale.","explanation":"WorldScore is a world-generation evaluation benchmark released in April 2025 by Fei-Fei Li and Jiajun Wu's group at Stanford, included in ICCV 2025. World generation means producing, from an image or text, a scene that can be explored along a camera path; up to that point, 3D scene generation, 4D scene generation, and video generation models were each evaluated separately, with no way to compare across them directly. WorldScore unifies the task into a chain of “generate the next scene” steps, with each step's layout specified by a camera trajectory, over a dataset of 3,000 examples. Metrics fall into three groups: controllability (camera, object, and content alignment), quality (3D consistency, photometric and style consistency, subjective quality), and dynamics (motion accuracy, magnitude, and smoothness), combined into two overall scores, WorldScore-Static and WorldScore-Dynamic.","example":"The paper used the same set of examples to evaluate 19 open- and closed-source models across four categories — 3D scene generation, 4D scene generation, image-to-video, and text-to-video — and maintains a public leaderboard.","related":["World Model","Video Generation Model","4D World Model","Marble (World Labs)","WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models","VBench: Comprehensive Benchmark Suite for Video Generative Models"]},{"id":"physics-iq","category":"sim","sec":11,"tier":3,"sources":[{"title":"Do generative video models understand physical principles? (arXiv 2501.09038)","url":"https://arxiv.org/abs/2501.09038"},{"title":"Physics-IQ benchmark GitHub（含 Original / Verified 排行榜）","url":"https://github.com/google-deepmind/physics-IQ-benchmark"},{"title":"Physics-IQ 项目主页","url":"https://physics-iq.github.io/"}],"as_of":"2026-09","related_ids":["intuitive-physics","video-generation-model","world-model","vbench-comprehensive-benchmark-suite-for-video-generative-mo","sora","worldscore-a-unified-evaluation-benchmark-for-world-generati"],"name":"Physics-IQ (Do generative video models understand physical principles?)","alt":"Physics-IQ 物理理解基准","abbr":"","aliases":["Physics-IQ Benchmark","Physics-IQ Verified"],"one_liner":"A benchmark of real filmed footage that tests whether video-generation models actually understand physics, not just look realistic.","explanation":"Physics-IQ was released in January 2025 by researchers at INSAIT (Sofia University, Bulgaria) and Google DeepMind, with the paper appearing at WACV 2026. Every clip is real footage: 66 physical scenarios, each filmed twice from 3 camera angles, for 396 eight-second videos spanning solid mechanics, fluids, optics, thermodynamics, and magnetism. At test time a model only sees the beginning (one frame for image-to-video models, three seconds for video-continuation models) and has to generate the next 5 seconds, which is then compared against the real footage using spatial IoU (intersection-over-union, checking whether things moved to the right place), spatiotemporal IoU, weighted spatial IoU, and mean squared error, combined into a Physics-IQ score; the difference between two real takes of the same scenario is calibrated to 100%. In the original paper Sora scored only 10%, leading to the conclusion that looking realistic is not the same as understanding physics. It is often used to test whether video models can serve as world models; an improved-prompt, improved-metric Physics-IQ Verified version followed in 2026, with a leaderboard updated continuously.","example":"Show a model the moment the first domino has just been tipped over but hasn't yet touched the second one, and ask it to generate the next 5 seconds; check whether the dominoes fall in order, matching the real filmed video pixel by pixel.","related":["Intuitive Physics","Video Generation Model","World Model","VBench: Comprehensive Benchmark Suite for Video Generative Models","Sora","WorldScore: A Unified Evaluation Benchmark for World Generation"]},{"id":"ewmbench","category":"sim","sec":11,"tier":3,"sources":[{"title":"EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models (arXiv 2505.09694)","url":"https://arxiv.org/abs/2505.09694"}],"as_of":"2025-05","related_ids":["world-model","video-generation-model","frechet-video-distance","vbench-comprehensive-benchmark-suite-for-video-generative-mo","worldarena-a-unified-benchmark-for-evaluating-perception-and","agibot-world"],"name":"EWMBench","alt":"EWMBench 具身世界模型评测","abbr":"","aliases":["Evaluating Scene, Motion, and Semantic Quality in Embodied World Models"],"one_liner":"A benchmark from AgiBot and others that specifically evaluates whether robot-manipulation video generation models get the details right.","explanation":"EWMBench is an evaluation benchmark released in May 2025 by AgiBot together with Shanghai Jiao Tong University, CUHK MMLab, and the Harbin Institute of Technology. It targets embodied world models: models that generate a video of robot manipulation given a starting frame and a task instruction. The authors argue that metrics like FVD (Fréchet Video Distance) and VBench mainly assess visual quality and can't tell whether the robot arm's motion is actually plausible, so they score along three dimensions instead: scene consistency (comparing inter-frame features with a DINOv2 model fine-tuned on embodied data), motion correctness (Hausdorff distance and normalized dynamic time warping between the end-effector trajectory and ground truth, plus velocity and acceleration distributions), and semantic alignment (having a multimodal large model write a description of the generated video, checking for logical errors, and comparing it against ground truth). Test data is drawn from 10 tasks in the AgiBot World dataset, and both the dataset and evaluation code are open-sourced on GitHub.","example":"The paper used EWMBench to compare 7 models — OpenSora 2.0, LTX, COSMOS-7B, Kling-1.6, Hailuo, EnerVerse, and others — concluding that current video generation models still have clear shortcomings when applied to embodied tasks.","related":["World Model","Video Generation Model","Fréchet Video Distance","VBench: Comprehensive Benchmark Suite for Video Generative Models","WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models","AgiBot World"]},{"id":"worldarena-a-unified-benchmark-for-evaluating-perception-and","category":"sim","sec":11,"tier":3,"sources":[{"title":"WorldArena (arXiv 2602.08971)","url":"https://arxiv.org/abs/2602.08971"},{"title":"WorldArena 2.0 (arXiv 2605.17912)","url":"https://arxiv.org/abs/2605.17912"}],"as_of":"2026-05","related_ids":["world-model","world-model-based-policy-evaluation","ewmbench","robotwin","ctrl-world","worldscore-a-unified-evaluation-benchmark-for-world-generati"],"name":"WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models","alt":"WorldArena 具身世界模型基准","abbr":"","aliases":["WorldArena","WorldArena 2.0"],"one_liner":"A Tsinghua-led embodied world-model benchmark scoring both how good the generated video looks and how useful it actually is.","explanation":"WorldArena is an embodied world-model evaluation benchmark led by Tsinghua University with several other institutions, released in February 2026. An embodied world model predicts future frames given the current frame plus an action or instruction. The authors argue existing evaluations only ask whether a video looks good, not whether it actually helps decision-making, so they score along three axes: video quality (6 sub-dimensions and 16 metrics in total), embodied-task utility (testing the world model as a data engine, a policy evaluator, and an action planner), and human evaluation, combined into an overall EWMScore. After evaluating 14 models on RoboTwin 2.0 dual-arm tasks, the authors found that good-looking video does not imply strong task performance. A 2.0 version from May 2026 added visuotactile input and real-robot platforms.","example":"The first version evaluated general video models such as Wan 2.2 and Veo 3.1, alongside embodied world models such as Cosmos-Predict 2.5, Genie Envisioner, and Ctrl-World, with a public leaderboard at world-arena.ai.","related":["World Model","World-Model-based Policy Evaluation","EWMBench","RoboTwin","Ctrl-World","WorldScore: A Unified Evaluation Benchmark for World Generation"]},{"id":"demonstration-data","category":"data","sec":0,"tier":1,"sources":[{"title":"ALOHA: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","url":"https://tonyzhaozh.github.io/aloha/"},{"title":"Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"LeRobot Docs: Imitation Learning on Real-World Robots","url":"https://huggingface.co/docs/lerobot/il_robots"}],"as_of":"","related_ids":["imitation-learning","behavior-cloning","teleoperation","observation-action-pair","episode","real-robot-data"],"name":"Demonstration Data","alt":"演示数据","abbr":"","aliases":["Demonstrations","Expert Demonstration","Demo Data"],"one_liner":"The recorded observations and actions from a human performing a task, used as raw material for robot imitation learning.","explanation":"Demonstration data is the raw material of imitation learning: a human performs a task once, through teleoperation, kinesthetic teaching (physically guiding the robot's arm), or a handheld capture device, and the system records, at a fixed rate, each frame's observation — camera images, joint angles, and other proprioceptive state — together with the matching action. One complete run from start to finish is stored as a trajectory, usually called an episode in a dataset; an episode recorded while the robot executes on its own does not count as a demonstration. A policy network learns exactly this mapping: given this observation, do this action. Quality drives results directly — whether the motion is consistent, the camera stays fixed, and the object stays in view all matter. A single task can reach a usable policy from a few dozen demonstrations, while general-purpose VLA (vision-language-action) models need pretraining on thousands of hours of demonstrations across many tasks and scenes.","example":"In the ALOHA paper, most bimanual tasks used just 50 demonstrations (roughly 10 to 20 minutes of data per task); trained with the ACT algorithm, the task of inserting a battery into a remote control reached a 96% success rate. The LeRobot tutorials recommend recording at least 50 demonstrations for a beginner grasping task.","related":["Imitation Learning","Behavior Cloning","Teleoperation","Observation-Action Pair","Episode","Real-Robot Data"]},{"id":"observation-action-pair","category":"data","sec":0,"tier":2,"sources":[{"title":"LeRobotDataset v3.0 (Hugging Face LeRobot docs)","url":"https://huggingface.co/docs/lerobot/lerobot-dataset-v3"},{"title":"google-research/rlds (GitHub)","url":"https://github.com/google-research/rlds"}],"as_of":"","related_ids":["observation","action-label","behavior-cloning","trajectory","demonstration-data","compounding-error"],"name":"Observation-Action Pair","alt":"观测-动作对","abbr":"","aliases":["State-Action Pair"],"one_liner":"One training sample: what the robot 'saw' at a given moment paired with what it 'did' next.","explanation":"This is the basic unit of imitation-learning data. At every timestep, an observation o (camera images, joint angles and other proprioceptive state, often with a language instruction attached) and the action a executed at that moment (target joint angles, end-effector displacement, gripper opening) are recorded together; a full demonstration trajectory is a time-ordered sequence of these (o, a) pairs. Behavior cloning treats this as supervised learning: given o, predict a. Strictly speaking, “state” is a complete description of the environment, while “observation” is only the part sensors can actually see, and on a real robot usually only observations are available. Because these samples are sequentially correlated, once a policy drifts even slightly it runs into observations it never saw in training, and the error compounds — this is called compounding error. In a LeRobot dataset, each frame's observation.state, observation.images, and action fields are exactly this structure.","example":"Recording a demonstration of an SO-101 arm grasping a block at 30 frames per second in LeRobot, every frame stores the camera image, the follower arm's joint readings, and the leader arm's target joint angles at that moment as the action.","related":["Observation","Action Label","Behavior Cloning","Trajectory","Demonstration Data","Compounding Error"]},{"id":"action-label","category":"data","sec":0,"tier":2,"sources":[{"title":"Baker et al. 2022: Video PreTraining (VPT) (arXiv 2206.11795)","url":"https://arxiv.org/abs/2206.11795"},{"title":"Open X-Embodiment 项目页","url":"https://robotics-transformer-x.github.io/"},{"title":"Ye et al. 2024: Latent Action Pretraining from Videos (LAPA)","url":"https://arxiv.org/abs/2410.11758"}],"as_of":"","related_ids":["observation-action-pair","behavior-cloning","action-space","pseudo-action-labels","inverse-dynamics-model","action-free-video"],"name":"Action Label","alt":"动作标签","abbr":"","aliases":["Action Annotation"],"one_liner":"The action value a robot executed at each moment in a dataset — the “answer” imitation learning tries to fit.","explanation":"An action label is the action value aligned with each frame of observation in robot data — end-effector displacement and rotation, target joint angles, gripper opening. During teleoperation, the system records it directly; handheld capture methods like UMI instead recover it afterward with algorithms such as SLAM. Behavior cloning treats the observation as input and the action label as the supervision signal, so a label's frequency, coordinate frame, and whether it's absolute or incremental all directly shape what the model learns; RT-1-X and similar models unify data from different sources into a 7-dimensional action (3 for position, 3 for rotation, 1 for the gripper). Action labels can only be captured on a robot or with special equipment, which is the main reason robot data is expensive — hence approaches that use an inverse dynamics model to add pseudo action labels to video, or that learn latent actions instead.","example":"OpenAI's VPT first trained an inverse dynamics model on 1,962 hours of Minecraft footage recorded by contractors along with their keyboard and mouse input, then used it to add pseudo action labels to about 70,000 hours of web video, ultimately training an agent capable of crafting a diamond tool.","related":["Observation-Action Pair","Behavior Cloning","Action Space","Pseudo Action Labels","Inverse Dynamics Model","Action-free Video"]},{"id":"multimodal-data","category":"data","sec":0,"tier":2,"sources":[{"title":"RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation (arXiv HTML)","url":"https://arxiv.org/html/2412.13877v3"},{"title":"RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence","url":"https://arxiv.org/abs/2512.24653"}],"as_of":"2025-12","related_ids":["multimodal-fusion","tactile-data","proprioception","multi-sensor-time-synchronization-timestamp-alignment","hierarchical-data-format-version-5","robomind"],"name":"Multimodal Data","alt":"多模态数据","abbr":"","aliases":["Omni-modal Data"],"one_liner":"Images, depth, joint state, force and touch, language, and other signals recorded in sync while a robot works.","explanation":"In embodied AI, multimodal data means multiple signals recorded in sync while a robot performs a task: multi-view RGB images, depth maps, proprioception (the robot's own joint angles, end-effector pose, and other internal state), force or tactile readings, language instructions, and sometimes sound. A camera alone can't see a contact point hidden by the hand, or how much force was used, so fine manipulation often needs signals beyond vision recorded together. RoboMIND, from the Beijing Humanoid Robot Innovation Center and Peking University, is an example: each trajectory is stored as one HDF5 file containing multi-view RGB-D, proprioceptive state, end-effector state, and the teleoperator's own body state; RoboMIND 2.0 added 12,000 tactile-augmented segments. The difficulties are synchronizing all these signals in time, storage size, and how to train when one modality is missing.","example":"In RoboMIND 2.0's tactile-augmented segments, the same instant carries multi-view RGB-D footage and joint state together with normal and shear forces measured by a Tashan tactile sensor.","related":["Multimodal Fusion","Tactile Data","Proprioception","Multi-sensor Time Synchronization / Timestamp Alignment","Hierarchical Data Format version 5","RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation)"]},{"id":"tactile-data","category":"data","sec":0,"tier":2,"sources":[{"title":"Sparsh: Self-supervised Touch Representations for Vision-based Tactile Sensing","url":"https://arxiv.org/abs/2410.24090"},{"title":"Touch and Go: Learning from Human-Collected Vision and Touch","url":"https://arxiv.org/abs/2211.12498"},{"title":"RoboMIND 2.0 (arXiv HTML)","url":"https://arxiv.org/html/2512.24653v1"}],"as_of":"2025-12","related_ids":["tactile-sensor","vision-based-tactile-sensor","tactile-representation-learning","visuo-tactile-fusion","multimodal-data","slip-detection"],"name":"Tactile Data","alt":"触觉数据","abbr":"","aliases":["Visuo-tactile Data"],"one_liner":"Pressure, shear force, and contact-surface deformation recorded by a tactile sensor when it touches an object.","explanation":"Tactile data comes from tactile sensors mounted on a fingertip, palm, or gripper, recording contact location, normal force (how hard something is pressed), shear force (related to slipping), and how the contact surface deforms. Roughly two kinds of sensors exist: visuotactile sensors (such as GelSight and DIGIT) use a built-in camera to photograph the deformation of an elastic gel pad and output something like a photograph, a tactile image; array-style sensors — piezoresistive, capacitive, or magnetic — output a numeric reading per taxel. When the hand blocks the camera's view, or force needs to be controlled precisely (peg insertion, unscrewing a cap, holding something fragile), touch supplies information a camera can't. The difficulty is that sensor models vary and their data doesn't transfer between them, and ground-truth force and slip are hard to label, so Meta's Sparsh used self-supervised pretraining on more than 460,000 unlabeled tactile images. Datasets such as RoboMIND 2.0 now record touch in sync with vision and joint state.","example":"Pinching a strawberry with a GelSight-equipped gripper, the size of the contact area and the displacement of the tactile image's surface markers reveal how tightly it's being held and whether it has started to slip.","related":["Tactile Sensor","Vision-Based Tactile Sensor","Tactile Representation Learning","Visuo-Tactile Fusion","Multimodal Data","Slip Detection"]},{"id":"multi-sensor-time-synchronization-timestamp-alignment","category":"data","sec":0,"tier":2,"sources":[{"title":"Universal Manipulation Interface (arXiv 2402.10329, HTML)","url":"https://arxiv.org/html/2402.10329"},{"title":"Wikipedia: Precision Time Protocol","url":"https://en.wikipedia.org/wiki/Precision_Time_Protocol"},{"title":"ros2/message_filters 文档","url":"https://raw.githubusercontent.com/ros2/message_filters/rolling/doc/index.rst"}],"as_of":"","related_ids":["precision-time-protocol","multi-sensor-fusion","camera-imu-calibration","observation-action-pair","ros-bag","control-latency"],"name":"Multi-sensor Time Synchronization / Timestamp Alignment","alt":"多传感器时间同步（时间戳对齐）","abbr":"","aliases":["Timestamp Alignment","Time Sync","Hardware Sync"],"one_liner":"Aligning readings from cameras, encoders, IMUs, and other sensors onto one shared clock so simultaneous data lines up correctly.","explanation":"A robot's cameras, joint encoders, force sensors, and IMU each run at their own frequency and with their own latency — a camera's latency, from exposure, encoding, and transmission, is usually larger than a joint reading's. Time synchronization puts all this data on one shared timeline: on the hardware side, a shared trigger signal can make multiple cameras expose at the same instant, or Precision Time Protocol (PTP, IEEE 1588) can align device clocks to sub-microsecond precision; on the software side, messages get paired by matching their timestamps to the closest match, as with ROS's message_filters, which offers both exact and approximate matching. For imitation learning, misalignment means the model ends up learning from mismatched observation-action pairs.","example":"UMI first measures the latency of each data stream — camera, gripper width, and so on — separately, then aligns every observation to whichever stream has the largest latency, usually the camera; at deployment time it issues action commands slightly early to cancel out execution latency.","related":["Precision Time Protocol","Multi-Sensor Fusion","Camera-IMU Calibration","Observation-Action Pair","ROS Bag","Control Latency"]},{"id":"real-robot-data","category":"data","sec":0,"tier":1,"sources":[{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"},{"title":"AgiBot World Colosseo (arXiv 2503.06669)","url":"https://arxiv.org/abs/2503.06669"}],"as_of":"2025-03","related_ids":["teleoperation","simulation-data","data-pyramid","droid","agibot-world","real-robot-data-camp-vs-sim-data-camp"],"name":"Real-Robot Data","alt":"真机数据","abbr":"","aliases":["Real-World Robot Data"],"one_liner":"Data recorded from an actual robot physically performing tasks in the real world.","explanation":"Real-robot data refers to data recorded from an actual robot body performing tasks in the physical world, including camera footage, joint state, and action commands, usually collected through teleoperation, though it also includes data recovered when a robot acts autonomously. It carries no sim-to-real gap (the physical and visual mismatch between simulation and reality), and its actions can be replayed directly on the same model of robot, which is generally considered the highest-quality data available. Its downside is that it is expensive and slow: it requires buying robots, building scenes, and hiring data collectors to record data one demonstration at a time. The industry often describes this with a “data pyramid”: massive internet and human video data at the bottom, simulation and synthetic data in the middle, and a small amount of precious real-robot data at the top. Notable datasets include DROID, AgiBot World, and Open X-Embodiment.","example":"DROID was collected by 50 people across 13 institutions teleoperating Franka arms with Quest 2 headsets over 12 months, gathering 76,000 trajectories totaling 350 hours across 564 scenes.","related":["Teleoperation","Simulation Data","Data Pyramid","DROID (Distributed Robot Interaction Dataset)","AgiBot World","Real-Robot-Data Camp vs. Sim-Data Camp"]},{"id":"data-scarcity","category":"data","sec":0,"tier":2,"sources":[{"title":"UC Berkeley CDSS: Humanoid robots face challenges in gaining real-world skills","url":"https://cdss.berkeley.edu/news/humanoid-robots-face-challenges-gaining-real-world-skills-says-berkeley-expert"},{"title":"Rockingrobots: Humanoid robots are advancing but face a massive data gap","url":"https://www.rockingrobots.com/humanoid-robots-are-advancing-but-face-a-massive-data-gap"},{"title":"钛媒体：机器人还没学会做家务，卖数据的已经先赚到了钱","url":"https://www.tmtpost.com/8062934.html"}],"as_of":"2026-07","related_ids":["data-pyramid","simulation-data","synthetic-data","human-video-data","data-flywheel","teleoperation"],"name":"Data Scarcity","alt":"数据荒","abbr":"","aliases":["Data Bottleneck","Data Gap"],"one_liner":"The core bottleneck in embodied AI: usable real robot-interaction data falls far short of what's needed.","explanation":"Data scarcity refers to embodied AI's lack of enough high-quality training data. Large language models can draw on the entire text internet, but what a robot needs is data about “how to actually move” — joint angles, force, touch — which is almost nowhere on the internet and can only be gathered one demonstration at a time through teleoperation or wearable devices, which is expensive, slow, and doesn't transfer easily between different robots. UC Berkeley professor Ken Goldberg wrote in Science Robotics in August 2025 that the text used to train large models is equivalent to what a person would need 100,000 years to read, and argued that robotics correspondingly faces a “100,000-year data gap.” Responses include collecting more real-robot data, generating data with simulation and generative models, learning from human video, sharing data across robot embodiments, and letting robots accumulate data during deployment.","example":"According to a July 2026 report from tech outlet TMTPost, the CEO of Mifeng Technology estimated embodied AI would need about 100 million hours of training data to reach GPT-3.5-level capability, while as of early 2026 the world's available high-quality real physical-interaction data totaled only a few hundred thousand hours.","related":["Data Pyramid","Simulation Data","Synthetic Data","Human Video Data","Data Flywheel","Teleoperation"]},{"id":"simulation-data","category":"data","sec":0,"tier":1,"sources":[{"title":"GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data","url":"https://pku-epic.github.io/GraspVLA-web/"},{"title":"MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations","url":"https://mimicgen.github.io/"}],"as_of":"2025","related_ids":["simulator","synthetic-data","sim-to-real-gap","domain-randomization","sim-to-real-transfer","syngrasp-1b"],"name":"Simulation Data","alt":"仿真数据","abbr":"","aliases":["Sim Data"],"one_liner":"Training data generated automatically by having a virtual robot perform tasks inside a physics simulator.","explanation":"Simulation data is generated inside a physics simulator such as Isaac Sim or MuJoCo: the scene, objects, and robot are all virtual, and a script, motion planner, or reinforcement-learning policy completes tasks on its own while images, depth, joint state, and actions are all recorded. The advantages are that it's cheap, runs at massive parallel scale, comes with perfectly accurate labels for free, and can randomize lighting, materials, and object placement — domain randomization — to cover long-tail situations that are hard to capture with a real robot. The main obstacle is the sim-to-real gap: simulated contact physics and rendered images never perfectly match reality, and deformable objects and fine-grained contact are especially hard to get right. A common approach is to pretrain at scale on simulation data first, then fine-tune with a small amount of real-robot data, or train on a mixture of both together.","example":"Teams such as Galbot pretrained GraspVLA using SynGrasp-1B, a billion-frame simulated grasping dataset (over 10,000 objects across 240 categories), for all of its action data, combined with internet image-text grounding data for joint pretraining; without using any real-robot data at all, it transferred zero-shot to real-world grasping.","related":["Simulator","Synthetic Data","Sim-to-Real Gap (Reality Gap)","Domain Randomization","Sim-to-Real Transfer","SynGrasp-1B"]},{"id":"synthetic-data","category":"data","sec":0,"tier":1,"sources":[{"title":"NVIDIA Glossary: What Is Synthetic Data Generation?","url":"https://www.nvidia.com/en-us/glossary/synthetic-data-generation/"},{"title":"MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations","url":"https://mimicgen.github.io/"},{"title":"DreamGen: Unlocking Generalization in Robot Learning through Video World Models (NVIDIA GEAR)","url":"https://research.nvidia.com/labs/gear/dreamgen/"}],"as_of":"2025-05","related_ids":["simulation-data","mimicgen","dreamgen","neural-trajectories","pseudo-action-labels","generative-data-augmentation"],"name":"Synthetic Data","alt":"合成数据","abbr":"","aliases":["Generated Data"],"one_liner":"Training data manufactured by computers, via simulation or generative models, instead of collected from the real world.","explanation":"Synthetic data is data produced by a computer rather than collected from reality — NVIDIA defines it as text, images, and video generated using computer simulation, generative AI models, or a combination of both. In embodied AI there are mainly three routes: rendering and executing tasks inside a simulator (simulation data); automatically transforming and expanding a small number of human demonstrations into many new ones, as in MimicGen; and directly generating videos of a robot working with a video-generation model or world model, then filling in pseudo action labels with an inverse dynamics model (a model that infers actions from before-and-after frames), as in NVIDIA's DreamGen. It helps relieve the scarcity of real-robot data, but whether the generated content is physically plausible and close to the real distribution ultimately still has to be checked with real-robot evaluation.","example":"MimicGen automatically expands fewer than 200 human demonstrations into more than 50,000, covering 18 tasks; DreamGen, using only teleoperation data from a single pick-and-place task, taught a humanoid robot 22 new skills across 10 new environments.","related":["Simulation Data","MimicGen","DreamGen","Neural Trajectories","Pseudo Action Labels","Generative Data Augmentation"]},{"id":"human-video-data","category":"data","sec":0,"tier":1,"sources":[{"title":"EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (NVIDIA GEAR)","url":"https://research.nvidia.com/labs/gear/egoscale/"},{"title":"EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video (arXiv 2505.11709)","url":"https://arxiv.org/abs/2505.11709"}],"as_of":"2026-02","related_ids":["egocentric-video","internet-video-data","action-free-video","latent-action-pretraining","embodiment-gap","pretraining-on-human-videos"],"name":"Human Video Data","alt":"人类视频数据","abbr":"","aliases":["Human Video"],"one_liner":"Video of people doing everyday tasks, collectable without any robot, used to teach robots to understand and copy manipulation.","explanation":"Human video data means video recording people performing all kinds of manipulation, including egocentric video, third-person (exocentric) video, and instructional videos pulled from the internet. Its biggest advantage is volume, low cost, and scene diversity, since collecting it doesn't depend on an expensive robot. The difficulty is that the video carries no action labels a robot can directly execute, and a human hand differs structurally from a robot hand or gripper — the embodiment gap. Common approaches: use hand-pose estimation to extract wrist and finger trajectories as actions; train a latent-action model to learn an abstract action from consecutive frames; or use the video only to pretrain visual representations and world models, then fine-tune with a small amount of real-robot data. NVIDIA's 2026 EgoScale pretrained a VLA (vision-language-action) model on more than 20,000 hours of action-annotated egocentric video, and found that data volume scales log-linearly with validation loss.","example":"EgoScale first pretrains on human video, then runs mid-training — a transitional stage between pretraining and task fine-tuning — on a small amount of paired human-robot data; on a 22-degree-of-freedom dexterous hand, its average success rate came out 54% higher than skipping the pretraining step.","related":["Egocentric Video","Internet Video Data","Action-free Video","Latent Action Pretraining","Embodiment Gap","Pretraining on Human Videos"]},{"id":"web-scale-vision-language-data","category":"data","sec":0,"tier":2,"sources":[{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","url":"https://robotics-transformer2.github.io/"},{"title":"LAION-5B: An open large-scale dataset for training next generation image-text models","url":"https://arxiv.org/abs/2210.08402"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization","url":"https://arxiv.org/abs/2504.16054"}],"as_of":"2025-04","related_ids":["vision-language-model","co-training","vision-language-action-model","rt-2","catastrophic-forgetting","knowledge-insulation"],"name":"Web-scale Vision-Language Data","alt":"互联网图文数据","abbr":"","aliases":["Internet-scale Vision-Language Data","Web Data"],"one_liner":"The vast supply of paired images and text, plus image-based Q&A, collected from the web, teaching a model about the world.","explanation":"This refers to large-scale image-text pairs collected from web pages, and the image captioning, visual question answering, and object detection data derived from them — for instance LAION-5B, which used CLIP to filter Common Crawl web pages down to 5.85 billion image-text pairs. Vision-language models (VLMs) pretrain on data like this, which is why they recognize a huge range of objects and carry a fair amount of common sense. Robot data is far smaller in scale, covering a limited set of objects and scenes, so VLA models mainly draw on internet-scale knowledge two ways: using a VLM backbone that was itself pretrained on web-scale data, and co-training on a mix of web data and robot data during fine-tuning, so the fine-tuning process doesn't wash out what the model already knew. RT-2 was the first to systematically validate this approach, and π0.5 also lists web data as one of its co-training sources.","example":"RT-2 fine-tunes on a mix of robot trajectories and web data such as visual question answering, so the model can apply objects and concepts it learned from the web to actions, executing instructions that never appeared in the robot data.","related":["Vision-Language Model","Co-training","Vision-Language-Action Model","RT-2","Catastrophic Forgetting","Knowledge Insulation"]},{"id":"data-pyramid","category":"data","sec":0,"tier":2,"sources":[{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"}],"as_of":"2025-03","related_ids":["real-robot-data","simulation-data","synthetic-data","human-video-data","neural-trajectories","nvidia-isaac-gr00t-n1"],"name":"Data Pyramid","alt":"数据金字塔","abbr":"","aliases":["Robot Data Pyramid"],"one_liner":"A three-tier framework organizing robot training data by volume and how closely it matches a real robot.","explanation":"The data pyramid is a framework NVIDIA used to organize training data in its March 2025 GR00T N1 paper: the base is web data and human video, by far the largest in volume; the middle tier is data generated by physical simulation or synthesized by neural networks such as video-generation models; and the top tier is data collected on real robots. Moving up the pyramid, data volume shrinks while how closely it matches the specific embodiment (a robot's physical hardware) grows. The base supplies general common sense and behavioral priors, while the top ensures actions can actually be executed on the real robot. It targets the reality that real-robot data is expensive and scarce: build on a large, cheap base and calibrate with a small amount of real-robot data at the top.","example":"GR00T N1's base tier used human video such as Ego4D and EPIC-KITCHENS; its middle tier had about 827 hours of neural trajectories generated by a video model plus 780,000 DexMimicGen simulated trajectories; its top tier was real-robot data such as Fourier GR-1 teleoperation data, Open X-Embodiment, and AgiBot World.","related":["Real-Robot Data","Simulation Data","Synthetic Data","Human Video Data","Neural Trajectories","NVIDIA Isaac GR00T N1"]},{"id":"teleoperation","category":"data","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Teleoperation","url":"https://en.wikipedia.org/wiki/Teleoperation"},{"title":"ALOHA / ACT 项目页: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","url":"https://tonyzhaozh.github.io/aloha/"},{"title":"Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["demonstration-data","leader-follower-teleoperation","vr-teleoperation","data-collector","imitation-learning","universal-manipulation-interface"],"name":"Teleoperation","alt":"遥操作","abbr":"","aliases":["teleop","Remote Operation"],"one_liner":"A human controls a robot in real time while the process is recorded to produce training data.","explanation":"Teleoperation originally just means operating a machine from a distance — the everyday term “remote control” in academic and technical usage. In embodied AI it is the primary way to collect demonstration data: an operator issues commands through a leader-follower arm, a VR headset, a 3D mouse, or a motion-capture glove, the robot follows along, and the system simultaneously logs camera footage, joint state, and the action taken at every step, producing the “observation-action” pairs that imitation learning needs. Because the data is recorded directly on the robot's own body, the resulting policy can be deployed as-is, with no mismatch between a human hand and a robotic one; the trade-off is that it needs a real robot and a skilled data collector, and is slow and expensive. Handheld capture devices like UMI, and learning from human video, both exist to get around this bottleneck.","example":"The ALOHA bimanual platform uses leader-follower teleoperation: an operator moves a pair of leader arms while a matching pair of follower arms tracks them in real time, with the whole hardware setup costing about $20,000. On most tasks, ACT needed just 50 such recorded demonstrations (about 10–20 minutes) to learn fine-grained skills like uncapping a condiment jar or inserting a battery into a remote control.","related":["Demonstration Data","Leader-Follower Teleoperation","VR Teleoperation","Data Collector (Teleoperator)","Imitation Learning","Universal Manipulation Interface"]},{"id":"kinesthetic-teaching","category":"data","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Programming by demonstration","url":"https://en.wikipedia.org/wiki/Programming_by_demonstration"},{"title":"Kinesthetic Teaching in Robotics: a Mixed Reality Approach (arXiv 2409.02305)","url":"https://arxiv.org/abs/2409.02305"},{"title":"DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection (arXiv 2607.27784)","url":"https://arxiv.org/abs/2607.27784"}],"as_of":"","related_ids":["zero-force-drag","gravity-compensation","teach-and-playback-programming","teleoperation","demonstration-data","impedance-control"],"name":"Kinesthetic Teaching","alt":"拖动示教","abbr":"KT","aliases":["KT","Hand Guiding","Direct Teaching"],"one_liner":"A human physically pushes a robot arm through a motion by hand, and the robot records it as a demonstration.","explanation":"Kinesthetic teaching is the most direct form of demonstration: the arm is switched into gravity-compensation or zero-force (free-drive) mode, in which the motors only cancel out the arm's own weight, so a person can move it around easily, then an operator guides the arm by hand through the whole task while its joint encoders record the entire trajectory. Early industrial robots' “teach and playback” programming worked this way — remembering a sequence of positions and replaying them; in robot learning, it supplies demonstrations for imitation learning. The benefit is that the motion happens directly on the robot's own joints, so there's no need to map a human's action onto the robot, and no extra teleoperation hardware is required. The drawbacks are that the arm has to support being pushed by hand, the human's body enters the camera view and interferes with vision-based policies, it's hard to guide two arms or a multi-fingered hand at once, and it is more physically tiring than teleoperation.","example":"Put an arm with joint-torque sensors into hand-guiding mode, grab the end effector, guide the gripper to a cup's handle, close the gripper, and lift — the recorded sequence of joint angles is one demonstration.","related":["Zero-Force Drag","Gravity Compensation","Teach-and-Playback Programming","Teleoperation","Demonstration Data","Impedance Control"]},{"id":"leader-follower-teleoperation","category":"data","sec":1,"tier":1,"sources":[{"title":"GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators","url":"https://wuphilipp.github.io/gello_site/"},{"title":"ALOHA: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware","url":"https://tonyzhaozh.github.io/aloha/"},{"title":"LeRobot Docs: Imitation Learning on Real-World Robots","url":"https://huggingface.co/docs/lerobot/il_robots"}],"as_of":"","related_ids":["teleoperation","aloha","gello","bilateral-teleoperation","demonstration-data","so-100-so-101-arm"],"name":"Leader-Follower Teleoperation","alt":"主从臂遥操作","abbr":"","aliases":["Leader-Follower Arms","Isomorphic Leader Arm"],"one_liner":"A human moves a leader arm, and the robot's follower arm copies each joint angle in real time.","explanation":"Leader-follower teleoperation is a common way to collect robot data: the leader arm is a hand-operated mechanical arm with the same kinematic structure as the robot's follower arm, or a scaled-down version of it, with encoders or servos at each joint. The operator moves the leader arm, and the system reads out each joint angle and sends it straight to the follower arm. Because it is a direct joint-to-joint mapping, there is no need to solve inverse kinematics (working backward from end-effector position to joint angles), so latency is low and the motion is intuitive, which suits collecting fine bimanual manipulation data. The trade-off is that each robot needs its own matching leader arm, and there is usually no force feedback. ALOHA uses a WidowX leader arm with a ViperX follower arm for bimanual teleoperation; Berkeley's GELLO uses 3D-printed parts and off-the-shelf motors, at a cost under $300.","example":"In LeRobot, teleoperate an SO-101 follower arm with an SO-101 leader arm, recording camera footage and joint angles as you go; once you've recorded 50 demonstrations, that's enough to train an ACT policy.","related":["Teleoperation","ALOHA","GELLO","Bilateral Teleoperation","Demonstration Data","SO-100 / SO-101 Arm"]},{"id":"gello","category":"data","sec":1,"tier":3,"sources":[{"title":"GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators (arXiv)","url":"https://arxiv.org/abs/2309.13037"},{"title":"GELLO 项目主页","url":"https://wuphilipp.github.io/gello_site/"},{"title":"gello_mechanical (GitHub)","url":"https://github.com/wuphilipp/gello_mechanical"}],"as_of":"2024-07","related_ids":["leader-follower-teleoperation","teleoperation","aloha","demonstration-data","spacemouse-teleoperation","vr-teleoperation"],"name":"GELLO","alt":"GELLO","abbr":"GELLO","aliases":["GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators","GELLO Teleoperation Arm"],"one_liner":"Berkeley's open-source low-cost leader-follower teleoperation device, using a scaled-down replica arm to control the real robot.","explanation":"GELLO is a teleoperation framework proposed in 2023 by Pieter Abbeel's group at UC Berkeley. The approach is to 3D-print a scaled-down replica arm with the same joint structure as the target robot arm, fitting each joint with an off-the-shelf servo that reads its angle; the operator moves the replica arm, and the real robot follows along joint angle by joint angle — a form of leader-follower teleoperation. Because the two arms share the same kinematics, joint angles map directly across, unlike a VR controller, which has to convert a hand's pose into a robot-arm pose. Parts cost under $300, and all the hardware and software designs are open-sourced. The paper builds versions for the Franka, UR5, and xArm arms; user studies show it collects demonstrations more stably and faster than a VR controller or a 3D mouse, and it also supports bimanual and contact-rich tasks. The community has since contributed designs for other arms, including the FR3 and YAM, and it's commonly used to collect imitation-learning data.","example":"An operator holds the GELLO replica arm on the desk and performs a pick-and-place motion; the real Franka arm follows along in sync, and the camera footage and joint angles are recorded as one training demonstration.","related":["Leader-Follower Teleoperation","Teleoperation","ALOHA","Demonstration Data","SpaceMouse Teleoperation","VR Teleoperation"]},{"id":"bilateral-teleoperation","category":"data","sec":1,"tier":3,"sources":[{"title":"Fast Bilateral Teleoperation and Imitation Learning Using Sensorless Force Control via Accurate Dynamics Model (arXiv 2507.06174)","url":"https://arxiv.org/html/2507.06174v2"},{"title":"ALPHA-α and Bi-ACT Are All You Need（大阪大学项目页）","url":"https://mertcookimg.github.io/alpha-biact"}],"as_of":"2025-07","related_ids":["teleoperation","leader-follower-teleoperation","haptic-glove","force-control","contact-rich-manipulation","action-chunking-with-transformers"],"name":"Bilateral Teleoperation","alt":"双边遥操作","abbr":"","aliases":["Force-Feedback Teleoperation","Bilateral Control"],"one_liner":"Teleoperation where position and force are exchanged in both directions between leader and follower, so the operator feels what the robot touches.","explanation":"Bilateral teleoperation is a form of leader-follower teleoperation: a human moves a local leader device, a remote follower robot tracks that motion, and the contact forces the follower encounters are sent back to the leader so the operator feels resistance in their hand. Unilateral teleoperation, by contrast, only sends the leader's position to the follower — low-cost leader-follower arm setups like ALOHA work this way. Bilateral teleoperation dates back more than fifty years; its control design has to balance stability and transparency (how closely the operator's sensation matches actually doing the task by hand), and communication latency is the main obstacle. A four-channel architecture exchanges both position and force in both directions, which can in principle keep the two sides synchronized. For data collection, bilateral teleoperation helps a human perform contact-rich tasks like plugging in connectors or wiping surfaces, and records the force signal alongside the demonstration for a policy to learn from — for example, Bi-ACT adds force information on top of the ACT architecture.","example":"A team at the University of Tsukuba and collaborators used four-channel bilateral control to collect demonstrations for policy training: without force information, the policy could barely pick up small 10–20mm objects; adding force to both the input and output raised the pick success rate to 100%.","related":["Teleoperation","Leader-Follower Teleoperation","Haptic Glove","Force Control","Contact-rich Manipulation","Action Chunking with Transformers"]},{"id":"exoskeleton-teleoperation","category":"data","sec":1,"tier":2,"sources":[{"title":"AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild (arXiv 2309.14975)","url":"https://arxiv.org/abs/2309.14975"},{"title":"HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit (arXiv 2502.13013)","url":"https://arxiv.org/abs/2502.13013"}],"as_of":"2025-04","related_ids":["teleoperation","leader-follower-teleoperation","whole-body-teleoperation","airexo","homie","exoskeleton"],"name":"Exoskeleton Teleoperation","alt":"外骨骼遥操作","abbr":"","aliases":["Exoskeleton Collection"],"one_liner":"The operator wears an exoskeleton matching the robot's joints, controlling it directly with their own arm motion.","explanation":"Exoskeleton teleoperation is a form of teleoperation: the operator wears an exoskeleton whose joints each carry an encoder (a sensor that measures rotation angle), built with the same structure as the robot's arm, or a proportionally matched one. However the person's arm moves, the readings map straight onto the robot's joint angles, with no need to first solve inverse kinematics (working backward from end-effector pose to joint angles) the way a VR controller requires. The benefits are that the whole arm's pose is controllable, latency is low, it feels intuitive, and it can also be used disconnected from any robot to collect data “in the wild”; the drawbacks are that the structure has to be adapted to each robot, and wearing it for a long time is tiring. Notable examples include Shanghai Jiao Tong University's Cewu Lu group's AirExo family of bimanual exoskeletons, and HOMIE, a 2025 exoskeleton cockpit for humanoid robots.","example":"In the AirExo paper, a policy trained on just 3 minutes of teleoperation data plus a large amount of in-the-wild exoskeleton demonstrations matched or exceeded the performance of a policy trained on more than 20 minutes of teleoperation data alone; the full HOMIE cockpit costs about $500.","related":["Teleoperation","Leader-Follower Teleoperation","Whole-Body Teleoperation","AirExo","HOMIE","Exoskeleton"]},{"id":"airexo","category":"data","sec":1,"tier":3,"sources":[{"title":"AirExo: Low-Cost Exoskeletons for Learning Whole-Arm Manipulation in the Wild (arXiv 2309.14975)","url":"https://arxiv.org/abs/2309.14975"},{"title":"AirExo 项目主页","url":"https://airexo.github.io/"},{"title":"AirExo-2 项目主页","url":"https://airexo.tech/airexo2/"}],"as_of":"2025-09","related_ids":["exoskeleton-teleoperation","robot-free-data-collection","leader-follower-teleoperation","in-the-wild-data","sjtu-mvig-lab","flexiv-rizon"],"name":"AirExo","alt":"AirExo 外骨骼","abbr":"","aliases":["AirExo-2"],"one_liner":"A low-cost bimanual exoskeleton from Shanghai Jiao Tong's Cewu Lu group that can both teleoperate a robot and collect data without one.","explanation":"AirExo is an open-source bimanual exoskeleton proposed by Cewu Lu's group at Shanghai Jiao Tong University in September 2023, published at ICRA 2024. A person wears it on both arms; the exoskeleton's joints correspond one-to-one with the robot's joints, so it can do joint-level teleoperation, and it can also be worn away from the robot to record “in-the-wild” demonstrations directly in real environments. Its structural parts can all be 3D-printed, costing about $300 per arm, originally designed for a Flexiv Rizon dual-arm setup but adaptable to UR5, Franka, and others. AirExo-2 (CoRL 2025) adds a vision adapter that converts footage of a person wearing the exoskeleton into “pseudo-robot demonstrations,” paired with a policy called RISE-2 that fuses 3D point clouds with 2D semantic features; using no real-robot data at all, it approaches the performance of a policy trained on teleoperation data. The official cost is about $600 per full kit, compared with roughly $60,000 for a teleoperation platform.","example":"In the AirExo paper's experiments, a policy trained on just 3 minutes of teleoperation demonstrations plus a large amount of in-the-wild exoskeleton demonstrations matched or exceeded one trained on more than 20 minutes of teleoperation data alone.","related":["Exoskeleton Teleoperation","Robot-free (Embodiment-free) Data Collection","Leader-Follower Teleoperation","In-the-wild Data","SJTU MVIG Lab","Flexiv Rizon"]},{"id":"spacemouse-teleoperation","category":"data","sec":1,"tier":3,"sources":[{"title":"robosuite Docs: I/O Devices (3Dconnexion SpaceMouse)","url":"https://robosuite.ai/docs/modules/devices.html"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop RL (HIL-SERL, arXiv 2410.21845)","url":"https://arxiv.org/html/2410.21845"},{"title":"Wikipedia: 3Dconnexion","url":"https://en.wikipedia.org/wiki/3Dconnexion"}],"as_of":"","related_ids":["teleoperation","human-in-the-loop","hil-serl","human-intervention-data","robosuite","leader-follower-teleoperation"],"name":"SpaceMouse Teleoperation","alt":"3D 鼠标遥操作","abbr":"","aliases":["SpaceMouse","3D Mouse"],"one_liner":"Controlling a robot arm's end effector in six degrees of freedom by pushing, pulling, twisting, and tilting a 3D mouse.","explanation":"The SpaceMouse is a 6-DoF input device from 3Dconnexion, originally built for panning, zooming, and rotating 3D models in CAD software; its knob can be pushed sideways, pressed up and down, twisted, and tilted. When used for teleoperation, these 6 quantities are mapped onto incremental translation and rotation of the robot arm's end effector, with a button controlling the gripper's open/close. It's cheap and plug-and-play, needing no leader-follower arm or VR headset, which makes it well suited to single-arm tabletop tasks; simulation frameworks like robosuite support it natively. The downside is that it only controls end-effector pose, so it doesn't suit bimanual or dexterous-hand setups, feels less intuitive, and produces slower, less natural motion than a leader-follower arm. In human-in-the-loop reinforcement learning, it's commonly used as the tool a human uses to take over and correct the robot at any moment.","example":"During HIL-SERL training, a person holds a SpaceMouse while supervising the robot; whenever the policy makes a mistake, the person takes over directly through the SpaceMouse and provides a corrective action, and these intervention episodes are stored in the replay buffer to continue training.","related":["Teleoperation","Human-in-the-Loop","HIL-SERL","Human Intervention Data","robosuite","Leader-Follower Teleoperation"]},{"id":"vr-teleoperation","category":"data","sec":1,"tier":1,"sources":[{"title":"Cheng et al. 2024: Open-TeleVision (arXiv 2407.01512)","url":"https://arxiv.org/abs/2407.01512"},{"title":"GitHub: unitreerobotics/xr_teleoperate","url":"https://github.com/unitreerobotics/xr_teleoperate"},{"title":"AgiBot World Colosseo 技术报告 (arXiv 2503.06669)","url":"https://arxiv.org/html/2503.06669"}],"as_of":"2026-09","related_ids":["teleoperation","open-television","apple-vision-pro","motion-retargeting","unitree-xr-teleoperate","vr-headset"],"name":"VR Teleoperation","alt":"VR 遥操作","abbr":"","aliases":["XR Teleoperation","Headset Teleoperation"],"one_liner":"Controlling a robot in real time by wearing a VR/XR headset and using hand tracking or controllers.","explanation":"VR teleoperation is a form of teleoperation: the operator wears a VR/XR headset, which displays the video streamed back from the robot's camera; the headset's hand tracking or controllers read out the human's wrist and finger poses, which are then converted into robot commands through inverse kinematics and motion retargeting (translating human hand pose into robot joint angles). It is cheaper and more portable than a leader-follower arm, and suits humanoid robots and dexterous hands, which have many more degrees of freedom — datasets such as BridgeData V2, DROID, and AgiBot World have all used it. Its downside is that it usually has no force feedback, and hand tracking has latency and jitter, making it less stable than a leader-follower arm for fine-contact tasks; AgiBot World's paper also notes that its VR setup could only produce a few preset dexterous-hand gestures, switching to motion capture for more complex tasks.","example":"Open-TeleVision uses an Apple Vision Pro to control a Unitree H1: the robot's head follows the operator's head movements, its stereo camera feed streams back to the headset, and both hands follow via dex-retargeting; Unitree's open-source xr_teleoperate supports Apple Vision Pro, PICO 4 Ultra Enterprise, and Meta Quest 3, and saves each episode for imitation learning.","related":["Teleoperation","Open-TeleVision","Apple Vision Pro","Motion Retargeting","Unitree xr_teleoperate","VR Headset"]},{"id":"apple-vision-pro","category":"data","sec":1,"tier":2,"sources":[{"title":"Apple Newsroom: Apple Vision Pro available in the U.S. on February 2","url":"https://www.apple.com/newsroom/2024/01/apple-vision-pro-available-in-the-us-on-february-2/"},{"title":"Wikipedia: Apple Vision Pro","url":"https://en.wikipedia.org/wiki/Apple_Vision_Pro"},{"title":"Hoque et al. 2025: EgoDex (arXiv 2505.11709)","url":"https://arxiv.org/abs/2505.11709"}],"as_of":"2025-10","related_ids":["vr-teleoperation","vr-headset","open-television","egodex","egocentric-video","meta-quest-3-pico-4-ultra-xr-headsets"],"name":"Apple Vision Pro","alt":"Apple Vision Pro","abbr":"AVP","aliases":["AVP"],"one_liner":"Apple's head-mounted display, whose hand tracking embodied-AI researchers commonly use for teleoperation and collecting human data.","explanation":"Apple Vision Pro is a headset Apple unveiled at WWDC in June 2023 and began selling in the US on February 2, 2024, starting at $3,499; it originally shipped with an M2 chip plus an R1 chip dedicated to processing sensor input, and has 12 cameras, 5 sensors, and 6 microphones, with an M5 version following in October 2025. It is not a robotics product, but its system tracks the wrist and fingers' 3D pose in real time and has a high-resolution stereo display, so embodied-AI research has adopted it for teleoperation: a human's hand motion, after retargeting, drives the robot, and the robot's head camera streams back to the headset. It can also record egocentric video and 3D hand trajectories in sync while a person does chores, which is used to collect human manipulation data.","example":"Open-TeleVision used a Vision Pro to remotely control a Unitree H1 — the author, at MIT, controlled a robot physically located at UC San Diego; Apple itself used Vision Pro to record EgoDex, an 829-hour, 194-task egocentric dataset of human hand manipulation.","related":["VR Teleoperation","VR Headset","Open-TeleVision","EgoDex","Egocentric Video","Meta Quest 3 / PICO 4 Ultra XR Headsets"]},{"id":"data-glove","category":"data","sec":1,"tier":1,"sources":[{"title":"Wikipedia: Wired glove","url":"https://en.wikipedia.org/wiki/Wired_glove"},{"title":"MANUS 官网（Metagloves 产品）","url":"https://www.manus-meta.com/"},{"title":"DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation (arXiv 2403.07788)","url":"https://arxiv.org/abs/2403.07788"}],"as_of":"2026-09","related_ids":["motion-capture","haptic-glove","tactile-glove","teleoperation","motion-retargeting","dexterous-hand"],"name":"Data Glove","alt":"数据手套","abbr":"","aliases":["Wired Glove","CyberGlove"],"one_liner":"A wearable sensing glove that tracks each finger's bend and the hand's pose in real time.","explanation":"A data glove is a wearable input device that measures each finger joint's angle using bend sensors, an inertial measurement unit (IMU — a chip that senses angular velocity and acceleration), or electromagnetic sensing, usually paired with a tracker to also capture the whole hand's pose in space. It was originally used mainly for human-computer interaction, virtual reality, and animation; the 1977 Sayre Glove is considered the first one. In embodied AI it has two main uses: teleoperating a dexterous hand by mapping a human hand's motion onto the robot hand in real time, and collecting human hand-manipulation data directly, without any robot present, which is then converted through motion retargeting — translating human motion into robot joint commands — for training. Versions with force feedback can also relay the sensation of contact back to the operator.","example":"MANUS markets its Metagloves Pro for robot-training data collection and teleoperation. The DexCap system pairs an electromagnetic motion-capture glove with a SLAM camera — one that can compute its own position while moving — to track the wrist, giving a portable way to record human manipulation for dexterous hands to learn from through imitation learning.","related":["Motion Capture","Haptic Glove","Tactile Glove","Teleoperation","Motion Retargeting","Dexterous Hand"]},{"id":"haptic-glove","category":"data","sec":1,"tier":3,"sources":[{"title":"Wired glove - Wikipedia","url":"https://en.wikipedia.org/wiki/Wired_glove"},{"title":"HaptX 官网","url":"https://haptx.com/"},{"title":"SenseGlove 官网","url":"https://www.senseglove.com/"}],"as_of":"2026-09","related_ids":["data-glove","tactile-glove","teleoperation","bilateral-teleoperation","exoskeleton-teleoperation","dexterous-hand"],"name":"Haptic Glove","alt":"力反馈手套","abbr":"","aliases":["Force-Feedback Glove","Haptic Feedback Glove"],"one_liner":"A teleoperation glove that both records hand motion and pushes force back against the operator's fingers.","explanation":"An ordinary data glove only records: it measures finger bend and hand position. A haptic (force-feedback) glove adds a return channel that sends the force a virtual object or a remote robot hand experiences back to the operator's own hand. The usual approach mounts an exoskeleton-like mechanism on the back of the hand and fingers, using motors, brakes, or pneumatic/hydraulic actuators to resist further bending of the fingers, so the wearer feels an object's size and stiffness; some designs add vibration or microfluidic contact points to simulate the feel of contact — HaptX gloves, for instance, press against the skin using hundreds of actuators embedded in a microfluidic fabric. An early commercial product was the CyberGrasp from Immersion Corporation. In embodied AI, this is mainly used for dexterous-hand teleoperation: without force feedback, the operator can't tell how hard they're gripping, and objects easily get crushed or dropped; with feedback, the resulting demonstration data is more natural. The tradeoffs are cost, weight, and the hassle of fitting and calibration.","example":"SenseGlove has released an exoskeleton glove for humanoid robots called the R1, which the company's website says includes active force feedback and vibration feedback, used for dexterous-hand teleoperation and imitation-learning data collection.","related":["Data Glove","Tactile Glove","Teleoperation","Bilateral Teleoperation","Exoskeleton Teleoperation","Dexterous Hand"]},{"id":"tactile-glove","category":"data","sec":1,"tier":3,"sources":[{"title":"MIT News: Sensor-packed glove learns signatures of the human grasp (2019)","url":"https://news.mit.edu/2019/sensor-glove-human-grasp-robotics-0529"},{"title":"OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer (arXiv 2512.08920)","url":"https://arxiv.org/abs/2512.08920"}],"as_of":"2025-12","related_ids":["tactile-data","data-glove","haptic-glove","tactile-sensor","tactile-representation-learning","contact-rich-manipulation"],"name":"Tactile Glove","alt":"触觉手套","abbr":"","aliases":["Pressure-Sensing Glove","Tactile Sensing Glove"],"one_liner":"A glove covered in pressure or tactile sensors on the fingers and palm, recording the contact forces of a human hand.","explanation":"A tactile glove is a device with pressure or tactile sensors arranged across the glove, recording the distribution of contact force at each point on the hand while grasping or manipulating objects. It differs in focus from a data glove (which mainly measures finger joint angles) and a haptic glove (which pushes force back onto the wearer), though the three are often combined. It's hard to tell how tightly something is being pinched, or whether it's slipping, from vision alone, and this information is critical for contact-rich tasks like wiping, plugging in connectors, or handling fragile items — so tactile gloves are used to collect force-annotated human demonstrations, study patterns of human grasping, or record touch data synchronously during teleoperation. The challenges are sensor durability and calibration, and the difference in tactile sensing between a human hand and a robot hand.","example":"MIT's STAG glove, published in Nature in 2019, has about 550 sensors and costs roughly $10 to make, and recorded about 135,000 frames of human interaction with 26 object types; the OSMO glove, open-sourced in December 2025, places 12 three-axis tactile sensors on the fingertips and palm, with humans and the robot wearing matching gloves — a wiping policy trained only on human demonstrations reached 72% success.","related":["Tactile Data","Data Glove","Haptic Glove","Tactile Sensor","Tactile Representation Learning","Contact-rich Manipulation"]},{"id":"motion-retargeting","category":"data","sec":1,"tier":2,"sources":[{"title":"GMR: General Motion Retargeting (GitHub)","url":"https://github.com/YanjieZe/GMR"},{"title":"dex-retargeting (GitHub)","url":"https://github.com/dexsuite/dex-retargeting"}],"as_of":"","related_ids":["general-motion-retargeting","dex-retargeting","inverse-kinematics","motion-capture","motion-tracking","omniretarget"],"name":"Motion Retargeting","alt":"动作重定向","abbr":"","aliases":["Retargeting","Hand Retargeting"],"one_liner":"Converting a human's (or another robot's) motion into the joint motion a target robot can actually perform.","explanation":"Motion retargeting originated as an animation technique for transferring a character's motion onto another character with different bone proportions. In embodied AI it means converting human-body or hand motion — from motion capture, video pose estimation, or a VR device — into joint angles a robot can execute. The difficulty is that limb lengths, joint counts, and ranges of motion differ between a human and a robot, so joint angles cannot simply be copied over. The usual approach frames it as an optimization problem: make key points such as the wrist and fingertips match the human's position or orientation as closely as possible, while respecting the robot's joint limits — essentially a constrained inverse-kinematics problem. It is a prerequisite step for teleoperation, learning from human video, and training humanoid robots' motion tracking; common open-source tools include GMR and dex-retargeting.","example":"GMR can convert human motion-capture data from AMASS or LAFAN1 into joint trajectories for humanoid robots like the Unitree G1, runs in real time on a CPU, and is also used by TWIST for whole-body teleoperation.","related":["General Motion Retargeting","dex-retargeting","Inverse Kinematics (IK)","Motion Capture","Motion Tracking","OmniRetarget"]},{"id":"anyteleop-a-general-vision-based-dexterous-robot-arm-hand-te","category":"data","sec":1,"tier":3,"sources":[{"title":"AnyTeleop (arXiv 2307.04577)","url":"https://arxiv.org/abs/2307.04577"},{"title":"AnyTeleop 项目主页","url":"https://yzqin.github.io/anyteleop/"}],"as_of":"2023-07","related_ids":["teleoperation","motion-retargeting","dex-retargeting","dexterous-manipulation","open-television","bunny-visionpro"],"name":"AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System","alt":"AnyTeleop","abbr":"","aliases":[],"one_liner":"A general vision-based teleoperation system that uses an ordinary camera to capture hand motion and drive many arms and dexterous hands.","explanation":"AnyTeleop was proposed in July 2023 by Yuzhe Qin and colleagues at UC San Diego and NVIDIA, published at RSS 2023. The operator wears no glove or exoskeleton — they simply move in front of a camera: the system estimates hand and wrist pose from the image, then converts it through motion retargeting (mapping human hand joints onto a structurally different robot hand) into commands for a dexterous hand and arm, with collision avoidance built in. Earlier vision-based teleoperation systems were usually tied to one specific piece of hardware; AnyTeleop instead supports many arms, dexterous hands, and camera configurations within the same system, connecting to a real robot or to simulators such as SAPIEN and Isaac Gym, and also supporting remote browser viewing and multi-robot collaboration. The paper reports it performs better than systems custom-built for the same hardware. Its retargeting module has been open-sourced as the dex-retargeting library.","example":"Inside a simulator, an operator performs a grasping motion in front of a camera, and AnyTeleop drives a virtual arm and dexterous hand in real time to grasp the object, with the recorded trajectory usable directly for training an imitation-learning policy.","related":["Teleoperation","Motion Retargeting","dex-retargeting","Dexterous Manipulation","Open-TeleVision","Bunny-VisionPro"]},{"id":"dex-retargeting","category":"data","sec":1,"tier":3,"sources":[{"title":"dex-retargeting GitHub 仓库","url":"https://github.com/dexsuite/dex-retargeting"},{"title":"AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System (arXiv)","url":"https://arxiv.org/abs/2307.04577"}],"as_of":"2026-09","related_ids":["anyteleop-a-general-vision-based-dexterous-robot-arm-hand-te","motion-retargeting","mano","dexterous-hand","mediapipe","dexycb"],"name":"dex-retargeting","alt":"dex-retargeting","abbr":"","aliases":["dex_retargeting","AnyTeleop Hand Retargeting Library"],"one_liner":"An open-source library that converts human hand keypoint poses into joint angles for many different robotic dexterous hands.","explanation":"dex-retargeting is an open-source Python library, MIT-licensed and pip-installable, spun out of the AnyTeleop project (UC San Diego and NVIDIA, RSS 2023). Human hands and robot hands differ in finger count, length, and joint structure, so joint angles can't just be copied over — they need retargeting: the library takes hand keypoints from a MediaPipe or MANO hand model as input, and solves an optimization problem for the robot hand's joint angles that keeps fingertip positions, or the vectors between fingertips, as close as possible to the human hand's. It offers three retargeting modes — vector retargeting (suited to real-time video teleoperation), position retargeting (suited to offline conversion of hand-object interaction datasets), and DexPilot — along with a model library covering hands such as Allegro, Shadow, LEAP, and Inspire. Many dexterous-hand teleoperation and human-hand data-collection projects use it to convert human hand motion into robot hand motion.","example":"An official example uses position retargeting to offline-convert the MANO hand poses of people grasping objects in the DexYCB dataset into grasping trajectories for robot hands such as Allegro and Shadow.","related":["AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System","Motion Retargeting","MANO","Dexterous Hand","MediaPipe","DexYCB"]},{"id":"open-television","category":"data","sec":1,"tier":3,"sources":[{"title":"Open-TeleVision: Teleoperation with Immersive Active Visual Feedback (arXiv 2407.01512)","url":"https://arxiv.org/abs/2407.01512"},{"title":"Open-TeleVision 项目主页","url":"https://robot-tv.github.io/"},{"title":"OpenTeleVision/TeleVision (GitHub)","url":"https://github.com/OpenTeleVision/TeleVision"}],"as_of":"2024-07","related_ids":["vr-teleoperation","apple-vision-pro","dex-retargeting","action-chunking-with-transformers","unitree-h1","ph2d"],"name":"Open-TeleVision","alt":"Open-TeleVision","abbr":"","aliases":["TeleVision","OpenTeleVision","OpenTV"],"one_liner":"An open-source system for teleoperating a humanoid robot by wearing a VR headset and seeing stereo video from the robot's own viewpoint.","explanation":"Open-TeleVision is an open-source teleoperation system from Xiaolong Wang's group at UC San Diego and MIT (CoRL 2024). The operator wears an Apple Vision Pro headset (the code also supports Meta Quest 3); when they turn their head, the robot's pan-tilt unit or neck turns to match, and a head-mounted ZED Mini stereo camera streams stereo video back to the headset in real time, so the operator sees a 3D view from the robot's own viewpoint; arm and finger motion is mapped onto the robot through the dex-retargeting library. It targets the limitations of ordinary teleoperation — a fixed camera view, no sense of depth, and difficulty with fine manipulation. The authors collected demonstrations on a Unitree H1 and a Fourier GR-1, training policies with a modified version of ACT to perform tasks like sorting cans and folding clothes, and also demonstrated teleoperation across a distance of about 3,000 miles.","example":"The operator wears an Apple Vision Pro and turns their head toward the table; the Unitree H1's head unit turns to match. As the operator reaches out to grab something, the robot's dexterous hand sorts cans into different boxes in sync, and the whole process is recorded as imitation-learning demonstration data.","related":["VR Teleoperation","Apple Vision Pro","dex-retargeting","Action Chunking with Transformers","Unitree H1","PH2D"]},{"id":"unitree-xr-teleoperate","category":"data","sec":1,"tier":3,"sources":[{"title":"unitreerobotics/xr_teleoperate (GitHub)","url":"https://github.com/unitreerobotics/xr_teleoperate"}],"as_of":"2026-07","related_ids":["teleoperation","vr-teleoperation","open-television","unitree-g1","motion-retargeting","vuer"],"name":"Unitree xr_teleoperate","alt":"宇树 xr_teleoperate（XR 遥操作）","abbr":"","aliases":["avp_teleoperate"],"one_liner":"Unitree's open-source XR headset teleoperation program, used to control its humanoid robots and record training data.","explanation":"This is Unitree Robotics's open-source teleoperation codebase on GitHub, adapted from Open-TeleVision. The operator wears an Apple Vision Pro, PICO 4 Ultra Enterprise, or Meta Quest 3; the headset captures head and hand pose through a browser page (the Vuer interface, streaming video over WebRTC), and the program then uses inverse kinematics (computing joint angles from a target hand position) to drive the arms of humanoid robots such as the G1 and H1, mapping finger motion onto dexterous hands or grippers like the Dex3-1 or BrainCo, while streaming the robot's head-camera view back to the headset. It's mainly used to collect imitation-learning data, recording timestamped joint states, end-effector poses, and images for every episode. According to the repository, version 1.6 from July 2026 added support for the H2, R1, and BrainCo dexterous hands.","example":"A researcher wears an Apple Vision Pro to remotely control a Unitree G1 folding a towel, recording several dozen demonstrations along the way, then uses them to train ACT or a diffusion policy.","related":["Teleoperation","VR Teleoperation","Open-TeleVision","Unitree G1","Motion Retargeting","Vuer"]},{"id":"bunny-visionpro","category":"data","sec":1,"tier":3,"sources":[{"title":"Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning (arXiv 2407.03162)","url":"https://arxiv.org/html/2407.03162v1"},{"title":"Bunny-VisionPro 项目主页","url":"https://dingry.github.io/projects/bunny_visionpro.html"}],"as_of":"2024-07","related_ids":["vr-teleoperation","apple-vision-pro","anyteleop-a-general-vision-based-dexterous-robot-arm-hand-te","bimanual-manipulation","dexterous-manipulation","motion-retargeting"],"name":"Bunny-VisionPro","alt":"Bunny-VisionPro","abbr":"","aliases":["Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning"],"one_liner":"A teleoperation system that uses Apple Vision Pro to control two dexterous hands in real time, with vibration feedback, for imitation-learning data collection.","explanation":"Bunny-VisionPro is an open-source teleoperation system released in July 2024 by researchers at the University of Hong Kong and Xiaolong Wang's lab at UC San Diego, built for collecting bimanual dexterous-manipulation demonstrations for imitation learning. It uses an Apple Vision Pro headset to track the operator's hands and wrists in real time and maps that motion onto two xArm-7 robot arms fitted with two 6-DoF Ability dexterous hands (24 degrees of freedom in total). Finger retargeting uses an optimization method that aligns human and robot fingertip keypoints. The system includes built-in avoidance of collisions and singular configurations (poses where the arm loses the ability to move in some direction), and uses low-cost vibration motors to feed the robot hand's fingertip touch signals back to the operator. The paper used data collected with the system to train ACT, Diffusion Policy, and DP3, reaching higher success rates than data collected with baseline teleoperation systems.","example":"An operator wearing a Vision Pro headset remotely controls both arms and hands to complete long, multi-step tasks like wiping a window, sweeping the floor, or making coffee; when the robot's fingertips touch something, small vibration motors give the operator real-time touch feedback.","related":["VR Teleoperation","Apple Vision Pro","AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System","Bimanual Manipulation","Dexterous Manipulation","Motion Retargeting"]},{"id":"xrobotoolkit","category":"data","sec":1,"tier":3,"sources":[{"title":"arXiv 2508.00097: XRoboToolkit","url":"https://arxiv.org/abs/2508.00097"},{"title":"XRoboToolkit 项目主页","url":"https://xr-robotics.github.io"},{"title":"XR-Robotics/XRoboToolkit-Teleop-Sample-Python (GitHub)","url":"https://github.com/XR-Robotics/XRoboToolkit-Teleop-Sample-Python"}],"as_of":"2026-09","related_ids":["teleoperation","vr-teleoperation","unitree-xr-teleoperate","meta-quest-3-pico-4-ultra-xr-headsets","inverse-kinematics","demonstration-data"],"name":"XRoboToolkit","alt":"XRoboToolkit","abbr":"","aliases":["XRoboToolkit: A Cross-Platform Framework for Robot Teleoperation"],"one_liner":"PICO's open-source cross-platform XR teleoperation framework, controlling many kinds of robots with a headset and controllers to collect data.","explanation":"XRoboToolkit was developed by a team at PICO, the XR company owned by ByteDance (Zhigen Zhao, Liuchuan Yu, Ke Jing, and Ning Yang), with the paper posted to arXiv in July 2025, and won the Best Paper award at SII 2026. It targets the large amount of demonstration data VLA models need: on the headset side, it follows the OpenXR standard (a cross-vendor interface specification for XR devices), supporting head, controller, and hand tracking plus additional motion trackers (such as one strapped to the elbow to control elbow position); on the robot side, it provides Python and C++ interfaces, uses optimization-based inverse kinematics to convert the operator's hand pose into joint commands, and streams the robot's stereo camera view back to the headset with low latency. The framework is modular and has already been integrated with MuJoCo simulation, dual UR5 arms, mobile robots, and dexterous hands. The paper used data collected with it to train a VLA, validating its effectiveness on fine manipulation tasks.","example":"An operator wears a PICO 4 Ultra headset and uses two controllers to simultaneously control dual UR5 arms performing a precise connector-insertion task, seeing stereo video streamed back from the robot's cameras inside the headset.","related":["Teleoperation","VR Teleoperation","Unitree xr_teleoperate","Meta Quest 3 / PICO 4 Ultra XR Headsets","Inverse Kinematics (IK)","Demonstration Data"]},{"id":"nvidia-isaac-teleop","category":"data","sec":1,"tier":3,"sources":[{"title":"NVIDIA Isaac Teleop（GitHub，现为 NVIDIA/IsaacCapture）","url":"https://github.com/NVIDIA/IsaacCapture"},{"title":"Isaac Teleop Architecture 文档","url":"https://github.com/NVIDIA/IsaacCapture/blob/main/docs/source/overview/architecture.rst"}],"as_of":"2026-09","related_ids":["teleoperation","vr-teleoperation","motion-retargeting","nvidia-isaac-lab","mcap","homie"],"name":"NVIDIA Isaac Teleop","alt":"Isaac Teleop 遥操作框架","abbr":"","aliases":["Isaac Teleop","IsaacCapture"],"one_liner":"NVIDIA's open-source teleoperation and data-collection framework, with unified support for XR headsets, gloves, and other input devices.","explanation":"Isaac Teleop is NVIDIA's open-source (Apache 2.0) teleoperation and demonstration-data-collection framework; its first version was released in November 2025, bundled with Isaac Lab starting at Isaac Lab 3.0 Beta, and it also supports Isaac Sim, ROS 2, and real robots. It addresses the problem that every headset, glove, and robot combination otherwise needs its own integration code and produces data in its own format: it provides unified access to XR headsets such as Apple Vision Pro, PICO, and Quest, plus data gloves, foot pedals, and body trackers, all with synchronized timestamps; it uses composable retargeting modules to map human motion onto different robots; and it records and replays using MCAP, which interoperates with the LeRobot format. It already supports XR-controlled grippers or dexterous hands, seated whole-body teleoperation (HOMIE), SONIC-based whole-body teleoperation, and embodiment-free first-person data collection. According to GitHub, the repository was renamed NVIDIA/IsaacCapture in 2026, though the documentation still refers to it as Isaac Teleop.","example":"Wearing an XR headset such as Apple Vision Pro inside Isaac Lab, an operator teleoperates a simulated robot arm through a demonstration using hand motion; the recorded MCAP data is then converted into the LeRobot format for training a policy.","related":["Teleoperation","VR Teleoperation","Motion Retargeting","NVIDIA Isaac Lab","MCAP","HOMIE"]},{"id":"whole-body-teleoperation","category":"data","sec":1,"tier":2,"sources":[{"title":"Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation","url":"https://arxiv.org/abs/2401.02117"},{"title":"OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning","url":"https://arxiv.org/abs/2406.08858"},{"title":"TWIST: Teleoperated Whole-Body Imitation System","url":"https://arxiv.org/abs/2505.02833"}],"as_of":"2025-05","related_ids":["teleoperation","motion-retargeting","whole-body-control","omnih2o","twist","mobile-aloha"],"name":"Whole-Body Teleoperation","alt":"全身遥操作","abbr":"","aliases":["Humanoid Whole-body Teleop"],"one_liner":"One operator simultaneously controls a robot's arms, torso, and legs or base — its whole body at once.","explanation":"Ordinary teleoperation usually only controls the arm; whole-body teleoperation lets one operator simultaneously direct the robot's upper body, torso, and legs or mobile base, used to collect data for tasks that need the whole body to work together, such as reaching for something while walking or crouching to pick something up. There are two common forms: a mobile manipulation platform, such as Stanford's 2024 Mobile ALOHA, where the operator is physically tethered to the base, controlling both arms with a leader-follower setup and steering the base with their own body; and a humanoid robot, where mocap gear, a VR headset, or an exoskeleton captures the operator's whole-body motion, which is mapped through motion retargeting onto the robot and tracked by a reinforcement-learning-trained whole-body controller that keeps balance, as in OmniH2O (2024) and Stanford's TWIST (2025). The difficulty is that body proportions and balance capability differ between human and robot, and the operator usually has no force feedback either.","example":"In OmniH2O, an operator can wear just a single VR headset, driving a full-size humanoid's whole-body motion with head and hand pose alone; the recorded data forms the OmniH2O-6 dataset.","related":["Teleoperation","Motion Retargeting","Whole-Body Control","OmniH2O","Twist","Mobile ALOHA"]},{"id":"robot-free-data-collection","category":"data","sec":2,"tier":2,"sources":[{"title":"Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots","url":"https://arxiv.org/abs/2402.10329"},{"title":"UMI 项目主页","url":"https://umi-gripper.github.io/"},{"title":"DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation","url":"https://arxiv.org/abs/2505.21864"}],"as_of":"2025-10","related_ids":["universal-manipulation-interface","handheld-gripper-data-collection","dexumi","wearable-data-collection","capture-execution-isomorphism","embodiment-gap"],"name":"Robot-free (Embodiment-free) Data Collection","alt":"无本体采集","abbr":"","aliases":["Embodiment-free Collection"],"one_liner":"Collecting data with no real robot present, by having a human hold or wear a device shaped like the robot's end effector.","explanation":"Ordinary teleoperation requires a human to operate an actual robot, which is expensive equipment that's hard to bring into a real home. Robot-free collection instead has a person hold or wear a device designed to match the robot's end effector and perform the task directly; a camera and sensors on the device record the footage and motion, which is later converted into robot actions for training. The best example is UMI (Universal Manipulation Interface), proposed by Cheng Chi, Shuran Song, and colleagues in 2024: a handheld gripper with a wrist-mounted GoPro, each demonstration taking about 30 seconds, claimed to let collection start in any home or restaurant within two minutes, with the resulting policy deployable to different robot arms. DexUMI pushes the same idea to dexterous hands, using a hand exoskeleton to align kinematics and then inpainting a robot hand over the human hand in the video. The key is keeping the device's shape, viewpoint, and action capability as close as possible to the real robot, or an embodiment gap is left behind.","example":"UMI's cup-placement task: collectors carry a handheld gripper and demonstrate placing a cup on a saucer in different locations; back at the lab, that data trains a policy that deploys straight onto a robot arm.","related":["Universal Manipulation Interface","Handheld Gripper Data Collection","DexUMI","Wearable Data Collection","Capture-Execution Isomorphism","Embodiment Gap"]},{"id":"handheld-gripper-data-collection","category":"data","sec":2,"tier":2,"sources":[{"title":"Universal Manipulation Interface (UMI) 项目主页","url":"https://umi-gripper.github.io/"},{"title":"Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots (arXiv 2402.10329)","url":"https://arxiv.org/abs/2402.10329"}],"as_of":"","related_ids":["universal-manipulation-interface","fastumi","agilex-pika","robot-free-data-collection","capture-execution-isomorphism","dexumi"],"name":"Handheld Gripper Data Collection","alt":"手持夹爪采集","abbr":"","aliases":["UMI-style Collection"],"one_liner":"A human holds a camera-equipped gripper and performs tasks directly, recording footage and gripper motion as robot training data.","explanation":"Handheld gripper data collection is a way to gather data without a real robot present: a collector holds a gripper matching the exact model used on the robot's end effector, with a camera mounted on it, and performs the task directly in the real world; afterward, algorithms recover the gripper's pose trajectory and opening width from the video to serve as action labels. The best-known example is UMI, from Stanford, Columbia, and the Toyota Research Institute at RSS 2024: a wrist-mounted GoPro with a 155° fisheye lens and side mirrors uses visual-inertial SLAM (simultaneous localization and mapping) to track the trajectory. It is cheaper and more portable than teleoperation, letting data collection leave the lab entirely — UMI reports collection speeds more than 3 times faster than teleoperation. The precondition is that the gripper and camera viewpoint must match between collection and deployment.","example":"UMI uses a handheld gripper to collect demonstrations for dishwashing, two-arm cloth folding, and throwing objects into bins; the resulting policy deploys directly onto a robot arm and can complete tasks in environments and with objects it never saw during collection.","related":["Universal Manipulation Interface","FastUMI","AgileX Pika","Robot-free (Embodiment-free) Data Collection","Capture-Execution Isomorphism","DexUMI"]},{"id":"universal-manipulation-interface","category":"data","sec":2,"tier":1,"sources":[{"title":"Chi et al. 2024: Universal Manipulation Interface (arXiv 2402.10329)","url":"https://arxiv.org/abs/2402.10329"},{"title":"UMI 项目页","url":"https://umi-gripper.github.io/"}],"as_of":"2024-03","related_ids":["handheld-gripper-data-collection","robot-free-data-collection","fastumi","umi-on-legs","dexumi","rdt2"],"name":"Universal Manipulation Interface","alt":"通用操作接口","abbr":"UMI","aliases":["UMI","UMI Gripper","Handheld UMI Collection"],"one_liner":"A handheld gripper for recording demonstrations without a real robot, producing data that trains policies deployable on actual arms.","explanation":"UMI was proposed in February 2024 by researchers at Stanford, Columbia, and the Toyota Research Institute (Cheng Chi, Shuran Song, and colleagues). A person holds a 3D-printed parallel gripper and performs tasks directly; the gripper carries a 155° fisheye GoPro plus small mirrors on either side for extra viewpoints. Afterward, the visual SLAM algorithm ORB-SLAM3, combined with the GoPro's built-in IMU (inertial measurement unit), recovers the gripper's 6-degree-of-freedom pose in space, while marker tracking recovers how far the fingers opened, together yielding action labels. Because the handheld gripper matches the position of the gripper and camera mounted on a real robot, collection needs no robot present at all, so data can be gathered at home, outdoors, or anywhere in the real world. Paired with latency matching (compensating for the different camera and motor delays across robots) and a relative-trajectory action representation, the same policy can be deployed on both a UR5e and a Franka.","example":"The paper trained policies on UMI data that could set a cup down, wash dishes, toss objects into matching bins, and fold clothes with two arms. The gripper hardware costs about $73, the GoPro and accessories about $298, and on the cup-placing task it collected about 111 demonstrations per hour — more than 3 times faster than SpaceMouse (3D mouse) teleoperation.","related":["Handheld Gripper Data Collection","Robot-free (Embodiment-free) Data Collection","FastUMI","UMI on Legs","DexUMI","RDT2"]},{"id":"umi-on-legs","category":"data","sec":2,"tier":3,"sources":[{"title":"UMI on Legs 项目主页","url":"https://umi-on-legs.github.io/"},{"title":"arXiv 2407.10353: UMI on Legs","url":"https://arxiv.org/abs/2407.10353"}],"as_of":"2024-07","related_ids":["universal-manipulation-interface","handheld-gripper-data-collection","whole-body-control","diffusion-policy","mobile-manipulation","quadruped-robot"],"name":"UMI on Legs","alt":"UMI on Legs","abbr":"","aliases":["UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers"],"one_liner":"Training a manipulation policy on handheld-gripper human demonstrations, then deploying it on a quadruped fitted with a robot arm.","explanation":"UMI on Legs is a CoRL 2024 paper from July 2024 by Shuran Song's group at Stanford and Columbia University (Huy Ha, Yihuai Gao, and others). The hardware is a Unitree Go2 quadruped carrying a 6-axis ARX5 robot arm on its back. The approach splits into two halves: the manipulation skill is trained on demonstrations collected in the real world with a UMI handheld gripper (a gripper fitted with a GoPro, which needs no robot to collect data), producing a diffusion policy; coordinating the legs and arm is handled by a whole-body controller (a low-level controller driving both legs and arm together) trained with reinforcement learning in simulation. Only the “end-effector trajectory in task coordinates” is passed between the two halves, so a policy trained for a fixed-base robot arm can be mounted directly onto the quadruped and used as-is. The paper reports over 70% success across three task categories: throwing a ball, pushing a kettlebell, and arranging cups.","example":"For the quadruped to throw a ball into a bucket: the upper-level policy, trained on UMI data, outputs the gripper's target trajectory, while the whole-body controller adjusts the legs' posture, keeps balance, and carries out the throwing motion.","related":["Universal Manipulation Interface","Handheld Gripper Data Collection","Whole-Body Control","Diffusion Policy","Mobile Manipulation","Quadruped Robot"]},{"id":"fastumi","category":"data","sec":2,"tier":3,"sources":[{"title":"FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset (arXiv)","url":"https://arxiv.org/abs/2409.19499"},{"title":"FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset (arXiv)","url":"https://arxiv.org/abs/2510.08022"}],"as_of":"2025-10","related_ids":["universal-manipulation-interface","handheld-gripper-data-collection","robot-free-data-collection","agilex-pika","demonstration-data","visual-inertial-odometry"],"name":"FastUMI","alt":"FastUMI","abbr":"","aliases":["FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface","FastUMI-100K"],"one_liner":"A hardware-independent handheld-gripper data-collection system and dataset from Shanghai AI Lab, improving on UMI.","explanation":"FastUMI was proposed in September 2024, led by the Shanghai Artificial Intelligence Laboratory together with Shanghai Jiao Tong University, Fudan University, the University of Hong Kong, and others, as an improvement on UMI (Universal Manipulation Interface: a person holds a camera-equipped gripper to perform demonstrations, and the data is then transferred to a robot). The original UMI computes the gripper's pose by running visual-inertial odometry on GoPro footage, which is a fairly complex pipeline; FastUMI instead uses an off-the-shelf RealSense T265 tracking module to output pose directly at 200Hz, leaving the GoPro fisheye camera to handle only video capture, and adds standardized, swappable fingertips and camera mounts so the same device works across different robot arms and grippers. The team also open-sourced a dataset of 22 everyday tasks with over 10,000 trajectories; in October 2025 they released FastUMI-100K, covering 54 household tasks with more than 100,000 trajectories.","example":"A collector holds FastUMI and performs a single pick-and-place demonstration in a real home; the T265 records the gripper's trajectory, the GoPro records the video, and the resulting demonstration can be used to train any robot arm fitted with the matching fingertip.","related":["Universal Manipulation Interface","Handheld Gripper Data Collection","Robot-free (Embodiment-free) Data Collection","AgileX Pika","Demonstration Data","Visual-Inertial Odometry"]},{"id":"agilex-pika","category":"data","sec":2,"tier":3,"sources":[{"title":"agilexrobotics/pika_ros (GitHub)","url":"https://github.com/agilexrobotics/pika_ros"},{"title":"agilexrobotics/pika_sdk (GitHub)","url":"https://github.com/agilexrobotics/pika_sdk"},{"title":"AgileX Pika 产品页","url":"https://global.agilex.ai/products/pika"}],"as_of":"2026-09","related_ids":["handheld-gripper-data-collection","universal-manipulation-interface","robot-free-data-collection","capture-execution-isomorphism","agilex-robotics","fastumi"],"name":"AgileX Pika","alt":"松灵 Pika 采集套件","abbr":"","aliases":["Pika Sense","Pika Gripper","Pika Pro"],"one_liner":"AgileX Robotics' handheld data-collection kit that records manipulation demonstrations without needing a real robot.","explanation":"Pika is an embodied-AI data-collection product from AgileX Robotics, similar in concept to UMI: rather than using a real robot, a person holds a gripper-equipped capture device to perform the task, recording footage and motion at the same time. According to the official open-source repository, it consists of the Pika Sense capture device, the Pika Gripper (the same gripper used for inference deployment), a positioning base station, and a data backpack, recording 6-DoF pose, depth, ultra-wide RGB images, and gripper opening state; the SDK documentation shows Sense carries a fisheye camera, a RealSense depth camera, and a Vive Tracker for localization. The collection and deployment grippers share the same camera layout, so once a policy is trained on collected data, the Pika Gripper can be mounted on a robot arm's end effector to execute it directly. The newer Pika Pro kit adds an egocentric head-mounted capture device called Pika EGO.","example":"A typical workflow: a collector holds Pika Sense and repeatedly performs “put the cup in the sink” in a real kitchen, trains a diffusion policy on the recorded wrist footage and gripper pose, then mounts the Pika Gripper on a robot arm to execute it.","related":["Handheld Gripper Data Collection","Universal Manipulation Interface","Robot-free (Embodiment-free) Data Collection","Capture-Execution Isomorphism","AgileX Robotics","FastUMI"]},{"id":"wearable-data-collection","category":"data","sec":2,"tier":3,"sources":[{"title":"无本体数据采集技术演进：从 UMI、可穿戴采集到规模化交付（新浪科技，2026-09）","url":"https://finance.sina.com.cn/tech/roll/2026-09-17/doc-iniscqyc0291536.shtml"},{"title":"arXiv 2512.24310: World In Your Hands","url":"https://arxiv.org/abs/2512.24310"}],"as_of":"2026-09","related_ids":["robot-free-data-collection","handheld-gripper-data-collection","egocentric-video","data-glove","motion-retargeting","human-video-data"],"name":"Wearable Data Collection","alt":"可穿戴采集","abbr":"","aliases":[],"one_liner":"Having a person wear cameras, gloves, and similar devices while doing their normal job, recording their motion as robot training data.","explanation":"Wearable data collection is one of the main forms of embodiment-free data collection (collecting data without needing an actual robot). A collector wears a head- or chest-mounted camera and a wrist-mounted camera, often paired with a data glove, an IMU (inertial measurement unit, measuring acceleration and angular velocity), or tactile sensors, and does their normal job in real settings like factories, supermarkets, or homes, while the devices synchronously record first-person video, depth, and hand and body pose. Compared with teleoperation, it doesn't require moving a robot to the site, is cheaper, produces more natural motion, and scales up easily; the tradeoff is that a human hand's structure differs from a robot hand's, so the motion needs retargeting or additional modeling before a robot can use it. Compared with a handheld gripper like UMI, it captures richer hand motion, but the correspondence to a robot's end effector is weaker. The engineering challenges are multi-sensor time synchronization, calibration, recovering from occlusion, and unifying coordinate frames.","example":"TARS Robotics has collectors wear its own capture kit while working in factories, supermarkets, hotels, and other settings, recording over 1,000 hours of human manipulation data, organized into the WIYH dataset.","related":["Robot-free (Embodiment-free) Data Collection","Handheld Gripper Data Collection","Egocentric Video","Data Glove","Motion Retargeting","Human Video Data"]},{"id":"dexcap","category":"data","sec":2,"tier":3,"sources":[{"title":"DexCap 项目主页","url":"https://dex-cap.github.io"},{"title":"DexCap (arXiv)","url":"https://arxiv.org/abs/2403.07788"},{"title":"DexCap GitHub 仓库（RSS 2024）","url":"https://github.com/j96w/DexCap"}],"as_of":"2024-03","related_ids":["motion-capture","data-glove","robot-free-data-collection","dexumi","leap-hand","human-in-the-loop"],"name":"DexCap","alt":"DexCap","abbr":"","aliases":["DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation"],"one_liner":"Stanford's wearable hand motion-capture system that collects data by having a human use their own hands, for teaching robot dexterous hands.","explanation":"DexCap is a portable hand-motion-capture data collection system from Fei-Fei Li's group at Stanford (first author Chen Wang), published at RSS 2024, paired with DexIL, an algorithm for learning policies from human hand data. Teleoperating a dexterous hand to collect data is slow, expensive, and requires a robot to be physically present; DexCap instead lets a person wear a device and perform tasks directly with their own hands. Rokoko electromagnetic motion-capture gloves measure each finger's position relative to the palm and aren't disrupted by objects blocking the view; a chest-mounted rig carries an RGB-D LiDAR camera plus three SLAM tracking cameras that record wrist pose and a point cloud of the scene; a mini PC and battery pack in a backpack support about 40 minutes of continuous collection. DexIL uses inverse kinematics to convert the human hand motion into motion for a LEAP robot hand, then trains a policy with point-cloud-based imitation learning; a human can also step in to correct the policy during deployment.","example":"The paper demonstrates policies trained with only about 30 minutes of human hand motion-capture data and no teleoperation at all, including two-handed tasks like making tea and cutting things with scissors.","related":["Motion Capture","Data Glove","Robot-free (Embodiment-free) Data Collection","DexUMI","LEAP Hand","Human-in-the-Loop"]},{"id":"dexwild","category":"data","sec":2,"tier":3,"sources":[{"title":"DexWild 项目主页","url":"https://dexwild.github.io"},{"title":"DexWild (arXiv)","url":"https://arxiv.org/abs/2505.07813"}],"as_of":"2025-05","related_ids":["in-the-wild-data","co-training","data-glove","cross-embodiment","leap-hand","robot-free-data-collection"],"name":"DexWild","alt":"DexWild","abbr":"","aliases":["DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies","DexWild-System"],"one_liner":"CMU's portable hand-data-collection system that lets ordinary people gather dexterous-manipulation data with their own hands in real-world settings.","explanation":"DexWild is a dexterous-manipulation data-collection and training approach from Deepak Pathak's group at Carnegie Mellon University, published at RSS 2025. Teleoperated data is high quality but expensive, and it's hard to cover enough different environments this way; DexWild instead has ordinary people wear a low-cost, portable device called the DexWild-System and collect data with their own hands in real-world settings. A motion-capture glove measures finger pose; markers on the glove are located by a tracking camera to capture wrist position; two cameras mounted on the palm capture a close-up view in which the hand itself is barely visible, which makes the footage easier to reuse across different robot embodiments. Ten untrained collectors gathered 9,290 demonstrations across 93 environments, at roughly 4.6 times the speed of teleoperation. After co-training on this human hand data together with a smaller amount of robot data, the resulting policy reaches 68.5% success in unseen environments — about 4 times the success rate of training on robot data alone.","example":"A clothes-folding task was co-trained on about 1,124 human hand demonstrations plus 290 robot demonstrations; the training used a LEAP hand mounted on an xArm, and the resulting policy also transferred zero-shot to a Franka arm.","related":["In-the-wild Data","Co-training","Data Glove","Cross-Embodiment","LEAP Hand","Robot-free (Embodiment-free) Data Collection"]},{"id":"capture-execution-isomorphism","category":"data","sec":2,"tier":3,"sources":[{"title":"帕西尼 PXCap III 三指数据采集手套（官网）","url":"https://paxini.com/cn/ax/pxcap3"},{"title":"数据决定上限：25家国内具身智能数据采集厂商盘点（艾邦机器人）","url":"https://www.aibangbots.com/a/11921"},{"title":"UMI: Universal Manipulation Interface 项目主页","url":"https://umi-gripper.github.io/"}],"as_of":"2026-09","related_ids":["robot-free-data-collection","universal-manipulation-interface","motion-retargeting","embodiment-gap","leader-follower-teleoperation","paxini-tech"],"name":"Capture-Execution Isomorphism","alt":"采执同构","abbr":"","aliases":["Collection-Execution Isomorphism"],"one_liner":"A term from China's robotics-data industry: building the capture device to match the robot's end-effector so recorded data transfers directly, with no retargeting.","explanation":"Capture-execution isomorphism (采执同构) is a term that has become common in China's embodied-data industry in recent years. It means the capture device — a gripper or glove a human holds or wears, or a leader arm used for teleoperation — is built with the same kinematic structure, degrees of freedom, sensor layout, and even camera viewpoint as the robot's actual end-effector (its gripper or dexterous hand). That way the recorded joint angles, touch signals, and video can be used directly as the robot's observations and actions, skipping motion retargeting (converting human motion into robot motion) and the errors that conversion introduces, and narrowing the embodiment gap between different robots. ALOHA's matched leader-follower arms, and UMI's design that lets its handheld gripper share the wrist-camera viewpoint with the robot, reflect a similar idea. Companies such as PaXini Technology market it as a selling point, advertising a 1:1 physical match between their capture glove and their end-effector — “what you capture is what you deploy.” The tradeoff is that the capture device is tied to one specific end-effector; switching to a different robotic hand usually means switching capture devices too.","example":"PaXini's PXCap III three-fingered capture glove and its PXDex III three-fingered end-effector are designed with matching sensor layout, linkages, and joint degrees of freedom, so touch and joint data captured with the glove can be used directly on a robot fitted with the matching end-effector.","related":["Robot-free (Embodiment-free) Data Collection","Universal Manipulation Interface","Motion Retargeting","Embodiment Gap","Leader-Follower Teleoperation","PaXini Tech"]},{"id":"skill-capture-glove","category":"data","sec":2,"tier":3,"sources":[{"title":"Sunday Robotics 官网","url":"https://www.sunday.ai/"},{"title":"Sunday: ACT-1, a robot foundation model trained on zero robot data","url":"https://www.sunday.ai/journal/no-robot-data"},{"title":"Humanoids Daily: Sunday Unveils Memo, a Wheeled Domestic Robot That Learns From $200 Gloves","url":"https://www.humanoidsdaily.com/news/sunday-unveils-memo-a-wheeled-domestic-robot-that-learns-from-200-gloves"}],"as_of":"2026-09","related_ids":["sunday-robotics","sunday-robotics-memo","sunday-robotics-act-1","robot-free-data-collection","capture-execution-isomorphism","universal-manipulation-interface"],"name":"Skill Capture Glove","alt":"技能采集手套","abbr":"","aliases":["Sunday Robotics Capture Glove"],"one_liner":"Sunday Robotics's glove that lets a person do housework wearing it and directly produce robot training data.","explanation":"The Skill Capture Glove is a data-collection device from the U.S. home-robotics company Sunday Robotics, made public in November 2025 alongside its household robot, Memo; the company was founded by Tony Zhao (an author of ACT/ALOHA) and Cheng Chi (an author of Diffusion Policy/UMI). The glove shares the same geometry and sensor layout as Memo's robotic hand, so when a person wears it to do housework, the recorded motion and force data can be used almost directly as robot data; the remaining embodiment gap is handled by a process called Skill Transform, which the company says succeeds about 90% of the time. Each glove reportedly costs about $200, with a full teleoperation setup costing around $20,000. It's a representative example of embodiment-free, capture-execution-isomorphic data collection.","example":"Sunday says it has sent thousands of gloves to “Memory Developers” — people who wear them at home to record household chores — with data reportedly coming from about 500 households, used to train its foundation model, ACT-1, so that Memo learns chores like washing dishes, making coffee, and folding laundry.","related":["Sunday Robotics","Sunday Robotics Memo","Sunday Robotics ACT-1","Robot-free (Embodiment-free) Data Collection","Capture-Execution Isomorphism","Universal Manipulation Interface"]},{"id":"dexumi","category":"data","sec":2,"tier":3,"sources":[{"title":"DexUMI 项目主页","url":"https://dex-umi.github.io"},{"title":"DexUMI (arXiv)","url":"https://arxiv.org/abs/2505.21864"}],"as_of":"2025-05","related_ids":["universal-manipulation-interface","exoskeleton","embodiment-gap","dexcap","dexterous-hand","dexop"],"name":"DexUMI","alt":"DexUMI","abbr":"","aliases":["DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation"],"one_liner":"A wearable exoskeleton that lets a human hand collect dexterous-hand data directly, then digitally replaces the hand in footage with the robot hand.","explanation":"DexUMI is a dexterous-hand data collection and policy-learning framework from Shuran Song's group at Stanford, together with Columbia University, NVIDIA, and others, from 2025; it was a best-paper finalist at CoRL 2025. It extends the idea behind UMI (Universal Manipulation Interface, which collects data with a handheld gripper) to multi-fingered dexterous hands, using the human hand itself as the collection interface, and narrows the embodiment gap between human and robot hands in two ways. On the hardware side, a wearable exoskeleton is custom-built for the target robot hand, constraining human hand motion to what the robot hand can actually do while letting the operator feel contact directly; wrist pose is recorded with an iPhone, a wide-angle camera sits below the wrist, and the exoskeleton carries tactile sensors. On the software side, the human hand and exoskeleton are digitally erased from the footage and the background is inpainted, then an image of the robot hand in the matching pose is composited in, so the training footage matches what the robot will actually see at deployment. Across two robot hands, Inspire and XHand, it reaches 86% average success.","example":"On a tea-leaf-scooping task, the paper reports DexUMI's data-collection throughput is about 3.2 times that of conventional teleoperation.","related":["Universal Manipulation Interface","Exoskeleton","Embodiment Gap","DexCap","Dexterous Hand","DEXOP"]},{"id":"dexop","category":"data","sec":2,"tier":3,"sources":[{"title":"DEXOP 项目主页","url":"https://dex-op.github.io"},{"title":"DEXOP (arXiv)","url":"https://arxiv.org/abs/2509.04441"}],"as_of":"2025-09","related_ids":["exoskeleton","tactile-data","robot-free-data-collection","dexumi","teleoperation","contact-rich-manipulation"],"name":"DEXOP","alt":"DEXOP 数据采集装置","abbr":"","aliases":["DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation"],"one_liner":"MIT's passive hand exoskeleton that lets a human's own hand directly drive a robot hand to collect vision and touch data.","explanation":"DEXOP is a dexterous-hand data-collection device made public in September 2025 by the MIT Improbable AI Lab (Haozhi Qi, Pulkit Agrawal, and others). The authors propose a collection paradigm called perioperation: record real human manipulation while keeping the data as directly transferable to a robot as possible. DEXOP is a passive hand exoskeleton: a person's fingers are mechanically linked to the fingers of a robot hand, so whatever the person's hand does, the robot hand mirrors the same pose; the robot hand itself carries cameras and tactile sensors, so the data collected is the robot's own vision and touch data from the start. Compared with teleoperation, the operator feels contact force directly through their fingers, making the motion more natural, faster, and more accurate. The paper reports that, measured per unit of collection time, policies trained on DEXOP data clearly outperform those trained on teleoperation data.","example":"An operator wearing DEXOP performs contact-rich dexterous tasks while the cameras and tactile sensors on the robot hand record in sync; this data is used directly to train a policy for that same robot hand.","related":["Exoskeleton","Tactile Data","Robot-free (Embodiment-free) Data Collection","DexUMI","Teleoperation","Contact-rich Manipulation"]},{"id":"motion-capture","category":"data","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Motion capture","url":"https://en.wikipedia.org/wiki/Motion_capture"},{"title":"AMASS: Archive of Motion Capture as Surface Shapes","url":"https://amass.is.tue.mpg.de/"}],"as_of":"","related_ids":["optical-motion-capture","inertial-motion-capture","markerless-motion-capture-2","motion-retargeting","amass","data-glove"],"name":"Motion Capture","alt":"动作捕捉","abbr":"MoCap","aliases":["MoCap"],"one_liner":"Recording the motion of a person or object into a computer with high precision, using cameras or wearable sensors.","explanation":"Motion capture (mocap) is technology for recording human or object motion into a computer at high precision; it was first used heavily in film, game animation, and sports analysis. There are two main approaches. Optical mocap sticks reflective markers on the body and triangulates their positions with multiple infrared cameras, reaching millimeter-level accuracy or better — Vicon and OptiTrack are leading vendors. Inertial mocap instead straps IMUs (sensors that measure angular velocity and acceleration) to the body, needs no external cameras, and suits outdoor or large-scale use — Xsens is a leading example. There's also markerless mocap, which estimates pose from ordinary video alone. In embodied AI, mocap data, once passed through motion retargeting, can train whole-body motion tracking for humanoid robots, and is also used for teleoperation and to provide ground-truth pose for datasets.","example":"AMASS unifies 15 optical mocap datasets into the SMPL format — a human body mesh model that represents shape and pose with a small number of parameters — covering more than 300 subjects and over 10,000 motion sequences; it's commonly used to train humanoid robots' motion-tracking policies.","related":["Optical Motion Capture","Inertial Motion Capture","Markerless (Video-Based) Motion Capture","Motion Retargeting","AMASS (Archive of Motion Capture as Surface Shapes)","Data Glove"]},{"id":"optical-motion-capture","category":"data","sec":3,"tier":3,"sources":[{"title":"Motion capture - Wikipedia（Optical systems 一节）","url":"https://en.wikipedia.org/wiki/Motion_capture"},{"title":"Object Motion Guided Human Motion Synthesis（OMOMO 采集设置）","url":"https://arxiv.org/html/2309.16237"}],"as_of":"","related_ids":["motion-capture","inertial-motion-capture","markerless-motion-capture","vicon","optitrack","amass"],"name":"Optical Motion Capture","alt":"光学动捕","abbr":"","aliases":["Marker-Based Motion Capture","Optical MoCap"],"one_liner":"Tracking markers on the body with multiple infrared cameras and triangulating them into 3D motion.","explanation":"Optical motion capture is one of the two main approaches to motion capture: reflective markers (passive) or light-emitting LEDs (active) are attached to a person or object, surrounded by several calibrated infrared cameras; each camera sees the 2D position of the markers, and triangulation across cameras recovers 3D coordinates, which are then fit to a skeleton or body model to reconstruct motion. Leading vendors include Vicon and OptiTrack. It's highly accurate (active systems can reach about 0.1mm) with frame rates commonly above 120 fps, and is often treated as the “gold standard” for human motion data; the downsides are expensive equipment, the need for a dedicated space, and dropped markers when they're occluded. It's complementary to inertial motion capture (IMU-based, immune to occlusion but prone to drift) and markerless motion capture (estimating pose directly from video). Datasets such as AMASS and OMOMO were both captured with optical motion capture.","example":"The OMOMO dataset used 12 Vicon cameras at 120 frames per second to record subjects carrying objects, with 5 markers attached to each object so both the person's and the object's motion were tracked at once.","related":["Motion Capture","Inertial Motion Capture","Markerless Motion Capture","Vicon","OptiTrack","AMASS (Archive of Motion Capture as Surface Shapes)"]},{"id":"inertial-motion-capture","category":"data","sec":3,"tier":3,"sources":[{"title":"Motion capture - Wikipedia","url":"https://en.wikipedia.org/wiki/Motion_capture"},{"title":"Xsens Motion Capture","url":"https://www.xsens.com/products/motion-capture"}],"as_of":"","related_ids":["motion-capture","optical-motion-capture","inertial-measurement-unit","motion-retargeting","whole-body-teleoperation","xsens"],"name":"Inertial Motion Capture","alt":"惯性动捕","abbr":"","aliases":["IMU-Based Motion Capture","MoCap Suit"],"one_liner":"A motion-capture method that straps a set of IMUs to the body to measure each segment's rotation and reconstruct full-body pose.","explanation":"Inertial motion capture attaches more than a dozen inertial measurement units (IMUs, each containing a gyroscope, accelerometer, and magnetometer) to different segments of the body, measures each segment's orientation and acceleration, and then uses software to assemble these readings onto a human skeleton to reconstruct full-body pose. Xsens's system uses 17 wireless sensors; the Chinese company Noitom (诺亦腾) makes similar products. Compared with optical motion capture, which relies on multiple cameras tracking reflective markers, inertial motion capture needs no camera setup, isn't disrupted by occlusion, and can be used outdoors, in factories, or in any other setting; it's also cheaper. The tradeoff is that it can only directly compute relative pose — global position drifts over time — so it's less accurate than optical motion capture. In embodied AI, it's commonly used for whole-body teleoperation of humanoid robots and for collecting human-motion data: captured human motion goes through motion retargeting (converting human joint angles into robot joint angles) before either driving a robot to follow in real time or being used to train a policy.","example":"An operator wears an Xsens motion-capture suit; their full-body motion streams in real time into ROS or MuJoCo, and after retargeting it drives a humanoid robot to perform the same motion synchronously.","related":["Motion Capture","Optical Motion Capture","Inertial Measurement Unit","Motion Retargeting","Whole-Body Teleoperation","Xsens"]},{"id":"markerless-motion-capture-2","category":"data","sec":3,"tier":3,"sources":[{"title":"Motion capture - Wikipedia（Markerless 部分）","url":"https://en.wikipedia.org/wiki/Motion_capture"},{"title":"HumanPlus: Humanoid Shadowing and Imitation from Humans (arXiv 2406.10454)","url":"https://arxiv.org/abs/2406.10454"},{"title":"VideoMimic 项目主页","url":"https://www.videomimic.net/"}],"as_of":"2025-09","related_ids":["motion-capture","optical-motion-capture","human-mesh-recovery","gvhmr","humanplus","videomimic"],"name":"Markerless (Video-Based) Motion Capture","alt":"视频动捕（无标记动捕）","abbr":"","aliases":["Markerless Mocap","Monocular Motion Capture","Vision-Based Motion Capture"],"one_liner":"Estimating a person's 3D motion directly from ordinary video, with no markers or motion-capture suit needed.","explanation":"Video-based (markerless) motion capture doesn't rely on reflective markers or inertial sensors; it uses just one or a few ordinary cameras, with a computer-vision algorithm estimating the 3D pose and motion trajectory of the human body, including the hands. Compared with optical motion capture, it needs no dedicated studio or special clothing, and can work with footage shot on a phone or pulled from the internet, so it can gather human motion data at scale; the tradeoff is lower accuracy and more noise, and monocular video adds depth and scale ambiguity plus occlusion problems, usually requiring post-processing or physical constraints to clean up. Commonly used models include WHAM and GVHMR (from Zhejiang University, which recovers human motion in world coordinates from monocular video), with HaMeR commonly used for hands. In embodied AI, this is an entry point for humanoid robots learning motion from human video and for real-time teleoperation: the estimated human motion is mapped onto the robot through motion retargeting, then used to train a motion-tracking policy.","example":"HumanPlus uses a single RGB camera to estimate a person's body (with WHAM) and hand (with HaMeR) pose in real time, letting a humanoid robot follow the operator's movements; UC Berkeley's VideoMimic reconstructs both human motion and scene geometry from casually shot monocular video, training a humanoid robot to climb stairs, sit down, and stand up.","related":["Motion Capture","Optical Motion Capture","Human Mesh Recovery","GVHMR","HumanPlus","VideoMimic"]},{"id":"bvh-fbx-motion-capture-file-formats","category":"data","sec":3,"tier":3,"sources":[{"title":"Biovision BVH（威斯康星大学课程资料）","url":"https://research.cs.wisc.edu/graphics/Courses/cs-838-1999/Jeff/BVH.html"},{"title":"FBX - Wikipedia","url":"https://en.wikipedia.org/wiki/FBX"},{"title":"GMR: General Motion Retargeting (GitHub)","url":"https://github.com/YanjieZe/GMR"}],"as_of":"","related_ids":["motion-capture","motion-retargeting","general-motion-retargeting","lafan1","amass","smpl"],"name":"BVH / FBX Motion Capture File Formats","alt":"BVH / FBX 动捕文件格式","abbr":"","aliases":["BioVision Hierarchy","Filmbox",".bvh",".fbx"],"one_liner":"Two common motion-capture file formats that store a human skeleton's structure and its joint rotations frame by frame.","explanation":"BVH (BioVision Hierarchy) was created by the motion-capture company Biovision and is a plain-text format: a HIERARCHY section describes the skeleton as a tree of ROOT and JOINT nodes, giving each joint's offset and rotation channels relative to its parent, while a MOTION section gives the frame count and frame time, followed by one line per frame listing each joint's rotation (the root joint also gets a translation). BVH stores only skeletal motion, not mesh or material data. FBX was originally designed by Kaydara for its Filmbox motion-capture software and is now owned by Autodesk; it can store geometry, skeleton, animation, and materials together, but the format itself isn't publicly documented, so it's normally read and written through Autodesk's official SDK. When humanoid robots imitate human motion, researchers commonly retarget human motion from these two file formats onto the robot's joints — for example, GMR supports both BVH from the LAFAN1 dataset and FBX exported from OptiTrack.","example":"The publicly released LAFAN1 motion-capture dataset from Ubisoft ships as BVH files; researchers use GMR to retarget walking and dancing clips from it onto the joints of the Unitree G1 humanoid, as reference motion for a motion-tracking policy.","related":["Motion Capture","Motion Retargeting","General Motion Retargeting","LAFAN1","AMASS (Archive of Motion Capture as Surface Shapes)","SMPL"]},{"id":"amass","category":"data","sec":3,"tier":2,"sources":[{"title":"AMASS 官网","url":"https://amass.is.tue.mpg.de/"},{"title":"Mahmood et al. 2019: AMASS (arXiv 1904.03278)","url":"https://arxiv.org/abs/1904.03278"},{"title":"He et al. 2024: H2O: Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation","url":"https://arxiv.org/html/2403.04436"}],"as_of":"2019-10","related_ids":["smpl","motion-retargeting","motion-tracking","optical-motion-capture","lafan1","h2o"],"name":"AMASS (Archive of Motion Capture as Surface Shapes)","alt":"AMASS 人体动捕数据集","abbr":"AMASS","aliases":["AMASS"],"one_liner":"A large human-motion library that unifies 15 optical motion-capture datasets into the SMPL body-model format.","explanation":"AMASS was proposed by Mahmood, Black, and colleagues at the Max Planck Institute for Intelligent Systems in Germany, published at ICCV 2019. Before it, different optical mocap datasets each used their own marker layouts and skeleton definitions, making them hard to combine. The authors used a method called MoSh++ to fit raw marker data to sequences of the SMPL parametric body model (which describes a person's pose and shape with a few dozen parameters), unifying 15 datasets into more than 300 subjects, over 11,000 motion sequences, and more than 40 hours in total. In embodied AI it is a standard source of human motion for training humanoid robots' motion tracking and whole-body control: SMPL motions are first retargeted into robot joint trajectories, then a policy is trained in simulation with reinforcement learning to track them. It requires registration and is intended for research use.","example":"H2O retargeted about 13,000 motions from AMASS onto the Unitree H1, yielding 10,000 retargeted sequences, then used an imitation policy in simulation to filter out motions the robot physically couldn't perform, leaving about 8,500 sequences for training a whole-body humanoid teleoperation policy.","related":["SMPL","Motion Retargeting","Motion Tracking","Optical Motion Capture","LAFAN1","H2O"]},{"id":"lafan1","category":"data","sec":3,"tier":3,"sources":[{"title":"Ubisoft La Forge Animation Dataset（LAFAN1）GitHub","url":"https://github.com/ubisoft/ubisoft-laforge-animation-dataset"},{"title":"GMR: General Motion Retargeting（GitHub）","url":"https://github.com/YanjieZe/GMR"},{"title":"BeyondMimic whole_body_tracking（GitHub）","url":"https://github.com/HybridRobotics/whole_body_tracking"}],"as_of":"2020","related_ids":["amass","optical-motion-capture","motion-retargeting","general-motion-retargeting","beyondmimic","motion-tracking"],"name":"LAFAN1","alt":"LAFAN1 动捕数据集","abbr":"","aliases":["Ubisoft La Forge Animation Dataset","LaFAN1"],"one_liner":"Ubisoft La Forge's public human motion-capture dataset, commonly used as a source of motion for humanoid robots to imitate.","explanation":"LAFAN1 is an optical motion-capture dataset (optical motion capture: reflective markers on the body tracked by multiple cameras) released by Ubisoft's research division, La Forge, alongside its SIGGRAPH 2020 paper Robust Motion In-betweening. It contains 77 sequences from 5 actors, about 497,000 frames at 30 frames per second, totaling roughly 4.6 hours, with motions including walking, running, jumping, fighting, crawling, getting up after a fall, dancing, and climbing over obstacles; it's released as BVH skeletal animation files under the CC BY-NC-ND 4.0 license (attribution, non-commercial, no derivatives). It was originally meant for filling in in-between frames in game animation, but has been widely used in humanoid robotics in recent years: human motion is first mapped onto robot joints through motion retargeting, and a motion-tracking policy is then trained with reinforcement learning in simulation. Unitree has released a version of LAFAN1 retargeted onto its own humanoid robots, and the GMR retargeting tool directly supports its BVH files.","example":"UC Berkeley's open-source BeyondMimic code trains its tracking policy directly on Unitree's retargeted LAFAN1 motions; the authors say any motion in the dataset suitable for a real robot can be trained directly with no parameter tuning.","related":["AMASS (Archive of Motion Capture as Surface Shapes)","Optical Motion Capture","Motion Retargeting","General Motion Retargeting","BeyondMimic","Motion Tracking"]},{"id":"humanml3d","category":"data","sec":3,"tier":3,"sources":[{"title":"HumanML3D GitHub","url":"https://github.com/EricGuo5513/HumanML3D"}],"as_of":"","related_ids":["text-to-motion","amass","mdm","motion-retargeting","smpl","motion-tracking"],"name":"HumanML3D","alt":"HumanML3D 数据集","abbr":"","aliases":["Text-Annotated 3D Human Motion Dataset"],"one_liner":"A text-to-motion dataset pairing 14,000 clips of 3D human motion with 45,000 natural-language descriptions.","explanation":"HumanML3D comes from Guo et al.'s CVPR 2022 paper, Generating Diverse and Natural 3D Human Motions From Text. The authors drew 14,616 motion clips from two human motion-capture datasets, AMASS and HumanAct12, and had people write 3–4 English sentences describing each one, yielding 44,970 sentences totaling about 28.59 hours of motion, ranging from everyday actions to sports and dancing. Motion is standardized to a 22-joint skeleton at 20 frames per second, and the dataset is doubled in size through left-right mirroring. It's the most commonly used training and evaluation benchmark for text-to-motion generation (input a sentence, output a 3D motion clip); motion-diffusion models such as MDM all report results on it. For humanoid robots, one path to making a robot perform actions on command is to first generate human motion from text, then retarget that motion onto the robot.","example":"Given the input “a person walks forward and then sits down,” a model trained on HumanML3D generates a 3D skeletal motion of walking followed by sitting, which is then retargeted for a humanoid robot to track and execute.","related":["Text-to-Motion","AMASS (Archive of Motion Capture as Surface Shapes)","MDM (Motion Diffusion Model)","Motion Retargeting","SMPL","Motion Tracking"]},{"id":"omomo","category":"data","sec":3,"tier":3,"sources":[{"title":"Object Motion Guided Human Motion Synthesis (arXiv 2309.16237)","url":"https://arxiv.org/abs/2309.16237"},{"title":"Object Motion Guided Human Motion Synthesis (arXiv HTML 全文)","url":"https://arxiv.org/html/2309.16237"}],"as_of":"2023-09","related_ids":["human-object-interaction","optical-motion-capture","smpl","diffusion-model","omniretarget","motion-retargeting"],"name":"OMOMO","alt":"OMOMO 人-物交互动作数据集","abbr":"","aliases":["OMOMO Dataset","Object Motion Guided Human Motion Synthesis"],"one_liner":"About 10 hours of optical motion-capture data from Stanford, recording full-body motion while carrying everyday objects.","explanation":"The OMOMO dataset comes from a SIGGRAPH Asia 2023 paper by Jiaman Li, Jiajun Wu, and C. Karen Liu at Stanford. The authors used a 12-camera Vicon optical motion-capture rig to record 17 subjects carrying and dragging 15 everyday objects (a mop, a floor lamp, a chair, a table, boxes, and more), capturing full-body motion for about 10 hours in total, along with each object's 3D geometry, its own motion, and the human motion in SMPL-X format. The paper's method takes only the object's motion as input, and uses a conditional diffusion model to first predict hand positions, then generate full-body pose. This is one of the few datasets capturing full-body interaction with large objects, and it's now commonly used as reference motion for teaching humanoid robots whole-body manipulation like carrying, pushing, and pulling — OmniRetarget, for example, retargets it onto the Unitree G1.","example":"Given the motion trajectory of a chair being dragged, the OMOMO method generates a full-body human motion of bending down to grab the chair back and dragging it while walking.","related":["Human-Object Interaction","Optical Motion Capture","SMPL","Diffusion Model","OmniRetarget","Motion Retargeting"]},{"id":"foot-skating-floating-penetration","category":"data","sec":3,"tier":3,"sources":[{"title":"PHUMA: Physically Reliable Humanoid Locomotion Dataset (arXiv)","url":"https://arxiv.org/abs/2510.26236"},{"title":"Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking (arXiv)","url":"https://arxiv.org/abs/2510.02252"}],"as_of":"","related_ids":["motion-retargeting","general-motion-retargeting","phuma","motion-tracking","motion-capture","interpenetration"],"name":"Foot Skating / Floating / Penetration","alt":"脚滑 / 漂浮 / 穿地（动作数据伪影）","abbr":"","aliases":["Motion Retargeting Artifacts","Foot Sliding","Ground Penetration"],"one_liner":"Three common physically-impossible foot errors that appear when human motion is retargeted onto a robot.","explanation":"These are the three most common physically implausible artifacts that show up when human motion — from motion capture or video pose estimation — is retargeted (converted into robot joint trajectories). Foot skating is when a foot that should be planted firmly on the ground instead slides horizontally; floating is when a foot that should be touching the ground instead hovers above it; penetration is when a foot sinks below the ground surface. These mostly come from scaling errors caused by differences in height and leg-length proportions between the human and the robot, plus errors in the video pose estimation itself. They're not always obvious to the eye, but when a motion-tracking policy is trained in physics simulation, the robot simply can't reproduce these motions, which hurts balance learning and increases tracking error. The usual fix is to add constraints during retargeting — forcing foot height to match the ground closely on contact frames, damping horizontal foot velocity during contact, or filtering out problematic segments by threshold. Works such as GMR and PHUMA specifically handle these artifacts.","example":"PHUMA measures this with three metrics: a foot within 1cm of the ground on a contact frame counts as not floating, sinking less than 1cm counts as not penetrating, and horizontal velocity below 10cm/s counts as not skating.","related":["Motion Retargeting","General Motion Retargeting","PHUMA","Motion Tracking","Motion Capture","Interpenetration"]},{"id":"general-motion-retargeting","category":"data","sec":3,"tier":3,"sources":[{"title":"Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking (arXiv)","url":"https://arxiv.org/abs/2510.02252"},{"title":"GMR (GitHub)","url":"https://github.com/YanjieZe/GMR"}],"as_of":"2026-09","related_ids":["motion-retargeting","foot-skating-floating-penetration","twist","lafan1","amass","motion-tracking"],"name":"General Motion Retargeting","alt":"GMR 通用动作重定向","abbr":"GMR","aliases":["GMR","Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking"],"one_liner":"Stanford's open-source tool that retargets human motion onto many different humanoid robots in real time.","explanation":"GMR is a motion-retargeting method and codebase open-sourced in 2025 by Jiajun Wu and C. Karen Liu's group at Stanford (with Araújo, Yanjie Ze, and others). Motion retargeting converts human motion into robot joint angles, and the hard part is that humans and robots differ in body proportions and joint structure. GMR first specifies which body parts on the human correspond to which parts on the robot, aligns the two bodies' rest poses, scales each body part separately, and then solves joint angles with a two-stage inverse-kinematics optimization, which reduces foot skating, self-penetration, and sudden jumps in joint angle. Supported inputs include SMPL-X-format data such as AMASS, BVH files such as LAFAN1, OptiTrack FBX exports, live Xsens streams, and motion extracted from monocular video via GVHMR. The README lists support for 18 humanoid robots, running at 60–70 frames per second on an ordinary CPU. The paper shows it produces better results for training tracking policies than open-source alternatives like PHC and ProtoMotions, and it also serves as the retargeting module in the TWIST teleoperation system.","example":"A clip of a human running and jumping from the LAFAN1 motion-capture dataset is retargeted with GMR onto the Unitree G1 humanoid, and the result is used to train a motion-tracking policy in simulation to imitate it.","related":["Motion Retargeting","Foot Skating / Floating / Penetration","Twist","LAFAN1","AMASS (Archive of Motion Capture as Surface Shapes)","Motion Tracking"]},{"id":"omniretarget","category":"data","sec":3,"tier":3,"sources":[{"title":"OmniRetarget (arXiv 2509.26633)","url":"https://arxiv.org/abs/2509.26633"},{"title":"OmniRetarget 项目主页","url":"https://omniretarget.github.io/"}],"as_of":"2026-06","related_ids":["motion-retargeting","general-motion-retargeting","omomo","lafan1","loco-manipulation","unitree-g1"],"name":"OmniRetarget","alt":"OmniRetarget","abbr":"","aliases":["OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction"],"one_liner":"A data-generation method that retargets a human's interaction with objects and terrain onto a humanoid robot together, not just body pose.","explanation":"OmniRetarget is a motion-retargeting and data-generation method proposed in 2025 by Amazon FAR (Frontier AI and Robotics), together with MIT, UC Berkeley, Stanford, and CMU. Retargeting converts human motion into robot joint motion; when only body keypoints are aligned, tasks like carrying a box or climbing onto a platform often end up with the hands missing the box or the feet sinking into the floor. OmniRetarget uses an “interaction mesh” to jointly model the spatial and contact relationships among the human, the object, and the terrain, preserving those relationships while respecting the robot's joint limits, and it can also swap in different robots, terrains, or objects for data augmentation. After generating over 8 hours of trajectories from OMOMO, LAFAN1, and its own motion-capture data, only 5 reward terms were needed to train motions like box-carrying and climbing, up to about 30 seconds long, on a Unitree G1. The project page states it received the Best Conference Paper award at ICRA 2026.","example":"Motion-capture data of a person carrying a box is retargeted onto a Unitree G1 while keeping both hands pressed against the box and both feet planted on the ground; the resulting trajectories are then used to train a reinforcement-learning policy that can carry boxes on the real robot.","related":["Motion Retargeting","General Motion Retargeting","OMOMO","LAFAN1","Loco-manipulation","Unitree G1"]},{"id":"humanoid-x","category":"data","sec":3,"tier":3,"sources":[{"title":"Learning from Massive Human Videos for Universal Humanoid Pose Control (arXiv)","url":"https://arxiv.org/abs/2412.14172"},{"title":"UH-1 项目主页（PSI Lab）","url":"https://psi-lab.ai/UH-1/"}],"as_of":"2025-10","related_ids":["human-video-data","internet-video-data","motion-retargeting","text-to-motion","humanoid-robot","uh-1"],"name":"Humanoid-X","alt":"Humanoid-X 数据集","abbr":"","aliases":["UH-1 Dataset","Learning from Massive Human Videos for Universal Humanoid Pose Control"],"one_liner":"A humanoid-motion dataset built by extracting and captioning motion from massive amounts of internet human videos.","explanation":"Humanoid-X is a dataset built by the University of Southern California, UC Berkeley, and the Toyota Research Institute, described in the paper Learning from Massive Human Videos for Universal Humanoid Pose Control, made public in December 2024, with the paper as an oral presentation at Humanoids 2025. The pipeline: mine human-motion videos from the internet, automatically generate text descriptions, estimate the 3D pose of the person in each video, retarget that motion into joint targets for a humanoid robot, and train a control policy to turn those targets into motion the robot can actually execute. The final dataset has 163,800 samples and over 20 million humanoid robot poses, each entry carrying the source video, text, human pose, robot keypoints, and the resulting action. The UH-1 model trained on it takes a text instruction as input and outputs humanoid robot motion. Its significance is bypassing expensive teleoperation and motion capture altogether, expanding humanoid motion data directly from human video already available online.","example":"Given the instruction “wave hello” as input, UH-1 outputs a sequence of humanoid robot joint motions that execute a wave, in simulation or on a real robot.","related":["Human Video Data","Internet Video Data","Motion Retargeting","Text-to-Motion","Humanoid Robot","UH-1"]},{"id":"phuma","category":"data","sec":3,"tier":3,"sources":[{"title":"PHUMA: Physically Reliable Humanoid Locomotion Dataset (arXiv 2510.26236)","url":"https://arxiv.org/abs/2510.26236"},{"title":"PHUMA 论文 HTML 全文","url":"https://arxiv.org/html/2510.26236"}],"as_of":"2026-06","related_ids":["foot-skating-floating-penetration","humanoid-x","amass","motion-retargeting","motion-tracking","unitree-g1"],"name":"PHUMA","alt":"PHUMA 人形动作数据集","abbr":"","aliases":["Physically Reliable Humanoid Locomotion Dataset"],"one_liner":"A large-scale humanoid reference-motion dataset, filtered and retargeted under physical constraints.","explanation":"PHUMA is a humanoid-robot motion dataset released in October 2025 by Jaegul Choo's group at KAIST (the Korea Advanced Institute of Science and Technology). Motion extracted from internet video and then retargeted (such as Humanoid-X) often contains physically impossible artifacts like floating, ground penetration, and foot skating. PHUMA first applies physics-aware filtering, removing clips with excessive jitter, an unstable center of mass, or feet that never touch the ground; it then uses PhySINK (physically constrained inverse-kinematics retargeting) to stay close to the original motion while satisfying joint limits, keeping feet planted on the ground, and preventing the supporting foot from sliding. The data merges motion capture such as AMASS with several video sources, totaling 76,010 clips and about 73 hours, released in versions for both the Unitree G1 and the H1-2. Motion-tracking policies trained on it achieve higher success rates than ones trained on AMASS or Humanoid-X, and transfer zero-shot to a real G1.","example":"A general-purpose motion-tracking policy trained on PHUMA's G1 version lets a real Unitree G1 reproduce walking, turning, and other motions performed by a person in video.","related":["Foot Skating / Floating / Penetration","Humanoid-X","AMASS (Archive of Motion Capture as Surface Shapes)","Motion Retargeting","Motion Tracking","Unitree G1"]},{"id":"egocentric-video","category":"data","sec":4,"tier":1,"sources":[{"title":"Ego4D 官网","url":"https://ego4d-data.org/"},{"title":"EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video (arXiv 2505.11709)","url":"https://arxiv.org/abs/2505.11709"}],"as_of":"2025-05","related_ids":["exocentric-video","human-video-data","ego4d","egodex","project-aria-glasses","hand-pose-estimation"],"name":"Egocentric Video","alt":"第一人称视频","abbr":"Ego","aliases":["Ego","First-person Video","Egocentric View"],"one_liner":"Video filmed from a head- or glasses-mounted camera, showing the wearer's own point of view with the hands in frame.","explanation":"Egocentric video is filmed by a camera worn on the head or on a pair of glasses, so the frame looks like what the wearer's own eyes would see, with the hands and whatever they're handling usually near the center — the counterpart to third-person (exocentric) video. It's useful for robotics because the viewpoint is close to a robot's head-mounted camera and shows hand-manipulation detail clearly. Notable datasets include Ego4D (released 2022, led by Meta with 13 universities, covering more than 3,670 hours of daily activity) and Apple's EgoDex, captured with the Vision Pro and annotated with 3D hand-joint positions (829 hours). This kind of video carries no robot action labels, so it first has to be converted — through hand-pose estimation, or through latent-action methods that infer an abstract action from a pair of consecutive frames — before it can be used to train a policy.","example":"EgoDex records 194 kinds of tabletop manipulation tasks with an Apple Vision Pro, saving the 3D position of every finger joint at the same time, so the recording can be used directly as a hand-motion trajectory.","related":["Exocentric Video","Human Video Data","Ego4D","EgoDex","Project Aria Glasses","Hand Pose Estimation"]},{"id":"exocentric-video","category":"data","sec":4,"tier":3,"sources":[{"title":"Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives (arXiv)","url":"https://arxiv.org/abs/2311.18259"},{"title":"Ego-Exo4D 官网","url":"https://ego-exo4d-data.org/"}],"as_of":"","related_ids":["egocentric-video","third-person-camera","ego-exo4d","internet-video-data","human-pose-estimation","motion-retargeting"],"name":"Exocentric Video","alt":"第三视角视频","abbr":"Exo","aliases":["Third-Person Video","Exo"],"one_liner":"Video shot from a camera outside the performer's own body, looking at a person or robot from a bystander's vantage point.","explanation":"Exocentric video is the counterpart to first-person (egocentric) video, where the camera is worn on the performer's head or chest. In exocentric video, the camera sits outside the performer's body — a tripod-mounted camera, a security camera, or handheld footage shot by someone nearby. Most instructional videos and sports broadcasts on the internet fall into this category, and there's far more exocentric footage available online than egocentric footage. It shows full-body posture and the overall relationship between a person and their environment clearly, but hand detail is often blocked by the body or by objects, and the viewpoint doesn't match what a robot's own onboard camera sees. In embodied AI, it's commonly used to estimate human body motion, which is then retargeted for a humanoid robot to imitate, or mined for task procedures; the third-person cameras fixed beside a table during robot data collection also produce this kind of footage. Datasets like Ego-Exo4D record both viewpoints at once specifically to study the correspondence between them.","example":"GVHMR is used to estimate human body motion from a monocular exocentric video, and GMR then retargets that motion into joint trajectories for a humanoid robot.","related":["Egocentric Video","Third-Person Camera","Ego-Exo4D","Internet Video Data","Human Pose Estimation","Motion Retargeting"]},{"id":"internet-video-data","category":"data","sec":4,"tier":2,"sources":[{"title":"LAPA: Latent Action Pretraining from Videos (arXiv 2410.11758)","url":"https://arxiv.org/abs/2410.11758"},{"title":"Ego4D 官网","url":"https://ego4d-data.org/"}],"as_of":"","related_ids":["human-video-data","action-free-video","latent-action","lapa","egocentric-video","ego4d"],"name":"Internet Video Data","alt":"互联网视频数据","abbr":"","aliases":["Web Video Data"],"one_liner":"The vast supply of publicly available online video — large and cheap, but without action labels a robot can use.","explanation":"Internet video data refers to video from video platforms and public video datasets, mostly showing people doing chores, cooking, and using tools. Compared with robot data, it is orders of magnitude larger in scale, covers a far wider range of scenes and objects, and carries a great deal of knowledge about how objects move and how tasks break into steps. The difficulty is that it has no action labels: the frame shows a human hand rather than a robot arm, and there is no way to know what joint command each frame corresponds to. Common uses fall into three groups: pretraining visual representations (as in R3M); learning latent actions, which automatically infer an action-like code from the change between adjacent frames (as in LAPA); and training video-generation or world models that later infer actions from predicted frames. Egocentric video collections such as Ego4D are commonly used this way too.","example":"LAPA first uses a VQ-VAE to learn discrete latent actions from adjacent frames of unlabeled video, pretrains a VLA on those latent actions, then maps them to actual robot commands using only a small amount of real-robot data.","related":["Human Video Data","Action-free Video","Latent Action","LAPA","Egocentric Video","Ego4D"]},{"id":"action-free-video","category":"data","sec":4,"tier":2,"sources":[{"title":"Ko et al. 2023: Learning to Act from Actionless Videos through Dense Correspondences (AVDC)","url":"https://arxiv.org/abs/2310.08576"},{"title":"Ye et al. 2024: Latent Action Pretraining from Videos (LAPA)","url":"https://arxiv.org/abs/2410.11758"}],"as_of":"","related_ids":["latent-action-pretraining","human-video-data","internet-video-data","pseudo-action-labels","latent-action-model","imitation-from-observation"],"name":"Action-free Video","alt":"无动作标签视频","abbr":"","aliases":["Actionless Video"],"one_liner":"Video that has only images, no frame-by-frame action record, like human videos and web videos.","explanation":"Action-free video means an image sequence with no frame-by-frame action labels, commonly sourced from human-manipulation videos on the web, egocentric video, and robot footage where only the images were kept. It vastly outnumbers robot data with actions attached, and it carries a great deal of information about how objects move and how tasks are broken down, but it cannot be used directly for behavior cloning. Common uses fall into a few groups: training an inverse dynamics model to fill in pseudo action labels; learning discrete “latent actions” from adjacent-frame changes for pretraining, as in LAPA and Genie; first generating future frames and then inferring the action, as in UniPi and AVDC; or using it only to pretrain visual representations and world models. LAPA reports that a model pretrained this way outperformed a VLA trained with robot action labels on real-robot tasks.","example":"AVDC trains using only action-free RGB video: it first synthesizes video of a robot completing the task with a video-generation model, then infers how the robot should move from the dense correspondence (optical flow) between adjacent frames, validated on tabletop manipulation and navigation tasks.","related":["Latent Action Pretraining","Human Video Data","Internet Video Data","Pseudo Action Labels","Latent Action Model","Imitation from Observation"]},{"id":"pseudo-action-labels","category":"data","sec":4,"tier":3,"sources":[{"title":"Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (arXiv:2206.11795)","url":"https://arxiv.org/abs/2206.11795"},{"title":"DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv:2505.12705)","url":"https://arxiv.org/abs/2505.12705"}],"as_of":"","related_ids":["inverse-dynamics-model","latent-action-model","action-free-video","neural-trajectories","dreamgen","vpt"],"name":"Pseudo Action Labels","alt":"伪动作标签","abbr":"","aliases":["Pseudo-Action Labeling","Pseudo-Actions"],"one_liner":"Action labels inferred by a model from video, standing in for actions that were never actually recorded.","explanation":"When video itself doesn't record what action a human or robot took, a model can “guess” the action at each step from the surrounding frames, and those guesses can be used as labels to train a policy — this is a pseudo action label. There are two common approaches. One trains an inverse-dynamics model (which infers the action between two adjacent frames) on a small amount of action-labeled data, then uses it to label a huge amount of video in bulk — OpenAI's VPT did exactly this to add keyboard-and-mouse action labels to a large collection of Minecraft videos from the internet. The other uses a latent-action model to learn an abstract action encoding directly from video. NVIDIA's DreamGen uses both approaches to add pseudo actions to video generated by a world model, producing trainable “neural trajectories.” This lets video with no action labels be used for training too, though the labels carry error and usually still need fine-tuning on real data.","example":"VPT first had people play Minecraft while recording their keyboard and mouse actions, trained an inverse-dynamics model on that data, then used it to label pseudo actions across a large collection of internet videos, and finally ran behavioral cloning on that labeled data.","related":["Inverse Dynamics Model","Latent Action Model","Action-free Video","Neural Trajectories","DreamGen","VPT"]},{"id":"robotizing-human-videos-human-to-robot-video-translation","category":"data","sec":4,"tier":2,"sources":[{"title":"Phantom: Training Robots Without Robots Using Only Human Videos","url":"https://arxiv.org/abs/2503.00779"},{"title":"Masquerade: Learning from In-the-wild Human Videos using Data-Editing","url":"https://arxiv.org/abs/2508.09976"},{"title":"H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos","url":"https://arxiv.org/abs/2512.09406"}],"as_of":"2025-12","related_ids":["human-video-data","cross-painting","embodiment-gap","hand-pose-estimation","egocentric-video","phantom"],"name":"Robotizing Human Videos / Human-to-Robot Video Translation","alt":"人类视频机器人化（人→机视频转换）","abbr":"","aliases":["Human-to-Robot Video Translation"],"one_liner":"Editing footage of a human doing a task so it looks like a robot doing it, turning it into usable training data.","explanation":"Human video is plentiful and cheap, but it shows a human hand and arm rather than a robot arm — a large visual embodiment gap — so training a policy on it directly performs poorly. Robotizing human video uses image editing or video generation to swap the human for a robot: first estimate the hand's 3D pose to serve as an action label, then erase the human hand and arm and inpaint the background, and finally overlay a rendered robot arm or gripper following that same trajectory. Stanford's Bohg group published Phantom in 2025, training a policy deployable directly on a real robot using only such edited human video; the same group's Masquerade pretrains a visual encoder on 675,000 frames of edited video. A late-2025 method called H2R-Grounder instead uses a fine-tuned video diffusion model to generate the robot footage directly, requiring no paired human-robot data at all.","example":"Phantom records video of a human hand sweeping several objects together on a table, erases the hand, overlays a rendered robot arm, and trains a policy directly on this edited video that deploys straight to a real robot.","related":["Human Video Data","Cross-Painting","Embodiment Gap","Hand Pose Estimation","Egocentric Video","Phantom"]},{"id":"cross-painting","category":"data","sec":4,"tier":3,"sources":[{"title":"Mirage: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting (arXiv 2402.19249)","url":"https://arxiv.org/abs/2402.19249"},{"title":"BerkeleyAutomation/mirage (GitHub)","url":"https://github.com/BerkeleyAutomation/mirage"}],"as_of":"2024-07","related_ids":["cross-embodiment","embodiment-gap","zero-shot","robotizing-human-videos-human-to-robot-video-translation","generative-data-augmentation","forward-dynamics-model"],"name":"Cross-Painting","alt":"跨本体图像替换","abbr":"","aliases":["Mirage","Robot Embodiment Visual Swap"],"one_liner":"Digitally erasing the deployed robot from the camera feed and painting in the robot the policy was trained on, so a vision policy transfers without retraining.","explanation":"Cross-painting was introduced by UC Berkeley's Ken Goldberg lab together with Google DeepMind researchers in the Mirage method (RSS 2024). A vision-based policy that has only ever seen a “source” robot tends to fail once deployed on a differently shaped “target” robot, because the camera image looks different. Cross-painting fixes this at run time: it uses a segmentation mask to erase the target robot from the video feed and inpaint the background, then uses the robot's URDF (the file format that describes a robot's links and joints) to compute the joint angles that would place the source robot's end effector at the same pose, and renders the source robot into the scene — so the policy always “sees” the robot it was trained on. Differences in how the two robots actually move are compensated separately with a forward-dynamics model (which predicts where the end effector ends up after a given action). Mirage achieved zero-shot transfer across a Franka arm, a UR5 arm, and different grippers. A similar segment-and-render pipeline has also been used to turn human hands in video into robot hands.","example":"A policy trained only on data collected with a Franka arm is deployed on a UR5; at run time, the UR5 in the camera feed is replaced with a rendered Franka, and the policy completes tasks like grasping with no retraining.","related":["Cross-Embodiment","Embodiment Gap","Zero-shot","Robotizing Human Videos / Human-to-Robot Video Translation","Generative Data Augmentation","Forward Dynamics Model"]},{"id":"something-something-v2","category":"data","sec":4,"tier":3,"sources":[{"title":"The something something video database for learning and evaluating visual common sense (arXiv 1706.04261)","url":"https://arxiv.org/abs/1706.04261"},{"title":"Hugging Face: something_something_v2 数据集卡片","url":"https://huggingface.co/datasets/HuggingFaceM4/something_something_v2"},{"title":"LAPA: Latent Action Pretraining from Videos（项目页）","url":"https://latentactionpretraining.github.io/"}],"as_of":"2024-10","related_ids":["human-video-data","action-free-video","latent-action-pretraining","lapa","ego4d","epic-kitchens"],"name":"Something-Something V2","alt":"Something-Something V2 数据集","abbr":"SSv2","aliases":["20BN-Something-Something V2","Sth-Sth V2","SSv2"],"one_liner":"About 220,000 short videos of hands manipulating everyday objects, labeled with 174 action templates.","explanation":"Something-Something V2 is an action-recognition video dataset released by the company TwentyBN (20BN), now distributed for download by Qualcomm. It contains 220,847 short videos, each filmed by a crowdworker following a given action template such as “putting something into something” or “turning something upside down,” across 174 categories in total. The object in each template is abstracted as “something,” so a model can't guess the answer just by recognizing the object — it has to actually understand the motion and sequencing between the hand and the object — which is why it's commonly used to test a video model's understanding of temporal order. In embodied AI, it's treated as a source of human manipulation video: the videos carry no robot action labels, but they contain large amounts of hand-object interaction, useful for pretraining visual representations and for latent-action pretraining.","example":"LAPA does latent-action pretraining using only the roughly 220,000 human videos in SSv2, then fine-tunes on robot data; the project page reports it outperforms OpenVLA, which was trained on Bridge data, on average.","related":["Human Video Data","Action-free Video","Latent Action Pretraining","LAPA","Ego4D","EPIC-KITCHENS"]},{"id":"epic-kitchens","category":"data","sec":4,"tier":3,"sources":[{"title":"Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100 (arXiv)","url":"https://arxiv.org/abs/2006.13256"},{"title":"The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines (arXiv)","url":"https://arxiv.org/abs/2005.00343"},{"title":"EPIC-KITCHENS 官网","url":"https://epic-kitchens.github.io/"}],"as_of":"2025-06","related_ids":["egocentric-video","ego4d","human-video-data","hand-object-interaction","affordance-detection","data-annotation"],"name":"EPIC-KITCHENS","alt":"EPIC-KITCHENS 数据集","abbr":"","aliases":["EPIC-KITCHENS-100","EPIC-KITCHENS-55"],"one_liner":"A classic first-person video dataset of unscripted everyday cooking and kitchen tasks, filmed with head-mounted cameras.","explanation":"EPIC-KITCHENS is a first-person video dataset led by the University of Bristol in the UK. The first version, EPIC-KITCHENS-55, was released in 2018 with 55 hours from 32 participants. It was expanded in 2020 into EPIC-KITCHENS-100: participants wore head-mounted cameras while cooking and washing dishes in their own kitchens, unscripted, across 45 kitchens, totaling 100 hours, 20 million frames, and about 90,000 action segments. Annotations come from participants narrating afterward what they were doing, which is then organized into “verb + noun” action labels. Companion benchmarks cover action recognition, detection, anticipation, cross-modal retrieval, and unsupervised domain adaptation. Later additions layered on hand-object segmentation (VISOR), sound events (EPIC-Sounds), and 3D camera information (EPIC-Fields). In robotics, it's commonly used to study hand-object interaction, affordances, and pretraining visual representations.","example":"One action segment is labeled “take plate”: the video shows the participant lifting a plate off a shelf, and the label is built from the verb take and the noun plate.","related":["Egocentric Video","Ego4D","Human Video Data","Hand-Object Interaction","Affordance Detection","Data Annotation"]},{"id":"ego4d","category":"data","sec":4,"tier":2,"sources":[{"title":"Ego4D: Around the World in 3,000 Hours of Egocentric Video (arXiv 2110.07058)","url":"https://arxiv.org/abs/2110.07058"},{"title":"Ego4D 官网","url":"https://ego4d-data.org/"},{"title":"R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)","url":"https://arxiv.org/abs/2203.12601"}],"as_of":"2022-02","related_ids":["egocentric-video","human-video-data","ego-exo4d","r3m","egodex","pre-trained-visual-representation"],"name":"Ego4D","alt":"Ego4D 数据集","abbr":"","aliases":[],"one_liner":"A roughly 3,670-hour egocentric daily-video dataset collected by Meta with 13 universities.","explanation":"Ego4D was collected by Facebook AI (now Meta) together with a consortium of 13 universities; the paper was posted in October 2021 and published at CVPR 2022. More than 900 participants across 9 countries and 74 locations wore head-mounted cameras to record daily activities such as cooking, repairs, shopping, and socializing, totaling 3,670 hours, with some sequences also including audio, eye tracking, 3D scene meshes, and synchronized multi-camera video, plus benchmark tasks for episodic memory, hand-object interaction, social behavior, and future prediction. For robotics, it is one of the largest sources of human egocentric video and is commonly used to pretrain visual representations; but it has no native 3D hand-pose annotation and no robot actions, so it cannot be used for imitation directly. A license agreement is required before use.","example":"R3M pretrains on Ego4D video using time-contrastive learning and video-language alignment; once frozen, it serves as a visual module for a Franka arm, which learned several manipulation tasks from just 20 demonstrations.","related":["Egocentric Video","Human Video Data","Ego-Exo4D","R3M","EgoDex","Pre-trained Visual Representation"]},{"id":"ego-exo4d","category":"data","sec":4,"tier":3,"sources":[{"title":"Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives (arXiv)","url":"https://arxiv.org/abs/2311.18259"},{"title":"Ego-Exo4D 官网","url":"https://ego-exo4d-data.org/"}],"as_of":"2024-09","related_ids":["egocentric-video","exocentric-video","ego4d","project-aria-glasses","human-video-data","hand-pose-estimation"],"name":"Ego-Exo4D","alt":"Ego-Exo4D 数据集","abbr":"","aliases":[],"one_liner":"A large-scale video dataset led by Meta, capturing skilled human activities simultaneously from first-person and third-person viewpoints.","explanation":"Ego-Exo4D is a multi-view video dataset released in 2023, led by Meta FAIR together with more than a dozen universities worldwide: about 1,286 hours of footage from 740 participants across 13 cities. During recording, participants wore Aria smart glasses to capture first-person (egocentric) video while 4–5 surrounding GoPro cameras captured third-person (exocentric) footage at the same time, covering 8 categories of skilled activity such as cooking, bike repair, dancing, and rock climbing, along with gaze data, IMU (inertial measurement unit) readings, 3D point clouds, camera poses, and expert narration transcripts. It's used to study what the same action looks like from the performer's own eyes versus from a bystander's viewpoint. For embodied AI, a robot's head- and wrist-mounted cameras see something close to a first-person view, while most instructional video on the internet is shot in third person, so this dataset can be used to learn the correspondence between the two viewpoints, as well as for hand and 3D human pose estimation.","example":"The same “bike repair” session is filmed simultaneously by Aria glasses and 4 GoPro cameras, letting researchers train a model to translate third-person footage into a first-person view — one of the cross-view translation benchmarks in the paper.","related":["Egocentric Video","Exocentric Video","Ego4D","Project Aria Glasses","Human Video Data","Hand Pose Estimation"]},{"id":"project-aria-glasses","category":"data","sec":4,"tier":3,"sources":[{"title":"Project Aria Glasses（官网规格页）","url":"https://www.projectaria.com/glasses/"},{"title":"Project Aria: A New Tool for Egocentric Multi-Modal AI Research (arXiv:2308.13561)","url":"https://arxiv.org/abs/2308.13561"},{"title":"EgoMimic: Scaling Imitation Learning via Egocentric Video (arXiv:2410.24221)","url":"https://arxiv.org/abs/2410.24221"}],"as_of":"2026-09","related_ids":["egocentric-video","human-video-data","egomimic","nymeria","hot3d","wearable-data-collection"],"name":"Project Aria Glasses","alt":"Project Aria 眼镜","abbr":"","aliases":["Aria Glasses","Aria Gen 2"],"one_liner":"Meta's research eyewear for collecting first-person data; it only records, it doesn't display anything to the wearer.","explanation":"Project Aria is a research project from Meta Reality Labs, centered on a pair of sensor-laden glasses that look close to ordinary eyewear, used to collect multimodal first-person data. The first generation carries 1 RGB camera, 2 SLAM cameras (for localization and mapping), 2 eye-tracking cameras, plus an IMU, microphones, and more; the second generation increases the SLAM cameras to 4, adds a heart-rate (PPG) sensor and a contact microphone, runs 6–8 hours on a charge, and can perform localization, hand tracking, and eye tracking directly on the device. Meta provides it to academic and industry partners by application, and its website states more than 200 partners so far. In embodied AI, it's mainly used to cheaply collect first-person human manipulation video and hand trajectories for robot learning; datasets such as Nymeria and HOT3D were both collected with it.","example":"Georgia Tech's EgoMimic has people wear Aria glasses while performing tabletop manipulation, then trains a policy on the recorded first-person video and hand trajectories together with a small amount of robot data.","related":["Egocentric Video","Human Video Data","EgoMimic","Nymeria","HOT3D","Wearable Data Collection"]},{"id":"nymeria","category":"data","sec":4,"tier":3,"sources":[{"title":"Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild (arXiv 2406.09905)","url":"https://arxiv.org/abs/2406.09905"},{"title":"Project Aria - Nymeria Dataset","url":"https://www.projectaria.com/datasets/nymeria/"},{"title":"facebookresearch/nymeria_dataset (GitHub)","url":"https://github.com/facebookresearch/nymeria_dataset"}],"as_of":"2026-06","related_ids":["egocentric-video","project-aria-glasses","inertial-motion-capture","ego-exo4d","human-video-data","motion-capture"],"name":"Nymeria","alt":"Nymeria 数据集","abbr":"","aliases":["Nymeria Dataset","NymeriaPlus"],"one_liner":"Meta's large-scale first-person dataset of everyday human motion, recorded with Aria glasses and a motion-capture suit.","explanation":"Nymeria is a first-person human-motion dataset released by Meta's Project Aria team, published at ECCV 2024. 264 participants performed everyday activities at 50 real-world locations, totaling about 300 hours: Aria smart glasses on the head recorded first-person video, eye gaze, and IMU data; a miniAria wristband was worn on the wrist; an XSens inertial motion-capture suit provided full-body ground-truth motion; observers filmed a third-person view as well, and everything is paired with hierarchical language descriptions. It fills a data gap around reconstructing what a person's entire body is doing from head-worn devices alone, and can be used for full-body pose estimation, motion generation, and action recognition. It's released under CC BY-NC 4.0 (non-commercial); an expanded version, NymeriaPlus, followed later.","example":"First-person video and IMU signals recorded with Aria glasses are used to train a model to estimate the wearer's full-body pose, with the synchronized XSens motion-capture data serving as ground truth to measure the error.","related":["Egocentric Video","Project Aria Glasses","Inertial Motion Capture","Ego-Exo4D","Human Video Data","Motion Capture"]},{"id":"dexycb","category":"data","sec":4,"tier":3,"sources":[{"title":"DexYCB 项目主页","url":"https://dex-ycb.github.io"},{"title":"DexYCB: A Benchmark for Capturing Hand Grasping of Objects (arXiv)","url":"https://arxiv.org/abs/2104.04631"}],"as_of":"2021-04","related_ids":["ycb-object-and-model-set","mano","hand-pose-estimation","6d-object-pose-estimation","human-robot-handover","dex-retargeting"],"name":"DexYCB","alt":"DexYCB 数据集","abbr":"","aliases":["DexYCB: A Benchmark for Capturing Hand Grasping of Objects"],"one_liner":"NVIDIA's multi-view dataset of human hands grasping objects, with 3D pose annotations for both the hand and the object.","explanation":"DexYCB is a real human-hand object-grasping dataset released by NVIDIA and the University of Washington, published at CVPR 2021. The authors used 8 synchronized RealSense D415 depth cameras to film 10 subjects grasping 20 YCB objects (YCB is a standard set of everyday objects widely used in robotics research) from multiple angles, yielding 1,000 sequences and about 580,000 RGB-D frames in total. Hand pose is annotated using the MANO parametric hand model, and each object is given a 6D pose (3D position plus 3D orientation). It's used as a benchmark for 2D detection, 6D object pose estimation, and 3D hand pose estimation, and the paper also proposes an evaluation for generating safe grasps when a person hands an object to a robot. In embodied AI, it's commonly used as a source of human grasping demonstrations, which are converted into robot dexterous-hand trajectories through motion retargeting.","example":"Using dex-retargeting's position-retargeting mode, the MANO hand poses of people grasping objects in DexYCB are converted into grasping trajectories for robot hands such as Allegro and Shadow.","related":["YCB Object and Model Set","MANO","Hand Pose Estimation","6D Object Pose Estimation","Human-Robot Handover","dex-retargeting"]},{"id":"hoi4d","category":"data","sec":4,"tier":3,"sources":[{"title":"HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction (arXiv)","url":"https://arxiv.org/abs/2203.01577"},{"title":"HOI4D 项目主页","url":"https://hoi4d.github.io/"}],"as_of":"","related_ids":["hand-object-interaction","egocentric-video","category-level-pose-estimation","hand-pose-estimation","articulated-object","hot3d"],"name":"HOI4D","alt":"HOI4D 数据集","abbr":"","aliases":["HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction"],"one_liner":"A first-person 4D hand-object interaction dataset with 2.4 million RGB-D frames, from Tsinghua and collaborators.","explanation":"HOI4D was released jointly by Tsinghua University, Peking University, and the Shanghai Qi Zhi Institute, published at CVPR 2022. Collectors wore a helmet fitted with two RGB-D cameras, a Kinect v2 and a RealSense D455, recording themselves manipulating objects from a first-person viewpoint — 4,000 sequences and 2.4 million frames in total, covering 16 categories and 800 object instances (7 rigid-body categories and 9 categories of articulated objects with moving parts, such as laptops, cabinets, and scissors), across 610 indoor rooms. Frame-by-frame annotations include panoptic segmentation, motion segmentation, 3D hand pose, category-level object pose (estimating pose when only the object's category, not the specific instance, has been seen before), and action labels, along with object meshes and scene point clouds. Its purpose is to let a model learn “how a hand interacts with a category of object” from a human viewpoint, useful for hand-object interaction understanding and pose tracking, and it also serves as material for robots learning manipulation from human data.","example":"Using the “open a laptop” sequences in HOI4D to train category-level pose tracking lets a model keep estimating the pose of a laptop it has never seen before, even while the hand partly blocks it from view.","related":["Hand-Object Interaction","Egocentric Video","Category-Level Pose Estimation","Hand Pose Estimation","Articulated Object","HOT3D"]},{"id":"oakink","category":"data","sec":4,"tier":3,"sources":[{"title":"OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object Interaction (arXiv 2203.15709)","url":"https://arxiv.org/abs/2203.15709"},{"title":"OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion (arXiv 2403.19417)","url":"https://arxiv.org/abs/2403.19417"},{"title":"OakInk2 项目主页","url":"https://oakink.net/v2/"}],"as_of":"2024-12","related_ids":["hand-object-interaction","mano","affordance","bimanual-manipulation","dexycb","arctic-a-dataset-for-dexterous-bimanual-hand-object-manipula"],"name":"OakInk","alt":"OakInk 数据集","abbr":"","aliases":["OakInk2","OakInk Dataset"],"one_liner":"A Shanghai Jiao Tong University hand-object interaction dataset, annotated with object use and two-handed manipulation.","explanation":"OakInk is a series of hand-object interaction datasets from Cewu Lu's group at Shanghai Jiao Tong University. In the first version (CVPR 2022), the “Oak” part annotates affordances (how an object can be used) for 1,800 everyday objects, while the “Ink” part records real human grasps on 100 of those objects and transfers them to virtual objects, yielding about 50,000 interactions labeled with intended use. The second version, OakInk2 (CVPR 2024), shifts to two-handed completion of complex tasks: 627 sequences and 4.01 million multi-view image frames, annotating the 3D pose of the body, both hands, and the objects, and organizing tasks into three layers — affordance, atomic task, and complex task. It's commonly used for hand pose estimation, grasp generation, and synthesizing two-handed motion.","example":"One OakInk2 sequence breaks a complex task down into atomic subtasks like picking up, pouring, and setting down; every frame is labeled with the 3D pose of both hands (as MANO parameters) and the object.","related":["Hand-Object Interaction","MANO","Affordance","Bimanual Manipulation","DexYCB","ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation"]},{"id":"arctic-a-dataset-for-dexterous-bimanual-hand-object-manipula","category":"data","sec":4,"tier":3,"sources":[{"title":"ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation (arXiv 2204.13662)","url":"https://arxiv.org/abs/2204.13662"},{"title":"ARCTIC 项目主页","url":"https://arctic.is.tue.mpg.de/"}],"as_of":"2023-06","related_ids":["hand-object-interaction","mano","smpl","optical-motion-capture","articulated-object","hoi4d"],"name":"ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation","alt":"ARCTIC 数据集","abbr":"","aliases":[],"one_liner":"A motion-capture dataset of two-handed dexterous manipulation of articulated objects, with precise 3D meshes for hands and objects every frame.","explanation":"ARCTIC was released in 2022 by ETH Zurich, the Max Planck Institute for Intelligent Systems, and the University of Amsterdam, published at CVPR 2023. Ten subjects used both hands to manipulate 11 articulated objects (objects with moving parts, such as scissors and a laptop), producing 339 sequences and 2.1 million frames of imagery, filmed synchronously from 8 fixed viewpoints and 1 egocentric viewpoint. The team used a 54-camera infrared Vicon optical motion-capture system to track small markers attached to the hands and objects, and solved for a MANO hand mesh, an SMPL-X body mesh, object pose and joint angle, and hand-object contact information for every frame. Earlier hand-object interaction datasets were mostly single-handed, rigid-object, and slow-motion; ARCTIC fills the gap with precise 3D ground truth for fast, two-handed manipulation of moving objects, and can also serve as a reference for learning dexterous manipulation from human data.","example":"ARCTIC contains sequences of both hands opening a laptop and opening and closing scissors, with the 3D mesh of both hands and the rotation angle of the object's moving part given for every frame.","related":["Hand-Object Interaction","MANO","SMPL","Optical Motion Capture","Articulated Object","HOI4D"]},{"id":"hot3d","category":"data","sec":4,"tier":3,"sources":[{"title":"HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos (arXiv)","url":"https://arxiv.org/abs/2411.19167"},{"title":"HOT3D 项目主页","url":"https://facebookresearch.github.io/hot3d/"}],"as_of":"2025-06","related_ids":["hand-object-interaction","egocentric-video","project-aria-glasses","mano","optical-motion-capture","hand-pose-estimation"],"name":"HOT3D","alt":"HOT3D 数据集","abbr":"","aliases":["HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos"],"one_liner":"Meta's first-person hand-object interaction dataset, recorded with Aria glasses and Quest 3, with motion-capture ground truth.","explanation":"HOT3D is a first-person hand-object interaction dataset released by Meta, with the paper appearing at CVPR 2025. 19 participants manipulated 33 rigid objects in staged kitchen, office, and living-room settings, recorded with two kinds of device, Project Aria research glasses and a Quest 3 headset, for a total of over 833 minutes and 3.7 million images. The data includes multi-view RGB and monochrome images, eye gaze, scene point clouds, and 3D poses for the cameras, both hands, and the objects; the pose ground truth comes from a motion-capture system using optical markers, which is more accurate than pose estimated algorithmically. Hand annotations are provided in both UmeTrack and MANO format (a commonly used parametric hand model), and objects come with 3D meshes carrying PBR (physically based rendering) materials. It's mainly used for research on 3D hand and object tracking and pose estimation, and also serves as a data source for learning dexterous manipulation from first-person human video. Using it requires agreeing to the HOT3D license.","example":"","related":["Hand-Object Interaction","Egocentric Video","Project Aria Glasses","MANO","Optical Motion Capture","Hand Pose Estimation"]},{"id":"egomimic","category":"data","sec":4,"tier":3,"sources":[{"title":"EgoMimic: Scaling Imitation Learning via Egocentric Video (arXiv)","url":"https://arxiv.org/abs/2410.24221"},{"title":"EgoMimic 项目主页","url":"https://egomimic.github.io/"}],"as_of":"2024-10","related_ids":["egocentric-video","project-aria-glasses","co-training","human-video-data","embodiment-gap","egoverse"],"name":"EgoMimic","alt":"EgoMimic","abbr":"","aliases":["EgoMimic: Scaling Imitation Learning via Egocentric Video"],"one_liner":"A framework that co-trains a single policy on Project Aria human-hand data and robot data together.","explanation":"EgoMimic is an imitation-learning framework proposed in 2024 by Danfei Xu's group at Georgia Tech. A human performs tasks while wearing Project Aria smart glasses, which record first-person video and 3D hand trajectories; this is paired with a low-cost bimanual robot deliberately designed to have kinematics close to human arms. During training, the human hand trajectories and robot actions are normalized into a shared representation, masks are applied over the human hands and the robot arms in the video to reduce visual appearance differences, and a single policy network is co-trained on both. In experiments, it outperforms imitation-learning methods trained on robot data alone on long-horizon, single-arm, and bimanual tasks, and generalizes to new scenes; the authors also found that adding 1 more hour of human hand data helps more than adding 1 more hour of robot data. It's a representative example of using cheap human first-person data to replace part of the teleoperation data needed; the same group later led the EgoVerse data platform.","example":"On a place-object-in-bowl task, EgoMimic trained on 2 hours of robot data plus 1 hour of human hand data clearly outperforms ACT trained on 3 hours of robot data alone.","related":["Egocentric Video","Project Aria Glasses","Co-training","Human Video Data","Embodiment Gap","EgoVerse"]},{"id":"ph2d","category":"data","sec":4,"tier":3,"sources":[{"title":"Humanoid Policy ~ Human Policy (arXiv 2503.13441)","url":"https://arxiv.org/abs/2503.13441"},{"title":"Humanoid Policy ~ Human Policy 项目主页","url":"https://human-as-robot.github.io/"}],"as_of":"2025-09","related_ids":["egocentric-video","human-video-data","co-training","embodiment-gap","open-television","unitree-h1"],"name":"PH2D","alt":"PH2D 数据集（HAT）","abbr":"PH2D","aliases":["HAT (Human Action Transformer)","Humanoid Policy ~ Human Policy"],"one_liner":"First-person human manipulation data collected with a VR headset, co-trained with humanoid robot data to train one policy.","explanation":"PH2D (Physical Human-Humanoid Data) comes from the paper Humanoid Policy ~ Human Policy, led by UC San Diego together with CMU, MIT, Apple, and others, published at CoRL 2025. The authors had people perform manipulation tasks directly while wearing consumer VR headsets like Apple Vision Pro or Meta Quest 3, automatically recording first-person video plus the 3D position of the head, wrists, and fingertips — about 27,000 language-annotated demonstrations in total. The companion model, HAT (Human Action Transformer), places humans and humanoid robots into one shared state-action space, and its output is then retargeted onto the robot's joints. The underlying idea: a person collecting their own data is far faster than teleoperating a robot, and as long as the representation is aligned, this data can be co-trained with a smaller amount of robot data to improve generalization and robustness.","example":"A person wearing a Vision Pro quickly records a large number of pick-and-place demonstrations, which are combined with a smaller set of Unitree H1 teleoperation data to train HAT; the resulting policy is deployed directly on the H1.","related":["Egocentric Video","Human Video Data","Co-training","Embodiment Gap","Open-TeleVision","Unitree H1"]},{"id":"egodex","category":"data","sec":4,"tier":2,"sources":[{"title":"EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video (arXiv 2505.11709)","url":"https://arxiv.org/abs/2505.11709"},{"title":"apple/ml-egodex (GitHub)","url":"https://github.com/apple/ml-egodex"}],"as_of":"2026-03","related_ids":["egocentric-video","human-video-data","apple-vision-pro","hand-pose-estimation","ego4d","dexterous-manipulation"],"name":"EgoDex","alt":"EgoDex 数据集","abbr":"","aliases":[],"one_liner":"An Apple dataset of egocentric manipulation video captured with Vision Pro, annotated with precise 3D hand joints.","explanation":"EgoDex was released by Apple's research team in May 2025, and the paper was accepted at ICLR 2026. Collectors wore an Apple Vision Pro to perform everyday tabletop manipulation; the headset's multiple calibrated cameras and on-device SLAM (simultaneous localization and mapping) computed the 3D position and orientation of the head, upper body, and 25 joints per hand, in sync, while recording. The dataset totals 829 hours, about 90 million frames, and 338,000 demonstrations, covering 194 kinds of tabletop tasks, with language descriptions generated by GPT-4. It fills a gap left by video datasets like Ego4D, which lack precise hand pose, and can be used to train hand-trajectory prediction models that then transfer to dexterous-hand robots. The data is released under a CC BY-NC-ND license, for non-commercial use only.","example":"The paper trains an imitation-learning policy on EgoDex that takes an egocentric frame as input and predicts the 3D motion trajectory of both hands over the following span of time, and uses this to build a hand-trajectory-prediction benchmark.","related":["Egocentric Video","Human Video Data","Apple Vision Pro","Hand Pose Estimation","Ego4D","Dexterous Manipulation"]},{"id":"unihand","category":"data","sec":4,"tier":3,"sources":[{"title":"arXiv 2507.15597: Being-H0","url":"https://arxiv.org/html/2507.15597v1"},{"title":"BeingBeyond/UniHand_Preview (Hugging Face)","url":"https://huggingface.co/datasets/BeingBeyond/UniHand_Preview"},{"title":"Being-H0.5 项目页","url":"https://research.beingbeyond.com/being-h05"}],"as_of":"2026-01","related_ids":["being-h0","mano","human-video-data","pretraining-on-human-videos","dexterous-manipulation","egodex"],"name":"UniHand","alt":"UniHand 数据集","abbr":"","aliases":["UniHand-1.0","UniHand_Preview"],"one_liner":"BeingBeyond's dataset that unifies many kinds of human hand video into hand-motion annotations, for pretraining dexterous-hand VLA models.","explanation":"UniHand was released in July 2025 by BeingBeyond (智在无界), together with Peking University and Renmin University of China, alongside the Being-H0 model. It unifies human hand data from 11 sources: motion-capture data (such as ARCTIC, HOI4D, and DexYCB), VR-recorded data (such as EgoDex), and RGB-only video (with hand pose estimated using the HaMeR model) — all converted into MANO parameters (a parametric hand model that represents hand shape and pose with a few dozen numbers), totaling over 1,100 hours of video and more than 150 million instruction samples. It addresses the scarcity of real dexterous-hand robot data: a model first learns “how a hand should move” from human hands, then transfers that to a robot. The January 2026 Being-H0.5 expanded it into UniHand-2.0; UniHand_Preview on Hugging Face is a subset of this pretraining data, packaged in WebDataset format.","example":"After pretraining on UniHand, Being-H0 can be given an image and the instruction “unplug the AirPods charging cable” and generate a sequence of hand motions for the next few seconds for both hands.","related":["Being-H0","MANO","Human Video Data","Pretraining on Human Videos","Dexterous Manipulation","EgoDex"]},{"id":"project-go-big","category":"data","sec":4,"tier":3,"sources":[{"title":"Project Go-Big: Internet-Scale Humanoid Pretraining（Figure 官方）","url":"https://www.figure.ai/news/project-go-big"}],"as_of":"2025-09","related_ids":["figure-helix","figure-ai","human-video-data","egocentric-video","zero-shot","vision-language-action-model"],"name":"Project Go-Big","alt":"Figure Project Go-Big","abbr":"","aliases":["Go-Big"],"one_liner":"Figure's plan to pretrain its humanoid robot using massive amounts of first-person human video.","explanation":"Project Go-Big is a data initiative the U.S. humanoid robot company Figure AI announced on September 18, 2025, aimed at building “internet-scale” humanoid pretraining data for its VLA (vision-language-action) model, Helix. Figure partnered with the asset management firm Brookfield, which provides access to more than 100,000 residential units along with large amounts of office and logistics space, used to collect first-person video of people doing everyday things in real homes. Figure says Helix learned to navigate a cluttered home from language instructions — going straight from images and language to chassis velocity commands — using human video alone, and describes this as the first time a humanoid robot has learned this capability end-to-end purely from human video; navigation and manipulation were also merged into the same Helix network. It represents the strategy of substituting human video for part of the teleoperation data a humanoid would otherwise need.","example":"A user says “walk over to the kitchen table,” and Helix outputs walking-speed commands directly from the camera feed; the training data behind this capability came entirely from first-person human video, with no robot demonstrations involved.","related":["Figure Helix","Figure AI","Human Video Data","Egocentric Video","Zero-shot","Vision-Language-Action Model"]},{"id":"world-in-your-hands","category":"data","sec":4,"tier":3,"sources":[{"title":"arXiv 2512.24310: World In Your Hands","url":"https://arxiv.org/abs/2512.24310"},{"title":"全球首个真实世界具身多模态数据集，它石智航交卷（量子位 / 新浪财经，2025-10）","url":"https://finance.sina.com.cn/roll/2025-10-10/doc-inftmhce1065095.shtml"}],"as_of":"2026-09","related_ids":["wearable-data-collection","robot-free-data-collection","human-video-data","tactile-data","tars-robotics","multimodal-data"],"name":"World In Your Hands (WIYH)","alt":"它石 WIYH 数据集","abbr":"WIYH","aliases":[],"one_liner":"TARS Robotics's open-source dataset of over a thousand hours of real-world human manipulation, with vision, language, touch, and action.","explanation":"World In Your Hands (WIYH) was released by TARS Robotics (它石智航) in October 2025, formally open-sourced on December 26, with the paper posted to arXiv in December 2025. Its focus is “human-centered”: rather than using a robot to collect data, people wear the company's own capture kit (called the Oracle Suite in the paper; the commercial version is called SenseHub) while working in real settings like factories, supermarkets, hotels, and restaurants, paired with an automatic annotation pipeline that produces millimeter-accurate ground-truth motion. The total is over 1,000 hours, covering hundreds of skills. The data includes multi-view images, camera calibration parameters, depth maps, 3D hand pose, 6D wrist trajectories, and task annotations; the company describes it as the first large-scale real-world vision-language-tactile-action (VLTA) dataset. The paper reports that adding WIYH's human data to training raised robot manipulation success in cluttered scenes from 8% to 60%.","example":"TARS trained its own AWE model on WIYH data, and demonstrated a robot capable of embroidery at its technology debut event.","related":["Wearable Data Collection","Robot-free (Embodiment-free) Data Collection","Human Video Data","Tactile Data","TARS Robotics","Multimodal Data"]},{"id":"egocentric-10k","category":"data","sec":4,"tier":3,"sources":[{"title":"Egocentric-10K (Hugging Face dataset card)","url":"https://huggingface.co/datasets/builddotai/Egocentric-10K"},{"title":"Build AI 官网","url":"https://www.build.ai/"}],"as_of":"2026-03","related_ids":["egocentric-video","human-video-data","action-free-video","pretraining-on-human-videos","latent-action-pretraining","in-the-wild-data"],"name":"Egocentric-10K","alt":"Egocentric-10K 数据集","abbr":"","aliases":[],"one_liner":"Build AI's open-source dataset of roughly ten thousand hours of first-person video recorded by factory workers.","explanation":"Egocentric-10K is a first-person video dataset released on Hugging Face in 2025 by the U.S. startup Build AI, under the Apache 2.0 license. It was recorded by workers wearing monocular head-mounted cameras while working in 85 real factories, totaling about 10,000 hours across 192,900 video clips and 1.08 billion frames, at 1080p and 30 frames per second, with fisheye camera intrinsics included. It emphasizes footage where the hands stay visible and are actively manipulating things, and is positioned specifically as data for robot learning rather than for general video understanding. The dataset carries no action labels, so it's typically used for human-video pretraining or latent-action pretraining: a model first learns visual representations and action priors from large amounts of human hand manipulation, then gets fine-tuned on a small amount of robot data. According to its website, Build AI has since released larger follow-up versions, Egocentric-100K and Egocentric-1M.","example":"A typical clip in the dataset shows a worker wearing a head-mounted camera at their station, fitting a part into a jig with both hands visible throughout — well suited to learning hand manipulation.","related":["Egocentric Video","Human Video Data","Action-free Video","Pretraining on Human Videos","Latent Action Pretraining","In-the-wild Data"]},{"id":"xperience-10m","category":"data","sec":4,"tier":3,"sources":[{"title":"ropedia-ai/xperience-10m (Hugging Face)","url":"https://huggingface.co/datasets/ropedia-ai/xperience-10m"},{"title":"Xperience-10M Dataset (Ropedia Blog)","url":"https://ropedia.com/blog/20260316_xperience_10m"}],"as_of":"2026-03","related_ids":["egocentric-video","human-video-data","wearable-data-collection","ego4d","world-model","motion-capture"],"name":"Xperience-10M","alt":"Xperience-10M 数据集","abbr":"","aliases":["Ropedia Xperience-10M"],"one_liner":"Ropedia's open dataset of 10,000 hours of first-person, multimodal human experience, aimed at embodied AI.","explanation":"Xperience-10M is an open dataset from the physical-AI data company Ropedia, published on Hugging Face according to the company's blog in March 2026. The data was recorded by people wearing a multi-camera capture device in real daily life; the “10M” in the name refers to 10 million “experiences” (one complete interaction), totaling 10,000 hours. Every experience synchronously records 6 video streams (4 fisheye, 2 stereo), audio, stereo depth, camera pose and SLAM trajectory, hand motion capture, full-body motion capture, and IMU data, along with multi-level language annotations for task, subtask, action, interaction, and object — about 2.88 billion RGB frames in total, roughly 1PB of data, which the company describes as the largest first-person dataset with structured 3D/4D annotation. It can be used for pretraining world models, robot learning, and first-person perception. It's restricted to research and non-commercial use, requiring an application, review, and signed agreement to download; a separate enterprise version exists for commercial use.","example":"A researcher pretrains a hand-motion prediction model on the dataset's hand motion-capture and first-person video, then fine-tunes it on a small amount of robot data.","related":["Egocentric Video","Human Video Data","Wearable Data Collection","Ego4D","World Model","Motion Capture"]},{"id":"egoverse","category":"data","sec":4,"tier":3,"sources":[{"title":"EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World (arXiv)","url":"https://arxiv.org/abs/2604.07607"},{"title":"EgoVerse 官网","url":"https://egoverse.ai/"}],"as_of":"2026-09","related_ids":["egomimic","egocentric-video","human-video-data","cross-embodiment-data","data-scaling-laws-in-imitation-learning","crowdsourced-data-collection"],"name":"EgoVerse","alt":"EgoVerse 数据集","abbr":"","aliases":["EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World"],"one_liner":"A continually growing platform of first-person human demonstration data for robot learning, built jointly by universities and companies.","explanation":"EgoVerse is an open dataset and platform led by Danfei Xu's group at Georgia Tech, together with Stanford, UC San Diego, ETH Zurich, and other universities, plus companies including Meta, Scale AI, and LightWheel AI (光轮智能); its paper was released in April 2026. The version described in the paper contains 1,362 hours and about 80,000 human demonstrations, covering 1,965 tasks, 240 scenes, and 2,087 demonstrators; the project's website describes it as a continually growing dataset, with the current version at roughly 4,000 hours. The data uses a unified format, with camera pose, 3D head tracking, and dense language annotations included. The paper runs human-to-robot transfer experiments with a shared pipeline across multiple labs and robot types, and concludes that more human data generally improves the resulting policy — but only when the human data actually matches the task the robot needs to learn.","example":"In the paper's “place cup on saucer” bimanual task, multiple labs each running their own robot use the same pipeline to compare success rates as different amounts of human data are added.","related":["EgoMimic","Egocentric Video","Human Video Data","Cross-Embodiment Data","Data Scaling Laws in Imitation Learning (Robotic Manipulation)","Crowdsourced Data Collection"]},{"id":"ycb-object-and-model-set","category":"data","sec":5,"tier":2,"sources":[{"title":"Benchmarking in Manipulation Research: The YCB Object and Model Set and Benchmarking Protocols (arXiv 1502.03143)","url":"https://arxiv.org/abs/1502.03143"},{"title":"PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes (arXiv 1711.00199)","url":"https://arxiv.org/abs/1711.00199"}],"as_of":"2015-02","related_ids":["grasping","simulation-assets","6d-object-pose-estimation","benchmark","google-scanned-objects","bop"],"name":"YCB Object and Model Set","alt":"YCB 物体集","abbr":"YCB","aliases":["YCB","YCB Objects","Yale-CMU-Berkeley Object Set"],"one_liner":"A standard set of purchasable everyday objects with 3D scans, letting grasping and manipulation research compare results directly.","explanation":"YCB was proposed in 2015 by researchers at Yale, Carnegie Mellon, and Berkeley (its name comes from the three schools' initials), with the paper published in IEEE Robotics & Automation Magazine. It includes everyday objects numbered 1 through 73, split into five categories — food, kitchen items, tools, shape primitives, and task objects — such as cans, mugs, a power drill, wooden blocks, and rope. Every object was scanned on a multi-camera turntable, yielding 600 RGB-D images, 600 high-resolution color photos, segmentation masks, and a textured 3D mesh. Before YCB, each lab picked its own objects and results couldn't be compared directly; YCB let everyone use the same physical objects and the same models in simulation, and came with template manipulation-benchmark protocols. Today it is commonly imported into simulators as grasping targets, and is also the object source behind several pose-estimation datasets.","example":"PoseCNN's YCB-Video pose-estimation dataset (2017) was built by filming YCB objects on cluttered tabletops, totaling 133,827 frames.","related":["Grasping","Simulation Assets","6D Object Pose Estimation","Benchmark","Google Scanned Objects","BOP (Benchmark for 6D Object Pose Estimation)"]},{"id":"google-scanned-objects","category":"data","sec":5,"tier":3,"sources":[{"title":"Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items (arXiv)","url":"https://arxiv.org/abs/2204.11918"},{"title":"Scanned Objects by Google Research: A Dataset of 3D-Scanned Common Household Items (Google Research Blog)","url":"https://research.google/blog/scanned-objects-by-google-research-a-dataset-of-3d-scanned-common-household-items/"}],"as_of":"2022-06","related_ids":["simulation-assets","ycb-object-and-model-set","objaverse","shapenet","synthetic-data","gazebo"],"name":"Google Scanned Objects","alt":"Google Scanned Objects（GSO）","abbr":"GSO","aliases":["GSO","Scanned Objects by Google Research"],"one_liner":"Google's open-source library of high-precision 3D scans of 1,030 real household objects.","explanation":"Google Scanned Objects (GSO) is a library of 3D object assets that Google Research released in 2022, containing 1,030 real household items — shoes, toys, tableware, packaging boxes, and more. Google built a dedicated scanning rig that uses structured light (projecting a pattern onto the object and inferring its shape from camera images) combined with DSLR photography for high-dynamic-range color, producing mesh models with realistic texture. The models are pre-processed into formats that load directly into Gazebo and PyBullet, hosted on Gazebo Fuel, and released under the CC-BY 4.0 license. It solves the problem of simulators having too few objects, and objects that look too fake: grasping, placement, and synthetic perception data all need large numbers of objects that resemble real things in both shape and appearance, and GSO is one of the commonly used ready-made sources for this, often combined with asset libraries like YCB and Objaverse.","example":"Shoe and toy models from GSO are downloaded from Gazebo Fuel and scattered randomly across a PyBullet tabletop scene to generate grasping training data in bulk.","related":["Simulation Assets","YCB Object and Model Set","Objaverse","ShapeNet","Synthetic Data","Gazebo"]},{"id":"shapenet","category":"data","sec":5,"tier":2,"sources":[{"title":"ShapeNet: An Information-Rich 3D Model Repository","url":"https://arxiv.org/abs/1512.03012"},{"title":"ShapeNet/ShapeNetCore (Hugging Face)","url":"https://huggingface.co/datasets/ShapeNet/ShapeNetCore"},{"title":"ACRONYM: A Large-Scale Grasp Dataset Based on Simulation","url":"https://arxiv.org/abs/2011.09584"}],"as_of":"","related_ids":["objaverse","acronym-a-large-scale-grasp-dataset-based-on-simulation","simulation-assets","ycb-object-and-model-set","partnet-mobility","domain-randomization"],"name":"ShapeNet","alt":"ShapeNet 3D 模型库","abbr":"","aliases":["ShapeNetCore","ShapeNetSem"],"one_liner":"A large-scale library of 3D object meshes organized by WordNet categories, with semantic annotations.","explanation":"ShapeNet was released in 2015 by researchers at Stanford, Princeton, and other institutions, containing more than 3 million 3D models, about 220,000 of which are sorted into 3,135 WordNet categories with annotations for canonical orientation, parts, symmetry planes, and real-world size. The most commonly used subset, ShapeNetCore, has about 51,300 models across 55 common categories; ShapeNetSem has fewer models but finer annotation. It originally served graphics and vision research, and later became a standard object source for robot grasping and simulation — populating a simulator with objects of varied shape at scale, or automatically generating grasp annotations. NVIDIA's ACRONYM grasp dataset uses ShapeNetSem meshes, annotating 17.7 million parallel-jaw grasps inside physics simulation. It is now available through Hugging Face and restricted to non-commercial research and educational use.","example":"ACRONYM takes cup and bowl meshes from ShapeNetSem and tests a large number of gripper grasp poses on each in simulation, using the successes and failures as training labels for a grasp network.","related":["Objaverse","ACRONYM: A Large-Scale Grasp Dataset Based on Simulation","Simulation Assets","YCB Object and Model Set","PartNet-Mobility","Domain Randomization"]},{"id":"objaverse","category":"data","sec":5,"tier":3,"sources":[{"title":"Objaverse: A Universe of Annotated 3D Objects (arXiv 2212.08051)","url":"https://arxiv.org/abs/2212.08051"},{"title":"Objaverse-XL: A Universe of 10M+ 3D Objects (arXiv 2307.05663)","url":"https://arxiv.org/abs/2307.05663"},{"title":"Holodeck: Language Guided Generation of 3D Embodied AI Environments (arXiv 2312.09067)","url":"https://arxiv.org/abs/2312.09067"}],"as_of":"2023-07","related_ids":["simulation-assets","holodeck","shapenet","single-image-3d-reconstruction","partnet-mobility","procedural-generation"],"name":"Objaverse","alt":"Objaverse 3D 资产库","abbr":"","aliases":["Objaverse 1.0","Objaverse-XL"],"one_liner":"A massive open library of 3D object models, led by the Allen Institute for AI.","explanation":"Objaverse is an open 3D object dataset led by the Allen Institute for AI (Ai2). The 2022 version 1.0 contains more than 800,000 models with titles and tags; the 2023 Objaverse-XL, built with several partner institutions, expanded this to over 10 million, sourced from human-made models, photogrammetry, and scans of physical artifacts. 3D data had previously been far smaller in scale than image and text data, which limited 3D vision and generative models. In embodied AI, it's mainly used as a source of simulation assets — for example, Holodeck uses a large language model to plan a room's layout, then retrieves objects from Objaverse to populate the scene. Most models come with only shape and texture, so before they can be used in a simulator, they usually need collision geometry and physical parameters like mass and friction added.","example":"Given a text description like “a study crammed with books,” Holodeck retrieves models such as bookshelves, a desk, and a lamp from Objaverse and arranges them in an AI2-THOR scene for navigation and manipulation training.","related":["Simulation Assets","Holodeck","ShapeNet","Single-Image 3D Reconstruction","PartNet-Mobility","Procedural Generation"]},{"id":"partnet-mobility","category":"data","sec":5,"tier":3,"sources":[{"title":"SAPIEN: A SimulAted Part-based Interactive ENvironment (arXiv 2003.08515)","url":"https://arxiv.org/abs/2003.08515"},{"title":"SAPIEN 论文 HTML 全文（PartNet-Mobility 统计）","url":"https://arxiv.org/html/2003.08515"}],"as_of":"2020-03","related_ids":["sapien","articulated-object","articulated-object-manipulation","unified-robot-description-format","maniskill","articulation-estimation"],"name":"PartNet-Mobility","alt":"PartNet-Mobility 数据集","abbr":"","aliases":["PartNet-Mobility Dataset"],"one_liner":"A library of articulated 3D object models with joint-motion annotations, ready to load directly into a simulator.","explanation":"PartNet-Mobility is an articulated-object dataset released by Hao Su's group at UC San Diego alongside the SAPIEN simulation environment (CVPR 2020). Articulated objects are things like cabinet doors, drawers, and laptops, made of parts connected by joints that can move relative to each other. The dataset contains 46 categories, 2,346 objects, and 14,068 moveable parts; each joint is annotated with its type (revolute, prismatic, or screw), its range of motion, and its parent-child relationship, and comes with textures and a URDF file that loads directly into a physics simulator. Earlier part datasets were mostly static geometry with no way to actually “pull open a drawer.” It's the most commonly used asset source for research on articulated-object manipulation, and benchmarks like ManiSkill build their door-opening and drawer-opening tasks on top of it.","example":"Loading a PartNet-Mobility cabinet's URDF file into SAPIEN or ManiSkill, a robot arm is trained to grasp the handle and pull the drawer open to a target distance.","related":["SAPIEN (SimulAted Part-based Interactive ENvironment)","Articulated Object","Articulated Object Manipulation","Unified Robot Description Format","ManiSkill","Articulation Estimation"]},{"id":"objectfolder","category":"data","sec":5,"tier":3,"sources":[{"title":"ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations (arXiv 2109.07991)","url":"https://arxiv.org/abs/2109.07991"},{"title":"ObjectFolder 2.0: A Multisensory Object Dataset for Sim2Real Transfer (arXiv 2204.02389)","url":"https://arxiv.org/abs/2204.02389"},{"title":"The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects (arXiv 2306.00956)","url":"https://arxiv.org/abs/2306.00956"}],"as_of":"2023-06","related_ids":["multimodal-perception","tactile-data","visuo-tactile-fusion","tactile-simulation","ycb-object-and-model-set","sim-to-real-transfer"],"name":"ObjectFolder","alt":"ObjectFolder 多感官物体数据集","abbr":"","aliases":["ObjectFolder 2.0","ObjectFolder Real","ObjectFolder Benchmark"],"one_liner":"A multisensory object dataset providing an object's appearance, tapping sound, and tactile readings all together.","explanation":"ObjectFolder is a series of multisensory object datasets released by Jiajun Wu, Fei-Fei Li, and others at Stanford. Version 1.0 (CoRL 2021) contains 100 virtual objects, using implicit neural representations (a small network that stores one kind of sensory data and outputs it on query) to encode vision, sound, and touch in a unified way; version 2.0 (CVPR 2022) expanded to 1,000 objects and improved rendering speed and quality; ObjectFolder Real, from 2023, captured meshes, video, tapping sounds, and tactile readings for 100 real household objects, along with an evaluation benchmark of 10 tasks. It fills a gap in most object datasets, which lack sound and touch entirely, and can be used for research on cross-sensory retrieval, contact localization, shape reconstruction, and sim-to-real transfer.","example":"Given a rendered image of a cup, retrieve the sound it makes when tapped, or the tactile image recorded by a vision-based tactile sensor pressing on it — an example of cross-sensory retrieval.","related":["Multimodal Perception","Tactile Data","Visuo-Tactile Fusion","Tactile Simulation","YCB Object and Model Set","Sim-to-Real Transfer"]},{"id":"coco-lvis","category":"data","sec":5,"tier":3,"sources":[{"title":"Microsoft COCO: Common Objects in Context (arXiv 1405.0312)","url":"https://arxiv.org/abs/1405.0312"},{"title":"LVIS: A Dataset for Large Vocabulary Instance Segmentation (arXiv 1908.03195)","url":"https://arxiv.org/abs/1908.03195"},{"title":"COCO 官网数据集介绍","url":"https://cocodataset.org/dataset/home.htm"}],"as_of":"","related_ids":["object-detection","instance-segmentation","open-vocabulary-object-detection","mean-average-precision","yolo","grounding-dino"],"name":"COCO / LVIS","alt":"COCO / LVIS 数据集","abbr":"","aliases":["Common Objects in Context","Large Vocabulary Instance Segmentation","MS COCO","LVIS v1.0"],"one_liner":"The most widely used object-detection and segmentation benchmark; LVIS extends COCO's images to over a thousand long-tail categories.","explanation":"COCO was released by Microsoft and collaborators in 2014: about 330,000 everyday-scene images with 1.5 million object instances, using 80 categories for detection and instance segmentation, plus annotations for image captioning, human keypoints, and more. It's the most common training and evaluation set for detection and segmentation models, and mAP (mean average precision) is usually reported on it. LVIS was introduced by Agrim Gupta, Piotr Dollár, and Ross Girshick at Facebook AI Research at CVPR 2019; it reuses COCO's images but re-annotates them with 1,203 categories and about 2 million instance masks, following a long-tail distribution where a few categories are common and most have very few examples, specifically to test recognition of rare categories. Open-vocabulary detection (finding objects from an arbitrary text description) commonly uses LVIS's rare classes to measure zero-shot ability. Most detection and segmentation models used in robot perception are pretrained or evaluated on one or both of these datasets.","example":"New releases in the YOLO family typically report mAP on the COCO val2017 split; open-vocabulary detectors like YOLO-World and Grounding DINO instead report zero-shot AP on LVIS.","related":["Object Detection","Instance Segmentation","Open-Vocabulary Object Detection","Mean Average Precision","YOLO","Grounding DINO"]},{"id":"scannet","category":"data","sec":5,"tier":3,"sources":[{"title":"ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes (arXiv 1702.04405)","url":"https://arxiv.org/abs/1702.04405"},{"title":"ScanNet++ 官网","url":"https://scannetpp.mlsg.cit.tum.de/scannetpp/"},{"title":"ScanRefer (arXiv 1912.08830)","url":"https://arxiv.org/abs/1912.08830"}],"as_of":"2024-12","related_ids":["3d-visual-grounding","depth-camera","point-cloud","3d-scene-graph","habitat-matterport-3d-dataset","matterport3d"],"name":"ScanNet","alt":"ScanNet 数据集","abbr":"","aliases":["ScanNet++","ScanNet v2"],"one_liner":"A large-scale indoor RGB-D scan dataset with 3D reconstruction and semantic annotations.","explanation":"ScanNet is an indoor-scene dataset released by Angela Dai, Matthias Nießner, and colleagues at CVPR 2017: 1,513 indoor scenes were scanned with consumer RGB-D cameras (capturing color and depth simultaneously), totaling 2.5 million frames, along with camera poses, 3D surface reconstructions, and crowdsourced semantic segmentation labels, providing a unified training and evaluation set for 3D semantic segmentation and 3D detection. In embodied AI, many 3D-language tasks are built on top of it, such as 3D visual grounding (locating an object in a 3D scene from a sentence) and 3D question answering. A later version, ScanNet++, switched to a laser scanner, a DSLR, and an iPhone for capture, achieving higher precision, and v2, released in December 2024, expanded coverage to more than 1,000 scenes.","example":"ScanRefer writes 51,583 natural-language descriptions for 11,046 objects across 800 ScanNet scenes; a model must use the description to locate the target object within the 3D scan.","related":["3D Visual Grounding","Depth Camera","Point Cloud","3D Scene Graph","Habitat-Matterport 3D Dataset","Matterport3D"]},{"id":"matterport3d","category":"data","sec":5,"tier":3,"sources":[{"title":"Matterport3D 项目主页","url":"https://niessner.github.io/Matterport/"},{"title":"Room-to-Room (R2R) 数据集主页","url":"https://bringmeaspoon.org/"}],"as_of":"2017","related_ids":["habitat-matterport-3d-dataset","habitat","room-to-room","vision-and-language-navigation","object-goal-navigation","scannet"],"name":"Matterport3D","alt":"Matterport3D 数据集","abbr":"MP3D","aliases":["MP3D"],"one_liner":"A dataset of RGB-D indoor scans from 90 real buildings, a standard scene library for indoor navigation research.","explanation":"Matterport3D (abbreviated MP3D) was published at the 3DV conference in 2017 by researchers from Princeton University, Stanford University, and other institutions, built from data captured with Matterport's 3D scanning cameras. It covers 90 building-scale real indoor scenes, with 194,400 RGB-D images (color plus depth) stitched into 10,800 panoramas, along with surface reconstruction meshes, camera poses, and 2D/3D semantic segmentation labels. Its significance is letting researchers train and evaluate agents inside digital copies of real houses, without physically moving a robot into a real home every time: the vision-and-language navigation benchmark R2R is built on top of it, and simulation platforms such as Habitat use it as a standard set of scenes for tasks like point-goal and object-goal navigation. The later HM3D dataset expanded the scene count to 1,000. Access requires signing a terms-of-use agreement.","example":"The R2R benchmark annotates about 22,000 navigation instructions (averaging 29 words) across Matterport3D's 90 buildings; an agent must understand a description like “go through the kitchen and stop at the doorway by the stairs” and walk to the destination.","related":["Habitat-Matterport 3D Dataset","Habitat","Room-to-Room","Vision-and-Language Navigation","Object-Goal Navigation","ScanNet"]},{"id":"habitat-matterport-3d-dataset","category":"data","sec":5,"tier":3,"sources":[{"title":"Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI (arXiv)","url":"https://arxiv.org/abs/2109.08238"},{"title":"HM3D - AI Habitat","url":"https://aihabitat.org/datasets/hm3d/"},{"title":"HM3D-Semantics - AI Habitat","url":"https://aihabitat.org/datasets/hm3d-semantics/"}],"as_of":"","related_ids":["habitat","matterport3d","object-goal-navigation","point-goal-navigation","navigation","scannet"],"name":"Habitat-Matterport 3D Dataset","alt":"HM3D 数据集","abbr":"HM3D","aliases":["HM3D","HM3D-Semantics (HM3DSem)"],"one_liner":"A dataset of 1,000 real buildings, 3D-reconstructed by Meta and Matterport, used to train navigation agents.","explanation":"HM3D is an indoor scene dataset released jointly in 2021 by Meta AI (FAIR) and the property-scanning company Matterport. It contains 1,000 building-scale, textured 3D meshes covering homes, stores, and public buildings, with 112,500 square meters of navigable area combined — 1.4 to 3.7 times more than earlier datasets like Matterport3D and Gibson — and with fewer reconstruction defects. It's mainly used together with the Habitat simulator: an agent practices tasks like point-goal navigation and object-goal navigation inside digital copies of these real buildings. The paper shows that a point-goal navigation agent trained on HM3D also performs best when tested on other datasets. A later release, HM3D-Semantics v0.2, added 142,646 object-instance annotations across 216 of the spaces; v0.1 was the basis for the Habitat 2022 Object-Goal Navigation Challenge. The data is licensed for academic, non-commercial use only.","example":"Loading one HM3D house into Habitat, a robot agent starts at the front door and must find the location of the “sofa” using only its camera feed.","related":["Habitat","Matterport3D","Object-Goal Navigation","Point-Goal Navigation","Navigation","ScanNet"]},{"id":"graspnet-1billion","category":"data","sec":5,"tier":3,"sources":[{"title":"GraspNet-1Billion 官网","url":"https://graspnet.net/"}],"as_of":"2023","related_ids":["grasp-pose-detection","anygrasp","antipodal-grasp","6d-object-pose-estimation","bin-picking","sjtu-mvig-lab"],"name":"GraspNet-1Billion","alt":"GraspNet-1Billion 数据集","abbr":"","aliases":["GraspNet","GraspNet-1B"],"one_liner":"A Shanghai Jiao Tong University dataset of real cluttered-scene grasping data, with over 1.1 billion annotated grasp poses.","explanation":"GraspNet-1Billion is a general object-grasping benchmark from the MVIG lab at Shanghai Jiao Tong University (Cewu Lu's group), published at CVPR 2020, with an extended version published in the robotics journal IJRR in 2023. It contains 190 real cluttered scenes with 88 object types, captured using two RGB-D cameras, a RealSense D435 and an Azure Kinect, for a total of 97,280 images. Every image is annotated with each object's 6D pose (3D position plus 3D orientation), instance masks, and dense 6-DoF grasp poses, totaling more than 1.1 billion grasps. Earlier grasping datasets were mostly single-object or synthetically generated in simulation; this dataset provides large-scale annotation on real, cluttered tabletops, and ships with an open-source evaluation API and baseline network so different grasp-detection algorithms can be compared under the same standard. The same group's later work, AnyGrasp, builds on this line of research.","example":"To train a grasp-detection network: feed in one frame of depth point cloud, output several gripper poses with scores, then compute average precision on test scenes using the graspnetAPI.","related":["Grasp Pose Detection","AnyGrasp","Antipodal Grasp","6D Object Pose Estimation","Bin Picking","SJTU MVIG Lab"]},{"id":"acronym-a-large-scale-grasp-dataset-based-on-simulation","category":"data","sec":5,"tier":3,"sources":[{"title":"ACRONYM: A Large-Scale Grasp Dataset Based on Simulation (arXiv 2011.09584)","url":"https://arxiv.org/abs/2011.09584"},{"title":"NVlabs/acronym (GitHub)","url":"https://github.com/NVlabs/acronym"},{"title":"Contact-GraspNet (arXiv 2103.14127)","url":"https://arxiv.org/abs/2103.14127"}],"as_of":"2020-11","related_ids":["grasping","contact-graspnet","simulation-data","antipodal-grasp","grasp-pose-detection","shapenet"],"name":"ACRONYM: A Large-Scale Grasp Dataset Based on Simulation","alt":"ACRONYM 抓取数据集","abbr":"","aliases":[],"one_liner":"An NVIDIA grasp dataset labeled in bulk through physics simulation, containing about 17.7 million parallel-jaw grasps.","explanation":"ACRONYM was released in November 2020 by NVIDIA's Clemens Eppner, Arsalan Mousavian, and Dieter Fox. It selected 8,872 objects across 262 categories from the ShapeNetSem 3D model library, generated 2,000 candidate grasps per object using antipodal sampling (finding two opposing contact points on the surface), and then simulated each one in NVIDIA's FleX physics engine — closing a Franka Panda gripper and shaking the object to see whether it stayed held — yielding about 17.7 million labeled grasps, roughly 59% of them successful. Testing grasps one by one on a real robot is far too costly, so this kind of bulk simulated labeling gave learned grasping networks enough training data. The release also includes a tool for generating cluttered scenes with multiple objects randomly placed on support surfaces, capable of rendering depth images and point clouds.","example":"NVIDIA's Contact-GraspNet, trained on about 17 million of these simulated grasps, exceeded a 90% success rate grasping previously unseen objects in real cluttered scenes.","related":["Grasping","Contact-GraspNet","Simulation Data","Antipodal Grasp","Grasp Pose Detection","ShapeNet"]},{"id":"dexgraspnet","category":"data","sec":5,"tier":3,"sources":[{"title":"DexGraspNet 项目主页","url":"https://pku-epic.github.io/DexGraspNet/"},{"title":"DexGraspNet (arXiv)","url":"https://arxiv.org/abs/2210.02697"},{"title":"DexGraspNet 2.0 (arXiv)","url":"https://arxiv.org/abs/2410.23004"}],"as_of":"2024-10","related_ids":["dexterous-manipulation","force-closure","shadow-dexterous-hand","isaac-gym","synthetic-data","grasp-pose-detection"],"name":"DexGraspNet","alt":"DexGraspNet 数据集","abbr":"","aliases":["DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation"],"one_liner":"A simulation-generated dexterous-hand grasping dataset from Peking University, with 1.32 million grasps across 5,355 objects.","explanation":"DexGraspNet is a large-scale dexterous-hand grasping dataset released by He Wang's group at Peking University, together with the Beijing Institute for General Artificial Intelligence (BIGAI) and Tsinghua University, published at ICRA 2023. Large datasets already existed for two-finger parallel-jaw grasping, but dexterous-hand grasping had long lacked data. The authors used an accelerated, differentiable force-closure estimator (force closure means the fingers' contact forces can resist an external force from any direction) to synthesize grasp poses at scale, generating 1.32 million grasps for a Shadow dexterous hand across 5,355 objects in more than 133 categories — over 200 grasps per object — all verified in the Isaac Gym simulator. It's mainly used to train and evaluate dexterous grasp-generation algorithms. In 2024 the team released DexGraspNet 2.0, expanding to 8,270 cluttered scenes and 427 million grasps, and demonstrating zero-shot transfer to a real robot.","example":"A generative model that takes an object's point cloud as input and outputs a Shadow-hand grasp pose can be trained using DexGraspNet's grasps as supervision.","related":["Dexterous Manipulation","Force Closure","Shadow Dexterous Hand","Isaac Gym","Synthetic Data","Grasp Pose Detection"]},{"id":"scripted-demonstrations","category":"data","sec":5,"tier":3,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/html/2304.13705"},{"title":"tonyzhaozh/act (GitHub)","url":"https://github.com/tonyzhaozh/act"}],"as_of":"","related_ids":["demonstration-data","simulation-data","privileged-information","action-multimodality","mimicgen","aloha-sim"],"name":"Scripted Demonstrations","alt":"脚本化演示","abbr":"","aliases":["Scripted Policy Demonstrations","Programmatic Demonstrations"],"one_liner":"Demonstration data generated automatically by a hand-written rule-based program controlling the robot, instead of a human.","explanation":"Scripted demonstrations are produced by a hand-written rule-based program (a scripted policy) that controls the robot through a task automatically, with the process recorded as demonstration data. The script typically reads privileged information directly in simulation, such as an object's exact pose (precise state that isn't available in real deployment), and computes a trajectory using preset waypoints or a motion planner. The advantage is low cost, unlimited generation, and clean, consistent motion; the disadvantage is a narrow range of motion patterns, very different from the diverse, hesitant style of real human demonstrations — complex tasks are hard to script at all, and precise state is hard to obtain on a real robot. It's commonly used for simulation benchmarks, algorithm debugging, and large-scale synthetic data. For the same number of demonstrations, a policy trained on scripted data often reaches a higher success rate than one trained on human teleoperation data, simply because scripted data lacks the multimodality and noise of human motion — so methods validated only on scripted data should be viewed cautiously.","example":"The ACT paper recorded 50 scripted demonstrations and 50 human-teleoperated demonstrations for the same block-transfer task in ALOHA simulation; ACT reached 97% success on the scripted data, dropping to 82% on the human data.","related":["Demonstration Data","Simulation Data","Privileged Information","Action Multimodality","MimicGen","ALOHA Sim (Transfer Cube / Insertion)"]},{"id":"mimicgen","category":"data","sec":5,"tier":2,"sources":[{"title":"MimicGen 项目主页","url":"https://mimicgen.github.io/"},{"title":"Isaac Lab Docs: Teleoperation and Imitation Learning with Isaac Lab Mimic","url":"https://isaac-sim.github.io/IsaacLab/main/source/overview/imitation-learning/teleop_imitation.html"},{"title":"NVIDIA Technical Blog: Building a Synthetic Motion Generation Pipeline for Humanoid Robot Learning","url":"https://developer.nvidia.com/blog/building-a-synthetic-motion-generation-pipeline-for-humanoid-robot-learning/"}],"as_of":"","related_ids":["dexmimicgen","demogen","synthetic-data","demonstration-data","nvidia-isaac-lab","behavior-cloning"],"name":"MimicGen","alt":"MimicGen","abbr":"","aliases":[],"one_liner":"A system that segments a few human demonstrations by object, transforms and stitches them, and auto-generates many new demonstrations.","explanation":"MimicGen is an automatic data-generation system proposed by NVIDIA and UT Austin at CoRL 2023. It splits a human demonstration into object-centric subtask segments (such as “grasp the cup” and “place it on the plate”); when applied to a new scene where the object is in a different position, it transforms each segment's end-effector trajectory to match that object's new pose, stitches the segments back together, executes the result in simulation, and keeps only the successful runs. Using fewer than 200 human demonstrations, the paper generated more than 50,000 demonstrations across 18 tasks, easing imitation learning's dependence on human-collected demonstrations. NVIDIA's Isaac Lab Mimic, built into Isaac Lab, follows the same idea — annotate subtasks, then generate automatically — and also supports humanoid robots.","example":"For one simulated manipulation task, giving MimicGen just 10 human demonstrations lets it generate 1,000 demonstrations across a wider range of object placements, which are then used to train a policy.","related":["DexMimicGen","DemoGen","Synthetic Data","Demonstration Data","NVIDIA Isaac Lab","Behavior Cloning"]},{"id":"dexmimicgen","category":"data","sec":5,"tier":3,"sources":[{"title":"DexMimicGen 项目主页","url":"https://dexmimicgen.github.io"},{"title":"DexMimicGen (arXiv)","url":"https://arxiv.org/abs/2410.24185"}],"as_of":"2024-10","related_ids":["mimicgen","demogen","bimanual-manipulation","real-to-sim-to-real","synthetic-data","behavior-cloning"],"name":"DexMimicGen","alt":"DexMimicGen","abbr":"","aliases":["DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning"],"one_liner":"NVIDIA's automatic data-generation system that expands a few dozen demonstrations into tens of thousands of bimanual dexterous-hand demonstrations.","explanation":"DexMimicGen is an automated data-generation system from NVIDIA Research, together with the University of Texas at Austin and UC San Diego, published at ICRA 2025; it extends MimicGen to bimanual dexterous manipulation, such as the upper body of a humanoid robot. MimicGen's approach is to split a small number of human demonstrations into per-object subtask segments, transform those segments to match new object positions, replay them in simulation, and keep only the trajectories that succeed. In two-arm tasks, the two hands sometimes act independently, sometimes need to move in sync, and sometimes must act in a specific order, so DexMimicGen adds per-arm subtask segmentation along with synchronization and ordering constraints. The paper generates about 21,000 demonstrations from 60 human demonstrations across 9 tasks, and validates the approach on a real-robot can-sorting task using a real-to-sim-to-real pipeline.","example":"Starting from 60 human teleoperated demonstrations, the system automatically generates about 21,000 bimanual dexterous-hand demonstrations in simulators such as robosuite, which are then used to train a policy with behavioral cloning.","related":["MimicGen","DemoGen","Bimanual Manipulation","Real-to-Sim-to-Real","Synthetic Data","Behavior Cloning"]},{"id":"demogen","category":"data","sec":5,"tier":3,"sources":[{"title":"DemoGen 项目主页","url":"https://demo-generation.github.io"},{"title":"DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning (arXiv)","url":"https://arxiv.org/abs/2502.16932"}],"as_of":"2025-06","related_ids":["mimicgen","dexmimicgen","3d-diffusion-policy","synthetic-data","spatial-generalization","point-cloud"],"name":"DemoGen","alt":"DemoGen","abbr":"","aliases":["DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning"],"one_liner":"A method that automatically turns one real-robot demonstration into many synthetic demonstrations with objects in new positions.","explanation":"DemoGen is a synthetic-demonstration-generation method from Huazhe Xu's group at Tsinghua University, in collaboration with the Shanghai Qi Zhi Institute and Shanghai AI Lab, published at RSS 2025. Imitation-learning policies generalize poorly across space — moving an object to a new position can make them fail — but collecting demonstrations by hand at every possible position is expensive. DemoGen needs only one human demonstration per task: it splits the trajectory into free-space transport segments and contact-based manipulation segments; when the object moves to a new position, the contact segment is transformed along with it, and the transport segment is reconnected using motion planning. Observations are produced by directly editing the point cloud, moving the points belonging to the object and the end effector together — no simulator is needed, and there's no need to re-collect data on the real robot. The paper reports generating a full dataset in about 22 seconds, versus roughly 83.7 hours for the MimicGen approach. Training pairs this with 3D Diffusion Policy (DP3), which takes point clouds as input.","example":"On a jar-opening task, with only 1 collected demonstration, DemoGen generates demonstrations with the jar in different positions; the paper reports the resulting policy reaches 100% success even at out-of-distribution positions.","related":["MimicGen","DexMimicGen","3D Diffusion Policy","Synthetic Data","Spatial Generalization","Point Cloud"]},{"id":"robogen","category":"data","sec":5,"tier":3,"sources":[{"title":"RoboGen (arXiv:2311.01455)","url":"https://arxiv.org/abs/2311.01455"},{"title":"RoboGen 项目主页","url":"https://robogen-ai.github.io/"}],"as_of":"2024-06","related_ids":["generative-simulation","genesis","procedural-generation","llm-based-task-planning","reward-function","simulation-data"],"name":"RoboGen","alt":"RoboGen","abbr":"","aliases":["RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation"],"one_liner":"A generative simulation framework where a large model proposes tasks, builds the simulated scene, and learns the skill.","explanation":"RoboGen is a robot-learning framework proposed in November 2023 by researchers at CMU, MIT, the Tsinghua-affiliated IIIS, UMass Amherst, and other institutions, published at ICML 2024. Rather than having a large model output actions directly, it runs a “propose–generate–learn” loop: a large model first proposes a skill to learn, then automatically builds the matching simulated scene and objects, breaks the task into subtasks, and writes supervision signals such as reward functions; finally, depending on the task type, it picks reinforcement learning, motion planning, or trajectory optimization to learn the skill. This needs almost no manual task design, yet can continuously produce demonstration data for a wide range of skills, spanning articulated objects, deformable-object manipulation, and legged locomotion. The project page states it uses the Genesis simulation engine for physics and rendering. It's a representative example of “generative simulation.”","example":"A large model proposes the task “pull out a suitcase's handle”; RoboGen automatically places a suitcase model into the scene, sets up the initial state, generates a reward function, and trains the skill with reinforcement learning.","related":["Generative Simulation","Genesis","Procedural Generation","LLM-based Task Planning","Reward Function","Simulation Data"]},{"id":"nvidia-physical-ai-dataset","category":"data","sec":5,"tier":3,"sources":[{"title":"NVIDIA 博客：NVIDIA Releases Open Physical AI Dataset","url":"https://blogs.nvidia.com/blog/open-physical-ai-dataset/"},{"title":"nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim（Hugging Face）","url":"https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim"}],"as_of":"2026-09","related_ids":["nvidia","nvidia-isaac-gr00t-n1","synthetic-data","simready-assets","autonomous-driving","physical-ai"],"name":"NVIDIA Physical AI Dataset","alt":"NVIDIA 物理 AI 数据集","abbr":"","aliases":["PhysicalAI Dataset"],"one_liner":"NVIDIA's open collection of robotics and autonomous-driving datasets hosted on Hugging Face.","explanation":"The NVIDIA Physical AI Dataset is an open collection of datasets that NVIDIA announced at its GTC conference on March 18, 2025, hosted on Hugging Face, aimed at developing physical AI for robotics, autonomous driving, and related fields, combining both real and synthetic data. The initial release totaled about 15TB, including more than 320,000 robot training trajectories and up to 1,000 OpenUSD assets. It isn't a single file but a series of datasets under NVIDIA's Hugging Face account, all prefixed PhysicalAI-, such as the simulated trajectories used for GR00T N1 post-training, GR00T teleoperation data, SimReady asset repositories, about 1,700 hours of multi-sensor autonomous-driving data, and, added in 2026, synthetic world-model scene data. As of September 2026, more than 40 datasets exist under this prefix, each with its own license, so licenses need to be checked individually before use.","example":"The GR00T-X-Embodiment-Sim subset provides the simulated trajectories used for GR00T N1 post-training; the GR-1 humanoid's tabletop manipulation portion alone has 240,000 trajectories, downloadable directly for fine-tuning GR00T.","related":["NVIDIA","NVIDIA Isaac GR00T N1","Synthetic Data","SimReady Assets","Autonomous Driving","Physical AI"]},{"id":"syngrasp-1b","category":"data","sec":5,"tier":3,"sources":[{"title":"GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data (arXiv 2505.03233)","url":"https://arxiv.org/abs/2505.03233"},{"title":"GraspVLA 项目页","url":"https://pku-epic.github.io/GraspVLA-web/"},{"title":"PKU-EPIC/GraspVLA (GitHub)","url":"https://github.com/PKU-EPIC/GraspVLA"}],"as_of":"2026-08","related_ids":["graspvla","synthetic-data","domain-randomization","objaverse","curobo","sim-to-real-transfer"],"name":"SynGrasp-1B","alt":"SynGrasp-1B 数据集","abbr":"","aliases":["SynGrasp"],"one_liner":"A billion-frame simulated synthetic grasping dataset built by Galbot and others, used to pretrain GraspVLA.","explanation":"SynGrasp-1B is a synthetic grasping dataset built by Galbot (银河通用) together with Peking University, the University of Hong Kong, and the Beijing Academy of Artificial Intelligence, totaling about 1 billion frames, all generated in simulation: 240 categories and more than 10,000 object models were selected from Objaverse and randomly placed on tables, BoDex generated stable grasp poses, cuRobo planned the grasping trajectories, and the results were rendered after domain randomization of materials, lighting, camera viewpoint, background, and initial pose. It's used to pretrain the grasping VLA model GraspVLA (CoRL 2025), testing whether synthetic action data alone can transfer zero-shot to real-world grasping. The dataset was made public on Hugging Face in August 2026.","example":"GraspVLA, pretrained on SynGrasp-1B with no real-robot fine-tuning at all, can grasp objects it never saw during training on a real tabletop, following a language instruction.","related":["GraspVLA","Synthetic Data","Domain Randomization","Objaverse","cuRobo (NVIDIA GPU-accelerated motion planning)","Sim-to-Real Transfer"]},{"id":"interndata-a1","category":"data","sec":5,"tier":3,"sources":[{"title":"InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy (arXiv 2511.16651)","url":"https://arxiv.org/abs/2511.16651"},{"title":"InternData-A1 数据集页面（Hugging Face）","url":"https://huggingface.co/datasets/InternRobotics/InternData-A1"}],"as_of":"2026-01","related_ids":["synthetic-data","simulation-data","pi0","internvla","curobo","domain-randomization"],"name":"InternData-A1","alt":"InternData-A1 数据集","abbr":"","aliases":["InternData A1"],"one_liner":"A large-scale simulated synthetic dataset from Shanghai AI Lab, built for pretraining general-purpose robot policies.","explanation":"InternData-A1 is a synthetic robot-manipulation dataset released in November 2025 by the Shanghai Artificial Intelligence Laboratory (the InternRobotics team, with participation from Peking University), generated entirely in simulation: over 630,000 trajectories totaling 7,433 hours, covering 4 robot embodiments (ARX Lift-2, AgileX Split ALOHA, A2D, and Franka), 70 tasks, and 227 scenes, including manipulation of rigid, articulated, and deformable objects as well as fluids. The pipeline decouples skills, tasks, and embodiments so they can be freely recombined, uses cuRobo for motion planning, and applies domain randomization (randomly varying things like camera viewpoint, lighting, and object placement so the model doesn't depend on any one fixed appearance). The paper's conclusion: pretraining π0 on this synthetic data alone matches the performance of the official π0, which was pretrained on π's real-robot data, across 49 simulated tasks, 5 real-robot tasks, and 4 long-horizon dexterous tasks. The data is released on Hugging Face under CC BY-NC-SA 4.0.","example":"The paper reports that for some tasks, fewer than 1,600 simulated samples match the effect of 200 real-robot samples; for 10 of the 70 tasks, the policy achieves strong real-robot success without using any real-robot data at all.","related":["Synthetic Data","Simulation Data","π0","InternVLA (Shanghai AI Laboratory)","cuRobo (NVIDIA GPU-accelerated motion planning)","Domain Randomization"]},{"id":"generative-data-augmentation","category":"data","sec":5,"tier":2,"sources":[{"title":"ROSIE: Scaling Robot Learning with Semantically Imagined Experience (arXiv 2302.11550)","url":"https://arxiv.org/abs/2302.11550"},{"title":"GenAug: Retargeting behaviors to unseen situations via Generative Augmentation (arXiv 2302.06671)","url":"https://arxiv.org/abs/2302.06671"}],"as_of":"","related_ids":["data-augmentation","synthetic-data","diffusion-model","nvidia-cosmos-transfer","visual-generalization","cross-painting"],"name":"Generative Data Augmentation","alt":"生成式数据增强","abbr":"","aliases":["Semantic Data Augmentation"],"one_liner":"Using image or video generation models to rewrite existing robot data, producing new-scene training examples from old demonstrations.","explanation":"Generative data augmentation means using generative models — text-to-image or video generation — to swap in new objects, backgrounds, distractors, or lighting on top of existing robot demonstrations, producing new training examples while reusing the original action labels. Notable examples include Google's 2023 ROSIE, which uses a text-guided diffusion model to inpaint new objects and backgrounds into images, and the University of Washington and collaborators' GenAug. It targets the problem that real-robot data is expensive and each demonstration only covers one tabletop: after augmentation, the same motion can be paired with many different appearances, improving visual generalization and robustness to distractors. Its limitation is that it mainly changes appearance, not the physical process or the action itself. Video world models such as Cosmos Transfer are also commonly used for this kind of augmentation.","example":"ROSIE uses a diffusion model to inpaint new objects and backgrounds onto Google's existing robot data; a policy trained on these images can handle new objects and distractors that never appeared in the original data.","related":["Data Augmentation","Synthetic Data","Diffusion Model","NVIDIA Cosmos Transfer","Visual Generalization","Cross-Painting"]},{"id":"neural-trajectories","category":"data","sec":5,"tier":3,"sources":[{"title":"DreamGen 项目主页（NVIDIA GEAR）","url":"https://research.nvidia.com/labs/gear/dreamgen/"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"},{"title":"DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv 2505.12705)","url":"https://arxiv.org/abs/2505.12705"}],"as_of":"2025-06","related_ids":["dreamgen","nvidia-isaac-gr00t-n1","synthetic-data","pseudo-action-labels","inverse-dynamics-model","latent-action-model"],"name":"Neural Trajectories","alt":"神经轨迹","abbr":"","aliases":[],"one_liner":"Synthetic robot training data generated by a video world model, with pseudo action labels added afterward.","explanation":"Neural trajectories is a term used by NVIDIA's GEAR lab in GR00T N1 (March 2025) and DreamGen (May 2025) for synthetic robot data generated by a video world model. The pipeline has four steps: first, fine-tune a video generation model on the target robot's real data, so it learns that robot's appearance and how it moves; then, given a starting frame and a language instruction, generate video of the robot performing a task — which can involve entirely new motions or new environments never actually collected; next, use a latent-action model or an inverse-dynamics model (a model that infers the action from a pair of before-and-after frames) to add pseudo action labels to the generated video; finally, train a visuomotor policy on these pseudo-labeled videos together with real data. Compared with synthetic data from a physics simulator, this approach needs no simulated scenes or assets to be built, but the physical plausibility of the generated footage isn't guaranteed and needs dedicated evaluation — which is what DreamGen Bench was built for.","example":"GR00T N1 generated about 827 hours of neural trajectories (roughly 10x) from about 88 hours of real-robot data, using 3,600 L40 GPUs over about 1.5 days; co-training with the real data raised the GR-1 humanoid's average success rate across 8 real-robot tasks by 5.8 percentage points.","related":["DreamGen","NVIDIA Isaac GR00T N1","Synthetic Data","Pseudo Action Labels","Inverse Dynamics Model","Latent Action Model"]},{"id":"real2render2real","category":"data","sec":5,"tier":3,"sources":[{"title":"Real2Render2Real (arXiv:2505.09601)","url":"https://arxiv.org/abs/2505.09601"},{"title":"Real2Render2Real 项目主页","url":"https://real2render2real.com/"}],"as_of":"2025-05","related_ids":["synthetic-data","3d-gaussian-splatting","real-to-sim-to-real","human-video-data","teleoperation","diffusion-policy"],"name":"Real2Render2Real","alt":"Real2Render2Real（R2R2R）","abbr":"R2R2R","aliases":["R2R2R"],"one_liner":"A method that renders large amounts of robot training data from a phone scan plus one video of a human demonstration.","explanation":"Real2Render2Real is a data-generation method released in May 2025 by UC Berkeley (Ken Goldberg's group) and the Toyota Research Institute. It needs only two inputs: a 3D scan of an object taken with a phone, and one video of a person performing the task by hand. The system reconstructs the object's shape with 3D Gaussian splatting, tracks the object's 6-DoF motion throughout the video, converts it into a mesh, pairs it with a robot model, and re-renders thousands of demonstrations with varied object positions and camera viewpoints. It only renders images and does no dynamics simulation at all (collisions are turned off), so there's no need to tune physical parameters or use a real robot. The paper reports that a model trained on data generated from just 1 human demonstration matches the performance of one trained on 150 teleoperated demonstrations, and that generation takes about 1/27th the time of teleoperation.","example":"Scanning a cup with a phone and recording one video of a person placing the cup onto a plate, R2R2R can render thousands of image-action demonstrations of a robot arm performing the same motion, used to train π0-FAST or a diffusion policy.","related":["Synthetic Data","3D Gaussian Splatting","Real-to-Sim-to-Real","Human Video Data","Teleoperation","Diffusion Policy"]},{"id":"cross-embodiment-data","category":"data","sec":6,"tier":2,"sources":[{"title":"Open X-Embodiment Collaboration 2023: Open X-Embodiment: Robotic Learning Datasets and RT-X Models","url":"https://arxiv.org/abs/2310.08864"},{"title":"Open X-Embodiment 项目页","url":"https://robotics-transformer-x.github.io/"}],"as_of":"","related_ids":["cross-embodiment","open-x-embodiment","embodiment-gap","unified-action-space","embodiment-specific-head","positive-negative-transfer"],"name":"Cross-Embodiment Data","alt":"跨本体数据","abbr":"","aliases":["Multi-Embodiment Data","X-Embodiment Data"],"one_liner":"Data collected from many different kinds of robots and pooled together for training.","explanation":"Cross-embodiment data refers to a collection of data drawn from multiple robot embodiments (a robot's specific physical hardware form) whose shape, degrees of freedom, camera placement, and action space all differ — single arms, dual arms, mobile bases, and humanoids among them. A single lab's data is limited in volume, so pooling data from many labs is meant to help a model learn skills that don't depend on any one particular robot, benefiting every embodiment involved. The leading example is Open X-Embodiment, led by Google DeepMind in 2023: 60 datasets from 34 labs, 22 kinds of robots, and more than 1 million trajectories, unified into RLDS format. The difficulty is that different labs' actions and observations don't line up, requiring normalization, a unified action space, or an embodiment-specific output head for each robot; a poorly chosen mixture can even cause negative transfer.","example":"RT-1-X, trained on Open X-Embodiment, averaged about a 50% higher success rate than the original methods each lab trained using only its own data, on several tasks from labs with small datasets.","related":["Cross-Embodiment","Open X-Embodiment","Embodiment Gap","Unified Action Space","Embodiment-specific Head","Positive / Negative Transfer"]},{"id":"open-x-embodiment","category":"data","sec":6,"tier":1,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models（项目主页）","url":"https://robotics-transformer-x.github.io/"},{"title":"google-deepmind/open_x_embodiment（GitHub）","url":"https://github.com/google-deepmind/open_x_embodiment"}],"as_of":"2023-10","related_ids":["cross-embodiment-data","rt-x","rlds","octo","openvla","oxe-magic-soup"],"name":"Open X-Embodiment","alt":"Open X-Embodiment 数据集","abbr":"OXE","aliases":["OXE","Open X","RT-X Dataset"],"one_liner":"A Google-led open dataset pooling more than a million real-robot trajectories from 22 kinds of robots.","explanation":"Open X-Embodiment (OXE) is a cross-embodiment robot dataset released in October 2023 by Google DeepMind together with 21 other institutions. It consolidates 60 existing datasets from 34 labs into a unified RLDS format — a format Google proposed for storing robot data by episode — totaling more than 1 million real-robot trajectories, 22 robot embodiments, and 527 skills. Before this, each lab used its own data format and action definitions, making it hard to combine data for training. The team trained RT-1-X and RT-2-X on the pooled data: on robots that had less data of their own, RT-1-X's success rate was on average about 50% higher than a model trained only on that robot's own data, showing that mixing data from many kinds of robots helps them all. OXE went on to become the main pretraining data for open generalist policies such as Octo and OpenVLA.","example":"After training on OXE's mixed data, RT-2-X's success rate on an emergent-skills evaluation — skills absent from Google's own robot data, learned only from other robots' data — was roughly 3 times that of RT-2, which was trained only on Google's own data.","related":["Cross-Embodiment Data","RT-X","RLDS (Reinforcement Learning Datasets)","Octo","OpenVLA","OXE Magic Soup"]},{"id":"heterogeneous-data","category":"data","sec":6,"tier":3,"sources":[{"title":"Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (arXiv)","url":"https://arxiv.org/abs/2409.20537"},{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv)","url":"https://arxiv.org/abs/2310.08864"}],"as_of":"","related_ids":["cross-embodiment-data","cross-embodiment","heterogeneous-pre-trained-transformers","unified-action-space","embodiment-specific-head","data-mixture"],"name":"Heterogeneous Data","alt":"异构数据","abbr":"","aliases":["Multi-Source Heterogeneous Data"],"one_liner":"Robot data from different robots, sensors, and collection methods, with mismatched formats and meanings.","explanation":"“Heterogeneity” in robot data shows up on several levels at once: different embodiments (single-arm, bimanual, humanoid, each with its own joint count and action space); different sensors (number and placement of cameras, presence or absence of depth or touch); different control schemes (joint angles versus end-effector pose, different control frequencies); and different sources (real-robot teleoperation, simulation, human video). Any single lab's own data is too small on its own to train a general-purpose policy, so labs need to mix this kind of data together — but simply concatenating it confuses a model about what the same number means on different robots. A common fix gives each embodiment its own input encoder and output head, while sharing one large backbone in the middle — for instance, Kaiming He's group used this structure in HPT to pool 52 datasets together. Other approaches unify actions into one shared space, or adjust the data mixing ratio by source. Open X-Embodiment, which pools data from 22 kinds of robots, is a textbook example of a heterogeneous dataset.","example":"HPT combines 52 datasets — real robots, multiple simulators, and human video — into one pretraining run; each embodiment has its own “stem” that compresses vision and proprioception into a small number of tokens, which are then fed into a shared Transformer backbone.","related":["Cross-Embodiment Data","Cross-Embodiment","Heterogeneous Pre-trained Transformers","Unified Action Space","Embodiment-specific Head","Data Mixture"]},{"id":"robonet","category":"data","sec":6,"tier":3,"sources":[{"title":"RoboNet: Large-Scale Multi-Robot Learning (arXiv:1910.11215)","url":"https://arxiv.org/abs/1910.11215"},{"title":"RoboNet 项目主页","url":"https://www.robonet.wiki/"}],"as_of":"2020-01","related_ids":["cross-embodiment-data","visual-foresight","video-prediction-model","autonomous-data-collection","open-x-embodiment","pre-training"],"name":"RoboNet","alt":"RoboNet 数据集","abbr":"","aliases":["RoboNet: Large-Scale Multi-Robot Learning"],"one_liner":"A 2019 dataset of multi-robot interaction video, about 15 million frames across 7 kinds of robots.","explanation":"RoboNet is an early large-scale multi-robot dataset released in 2019 by researchers at Berkeley, Stanford, Penn, and CMU, including Sergey Levine and Chelsea Finn. It pools about 15 million frames of video and corresponding actions from 7 robot platforms, ranging from the industrial Kuka arm to the low-cost WidowX, as well as the Sawyer, Franka, Baxter, Fetch, and Google's R3. The data was collected autonomously by having robots execute random actions, requiring almost no human demonstration and carrying no task labels; it's mainly used to train action-conditioned video-prediction models for visual foresight planning. The paper found that pretraining on RoboNet and then transferring to a new robot outperformed training from scratch on that robot with 4 to 20 times as much data. It was an early attempt at a cross-embodiment dataset.","example":"A researcher collects only 300–400 random trajectories on a new robot arm, fine-tunes a video-prediction model pretrained on RoboNet, and then uses it to plan pushing actions on objects.","related":["Cross-Embodiment Data","Visual Foresight","Video Prediction Model","Autonomous Data Collection","Open X-Embodiment","Pre-training"]},{"id":"bc-z","category":"data","sec":6,"tier":3,"sources":[{"title":"BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning (CoRL 2021, PMLR)","url":"https://proceedings.mlr.press/v164/jang22a/jang22a.pdf"},{"title":"BC-Z 项目主页","url":"https://sites.google.com/view/bc-z/home"}],"as_of":"2022-02","related_ids":["open-x-embodiment","rt-1","human-gated-dagger","language-conditioned-policy","zero-shot","feature-wise-linear-modulation"],"name":"BC-Z","alt":"BC-Z 数据集","abbr":"","aliases":["BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning","BC-Z Robot Dataset"],"one_liner":"A 2021 Google dataset of real-robot demonstrations across 100 tasks, collected to study zero-shot task generalization.","explanation":"BC-Z is a project from Google and X, Alphabet's “moonshot factory” division, presented at the robot-learning conference CoRL 2021; the name stands for “behavioral cloning + zero-shot.” Using 12 robots and 7 operators, the team collected demonstrations through VR teleoperation combined with “shared autonomy” — the policy acts on its own, and a human takes over just before it would fail, a method called HG-DAgger — gathering 25,877 demonstrations totaling 125 hours across 100 tasks, plus 18,726 videos of humans performing the same tasks. The policy receives a language or human-video task embedding through FiLM layers (a way to inject a conditioning signal into a neural network) and reached 44% average success on 24 held-out tasks it was never trained on. It was an early demonstration that multi-task data plus language conditioning produces zero-shot generalization; the dataset was later folded into the Open X-Embodiment collection.","example":"At test time, the policy is given instructions involving object combinations it never saw paired together during training — for instance, placing an object into a container it was never matched with — and must complete the task with zero additional demonstrations. Across 24 such novel tasks, it reached 44% average success.","related":["Open X-Embodiment","RT-1","Human-Gated DAgger","Language-conditioned Policy","Zero-shot","Feature-wise Linear Modulation"]},{"id":"language-table","category":"data","sec":6,"tier":3,"sources":[{"title":"Interactive Language 项目主页","url":"https://interactive-language.github.io/"},{"title":"google-research/language-table（GitHub）","url":"https://github.com/google-research/language-table"},{"title":"Interactive Language: Talking to Robots in Real Time (arXiv 2210.06407)","url":"https://arxiv.org/abs/2210.06407"}],"as_of":"2022-10","related_ids":["tabletop-manipulation","language-conditioned-policy","instruction-following","hindsight-relabeling","open-x-embodiment","palm-e"],"name":"Language-Table","alt":"Language-Table 数据集","abbr":"","aliases":["Language Table","language_table"],"one_liner":"Google's tabletop block-pushing dataset, with nearly 600,000 trajectories paired with natural-language instructions.","explanation":"Language-Table is a dataset, simulation environment, and benchmark that the Google robotics team open-sourced alongside its 2022 paper Interactive Language: Talking to Robots in Real Time. The setup is simple: a UFACTORY xArm6 arm fitted with a cylindrical end effector pushes colored blocks around on a tabletop; the action space is just 2D planar displacement, at 5Hz. The data was collected using 4 robots and 10 teleoperators, and crowdworkers then reviewed the videos afterward, marking the start and end of each behavior and writing an instruction for it — that is, hindsight relabeling. In total there are nearly 600,000 language-labeled trajectories: 442,000 from the real robot and 181,000 from human-operated simulation, plus several script-generated simulation subsets. The resulting policy can execute roughly 87,000 different instruction phrasings, with an estimated success rate of 93.5%. It's mainly used to study real-time language interaction and instruction following, and was later folded into Open X-Embodiment.","example":"An operator can talk to the robot as they watch: first asking it to push one block next to another, seeing the result, then giving the next instruction, gradually building up a more complex long-horizon goal, with the robot responding to each instruction in real time.","related":["Tabletop Manipulation","Language-conditioned Policy","Instruction Following","Hindsight Relabeling","Open X-Embodiment","PaLM-E"]},{"id":"rt-1-robot-action-dataset","category":"data","sec":6,"tier":3,"sources":[{"title":"RT-1: Robotics Transformer for Real-World Control at Scale（项目页）","url":"https://robotics-transformer1.github.io/"},{"title":"TFDS Catalog: fractal20220817_data","url":"https://www.tensorflow.org/datasets/catalog/fractal20220817_data"}],"as_of":"2022-12","related_ids":["rt-1","open-x-embodiment","simplerenv","everyday-robots-mobile-manipulator","rlds","real-robot-data"],"name":"RT-1 Robot Action Dataset","alt":"RT-1 数据集","abbr":"","aliases":["Fractal","fractal20220817_data","Google Robot Dataset"],"one_liner":"Google's real-robot manipulation dataset, collected over 17 months using 13 robots.","explanation":"The RT-1 dataset is the real-robot data Google collected to train its RT-1 model: 13 Everyday Robots single-arm mobile robots, operating in office-kitchen-like settings, recorded via human teleoperation over 17 months, producing more than 130,000 episodes (one complete task execution) covering over 700 language instructions. Inside Open X-Embodiment it's named fractal20220817_data; the TFDS version contains 87,212 episodes, stored in RLDS format, with each step carrying an image, a language instruction, and an action (end-effector displacement, gripper open/close, base movement, and a termination signal). It was among the largest public real-robot datasets of its time, and it's also the training data that SimplerEnv's “Google Robot” evaluation tasks are built to match.","example":"Loading fractal20220817_data from Open X-Embodiment gives frame-by-frame images and actions for the robot picking up a soda can from a table under an instruction like “pick coke can”; SimplerEnv's Google Robot “pick up the coke can” task was recreated directly from this setup.","related":["RT-1","Open X-Embodiment","SimplerEnv","Everyday Robots Mobile Manipulator","RLDS (Reinforcement Learning Datasets)","Real-Robot Data"]},{"id":"rh20t","category":"data","sec":6,"tier":3,"sources":[{"title":"RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot (arXiv:2307.00595)","url":"https://arxiv.org/abs/2307.00595"},{"title":"RH20T 项目主页","url":"https://rh20t.github.io/"}],"as_of":"2023-09","related_ids":["sjtu-mvig-lab","multimodal-data","contact-rich-manipulation","six-axis-force-torque-sensor","real-robot-data","one-shot-imitation-learning"],"name":"RH20T","alt":"RH20T 数据集","abbr":"","aliases":["RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot"],"one_liner":"A multimodal real-robot manipulation dataset from Shanghai Jiao Tong University, with force, audio, and human demonstration video.","explanation":"RH20T is a real-robot manipulation dataset released in 2023 by Cewu Lu's group (the MVIG lab) at Shanghai Jiao Tong University. It contains over 110,000 contact-rich manipulation sequences covering 147 tasks (48 drawn from RLBench, 29 from Meta-World, and 70 designed by the authors), collected across 7 different robot configurations. Beyond multi-view RGB and depth images, each sequence also includes 6-axis force/torque readings, audio, and action records, with fingertip tactile data on some configurations; every robot sequence is also paired with a video of a human performing the same task plus a language description, aimed at supporting one-shot imitation learning — learning a task from watching a single human demonstration. The full dataset totals about 40TB.","example":"One RH20T sequence provides synchronized RGB-D footage from multiple cameras, 100Hz force/torque readings, dual-microphone audio, and a video of a human performing the same task.","related":["SJTU MVIG Lab","Multimodal Data","Contact-rich Manipulation","Six-Axis Force/Torque Sensor","Real-Robot Data","One-shot Imitation Learning"]},{"id":"bridgedata-v2","category":"data","sec":6,"tier":2,"sources":[{"title":"Walke et al. 2023: BridgeData V2 (arXiv 2308.12952)","url":"https://arxiv.org/abs/2308.12952"},{"title":"BridgeData V2 项目页","url":"https://rail-berkeley.github.io/bridgedata/"},{"title":"GitHub: simpler-env/SimplerEnv","url":"https://github.com/simpler-env/SimplerEnv"}],"as_of":"2024-01","related_ids":["open-x-embodiment","trossen-robotics-widowx-250","simplerenv","openvla","octo","teleoperation"],"name":"BridgeData V2","alt":"BridgeData V2 数据集","abbr":"","aliases":["Bridge Dataset","Bridge V2"],"one_liner":"A roughly 60,000-trajectory multi-task manipulation dataset collected by UC Berkeley on a low-cost WidowX arm.","explanation":"BridgeData V2 was released in August 2023 by UC Berkeley's Sergey Levine group and others, collected entirely on a 6-DoF WidowX 250 arm, totaling 60,096 trajectories: 50,365 are demonstrations from VR-controller teleoperation, and 9,731 come from a scripted pick-and-place policy. It covers 24 environments — a toy kitchen, a sink, a tabletop — and 13 skill categories such as pick-and-place, pushing, wiping, opening and closing drawers, and folding cloth, recorded at a 5 Hz control frequency, with every trajectory paired with a language instruction and released under a CC BY 4.0 license. Because the hardware is cheap and the data is open, it was folded into Open X-Embodiment and is a staple for pretraining and evaluating open models like Octo and OpenVLA.","example":"SimplerEnv recreates Bridge's WidowX setup in simulation, with four tasks — put the spoon on the towel, put the carrot on the plate, stack the green block on the yellow one, and put the eggplant in the basket — designed specifically to evaluate policies trained on Bridge data.","related":["Open X-Embodiment","Trossen Robotics WidowX 250","SimplerEnv","OpenVLA","Octo","Teleoperation"]},{"id":"roboset","category":"data","sec":6,"tier":3,"sources":[{"title":"RoboSet 数据集页面","url":"https://robopen.github.io/roboset/"},{"title":"RoboAgent 项目主页","url":"https://robopen.github.io/"}],"as_of":"2023-09","related_ids":["roboagent","action-chunking","teleoperation","kinesthetic-teaching","multi-view","hierarchical-data-format-version-5"],"name":"RoboSet","alt":"RoboSet 数据集","abbr":"","aliases":[],"one_liner":"A real-robot kitchen-task dataset from CMU and Meta, with multi-task demonstrations and 4 camera views per frame.","explanation":"RoboSet is a real-robot dataset that Carnegie Mellon University and Meta AI made public as part of the RoboAgent project (2023), set mostly in home-kitchen activities. According to the dataset's page, it totals 28,500 trajectories: 9,500 collected through teleoperation with an Oculus Quest 2 controller, and 19,000 obtained through kinesthetic-teaching replay (replaying an existing trajectory in a rearranged version of the scene). Every frame carries 4 camera views, and tasks are organized by “activity,” with each activity containing 4–6 language-instructed subtasks. RoboAgent used about 7,500 of these trajectories to train its MT-ACT policy, which learned 12 manipulation skills. The data is released in HDF5 format under the MIT license.","example":"Opening one RoboSet trajectory with h5py gives images from 4 camera views, the robot arm's state, the corresponding actions, and the language instruction for that task.","related":["RoboAgent","Action Chunking","Teleoperation","Kinesthetic Teaching","Multi-View","Hierarchical Data Format version 5"]},{"id":"droid","category":"data","sec":6,"tier":2,"sources":[{"title":"DROID 项目主页","url":"https://droid-dataset.github.io/"},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (arXiv 2403.12945)","url":"https://arxiv.org/html/2403.12945v2"},{"title":"The DROID Dataset (DROID Docs)","url":"https://droid-dataset.github.io/droid/the-droid-dataset"}],"as_of":"2025-04","related_ids":["open-x-embodiment","real-robot-data","vr-teleoperation","data-diversity","rlds","franka-emika-panda-franka-research-3"],"name":"DROID (Distributed Robot Interaction Dataset)","alt":"DROID 数据集","abbr":"DROID","aliases":["DROID"],"one_liner":"A real-robot manipulation dataset collected by 50 people across 564 scenes on three continents using Franka arms.","explanation":"DROID was collected jointly by 13 institutions including Stanford and UC Berkeley, and released in March 2024. Every site used identical hardware: a 7-DoF Franka Panda arm mounted on a height-adjustable mobile cart, with two third-person ZED 2 stereo cameras and a wrist-mounted ZED Mini, with collectors using a Quest 2 VR controller for teleoperation. Over 12 months, 50 collectors across labs, offices, and homes in North America, Asia, and Europe gathered 76,000 successful trajectories totaling about 350 hours, spanning 564 scenes, 86 task categories, and 1,417 camera viewpoints, along with roughly 16,000 trajectories labeled as failures. It emphasizes “in-the-wild” scene diversity, and models such as π0 include DROID in their pretraining data.","example":"For training a policy, one can download the full RLDS-format DROID directly (about 1.7 TB); when high-resolution stereo footage or depth data is needed, use the raw-data release instead (about 8.7 TB).","related":["Open X-Embodiment","Real-Robot Data","VR Teleoperation","Data Diversity","RLDS (Reinforcement Learning Datasets)","Franka Emika Panda / Franka Research 3"]},{"id":"all-robots-in-one","category":"data","sec":6,"tier":3,"sources":[{"title":"All Robots in One: A New Standard and Unified Dataset for Versatile, General-Purpose Embodied Agents (arXiv 2408.10899)","url":"https://arxiv.org/abs/2408.10899"},{"title":"ARIO 项目主页","url":"https://imaei.github.io/project_pages/ario/"}],"as_of":"2024-08","related_ids":["cross-embodiment-data","open-x-embodiment","multimodal-data","heterogeneous-data","agilex-cobot-magic","rlds"],"name":"All Robots In One","alt":"ARIO 数据集","abbr":"ARIO","aliases":["ARIO","ARIO Data Standard"],"one_liner":"A unified embodied-data format standard from Peng Cheng Laboratory and partners, and a roughly 3-million-item dataset built to that standard.","explanation":"ARIO was released in August 2024 by the Institute of Multi-Agent and Embodied Intelligence at Peng Cheng Laboratory, together with Southern University of Science and Technology, Sun Yat-sen University, and others; it is both a data-format standard and a large dataset assembled under it. The authors argue that aggregating datasets such as Open X-Embodiment still leave formats inconsistent and modalities incomplete, so ARIO specifies: control data from robots of different forms recorded in one unified format, sensors of differing frequency aligned by timestamp, organization by a three-level “series–task–episode” hierarchy, and support for five modalities — image, 3D, sound, text, and touch. The dataset contains about 3 million episodes, 258 series, and more than 320,000 tasks, drawn from three sources: 3,662 episodes from a self-built real-robot platform; about 700,000 from simulators such as Habitat and MuJoCo; and roughly 2.33 million converted from existing open datasets.","example":"ARIO's real-robot portion used the AgileX Cobot Magic dual-arm platform to collect more than 30 kinds of manipulation tasks in real household scenes.","related":["Cross-Embodiment Data","Open X-Embodiment","Multimodal Data","Heterogeneous Data","AgileX Cobot Magic","RLDS (Reinforcement Learning Datasets)"]},{"id":"robomind","category":"data","sec":6,"tier":2,"sources":[{"title":"RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation","url":"https://arxiv.org/abs/2412.13877"},{"title":"RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence","url":"https://arxiv.org/abs/2512.24653"},{"title":"RoboMIND 项目主页","url":"https://x-humanoid-robomind.github.io/"}],"as_of":"2025-12","related_ids":["cross-embodiment-data","real-robot-data","failure-data","agibot-world","open-x-embodiment","beijing-humanoid-robot-innovation-center"],"name":"RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation)","alt":"RoboMIND 数据集","abbr":"","aliases":["RoboMIND 2.0"],"one_liner":"A multi-embodiment real-robot manipulation dataset from the Beijing Humanoid Robot Innovation Center and Peking University, collected under a unified protocol.","explanation":"RoboMIND was released in December 2024 by the Beijing Humanoid Robot Innovation Center, Peking University's School of Computer Science, and the Beijing Academy of Artificial Intelligence, with the paper accepted at RSS 2025. Under a unified collection protocol, it gathered 107,000 demonstrations on four embodiments — Franka, UR5e, an AgileX dual-arm platform, and a humanoid with two dexterous hands — covering 479 tasks and 96 object categories, plus 5,000 failed demonstrations with the reason labeled, and it built a digital-twin environment in Isaac Sim; every trajectory is stored as one HDF5 file. RoboMIND 2.0, from December 2025, expanded to six embodiments, 739 tasks, and more than 310,000 dual-arm trajectories, adding 12,000 tactile segments, 20,000 mobile-manipulation trajectories, and 20,000 simulated trajectories, and proposed a hierarchical dual-system framework called MIND-2. The data is downloadable on ModelScope and Hugging Face.","example":"A researcher could take only RoboMIND's UR5e single-arm trajectories to train an imitation-learning policy, then compare against the failed demonstrations for the same task to see where the policy tends to go wrong.","related":["Cross-Embodiment Data","Real-Robot Data","Failure Data","AgiBot World","Open X-Embodiment","Beijing Humanoid Robot Innovation Center"]},{"id":"agibot-world","category":"data","sec":6,"tier":2,"sources":[{"title":"AgiBot World Colosseo 技术报告 (arXiv 2503.06669)","url":"https://arxiv.org/abs/2503.06669"},{"title":"GitHub: OpenDriveLab/AgiBot-World","url":"https://github.com/OpenDriveLab/AgiBot-World"},{"title":"Hugging Face: agibot-world/AgiBotWorld2026","url":"https://huggingface.co/datasets/agibot-world/AgiBotWorld2026"}],"as_of":"2026-09","related_ids":["agibot","agibot-go-1","open-x-embodiment","embodied-ai-training-ground","real-robot-data","agibot-genie-g1"],"name":"AgiBot World","alt":"AgiBot World 数据集","abbr":"","aliases":["AgiBot World Colosseo","AgiBot World Alpha","AgiBot World Beta","AgiBot World 2026"],"one_liner":"A large-scale real-robot manipulation dataset from AgiBot and partners, with the Beta release exceeding a million trajectories.","explanation":"AgiBot World is a real-robot manipulation dataset released by AgiBot Robotics together with the University of Hong Kong, the Shanghai AI Laboratory, and others. The Alpha release, in December 2024, contained about 92,000 trajectories; the Beta release, in March 2025, expanded to over 1 million trajectories (about 2,976 hours), covering 217 tasks across 5 scene categories, recorded by 100 AgiBot G1 robots teleoperated with VR and motion capture inside roughly 4,000 square meters of collection space, with task- and substep-level language annotations checked by human quality control; AgiBot's own GO-1 model was trained on it. AgiBot World 2026, a preview released in 2026, switches to the G2 robot collecting in real commercial spaces and homes. It is licensed CC BY-NC-SA 4.0 and cannot be used commercially.","example":"The technical report states that a policy pretrained on AgiBot World outperformed one pretrained on Open X-Embodiment by about 30% on average, both in-distribution and out-of-distribution.","related":["AgiBot","AgiBot GO-1","Open X-Embodiment","Embodied AI Training Ground (Robot Data Collection Center)","Real-Robot Data","AgiBot Genie G1"]},{"id":"fourier-actionnet","category":"data","sec":6,"tier":3,"sources":[{"title":"Fourier ActionNet 官网","url":"https://action-net.org/"}],"as_of":"2025","related_ids":["fourier","fourier-gr-2","vr-teleoperation","dexterous-hand","lerobotdataset","real-robot-data"],"name":"Fourier ActionNet","alt":"傅利叶 ActionNet 数据集","abbr":"","aliases":["ActionNet"],"one_liner":"Fourier's open-source teleoperated dataset of about 30,000 humanoid-plus-dexterous-hand manipulation trajectories.","explanation":"Fourier ActionNet is a real-robot dataset open-sourced in 2025 by the humanoid robot company Fourier Intelligence. The data was collected by operators performing VR teleoperation while wearing an Apple Vision Pro headset, driving Fourier's GR1-T1, GR1-T2, and GR2 humanoid robots, fitted with either 6-DoF or 12-DoF dexterous hands. It totals over 30,000 trajectories and about 140 hours, mostly tabletop bimanual manipulation such as pick-and-place, pouring water, opening and closing cabinet doors, and precise placement. Each trajectory's language instruction was first auto-generated by Qwen2.5-VL-7B and then manually checked in full. The data uses the LeRobot dataset format and can be used directly to train methods like ACT, Diffusion Policy, and iDP3; the dataset itself is released under a non-commercial license, while the training code is Apache 2.0.","example":"One water-pouring trajectory: a GR2 robot picks up a container with its dexterous hand and pours water into a cup, while first-person camera footage, joint states, and a language instruction are all recorded together.","related":["Fourier","Fourier GR-2","VR Teleoperation","Dexterous Hand","LeRobotDataset","Real-Robot Data"]},{"id":"galaxea-open-world-dataset","category":"data","sec":6,"tier":3,"sources":[{"title":"Galaxea Open-World Dataset and G0 Dual-System VLA Model (arXiv)","url":"https://arxiv.org/abs/2509.00576"},{"title":"Galaxea Open-World Dataset (Hugging Face)","url":"https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset"},{"title":"GalaxeaVLA 项目主页","url":"https://opengalaxea.github.io/GalaxeaVLA/"}],"as_of":"2025-08","related_ids":["galaxea-ai","galaxea-r1","galaxea-g0-dual-system-vla","real-robot-data","subtask-segmentation","lerobotdataset"],"name":"Galaxea Open-World Dataset","alt":"星海图开放世界数据集","abbr":"","aliases":[],"one_liner":"A roughly 500-hour dataset Galaxea collected at real-world locations using a single model of mobile bimanual robot.","explanation":"The Galaxea Open-World Dataset is a real-robot dataset the embodied-AI company Galaxea (星海图) open-sourced in August 2025 alongside its G0 model paper. All the data was collected using the company's own R1 Lite mobile bimanual robot (23 degrees of freedom) at 11 real-world locations — homes, restaurants, retail stores, and offices — totaling about 100,000 trajectories, 150 task categories, 50 scenes, more than 1,600 object types, and 58 manipulation skills, adding up to roughly 500 hours. Each trajectory is split into atomic subtasks, with annotators choosing wording from standardized descriptions, and every subtask is labeled in both Chinese and English. The data is in LeRobot format, includes 4 camera streams (head plus both wrists), joint states, and IMU data, and can be downloaded from Hugging Face and the ModelScope community. Its distinguishing feature is using a single robot embodiment in real environments; the paper found that pretraining on this single-embodiment data was critical to the G0 model's performance.","example":"In the “tidy up fruit” task, the R1 Lite arranges fruit in a real environment; the whole trajectory is split into several subtasks, each labeled with a description in both Chinese and English.","related":["Galaxea AI","Galaxea R1","Galaxea G0 Dual-System VLA","Real-Robot Data","Subtask Segmentation","LeRobotDataset"]},{"id":"humanoid-everyday","category":"data","sec":6,"tier":3,"sources":[{"title":"Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation (arXiv)","url":"https://arxiv.org/abs/2510.08807"},{"title":"Humanoid Everyday 项目主页","url":"https://humanoideveryday.github.io/"}],"as_of":"2025-10","related_ids":["humanoid-robot","unitree-g1","unitree-h1","vr-teleoperation","apple-vision-pro","real-world-evaluation"],"name":"Humanoid Everyday","alt":"Humanoid Everyday 数据集","abbr":"","aliases":["Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation"],"one_liner":"A real-robot humanoid manipulation dataset from USC and the Toyota Research Institute, covering 260 tasks.","explanation":"Humanoid Everyday was released by the University of Southern California and the Toyota Research Institute in October 2025. Most existing robot datasets come from fixed robot arms, and humanoid-specific data has been scarce and narrow in task coverage; this dataset was built to fill that gap. The data was collected using a Unitree G1 (29 degrees of freedom, fitted with a Dex3-1 three-fingered dexterous hand) and a Unitree H1 (27 degrees of freedom, fitted with an Inspire dexterous hand), with operators teleoperating via Apple Vision Pro headsets. It totals 10,300 trajectories, over 3 million frames, and 260 tasks, organized into 7 categories: basic manipulation, deformable objects, articulated objects, tool use, high-precision manipulation, human-robot interaction, and mobile manipulation. Each trajectory synchronously records RGB, depth, LiDAR, tactile, IMU, and language description data at 30Hz. The authors also built a cloud evaluation platform where researchers can submit a policy and have it tested in their controlled real-robot setup. The data is released under the CC BY 4.0 license.","example":"","related":["Humanoid Robot","Unitree G1","Unitree H1","VR Teleoperation","Apple Vision Pro","Real-World Evaluation"]},{"id":"robocoin","category":"data","sec":6,"tier":3,"sources":[{"title":"RoboCOIN (arXiv:2511.17441)","url":"https://arxiv.org/abs/2511.17441"},{"title":"RoboCOIN 项目主页","url":"https://flagopen.github.io/RoboCOIN/"}],"as_of":"2026-04","related_ids":["beijing-academy-of-artificial-intelligence","bimanual-manipulation","cross-embodiment-data","lerobotdataset","subtask-segmentation","agibot-world"],"name":"RoboCOIN","alt":"RoboCOIN 数据集","abbr":"","aliases":["RoboCoin","RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation"],"one_liner":"A cross-embodiment bimanual manipulation dataset led by BAAI, open-sourced across 15 kinds of robots.","explanation":"RoboCOIN is a bimanual-manipulation dataset led by the Beijing Academy of Artificial Intelligence (BAAI), open-sourced in November 2025 together with more than twenty universities and robotics companies. It contains over 180,000 demonstrations from 15 robot platforms, including bimanual robots, half-body humanoids, and full-body humanoids (models from manufacturers such as AgileX, Galaxea, AgiBot, Galbot, Leju Robotics, Unitree, and others), covering 16 environments, 421 bimanual tasks, and 432 objects, with tasks organized into 39 categories of bimanual coordination behavior. Annotation is done at three levels: a task concept for the whole trajectory, segmented subtasks, and frame-by-frame kinematic information. A companion processing pipeline, CoRobot, built on LeRobot, handles data quality control, automatic annotation, and unified management of cross-embodiment data.","example":"A single bimanual demonstration carries three layers of annotation: the whole trajectory states what task is being done, segments mark the start and end of each subtask, and each frame carries kinematic data; the data can be loaded directly for training using tools from the LeRobot ecosystem.","related":["Beijing Academy of Artificial Intelligence","Bimanual Manipulation","Cross-Embodiment Data","LeRobotDataset","Subtask Segmentation","AgiBot World"]},{"id":"molmoact2-bimanualyam","category":"data","sec":6,"tier":3,"sources":[{"title":"Ai2 博客：MolmoAct 2","url":"https://allenai.org/blog/molmoact2"},{"title":"allenai/MolmoAct2-BimanualYAM-Dataset（Hugging Face）","url":"https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset"},{"title":"MolmoAct2: Action Reasoning Models for Real-world Deployment (arXiv 2605.02881)","url":"https://arxiv.org/abs/2605.02881"}],"as_of":"2026-06","related_ids":["molmoact","i2rt-yam-arm","bimanual-manipulation","teleoperation","lerobotdataset","allen-institute-for-ai"],"name":"MolmoAct2-BimanualYAM","alt":"BimanualYAM 数据集","abbr":"","aliases":["MolmoAct2-BimanualYAM Dataset","MolmoAct2-Bimanual YAM"],"one_liner":"Ai2's open-source bimanual teleoperation dataset for MolmoAct2, totaling more than 720 hours.","explanation":"MolmoAct2-BimanualYAM is a bimanual-manipulation dataset the Allen Institute for AI (Ai2) open-sourced in May 2026 alongside its MolmoAct2 model, collected and curated with support from Cortex AI. The data consists of teleoperated demonstrations using two I2RT YAM robot arms (6 joints plus 1 gripper per arm, giving a 14-dimensional joint-position action space), totaling more than 720 hours; the merged version on Hugging Face contains 32,246 trajectories and about 76 million frames (at 30 frames per second), mostly tabletop bimanual tasks such as folding a towel, scanning items, and charging a phone, with each trajectory carrying an annotated language instruction. Ai2 describes it as the largest open bimanual dataset to date, more than 30 times the amount of robot data used for the original MolmoAct. The data uses the LeRobot v3.0 format under the Apache 2.0 license, and can be used directly to pretrain or fine-tune a bimanual VLA (vision-language-action model).","example":"MolmoAct2 used this dataset as one of its main training sources, mixed with other robot datasets; Ai2 says the resulting model can perform bimanual manipulation without needing separate fine-tuning for each task.","related":["MolmoAct","I2RT YAM Arm","Bimanual Manipulation","Teleoperation","LeRobotDataset","Allen Institute for AI"]},{"id":"robovqa","category":"data","sec":6,"tier":3,"sources":[{"title":"RoboVQA: Multimodal Long-Horizon Reasoning for Robotics (arXiv:2311.00899)","url":"https://arxiv.org/abs/2311.00899"}],"as_of":"2023-11","related_ids":["visual-question-answering","embodied-reasoning","long-horizon-task","vision-language-model","google-deepmind","language-annotation"],"name":"RoboVQA","alt":"RoboVQA 数据集","abbr":"","aliases":["RoboVQA: Multimodal Long-Horizon Reasoning for Robotics"],"one_liner":"A video question-answering dataset from Google DeepMind, built around long-horizon robot tasks.","explanation":"RoboVQA is a dataset and modeling effort released by Google DeepMind (Pierre Sermanet and others) in November 2023. The dataset contains about 830,000 video-text pairs across 29,500 distinct instructions, with questions centered on long-horizon tasks: what to do next, whether the current step is finished, or whether a certain action can be performed right now. Collection follows a “bottom-up” crowdsourcing approach: open-ended long-horizon tasks are collected first, then executed and annotated by a robot, a person, or a person holding a handheld grasping tool, which the paper reports gives 2.2 times the throughput of traditional step-by-step collection. The authors used it to train a video vision-language model, RoboVQA-VideoCoCa, measuring performance with a unified metric based on the human-intervention rate; the results show that models taking video input have a 19% lower average error rate than models that only see a single image.","example":"Shown a video of a robot in a kitchen and asked “is the task done?” or “what should happen next?”, the model answers in text, for instance “put the sponge in the sink.”","related":["Visual Question Answering","Embodied Reasoning","Long-horizon Task","Vision-Language Model","Google DeepMind","Language Annotation"]},{"id":"baihu-vtouch-visuo-tactile-dataset","category":"data","sec":6,"tier":3,"sources":[{"title":"国地中心发布全球首个跨本体视触觉多模态数据集「白虎-VTouch」（腾讯新闻）","url":"https://news.qq.com/rain/a/20260126A04W2D00"}],"as_of":"2026-01","related_ids":["vision-based-tactile-sensor","tactile-data","cross-embodiment-data","weitai-robotics","qinglong","national-and-local-co-built-humanoid-robotics-innovation-cen"],"name":"Baihu-VTouch Visuo-Tactile Dataset","alt":"白虎-VTouch 视触觉数据集","abbr":"","aliases":["VTouch"],"one_liner":"A cross-embodiment visuotactile dataset released in 2026 by China's National Humanoid Robot Innovation Center and Weitai Robotics.","explanation":"Baihu-VTouch was released on January 26, 2026 by the National and Local Co-built Humanoid Robotics Innovation Center (Shanghai) together with Shanghai Weitai Robotics, a visuotactile-sensor company. A visuotactile sensor uses a built-in camera to photograph the deformation of a soft gel surface as it's pressed, sensing the shape and force of contact. According to the release, the dataset totals over 60,000 minutes and about 90.72 million real object-contact samples, recording visuotactile images, RGB-D footage, and joint poses in sync; the embodiments used include wheeled-arm robots, the bipedal humanoid Qinglong, and handheld smart terminals. Tasks span 4 scene categories — household chores, industrial manufacturing, food service, and specialized operations — with more than 380 tasks, over 100 atomic skills, and more than 500 real objects in total. Most existing large-scale robot data has only vision and action; this dataset adds the missing dimension of contact. The first 6,000 minutes have already been released on the OpenLoong open-source community.","example":"The dataset includes more than 260 contact-intensive tasks; when training a manipulation policy that needs a sense of “feel,” these visuotactile images can be used as input alongside vision and pose.","related":["Vision-Based Tactile Sensor","Tactile Data","Cross-Embodiment Data","Weitai Robotics","Qinglong (OpenLoong)","National and Local Co-built Humanoid Robotics Innovation Center (Shanghai)"]},{"id":"daimon-infinity","category":"data","sec":6,"tier":3,"sources":[{"title":"戴盟联合数十家头部机构，发布全球最大规模含触觉全模态物理世界数据集（魔搭社区）","url":"https://modelscope.csdn.net/69e03dd70a2f6a37c5a04836.html"},{"title":"Daimon-Infinity 数据集（戴盟官网）","url":"https://www.dmrobot.com/daas/daimon-infinity.html"},{"title":"中国移动与戴盟机器人共建含触觉全模态物理世界数据集（新华网）","url":"http://www.zj.xinhuanet.com/20260421/603839d919c34af49dcbd0d7677e7786/c.html"}],"as_of":"2026-09","related_ids":["tactile-data","vision-based-tactile-sensor","robot-free-data-collection","daimon-robotics","daimon-dm-tac-visuotactile-sensor","modelscope"],"name":"Daimon-Infinity","alt":"戴盟 Daimon-Infinity 数据集","abbr":"","aliases":["Daimon Infinity Dataset"],"one_liner":"A large-scale real-robot manipulation dataset with vision-and-touch signals, released by Daimon Robotics in 2026.","explanation":"Daimon-Infinity is an embodied-AI dataset released on April 15, 2026 by Shenzhen-based Daimon Robotics (戴盟机器人), a company whose core business is vision-based tactile sensors, together with dozens of partner institutions; its selling point is full-modality data that includes touch. Most of the data comes from Daimon's own embodiment-free two-finger grippers and five-finger gloves, fitted with a vision-based tactile sensor — one that measures contact by filming the deformation of an elastic surface — running at 120Hz with 110,000 sensing units, plus fisheye cameras, encoders, an IMU (inertial measurement unit), and stereo cameras; the release also includes head-mounted first-person, teleoperated, and simulated data. At launch, Daimon said it planned to open-source 10,000 hours covering 16 industries, 80 scenes, and more than 2,000 task categories, with the first 1,000 hours already posted on the ModelScope community; the company's website now states the total dataset exceeds a million hours. It fills a gap in most manipulation datasets, which lack signals like contact force and slip.","example":"The open-sourced portion includes more than 1,400 long-sequence tasks lasting over 40 seconds each, such as grasping, inserting, and stacking using the gripper or glove; the data records vision-tactile signals, video, and action trajectories together, and can be used to train manipulation policies that take tactile input.","related":["Tactile Data","Vision-Based Tactile Sensor","Robot-free (Embodiment-free) Data Collection","Daimon Robotics","Daimon DM-Tac Visuotactile Sensor","ModelScope"]},{"id":"hierarchical-data-format-version-5","category":"data","sec":7,"tier":2,"sources":[{"title":"The HDF Group: HDF5","url":"https://www.hdfgroup.org/solutions/hdf5/"},{"title":"robomimic Docs: Datasets Overview","url":"https://robomimic.github.io/docs/datasets/overview.html"},{"title":"ACT (tonyzhaozh/act) GitHub","url":"https://github.com/tonyzhaozh/act"}],"as_of":"","related_ids":["lerobotdataset","rlds","zarr","apache-parquet","robomimic","action-chunking-with-transformers"],"name":"Hierarchical Data Format version 5","alt":"HDF5 格式","abbr":"HDF5","aliases":["HDF5",".h5 File",".hdf5 File"],"one_liner":"A general-purpose file format that stores many large arrays hierarchically in one file, commonly used for robot demonstration data.","explanation":"HDF5 is an open-source file format and software library maintained by the HDF Group, usually with a .h5 or .hdf5 extension. It works like a folder system inside a single file: “groups” provide the hierarchy, each group holds multidimensional arrays (called datasets), metadata can be attached anywhere, and any part of the file can be read without loading the rest. A robot demonstration contains multiple camera streams, joint state, and actions as time series, which fits naturally into HDF5. robomimic stores a whole dataset this way, with each trajectory as a group like data/demo_0; the ACT/ALOHA collection scripts likewise save each episode as its own HDF5 file. Storing images verbatim makes files large, so newer open datasets increasingly use the LeRobot format or RLDS instead.","example":"In a robomimic dataset file, data/demo_0/actions is one trajectory's action array, shaped (number of steps, action dimension); data/demo_0/obs/agentview_image holds the camera image for each of those same time steps.","related":["LeRobotDataset","RLDS (Reinforcement Learning Datasets)","Zarr","Apache Parquet","robomimic","Action Chunking with Transformers"]},{"id":"zarr","category":"data","sec":7,"tier":3,"sources":[{"title":"Zarr 官网","url":"https://zarr.dev/"},{"title":"UMI Robot Dataset Community: Zarr data format","url":"https://umi-data.github.io"}],"as_of":"","related_ids":["hierarchical-data-format-version-5","webdataset","apache-parquet","lerobotdataset","diffusion-policy","universal-manipulation-interface"],"name":"Zarr","alt":"Zarr 格式","abbr":"","aliases":[".zarr.zip"],"one_liner":"An open-source format for storing large multidimensional arrays in compressed chunks; Diffusion Policy and UMI use it for training data.","explanation":"Zarr is an open-source, chunked, compressed N-dimensional array storage format, community-maintained with financial support from NumFOCUS, with both v2 and v3 specifications; implementations exist in about ten languages including Python, Rust, C++, and JavaScript, and it was first popular in scientific data fields like climate science and bioimaging. It splits a large array into many small chunks, compressing each one separately, so reading one section only needs to decompress the relevant chunks; the data can live in a local directory, a zip file, or cloud object storage, and supports parallel reads and writes. Conceptually it can be thought of as a “nested dictionary of NumPy arrays,” playing a role close to HDF5, but with a more flexible choice of storage backend, chunking, and compression. In robot learning, the Diffusion Policy codebase uses Zarr to store training data, and UMI follows this same approach, packaging one collection session into a single dataset.zarr.zip file for fast random access during training.","example":"UMI's example dataset.zarr.zip contains arrays such as camera0_rgb (images sized 2315×224×224×3), robot0_eef_pos (end-effector position), and robot0_gripper_width (gripper opening width), plus an episode_ends array recording which frame each demonstration ends on.","related":["Hierarchical Data Format version 5","WebDataset","Apache Parquet","LeRobotDataset","Diffusion Policy","Universal Manipulation Interface"]},{"id":"rlds","category":"data","sec":7,"tier":2,"sources":[{"title":"RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning","url":"https://arxiv.org/abs/2111.02767"},{"title":"google-research/rlds (GitHub)","url":"https://github.com/google-research/rlds"},{"title":"google-deepmind/open_x_embodiment (GitHub)","url":"https://github.com/google-deepmind/open_x_embodiment"}],"as_of":"2025-11","related_ids":["open-x-embodiment","tensorflow-datasets","tfrecord","lerobotdataset","hierarchical-data-format-version-5","episode"],"name":"RLDS (Reinforcement Learning Datasets)","alt":"RLDS 格式","abbr":"RLDS","aliases":["RLDS"],"one_liner":"A Google-proposed standard format and toolchain for organizing sequential-decision data as episodes made of steps.","explanation":"RLDS is a data ecosystem Google Research released in 2021 for recording, sharing, and processing reinforcement learning, offline RL, and imitation learning data. Data has two levels: an episode carries metadata such as an episode_id and contains a sequence of steps; every step has two required flags, is_first and is_last, plus optional fields such as observation, action, reward, and discount. A companion tool called EnvLogger records the data, and it is published and read through TensorFlow Datasets (TFDS). This lets data from different labs be read with the same code: every sub-dataset in Open X-Embodiment is provided in RLDS episode format, and the training code for Octo and OpenVLA reads it directly. The PyTorch ecosystem has increasingly moved toward LeRobotDataset in recent years, and the RLDS GitHub repository was archived in November 2025.","example":"Use tfds.load to load the BridgeData V2 subset inside Open X-Embodiment, iterate episode by episode over its steps, and pull out each step's image, language instruction, and action for training.","related":["Open X-Embodiment","TensorFlow Datasets (TFDS)","TFRecord","LeRobotDataset","Hierarchical Data Format version 5","Episode"]},{"id":"tensorflow-datasets","category":"data","sec":7,"tier":3,"sources":[{"title":"TensorFlow Datasets Overview","url":"https://www.tensorflow.org/datasets/overview"},{"title":"google-deepmind/open_x_embodiment (GitHub)","url":"https://github.com/google-deepmind/open_x_embodiment"},{"title":"google-research/rlds (GitHub)","url":"https://github.com/google-research/rlds"}],"as_of":"","related_ids":["rlds","tfrecord","open-x-embodiment","tensorflow","lerobotdataset"],"name":"TensorFlow Datasets (TFDS)","alt":"TFDS（TensorFlow Datasets）","abbr":"TFDS","aliases":["tensorflow_datasets"],"one_liner":"Google's open-source dataset library that downloads and loads a dataset into a trainable pipeline in one line of code.","explanation":"TFDS is an open-source dataset library maintained by Google, cataloging thousands of ready-to-use datasets spanning images, text, audio, robotics, and more. Every dataset has a builder, responsible for downloading the raw data, converting it through a fixed pipeline, and storing it as sharded TFRecord or ArrayRecord files; calling tfds.load returns a tf.data.Dataset, which can also be converted to NumPy for use with JAX or PyTorch, along with a version number and metadata. It's a layer built on top of tf.data, not tf.data itself. In embodied AI, the RLDS format and the Open X-Embodiment dataset are both released following the TFDS convention, and the training code for Octo and OpenVLA reads data through it as well.","example":"Pointing tfds.builder_from_directory at Open X-Embodiment's fractal20220817_data directory on Google Cloud Storage lets you iterate through the RT-1 dataset episode by episode, getting images, language instructions, and actions.","related":["RLDS (Reinforcement Learning Datasets)","TFRecord","Open X-Embodiment","TensorFlow","LeRobotDataset"]},{"id":"tfrecord","category":"data","sec":7,"tier":3,"sources":[{"title":"TensorFlow Tutorial: TFRecord and tf.train.Example","url":"https://www.tensorflow.org/tutorials/load_data/tfrecord"},{"title":"TensorFlow Datasets Overview","url":"https://www.tensorflow.org/datasets/overview"}],"as_of":"","related_ids":["tensorflow-datasets","rlds","protocol-buffers","apache-parquet","hierarchical-data-format-version-5","webdataset"],"name":"TFRecord","alt":"TFRecord 格式","abbr":"","aliases":[".tfrecord","tf.train.Example"],"one_liner":"TensorFlow's binary file format, storing a sequence of serialized records back to back.","explanation":"TFRecord is TensorFlow's data storage format: a file consists of binary records placed one after another, each with a length field and a CRC checksum, and the content is usually a tf.train.Example encoded with protobuf (Google's serialization protocol) — essentially a dictionary mapping field names to lists of values. It's well suited to reading large amounts of data sequentially, and the official guidance recommends splitting data into multiple shards, ideally over 100MB each, to support parallel reads. The downside is that it can only be scanned sequentially — it's awkward to randomly fetch one record by index — and parsing depends on TensorFlow. TFDS uses it as the default storage format, so datasets like RLDS and Open X-Embodiment download as a set of TFRecord shards; newer datasets in the PyTorch ecosystem have largely shifted to Parquet plus video, HDF5, or Zarr instead.","example":"An RLDS dataset directory typically contains dataset_info.json and features.json, plus shard files named something like xxx-train.tfrecord-00000-of-01024, where each record stores one complete episode.","related":["TensorFlow Datasets (TFDS)","RLDS (Reinforcement Learning Datasets)","Protocol Buffers","Apache Parquet","Hierarchical Data Format version 5","WebDataset"]},{"id":"lerobotdataset","category":"data","sec":7,"tier":2,"sources":[{"title":"Hugging Face LeRobot Docs: LeRobotDataset v3.0","url":"https://huggingface.co/docs/lerobot/lerobot-dataset-v3"},{"title":"Hugging Face Blog: LeRobotDataset v3.0","url":"https://huggingface.co/blog/lerobot-datasets-v3"}],"as_of":"2025-09","related_ids":["lerobot","hugging-face","apache-parquet","hierarchical-data-format-version-5","rlds","so-100-so-101-arm"],"name":"LeRobotDataset","alt":"LeRobot 数据集格式","abbr":"","aliases":["LeRobot Dataset","LeRobotDataset v3.0"],"one_liner":"LeRobot's standard dataset format, storing state and actions in tables and camera footage as video.","explanation":"LeRobotDataset is the dataset format used by LeRobot, Hugging Face's open-source robotics library, and can be shared directly on the Hub. It splits data into three parts: low-dimensional signals such as state, actions, and timestamps are stored as Parquet tables; camera footage is encoded as MP4 video; and a meta directory records field definitions, frame rate, normalization statistics, task text, and each episode's start and end position. Version 3.0, released in September 2025, packs multiple episodes into the same file (v2.1 used one file per episode) and locates them by metadata, supporting much larger scale, and it can also be streamed directly from the Hub without downloading. Models in LeRobot such as ACT, Diffusion Policy, and SmolVLA all read this format directly.","example":"Use lerobot-record to teleoperate an SO-101 arm while recording; the data is saved in v3.0 format and pushed to the Hub. During training, LeRobotDataset loads it, and each sample is a tensor dictionary containing observation.state, action, and camera images.","related":["LeRobot","Hugging Face","Apache Parquet","Hierarchical Data Format version 5","RLDS (Reinforcement Learning Datasets)","SO-100 / SO-101 Arm"]},{"id":"apache-parquet","category":"data","sec":7,"tier":3,"sources":[{"title":"Apache Parquet 官方文档 Overview","url":"https://parquet.apache.org/docs/overview/"},{"title":"LeRobotDataset v3.0 文档（Hugging Face）","url":"https://huggingface.co/docs/lerobot/lerobot-dataset-v3"},{"title":"Apache Parquet - Wikipedia","url":"https://en.wikipedia.org/wiki/Apache_Parquet"}],"as_of":"2026-09","related_ids":["lerobotdataset","hierarchical-data-format-version-5","rlds","zarr","lerobot","mcap"],"name":"Apache Parquet","alt":"Parquet 格式","abbr":"","aliases":["Parquet"],"one_liner":"An open-source column-oriented file format; LeRobot datasets use it to store numeric data like state and actions.","explanation":"Apache Parquet is an open-source column-oriented data file format, introduced by Twitter and Cloudera in 2013 and made an Apache top-level project in 2015. “Column-oriented” means values from the same column are stored contiguously rather than row by row; grouping similar data together compresses better and lets a reader pull out only the columns it needs. Tools such as Spark, pandas, and DuckDB can all read and write it directly. In embodied AI, Hugging Face's LeRobot dataset format uses Parquet to store low-dimensional, high-frequency data such as joint state, actions, and timestamps, while camera footage is separately encoded as MP4 video, with JSON/Parquet metadata recording where each episode starts and ends within the files. Version 3 packs multiple episodes into the same batch of Parquet files and supports streaming directly from the Hub.","example":"Open a LeRobot dataset's data/ directory files with pandas.read_parquet to see columns such as observation.state, action, and timestamp for every frame.","related":["LeRobotDataset","Hierarchical Data Format version 5","RLDS (Reinforcement Learning Datasets)","Zarr","LeRobot","MCAP"]},{"id":"webdataset","category":"data","sec":7,"tier":3,"sources":[{"title":"webdataset/webdataset (GitHub)","url":"https://github.com/webdataset/webdataset"},{"title":"Efficient PyTorch I/O library for Large Datasets, Many Files, Many GPUs (PyTorch Blog)","url":"https://pytorch.org/blog/efficient-pytorch-io-library-for-large-datasets-many-files-many-gpus"},{"title":"BeingBeyond/UniHand_Preview (Hugging Face)","url":"https://huggingface.co/datasets/BeingBeyond/UniHand_Preview"}],"as_of":"","related_ids":["apache-parquet","hierarchical-data-format-version-5","zarr","tfrecord","lerobotdataset","rlds"],"name":"WebDataset","alt":"WebDataset 格式","abbr":"","aliases":["wds"],"one_liner":"A format that packages training samples into same-named files inside a series of tar shards, for fast large-scale sequential reads.","explanation":"WebDataset is an open-source deep-learning data format and Python library of the same name, developed by Thomas Breuel; the official PyTorch blog has featured it, and it pairs naturally with NVIDIA's AIStore storage service. It doesn't invent a new file format — it uses standard tar archives directly: the multiple files belonging to one sample (say, 000042.jpg and 000042.json) share the same base name minus extension and sit next to each other, and the whole dataset is split into consecutively numbered shards. During training, shards are read in order with no need to decompress the whole archive, and can be streamed directly from local disk, a web server, or cloud storage, which suits huge numbers of small files and multi-GPU training while avoiding the filesystem overhead of random access. It implements PyTorch's IterableDataset interface, so it plugs straight into a DataLoader. In embodied AI, it's commonly used to package large-scale video and image data — for instance, BeingBeyond's UniHand_Preview, released on Hugging Face, uses this format.","example":"One million first-person videos are split into 1,000 tar shards, each containing pairs like number.mp4 plus number.json; during training, multiple machines each read a different subset of shards in parallel.","related":["Apache Parquet","Hierarchical Data Format version 5","Zarr","TFRecord","LeRobotDataset","RLDS (Reinforcement Learning Datasets)"]},{"id":"mcap","category":"data","sec":7,"tier":3,"sources":[{"title":"MCAP 官网","url":"https://mcap.dev/"},{"title":"Foxglove: MCAP as the ROS 2 default bag format","url":"https://foxglove.dev/blog/mcap-as-the-ros2-default-bag-format"},{"title":"ros2/rosbag2（GitHub）","url":"https://github.com/ros2/rosbag2"}],"as_of":"2026-09","related_ids":["ros-bag","robot-operating-system-2","foxglove-studio","hierarchical-data-format-version-5","lerobotdataset","nvidia-isaac-teleop"],"name":"MCAP","alt":"MCAP 格式","abbr":"MCAP","aliases":[".mcap"],"one_liner":"An open-source multimodal logging file format from Foxglove, the default recording format for ROS 2 bags.","explanation":"MCAP is an open-source container file format designed by the robotics software company Foxglove, used to store multiple timestamped data streams — camera images, point clouds, joint states, control commands, and so on — into a single file, organized by publish/subscribe channels. It's agnostic to serialization method, so ROS messages, Protobuf, JSON, and other formats can all be stored, and it writes each message's schema (structure definition) into the file itself, so the file stays readable even years after the code that wrote it has changed. It uses append-only writes, so data already written survives if a program crashes unexpectedly; it also supports chunked LZ4/Zstd compression and indexing for fast reads over a specific time range. Starting with ROS 2 Iron in May 2023, MCAP replaced SQLite as rosbag2's default storage format. Official read/write libraries exist for C++, Python, Go, Rust, Swift, and TypeScript. In embodied-AI data collection, it's commonly used as the raw recording format, later converted into training formats such as LeRobot or RLDS.","example":"NVIDIA's Isaac Teleop framework uses MCAP to record and replay collected data, and can interoperate with the LeRobot dataset format; a recorded .mcap file can also be opened and played back directly in the Foxglove visualization tool.","related":["ROS Bag","Robot Operating System 2","Foxglove Studio","Hierarchical Data Format version 5","LeRobotDataset","NVIDIA Isaac Teleop"]},{"id":"data-cleaning","category":"data","sec":8,"tier":2,"sources":[{"title":"Wikipedia: Data cleansing","url":"https://en.wikipedia.org/wiki/Data_cleansing"},{"title":"Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/abs/2406.09246"},{"title":"AgiBot World Colosseo 技术报告 (arXiv 2503.06669)","url":"https://arxiv.org/html/2503.06669"}],"as_of":"","related_ids":["data-quality-control","data-curation","no-op-action-filtering","valid-data","failure-data","multi-sensor-time-synchronization-timestamp-alignment"],"name":"Data Cleaning","alt":"数据清洗","abbr":"","aliases":["Data Cleansing"],"one_liner":"Finding and fixing or removing bad data before training, such as failed trajectories, empty actions, or misaligned frames.","explanation":"Data cleaning means identifying and fixing or removing corrupted, wrong, or irrelevant records in a dataset — a general step in data processing. Common problems in robot data include: trajectories that were never finished or where the operator made a mistake; idle frames with no motion at the start or end; all-zero actions; timestamps from multiple cameras and joint sensors that aren't aligned; dropped frames; sensor readings that jump unexpectedly; and language instructions that don't match what actually happened. This kind of noise teaches imitation learning to copy the pauses and jitter along with everything else. Cleaning methods include rule-based filtering, statistical anomaly detection, and manual spot-checking, or training a policy on a small cleaned batch first and seeing how it performs on the robot. Not all failure data should be deleted — failures and correction segments with the reason labeled are useful for learning to recover from mistakes.","example":"The OpenVLA paper attributes part of its improvement over RT-2-X to more careful data cleaning, such as removing all-zero actions from the Bridge data; AgiBot World's ablation study also shows that human-verified data raised the task-completion score by 0.18.","related":["Data Quality Control","Data Curation","No-op (Idle) Action Filtering","Valid (Usable) Data","Failure Data","Multi-sensor Time Synchronization / Timestamp Alignment"]},{"id":"no-op-action-filtering","category":"data","sec":8,"tier":3,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"},{"title":"OpenVLA regenerate_libero_dataset.py（GitHub）","url":"https://github.com/openvla/openvla/blob/main/experiments/robot/libero/regenerate_libero_dataset.py"},{"title":"openpi convert_libero_data_to_lerobot.py（GitHub）","url":"https://github.com/Physical-Intelligence/openpi/blob/main/examples/libero/convert_libero_data_to_lerobot.py"}],"as_of":"2024-06","related_ids":["data-cleaning","openvla","libero-benchmark","bridgedata-v2","behavior-cloning","data-quality-control"],"name":"No-op (Idle) Action Filtering","alt":"空操作 / 静止帧过滤","abbr":"","aliases":["No-op Filtering","no_noops"],"one_liner":"Removing timesteps where the robot doesn't move from demonstrations before training, so the model doesn't learn to freeze in place.","explanation":"A no-op (idle action) is a timestep in a demonstration trajectory where the action is essentially zero and the robot's state doesn't change — common at the start of teleoperation while waiting, during a pause by the operator, or because a recording script fills in a zero action for the first step. No-op filtering removes these timesteps before training. The reason this matters is that imitation learning takes everything it's given at face value: an expressive single-step policy that learns from these samples may, at deployment, end up outputting a zero action repeatedly in some state, leaving the robot stuck in place. The OpenVLA (2024) paper documented this issue: in the original BridgeData V2, the first step of every demonstration was an all-zero action, and training on it unmodified caused the policy to frequently output zero actions and freeze; on LIBERO, the authors removed actions whose translation and rotation components were near zero and that didn't change the gripper state, and describe this step as critical for models like OpenVLA. The gripper must be checked separately during filtering, or valid actions where the arm holds still but the gripper opens or closes get mistakenly deleted too.","example":"OpenVLA's open-source script uses this rule: an action counts as a no-op if the norm of every dimension except the gripper is below 1e-4 and the gripper command matches the previous step. The processed LIBERO data is saved with a no_noops suffix, and π0's official codebase, openpi, uses this same processed data directly when fine-tuning on LIBERO.","related":["Data Cleaning","OpenVLA","LIBERO Benchmark","BridgeData V2","Behavior Cloning","Data Quality Control"]},{"id":"data-quality-control","category":"data","sec":8,"tier":3,"sources":[{"title":"首个具身智能数据集质量标准发布（人民邮电报，数字中国建设峰会网站转载）","url":"https://www.szzg.gov.cn/2026/xwzx/szkx/202608/t20260804_5354035.htm"},{"title":"具身智能迈向2.0：数据采集从训练场走向真实世界（科学网转澎湃新闻）","url":"https://news.sciencenet.cn/htmlnews/2026/9/570778.shtm"}],"as_of":"2026-08","related_ids":["data-cleaning","data-curation","valid-data","data-collection-sop","multi-sensor-time-synchronization-timestamp-alignment","failure-data"],"name":"Data Quality Control","alt":"数据质检","abbr":"","aliases":["Data Quality Inspection","Data Quality Review"],"one_liner":"Checking each piece of collected robot data before training to catch and discard unusable or flawed samples.","explanation":"Data quality control is the step, after collection and before annotation and training, where each piece of robot data is checked and judged usable or not. Common checks include whether the camera dropped frames, whether the different sensors' timestamps are aligned, whether the joint and action recordings are complete, whether the trajectory has abnormal jumps, whether the task was actually completed, and whether the language annotation matches the video. The usual approach runs automated rule-based scripts first, followed by manual spot checks. Imitation learning will happily learn from bad demonstrations along with good ones, so quality control has a direct effect on the resulting model. In July 2026, China's Ministry of Industry and Information Technology approved industry standard YD/T 6771-2026, drafted under the lead of the China Academy of Information and Communications Technology (CAICT), which scores embodied-AI datasets across eight dimensions including completeness, consistency, and authenticity; it takes effect on November 1, 2026.","example":"In a teleoperated clothes-folding demonstration, the wrist camera drops frames partway through and the gripper's open/close timing no longer lines up with the video; a quality-control script flags it as unusable, so it's excluded from the training set.","related":["Data Cleaning","Data Curation","Valid (Usable) Data","Data Collection SOP","Multi-sensor Time Synchronization / Timestamp Alignment","Failure Data"]},{"id":"trajectory-episode-replay","category":"data","sec":8,"tier":2,"sources":[{"title":"Imitation Learning on Real-World Robots: Replay an episode (LeRobot docs)","url":"https://huggingface.co/docs/lerobot/il_robots"},{"title":"robomimic: Dataset Contents and Visualization","url":"https://robomimic.github.io/docs/tutorials/dataset_contents.html"}],"as_of":"","related_ids":["trajectory","episode","data-quality-control","open-loop-control","lerobot","robomimic"],"name":"Trajectory / Episode Replay","alt":"轨迹回放","abbr":"","aliases":["Episode Replay","Action Replay"],"one_liner":"Re-sending a recorded action sequence to a robot or simulator to check whether the data can be reproduced.","explanation":"Trajectory replay means reading a previously recorded trajectory from a dataset and sending its actions, frame by frame, at the original frequency, to a real robot or a simulator, to see whether the outcome matches what was recorded. It serves three main purposes: quality checking, to confirm that action labels, timestamps, and coordinate frames were not recorded incorrectly — being able to reproduce the original motion means the data is usable for training; checking consistency between robots of the same model; and, in simulation, re-rendering an observation from a different camera viewpoint using the recorded state. Hugging Face's LeRobot provides a lerobot-replay command, which its documentation says is meant to test whether actions are reproducible and whether they transfer between robots of the same model; robomimic's playback_dataset.py can both re-render from state and replay from actions. Note that real-robot replay is open-loop — a small shift in an object's position can make it fail, which does not by itself mean the data is wrong.","example":"Use lerobot-replay to replay your own recorded demonstration #0 of grasping a block on an SO-101 follower arm; if the motion it produces looks noticeably different from the original recording, that's a sign to check calibration or recording frame rate.","related":["Trajectory","Episode","Data Quality Control","Open-loop Control","LeRobot","robomimic"]},{"id":"data-anonymization","category":"data","sec":8,"tier":3,"sources":[{"title":"中华人民共和国个人信息保护法（中央网信办）","url":"https://www.cac.gov.cn/2021-08/20/c_1631050028355286.htm"},{"title":"EGO4D's approach to privacy and ethics in data collection","url":"https://ego4d-data.org/pdfs/Ego4D-Privacy-and-ethics-consortium-statement.pdf"},{"title":"EgoBlur（Project Aria）","url":"https://www.projectaria.com/tools/egoblur"}],"as_of":"","related_ids":["crowdsourced-data-collection","egocentric-video","ego4d","project-aria-glasses","data-cleaning","data-quality-control"],"name":"Data Anonymization","alt":"数据脱敏","abbr":"","aliases":["De-identification","Data De-identification"],"one_liner":"Blurring or removing personal information like faces, license plates, and voices from data so specific individuals can't be identified.","explanation":"Data anonymization means removing or masking information that could identify a specific person before data is released or used for training. China's Personal Information Protection Law distinguishes two levels: de-identification means the data can no longer identify a specific natural person without additional information; anonymization means identification is impossible and irreversible, and anonymized information is no longer treated as personal information at all. Embodied-AI data increasingly comes from first-person video and crowdsourced collection in homes, shops, and factories, where footage captures bystanders' and operators' faces, house numbers, and screen content, and audio may capture conversations — so anonymization is a required step before releasing a dataset. The common approach is to use a detection model to automatically find and blur faces and license plates, remove or process audio, and then spot-check manually. Ego4D, for example, blurred the faces of bystanders and the license plates of passing cars in its videos, and stripped the audio from many clips entirely.","example":"Meta open-sourced a model called EgoBlur for its Project Aria glasses, which automatically detects and blurs faces and license plates in first-person video; it's released under the Apache 2.0 license, which permits commercial use.","related":["Crowdsourced Data Collection","Egocentric Video","Ego4D","Project Aria Glasses","Data Cleaning","Data Quality Control"]},{"id":"data-annotation","category":"data","sec":8,"tier":2,"sources":[{"title":"Wikipedia: Labeled data","url":"https://en.wikipedia.org/wiki/Labeled_data"},{"title":"Hugging Face: agibot-world/AgiBotWorld2026","url":"https://huggingface.co/datasets/agibot-world/AgiBotWorld2026"},{"title":"DROID 数据集项目页","url":"https://droid-dataset.github.io/"}],"as_of":"","related_ids":["language-annotation","subtask-segmentation","auto-labeling","action-label","data-quality-control","hindsight-relabeling"],"name":"Data Annotation","alt":"数据标注","abbr":"","aliases":["Data Labeling"],"one_liner":"Adding text descriptions, segment boundaries, bounding boxes, and other labels to raw data to tell a model what it is.","explanation":"Data annotation means attaching labels to raw data that a human can understand and a model can use as a supervision signal — the classic example is labeling an image's category or drawing a bounding box. In robot data, actions are already recorded automatically during collection, so what mostly needs adding is: the language instruction for the whole task, where a task splits into subtasks and their start/end times, bounding boxes around target objects, and whether the attempt succeeded or failed and why. These labels determine whether a language-conditioned policy can actually understand instructions, and whether a long-horizon task can be learned as a sequence of steps. Annotation can be done entirely by hand, or generated automatically by a vision-language model and then checked by a human. Labeled data costs far more than raw data, and inconsistency between annotators can directly hurt model performance.","example":"The AgiBot World 2026 dataset provides three layers of annotation: task-frame segmentation with subtask instructions, 2D bounding boxes for object interactions, and step-level instruction segmentation with atomic skills; DROID's December 2024 update added 3 natural-language descriptions to each of about 75,000 successful trajectories.","related":["Language Annotation","Subtask Segmentation","Auto-labeling","Action Label","Data Quality Control","Hindsight Relabeling"]},{"id":"language-annotation","category":"data","sec":8,"tier":2,"sources":[{"title":"Interactive Language / Language-Table 项目主页","url":"https://interactive-language.github.io/"},{"title":"DROID 数据集主页","url":"https://droid-dataset.github.io/"}],"as_of":"2024-12","related_ids":["data-annotation","hindsight-relabeling","subtask-segmentation","auto-labeling","language-conditioned-policy","instruction-augmentation"],"name":"Language Annotation","alt":"语言标注","abbr":"","aliases":["Instruction Annotation"],"one_liner":"Attaching a natural-language description to a robot trajectory, such as 'put the red cup in the sink,' so a model can follow instructions.","explanation":"Language annotation means attaching a text description to robot data, most commonly one task instruction per trajectory, though it can go finer, with one sentence per subtask. It is what makes it possible to train language-conditioned policies and VLA models that act on instructions. Sources fall roughly into three groups: deciding the instruction before collection and having the operator follow it; writing it in afterward by a human watching the video, called hindsight relabeling; and generating or rewriting it automatically with a vision-language model. Annotation quality directly affects how well a model understands instructions — if the same motion is only ever described one way, the model tends to latch onto that fixed phrasing. Large datasets therefore commonly pair one trajectory with several different phrasings, and split long tasks into text-labeled subtask segments to help hierarchical models learn.","example":"Google's Language-Table used hindsight relabeling to produce nearly 600,000 language-labeled tabletop pushing trajectories; DROID's December 2024 update added 3 language descriptions to each of about 75,000 successful trajectories.","related":["Data Annotation","Hindsight Relabeling","Subtask Segmentation","Auto-labeling","Language-conditioned Policy","Instruction Augmentation"]},{"id":"subtask-segmentation","category":"data","sec":8,"tier":3,"sources":[{"title":"Hugging Face: agibot-world/AgiBotWorld2026 数据集卡片","url":"https://huggingface.co/datasets/agibot-world/AgiBotWorld2026"},{"title":"Universal Visual Decomposer (arXiv 2310.08581)","url":"https://arxiv.org/abs/2310.08581"}],"as_of":"","related_ids":["long-horizon-task","language-annotation","auto-labeling","skill-primitive","hierarchical-architecture","agibot-world"],"name":"Subtask Segmentation","alt":"子任务切分","abbr":"","aliases":["Trajectory Segmentation","Skill Segmentation"],"one_liner":"Splitting one long demonstration into steps, marking each segment's start and end frame with a description.","explanation":"Subtask segmentation is a data-annotation step: a complete long-horizon demonstration (such as “clear the table”) is split by meaning into subtask segments (“pick up the plate,” “put it in the sink”), with the start and end frame of each recorded, usually along with a language description and a skill type. It lets long-horizon task data be used step by step: when training a hierarchical policy, the high-level model learns to predict which subtask comes next, while the low-level model learns to execute a single subtask; segmentation can also be used to compute per-skill success rates, build a progress-based reward, or discard failed segments. Segmentation methods include manual annotation, rule-based automatic segmentation (for example, at points where the gripper opens or closes, or velocity drops near zero), and automatically finding boundaries using visual representations or a VLM — UVD, for instance, discovers subgoals by detecting sudden shifts in a pretrained visual representation.","example":"The AgiBot World 2026 dataset provides step-level instruction_segments annotations for every trajectory, with each segment recording a skill type (such as Pick), one instruction sentence, and its start/end frames — for example, “left arm picks up the red-capped drink from the shopping cart” corresponds to frames 284–493.","related":["Long-horizon Task","Language Annotation","Auto-labeling","Skill Primitive","Hierarchical Architecture","AgiBot World"]},{"id":"auto-labeling","category":"data","sec":8,"tier":3,"sources":[{"title":"Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (DIAL, arXiv 2211.11736)","url":"https://arxiv.org/abs/2211.11736"},{"title":"Robotic Control via Embodied Chain-of-Thought Reasoning (ECoT, arXiv 2407.08693)","url":"https://arxiv.org/abs/2407.08693"}],"as_of":"","related_ids":["data-annotation","language-annotation","hindsight-relabeling","instruction-augmentation","subtask-segmentation","data-quality-control"],"name":"Auto-labeling","alt":"自动标注","abbr":"","aliases":["Automated Annotation"],"one_liner":"Using pretrained models or scripts, instead of humans, to automatically add language instructions, bounding boxes, and other labels to robot data.","explanation":"Auto-labeling means using a program or a pretrained model — a vision-language model, an object detector, a large language model — to automatically add labels to collected data instead of a human doing it. Robot data often lacks several kinds of information: what a trajectory is doing (a language instruction), where it splits into subtasks, and where objects are in the frame (bounding boxes, masks). Manual annotation one item at a time is expensive and slow, and can't keep up once data volume grows. Two typical approaches: using an image-text model like CLIP to match un-annotated demonstrations to instructions, as Google's DIAL does; or, like Berkeley's ECoT, using Grounding DINO to box objects and a large model to write out task decomposition and reasoning steps as extra supervision for training a VLA with reasoning ability. Auto-labeling is less reliable than human annotation, so it is usually spot-checked or filtered with rules, and commonly paired with data quality inspection, hindsight relabeling, and instruction augmentation.","example":"Google's 2022 DIAL used CLIP to automatically add language labels to about 80,000 demonstrations (96.5% of which originally had no crowd-sourced label), and the resulting policy could execute 60 new instructions absent from the original data.","related":["Data Annotation","Language Annotation","Hindsight Relabeling","Instruction Augmentation","Subtask Segmentation","Data Quality Control"]},{"id":"play-data","category":"data","sec":8,"tier":3,"sources":[{"title":"Learning Latent Plans from Play (arXiv:1903.01973)","url":"https://arxiv.org/abs/1903.01973"}],"as_of":"","related_ids":["goal-conditioned-policy","hindsight-relabeling","teleoperation","demonstration-data","calvin-benchmark","mimicplay"],"name":"Play Data","alt":"玩耍数据","abbr":"","aliases":["Free Play Data"],"one_liner":"Robot data recorded while an operator freely manipulates a scene with no specific task in mind.","explanation":"Play data is a collection approach introduced by Corey Lynch and colleagues in the 2019 paper Learning Latent Plans from Play (CoRL 2019): an operator teleoperates a robot around a scene full of objects, opening drawers, pushing sliders, and grabbing blocks out of curiosity, with no task specified in advance. Its advantage is cost: there's no need to segment by task, label anything, or reset the scene between attempts. The paper reports that, for the same amount of collection time, this approach covers roughly 4 times the range of interactions that task-by-task demonstrations do. The tradeoff is that there are no task labels, so this data is usually paired with hindsight relabeling (treating the endpoint of a segment as its goal) to train a goal-conditioned policy, or given a language description after the fact. The CALVIN benchmark and MimicPlay both build on this idea.","example":"In the Play-LMP paper, an operator teleoperates a simulated robot through VR, freely manipulating a tabletop scene with drawers, sliding doors, and blocks; the resulting continuous data, with no segmentation or task labels, is used to train a policy that can complete various tasks given a goal image.","related":["Goal-conditioned Policy","Hindsight Relabeling","Teleoperation","Demonstration Data","CALVIN Benchmark","MimicPlay"]},{"id":"hindsight-relabeling","category":"data","sec":8,"tier":3,"sources":[{"title":"Hindsight Experience Replay (arXiv)","url":"https://arxiv.org/abs/1707.01495"},{"title":"Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (DIAL, arXiv)","url":"https://arxiv.org/abs/2211.11736"}],"as_of":"","related_ids":["hindsight-experience-replay","goal-conditioned-reinforcement-learning","sparse-reward","language-annotation","auto-labeling","instruction-augmentation"],"name":"Hindsight Relabeling","alt":"事后重标注","abbr":"","aliases":["Hindsight Goal Relabeling"],"one_liner":"After data is collected, relabeling a trajectory's goal or instruction to match whatever it actually achieved.","explanation":"The idea behind hindsight relabeling is: a trajectory that failed to reach its original goal still ended up somewhere, so that actual outcome can be relabeled as the goal after the fact — turning a failed sample into a success at “achieving a different goal.” The technique became popular through OpenAI's 2017 Hindsight Experience Replay (HER), which addresses how little a policy can learn under sparse rewards (a reward given only on success), and was validated on robot-arm tasks like pushing, sliding, and pick-and-place. It was later extended to imitation learning and language-conditioned policies: a large amount of manipulation data is collected with no task label at all, and afterward it's labeled with an image goal or a language instruction based on what the footage shows. Google's DIAL, for instance, uses a vision-language model like CLIP to automatically add language instructions to 80,000 demonstrations, 96.5% of which had no human annotation to begin with. The precondition for this technique is that the policy is conditioned on a goal or instruction — otherwise there's nothing to relabel.","example":"A robot arm meant to push a block to point A instead pushes it to point B; after relabeling, this trajectory is stored as a successful demonstration with “goal = point B,” and used to train a goal-conditioned policy.","related":["Hindsight Experience Replay","Goal-Conditioned Reinforcement Learning","Sparse Reward","Language Annotation","Auto-labeling","Instruction Augmentation"]},{"id":"instruction-augmentation","category":"data","sec":8,"tier":3,"sources":[{"title":"Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (arXiv 2211.11736)","url":"https://arxiv.org/abs/2211.11736"},{"title":"DIAL 项目主页","url":"https://instructionaugmentation.github.io/"}],"as_of":"","related_ids":["language-annotation","hindsight-relabeling","auto-labeling","data-augmentation","language-conditioned-policy","clip"],"name":"Instruction Augmentation","alt":"指令增强","abbr":"","aliases":["Language Augmentation"],"one_liner":"Automatically writing or rewriting language instructions for robot trajectories, so the same motion gets many phrasings.","explanation":"Instruction augmentation is a family of data-processing methods for language-conditioned policies (policies that act based on text instructions): rather than collecting new robot motion, it adds or rewrites the language label on trajectories that already exist. Writing an instruction by hand for every demonstration is expensive, and the wording tends to be narrow, so a model can fail to understand the same task described a different way. A representative example is DIAL from the Google robotics team, proposed in 2022 and published at RSS 2023: it first fine-tunes CLIP (an image-text matching model) on a small amount of human-labeled data, then has a large language model propose candidate instructions, uses CLIP to score and pick the ones that match what the trajectory's footage actually shows, and relabels 80,000 demonstrations this way (96.5% of which had no human language label at all), before running behavioral cloning. A simpler version just has a large model rewrite the original instruction into several synonymous phrasings. It follows the same underlying idea as hindsight relabeling and automatic annotation.","example":"The policy trained with DIAL was tested on new instructions that never appeared in the original 60-instruction dataset, spanning three categories: spatial-relationship descriptions, differently worded descriptions, and entirely new semantic skills.","related":["Language Annotation","Hindsight Relabeling","Auto-labeling","Data Augmentation","Language-conditioned Policy","CLIP"]},{"id":"data-diversity","category":"data","sec":8,"tier":2,"sources":[{"title":"Data Scaling Laws in Imitation Learning for Robotic Manipulation (arXiv 2410.18647)","url":"https://arxiv.org/abs/2410.18647"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054)","url":"https://arxiv.org/html/2504.16054"}],"as_of":"2025-04","related_ids":["generalization","scene-generalization","object-generalization","data-scaling-laws-in-imitation-learning","in-the-wild-data","data-mixture"],"name":"Data Diversity","alt":"数据多样性","abbr":"","aliases":["Data Coverage"],"one_liner":"How broadly training data varies across scenes, objects, tasks, and viewpoints.","explanation":"Data diversity describes how much a dataset varies along dimensions such as environment, object, lighting, camera viewpoint, task, operator, and robot embodiment — a different question from “how many demonstrations there are.” A robot policy easily memorizes the specific table and objects it trained on and fails in a new room, so insufficient diversity is one of the main reasons generalization is poor. A 2024 data-scaling-law study by Tsinghua's Yang Gao and colleagues found that a policy's performance in new environments and on new objects scales mainly with the number of distinct environments and objects, and that collecting more demonstrations in the same environment quickly stops helping. Physical Intelligence's π0.5 trained on mobile-manipulation data from roughly 100 real homes and was still able to do tidying-type tasks in homes it had never seen. Deliberately varying the scene, the objects, and the person collecting, during data collection, is what raises diversity.","example":"The π0.5 paper's experiments show that the more homes the training data came from, the better the policy performed in test homes; training on data from 104 locations came close to matching a control model trained directly on the test home's own data.","related":["Generalization","Scene Generalization","Object Generalization","Data Scaling Laws in Imitation Learning (Robotic Manipulation)","In-the-wild Data","Data Mixture"]},{"id":"data-scaling-laws-in-imitation-learning","category":"data","sec":8,"tier":2,"sources":[{"title":"Data Scaling Laws in Imitation Learning for Robotic Manipulation (arXiv 2410.18647)","url":"https://arxiv.org/abs/2410.18647"},{"title":"Data Scaling Laws in Imitation Learning 项目主页","url":"https://data-scaling-laws.github.io/"}],"as_of":"2025-01","related_ids":["scaling-law","data-diversity","imitation-learning","universal-manipulation-interface","generalization","diffusion-policy"],"name":"Data Scaling Laws in Imitation Learning (Robotic Manipulation)","alt":"数据缩放律（模仿学习）","abbr":"","aliases":[],"one_liner":"Research into how, in robot imitation learning, generalization improves as training data grows.","explanation":"This is a paper from Tsinghua University's Yang Gao group (with partners including the Shanghai Qi Zhi Institute and the Shanghai AI Laboratory), released in October 2024 and selected as an oral presentation at ICLR 2025. Using the handheld capture device UMI (Universal Manipulation Interface), the team collected more than 40,000 demonstrations and ran more than 15,000 real-robot tests to study how a single-task policy's generalization to new environments and new objects scales with data. The finding: generalization follows roughly a power law with the number of distinct training environments and objects; diversity of environments and objects matters far more than simply adding more demonstrations, and returns drop off sharply once a given environment or object has enough demonstrations. Based on this, they recommend an efficient collection strategy: change environments often, pair each environment with a different object, and collect about 50 demonstrations per environment.","example":"Following this recipe, data collected by 4 people in a single afternoon was enough for policies on two new tasks to reach about a 90% success rate in environments and on objects they had never seen.","related":["Scaling Law","Data Diversity","Imitation Learning","Universal Manipulation Interface","Generalization","Diffusion Policy"]},{"id":"in-the-wild-data","category":"data","sec":8,"tier":2,"sources":[{"title":"DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset","url":"https://droid-dataset.github.io/"},{"title":"DexWild 项目主页","url":"https://dexwild.github.io/"},{"title":"Universal Manipulation Interface (UMI) 项目主页","url":"https://umi-gripper.github.io/"}],"as_of":"","related_ids":["droid","universal-manipulation-interface","dexwild","scene-generalization","data-diversity","robot-free-data-collection"],"name":"In-the-wild Data","alt":"野外数据","abbr":"","aliases":["Real-world Scene Data"],"one_liner":"Data collected in homes, offices, outdoors, and other uncontrolled real-world settings, rather than a fixed lab bench.","explanation":"In-the-wild data borrows a term from computer vision, meaning data collected in real settings outside the lab, where lighting, background, and object placement are all uncontrolled and every scene differs from the last. Most robot data used to come from a handful of fixed lab tabletops, so models easily failed in a new room; researchers have since moved collection into real environments to improve scene generalization. The difficulty is that a real robot isn't easy to move around, so portable alternatives have emerged: DROID mounts a Franka arm on a movable height-adjustable cart and collected 76,000 trajectories across 564 scenes; UMI uses a handheld gripper, and DexWild uses a bare human hand plus a palm-mounted camera — neither needs a robot physically on-site.","example":"Carnegie Mellon's DexWild collected 9,290 human demonstrations across 93 different environments; co-trained with robot data, the resulting policy's success rate in new environments was about 4 times that of a policy trained on robot data alone.","related":["DROID (Distributed Robot Interaction Dataset)","Universal Manipulation Interface","DexWild","Scene Generalization","Data Diversity","Robot-free (Embodiment-free) Data Collection"]},{"id":"suboptimal-demonstrations","category":"data","sec":8,"tier":3,"sources":[{"title":"robomimic v0.1 Datasets (PH / MH / MG)","url":"https://robomimic.github.io/docs/datasets/robomimic_v0.1.html"},{"title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (arXiv 2108.03298)","url":"https://arxiv.org/abs/2108.03298"}],"as_of":"","related_ids":["behavior-cloning","data-curation","data-quality-control","advantage-weighted-regression","offline-reinforcement-learning","recap"],"name":"Suboptimal (Noisy) Demonstrations","alt":"次优演示","abbr":"","aliases":["Imperfect Demonstrations","Mixed-Quality Demonstrations"],"one_liner":"Human demonstration data that includes wasted motion, hesitation, mistakes, or inconsistent quality across operators.","explanation":"Suboptimal demonstrations are demonstrations that fall short of ideal: slow motion, pauses, back-and-forth trial and error, mistakes followed by recovery, or inconsistency because different operators have different skill levels and habits. Almost all real-world collected data has some of these issues. Behavioral cloning (directly imitating the demonstrated actions) learns from the good and bad actions alike, so the more mixed the data, the more hesitant and unstable the resulting policy tends to be. Common countermeasures include: filtering and quality-checking the data after collection; weighting samples by reward or advantage (how much better an action is than average), as in advantage-weighted regression; using offline reinforcement learning to extract better behavior from mixed-quality data; or feeding quality information in as a conditioning signal, so that only the “good” behavior is requested at inference time, as in RECAP's advantage conditioning.","example":"robomimic's Multi-Human dataset was collected by 6 operators of varying skill (2 each rated “worse,” “okay,” and “better”), each recording 50 successful trajectories for 300 total, specifically to study how mixed-quality data affects imitation learning.","related":["Behavior Cloning","Data Curation","Data Quality Control","Advantage-Weighted Regression","Offline Reinforcement Learning","RECAP"]},{"id":"data-curation","category":"data","sec":8,"tier":2,"sources":[{"title":"Robot Data Curation with Mutual Information Estimators (arXiv 2502.08623)","url":"https://arxiv.org/abs/2502.08623"},{"title":"CUPID: Curating Data your Robot Loves with Influence Functions (arXiv 2506.19121)","url":"https://arxiv.org/abs/2506.19121"}],"as_of":"2025-09","related_ids":["data-cleaning","data-quality-control","data-mixture","valid-data","suboptimal-demonstrations","imitation-learning"],"name":"Data Curation","alt":"数据筛选","abbr":"","aliases":["Data Filtering"],"one_liner":"Picking out the subset of a large robot dataset that actually helps training, and discarding the harmful part.","explanation":"Data curation means selecting and weighting a dataset before training: dropping demonstrations that are error-prone, hesitant, or inconsistent in technique, and keeping the high-quality, broadly covering portion. It differs from data cleaning (fixing formats, removing bad frames) in that it asks which data will actually make the policy better. Robot demonstrations are often collected by many people and vary widely in quality, and imitation learning will copy bad habits right along with good ones, so more data is not automatically better. Notable methods include DemInf, proposed by researchers at Stanford and Google DeepMind in 2025, which scores each demonstration using the mutual information between state and action, and CUPID, from CoRL 2025, which uses influence functions to estimate each demonstration's contribution to policy success rate — the paper reports that using less than 33% of the curated data trained a then-state-of-the-art diffusion policy on the RoboMimic benchmark.","example":"In the DemInf paper, on real ALOHA and Franka setups, demonstrations collected by multiple people were scored one by one, the lowest-scoring batch was removed, and the policy trained on the remainder outperformed one trained on the full dataset.","related":["Data Cleaning","Data Quality Control","Data Mixture","Valid (Usable) Data","Suboptimal (Noisy) Demonstrations","Imitation Learning"]},{"id":"retrieval-based-data-selection","category":"data","sec":8,"tier":3,"sources":[{"title":"Behavior Retrieval: Few-Shot Imitation Learning by Querying Unlabeled Datasets (arXiv)","url":"https://arxiv.org/abs/2304.08742"},{"title":"STRAP: Robot Sub-Trajectory Retrieval for Augmented Policy Learning (arXiv)","url":"https://arxiv.org/abs/2412.15182"},{"title":"Data Retrieval with Importance Weights for Few-Shot Imitation Learning (arXiv)","url":"https://arxiv.org/abs/2509.01657"}],"as_of":"2025-09","related_ids":["data-curation","data-mixture","few-shot","positive-negative-transfer","open-x-embodiment","behavior-cloning"],"name":"Retrieval-Based Data Selection","alt":"数据检索","abbr":"","aliases":["Data Retrieval"],"one_liner":"Using a handful of demonstrations for a target task to search a large dataset for similar clips, then training on the combined set.","explanation":"Retrieval-based data selection is a family of methods for choosing training data: collect a few demonstrations for the target task, then search a large existing robot dataset for trajectories or segments that are visually or behaviorally similar, and train on the retrieved data together with the target demonstrations. It addresses a tradeoff: training on all available data mixed together risks negative transfer, where unrelated data interferes with learning, while training on just a handful of demonstrations isn't enough data on its own. Representative work includes Stanford's Behavior Retrieval (2023); the University of Washington's STRAP (ICLR 2025), which retrieves at the sub-trajectory level using vision-foundation-model features and dynamic time warping for matching; and Stanford's IWR (CoRL 2025), which uses importance weighting to correct for retrieval bias. It addresses the same underlying problem as data filtering and data mixing ratios.","example":"To teach a robot to pick up a cup, put it in a drawer, and close the drawer, only a few demonstrations were collected; STRAP retrieves sub-segments involving “picking up a cup” and “opening/closing a drawer” from a large offline dataset and trains the policy on those together with the target demonstrations.","related":["Data Curation","Data Mixture","Few-shot","Positive / Negative Transfer","Open X-Embodiment","Behavior Cloning"]},{"id":"data-mixture","category":"data","sec":8,"tier":2,"sources":[{"title":"Re-Mix: Optimizing Data Mixtures for Large Scale Imitation Learning (arXiv 2408.14037)","url":"https://arxiv.org/abs/2408.14037"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"2024-10","related_ids":["oxe-magic-soup","co-training","heterogeneous-data","cross-embodiment-data","data-curation","open-x-embodiment"],"name":"Data Mixture","alt":"数据配比","abbr":"","aliases":["Data Recipe","Data Mixture Weights"],"one_liner":"The proportion each data source is given when training on a mix of multiple sources at once.","explanation":"Data mixture refers to how often each source gets sampled when training on a combination of data from different origins — different robots, different tasks, simulation versus real data, web text and images, human video. Source sizes can differ enormously; mixing purely in proportion to raw counts lets big datasets drown out small ones, while poorly tuned weights instead make the model lopsided. Common approaches include hand-tuned weights, such as the recipe Octo and OpenVLA use on Open X-Embodiment (informally called the Magic Soup); π0's rule of weighting each task-robot combination by n^0.43 (n being that combination's sample count), which down-weights combinations with too much data; and Stanford's Re-Mix, which learns weights automatically with distributionally robust optimization, reporting a 38% average improvement over uniform weighting. Mixtures are also commonly changed across training stages, leaning toward diversity during pretraining and toward quality during post-training.","example":"In π0's pretraining data, 9.1% comes from open datasets such as OXE, Bridge v2, and DROID, with the rest from Physical Intelligence's own collection, and each task-robot combination is then weighted by the n^0.43 rule.","related":["OXE Magic Soup","Co-training","Heterogeneous Data","Cross-Embodiment Data","Data Curation","Open X-Embodiment"]},{"id":"oxe-magic-soup","category":"data","sec":8,"tier":3,"sources":[{"title":"octo-models/octo: oxe_dataset_mixes.py (GitHub)","url":"https://github.com/octo-models/octo/blob/main/octo/data/oxe/oxe_dataset_mixes.py"},{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/html/2405.12213"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"2024-06","related_ids":["open-x-embodiment","data-mixture","octo","openvla","droid","rlds"],"name":"OXE Magic Soup","alt":"Magic Soup（OXE 数据配方）","abbr":"","aliases":["oxe_magic_soup","Magic Soup++","oxe_magic_soup_plus"],"one_liner":"The recipe Octo and OpenVLA used to select and weight subsets of Open X-Embodiment for pretraining.","explanation":"Magic Soup is the name for a particular robot-data mixing recipe, taken from the Octo codebase's oxe_magic_soup configuration. The individual datasets inside Open X-Embodiment (OXE, a cross-embodiment robot dataset pooled from many institutions) vary hugely in size, quality, robot type, and camera setup, so mixing them together naively doesn't work well. Octo's (2024) training mixture drew on 25 of these datasets, about 800,000 trajectories in total, with weights roughly proportional to sample count, then doubled the weight of datasets with richer scenes and tasks and down-weighted ones with many repetitive clips. OpenVLA reused this same weighting, and its codebase also has an expanded version, Magic Soup++, which adds DROID for about 970,000 demonstrations total. It isn't new data — it's an empirically tuned recipe for which datasets to include and how much weight to give each; even the Octo paper admits it still lacks a systematic analysis of the choice.","example":"In OpenVLA's training mixture, Fractal (RT-1 data) and Kuka each make up about 12.7%, while UCSD Kitchen is under 0.1%; DROID was included at 10% weight but removed for the final third of training because the model was learning slowly from it.","related":["Open X-Embodiment","Data Mixture","Octo","OpenVLA","DROID (Distributed Robot Interaction Dataset)","RLDS (Reinforcement Learning Datasets)"]},{"id":"data-leakage-test-set-contamination","category":"data","sec":8,"tier":3,"sources":[{"title":"What is Data Leakage in Machine Learning?（IBM）","url":"https://www.ibm.com/think/topics/data-leakage-machine-learning"},{"title":"LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (arXiv 2510.03827)","url":"https://arxiv.org/abs/2510.03827"}],"as_of":"2026-05","related_ids":["training-validation-test-set","overfitting","out-of-distribution","libero-pro","benchmark-saturation","leaderboard-chasing"],"name":"Data Leakage / Test-Set Contamination","alt":"数据泄漏 / 测试集污染","abbr":"","aliases":["Train-Test Contamination","Test-Set Leakage"],"one_liner":"When information that should only appear at test time leaks into training, inflating evaluation scores beyond real-world performance.","explanation":"Data leakage means a model was exposed during training to information it shouldn't have access to at test or deployment time, so it looks good on offline evaluation but performs much worse in real use. IBM divides this into two types: target leakage, where a feature secretly contains the “answer” that wouldn't be available at prediction time, and train-test contamination, where test data leaks into training, or where preprocessing steps like normalization are computed on the full dataset before splitting. In robot learning, common cases include splitting train and validation sets by individual frame rather than by whole trajectory, training on simulation benchmarks using scenes and initial states nearly identical to the test set, and letting evaluation questions leak into a large model's pretraining corpus. LIBERO-PRO found that VLA (vision-language-action) models scoring over 90% success on the original LIBERO benchmark could drop to 0% once objects, initial positions, or instructions were changed — evidence that they had mostly memorized the training set. The fix is to split data by trajectory or scene, and evaluate under out-of-distribution settings.","example":"If all the frames from a 30-second demonstration are shuffled randomly before being split into train and validation sets, the action error measured on the validation set will look very low, because the model has essentially seen the neighboring frames of nearly every validation frame; splitting by whole trajectory instead gives an error that reflects real generalization ability.","related":["Training / Validation / Test Set","Overfitting","Out-of-Distribution","LIBERO-PRO","Benchmark Saturation","Leaderboard Chasing"]},{"id":"embodied-ai-training-ground","category":"data","sec":9,"tier":1,"sources":[{"title":"AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv 2503.06669)","url":"https://arxiv.org/abs/2503.06669"},{"title":"AgiBot World GitHub","url":"https://github.com/OpenDriveLab/AgiBot-World"},{"title":"中新网：全国首个异构人形机器人训练场在上海启用（2025-01-21）","url":"https://www.chinanews.com.cn/cj/2025/01-21/10357409.shtml"}],"as_of":"2025-03","related_ids":["real-robot-data","data-collector","data-collection-sop","agibot-world","teleoperation","data-quality-control"],"name":"Embodied AI Training Ground (Robot Data Collection Center)","alt":"具身智能训练场","abbr":"","aliases":["Data Factory","Data Collection Factory","Robot Training Ground"],"one_liner":"A facility built to replicate real-world settings, where large numbers of robots are teleoperated to mass-produce training data.","explanation":"An embodied AI training ground — also called a data factory or data collection center — is a term common in China's robotics industry for a site that recreates real environments such as homes, supermarkets, factories, and restaurants at 1:1 scale. Data collectors teleoperate many robots there, running through tasks repeatedly to mass-produce real-robot data; the same site also serves as a place to test and evaluate models. It exists because robot data can't be scraped from the internet — it has to be collected one trajectory at a time — so standardizing equipment, scenes, collection procedures (SOPs), and quality checks is what makes volume and quality possible at all. AgiBot built roughly 4,000 square meters of collection space for AgiBot World, using 100 real robots to gather more than a million trajectories. Sites like this have been opening across China since 2025: that January, the state-backed National and Local Co-built Humanoid Robot Innovation Center opened a heterogeneous humanoid training ground in Shanghai — hosting robots from multiple vendors and models — with room in its first phase for over 100 humanoid robots to train at once.","example":"AgiBot World's collection site covers five categories — home, retail, industrial, dining, and office — spanning more than 100 real-world scenes in total, of which the industrial and retail scenes are replicated 1:1; every piece of data carries multi-view camera footage, depth, camera calibration, and sub-step language annotations.","related":["Real-Robot Data","Data Collector (Teleoperator)","Data Collection SOP","AgiBot World","Teleoperation","Data Quality Control"]},{"id":"data-collector","category":"data","sec":9,"tier":2,"sources":[{"title":"北京人形：具身智能机器人应用技术员进入国家新职业序列","url":"https://www.x-humanoid.com/news-view-330.html"},{"title":"钛媒体：机器人还没学会做家务，卖数据的已经先赚到了钱","url":"https://www.tmtpost.com/8062934.html"},{"title":"中证网：实探北京人形机器人数据基地，月产数据1.5万小时","url":"https://jnzstatic.cs.com.cn/zzb/htmlInfo/117207.html"}],"as_of":"2026-09","related_ids":["teleoperation","data-collection-sop","embodied-ai-training-ground","demonstration-data","valid-data","embodied-ai-robot-application-technician"],"name":"Data Collector (Teleoperator)","alt":"数采员","abbr":"","aliases":["Robot Trainer","AI Trainer"],"one_liner":"A front-line worker who records robot training data through teleoperation or wearable demonstration devices.","explanation":"Data collector is a front-line job at embodied-AI companies and data-collection bases, producing training data. Workers follow a collection SOP (standard operating procedure): they wear a VR headset and operate a leader arm or exoskeleton to teleoperate a real robot, or wear a mocap suit, a head-mounted camera, and data gloves to perform the demonstration with their own hands directly, while the system logs video, joint angles, force, and other signals in sync. Imitation learning depends on large numbers of human demonstrations, so how consistent and correct a data collector's motions are directly determines whether the data is usable. In September 2026, China's Ministry of Human Resources and Social Security and other agencies announced a new occupation, “Embodied AI Robot Application Technician,” which lists data collector and trainer as sub-roles under it. Reports describe the work as repetitive and tedious, with equipment setup, resetting scenes, and re-recording taking up a large share of the time.","example":"According to a China Securities Journal report from March 2026, the Beijing Humanoid Robot Innovation Center's data base has more than 120 robots and over 30 scenes — home, retail, production lines — where data collectors do real-robot teleoperation and motion-capture collection, producing about 15,000 hours of data a month.","related":["Teleoperation","Data Collection SOP","Embodied AI Training Ground (Robot Data Collection Center)","Demonstration Data","Valid (Usable) Data","Embodied AI Robot Application Technician"]},{"id":"data-collection-sop","category":"data","sec":9,"tier":3,"sources":[{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (arXiv 2403.12945)","url":"https://arxiv.org/html/2403.12945v1"},{"title":"国地共建具身智能机器人创新中心打造具身智能标准体系","url":"https://www.x-humanoid.com/news-view-50.html"}],"as_of":"2024-11","related_ids":["data-collector","valid-data","data-quality-control","demonstration-data","droid","crowdsourced-data-collection"],"name":"Data Collection SOP","alt":"采集 SOP","abbr":"SOP","aliases":["Data Collection Standard Operating Procedure","Data Collection Protocol"],"one_liner":"The written procedure that specifies exactly how each demonstration should be collected and what counts as an acceptable one.","explanation":"A data collection SOP is the operating manual a data team writes for its collection workers. It typically spells out: equipment checks and camera calibration; how the scene should be arranged and how often it should be changed; where task instructions come from; how the start and end of one demonstration are defined; the pace and technique for the motion; how to handle and flag failures; file naming and upload; and quality-control criteria. Robot learning is sensitive to inconsistency in the data — if different people perform a task differently, start and end points aren't standardized, or long pauses creep in, the model struggles to learn — so an SOP is the foundation for keeping the proportion of valid data high. A good SOP also builds in diversity: DROID's shared collection protocol has operators change scenes roughly every 20 minutes and randomly prompts them to adjust lighting, move the camera, or add and remove objects. China is also pushing for standardization: the Beijing Humanoid Robot Innovation Center is leading a proposed industry standard for the Ministry of Industry and Information Technology, titled “Artificial Intelligence — Embodied Intelligence — Data Collection Specification.”","example":"DROID's 50 collection workers are spread across multiple universities and all use the same collection protocol and GUI: the interface randomly draws an instruction from a list of feasible tasks the worker has entered, and after each trajectory the worker marks it as a success or failure.","related":["Data Collector (Teleoperator)","Valid (Usable) Data","Data Quality Control","Demonstration Data","DROID (Distributed Robot Interaction Dataset)","Crowdsourced Data Collection"]},{"id":"valid-data","category":"data","sec":9,"tier":3,"sources":[{"title":"规模化「上课」为具身智能加注「数据燃料」（证券时报网，2026-04）","url":"https://www.stcn.com/article/detail/3731496.html"},{"title":"具身智能迈向 2.0：数据采集从训练场走向真实世界（科学网转澎湃新闻，2026-09）","url":"https://news.sciencenet.cn/htmlnews/2026/9/570778.shtm"}],"as_of":"2026-09","related_ids":["data-quality-control","data-collector","data-collection-sop","data-cleaning","data-curation","embodied-ai-training-ground"],"name":"Valid (Usable) Data","alt":"有效数据","abbr":"","aliases":["Valid Data Rate","Data Pass Rate"],"one_liner":"The portion of raw collected data that passes quality control and can actually be used to train a model.","explanation":"This is a common term in the embodied-AI data-collection industry, without a single standardized definition. Some fraction of raw recorded data is always unusable: the task failed, the arm touched a prop it shouldn't have, the multiple sensor streams weren't time-synchronized, calibration was off, or there was occlusion or dropped frames. What remains after quality control filters these out is called valid (usable) data, and the fraction of raw data it represents is called the valid data rate (also called the pass rate). The head of the Beijing Humanoid Robot Innovation Center's data facility told media in April 2026 that the facility's pass rate had once been only about 50%, but stabilized above 95% after operator training and improvements to quality-control standards. The industry is also increasingly wary of judging data purely by hour count: frame rate and calibration have to meet a bar, scenes/objects/tasks need enough diversity, and ultimately the data has to pass client acceptance and prove itself in actual model performance.","example":"A data collection worker records 100 demonstrations of “put the cup in the cabinet”; 8 have the cup fall, and 5 have the wrist camera drop frames. After quality control removes these, 87 remain as valid data — a valid data rate of 87%.","related":["Data Quality Control","Data Collector (Teleoperator)","Data Collection SOP","Data Cleaning","Data Curation","Embodied AI Training Ground (Robot Data Collection Center)"]},{"id":"crowdsourced-data-collection","category":"data","sec":9,"tier":3,"sources":[{"title":"RoboTurk - Crowdsourcing Robotics（斯坦福）","url":"https://roboturk.stanford.edu"},{"title":"机器人开始向更多人类买数据（新京报）","url":"https://www.bjnews.com.cn/detail/1790255981129859.html"},{"title":"日薪120元全民数采：谁在训练下一个机器人保姆？（36氪）","url":"https://eu.36kr.com/zh/p/3810340908817928"}],"as_of":"2026-09","related_ids":["data-collector","robot-free-data-collection","egocentric-video","valid-data","data-anonymization","data-quality-control"],"name":"Crowdsourced Data Collection","alt":"众包采集","abbr":"","aliases":["Crowdsourcing"],"one_liner":"Distributing robot data-collection tasks to large numbers of non-professional workers, who are paid for each validated, usable clip.","explanation":"Crowdsourced data collection doesn't rely on a small team of trained operators. Instead, a website or app distributes tasks to large numbers of ordinary people, who either teleoperate a robot remotely or record themselves performing tasks at home or at work; the platform reviews submissions and pays based on validated, usable duration. An early example is Stanford's RoboTurk (2018), which let users control a robot arm through a browser using their phone as a 6-DoF controller. Between 2025 and 2026, as first-person video and embodiment-free capture devices became more common, crowdsourcing shifted toward real-world human data: Figure reportedly opened a paid project called “Index” in August 2026 to collect users' task videos, and the Chinese startup Mifeng Technology (觅蜂科技) launched a crowdsourced app called “Mifeng Pai” (觅蜂派) in September, where users take on data-collection jobs while wearing its MEgo capture device. The advantages are low cost and broad scene coverage; the challenges are inconsistent quality, privacy protection, and the cost of cleaning and labeling the results.","example":"Stanford's 2019 real-robot RoboTurk dataset was collected by 54 non-expert users teleoperating robots remotely over the course of a week, yielding 2,144 demonstrations totaling 111 hours.","related":["Data Collector (Teleoperator)","Robot-free (Embodiment-free) Data Collection","Egocentric Video","Valid (Usable) Data","Data Anonymization","Data Quality Control"]},{"id":"roboturk","category":"data","sec":9,"tier":3,"sources":[{"title":"RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation (arXiv:1811.02790)","url":"https://arxiv.org/abs/1811.02790"}],"as_of":"2018-11","related_ids":["crowdsourced-data-collection","teleoperation","demonstration-data","robomimic","imitation-learning","stanford-artificial-intelligence-laboratory"],"name":"RoboTurk","alt":"RoboTurk","abbr":"","aliases":["RoboTurk Crowdsourcing Platform"],"one_liner":"Stanford's 2018 platform for crowdsourced robot demonstrations, teleoperated remotely from a smartphone.","explanation":"RoboTurk is a crowdsourced data-collection platform published at CoRL 2018 by Fei-Fei Li, Silvio Savarese, and colleagues at Stanford, named after the crowdsourcing platform Amazon Mechanical Turk. It lets ordinary people perform 6-DoF teleoperation using a smartphone like an iPhone: moving the phone through the air makes the remote robot's end effector move the same way, with no VR headset or special equipment needed, so collection tasks can be distributed to remote workers anywhere. In an initial trial totaling 22 hours of system usage, it collected 137.5 hours of operation data and over 2,200 successful demonstrations; the paper also found that remote users could complete demonstrations successfully even with low bandwidth and high latency. It's a representative example of crowdsourced collection, and the multi-operator demonstration data in the robomimic benchmark was also collected with it.","example":"A remote worker opens the app at home, watching a live video feed streamed back over the web while moving their phone to control a robot arm through a pick-and-place task; the data uploads automatically for imitation learning.","related":["Crowdsourced Data Collection","Teleoperation","Demonstration Data","robomimic","Imitation Learning","Stanford Artificial Intelligence Laboratory"]},{"id":"autonomous-data-collection","category":"data","sec":9,"tier":3,"sources":[{"title":"AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents (arXiv 2401.12963)","url":"https://arxiv.org/abs/2401.12963"},{"title":"Autonomous Improvement of Instruction Following Skills via Foundation Models (SOAR, arXiv 2407.20635)","url":"https://arxiv.org/abs/2407.20635"},{"title":"Deep Learning for Robots: Learning from Large-Scale Interaction (Google Research Blog)","url":"https://research.google/blog/deep-learning-for-robots-learning-from-large-scale-interaction/"}],"as_of":"","related_ids":["google-arm-farm","autort","data-flywheel","self-improvement","real-world-reinforcement-learning","success-detector"],"name":"Autonomous Data Collection","alt":"自主数据采集","abbr":"","aliases":["Robot Self-Collection"],"one_liner":"Letting a robot propose its own tasks, attempt them, and log the results, with little or no human teleoperation.","explanation":"Autonomous data collection means a robot decides for itself what to do, executes it, and records the data with no human directing it step by step — a human only intervenes occasionally, mostly for safety oversight. Teleoperated collection needs one person per robot, so cost scales linearly with data volume; letting robots collect on their own means more robots simply means more data. Three difficulties stand out: the robot needs to propose meaningful and diverse tasks; it needs to automatically judge success or failure; and autonomously collected data tends to have a lower success rate and more variable quality, requiring methods that can learn from imperfect data. An early example is Google's 2016 Arm Farm, which used automatically judged grasp outcomes as labels; Google DeepMind's AutoRT uses a vision-language model to understand a scene and a large language model to propose tasks, orchestrating more than 20 robots; and Berkeley's SOAR uses a VLM to propose and judge tasks, autonomously collecting more than 30,000 trajectories across 5 tabletop environments to improve a policy.","example":"AutoRT ran more than 20 robots across multiple office buildings, with a large language model proposing tasks, collecting about 77,000 real-robot trajectories total through a mix of teleoperation and autonomous execution.","related":["Google Arm Farm","AutoRT","Data Flywheel","Self-improvement","Real-World Reinforcement Learning","Success Detector"]},{"id":"failure-data","category":"data","sec":9,"tier":2,"sources":[{"title":"π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"},{"title":"DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (arXiv 2403.12945)","url":"https://arxiv.org/html/2403.12945v2"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"2025-11","related_ids":["recovery-and-correction-data","human-intervention-data","success-detector","reward-model","recap","offline-reinforcement-learning"],"name":"Failure Data","alt":"失败数据","abbr":"","aliases":["Failed Trajectories","Failed Demonstrations"],"one_liner":"Trajectories that didn't complete the task, useful for learning to recover from mistakes and for training reward models.","explanation":"Failure data refers to trajectories that failed to complete the task, coming from teleoperation mistakes, a policy's own autonomous failures, or segments recorded before a human stepped in to correct things. Pure imitation learning usually throws these away, since copying them would teach the policy to fail the same way, but they record exactly what conditions lead to errors, which has real uses: training success detectors, reward models, and value functions, serving as negative feedback in reinforcement learning, and, paired with correction data, teaching a model to recover from mistakes. The π0 paper points out that a model trained only on high-quality data never learns to recover, because this kind of data rarely contains mistakes at all. Physical Intelligence's RECAP method trains π*0.6 on a mix of failure trajectories, autonomous-run data, and human corrections, reporting more than double the throughput and roughly half the failure rate on harder tasks.","example":"Besides its 76,000 successful trajectories, the DROID dataset also separately releases about 16,000 trajectories the collectors labeled “unsuccessful,” usable for training a success detector or for offline reinforcement learning.","related":["Recovery and Correction Data","Human Intervention Data","Success Detector","Reward Model","RECAP","Offline Reinforcement Learning"]},{"id":"human-intervention-data","category":"data","sec":9,"tier":2,"sources":[{"title":"HG-DAgger: Interactive Imitation Learning with Human Experts (arXiv 1810.02890)","url":"https://arxiv.org/abs/1810.02890"},{"title":"HIL-SERL 项目主页","url":"https://hil-serl.github.io/"},{"title":"π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"","related_ids":["human-in-the-loop","human-gated-dagger","recovery-and-correction-data","hil-serl","recap","intervention-rate"],"name":"Human Intervention Data","alt":"干预数据","abbr":"","aliases":["Takeover Data"],"one_liner":"Data recorded when a human takes over from an autonomously running policy just as it's about to make a mistake.","explanation":"Human intervention data is recorded when, during a policy's autonomous run on a real robot, a human sees it about to go wrong and takes over by teleoperation, steering the robot back on track; the observations and the human's actions during that takeover are what get recorded. The idea traces back to interactive imitation learning methods like DAgger: pure behavior cloning has only ever seen an expert's “standard route,” so once it drifts off that path it has no idea what to do, and the error compounds. Intervention data fills in exactly “how to recover once things have gone wrong,” concentrated at precisely the states where the policy is weakest. HG-DAgger (2018) has a human take over whenever they judge the situation unsafe; Berkeley's HIL-SERL uses human intervention to guide exploration during real-robot reinforcement learning; and Physical Intelligence's π*0.6 also trains on a mix of expert interventions and autonomous experience.","example":"Early in HIL-SERL training, a human takes over frequently, demonstrating how to complete the task from a wide range of states; once the policy's success rate rises, the amount of intervention is gradually reduced.","related":["Human-in-the-Loop","Human-Gated DAgger","Recovery and Correction Data","HIL-SERL","RECAP","Intervention Rate"]},{"id":"recovery-and-correction-data","category":"data","sec":9,"tier":2,"sources":[{"title":"RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction","url":"https://arxiv.org/abs/2509.07953"},{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger)","url":"https://arxiv.org/abs/1011.0686"}],"as_of":"2025-09","related_ids":["human-intervention-data","failure-recovery","dagger","human-gated-dagger","compounding-error","rac"],"name":"Recovery and Correction Data","alt":"纠偏数据","abbr":"","aliases":["Recovery Data","Correction Data"],"one_liner":"Demonstration data specifically recorded to show how a robot gets back on track after starting to go wrong.","explanation":"Most teleoperated demonstrations are “perfect trajectories” done right the first time, so a policy never sees what a drifted-off state looks like and has no idea how to save itself once it starts going wrong, and the error compounds. Recovery and correction data fills exactly this gap: a policy is left to run on its own, and just as it's about to fail, a human takes over, steers the robot back on track, and finishes the task, with that takeover segment added to the training set. The idea traces back to DAgger (Dataset Aggregation) from 2011, with later variants such as Human-Gated DAgger. A 2025 method called RaC turns it into a dedicated stage after imitation learning: the operator first rewinds the robot to a familiar state, then demonstrates a correction, and on long-horizon tasks such as hanging a shirt or sealing a lunchbox it beats prior methods using roughly one-tenth the collection time.","example":"In RaC's shirt-hanging task, just as the policy is about to hang the shirt crookedly, the operator takes over, first pulls the arm back to a normal prior pose, then demonstrates hanging it straight on the hanger — that takeover becomes one piece of recovery and correction data.","related":["Human Intervention Data","Failure Recovery","DAgger","Human-Gated DAgger","Compounding Error","RaC"]},{"id":"deployment-data-backflow","category":"data","sec":9,"tier":3,"sources":[{"title":"中国信通院《具身智能发展报告（2025年）》","url":"http://www.caict.ac.cn/kxyj/qwfb/bps/202601/P020260130541978285206.pdf"},{"title":"具身智能迈向2.0：数据采集从训练场走向真实世界（科学网转澎湃新闻）","url":"https://news.sciencenet.cn/htmlnews/2026/9/570778.shtm"},{"title":"π*0.6: a VLA That Learns From Experience (arXiv)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"2026-09","related_ids":["data-flywheel","human-intervention-data","recovery-and-correction-data","fleet-learning","recap","pi-star-0-6"],"name":"Deployment Data Backflow","alt":"数据回流","abbr":"","aliases":["Data Backflow","Field Data Feedback Loop"],"one_liner":"A term from China's robotics industry: feeding data generated during real-world robot deployment back into training, then redeploying the updated model.","explanation":"Data backflow is a common term in China's embodied-AI industry for a loop: data generated while a robot works in the real world — autonomously executed trajectories, failure cases, and clips where a human took over and corrected it — is logged, sent back, cleaned, and annotated, then added to training; the updated model is redeployed, and the cycle repeats. It's the key link that gets a “data flywheel” actually spinning: teleoperation data collected in a training facility covers a limited distribution, and errors that show up in the field are the best way to expose a model's real weaknesses. CAICT's Embodied Intelligence Development Report (2025) argues that robots need to move past pure “demonstration” data as quickly as possible and establish a continuous data-backflow loop from real deployment. The hard part is cost: if every robot needs an operator watching and ready to take over the whole time, it's difficult to make commercially viable.","example":"When training π*0.6, Physical Intelligence used the RECAP method to keep training on a mix of data autonomously generated while robots made espresso, folded laundry, and assembled boxes, together with human correction data; the company reported that throughput on some of the hardest tasks more than doubled, while the failure rate dropped by roughly half.","related":["Data Flywheel","Human Intervention Data","Recovery and Correction Data","Fleet Learning (Learning While Deploying)","RECAP","π*0.6"]},{"id":"data-flywheel","category":"data","sec":9,"tier":1,"sources":[{"title":"NVIDIA Glossary: What Is a Data Flywheel?","url":"https://www.nvidia.com/en-us/glossary/data-flywheel/"},{"title":"π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"2025-11","related_ids":["deployment-data-backflow","human-intervention-data","recap","self-improvement","human-in-the-loop","fleet-learning"],"name":"Data Flywheel","alt":"数据飞轮","abbr":"","aliases":["Data Closed Loop","Data Engine"],"one_liner":"A self-reinforcing loop: deployment produces data, the data improves the model, and the better model gets deployed even more widely.","explanation":"A data flywheel is a self-reinforcing loop: a model is deployed, generates new data through real use — successes, failures, human corrections — and that data, once filtered and labeled, is used to improve the model. A better model can then take on more tasks and reach more deployments, which brings back still more data. NVIDIA defines it as a self-improving loop that keeps refining a model using data collected from its own interactions. The self-driving industry adopted this kind of approach early — Tesla calls its version the data engine. Robot data is expensive to collect, so companies generally want real-world deployment to help spread out that cost; but the flywheel can only start turning once a robot is already useful enough to be deployed in the real world.","example":"Physical Intelligence's π*0.6 uses a method called RECAP, which feeds a robot's own autonomous-execution data — from folding laundry, brewing espresso on a professional machine, and assembling boxes in real households — back into training, alongside expert teleoperation corrections. The paper reports throughput on the hardest tasks more than doubling, with the failure rate roughly cut in half.","related":["Deployment Data Backflow","Human Intervention Data","RECAP","Self-improvement","Human-in-the-Loop","Fleet Learning (Learning While Deploying)"]},{"id":"supervised-learning","category":"training","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Supervised learning","url":"https://en.wikipedia.org/wiki/Supervised_learning"}],"as_of":"","related_ids":["behavior-cloning","unsupervised-learning","reinforcement-learning","loss-function","ground-truth","supervised-fine-tuning"],"name":"Supervised Learning","alt":"监督学习","abbr":"","aliases":[],"one_liner":"Training a model on input–correct-answer pairs so it learns to predict the answer from the input alone.","explanation":"Supervised learning is the most common machine-learning paradigm. Every training input comes paired with the correct output — a label, also called ground truth — and the model repeatedly shrinks the gap between its own prediction and that label, measured by a loss function. The ultimate goal is for the model to predict correctly on new data it has never seen, a property called generalization. When the output is a category, this is called classification; when it's a continuous number, it's regression. The contrasting paradigms are unsupervised learning, which uses no labels at all, and reinforcement learning, which learns from a reward signal instead. In embodied AI, behavior cloning is essentially supervised learning: the observations recorded during a human demonstration become the input, and the human's actions become the label, and a policy is trained to imitate them.","example":"Collect 1,000 paired examples of “camera image → arm joint angles” through teleoperation, then train a network that takes the image as input and outputs joint angles, with the loss defined as the mean squared error between predicted and demonstrated angles. That is behavior cloning framed as supervised learning.","related":["Behavior Cloning","Unsupervised Learning","Reinforcement Learning","Loss Function","Ground Truth","Supervised Fine-Tuning"]},{"id":"unsupervised-learning","category":"training","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Unsupervised learning","url":"https://en.wikipedia.org/wiki/Unsupervised_learning"}],"as_of":"","related_ids":["supervised-learning","self-supervised-learning","representation-learning","pre-training","unsupervised-skill-discovery","autoencoder"],"name":"Unsupervised Learning","alt":"无监督学习","abbr":"","aliases":[],"one_liner":"Using only unlabeled data and letting the model discover structure in it on its own.","explanation":"Unsupervised learning is the machine-learning paradigm contrasted with supervised learning: the training data has only inputs, no human-annotated answers, and the model must find patterns on its own. Classic tasks include clustering (such as k-means, which groups similar samples together), dimensionality reduction (such as principal component analysis, PCA), and density estimation (learning the data's underlying probability distribution); generative models like autoencoders are also often grouped here. Self-supervised learning (constructing a supervisory signal from the data itself, such as predicting a hidden part) is sometimes considered a branch of it. It matters because labeling is expensive while unlabeled data is abundant, and large-model pretraining relies heavily on unlabeled text, images, and video. In embodied AI, learning latent actions from action-free human video, and letting a robot discover a variety of behaviors through unsupervised skill discovery, both follow this line of thinking.","example":"Running k-means clustering on a batch of unlabeled tabletop scene images: the algorithm doesn't know any category names, it just groups the images by feature similarity.","related":["Supervised Learning","Self-Supervised Learning","Representation Learning","Pre-training","Unsupervised Skill Discovery","Autoencoder"]},{"id":"self-supervised-learning","category":"training","sec":0,"tier":2,"sources":[{"title":"Self-supervised learning: The dark matter of intelligence (Meta AI, 2021)","url":"https://ai.meta.com/blog/self-supervised-learning-the-dark-matter-of-intelligence/"},{"title":"Masked Autoencoders Are Scalable Vision Learners (arXiv 2111.06377)","url":"https://arxiv.org/abs/2111.06377"},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)","url":"https://arxiv.org/abs/2506.09985"}],"as_of":"","related_ids":["contrastive-learning","masked-autoencoder","pre-training","representation-learning","next-token-prediction","v-jepa-2"],"name":"Self-Supervised Learning","alt":"自监督学习","abbr":"SSL","aliases":["SSL","Self-Supervised Pretraining"],"one_liner":"Training a model on “questions and answers” constructed from the data itself, with no human labeling needed.","explanation":"Self-supervised learning is training where the supervisory signal comes from the data itself rather than human labels: part of the input is hidden, and the model is trained to predict it from the rest. Yann LeCun and colleagues called it the “dark matter” of intelligence and championed it heavily. Common forms include masking out words in text for the model to fill in (BERT; GPT's next-token prediction is also often grouped under this umbrella), masking out 75% of an image's patches and reconstructing them (MAE, masked autoencoders), and pulling representations of different views of the same content closer together (contrastive learning). It solves the problem that human labeling is expensive and caps how much data can be used, and it underlies the pretraining of today's large models. In embodied AI, since robot data with action labels is scarce, researchers often first self-supervise a visual representation or a world model on huge amounts of human video, then fine-tune it with a small amount of robot data.","example":"Meta's V-JEPA 2 first self-supervises on over 1 million hours of internet video, then trains an action-conditioned world model, V-JEPA 2-AC, on fewer than 62 hours of unlabeled DROID robot video, letting a Franka arm in labs it never trained on do pick-and-place by planning against an image goal, with no task-specific training.","related":["Contrastive Learning","Masked Autoencoder","Pre-training","Representation Learning","Next-Token Prediction","V-JEPA 2"]},{"id":"imitation-learning","category":"training","sec":0,"tier":1,"sources":[{"title":"Osa et al. 2018: An Algorithmic Perspective on Imitation Learning","url":"https://arxiv.org/abs/1811.06711"},{"title":"Wikipedia: Imitation learning","url":"https://en.wikipedia.org/wiki/Imitation_learning"},{"title":"Fu et al. 2024: Mobile ALOHA","url":"https://arxiv.org/abs/2401.02117"}],"as_of":"","related_ids":["behavior-cloning","inverse-reinforcement-learning","generative-adversarial-imitation-learning","dagger","demonstration-data","teleoperation"],"name":"Imitation Learning","alt":"模仿学习","abbr":"IL","aliases":["Learning from Demonstration (LfD)","Programming by Demonstration","IL"],"one_liner":"Teaching a robot a skill by having it imitate an expert's — usually a human's — demonstrated actions.","explanation":"Imitation learning studies how to get an agent to learn behavior from expert demonstrations, and it's well suited to situations where “just show it once” is easier than writing rules or designing a reward. Demonstrations are usually recorded from a human through teleoperation, kinesthetic teaching (physically guiding the robot's arm), or motion capture. There are three main approaches: behavioral cloning, which directly uses supervised learning to fit the demonstrated actions; inverse reinforcement learning, which first infers a reward function from the demonstrations and then optimizes a policy against it with reinforcement learning; and adversarial imitation methods such as GAIL, which push the policy's behavior distribution to match the expert's. The core difficulty is distribution shift: during execution, the policy ends up in states the demonstrations never covered, and methods like DAgger address this by asking the expert to label additional actions in exactly those states. The core training of virtually every mainstream VLA model today is, at heart, large-scale imitation learning.","example":"Mobile ALOHA collected just 50 human teleoperated demonstrations per task, co-trained with existing static-ALOHA data, and learned mobile manipulation tasks like sautéing shrimp and plating it, calling an elevator and riding it, and opening a cabinet to put away a pot.","related":["Behavior Cloning","Inverse Reinforcement Learning","Generative Adversarial Imitation Learning","DAgger","Demonstration Data","Teleoperation"]},{"id":"reinforcement-learning","category":"training","sec":0,"tier":1,"sources":[{"title":"Wikipedia: Reinforcement learning","url":"https://en.wikipedia.org/wiki/Reinforcement_learning"},{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"OpenAI et al. 2018: Learning Dexterous In-Hand Manipulation","url":"https://arxiv.org/abs/1808.00177"}],"as_of":"","related_ids":["markov-decision-process","reward-function","policy","proximal-policy-optimization","sim-to-real-transfer","reinforcement-fine-tuning"],"name":"Reinforcement Learning","alt":"强化学习","abbr":"RL","aliases":[],"one_liner":"A machine-learning approach where an agent improves its behavior through repeated trial and error, guided by a reward signal.","explanation":"Reinforcement learning is a machine-learning paradigm alongside supervised and unsupervised learning. An agent observes a state in an environment, takes an action, and the environment returns a new state along with a reward (a number measuring how good that step was); the goal is to learn a policy that maximizes long-term cumulative reward. It doesn't require labeling the “correct action” at every step — only a reward needs to be defined — but the agent has to balance trying new actions (exploration) against using actions already known to be good (exploitation). The problem is usually formalized as a Markov decision process. In embodied AI, reinforcement learning is mainly used in two places: training legged and humanoid robot locomotion control at massive scale in parallel simulation, then transferring it to the real robot; and fine-tuning a VLA model with reinforcement learning after an imitation-learning base has been trained, to improve success rate and error recovery.","example":"In 2018, OpenAI trained a Shadow dexterous hand entirely in simulation using reinforcement learning, randomizing physical parameters like friction coefficients; the resulting policy transferred directly to a real dexterous hand and performed vision-based in-hand object reorientation.","related":["Markov Decision Process","Reward Function","Policy","Proximal Policy Optimization","Sim-to-Real Transfer","Reinforcement Fine-Tuning (RL Fine-Tuning)"]},{"id":"training-validation-test-set","category":"training","sec":1,"tier":1,"sources":[{"title":"Google Machine Learning Crash Course: Dividing the original dataset","url":"https://developers.google.com/machine-learning/crash-course/overfitting/dividing-datasets"},{"title":"Google Machine Learning Glossary: validation set","url":"https://developers.google.com/machine-learning/glossary#validation-set"}],"as_of":"","related_ids":["overfitting","hyperparameter","checkpoint","generalization","out-of-distribution","data-leakage-test-set-contamination"],"name":"Training / Validation / Test Set","alt":"训练集 / 验证集 / 测试集","abbr":"","aliases":["Train/Val/Test Split","Training Set","Validation Set","Test Set","Held-out Set"],"one_liner":"Splitting a dataset into three parts: one to learn from, one to tune and select a model, and one saved for a final, honest check.","explanation":"This is the standard way machine learning splits its data. The training set is used to update the model's parameters. The validation set evaluates the model during development, guiding the choice of hyperparameters and which checkpoint (a saved snapshot of the model's weights partway through training) to keep. The test set is touched only at the very end, to report how the model performs on data it has truly never seen. Google's Machine Learning Crash Course uses a 70/15/15 split as an example, but the exact ratio isn't fixed; what matters is that the validation and test sets are large enough to give reliable conclusions. The three sets must not share samples — overlap is equivalent to peeking at the answer key and inflates the reported score — and a validation set that gets reused for many rounds of decisions can gradually lose its value as an honest check. In robotics, real-robot or simulation evaluation often plays the role of the test set, and it's just as important that the test scenes and objects never appeared in training.","example":"To train a grasping policy: train on data collected in 8 kitchens, use data from a 9th kitchen as validation to decide how long to train, then run real-robot trials in a never-before-seen 10th kitchen and report that success rate.","related":["Overfitting","Hyperparameter","Checkpoint","Generalization","Out-of-Distribution","Data Leakage / Test-Set Contamination"]},{"id":"ground-truth","category":"training","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Ground truth","url":"https://en.wikipedia.org/wiki/Ground_truth"},{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"}],"as_of":"","related_ids":["data-annotation","privileged-information","supervised-learning","action-label","synthetic-data","motion-capture"],"name":"Ground Truth","alt":"真值","abbr":"GT","aliases":["GT","Gold Label"],"one_liner":"The real data treated as the “correct answer” for training and evaluating a model.","explanation":"Ground truth originally comes from remote sensing, where it referred to data collected on the ground to calibrate satellite measurements; machine learning and statistical modeling borrowed the term to mean whatever data is treated as the correct answer during training and evaluation, such as an image's class label or an object's true position. Ground truth isn't necessarily perfect — human annotators make mistakes and sensors have error — so data quality directly caps how good a model can get. In embodied AI, imitation learning treats the actions recorded during human teleoperation as the action ground truth; a simulator can read out an object's exact pose, contact forces, and other precise state directly, so it's often used as ground truth for training a perception module, or as privileged information for a teacher policy; real-robot experiments often use poses measured by a motion-capture system as ground truth for evaluating an algorithm.","example":"When training a 6D pose-estimation model on simulator-rendered images, each image's true object pose comes straight from the simulator — that's ground truth, and it needs no human labeling.","related":["Data Annotation","Privileged Information","Supervised Learning","Action Label","Synthetic Data","Motion Capture"]},{"id":"loss-function","category":"training","sec":1,"tier":1,"sources":[{"title":"Google Machine Learning Glossary: loss function","url":"https://developers.google.com/machine-learning/glossary#loss-function"},{"title":"Wikipedia: Loss function","url":"https://en.wikipedia.org/wiki/Loss_function"},{"title":"Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["mean-squared-error","l1-loss","cross-entropy","denoising-loss","gradient-descent","backpropagation"],"name":"Loss Function","alt":"损失函数","abbr":"","aliases":["Cost Function","Objective Function"],"one_liner":"A single number measuring how far a model's prediction is from the correct answer; training just means shrinking it.","explanation":"A loss function measures how badly wrong a single prediction is, producing a real number that grows as the gap between prediction and target grows. During training, backpropagation computes the gradient of the loss with respect to every parameter, and the optimizer updates the parameters in the direction that shrinks the loss, repeating this until the loss stops improving meaningfully. Regression over continuous values commonly uses mean squared error (MSE) or L1 loss; classification, or predicting the next token, commonly uses cross-entropy. “Objective function” is a broader term — it can be a loss to minimize, or a quantity to maximize instead, such as cumulative reward in reinforcement learning. In embodied models, the choice of loss directly affects action quality: ACT uses L1 loss to regress actions, while diffusion policies and flow-matching models use a denoising loss or a flow-matching loss instead.","example":"The ACT paper found that L1 loss modeled the action sequence more precisely than the more common L2 (mean squared error) loss when regressing actions, so they switched to L1.","related":["Mean Squared Error","L1 Loss","Cross-Entropy","Denoising Loss (Diffusion Loss)","Gradient Descent","Backpropagation"]},{"id":"gradient-descent","category":"training","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Gradient descent","url":"https://en.wikipedia.org/wiki/Gradient_descent"},{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"}],"as_of":"","related_ids":["backpropagation","learning-rate","optimizer","loss-function","batch-size","adamw"],"name":"Gradient Descent","alt":"梯度下降","abbr":"","aliases":["Stochastic Gradient Descent","SGD","Mini-batch Gradient Descent"],"one_liner":"An optimization method that nudges parameters step by step opposite the loss function's gradient, making the loss smaller each time.","explanation":"Gradient descent is the core optimization method for training neural networks, generally credited to the mathematician Augustin-Louis Cauchy in 1847. The gradient is the loss function's partial derivative with respect to each parameter, pointing in the direction the loss increases fastest; each step moves the parameters a bit in the opposite direction, and how far is set by the learning rate — too small and convergence is slow, too large and the loss oscillates or even diverges. Because deep-learning datasets are large, each step usually estimates the gradient from just a small batch of data, called stochastic (mini-batch) gradient descent, or SGD. The gradient itself is computed by backpropagation. The Adam and AdamW optimizers commonly used to train large models like VLAs today are refinements of gradient descent that add momentum and adaptive step sizes.","example":"Training a behavior-cloning policy: each step takes 64 frames of demonstration, computes the mean squared error between predicted and demonstrated actions, runs backpropagation to get the gradient, then updates once via “parameter ← parameter − learning rate × gradient.”","related":["Backpropagation","Learning Rate","Optimizer","Loss Function","Batch Size","AdamW"]},{"id":"backpropagation","category":"training","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Backpropagation","url":"https://en.wikipedia.org/wiki/Backpropagation"},{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"}],"as_of":"","related_ids":["gradient-descent","optimizer","loss-function","vanishing-exploding-gradients","stop-gradient","neural-network"],"name":"Backpropagation","alt":"反向传播","abbr":"BP","aliases":["BP","Backprop","Error Backpropagation"],"one_liner":"The chain-rule algorithm that computes gradients layer by layer from the output backward; the core algorithm for training neural networks.","explanation":"Backpropagation is how a neural network's gradients get computed. Training first runs a forward pass to get an output and a loss, then works backward from that loss, applying the chain rule layer by layer from the last layer to the first, computing the gradient of the loss with respect to every parameter in one pass. Those gradients then go to an optimizer like gradient descent or AdamW to update the parameters, avoiding the huge redundant computation of differentiating each parameter separately. The idea traces back to Seppo Linnainmaa's automatic-differentiation work in 1970; Paul Werbos applied it to neural networks in 1974; and a 1986 Nature paper by Rumelhart, Hinton, and Williams brought it wide attention. In PyTorch, calling loss.backward() performs backpropagation.","example":"Training a behavior-cloning policy: each step runs a forward pass to compute the mean squared error between predicted and demonstrated actions, calls loss.backward() to get the gradient for every weight, then optimizer.step() to update the parameters.","related":["Gradient Descent","Optimizer","Loss Function","Vanishing / Exploding Gradients","Stop-Gradient","Neural Network"]},{"id":"optimizer","category":"training","sec":1,"tier":2,"sources":[{"title":"Adam: A Method for Stochastic Optimization","url":"https://arxiv.org/abs/1412.6980"},{"title":"Decoupled Weight Decay Regularization (AdamW)","url":"https://arxiv.org/abs/1711.05101"},{"title":"PyTorch Documentation: torch.optim","url":"https://docs.pytorch.org/docs/main/optim.html"}],"as_of":"","related_ids":["gradient-descent","backpropagation","learning-rate","adamw","learning-rate-schedule","gradient-clipping"],"name":"Optimizer","alt":"优化器","abbr":"","aliases":["Optimization Algorithm"],"one_liner":"The algorithm that turns gradients into parameter updates during training, such as SGD, Adam, or AdamW.","explanation":"When training a neural network, backpropagation computes the loss's gradient with respect to every parameter, and the optimizer decides how to turn that gradient into an actual update: which direction to move, and how big a step to take. The most basic is stochastic gradient descent (SGD), smoother once momentum is added; Adam (Kingma and Ba, 2014) tracks both the mean and the squared mean of the gradient, giving each parameter its own adaptive step size, and needs little tuning while converging fast; AdamW (Loshchilov and Hutter, ICLR 2019) pulls weight decay out of the gradient update, generalizing better, and is common when training Transformers, diffusion policies, and similar models. Optimizers are usually configured together with a learning rate schedule (warmup, cosine decay) and gradient clipping. Note that PNDbotics' robot named “Adam” is an unrelated humanoid robot product, not the Adam optimizer.","example":"Diffusion Policy's official training config uses torch.optim.AdamW with a learning rate of 1e-4, weight decay of 1e-6, and 500 warmup steps followed by cosine decay; in PyTorch, each training step calls zero_grad(), loss.backward(), and optimizer.step() in turn.","related":["Gradient Descent","Backpropagation","Learning Rate","AdamW","Learning Rate Schedule (Warmup and Cosine Decay)","Gradient Clipping"]},{"id":"adamw","category":"training","sec":1,"tier":2,"sources":[{"title":"Decoupled Weight Decay Regularization (arXiv 1711.05101)","url":"https://arxiv.org/abs/1711.05101"},{"title":"PyTorch docs: torch.optim.AdamW","url":"https://docs.pytorch.org/docs/2.14/generated/torch.optim.AdamW.html"}],"as_of":"","related_ids":["optimizer","gradient-descent","learning-rate","regularization","learning-rate-schedule"],"name":"AdamW","alt":"AdamW 优化器","abbr":"","aliases":["Adam with Decoupled Weight Decay"],"one_liner":"A version of the Adam optimizer that separates weight decay from the gradient update; the default choice for training most large models.","explanation":"AdamW comes from Ilya Loshchilov and Frank Hutter's paper “Decoupled Weight Decay Regularization” (ICLR 2019). An optimizer is the algorithm that updates a model's parameters based on gradients, and Adam adapts the step size for each parameter individually. The older way to regularize Adam was to fold an L2 penalty into the gradient, but that penalty then gets rescaled by Adam's adaptive step sizes along with everything else, which weakens its effect. The authors showed that this L2-in-the-gradient approach is equivalent to weight decay (shrinking every parameter by a small proportion at each step) for plain SGD, but not for Adam — so they pulled the decay step out and applied it separately from the gradient update. This improves generalization and decouples the best decay value from the learning rate. AdamW is now one of the most common optimizers for training Transformer-based models.","example":"In PyTorch, torch.optim.AdamW(model.parameters(), lr=1e-4, weight_decay=0.01) uses AdamW directly; PyTorch's own defaults for it are a learning rate of 1e-3, betas of (0.9, 0.999), and a weight decay of 0.01.","related":["Optimizer","Gradient Descent","Learning Rate","Regularization","Learning Rate Schedule (Warmup and Cosine Decay)"]},{"id":"hyperparameter","category":"training","sec":1,"tier":1,"sources":[{"title":"Google Machine Learning Glossary: hyperparameter","url":"https://developers.google.com/machine-learning/glossary#hyperparameter"},{"title":"Wikipedia: Hyperparameter (machine learning)","url":"https://en.wikipedia.org/wiki/Hyperparameter_(machine_learning)"}],"as_of":"","related_ids":["learning-rate","batch-size","epoch","overfitting","alchemy","population-based-training"],"name":"Hyperparameter","alt":"超参数","abbr":"","aliases":[],"one_liner":"A training setting chosen by a person beforehand rather than learned by the model itself, like the learning rate or batch size.","explanation":"The numbers inside a model are called parameters (or weights), and training learns them automatically. Hyperparameters, by contrast, are settings chosen by a person before training starts and usually left unchanged during it — they determine how the learning happens. Common ones include the learning rate, batch size, number of training steps or epochs, the network's depth and width, and the strength of regularization; reinforcement learning adds things like the discount factor and PPO's clipping coefficient, while embodied models add things like the action-chunk length or the number of denoising steps in a diffusion model. Poorly chosen hyperparameters can keep a model from converging, cause overfitting, or just leave performance well below what's possible. The process of searching for a good combination is called hyperparameter tuning, commonly done with grid search, random search, or Bayesian optimization; in practice, people also often just start from a paper's default values and adjust a little — informally called “tuning knobs,” or in Chinese ML slang, “alchemy” (炼丹).","example":"When training ACT, the action-chunk length k, the learning rate, and β (the weight on the KL term in the CVAE loss) are all hyperparameters that have to be set by hand.","related":["Learning Rate","Batch Size","Epoch","Overfitting","Alchemy (Deep-Learning Slang)","Population-Based Training"]},{"id":"learning-rate","category":"training","sec":1,"tier":2,"sources":[{"title":"Learning rate（Wikipedia）","url":"https://en.wikipedia.org/wiki/Learning_rate"},{"title":"Visual Instruction Tuning (LLaVA, arXiv 2304.08485)","url":"https://arxiv.org/abs/2304.08485"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["learning-rate-schedule","gradient-descent","optimizer","adamw","batch-size","hyperparameter"],"name":"Learning Rate","alt":"学习率","abbr":"LR","aliases":["LR","Step Size"],"one_liner":"The size of the step taken at each parameter update in gradient descent; one of the most important hyperparameters.","explanation":"The learning rate is the hyperparameter in an optimization algorithm that controls how big each parameter update is: parameters move in the direction opposite the gradient, by an amount equal to the learning rate times the gradient (adaptive optimizers like Adam further rescale this per parameter, based on that parameter's gradient history). Set it too high and parameters overshoot the minimum, making the loss oscillate or even diverge to NaN; set it too low and convergence is slow, or training gets stuck at a mediocre point. It's usually the first thing tuned, and it's tied to batch size — when batch size goes up, the learning rate is often scaled up proportionally too (the linear scaling rule). In practice, training rarely uses one fixed value throughout; instead it follows a learning rate schedule, typically warming up first and then decaying gradually. Fine-tuning a pretrained large model usually uses a much smaller learning rate than training from scratch, to avoid damaging what it already learned.","example":"LLaVA uses a learning rate of 2e-3 in its first stage, when only the projection layer is trained, and drops to 2e-5 in the second stage, when the language model is fine-tuned along with it. OpenVLA trains with a fixed 2e-5 throughout, and its authors found adding warmup gave no benefit; ACT uses 1e-5.","related":["Learning Rate Schedule (Warmup and Cosine Decay)","Gradient Descent","Optimizer","AdamW","Batch Size","Hyperparameter"]},{"id":"learning-rate-schedule","category":"training","sec":1,"tier":2,"sources":[{"title":"SGDR: Stochastic Gradient Descent with Warm Restarts (arXiv 1608.03983)","url":"https://arxiv.org/abs/1608.03983"},{"title":"Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (arXiv 1706.02677)","url":"https://arxiv.org/abs/1706.02677"},{"title":"openpi optimizer.py（CosineDecaySchedule 默认参数）","url":"https://github.com/Physical-Intelligence/openpi/blob/main/src/openpi/training/optimizer.py"}],"as_of":"","related_ids":["learning-rate","optimizer","adamw","batch-size","convergence","fine-tuning"],"name":"Learning Rate Schedule (Warmup and Cosine Decay)","alt":"学习率调度（预热与余弦退火）","abbr":"","aliases":["Warmup","Cosine Decay","Cosine Annealing","Learning Rate Decay"],"one_liner":"Changing the learning rate over the course of training: ramping it up during warmup, then easing it down along a cosine curve.","explanation":"A learning rate schedule changes the learning rate during training according to a predefined rule, rather than keeping it fixed throughout. The most common combination is “warmup plus cosine decay.” Warmup linearly raises the learning rate from near 0 up to its peak value over the first stretch of training: at that point the parameters are freshly initialized or a new module has just been attached, so gradients are unstable, and jumping straight to a large learning rate risks divergence. Goyal and colleagues (2017) used gradual warmup together with a linear scaling rule to scale ResNet-50 training up to a batch size of 8192 without losing accuracy. After warmup, the learning rate eases down to a very small value along a cosine curve, letting the model converge carefully in the later stages; this cosine shape comes from Loshchilov and Hutter's SGDR paper (ICLR 2017). Other common schedules include step decay, exponential decay, and linear decay.","example":"openpi's default recipe for fine-tuning π0 linearly ramps the learning rate up to 2.5e-5 over the first 1,000 steps, then eases it down to 2.5e-6 along a cosine curve over 30,000 steps.","related":["Learning Rate","Optimizer","AdamW","Batch Size","Convergence","Fine-tuning"]},{"id":"batch-size","category":"training","sec":1,"tier":2,"sources":[{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"2024-06","related_ids":["learning-rate","epoch","gradient-accumulation","distributed-training","hyperparameter","gradient-descent"],"name":"Batch Size","alt":"批大小","abbr":"","aliases":["Mini-batch Size"],"one_liner":"The number of samples fed through the model together for each single parameter update.","explanation":"Training doesn't use the whole dataset at once; instead the data is split into small mini-batches, and each one produces one average loss, one gradient, and one parameter update. The number of samples in each of those batches is the batch size. A larger batch size gives a more stable gradient estimate and better GPU utilization, but uses more memory; too small a batch makes the gradient noisier and training slower. Batch size is closely tied to the learning rate, so changing one usually means retuning the other. When memory is too limited for a large batch, two common workarounds are gradient accumulation (summing gradients over several small batches before updating) and multi-GPU data parallelism. It's worth distinguishing batch size from a training epoch, which is one full pass through the entire dataset and consists of many batches.","example":"OpenVLA trained for 14 days on 64 A100 GPUs, using a global batch size of 2048 and a fixed learning rate of 2e-5, passing over the training set 27 times in total.","related":["Learning Rate","Epoch","Gradient Accumulation","Distributed Training","Hyperparameter","Gradient Descent"]},{"id":"epoch","category":"training","sec":1,"tier":2,"sources":[{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"},{"title":"GeeksforGeeks: Epoch in Machine Learning","url":"https://www.geeksforgeeks.org/machine-learning/epoch-in-machine-learning/"},{"title":"legged_gym: legged_robot_config.py (num_learning_epochs = 5)","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["batch-size","gradient-descent","early-stopping","overfitting","underfitting","proximal-policy-optimization"],"name":"Epoch","alt":"训练轮次","abbr":"","aliases":["Training Epoch"],"one_liner":"One complete pass of the model through the entire training set.","explanation":"An epoch is a unit for measuring training progress. Training data is split into batches, and processing one batch and applying one parameter update is called an iteration (also called a step); once every sample in the training set has been processed once, that's one epoch, so the number of iterations per epoch is roughly the total sample count divided by the batch size. Too few epochs causes underfitting, too many causes overfitting, and early stopping is often used to decide where to stop. Large-model pretraining datasets are so enormous that the data is often passed over only once or a few times, so progress there is more often measured in training steps or tokens seen; small-scale robot demonstration datasets, by contrast, are often trained for many epochs. Reinforcement learning uses the term slightly differently: PPO reuses the same batch of rollout data several times, and each pass is also called an epoch.","example":"With 50 demonstrations totaling 20,000 frames and a batch size of 64, one epoch is roughly 313 steps. legged_gym's PPO config trains 5 epochs on each batch of rollout data (num_learning_epochs = 5).","related":["Batch Size","Gradient Descent","Early Stopping","Overfitting","Underfitting","Proximal Policy Optimization"]},{"id":"convergence","category":"training","sec":1,"tier":2,"sources":[{"title":"Google Machine Learning Glossary: convergence","url":"https://developers.google.com/machine-learning/glossary"},{"title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic, arXiv 2108.03298)","url":"https://arxiv.org/abs/2108.03298"}],"as_of":"","related_ids":["loss-function","learning-rate","gradient-descent","overfitting","checkpoint","early-stopping"],"name":"Convergence","alt":"收敛","abbr":"","aliases":["Training Convergence","Converge"],"one_liner":"The point in training where the loss or return stops changing much, meaning the model has settled into a stable state.","explanation":"Convergence describes training reaching a point where the loss (a number measuring the gap between prediction and target) or, in reinforcement learning, the return, stops changing much across iterations, so further training gives little benefit. Gradient-descent-style methods are only guaranteed to converge to a local optimum or stationary point under certain conditions; in practice, deep networks are judged by watching the training curve. Too large a learning rate can make the loss oscillate or even diverge, and exploding gradients or data problems can also prevent convergence. In robot learning, a converged loss doesn't necessarily mean the best policy: large-scale robomimic experiments found that the training objective and the real evaluation objective don't always agree, so picking a checkpoint by validation loss alone is unreliable — it's usually necessary to actually roll out the policy and check its success rate before choosing a model.","example":"When training a legged locomotion policy, the average-return curve rises quickly and then flattens out, which is taken as roughly converged. When training a behavior-cloning policy, even after the loss curve flattens, several checkpoints still get evaluated for success rate before deciding which one to use.","related":["Loss Function","Learning Rate","Gradient Descent","Overfitting","Checkpoint","Early Stopping"]},{"id":"checkpoint","category":"training","sec":1,"tier":1,"sources":[{"title":"PyTorch Tutorials: Saving and Loading Models","url":"https://docs.pytorch.org/tutorials/beginner/saving_loading_models.html"},{"title":"Google Machine Learning Glossary: checkpoint","url":"https://developers.google.com/machine-learning/glossary#checkpoint"},{"title":"GitHub: Physical-Intelligence/openpi","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["fine-tuning","pre-training","open-weight-model","safetensors","inference-deployment","epoch"],"name":"Checkpoint","alt":"检查点","abbr":"ckpt","aliases":["Model Weights","Model Checkpoint","ckpt"],"one_liner":"A saved snapshot of a model's parameters, taken during or after training, that can be loaded to run inference or resume training.","explanation":"A checkpoint is a saved snapshot of a model's parameters at a given moment. Training a large model often takes days or even weeks, so a program will save the parameters to a file every so often, both to avoid starting over after a crash and to make it easy to later pick the version that performed best on validation. If a checkpoint is only meant for inference, saving the model weights is enough; to resume training exactly where it left off, the optimizer state, current epoch, and similar information need to be saved alongside it — PyTorch's documentation notes that this kind of full checkpoint is typically 2 to 3 times the size of the weights alone. Common file formats include .pt, .pth, .ckpt, and safetensors. When an open-source VLA project says it has “released weights,” it means it has published a checkpoint, which anyone can download to run inference on directly or continue fine-tuning from.","example":"Physical Intelligence's openpi repository provides base checkpoints such as pi0_base and pi05_base for fine-tuning, as well as already fine-tuned checkpoints like pi05_libero and pi0_aloha_towel that can be downloaded and run for inference right away.","related":["Fine-tuning","Pre-training","Open-weight Model","safetensors","Inference Deployment","Epoch"]},{"id":"alchemy","category":"training","sec":1,"tier":2,"sources":[{"title":"Reflections on Random Kitchen Sinks (Ali Rahimi & Ben Recht, NIPS 2017 Test-of-Time talk)","url":"https://archives.argmin.net/2017/12/05/kitchen-sinks/"},{"title":"Wikipedia: Hyperparameter optimization","url":"https://en.wikipedia.org/wiki/Hyperparameter_optimization"}],"as_of":"","related_ids":["hyperparameter","learning-rate","batch-size","population-based-training","ablation-study","random-seed-and-reproducibility"],"name":"Alchemy (Deep-Learning Slang)","alt":"炼丹 / 调参（黑话）","abbr":"","aliases":["Elixir Refining","Tuning Wizard (slang)","Alchemist (slang)"],"one_liner":"Chinese deep-learning slang for training models and tuning hyperparameters by trial and error, likening it to alchemy.","explanation":"This is slang from China's deep-learning community that compares training a model to alchemy. Data and a model go into a GPU server (jokingly the “alchemy furnace”), you set hyperparameters — configuration choices like learning rate and batch size fixed before training starts, nicknamed the “heat” of the fire — then wait hours or days for a result whose cause is often hard to pin down, so all you can do is try again. Adjusting these hyperparameters is called “炼丹” (elixir-refining) or “调参” (tuning), and the people who do it jokingly call themselves “tuning wizards” or “alchemists.” English-speaking ML has voiced a similar complaint: in his 2017 NeurIPS Test of Time talk, Ali Rahimi said outright that “machine learning has become alchemy,” criticizing the field's reliance on techniques nobody fully understands. More systematic approaches to tuning include grid search, random search, Bayesian optimization, and population-based training.","example":"A VLA fine-tune isn't working well, so a student changes the learning rate from 1e-4 to 2e-5, the batch size from 32 to 128, trains for 20,000 more steps, and checks again — round after round of this trial and error is what “alchemy” refers to.","related":["Hyperparameter","Learning Rate","Batch Size","Population-Based Training","Ablation Study","Random Seed and Reproducibility"]},{"id":"random-seed-and-reproducibility","category":"training","sec":1,"tier":2,"sources":[{"title":"Deep Reinforcement Learning that Matters (Henderson et al., AAAI 2018)","url":"https://arxiv.org/abs/1709.06560"},{"title":"PyTorch Docs: Reproducibility","url":"https://docs.pytorch.org/docs/stable/notes/randomness.html"},{"title":"Deep Reinforcement Learning at the Edge of the Statistical Precipice (NeurIPS 2021)","url":"https://arxiv.org/abs/2108.13264"}],"as_of":"","related_ids":["hyperparameter","ablation-study","statistical-rigor-in-policy-evaluation","benchmark","training-validation-test-set"],"name":"Random Seed and Reproducibility","alt":"随机种子与可复现性","abbr":"","aliases":["Random Seed","Reproducibility","Multi-seed Experiments"],"one_liner":"Fixing the random number generator's starting point so runs can be repeated, and checking results across several seeds to rule out luck.","explanation":"Many parts of deep learning involve randomness: network initialization, data shuffling, dropout, reinforcement learning's exploration noise, and a simulator's domain randomization. A random seed is the starting point for a random number generator; fixing the seed in Python, NumPy, and PyTorch lets the same code produce close to the same result, though PyTorch's own documentation notes exact reproducibility isn't guaranteed across versions or platforms. Reproducibility also has a second meaning: whether a conclusion can be trusted. Henderson and colleagues (2018) found that in deep reinforcement learning, changing only the random seed can shift results enough to flip which algorithm looks better. Because of this, results should be reported as a mean and confidence interval across multiple seeds; Agarwal and colleagues (2021) further recommend using the interquartile mean (IQM) and bootstrap confidence intervals.","example":"In PyTorch, fixing seeds looks like torch.manual_seed(0), np.random.seed(0), random.seed(0), plus setting torch.backends.cudnn.benchmark=False and torch.use_deterministic_algorithms(True); when comparing two reinforcement-learning algorithms, each is run once per seed across 5 different seeds, and the mean and confidence interval are reported.","related":["Hyperparameter","Ablation Study","Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)","Benchmark","Training / Validation / Test Set"]},{"id":"overfitting","category":"training","sec":1,"tier":1,"sources":[{"title":"Google Machine Learning Glossary: overfitting","url":"https://developers.google.com/machine-learning/glossary#overfitting"},{"title":"Wikipedia: Overfitting","url":"https://en.wikipedia.org/wiki/Overfitting"}],"as_of":"","related_ids":["underfitting","generalization","regularization","early-stopping","data-augmentation","training-validation-test-set"],"name":"Overfitting","alt":"过拟合","abbr":"","aliases":[],"one_liner":"When a model memorizes the training data too closely, so its performance drops on new data it hasn't seen.","explanation":"Overfitting means a model has fit the training data too tightly, memorizing even its noise and incidental details, so its predictions become unreliable on data it hasn't seen before. The classic sign is that training loss keeps falling while validation loss starts rising; it's most likely to happen when data is scarce, the model is large, or training runs for too many epochs. Common remedies include adding more data or using data augmentation, regularization, dropout, early stopping, and picking the checkpoint that performs best on a validation set. This problem is especially prominent in embodied AI: real-robot demonstrations often number only in the dozens to low hundreds, so a policy can end up memorizing the object placement, lighting, and background it saw during data collection — and fail once an object moves a few centimeters or the tablecloth changes — which is why papers specifically test positional and scene generalization. The opposite problem, where a model fails to fit even the training data well, is called underfitting.","example":"For instance, training a cup-grasping policy on just 50 demonstrations where the cup always sits in the same spot: after enough training, it succeeds almost every time at that original spot but grasps at empty air once the cup moves 10cm away — the policy memorized that one fixed trajectory instead of learning to “find the cup, then grasp it.”","related":["Underfitting","Generalization","Regularization","Early Stopping","Data Augmentation","Training / Validation / Test Set"]},{"id":"underfitting","category":"training","sec":1,"tier":2,"sources":[{"title":"Google Machine Learning Crash Course: Overfitting","url":"https://developers.google.com/machine-learning/crash-course/overfitting/overfitting"},{"title":"Wikipedia: Overfitting（含 Underfitting 一节）","url":"https://en.wikipedia.org/wiki/Overfitting"}],"as_of":"","related_ids":["overfitting","regularization","training-validation-test-set","loss-function","parameter-count","convergence"],"name":"Underfitting","alt":"欠拟合","abbr":"","aliases":[],"one_liner":"A model too weak or too undertrained to learn even the patterns already present in the training data.","explanation":"Underfitting is a basic machine-learning concept, the opposite of overfitting: the model's error is already high on the training set, and it does just as poorly on validation and test data. Common causes include insufficient model capacity (too few parameters, too simple an architecture), too little training, a poorly chosen learning rate, too much regularization, or inputs that are missing information the task actually needs. It's diagnosed by watching the training loss: training loss that won't come down is underfitting, while training loss that's low but validation loss that rises again is overfitting. In robot learning, a typical symptom is a policy that fails even on the exact demonstration scenes it trained on — the right fix is usually a bigger model or longer training, or checking whether the observation and action data (camera images, action normalization) is broken, rather than reaching for anti-overfitting tools like data augmentation.","example":"Training a very small multilayer perceptron to imitate demonstrations of bimanual clothes-folding: the training loss plateaus at a high value early on, and the policy still can't grasp the fabric correctly even when replayed on the exact scenes it trained on — that's underfitting.","related":["Overfitting","Regularization","Training / Validation / Test Set","Loss Function","Parameter Count (Model Size)","Convergence"]},{"id":"regularization","category":"training","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Regularization (mathematics)","url":"https://en.wikipedia.org/wiki/Regularization_(mathematics)"},{"title":"Dive into Deep Learning: Weight Decay","url":"https://d2l.ai/chapter_linear-regression/weight-decay.html"}],"as_of":"","related_ids":["overfitting","dropout","early-stopping","data-augmentation","entropy-regularization","kl-regularization"],"name":"Regularization","alt":"正则化","abbr":"","aliases":["Regularization Term"],"one_liner":"Adding constraints or penalties during training to keep a model from memorizing training data, improving performance on new data.","explanation":"Regularization is a broad category of techniques for preventing overfitting: whenever a model keeps improving on the training set but gets worse on new data, regularization is used to limit the model's effective complexity. Explicit regularization adds a penalty term to the loss function, such as L2 regularization (also called weight decay, which penalizes the sum of squared weights and shrinks them overall) and L1 regularization (which penalizes absolute value and pushes many weights to exactly zero); implicit regularization includes early stopping, dropout (randomly zeroing some neurons during training), and data augmentation. Reinforcement learning has several regularizers of its own: entropy regularization encourages the policy to stay somewhat random to keep exploring, KL regularization keeps a fine-tuned policy from drifting too far from the original model, and behavior regularization keeps an offline reinforcement-learning policy close to the actions in the dataset.","example":"Training a visuomotor policy applies random cropping to camera images (data augmentation) and sets weight decay in the AdamW optimizer; RLHF subtracts a KL penalty term from the reward, keeping the model from drifting away from the original language model just to rack up a higher score.","related":["Overfitting","Dropout","Early Stopping","Data Augmentation","Entropy Regularization","KL Regularization"]},{"id":"dropout","category":"training","sec":1,"tier":2,"sources":[{"title":"Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR 2014)","url":"https://jmlr.org/papers/v15/srivastava14a.html"},{"title":"ACT 官方代码 detr/main.py（--dropout 默认 0.1）","url":"https://github.com/tonyzhaozh/act/blob/main/detr/main.py"},{"title":"Dropout Q-Functions for Doubly Efficient Reinforcement Learning (DroQ)","url":"https://arxiv.org/abs/2110.02034"}],"as_of":"","related_ids":["regularization","overfitting","early-stopping","data-augmentation","normalization-layers"],"name":"Dropout","alt":"随机失活","abbr":"","aliases":["Dropout Regularization"],"one_liner":"Randomly “turning off” some neurons during training, a regularization technique that keeps a network from memorizing the training data.","explanation":"Dropout comes from Hinton's group; Srivastava, Hinton, and colleagues published the systematic paper in JMLR in 2014. At every training step, each neuron's output is randomly zeroed out with some set probability, which is equivalent to training a different “sub-network” each time; at test time dropout is turned off, the full network is used, and the weights are rescaled accordingly to approximate the average over all those sub-networks. This keeps neurons from relying too heavily on each other, which reduces overfitting — doing well on training data but poorly on new data. It's commonly combined with weight decay, data augmentation, and early stopping; a dropout probability of 0.1 is common in Transformers. Robot demonstration datasets often have only dozens to a few hundred examples and overfit easily, so dropout is a common setting there.","example":"The official ACT (Action Chunking with Transformers) code uses a default Transformer dropout of 0.1. The reinforcement-learning algorithm DroQ adds dropout and layer normalization inside its Q-network, matching the sample efficiency of the larger ensemble method REDQ while using computation close to plain SAC.","related":["Regularization","Overfitting","Early Stopping","Data Augmentation","Normalization Layers"]},{"id":"early-stopping","category":"training","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Early stopping","url":"https://en.wikipedia.org/wiki/Early_stopping"},{"title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic)","url":"https://arxiv.org/abs/2108.03298"},{"title":"Google Machine Learning Glossary","url":"https://developers.google.com/machine-learning/glossary"}],"as_of":"","related_ids":["overfitting","training-validation-test-set","checkpoint","regularization","epoch","early-termination"],"name":"Early Stopping","alt":"早停","abbr":"","aliases":[],"one_liner":"Stopping training once validation performance stops improving, and keeping the checkpoint that performed best.","explanation":"Early stopping is one of the most common regularization techniques. During training, the model is periodically evaluated on a validation set (data held out from training, used specifically to pick a model): once training loss keeps falling but validation loss starts rising, the model has begun to overfit, so training stops and the checkpoint with the best validation performance is kept instead of the final one. In practice, a “patience” value is often set — stop only after several rounds in a row with no improvement — to avoid being misled by noise. In robot imitation learning this needs extra care: low validation loss doesn't guarantee a high real-robot success rate. robomimic's research found the training objective and the evaluation objective don't always agree, so which step training stops at matters a great deal, and many projects save several checkpoints and pick among them using simulation or real-robot rollouts instead. Don't confuse this with “early termination” in reinforcement learning, which ends a single episode early, for instance when the robot falls.","example":"Say a grasping policy is trained with validation action error computed once per epoch: the error is lowest at epoch 30 and doesn't improve for 10 epochs after that, so training stops and the epoch-30 checkpoint is kept.","related":["Overfitting","Training / Validation / Test Set","Checkpoint","Regularization","Epoch","Early Termination"]},{"id":"vanishing-exploding-gradients","category":"training","sec":1,"tier":2,"sources":[{"title":"Dive into Deep Learning: Numerical Stability and Initialization","url":"https://d2l.ai/chapter_multilayer-perceptrons/numerical-stability-and-init.html"},{"title":"On the difficulty of training Recurrent Neural Networks (Pascanu et al., arXiv 1211.5063)","url":"https://arxiv.org/abs/1211.5063"}],"as_of":"","related_ids":["backpropagation","gradient-clipping","residual-network","activation-function","normalization-layers","long-short-term-memory-gated-recurrent-unit"],"name":"Vanishing / Exploding Gradients","alt":"梯度消失 / 梯度爆炸","abbr":"","aliases":["Vanishing Gradient","Exploding Gradient"],"one_liner":"Gradients shrinking to near zero or blowing up as they're multiplied across many layers during backpropagation, stalling deep-network training.","explanation":"A deep network computes gradients through backpropagation, and the gradient reaching an early layer is a product of the derivatives (matrices) of every layer after it. When those factors tend to be small, the gradient shrinks toward zero by the time it reaches shallow layers, so the early layers barely learn at all — vanishing gradients; when the factors tend to be large, the product grows exponentially, and a single update can send parameters diverging and the loss to NaN — exploding gradients. Bengio and colleagues analyzed this problem in recurrent neural networks in 1994, and Pascanu and colleagues proposed gradient-norm clipping to handle explosion in 2013. Other common fixes include activation functions that don't saturate easily, like ReLU (sigmoid's derivative approaches zero when its input is very large or very small), Xavier/He initialization, residual connections, normalization layers, and LSTM's gating structure. The gradient clipping commonly set when training VLAs and diffusion policies is also there to prevent this kind of numerical instability.","example":"Stacking dozens of fully connected layers with sigmoid activations, the gradient in the first few layers is often close to 0 and their parameters barely move; switching to ReLU and adding residual connections lets a network of the same depth train normally.","related":["Backpropagation","Gradient Clipping","Residual Network","Activation Function","Normalization Layers","Long Short-Term Memory / Gated Recurrent Unit"]},{"id":"gradient-clipping","category":"training","sec":1,"tier":3,"sources":[{"title":"On the difficulty of training Recurrent Neural Networks (arXiv 1211.5063)","url":"https://arxiv.org/abs/1211.5063"},{"title":"torch.nn.utils.clip_grad_norm_ (PyTorch 文档)","url":"https://docs.pytorch.org/docs/main/generated/torch.nn.utils.clip_grad_norm_.html"}],"as_of":"","related_ids":["vanishing-exploding-gradients","learning-rate","optimizer","backpropagation","recurrent-neural-network","proximal-policy-optimization"],"name":"Gradient Clipping","alt":"梯度裁剪","abbr":"","aliases":["Gradient Norm Clipping"],"one_liner":"Scaling down an overly large gradient to stay within a threshold, so one bad update can't wreck the model.","explanation":"Training occasionally hits exploding gradients: one step's gradient is unusually large, pushing the parameters much too far and sending the loss spiking or even to NaN. Gradient clipping checks the gradient's magnitude before the optimizer applies it, and shrinks it if it exceeds a threshold. The most common form is norm clipping: treat every parameter's gradient as one long vector, compute its total norm, and if that exceeds max_norm, multiply the whole vector by max_norm divided by the total norm, which keeps the direction the same and only shortens its length; an alternative, value clipping, clamps each individual component to a fixed range. Pascanu, Mikolov, and Bengio proposed gradient-norm clipping in 2013 while analyzing exploding gradients in recurrent neural networks. It's now essentially a default setting when training Transformers and diffusion policies, and common PPO implementations generally include it too.","example":"In PyTorch, calling torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0) after loss.backward() and before optimizer.step().","related":["Vanishing / Exploding Gradients","Learning Rate","Optimizer","Backpropagation","Recurrent Neural Network","Proximal Policy Optimization"]},{"id":"stop-gradient","category":"training","sec":1,"tier":3,"sources":[{"title":"Chen & He 2020: Exploring Simple Siamese Representation Learning (SimSiam)","url":"https://arxiv.org/abs/2011.10566"},{"title":"Driess et al. 2025: Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better","url":"https://arxiv.org/abs/2505.23705"}],"as_of":"2025-05","related_ids":["knowledge-insulation","target-network","representation-collapse","backbone-freezing","backpropagation","vector-quantized-variational-autoencoder"],"name":"Stop-Gradient","alt":"梯度阻断","abbr":"","aliases":["stop-grad","detach","sg(·)"],"one_liner":"A value is used normally in the forward pass but treated as a constant during backpropagation, blocking its gradient.","explanation":"Stop-gradient is a basic deep-learning operation: some quantity is used normally in the forward computation, but during backpropagation it is treated as a constant, so no gradient flows back through it to earlier parameters. Papers often write it as sg(·); PyTorch uses .detach(), and JAX uses jax.lax.stop_gradient. It has many uses: in reinforcement learning it stabilizes temporal-difference learning by blocking gradients into the target value; VQ-VAE uses it so the codebook and the encoder can be updated separately; and SimSiam found it to be the key ingredient that stops self-supervised learning's representations from collapsing. In VLAs, Physical Intelligence's 2025 knowledge insulation blocks the gradient flowing from the action expert back into the VLM backbone, so a newly initialized action expert doesn't disrupt the backbone's pretrained knowledge, and the backbone instead learns from discretized action tokens. Note this is not the same as gradient clipping, which limits gradient magnitude rather than blocking it.","example":"In knowledge-insulation training, the continuous action expert reads features from the VLM backbone to generate actions, but its loss is never backpropagated into the backbone; the backbone adapts to robot data only through the next-token-prediction loss on discretized action tokens.","related":["Knowledge Insulation","Target Network","Representation Collapse","Backbone Freezing","Backpropagation","Vector-Quantized Variational Autoencoder"]},{"id":"exponential-moving-average","category":"training","sec":1,"tier":3,"sources":[{"title":"Denoising Diffusion Probabilistic Models (arXiv:2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"diffusion_policy: train_diffusion_unet_image_workspace.yaml","url":"https://github.com/real-stanford/diffusion_policy/blob/main/diffusion_policy/config/train_diffusion_unet_image_workspace.yaml"},{"title":"OpenAI Spinning Up: Soft Actor-Critic","url":"https://spinningup.openai.com/en/latest/algorithms/sac.html"}],"as_of":"","related_ids":["target-network","diffusion-policy","checkpoint","optimizer","denoising-diffusion-probabilistic-model","soft-actor-critic"],"name":"Exponential Moving Average","alt":"指数移动平均","abbr":"EMA","aliases":["EMA","Polyak Averaging"],"one_liner":"A weighted average of past values that decays exponentially, often used to get smoother, more stable model weights.","explanation":"EMA is a weighted average updated at every step as “average ← β × average + (1 − β) × current value,” where β is the decay rate — the closer to 1, the longer the effective averaging window. The most common use in deep learning is keeping an EMA copy of the model's weights: training keeps updating the original weights as usual, while evaluation and deployment switch to the smoother EMA weights, which are usually more stable; diffusion models depend on this especially heavily, with the DDPM paper using a decay rate of 0.9999. In reinforcement learning, the soft update of a target network (Polyak averaging) is also an EMA, keeping the Q-value target changing more smoothly. The Adam optimizer's running estimates of a gradient's first and second moments are EMAs too.","example":"Diffusion Policy's official code enables EMA by default, with the decay rate ramping up from 0 to a cap of 0.9999 over training, and evaluation rollouts use the EMA model rather than the raw weights.","related":["Target Network","Diffusion Policy","Checkpoint","Optimizer","Denoising Diffusion Probabilistic Model","Soft Actor-Critic"]},{"id":"mean-squared-error","category":"training","sec":2,"tier":2,"sources":[{"title":"Linear regression: Loss（Google Machine Learning Crash Course）","url":"https://developers.google.com/machine-learning/crash-course/linear-regression/loss"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["l1-loss","loss-function","maximum-likelihood-estimation","action-multimodality","continuous-action-regression","denoising-loss"],"name":"Mean Squared Error","alt":"均方误差","abbr":"MSE","aliases":["MSE","L2 Loss","Squared Error Loss"],"one_liner":"The average of the squared difference between predictions and ground truth; the most common loss for regression.","explanation":"Mean squared error squares the difference between each sample's prediction and its ground truth, then averages over all samples; the sum without averaging is often called L2 loss. It's differentiable everywhere, and its gradient shrinks as the error shrinks, giving smooth optimization, which makes it the default loss for regression tasks. From a probabilistic view, minimizing MSE is equivalent to maximum likelihood estimation under the assumption that errors are Gaussian, with the optimal solution being the conditional mean. Because it squares the error, large errors get amplified, so MSE is more easily thrown off by outliers than L1 loss. In robot imitation learning it has a well-known failure mode: if a demonstrator sometimes goes around an obstacle from the left and sometimes from the right, regressing actions directly with MSE learns the average of the two, producing a path down the middle that matches neither — this is the action-multimodality problem, and it's part of why generative policies like diffusion policies and flow matching became popular. A diffusion model's own training objective is also MSE, just applied to the noise that was added rather than to the action itself.","example":"Diffusion Policy adds random noise to an expert action sequence during training and has the network predict that added noise; the loss is the MSE between predicted and true noise, and the paper notes that minimizing this loss also minimizes a variational bound on the KL divergence between the data distribution and the model's distribution.","related":["L1 Loss","Loss Function","Maximum Likelihood Estimation","Action Multimodality","Continuous Action Regression","Denoising Loss (Diffusion Loss)"]},{"id":"l1-loss","category":"training","sec":2,"tier":2,"sources":[{"title":"Linear regression: Loss（Google Machine Learning Crash Course）","url":"https://developers.google.com/machine-learning/crash-course/linear-regression/loss"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"","related_ids":["mean-squared-error","loss-function","continuous-action-regression","action-chunking-with-transformers","openvla-oft","action-chunking"],"name":"L1 Loss","alt":"L1 损失","abbr":"","aliases":["Mean Absolute Error","MAE","Absolute Error Loss"],"one_liner":"A loss function that measures error as the absolute value of the difference between a prediction and the ground truth.","explanation":"L1 loss takes the absolute value of the difference between each prediction and its ground truth and sums it up; averaged per sample, it's called mean absolute error (MAE). Compared with L2 loss (mean squared error), which squares the error, L1 penalizes large errors linearly rather than blowing them up quadratically, so it's less thrown off by outliers — but its gradient has constant magnitude, so it doesn't automatically shrink as the prediction nears the target. From a probabilistic view, minimizing L1 loss corresponds to maximum likelihood estimation under the assumption that errors follow a Laplace distribution, and the optimal solution trends toward the median rather than the mean. In embodied AI, L1 is common for continuous action regression: ACT uses L1 for reconstructing action sequences, with the authors reporting it models action sequences more precisely than the more common L2; OpenVLA-OFT also switched to L1 regression to output continuous actions directly.","example":"OpenVLA-OFT replaced OpenVLA's approach of generating discrete action tokens one at a time with parallel decoding, action chunking, and L1 regression for continuous action output, raising action-generation throughput 26-fold and lifting the average success rate across four LIBERO task suites from 76.5% to 97.1%.","related":["Mean Squared Error","Loss Function","Continuous Action Regression","Action Chunking with Transformers","OpenVLA-OFT","Action Chunking"]},{"id":"cross-entropy","category":"training","sec":2,"tier":2,"sources":[{"title":"Google Machine Learning Glossary: cross-entropy","url":"https://developers.google.com/machine-learning/glossary"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"","related_ids":["loss-function","next-token-prediction","action-binning","mean-squared-error","kullback-leibler-divergence","maximum-likelihood-estimation"],"name":"Cross-Entropy","alt":"交叉熵","abbr":"CE","aliases":["CE","Cross-Entropy Loss","Log Loss"],"one_liner":"A measure of how far a model's predicted probability distribution is from the correct answer; the standard loss for classification.","explanation":"Cross-entropy comes from information theory, where it measures the difference between two probability distributions. In deep learning it's used as a classification loss: the model outputs a probability for each class (usually normalized with softmax), and the loss grows the lower the predicted probability assigned to the correct class is; when there's exactly one correct class per label, it's equivalent to negative log-likelihood (NLL). A large language model's next-token prediction is cross-entropy computed over the entire vocabulary. In embodied AI, VLA models that discretize continuous actions into tokens, such as RT-2 and OpenVLA, also train their action outputs with cross-entropy; policies that regress continuous actions directly usually use mean squared error or L1 loss instead, and diffusion and flow-matching policies use their own denoising losses.","example":"OpenVLA divides each action dimension into 256 bins spanning the 1st to 99th percentile of the training data, maps them onto the 256 least-used tokens in the Llama vocabulary, and during training computes cross-entropy loss only on those action tokens.","related":["Loss Function","Next-Token Prediction","Action Binning","Mean Squared Error","Kullback-Leibler Divergence","Maximum Likelihood Estimation"]},{"id":"maximum-likelihood-estimation","category":"training","sec":2,"tier":2,"sources":[{"title":"Maximum likelihood estimation（Wikipedia）","url":"https://en.wikipedia.org/wiki/Maximum_likelihood_estimation"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["cross-entropy","kullback-leibler-divergence","mean-squared-error","behavior-cloning","loss-function","gaussian-policy"],"name":"Maximum Likelihood Estimation","alt":"最大似然估计 / 负对数似然","abbr":"MLE / NLL","aliases":["MLE","Negative Log-Likelihood","NLL","NLL Loss"],"one_liner":"Finding the parameters that make the observed data most probable; taking the negative log of that probability gives a loss to minimize.","explanation":"Maximum likelihood estimation (MLE) is statistics' most basic method for estimating parameters: assume the data comes from some probability model with unknown parameters, and pick the parameters that make the probability of the observed data — the likelihood — as large as possible. In practice, a logarithm turns the product of probabilities into a sum, and a minus sign turns it into a minimization problem — this is exactly the negative log-likelihood (NLL) loss used in deep learning. Many familiar losses are special cases of it: cross-entropy for classification is the negative log-likelihood under a categorical distribution; when errors are assumed Gaussian, MLE is equivalent to minimizing mean squared error; when errors are assumed to follow a Laplace distribution, it corresponds to L1 loss. It's also equivalent to minimizing the KL divergence between the data distribution and the model's distribution. In embodied AI, behavior cloning is essentially maximum likelihood estimation on expert actions: a VLA that discretizes actions predicts action tokens with cross-entropy, while a Gaussian policy directly minimizes the negative log-likelihood of the action.","example":"OpenVLA divides each action dimension into 256 evenly spaced bins between the 1st and 99th percentile of the training data, then uses the standard next-token-prediction objective, computing cross-entropy only on the action tokens — which is maximum likelihood estimation applied to expert actions.","related":["Cross-Entropy","Kullback-Leibler Divergence","Mean Squared Error","Behavior Cloning","Loss Function","Gaussian Policy"]},{"id":"kullback-leibler-divergence","category":"training","sec":2,"tier":2,"sources":[{"title":"Kullback–Leibler divergence（Wikipedia）","url":"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["cross-entropy","kl-regularization","variational-autoencoder","proximal-policy-optimization","knowledge-distillation","maximum-likelihood-estimation"],"name":"Kullback-Leibler Divergence","alt":"KL 散度","abbr":"KL","aliases":["KL Divergence","KL","Relative Entropy"],"one_liner":"A measure of how far one probability distribution is from another; not symmetric between the two.","explanation":"KL divergence was introduced by Solomon Kullback and Richard Leibler in 1951; in the discrete case, D_KL(P‖Q) = Σ P(x)·log(P(x)/Q(x)), representing the extra information cost of approximating the true distribution P with a distribution Q. It's always non-negative, and equals zero only when the two distributions are identical, but it isn't symmetric — D_KL(P‖Q) generally doesn't equal D_KL(Q‖P) — so it isn't a distance in the strict mathematical sense. It shows up everywhere in machine learning: cross-entropy equals KL divergence plus P's own entropy, so minimizing cross-entropy is equivalent to minimizing KL; variational autoencoders use a KL term to pull the latent variable toward a standard normal distribution; PPO, RLHF, and some offline reinforcement-learning methods use a KL penalty to keep a new policy from drifting too far from a reference policy; and knowledge distillation uses it to pull the student's output distribution toward the teacher's.","example":"ACT's training loss is the action-reconstruction error plus β times a KL term (the paper uses β = 10); the KL term keeps the style variable z produced by its conditional variational autoencoder close to a standard normal distribution, and at inference time z is simply set to the prior's mean of 0.","related":["Cross-Entropy","KL Regularization","Variational Autoencoder","Proximal Policy Optimization","Knowledge Distillation","Maximum Likelihood Estimation"]},{"id":"next-token-prediction","category":"training","sec":2,"tier":2,"sources":[{"title":"Improving Language Understanding by Generative Pre-Training (GPT-1, OpenAI 2018)","url":"https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/html/2406.09246"},{"title":"Humanoid Locomotion as Next Token Prediction","url":"https://arxiv.org/abs/2402.19469"}],"as_of":"","related_ids":["autoregressive-decoding","action-tokenizer","action-binning","teacher-forcing","cross-entropy","vision-language-action-model"],"name":"Next-Token Prediction","alt":"下一个 token 预测","abbr":"NTP","aliases":["NTP","Autoregressive Pretraining","Causal Language Modeling"],"one_liner":"The training objective of predicting the next token in a sequence given everything that came before it.","explanation":"This is the core training objective behind the GPT family of large language models. Text is cut into tokens; the model reads all the preceding tokens and outputs a probability distribution over the next one, and cross-entropy loss pushes up the probability assigned to the actual next token. It needs only raw text, no human labeling, so it scales to enormous pretraining datasets — OpenAI's GPT-1 (2018) already used it for generative pretraining. At inference time, the model generates tokens one after another, called autoregressive decoding. In embodied AI, RT-2 and OpenVLA discretize continuous actions into tokens and place them in the same sequence as text, reusing this same objective to train a VLA; a UC Berkeley team in 2024 even modeled humanoid-robot walking as next-token prediction. The cost is that discretization loses precision and generating tokens one at a time is slow, which is why some models switch to diffusion or flow-matching action heads instead.","example":"OpenVLA divides each action dimension into 256 bins based on the training-data distribution, represents them using the 256 least-used tokens in the Llama vocabulary, and trains with the standard next-token-prediction objective, computing cross-entropy only on the action tokens.","related":["Autoregressive Decoding","Action Tokenizer","Action Binning","Teacher Forcing","Cross-Entropy","Vision-Language-Action Model"]},{"id":"teacher-forcing","category":"training","sec":2,"tier":3,"sources":[{"title":"Wikipedia: Teacher forcing","url":"https://en.wikipedia.org/wiki/Teacher_forcing"},{"title":"Bengio et al. 2015: Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks","url":"https://arxiv.org/abs/1506.03099"},{"title":"Huang et al. 2025: Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","url":"https://arxiv.org/abs/2506.08009"}],"as_of":"","related_ids":["next-token-prediction","autoregressive-decoding","exposure-bias","self-forcing","diffusion-forcing","compounding-error"],"name":"Teacher Forcing","alt":"教师强制","abbr":"","aliases":[],"one_liner":"Training a sequence model by feeding it the real previous step at every step, instead of its own prediction.","explanation":"Teacher forcing is the standard way to train autoregressive sequence models, named by Williams and Zipser in 1989: when predicting step t, the input uses the real first t−1 steps from the data, not the model's own generated output. This lets every step's loss be computed in parallel, so training is fast and stable, and it's how large language models and VLAs that discretize actions into tokens, such as RT-2 and OpenVLA, are trained. The cost is a mismatch between training and inference: at inference the model can only keep conditioning on its own output, so small early errors accumulate, a problem called exposure bias. Scheduled sampling (2015) gradually mixes in the model's own predictions during training; Self Forcing (2025), in video generation, instead unrolls the model autoregressively during training itself.","example":"When training OpenVLA, each action is discretized into 7 tokens, and the k-th token is predicted conditioned on the true first k−1 tokens; at deployment, though, it can only keep decoding after the tokens it just generated itself.","related":["Next-Token Prediction","Autoregressive Decoding","Exposure Bias","Self Forcing","Diffusion Forcing","Compounding Error"]},{"id":"self-forcing","category":"training","sec":2,"tier":3,"sources":[{"title":"Huang et al. 2025: Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion","url":"https://arxiv.org/abs/2506.08009"},{"title":"Self Forcing 项目页","url":"https://self-forcing.github.io/"},{"title":"GitHub: guandeh17/Self-Forcing","url":"https://github.com/guandeh17/Self-Forcing"}],"as_of":"2025-11","related_ids":["teacher-forcing","diffusion-forcing","exposure-bias","autoregressive-video-generation","key-value-cache","diffusion-step-distillation"],"name":"Self Forcing","alt":"自强制","abbr":"","aliases":["Self-Forcing Training"],"one_liner":"Training a video model by having it keep generating from its own previously generated frames, removing the train-test mismatch.","explanation":"Self Forcing was proposed by Adobe Research and UT Austin's Xun Huang and colleagues in June 2025 (NeurIPS 2025 Spotlight), as an autoregressive video-diffusion training method. Autoregressive video models generate a segment at a time; earlier training used teacher forcing, conditioning on real preceding frames, or diffusion forcing, conditioning on noised real preceding frames, but at inference time the model can only condition on its own previously generated, imperfect frames, a mismatch called exposure bias that makes long videos degrade further and further. Self Forcing instead unrolls the model autoregressively during training the same way it will run at inference, using a KV cache and conditioning on its own generated frames, then computes one overall loss across the whole video. Built on Wan2.1-T2V-1.3B, it achieves sub-second-latency, real-time streaming generation on a single GPU, which matters for interactive world models that need to generate video while responding to actions.","example":"According to the project page, the Self Forcing model generates 480p video at about 16 frames per second on a single H100, with a first-frame latency of about 0.8 seconds, and can also stream in real time on a single RTX 4090.","related":["Teacher Forcing","Diffusion Forcing","Exposure Bias","Autoregressive Video Generation","Key-Value Cache","Diffusion Step Distillation"]},{"id":"evidence-lower-bound","category":"training","sec":2,"tier":3,"sources":[{"title":"Auto-Encoding Variational Bayes (arXiv:1312.6114)","url":"https://arxiv.org/abs/1312.6114"},{"title":"Denoising Diffusion Probabilistic Models (arXiv:2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv:2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["variational-autoencoder","conditional-variational-autoencoder","kullback-leibler-divergence","denoising-diffusion-probabilistic-model","maximum-likelihood-estimation","action-chunking-with-transformers"],"name":"Evidence Lower Bound","alt":"证据下界","abbr":"ELBO","aliases":["ELBO","Variational Lower Bound","VLB"],"one_liner":"A tractable lower bound on the log-likelihood of data; VAEs and diffusion models both train by maximizing it.","explanation":"A generative model wants to maximize the data's log-likelihood, log p(x) (also called the “evidence”), but this is usually intractable when the model has latent variables. The ELBO introduces an approximate posterior q(z|x), giving log p(x) ≥ E_q[log p(x|z)] − KL(q(z|x) ‖ p(z)): the first term measures reconstruction quality, and the second pulls the encoding distribution toward the prior, with the gap between the two sides equal to the KL divergence between q and the true posterior. Kingma and Welling's 2013 VAE paper used the reparameterization trick to let the ELBO be optimized directly with gradient descent; Ho and colleagues' 2020 DDPM also starts from a variational lower bound, simplifying and reweighting it into the familiar noise-prediction loss. In robotics, ACT is trained as a conditional VAE.","example":"ACT's training loss is the action-chunk reconstruction error plus β times the KL divergence between the encoding distribution and a standard normal distribution — exactly a β-weighted ELBO.","related":["Variational Autoencoder","Conditional Variational Autoencoder","Kullback-Leibler Divergence","Denoising Diffusion Probabilistic Model","Maximum Likelihood Estimation","Action Chunking with Transformers"]},{"id":"denoising-loss","category":"training","sec":2,"tier":3,"sources":[{"title":"Denoising Diffusion Probabilistic Models (arXiv 2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["diffusion-model","denoising-diffusion-probabilistic-model","noise-schedule","prediction-target-parameterization","flow-matching-loss","diffusion-policy"],"name":"Denoising Loss (Diffusion Loss)","alt":"去噪损失","abbr":"","aliases":["Diffusion Loss","Noise-Prediction Loss","L_simple"],"one_liner":"A diffusion model's training objective: add noise to clean data and have the network predict exactly the noise that was added.","explanation":"Denoising loss is the loss function used to train diffusion models; its most common form comes from Ho, Jain, and Abbeel's 2020 DDPM paper: pick a random noising step t, mix Gaussian noise into a real sample according to a noise schedule, and have the network predict the added noise from the noisy sample and t, with the loss being the mean squared error between predicted and true noise. The paper calls this the simplified objective, L_simple; it's a weighted form of the variational lower bound and corresponds to denoising score matching, and in practice it produces better generation quality than optimizing the full lower bound directly. The network can instead be reparameterized to predict the clean sample x0 or a velocity term v — different prediction-target parameterizations. In robotics, Diffusion Policy treats an action sequence as the data to generate, conditioned on observations, and trains it with the same noise-prediction mean squared error; flow-matching models like π0 use a similarly-shaped flow-matching loss instead.","example":"Training Diffusion Policy: take a snippet of future actions from a demonstration, pick a random noise step and add noise to it; the network looks at the current camera image and the noisy actions and outputs what it thinks the added noise was, and the mean squared error against the true noise is backpropagated.","related":["Diffusion Model","Denoising Diffusion Probabilistic Model","Noise Schedule","Prediction Target Parameterization","Flow Matching Loss","Diffusion Policy"]},{"id":"score-matching","category":"training","sec":2,"tier":3,"sources":[{"title":"Hyvärinen 2005: Estimation of Non-Normalized Statistical Models by Score Matching (JMLR)","url":"https://www.jmlr.org/papers/v6/hyvarinen05a.html"},{"title":"Song & Ermon 2019: Generative Modeling by Estimating Gradients of the Data Distribution","url":"https://arxiv.org/abs/1907.05600"},{"title":"Song et al. 2020: Score-Based Generative Modeling through Stochastic Differential Equations","url":"https://arxiv.org/abs/2011.13456"}],"as_of":"","related_ids":["diffusion-model","denoising-diffusion-probabilistic-model","denoising-loss","flow-matching","energy-based-model","prediction-target-parameterization"],"name":"Score Matching","alt":"分数匹配","abbr":"","aliases":["Score Function","Denoising Score Matching"],"one_liner":"Training a network to estimate the gradient of a data distribution's log-density, without needing its normalizing constant.","explanation":"The score is the gradient of the log probability density with respect to the input, ∇x log p(x), and it points toward where data is more likely to occur. Hyvärinen proposed score matching in 2005: train a model's score to fit the data's score, which sidesteps the normalizing constant that is hard to compute in energy-based models. Later, denoising score matching turned the problem into “add noise to data, then learn to denoise it.” In 2019, Song and Ermon used score estimation under multiple noise levels plus Langevin dynamics sampling for image generation; in 2020, Song and colleagues unified score-based models and diffusion models using stochastic differential equations, where the reverse denoising process depends only on the score at each noise level. DDPM's noise prediction and flow matching's velocity prediction can both be converted into score estimates, so the training behind generative robot policies such as Diffusion Policy is, underneath, score matching.","example":"Diffusion Policy trains by adding Gaussian noise of varying strength to expert actions and having the network predict the noise that was added; dividing that prediction by the noise's standard deviation and negating it gives a score estimate for the noisy-action distribution.","related":["Diffusion Model","Denoising Diffusion Probabilistic Model","Denoising Loss (Diffusion Loss)","Flow Matching","Energy-Based Model","Prediction Target Parameterization"]},{"id":"flow-matching-loss","category":"training","sec":2,"tier":3,"sources":[{"title":"Flow Matching for Generative Modeling (arXiv:2210.02747)","url":"https://arxiv.org/abs/2210.02747"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["flow-matching","velocity-field","rectified-flow","mean-squared-error","denoising-loss","pi0"],"name":"Flow Matching Loss","alt":"流匹配损失","abbr":"","aliases":["Conditional Flow Matching Loss","CFM Loss","Velocity Field Regression Loss"],"one_liner":"Training a network with mean squared error to predict the velocity that points from noise toward the real data.","explanation":"Flow matching was introduced by Lipman and colleagues in 2022, treating “noise gradually turning into data” as continuous flow along a path, with the network learning the velocity field at every point along that path. Training doesn't need to simulate the whole trajectory: pick a random time τ, linearly interpolate between a real sample and Gaussian noise to get an intermediate point, and use mean squared error to have the network's predicted velocity match the direction of that straight-line path (the real sample minus the noise) — that's the flow matching loss. It looks a lot like a diffusion model's denoising loss in form, but the path is straighter, so fewer sampling steps are needed. At inference, starting from pure noise, a few steps of Euler integration along the predicted velocity produce a sample. The π0 family of VLAs uses it to train their action expert, generating continuous action chunks directly.","example":"π0 samples τ during training from a Beta distribution skewed toward the high-noise end, forms the noisy action τ·A + (1 − τ)·ε, and has the action expert regress toward A − ε; at inference it starts from pure noise and generates an action chunk with 10 steps of Euler integration.","related":["Flow Matching","Velocity Field","Rectified Flow","Mean Squared Error","Denoising Loss (Diffusion Loss)","π0"]},{"id":"auxiliary-loss-auxiliary-task","category":"training","sec":2,"tier":3,"sources":[{"title":"Reinforcement Learning with Unsupervised Auxiliary Tasks (UNREAL, arXiv 1611.05397)","url":"https://arxiv.org/abs/1611.05397"},{"title":"Learning to Navigate in Complex Environments (arXiv 1611.03673)","url":"https://arxiv.org/abs/1611.03673"}],"as_of":"","related_ids":["loss-function","representation-learning","multi-task-learning","reinforcement-learning","sparse-reward","backbone-network"],"name":"Auxiliary Loss / Auxiliary Task","alt":"辅助损失 / 辅助任务","abbr":"","aliases":["Auxiliary Objective"],"one_liner":"An extra prediction task and loss term added alongside the main training objective, to help the model learn better features.","explanation":"An auxiliary task means having the same network make additional, related predictions beyond the main training objective (such as outputting an action or maximizing reward); each prediction has its own auxiliary loss, weighted and added into the total loss to optimize together. These tasks share the backbone (the part of the network responsible for extracting features) with the main task, and the extra supervisory signal helps features get learned faster and more stably; the extra branches are usually simply discarded at inference. In reinforcement learning, where rewards are often sparse and the learning signal weak, auxiliary tasks are especially useful: DeepMind's UNREAL agent (2016) added pixel-control, reward-prediction, and value-replay auxiliary tasks, learning roughly 10 times faster in the 3D maze environment Labyrinth. Robot imitation learning and VLAs also commonly have the model predict future images, object locations, or sub-task text on the side. If an auxiliary loss is weighted too heavily, it competes with the main task for model capacity, so the weight needs tuning.","example":"DeepMind's 2016 navigation agent, while learning to find goals in a 3D maze, also had the same network predict scene depth and judge whether it had returned to a previously visited location (loop-closure classification); these two auxiliary losses noticeably improved navigation performance.","related":["Loss Function","Representation Learning","Multi-Task Learning","Reinforcement Learning","Sparse Reward","Backbone Network"]},{"id":"behavior-cloning","category":"training","sec":3,"tier":1,"sources":[{"title":"Wikipedia: Imitation learning","url":"https://en.wikipedia.org/wiki/Imitation_learning"},{"title":"Pomerleau 1988: ALVINN: An Autonomous Land Vehicle in a Neural Network (NeurIPS)","url":"https://proceedings.neurips.cc/paper/1988/hash/812b4ba287f5ee0bc9d43bbf5bbe87fb-Abstract.html"},{"title":"Zhao et al. 2023: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["imitation-learning","compounding-error","dagger","demonstration-data","action-chunking-with-transformers","diffusion-policy"],"name":"Behavior Cloning","alt":"行为克隆","abbr":"BC","aliases":["Behavioral Cloning","Behavioural Cloning","BC"],"one_liner":"Treating expert demonstrations as labeled data and using supervised learning to directly map observations to actions.","explanation":"Behavior cloning is the most basic form of imitation learning. You collect a large set of observation-action pairs — for example, the camera images and joint commands recorded while a human teleoperates a robot — treat the observations as inputs and the expert's actions as labels, and train a policy network with ordinary supervised learning to fit them. An early example is Pomerleau's 1988 ALVINN, which used a neural network to output a steering direction directly from images. Behavior cloning is simple and stable, and needs no reward function to be designed, so the core training of ACT, Diffusion Policy, and most VLA (vision-language-action) models is essentially behavior cloning. Its main weakness is compounding error: once the policy drifts even slightly off the demonstrated trajectory, it enters states it never saw during training, and small mistakes snowball from there. DAgger, correction data, and reinforcement-learning fine-tuning are all ways of addressing this weakness.","example":"ACT used behavior cloning on just about 10 minutes — 50 human teleoperated demonstrations — to get a low-cost bimanual robot to succeed 80–90% of the time at fine-grained tasks like opening the lid of a translucent condiment cup or putting batteries into a remote control.","related":["Imitation Learning","Compounding Error","DAgger","Demonstration Data","Action Chunking with Transformers","Diffusion Policy"]},{"id":"compounding-error","category":"training","sec":3,"tier":2,"sources":[{"title":"Efficient Reductions for Imitation Learning (Ross & Bagnell, AISTATS 2010)","url":"https://proceedings.mlr.press/v9/ross10a.html"},{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (AISTATS 2011)","url":"https://proceedings.mlr.press/v15/ross11a/ross11a.pdf"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["distribution-shift","dagger","behavior-cloning","action-chunking","recovery-and-correction-data","exposure-bias"],"name":"Compounding Error","alt":"复合误差","abbr":"","aliases":["Error Compounding","Error Accumulation","Compounding Errors"],"one_liner":"Small per-step deviations in an imitation-learned policy add up, pushing the robot into unfamiliar states where it fails.","explanation":"Compounding error is the central problem with behavior cloning (imitating an expert's actions directly via supervised learning). The policy makes a small error at every step, and once it drifts off the expert's trajectory, it ends up in states that never appeared in the training data, where it's even more likely to err — the deviation snowballs. Ross and Bagnell's 2010 analysis showed that with a per-step error rate of ε over a task of length T, pure supervised imitation can incur cost as bad as order T²ε relative to the expert, growing quadratically with task length. The underlying cause is covariate shift. Common mitigations include DAgger, which has the expert label the states the policy actually wanders into; deliberately collecting recovery data for correcting mistakes; and action chunking, as in ACT, which predicts a whole block of actions at once and reduces the number of decision points.","example":"The ACT paper notes that in high-precision bimanual fine manipulation, an imitation-learned policy's error accumulates over time, so instead it predicts an entire action sequence at once, reaching 80–90% success on several tasks using just 10 minutes of human demonstrations.","related":["Distribution Shift","DAgger","Behavior Cloning","Action Chunking","Recovery and Correction Data","Exposure Bias"]},{"id":"dagger","category":"training","sec":3,"tier":2,"sources":[{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (PMLR v15)","url":"https://proceedings.mlr.press/v15/ross11a.html"},{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (PDF, 含 DAgger 算法与 T²ε 分析)","url":"https://proceedings.mlr.press/v15/ross11a/ross11a.pdf"}],"as_of":"","related_ids":["compounding-error","distribution-shift","behavior-cloning","human-gated-dagger","interactive-imitation-learning","human-in-the-loop"],"name":"DAgger","alt":"DAgger（数据集聚合）","abbr":"DAgger","aliases":["Dataset Aggregation"],"one_liner":"Letting the policy run itself, having an expert label the correct action at the states it actually reaches, then retraining on the combined data.","explanation":"DAgger (Dataset Aggregation) is an interactive imitation-learning algorithm introduced by Ross, Gordon, and Bagnell at AISTATS 2011. It targets behavior cloning's compounding-error problem: a policy trained only on an expert's trajectory has no idea what to do once it drifts off that path. The procedure iterates: train an initial policy on expert data; let the current policy run and record the states it actually visits; ask the expert to label what should happen at those states; merge the new labels into the growing dataset and retrain; repeat. The paper proves this reduces the growth of error with task length from quadratic to close to linear. The downside is that the expert must be available to label on demand; on real robots, variants like HG-DAgger are common, where a human supervises and takes over when needed.","example":"The original paper learned steering in the racing game Super Tux Kart and level completion in Super Mario Bros., and DAgger outperformed plain supervised learning on expert demonstrations alone in both.","related":["Compounding Error","Distribution Shift","Behavior Cloning","Human-Gated DAgger","Interactive Imitation Learning","Human-in-the-Loop"]},{"id":"causal-confusion","category":"training","sec":3,"tier":3,"sources":[{"title":"Causal Confusion in Imitation Learning (arXiv 1905.11979)","url":"https://arxiv.org/abs/1905.11979"}],"as_of":"","related_ids":["behavior-cloning","imitation-learning","distribution-shift","compounding-error","dagger","overfitting"],"name":"Causal Confusion (Causal Misidentification)","alt":"因果混淆","abbr":"","aliases":["Shortcut Learning (in Imitation Learning)","Causal Misidentification"],"one_liner":"Imitation learning latching onto a cue correlated with the expert's actions but not actually its cause, then failing at deployment.","explanation":"Causal confusion was systematically identified by de Haan, Jayaraman, and Levine in a NeurIPS 2019 paper. Behavior cloning treats imitation as supervised learning, regressing directly from observation to the expert's action, learning only correlation without distinguishing cause from effect; if some cue in the training data happens to correlate strongly with the expert's actions, the model will latch onto it. That cue is present the whole time during training, so the loss stays low; at deployment, the states the policy wanders into differ from the expert's (distribution shift), the cue stops being reliable, and the policy fails. A counterintuitive symptom: giving the model more input information can actually make performance worse, especially common when the input includes history. The paper proposes targeted interventions — trying an action in the environment or asking the expert — to recover the correct causal structure, outperforming baselines like DAgger. It's a reminder that more input isn't automatically better for a robot policy.","example":"The paper's driving example: model A sees the full dashboard view, model B has the dashboard masked out. Both reach low training loss, but on the road B drives well and A doesn't — the dashboard has a brake-light indicator that lights up whenever the brake is pressed, and A learned “brake when the light is on,” mistaking an effect of braking for its cause.","related":["Behavior Cloning","Imitation Learning","Distribution Shift","Compounding Error","DAgger","Overfitting"]},{"id":"action-state-normalization","category":"training","sec":3,"tier":3,"sources":[{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)","url":"https://arxiv.org/abs/2501.09747"},{"title":"openpi（Physical Intelligence GitHub）","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["action-tokenizer","pi0-fast","proprioception","cross-embodiment-data","openpi","fine-tuning"],"name":"Action / State Normalization","alt":"动作与状态归一化","abbr":"","aliases":["Norm Stats","Quantile Normalization"],"one_liner":"Scaling each action and state dimension to a common range before training, then converting model output back to real units at inference.","explanation":"A robot's actions and proprioceptive state (joint angles, gripper opening, end-effector position, and so on) span very different units and scales across dimensions, and feeding them into a network unscaled lets the largest-magnitude dimensions dominate the loss and destabilize training. A common fix is to compute, from the training data, each dimension's mean and standard deviation, or its 1st and 99th percentiles (quantile normalization), and rescale the data to zero mean and unit variance, or to the range [-1, 1]. These statistics are called norm stats and must be saved alongside the model weights, since inference needs to de-normalize the output back into real commands; percentiles are more robust to outliers than a raw min/max. Physical Intelligence's FAST paper, for instance, normalizes actions using the 1st/99th percentiles; openpi requires running compute_norm_stats.py before fine-tuning, though a new task on a robot already present in the pretraining data can reuse the pretrained statistics. Mismatched normalization statistics are a common cause of erratic real-robot behavior.","example":"π0-FAST maps each action dimension's 1st and 99th training-set percentiles onto [-1, 1] before tokenizing actions, so data from different robots can share the same tokenizer.","related":["Action Tokenizer","π0-FAST","Proprioception","Cross-Embodiment Data","openpi (Physical Intelligence)","Fine-tuning"]},{"id":"goal-conditioned-behavior-cloning","category":"training","sec":3,"tier":3,"sources":[{"title":"Learning Latent Plans from Play（项目页，含 Play-GCBC 基线）","url":"https://learning-from-play.github.io/"},{"title":"Learning to Reach Goals via Iterated Supervised Learning (GCSL, arXiv 1912.06088)","url":"https://arxiv.org/abs/1912.06088"}],"as_of":"","related_ids":["behavior-cloning","goal-conditioned-policy","hindsight-relabeling","play-data","goal-conditioned-reinforcement-learning","action-multimodality"],"name":"Goal-Conditioned Behavior Cloning","alt":"目标条件模仿学习","abbr":"GCBC","aliases":["GCBC","Goal-Conditioned Imitation Learning"],"one_liner":"Behavior cloning that feeds the policy both the current observation and the goal it's meant to reach.","explanation":"This builds on ordinary behavior cloning (copying demonstrated actions via supervised learning) by giving the policy one more input: a goal, usually a goal image or goal state, though it can also be language. Training data is often produced through hindsight relabeling: cut a segment out of a longer demonstration and treat its last frame as the goal, so the actions leading up to it automatically become a demonstration of reaching that goal — which means even task-unlabeled “play” data can be used. Google's Lynch and colleagues used it as the baseline Play-GCBC in their 2019 Learning from Play paper; GCSL, from the same year, has the agent collect its own data and repeatedly relabel it before running supervised learning. It's structurally simple and the starting point for many goal-conditioned and language-conditioned policies, but its weak point is that when a single goal has multiple valid routes to it (action multimodality), it tends to learn the average of them instead of any one of them.","example":"Randomly cutting a short segment out of teleoperated play data and using its last frame as the goal, a policy is trained to output an action given the current image and the goal image.","related":["Behavior Cloning","Goal-conditioned Policy","Hindsight Relabeling","Play Data","Goal-Conditioned Reinforcement Learning","Action Multimodality"]},{"id":"imitation-from-observation","category":"training","sec":3,"tier":3,"sources":[{"title":"Recent Advances in Imitation Learning from Observation (IJCAI 2019 survey)","url":"https://arxiv.org/abs/1905.13566"},{"title":"Behavioral Cloning from Observation (IJCAI 2018)","url":"https://arxiv.org/abs/1805.01954"},{"title":"Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation (ICRA 2018)","url":"https://arxiv.org/abs/1707.03374"}],"as_of":"","related_ids":["action-free-video","inverse-dynamics-model","human-video-data","latent-action","pseudo-action-labels","lapa"],"name":"Imitation from Observation","alt":"从观测中模仿学习","abbr":"IfO","aliases":["IfO","Imitation Learning from Observation","Learning from Video"],"one_liner":"Learning to imitate from only the demonstrator's states or video, with no recorded action labels at all.","explanation":"Ordinary imitation learning needs observation-action pairs; imitation from observation gets only a sequence of the demonstrator's states — video of a person doing something, say — with no idea what control command produced each step. This opens the door to using internet video and other huge collections of human video, but it also has to deal with differing viewpoints and body structures. Notable examples: UC Berkeley's Liu and colleagues (2017) used context translation to convert human video into a robot's viewpoint before running reinforcement learning; UT Austin's Torabi and colleagues (2018) proposed BCO, which first lets the agent explore on its own to learn an inverse dynamics model (inferring the action from a pair of consecutive frames), then uses it to fill in action labels for expert video so behavior cloning can be applied; adversarial approaches exist too. Today's embodied-AI work that pretrains on human video, or uses latent actions or pseudo-action labels, is solving exactly this same problem.","example":"In Liu and colleagues' 2017 paper, a robot watches only video of a person sweeping, scooping almonds, and pushing objects — no joint recordings at all — and learns to perform the same actions with tools.","related":["Action-free Video","Inverse Dynamics Model","Human Video Data","Latent Action","Pseudo Action Labels","LAPA"]},{"id":"markov-decision-process","category":"training","sec":4,"tier":2,"sources":[{"title":"Markov decision process（Wikipedia）","url":"https://en.wikipedia.org/wiki/Markov_decision_process"}],"as_of":"","related_ids":["reinforcement-learning","partially-observable-markov-decision-process","bellman-equation","reward-function","discount-factor","policy"],"name":"Markov Decision Process","alt":"马尔可夫决策过程","abbr":"MDP","aliases":["MDP"],"one_liner":"The standard mathematical framework for an agent's loop of seeing a state, acting, getting a reward, and moving to a new state.","explanation":"The Markov decision process is reinforcement learning's standard mathematical model, usually written as the tuple (S, A, P, R, γ): the state space, the action space, the state-transition probability (the chance of reaching a given next state after taking action a in state s), the reward function, and a discount factor that discounts distant rewards. Its core assumption is the Markov property: the next state depends only on the current state and action, not on earlier history. Its mathematical foundation comes from Richard Bellman's dynamic-programming work around 1957, and the Bellman equation, value functions, and policy gradients are all built on top of it. Reinforcement learning's goal is to find a policy within an MDP that maximizes expected cumulative discounted return. Real robots often can't observe the full state (an object might be occluded), which calls for a partially observable Markov decision process (POMDP) instead; in practice this is often handled by feeding in a history of past observations.","example":"When training a quadruped to walk, the state can include joint angles, joint velocities, body orientation, and the commanded velocity; the action is the 12 joints' target angles; the reward rewards tracking the commanded velocity and penalizes energy use and falling; and the policy outputs one action per control cycle based on the current state.","related":["Reinforcement Learning","Partially Observable Markov Decision Process","Bellman Equation","Reward Function","Discount Factor","Policy"]},{"id":"reward-function","category":"training","sec":4,"tier":1,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Eureka: Human-Level Reward Design via Coding Large Language Models (arXiv 2310.12931)","url":"https://arxiv.org/abs/2310.12931"}],"as_of":"","related_ids":["reinforcement-learning","return","reward-shaping","sparse-reward","reward-hacking","eureka"],"name":"Reward Function","alt":"奖励函数","abbr":"","aliases":["Reward","Reward Signal"],"one_liner":"The scoring rule in reinforcement learning that rates an agent's behavior at each step, defining what counts as doing well.","explanation":"The reward function is a core piece of reinforcement learning (RL): after every action the agent takes, the environment returns a scalar score based on the current state, the action, and the resulting next state, typically written r = R(s, a, s′). The agent's objective is to maximize the sum of these rewards over time (the return), so the reward function effectively defines what the task is. For robot tasks, reward functions are usually hand-written, which is harder than it sounds. Giving credit only on full success (a sparse reward) makes learning slow, but a carelessly designed reward is easy for a policy to exploit — it drives the score up without actually solving the task, a failure mode called reward hacking. This has led to techniques like reward shaping (adding extra intermediate rewards to guide learning), learned reward models, and even having a large language model write reward code automatically.","example":"A quadruped locomotion reward is usually a weighted sum of terms: a bonus for tracking the target velocity, a penalty for excessive joint torque, and a penalty for falling. NVIDIA and collaborators' Eureka has GPT-4 write reward-function code directly, and it beat reward functions written by human experts on 83% of 29 tasks.","related":["Reinforcement Learning","Return","Reward Shaping","Sparse Reward","Reward Hacking","Eureka"]},{"id":"return","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Hugging Face Deep RL Course: The Reinforcement Learning Framework","url":"https://huggingface.co/learn/deep-rl-course/unit1/rl-framework"}],"as_of":"","related_ids":["reward-function","discount-factor","value-function","q-function","episode","return-conditioning"],"name":"Return","alt":"回报","abbr":"","aliases":["Cumulative Reward","Discounted Return","G_t"],"one_liner":"The sum of all rewards from a given moment onward, usually discounted so that later rewards count for less.","explanation":"Return is the quantity reinforcement learning aims to maximize. A reward is the immediate score the environment gives at each step; the return is the total of all rewards from the current moment onward, often written G_t or R(τ). When an episode has a fixed length, these can just be added up directly; for long or never-ending tasks, a discounted return is normally used instead: a reward k steps in the future is multiplied by the discount factor γ raised to the k-th power (γ is between 0 and 1, commonly 0.95–0.99), so more distant rewards count for less — this both keeps the sum finite and makes the agent weigh near-term outcomes more. Reinforcement learning's objective is to maximize expected return; the value function and Q-function are exactly estimates of expected future return. Methods like return conditioning even feed a target return in as an input, so the policy acts at a specified performance level.","example":"A grasping task gives +1 only on success and 0 otherwise, with γ = 0.99: succeeding at step 10 gives a discounted return from the start of about 0.99^10 ≈ 0.904, while dragging it out to step 50 gives only about 0.605 — so the policy is pushed toward finishing faster.","related":["Reward Function","Discount Factor","Value Function","Q-Function","Episode","Return Conditioning"]},{"id":"discount-factor","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"legged_gym: legged_robot_config.py (PPO gamma = 0.99)","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["return","reward-function","value-function","bellman-equation","generalized-advantage-estimation","markov-decision-process"],"name":"Discount Factor","alt":"折扣因子","abbr":"γ","aliases":["γ (gamma)","Discount Rate"],"one_liner":"The coefficient in reinforcement learning that discounts future rewards, controlling how much the agent weighs long-term payoff.","explanation":"The discount factor is a basic reinforcement-learning hyperparameter, written γ, with a value between 0 and 1. When computing the return (the sum of rewards accumulated from the current moment onward), a reward k steps in the future gets multiplied by γ raised to the k-th power. It serves two purposes: expressing that a reward received sooner is more certain and more valuable than one received later, and turning an infinitely long sum of rewards into a finite, well-behaved value, which makes it solvable with the Bellman equation (the recursive relationship that defines a value function). A γ closer to 1 makes the agent weigh long-term outcomes more heavily, but also makes the value estimate harder to learn and noisier; a smaller γ makes the agent more shortsighted. Robot reinforcement learning commonly uses a γ around 0.99. It appears in the definitions of return, the value function, and the advantage function.","example":"legged_gym (ETH's open-source reinforcement-learning framework for training legged robots) sets gamma = 0.99 in its PPO config: a reward 100 steps away is weighted by roughly 0.99 to the 100th power, about 0.37.","related":["Return","Reward Function","Value Function","Bellman Equation","Generalized Advantage Estimation","Markov Decision Process"]},{"id":"exploration-vs-exploitation","category":"training","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Exploration–exploitation dilemma","url":"https://en.wikipedia.org/wiki/Exploration%E2%80%93exploitation_dilemma"},{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"","related_ids":["intrinsic-motivation","entropy-regularization","reinforcement-learning","sample-efficiency","safe-reinforcement-learning","real-world-reinforcement-learning"],"name":"Exploration vs. Exploitation","alt":"探索与利用","abbr":"","aliases":["Exploration-Exploitation Trade-off","Exploration-Exploitation Dilemma"],"one_liner":"The trade-off between trying new actions to find something better and using the best-known action to collect reward now.","explanation":"This is a fundamental tension in reinforcement learning and sequential decision-making. Exploitation means picking the action that currently looks best given what the agent already knows; exploration means trying an uncertain, unfamiliar action, which may pay off worse in the short term but could reveal a better strategy. Pure exploitation easily gets stuck on a suboptimal solution, while pure exploration never cashes in on the reward it finds. The classic framework for studying this is the multi-armed bandit problem; common strategies include ε-greedy (pick a random action with small probability ε), UCB (upper confidence bound, which favors options with more uncertainty), and Thompson sampling. Deep reinforcement learning also uses entropy regularization (encouraging the policy to stay somewhat random) and intrinsic reward (rewarding “curiosity”). Exploration is harder on real robots: unconstrained trial and error can damage the hardware, so policies are often first brought into a reasonable region using demonstration data or simulation, and only then explore in a controlled way.","example":"When SimpleVLA-RL fine-tunes a VLA with reinforcement learning, it raises the rollout sampling temperature from 1.0 to 1.6 and widens the clipping range from 0.2 to 0.28, so the model generates more varied trajectories and explores more.","related":["Intrinsic Motivation","Entropy Regularization","Reinforcement Learning","Sample Efficiency","Safe Reinforcement Learning","Real-World Reinforcement Learning"]},{"id":"credit-assignment","category":"training","sec":4,"tier":3,"sources":[{"title":"A Survey of Temporal Credit Assignment in Deep Reinforcement Learning (arXiv 2312.01072)","url":"https://arxiv.org/abs/2312.01072"},{"title":"Minsky (1961): Steps Toward Artificial Intelligence (Proceedings of the IRE)","url":"https://courses.csail.mit.edu/6.803/pdf/steps.pdf"}],"as_of":"","related_ids":["sparse-reward","temporal-difference-learning","advantage-function","generalized-advantage-estimation","reward-shaping","long-horizon-task"],"name":"Credit Assignment","alt":"信用分配","abbr":"","aliases":["Credit Assignment Problem","CAP"],"one_liner":"Figuring out, once a final reward or penalty arrives, which of the earlier actions deserve the credit or the blame.","explanation":"The credit assignment problem was already discussed specifically in Marvin Minsky's 1961 survey “Steps Toward Artificial Intelligence”: once a complex strategy succeeds, how should the credit be divided among the many decisions involved? His example was that winning a game of chess might involve a million decisions, and asked whether each one could just get an equal millionth of the credit. In reinforcement learning it mainly refers to credit assignment across time: reward is often delayed until a task finishes, mixed in with noise and chance along the way, and the agent has to learn each action's actual contribution to the final outcome from limited experience. Common tools include temporal-difference learning with bootstrapping, the discount factor, the advantage function and Generalized Advantage Estimation (GAE), reward shaping (adding intermediate reward), and hierarchical reinforcement learning. Robot long-horizon tasks have many action steps and sparse reward, which makes credit assignment one of the main reasons reinforcement learning is hard to apply there.","example":"A robot arm stacks 5 blocks and gets 1 point only if all 5 end up stacked correctly. If one episode fails, the cause might be block 2 placed crooked, or block 4 released too early; credit assignment has to learn how much blame each step's action deserves from a large number of episodes like this.","related":["Sparse Reward","Temporal-Difference Learning","Advantage Function","Generalized Advantage Estimation","Reward Shaping","Long-horizon Task"]},{"id":"value-function","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL（Value Functions）","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Hugging Face Deep RL Course: Advantage Actor-Critic (A2C)","url":"https://huggingface.co/learn/deep-rl-course/unit6/advantage-actor-critic"},{"title":"π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"","related_ids":["q-function","advantage-function","bellman-equation","temporal-difference-learning","return","asymmetric-actor-critic"],"name":"Value Function","alt":"价值函数","abbr":"","aliases":["State-Value Function","V-Function","Critic"],"one_liner":"A function estimating how much total reward will follow from a given state if the agent keeps following a given policy.","explanation":"The value function is a core concept in reinforcement learning. The state-value function V(s) is the expected return — the discounted sum of future rewards — from state s onward while following policy π the whole time; the Q-function Q(s,a) is the expected return from taking action a first and then following the policy; A = Q − V is called the advantage function, measuring how much better a given action is than average. With it, an agent can judge which states and actions are worth pursuing. The actor-critic architecture combines a policy with a value function: the actor is the policy network, responsible for producing actions, and the critic is the value network, scoring actions and supplying the advantage signal used to update the actor — PPO, SAC, and TD3 all belong to this family. In embodied AI, value functions are also used to filter data and to run reinforcement learning on VLAs, as in Physical Intelligence's RECAP method for π*0.6.","example":"π*0.6 trains a value function that predicts “how many steps remain until the task succeeds” (a failed trajectory gets a very low value), and uses it to compute each action's advantage, telling the policy which actions are better.","related":["Q-Function","Advantage Function","Bellman Equation","Temporal-Difference Learning","Return","Asymmetric Actor-Critic"]},{"id":"bellman-equation","category":"training","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Bellman equation","url":"https://en.wikipedia.org/wiki/Bellman_equation"},{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"}],"as_of":"","related_ids":["value-function","q-function","discount-factor","temporal-difference-learning","q-learning","markov-decision-process"],"name":"Bellman Equation","alt":"贝尔曼方程","abbr":"","aliases":["Bellman Optimality Equation","Bellman Expectation Equation"],"one_liner":"A recursive formula that breaks a state's value into the immediate reward plus the discounted value of the next state.","explanation":"The Bellman equation is named after the American mathematician Richard Bellman, and it comes from the dynamic programming method he introduced. It states that the value of a state equals the immediate reward earned there plus a discount factor times the expected value of the next state. This turns the hard problem of “how good is this in the long run” into a step-by-step recursion. The version written for a fixed policy is called the Bellman expectation equation; the version that takes the maximum over actions is the Bellman optimality equation. Nearly all value-based reinforcement learning builds on this idea — Q-learning, Deep Q-Networks (DQN), and temporal-difference learning all use it to construct their training targets, pushing the network's predictions to satisfy this self-consistent relationship between a state and what follows it.","example":"Q-learning's training target is r + γ·max Q(s′, a′): when a robot takes action a in state s, gets reward r, and lands in state s′, it updates Q(s, a) toward that target — which is exactly the Bellman optimality equation in use.","related":["Value Function","Q-Function","Discount Factor","Temporal-Difference Learning","Q-Learning","Markov Decision Process"]},{"id":"q-function","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Hugging Face Deep RL Course: Introducing Q-Learning","url":"https://huggingface.co/learn/deep-rl-course/unit2/q-learning"}],"as_of":"","related_ids":["value-function","advantage-function","bellman-equation","q-learning","deep-q-network","return"],"name":"Q-Function","alt":"Q 函数","abbr":"","aliases":["Action-Value Function","Q-value","Q(s,a)"],"one_liner":"The expected return of taking a specific action in a state and then following the policy from then on.","explanation":"The Q-function is written Q(s,a): the expected return — the discounted sum of future rewards — from taking action a in state s and then following policy π forever after. Compared with the value function V(s), it takes an extra action input, so it can directly compare how good different actions are in the same state; once the optimal Q-function is known, picking the action with the highest Q-value at every step is the optimal policy. The Q-function satisfies the Bellman equation: the current Q-value equals the immediate reward plus the discounted Q-value of what follows. Q-learning, DQN, SAC, and TD3 all learn it; the critic in an actor-critic algorithm is often literally a Q-network, and the advantage function A(s,a) = Q(s,a) − V(s) is derived from it too.","example":"QT-Opt uses a convolutional network to estimate Q-values: it takes the current camera image and a candidate gripper action as input and outputs the probability that this action ultimately leads to a successful grasp; at every step it searches for the action with the highest Q-value using a cross-entropy method.","related":["Value Function","Advantage Function","Bellman Equation","Q-Learning","Deep Q-Network","Return"]},{"id":"advantage-function","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"High-Dimensional Continuous Control Using Generalized Advantage Estimation (arXiv 1506.02438)","url":"https://arxiv.org/abs/1506.02438"}],"as_of":"","related_ids":["value-function","q-function","generalized-advantage-estimation","policy-gradient","proximal-policy-optimization","recap"],"name":"Advantage Function","alt":"优势函数","abbr":"","aliases":["Advantage","A(s,a)"],"one_liner":"A measure of how much better an action is than the policy's average, equal to the Q-value minus the value function.","explanation":"The advantage function is defined as A(s,a) = Q(s,a) − V(s). Q is the expected return of taking action a in state s and then following the policy afterward; V is the expected return of just following the policy from state s directly. Subtracting the two tells you how much better than average that specific action is: positive means do more of it, negative means do less. Policy-gradient methods use the advantage in place of the raw return because it substantially reduces the variance (noisiness) of the gradient estimate, which makes training more stable. Algorithms like PPO typically compute it with Generalized Advantage Estimation (GAE), introduced by Schulman and colleagues in 2015; Physical Intelligence's RECAP instead feeds the advantage into a VLA as a conditioning input, letting the model distinguish good experience from bad.","example":"When an arm reaches for a cup in state s, the current policy earns an average return of 0.6. If leading with a straight vertical approach raises the expected return to 0.8, that action's advantage is +0.2, and training will increase the probability of picking it.","related":["Value Function","Q-Function","Generalized Advantage Estimation","Policy Gradient","Proximal Policy Optimization","RECAP"]},{"id":"policy-iteration-value-iteration","category":"training","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Markov decision process（Algorithms 一节）","url":"https://en.wikipedia.org/wiki/Markov_decision_process"},{"title":"Wikipedia: Bellman equation","url":"https://en.wikipedia.org/wiki/Bellman_equation"}],"as_of":"","related_ids":["markov-decision-process","bellman-equation","value-function","q-learning","model-based-reinforcement-learning","policy-gradient"],"name":"Policy Iteration / Value Iteration","alt":"策略迭代 / 价值迭代","abbr":"","aliases":["Dynamic Programming (RL)","Value Iteration"],"one_liner":"Two classic dynamic-programming algorithms that repeatedly update values or the policy to solve for an optimal policy, given a known model.","explanation":"Policy iteration and value iteration are two classic dynamic-programming algorithms for solving a Markov decision process, the mathematical framework describing states, actions, rewards, and transitions, assuming both the transition probabilities and the reward function are already known. Value iteration, introduced by Bellman in 1957, repeatedly applies the Bellman optimality equation to update every state's value, and once it converges, picks actions by value. Policy iteration, introduced by Howard in 1960, alternates two steps: policy evaluation, computing every state's value under the current policy, and policy improvement, switching each state to its highest-value action, until the policy stops changing. Real robots have continuous states and an unknown model, so these cannot be applied directly, but algorithms such as Q-learning and actor-critic methods can be viewed as sampled, function-approximated variants of them.","example":"In a 4×4 grid maze where every step costs a reward of −1, the episode ends at the goal, and movement is deterministic, value iteration repeatedly updates each cell's value; once it converges, each cell's value equals the negative of its shortest distance to the goal, and moving toward higher value traces out the shortest path.","related":["Markov Decision Process","Bellman Equation","Value Function","Q-Learning","Model-Based Reinforcement Learning","Policy Gradient"]},{"id":"monte-carlo-methods","category":"training","sec":4,"tier":3,"sources":[{"title":"Wikipedia: Reinforcement learning (Monte Carlo methods)","url":"https://en.wikipedia.org/wiki/Reinforcement_learning"},{"title":"Physical Intelligence 2025: π*0.6: a VLA That Learns From Experience (RECAP)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"","related_ids":["return","temporal-difference-learning","value-function","discount-factor","monte-carlo-tree-search","recap"],"name":"Monte Carlo Methods","alt":"蒙特卡洛方法（蒙特卡洛回报）","abbr":"MC","aliases":["MC","Monte Carlo Return"],"one_liner":"Estimating a value by averaging over many random samples; in RL, using a whole episode's actual return to estimate value.","explanation":"Monte Carlo methods broadly refers to any technique that approximates a computation through random sampling and averaging. In reinforcement learning it refers specifically to this: let the agent run a full episode to completion, sum up the actual (discounted) rewards received from some step onward to get the “Monte Carlo return,” then average that over many episodes to estimate the value of a state or action. It does not rely on any estimate of later states' value, so it avoids the bias that bootstrapping introduces, but variance grows with episode length, and it can only update once an episode ends, so it only applies to tasks that actually terminate. Its counterpart is temporal difference learning, which can update at every single step; TD(λ) lets you tune continuously between the two. Monte Carlo tree search borrows the same idea of random simulation.","example":"π*0.6's RECAP training uses the Monte Carlo return directly for its value function: each step costs a reward of −1, a successful finish scores 0, and a failure subtracts a large penalty, so the value roughly equals the negative of how many steps remain to complete the task.","related":["Return","Temporal-Difference Learning","Value Function","Discount Factor","Monte Carlo Tree Search","RECAP"]},{"id":"temporal-difference-learning","category":"training","sec":4,"tier":2,"sources":[{"title":"Temporal difference learning (Wikipedia)","url":"https://en.wikipedia.org/wiki/Temporal_difference_learning"},{"title":"Soft Actor-Critic (OpenAI Spinning Up)","url":"https://spinningup.openai.com/en/latest/algorithms/sac.html"}],"as_of":"","related_ids":["value-function","bellman-equation","bootstrapping","q-learning","monte-carlo-methods","generalized-advantage-estimation"],"name":"Temporal-Difference Learning","alt":"时序差分学习","abbr":"TD","aliases":["TD Learning","TD Error"],"one_liner":"Updating a value estimate using “this step's reward plus the estimated value of the next state,” without waiting for the episode to end.","explanation":"Temporal-difference (TD) learning was systematically introduced by Richard Sutton in a 1988 paper and is the core method for estimating a value function (how much total return follows from a given state or action) in reinforcement learning. It doesn't need to wait for an episode to finish to get the true return; instead it updates after every single step, treating “the reward r actually received, plus the discounted estimated value of the next state, γV(s′)” as a target, calling the gap between that target and the current estimate V(s) the TD error, and nudging V(s) toward the target by that amount. This “using an estimate to update an estimate” approach is called bootstrapping, and it has lower variance than Monte Carlo methods, which wait for the full episode, letting it learn while still acting — at the cost of introducing some bias. Q-learning, DQN, and the critic networks in SAC and TD3 are all trained with TD targets; the early landmark system TD-Gammon used it to reach expert-level backgammon play.","example":"A robot arm takes one step and gets a reward of 0; the critic estimates the next state's value at 0.8, and with discount factor γ = 0.99, the TD target is 0 + 0.99 × 0.8 = 0.792. If the current state's estimated value is 0.5, the TD error is 0.292, and the network nudges its estimate toward 0.792.","related":["Value Function","Bellman Equation","Bootstrapping (in Reinforcement Learning)","Q-Learning","Monte Carlo Methods","Generalized Advantage Estimation"]},{"id":"bootstrapping","category":"training","sec":4,"tier":3,"sources":[{"title":"Lilian Weng: A (Long) Peek into Reinforcement Learning","url":"https://lilianweng.github.io/posts/2018-02-19-rl-overview/"},{"title":"Wikipedia: Temporal difference learning","url":"https://en.wikipedia.org/wiki/Temporal_difference_learning"}],"as_of":"","related_ids":["temporal-difference-learning","value-function","bellman-equation","target-network","monte-carlo-methods","overestimation-bias"],"name":"Bootstrapping (in Reinforcement Learning)","alt":"自举","abbr":"","aliases":["Bootstrap"],"one_liner":"Updating a state's estimated value using the agent's own current estimate of the next state's value.","explanation":"Bootstrapping is a basic technique for updating a value function in reinforcement learning: the update partly relies on an existing value estimate, rather than only on rewards actually received. Temporal-difference (TD) learning is the classic example: when updating the value of a given step, the target is “this step's reward plus the discounted estimated value of the next state,” with no need to wait for the episode to end; by contrast, Monte Carlo methods must run a full episode to completion and use the actual return to update. The benefit of bootstrapping is that it can learn at every step, has low variance, and works for tasks with no natural endpoint; the cost is that the target itself carries bias, and estimation errors can propagate forward. When bootstrapping, function approximation (such as a neural network), and off-policy learning all appear together, training tends to become unstable — this combination is called the “deadly triad,” and DQN's experience replay and target network were both designed to stabilize training against it. This is unrelated to the bootstrap resampling technique in statistics.","example":"A robot arm opening a drawer gets 0 reward at every step and 1 when the drawer is fully open. When TD learning updates the value of the “gripper is holding the handle” state, it simply takes the current value estimate of the “drawer half-open” state, multiplies it by the discount factor, and uses that as the target — no need to wait for the episode to finish.","related":["Temporal-Difference Learning","Value Function","Bellman Equation","Target Network","Monte Carlo Methods","Overestimation Bias"]},{"id":"sample-efficiency","category":"training","sec":4,"tier":2,"sources":[{"title":"SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (arXiv 2401.16013)","url":"https://arxiv.org/abs/2401.16013"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL, arXiv 2410.21845)","url":"https://arxiv.org/abs/2410.21845"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["off-policy","experience-replay","real-world-reinforcement-learning","model-based-reinforcement-learning","massively-parallel-reinforcement-learning","serl"],"name":"Sample Efficiency","alt":"样本效率","abbr":"","aliases":["Data Efficiency"],"one_liner":"How much data or environment interaction is needed to reach a given performance level; using less means being more efficient.","explanation":"Sample efficiency measures how many samples a learning algorithm needs to reach a certain level of performance: in reinforcement learning, this means the number of steps of interaction with the environment; in imitation learning, the number of demonstrations. It matters enormously in robotics. In simulation it can be offset with GPU parallelism — Rudin and colleagues (2021), for instance, ran thousands of ANYmal quadrupeds at once on a single GPU, training flat-ground walking in under 4 minutes. A real robot can only act step by step, with wear and the cost of manually resetting the scene on top, so real-robot reinforcement learning must use highly sample-efficient methods. Common ways to improve it include off-policy algorithms with experience replay that reuse old data repeatedly, adding human demonstrations or corrections, model-based reinforcement learning that trains partly “in imagination,” and using pretrained visual representations. On-policy PPO is generally less sample-efficient and suits large-scale parallel simulation, while off-policy methods like SAC use samples more sparingly and are common on real robots.","example":"UC Berkeley's SERL uses off-policy reinforcement learning on a real robot, learning tasks like PCB insertion and cable routing in 25–50 minutes of training on average per task; its successor HIL-SERL adds human correction and reaches near-perfect success in 1–2.5 hours.","related":["Off-Policy","Experience Replay","Real-World Reinforcement Learning","Model-Based Reinforcement Learning","Massively Parallel Reinforcement Learning","SERL"]},{"id":"model-free-reinforcement-learning","category":"training","sec":4,"tier":2,"sources":[{"title":"Wikipedia: Model-free (reinforcement learning)","url":"https://en.wikipedia.org/wiki/Model-free_(reinforcement_learning)"},{"title":"OpenAI Spinning Up: Kinds of RL Algorithms","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html"},{"title":"RSL-RL (GitHub, leggedrobotics)","url":"https://github.com/leggedrobotics/rsl_rl"}],"as_of":"","related_ids":["model-based-reinforcement-learning","policy-gradient","proximal-policy-optimization","deep-q-network","sample-efficiency","massively-parallel-reinforcement-learning"],"name":"Model-Free Reinforcement Learning","alt":"无模型强化学习","abbr":"","aliases":["Model-Free RL"],"one_liner":"Reinforcement learning that skips building an environment model and learns a policy or value function directly from trial-and-error data.","explanation":"The counterpart to model-based reinforcement learning: the algorithm never tries to estimate the environment's state-transition or reward dynamics, and learns a policy or value function directly from (state, action, reward) samples gathered through interaction — essentially pure trial-and-error learning. It splits into two main families: methods that directly optimize the policy, such as policy gradient and PPO; and methods that learn a Q-function (estimating how much return ultimately follows from taking a given action in a given state), such as Q-learning and DQN, with DDPG and SAC sitting somewhere between the two. The advantage is simplicity and immunity to model error; the disadvantage is needing a huge amount of interaction, since sample efficiency is low. Robotics commonly compensates for this with GPU-parallel simulation: running thousands of environments at once in Isaac Gym or Isaac Lab to collect data, then transferring the trained policy to a real robot.","example":"The mainstream approach to legged-robot locomotion control: train a walking policy with PPO from the rsl_rl library inside Isaac Lab, never building an environment model at all, then deploy the trained policy to the real robot.","related":["Model-Based Reinforcement Learning","Policy Gradient","Proximal Policy Optimization","Deep Q-Network","Sample Efficiency","Massively Parallel Reinforcement Learning"]},{"id":"model-based-reinforcement-learning","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Kinds of RL Algorithms","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html"},{"title":"Model-based Reinforcement Learning: A Survey (Moerland et al.)","url":"https://arxiv.org/abs/2006.16712"},{"title":"DayDreamer: World Models for Physical Robot Learning","url":"https://arxiv.org/abs/2206.14176"}],"as_of":"","related_ids":["model-free-reinforcement-learning","world-model","learning-in-imagination","sample-efficiency","dreamerv3","model-predictive-control"],"name":"Model-Based Reinforcement Learning","alt":"基于模型的强化学习","abbr":"MBRL","aliases":["MBRL","Model-Based RL"],"one_liner":"Learning a model that predicts how the environment will respond, then using it to plan or “imagine” training data for a policy.","explanation":"This is a major branch of reinforcement learning. Beyond learning a policy, the agent also learns a dedicated environment model (also called a dynamics model or world model): given the current state and action, it predicts the next state and reward. With a model in hand, an agent can plan by rolling the model forward before acting (as in model predictive control), or generate large numbers of “imagined” trajectories from the model to train a policy — Sutton's Dyna architecture is an early example of the latter. The benefit is saving real interaction and improving sample efficiency, which matters a lot when trial and error on a real robot is costly; the risk is that once the model is inaccurate, the policy learns to exploit the model's blind spots, called model bias. The Dreamer series and TD-MPC2 both belong to this family, and it's also where world-model research and robot reinforcement learning meet.","example":"DayDreamer (2022) ran the Dreamer algorithm directly on real hardware: a quadruped robot learned to roll over, stand up, and walk from scratch in just 1 hour, with no simulator involved.","related":["Model-Free Reinforcement Learning","World Model","Learning in Imagination","Sample Efficiency","DreamerV3","Model Predictive Control"]},{"id":"learning-in-imagination","category":"training","sec":4,"tier":3,"sources":[{"title":"Dream to Control: Learning Behaviors by Latent Imagination (Dreamer, arXiv:1912.01603)","url":"https://arxiv.org/abs/1912.01603"},{"title":"DayDreamer: World Models for Physical Robot Learning (arXiv:2206.14176)","url":"https://arxiv.org/abs/2206.14176"},{"title":"Training Agents Inside of Scalable World Models (Dreamer 4, arXiv:2509.24527)","url":"https://arxiv.org/abs/2509.24527"}],"as_of":"2025-09","related_ids":["world-model","model-based-reinforcement-learning","dreamerv3","dreamer-4","daydreamer","world-models"],"name":"Learning in Imagination","alt":"想象中学习","abbr":"","aliases":["World-Model-Based RL","Training Inside a World Model","Dream Training"],"one_liner":"Learning a world model first, then training the policy inside trajectories the model “imagines”.","explanation":"Learning in imagination is a form of model-based reinforcement learning: first learn a world model from real interaction data — a model that predicts the next state and reward given the current state and action — then let the policy learn by trial and error inside trajectories the world model generates, touching the real environment as little as possible. David Ha and Jürgen Schmidhuber's 2018 paper “World Models” was an early demonstration of this idea. Danijar Hafner and colleagues' Dreamer series turned it into a general-purpose algorithm, imagining trajectories in a compressed latent space and backpropagating value-estimate gradients along the trajectory to train the policy. The upside is fewer real-world interactions and greater safety; the risk is that if the world model is inaccurate, the policy can learn tricks that only work “in the dream.”","example":"DayDreamer (2022) ran Dreamer on a real robot: a quadruped learned to stand and walk from about an hour of real interaction. Dreamer 4 (2025), trained purely on offline data inside a world model, became the first such agent to obtain a diamond in Minecraft.","related":["World Model","Model-Based Reinforcement Learning","DreamerV3","Dreamer 4","DayDreamer","World Models"]},{"id":"on-policy","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Kinds of RL Algorithms","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html"},{"title":"Proximal Policy Optimization Algorithms","url":"https://arxiv.org/abs/1707.06347"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["off-policy","proximal-policy-optimization","policy-gradient","massively-parallel-reinforcement-learning","sample-efficiency","rsl-rl"],"name":"On-Policy","alt":"同策略","abbr":"","aliases":["On-Policy Learning"],"one_liner":"Updating only with data the current policy just collected itself, and discarding old data once it's used.","explanation":"The counterpart to off-policy: the behavior policy producing the data and the target policy being optimized are the same one. Once the parameters are updated, data collected by the old policy no longer matches the current policy's distribution, so it must be discarded and fresh data collected. SARSA, REINFORCE, A2C/A3C, TRPO, and PPO are all on-policy algorithms; PPO does run several mini-batch updates on the same batch of data, but it's still considered on-policy. The advantage is stable, simple training; the disadvantage is low data efficiency, requiring huge amounts of interaction. Once GPU-based massively parallel simulation made sampling cheap, PPO became the mainstream algorithm for training legged and humanoid locomotion control.","example":"Rudin and colleagues (2021) ran thousands of ANYmal quadrupeds in parallel simulation on a single GPU, training a flat-ground walking policy in under 4 minutes and a rough-terrain policy in about 20, then transferred it to the real robot; their open-source legged_gym pairs with rsl_rl, a PPO implementation.","related":["Off-Policy","Proximal Policy Optimization","Policy Gradient","Massively Parallel Reinforcement Learning","Sample Efficiency","rsl_rl"]},{"id":"off-policy","category":"training","sec":4,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Kinds of RL Algorithms","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html"},{"title":"Wikipedia: State–action–reward–state–action (SARSA vs Q-learning)","url":"https://en.wikipedia.org/wiki/State%E2%80%93action%E2%80%93reward%E2%80%93state%E2%80%93action"},{"title":"SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning","url":"https://arxiv.org/html/2401.16013"}],"as_of":"","related_ids":["on-policy","experience-replay","soft-actor-critic","q-learning","offline-reinforcement-learning","serl"],"name":"Off-Policy","alt":"异策略","abbr":"","aliases":["Off-Policy Learning"],"one_liner":"Reinforcement learning that can use data collected by a different policy, not only data the current policy gathered itself.","explanation":"Reinforcement learning distinguishes two roles: the behavior policy, which interacts with the environment and produces data, and the target policy, the one actually being learned and improved. When these can differ, it's called off-policy learning. This means experience from older versions of the policy, human demonstrations, or data from other algorithms can all be stored in an experience-replay buffer and reused repeatedly, giving good sample efficiency; Q-learning, DQN, DDPG, TD3, and SAC all belong to this family. The cost is that the data distribution doesn't match the current policy, which makes training more prone to instability. Because real-robot interaction is expensive, real-robot reinforcement learning often chooses off-policy algorithms. Note that although this idea is sometimes loosely referred to with the same Chinese phrase as “offline,” off-policy is not the same thing as offline reinforcement learning — an off-policy algorithm is usually still collecting new data while it trains.","example":"SERL uses RLPD, an off-policy algorithm built on SAC, with each training batch drawn half from human demonstrations and half from replay data the robot collects online, training policies for tasks like PCB insertion and cable routing in 25 to 50 minutes on average.","related":["On-Policy","Experience Replay","Soft Actor-Critic","Q-Learning","Offline Reinforcement Learning","SERL"]},{"id":"policy-gradient","category":"training","sec":5,"tier":2,"sources":[{"title":"OpenAI Spinning Up: Intro to Policy Optimization","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html"},{"title":"Policy Gradient Methods for Reinforcement Learning with Function Approximation (Sutton et al., NIPS 1999)","url":"https://papers.nips.cc/paper_files/paper/1999/hash/464d828b85b0bed98e80ade0a5c43b0f-Abstract.html"}],"as_of":"","related_ids":["reinforce","advantage-function","proximal-policy-optimization","trust-region-policy-optimization","on-policy","generalized-advantage-estimation"],"name":"Policy Gradient","alt":"策略梯度","abbr":"PG","aliases":["PG","Policy Gradient Methods"],"one_liner":"Computing the gradient of expected return with respect to the policy's parameters directly, and improving the policy along that gradient.","explanation":"This is a major family of reinforcement-learning methods that optimize the policy directly. The policy is represented by a parameterized network, and the goal is to maximize expected return; the policy gradient gives the gradient of that objective with respect to the parameters, by multiplying the gradient of the log-probability of each action taken by the return that followed it — so actions that led to high return get their probability raised, and actions that led to low return get it lowered — and this can be estimated from sampled trajectories. Williams's 1992 REINFORCE is an early example, and Sutton and colleagues gave the policy gradient theorem for use with function approximation in 1999. In practice, a baseline is usually subtracted, using the advantage function (how much better this action is than average) in place of the raw return to reduce variance, which led to actor-critic methods, TRPO, and PPO. It handles continuous actions naturally, suiting robot control, but is typically on-policy and not very sample-efficient.","example":"Training a robot arm to push a box: sample a batch of trajectories with the current policy, raise the probability of actions taken in trajectories that pushed the box to the goal, lower it for actions in failed trajectories, and repeat — the policy gradually improves.","related":["REINFORCE","Advantage Function","Proximal Policy Optimization","Trust Region Policy Optimization","On-Policy","Generalized Advantage Estimation"]},{"id":"reinforce","category":"training","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Policy gradient method","url":"https://en.wikipedia.org/wiki/Policy_gradient_method"},{"title":"Ahmadian et al. 2024: Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs","url":"https://arxiv.org/abs/2402.14740"}],"as_of":"","related_ids":["policy-gradient","monte-carlo-methods","return","advantage-function","proximal-policy-optimization","group-relative-policy-optimization"],"name":"REINFORCE","alt":"REINFORCE 算法","abbr":"","aliases":["Monte Carlo Policy Gradient"],"one_liner":"The earliest policy-gradient algorithm: raise the probability of actions that led to high return across a full trajectory.","explanation":"REINFORCE was proposed by Ronald Williams in 1992, and is the earliest policy-gradient algorithm. It runs the current policy through several full episodes and, for every action taken, weights the gradient of that action's log-probability by the cumulative return that followed it, so actions that led to higher return get pushed to be more likely. Because the return comes directly from sampling a whole trajectory rather than from a learned value estimate, it is also called the Monte Carlo policy gradient; the estimate is unbiased but high-variance and sample-inefficient. Subtracting a baseline, such as the average return or a value function, noticeably reduces variance, and this idea later grew into actor-critic methods. PPO also builds on policy gradients; in large-language-model post-training, RLOO and GRPO use the average return within a group as a baseline, carrying REINFORCE's basic idea forward.","example":"Training an arm to push a block: an episode that reaches the goal scores 1, otherwise 0. REINFORCE raises the probability of every action in a successful episode, while a failed episode gets zero gradient; with a baseline added, actions in below-average episodes get pushed down instead.","related":["Policy Gradient","Monte Carlo Methods","Return","Advantage Function","Proximal Policy Optimization","Group Relative Policy Optimization"]},{"id":"generalized-advantage-estimation","category":"training","sec":5,"tier":3,"sources":[{"title":"High-Dimensional Continuous Control Using Generalized Advantage Estimation (arXiv:1506.02438)","url":"https://arxiv.org/abs/1506.02438"},{"title":"legged_gym: legged_robot_config.py","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["advantage-function","proximal-policy-optimization","temporal-difference-learning","discount-factor","value-function","policy-gradient"],"name":"Generalized Advantage Estimation","alt":"广义优势估计","abbr":"GAE","aliases":["GAE","GAE(λ)"],"one_liner":"Estimating the advantage function by summing multi-step temporal-difference errors with exponentially decaying λ weights.","explanation":"GAE was introduced by Schulman, Levine, Abbeel, and colleagues in 2015. Policy gradient methods need to know how much better than average a given action was — the advantage function. Estimating it from just a single-step temporal-difference (TD) error, δ = r + γV(s′) − V(s), gives low variance but high bias; using the full episode's Monte Carlo return gives low bias but high variance. GAE sums the TD errors of future steps, each weighted by (γλ)^k, with λ between 0 and 1 tuning the trade-off between the two: λ = 0 recovers single-step TD, and λ = 1 recovers Monte Carlo — the same idea as TD(λ). It's PPO's standard configuration, and it's the near-universal way PPO computes advantages when training locomotion control for legged and humanoid robots.","example":"legged_gym's default PPO configuration uses a discount factor γ = 0.99 and a GAE parameter λ = 0.95.","related":["Advantage Function","Proximal Policy Optimization","Temporal-Difference Learning","Discount Factor","Value Function","Policy Gradient"]},{"id":"importance-sampling","category":"training","sec":5,"tier":3,"sources":[{"title":"Importance sampling (Wikipedia)","url":"https://en.wikipedia.org/wiki/Importance_sampling"},{"title":"Policy Gradient Algorithms (Lilian Weng)","url":"https://lilianweng.github.io/posts/2018-04-08-policy-gradient/"}],"as_of":"","related_ids":["off-policy","proximal-policy-optimization","off-policy-evaluation","policy-gradient","trust-region-policy-optimization","monte-carlo-methods"],"name":"Importance Sampling","alt":"重要性采样","abbr":"","aliases":["Importance Weighting"],"one_liner":"Estimating an expectation under one distribution using samples drawn from another, weighted by the ratio of the two probabilities.","explanation":"This is a Monte Carlo estimation technique. Wanting the expectation of some quantity under a distribution p, but having only samples drawn from a different distribution q, each sample is weighted by p(x)/q(x); the weighted average is still an unbiased estimate of the expectation under p (provided q can actually produce every value p can). In reinforcement learning, p and q are usually two policies: data collected under an old or behavior policy is used to evaluate or update a new policy, with the weight being the ratio of the new and old policies' probabilities for the same action. Both off-policy learning and off-policy evaluation depend on it; the probability ratio r(θ) in the TRPO and PPO objectives is exactly this, and PPO clips that ratio to keep any single update from being too large. Its main problem is that the further apart the two distributions are, the higher the weights' variance gets, so in practice the weights are often truncated or clipped.","example":"Each round, PPO collects a batch of data with the old policy and then updates on that same batch several times; every sample's loss is multiplied by the new-to-old policy probability ratio, which is clipped to the range [1−ε, 1+ε].","related":["Off-Policy","Proximal Policy Optimization","Off-Policy Evaluation","Policy Gradient","Trust Region Policy Optimization","Monte Carlo Methods"]},{"id":"trust-region-policy-optimization","category":"training","sec":5,"tier":3,"sources":[{"title":"Schulman et al. 2015: Trust Region Policy Optimization (ICML 2015)","url":"https://arxiv.org/abs/1502.05477"},{"title":"OpenAI Spinning Up: Trust Region Policy Optimization","url":"https://spinningup.openai.com/en/latest/algorithms/trpo.html"},{"title":"Schulman et al. 2017: Proximal Policy Optimization Algorithms","url":"https://arxiv.org/abs/1707.06347"}],"as_of":"","related_ids":["proximal-policy-optimization","policy-gradient","kullback-leibler-divergence","on-policy","generalized-advantage-estimation","reinforcement-learning"],"name":"Trust Region Policy Optimization","alt":"信赖域策略优化","abbr":"TRPO","aliases":["TRPO"],"one_liner":"A policy-gradient algorithm that bounds the KL divergence between old and new policy at every update; PPO's predecessor.","explanation":"Trust region policy optimization was proposed by Schulman, Levine, Abbeel, and colleagues at ICML 2015. An ordinary policy-gradient step can, if too large, cause performance to collapse, and even a small change in parameters can shift the action distribution a great deal. TRPO instead puts the constraint on the policy's distribution: it maximizes a surrogate objective within a “trust region” where the average KL divergence between old and new policy stays under a threshold, which in theory guarantees monotonic improvement. In practice it uses conjugate gradient to find an approximate second-order update direction, then a backtracking line search to check the constraint and the improvement. It is an on-policy algorithm and fairly involved to implement; PPO, proposed by the same authors in 2017, approximates the same constraint far more simply and has become the mainstream choice for robot reinforcement learning.","example":"In OpenAI's Spinning Up implementation of TRPO, each iteration first computes an update direction with conjugate gradient, then keeps shrinking the step size by a backtracking factor until the new policy stays within the KL limit and improves the surrogate objective.","related":["Proximal Policy Optimization","Policy Gradient","Kullback-Leibler Divergence","On-Policy","Generalized Advantage Estimation","Reinforcement Learning"]},{"id":"proximal-policy-optimization","category":"training","sec":5,"tier":1,"sources":[{"title":"Schulman et al. 2017: Proximal Policy Optimization Algorithms","url":"https://arxiv.org/abs/1707.06347"},{"title":"OpenAI Spinning Up: Proximal Policy Optimization","url":"https://spinningup.openai.com/en/latest/algorithms/ppo.html"},{"title":"Rudin et al. 2021: Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["reinforcement-learning","policy-gradient","trust-region-policy-optimization","on-policy","generalized-advantage-estimation","rsl-rl"],"name":"Proximal Policy Optimization","alt":"近端策略优化","abbr":"PPO","aliases":["PPO-Clip","PPO"],"one_liner":"OpenAI's 2017 reinforcement-learning algorithm that limits how much each policy update can change the policy, making training stable.","explanation":"PPO was proposed in 2017 by Schulman and colleagues at OpenAI. It's a policy-gradient method (one that directly adjusts a policy network so that actions leading to higher return get chosen more often), and it's on-policy: it only trains on data just collected by the current version of the policy. Its core idea is an objective function with “clipping”: once the new policy changes an action's probability beyond a certain range relative to the old policy, the objective stops rewarding further change in that direction. This keeps a single update from wrecking the policy, while still allowing multiple passes of training over the same batch of data. It's about as stable as TRPO (Trust Region Policy Optimization) but far simpler to implement. PPO is the workhorse algorithm for training legged and humanoid robot locomotion in simulation — ETH Zurich's open-source robot RL library rsl_rl is built around it — and it's also the algorithm behind InstructGPT's RLHF.","example":"ETH Zurich's Rudin and colleagues used PPO in legged_gym, running thousands of simulated environments in parallel on a single GPU, to train an ANYmal quadruped: a flat-ground walking policy trained in under 4 minutes and a rough-terrain policy in about 20, both then transferred to the real robot.","related":["Reinforcement Learning","Policy Gradient","Trust Region Policy Optimization","On-Policy","Generalized Advantage Estimation","rsl_rl"]},{"id":"group-relative-policy-optimization","category":"training","sec":5,"tier":2,"sources":[{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","url":"https://arxiv.org/abs/2402.03300"},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","url":"https://arxiv.org/abs/2501.12948"},{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"2025-09","related_ids":["proximal-policy-optimization","advantage-function","reinforcement-fine-tuning","reinforcement-learning-with-verifiable-rewards","kl-regularization","simplevla-rl"],"name":"Group Relative Policy Optimization","alt":"组相对策略优化","abbr":"GRPO","aliases":["GRPO"],"one_liner":"A reinforcement-learning algorithm that samples a group of outputs for the same input and uses their relative scores instead of a value network.","explanation":"GRPO was introduced by the DeepSeek team in the DeepSeekMath paper (February 2024) as a variant of PPO (Proximal Policy Optimization). PPO needs a separately trained value network (a “critic,” roughly as large as the policy itself) to estimate a baseline; GRPO removes it. For the same input, it samples a group of outputs, scores each one, and uses “score minus the group's mean, divided by the group's standard deviation” as each output's advantage, saving a large amount of memory. It also adds the KL divergence (a measure of how far the policy has drifted) from a reference model directly into the loss, to constrain how much the policy changes per update. GRPO became widely adopted after DeepSeek-R1 used it to train reasoning ability. The embodied-AI field has begun applying it to RL fine-tuning of VLAs too: several rollouts are sampled for the same task and given a reward of 0 or 1 depending on success.","example":"SimpleVLA-RL trains OpenVLA-OFT with GRPO: it samples 8 rollouts per task, scores success as 1 and failure as 0, drops the KL term, and filters out groups that are all successes or all failures, raising the average LIBERO success rate from 91.0% to 99.1%.","related":["Proximal Policy Optimization","Advantage Function","Reinforcement Fine-Tuning (RL Fine-Tuning)","Reinforcement Learning with Verifiable Rewards","KL Regularization","SimpleVLA-RL"]},{"id":"q-learning","category":"training","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Q-learning","url":"https://en.wikipedia.org/wiki/Q-learning"},{"title":"Hugging Face Deep RL Course: Introducing Q-Learning","url":"https://huggingface.co/learn/deep-rl-course/unit2/q-learning"},{"title":"QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation","url":"https://arxiv.org/abs/1806.10293"}],"as_of":"","related_ids":["q-function","temporal-difference-learning","deep-q-network","off-policy","overestimation-bias","qt-opt"],"name":"Q-Learning","alt":"Q 学习","abbr":"","aliases":["Tabular Q-Learning"],"one_liner":"A classic model-free reinforcement-learning algorithm that repeatedly corrects its Q-values toward the immediate reward plus the next state's best Q-value.","explanation":"Q-learning was introduced by Chris Watkins in his 1989 doctoral thesis, with a convergence proof given with Peter Dayan in 1992. It needs no environment model: at every step, it uses “immediate reward plus discount times the next state's maximum Q-value” as a target and nudges the current Q-value toward it, a self-referential update that makes it a form of temporal-difference learning. It's an off-policy algorithm: it explores randomly with ε-greedy while collecting data, but learns the greedy, optimal policy anyway, so old data can be reused repeatedly. Early implementations stored Q-values in a table, which only worked for small discrete problems; DeepMind's DQN replaced the table with a neural network. The max operation tends to bias Q-values upward, and double Q-learning was designed specifically to correct this.","example":"QT-Opt (Google, 2018) trained a Q-learning-based visual grasping policy on more than 580,000 real grasp attempts, reaching a 96% success rate on objects it had never seen during training.","related":["Q-Function","Temporal-Difference Learning","Deep Q-Network","Off-Policy","Overestimation Bias","QT-Opt"]},{"id":"deep-q-network","category":"training","sec":5,"tier":2,"sources":[{"title":"Playing Atari with Deep Reinforcement Learning (arXiv 1312.5602)","url":"https://arxiv.org/abs/1312.5602"},{"title":"Google DeepMind: Deep Reinforcement Learning","url":"https://deepmind.google/discover/blog/deep-reinforcement-learning/"},{"title":"QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (arXiv 1806.10293)","url":"https://arxiv.org/abs/1806.10293"}],"as_of":"","related_ids":["q-learning","q-function","experience-replay","target-network","deep-deterministic-policy-gradient","qt-opt"],"name":"Deep Q-Network","alt":"深度 Q 网络","abbr":"DQN","aliases":["DQN","Deep Q-Learning"],"one_liner":"A reinforcement-learning algorithm that uses a deep neural network to estimate each action's long-term value, its Q-value.","explanation":"DQN comes from DeepMind's Mnih and colleagues: a 2013 preprint validated it on 7 Atari games, and a 2015 Nature paper extended it to 49 games, using only raw screen pixels and score, reaching performance comparable to a professional human tester overall. It replaces the lookup table in Q-learning (learning how much total return follows from taking a given action in a given state) with a convolutional neural network, and stabilizes training with experience replay (storing past experience and sampling it randomly for training) and a target network (a periodically synced copy used to compute the training target). This launched the deep-reinforcement-learning boom. Because it takes the maximum Q-value over all actions, DQN suits discrete action spaces; for a robot arm's continuous actions, methods like DDPG and SAC are used instead.","example":"Google's QT-Opt (2018) brought Q-learning to real robot-arm grasping: after more than 580,000 real-world grasp attempts, the resulting closed-loop, vision-based grasping policy reached 96% success on objects it had never seen.","related":["Q-Learning","Q-Function","Experience Replay","Target Network","Deep Deterministic Policy Gradient","QT-Opt"]},{"id":"experience-replay","category":"training","sec":5,"tier":2,"sources":[{"title":"Playing Atari with Deep Reinforcement Learning (DQN, 2013)","url":"https://arxiv.org/abs/1312.5602"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/abs/2410.21845"}],"as_of":"","related_ids":["off-policy","deep-q-network","soft-actor-critic","hindsight-experience-replay","sample-efficiency","hil-serl"],"name":"Experience Replay","alt":"经验回放","abbr":"","aliases":["Replay Buffer"],"one_liner":"Storing an agent's past interactions in a buffer and sampling them randomly during training, instead of using each one only once.","explanation":"Experience replay is a data-reuse mechanism in reinforcement learning, proposed by Long-Ji Lin in the early 1990s and popularized when DeepMind's DQN (Deep Q-Network) combined it with deep networks in 2013. At every step, the agent stores the transition — state, action, reward, next state — into a buffer of limited capacity, overwriting the oldest entries once it's full; training then samples random mini-batches from this buffer to update the network. This has two benefits: each piece of data can be reused many times, improving sample efficiency; and random sampling breaks the strong correlation between consecutive steps, making training more stable. Because the sampled data comes from an older policy, any algorithm using it must be off-policy, such as DQN or SAC. Real-robot reinforcement learning, where samples are expensive, relies on this especially heavily.","example":"DQN playing Atari keeps the most recent 1 million frames in its replay buffer. HIL-SERL uses two buffers — one for human demonstrations and intervention data, one for the policy's own interactions — and samples half of each training batch from each.","related":["Off-Policy","Deep Q-Network","Soft Actor-Critic","Hindsight Experience Replay","Sample Efficiency","HIL-SERL"]},{"id":"target-network","category":"training","sec":5,"tier":3,"sources":[{"title":"OpenAI Spinning Up: Deep Deterministic Policy Gradient（Target Networks 一节）","url":"https://spinningup.openai.com/en/latest/algorithms/ddpg.html"},{"title":"Wikipedia: Q-learning（Deep Q-learning 一节）","url":"https://en.wikipedia.org/wiki/Q-learning"},{"title":"Fujimoto et al. 2018: Addressing Function Approximation Error in Actor-Critic Methods (TD3)","url":"https://arxiv.org/abs/1802.09477"}],"as_of":"","related_ids":["deep-q-network","q-function","temporal-difference-learning","experience-replay","overestimation-bias","soft-actor-critic"],"name":"Target Network","alt":"目标网络","abbr":"","aliases":["Target Q-Network"],"one_liner":"A slowly updated copy of the main network, used only to compute training targets, making Q-learning more stable.","explanation":"In temporal-difference methods like Q-learning, the training target is immediate reward plus the discounted Q-value of the next state, and that Q-value is computed by the very same network being trained, effectively chasing a target that moves as you move, which is prone to diverge with neural networks. DeepMind's 2015 DQN paper, published in Nature, introduced the target network: a copy of the main network dedicated to computing targets, synchronized only every fixed number of steps. Continuous-control algorithms such as DDPG, TD3, and SAC switched to soft updates instead, nudging the target network's parameters toward the main network's a little every step, called Polyak averaging, with a coefficient close to 1. Together with experience replay, it is a standard component of off-policy deep reinforcement learning; the TD3 paper also analyzes its relationship to Q-value overestimation.","example":"In OpenAI's Spinning Up documentation, DDPG's target network is soft-updated as φ_targ ← ρ·φ_targ + (1−ρ)·φ, with an example ρ of 0.995, while DQN-style algorithms instead copy the whole main network over every fixed number of steps.","related":["Deep Q-Network","Q-Function","Temporal-Difference Learning","Experience Replay","Overestimation Bias","Soft Actor-Critic"]},{"id":"overestimation-bias","category":"training","sec":5,"tier":3,"sources":[{"title":"van Hasselt, Guez, Silver 2015: Deep Reinforcement Learning with Double Q-learning","url":"https://arxiv.org/abs/1509.06461"},{"title":"Fujimoto et al. 2018: Addressing Function Approximation Error in Actor-Critic Methods (TD3)","url":"https://arxiv.org/abs/1802.09477"}],"as_of":"","related_ids":["q-learning","deep-q-network","twin-delayed-ddpg","target-network","extrapolation-error","conservative-q-learning"],"name":"Overestimation Bias","alt":"Q 值高估","abbr":"","aliases":["Q-value Overestimation","Maximization Bias"],"one_liner":"Taking a max over noisy Q-value estimates systematically inflates them, biasing an agent toward overrated actions.","explanation":"Q-learning's update target includes a step that takes the maximum Q-value over all actions in the next state. Because Q-values are themselves noisy estimates, taking the max of a set of noisy numbers is biased high in expectation, and that bias then compounds across many rounds of bootstrapping, pushing the agent to favor overrated actions and degrading the policy. Thrun and Schwartz pointed this out as early as 1993. Hado van Hasselt proposed Double Q-learning, using one set of estimates to select the action and another to evaluate it, and in 2015 turned it into Double DQN with DeepMind colleagues, confirming that the original DQN showed clear overestimation on several Atari games. In continuous control, TD3 takes the smaller of two critic networks' values to suppress overestimation, and SAC does the same. In offline RL, overestimation is even worse for actions never seen in the data, which is exactly what methods like conservative Q-learning are built to handle.","example":"TD3 trains two critic networks at once and, when computing the target value, takes the smaller of the two (clipped double Q-learning), preferring a slight underestimate over an overestimate to make training more stable.","related":["Q-Learning","Deep Q-Network","Twin Delayed DDPG","Target Network","Extrapolation Error (OOD Actions in Offline RL)","Conservative Q-Learning"]},{"id":"distributional-value-function","category":"training","sec":5,"tier":3,"sources":[{"title":"A Distributional Perspective on Reinforcement Learning (arXiv:1707.06887)","url":"https://arxiv.org/abs/1707.06887"},{"title":"π*0.6: a VLA That Learns From Experience (arXiv:2511.14759)","url":"https://arxiv.org/abs/2511.14759"},{"title":"Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (arXiv:2605.00416)","url":"https://arxiv.org/abs/2605.00416"}],"as_of":"2026-09","related_ids":["value-function","bellman-equation","return","recap","pi-star-0-6","cross-entropy"],"name":"Distributional Value Function","alt":"分布式价值函数","abbr":"","aliases":["Distributional Reinforcement Learning","Distributional RL","Value Distribution"],"one_liner":"A value function that predicts the full probability distribution of future return, not just its average.","explanation":"An ordinary value function outputs only the expected value of future return (cumulative reward); a distributional value function outputs the entire distribution of possible return values instead. DeepMind's Bellemare, Dabney, and Munos systematically introduced this view in a 2017 ICML paper, giving a distributional Bellman equation and the C51 algorithm (which represents the return distribution with 51 discrete support points), reaching state-of-the-art results on Atari at the time. A common implementation splits the return into several bins, has the network output a probability for each bin, and trains with cross-entropy, which is more stable than regressing a single number directly and also captures uncertainty. Recent VLA reinforcement-learning methods commonly use this as the critic.","example":"Physical Intelligence's π*0.6 (RECAP) trains a multi-task distributional value function: it discretizes the “steps remaining until success” return into bins and trains with cross-entropy, then uses it to estimate advantages for filtering good actions; AgiBot's LWD (2026) similarly uses distributional implicit value learning (DIVL) to handle sparse-reward data collected by a robot fleet.","related":["Value Function","Bellman Equation","Return","RECAP","π*0.6","Cross-Entropy"]},{"id":"deterministic-vs-stochastic-policy","category":"training","sec":5,"tier":3,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL (Policies)","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"Deterministic Policy Gradient Algorithms (ICML 2014)","url":"https://proceedings.mlr.press/v32/silver14.html"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["policy","gaussian-policy","action-multimodality","deep-deterministic-policy-gradient","diffusion-policy","exploration-vs-exploitation"],"name":"Deterministic vs. Stochastic Policy","alt":"确定性策略 / 随机策略","abbr":"","aliases":["Deterministic Policy","Stochastic Policy"],"one_liner":"A deterministic policy always gives the same action for the same state; a stochastic policy gives a probability distribution to sample from.","explanation":"This is a classification of policies (mappings from observation to action) by the form of their output. A deterministic policy is written a = μ(s), always outputting the same action for the same state; a stochastic policy is written a ~ π(·|s), outputting a probability distribution over actions, which may sample differently each time. Discrete actions commonly use a categorical distribution, giving a probability per action via softmax like a classifier; continuous actions commonly use a diagonal Gaussian, with the network outputting a mean and a log standard deviation. In reinforcement learning, stochastic policies carry exploration built in, and PPO and SAC both use them; DDPG and TD3 use deterministic policies and add noise separately during training for exploration, which makes the policy-gradient estimate more efficient. In imitation learning, behavior cloning trained with mean squared error regression is essentially deterministic, and when the same scene in the demonstrations has multiple valid solutions, it gets averaged into one wrong action; stochastic policies like diffusion policies or Gaussian mixture models can represent this action multimodality instead.","example":"A robot arm reaching around an obstacle to grab a cup, with half the demonstrations going left and half going right: a deterministic policy trained with mean squared error regression outputs the average of the two and drives straight into the obstacle, while a stochastic policy like a diffusion policy lands on either the left or the right route on any given sample.","related":["Policy","Gaussian Policy","Action Multimodality","Deep Deterministic Policy Gradient","Diffusion Policy","Exploration vs. Exploitation"]},{"id":"deep-deterministic-policy-gradient","category":"training","sec":5,"tier":3,"sources":[{"title":"Continuous control with deep reinforcement learning (arXiv 1509.02971)","url":"https://arxiv.org/abs/1509.02971"},{"title":"OpenAI Spinning Up: Deep Deterministic Policy Gradient","url":"https://spinningup.openai.com/en/latest/algorithms/ddpg.html"},{"title":"Hindsight Experience Replay (arXiv 1707.01495)","url":"https://arxiv.org/abs/1707.01495"}],"as_of":"","related_ids":["twin-delayed-ddpg","soft-actor-critic","deep-q-network","deterministic-vs-stochastic-policy","off-policy","hindsight-experience-replay"],"name":"Deep Deterministic Policy Gradient","alt":"深度确定性策略梯度","abbr":"DDPG","aliases":["DDPG"],"one_liner":"An actor-critic algorithm that extends DQN to continuous actions, with the policy outputting one deterministic action directly.","explanation":"DDPG was introduced by DeepMind's Lillicrap and colleagues in 2015 (ICLR 2016), building on the deterministic policy gradient (DPG) theory from Silver and colleagues (2014). DQN can only handle discrete actions, because it needs to search over every action for the one with the highest Q-value, and continuous actions like a robot arm's joint angles can't be enumerated one by one. DDPG uses two networks: a critic that learns the Q-function, and an actor (the policy network) that outputs one deterministic action directly, updated along the gradient direction that increases the Q-value. It's off-policy, reusing DQN's experience replay and target network, and adds noise to the action during training for exploration. The original paper validated it on more than 20 simulated physics tasks, several of which could be learned directly from pixels. DDPG is sensitive to hyperparameters and prone to overestimating Q-values; the later TD3 (Twin Delayed DDPG) specifically fixes the overestimation problem, and algorithms like SAC are more commonly used in practice today.","example":"OpenAI's 2017 hindsight experience replay (HER) experiments used DDPG to train a 7-DOF Fetch arm in simulation to push, slide, and pick-and-place objects, and deployed the resulting policy to a real robot.","related":["Twin Delayed DDPG","Soft Actor-Critic","Deep Q-Network","Deterministic vs. Stochastic Policy","Off-Policy","Hindsight Experience Replay"]},{"id":"twin-delayed-ddpg","category":"training","sec":5,"tier":2,"sources":[{"title":"Addressing Function Approximation Error in Actor-Critic Methods (TD3, arXiv 1802.09477)","url":"https://arxiv.org/abs/1802.09477"},{"title":"OpenAI Spinning Up: Twin Delayed DDPG","url":"https://spinningup.openai.com/en/latest/algorithms/td3.html"},{"title":"A Minimalist Approach to Offline Reinforcement Learning (TD3+BC, arXiv 2106.06860)","url":"https://arxiv.org/abs/2106.06860"}],"as_of":"","related_ids":["deep-deterministic-policy-gradient","overestimation-bias","target-network","soft-actor-critic","off-policy","value-function"],"name":"Twin Delayed DDPG","alt":"双延迟深度确定性策略梯度","abbr":"TD3","aliases":["TD3","Twin Delayed Deep Deterministic Policy Gradient"],"one_liner":"A continuous-action reinforcement-learning algorithm that adds three fixes to DDPG specifically to curb Q-value overestimation.","explanation":"TD3 was introduced by Scott Fujimoto, Herke van Hoof, and David Meger at ICML 2018, as an improved version of DDPG (an actor-critic algorithm that outputs deterministic actions). It's off-policy (able to reuse old data repeatedly) and applies only to continuous actions. DDPG's Q-network tends to overestimate action values, and the policy learns to exploit those inflated estimates, which often makes training unstable. TD3 addresses this with three changes: training two Q-networks and taking the smaller one when computing the target (twin); updating the policy and target networks only once for every two Q-network updates (delayed); and adding clipped noise to the target action, smoothing how the Q-value changes with the action. TD3 and SAC are the two most common baselines for continuous control, and the offline reinforcement-learning method TD3+BC is also built on it.","example":"The original paper tests on OpenAI Gym's MuJoCo continuous-control tasks (such as HalfCheetah, Hopper, and Walker2d), where TD3 outperformed DDPG and other leading algorithms of the time.","related":["Deep Deterministic Policy Gradient","Overestimation Bias","Target Network","Soft Actor-Critic","Off-Policy","Value Function"]},{"id":"soft-actor-critic","category":"training","sec":5,"tier":2,"sources":[{"title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep RL with a Stochastic Actor (arXiv 1801.01290)","url":"https://arxiv.org/abs/1801.01290"},{"title":"Soft Actor Critic—Deep Reinforcement Learning with Real-World Robots (BAIR Blog, 2018)","url":"https://bair.berkeley.edu/blog/2018/12/14/sac/"},{"title":"Soft Actor-Critic (OpenAI Spinning Up)","url":"https://spinningup.openai.com/en/latest/algorithms/sac.html"}],"as_of":"","related_ids":["off-policy","entropy-regularization","q-function","experience-replay","twin-delayed-ddpg","serl"],"name":"Soft Actor-Critic","alt":"软演员-评论家","abbr":"SAC","aliases":["SAC","Maximum Entropy Actor-Critic"],"one_liner":"An off-policy reinforcement-learning algorithm that pursues high return while also encouraging the policy to stay somewhat random.","explanation":"SAC was introduced by UC Berkeley's Tuomas Haarnoja, Sergey Levine, and colleagues in 2018 (ICML 2018). It's an actor-critic method: the “actor” is the policy network that outputs actions, and the “critic” is a Q-function estimating how good an action is. Its core idea is the maximum-entropy objective: maximize return while also keeping the policy's entropy — how random it is — as high as reasonably possible, which encourages exploration and avoids collapsing onto a suboptimal solution too early; the degree of randomness is set by a temperature coefficient α, which later versions can tune automatically. It's off-policy, able to reuse old data from an experience-replay buffer repeatedly, which gives it good sample efficiency and low sensitivity to hyperparameters, though it applies only to continuous action spaces. Implementation-wise, it trains two Q-networks and takes the smaller estimate, to counter Q-value overestimation. SAC is a common foundation for real-robot reinforcement learning; the RLPD algorithm used by SERL and HIL-SERL is itself a refinement built on SAC.","example":"The SAC extension paper applies it directly on real hardware: a Minitaur quadruped learns to walk in about 2 hours, a Sawyer arm learns to stack blocks in about 2 hours, and a dexterous hand learns to turn a valve directly from images in about 20 hours.","related":["Off-Policy","Entropy Regularization","Q-Function","Experience Replay","Twin Delayed DDPG","SERL"]},{"id":"entropy-regularization","category":"training","sec":5,"tier":3,"sources":[{"title":"OpenAI Spinning Up: Soft Actor-Critic (Entropy-Regularized RL)","url":"https://spinningup.openai.com/en/latest/algorithms/sac.html"},{"title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep RL with a Stochastic Actor (arXiv:1801.01290)","url":"https://arxiv.org/abs/1801.01290"},{"title":"legged_gym: legged_robot_config.py","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"}],"as_of":"","related_ids":["soft-actor-critic","proximal-policy-optimization","exploration-vs-exploitation","entropy-collapse-mode-collapse","kl-regularization","deterministic-vs-stochastic-policy"],"name":"Entropy Regularization","alt":"熵正则化","abbr":"","aliases":["Entropy Bonus","Maximum Entropy RL"],"one_liner":"Adding a bonus for policy entropy to the training objective, to encourage the policy to stay random and keep exploring.","explanation":"Entropy measures how random a probability distribution is: a uniform distribution has high entropy, and a nearly certain one has low entropy. Entropy regularization adds α times the policy's entropy to the reinforcement-learning objective (α is a weighting coefficient), so the agent pursues return while retaining some randomness, avoiding collapsing onto a suboptimal behavior too early. There are two common forms: adding a small entropy bonus term to the loss in policy-gradient algorithms like PPO, or writing entropy directly into the optimization objective, giving maximum-entropy reinforcement learning — the leading example is Haarnoja and colleagues' 2018 SAC (Soft Actor-Critic), whose Q-value target also carries an entropy term. Too large an α and the policy stays erratic forever; too small and it risks entropy collapse, stopping exploration too soon.","example":"legged_gym's default PPO config sets the entropy coefficient entropy_coef to 0.01, adding 0.01 times the policy's entropy as a bonus in the loss.","related":["Soft Actor-Critic","Proximal Policy Optimization","Exploration vs. Exploitation","Entropy Collapse / Mode Collapse","KL Regularization","Deterministic vs. Stochastic Policy"]},{"id":"entropy-collapse-mode-collapse","category":"training","sec":5,"tier":3,"sources":[{"title":"NIPS 2016 Tutorial: Generative Adversarial Networks (arXiv:1701.00160)","url":"https://arxiv.org/abs/1701.00160"},{"title":"The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (arXiv:2505.22617)","url":"https://arxiv.org/abs/2505.22617"},{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv:2509.09674)","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"","related_ids":["entropy-regularization","exploration-vs-exploitation","generative-adversarial-network","action-multimodality","group-relative-policy-optimization","simplevla-rl"],"name":"Entropy Collapse / Mode Collapse","alt":"熵坍缩 / 模式坍缩","abbr":"","aliases":["Policy Entropy Collapse","Helvetica Scenario"],"one_liner":"A model's output diversity collapsing, so it only ever produces a handful of answers or actions.","explanation":"These two terms describe the same broad phenomenon. Mode collapse was first used for GANs (generative adversarial networks): Goodfellow's 2016 tutorial describes it as the generator mapping many different input noise vectors to the same output, covering only one or two “modes” of the data distribution. Entropy collapse is used more in reinforcement learning: a policy's entropy (how random it is) drops rapidly early in training, the model settles on almost one single behavior, stops exploring, and performance plateaus; Cui and colleagues (2025) analyzed this systematically in large-model reasoning reinforcement learning and proposed Clip-Cov and KL-Cov to maintain entropy. Robot actions are naturally multimodal, and when RL fine-tuning a VLA, common fixes for collapse include raising the sampling temperature, widening the upper clipping bound, and adding an entropy bonus.","example":"When running GRPO reinforcement learning on OpenVLA-OFT, SimpleVLA-RL strengthens exploration with three changes: dynamic sampling, raising the upper clipping bound of PPO-style clipping (following DAPO's Clip-Higher), and raising the rollout sampling temperature.","related":["Entropy Regularization","Exploration vs. Exploitation","Generative Adversarial Network","Action Multimodality","Group Relative Policy Optimization","SimpleVLA-RL"]},{"id":"update-to-data-ratio","category":"training","sec":5,"tier":3,"sources":[{"title":"Chen et al. 2021: Randomized Ensembled Double Q-Learning: Learning Fast Without a Model (REDQ, ICLR 2021)","url":"https://arxiv.org/abs/2101.05982"},{"title":"Smith, Kostrikov, Levine 2022: A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning","url":"https://arxiv.org/abs/2208.07860"},{"title":"Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning","url":"https://arxiv.org/abs/2401.16013"}],"as_of":"","related_ids":["sample-efficiency","experience-replay","off-policy","real-world-reinforcement-learning","serl","overestimation-bias"],"name":"Update-to-Data Ratio","alt":"更新-数据比","abbr":"UTD","aliases":["UTD","UTD Ratio","Replay Ratio"],"one_liner":"How many gradient updates are done per step of environment data collected; higher saves data but costs more compute.","explanation":"The update-to-data ratio, in off-policy reinforcement learning, is how many gradient updates the network does for every one step of interaction added to the experience replay buffer. Standard SAC usually uses 1. Real-robot RL collects data slowly and expensively while compute is comparatively cheap, so there is an incentive to raise this ratio and squeeze more learning out of the same batch of data. But raising it naively tends to make the Q-function overfit and overestimate, hurting training. REDQ (ICLR 2021) used an ensemble of Q-networks to let a model-free algorithm run stably at a UTD far above 1 for the first time; DroQ instead used dropout and layer normalization for a cheaper version. Real-robot RL frameworks such as SERL rely on a high UTD for sample efficiency, and split data collection and training into two separate threads.","example":"“A Walk in the Park” (2022) built on SAC for a Unitree A1 quadruped, adding layer normalization and raising the UTD to 20, that is, 20 critic updates per step of data collected, learning to walk from about 20 minutes of real-robot training.","related":["Sample Efficiency","Experience Replay","Off-Policy","Real-World Reinforcement Learning","SERL","Overestimation Bias"]},{"id":"multi-agent-reinforcement-learning","category":"training","sec":5,"tier":3,"sources":[{"title":"Albrecht, Christianos, Schäfer: Multi-Agent Reinforcement Learning: Foundations and Modern Approaches (MIT Press, 2024)","url":"https://www.marl-book.com/"},{"title":"Wikipedia: Multi-agent reinforcement learning","url":"https://en.wikipedia.org/wiki/Multi-agent_reinforcement_learning"},{"title":"Yu et al. 2021: The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games (MAPPO)","url":"https://arxiv.org/abs/2103.01955"}],"as_of":"","related_ids":["reinforcement-learning","multi-robot-collaboration","self-play","swarm-intelligence","partially-observable-markov-decision-process","proximal-policy-optimization"],"name":"Multi-Agent Reinforcement Learning","alt":"多智能体强化学习","abbr":"MARL","aliases":["MARL"],"one_liner":"Reinforcement learning where multiple agents learn at once in a shared environment, cooperating or competing.","explanation":"Multi-agent reinforcement learning studies settings where several learning decision-makers share one environment; based on the reward relationship, it splits into fully cooperative (a shared reward), fully competitive (a zero-sum game, like chess), and mixed settings, such as self-driving cars that each want to reach their own destination while all avoiding collisions. The biggest difficulty compared to single-agent RL is non-stationarity: every other agent keeps changing its policy too, so from any one agent's perspective the environment never stops shifting, which breaks the convergence guarantees single-agent algorithms rely on; credit assignment and partial observability are additional challenges. A common approach is “centralized training, decentralized execution”: during training a critic sees global information, while at execution time each agent sees only its own local observation, as in 2017's MADDPG and 2021's MAPPO. Multi-robot collaboration, robot soccer, and adversarial self-play training all draw on this field.","example":"MAPPO controls each unit in a battle with its own PPO-based agent, and reaches results on par with or better than off-policy methods on benchmarks such as the StarCraft Multi-Agent Challenge (SMAC) and Google Research Football.","related":["Reinforcement Learning","Multi-robot Collaboration","Self-Play","Swarm Intelligence","Partially Observable Markov Decision Process","Proximal Policy Optimization"]},{"id":"self-play","category":"training","sec":5,"tier":3,"sources":[{"title":"Google DeepMind Blog: AlphaGo Zero: Starting from scratch","url":"https://deepmind.google/discover/blog/alphago-zero-starting-from-scratch/"},{"title":"Haarnoja et al. 2023: Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning","url":"https://arxiv.org/abs/2304.13653"}],"as_of":"","related_ids":["multi-agent-reinforcement-learning","reinforcement-learning","monte-carlo-tree-search","curriculum-learning","population-based-training","op3-soccer"],"name":"Self-Play","alt":"自博弈","abbr":"","aliases":["Self-Competition"],"one_liner":"Having an agent compete against itself, or past versions of itself, to keep improving through the outcomes.","explanation":"Self-play is a training approach within multi-agent reinforcement learning where the opponent is not a human or a fixed script but the agent's own current or past self. As the agent gets stronger, its opponent gets stronger in lockstep, which amounts to an automatically generated curriculum of rising difficulty, and it needs no human match data at all. The most famous example is DeepMind's 2017 AlphaGo Zero: starting from random moves and using only self-play plus Monte Carlo tree search, after three days of training it beat the earlier AlphaGo version that had defeated the human champion, 100 games to 0. In robotics, adversarial tasks commonly use it: DeepMind trained the small humanoid robot OP3 to play one-on-one soccer by first training standing-up and scoring skills separately, distilling both into one policy, and then having it compete against snapshots of its own past self. Playing only against the very latest version of itself tends to cause cycling, so opponents are usually drawn from a pool of historical versions instead.","example":"In OP3's soccer training, opponents were drawn from a pool of the agent's own periodically saved historical snapshots; an ablation study showed that an agent trained without self-play performed worse even against a fixed opponent.","related":["Multi-Agent Reinforcement Learning","Reinforcement Learning","Monte Carlo Tree Search","Curriculum Learning","Population-Based Training","OP3 Soccer (DeepMind)"]},{"id":"sparse-reward","category":"training","sec":6,"tier":2,"sources":[{"title":"Hindsight Experience Replay (arXiv 1707.01495)","url":"https://arxiv.org/abs/1707.01495"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/html/2410.21845"}],"as_of":"","related_ids":["dense-reward","reward-shaping","hindsight-experience-replay","success-detector","exploration-vs-exploitation","hil-serl"],"name":"Sparse Reward","alt":"稀疏奖励","abbr":"","aliases":["Binary Reward","Success Reward"],"one_liner":"A reward given only at a few key moments, such as task completion, with zero reward the rest of the time.","explanation":"Sparse reward means the environment gives zero reward almost all the time, with a signal only when the goal is reached — the most typical form is a binary reward: 1 for success, 0 otherwise. The upside is that it's simple to define and hard to game, since it directly reflects whether the task was actually completed; the difficulty is that an agent exploring randomly rarely stumbles into success by chance, so it can go a long time with no learning signal at all, and this gets worse as tasks get longer. Common countermeasures include reward shaping to add intermediate hints (the opposite approach is called dense reward), using human demonstrations to supply examples of success, hindsight experience replay (HER, which relabels a failed trajectory's actual endpoint as the goal, turning it into a “success” example), and curriculum learning from easy to hard. Real-robot reinforcement learning often trains a success detector to generate this kind of 0/1 reward automatically.","example":"HIL-SERL collects about 200 success images and 1,000 failure images per task through teleoperation and trains a binary classifier as the reward: a positive reward is given only when the classifier judges the task complete, and 0 otherwise; the classifier's accuracy on held-out data is typically above 95%.","related":["Dense Reward","Reward Shaping","Hindsight Experience Replay","Success Detector","Exploration vs. Exploitation","HIL-SERL"]},{"id":"dense-reward","category":"training","sec":6,"tier":2,"sources":[{"title":"Gymnasium-Robotics: Fetch Reach","url":"https://robotics.farama.org/envs/fetch/reach/"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/html/2109.11978"}],"as_of":"","related_ids":["sparse-reward","reward-shaping","reward-function","reward-engineering","reward-hacking","eureka"],"name":"Dense Reward","alt":"稠密奖励","abbr":"","aliases":[],"one_liner":"A reward design that gives informative feedback at nearly every step, showing whether the agent is getting closer to or further from the goal.","explanation":"Dense reward means a reinforcement-learning environment gives an informative reward at almost every timestep, as opposed to sparse reward, which pays off only on task success. It's usually built by a human, breaking the task into several weighted terms added together — distance to the goal, velocity-tracking error, an energy penalty, and so on. The benefit is that the agent knows which direction to improve in from the very start, so learning is faster; legged-locomotion reinforcement learning almost always uses dense reward. The cost is that it's laborious to design, and if the weights aren't tuned well, the agent will find a shortcut that racks up score without actually accomplishing the real goal — reward hacking. Projects like Eureka try to have a large language model write this kind of reward-function code automatically.","example":"Gymnasium-Robotics' FetchReach task has two versions: the sparse version gives −1 per step while the end effector is more than 5 cm from the target and 0 on arrival, while the dense version, FetchReachDense, gives the negative Euclidean distance to the target at every step. legged_gym's quadruped-walking reward likewise combines velocity tracking, a joint-torque penalty, a collision penalty, and more, computed at every simulation step.","related":["Sparse Reward","Reward Shaping","Reward Function","Reward Engineering","Reward Hacking","Eureka"]},{"id":"reward-shaping","category":"training","sec":6,"tier":2,"sources":[{"title":"Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping (Ng, Harada, Russell, ICML 1999)","url":"https://ai.stanford.edu/~ang/papers/shaping-icml99.pdf"},{"title":"Reward Hacking in Reinforcement Learning (Lilian Weng, 2024)","url":"https://lilianweng.github.io/posts/2024-11-28-reward-hacking/"}],"as_of":"","related_ids":["sparse-reward","dense-reward","reward-function","reward-engineering","reward-hacking","eureka"],"name":"Reward Shaping","alt":"奖励塑形","abbr":"","aliases":["Potential-based Reward Shaping"],"one_liner":"Adding intermediate guidance rewards on top of a task's original reward, so the agent learns the task faster.","explanation":"Reward shaping means adding extra reward for intermediate progress on top of the task's own reward (for instance, only +1 on success), such as scoring higher the closer a robot arm gets to its target. It mainly addresses the problem that under a sparse reward, an agent can go a very long time with no feedback and simply fail to learn. The risk is that shaping it wrong changes what the “optimal behavior” actually is: a classic example is a simulated cyclist that learned to circle endlessly near a goal because getting closer was rewarded and moving away wasn't penalized. Ng, Harada, and Russell proved in a 1999 ICML paper that as long as the extra reward is written as the difference of a potential function, F = γΦ(s′) − Φ(s) (Φ scores each state, γ is the discount factor), the optimal policy is provably unchanged — this is called potential-based reward shaping. In embodied AI, a legged-locomotion reward is usually a weighted combination of velocity tracking, posture, energy use, and foot airtime, which is essentially a large amount of hand-tuned shaping; projects like Eureka try to have a large model write this kind of reward code automatically.","example":"Training a robot arm to push a block to a target: the raw reward gives +1 only when the block arrives; after shaping, each step also rewards “how much closer did the gripper get to the block” and “how much closer did the block get to the target,” giving the policy a learning signal from early in training.","related":["Sparse Reward","Dense Reward","Reward Function","Reward Engineering","Reward Hacking","Eureka"]},{"id":"reward-engineering","category":"training","sec":6,"tier":2,"sources":[{"title":"Eureka: Human-Level Reward Design via Coding Large Language Models","url":"https://arxiv.org/abs/2310.12931"},{"title":"legged_gym: legged_robot_config.py (reward scales)","url":"https://raw.githubusercontent.com/leggedrobotics/legged_gym/master/legged_gym/envs/base/legged_robot_config.py"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning","url":"https://arxiv.org/abs/2109.11978"}],"as_of":"","related_ids":["reward-function","reward-shaping","sparse-reward","dense-reward","reward-hacking","eureka"],"name":"Reward Engineering","alt":"奖励工程","abbr":"","aliases":["Reward Design"],"one_liner":"Designing and debugging a reward function for reinforcement learning so the robot actually learns the intended behavior.","explanation":"Reward engineering covers the whole job of designing a reward function for a reinforcement-learning task: which terms to include, how heavily to weight each one, when to award them, and how to prevent the agent from gaming them. With only a sparse reward like “+1 on success,” the agent rarely stumbles onto positive feedback by chance, so dense intermediate rewards are often added instead — reward shaping — such as scoring higher the closer the agent gets to the goal. A robot locomotion reward often has a dozen or more terms: rewarding tracking of a commanded velocity while penalizing excessive torque, jerky motion, and body collisions, with weights that need extensive trial and error — get it wrong and reward hacking follows. Recent work also has large models write rewards automatically: Eureka has GPT-4 iteratively generate reward code, beating reward functions written by human experts on 83% of 29 tasks.","example":"legged_gym's default quadruped-walking reward: tracking linear-velocity commands (weight 1.0), tracking angular velocity (0.5), penalizing vertical velocity (−2.0), penalizing joint torque and acceleration, penalizing the rate of change of actions (−0.01) and body collisions (−1), and rewarding time each foot spends airborne (1.0).","related":["Reward Function","Reward Shaping","Sparse Reward","Dense Reward","Reward Hacking","Eureka"]},{"id":"reward-hacking","category":"training","sec":6,"tier":2,"sources":[{"title":"Concrete Problems in AI Safety (Amodei et al., 2016)","url":"https://arxiv.org/abs/1606.06565"},{"title":"Google DeepMind: Specification gaming: the flip side of AI ingenuity","url":"https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/"},{"title":"Lilian Weng: Reward Hacking in Reinforcement Learning","url":"https://lilianweng.github.io/posts/2024-11-28-reward-hacking/"}],"as_of":"","related_ids":["reward-engineering","reward-shaping","reward-model","reinforcement-learning-from-human-feedback","safe-reinforcement-learning"],"name":"Reward Hacking","alt":"奖励黑客","abbr":"","aliases":["Reward Exploitation","Specification Gaming"],"one_liner":"An agent finding a loophole in the reward function to score high without actually accomplishing what the designer intended.","explanation":"Reward hacking means a reinforcement-learning agent exploits a flaw in the reward function to earn high reward without truly learning the intended task. Amodei and colleagues' 2016 paper “Concrete Problems in AI Safety” lists it as one of five concrete problems in AI safety; DeepMind called the same phenomenon “specification gaming” in 2020: satisfying the literal specification of the goal without achieving what was actually meant. The root cause is that the reward is only an approximation of the real objective, and the harder that approximation gets optimized, the more the gap tends to widen — a pattern known as Goodhart's law — with more capable agents typically better at finding the loopholes. In robot simulation this often shows up as exploiting a quirk in the physics engine; in RLHF, a model may learn to flatter the reward model instead. Responses include repeatedly auditing the reward, watching rollout videos by hand, and adding constraint terms.","example":"A case DeepMind lists: told to stack a red block on a blue one with reward based on the red block's bottom-face height, a robot arm learned to simply flip the red block over instead; in another case, a simulated robot hand learned to hover between the camera and the object so it merely looked like a successful grasp.","related":["Reward Engineering","Reward Shaping","Reward Model","Reinforcement Learning from Human Feedback","Safe Reinforcement Learning"]},{"id":"inverse-reinforcement-learning","category":"training","sec":6,"tier":2,"sources":[{"title":"Apprenticeship learning（Wikipedia，含 Inverse reinforcement learning 一节）","url":"https://en.wikipedia.org/wiki/Apprenticeship_learning"},{"title":"A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress (arXiv 1806.06877)","url":"https://arxiv.org/abs/1806.06877"}],"as_of":"","related_ids":["reinforcement-learning","reward-function","imitation-learning","generative-adversarial-imitation-learning","adversarial-motion-priors","behavior-cloning"],"name":"Inverse Reinforcement Learning","alt":"逆强化学习","abbr":"IRL","aliases":["IRL"],"one_liner":"Working backward from an expert's demonstrated behavior to infer the reward function it's implicitly optimizing.","explanation":"Inverse reinforcement learning runs reinforcement learning in reverse: ordinary RL starts from a given reward function (the rule that scores behavior) and learns a policy, while IRL instead observes expert demonstrations and infers what reward the expert must be optimizing. Stuart Russell posed this problem in 1998, Andrew Ng and Russell gave the first algorithms in 2000, and Abbeel and Ng applied it to “apprenticeship learning” in 2004: infer the reward first, then use reinforcement learning to find a policy from it. Its value is that many tasks have rewards that are hard to hand-write, and a learned reward often transfers to new environments better than copying actions directly does. The difficulty is that the same behavior can be explained by many different reward functions, so extra assumptions such as maximum entropy are needed to resolve the ambiguity. Later methods like generative adversarial imitation learning (GAIL) and adversarial motion priors (AMP) continue this line of thinking.","example":"Abbeel, Coates, and Ng used apprenticeship learning to teach an autonomous helicopter aerobatic maneuvers — flips, loops, autorotation landings — by learning from a human pilot's demonstrations, instead of hand-writing a reward function.","related":["Reinforcement Learning","Reward Function","Imitation Learning","Generative Adversarial Imitation Learning","Adversarial Motion Priors","Behavior Cloning"]},{"id":"generative-adversarial-imitation-learning","category":"training","sec":6,"tier":3,"sources":[{"title":"Generative Adversarial Imitation Learning (Ho & Ermon, arXiv 1606.03476)","url":"https://arxiv.org/abs/1606.03476"}],"as_of":"","related_ids":["inverse-reinforcement-learning","imitation-learning","behavior-cloning","generative-adversarial-network","adversarial-motion-priors","trust-region-policy-optimization"],"name":"Generative Adversarial Imitation Learning","alt":"生成对抗模仿学习","abbr":"GAIL","aliases":["GAIL"],"one_liner":"Using a discriminator to tell expert actions from policy actions, forcing the policy to act more and more like the expert.","explanation":"Introduced by Stanford's Jonathan Ho and Stefano Ermon in 2016. The traditional route was two separate steps: recover the expert's reward function with inverse reinforcement learning (working backward from demonstrations), then train a policy with reinforcement learning on that reward — slow and indirect. GAIL borrows the structure of a generative adversarial network instead: a discriminator learns to tell whether a state-action pair came from an expert demonstration or from the current policy, and the policy treats the discriminator's output as its reward, updating with TRPO (a policy-gradient algorithm that limits how much the policy can change per step) until the discriminator can no longer tell the difference. It requires repeated interaction with the environment, so it's mostly used in simulation, but it's less prone to compounding error (small mistakes snowballing) than behavior cloning. Adversarial motion priors (AMP), commonly used for humanoid robots and character animation, grew out of this same adversarial-imitation idea.","example":"The original paper, given only a handful of expert trajectories and no reward function, trains GAIL policies on MuJoCo-simulated tasks like humanoid walking that mostly reach over 70% of expert-level performance.","related":["Inverse Reinforcement Learning","Imitation Learning","Behavior Cloning","Generative Adversarial Network","Adversarial Motion Priors","Trust Region Policy Optimization"]},{"id":"adversarial-motion-priors","category":"training","sec":6,"tier":3,"sources":[{"title":"AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control (arXiv 2104.02180)","url":"https://arxiv.org/abs/2104.02180"},{"title":"Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions (arXiv 2203.15103)","url":"https://arxiv.org/abs/2203.15103"}],"as_of":"","related_ids":["generative-adversarial-imitation-learning","reward-engineering","deepmimic","ase","motion-tracking","sim-to-real-transfer"],"name":"Adversarial Motion Priors","alt":"对抗运动先验","abbr":"AMP","aliases":["AMP","AMP Style Reward"],"one_liner":"Using a discriminator that judges how much motion resembles motion-capture data, and turning that resemblance into a reward for natural movement.","explanation":"AMP was introduced by Xue Bin Peng, Pieter Abbeel, Sergey Levine, Angjoo Kanazawa, and colleagues in 2021, originally for simulated character animation. It borrows from generative adversarial imitation learning: a discriminator is trained on pairs of consecutive-frame states, judging whether a snippet of motion came from a reference motion-capture dataset or from the policy; the policy treats how well it fools the discriminator as a style reward, added to the task reward (such as moving forward at a commanded speed) and optimized together with reinforcement learning. The benefit is not having to hand-write an imitation objective or pick specific motion clips to track — an unstructured pile of mocap clips is enough to make the motion look natural. In 2022, Escontrela and colleagues applied it to the Unitree A1 quadruped, learning a natural, energy-efficient gait from only about 4.5 seconds of German shepherd motion-capture data, and it's often used as a substitute for laborious reward engineering.","example":"Escontrela and colleagues set the style-reward weight to 0.65 and the task-reward weight to 0.35 on the A1; the robot dog learned a gait resembling a real dog's, naturally switching gaits with speed, and transferred directly to the real robot.","related":["Generative Adversarial Imitation Learning","Reward Engineering","DeepMimic","ASE","Motion Tracking","Sim-to-Real Transfer"]},{"id":"reward-model","category":"training","sec":6,"tier":2,"sources":[{"title":"Deep reinforcement learning from human preferences (Christiano et al., arXiv 1706.03741)","url":"https://arxiv.org/abs/1706.03741"},{"title":"Training language models to follow instructions with human feedback (InstructGPT, arXiv 2203.02155)","url":"https://arxiv.org/abs/2203.02155"},{"title":"Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons (arXiv 2603.02115)","url":"https://arxiv.org/abs/2603.02115"}],"as_of":"2026-05","related_ids":["reward-function","reinforcement-learning-from-human-feedback","progress-reward-model","success-detector","vlm-as-reward","reward-hacking"],"name":"Reward Model","alt":"奖励模型","abbr":"RM","aliases":["RM","Learned Reward Model"],"one_liner":"A trained scoring network that outputs a reward telling how good a given behavior or outcome is.","explanation":"A reward model is a trained neural network: it takes a piece of behavior — a robot trajectory video, a large model's answer — as input and outputs a reward score, standing in for a hand-written reward function. A landmark example is Christiano and colleagues' 2017 work, which trained a reward model from human preferences between pairs of trajectories, teaching a simulated robot new behaviors from roughly an hour of human feedback; OpenAI's InstructGPT (2022) put this inside the RLHF (reinforcement learning from human feedback) pipeline, making it a standard step in post-training large models. Many robot tasks resist a precise, hand-written reward (was the shirt actually folded neatly?), and a reward model gives reinforcement learning, failure detection, and data filtering a usable signal instead. Common forms include success detectors, progress reward models, and simply having a VLM score the outcome; the risk is that a policy learns to exploit the model's blind spots — reward hacking.","example":"Robometer (RSS 2026) trains on RBM-1M, a dataset of over 1 million trajectories including many failures and suboptimal attempts, learning both “task progress at each frame” and “which of two trajectories on the same task is better” at once, producing a general-purpose reward model usable across many robots.","related":["Reward Function","Reinforcement Learning from Human Feedback","Progress Reward Model","Success Detector","VLM-as-Reward","Reward Hacking"]},{"id":"progress-reward-model","category":"training","sec":6,"tier":3,"sources":[{"title":"Ma et al. 2024: Vision Language Models are In-Context Value Learners (GVL)","url":"https://arxiv.org/abs/2411.04549"},{"title":"Liang et al. 2026: Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons","url":"https://arxiv.org/abs/2603.02115"},{"title":"Ayalew et al. 2024: PROGRESSOR","url":"https://arxiv.org/abs/2411.17764"}],"as_of":"2026-03","related_ids":["reward-model","sparse-reward","dense-reward","success-detector","generative-value-learning","robometer"],"name":"Progress Reward Model","alt":"进度奖励模型","abbr":"","aliases":["Progress Estimator","Task Progress Prediction"],"one_liner":"A model that looks at the current frame and estimates how much of a task is done, used as a reward.","explanation":"A progress reward model is a class of robot reward model: given a task instruction and the current frame, sometimes also the start frame or recent history, it outputs how far along the task is, usually normalized to between 0 and 1. Real manipulation tasks often only offer a sparse succeeded-or-not reward, which reinforcement learning struggles to learn from; a progress score can serve as a dense reward instead, letting the policy know at every step whether it is getting closer to or farther from the goal, and it can also be used to filter data or judge success. There are two main approaches: having a vision-language model estimate progress zero-shot, as in Google DeepMind's GVL, or training one specifically on large amounts of trajectory data, as in 2026's Robometer. The risk is reward hacking, where the policy learns to make the footage merely look like it is progressing.","example":"GVL shuffles the frame order of a robot video and asks a VLM to estimate the completion percentage of each frame; with no training at all, it gives usable progress values across more than 300 real tasks, and is used for data filtering and success detection.","related":["Reward Model","Sparse Reward","Dense Reward","Success Detector","Generative Value Learning (GVL)","Robometer"]},{"id":"success-detector","category":"training","sec":6,"tier":3,"sources":[{"title":"Du et al. 2023: Vision-Language Models as Success Detectors","url":"https://arxiv.org/abs/2303.07280"},{"title":"Luo et al. 2024: Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/abs/2410.21845"}],"as_of":"","related_ids":["reward-model","vlm-as-reward","sparse-reward","progress-reward-model","real-world-reinforcement-learning","hil-serl"],"name":"Success Detector","alt":"成功检测器","abbr":"","aliases":["Success Classifier","Reward Classifier"],"one_liner":"A model that judges whether a robot completed a task in a given episode, often used to give reward in RL.","explanation":"A success detector takes an observation, usually a camera image, sometimes with the task instruction, and outputs success or failure, or a probability of success. Reinforcement learning, automatic data collection, and automatic evaluation all need to know whether a task got done, and the real world has no simulator-style ground truth for that, so people either watch it themselves or train a model to watch it. Two approaches are common: training a binary classifier for a single task, as HIL-SERL does, teleoperating roughly 200 successful and 1,000 failed examples per task; or directly querying a vision-language model, as in DeepMind's 2023 SuccessVQA, which reframes the judgment as the visual question “did the task complete?” and fine-tunes it on Flamingo. It provides a sparse reward, and misjudgments can be exploited by the policy, a failure mode called reward hacking.","example":"HIL-SERL trains a binary classifier per task on wrist and side-camera images to judge success, and only gives positive reward when it judges success; the paper reports the classifier's accuracy on held-out evaluation generally exceeds 95%.","related":["Reward Model","VLM-as-Reward","Sparse Reward","Progress Reward Model","Real-World Reinforcement Learning","HIL-SERL"]},{"id":"vlm-as-reward","category":"training","sec":6,"tier":3,"sources":[{"title":"Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning (ICLR 2024)","url":"https://arxiv.org/abs/2310.12921"},{"title":"RoboCLIP: One Demonstration is Enough to Learn Robot Policies","url":"https://arxiv.org/abs/2310.07899"},{"title":"RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback (ICML 2024)","url":"https://arxiv.org/abs/2402.03681"}],"as_of":"2024-07","related_ids":["reward-model","vision-language-model","clip","success-detector","progress-reward-model","generative-value-learning"],"name":"VLM-as-Reward","alt":"VLM 作奖励模型","abbr":"","aliases":["VLM Reward","VLM Reward Model","VLM-RM","Vision-Language Models as Reward Models"],"one_liner":"Having a vision-language model watch footage against a task description and score the robot's performance as a reward.","explanation":"This approach uses a vision-language model, a large model that can both look at images and read text, as a reinforcement-learning reward model: given a task description and footage the robot captured, it judges how well the robot is doing. Representative work includes 2023's VLM-RM, which uses CLIP's image-text similarity as the reward; RoboCLIP, which compares an agent's video to a demonstration video; and ICML 2024's RL-VLM-F, which has a VLM give a preference between two images and then learns a reward function from those preferences. It targets the problem that reward functions are hard to hand-write: a task like folding clothes is very hard to define “success” for with a formula. Its limitations are that VLMs are weak at spatial reasoning and give noisy scores, so a policy can learn to exploit the gaps, a failure mode called reward hacking. It is commonly used for RL fine-tuning, success detection, and data filtering.","example":"VLM-RM used only a single English sentence describing a target pose, with CLIP's similarity between that sentence and a rendered simulation frame as the reward, and taught a simulated humanoid to kneel, do a split, and sit in lotus position with no hand-written reward function at all; the paper also found that a bigger VLM makes a better reward model.","related":["Reward Model","Vision-Language Model","CLIP","Success Detector","Progress Reward Model","Generative Value Learning (GVL)"]},{"id":"intrinsic-motivation","category":"training","sec":6,"tier":3,"sources":[{"title":"Curiosity-driven Exploration by Self-supervised Prediction (ICM, arXiv:1705.05363)","url":"https://arxiv.org/abs/1705.05363"},{"title":"Exploration by Random Network Distillation (RND, arXiv:1810.12894)","url":"https://arxiv.org/abs/1810.12894"}],"as_of":"","related_ids":["exploration-vs-exploitation","sparse-reward","unsupervised-skill-discovery","reward-shaping","reinforcement-learning","active-exploration"],"name":"Intrinsic Motivation","alt":"内在奖励","abbr":"","aliases":["Curiosity-Driven Exploration","Intrinsic Reward"],"one_liner":"A reward the agent generates for itself from novel or hard-to-predict states, used to drive exploration.","explanation":"Intrinsic reward is a reward signal computed internally by the agent rather than coming from the task itself, usually measuring how novel a state is or how poorly the agent can predict it; it is typically added on top of the environment's extrinsic reward. It mainly addresses exploration under sparse rewards, when task reward shows up so rarely that random trial and error almost never stumbles onto it. Representative work includes Pathak and colleagues' 2017 ICM, which treats the error in predicting the consequences of one's own actions as curiosity, and OpenAI's 2018 RND, which uses the error in predicting a fixed random network's output as the reward — becoming the first method to beat average human performance on the Atari game Montezuma's Revenge without demonstrations. In robotics it is commonly used for exploration and for unsupervised skill discovery.","example":"In Super Mario Bros., ICM gives the agent no game score at all, only a reward for failing to predict the next frame's features, and the agent still actively explores deeper into the level.","related":["Exploration vs. Exploitation","Sparse Reward","Unsupervised Skill Discovery","Reward Shaping","Reinforcement Learning","Active Exploration"]},{"id":"unsupervised-skill-discovery","category":"training","sec":6,"tier":3,"sources":[{"title":"Eysenbach et al. 2018: Diversity is All You Need: Learning Skills without a Reward Function","url":"https://arxiv.org/abs/1802.06070"},{"title":"Sharma et al. 2020: Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning","url":"https://arxiv.org/abs/2004.12974"},{"title":"Park, Rybkin, Levine 2023: METRA: Scalable Unsupervised RL with Metric-Aware Abstraction (ICLR 2024)","url":"https://arxiv.org/abs/2310.08887"}],"as_of":"","related_ids":["intrinsic-motivation","hierarchical-reinforcement-learning","exploration-vs-exploitation","behavior-foundation-model","bfm-zero","entropy-regularization"],"name":"Unsupervised Skill Discovery","alt":"无监督技能发现","abbr":"","aliases":["DIAYN","Reward-Free Skill Learning"],"one_liner":"With no task reward, letting an agent practice into a set of distinguishable skills it can call on later.","explanation":"Unsupervised skill discovery lets an agent spontaneously learn a variety of different behaviors with no external reward at all. The best known method, DIAYN (“Diversity is All You Need”), was proposed by Eysenbach, Levine, and colleagues in 2018: the policy is given a skill number as input, and a discriminator is trained at the same time to guess which skill was used just from the states reached, so the intrinsic reward rises the more accurately it can be guessed, plus a max-entropy term to encourage varied actions; underneath, this maximizes the mutual information between skill and state. Simulated robots have spontaneously learned to walk, hop, and more this way. Skills can serve as a pretraining initialization, or be composed by a higher-level policy to solve sparse-reward tasks. METRA (ICLR 2024) instead learns skills in a latent space that preserves temporal distance, easing the weak-exploration problem that mutual-information methods tend to have.","example":"Sharma and colleagues (2020) turned the skill-discovery algorithm DADS into an off-policy version and, with no reward and no demonstrations at all, had a real quadruped robot learn distinct gaits and headings, then chained them together with model-predictive control to perform navigation.","related":["Intrinsic Motivation","Hierarchical Reinforcement Learning","Exploration vs. Exploitation","Behavior Foundation Model","BFM-Zero","Entropy Regularization"]},{"id":"goal-conditioned-reinforcement-learning","category":"training","sec":6,"tier":3,"sources":[{"title":"Goal-Conditioned Reinforcement Learning: Problems and Solutions (IJCAI 2022 survey)","url":"https://arxiv.org/abs/2201.08299"},{"title":"Universal Value Function Approximators (ICML 2015)","url":"https://proceedings.mlr.press/v37/schaul15.html"}],"as_of":"","related_ids":["goal-conditioned-policy","hindsight-experience-replay","sparse-reward","value-function","goal-conditioned-behavior-cloning","hierarchical-reinforcement-learning"],"name":"Goal-Conditioned Reinforcement Learning","alt":"目标条件强化学习","abbr":"GCRL","aliases":["GCRL"],"one_liner":"Reinforcement learning where both the policy and the value function take the goal as input, letting one model reach many goals.","explanation":"Ordinary reinforcement learning usually learns just one fixed task; goal-conditioned reinforcement learning writes the policy as π(a|s,g), so the same network carries out different tasks depending on the goal g, with the reward usually just “did it reach the goal or not.” A landmark example is DeepMind's 2015 Universal Value Function Approximators (UVFA), which extends the value function to V(s,g), able to generalize to goals never seen during training. Its biggest difficulty is sparse reward: most attempts never reach the goal, so there's no learning signal, which is why it's often paired with hindsight experience replay (relabeling a failed trajectory as having reached whatever state it actually ended up at). Goals can be coordinates, images, or language, and robot grasping, pushing, and navigation are all commonly framed this way; the low-level policy in hierarchical reinforcement learning is also often goal-conditioned.","example":"A robot arm pushing a block: the goal is any position on the table, and success means the block ends up near that goal; the same policy learns to push the block to many different positions.","related":["Goal-conditioned Policy","Hindsight Experience Replay","Sparse Reward","Value Function","Goal-Conditioned Behavior Cloning","Hierarchical Reinforcement Learning"]},{"id":"hindsight-experience-replay","category":"training","sec":6,"tier":3,"sources":[{"title":"Hindsight Experience Replay (arXiv 1707.01495)","url":"https://arxiv.org/abs/1707.01495"}],"as_of":"","related_ids":["sparse-reward","goal-conditioned-reinforcement-learning","experience-replay","hindsight-relabeling","deep-deterministic-policy-gradient","off-policy"],"name":"Hindsight Experience Replay","alt":"后见之明经验回放","abbr":"HER","aliases":["HER"],"one_liner":"Relabeling a failed attempt as having succeeded at whatever goal it actually reached, so even sparse rewards can be learned from.","explanation":"Introduced by OpenAI's Andrychowicz and colleagues in 2017 (NIPS 2017). In goal-conditioned tasks with only a binary success/failure reward, a robot almost never succeeds early on, so the replay buffer fills up with nothing but failures and there's nothing to learn from. HER's fix is to store a second copy of every failed trajectory, with its goal swapped out for whatever state that trajectory actually reached (such as its final state, or a state at some later point) — which turns it into a successful example. It can be attached after any off-policy algorithm (one that can learn from old data, such as DDPG), and in effect works like an automatically generated curriculum. The original paper completed pushing, sliding, and pick-and-place tasks on a Fetch arm and deployed a policy trained in simulation to the real robot. This hindsight-relabeling idea was later widely borrowed by methods like goal-conditioned behavior cloning.","example":"A robot arm tries to push a block to a red dot but ends up pushing it 10 cm to the side instead; HER rewrites that trajectory's goal to be that 10-cm-off position, turning it into a successful experience.","related":["Sparse Reward","Goal-Conditioned Reinforcement Learning","Experience Replay","Hindsight Relabeling","Deep Deterministic Policy Gradient","Off-Policy"]},{"id":"hierarchical-reinforcement-learning","category":"training","sec":6,"tier":3,"sources":[{"title":"FeUdal Networks for Hierarchical Reinforcement Learning (arXiv 1703.01161)","url":"https://arxiv.org/abs/1703.01161"},{"title":"Data-Efficient Hierarchical Reinforcement Learning (HIRO, NeurIPS 2018)","url":"https://arxiv.org/abs/1805.08296"}],"as_of":"","related_ids":["hierarchical-architecture","dual-system-architecture","skill-primitive","long-horizon-task","goal-conditioned-reinforcement-learning","credit-assignment"],"name":"Hierarchical Reinforcement Learning","alt":"分层强化学习","abbr":"HRL","aliases":["HRL"],"one_liner":"Reinforcement learning split into layers: a high level sets sub-goals, and a low level executes the concrete actions to reach them.","explanation":"If a long task is learned directly at the level of low-level actions, reward only ever shows up after a very long delay, making credit assignment (figuring out which step actually caused the outcome) extremely hard. Hierarchical reinforcement learning has a high-level policy pick sub-goals or skills at a lower frequency, while a low-level policy outputs concrete actions at a higher frequency to carry them out. Classic frameworks include Sutton and colleagues' 1999 options framework and Dayan and Hinton's feudal reinforcement learning; in the deep-learning era there's DeepMind's 2017 FeUdal Networks (a manager sets goals, a worker outputs actions) and Nachum, Levine, and colleagues' 2018 HIRO (the high level hands the low level a goal, with an off-policy correction for reusing old data). Today's embodied-AI architectures that combine large-model planning with low-level skills, or fast-slow dual systems, follow a similar layered logic, though they aren't always trained with reinforcement learning.","example":"HIRO, on tasks like a simulated quadruped “ant” robot navigating a maze, has the high level issue a desired state as a sub-goal every fixed number of steps, while the low level controls each leg to reach it.","related":["Hierarchical Architecture","Dual-System Architecture (System 1 / System 2)","Skill Primitive","Long-horizon Task","Goal-Conditioned Reinforcement Learning","Credit Assignment"]},{"id":"online-reinforcement-learning","category":"training","sec":7,"tier":2,"sources":[{"title":"Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (Fig. 1)","url":"https://arxiv.org/html/2005.01643"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/abs/2410.21845"}],"as_of":"","related_ids":["offline-reinforcement-learning","offline-to-online-reinforcement-learning","real-world-reinforcement-learning","reinforcement-fine-tuning","on-policy","off-policy"],"name":"Online Reinforcement Learning","alt":"在线强化学习","abbr":"Online RL","aliases":["Online RL"],"one_liner":"Learning while interacting with the environment, continuously collecting new data with the latest policy during training.","explanation":"This is reinforcement learning's most classic setup: the agent interacts with the environment using its current policy, updates the policy once new data comes in, then keeps collecting with the updated policy, repeating the cycle. It covers both on-policy methods that use only the freshest data, like PPO, and off-policy methods that store history in a replay buffer, like SAC; what defines it is that training keeps getting new data throughout, which is exactly the line separating it from offline reinforcement learning. Online learning can explore actively and improve on its own mistakes, but it demands a lot of interaction: simulation speeds this up through parallelism, while real robots are limited by safety, time, and how often a scene can be reset. A common recipe now is to start from demonstrations or offline data and then fine-tune with online reinforcement learning — for example, using RL to further improve a VLA.","example":"HIL-SERL runs online reinforcement learning on a real robot with a human correcting it as needed, bringing an arm to near-perfect success on tasks like precision assembly and bimanual coordination within 1 to 2.5 hours.","related":["Offline Reinforcement Learning","Offline-to-Online Reinforcement Learning","Real-World Reinforcement Learning","Reinforcement Fine-Tuning (RL Fine-Tuning)","On-Policy","Off-Policy"]},{"id":"offline-reinforcement-learning","category":"training","sec":7,"tier":2,"sources":[{"title":"Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (Levine et al., 2020)","url":"https://arxiv.org/abs/2005.01643"},{"title":"Conservative Q-Learning for Offline Reinforcement Learning","url":"https://arxiv.org/abs/2006.04779"}],"as_of":"","related_ids":["online-reinforcement-learning","off-policy","conservative-q-learning","implicit-q-learning","offline-to-online-reinforcement-learning","d4rl"],"name":"Offline Reinforcement Learning","alt":"离线强化学习","abbr":"Offline RL","aliases":["Offline RL","Batch Reinforcement Learning","Batch RL"],"one_liner":"Training a policy using only a fixed, previously collected dataset, with no further interaction with the environment during training.","explanation":"Earlier also called batch reinforcement learning. A 2020 survey by Sergey Levine and colleagues defines it as reinforcement learning that uses only pre-collected data, with no additional online data collection at all. The data can come from human teleoperation, an older policy, or deployment logs, and training never touches the real environment, which suits robotics, healthcare, and other settings where trial and error is expensive or dangerous. Unlike imitation learning, it makes use of the reward signal, so it has a chance of learning a policy better than the data itself from a mix of good and mediocre demonstrations. Its central difficulty is distribution shift: once the policy picks an action that never appears in the dataset, the Q-function's estimate for it tends to be inflated (extrapolation error), and these errors can compound. Conservative Q-learning (CQL), implicit Q-learning (IQL), and various policy-constraint methods were all designed to address this.","example":"Conservative Q-learning (CQL) adds a regularization term to ordinary Q-learning that deliberately pushes down the Q-values of actions not seen in the dataset, so the learned value becomes a lower bound on the true value and the policy isn't misled by inflated estimates.","related":["Online Reinforcement Learning","Off-Policy","Conservative Q-Learning","Implicit Q-Learning","Offline-to-Online Reinforcement Learning","D4RL"]},{"id":"extrapolation-error","category":"training","sec":7,"tier":3,"sources":[{"title":"Off-Policy Deep Reinforcement Learning without Exploration (BCQ, arXiv:1812.02900)","url":"https://arxiv.org/abs/1812.02900"},{"title":"Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (arXiv:2005.01643)","url":"https://arxiv.org/abs/2005.01643"}],"as_of":"","related_ids":["offline-reinforcement-learning","overestimation-bias","conservative-q-learning","implicit-q-learning","policy-constraint","out-of-distribution"],"name":"Extrapolation Error (OOD Actions in Offline RL)","alt":"外推误差","abbr":"","aliases":["Out-of-Distribution Action Problem","OOD Action Overestimation"],"one_liner":"The error in offline reinforcement learning that comes from a Q-network guessing wildly at the value of actions absent from the data.","explanation":"This concept was systematically laid out by Fujimoto, Meger, and Precup in their 2019 ICML paper introducing the BCQ algorithm. Offline reinforcement learning can only train on a fixed dataset, with no further environment interaction. Q-learning's update needs max_a Q(s′, a) over the next state, but the network's estimate for actions absent from the data (out-of-distribution actions) is just an extrapolation, and can be badly inflated; the policy then gravitates toward exactly those actions, the error compounds through repeated Bellman backups, and the resulting policy ends up poor. Online training can correct this by actually trying the action; offline training cannot. The main countermeasures are constraining the policy to stay close to the data's behavior (BCQ, policy constraints), pushing down the Q-values of out-of-distribution actions (CQL), and avoiding querying out-of-distribution actions altogether (IQL).","example":"In the BCQ paper, an offline DDPG agent trained on exactly the same batch of data as an online DDPG agent falls clearly behind on every task, with value estimates that are unstable or even diverge — the authors attribute this to extrapolation error.","related":["Offline Reinforcement Learning","Overestimation Bias","Conservative Q-Learning","Implicit Q-Learning","Policy Constraint","Out-of-Distribution"]},{"id":"policy-constraint","category":"training","sec":7,"tier":3,"sources":[{"title":"Wu, Tucker, Nachum 2019: Behavior Regularized Offline Reinforcement Learning (BRAC)","url":"https://arxiv.org/abs/1911.11361"},{"title":"Fujimoto & Gu 2021: A Minimalist Approach to Offline Reinforcement Learning (TD3+BC)","url":"https://arxiv.org/abs/2106.06860"},{"title":"Levine et al. 2020: Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems","url":"https://arxiv.org/abs/2005.01643"}],"as_of":"","related_ids":["offline-reinforcement-learning","extrapolation-error","kl-regularization","behavior-cloning","conservative-q-learning","advantage-weighted-regression"],"name":"Policy Constraint","alt":"策略约束","abbr":"","aliases":["Behavior Regularization","Behavior Constraint"],"one_liner":"In offline RL, keeping the new policy's actions close to what the behavior policy in the dataset actually did.","explanation":"Policy constraint is a major family of methods in offline reinforcement learning, which trains only on existing data with no further environment interaction. The Q-function, which estimates how much return an action will earn, tends to be overestimated for actions that never appear in the data, and a policy that chases those inflated values falls apart — a failure mode called extrapolation error. Policy constraint methods add a limit: the policy's output actions must stay close to the behavior policy that collected the data, enforced with a distance measure such as KL divergence or MMD, or by adding a behavior-cloning term to the objective. BCQ, BEAR, and TD3+BC all belong to this line; in 2019, Wu, Tucker, and Nachum's BRAC paper unified them under the name “behavior regularization.” Together with conservative-Q-learning-style methods, which instead suppress the value of unfamiliar actions, this forms one of the two main lines of offline RL.","example":"TD3+BC just adds a behavior-cloning term to the online algorithm TD3's policy update — pulling the output action toward the dataset's actions — and normalizes the states, yet it matched the performance of far more complex offline RL algorithms of its time.","related":["Offline Reinforcement Learning","Extrapolation Error (OOD Actions in Offline RL)","KL Regularization","Behavior Cloning","Conservative Q-Learning","Advantage-Weighted Regression"]},{"id":"conservative-q-learning","category":"training","sec":7,"tier":3,"sources":[{"title":"Conservative Q-Learning for Offline Reinforcement Learning (arXiv 2006.04779)","url":"https://arxiv.org/abs/2006.04779"},{"title":"NeurIPS 2020 Proceedings: Conservative Q-Learning for Offline Reinforcement Learning","url":"https://proceedings.neurips.cc/paper/2020/hash/0d2b2061826a5df3221116a5085a6052-Abstract.html"},{"title":"Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions (arXiv 2309.10150)","url":"https://arxiv.org/abs/2309.10150"}],"as_of":"","related_ids":["offline-reinforcement-learning","calibrated-q-learning","q-function","overestimation-bias","extrapolation-error","q-transformer"],"name":"Conservative Q-Learning","alt":"保守 Q 学习","abbr":"CQL","aliases":["CQL"],"one_liner":"Deliberately pushing down the Q-values of actions absent from the dataset, so offline reinforcement learning isn't misled by inflated estimates.","explanation":"CQL was introduced by Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine in 2020 (NeurIPS 2020), and is a landmark algorithm for offline reinforcement learning (using only a fixed dataset, with no further environment interaction). Ordinary Q-learning fails when applied offline as-is: the policy tends to favor actions absent from the data, whose Q-values (the estimated long-term return of taking a given action in a given state) were never corrected by real data and are often overestimated, so the policy drifts toward these inflated values. CQL adds a regularization term on top of the standard Bellman error: it pushes down the Q-values of actions the policy is likely to pick, while pushing up the Q-values of actions actually taken in the dataset, so the learned value becomes a lower bound on the true value. It's easy to bolt onto existing deep Q-learning or actor-critic algorithms, and the paper reports final returns often 2–5 times those of prior offline methods. Later methods like Cal-QL build directly on it.","example":"Google DeepMind's Q-Transformer (2023) represents a multi-task robot Q-function with a Transformer, training on offline real-robot data combining human demonstrations and autonomously collected data, using an adapted version of CQL's conservative regularizer.","related":["Offline Reinforcement Learning","Calibrated Q-Learning","Q-Function","Overestimation Bias","Extrapolation Error (OOD Actions in Offline RL)","Q-Transformer"]},{"id":"implicit-q-learning","category":"training","sec":7,"tier":3,"sources":[{"title":"Offline Reinforcement Learning with Implicit Q-Learning (arXiv 2110.06169)","url":"https://arxiv.org/abs/2110.06169"}],"as_of":"","related_ids":["offline-reinforcement-learning","advantage-weighted-regression","conservative-q-learning","extrapolation-error","offline-to-online-reinforcement-learning","q-function"],"name":"Implicit Q-Learning","alt":"隐式 Q 学习","abbr":"IQL","aliases":["IQL"],"one_liner":"An offline reinforcement-learning algorithm that only ever values actions already in the dataset, never querying an action it hasn't seen.","explanation":"Introduced by UC Berkeley's Kostrikov, Nair, and Levine in 2021. Offline reinforcement learning trains only on a fixed dataset, and its central difficulty is that the Q-function (an action's estimated value) tends to wildly overestimate actions absent from the data (extrapolation error), and the policy collapses as soon as it chases them. IQL's solution is to never evaluate a new action at all: it first uses expectile regression (a form of regression that leans toward the upper quantiles) to fit a state value V, from the dataset's own actions, that approximates the value of a close-to-the-best action; it then uses that V to do a temporal-difference update of Q; and finally extracts the policy with advantage-weighted regression (weighting behavior cloning by the size of the advantage). It's simple to implement, performs strongly on the D4RL offline benchmark, suits pretraining offline and then fine-tuning online, and is commonly used as a baseline in robot offline-reinforcement-learning work.","example":"Given a batch of historical robot-arm data mixing good and bad manipulation attempts, with no further environment interaction, IQL trains a policy that ends up better than the average quality of the data itself.","related":["Offline Reinforcement Learning","Advantage-Weighted Regression","Conservative Q-Learning","Extrapolation Error (OOD Actions in Offline RL)","Offline-to-Online Reinforcement Learning","Q-Function"]},{"id":"advantage-weighted-regression","category":"training","sec":7,"tier":3,"sources":[{"title":"Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning (arXiv 1910.00177)","url":"https://arxiv.org/abs/1910.00177"},{"title":"Offline Reinforcement Learning with Implicit Q-Learning (arXiv 2110.06169)","url":"https://arxiv.org/abs/2110.06169"}],"as_of":"","related_ids":["advantage-function","offline-reinforcement-learning","implicit-q-learning","behavior-cloning","policy-constraint","advantage-conditioning"],"name":"Advantage-Weighted Regression","alt":"优势加权回归","abbr":"AWR","aliases":["AWR","Advantage-Weighted Behavioral Cloning"],"one_liner":"Doing imitation learning weighted by each action's advantage, so higher-advantage actions get imitated more.","explanation":"AWR was introduced by Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine in 2019, aiming to do reinforcement learning using only supervised-learning regression steps. Each round has two steps: fit the value function by regression, then do weighted behavior cloning, where each action in the data is weighted by exp(advantage/β), so higher-advantage actions get imitated more heavily, with β a temperature coefficient. It can reuse old data from a replay buffer (off-policy), and can also learn from a fixed dataset alone; it needs only a few lines of code to implement, and supports both continuous and discrete actions. This “estimate the advantage, then imitate with weighting” style of policy extraction has been widely adopted by later offline reinforcement-learning methods — implicit Q-learning (IQL), for instance, uses advantage-weighted behavior cloning as its final policy-extraction step.","example":"Implicit Q-learning first learns a Q-function and a value function from offline data, then extracts the final policy with advantage-weighted behavior cloning: actions in the data with a larger positive advantage get a higher weight during imitation.","related":["Advantage Function","Offline Reinforcement Learning","Implicit Q-Learning","Behavior Cloning","Policy Constraint","Advantage Conditioning"]},{"id":"return-conditioning","category":"training","sec":7,"tier":3,"sources":[{"title":"Chen et al. 2021: Decision Transformer: Reinforcement Learning via Sequence Modeling","url":"https://arxiv.org/abs/2106.01345"},{"title":"Schmidhuber 2019: Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions","url":"https://arxiv.org/abs/1912.02875"},{"title":"Brandfonbrener et al. 2022: When does return-conditioned supervised learning work for offline reinforcement learning?","url":"https://arxiv.org/abs/2206.01079"}],"as_of":"","related_ids":["decision-transformer","advantage-conditioning","return","offline-reinforcement-learning","recap","goal-conditioned-policy"],"name":"Return Conditioning","alt":"回报条件化","abbr":"","aliases":["Return-Conditioned Policy","Return-Conditioned Supervised Learning (RCSL)"],"one_liner":"Feeding a policy the return you want it to achieve, so it acts toward that target return.","explanation":"Return conditioning rewrites reinforcement learning as supervised learning: during training, the remaining cumulative reward from each step onward in a trajectory, the return-to-go, is fed into the policy alongside the observation, and the policy is trained to imitate the action actually taken; at test time, a high target return is given, hoping the policy produces high-return behavior. Schmidhuber's 2019 Upside-Down RL and 2021's Decision Transformer are representative: the latter arranges return, state, and action into a sequence for a Transformer to predict actions from, matching mainstream methods on offline RL benchmarks. It needs no learned value function, so training is stable, but Brandfonbrener and colleagues (2022) showed it can fail to find the optimal policy when the environment is very stochastic or data coverage is poor. The advantage conditioning used in π*0.6's RECAP is a variant of the same idea.","example":"At test time, Decision Transformer is given a target return, such as 1 for task success, and after every step it subtracts the actual reward received from the target, then predicts the next action conditioned on the remaining target return.","related":["Decision Transformer","Advantage Conditioning","Return","Offline Reinforcement Learning","RECAP","Goal-conditioned Policy"]},{"id":"advantage-conditioning","category":"training","sec":7,"tier":3,"sources":[{"title":"π*0.6: a VLA That Learns From Experience (RECAP, arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"},{"title":"Diffusion Guidance Is a Controllable Policy Improvement Operator (CFGRL, arXiv 2505.23458)","url":"https://arxiv.org/abs/2505.23458"}],"as_of":"2025-11","related_ids":["recap","pi-star-0-6","advantage-function","return-conditioning","classifier-free-guidance","advantage-weighted-regression"],"name":"Advantage Conditioning","alt":"优势条件化","abbr":"","aliases":["Advantage-Conditioned Policy"],"one_liner":"Telling the policy how “good” each action was during training, then at deployment asking it only to generate “good” actions.","explanation":"Advantage conditioning turns reinforcement learning into conditional supervised learning: first train a value function and use it to compute each action's advantage in the dataset (how much better than average it was), then feed that advantage, or a binarized version of it, into the policy as an extra input and train with ordinary imitation learning. This way both good and bad data contribute to training, and at inference, setting the condition to “good” biases the policy toward high-advantage actions. The idea is close to return-conditioned methods like Decision Transformer, and related to CFGRL, proposed by Frans and colleagues in 2025, which treats a diffusion model's classifier-free guidance as a form of policy improvement. Physical Intelligence's π*0.6 (November 2025) builds its RECAP method on this: it binarizes the advantage by a threshold, feeds it to the policy as the text “Advantage: positive / negative,” randomly drops the condition during training, and sets it to positive at inference — or amplifies the effect further with classifier-free guidance.","example":"After training with RECAP, π*0.6 can fold laundry in real homes, reliably assemble cardboard boxes, and make espresso on a professional machine; the paper reports throughput more than doubling on some of the hardest tasks and failure rate roughly halving.","related":["RECAP","π*0.6","Advantage Function","Return Conditioning","Classifier-Free Guidance","Advantage-Weighted Regression"]},{"id":"offline-to-online-reinforcement-learning","category":"training","sec":7,"tier":3,"sources":[{"title":"Nakamoto et al. 2023: Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning","url":"https://arxiv.org/abs/2303.05479"},{"title":"Ball et al. 2023: Efficient Online Reinforcement Learning with Offline Data (RLPD)","url":"https://arxiv.org/abs/2302.02948"},{"title":"Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning","url":"https://arxiv.org/abs/2401.16013"}],"as_of":"","related_ids":["offline-reinforcement-learning","online-reinforcement-learning","reinforcement-learning-with-prior-data","calibrated-q-learning","real-world-reinforcement-learning","serl"],"name":"Offline-to-Online Reinforcement Learning","alt":"离线到在线强化学习","abbr":"O2O RL","aliases":["O2O RL","Offline-to-Online Fine-tuning"],"one_liner":"Pretraining with offline reinforcement learning on existing data, then letting the robot keep improving through live interaction.","explanation":"Offline-to-online reinforcement learning happens in two stages: first, offline RL on data that already exists — human demonstrations, historical logs, or an older policy's trajectories — produces an initial policy and value function that does not have to explore from scratch; then the agent interacts with the real environment to collect new data and fine-tunes online. The benefit is skipping a lot of dangerous, expensive random exploration, which matters enormously for real robots. The hard part is the handoff between the two stages: the offline stage is deliberately conservative to avoid overestimation, so its value scale is often inaccurate, and performance commonly dips right after switching to online training. Representative methods include AWAC (2020), Calibrated Q-Learning (Cal-QL, 2023), and RLPD (2023), which mixes offline data directly into the online replay buffer and trains from scratch. Real-robot RL systems such as SERL and HIL-SERL, and π*0.6's RECAP, all follow this same two-stage idea.","example":"SERL uses RLPD as its core algorithm: it starts with 20 demonstrations teleoperated with a SpaceMouse, then trains online on the real robot, learning tasks such as PCB assembly or cable routing in 25 to 50 minutes on average.","related":["Offline Reinforcement Learning","Online Reinforcement Learning","Reinforcement Learning with Prior Data","Calibrated Q-Learning","Real-World Reinforcement Learning","SERL"]},{"id":"calibrated-q-learning","category":"training","sec":7,"tier":3,"sources":[{"title":"Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning (arXiv 2303.05479)","url":"https://arxiv.org/abs/2303.05479"}],"as_of":"","related_ids":["conservative-q-learning","offline-to-online-reinforcement-learning","offline-reinforcement-learning","q-function","reinforcement-fine-tuning","reinforcement-learning-with-prior-data"],"name":"Calibrated Q-Learning","alt":"校准 Q 学习","abbr":"Cal-QL","aliases":["Cal-QL","Calibrated Offline RL Pre-Training"],"one_liner":"Adding a “calibration” constraint to conservative Q-learning, so offline pretraining can transition smoothly into online fine-tuning.","explanation":"Cal-QL was introduced by Nakamoto, Chelsea Finn, Aviral Kumar, Sergey Levine, and colleagues in 2023, published at NeurIPS 2023, targeting offline-to-online reinforcement learning: pretrain offline on a fixed dataset first, then fine-tune online in the actual environment. The authors found that Q-values learned by pretraining with conservative Q-learning (CQL) are pushed down so far that they end up smaller than the true return of any reasonable policy; as soon as online fine-tuning starts, the Q-values swing wildly in scale, and the policy effectively “forgets” what it learned offline first, with performance dropping before slowly climbing back. Cal-QL adds a calibration constraint: the learned Q-value must still be a lower bound on the current policy's true value, but it may not fall below the value of some reference policy — in practice, the reference policy is just the dataset's behavior policy, with its value estimated from Monte Carlo returns. Implementation-wise, this needs only a one-line change to CQL's code, and it beats existing methods on 9 of 11 fine-tuning benchmark tasks.","example":"Cal-QL's test tasks include FrankaKitchen (controlling a 9-DOF Franka arm through a sequence of kitchen sub-tasks) and Adroit (a 28-DOF five-fingered hand spinning a pen, opening a door), all following the same offline-pretrain-then-online-fine-tune pipeline.","related":["Conservative Q-Learning","Offline-to-Online Reinforcement Learning","Offline Reinforcement Learning","Q-Function","Reinforcement Fine-Tuning (RL Fine-Tuning)","Reinforcement Learning with Prior Data"]},{"id":"reinforcement-learning-with-prior-data","category":"training","sec":7,"tier":3,"sources":[{"title":"Ball et al. 2023: Efficient Online Reinforcement Learning with Offline Data (RLPD)","url":"https://arxiv.org/abs/2302.02948"},{"title":"GitHub: ikostrikov/rlpd","url":"https://github.com/ikostrikov/rlpd"},{"title":"Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning","url":"https://arxiv.org/abs/2401.16013"}],"as_of":"","related_ids":["offline-to-online-reinforcement-learning","soft-actor-critic","update-to-data-ratio","experience-replay","serl","hil-serl"],"name":"Reinforcement Learning with Prior Data","alt":"利用先验数据的强化学习","abbr":"RLPD","aliases":["RLPD","Efficient Online Reinforcement Learning with Offline Data"],"one_liner":"A simple, efficient way to do online RL by mixing offline data half-and-half with freshly collected data in every batch.","explanation":"RLPD was proposed by Ball, Smith, Kostrikov, and Levine at ICML 2023, addressing how to make the best use of existing offline data — expert demonstrations or a large amount of suboptimal trajectories — once online interaction begins. The authors found that no elaborate offline pretraining was needed: just a few small changes to the off-policy algorithm SAC were enough. Every training batch draws half from the offline data and half from the online replay buffer (symmetric sampling); the critic network gets layer normalization to stop it from overestimating the value of unseen actions; and the critic is an ensemble of 10 networks trained with a higher update-to-data ratio. The paper reports roughly a 2.5x improvement over prior methods across several benchmarks. The real-robot RL framework SERL uses RLPD as its core algorithm.","example":"SERL trains on a real robot using RLPD: human demonstrations go into the offline buffer, and data the robot generates through its own interaction goes into the online buffer, with every training batch drawing half from each.","related":["Offline-to-Online Reinforcement Learning","Soft Actor-Critic","Update-to-Data Ratio","Experience Replay","SERL","HIL-SERL"]},{"id":"q-chunking","category":"training","sec":7,"tier":3,"sources":[{"title":"Li, Zhou, Levine 2025: Reinforcement Learning with Action Chunking","url":"https://arxiv.org/abs/2507.07969"}],"as_of":"2025-07","related_ids":["action-chunking","offline-to-online-reinforcement-learning","temporal-difference-learning","reinforcement-learning-with-prior-data","flow-matching","sparse-reward"],"name":"Q-Chunking","alt":"动作分块强化学习","abbr":"QC","aliases":["Reinforcement Learning with Action Chunking","QC"],"one_liner":"Doing reinforcement learning where both the policy and the Q-function operate on a whole chunk of actions at once.","explanation":"Q-chunking is a reinforcement learning method proposed by Sergey Levine's team at Berkeley in 2025, aimed at long-horizon, sparse-reward, offline-to-online tasks, where a policy is pretrained on existing data and then goes online. It brings action chunking, a technique common in imitation learning, into reinforcement learning: the policy outputs a chunk of h consecutive actions at once, and the Q-function, which judges how good an action is, also takes the state plus the whole action chunk as input. This has two benefits: acting in chunks makes exploration more coherent and lets the agent carry over behavioral habits from the offline data, and scoring a whole chunk at once enables unbiased multi-step temporal-difference updates, so value information propagates faster. To keep the policy from drifting too far from the data, it samples several action chunks from a flow-matching policy and picks the one with the highest Q-value, or adds a distillation constraint.","example":"On OGBench's block- and puzzle-manipulation tasks and on robomimic tasks, with 1 million steps of offline pretraining followed by 1 million steps of online interaction, Q-chunking with a chunk length of 5 clearly beats offline-to-online baselines such as RLPD, and the harder the task, the bigger the gap.","related":["Action Chunking","Offline-to-Online Reinforcement Learning","Temporal-Difference Learning","Reinforcement Learning with Prior Data","Flow Matching","Sparse Reward"]},{"id":"pre-training","category":"training","sec":8,"tier":1,"sources":[{"title":"Google Machine Learning Glossary: pre-trained model","url":"https://developers.google.com/machine-learning/glossary#pre-trained-model"},{"title":"Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/abs/2406.09246"},{"title":"Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"","related_ids":["post-training","fine-tuning","foundation-model","downstream-task","scaling-law","self-supervised-learning"],"name":"Pre-training","alt":"预训练","abbr":"","aliases":["Pretraining"],"one_liner":"Training a foundation model on massive general-purpose data first, before adapting it to any specific task.","explanation":"Pretraining is the first stage of building a foundation model: the model is trained on data that's both large in scale and broad in coverage — web text, image-text pairs, video, manipulation data from many different robots — so it learns general representations and capabilities. The result is called a pretrained model, or base model, which is later adapted to a specific downstream task through fine-tuning or post-training. Its value is in doing the expensive part of general learning once, so that downstream tasks only need a small amount of additional data. In embodied AI, pretraining typically happens in two layers: a VLA model first inherits a vision-language model that was already pretrained on internet image-text data, then continues training on large-scale data from many different robots — π0, for example, used over 10,000 hours of robot data for this second layer.","example":"OpenVLA (7B parameters), built on a Llama 2 language model plus DINOv2 and SigLIP vision encoders, was pretrained on 970,000 real-robot demonstrations from Open X-Embodiment, and can then be fine-tuned to a new task with LoRA on a consumer GPU.","related":["Post-training","Fine-tuning","Foundation Model","Downstream Task","Scaling Law","Self-Supervised Learning"]},{"id":"downstream-task","category":"training","sec":8,"tier":1,"sources":[{"title":"Bommasani et al. 2021: On the Opportunities and Risks of Foundation Models","url":"https://arxiv.org/abs/2108.07258"},{"title":"Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"","related_ids":["pre-training","post-training","fine-tuning","foundation-model","zero-shot","benchmark"],"name":"Downstream Task","alt":"下游任务","abbr":"","aliases":["Target Task"],"one_liner":"The specific task a pretrained model is ultimately meant to solve, usually reached by further adaptation or fine-tuning.","explanation":"“Downstream” is defined relative to “upstream” pretraining: a foundation model is first trained on massive, general-purpose data, and is then applied to some specific task — that specific task is the downstream task. A 2021 survey on foundation models led by Stanford defines a foundation model as one trained on broad, large-scale data that can be adapted to a wide range of downstream tasks. Adaptation can mean using the model zero-shot, giving it a few examples, or fine-tuning it. In embodied AI, a VLA foundation model's downstream task is usually a concrete manipulation job on a specific robot — folding laundry, clearing a table — or a task from a benchmark like LIBERO; papers commonly report “success rate on the downstream task after fine-tuning” as the measure of how valuable the pretraining was.","example":"π0 is first pretrained on over 10,000 hours of data spanning 7 robot configurations, then separately post-trained on downstream tasks like folding laundry, clearing a table, and assembling a cardboard box, using anywhere from 5 to over 100 hours of data per task.","related":["Pre-training","Post-training","Fine-tuning","Foundation Model","Zero-shot","Benchmark"]},{"id":"transfer-learning","category":"training","sec":8,"tier":2,"sources":[{"title":"Transfer learning (Wikipedia)","url":"https://en.wikipedia.org/wiki/Transfer_learning"},{"title":"How transferable are features in deep neural networks? (Yosinski et al., arXiv 1411.1792)","url":"https://arxiv.org/abs/1411.1792"},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (project page)","url":"https://robotics-transformer2.github.io/"}],"as_of":"","related_ids":["pre-training","fine-tuning","domain-adaptation","sim-to-real-transfer","positive-negative-transfer","rt-2"],"name":"Transfer Learning","alt":"迁移学习","abbr":"","aliases":["Knowledge Transfer"],"one_liner":"Using knowledge learned on one task or domain to help learn another, related task.","explanation":"Transfer learning means applying a model or representation learned on a source task (usually data-rich) to a target task (usually data-scarce), most commonly by pretraining and then fine-tuning. Related research traces back to Bozinovski and colleagues' 1976 neural-network work; Yosinski and colleagues (2014) found that a network's lower layers learn more general features while higher layers grow more task-specific, and that initializing with transferred weights improves generalization even after further fine-tuning. Transfer isn't automatically beneficial: when the source and target are too dissimilar, it can actually hurt performance, called negative transfer. Embodied AI relies on transfer almost everywhere: turning an internet-pretrained VLM into a VLA, sim-to-real transfer, cross-embodiment transfer, and transferring from human video to a robot; domain adaptation is one sub-area of it.","example":"Google DeepMind's RT-2 fine-tunes a vision-language model pretrained on web-scale image-text data together with robot trajectory data, writing actions as text tokens; the model can carry out instructions absent from its training data, such as “pick up the extinct animal,” performing about twice as well as baselines like RT-1 on generalization tests.","related":["Pre-training","Fine-tuning","Domain Adaptation","Sim-to-Real Transfer","Positive / Negative Transfer","RT-2"]},{"id":"training-from-scratch","category":"training","sec":8,"tier":2,"sources":[{"title":"Rethinking ImageNet Pre-training (He et al., arXiv 1811.08883)","url":"https://arxiv.org/abs/1811.08883"},{"title":"R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)","url":"https://arxiv.org/abs/2203.12601"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/html/2303.04137"}],"as_of":"","related_ids":["pre-training","fine-tuning","baseline","overfitting","transfer-learning","pre-trained-visual-representation"],"name":"Training from Scratch","alt":"从零训练","abbr":"","aliases":["Random Initialization Training","From Scratch"],"one_liner":"Training a model directly on the target data with randomly initialized parameters, borrowing no pretrained weights at all.","explanation":"Training from scratch means starting a network's weights from random values and training only on data for the current task, as opposed to loading a pretrained model and fine-tuning it. The benefit is not being constrained by the pretraining data's distribution or the pretrained model's architecture, and a simpler pipeline; the downside is needing more data and training time, and being more prone to overfitting when data is scarce. Kaiming He and colleagues' 2018 paper “Rethinking ImageNet Pre-training” found that for object detection, training from scratch for long enough can match ImageNet pretraining, showing that one of pretraining's main benefits is simply faster convergence. Because robot data is scarce, papers often use training from scratch as a baseline, to measure how much benefit a pretrained representation or large-scale pretraining actually provides — though pretraining isn't always better: Diffusion Policy's visual encoder, for instance, is an un-pretrained ResNet-18 trained end-to-end from scratch.","example":"The R3M paper compares across 12 simulated manipulation tasks: a visual representation pretrained on Ego4D human video beats a visual encoder trained from scratch by more than 20 percentage points in success rate.","related":["Pre-training","Fine-tuning","Baseline","Overfitting","Transfer Learning","Pre-trained Visual Representation"]},{"id":"scaling-law","category":"training","sec":8,"tier":1,"sources":[{"title":"Scaling Laws for Neural Language Models (arXiv 2001.08361)","url":"https://arxiv.org/abs/2001.08361"},{"title":"Training Compute-Optimal Large Language Models (arXiv 2203.15556)","url":"https://arxiv.org/abs/2203.15556"},{"title":"Data Scaling Laws in Imitation Learning for Robotic Manipulation (arXiv 2410.18647)","url":"https://arxiv.org/abs/2410.18647"}],"as_of":"2024-10","related_ids":["data-scaling-laws-in-imitation-learning","parameter-count","large-language-model","data-diversity","the-bitter-lesson","emergent-abilities"],"name":"Scaling Law","alt":"缩放定律","abbr":"","aliases":["Neural Scaling Law","Robot Scaling Law","Data Scaling Law"],"one_liner":"The empirical pattern that model performance improves smoothly, as a power law, with more parameters, data, and compute.","explanation":"Scaling laws quantify the intuition that bigger models, more data, and more compute produce better results. In 2020, OpenAI's Kaplan and colleagues found that a language model's test loss follows power-law relationships with parameter count, dataset size, and training compute, with some of these trends holding across seven or more orders of magnitude. In 2022, DeepMind's Chinchilla study refined this, showing that for a fixed compute budget, model size and the number of training tokens should be scaled up in roughly equal proportion. Scaling laws matter because they let researchers run small, cheap experiments and extrapolate the payoff of a much larger training run, which guides where to spend a limited budget. The embodied-AI field is now testing whether robot data obeys a similar law, and this is part of the motivation behind the industry's large-scale data-collection efforts.","example":"Tsinghua's Yang Gao lab published “Data Scaling Laws in Imitation Learning for Robotic Manipulation” (2024), collecting over 40,000 demonstrations and running more than 15,000 real-robot trials. They found that a policy's ability to generalize to new environments and objects scales roughly as a power law with the number of training environments and objects, while adding more demonstrations within a single environment gives diminishing returns past a certain point.","related":["Data Scaling Laws in Imitation Learning (Robotic Manipulation)","Parameter Count (Model Size)","Large Language Model","Data Diversity","The Bitter Lesson","Emergent Abilities"]},{"id":"ossification","category":"training","sec":8,"tier":3,"sources":[{"title":"Generalist AI 2025-11-04: GEN-0: Embodied Foundation Models That Scale with Physical Interaction","url":"https://generalistai.com/blog/gen-0"},{"title":"Hernandez et al. 2021: Scaling Laws for Transfer","url":"https://arxiv.org/abs/2102.01293"}],"as_of":"2025-11","related_ids":["scaling-law","gen-0","parameter-count","pre-training","transfer-learning","catastrophic-forgetting"],"name":"Ossification","alt":"骨化（模型骨化）","abbr":"","aliases":["Model Ossification","Weight Ossification"],"one_liner":"A model's weights seem to “freeze up,” unable to absorb new information even as more training data is added.","explanation":"The term “ossification” comes from OpenAI's Hernandez, Kaplan, and colleagues, in their 2021 paper “Scaling Laws for Transfer”: they found that with enough fine-tuning data, a small model that had been pretrained could actually fall behind a same-sized model trained from scratch, as if pretraining had “ossified” the weights into a bad initialization that was hard to escape. Embodied AI picked up the term after Generalist AI's GEN-0, released in November 2025: when pretraining at scale on real-robot data, a 1-billion-parameter model struggled to absorb complex, diverse sensorimotor data, its weights becoming unable to learn new information over time, while 6-billion- and 7-billion-parameter models kept improving. The team describes this as the first time ossification has been observed in robotics; previously it had only been seen in language models, and only at a much smaller, tens-of-millions-of-parameters scale. It is a reminder that as data scale grows, model capacity has to keep up.","example":"In GEN-0's comparison, a 1-billion-parameter model shows ossification-like stagnation when faced with large, diverse sensorimotor data, while 6-billion- and 7-billion-parameter models keep improving as pretraining data grows.","related":["Scaling Law","GEN-0","Parameter Count (Model Size)","Pre-training","Transfer Learning","Catastrophic Forgetting"]},{"id":"distributed-training","category":"training","sec":8,"tier":2,"sources":[{"title":"PyTorch Tutorial: Getting Started with Distributed Data Parallel","url":"https://docs.pytorch.org/tutorials/intermediate/ddp_tutorial.html"},{"title":"PyTorch Tutorial: Getting Started with Fully Sharded Data Parallel (FSDP)","url":"https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html"},{"title":"Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism","url":"https://arxiv.org/abs/1909.08053"}],"as_of":"","related_ids":["distributed-data-parallel","fully-sharded-data-parallel","deepspeed","mixed-precision-training","gradient-accumulation","batch-size"],"name":"Distributed Training","alt":"分布式训练（数据并行 / 模型并行）","abbr":"","aliases":["Data Parallelism","Model Parallelism","Tensor Parallelism","Pipeline Parallelism","Multi-GPU Training"],"one_liner":"Splitting a single training run across multiple GPUs or machines so larger models and more data can be trained.","explanation":"When a model or dataset is too large for one GPU to hold or process, training gets split across multiple GPUs or machines. There are two main approaches. Data parallelism puts a full copy of the model on each GPU, has each one process a different slice of the data, and after backpropagation averages the gradients across GPUs before applying a synchronized update — PyTorch's DDP (Distributed Data Parallel) works this way. Model parallelism instead splits the model itself across GPUs; NVIDIA's Megatron-LM, for instance, splits each Transformer layer's matrices across multiple GPUs (tensor parallelism), and there's also pipeline parallelism, which assigns different layers to different GPUs. FSDP (Fully Sharded Data Parallel) still splits by data, but also shards the parameters, gradients, and optimizer state across GPUs, substantially cutting memory use per device. Pretraining and full fine-tuning of multi-billion-parameter VLA models depend on these techniques.","example":"OpenVLA (7B parameters) was pretrained on 64 A100 GPUs over 14 days; the paper's full fine-tuning also needs 8 A100s running 5–15 hours, and its memory-usage tests split the model across 2 GPUs with FSDP.","related":["Distributed Data Parallel (DDP)","Fully Sharded Data Parallel (FSDP)","DeepSpeed","Mixed-Precision Training","Gradient Accumulation","Batch Size"]},{"id":"mixed-precision-training","category":"training","sec":8,"tier":3,"sources":[{"title":"Micikevicius et al. 2017: Mixed Precision Training","url":"https://arxiv.org/abs/1710.03740"},{"title":"NVIDIA Docs: Train With Mixed Precision","url":"https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html"},{"title":"GitHub: Physical-Intelligence/openpi (Precision Settings)","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["numerical-precision-formats","gpu-memory","distributed-training","gradient-checkpointing","quantization-aware-training","openpi"],"name":"Mixed-Precision Training","alt":"混合精度训练","abbr":"","aliases":["bf16 Training","Half-Precision Training","Automatic Mixed Precision (AMP)"],"one_liner":"Running most computation in 16-bit floating point and keeping only the sensitive parts in 32-bit, to save memory and gain speed.","explanation":"Deep learning defaults to 32-bit floating point (FP32) for storing parameters and doing computation. Mixed-precision training switches operations like matrix multiplication and convolution to a 16-bit format (FP16 or BF16), while reductions, normalization, and other numerically sensitive operations, plus the master copy of the parameters, stay in FP32. Researchers at NVIDIA and Baidu systematically proposed this approach in 2017 and introduced “loss scaling” to stop small FP16 gradients from underflowing to zero; memory use drops by nearly half, and GPU tensor cores have higher throughput for 16-bit math, so training runs faster too. BF16 has the same numeric range as FP32 and usually needs no loss scaling, so most large-model and VLA training today uses BF16. PyTorch's autocast handles most of the precision switching automatically.","example":"openpi trains the π0 family in JAX with mixed precision by default: weights and gradients are kept in FP32, while most activations and computation run in BF16; training in full FP32 is just a matter of changing the config's dtype to float32.","related":["Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)","GPU Memory (VRAM)","Distributed Training","Gradient Checkpointing (Activation Recomputation)","Quantization-Aware Training","openpi (Physical Intelligence)"]},{"id":"gradient-accumulation","category":"training","sec":8,"tier":3,"sources":[{"title":"Performing gradient accumulation with Accelerate (Hugging Face 文档)","url":"https://huggingface.co/docs/accelerate/usage_guides/gradient_accumulation"}],"as_of":"","related_ids":["batch-size","gradient-checkpointing","mixed-precision-training","distributed-training","optimizer","gpu-memory"],"name":"Gradient Accumulation","alt":"梯度累积","abbr":"","aliases":[],"one_liner":"Accumulating gradients over several small batches before applying one combined parameter update, simulating a larger batch size.","explanation":"This is the most common trick for when there isn't enough memory for the batch size you want. Each small batch runs its forward and backward pass as usual, but the optimizer isn't called yet — the gradients simply accumulate on the parameters; only after N small batches have been processed does one parameter update actually happen, after which the gradients are zeroed. The effective batch size becomes the per-GPU batch size times the number of accumulation steps times the number of GPUs, at the cost of fewer updates and more time per update. Two things need care: the loss must be divided by the number of accumulation steps (or normalized by total token count, when computing a token-level loss), or the gradient ends up scaled up too much; and layers like BatchNorm that depend on batch statistics still only ever see the small batch. Training libraries like Hugging Face Accelerate offer this as a ready-made setting, and it's used often when fine-tuning a VLA on a single GPU or just a few.","example":"If a single GPU only fits 8 samples but the target batch size is 64, set the accumulation steps to 8: run backward on 8 small batches in a row, then call optimizer.step() once.","related":["Batch Size","Gradient Checkpointing (Activation Recomputation)","Mixed-Precision Training","Distributed Training","Optimizer","GPU Memory (VRAM)"]},{"id":"gradient-checkpointing","category":"training","sec":8,"tier":3,"sources":[{"title":"Training Deep Nets with Sublinear Memory Cost (arXiv 1604.06174)","url":"https://arxiv.org/abs/1604.06174"},{"title":"torch.utils.checkpoint (PyTorch 文档)","url":"https://docs.pytorch.org/docs/main/checkpoint.html"}],"as_of":"","related_ids":["backpropagation","gpu-memory","gradient-accumulation","mixed-precision-training","fully-sharded-data-parallel","deepspeed"],"name":"Gradient Checkpointing (Activation Recomputation)","alt":"梯度检查点（激活重计算）","abbr":"","aliases":["Activation Checkpointing"],"one_liner":"Storing fewer intermediate activations during the forward pass and recomputing them during backward, trading compute for memory.","explanation":"During training, the intermediate results (activations) produced by the forward pass are normally kept in memory until backpropagation is done with them, which is one of the biggest consumers of GPU memory. Gradient checkpointing saves activations only at a handful of checkpoint locations and discards the rest, recomputing them with a fresh forward pass from the nearest checkpoint whenever backpropagation needs them. Tianqi Chen and colleagues' 2016 paper “Training Deep Nets with Sublinear Memory Cost” systematically introduced this: an n-layer network needs only about O(√n) memory for activations, at the cost of roughly one extra forward pass per small batch; in the paper, a 1,000-layer residual network's memory usage dropped from 48GB to 7GB, with about a 30% increase in running time. PyTorch's torch.utils.checkpoint implements this, and it's commonly turned on together with mixed precision and gradient accumulation when training large models and VLAs.","example":"Fine-tuning a Transformer policy by wrapping each Transformer layer in torch.utils.checkpoint noticeably lowers memory usage, at the cost of each training step running somewhat slower.","related":["Backpropagation","GPU Memory (VRAM)","Gradient Accumulation","Mixed-Precision Training","Fully Sharded Data Parallel (FSDP)","DeepSpeed"]},{"id":"representation-learning","category":"training","sec":8,"tier":2,"sources":[{"title":"Representation Learning: A Review and New Perspectives (Bengio et al.)","url":"https://arxiv.org/abs/1206.5538"},{"title":"R3M: A Universal Visual Representation for Robot Manipulation","url":"https://arxiv.org/abs/2203.12601"}],"as_of":"","related_ids":["self-supervised-learning","contrastive-learning","pre-trained-visual-representation","embedding","latent-space","r3m"],"name":"Representation Learning","alt":"表征学习","abbr":"","aliases":["Feature Learning"],"one_liner":"Letting a model automatically learn useful feature vectors from raw data, instead of hand-designing the features.","explanation":"Representation learning studies how to turn raw data — images, text, sensor readings — into a set of numbers (a representation, also called a feature or embedding vector) that's easier for downstream tasks to use. Bengio and colleagues' 2013 survey treats this as deep learning's central problem: a good representation should separate out the different underlying factors of variation behind the data. There are many ways to learn one: supervised classification, self-supervised learning (contrastive learning, masked autoencoders), and image-text alignment (as in CLIP). In embodied AI, because robot data is scarce, it's common to first pretrain a visual encoder on large-scale images or human video, then freeze or fine-tune it while learning a policy, so only a small number of demonstrations are needed — R3M, VC-1, and DINOv2 all follow this path. A world model's latent space and its latent actions are also products of representation learning.","example":"R3M pretrains a visual representation on Ego4D first-person human video using time-contrastive learning and video-language alignment; once frozen and attached to a policy for a Franka arm, it learns manipulation tasks in a real, cluttered apartment from just 20 demonstrations.","related":["Self-Supervised Learning","Contrastive Learning","Pre-trained Visual Representation","Embedding","Latent Space","R3M"]},{"id":"contrastive-learning","category":"training","sec":8,"tier":2,"sources":[{"title":"Representation Learning with Contrastive Predictive Coding (arXiv 1807.03748)","url":"https://arxiv.org/abs/1807.03748"},{"title":"Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)","url":"https://arxiv.org/abs/2103.00020"},{"title":"R3M: A Universal Visual Representation for Robot Manipulation (arXiv 2203.12601)","url":"https://arxiv.org/abs/2203.12601"}],"as_of":"","related_ids":["self-supervised-learning","infonce-loss","clip","siglip","time-contrastive-networks","r3m"],"name":"Contrastive Learning","alt":"对比学习","abbr":"","aliases":["Contrastive Representation Learning"],"one_liner":"Learning features from unlabeled data by pulling similar samples' representations together and pushing dissimilar ones apart.","explanation":"Contrastive learning is a family of self-supervised representation-learning methods: construct “positive pairs” (two pieces of data that should be similar, such as two random augmentations of the same image, or an image and its caption) and “negative” examples, then train an encoder to pull positive pairs' vectors closer together and push negatives further apart. A common loss for this is InfoNCE, from the 2018 CPC paper. Well-known examples include SimCLR, MoCo, and CLIP. Because it needs no human labels, it can learn general-purpose visual features from huge amounts of images, video, and image-text pairs. In embodied AI, many VLA models' vision encoders (such as CLIP and SigLIP) are trained with a contrastive objective; R3M instead pretrains robot visual representations on human videos using time-contrastive learning.","example":"CLIP trains on 400 million image-text pairs scraped from the web, learning to tell which caption matches which image. R3M pretrains on Ego4D human videos with time-contrastive learning among other objectives, and once frozen and handed to a Franka arm, needs only 20 demonstrations to learn manipulation tasks in a real, cluttered apartment.","related":["Self-Supervised Learning","InfoNCE Loss","CLIP","SigLIP","Time-Contrastive Networks","R3M"]},{"id":"infonce-loss","category":"training","sec":8,"tier":3,"sources":[{"title":"Representation Learning with Contrastive Predictive Coding (arXiv:1807.03748)","url":"https://arxiv.org/abs/1807.03748"},{"title":"CPC 论文 HTML 版（第 2.3 节 InfoNCE Loss and Mutual Information Estimation）","url":"https://arxiv.org/html/1807.03748"}],"as_of":"","related_ids":["contrastive-learning","clip","self-supervised-learning","time-contrastive-networks","cross-entropy","representation-learning"],"name":"InfoNCE Loss","alt":"InfoNCE 损失","abbr":"","aliases":["Contrastive Loss","InfoNCE Objective"],"one_liner":"A contrastive-learning loss that trains a model to pick out the one true positive among many candidates.","explanation":"InfoNCE was introduced by Aaron van den Oord and colleagues at DeepMind in their 2018 Contrastive Predictive Coding (CPC) paper. It pairs an anchor sample with one positive (say, a different augmentation of the same image, or the next moment in a sequence) and several negatives, computes similarity scores, and applies softmax classification so the positive scores highest — which is really cross-entropy underneath. The paper shows that minimizing it is equivalent to maximizing a lower bound on mutual information (how much two variables share), and the bound gets tighter as the number of negatives grows. It is the core training objective behind contrastive methods such as CLIP and SimCLR, and robot visual representations like R3M use a similar time-contrastive version of it.","example":"When CLIP trains on a batch of N image-text pairs, each image treats only its own caption as the positive and the other N−1 captions as negatives, computing InfoNCE in both the image-to-text and text-to-image directions.","related":["Contrastive Learning","CLIP","Self-Supervised Learning","Time-Contrastive Networks","Cross-Entropy","Representation Learning"]},{"id":"time-contrastive-networks","category":"training","sec":8,"tier":3,"sources":[{"title":"Sermanet et al. 2017: Time-Contrastive Networks: Self-Supervised Learning from Video","url":"https://arxiv.org/abs/1704.06888"},{"title":"项目页：Time-Contrastive Networks","url":"https://sermanet.github.io/imitate/"},{"title":"Nair et al. 2022: R3M: A Universal Visual Representation for Robot Manipulation","url":"https://arxiv.org/abs/2203.12601"}],"as_of":"","related_ids":["contrastive-learning","self-supervised-learning","r3m","representation-learning","imitation-from-observation","pre-trained-visual-representation"],"name":"Time-Contrastive Networks","alt":"时间对比学习","abbr":"TCN","aliases":["TCN","Time-Contrastive Learning"],"one_liner":"Self-supervised video representation learning: pull same-moment multi-view frames together, push nearby-but-different moments apart.","explanation":"Time-Contrastive Networks were proposed by Pierre Sermanet and colleagues at Google Brain in 2017 (ICRA 2018), an early landmark in learning visual representations from unlabeled video. The multi-view version uses several cameras filming the same process at once: frames from different viewpoints at the same moment are pulled together in feature space, while frames from the same viewpoint that look similar but sit at different points in the task are pushed apart, trained with a triplet loss. The resulting features are insensitive to viewpoint and lighting, yet can still distinguish task-relevant state, such as how far a cup has tilted. The paper used it to let a robot imitate pouring water and human poses from a single third-person human video, treating feature distance to the demonstration as a reinforcement-learning reward. Later robot visual pretraining such as R3M also used time-contrastive learning as one of its objectives.","example":"A robot watches a third-person video of a person pouring water, uses the distance between its own footage and the demonstration in TCN feature space as a reward, and learns the pouring motion through reinforcement learning.","related":["Contrastive Learning","Self-Supervised Learning","R3M","Representation Learning","Imitation from Observation","Pre-trained Visual Representation"]},{"id":"masked-autoencoder","category":"training","sec":8,"tier":3,"sources":[{"title":"Masked Autoencoders Are Scalable Vision Learners (arXiv:2111.06377)","url":"https://arxiv.org/abs/2111.06377"},{"title":"Masked Visual Pre-training for Motor Control (MVP, arXiv:2203.06173)","url":"https://arxiv.org/abs/2203.06173"}],"as_of":"","related_ids":["self-supervised-learning","vision-transformer","pre-trained-visual-representation","mvp","vc-1","representation-learning"],"name":"Masked Autoencoder","alt":"掩码自编码器","abbr":"MAE","aliases":["MAE","Masked Image Modeling"],"one_liner":"Hiding most patches of an image and training a model to reconstruct them, as self-supervised visual pretraining.","explanation":"The masked autoencoder is a self-supervised visual-pretraining method proposed by Kaiming He and colleagues at Meta FAIR in 2021. An image is split into patches, about 75% of them are randomly masked, the encoder processes only the visible patches, and a lightweight decoder reconstructs the masked-out pixels from that encoding. Because the encoder only ever sees a quarter of the patches, training runs more than 3 times faster; the method needs no human labels and still learns strong features, reaching 87.8% accuracy on ImageNet-1K with a ViT-Huge. In robotics it is commonly used to pretrain visual encoders: MVP pretrains an encoder with MAE and then freezes it for learning motor control, and VC-1 also uses an MAE objective.","example":"MVP (2022) pretrains a ViT encoder with MAE on a large collection of real-world images, then freezes it and attaches a reinforcement-learning controller; on a set of arm-manipulation tasks it beats a supervised-pretrained encoder by up to 80 percentage points in success rate.","related":["Self-Supervised Learning","Vision Transformer","Pre-trained Visual Representation","MVP","VC-1","Representation Learning"]},{"id":"tactile-representation-learning","category":"training","sec":8,"tier":3,"sources":[{"title":"Higuera et al. 2024: Sparsh: Self-supervised Touch Representations for Vision-based Tactile Sensing (CoRL 2024)","url":"https://arxiv.org/abs/2410.24090"},{"title":"Zhao et al. 2024: Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks (T3)","url":"https://arxiv.org/abs/2406.13640"},{"title":"Feng et al. 2025: AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors (ICLR 2025)","url":"https://arxiv.org/abs/2502.12191"}],"as_of":"2025-04","related_ids":["vision-based-tactile-sensor","self-supervised-learning","sparsh","anytouch","visuo-tactile-fusion","representation-learning"],"name":"Tactile Representation Learning","alt":"触觉表征学习","abbr":"","aliases":["Touch Representation Learning","Tactile Pretraining"],"one_liner":"Encoding raw tactile-sensor readings into general-purpose features that downstream tasks like slip detection can reuse.","explanation":"Tactile representation learning studies how to encode raw signals from tactile sensors, such as the gel-deformation images captured by vision-based tactile sensors like GelSight or DIGIT, into compact features that downstream tasks such as slip detection, force estimation, and precision insertion can use directly. The difficulty is that tactile data is scarce, labels are even scarcer, and different sensors image very differently, so switching sensors often means retraining from scratch. Recent work borrows self-supervised learning from vision: Meta's Sparsh (CoRL 2024) pretrains on more than 460,000 tactile images using masking and self-distillation; MIT's T3 uses a shared Transformer with sensor-specific encoders; and AnyTouch (ICLR 2025) learns a unified representation across four sensor types. This is foundational to visuo-tactile fusion and tactile VLAs.","example":"Sparsh is pretrained with self-supervision on more than 460,000 unlabeled visuo-tactile images; on the six tasks of the authors' own TacBench, the paper reports it beats models trained end-to-end per task and per sensor by 95.1% on average.","related":["Vision-Based Tactile Sensor","Self-Supervised Learning","Sparsh","AnyTouch","Visuo-Tactile Fusion","Representation Learning"]},{"id":"representation-alignment","category":"training","sec":8,"tier":3,"sources":[{"title":"Yu et al. 2024: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think","url":"https://arxiv.org/abs/2410.06940"},{"title":"GitHub: sihyun-yu/REPA","url":"https://github.com/sihyun-yu/REPA"},{"title":"Li et al. 2025: Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model","url":"https://arxiv.org/abs/2510.12276"}],"as_of":"2025-10","related_ids":["representation-learning","diffusion-transformer","dinov2","auxiliary-loss-auxiliary-task","vggt","pre-trained-visual-representation"],"name":"Representation Alignment","alt":"表征对齐","abbr":"REPA","aliases":["REPA","REPA Regularization"],"one_liner":"Training a model so its intermediate features match those of an existing pretrained encoder.","explanation":"REPA was proposed by Sihyun Yu, Saining Xie, and colleagues in October 2024 (an ICLR 2025 Oral), targeting the slow training of diffusion Transformers such as DiT and SiT. It adds an auxiliary loss: the network's intermediate hidden state while processing a noisy image, after passing through a small projection layer, is aligned with that same clean image's features from a pretrained visual encoder such as DINOv2. The authors argue that one bottleneck in diffusion-model training is having to learn good visual representations from scratch, and that borrowing an already-good representation saves that effort — SiT's training sped up by more than 17.5x. This “REPA-style” alignment has since been extended to video generation and robotics: Spatial Forcing, for example, aligns a VLA's intermediate visual tokens with features from the 3D foundation model VGGT, letting a VLA that has only ever seen 2D data implicitly pick up spatial awareness.","example":"Spatial Forcing adds a cosine-similarity alignment loss against VGGT features on top of OpenVLA-OFT and π0, with no extra depth-map or point-cloud input, speeding up training by up to 3.8x and improving data efficiency.","related":["Representation Learning","Diffusion Transformer","DINOv2","Auxiliary Loss / Auxiliary Task","VGGT","Pre-trained Visual Representation"]},{"id":"modality-alignment","category":"training","sec":8,"tier":2,"sources":[{"title":"Visual Instruction Tuning (LLaVA, arXiv 2304.08485)","url":"https://arxiv.org/abs/2304.08485"}],"as_of":"","related_ids":["projector-connector","vision-language-model","multimodal-large-language-model","clip","backbone-freezing","knowledge-insulation"],"name":"Modality Alignment (Alignment Pretraining Stage)","alt":"模态对齐","abbr":"","aliases":["Alignment Pretraining","Feature Alignment Pretraining","Cross-modal Alignment"],"one_liner":"Mapping features from a new modality, like images, into a representation space the language model can already understand.","explanation":"Modality alignment means bringing representations of different modalities — image, language, robot state, action — into a space where they correspond to one another. Narrowly, it often refers to the first stage of training a multimodal large model, “alignment pretraining”: LLaVA (2023), for example, freezes both a CLIP vision encoder and a large language model and trains only the projection layer in between, using about 595,000 image-text pairs, so image features turn into visual tokens the language model can process; the second stage then fine-tunes the projection layer together with the language model on instruction data. Aligning first, on its own, lets the randomly initialized projection layer learn to produce reasonable visual tokens before the language model itself is unfrozen. More broadly, CLIP's use of contrastive learning to pull image and text embeddings into the same space also counts as modality alignment. VLA models face the same problem when attaching a state encoder and an action head to a vision-language model.","example":"LLaVA's first stage trains only the projection matrix for 1 epoch, at a learning rate of 2e-3, on a filtered subset of CC3M (595,000 image-text pairs); its second stage freezes the vision encoder and fine-tunes the projection layer and language model together on 158,000 multimodal instruction examples.","related":["Projector / Connector","Vision-Language Model","Multimodal Large Language Model","CLIP","Backbone Freezing","Knowledge Insulation"]},{"id":"pretraining-on-human-videos","category":"training","sec":8,"tier":2,"sources":[{"title":"R3M: A Universal Visual Representation for Robot Manipulation","url":"https://arxiv.org/abs/2203.12601"},{"title":"Latent Action Pretraining from Videos (LAPA)","url":"https://arxiv.org/abs/2410.11758"},{"title":"Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos","url":"https://arxiv.org/abs/2507.15597"}],"as_of":"2025-07","related_ids":["human-video-data","egocentric-video","latent-action-pretraining","embodiment-gap","r3m","lapa"],"name":"Pretraining on Human Videos","alt":"人类视频预训练","abbr":"","aliases":["Human Video Pretraining","Learning from Human Videos"],"one_liner":"Pretraining a robot model on large amounts of video of humans doing things, then fine-tuning it with a small amount of robot data.","explanation":"Real-robot data is expensive to collect and stays small in scale, while video of humans going about everyday tasks — first-person datasets like Ego4D, or general internet video — is abundant and covers far more scenes, so a common approach is pretraining on human video first and fine-tuning with a small amount of robot data. Two difficulties stand out: videos have no robot action labels, and human hands and robot grippers have different structures (an embodiment gap). There are three main approaches. One learns only a visual representation, as R3M (CoRL 2022) does, training a visual encoder on Ego4D with time-contrastive learning and video-language alignment. A second extracts a latent action as a pseudo-label from the change between consecutive frames, as LAPA does, using a VQ-VAE to get discrete latent actions for pretraining a VLA. A third estimates the human hand's 3D motion directly as the action, as Being-H0 does, treating the human hand as a general-purpose manipulator to pretrain a VLA.","example":"R3M pretrains a visual representation on human Ego4D videos, freezes it, and attaches a small policy network; with the representation frozen, a Franka arm learns a variety of manipulation tasks in a real, cluttered apartment from just 20 demonstrations.","related":["Human Video Data","Egocentric Video","Latent Action Pretraining","Embodiment Gap","R3M","LAPA"]},{"id":"latent-action-pretraining","category":"training","sec":8,"tier":3,"sources":[{"title":"Latent Action Pretraining from Videos (LAPA, arXiv:2410.11758)","url":"https://arxiv.org/abs/2410.11758"},{"title":"LAPA 论文 HTML 版","url":"https://arxiv.org/html/2410.11758"}],"as_of":"2024-10","related_ids":["latent-action","latent-action-model","lapa","action-free-video","pretraining-on-human-videos","vector-quantized-variational-autoencoder"],"name":"Latent Action Pretraining","alt":"潜在动作预训练","abbr":"","aliases":[],"one_liner":"Extracting unlabeled “latent actions” from video, then pretraining a robot model to predict them.","explanation":"Latent action pretraining first learns “latent actions” — an encoding of what changed between two adjacent frames — from video that has no action labels, then pretrains a VLA to predict these latent actions; a small amount of real-robot data is used at the end to map the latent actions onto real robot actions. The representative work is LAPA (October 2024), from researchers at the University of Washington, KAIST, Microsoft Research, NVIDIA, and others, which learns discrete latent actions with a VQ-VAE (a vector-quantized autoencoder). Its value is that it can exploit huge amounts of internet and human-manipulation video that carries no robot action labels at all. Genie's latent action model, UniVLA, and AgiBot's GO-1 all take a similar approach.","example":"The LAPA paper reports positive transfer even when pretraining only on Something-Something V2 human-manipulation videos; on real-robot tasks that require language conditioning and generalization, it beats OpenVLA, which was trained with real action labels, while using roughly one-thirtieth the pretraining compute.","related":["Latent Action","Latent Action Model","LAPA","Action-free Video","Pretraining on Human Videos","Vector-Quantized Variational Autoencoder"]},{"id":"multi-task-learning","category":"training","sec":8,"tier":2,"sources":[{"title":"Caruana (1997), Multitask Learning, Machine Learning 28","url":"https://www.cs.cornell.edu/~caruana/mlj97.pdf"},{"title":"An Overview of Multi-Task Learning in Deep Neural Networks (Ruder, 2017)","url":"https://arxiv.org/abs/1706.05098"},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (project page)","url":"https://robotics-transformer1.github.io/"}],"as_of":"","related_ids":["positive-negative-transfer","transfer-learning","language-conditioned-policy","generalist-policy","data-mixture","rt-1"],"name":"Multi-Task Learning","alt":"多任务学习","abbr":"MTL","aliases":["MTL","Multitask Learning"],"one_liner":"Training one model on several related tasks at once, so it shares knowledge and each task helps the others.","explanation":"Rich Caruana's 1997 paper “Multitask Learning” laid out this idea systematically: train several related tasks in parallel sharing one representation, so what one task learns helps the others learn better, with each task's training signal acting as an extra inductive bias (a model's built-in leaning toward what kind of solution is reasonable) for the rest. The most common deep-learning implementation is hard parameter sharing: several tasks share one backbone network, each with its own output head. In robotics, a multi-task policy is told what to do right now via a language instruction or a goal image, and one network covers many skills — opening a drawer, grasping, placing — and essentially every VLA today is a multi-task model. The difficulty is that tasks can interfere with each other (negative transfer), which requires tuning the data mix and network architecture.","example":"Google's RT-1 used 13 robots over 17 months to collect more than 130,000 demonstrations spanning over 700 language instructions, with one Transformer policy learning all of them and reaching 97% success on instructions it had seen before.","related":["Positive / Negative Transfer","Transfer Learning","Language-conditioned Policy","Generalist Policy","Data Mixture","RT-1"]},{"id":"co-training","category":"training","sec":8,"tier":2,"sources":[{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)","url":"https://arxiv.org/abs/2307.15818"},{"title":"Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation (arXiv 2401.02117)","url":"https://arxiv.org/abs/2401.02117"},{"title":"Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation (arXiv 2503.24361)","url":"https://arxiv.org/abs/2503.24361"}],"as_of":"2025-04","related_ids":["data-mixture","sim-to-real-transfer","cross-embodiment-data","rt-2","pi0-5","mobile-aloha"],"name":"Co-training","alt":"协同训练","abbr":"","aliases":["Joint Training","Mixed Training","Co-fine-tuning","Sim-and-Real Co-training"],"one_liner":"Mixing a small amount of target-robot data with data from other sources in fixed proportions to train one model.","explanation":"In embodied AI, co-training means mixing a small amount of target-robot data together with data from other sources, in a set ratio, into the same training batches for one shared model. The other sources can be web image-text QA data, data from other robots, or simulation data. Google's RT-2 (2023) fine-tuned on robot trajectories together with web tasks like visual question answering, calling it co-fine-tuning; π0.5 (2025) trains jointly on multi-robot data, high-level semantic predictions, and web data. The point is to work around having too little real-robot data: mixing in other data preserves visual and language common sense and improves generalization, and the key is getting the mixing ratio right. Note this is a different sense from the “co-training” Blum and Mitchell proposed in 1998, which is a semi-supervised learning method.","example":"Mobile ALOHA had only 50 mobile-manipulation demonstrations per task; co-training with existing static ALOHA bimanual data raised success rates by up to 90 percentage points. In a 2025 sim-and-real co-training study from NVIDIA and collaborators, mixing in simulation data raised real-robot task success rates by 38% on average.","related":["Data Mixture","Sim-to-Real Transfer","Cross-Embodiment Data","RT-2","π0.5","Mobile ALOHA"]},{"id":"positive-negative-transfer","category":"training","sec":8,"tier":2,"sources":[{"title":"Characterizing and Avoiding Negative Transfer (CVPR 2019)","url":"https://arxiv.org/abs/1811.09751"},{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models (project page)","url":"https://robotics-transformer-x.github.io/"},{"title":"Open X-Embodiment paper (arXiv HTML)","url":"https://arxiv.org/html/2310.08864"}],"as_of":"","related_ids":["transfer-learning","multi-task-learning","cross-embodiment","co-training","data-mixture","rt-x"],"name":"Positive / Negative Transfer","alt":"正迁移 / 负迁移","abbr":"","aliases":["Positive Transfer","Negative Transfer"],"one_liner":"When adding other tasks or data makes the target task better, that's positive transfer; when it makes it worse, that's negative transfer.","explanation":"This is a pair of concepts from transfer learning and multi-task learning: applying knowledge learned on a source task or source data to a target task, where the target task's performance improving is positive transfer, and it getting worse is negative transfer. Wang and colleagues gave a formal definition of negative transfer in a CVPR 2019 paper, noting it tends to happen when the source data is too weakly related to the target task. In embodied AI, whenever data from multiple robots, multiple tasks, or internet image-text sources is trained together, the question of whether the added data helps or hurts always comes up. Common causes of negative transfer include unrelated tasks, mismatched action spaces or coordinate frames across different robot embodiments, a poorly balanced data mix, and insufficient model capacity. Responses include tuning the data mixing ratio, giving each embodiment its own output head, and freezing or isolating parts of the model.","example":"In the Open X-Embodiment experiments, RT-1-X, trained on data from 22 robot types, beat single-robot training by 50% on average on robots with little of their own data — positive transfer. But on robots with plenty of data, RT-1-X underfit and did worse than RT-1 trained only on that robot's own data; only the much larger RT-2-X regained the advantage.","related":["Transfer Learning","Multi-Task Learning","Cross-Embodiment","Co-training","Data Mixture","RT-X"]},{"id":"mid-training","category":"training","sec":8,"tier":3,"sources":[{"title":"Mid-Training of Large Language Models: A Survey (arXiv 2510.06826)","url":"https://arxiv.org/abs/2510.06826"},{"title":"MolmoAct: Action Reasoning Models that can Reason in Space (arXiv 2508.07917)","url":"https://arxiv.org/abs/2508.07917"},{"title":"EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (arXiv 2602.16710)","url":"https://arxiv.org/abs/2602.16710"}],"as_of":"2026-02","related_ids":["pre-training","post-training","supervised-fine-tuning","vision-language-action-model","data-mixture","molmoact"],"name":"Mid-training","alt":"中训练","abbr":"","aliases":[],"one_liner":"An extra training stage between pretraining and post-training that uses more targeted data to fill in specific abilities.","explanation":"Mid-training first became popular with large language models: after large-scale pretraining and before instruction tuning and reinforcement learning, the model continues training for a while on higher-quality or more target-relevant data — math, code, long documents — often paired with learning-rate annealing and context-length extension, aiming to strengthen specific abilities without losing general foundations; a dedicated 2025 survey already maps out this stage. Embodied AI borrowed the term: an off-the-shelf vision-language model has never seen robot data, so fine-tuning it directly into a VLA gives limited results, so teams insert a transitional round in between using embodiment-related data, such as spatial reasoning, robot trajectories, or human motion aligned to robot motion. Ai2's MolmoAct released a robot dataset built specifically for mid-training, and the Chinese embodied-AI company Dexmal's DM0 likewise uses a three-stage pretraining, mid-training, post-training pipeline.","example":"EgoScale first pretrains a VLA on more than 20,000 hours of action-annotated egocentric human video, then runs one lightweight round of mid-training on a small amount of data that aligns human and robot actions; after that, it needs only minimal robot supervision to adapt to new dexterous-manipulation tasks.","related":["Pre-training","Post-training","Supervised Fine-Tuning","Vision-Language-Action Model","Data Mixture","MolmoAct"]},{"id":"post-training","category":"training","sec":9,"tier":1,"sources":[{"title":"Lambert et al. 2024: Tulu 3: Pushing Frontiers in Open Language Model Post-Training","url":"https://arxiv.org/abs/2411.15124"},{"title":"Black et al. 2024: π0: A Vision-Language-Action Flow Model for General Robot Control","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"","related_ids":["pre-training","fine-tuning","supervised-fine-tuning","reinforcement-fine-tuning","reinforcement-learning-from-human-feedback","mid-training"],"name":"Post-training","alt":"后训练","abbr":"","aliases":["Posttraining"],"one_liner":"The training stage that follows pretraining, using curated data or reinforcement learning to turn a foundation model into something actually useful.","explanation":"Post-training refers to the training steps that come after a foundation model's pretraining is done — the term became common alongside large language models. A large model is first pretrained on massive amounts of web data, then goes through post-training steps like supervised fine-tuning (SFT), preference optimization, and reinforcement learning before it becomes an assistant that can actually follow instructions; Ai2's Tulu 3, for instance, published a full post-training recipe combining SFT, DPO, and reinforcement learning with verifiable rewards. Embodied AI has adopted the same division of labor: the π0 paper states that the pretraining stage is responsible for broad capability and generalization, while the post-training stage uses narrower, more carefully curated data to make the model proficient at the actual target task. Post-training is often used interchangeably with fine-tuning, but it emphasizes the idea of a “stage,” which can itself contain multiple rounds of fine-tuning and reinforcement-learning fine-tuning.","example":"After pretraining, π0 goes through post-training for complex tasks like folding laundry: the simplest tasks need only about 5 hours of data, while the most complex ones use over 100 hours.","related":["Pre-training","Fine-tuning","Supervised Fine-Tuning","Reinforcement Fine-Tuning (RL Fine-Tuning)","Reinforcement Learning from Human Feedback","Mid-training"]},{"id":"fine-tuning","category":"training","sec":9,"tier":1,"sources":[{"title":"Wikipedia: Fine-tuning (deep learning)","url":"https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning)"},{"title":"Google Machine Learning Glossary: fine-tuning","url":"https://developers.google.com/machine-learning/glossary#fine-tuning"},{"title":"GitHub: Physical-Intelligence/openpi","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["pre-training","post-training","full-fine-tuning","lora","supervised-fine-tuning","catastrophic-forgetting"],"name":"Fine-tuning","alt":"微调","abbr":"","aliases":["Finetuning"],"one_liner":"Continuing to train an already-trained model on a small amount of new-task data, so it gets better at that specific job.","explanation":"Fine-tuning is the most common form of transfer learning: instead of starting from random parameters, you start from a pretrained model and keep training it on data from the target task, reusing the general knowledge the model already learned, which saves both data and compute. By how much gets updated, full fine-tuning changes every weight; alternatively, most layers can be frozen and only a subset trained, or a parameter-efficient method like LoRA (Low-Rank Adaptation, which inserts a small number of trainable low-rank matrices) can be used to train only a tiny fraction of the parameters. In embodied AI, the most common way to get started with an open-source VLA is to fine-tune it on demonstration data collected on your own robot. Too little data, or training for too long, easily leads to overfitting, and can also make the model forget capabilities it previously had — a problem called catastrophic forgetting.","example":"The openpi documentation gives this reference: fine-tuning π0 with LoRA needs at least 22.5GB of GPU memory (such as an RTX 4090), while full-parameter fine-tuning needs at least 70GB (such as an A100 or H100).","related":["Pre-training","Post-training","Full Fine-Tuning","LoRA","Supervised Fine-Tuning","Catastrophic Forgetting"]},{"id":"full-fine-tuning","category":"training","sec":9,"tier":2,"sources":[{"title":"LoRA: Low-Rank Adaptation of Large Language Models","url":"https://arxiv.org/abs/2106.09685"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/abs/2406.09246"},{"title":"openpi GitHub README（显存需求表）","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"2025","related_ids":["fine-tuning","parameter-efficient-fine-tuning","lora","catastrophic-forgetting","backbone-freezing","supervised-fine-tuning"],"name":"Full Fine-Tuning","alt":"全参数微调","abbr":"","aliases":["Full-parameter Fine-Tuning"],"one_liner":"Fine-tuning that updates every parameter of a pretrained model, rather than training just a small subset of them.","explanation":"Fine-tuning means continuing to train a pretrained model on data for a downstream task. Full fine-tuning updates every weight; the alternative is parameter-efficient fine-tuning (PEFT), which trains only a small number of new or selected parameters, such as LoRA (low-rank adaptation). Full fine-tuning has the most freedom to adapt, and when data is plentiful and the target task differs a lot from pretraining, it usually performs best — but it's expensive: every parameter needs its own stored gradient and optimizer state, pushing memory needs to several times that of plain inference, and a full copy of the model must be saved per task; with little data it's also more prone to overfitting or catastrophic forgetting, where the model loses old abilities while learning a new task. When adapting a VLA to a new robot or task, full fine-tuning and LoRA are the two most common choices, and picking between them comes down to available GPU memory and how much data there is.","example":"The OpenVLA (7B) paper reports full fine-tuning needing 8 A100s running 5–15 hours, reaching 69.7% success; LoRA trains only 1.4% of the parameters, fits on a single A100, and reaches 68.2%. The openpi repository lists π0 full fine-tuning as needing over 70 GB of memory versus about 22.5 GB for LoRA.","related":["Fine-tuning","Parameter-Efficient Fine-Tuning","LoRA","Catastrophic Forgetting","Backbone Freezing","Supervised Fine-Tuning"]},{"id":"backbone-freezing","category":"training","sec":9,"tier":2,"sources":[{"title":"PyTorch Tutorial: Transfer Learning for Computer Vision","url":"https://docs.pytorch.org/tutorials/beginner/transfer_learning_tutorial.html"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"}],"as_of":"2025-03","related_ids":["backbone-network","fine-tuning","parameter-efficient-fine-tuning","catastrophic-forgetting","knowledge-insulation","stop-gradient"],"name":"Backbone Freezing","alt":"冻结骨干网络","abbr":"","aliases":["Freezing the Backbone","Frozen Backbone","Frozen Pretrained Layers"],"one_liner":"Keeping a pretrained backbone network's parameters fixed during training and updating only the newly added parts.","explanation":"A backbone is the part of a model responsible for extracting features, usually a pretrained network such as a ResNet, a ViT, or an entire vision-language model (VLM). Freezing it means those parameters are not updated during training (in PyTorch, by setting requires_grad to False), and only newly added output heads or action modules are trained. The upside is lower memory use, faster training, less overfitting when data is scarce, and less catastrophic forgetting; the downside is that the backbone can't learn features the new task might need. The trade-off shows up concretely in VLA models: OpenVLA found that freezing the vision encoder loses fine-grained spatial information and hurts control performance, while NVIDIA's GR00T N1 freezes the VLM's language component during both pretraining and post-training and trains only the remaining modules.","example":"PyTorch's own transfer-learning tutorial fine-tunes an ImageNet-pretrained ResNet-18 to classify ants versus bees: every layer except the last is frozen, and only the newly swapped-in fully connected layer is trained.","related":["Backbone Network","Fine-tuning","Parameter-Efficient Fine-Tuning","Catastrophic Forgetting","Knowledge Insulation","Stop-Gradient"]},{"id":"parameter-efficient-fine-tuning","category":"training","sec":9,"tier":3,"sources":[{"title":"Hugging Face PEFT 文档","url":"https://huggingface.co/docs/peft/index"},{"title":"Han et al. 2024: Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey","url":"https://arxiv.org/abs/2403.14608"},{"title":"Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["lora","adapter","prompt-tuning-soft-prompt","full-fine-tuning","backbone-freezing","fine-tuning"],"name":"Parameter-Efficient Fine-Tuning","alt":"参数高效微调","abbr":"PEFT","aliases":["PEFT"],"one_liner":"Freezing most of a large model's parameters and training only a small added or selected subset for a new task.","explanation":"Parameter-efficient fine-tuning is an umbrella term for fine-tuning methods that keep a pretrained model's parameters mostly frozen and train only a small subset to adapt it to a new task. Four approaches are common: inserting small modules (adapters); prepending learnable vectors to the input (prompt tuning); representing weight changes with low-rank matrices (LoRA); and unfreezing only a few of the original parameters, such as just the biases. It addresses the high cost of full fine-tuning: memory, compute, and the storage cost of keeping a full copy of the weights per task all drop sharply, while performance often comes close to full fine-tuning. Hugging Face's PEFT library bundles many of these methods together. When adapting an open-source VLA to one's own robot, LoRA is the most commonly used choice.","example":"In the OpenVLA paper, using LoRA to update only about 1.4% of the parameters, trained for 10 to 15 hours on a single A100, reached a 68.2% task success rate, close to full fine-tuning's 69.7%, which needed two GPUs and about 163 GB of memory.","related":["LoRA","Adapter","Prompt Tuning / Soft Prompt","Full Fine-Tuning","Backbone Freezing","Fine-tuning"]},{"id":"lora","category":"training","sec":9,"tier":2,"sources":[{"title":"LoRA: Low-Rank Adaptation of Large Language Models (arXiv 2106.09685)","url":"https://arxiv.org/abs/2106.09685"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"},{"title":"openpi（Physical Intelligence 官方仓库，显存需求表）","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["parameter-efficient-fine-tuning","full-fine-tuning","adapter","fine-tuning","backbone-freezing","openvla"],"name":"LoRA","alt":"低秩适配","abbr":"LoRA","aliases":["Low-Rank Adaptation"],"one_liner":"Freezing a large model's original weights and fine-tuning it by training only two small matrices inserted alongside them.","explanation":"LoRA was introduced by Microsoft's Edward Hu and colleagues in 2021 and is the most widely used parameter-efficient fine-tuning method. It freezes the pretrained weight matrix W and adds a side path B·A next to selected layers (commonly the projection matrices inside a Transformer's attention blocks): A and B are two narrow matrices whose rank r is far smaller than the original matrix's dimensions, and only A and B are trained during fine-tuning. After training, B·A can be added straight back into W, adding no extra latency at inference. The paper reports roughly a 10,000-fold reduction in trainable parameters and about a third of the memory requirement compared with fully fine-tuning GPT-3 175B with Adam. For embodied AI, since most VLAs are built on multi-billion-parameter vision-language models, LoRA lets an ordinary lab adapt a model to its own robot and tasks on a single consumer GPU.","example":"In the OpenVLA paper, a rank-32 LoRA trains only about 1.4% of the parameters and reaches 68.2% success, close to full fine-tuning's 69.7%. The openpi repository lists π0 LoRA fine-tuning as needing 22.5 GB of memory or more, running on an RTX 4090, versus 70 GB or more for full fine-tuning.","related":["Parameter-Efficient Fine-Tuning","Full Fine-Tuning","Adapter","Fine-tuning","Backbone Freezing","OpenVLA"]},{"id":"adapter","category":"training","sec":9,"tier":3,"sources":[{"title":"Parameter-Efficient Transfer Learning for NLP (Houlsby et al., arXiv 1902.00751)","url":"https://arxiv.org/abs/1902.00751"},{"title":"TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models (arXiv 2310.05905)","url":"https://arxiv.org/abs/2310.05905"}],"as_of":"","related_ids":["parameter-efficient-fine-tuning","lora","backbone-freezing","catastrophic-forgetting","full-fine-tuning","prompt-tuning-soft-prompt"],"name":"Adapter","alt":"适配器","abbr":"","aliases":["Adapter Tuning","Bottleneck Adapter"],"one_liner":"Small modules inserted between a frozen large model's layers, with only those small modules trained to adapt it to a new task.","explanation":"Adapters are a family of parameter-efficient fine-tuning methods, introduced by Houlsby and colleagues in the 2019 paper “Parameter-Efficient Transfer Learning for NLP”: a small bottleneck network — project down, then back up, with a residual connection — is inserted into each Transformer layer, and fine-tuning freezes the original model and trains only these adapters. On 26 text-classification tasks with BERT, adding just 3.6% extra parameters per task matched full fine-tuning within 0.4%. Adapters save memory and storage, let one base model carry adapters for many tasks at once, and reduce catastrophic forgetting; LoRA, introduced later, continues the same basic idea. In robotics, TAIL (ICLR 2024) compares bottleneck adapters, P-Tuning, and LoRA for imitation learning, and this family of methods is commonly used to adapt a VLA to a new task or new robot.","example":"TAIL fine-tunes a pretrained decision-making model on a small number of demonstrations and finds that LoRA, training only about 1% of the parameters, matches full fine-tuning's performance while avoiding catastrophic forgetting during continual learning.","related":["Parameter-Efficient Fine-Tuning","LoRA","Backbone Freezing","Catastrophic Forgetting","Full Fine-Tuning","Prompt Tuning / Soft Prompt"]},{"id":"prompt-tuning-soft-prompt","category":"training","sec":9,"tier":3,"sources":[{"title":"Lester, Al-Rfou, Constant 2021: The Power of Scale for Parameter-Efficient Prompt Tuning","url":"https://arxiv.org/abs/2104.08691"},{"title":"Zheng et al. 2025: X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment VLA Model","url":"https://arxiv.org/abs/2510.10274"}],"as_of":"2025-10","related_ids":["parameter-efficient-fine-tuning","prompt-prompt-engineering","lora","adapter","x-vla","embedding"],"name":"Prompt Tuning / Soft Prompt","alt":"提示微调 / 软提示","abbr":"","aliases":["Soft Prompt","Soft Prompt Tuning"],"one_liner":"Freezing the whole model and training only a short learnable vector prepended to the input to adapt it to a task.","explanation":"Prompt tuning is a parameter-efficient fine-tuning method proposed by Google's Lester and colleagues in 2021. A hand-written text prompt is a “hard prompt”; a soft prompt, by contrast, is a sequence of vectors learned directly in embedding space that does not correspond to any actual words. During training, the model's parameters stay entirely frozen, and only this short vector sequence is updated by gradient descent, so each task needs only a tiny soft prompt stored for it, and one model can serve many tasks. The paper found that as models get larger, performance gets closer to full fine-tuning, essentially matching it at billions of parameters. Prefix tuning, which prepends a learnable prefix before every attention layer, and visual prompt tuning are related ideas. In robotics, soft prompts have also been used to mark different embodiments or data sources.","example":"X-VLA gives each data source — different robot embodiments and collection setups — its own learnable soft-prompt embedding on top of a standard Transformer backbone, letting a 0.9-billion-parameter model train across embodiments and get tested on multiple simulated and real-robot platforms.","related":["Parameter-Efficient Fine-Tuning","Prompt / Prompt Engineering","LoRA","Adapter","X-VLA","Embedding"]},{"id":"supervised-fine-tuning","category":"training","sec":9,"tier":2,"sources":[{"title":"Training language models to follow instructions with human feedback (InstructGPT, arXiv 2203.02155)","url":"https://arxiv.org/abs/2203.02155"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"","related_ids":["fine-tuning","behavior-cloning","post-training","instruction-tuning","reinforcement-fine-tuning","reinforcement-learning-from-human-feedback"],"name":"Supervised Fine-Tuning","alt":"监督微调","abbr":"SFT","aliases":["SFT"],"one_liner":"Continuing supervised training on a pretrained model using paired input–correct-answer data, to teach it a specific behavior.","explanation":"Supervised fine-tuning takes a pretrained model and keeps training it on a curated, high-quality labeled dataset to teach it some specific behavior; the training objective is simply to make its output match the given answer as closely as possible (cross-entropy for language output, often L1 or mean squared error for continuous actions). The term became popular through OpenAI's InstructGPT (2022): first run SFT on human-written example answers, then train a reward model and run RLHF — the now-classic three-step recipe for post-training a large model. In robotics, SFT usually means fine-tuning a foundation model like a VLA on teleoperation demonstrations for a specific robot and set of tasks, which is essentially behavior cloning. It's simple and stable, but it only imitates the demonstrations, so it tends to err on states the demonstrations never showed — which is why SFT is often followed by RL fine-tuning.","example":"Stanford's OpenVLA-OFT improves on how OpenVLA is fine-tuned — parallel decoding, action chunking, and continuous action output with L1 regression — raising the success rate on the LIBERO simulation benchmark from 76.5% to 97.1% while increasing action-generation throughput 26-fold.","related":["Fine-tuning","Behavior Cloning","Post-training","Instruction Tuning","Reinforcement Fine-Tuning (RL Fine-Tuning)","Reinforcement Learning from Human Feedback"]},{"id":"instruction-tuning","category":"training","sec":9,"tier":3,"sources":[{"title":"Finetuned Language Models Are Zero-Shot Learners (FLAN, arXiv:2109.01652)","url":"https://arxiv.org/abs/2109.01652"},{"title":"Visual Instruction Tuning (LLaVA, arXiv:2304.08485)","url":"https://arxiv.org/abs/2304.08485"}],"as_of":"","related_ids":["supervised-fine-tuning","large-language-model","vision-language-model","llava","zero-shot","post-training"],"name":"Instruction Tuning","alt":"指令微调","abbr":"","aliases":["Instruction Fine-tuning"],"one_liner":"Fine-tuning a model on many tasks written as instruction-and-answer pairs so it learns to follow instructions.","explanation":"Instruction tuning is supervised fine-tuning of a pretrained model on a large set of tasks phrased as natural-language instructions paired with the expected output. Google's 2021 FLAN work instruction-tuned a 137-billion-parameter model across more than 60 NLP tasks, and its zero-shot performance beat GPT-3 on 20 of 25 evaluation tasks. It turns a model from one that merely continues text into one that follows requests, making it a basic building block of chat assistants. In multimodal models, LLaVA (2023) used GPT-4 to generate image-text instruction data for “visual instruction tuning”; many VLA backbones go through a comparable step, and this kind of data is often mixed into VLA training to reduce forgetting.","example":"LLaVA (2023) used GPT-4 to rewrite image captions into “look at the image and answer the question”-style instruction data, then fine-tuned a vision-encoder-plus-language-model combination on it to produce an assistant that can converse about images.","related":["Supervised Fine-Tuning","Large Language Model","Vision-Language Model","LLaVA","Zero-shot","Post-training"]},{"id":"rejection-sampling-fine-tuning","category":"training","sec":9,"tier":3,"sources":[{"title":"Yuan et al. 2023: Scaling Relationship on Learning Mathematical Reasoning with Large Language Models","url":"https://arxiv.org/abs/2308.01825"},{"title":"Chen et al. 2021: Decision Transformer: Reinforcement Learning via Sequence Modeling","url":"https://arxiv.org/abs/2106.01345"},{"title":"Li et al. 2025: GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation","url":"https://arxiv.org/abs/2512.01801"}],"as_of":"2025-12","related_ids":["behavior-cloning","supervised-fine-tuning","self-improvement","best-of-n-sampling","suboptimal-demonstrations","advantage-conditioning"],"name":"Rejection Sampling Fine-Tuning","alt":"拒绝采样微调 / 过滤式行为克隆","abbr":"","aliases":["RFT","Filtered Behavior Cloning","Filtered BC","Percentile Behavior Cloning (%BC)"],"one_liner":"Letting a model attempt a task many times, keeping only the successful or high-scoring results, and training on those.","explanation":"Rejection sampling fine-tuning and filtered behavior cloning describe the same basic idea: let an existing model generate multiple results for a task, use answer-checking, success detection, or return ranking to throw out the bad ones, and use only the good samples as supervised fine-tuning or behavior-cloning data. In large-model work, Yuan and colleagues' 2023 paper is often cited, using it to expand math-reasoning training data; the 2021 Decision Transformer paper also used “percentile behavior cloning” (%BC), cloning only the top X% of data by return, as a baseline. It is simple to implement and trains stably, though information in the failed samples is simply thrown away. In robotics it is commonly used to keep only successful autonomous trajectories, or, as with ByteDance's GR-RL, to use a learned task-progress function to filter out demonstration segments that made no progress on the task. Note that the abbreviation RFT is also commonly used for RL fine-tuning.","example":"Yuan and colleagues had multiple models answer GSM8K math problems repeatedly, keeping only the reasoning traces that reached the correct answer for the training set; LLaMA-7B's accuracy rose from 35.9% with plain supervised fine-tuning to 49.3%.","related":["Behavior Cloning","Supervised Fine-Tuning","Self-improvement","Best-of-N Sampling","Suboptimal (Noisy) Demonstrations","Advantage Conditioning"]},{"id":"cold-start","category":"training","sec":9,"tier":2,"sources":[{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv 2501.12948)","url":"https://arxiv.org/html/2501.12948v1"},{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv 2509.09674)","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"2025-09","related_ids":["supervised-fine-tuning","reinforcement-fine-tuning","behavior-cloning","post-training","sparse-reward","simplevla-rl"],"name":"Cold Start","alt":"冷启动","abbr":"","aliases":["Cold-start SFT","Cold-start Data","Cold-start Fine-tuning"],"one_liner":"Running supervised fine-tuning on a small set of high-quality demonstrations before reinforcement learning, to give the model a decent starting point.","explanation":"In large-model and VLA training, a cold start means running supervised fine-tuning (SFT — training directly on correct-answer data) on a small amount of high-quality data before reinforcement learning begins, producing an initial model that can already attempt the task reasonably. The term became popular because of DeepSeek-R1 (January 2025): R1-Zero, trained with pure RL, produced hard-to-read output that mixed languages, so the team first fine-tuned on a few thousand long chain-of-thought examples before running RL. This matters because reinforcement learning relies on trial and error, and an initial model that almost never succeeds gets almost no reward signal, making training slow and unstable. The embodied-AI equivalent is training a base policy with behavior cloning on demonstration data first, then fine-tuning it with RL in simulation or on a real robot. Note that “cold start” in recommender systems — a new user with no history — is an unrelated meaning of the same term.","example":"SimpleVLA-RL (2025) first ran SFT with just one demonstration per task, reaching 17.3% success on LIBERO-Long; using that as the starting point for reinforcement learning then raised it to 91.7%.","related":["Supervised Fine-Tuning","Reinforcement Fine-Tuning (RL Fine-Tuning)","Behavior Cloning","Post-training","Sparse Reward","SimpleVLA-RL"]},{"id":"reinforcement-fine-tuning","category":"training","sec":9,"tier":2,"sources":[{"title":"OpenAI API Docs: Reinforcement fine-tuning","url":"https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning"},{"title":"πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models","url":"https://arxiv.org/abs/2510.25889"},{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"2026-01","related_ids":["post-training","supervised-fine-tuning","proximal-policy-optimization","group-relative-policy-optimization","pirl","simplevla-rl"],"name":"Reinforcement Fine-Tuning (RL Fine-Tuning)","alt":"强化学习微调","abbr":"RFT","aliases":["RFT","RL Fine-tuning","RL Post-training"],"one_liner":"Further improving a pretrained or supervised-fine-tuned model with reinforcement learning driven by a reward signal.","explanation":"Reinforcement fine-tuning takes a model that's already been pretrained or supervised-fine-tuned (SFT) as its starting policy, lets it generate its own answers or actions, scores them with a reward, and updates the parameters with an algorithm such as policy gradient. Unlike SFT, which only imitates demonstrations, this lets the model learn from its own successes and failures, discovering behaviors the demonstrations never showed. The term “RFT” became popular through OpenAI's reinforcement fine-tuning service for its o-series reasoning models: the user supplies a grader, and the model samples several answers per question and reinforces the higher-scoring ones. Since 2025, a wave of work has applied this to VLA models, usually with a binary success/failure reward and algorithms like PPO or GRPO; because a flow-matching action head can't compute an action's probability directly, it needs special adaptation, as in πRL.","example":"πRL fine-tunes flow-matching VLAs with online reinforcement learning: π0 first reaches 57.6% success on LIBERO after SFT on a handful of demonstrations, then rises to 97.6% after reinforcement learning; π0.5 goes from 77.1% to 98.3%. SimpleVLA-RL instead trains OpenVLA-OFT with just a “1 for success, 0 for failure” reward combined with GRPO.","related":["Post-training","Supervised Fine-Tuning","Proximal Policy Optimization","Group Relative Policy Optimization","πRL","SimpleVLA-RL"]},{"id":"reinforcement-learning-from-human-feedback","category":"training","sec":9,"tier":2,"sources":[{"title":"Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017)","url":"https://arxiv.org/abs/1706.03741"},{"title":"Training language models to follow instructions with human feedback (InstructGPT)","url":"https://arxiv.org/abs/2203.02155"},{"title":"Wikipedia: Reinforcement learning from human feedback","url":"https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback"}],"as_of":"","related_ids":["reward-model","direct-preference-optimization","proximal-policy-optimization","kl-regularization","reward-hacking","grape"],"name":"Reinforcement Learning from Human Feedback","alt":"基于人类反馈的强化学习","abbr":"RLHF","aliases":["RLHF","Preference Alignment"],"one_liner":"Having people compare pairs of model outputs to train a reward model, then using reinforcement learning to optimize toward human preference.","explanation":"RLHF was introduced by Christiano and colleagues (OpenAI and DeepMind) in 2017: a reward function can be learned just from people comparing which of two trajectories is better, needing human feedback on less than 1% of the total interactions. OpenAI's InstructGPT (2022) applied it to large language models in three steps: supervised fine-tuning; training a reward model on human rankings of multiple candidate answers; and optimizing the model with PPO, adding a KL penalty to keep it from drifting too far from the original model — a recipe ChatGPT continued to use. It suits goals that are hard to write as a formula but easy for a person to judge at a glance; the downsides are that labeling is expensive and the model can learn to exploit the reward model's blind spots. DPO skips the reward model and trains directly on preference data instead; in embodied AI, GRAPE aligns a VLA using trajectory-level preferences.","example":"In InstructGPT's human evaluations, a version with only 1.3 billion parameters that went through RLHF was preferred over the 175-billion-parameter GPT-3.","related":["Reward Model","Direct Preference Optimization","Proximal Policy Optimization","KL Regularization","Reward Hacking","GRAPE"]},{"id":"direct-preference-optimization","category":"training","sec":9,"tier":3,"sources":[{"title":"Direct Preference Optimization: Your Language Model is Secretly a Reward Model (arXiv:2305.18290)","url":"https://arxiv.org/abs/2305.18290"},{"title":"GRAPE: Generalizing Robot Policy via Preference Alignment (arXiv:2411.19309)","url":"https://arxiv.org/abs/2411.19309"}],"as_of":"","related_ids":["reinforcement-learning-from-human-feedback","reward-model","kl-regularization","grape","post-training","group-relative-policy-optimization"],"name":"Direct Preference Optimization","alt":"直接偏好优化","abbr":"DPO","aliases":["DPO"],"one_liner":"Aligning a model directly from paired “better/worse” samples, without training a reward model or running reinforcement learning.","explanation":"DPO was introduced by Stanford's Rafailov, Finn, and colleagues in 2023. Traditional RLHF (reinforcement learning from human feedback) first trains a reward model on preference data, then runs reinforcement learning with an algorithm like PPO — a long, often unstable pipeline. DPO proves that the optimal policy for a KL-constrained reward-maximization problem has a closed-form solution, which turns the problem into a classification-style loss: raise the probability of the preferred sample relative to a reference model, and lower the probability of the rejected one, with no need to sample from the model during training. It was first used to align large language models, and embodied AI has borrowed it for VLA post-training: pairing successful and failed trajectories into preferences and optimizing the policy directly.","example":"GRAPE (2024) extends DPO from single steps to whole trajectories (called TPO) on OpenVLA: it pairs successful and failed manipulation trajectories into preferences to fine-tune the model, using a step-wise DPO variant, OpenVLA-DPO, as a comparison baseline.","related":["Reinforcement Learning from Human Feedback","Reward Model","KL Regularization","GRAPE","Post-training","Group Relative Policy Optimization"]},{"id":"reinforcement-learning-with-verifiable-rewards","category":"training","sec":9,"tier":3,"sources":[{"title":"Lambert et al. 2024: Tülu 3: Pushing Frontiers in Open Language Model Post-Training","url":"https://arxiv.org/abs/2411.15124"},{"title":"DeepSeek-AI 2025: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","url":"https://arxiv.org/abs/2501.12948"},{"title":"Li et al. 2025: SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning","url":"https://arxiv.org/abs/2509.09674"}],"as_of":"2025-09","related_ids":["reinforcement-fine-tuning","reward-model","reinforcement-learning-from-human-feedback","group-relative-policy-optimization","sparse-reward","simplevla-rl"],"name":"Reinforcement Learning with Verifiable Rewards","alt":"基于可验证奖励的强化学习","abbr":"RLVR","aliases":["RLVR"],"one_liner":"Training a model with reinforcement learning using rewards a program can automatically check as correct or wrong.","explanation":"The name RLVR comes from the Allen Institute for AI's (Ai2) Tülu 3 paper in November 2024: for tasks such as math problems or automatically checkable instruction constraints, a rule or program judges whether the model's output is correct, gives a fixed reward for a correct answer and zero otherwise, and then updates the model with an algorithm such as PPO. It needs no separately trained reward model, the network that scores answers in RLHF, so the reward source is more reliable and harder for the model to game. DeepSeek-R1 (2025) used a rule-based accuracy reward to train reasoning ability, bringing this approach wide attention. Applied to robots, “did the task succeed or not” is a naturally verifiable reward, and work such as SimpleVLA-RL and VLA-RFT uses a binary success/failure reward to fine-tune VLAs with reinforcement learning.","example":"SimpleVLA-RL fine-tunes OpenVLA-OFT with reinforcement learning using only an outcome reward — success is rewarded, failure is not — reaching leading results on LIBERO, beating π0 on RoboTwin 1.0 and 2.0, and also outperforming a purely supervised-fine-tuned version on a real robot.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Reward Model","Reinforcement Learning from Human Feedback","Group Relative Policy Optimization","Sparse Reward","SimpleVLA-RL"]},{"id":"kl-regularization","category":"training","sec":9,"tier":3,"sources":[{"title":"Training language models to follow instructions with human feedback (InstructGPT, arXiv:2203.02155)","url":"https://arxiv.org/abs/2203.02155"},{"title":"Behavior Regularized Offline Reinforcement Learning (BRAC, arXiv:1911.11361)","url":"https://arxiv.org/abs/1911.11361"}],"as_of":"","related_ids":["kullback-leibler-divergence","policy-constraint","offline-reinforcement-learning","reinforcement-learning-from-human-feedback","proximal-policy-optimization","reward-hacking"],"name":"KL Regularization","alt":"KL 正则化","abbr":"","aliases":["KL Penalty","KL Constraint"],"one_liner":"Adding a KL-divergence penalty to a training objective so the new policy doesn't drift too far from a reference.","explanation":"KL regularization adds a term to the optimization objective — a coefficient times the KL divergence, a measure of how different two probability distributions are — that penalizes the current policy for drifting away from some reference policy. Three settings are common. In RLHF, InstructGPT adds a per-token KL penalty against the supervised fine-tuned model to stop the model from gaming the reward model. In offline reinforcement learning, methods such as BRAC use KL or similar divergences to keep the policy close to the behavior policy that generated the data, avoiding actions the data never covers. TRPO and PPO use KL to bound how much each update can change the policy. VLA reinforcement-learning fine-tuning commonly uses it too, to stop the policy from drifting and losing pretrained capabilities. Too large a coefficient stalls learning; too small and the constraint does nothing.","example":"InstructGPT's PPO objective is “reward-model score minus β times log(π_RL / π_SFT)”: the larger β is, the less the new model dares to drift from the supervised fine-tuned model.","related":["Kullback-Leibler Divergence","Policy Constraint","Offline Reinforcement Learning","Reinforcement Learning from Human Feedback","Proximal Policy Optimization","Reward Hacking"]},{"id":"catastrophic-forgetting","category":"training","sec":9,"tier":2,"sources":[{"title":"Wikipedia: Catastrophic interference","url":"https://en.wikipedia.org/wiki/Catastrophic_interference"},{"title":"Knowledge Insulating Vision-Language-Action Models (arXiv 2505.23705)","url":"https://arxiv.org/html/2505.23705"},{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054)","url":"https://arxiv.org/abs/2504.16054"}],"as_of":"2025-05","related_ids":["continual-learning","backbone-freezing","co-training","knowledge-insulation","stop-gradient","parameter-efficient-fine-tuning"],"name":"Catastrophic Forgetting","alt":"灾难性遗忘","abbr":"","aliases":["Catastrophic Interference"],"one_liner":"A neural network's sharp loss, or outright loss, of old abilities after it learns something new.","explanation":"Catastrophic forgetting, also called catastrophic interference, was first systematically reported by McCloskey and Cohen in 1989, when a backpropagation network trained on one set of addition problems and then another lost much of its performance on the first. Learning new material changes the same weights that encoded old knowledge, so performance on the old task drops sharply. Mitigations include replaying old data, elastic weight consolidation (EWC, which protects weights that matter most for old tasks), freezing part of the network, and parameter-efficient fine-tuning. In embodied AI, converting a pretrained VLM into a VLA and fine-tuning it only on robot data often erodes the model's original language understanding and general knowledge. That's why models like π0.5 co-train on a mix of web image-text data and robot data, and why Physical Intelligence's knowledge insulation approach blocks the action expert's gradients from flowing back into the VLM.","example":"Physical Intelligence's knowledge insulation paper found that directly attaching a continuous action expert to a VLM and training the whole thing together noticeably erodes the model's pretrained knowledge and weakens its ability to understand language instructions; they mitigate this with gradient blocking plus co-training on image-text QA data.","related":["Continual Learning","Backbone Freezing","Co-training","Knowledge Insulation","Stop-Gradient","Parameter-Efficient Fine-Tuning"]},{"id":"continual-learning","category":"training","sec":9,"tier":3,"sources":[{"title":"A Comprehensive Survey of Continual Learning: Theory, Method and Application (arXiv 2302.00487)","url":"https://arxiv.org/abs/2302.00487"},{"title":"Overcoming catastrophic forgetting in neural networks (EWC, arXiv 1612.00796)","url":"https://arxiv.org/abs/1612.00796"},{"title":"LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning (arXiv 2306.03310)","url":"https://arxiv.org/abs/2306.03310"}],"as_of":"","related_ids":["catastrophic-forgetting","transfer-learning","fine-tuning","multi-task-learning","libero-benchmark","knowledge-insulation"],"name":"Continual Learning","alt":"持续学习","abbr":"","aliases":["Lifelong Learning","Incremental Learning"],"one_liner":"Letting a model learn a sequence of new tasks and new data without forgetting what it already learned.","explanation":"Continual learning studies how a model can keep accumulating and updating knowledge as it continually receives new tasks and new data throughout its whole deployment lifetime. The biggest obstacle is catastrophic forgetting: a neural network's performance on old tasks often drops sharply once it keeps training on new data. The core trade-off is the stability-plasticity balance — protecting old knowledge while still being able to learn new things. Common approaches include regularization (such as DeepMind's 2017 EWC, which slows updates to weights important for old tasks), replay (mixing old data, or generated stand-ins for it, back into training), and structural expansion (adding separate parameters for each new task). For robots, deployment means constantly running into new objects and new scenes, and the ideal is learning on the fly rather than retraining from scratch every time; the 2023 LIBERO benchmark was designed specifically for lifelong robot learning, with 4 task suites totaling 130 tasks. A VLA losing its original VLM's language ability and common sense after fine-tuning is also a form of forgetting.","example":"A home robot first learns to fold towels, then later learns to open the refrigerator and wash dishes; continual learning requires that after it learns dishwashing, its success rate at folding towels doesn't drop noticeably.","related":["Catastrophic Forgetting","Transfer Learning","Fine-tuning","Multi-Task Learning","LIBERO Benchmark","Knowledge Insulation"]},{"id":"knowledge-insulation","category":"training","sec":9,"tier":3,"sources":[{"title":"Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better (arXiv:2505.23705)","url":"https://arxiv.org/abs/2505.23705"},{"title":"论文 HTML 版（方法与实验细节）","url":"https://arxiv.org/html/2505.23705"}],"as_of":"2025-05","related_ids":["stop-gradient","action-expert","pi0-5","pi0-fast","catastrophic-forgetting","co-training"],"name":"Knowledge Insulation","alt":"知识隔离","abbr":"KI","aliases":["KI","Knowledge Insulating VLA"],"one_liner":"A VLA training trick that blocks gradients from the action expert back into the VLM backbone, protecting its pretrained knowledge.","explanation":"Knowledge insulation is a VLA training method from a May 2025 Physical Intelligence paper by Danny Driess, Sergey Levine, and colleagues. Models such as π0 attach a newly initialized action expert to a VLM backbone and use flow matching to output continuous actions, but gradients flowing back from the action expert into the backbone disrupt it, which slows training and hurts language understanding. The fix has three parts: a stop-gradient cuts off gradients from the action expert to the backbone; the backbone instead learns actions through next-token prediction on discretized FAST action tokens, while co-training on ordinary VLM data such as visual question answering; and at inference time, the action expert still generates continuous action chunks quickly from the backbone's features.","example":"On a table-bussing task, the paper reports that the knowledge-insulated model converges about 7.5 times faster than a π0 trained purely with flow matching, while keeping better language-instruction-following ability.","related":["Stop-Gradient","Action Expert","π0.5","π0-FAST","Catastrophic Forgetting","Co-training"]},{"id":"model-merging","category":"training","sec":9,"tier":3,"sources":[{"title":"Yadav et al. 2025: Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging (RETAIN)","url":"https://arxiv.org/abs/2512.08333"},{"title":"Wortsman et al. 2022: Model soups","url":"https://arxiv.org/abs/2203.05482"},{"title":"Ilharco et al. 2022: Editing Models with Task Arithmetic","url":"https://arxiv.org/abs/2212.04089"}],"as_of":"2025-12","related_ids":["catastrophic-forgetting","fine-tuning","continual-learning","lora","exponential-moving-average","checkpoint"],"name":"Model Merging","alt":"模型合并","abbr":"","aliases":["Weight Interpolation","Weight Averaging","Model Soups","Task Arithmetic"],"one_liner":"Combining several related models' parameters by weighted averaging, with no extra training required.","explanation":"Model merging combines multiple models that share the same architecture — usually all fine-tuned from the same pretrained model — directly at the parameter level, needing no extra training and adding no inference cost. The simplest version is linear interpolation: new parameters equal (1−α) times the pretrained parameters plus α times the fine-tuned parameters. Notable examples include WiSE-FT, which interpolates a zero-shot model with a fine-tuned one to preserve robustness; Model Soups, which averages the results of fine-tuning runs with different hyperparameters; and task arithmetic, which treats “fine-tuned minus pretrained” as a task vector that can be added or subtracted. In robotics it is used to fight the catastrophic forgetting caused by fine-tuning: Sergey Levine's team's 2025 RETAIN interpolates a fine-tuned π0-FAST-DROID with the original model, keeping both the new skill and the model's original generalization. Other work, however, has found that directly merging VLAs each fine-tuned on a different task can push success rate close to zero.","example":"RETAIN fine-tunes π0-FAST-DROID on a “wipe the whiteboard with an eraser” task using a few dozen demonstrations, then interpolates it with the original model at an α between 0.25 and 0.75; the merged model clearly beats the plain fine-tuned version on out-of-distribution variations such as new objects and new camera angles.","related":["Catastrophic Forgetting","Fine-tuning","Continual Learning","LoRA","Exponential Moving Average","Checkpoint"]},{"id":"knowledge-distillation","category":"training","sec":9,"tier":2,"sources":[{"title":"Distilling the Knowledge in a Neural Network (arXiv 1503.02531)","url":"https://arxiv.org/abs/1503.02531"},{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Science Robotics 2020, arXiv 2010.11251)","url":"https://arxiv.org/abs/2010.11251"}],"as_of":"","related_ids":["teacher-student-distillation","privileged-information","policy-distillation","on-policy-distillation","kullback-leibler-divergence","diffusion-step-distillation"],"name":"Knowledge Distillation","alt":"知识蒸馏","abbr":"KD","aliases":["KD","Model Distillation"],"one_liner":"Training a small student model to mimic a large teacher model's outputs, compressing the teacher's ability into a smaller model.","explanation":"Knowledge distillation was systematically formalized by Hinton, Vinyals, and Dean in a 2015 paper: first train a large, strong teacher model, then train a small student model to match the teacher's output probability distribution (called soft labels) rather than only the ground-truth labels. The paper raises the softmax “temperature” to make the soft labels smoother, which exposes similarity relationships between classes so the student learns more from each example. This addresses the problem that large models are expensive to deploy and slow to run inference on. In embodied AI, distillation often takes a teacher-student form: a teacher trained in simulation with privileged information unavailable on a real robot (such as exact terrain shape or contact state), and a student that only sees real-robot sensors and learns to imitate the teacher's actions — this is one route to sim-to-real transfer. It's also used to distill a multi-step denoising diffusion policy into a model that needs far fewer steps, cutting inference latency.","example":"In Lee and colleagues' 2020 Science Robotics work on the ANYmal quadruped, the teacher policy sees ground-truth terrain and foot-contact information in simulation, while the student policy takes only a history of onboard proprioception and learns by imitating the teacher, eventually deploying zero-shot on mud, snow, and rubble in the wild.","related":["Teacher-Student Distillation","Privileged Information","Policy Distillation","On-Policy Distillation","Kullback-Leibler Divergence","Diffusion Step Distillation"]},{"id":"policy-distillation","category":"training","sec":9,"tier":3,"sources":[{"title":"Rusu et al. 2015: Policy Distillation","url":"https://arxiv.org/abs/1511.06295"},{"title":"Wan et al. 2023: UniDexGrasp++","url":"https://arxiv.org/abs/2304.00464"},{"title":"He et al. 2024: HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots","url":"https://arxiv.org/abs/2410.21229"}],"as_of":"","related_ids":["knowledge-distillation","teacher-student-distillation","specialist-policy","generalist-policy","dagger","on-policy-distillation"],"name":"Policy Distillation","alt":"策略蒸馏","abbr":"","aliases":["Specialist-to-Generalist Distillation"],"one_liner":"Having a student policy imitate one or more already-trained teacher policies, compressing or merging their skills.","explanation":"Policy distillation is knowledge distillation applied to decision-making: instead of learning directly from reward, the student policy fits the actions or action distributions the teacher policy outputs at each state. DeepMind's Rusu and colleagues systematically proposed this on Atari games in 2015, for two purposes: compressing a large network into a small one, and merging several single-game experts into one multi-task policy whose combined performance beats each expert trained alone. A common robotics pattern is “specialist to generalist”: train separate specialist policies for different objects or tasks, sometimes using privileged information (ground-truth state only available in simulation), and then distill them into one generalist policy, often paired with DAgger so the student queries the teacher for labels at the states it actually visits.","example":"Peking University's He Wang and colleagues' UniDexGrasp++ first groups thousands of objects by geometric features, trains a specialist grasping policy for each group, then iteratively distills them into one generalist policy, ultimately reaching grasp success rates of 85.4% on training objects and 78.2% on held-out ones.","related":["Knowledge Distillation","Teacher-Student Distillation","Specialist Policy","Generalist Policy","DAgger","On-Policy Distillation"]},{"id":"on-policy-distillation","category":"training","sec":9,"tier":3,"sources":[{"title":"Thinking Machines Lab (Kevin Lu), 2025-10-27: On-Policy Distillation","url":"https://thinkingmachines.ai/blog/on-policy-distillation/"},{"title":"Agarwal et al. 2023: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD)","url":"https://arxiv.org/abs/2306.13649"},{"title":"VLA-OPD: Bridging Offline SFT and Online RL for VLA Models via On-Policy Distillation (arXiv 2603.26666)","url":"https://arxiv.org/abs/2603.26666"}],"as_of":"2026-03","related_ids":["knowledge-distillation","policy-distillation","dagger","teacher-student-distillation","reinforcement-fine-tuning","vla-opd"],"name":"On-Policy Distillation","alt":"在线策略蒸馏","abbr":"OPD","aliases":["OPD","Generalized Knowledge Distillation (GKD)"],"one_liner":"Distillation where the student generates its own trajectories, and the teacher scores and corrects it at every step.","explanation":"Ordinary knowledge distillation has the student imitate data the teacher already generated, but once deployed the student runs into states that were never in the teacher's data, and errors accumulate — the same problem as compounding error in behavior cloning. On-policy distillation instead has the student generate its own samples, and the teacher gives a target distribution at every step the student actually reaches, whether that is every token or every action, usually measured with reverse KL divergence, an idea in the same spirit as DAgger. Google DeepMind's Agarwal and colleagues systematically proposed this under the name GKD in 2023; a Thinking Machines blog post in October 2025 made it widely known in large-model post-training, where its feedback is far denser than reinforcement learning's sparse reward. A 2026 paper, VLA-OPD, applies the idea to VLA post-training, replacing sparse environment reward with dense, step-by-step supervision from an expert teacher.","example":"According to the Qwen3 technical report cited in the Thinking Machines blog post, on-policy distillation scored 74.4 on AIME'24 using about 1,800 GPU-hours, while plain reinforcement learning scored only 67.6 using about 17,920 GPU-hours.","related":["Knowledge Distillation","Policy Distillation","DAgger","Teacher-Student Distillation","Reinforcement Fine-Tuning (RL Fine-Tuning)","VLA-OPD"]},{"id":"diffusion-step-distillation","category":"training","sec":9,"tier":3,"sources":[{"title":"One-step Diffusion with Distribution Matching Distillation (arXiv 2311.18828)","url":"https://arxiv.org/abs/2311.18828"},{"title":"Progressive Distillation for Fast Sampling of Diffusion Models (arXiv 2202.00512)","url":"https://arxiv.org/abs/2202.00512"},{"title":"One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation (arXiv 2410.21257)","url":"https://arxiv.org/abs/2410.21257"}],"as_of":"","related_ids":["diffusion-model","denoising-steps","consistency-model","one-step-generation","knowledge-distillation","consistency-policy"],"name":"Diffusion Step Distillation","alt":"扩散步数蒸馏","abbr":"DMD","aliases":["DMD","Distribution Matching Distillation","Progressive Distillation","Few-Step Distillation"],"one_liner":"Compressing a diffusion model that needs dozens to hundreds of denoising steps into a student that produces results in one or a few.","explanation":"Generating one sample from a diffusion model normally takes dozens to thousands of repeated denoising steps, which is slow. Step distillation uses a trained multi-step model as a teacher to train a student that needs far fewer steps. An early example is Google's Salimans and Ho's progressive distillation (ICLR 2022), which halves the number of steps each round, compressing from as many as 8,192 steps down to 4. Distribution Matching Distillation (DMD), from MIT and Adobe's Yin and colleagues (CVPR 2024), doesn't require the student to reproduce the teacher's denoising trajectory sample-by-sample; instead, it pulls the overall output distribution of a one-step generator toward the teacher's, approximating the gradient of the KL divergence using the difference between two diffusion models' scores (the gradient of the data distribution); a later version, DMD2, adds a GAN loss and supports multiple steps. In robotics, a diffusion policy's slow inference can bottleneck the control frequency, so this kind of distillation is commonly used to speed it up; autoregressive video-generation work like CausVid and Self Forcing also uses DMD-style losses to achieve real-time generation. Consistency distillation is another approach to the same problem.","example":"One-Step Diffusion Policy uses a trained diffusion policy as its teacher and distills a one-step generator at an extra cost of just 2%–10% of pretraining, raising the action output rate on real Franka-arm tasks from about 1.5 Hz to 62 Hz.","related":["Diffusion Model","Denoising Steps","Consistency Model","One-step Generation","Knowledge Distillation","Consistency Policy"]},{"id":"massively-parallel-reinforcement-learning","category":"training","sec":10,"tier":2,"sources":[{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/abs/2109.11978"},{"title":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (arXiv 2108.10470)","url":"https://arxiv.org/abs/2108.10470"}],"as_of":"","related_ids":["gpu-accelerated-parallel-simulation","vectorized-environments","proximal-policy-optimization","nvidia-isaac-lab","legged-gym","sim-to-real-transfer"],"name":"Massively Parallel Reinforcement Learning","alt":"大规模并行强化学习","abbr":"","aliases":["Parallel Environment Training","GPU-Parallel Reinforcement Learning"],"one_liner":"Running thousands of simulated environments at once on a single GPU to collect data, training a locomotion policy in just minutes.","explanation":"Massively parallel reinforcement learning means running both the physics simulation and the neural-network training on the GPU, with thousands of independent simulated environments running simultaneously, so a huge amount of interaction data can be collected in one pass to train a policy. Reinforcement learning needs enormous amounts of trial and error, and older CPU-based simulators with few parallel environments often took a dozen to a hundred-plus hours to train a legged-locomotion policy. NVIDIA's Isaac Gym (2021) kept simulation data on the GPU as PyTorch tensors the entire time, reportedly 2–3 orders of magnitude faster than CPU-based simulation; the same year, Rudin and colleagues ran 4,096 parallel ANYmal environments on a single GPU with PPO, training flat-ground walking in under 4 minutes and complex terrain in about 20. This paradigm made reinforcement learning the mainstream approach for legged and humanoid locomotion control, with Isaac Lab and legged_gym among the common tools.","example":"Rudin and colleagues' open-source legged_gym simulates 4,096 ANYmal robots at once on a single RTX A6000, uses a game-style terrain curriculum that adjusts difficulty automatically, trains a policy that walks complex terrain in about 20 minutes, and transfers it to a real ANYmal C.","related":["GPU-Accelerated Parallel Simulation","Vectorized Environments","Proximal Policy Optimization","NVIDIA Isaac Lab","legged_gym","Sim-to-Real Transfer"]},{"id":"actor-learner-architecture","category":"training","sec":10,"tier":3,"sources":[{"title":"IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures (arXiv 1802.01561)","url":"https://arxiv.org/abs/1802.01561"},{"title":"Distributed Prioritized Experience Replay (Ape-X, arXiv 1803.00933)","url":"https://arxiv.org/abs/1803.00933"},{"title":"SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (arXiv 2401.16013)","url":"https://arxiv.org/abs/2401.16013"}],"as_of":"","related_ids":["experience-replay","off-policy","importance-sampling","real-world-reinforcement-learning","serl","massively-parallel-reinforcement-learning"],"name":"Actor-Learner Architecture (Distributed RL)","alt":"Actor-Learner 分离架构","abbr":"","aliases":["Actor-Learner Architecture","Decoupled Sampling and Training"],"one_liner":"Splitting “interacting with the environment to collect data” and “updating the network's parameters” across separate processes or machines, run in parallel.","explanation":"This is a common system design in distributed reinforcement learning: several actors each hold a copy of the policy and interact with their own environment to generate experience; a learner centrally collects that experience and updates the parameters on a GPU, periodically syncing the new parameters back to the actors. DeepMind's Ape-X and IMPALA (2018) are notable examples: Ape-X has actors write experience into a shared replay buffer; IMPALA has actors send trajectories directly to the learner and uses V-trace, an importance-weighted correction, to handle the bias from the actors' policy lagging slightly behind the learner's. The benefit is that sampling and training never block each other, and the system can scale to thousands of machines. Real-robot reinforcement learning needs this same split: SERL runs the actor and learner on separate threads, so training being slow never drags down the robot's control frequency.","example":"SERL runs three processes in parallel on a real robot: the actor picks actions, the learner trains the network, and the robot's environment executes the actions — keeping a fixed control frequency and shortening total real-world training time.","related":["Experience Replay","Off-Policy","Importance Sampling","Real-World Reinforcement Learning","SERL","Massively Parallel Reinforcement Learning"]},{"id":"population-based-training","category":"training","sec":10,"tier":3,"sources":[{"title":"Jaderberg et al. 2017: Population Based Training of Neural Networks","url":"https://arxiv.org/abs/1711.09846"},{"title":"Google DeepMind Blog: Population based training of neural networks","url":"https://deepmind.google/discover/blog/population-based-training-of-neural-networks/"},{"title":"Petrenko et al. 2023: DexPBT","url":"https://arxiv.org/abs/2305.12127"}],"as_of":"","related_ids":["hyperparameter","massively-parallel-reinforcement-learning","exploration-vs-exploitation","gpu-accelerated-parallel-simulation","reinforcement-learning","isaac-gym"],"name":"Population-Based Training","alt":"基于群体的训练","abbr":"PBT","aliases":["PBT"],"one_liner":"Training a whole population of models at once, periodically copying good weights over bad ones and perturbing hyperparameters.","explanation":"Population-based training is a hyperparameter-optimization method proposed by DeepMind's Jaderberg and colleagues in 2017. It trains a batch of models in parallel and periodically compares their performance: worse-performing members directly copy the weights of better-performing ones (exploit), then have hyperparameters such as the learning rate randomly perturbed before continuing to train (explore). This way, there is no need to finish tuning hyperparameters before real training starts — a single run automatically discovers a hyperparameter schedule that changes over time, using roughly the same total compute as running the same number of ordinary parallel experiments. The original paper validated the approach on deep reinforcement learning, machine translation, and GANs. Reinforcement learning is especially sensitive to hyperparameters, so PBT is often paired with massively parallel GPU simulation.","example":"NVIDIA's DexPBT (RSS 2023) uses decentralized PBT in Isaac Gym to train single-arm and dual-arm robots fitted with multi-fingered dexterous hands on tasks such as regrasping, throwing after grasping, and object reorientation, exploring noticeably better than standard end-to-end training.","related":["Hyperparameter","Massively Parallel Reinforcement Learning","Exploration vs. Exploitation","GPU-Accelerated Parallel Simulation","Reinforcement Learning","Isaac Gym"]},{"id":"curriculum-learning","category":"training","sec":10,"tier":2,"sources":[{"title":"Curriculum Learning (Bengio et al., ICML 2009)","url":"https://ronan.collobert.com/pub/2009_curriculum_icml.pdf"},{"title":"Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)","url":"https://arxiv.org/html/2109.11978"}],"as_of":"","related_ids":["terrain-curriculum","reinforcement-learning","automatic-domain-randomization","sparse-reward","legged-gym","massively-parallel-reinforcement-learning"],"name":"Curriculum Learning","alt":"课程学习","abbr":"","aliases":["Curriculum","Automatic Curriculum Learning"],"one_liner":"A training strategy that starts a model on easy samples or tasks and gradually increases the difficulty.","explanation":"Curriculum learning was formally proposed by Bengio and colleagues in a 2009 ICML paper: like a student following a course syllabus, the model first sees easy samples or sub-tasks, and the difficulty is gradually increased. Their experiments showed this speeds up convergence and improves generalization. It's used heavily in robot reinforcement learning, because starting directly on the hardest version of a task means the policy almost never earns reward and can't learn at all. Common forms include terrain curricula (a legged robot practices on flat ground before stairs and slopes), gradually raising the commanded speed, and gradually widening the range of domain randomization; difficulty can also be adjusted automatically based on the policy's current performance, called an automatic curriculum.","example":"legged_gym (Rudin et al., 2021) simulates 4,096 ANYmal quadrupeds at once: a robot that walks past the edge of its current terrain gets a harder terrain next episode, and one that covers less than half the target distance gets an easier one; step height ranges from 5 to 20 cm and slope from 0° to 25° over the curriculum.","related":["Terrain Curriculum","Reinforcement Learning","Automatic Domain Randomization","Sparse Reward","legged_gym","Massively Parallel Reinforcement Learning"]},{"id":"terrain-curriculum","category":"training","sec":10,"tier":3,"sources":[{"title":"Rudin et al. 2021: Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (CoRL 2021)","url":"https://arxiv.org/abs/2109.11978"},{"title":"GitHub: leggedrobotics/legged_gym legged_robot.py（_update_terrain_curriculum）","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"},{"title":"GitHub: isaac-sim/IsaacLab locomotion velocity curriculums.py（terrain_levels_vel）","url":"https://github.com/isaac-sim/IsaacLab/blob/main/source/isaaclab_tasks/isaaclab_tasks/manager_based/locomotion/velocity/mdp/curriculums.py"}],"as_of":"","related_ids":["curriculum-learning","rl-based-locomotion-control","legged-gym","massively-parallel-reinforcement-learning","rough-terrain-locomotion","heightfield-terrain"],"name":"Terrain Curriculum","alt":"地形课程","abbr":"","aliases":["Game-inspired Curriculum"],"one_liner":"In legged reinforcement learning, a training schedule that gradually moves a robot onto harder terrain as it improves.","explanation":"Terrain curriculum is curriculum learning applied specifically to legged locomotion. Training from scratch directly on steep slopes or stairs makes the robot fall almost constantly, yielding little useful learning signal; training only on flat ground never teaches it complex terrain either. ETH's Rudin and colleagues proposed a “game-inspired” curriculum in 2021's legged_gym: terrain is arranged in a grid of rising difficulty, and after each episode, if the robot walked past half the length of its tile it moves up a level, if it didn't even cover half the commanded distance it moves down a level, and reaching the hardest level sends it back to a random earlier one to avoid forgetting easy terrain. Thousands of parallel environments each move up and down independently, so overall difficulty automatically tracks the policy's skill. Isaac Lab's velocity-tracking task follows the same rule.","example":"legged_gym by default splits terrain into 10 difficulty levels across 20 columns, with types including smooth slopes, rough slopes, stairs up, stairs down, and discrete obstacles; a robot starts at a random level between 0 and 5, then moves up or down based on how far it walks each episode.","related":["Curriculum Learning","RL-based Locomotion Control","legged_gym","Massively Parallel Reinforcement Learning","Rough-terrain Locomotion","Heightfield Terrain"]},{"id":"early-termination","category":"training","sec":10,"tier":3,"sources":[{"title":"DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills (arXiv:1804.02717)","url":"https://arxiv.org/abs/1804.02717"},{"title":"legged_gym: legged_robot.py (check_termination)","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot.py"}],"as_of":"","related_ids":["termination-vs-truncation","reference-state-initialization","reward-shaping","episode","deepmimic","massively-parallel-reinforcement-learning"],"name":"Early Termination","alt":"提前终止","abbr":"","aliases":["ET","Termination Condition"],"one_liner":"Ending an episode immediately and resetting the environment as soon as a failure state, like falling, occurs during training.","explanation":"Early termination is a common trick in reinforcement-learning training: an episode doesn't have to run for its full fixed length — as soon as a preset condition triggers, such as the torso touching the ground, the body tilting past some angle, or tracking error growing too large, the episode ends immediately and resets, with no more reward for the time that would have remained. Peng and colleagues' 2018 DeepMimic evaluated this specifically and found it, together with reference state initialization, to be key to letting a simulated character learn highly dynamic skills like backflips. It acts partly as an implicit penalty, since the policy learns to actively avoid failure, and partly as a way to avoid wasting samples on useless post-failure states. Implementations need to distinguish failure termination from timeout truncation — the latter isn't a failure, and value estimation typically still bootstraps through it.","example":"In legged_gym's legged-robot environments, an environment is judged to have fallen and reset as soon as the contact force on a designated body part exceeds a threshold; exceeding the maximum episode length is instead recorded separately as a time_out, with no termination penalty.","related":["Termination vs. Truncation","Reference State Initialization","Reward Shaping","Episode","DeepMimic","Massively Parallel Reinforcement Learning"]},{"id":"reference-state-initialization","category":"training","sec":10,"tier":3,"sources":[{"title":"Peng et al. 2018: DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills","url":"https://arxiv.org/abs/1804.02717"},{"title":"He et al. 2025: ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills","url":"https://arxiv.org/abs/2502.01143"},{"title":"Liao et al. 2025: BeyondMimic","url":"https://arxiv.org/abs/2508.08241"}],"as_of":"","related_ids":["deepmimic","early-termination","motion-tracking","asap","beyondmimic","exploration-vs-exploitation"],"name":"Reference State Initialization","alt":"参考状态初始化","abbr":"RSI","aliases":["RSI"],"one_liner":"In motion-imitation training, starting each episode from a random point in the reference motion instead of always frame one.","explanation":"Reference state initialization is a reinforcement-learning training trick introduced by Peng and colleagues in 2018's DeepMimic, for training a simulated character or robot to imitate a reference motion such as motion-capture data. If every episode starts from the beginning of the motion, the policy can only learn the first half before the second; for a backflip, without first learning how to land, the early takeoff phase alone tends to make the agent fall and score worse, so it never even gets to practice the landing. RSI instead picks a random point in the reference motion for each episode and sets the character's root position, orientation, velocity, and joint state directly to that frame, letting every phase of the motion get practiced in parallel. It is commonly paired with early termination, ending the episode on a fall, and has become standard practice for training humanoid motion tracking; BeyondMimic goes further and adaptively picks the starting point based on failure rate.","example":"In DeepMimic's ablation study, a backflip policy trained without RSI never learns the full flip and only manages a small backward hop; ASAP also used RSI when training a Unitree G1 to imitate highly dynamic motions such as Cristiano Ronaldo's signature jump celebration.","related":["DeepMimic","Early Termination","Motion Tracking","ASAP","BeyondMimic","Exploration vs. Exploitation"]},{"id":"data-augmentation","category":"training","sec":10,"tier":2,"sources":[{"title":"Google Machine Learning Glossary: data augmentation","url":"https://developers.google.com/machine-learning/glossary"},{"title":"Reinforcement Learning with Augmented Data (RAD, arXiv 2004.14990)","url":"https://arxiv.org/abs/2004.14990"}],"as_of":"","related_ids":["overfitting","generalization","domain-randomization","generative-data-augmentation","instruction-augmentation","visual-generalization"],"name":"Data Augmentation","alt":"数据增强","abbr":"","aliases":["Image Augmentation"],"one_liner":"Applying random transformations to existing training data to create new samples, without collecting anything new.","explanation":"Data augmentation means expanding a training set by applying random, meaning-preserving transformations to existing samples rather than collecting new ones — a standard way to reduce overfitting and improve robustness. Common image transformations include random cropping, translation, rotation, color jitter, and random occlusion. Because robot data is expensive and scarce, augmentation is especially valuable there, but the transformation must not break the correspondence between the image and the action label — for instance, flipping an image left-right means the left/right direction in the action must be flipped too. RAD (2020) and similar work showed that adding only simple augmentations like random crop and random translation noticeably improves the data efficiency of pixel-based reinforcement learning. More recent work also includes generative data augmentation, which uses generative models to swap backgrounds or objects, and instruction augmentation, which rewrites language instructions.","example":"RAD added random crop and random translation to the pixel inputs of the DeepMind Control Suite, reaching state-of-the-art data efficiency and final performance at the time, and it also clearly improved test-time generalization on ProcGen.","related":["Overfitting","Generalization","Domain Randomization","Generative Data Augmentation","Instruction Augmentation","Visual Generalization"]},{"id":"symmetry-augmentation","category":"training","sec":10,"tier":3,"sources":[{"title":"Yu, Turk, Liu 2018: Learning Symmetric and Low-energy Locomotion (SIGGRAPH 2018)","url":"https://arxiv.org/abs/1801.08093"},{"title":"Mittal et al. 2024: Symmetry Considerations for Learning Task Symmetric Robot Policies (ICRA 2024)","url":"https://arxiv.org/abs/2403.04359"},{"title":"GitHub: leggedrobotics/rsl_rl ppo.py（symmetry_cfg / use_mirror_loss）","url":"https://github.com/leggedrobotics/rsl_rl/blob/main/rsl_rl/algorithms/ppo.py"}],"as_of":"","related_ids":["data-augmentation","rl-based-locomotion-control","proximal-policy-optimization","rsl-rl","legged-locomotion","loss-function"],"name":"Symmetry Augmentation","alt":"对称性增强（镜像损失）","abbr":"","aliases":["Mirror Loss","Mirror-Symmetry Augmentation","Symmetry Loss"],"one_liner":"Using a robot's left-right symmetry, by mirroring data or adding a loss term, to make its learned motion symmetric.","explanation":"Most legged and humanoid robots are left-right symmetric, and an ideal policy should satisfy: mirror the state left-to-right, and the output should be exactly the mirror of the original action. Reinforcement learning from scratch, however, often learns an asymmetric gait, for instance dragging one leg. Two fixes are common. The mirror loss, proposed by Yu, Turk, and Liu at SIGGRAPH 2018, adds a term to the loss function penalizing the difference between the policy's output on a mirrored state and the mirror of its original output. Symmetry data augmentation instead mirrors every collected sample and trains on both copies. ETH's Mittal and colleagues compared the two at ICRA 2024 and found that data augmentation converges faster and reaches higher return within PPO. The rsl_rl library now has both options built in.","example":"Mittal and colleagues used symmetry data augmentation to train an ANYmal quadruped to climb boxes; it reached higher return than a plain PPO baseline with no symmetry handling, varied less across random seeds, and was deployed on the real robot.","related":["Data Augmentation","RL-based Locomotion Control","Proximal Policy Optimization","rsl_rl","Legged Locomotion","Loss Function"]},{"id":"domain-adaptation","category":"training","sec":10,"tier":3,"sources":[{"title":"Unsupervised Domain Adaptation by Backpropagation (arXiv:1409.7495)","url":"https://arxiv.org/abs/1409.7495"},{"title":"Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping (arXiv:1709.07857)","url":"https://arxiv.org/abs/1709.07857"}],"as_of":"","related_ids":["transfer-learning","sim-to-real-transfer","sim-to-real-gap","domain-randomization","distribution-shift","out-of-distribution"],"name":"Domain Adaptation","alt":"领域自适应","abbr":"","aliases":["Unsupervised Domain Adaptation"],"one_liner":"Making a model trained on one data distribution (the source domain) work well on a different distribution (the target domain).","explanation":"Domain adaptation is a branch of transfer learning: the source domain has plenty of labeled data, while the target domain has a different distribution and little or no labeling, and the goal is to make the model work on the target domain too. Two classic approaches exist: at the feature level, such as Ganin and Lempitsky's 2014 gradient reversal layer, which trains a classifier that tries to tell which domain the data came from and reverses its gradient, forcing the model to learn features that are indistinguishable across domains; and at the pixel level, using a GAN (generative adversarial network) to translate source-domain images into the target domain's visual style. In embodied AI, the most typical case is simulation-to-real: simulation data is cheap but differs from reality, and domain adaptation and domain randomization are the two main ways to narrow that sim-to-real gap.","example":"Google's GraspGAN (2017) used pixel-level domain adaptation to make simulated grasping images look more like real camera images; the paper reports this cut the number of real samples needed to match the same grasping performance by up to 50-fold.","related":["Transfer Learning","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Domain Randomization","Distribution Shift","Out-of-Distribution"]},{"id":"privileged-information","category":"training","sec":10,"tier":2,"sources":[{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Learning by Cheating (Chen et al., CoRL 2019)","url":"https://arxiv.org/abs/1912.12294"}],"as_of":"","related_ids":["teacher-student-distillation","asymmetric-actor-critic","sim-to-real-transfer","rapid-motor-adaptation","proprioception","knowledge-distillation"],"name":"Privileged Information","alt":"特权信息","abbr":"","aliases":["Privileged Observation","Learning Using Privileged Information","LUPI"],"one_liner":"Extra information available only during training, not at deployment, like a simulator's exact terrain shape or friction values.","explanation":"This term comes from Vapnik and Vashist's 2009 “learning using privileged information” (LUPI): extra information helps during training but is unavailable at test time. In robotics, it usually refers to quantities a simulator can read directly but a real robot's sensors cannot measure — terrain height, foot contact forces, friction coefficients, external disturbances. Training with reinforcement learning using only what a real robot can actually observe often fails to learn at all, so a common approach is to first train a teacher policy that can see the privileged information, then distill it into a student policy that uses only proprioception or camera input (teacher-student distillation); alternatively, only the critic (the network estimating value) is given privileged information, called an asymmetric actor-critic.","example":"ETH's ANYmal quadruped (Lee et al., Science Robotics 2020): the teacher policy sees terrain shape, foot contact state and contact forces, friction coefficients, and external disturbances in simulation; the student policy imitates the teacher using only a history of proprioception like joint states and IMU readings, and deploys directly to the real robot, walking across mud, snow, rubble, and dense vegetation.","related":["Teacher-Student Distillation","Asymmetric Actor-Critic","Sim-to-Real Transfer","Rapid Motor Adaptation","Proprioception","Knowledge Distillation"]},{"id":"asymmetric-actor-critic","category":"training","sec":10,"tier":3,"sources":[{"title":"Asymmetric Actor Critic for Image-Based Robot Learning (arXiv 1710.06542)","url":"https://arxiv.org/abs/1710.06542"},{"title":"Learning Dexterous In-Hand Manipulation (OpenAI Dactyl, arXiv 1808.00177)","url":"https://arxiv.org/abs/1808.00177"}],"as_of":"","related_ids":["privileged-information","teacher-student-distillation","value-function","partially-observable-markov-decision-process","sim-to-real-transfer","domain-randomization"],"name":"Asymmetric Actor-Critic","alt":"非对称演员-评论家","abbr":"","aliases":["Asymmetric AC"],"one_liner":"Letting the critic see the full simulated state while the actor sees only what a real robot could actually observe.","explanation":"Asymmetric actor-critic was introduced by Lerrel Pinto, Marcin Andrychowicz, Pieter Abbeel, and colleagues in 2017. In an actor-critic algorithm, the critic (the value network) is used only during training and discarded at deployment, so while training in simulation it can be given the full true state — exact object poses, velocities, and other privileged information — while the actor (the policy) receives only inputs a real robot can actually get, such as camera images. The critic's value estimates become more accurate, training is faster and more stable, and the policy that's left over can go straight onto a real robot. The original paper combined this with domain randomization to complete grasping and pushing tasks on a 7-DOF Fetch arm using no real-robot data at all. OpenAI's Dactyl dexterous-hand project used the same approach, and it's also common today to let the critic read privileged observations like terrain and friction when training legged locomotion with PPO.","example":"When training Dactyl, the value network could read extra information unavailable on the real robot, while the policy used only observations available on the real robot; once trained, the policy deployed directly to a Shadow dexterous hand to rotate a block.","related":["Privileged Information","Teacher-Student Distillation","Value Function","Partially Observable Markov Decision Process","Sim-to-Real Transfer","Domain Randomization"]},{"id":"teacher-student-distillation","category":"training","sec":10,"tier":2,"sources":[{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020, arXiv 2010.11251)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Learning by Cheating (Chen et al., CoRL 2019, arXiv 1912.12294)","url":"https://arxiv.org/abs/1912.12294"},{"title":"Distilling the Knowledge in a Neural Network (Hinton et al., arXiv 1503.02531)","url":"https://arxiv.org/abs/1503.02531"}],"as_of":"","related_ids":["knowledge-distillation","privileged-information","policy-distillation","asymmetric-actor-critic","sim-to-real-transfer","dagger"],"name":"Teacher-Student Distillation","alt":"教师-学生蒸馏","abbr":"","aliases":["Teacher-Student Training","Privileged Teacher Distillation"],"one_liner":"Training a teacher policy that can see privileged information first, then having a student that only uses real sensors imitate it.","explanation":"Teacher-student distillation grows out of knowledge distillation (introduced by Hinton and colleagues in 2015, which trains a small model on a large model's output as soft labels). In robotics it's usually done in two steps: first train a teacher policy in simulation, where it can read privileged information a real robot can't access directly, such as exact terrain height, friction coefficients, and object poses, which makes it much easier to train well with reinforcement learning; then train a student policy that only receives observations a real robot actually has — joint states, IMU readings, camera images — learning by imitating the teacher's actions or intermediate representations, often with DAgger-style online correction. This separates “hard-to-learn decisions” from “hard-to-learn perception,” and it's one of the mainstream routes to sim-to-real transfer. Notable examples include Learning by Cheating (2019) in autonomous driving and ETH's 2020 blind quadruped locomotion controller.","example":"ETH's Lee and colleagues (Science Robotics 2020) let the teacher read privileged terrain information in simulation while the student imitates it using only proprioceptive signals; the resulting ANYmal quadruped deploys zero-shot to mud, snow, rubble, dense vegetation, and swift-flowing water it never saw during training.","related":["Knowledge Distillation","Privileged Information","Policy Distillation","Asymmetric Actor-Critic","Sim-to-Real Transfer","DAgger"]},{"id":"real-world-reinforcement-learning","category":"training","sec":10,"tier":2,"sources":[{"title":"Challenges of Real-World Reinforcement Learning (Dulac-Arnold et al., 2019)","url":"https://arxiv.org/abs/1904.12901"},{"title":"SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning","url":"https://arxiv.org/abs/2401.16013"},{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/abs/2410.21845"}],"as_of":"2025-03","related_ids":["reinforcement-learning","sample-efficiency","human-in-the-loop","reset-free-reinforcement-learning","hil-serl","sim-to-real-transfer"],"name":"Real-World Reinforcement Learning","alt":"真机强化学习","abbr":"","aliases":["Real-world RL"],"one_liner":"Letting a robot learn by trial and error directly in the real environment, rather than only training in simulation.","explanation":"Real-world reinforcement learning means a policy's reinforcement-learning training happens, at least in part, directly on a physical robot. It skips simulator modeling and sidesteps the sim-to-real gap, which suits tasks with complex contact that's hard to simulate, like connector insertion or cable routing. The difficulties are that real-robot data is slow and expensive, so the algorithm must be highly sample-efficient; the scene has to be reset after every episode, by a person or a mechanism; rewards need to be judged automatically from camera images; and exploration can't be allowed to damage the robot. Dulac-Arnold and colleagues (2019) summarized this class of problems as nine challenges. Common countermeasures include off-policy algorithms with experience replay, warm-starting from demonstration data, training a success detector to serve as the reward, and keeping a human in the loop to correct mistakes; UC Berkeley's SERL and HIL-SERL are notable examples.","example":"SERL (ICRA 2024) learns PCB component insertion, cable routing, and object relocation on a real robot arm, training each policy in 25–50 minutes on average; its successor HIL-SERL adds human demonstration and correction, learning precision assembly, dynamic manipulation, and bimanual coordination in 1–2.5 hours with success rates near 100%.","related":["Reinforcement Learning","Sample Efficiency","Human-in-the-Loop","Reset-Free Reinforcement Learning","HIL-SERL","Sim-to-Real Transfer"]},{"id":"reset-free-reinforcement-learning","category":"training","sec":10,"tier":3,"sources":[{"title":"Eysenbach et al. 2017: Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning","url":"https://arxiv.org/abs/1711.06782"},{"title":"Gupta et al. 2021: Reset-Free Reinforcement Learning via Multi-Task Learning","url":"https://arxiv.org/abs/2104.11203"},{"title":"Sharma et al. 2021: Autonomous Reinforcement Learning: Formalism and Benchmarking","url":"https://arxiv.org/abs/2112.09605"}],"as_of":"","related_ids":["real-world-reinforcement-learning","episode","failure-recovery","autonomous-data-collection","multi-task-learning","online-reinforcement-learning"],"name":"Reset-Free Reinforcement Learning","alt":"无重置强化学习","abbr":"","aliases":["Autonomous RL","Autonomous Reinforcement Learning"],"one_liner":"Letting a robot keep learning through continuous interaction, without a person resetting the environment every episode.","explanation":"Standard reinforcement learning assumes the environment resets to its initial state after every episode, trivial in simulation, but on a real robot it usually means a person has to put the objects back, which is one of the main bottlenecks of real-robot RL. Reset-free reinforcement learning studies how to learn with as little manual resetting as possible. Eysenbach and colleagues' 2017 “Leave no Trace” learns a forward policy and a reset policy together, using the reset policy's value function to tell when the agent is about to enter an unrecoverable state. Gupta and colleagues (2021) had multiple tasks reset each other, treating it as a multi-task learning problem. That same year, Sharma, Finn, and colleagues formalized this as “autonomous reinforcement learning” and released the EARL benchmark, finding that ordinary episodic algorithms perform noticeably worse once manual intervention is reduced. It is closely related to failure recovery and autonomous data collection.","example":"Training an arm to open a drawer while also learning to close it: after opening, the “close” policy restores the environment, so the robot can practice continuously without anyone standing by to reset it.","related":["Real-World Reinforcement Learning","Episode","Failure Recovery","Autonomous Data Collection","Multi-Task Learning","Online Reinforcement Learning"]},{"id":"safe-reinforcement-learning","category":"training","sec":10,"tier":3,"sources":[{"title":"García & Fernández 2015: A Comprehensive Survey on Safe Reinforcement Learning (JMLR)","url":"https://www.jmlr.org/papers/v16/garcia15a.html"},{"title":"Achiam et al. 2017: Constrained Policy Optimization","url":"https://arxiv.org/abs/1705.10528"},{"title":"Chane-Sane et al. 2024: CaT: Constraints as Terminations for Legged Locomotion Reinforcement Learning","url":"https://arxiv.org/abs/2403.18765"}],"as_of":"","related_ids":["embodied-safety","control-barrier-function","safety-filter","real-world-reinforcement-learning","proximal-policy-optimization","reward-shaping"],"name":"Safe Reinforcement Learning","alt":"安全强化学习","abbr":"Safe RL","aliases":["Safe RL","Constrained Reinforcement Learning","Constrained RL"],"one_liner":"Maximizing reward while guaranteeing safety constraints are not violated, during both training and deployment.","explanation":"García and Fernández's 2015 survey defines safe reinforcement learning as maximizing expected return, during learning and/or deployment, in problems that require reasonable performance guarantees or must respect safety constraints. Two broad approaches are common. One changes the optimization objective, for example by formulating the problem as a constrained Markov decision process, which adds an upper limit on “cost” alongside reward, and solving it with Lagrangian multipliers or a method such as 2017's Constrained Policy Optimization (CPO). The other changes the exploration process itself, for example by injecting prior knowledge, or by using a safety filter or control barrier function to intercept dangerous actions before they execute. This matters enormously for real-robot RL and for humanoid or legged locomotion control, where falling, collisions, or exceeding a joint's limits can damage hardware or even injure people.","example":"The CaT method rewrites every constraint in a legged robot's locomotion control as “end the episode early, with some probability, the moment it's violated”; with only a small change to PPO, it learned to cross obstacles on a real Solo quadruped robot.","related":["Embodied Safety","Control Barrier Function","Safety Filter","Real-World Reinforcement Learning","Proximal Policy Optimization","Reward Shaping"]},{"id":"residual-reinforcement-learning","category":"training","sec":10,"tier":3,"sources":[{"title":"Johannink et al. 2018: Residual Reinforcement Learning for Robot Control","url":"https://arxiv.org/abs/1812.03201"},{"title":"Silver et al. 2018: Residual Policy Learning","url":"https://arxiv.org/abs/1812.06298"},{"title":"Ankile et al. 2024: From Imitation to Refinement -- Residual RL for Precise Assembly","url":"https://arxiv.org/abs/2407.16677"}],"as_of":"","related_ids":["residual-policy","behavior-cloning","real-world-reinforcement-learning","reinforcement-fine-tuning","action-chunking","noise-space-policy-steering"],"name":"Residual Reinforcement Learning","alt":"残差强化学习","abbr":"Residual RL","aliases":["Residual Policy Learning","RPL","Residual RL"],"one_liner":"Keeping an existing base controller and using reinforcement learning to learn only a correction on top of its output.","explanation":"Residual reinforcement learning was proposed in two nearly simultaneous papers at the end of 2018: UC Berkeley's Levine group working with Siemens on “Residual Reinforcement Learning for Robot Control,” and MIT's “Residual Policy Learning.” The final action equals the base policy's action plus a correction output by a residual policy; the base policy can be a hand-written controller, model-predictive control, or a policy learned through imitation, and reinforcement learning only has to fill in hard-to-model parts such as friction and contact. Because it starts from a point that already roughly works, exploration is safer and more sample-efficient, making it well suited to training directly on a real robot. In recent years it is often used to refine behavior-cloning policies: train and freeze a diffusion or action-chunking policy on demonstrations, then train a small closed-loop residual policy on top of it for real-time correction.","example":"MIT's ResiP freezes a demonstration-trained action-chunking policy and uses it as a trajectory planner, then trains a closed-loop residual policy with reinforcement learning for real-time correction, applied to precision-assembly tasks where adding more behavior-cloning data no longer raises success rate.","related":["Residual Policy","Behavior Cloning","Real-World Reinforcement Learning","Reinforcement Fine-Tuning (RL Fine-Tuning)","Action Chunking","Noise-Space Policy Steering"]},{"id":"noise-space-policy-steering","category":"training","sec":10,"tier":3,"sources":[{"title":"Wagenmaker et al. 2025: Steering Your Diffusion Policy with Latent Space Reinforcement Learning (DSRL)","url":"https://arxiv.org/abs/2506.15799"}],"as_of":"2025-06","related_ids":["diffusion-steering-via-reinforcement-learning","diffusion-policy","flow-matching","reinforcement-fine-tuning","residual-reinforcement-learning","pi0"],"name":"Noise-Space Policy Steering","alt":"噪声空间策略引导","abbr":"","aliases":["Diffusion Noise-Space RL","Latent Noise-Space RL","Diffusion Steering"],"one_liner":"Keeping a diffusion policy's weights fixed and using reinforcement learning to pick its input noise instead, changing its output.","explanation":"Diffusion and flow-matching policies generate actions by sampling a random noise vector and gradually denoising it into an action; feed the same model a different starting noise and it produces a different action. Noise-space policy steering treats that noise as a controllable “action”: it freezes the original policy and trains a small, separate reinforcement-learning policy that, given the current observation, outputs which noise to use so the original policy produces a better action. The representative method is DSRL, proposed by a UC Berkeley-led team in 2025. It only needs black-box calls to the original policy, with no backpropagation through the multi-step denoising process, and it never touches the large model's weights, so it is sample-efficient and well suited to online improvement on a real robot, including on-the-fly adaptation of general-purpose VLAs such as π0. Correspondingly, how much it can improve things is bounded by the range of actions the original policy is capable of producing in the first place.","example":"DSRL steers a public π0 checkpoint trained on DROID data to open a toaster with a Franka arm; after about 80 online episodes, success rate rises from 5/20 to 18/20.","related":["Diffusion Steering via Reinforcement Learning","Diffusion Policy","Flow Matching","Reinforcement Fine-Tuning (RL Fine-Tuning)","Residual Reinforcement Learning","π0"]},{"id":"human-in-the-loop","category":"training","sec":10,"tier":2,"sources":[{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)","url":"https://arxiv.org/abs/2410.21845"},{"title":"HG-DAgger: Interactive Imitation Learning with Human Experts","url":"https://arxiv.org/abs/1810.02890"},{"title":"Wikipedia: Human-in-the-loop","url":"https://en.wikipedia.org/wiki/Human-in-the-loop"}],"as_of":"","related_ids":["dagger","human-gated-dagger","hil-serl","human-intervention-data","intervention-rate","shared-autonomy"],"name":"Human-in-the-Loop","alt":"人在回路","abbr":"HITL","aliases":["HITL","Human Intervention"],"one_liner":"Keeping a person inside a robot's training or operating loop to correct, take over, or give feedback in real time.","explanation":"Human-in-the-loop broadly describes a system that needs a person involved while it runs: taking over control when the robot errs, providing a corrective action, scoring an outcome, or confirming a critical decision. In robot learning, its main job is covering the gap left by offline imitation learning: once a policy drifts into a state the demonstrations never covered, it tends to get further and further off track (compounding error), and human corrections at exactly those states supply the missing data. Notable methods include DAgger and 2018's HG-DAgger (a human takes over when they judge things are about to go wrong, and the segments they took over become new training data), and UC Berkeley's HIL-SERL (2024), where a person intervenes on real-robot reinforcement learning at any time using a 3D mouse, letting the robot learn precision assembly, bimanual coordination, and similar tasks in 1 to 2.5 hours. Human takeover during deployment and remote-teleoperation fallback also fall under this umbrella, and the intervention rate is a common way to measure how autonomous a system really is.","example":"During HIL-SERL training, an operator holds a SpaceMouse and watches over the robot, taking over just before it's about to fail; that intervention data goes into both the demonstration buffer and the reinforcement-learning buffer, speeding up learning.","related":["DAgger","Human-Gated DAgger","HIL-SERL","Human Intervention Data","Intervention Rate","Shared Autonomy"]},{"id":"interactive-imitation-learning","category":"training","sec":10,"tier":3,"sources":[{"title":"Interactive Imitation Learning in Robotics: A Survey (arXiv:2211.00600)","url":"https://arxiv.org/abs/2211.00600"},{"title":"A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger, arXiv:1011.0686)","url":"https://arxiv.org/abs/1011.0686"}],"as_of":"","related_ids":["dagger","human-gated-dagger","compounding-error","behavior-cloning","human-in-the-loop","human-intervention-data"],"name":"Interactive Imitation Learning","alt":"交互式模仿学习","abbr":"IIL","aliases":["IIL"],"one_liner":"Imitation learning where a human gives feedback while the robot acts, and the policy improves from it online.","explanation":"Interactive imitation learning is a branch of imitation learning in which a human gives intermittent feedback while the robot is executing — taking over to correct it, labeling the right action, or rating how well it did — and the policy improves online from that feedback. It mainly targets the compounding-error problem in behavior cloning: offline demonstrations only cover the states an expert visited, so once the robot drifts off that path there is no data to learn from, and small errors snowball. The flagship algorithm is DAgger, proposed by Stéphane Ross and colleagues in 2011: the expert labels the correct action for states the student policy actually runs into, those get folded into the dataset, and the policy is retrained; variants such as HG-DAgger let the human decide when to take over. A 2022 survey by Celemin and colleagues organizes this whole line of work.","example":"A typical pipeline: train an initial peg-insertion policy on demonstrations, let the arm run on its own, have an operator take over and record corrective actions whenever it drifts, fold those corrections into the training set, and retrain — repeating this for a few rounds.","related":["DAgger","Human-Gated DAgger","Compounding Error","Behavior Cloning","Human-in-the-Loop","Human Intervention Data"]},{"id":"human-gated-dagger","category":"training","sec":10,"tier":3,"sources":[{"title":"HG-DAgger: Interactive Imitation Learning with Human Experts (arXiv 1810.02890)","url":"https://arxiv.org/abs/1810.02890"}],"as_of":"","related_ids":["dagger","human-in-the-loop","compounding-error","human-intervention-data","recovery-and-correction-data","interactive-imitation-learning"],"name":"Human-Gated DAgger","alt":"人工门控 DAgger","abbr":"HG-DAgger","aliases":["HG-DAgger"],"one_liner":"A DAgger variant where a human takes over just before the robot is about to err, and only that takeover data gets used for retraining.","explanation":"Introduced by Stanford's Kelly and colleagues in 2018. The original DAgger (Dataset Aggregation) requires an expert to state, step by step, what action should be taken while the policy is actually driving the system — but it's hard for a person to give accurate labels, and unsafe, when they aren't the one actually in control of the car or robot. HG-DAgger instead lets a human decide when to intervene: the learned policy is normally in control, and the human takes over the moment they see it heading somewhere dangerous, steers the system back to a safe state, and then hands control back; only the observation-action pairs recorded during that takeover get added to the dataset, and the policy is retrained on the aggregated data. The paper also uses disagreement among an ensemble of networks to estimate the policy's uncertainty, and learns a risk threshold from the moments of human intervention. Experiments ran in both simulation and on a real self-driving car. This pattern of human takeover plus collecting correction data is now a common human-in-the-loop workflow in real-robot post-training.","example":"A self-driving policy is running on a test vehicle; just before it drifts over the lane line, the safety driver takes the wheel and pulls the car back to the center of the lane, and this correction data gets added to the training set.","related":["DAgger","Human-in-the-Loop","Compounding Error","Human Intervention Data","Recovery and Correction Data","Interactive Imitation Learning"]},{"id":"fleet-learning","category":"training","sec":10,"tier":3,"sources":[{"title":"Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision (arXiv:2206.14349)","url":"https://arxiv.org/abs/2206.14349"},{"title":"Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (arXiv:2605.00416)","url":"https://arxiv.org/abs/2605.00416"}],"as_of":"2026-09","related_ids":["data-flywheel","deployment-data-backflow","human-in-the-loop","dagger","offline-to-online-reinforcement-learning","real-world-reinforcement-learning"],"name":"Fleet Learning (Learning While Deploying)","alt":"机群学习 / 部署中学习","abbr":"","aliases":["Interactive Fleet Learning","IFL","LWD"],"one_liner":"Having a whole fleet of already-deployed robots collect data while they work, continually improving one shared policy together.","explanation":"Fleet learning means multiple simultaneously deployed robots pool their autonomous execution logs, failures, and human takeover data to update a shared policy, then push the new policy back out — turning deployment itself into part of training, in a data flywheel. UC Berkeley's Goldberg group proposed the “interactive fleet learning” setting with Fleet-DAgger (CoRL 2022): when a robot in the fleet is unsure, it asks a small pool of remote human operators for help, and their corrections feed imitation learning. “Learning while deploying” leans more toward reinforcement learning: in 2026, AgiBot (AGIBOT Finch) and the Shanghai Innovation Institute proposed the LWD framework, which uses autonomous rollouts and human-intervention data collected across a robot fleet for offline-to-online reinforcement learning, continually post-training a VLA.","example":"LWD was validated on a fleet of 16 bimanual robots across 8 real manipulation tasks, including semantic shelf restocking and long-horizon tasks lasting 3–5 minutes; a single generalist policy reached 95% average success as it accumulated experience across the fleet.","related":["Data Flywheel","Deployment Data Backflow","Human-in-the-Loop","DAgger","Offline-to-Online Reinforcement Learning","Real-World Reinforcement Learning"]},{"id":"self-improvement","category":"training","sec":10,"tier":3,"sources":[{"title":"Google DeepMind Blog: RoboCat: A self-improving robotic agent","url":"https://deepmind.google/discover/blog/robocat-a-self-improving-robotic-agent/"},{"title":"Bousmalis et al. 2023: RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation","url":"https://arxiv.org/abs/2306.11706"},{"title":"Ghasemipour et al. 2025: Self-Improving Embodied Foundation Models","url":"https://arxiv.org/abs/2509.15155"}],"as_of":"2025-09","related_ids":["robocat","data-flywheel","success-detector","real-world-reinforcement-learning","pi-star-0-6","rejection-sampling-fine-tuning"],"name":"Self-improvement","alt":"自我提升","abbr":"","aliases":["Self-Improving","Autonomous Improvement"],"one_liner":"A robot retrains itself on data from its own practice, getting steadily better with less reliance on human data.","explanation":"Self-improvement refers to a loop in which a model, after starting from a small amount of human data, generates new data by interacting with the environment on its own, automatically judges whether it did well, and retrains on that, aiming to stop data growth from depending entirely on human demonstrations. The key is automatically judging success or failure: common tools are success detectors, reward models, or value functions, paired with filtered behavior cloning or reinforcement learning to update the policy. Representative work includes DeepMind's 2023 RoboCat, which trains its next generation on data from its own practice, and Ghasemipour and colleagues' 2025 Self-Improving Embodied Foundation Models, which has the model predict how many steps remain, yielding both a reward and a success detector at once so a fleet of robots can practice autonomously, more sample-efficient than simply collecting more demonstrations. π*0.6, which learns from autonomous experience and human corrections using RECAP, belongs to this category too.","example":"One round of RoboCat's loop: humans teleoperate 100 to 1,000 new demonstrations for a task, a specialist branch model is fine-tuned from them, the branch practices autonomously about 10,000 times on average, and the new and old data are merged to train the next RoboCat generation.","related":["RoboCat","Data Flywheel","Success Detector","Real-World Reinforcement Learning","π*0.6","Rejection Sampling Fine-Tuning"]},{"id":"training-free","category":"training","sec":11,"tier":2,"sources":[{"title":"VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models (project page)","url":"https://voxposer.github.io/"},{"title":"ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation (arXiv 2409.01652)","url":"https://arxiv.org/abs/2409.01652"}],"as_of":"","related_ids":["zero-shot","foundation-model","voxposer","rekep","code-as-policies","visual-token-pruning"],"name":"Training-free","alt":"免训练","abbr":"","aliases":["No-training Method"],"one_liner":"Completing a new task by combining off-the-shelf models or algorithms, without updating any model parameters at all.","explanation":"Training-free means a method applied to a new task makes no gradient updates whatsoever: it calls already-trained foundation models — large language models, vision-language models, segmentation or pose-estimation models — directly, chaining them together with prompting, code generation, search, or an optimization solver. Its appeal is not needing to collect robot data at all; switching tasks only means changing the instruction, which suits data-scarce embodied settings — its limitation is that performance is capped by the off-the-shelf models' own abilities and by how the intermediate representation is designed, and it usually struggles with fine, contact-rich motion. It isn't quite the same as zero-shot: zero-shot emphasizes not having seen samples of the target task, though the underlying model may well have been specially trained; training-free emphasizes that the whole method does no training at all. Papers also use the term for plug-and-play inference speedups, such as training-free visual-token pruning.","example":"Stanford's VoxPoser has a large language model write code that calls a vision-language model, turning instructions like “hang the towel on the rack” or “close the top drawer” into a 3D value map, which a motion planner then turns into a trajectory; the paper states explicitly that the whole pipeline involves no additional training.","related":["Zero-shot","Foundation Model","VoxPoser","ReKep","Code as Policies","Visual Token Pruning"]},{"id":"inference-time-compute","category":"training","sec":11,"tier":3,"sources":[{"title":"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (arXiv:2408.03314)","url":"https://arxiv.org/abs/2408.03314"},{"title":"RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models (arXiv:2506.17811)","url":"https://arxiv.org/abs/2506.17811"}],"as_of":"2025-07","related_ids":["best-of-n-sampling","value-guided-sampling","robomonkey","chain-of-thought","inference-latency","scaling-law"],"name":"Inference-Time Compute","alt":"推理时计算","abbr":"","aliases":["Test-Time Scaling","Test-Time Compute"],"one_liner":"Improving results by spending more compute at inference time, instead of changing the trained model.","explanation":"Inference-time compute means spending more computation at inference time, after training is already finished, to improve results: generating a longer chain of thought, sampling many candidate answers and using a verifier (a model that scores candidates) to pick the best one, or repeatedly revising an answer. OpenAI's o1 (2024) and a paper by Snell and colleagues brought wide attention to this direction; the latter found that allocating inference compute according to a problem's difficulty let a small model outperform one 14 times larger on some problems. It offers a path to better performance besides simply making the model bigger. Robotics has an analogous technique: sampling a VLA's actions multiple times and using a verifier to select among them, at the cost of higher inference latency.","example":"RoboMonkey (2025) has a VLA sample multiple actions for the same observation, adds Gaussian perturbations, takes a majority vote, and then uses a VLM verifier to pick the best action; the paper reports about a 25-percentage-point absolute improvement in success rate on out-of-distribution tasks.","related":["Best-of-N Sampling","Value-Guided Sampling","RoboMonkey","Chain-of-Thought","Inference Latency","Scaling Law"]},{"id":"best-of-n-sampling","category":"training","sec":11,"tier":3,"sources":[{"title":"RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models (arXiv 2506.17811)","url":"https://arxiv.org/abs/2506.17811"},{"title":"RoboMonkey 项目页","url":"https://robomonkey-vla.github.io/"},{"title":"Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance (arXiv 2410.13816)","url":"https://arxiv.org/abs/2410.13816"}],"as_of":"2025-09","related_ids":["inference-time-compute","value-guided-sampling","robomonkey","v-gps","reward-model","vlm-as-reward"],"name":"Best-of-N Sampling","alt":"最优 N 采样","abbr":"BoN","aliases":["BoN","Test-Time Verifier","Best-of-N Reranking"],"one_liner":"Sampling N candidate outputs at once and using a scorer to pick the single highest-scoring one to actually execute.","explanation":"Best-of-N sampling is a form of test-time compute (spending extra computation at inference time in exchange for better results): a generative model samples N candidates for the same input, and a verifier — a reward model, a value function, or a VLM acting as judge — scores each one, keeping only the highest-scoring candidate. It's widely used with large language models, and its appeal is not needing to retrain the original model at all. Two notable robotics examples: V-GPS (CoRL 2024) reranks a generalist policy's candidate actions using a value function learned with offline reinforcement learning, and the same value function improves 5 different policies across 12 tasks in total; RoboMonkey (CoRL 2025, from Stanford, Berkeley, NVIDIA, and others) samples multiple actions, perturbs them with Gaussian noise and votes, then picks with a trained VLM verifier, improving out-of-distribution task performance by 25 percentage points. The cost is extra forward passes at every step, adding inference latency, and the ceiling on how much it helps depends entirely on how accurate the verifier's scores are.","example":"At every control step, RoboMonkey has a VLA model like OpenVLA generate multiple candidate actions for the same frame, has a VLM verifier score each one, and the robot executes only the highest-scoring action.","related":["Inference-Time Compute","Value-Guided Sampling","RoboMonkey","V-GPS","Reward Model","VLM-as-Reward"]},{"id":"value-guided-sampling","category":"training","sec":11,"tier":3,"sources":[{"title":"Nakamoto et al. 2024: Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance (V-GPS, CoRL 2024)","url":"https://arxiv.org/abs/2410.13816"},{"title":"Kwok et al. 2025: RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models","url":"https://arxiv.org/abs/2506.17811"}],"as_of":"2025-07","related_ids":["inference-time-compute","best-of-n-sampling","v-gps","q-function","offline-reinforcement-learning","robomonkey"],"name":"Value-Guided Sampling","alt":"价值引导采样","abbr":"","aliases":["Value-Guided Policy Reranking","Value-Guided Policy Steering"],"one_liner":"Sampling several candidate actions from a policy, then using a value function to score and pick the best one.","explanation":"Value-guided sampling improves a policy at inference time without touching its weights: at every step, sample several candidate actions from the policy, then score them with a separately trained value function, the Q-function, either taking the highest-scoring one or sampling via softmax over the scores. The representative work is Sergey Levine's group's V-GPS (CoRL 2024): it trains a language-conditioned Q-function with offline RL methods such as Cal-QL on the Bridge and RT-1 datasets, and uses it to rerank five different generalist policies, including Octo and OpenVLA, improving results across 12 tasks. Generalist policies are trained on data of mixed quality, so a value function's job is to pick out the actions that are actually good. This belongs to the same family of inference-time-compute techniques as best-of-N sampling and RoboMonkey's VLM-based action verification.","example":"In V-GPS's real-robot experiments, 50 candidate actions are sampled from the generalist policy at every step, and the Q-function picks the highest-scoring one to execute; the paper reports a relative improvement of 82.8% in average success rate across 6 tasks on a WidowX arm.","related":["Inference-Time Compute","Best-of-N Sampling","V-GPS","Q-Function","Offline Reinforcement Learning","RoboMonkey"]},{"id":"test-time-training","category":"training","sec":11,"tier":3,"sources":[{"title":"Sun et al. 2020: Test-Time Training with Self-Supervision for Generalization under Distribution Shifts (ICML 2020)","url":"https://arxiv.org/abs/1909.13231"},{"title":"Bai, Gao, Shou 2025: EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models","url":"https://arxiv.org/abs/2512.14666"},{"title":"Zhu et al. 2026: TTT-Parkour: Rapid Test-Time Training for Perceptive Robot Parkour","url":"https://arxiv.org/abs/2602.02331"}],"as_of":"2026-02","related_ids":["out-of-distribution","domain-adaptation","self-supervised-learning","inference-time-compute","continual-learning","rapid-motor-adaptation"],"name":"Test-Time Training","alt":"测试时训练","abbr":"TTT","aliases":["TTT","Test-Time Adaptation","TTA"],"one_liner":"Given new data at deployment, taking a few self-supervised update steps on it before predicting with the updated model.","explanation":"Test-time training was proposed by Sun and colleagues at ICML 2020: the model is trained with an auxiliary self-supervised task alongside its main one, for instance predicting how much an image was rotated, and at test time, given a new sample, it first takes a few gradient steps on that auxiliary task before making its real prediction, to cope with a mismatch between training and test distributions. A related idea, test-time adaptation, such as Tent (ICLR 2021), instead just minimizes prediction entropy on the test data and adjusts only the normalization layers' parameters. Embodied AI uses this to let a policy keep adapting on-site after deployment: EVOLVE-VLA uses automatically estimated task progress as feedback to keep a VLA learning at test time, and TTT-Parkour first scans and reconstructs unfamiliar terrain, then quickly fine-tunes a humanoid parkour policy on the reconstruction.","example":"TTT-Parkour (2026) uses an RGB-D camera to scan and reconstruct unfamiliar obstacles such as wedges and narrow beams, then fine-tunes a humanoid robot's parkour policy on the reconstructed terrain; the paper reports most terrain takes under 10 minutes from capture through reconstruction to test-time training.","related":["Out-of-Distribution","Domain Adaptation","Self-Supervised Learning","Inference-Time Compute","Continual Learning","Rapid Motor Adaptation"]},{"id":"in-context-learning","category":"training","sec":11,"tier":3,"sources":[{"title":"Language Models are Few-Shot Learners (GPT-3, arXiv:2005.14165)","url":"https://arxiv.org/abs/2005.14165"},{"title":"In-Context Imitation Learning via Next-Token Prediction (ICRT, arXiv:2408.15980)","url":"https://arxiv.org/abs/2408.15980"}],"as_of":"","related_ids":["few-shot","meta-learning","prompt-prompt-engineering","large-language-model","next-token-prediction","meta-reinforcement-learning"],"name":"In-Context Learning","alt":"上下文学习","abbr":"ICL","aliases":["ICL"],"one_liner":"Learning a new task from a few examples given in the prompt, without updating the model's weights.","explanation":"In-context learning means a model performs a new task at inference time using only the task description or a few examples given in the prompt, with no gradient updates (no parameter changes) at all. OpenAI's 2020 GPT-3 paper systematically demonstrated this ability: tasks like translation and question answering could be specified just with a text description plus a few examples, and the ability grew stronger as the model scaled up. Its significance is that switching tasks no longer requires retraining, only a different prompt. In embodied AI, researchers feed a policy model a few robot demonstration trajectories as a “prompt,” letting it imitate a new task on the spot; the idea is closely related to few-shot learning, meta-learning, and prompt engineering.","example":"ICRT (In-Context Robot Transformer, 2024) feeds a few human-teleoperated demonstration trajectories — images, states, and actions — into the model as a prompt at inference time, letting a Franka arm carry out a new task without updating any parameters.","related":["Few-shot","Meta-Learning","Prompt / Prompt Engineering","Large Language Model","Next-Token Prediction","Meta Reinforcement Learning"]},{"id":"meta-learning","category":"training","sec":11,"tier":3,"sources":[{"title":"Finn et al. 2017: Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (MAML)","url":"https://arxiv.org/abs/1703.03400"},{"title":"Lilian Weng: Meta-Learning: Learning to Learn Fast (2018)","url":"https://lilianweng.github.io/posts/2018-11-30-meta-learning/"},{"title":"Finn et al. 2017: One-Shot Visual Imitation Learning via Meta-Learning","url":"https://arxiv.org/abs/1709.04905"}],"as_of":"","related_ids":["few-shot","one-shot-imitation-learning","meta-reinforcement-learning","in-context-learning","transfer-learning","multi-task-learning"],"name":"Meta-Learning","alt":"元学习","abbr":"","aliases":["Learning to Learn"],"one_liner":"Training a model on many tasks so it gets good at quickly learning new tasks, not just one fixed task.","explanation":"Meta-learning, also called “learning to learn,” does not train toward a single task; instead the model repeatedly goes through “see a few examples, adapt, get tested” across many related small tasks, and what it learns is an initialization, update rule, or memory mechanism that lets it adapt quickly to a new task. Common approaches fall into three families: metric-based (such as prototypical networks), model-based (using external memory or fast weights), and optimization-based. The best known is MAML, proposed by Chelsea Finn, Pieter Abbeel, and Sergey Levine in 2017, which directly trains a set of initial parameters so the model performs well on a new task after just a few steps of gradient descent on a small amount of data. Because robot data is expensive, meta-learning was long the main approach to learning skills from very few examples; one-shot imitation learning and meta-reinforcement learning both build on it, and some researchers view large models' in-context learning as a form of implicit meta-learning.","example":"Finn and colleagues (2017) used MAML for “meta-imitation learning”: the robot is meta-trained on demonstrations from many tasks, and at test time it only needs to watch a single visual demonstration of a new task to learn it end to end.","related":["Few-shot","One-shot Imitation Learning","Meta Reinforcement Learning","In-Context Learning","Transfer Learning","Multi-Task Learning"]},{"id":"meta-reinforcement-learning","category":"training","sec":11,"tier":3,"sources":[{"title":"A Tutorial on Meta-Reinforcement Learning (arXiv:2301.08028)","url":"https://arxiv.org/abs/2301.08028"},{"title":"RL²: Fast Reinforcement Learning via Slow Reinforcement Learning (arXiv:1611.02779)","url":"https://arxiv.org/abs/1611.02779"},{"title":"Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning (arXiv:1803.11347)","url":"https://arxiv.org/abs/1803.11347"}],"as_of":"","related_ids":["meta-learning","in-context-learning","rapid-motor-adaptation","meta-world","few-shot","multi-task-learning"],"name":"Meta Reinforcement Learning","alt":"元强化学习","abbr":"Meta-RL","aliases":["Meta-RL","Learning to Reinforcement-Learn"],"one_liner":"Training across many similar tasks so an agent can adapt to a new one with only a little trial and error.","explanation":"Meta reinforcement learning treats “how to do reinforcement learning faster” itself as something to learn: train on a distribution of tasks — say, different payloads or different target positions — so the agent can adapt to a new task from that same distribution using only a handful of interactions. Two approaches are common. Methods like RL² (2016) use a recurrent network that carries memory across episodes, “learning” the new task online through its hidden state. Methods like MAML instead learn a set of initial parameters that can be fine-tuned to a new task in just a few gradient steps. This targets deep RL's poor sample efficiency and weak generalization. In robotics it is commonly used to handle changes in payload, terrain, or body damage, and Meta-World is a common manipulation benchmark built for it.","example":"Nagabandi and colleagues (2018) used meta-learning to train a dynamics model; a real legged mini-robot could use just its last few steps of observation to adapt the model online and keep moving when missing a leg, climbing a slope, or dragging a load.","related":["Meta-Learning","In-Context Learning","Rapid Motor Adaptation","Meta-World","Few-shot","Multi-Task Learning"]},{"id":"one-shot-imitation-learning","category":"training","sec":11,"tier":3,"sources":[{"title":"Duan et al. 2017: One-Shot Imitation Learning","url":"https://arxiv.org/abs/1703.07326"},{"title":"Finn et al. 2017: One-Shot Visual Imitation Learning via Meta-Learning","url":"https://arxiv.org/abs/1709.04905"}],"as_of":"","related_ids":["meta-learning","imitation-learning","few-shot","in-context-learning","demonstration-data","okami"],"name":"One-shot Imitation Learning","alt":"单样本模仿学习","abbr":"","aliases":["One-shot Imitation"],"one_liner":"The robot watches a new task demonstrated just once, then completes it starting from a different arrangement.","explanation":"One-shot imitation learning requires the robot, when faced with a new task, to succeed in a new situation — different object placement, different initial state — after seeing just a single demonstration, whether a teleoperated trajectory or a video. Rather than training from scratch on that one demonstration, the approach first trains a policy conditioned on a demonstration across many tasks: feed it a demonstration plus the current observation, and it outputs an action. OpenAI's Yan Duan and colleagues proposed and named this setting in 2017 on a block-stacking task; that same year Chelsea Finn and colleagues used MAML for meta-imitation learning, extending it to raw image input and verifying it on a real robot. Later work tried using human video directly as the demonstration, which also has to cross the embodiment gap between a human and a robot. In essence this is meta-learning applied to imitation learning, and it is conceptually close to large models' in-context learning.","example":"In Duan and colleagues' experiments, each task is stacking blocks on a table in some particular pattern, say all into one tower, or into several two-block towers; at test time, given one demonstration of a new pattern, the network has to reproduce that pattern with the blocks starting in new positions.","related":["Meta-Learning","Imitation Learning","Few-shot","In-Context Learning","Demonstration Data","OKAMI"]},{"id":"neural-network","category":"model","sec":0,"tier":1,"sources":[{"title":"Neural network (machine learning) (Wikipedia)","url":"https://en.wikipedia.org/wiki/Neural_network_(machine_learning)"}],"as_of":"","related_ids":["multilayer-perceptron","convolutional-neural-network","transformer","backpropagation","parameter-count","loss-function"],"name":"Neural Network","alt":"神经网络","abbr":"","aliases":["Deep Neural Network","Artificial Neural Network","ANN","DNN"],"one_liner":"A model made of many connected artificial neurons that learns by adjusting the strength of those connections.","explanation":"A neural network is a class of machine-learning model built from many simple computing units, artificial neurons, connected together: each unit takes a weighted sum of its inputs and passes it through a nonlinear function, the activation function, into the next layer. The strength of a connection is its weight, which is the model's parameter. Training measures the gap between the output and the correct answer with a loss function, then uses backpropagation to compute which direction to nudge each weight, repeating until the error shrinks. A network with at least two hidden layers between its input and output layers is generally called a deep neural network. AlexNet's decisive win over traditional methods at the 2012 ImageNet competition kicked off the deep-learning boom. Today's vision encoders, Transformers, VLAs, and reinforcement-learning policies in embodied AI are all, underneath, neural networks.","example":"A common reinforcement-learning walking policy for a legged robot is often just a multilayer perceptron a few layers deep: it takes in joint angles, IMU readings, and a velocity command, and outputs a target angle for every joint.","related":["Multilayer Perceptron","Convolutional Neural Network","Transformer","Backpropagation","Parameter Count (Model Size)","Loss Function"]},{"id":"parameter-count","category":"model","sec":0,"tier":1,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"RT-2: Vision-Language-Action Models (project page)","url":"https://robotics-transformer2.github.io/"},{"title":"SmolVLA: Efficient Vision-Language-Action Model trained on LeRobot Community Data (Hugging Face blog)","url":"https://huggingface.co/blog/smolvla"}],"as_of":"2025-06","related_ids":["neural-network","scaling-law","inference-latency","post-training-quantization","on-device-model","foundation-model"],"name":"Parameter Count (Model Size)","alt":"参数量","abbr":"","aliases":["Model Scale","Model Size","B (billion parameters)","M (million parameters)"],"one_liner":"The total number of trainable numbers, or weights, in a model, usually written with B for billion or M for million.","explanation":"Parameters are the numbers a neural network learns through training, mainly the weights and biases of each layer, and the parameter count is simply how many of them there are. 7B means 7 billion parameters; 300M means 300 million. Parameter count broadly determines a model's capacity, and how much compute, memory, and time training and deployment cost. Counting weights alone, memory can be roughly estimated as parameter count times bytes per parameter: OpenVLA, despite being called “7B,” actually has about 7.5 billion parameters, taking roughly 15GB loaded in bfloat16 (2 bytes per parameter), plus more at runtime. Robot models vary enormously in size: RT-2's largest version has 55 billion parameters, π0 has about 3.3 billion, and Hugging Face's SmolVLA has only 450 million. More parameters isn't always better: OpenVLA, with roughly a seventh of RT-2-X's 55 billion parameters, actually scored 16.5 percentage points higher in overall success rate across 29 tasks.","example":"Figure's Helix splits its two parts at very different scales: the 7-billion-parameter vision-language model handling understanding runs only 7 to 9 times per second, while the 80-million-parameter policy producing actions runs 200 times per second.","related":["Neural Network","Scaling Law","Inference Latency","Post-Training Quantization","On-device Model","Foundation Model"]},{"id":"multilayer-perceptron","category":"model","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Multilayer perceptron","url":"https://en.wikipedia.org/wiki/Multilayer_perceptron"},{"title":"legged_gym: legged_robot_config.py","url":"https://github.com/leggedrobotics/legged_gym/blob/master/legged_gym/envs/base/legged_robot_config.py"},{"title":"Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv:2310.03744)","url":"https://arxiv.org/abs/2310.03744"}],"as_of":"","related_ids":["neural-network","activation-function","transformer","projector-connector","action-head","backpropagation"],"name":"Multilayer Perceptron","alt":"多层感知机","abbr":"MLP","aliases":["MLP","Feedforward Neural Network"],"one_liner":"The most basic neural network: stacked fully-connected layers with nonlinear activations in between.","explanation":"A multilayer perceptron is the most basic feedforward neural network: an input layer, several hidden layers, and an output layer, with every neuron in one layer connected to every neuron in the next. Each layer applies a linear transformation followed by a nonlinear activation function (such as ReLU). It traces back to Rosenblatt's 1958 perceptron; only after the backpropagation algorithm spread in 1986 could multi-layer versions be trained effectively. A single-layer perceptron can only separate linearly separable data; adding hidden layers and nonlinearity lets it fit far more complex functions. MLPs are everywhere in today's models: the feedforward block inside each Transformer layer is itself a two-layer MLP; vision-language models such as LLaVA-1.5 use an MLP to project visual features into the language model's input space; and in robotics, reinforcement-learning locomotion policies with modest input dimensions, action heads, and state encoders are often just an MLP.","example":"The default policy network in the legged_gym reinforcement-learning framework for legged robots is an MLP with three hidden layers of width 512, 256, and 128, using ELU activations, taking in proprioception and velocity commands, and outputting target joint positions.","related":["Neural Network","Activation Function","Transformer","Projector / Connector","Action Head","Backpropagation"]},{"id":"activation-function","category":"model","sec":0,"tier":2,"sources":[{"title":"Gaussian Error Linear Units (GELUs) (arXiv:1606.08415)","url":"https://arxiv.org/abs/1606.08415"},{"title":"GLU Variants Improve Transformer (arXiv:2002.05202)","url":"https://arxiv.org/abs/2002.05202"},{"title":"LLaMA: Open and Efficient Foundation Language Models (arXiv:2302.13971)","url":"https://arxiv.org/abs/2302.13971"}],"as_of":"","related_ids":["multilayer-perceptron","transformer","neural-network","normalization-layers","vanishing-exploding-gradients","llama"],"name":"Activation Function","alt":"激活函数","abbr":"","aliases":["ReLU","GELU","SiLU","SwiGLU"],"one_liner":"The nonlinear function after each layer's linear transform, letting a network fit complex relationships.","explanation":"An activation function follows every linear transform in a neural network's layers; without it, stacking layers would still amount to just one linear transform, no matter how many layers there are. A few are common: ReLU zeroes out negative numbers and passes positive ones through unchanged, is cheap to compute, and was the default choice through the convolutional-network era. GELU, defined by Hendrycks and Gimpel in 2016 as x times Φ(x), where Φ is the standard normal distribution's cumulative distribution function, is common in Transformers such as BERT and ViT. SiLU, also called Swish, is x times sigmoid(x). SwiGLU, proposed by Shazeer in 2020, is a gated-linear-unit variant that multiplies two linear projections element-wise, with one of them passed through Swish first. Llama, for instance, replaced ReLU with SwiGLU in its Transformer feed-forward layers.","example":"For inputs −1, 0.5, and 2, ReLU outputs 0, 0.5, and 2; GELU outputs about −0.16, 0.35, and 1.95, no longer chopping negative numbers straight down to zero.","related":["Multilayer Perceptron","Transformer","Neural Network","Normalization Layers","Vanishing / Exploding Gradients","Llama"]},{"id":"softmax","category":"model","sec":0,"tier":2,"sources":[{"title":"动手学深度学习：softmax 回归","url":"https://zh.d2l.ai/chapter_linear-networks/softmax-regression.html"},{"title":"Dive into Deep Learning: Softmax Regression","url":"https://d2l.ai/chapter_linear-classification/softmax-regression.html"}],"as_of":"","related_ids":["cross-entropy","attention-mechanism","self-attention","decoding-strategies","spatial-softmax","action-binning"],"name":"Softmax","alt":"Softmax（归一化指数函数）","abbr":"","aliases":["Softmax Function","Temperature Softmax"],"one_liner":"A function that turns a set of arbitrary real numbers into a probability distribution: all positive, summing to 1.","explanation":"Softmax exponentiates each element of a real-valued vector, then divides by the sum of all the exponentials, producing a set of numbers that are all positive and sum to 1, so they can be treated as probabilities; larger input values get proportionally larger probabilities, and the exponential stretches the gaps between them. It has two main uses in deep learning. First, in a classification output layer, it turns the model's raw scores (logits) into per-class probabilities, which are then trained with cross-entropy loss. Second, inside attention mechanisms, it turns the similarity between a query and each key into attention weights. Dividing the scores by a temperature parameter before applying softmax lets you control how sharp the distribution is: a low temperature makes it close to picking only the largest value, while a high temperature makes it closer to uniform. RT-2 and OpenVLA, which discretize continuous actions into bins, also use softmax to give a probability to each bin.","example":"A classifier scores 'cup,' 'bowl,' and 'plate' as 2.0, 1.0, and 0.1; softmax turns these into probabilities of roughly 0.66, 0.24, and 0.10.","related":["Cross-Entropy","Attention Mechanism","Self-Attention","Decoding Strategies","Spatial Softmax","Action Binning"]},{"id":"normalization-layers","category":"model","sec":0,"tier":3,"sources":[{"title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift (arXiv:1502.03167)","url":"https://arxiv.org/abs/1502.03167"},{"title":"Layer Normalization (arXiv:1607.06450)","url":"https://arxiv.org/abs/1607.06450"},{"title":"Root Mean Square Layer Normalization (arXiv:1910.07467)","url":"https://arxiv.org/abs/1910.07467"}],"as_of":"","related_ids":["adaptive-layer-normalization","transformer","residual-network","convolutional-neural-network","llama","exponential-moving-average"],"name":"Normalization Layers","alt":"归一化层（层归一化 / RMSNorm / 批归一化）","abbr":"LN / RMSNorm / BN","aliases":["LayerNorm","RMSNorm","BatchNorm"],"one_liner":"Layers that rescale a network's intermediate features back into a stable numeric range, making deep networks train faster and more stably.","explanation":"A normalization layer rescales intermediate features by something like 'subtract the mean, divide by the standard deviation,' then multiplies by a learnable scale and adds a learnable shift, keeping values from spiraling out of control as depth increases and making training more stable. Batch normalization (BN, 2015) computes each channel's mean and variance across a batch, and became standard in convolutional networks, but it depends on batch size and behaves differently during training versus inference. Layer normalization (LN, 2016) instead computes statistics across all of a single sample's features, behaving the same way in training and inference, which made it the standard for Transformers. RMSNorm (2019) drops the mean-subtraction step and only divides by the root-mean-square, saving compute; large models like Llama use it, so VLAs built on them as a language backbone mostly use RMSNorm too. Diffusion Transformers also commonly use adaptive layer normalization to inject conditions like the timestep.","example":"Diffusion Policy replaces every batch normalization layer in its ResNet-18 vision encoder with group normalization (GroupNorm), because batch normalization was found to train unstably when combined with the exponential moving average (EMA) weights diffusion models commonly use.","related":["Adaptive Layer Normalization","Transformer","Residual Network","Convolutional Neural Network","Llama","Exponential Moving Average"]},{"id":"convolutional-neural-network","category":"model","sec":0,"tier":2,"sources":[{"title":"CS231n: Convolutional Neural Networks for Visual Recognition (Stanford)","url":"https://cs231n.github.io/convolutional-networks/"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"RT-1: Robotics Transformer (project page)","url":"https://robotics-transformer1.github.io/"}],"as_of":"","related_ids":["residual-network","efficientnet","vision-encoder","vision-transformer","backbone-network","spatial-softmax"],"name":"Convolutional Neural Network","alt":"卷积神经网络","abbr":"CNN","aliases":["CNN","ConvNet"],"one_liner":"A neural network that slides small filters across an image to extract local features; long the workhorse of vision.","explanation":"A convolutional neural network is designed specifically for grid-shaped data such as images. Its core is the convolutional layer: a set of small, learnable filters, such as 3×3, slide across the image, computing a local weighted sum at each position; the same filter's parameters are shared across every position, so the parameter count is far lower than a fully connected network, and it's naturally suited to recognizing edges, textures, and parts wherever they appear in the frame. Pooling layers then progressively shrink the resolution, letting the network build up from local detail to overall meaning. Historically, LeCun used LeNet to recognize handwritten digits in the 1990s, AlexNet won decisively on ImageNet in 2012, and ResNet used residual connections to go much deeper in 2015. Since Vision Transformers appeared, large models' vision encoders have mostly switched to Transformers, but CNNs remain common in smaller robot models — RT-1, for instance, uses EfficientNet, and Diffusion Policy uses ResNet-18 to process camera images.","example":"Diffusion Policy replaces a standard ResNet-18's global average pooling with a spatial softmax to preserve positional information, and replaces BatchNorm with GroupNorm to stabilize training, using it to encode each camera frame into features.","related":["Residual Network","EfficientNet","Vision Encoder","Vision Transformer","Backbone Network","Spatial Softmax"]},{"id":"residual-network","category":"model","sec":0,"tier":2,"sources":[{"title":"Deep Residual Learning for Image Recognition (arXiv 1512.03385)","url":"https://arxiv.org/abs/1512.03385"},{"title":"Dive into Deep Learning: Residual Networks (ResNet)","url":"https://d2l.ai/chapter_convolutional-modern/resnet.html"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"}],"as_of":"","related_ids":["convolutional-neural-network","backbone-network","vision-encoder","vanishing-exploding-gradients","action-chunking-with-transformers","diffusion-policy"],"name":"Residual Network","alt":"残差网络","abbr":"ResNet","aliases":["ResNet","ResNet-18","ResNet-50","Residual Connection","Skip Connection"],"one_liner":"A convolutional network with skip connections that let each layer just learn the difference between input and output, enabling very deep networks.","explanation":"Residual networks were introduced in 2015 by Kaiming He and colleagues at Microsoft Research Asia, and the paper won the CVPR 2016 Best Paper Award. Before this, adding more layers to a convolutional network actually made it harder to train and hurt accuracy. ResNet adds a skip connection inside each block that adds the input directly to the output, so the block only needs to learn the 'residual' between input and target, and gradients can also flow straight back along this shortcut. This design allowed networks as deep as 152 layers and won the ILSVRC 2015 image classification competition. Residual connections later became standard in nearly all deep networks — each Transformer layer has them too. In robot learning, small ResNets such as ResNet-18 are commonly used as image encoders; both ACT and Diffusion Policy use one to turn camera images into features.","example":"ACT uses a ResNet-18 to compress each 480×640 camera image into a 15×20×512 feature map, then flattens it into 300 tokens fed into the Transformer.","related":["Convolutional Neural Network","Backbone Network","Vision Encoder","Vanishing / Exploding Gradients","Action Chunking with Transformers","Diffusion Policy"]},{"id":"backbone-network","category":"model","sec":0,"tier":1,"sources":[{"title":"Mask R-CNN (arXiv 1703.06870)","url":"https://arxiv.org/abs/1703.06870"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["vision-encoder","vision-language-model","action-head","backbone-freezing","pre-training","residual-network"],"name":"Backbone Network","alt":"骨干网络","abbr":"","aliases":["Backbone"],"one_liner":"The main network that extracts general-purpose features from raw input, with task-specific heads attached after it.","explanation":"“Backbone network” originally comes from computer vision — 2017's Mask R-CNN paper, for instance, splits the network into a convolutional backbone that extracts features from the whole image, such as a ResNet-50, and heads that do classification, box regression, and mask prediction. A backbone is usually pretrained on large-scale data first and then reused across different tasks, and swapping in a stronger backbone tends to lift performance across the board. In embodied models, the backbone is usually an already-pretrained vision-language model: OpenVLA uses Llama 2 with DINOv2 and SigLIP vision encoders, π0 uses PaliGemma, and GR00T N1 uses NVIDIA's Eagle-2. The backbone supplies semantic knowledge, while an action head or action expert turns its features into actions; whether to freeze the backbone during fine-tuning is a common design choice.","example":"1.34 billion of GR00T N1's 2.2 billion parameters belong to the Eagle-2 vision-language backbone; it feeds the action module features from the backbone's 12th layer rather than its last layer, which the paper says makes inference faster and also raises the policy's success rate.","related":["Vision Encoder","Vision-Language Model","Action Head","Backbone Freezing","Pre-training","Residual Network"]},{"id":"efficientnet","category":"model","sec":0,"tier":3,"sources":[{"title":"EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks (arXiv:1905.11946)","url":"https://arxiv.org/abs/1905.11946"},{"title":"RT-1: Robotics Transformer project page","url":"https://robotics-transformer1.github.io/"}],"as_of":"","related_ids":["convolutional-neural-network","backbone-network","residual-network","rt-1","feature-wise-linear-modulation","tokenlearner"],"name":"EfficientNet","alt":"EfficientNet","abbr":"","aliases":["EfficientNet-B0-B7"],"one_liner":"Google's efficient convolutional-network family that scales depth, width, and resolution together by a single fixed ratio.","explanation":"EfficientNet is a family of convolutional neural networks proposed by Google's Mingxing Tan and Quoc V. Le at ICML 2019. Scaling up a CNN previously usually meant increasing just one of depth, width, or input resolution, and returns quickly leveled off. The paper introduced 'compound scaling': a single coefficient scales depth, width, and resolution together by a fixed ratio; a small baseline network, B0, is found with neural architecture search, and then scaled up step by step to get B1 through B7. The paper reports B7 reaching 84.3% top-1 accuracy on ImageNet while being 8.4x smaller and 6.1x faster at inference than the best convolutional networks at the time. Because it's small and fast, robot policies commonly used it as their vision backbone before ViT became widespread.","example":"Google's RT-1 uses an ImageNet-pretrained EfficientNet to extract image features, has the language instruction modulate them through FiLM layers, and then passes the result to TokenLearner and a Transformer to output discrete action tokens.","related":["Convolutional Neural Network","Backbone Network","Residual Network","RT-1","Feature-wise Linear Modulation","TokenLearner"]},{"id":"recurrent-neural-network","category":"model","sec":0,"tier":2,"sources":[{"title":"Dive into Deep Learning: Long Short-Term Memory (LSTM)","url":"https://d2l.ai/chapter_recurrent-modern/lstm.html"},{"title":"Dive into Deep Learning: Gated Recurrent Units (GRU)","url":"https://d2l.ai/chapter_recurrent-modern/gru.html"},{"title":"robomimic Documentation: Overview","url":"https://robomimic.github.io/docs/introduction/overview.html"}],"as_of":"","related_ids":["transformer","long-short-term-memory-gated-recurrent-unit","vanishing-exploding-gradients","state-space-model","recurrent-state-space-model","history-encoder"],"name":"Recurrent Neural Network","alt":"循环神经网络","abbr":"RNN","aliases":["RNN"],"one_liner":"A neural network that processes a sequence one step at a time, using a hidden state to remember the past.","explanation":"A recurrent neural network is a class of network for processing sequences: at each time step, it combines the new input with the hidden state from the previous step (a vector that records history) to compute a new state, reusing the same parameters at every step. Plain RNNs suffer from vanishing gradients on long sequences and struggle to remember anything from far back, which led to gated variants: Hochreiter and Schmidhuber proposed Long Short-Term Memory (LSTM) in 1997, and Cho and colleagues proposed the simpler Gated Recurrent Unit (GRU) in 2014. RNNs must be computed step by step and are hard to parallelize during training, so Transformers have largely replaced them in language tasks; but because an RNN only needs to update a single state at each inference step, which is cheap, they're still commonly used in robotics for policies that need to remember history.","example":"The imitation-learning library robomimic offers a BC-RNN algorithm: an LSTM reads in observations from the past several steps and outputs the current action, and it's often used as a comparison baseline for newer methods like Diffusion Policy.","related":["Transformer","Long Short-Term Memory / Gated Recurrent Unit","Vanishing / Exploding Gradients","State Space Model","Recurrent State-Space Model","History Encoder"]},{"id":"long-short-term-memory-gated-recurrent-unit","category":"model","sec":0,"tier":2,"sources":[{"title":"Wikipedia: Long short-term memory","url":"https://en.wikipedia.org/wiki/Long_short-term_memory"},{"title":"Wikipedia: Gated recurrent unit","url":"https://en.wikipedia.org/wiki/Gated_recurrent_unit"}],"as_of":"","related_ids":["recurrent-neural-network","transformer","state-space-model","history-encoder","vanishing-exploding-gradients","dactyl"],"name":"Long Short-Term Memory / Gated Recurrent Unit","alt":"长短期记忆网络 / 门控循环单元","abbr":"LSTM / GRU","aliases":["LSTM","GRU"],"one_liner":"Recurrent neural networks with learned 'gates' that let them retain longer histories, widely used for time-series data.","explanation":"LSTM was introduced by Hochreiter and Schmidhuber in 1997 as an improved recurrent neural network (RNN) — a network that reads a sequence one step at a time and stores information in a hidden state. In an ordinary RNN, gradients shrink exponentially over time steps during training, so the network struggles to remember anything from far in the past; LSTM adds a forget gate, an input gate, and an output gate that control what information is dropped, written, and output, which eases the problem. GRU, proposed by Cho and colleagues in 2014, has only an update gate and a reset gate and no output gate, so it has fewer parameters and runs faster, often matching LSTM's performance. Since Transformers became dominant, LSTM and GRU have taken a back seat in language tasks, but they're still common in robotics: legged locomotion policies often use them to infer terrain and body state from a history of proprioception, and some imitation-learning baselines for behavior cloning use RNNs too.","example":"OpenAI's Dactyl used policy-gradient reinforcement learning to train an LSTM policy that controlled a five-fingered dexterous hand, teaching it to reorient a block in-hand and solve a Rubik's cube.","related":["Recurrent Neural Network","Transformer","State Space Model","History Encoder","Vanishing / Exploding Gradients","Dactyl"]},{"id":"encoder-decoder","category":"model","sec":0,"tier":2,"sources":[{"title":"Dive into Deep Learning: The Encoder–Decoder Architecture","url":"https://d2l.ai/chapter_recurrent-modern/encoder-decoder.html"},{"title":"Sequence to Sequence Learning with Neural Networks (arXiv 1409.3215)","url":"https://arxiv.org/abs/1409.3215"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/html/2304.13705"}],"as_of":"","related_ids":["transformer","decoder-only-architecture","cross-attention","autoencoder","variational-autoencoder","action-chunking-with-transformers"],"name":"Encoder-Decoder","alt":"编码器-解码器","abbr":"","aliases":["Encoder-Decoder Architecture","Seq2Seq","Sequence-to-Sequence Model"],"one_liner":"A network design that compresses the input into an intermediate representation, then generates the output from that representation.","explanation":"Encoder-decoder is a general-purpose neural network design: an encoder turns the input (a sentence, an image, a set of sensor readings) into an intermediate representation, and a decoder generates the output from that representation, with input and output allowed to differ in length. Sutskever et al.'s 2014 machine-translation model, which used two LSTMs (a type of recurrent neural network) to translate English to French, is the classic example on sequence-to-sequence tasks. The original 2017 Transformer also used this design, with the decoder reading the encoder's output through cross-attention. Later, GPT-style large language models mostly switched to decoder-only designs, but encoder-decoder remains common in robotics: ACT uses a Transformer encoder to fuse multiple camera views and joint state, then has a decoder output an entire chunk of actions at once. Autoencoders and variational autoencoders also belong to this family.","example":"Controlling the bimanual ALOHA robot, ACT's encoder takes in 4 camera views and a 14-dimensional joint-angle reading; its decoder outputs the target 14-dimensional joint positions for the next k steps (typically 100).","related":["Transformer","Decoder-only Architecture","Cross-Attention","Autoencoder","Variational Autoencoder","Action Chunking with Transformers"]},{"id":"inductive-bias","category":"model","sec":0,"tier":3,"sources":[{"title":"Inductive bias（Wikipedia）","url":"https://en.wikipedia.org/wiki/Inductive_bias"},{"title":"An Image is Worth 16x16 Words (ViT, arXiv:2010.11929)","url":"https://arxiv.org/abs/2010.11929"},{"title":"Relational inductive biases, deep learning, and graph networks (arXiv:1806.01261)","url":"https://arxiv.org/abs/1806.01261"}],"as_of":"","related_ids":["convolutional-neural-network","vision-transformer","equivariant-policy-equivariant-neural-network","graph-neural-network","generalization","the-bitter-lesson"],"name":"Inductive Bias","alt":"归纳偏置","abbr":"","aliases":["Learning Bias","Structural Prior"],"one_liner":"The built-in assumptions a model or algorithm relies on to decide how it should generalize to unseen data.","explanation":"Any finite batch of training data can be explained by many different underlying rules; inductive bias is the set of assumptions a learning algorithm relies on to choose among them. Tom Mitchell pointed out in 1980 that without such assumptions, a model has no way to make predictions about inputs it hasn't seen. A common example: convolutional neural networks assume that nearby pixels are related and that features are translation-equivariant (shift an object, and its feature map shifts with it). The ViT paper notes that a Transformer has much less built-in image inductive bias than a CNN, so it scores slightly lower than a similarly sized ResNet when trained on medium-scale data like ImageNet, and needs pretraining on much larger datasets to catch up or surpass it. A strong inductive bias saves data but may cap performance; a weak one relies more on data scale. In robot learning, voxelizing observations into 3D and equivariant policies that build rotation symmetry into the network both trade a geometric inductive bias for better performance with fewer samples.","example":"PerAct voxelizes RGB-D observations before predicting actions, and the paper reports it outperforms a baseline that predicts actions directly from 2D images by 34x, which the authors attribute to the structural prior that 3D voxels provide.","related":["Convolutional Neural Network","Vision Transformer","Equivariant Policy / Equivariant Neural Network","Graph Neural Network","Generalization","The Bitter Lesson"]},{"id":"equivariant-policy-equivariant-neural-network","category":"model","sec":0,"tier":3,"sources":[{"title":"Equivariant Diffusion Policy (Wang et al., CoRL 2024, arXiv 2407.01812)","url":"https://arxiv.org/abs/2407.01812"},{"title":"Group Equivariant Convolutional Networks (Cohen & Welling, ICML 2016)","url":"https://arxiv.org/abs/1602.07576"}],"as_of":"","related_ids":["diffusion-policy","inductive-bias","data-augmentation","sample-efficiency","neural-descriptor-fields","convolutional-neural-network"],"name":"Equivariant Policy / Equivariant Neural Network","alt":"等变策略 / 等变网络","abbr":"","aliases":["Equivariant Diffusion Policy","EquiDiff","G-CNN"],"one_liner":"A network whose output rotates or shifts the same way as its input, building symmetry directly into the architecture.","explanation":"Equivariance means that when the input undergoes some transformation — rotation, translation, mirroring — the output transforms in a corresponding way. Cohen and Welling proposed the group-equivariant convolutional network (G-CNN) in 2016, building symmetry directly into the network's layers instead of relying on data augmentation to learn it slowly. Robot manipulation naturally has this kind of symmetry: if a cup on the table is rotated 90 degrees, the grasping action should rotate by the same 90 degrees. Wang and colleagues' Equivariant Diffusion Policy (EquiDiff, CoRL 2024) adds planar rotation symmetry around the vertical axis (SO(2)) into Diffusion Policy, reaching about 21.9% higher average success rate than the original across 12 MimicGen simulation tasks, and learning real-robot tasks from just 20-60 demonstrations. The cost is a more complex network implementation, and it only helps for the specific kind of symmetry it's designed for.","example":"EquiDiff reaches 80% success on a real-robot 'bagel toasting' task using 58 demonstrations, where the original Diffusion Policy only reaches 10% with the same data.","related":["Diffusion Policy","Inductive Bias","Data Augmentation","Sample Efficiency","Neural Descriptor Fields","Convolutional Neural Network"]},{"id":"graph-neural-network","category":"model","sec":0,"tier":3,"sources":[{"title":"Graph neural network - Wikipedia","url":"https://en.wikipedia.org/wiki/Graph_neural_network"},{"title":"Learning to Simulate Complex Physics with Graph Networks (Sanchez-Gonzalez et al., ICML 2020)","url":"https://arxiv.org/abs/2002.09405"},{"title":"NerveNet: Learning Structured Policy with Graph Neural Networks (project page)","url":"http://www.cs.toronto.edu/~tingwuwang/nervenet.html"}],"as_of":"","related_ids":["3d-scene-graph","neural-simulator","point-cloud-encoder","transformer","morphology-control-co-design","inductive-bias"],"name":"Graph Neural Network","alt":"图神经网络","abbr":"GNN","aliases":["GNN","Graph Network","Message-Passing Neural Network"],"one_liner":"A neural network built for graph-structured data, where each node repeatedly exchanges information with its neighbors along edges.","explanation":"A graph neural network processes data made of nodes and edges; its core operation is message passing: at each layer, every node aggregates information from its neighbors to update its own features, and stacking multiple layers lets information reach further-away nodes, with the result independent of how the nodes happen to be numbered. Scarselli and colleagues proposed the 'graph neural network model' in 2009, followed by variants such as graph convolutional networks (GCN) and graph attention networks (GAT). Embodied AI has three common uses for it: representing a robot's own body as a graph, with joints and limbs as nodes — as in NerveNet, which lets a policy transfer across agents with different body shapes; representing particles or objects as a graph to learn physics, as in DeepMind's 2020 Graph Network Simulator (GNS); and representing object relationships in a 3D scene graph for reasoning and planning.","example":"GNS treats every particle of a material like fluid or sand as a node, connects nearby particles with edges, and predicts each particle's next position through several rounds of message passing; at test time it can handle an order of magnitude more particles than it was trained with.","related":["3D Scene Graph","Neural Simulator","Point Cloud Encoder","Transformer","Morphology-Control Co-design","Inductive Bias"]},{"id":"token","category":"model","sec":1,"tier":1,"sources":[{"title":"Tokenizers (Hugging Face LLM Course, Chapter 2)","url":"https://huggingface.co/learn/llm-course/chapter2/4"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT, arXiv:2010.11929)","url":"https://arxiv.org/abs/2010.11929"}],"as_of":"","related_ids":["tokenizer","embedding","visual-token","action-tokenizer","next-token-prediction","context-length"],"name":"Token","alt":"token（词元）","abbr":"","aliases":["Marker"],"one_liner":"The basic unit a model processes data in; text, image patches, and actions can all be broken into tokens.","explanation":"A token is the basic unit that Transformer-style models read and produce. Text is first cut by a tokenizer into words or sub-word pieces; each piece has an index in a vocabulary, which is then turned into a vector, an embedding, and fed into the model — a large language model's training objective is exactly to predict the next token. In English, a token is often a word or part of a word; in Chinese, a single character may be its own token, or may get split up or merged with neighboring characters, depending on the tokenizer. The idea has since spread to other modalities: ViT cuts an image into 16×16-pixel patches, each one becoming a visual token; RT-2 and OpenVLA discretize each action dimension into 256 bins, each bin becoming a token, so a language model can output actions the same way it writes text. Context length is measured in tokens, and inference time grows with token count too.","example":"OpenVLA repurposes the 256 least-used tokens in the Llama tokenizer's vocabulary as action bins, so a 7-dimensional action — position, orientation, and gripper open/close — comes out as 7 tokens.","related":["Tokenizer","Embedding","Visual Token","Action Tokenizer","Next-Token Prediction","Context Length"]},{"id":"tokenizer","category":"model","sec":1,"tier":2,"sources":[{"title":"Hugging Face Transformers Docs: Tokenization algorithms","url":"https://huggingface.co/docs/transformers/tokenizer_summary"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"","related_ids":["token","byte-pair-encoding","action-tokenizer","video-tokenizer","visual-token","action-binning"],"name":"Tokenizer","alt":"分词器","abbr":"","aliases":[],"one_liner":"The preprocessing step that splits text, images, or actions into tokens and converts them into integer IDs.","explanation":"A tokenizer is the preprocessing module in front of a model: it splits the input into tokens (the smallest units a model works with), then converts them into integer IDs using a fixed vocabulary; it's also used in reverse, to turn the model's output IDs back into text. Large language models mostly use subword algorithms such as byte-pair encoding (BPE), which repeatedly merges the most frequent adjacent character pairs — common words stay as a single token, rare words get split into several — and GPT-2's vocabulary has 50,257 tokens. In embodied AI the term has been extended: a video tokenizer compresses video into tokens, and an action tokenizer turns continuous joint motion into discrete tokens, so a language model can output actions the same way it outputs words. How something is tokenized determines how long the sequence is and what the model can represent, making it a key design choice for a VLA.","example":"OpenVLA divides each action dimension into 256 bins and directly repurposes the 256 least-used tokens in the Llama tokenizer's vocabulary to represent them, so one step of a 7-dimensional action becomes 7 tokens.","related":["Token","Byte-Pair Encoding","Action Tokenizer","Video Tokenizer","Visual Token","Action Binning"]},{"id":"byte-pair-encoding","category":"model","sec":1,"tier":3,"sources":[{"title":"Neural Machine Translation of Rare Words with Subword Units (arXiv 1508.07909)","url":"https://arxiv.org/abs/1508.07909"},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)","url":"https://arxiv.org/html/2501.09747"},{"title":"Byte-pair encoding - Wikipedia","url":"https://en.wikipedia.org/wiki/Byte-pair_encoding"}],"as_of":"","related_ids":["tokenizer","token","action-tokenizer","discrete-cosine-transform","pi0-fast","large-language-model"],"name":"Byte-Pair Encoding","alt":"字节对编码","abbr":"BPE","aliases":["BPE","Byte-level BPE"],"one_liner":"A tokenization algorithm that builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair.","explanation":"Byte-pair encoding was first proposed by Philip Gage in 1994 as a text-compression algorithm; in 2016, Sennrich and colleagues applied it to tokenization for neural machine translation (ACL 2016), and it has since become the dominant scheme for tokenizers in large language models like GPT. The procedure starts from individual characters or bytes, counts the most frequent adjacent symbol pair in the corpus, merges it into a new symbol added to the vocabulary, and repeats until the vocabulary reaches a target size. This way, common words become a single token while rare words split into a few subwords, so the model never hits an out-of-vocabulary word and sequences don't get as long as character-by-character splitting would make them. Byte-level BPE first converts text into UTF-8 bytes before merging, so it can encode any text. In embodied AI, Physical Intelligence's FAST action tokenizer also uses BPE to compress action sequences.","example":"FAST first applies a discrete cosine transform to each action dimension, scales and rounds the result, flattens it, and trains BPE on top (the paper defaults to a scale of 10 and a vocabulary of 1,024), compressing a stretch of action into a short token sequence for a VLA to predict autoregressively.","related":["Tokenizer","Token","Action Tokenizer","Discrete Cosine Transform","π0-FAST","Large Language Model"]},{"id":"embedding","category":"model","sec":1,"tier":2,"sources":[{"title":"Embeddings | Machine Learning Crash Course (Google for Developers)","url":"https://developers.google.com/machine-learning/crash-course/embeddings"},{"title":"Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)","url":"https://arxiv.org/abs/2103.00020"}],"as_of":"","related_ids":["token","tokenizer","visual-token","latent-space","projector-connector","representation-learning"],"name":"Embedding","alt":"嵌入向量","abbr":"","aliases":["Vector Representation"],"one_liner":"Representing a word, image patch, or action as a string of real numbers, where similar things end up as nearby vectors.","explanation":"An embedding is the basic unit a neural network uses to represent information internally: a word, an image patch, a frame of robot state, or some other object gets mapped into a fixed-length string of real numbers, say 1,024 of them. It is learned during training, not hand-specified, and once trained, objects with similar meaning end up close together in vector space. Compared with one-hot encoding, where each category occupies its own dimension and only one entry is ever 1, embeddings use far fewer dimensions and can still express similarity; word2vec is an early, well-known example. Embeddings show up almost everywhere in embodied AI models: a tokenizer cuts text into tokens and looks up their word embeddings, a vision encoder turns each image patch into a visual-token embedding, and a projection layer then maps those into the language model's dimension; robot state and noisy actions also pass through a linear layer to become embeddings before joining other tokens into a Transformer.","example":"CLIP encodes a photo of a cat and the sentence “a photo of a cat” into embedding vectors separately, and the cosine similarity between them comes out clearly higher than between that same photo and “a photo of a dog.”","related":["Token","Tokenizer","Visual Token","Latent Space","Projector / Connector","Representation Learning"]},{"id":"attention-mechanism","category":"model","sec":1,"tier":1,"sources":[{"title":"Neural Machine Translation by Jointly Learning to Align and Translate (arXiv 1409.0473)","url":"https://arxiv.org/abs/1409.0473"},{"title":"Attention Is All You Need (arXiv 1706.03762)","url":"https://arxiv.org/abs/1706.03762"},{"title":"Attention (machine learning) - Wikipedia","url":"https://en.wikipedia.org/wiki/Attention_(machine_learning)"}],"as_of":"","related_ids":["transformer","self-attention","cross-attention","causal-attention","attention-mask","key-value-cache"],"name":"Attention Mechanism","alt":"注意力机制","abbr":"","aliases":["Attention"],"one_liner":"A way of weighting parts of the input by relevance and combining them, letting a model focus on what matters.","explanation":"Attention in deep learning is usually traced to Bahdanau, Cho, and Bengio's 2014 neural machine translation paper: when translating each word, the model automatically looks for the most relevant words in the source sentence, instead of compressing the whole sentence into one fixed-length vector. The now-common form has every position produce a query (Q), a key (K), and a value (V); the similarity between Q and each K is normalized with softmax into weights, and those weights are used to sum up V. The 2017 Transformer, by Vaswani and colleagues, replaced recurrence and convolution entirely with attention, which made training parallelizable and far easier to scale, and it became the shared foundation of large language models and VLAs. In embodied models, self-attention lets image patches, text, and action tokens exchange information with each other, and cross-attention is commonly used to let the action module read visual-language features.","example":"Given the instruction “put the red cup in the sink,” when the model processes the text tokens for “red cup,” attention weights concentrate on the image patches showing the red cup, tying the language to the picture.","related":["Transformer","Self-Attention","Cross-Attention","Causal Attention","Attention Mask","Key-Value Cache"]},{"id":"self-attention","category":"model","sec":1,"tier":2,"sources":[{"title":"Attention Is All You Need (arXiv 1706.03762)","url":"https://arxiv.org/abs/1706.03762"},{"title":"Dive into Deep Learning: Self-Attention and Positional Encoding","url":"https://d2l.ai/chapter_attention-mechanisms-and-transformers/self-attention-and-positional-encoding.html"}],"as_of":"","related_ids":["attention-mechanism","transformer","cross-attention","attention-mask","softmax","positional-encoding"],"name":"Self-Attention","alt":"自注意力","abbr":"","aliases":["Intra-Attention","Multi-Head Self-Attention"],"one_liner":"A mechanism where every element in a sequence gathers information from every other element, weighted by relevance.","explanation":"Self-attention is the core operation inside a Transformer, popularized by Vaswani and colleagues' 2017 paper 'Attention Is All You Need.' For each token in a sequence, learnable matrices first produce a query (Q), key (K), and value (V) vector; that token's query is dotted with every token's key, the results are turned into weights with a softmax, and those weights are used to compute a weighted sum of the value vectors, giving the token its new representation. 'Self' means the queries, keys, and values all come from the same sequence; when they come from two different sequences, it's called cross-attention instead. Self-attention can connect any two positions in a single step and is easy to parallelize, but its compute cost grows with the square of sequence length. In VLA models, image, text, and state tokens are often mixed into the same self-attention operation, with an attention mask specifying who is allowed to see whom.","example":"π0 concatenates image, language instruction, robot state, and noisy action tokens into one sequence, lets them exchange information via self-attention, and reads the action back out of the action tokens' output.","related":["Attention Mechanism","Transformer","Cross-Attention","Attention Mask","Softmax","Positional Encoding"]},{"id":"cross-attention","category":"model","sec":1,"tier":2,"sources":[{"title":"Attention Is All You Need (arXiv 1706.03762)","url":"https://arxiv.org/abs/1706.03762"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"}],"as_of":"","related_ids":["attention-mechanism","self-attention","transformer","multimodal-fusion","encoder-decoder","diffusion-transformer"],"name":"Cross-Attention","alt":"交叉注意力","abbr":"","aliases":["Encoder-Decoder Attention"],"one_liner":"Attention where one set of tokens queries information from a different set; a common way to fuse modalities.","explanation":"Cross-attention is a variant of the attention mechanism. In self-attention, the query, key, and value all come from the same sequence; in cross-attention, sequence A supplies the query while sequence B supplies the key and value, so every token in A can pull in relevant information from B, while B itself is left unchanged. It first appeared in the encoder-decoder structure of the 2017 Transformer paper, where the decoder uses it to read the encoder's representation of the source sentence for translation. It has since become the standard way to inject conditioning information into a model: text-to-image models use it to read text encodings, and robot policies use it to feed visual and language features into the action-generation part. An alternative is simply concatenating two sets of tokens and running self-attention over the combination; both approaches are common in VLAs, each with its own tradeoffs.","example":"NVIDIA's GR00T N1 action module, a DiT variant, stacks two kinds of blocks alternately: self-attention blocks process the noisy action tokens and robot state, while cross-attention blocks read the visual-language tokens output by the VLM.","related":["Attention Mechanism","Self-Attention","Transformer","Multimodal Fusion","Encoder-Decoder","Diffusion Transformer"]},{"id":"transformer","category":"model","sec":1,"tier":1,"sources":[{"title":"Attention Is All You Need (Vaswani et al., arXiv:1706.03762)","url":"https://arxiv.org/abs/1706.03762"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["attention-mechanism","self-attention","decoder-only-architecture","encoder-decoder","vision-transformer","rt-1"],"name":"Transformer","alt":"Transformer","abbr":"","aliases":["Transformer Architecture"],"one_liner":"A neural network architecture built entirely around attention, the shared backbone of large language models and VLAs.","explanation":"The Transformer was introduced by a Google team in the 2017 paper “Attention Is All You Need,” originally for machine translation. It does away with recurrence and convolution entirely, processing sequences purely through attention, which lets every token in a sequence directly weigh and pull in information from every other token by relevance. Compared with a recurrent network, which processes a sequence step by step, it can be computed in parallel, trains faster, and scales up far more easily, which made it the shared architecture behind large language models such as GPT and Llama, and vision models such as ViT. The original design is an encoder-decoder; GPT-style models use only the decoder half. Robotics started adopting it heavily around 2022, with Google's RT-1 (Robotics Transformer) a notable early example; today's VLAs, world models, and many diffusion policies are all built on a Transformer backbone.","example":"OpenVLA's backbone, Llama 2, is a decoder-only Transformer: image tokens and instruction tokens are arranged into one sequence as input, and the model outputs, one token at a time, the tokens that represent the action.","related":["Attention Mechanism","Self-Attention","Decoder-only Architecture","Encoder-Decoder","Vision Transformer","RT-1"]},{"id":"positional-encoding","category":"model","sec":1,"tier":2,"sources":[{"title":"Attention Is All You Need (arXiv:1706.03762)","url":"https://arxiv.org/html/1706.03762v7"},{"title":"RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv:2104.09864)","url":"https://arxiv.org/abs/2104.09864"},{"title":"Qwen2-VL (arXiv:2409.12191)","url":"https://arxiv.org/abs/2409.12191"}],"as_of":"","related_ids":["rotary-position-embedding","transformer","self-attention","vision-transformer","context-length","token"],"name":"Positional Encoding","alt":"位置编码","abbr":"PE","aliases":["PE","Position Embedding"],"one_liner":"Information added to each Transformer token telling the model which position in the sequence it's at.","explanation":"Self-attention inside a Transformer is order-blind on its own: shuffle the input tokens and the output just gets shuffled the same way. To give the model a sense of what came before what, the 2017 paper ‘Attention Is All You Need’ added a set of sine and cosine values at different frequencies to the input embeddings — this is positional encoding; the authors also tried learned position embeddings and found the results almost identical. Several variants have followed: Rotary Position Embedding (RoPE), proposed by Jianlin Su and colleagues in 2021, encodes position with a rotation matrix so that attention scores naturally reflect relative distance, and it has been adopted by mainstream large models such as Llama; Qwen2-VL further extends it to M-RoPE, which jointly encodes text order, an image's rows and columns, and a video's time dimension. Embodied models have to handle multiple frames, multiple camera views, and action sequences, so positional encoding determines whether the model can tell which frame came first and which patch belongs where.","example":"ViT splits a 224×224 image into 16×16-pixel patches, 196 in total, and adds a learned position embedding to each one so the model knows which patch is top-left and which is bottom-right.","related":["Rotary Position Embedding","Transformer","Self-Attention","Vision Transformer","Context Length","Token"]},{"id":"rotary-position-embedding","category":"model","sec":1,"tier":3,"sources":[{"title":"RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv 2104.09864)","url":"https://arxiv.org/abs/2104.09864"},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv 2409.12191)","url":"https://arxiv.org/html/2409.12191"},{"title":"LLaMA: Open and Efficient Foundation Language Models (arXiv 2302.13971)","url":"https://arxiv.org/abs/2302.13971"}],"as_of":"","related_ids":["positional-encoding","transformer","self-attention","context-length","qwen-vl","llama"],"name":"Rotary Position Embedding","alt":"旋转位置编码","abbr":"RoPE","aliases":["RoPE","M-RoPE"],"one_liner":"A positional encoding that writes position as a rotation angle applied to a vector, so attention naturally senses relative distance.","explanation":"Rotary Position Embedding was proposed by Jianlin Su and colleagues at Shenzhen's Zhuiyi Technology in the 2021 RoFormer paper. A Transformer has no inherent sense of token order, so it needs a positional encoding to supply one. RoPE treats every two dimensions of the query and key vectors as a 2D plane and rotates them by an angle set by the token's position, with different dimension-pairs rotating at different rates; as a result, when two tokens' vectors are dotted together, the result depends only on their relative distance. It adds no learnable parameters, attention naturally decays as distance grows, and context length can be extended by adjusting the rotation frequencies. After Meta's LLaMA adopted RoPE, it became the mainstream choice for large language models, and VLMs in the Qwen series, along with VLAs built on them, use it too. Qwen2-VL further introduces multimodal RoPE (M-RoPE), splitting position into three components — time, height, and width — for image and video tokens.","example":"Qwen2-VL's M-RoPE: a text token's three position indices are all the same, equivalent to ordinary RoPE; an image token's time index is fixed while its height and width indices vary with its position in the grid; each subsequent video frame increments the time index by one.","related":["Positional Encoding","Transformer","Self-Attention","Context Length","Qwen-VL","Llama"]},{"id":"causal-attention","category":"model","sec":1,"tier":2,"sources":[{"title":"Attention Is All You Need (arXiv:1706.03762)","url":"https://arxiv.org/abs/1706.03762"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["self-attention","attention-mask","decoder-only-architecture","autoregressive-decoding","key-value-cache","transformer"],"name":"Causal Attention","alt":"因果注意力","abbr":"","aliases":["Causal Mask","Masked Self-Attention"],"one_liner":"Attention where each position can only see itself and earlier positions, never anything that comes after.","explanation":"Causal attention adds a mask to self-attention that sets the attention score for any “future” position to negative infinity, so each token can only attend to itself and the tokens before it. It comes from the decoder in the original 2017 Transformer paper, and it is what guarantees autoregressive behavior: the whole sequence is fed in at once during training, but predicting position i still only ever depends on what comes before it. Decoder-only language models such as GPT and Llama all use it, and so do VLAs built on them when generating action tokens one at a time. Its counterpart is bidirectional attention, where every position can see every other position, used by BERT and ViT. The two are also often mixed: π0 uses a block-causal mask, splitting the input into three blocks — image/language, robot state, and noisy action — where tokens within a block can all see each other, but each block can only see itself and the blocks before it.","example":"In the sequence “pick up cup,” when the model computes the representation for “cup,” it can only attend to “pick,” “up,” and “cup” itself; anything after “cup” is blocked by the mask.","related":["Self-Attention","Attention Mask","Decoder-only Architecture","Autoregressive Decoding","Key-Value Cache","Transformer"]},{"id":"attention-mask","category":"model","sec":1,"tier":3,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["attention-mechanism","causal-attention","self-attention","key-value-cache","action-expert","pi0"],"name":"Attention Mask","alt":"注意力掩码","abbr":"","aliases":["Block-wise Causal Mask","Blockwise Causal Attention Mask"],"one_liner":"A matrix specifying which tokens each token in a sequence is allowed to 'see.'","explanation":"By default, self-attention in a Transformer lets every token attend to every other token in the sequence. An attention mask blocks out disallowed positions when computing attention scores — typically by adding negative infinity, which becomes a weight of zero after softmax — controlling how information can flow. Two kinds are most common: a causal mask, where each token can only see itself and earlier tokens, used in GPT-style models that generate one token at a time; and a padding mask, which blocks out empty positions added just to pad the sequence to a fixed length. VLA models commonly use a block-wise causal mask: the input is split into blocks, tokens within a block can all see each other, but a block can only see blocks before it. This both protects the input distribution the pretrained VLM originally saw and lets earlier blocks' KV cache be reused across multiple denoising steps, saving inference time.","example":"π0 splits its sequence into three blocks — 'image + language,' 'robot state,' and 'noisy action' — where the first block can't see the inputs added after it, to limit disruption to PaliGemma's pretrained distribution, and the state block can't see the action block, so its KV can be cached during sampling.","related":["Attention Mechanism","Causal Attention","Self-Attention","Key-Value Cache","Action Expert","π0"]},{"id":"foundation-model","category":"model","sec":1,"tier":1,"sources":[{"title":"On the Opportunities and Risks of Foundation Models (Bommasani et al., arXiv:2108.07258)","url":"https://arxiv.org/abs/2108.07258"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"2024-10","related_ids":["large-language-model","vision-language-model","embodied-foundation-model","pre-training","fine-tuning","vision-language-action-model"],"name":"Foundation Model","alt":"基础模型","abbr":"","aliases":["Foundation Models"],"one_liner":"A large model pretrained on massive data that can later adapt to many different downstream tasks.","explanation":"The term “foundation model” was coined by more than a hundred Stanford researchers in a long 2021 survey: a model trained on large-scale, diverse data, usually with self-supervised learning, which builds training targets from the data itself rather than from human labels, and that can then adapt to many downstream tasks through fine-tuning or prompting; the paper's examples are BERT, GPT-3, and DALL-E. It changed the old habit of training one model per task: spend a lot of compute once on a general-purpose base, then adjust it a little for each task, at the cost that any flaw in the base gets inherited by every downstream model. In embodied AI, VLAs are usually built on a vision-language model base, and a model pretrained on large amounts of robot data that can adapt to many robots and tasks is commonly called a robot foundation model or embodied foundation model.","example":"π0 is built on PaliGemma, Google's open-source vision-language model with 3 billion parameters, adding a 300-million-parameter action expert, and is trained on robot data into a robot foundation model that can then be fine-tuned further for specific tasks.","related":["Large Language Model","Vision-Language Model","Embodied Foundation Model","Pre-training","Fine-tuning","Vision-Language-Action Model"]},{"id":"large-language-model","category":"model","sec":1,"tier":1,"sources":[{"title":"Language Models are Few-Shot Learners (GPT-3, arXiv:2005.14165)","url":"https://arxiv.org/abs/2005.14165"},{"title":"Large language model (Wikipedia)","url":"https://en.wikipedia.org/wiki/Large_language_model"},{"title":"SayCan project page","url":"https://say-can.github.io/"}],"as_of":"","related_ids":["transformer","token","next-token-prediction","vision-language-model","llm-based-task-planning","saycan"],"name":"Large Language Model","alt":"大语言模型","abbr":"LLM","aliases":["LLM","Large Model","Language Large Model"],"one_liner":"A very large neural network trained on massive text that can understand and generate natural language.","explanation":"A large language model is a neural network trained on massive amounts of text, generally based on the Transformer architecture, with next-token prediction — predicting the next token, the smallest unit of text a model processes — as its main training objective. At sufficient scale, these models can perform new tasks just from a few examples in the prompt, without changing their weights: OpenAI's 2020 GPT-3, with 175 billion parameters, demonstrated this few-shot ability in its paper, and after ChatGPT launched in late 2022, large language models saw widespread adoption. In robotics, they are mainly used to understand human instructions, break long tasks into steps, as in SayCan, and generate control code, as in code-as-policies. Most VLA backbones are also large language models — OpenVLA, for instance, is built on Llama 2 with a vision encoder attached.","example":"Told “I spilled my Coke, can you bring me something to clean it up,” SayCan has a large language model score every skill the robot knows, combines that with an estimate of how likely each step is to succeed right now, and picks, in order, find the sponge, pick up the sponge, bring it over, done, which the robot then carries out step by step.","related":["Transformer","Token","Next-Token Prediction","Vision-Language Model","LLM-based Task Planning","SayCan"]},{"id":"autoregressive-decoding","category":"model","sec":1,"tier":1,"sources":[{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control（项目页）","url":"https://robotics-transformer2.github.io/"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["next-token-prediction","action-tokenizer","action-binning","parallel-decoding","rt-2","openvla"],"name":"Autoregressive Decoding","alt":"自回归解码","abbr":"AR","aliases":["AR","Autoregressive Model","Token-by-token Generation"],"one_liner":"Generating output one piece at a time, with each step conditioned on everything generated so far.","explanation":"Autoregressive decoding is how large language models such as GPT generate text: predict only the next token, append it to the input, predict the next one again, and repeat until done. Applying this to robots first requires discretizing continuous actions into tokens. Google DeepMind's 2023 RT-2 wrote actions as a string of number tokens and had the vision-language model output them the same way it outputs text; OpenVLA splits each action dimension into 256 bins, reusing the 256 least-used tokens in Llama's vocabulary. The upside is directly reusing a language model's architecture and training recipe; the downside is that generation has to happen one token at a time, which is slow — OpenVLA runs at only about 6Hz on an RTX 4090. This led to compressive action tokenizers such as FAST, parallel decoding, and alternatives that generate actions with diffusion or flow matching instead.","example":"One of RT-2's action outputs is a string of tokens like “1 128 91 241 5 101 127 217,” corresponding in order to whether the episode should end, end-effector translation and rotation, and gripper open/close, which then get converted back into continuous values sent to the robot.","related":["Next-Token Prediction","Action Tokenizer","Action Binning","Parallel Decoding","RT-2","OpenVLA"]},{"id":"decoding-strategies","category":"model","sec":1,"tier":3,"sources":[{"title":"How to generate text: using different decoding methods for language generation with Transformers (Hugging Face)","url":"https://huggingface.co/blog/how-to-generate"},{"title":"The Curious Case of Neural Text Degeneration (arXiv:1904.09751)","url":"https://arxiv.org/abs/1904.09751"},{"title":"openvla/openvla-7b model card","url":"https://huggingface.co/openvla/openvla-7b"}],"as_of":"","related_ids":["autoregressive-decoding","large-language-model","action-tokenizer","best-of-n-sampling","inference-time-compute","softmax"],"name":"Decoding Strategies","alt":"解码策略（贪心 / 温度采样 / Top-k / Top-p）","abbr":"","aliases":["Nucleus Sampling","Top-k Sampling","Top-p Sampling","Beam Search"],"one_liner":"The rule for picking an actual token once the model has given a probability distribution over what comes next.","explanation":"An autoregressive model only outputs a probability distribution at each step; a decoding strategy decides how to pick from it. Greedy decoding always takes the highest-probability token, giving deterministic but repetition-prone output; beam search keeps several candidate sequences at once; sampling draws randomly according to the probabilities. Temperature adjusts how sharp the distribution is — lower temperature pushes it toward greedy, higher makes it more random. Top-k only samples from the k highest-probability tokens, used by Fan and colleagues for story generation in 2018; top-p (nucleus sampling) samples from the smallest set of tokens whose cumulative probability just exceeds p, proposed by Holtzman and colleagues in 2019 to reduce the dull repetition that maximization-based decoding tends to produce. The same model can produce very different output quality depending on which strategy is used. VLAs that output action tokens usually turn off random sampling at deployment, so actions stay consistent.","example":"OpenVLA's official example calls predict_action with do_sample=False, meaning action tokens are decoded greedily, so the same image and instruction always produce the same action.","related":["Autoregressive Decoding","Large Language Model","Action Tokenizer","Best-of-N Sampling","Inference-Time Compute","Softmax"]},{"id":"decoder-only-architecture","category":"model","sec":1,"tier":3,"sources":[{"title":"Generating Wikipedia by Summarizing Long Sequences (arXiv:1801.10198)","url":"https://arxiv.org/abs/1801.10198"},{"title":"Generative pre-trained transformer - Wikipedia","url":"https://en.wikipedia.org/wiki/Generative_pre-trained_transformer"}],"as_of":"","related_ids":["transformer","encoder-decoder","causal-attention","autoregressive-decoding","large-language-model","next-token-prediction"],"name":"Decoder-only Architecture","alt":"仅解码器架构","abbr":"","aliases":["Decoder-only","Decoder-only Transformer"],"one_liner":"A Transformer design that keeps only the decoder, predicting the next token one at a time with causal attention.","explanation":"Decoder-only architecture removes the encoder from the original Transformer and stacks only decoder blocks. Google's Liu and colleagues used it to handle very long sequences in their 2018 work generating Wikipedia articles, and OpenAI's GPT-1 adopted the same design that same year; since then, the GPT series, Llama, and most other mainstream large language models have been decoder-only. It does next-token prediction using causal attention, where each position can only see the tokens before it, giving a simple, unified training objective that scales easily. Multimodal models feed it images by encoding them as visual tokens and placing them ahead of the text. Many VLAs reuse this language-model backbone directly, treating actions as just another kind of output token.","example":"OpenVLA is built on Llama 2 7B, a decoder-only language model: image features and the text instruction are concatenated into one sequence as input, and the model outputs discretized action tokens one at a time.","related":["Transformer","Encoder-Decoder","Causal Attention","Autoregressive Decoding","Large Language Model","Next-Token Prediction"]},{"id":"llama","category":"model","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Llama (language model)","url":"https://en.wikipedia.org/wiki/Llama_(language_model)"},{"title":"Meta AI Blog: The Llama 4 herd","url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"2026-09","related_ids":["large-language-model","open-weight-model","mixture-of-experts","openvla","prismatic-vlms","backbone-network"],"name":"Llama","alt":"Llama","abbr":"","aliases":["LLaMA","Llama 2","Llama 3","Llama 4"],"one_liner":"Meta's family of openly released large language models, widely used as the language backbone inside other models.","explanation":"Llama is Meta's family of large language models. The original LLaMA was released in February 2023, followed by Llama 2 (July 2023, the first version whose license allowed commercial use), Llama 3 / 3.1 (2024, up to 405B parameters), Llama 3.2 (September 2024, which added 11B/90B vision versions and small 1B/3B models), and Llama 4 (April 2025, which switched to a mixture-of-experts architecture and is natively multimodal). The weights can be downloaded, but the license carries usage restrictions, so ‘open-weight’ is a more accurate description than ‘open source’ in the strict sense. Many academic vision-language models and VLAs use Llama as their language backbone — OpenVLA, for instance, is built on Llama 2 7B. As of April 2026, Meta's own Meta AI assistant is reportedly powered instead by the closed-source Muse model family from its superintelligence lab.","example":"OpenVLA combines the Llama 2 language model with a visual encoder that fuses DINOv2 and SigLIP features, trained on about 970,000 real-robot demonstrations, to produce a 7B-parameter open VLA.","related":["Large Language Model","Open-weight Model","Mixture of Experts","OpenVLA","Prismatic VLMs","Backbone Network"]},{"id":"mixture-of-experts","category":"model","sec":1,"tier":2,"sources":[{"title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (arXiv:1701.06538)","url":"https://arxiv.org/abs/1701.06538"},{"title":"Wikipedia: Mixture of experts","url":"https://en.wikipedia.org/wiki/Mixture_of_experts"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"","related_ids":["mixture-of-transformers","action-expert","large-language-model","parameter-count","transformer","multilayer-perceptron"],"name":"Mixture of Experts","alt":"混合专家模型","abbr":"MoE","aliases":["MoE","Sparse MoE"],"one_liner":"A model made of many 'expert' subnetworks where only a few are activated per input, trading a little compute for a lot of capacity.","explanation":"The idea behind Mixture of Experts goes back to Jacobs, Jordan, Nowlan, and Hinton in 1991: several expert networks plus a gating network (also called a router) that decides which experts handle each input. In 2017, Shazeer and colleagues turned this into a sparsely-gated MoE layer for large-scale language models with hundreds of billions of parameters, where each input only runs through a small fraction of them. It solves a specific problem: you want a bigger model that holds more knowledge, but you don't want every inference to pay for the full parameter count. Large language models such as Mixtral 8x7B, DeepSeek-V3, and Llama 4 all use MoE. Worth noting: the π0 paper describes its own architecture as ‘an MoE with only two experts’ — images and text go through the VLM weights, and state and action go through the action-expert weights — a fixed division of labor by token type, not routing learned by a gating network.","example":"Mixtral 8x7B has 8 experts per layer and routes each token to 2 of them; total parameter count is about 46.7B, but each token actually only uses about 12.9B parameters worth of compute.","related":["Mixture-of-Transformers","Action Expert","Large Language Model","Parameter Count (Model Size)","Transformer","Multilayer Perceptron"]},{"id":"prompt-prompt-engineering","category":"model","sec":1,"tier":2,"sources":[{"title":"Claude Docs: Prompt engineering overview","url":"https://docs.claude.com/en/docs/build-with-claude/prompt-engineering/overview"},{"title":"openpi (Physical Intelligence) README","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"","related_ids":["large-language-model","chain-of-thought","in-context-learning","code-as-policies","saycan","prompt-tuning-soft-prompt"],"name":"Prompt / Prompt Engineering","alt":"提示词 / 提示工程","abbr":"","aliases":["Prompt","Prompting"],"one_liner":"A prompt is the text fed into a model; prompt engineering is designing it to get the output you want.","explanation":"A prompt is the text given as input to a large language model or multimodal model, which can include task instructions, background context, examples, and formatting requirements for the output. Prompt engineering is the practice of iterating on and testing prompts so the model reliably produces the desired result, without changing the model's parameters; common techniques include giving a few examples (few-shot), having the model write out its reasoning first (chain-of-thought), and specifying the output format. In embodied AI it shows up in two main ways. First, when using a large model for task planning, the prompt tells the model what skills the robot has and what's in the scene — SayCan and Code as Policies both depend on carefully designed prompts. Second, the natural-language instruction a VLA receives is also often just called the 'prompt' in the code. By contrast, a soft prompt is a trainable vector rather than text.","example":"Calling π0.5 through openpi, the input includes camera images plus a prompt field reading 'pick up the fork'; the model outputs an action chunk for picking up the fork based on it.","related":["Large Language Model","Chain-of-Thought","In-Context Learning","Code as Policies","SayCan","Prompt Tuning / Soft Prompt"]},{"id":"chain-of-thought","category":"model","sec":1,"tier":2,"sources":[{"title":"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (arXiv:2201.11903)","url":"https://arxiv.org/abs/2201.11903"},{"title":"Robotic Control via Embodied Chain-of-Thought Reasoning (arXiv:2407.08693)","url":"https://arxiv.org/abs/2407.08693"}],"as_of":"2025-01","related_ids":["embodied-chain-of-thought","visual-chain-of-thought","reasoning","embodied-reasoning","cot-vla","large-language-model"],"name":"Chain-of-Thought","alt":"思维链","abbr":"CoT","aliases":["CoT","Chain-of-Thought Reasoning","Chain-of-Thought Prompting"],"one_liner":"Having a model write out a series of intermediate reasoning steps before giving its final answer.","explanation":"Chain-of-thought means a model generates a series of intermediate reasoning steps before producing its answer. In 2022, Google's Wei and colleagues proposed “chain-of-thought prompting”: just putting a few examples with reasoning steps in the prompt gets a large model to imitate step-by-step thinking; prompting a 540-billion-parameter PaLM with 8 such examples reached the best result at the time on the math-word-problem benchmark GSM8K. Later it moved from a prompting trick to a training objective, with reasoning models such as OpenAI's o1 and DeepSeek-R1 using reinforcement learning to train longer chains of thought. Robotics has adopted the idea in VLAs: Embodied Chain-of-Thought (ECoT, 2024) has the model first write out a task plan, sub-task, action description, and object detection boxes before outputting the action, raising OpenVLA's success rate on generalization tasks by 28 percentage points absolute.","example":"Given “put the spoon on the towel,” an ECoT-style model first writes out a plan, grab the spoon, move it above the towel, release, the current sub-task, and the spoon's detection box, and only then outputs the action tokens.","related":["Embodied Chain-of-Thought","Visual Chain-of-Thought","Reasoning","Embodied Reasoning","CoT-VLA","Large Language Model"]},{"id":"context-length","category":"model","sec":1,"tier":3,"sources":[{"title":"What is a context window? (IBM Think)","url":"https://www.ibm.com/think/topics/context-window"},{"title":"PaliGemma – Google's Cutting-Edge Open Vision Language Model (Hugging Face Blog)","url":"https://huggingface.co/blog/paligemma"}],"as_of":"2024-10","related_ids":["token","key-value-cache","visual-token","embodied-memory","memory-augmented-vla","visual-token-pruning"],"name":"Context Length","alt":"上下文长度","abbr":"","aliases":["Context Window"],"one_liner":"The maximum number of tokens a model can process at once, which sets how much input it can 'see' at a time.","explanation":"Context length, also called the context window, is the upper limit on the total number of tokens a Transformer-style model can process at once — measured in tokens, not characters or words. Anything beyond that limit either gets truncated or has to be summarized before being fed in. The compute cost of self-attention grows with the square of sequence length, and the GPU memory used by the KV cache grows linearly with it, so a longer context means slower inference and more memory used. According to an IBM roundup from October 2024, GPT-4o and Llama 3.1 support 128K tokens, while Gemini 1.5 Pro supports up to 2 million. For a VLA, images eat up context especially fast: at 224×224 resolution, PaliGemma turns a single image into 256 visual tokens, and multiple cameras plus a history of several frames can make the sequence grow very quickly — one reason many VLAs only look at the current frame, or need a separate memory module or visual-token pruning instead.","example":"Encoding 3 camera views at 224×224 with PaliGemma uses 768 tokens for images alone; adding 4 frames of history per camera brings that to 12 images and 3,072 tokens.","related":["Token","Key-Value Cache","Visual Token","Embodied Memory","Memory-Augmented VLA","Visual Token Pruning"]},{"id":"state-space-model","category":"model","sec":1,"tier":3,"sources":[{"title":"Efficiently Modeling Long Sequences with Structured State Spaces (S4, arXiv:2111.00396)","url":"https://arxiv.org/abs/2111.00396"},{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)","url":"https://arxiv.org/abs/2312.00752"}],"as_of":"","related_ids":["mamba","transformer","recurrent-neural-network","robomamba","context-length","state-space"],"name":"State Space Model","alt":"状态空间模型","abbr":"SSM","aliases":["SSM","S4","Selective SSM"],"one_liner":"A network that processes long sequences with a hidden state updated over time, with compute growing linearly with sequence length.","explanation":"The state space model comes from control theory: a hidden state summarizes the past, and every new input updates that state and produces an output according to linear equations. In deep learning, an SSM discretizes those equations and uses them as a network layer, with the parameters learned through training. In 2021, Stanford's Gu, Goel, and Ré proposed S4 (ICLR 2022), which made SSMs work well on sequences tens of thousands of steps long for the first time; in 2023, Gu and Dao's Mamba made the parameters depend on the input itself (calling this 'selective'), letting the model decide what to remember or forget based on content. Compared with a Transformer, an SSM only needs to keep a fixed-size state at inference, so compute grows linearly with sequence length rather than quadratically like attention, which suits long history, high-frequency control, and on-device deployment. Don't confuse this with 'state space' in reinforcement learning, which means the set of all possible states.","example":"The Mamba paper reports about 5x the inference throughput of a similarly sized Transformer, with a 3B-parameter Mamba language model beating a Transformer of the same size and matching one twice as large; in embodied AI, RoboMamba uses a Mamba language model in place of a Transformer as the inference backbone for its VLA.","related":["Mamba","Transformer","Recurrent Neural Network","RoboMamba","Context Length","State Space"]},{"id":"mamba","category":"model","sec":1,"tier":3,"sources":[{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)","url":"https://arxiv.org/abs/2312.00752"},{"title":"Transformers are SSMs (Mamba-2, arXiv:2405.21060)","url":"https://arxiv.org/abs/2405.21060"},{"title":"RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation (arXiv:2406.04339)","url":"https://arxiv.org/abs/2406.04339"}],"as_of":"2024-05","related_ids":["state-space-model","transformer","recurrent-neural-network","robomamba","inference-latency","context-length"],"name":"Mamba","alt":"Mamba","abbr":"","aliases":["Selective State Space Model","Selective SSM","Mamba-2"],"one_liner":"A sequence model that replaces attention with an input-dependent state space model, scaling linearly with sequence length.","explanation":"Mamba was proposed by Albert Gu at Carnegie Mellon and Tri Dao at Princeton in December 2023. A Transformer's self-attention scales with the square of sequence length, which gets expensive for long sequences; a state space model (SSM) instead absorbs input step by step into a fixed-size hidden state, like a recurrent network, scaling linearly — but earlier SSMs had fixed parameters and couldn't decide what to remember or forget based on content. Mamba makes those parameters depend on the current input (making it 'selective'), paired with a GPU-oriented parallel scan algorithm, and the whole architecture uses no attention and no separate MLP block. The paper reports about 5x the inference throughput of a Transformer, with Mamba-3B matching the performance of a Transformer twice its size; in 2024 the same two authors released Mamba-2, 2-8x faster still. In robotics, work such as RoboMamba uses it to cut a VLA's inference latency.","example":"RoboMamba builds a VLA on a Mamba language-model backbone, learns manipulation skills while fine-tuning only about 0.1% of its parameters, and the paper reports 3x the inference speed of VLA models that existed at the time.","related":["State Space Model","Transformer","Recurrent Neural Network","RoboMamba","Inference Latency","Context Length"]},{"id":"vision-encoder","category":"model","sec":2,"tier":1,"sources":[{"title":"An Image is Worth 16x16 Words (ViT, arXiv:2010.11929)","url":"https://arxiv.org/abs/2010.11929"},{"title":"Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv:2103.00020)","url":"https://arxiv.org/abs/2103.00020"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"2024-06","related_ids":["vision-transformer","clip","siglip","dinov2","projector-connector","visual-token"],"name":"Vision Encoder","alt":"视觉编码器","abbr":"","aliases":["Image Encoder","Visual Backbone Network"],"one_liner":"A network that turns an image into a set of feature vectors, or visual tokens, for downstream models to use.","explanation":"A vision encoder turns pixels into features a model can use: given an image, it outputs a set of vectors, each summarizing the content of a small region. Early designs mostly used convolutional networks such as ResNet; the mainstream today is the Vision Transformer (ViT), which cuts an image into patches — 16×16 pixels in the original ViT — turns each patch into a vector, and uses attention to let the patches exchange information. An encoder's ability mostly comes from pretraining: OpenAI's 2021 CLIP trained on 400 million web image-text pairs to align images and text in the same space; SigLIP is Google's improved version built on that idea; and Meta's DINOv2 trains purely on images with self-supervision, preserving more spatial and geometric detail. VLAs usually use an off-the-shelf vision encoder as is, connecting its features to the language model through a projection layer, and may freeze it or fine-tune it jointly during training.","example":"OpenVLA feeds a 224×224 image into both SigLIP and DINOv2 at once, concatenates the two sets of features channel-wise, and passes them through a two-layer MLP projection layer to turn them into visual tokens the language model can read.","related":["Vision Transformer","CLIP","SigLIP","DINOv2","Projector / Connector","Visual Token"]},{"id":"vision-transformer","category":"model","sec":2,"tier":2,"sources":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv 2010.11929)","url":"https://arxiv.org/abs/2010.11929"}],"as_of":"","related_ids":["transformer","visual-token","self-attention","positional-encoding","convolutional-neural-network","vision-encoder"],"name":"Vision Transformer","alt":"视觉 Transformer","abbr":"ViT","aliases":["ViT"],"one_liner":"A network that cuts an image into small patches, treats them like a sequence of words, and processes them with a Transformer.","explanation":"The Vision Transformer was introduced by a Google team in the 2020 paper 'An Image is Worth 16x16 Words.' It cuts an image into fixed-size patches (commonly 16×16 pixels), flattens and linearly maps each patch into a vector, adds a positional encoding (telling the model where each patch was), and feeds the whole set as a sequence of tokens into a standard Transformer encoder; self-attention lets every patch look directly at every other patch in the image. Before this, image tasks relied mainly on convolutional neural networks; ViT showed that, given enough pretraining data, a pure Transformer could match that performance on benchmarks like ImageNet while using less training compute. Today, vision foundation models such as CLIP, SigLIP, and DINOv2 are almost all built on ViT, and VLA vision encoders mostly use it too.","example":"A 224×224 image cut into 16×16-pixel patches yields 14×14 = 196 patches — 196 visual tokens.","related":["Transformer","Visual Token","Self-Attention","Positional Encoding","Convolutional Neural Network","Vision Encoder"]},{"id":"visual-token","category":"model","sec":2,"tier":2,"sources":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv 2010.11929)","url":"https://arxiv.org/abs/2010.11929"},{"title":"PaliGemma: A versatile 3B VLM for transfer (arXiv 2407.07726)","url":"https://arxiv.org/html/2407.07726"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["vision-transformer","token","projector-connector","visual-token-pruning","spacetime-patches","vision-encoder"],"name":"Visual Token","alt":"视觉 token","abbr":"","aliases":["Image Token","Patch Token"],"one_liner":"One of the vectors produced by cutting an image into patches and encoding them — the basic unit a model uses for images.","explanation":"A visual token is the basic unit an image is broken into before entering a Transformer. The common approach traces back to ViT: the image is cut into fixed-size patches, and each patch is turned into a vector by a vision encoder — that vector is one visual token; a vision-language model then uses a projector to map these into the same space as text tokens, concatenating them with the instruction into a single sequence for joint processing. The token count grows with the square of resolution: PaliGemma produces 256 tokens for a 224-pixel image, 1,024 for 448 pixels, and 4,096 for 896 pixels. Robots typically have multiple cameras and need to run inference many times per second, so visual tokens often make up the bulk of the sequence and directly drive up inference latency — which is why compression methods like visual token pruning and resampling exist. The equivalent unit in video models is the spacetime patch.","example":"π0 is built on PaliGemma: each camera's image is first encoded into visual tokens, which are then concatenated with the language instruction's tokens into the same sequence fed into the model.","related":["Vision Transformer","Token","Projector / Connector","Visual Token Pruning","Spacetime Patches","Vision Encoder"]},{"id":"vision-foundation-model","category":"model","sec":2,"tier":2,"sources":[{"title":"On the Opportunities and Risks of Foundation Models (arXiv 2108.07258)","url":"https://arxiv.org/abs/2108.07258"},{"title":"DINOv2: Learning Robust Visual Features without Supervision (arXiv 2304.07193)","url":"https://arxiv.org/abs/2304.07193"},{"title":"Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)","url":"https://arxiv.org/abs/2103.00020"}],"as_of":"","related_ids":["foundation-model","vision-encoder","vision-transformer","clip","dinov2","segment-anything-model"],"name":"Vision Foundation Model","alt":"视觉基础模型","abbr":"VFM","aliases":["VFM","Visual Foundation Model"],"one_liner":"A general-purpose model pretrained on huge amounts of images, usable for many vision tasks with little or no adaptation.","explanation":"Foundation model refers to a model trained on large-scale data that can adapt to a wide range of downstream tasks, a term a Stanford team coined in 2021; a vision foundation model is the image-focused kind. Three approaches are common. OpenAI's CLIP uses contrastive learning on 400 million image-text pairs from the web to learn visual features aligned with language; Meta's DINOv2 uses self-supervised training on 142 million curated images, producing features especially good at spatial and geometric detail; and Meta's SAM (Segment Anything Model) can segment any object given a point or box prompt. Most of these are built on ViT. Embodied AI rarely trains its vision component from scratch — instead, it typically attaches an off-the-shelf VFM as the vision encoder and feeds its features into a language model or policy network.","example":"OpenVLA feeds the same camera image into two vision foundation models, SigLIP and DINOv2, concatenates the two sets of features, and projects them into Llama 2's input space with a two-layer MLP.","related":["Foundation Model","Vision Encoder","Vision Transformer","CLIP","DINOv2","Segment Anything Model"]},{"id":"clip","category":"model","sec":2,"tier":2,"sources":[{"title":"Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)","url":"https://arxiv.org/abs/2103.00020"},{"title":"CLIPort: What and Where Pathways for Robotic Manipulation (arXiv 2109.12098)","url":"https://arxiv.org/abs/2109.12098"}],"as_of":"","related_ids":["contrastive-learning","siglip","vision-encoder","open-vocabulary","embedding","cliport"],"name":"CLIP","alt":"CLIP","abbr":"CLIP","aliases":["Contrastive Language-Image Pre-training"],"one_liner":"An OpenAI model trained on 400 million image-text pairs that maps pictures and text into one shared vector space.","explanation":"CLIP is an image-text model OpenAI released in 2021, made of an image encoder and a text encoder trained together with contrastive learning on 400 million image-caption pairs collected from the internet: matching image-text pairs are pulled together in vector space, and mismatched ones are pushed apart. After training, images and text land in the same embedding space and their similarity can be computed directly, so it can do zero-shot classification with no further training at all — just write each category as a sentence and see which sentence the image is closest to. The paper reports this matched the zero-shot accuracy of a ResNet-50 trained on 1.28 million labeled images. In embodied AI, CLIP is often used as an open-vocabulary vision or language encoder: the early CLIPort used it to understand the semantics of objects named in an instruction, and many later VLMs and VLAs' vision encoders build on CLIP or its improved successor, SigLIP.","example":"Crop a few small patches from a tabletop photo and encode both them and the sentence “a red cup” with CLIP; the patch whose embedding is most similar to that sentence is the object the instruction is pointing to.","related":["Contrastive Learning","SigLIP","Vision Encoder","Open-vocabulary","Embedding","CLIPort"]},{"id":"siglip","category":"model","sec":2,"tier":2,"sources":[{"title":"Sigmoid Loss for Language Image Pre-Training (arXiv 2303.15343)","url":"https://arxiv.org/abs/2303.15343"},{"title":"SigLIP 2: Multilingual Vision-Language Encoders (arXiv 2502.14786)","url":"https://arxiv.org/abs/2502.14786"},{"title":"PaliGemma: A versatile 3B VLM for transfer (arXiv 2407.07726)","url":"https://arxiv.org/abs/2407.07726"}],"as_of":"2025-02","related_ids":["clip","contrastive-learning","vision-encoder","paligemma","dinov2","softmax"],"name":"SigLIP","alt":"SigLIP","abbr":"SigLIP","aliases":["Sigmoid Loss for Language-Image Pre-training","SigLIP 2","SigLIP-So400m"],"one_liner":"Google's image-text model trained with a sigmoid loss instead of softmax; its image encoder is used by many VLAs.","explanation":"SigLIP is an image-text pretraining method proposed in 2023 by Xiaohua Zhai and colleagues at Google (ICCV 2023). Like CLIP, it trains an image encoder and a text encoder so that matched image-text pairs end up close together as vectors; the difference is the loss function. CLIP uses a softmax normalized across the whole batch, while SigLIP treats each image-text pair independently as a binary 'do they match or not' classification problem with a sigmoid loss, which doesn't depend on batch-wide normalization, so it trains well even with small batches and uses less GPU memory. Its vision encoder (such as the roughly 400-million-parameter So400m) is the image input stage for many VLMs and VLAs. SigLIP 2, released in February 2025, added objectives like captioning and self-distillation during training, improving multilingual ability, localization, and dense features.","example":"π0's backbone, PaliGemma, is made of a SigLIP-So400m vision encoder and a Gemma-2B language model; OpenVLA instead concatenates SigLIP features together with DINOv2 features.","related":["CLIP","Contrastive Learning","Vision Encoder","PaliGemma","DINOv2","Softmax"]},{"id":"dinov2","category":"model","sec":2,"tier":2,"sources":[{"title":"DINOv2: Learning Robust Visual Features without Supervision (arXiv 2304.07193)","url":"https://arxiv.org/abs/2304.07193"},{"title":"facebookresearch/dinov2 (GitHub)","url":"https://github.com/facebookresearch/dinov2"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"}],"as_of":"","related_ids":["vision-foundation-model","self-supervised-learning","vision-transformer","dinov3","pre-trained-visual-representation","openvla"],"name":"DINOv2","alt":"DINOv2","abbr":"","aliases":["self-DIstillation with NO labels v2","DINO Features"],"one_liner":"Meta's 2023 self-supervised vision model that learns general-purpose image features with no labels at all.","explanation":"DINOv2 is a vision foundation model Meta AI released in April 2023, the second generation of 2021's DINO, whose name stands for “self-distillation with no labels.” It uses no human annotation at all: a student network is trained to match a teacher network's output across different crops of the same image, a form of self-supervised learning, and the team also curated a 142-million-image dataset, LVD-142M, with an automated filtering pipeline. The largest model, ViT-g/14, has about 1.1 billion parameters, and was distilled down into smaller versions from 21 million to 300 million parameters. Its features need no fine-tuning — just attach a linear layer to do classification, segmentation, or depth estimation — and they preserve relatively fine spatial and geometric detail, exactly what robots need. OpenVLA fuses DINOv2's features with SigLIP's as its vision encoder, and DINO-WM trains a world model directly on DINOv2's patch features.","example":"OpenVLA feeds the same camera image into both DINOv2 and SigLIP separately, fuses their features, and projects the result into the Llama 2 language model: DINOv2 supplies spatial and geometric detail, and SigLIP supplies semantics aligned with language.","related":["Vision Foundation Model","Self-Supervised Learning","Vision Transformer","DINOv3","Pre-trained Visual Representation","OpenVLA"]},{"id":"dinov3","category":"model","sec":2,"tier":2,"sources":[{"title":"DINOv3: Self-supervised learning for vision at unprecedented scale (Meta AI Blog)","url":"https://ai.meta.com/blog/dinov3-self-supervised-vision-model/"},{"title":"DINOv3 (arXiv 2508.10104)","url":"https://arxiv.org/abs/2508.10104"}],"as_of":"2025-08","related_ids":["dinov2","vision-foundation-model","self-supervised-learning","vision-encoder","pre-trained-visual-representation","semantic-segmentation"],"name":"DINOv3","alt":"DINOv3","abbr":"","aliases":[],"one_liner":"Meta's August 2025 third-generation DINO: 7 billion parameters, self-supervised on 1.7 billion images.","explanation":"DINOv3 is a self-supervised vision foundation model Meta released in August 2025, the successor to DINOv2. Its largest model has 7 billion parameters, trained on 1.7 billion images, roughly 7 times the model scale and 12 times the data of its predecessor. When large models train for a long time, their dense, per-patch features tend to degrade, so DINOv3 introduces Gram anchoring to keep those features stable. Meta reports that, with no fine-tuning and only a lightweight task head attached, it beats specially trained models on tasks such as detection and semantic segmentation. Besides the 7-billion-parameter flagship, several smaller ViT and ConvNeXt models were distilled from it, along with a version trained on satellite imagery. For robotics, it can serve as a stronger frozen vision encoder than DINOv2, providing high-resolution, dense features.","example":"Feed footage from a wrist-mounted camera into a frozen DINOv3, pull out each patch's feature vector, and train a lightweight head on top for object segmentation or grasp-point prediction, with no need to train a vision network from scratch.","related":["DINOv2","Vision Foundation Model","Self-Supervised Learning","Vision Encoder","Pre-trained Visual Representation","Semantic Segmentation"]},{"id":"florence-2","category":"model","sec":2,"tier":3,"sources":[{"title":"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks (arXiv 2311.06242)","url":"https://arxiv.org/abs/2311.06242"},{"title":"microsoft/Florence-2-large (Hugging Face model card)","url":"https://huggingface.co/microsoft/Florence-2-large"},{"title":"X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment VLA (arXiv 2510.10274)","url":"https://arxiv.org/html/2510.10274"}],"as_of":"2025-10","related_ids":["vision-foundation-model","vision-language-model","open-vocabulary-object-detection","visual-grounding","x-vla","backbone-network"],"name":"Florence-2","alt":"Florence-2","abbr":"","aliases":["Florence-2-base","Florence-2-large"],"one_liner":"Microsoft's small open-source vision foundation model that switches between captioning, detection, and segmentation via text prompts.","explanation":"Florence-2 is a vision foundation model Microsoft released in November 2023, open-sourced under the MIT license in two sizes: 0.23B (base) and 0.77B (large). It uses a sequence-to-sequence design — an image encoder plus a text decoder: given an image and a task prompt, such as <OD> for object detection or <CAPTION> for image captioning, the model outputs the result, including box coordinates, entirely as text, so a single model can do captioning, detection, phrase grounding, segmentation, OCR, and more. Its training data, FLD-5B, contains 126 million images and 5.4 billion annotations generated through an iterative automated pipeline. Because it's small and has strong spatial grounding ability, it's often used as the vision-language backbone for lightweight VLAs.","example":"FLOWER (CoRL 2025) uses only half of Florence-2-L's layers as its backbone; X-VLA-0.9B also uses Florence-Large to encode the main-view image and the language instruction.","related":["Vision Foundation Model","Vision-Language Model","Open-Vocabulary Object Detection","Visual Grounding","X-VLA","Backbone Network"]},{"id":"pre-trained-visual-representation","category":"model","sec":2,"tier":3,"sources":[{"title":"The Unsurprising Effectiveness of Pre-Trained Vision Models for Control (arXiv 2203.03580)","url":"https://arxiv.org/abs/2203.03580"},{"title":"Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? (VC-1, arXiv 2303.18240)","url":"https://arxiv.org/abs/2303.18240"}],"as_of":"","related_ids":["vision-encoder","r3m","vc-1","mvp","vision-foundation-model","backbone-freezing"],"name":"Pre-trained Visual Representation","alt":"预训练视觉表征","abbr":"PVR","aliases":["PVR","PVRs"],"one_liner":"A vision encoder pretrained on large-scale images or video, reused as the 'eyes' for a robot policy.","explanation":"A pre-trained visual representation is a vision encoder trained beforehand on large amounts of image or video data, which compresses a camera view into a feature vector for a robot policy or navigation agent to use, usually kept frozen or only lightly fine-tuned. It addresses the fact that robot data is scarce, making it inefficient to learn vision from scratch. In 2022, Parisi and colleagues found that representations pretrained only on generic vision data such as ImageNet let a control policy trained on top of them match or even beat one trained directly on ground-truth state, such as an object's exact position. Since then, PVRs built for embodied tasks have followed, including R3M (trained on Ego4D first-person human video), MVP (masked-autoencoder pretraining), and VC-1. The VC-1 paper compares these on CortexBench, a suite of 17 tasks, and finds no single PVR wins on all of them. VLAs today more often use vision foundation models like SigLIP or DINOv2 as the encoder directly instead.","example":"VC-1 trains a ViT with masked autoencoding on more than 4,000 hours of first-person video plus ImageNet, then freezes it and attaches a small policy network, evaluated on CortexBench's locomotion, navigation, dexterous manipulation, and mobile manipulation tasks.","related":["Vision Encoder","R3M","VC-1","MVP","Vision Foundation Model","Backbone Freezing"]},{"id":"spatial-softmax","category":"model","sec":2,"tier":3,"sources":[{"title":"End-to-End Training of Deep Visuomotor Policies (arXiv:1504.00702, JMLR 2016)","url":"https://arxiv.org/abs/1504.00702"},{"title":"Deep Spatial Autoencoders for Visuomotor Learning (arXiv:1509.06113)","url":"https://arxiv.org/abs/1509.06113"},{"title":"Diffusion Policy (arXiv:2303.04137, HTML)","url":"https://arxiv.org/html/2303.04137"}],"as_of":"","related_ids":["convolutional-neural-network","vision-encoder","visuomotor-policy","diffusion-policy","keypoint-detection","end-to-end-training-of-deep-visuomotor-policies"],"name":"Spatial Softmax","alt":"空间 Softmax","abbr":"","aliases":["Spatial Soft-Argmax","Soft-Argmax","Keypoint Pooling"],"one_liner":"Pooling each channel of a convolutional feature map down to the 2D coordinate of its 'brightest point,' preserving location.","explanation":"Spatial Softmax is a pooling layer that turns a convolutional feature map into coordinates, proposed by Berkeley's Levine, Finn, Darrell, and Abbeel in their 2015 work on end-to-end visuomotor policies. The procedure applies softmax across all pixel positions within each channel, producing a probability distribution for 'where this feature appears,' then takes the expectation over pixel coordinates to get one (x, y) pair — effectively a differentiable argmax. Ordinary classification networks use global average pooling, which erases location information, but robot manipulation specifically needs to know where things are; softmax also suppresses weak false activations, making it more robust to distractor objects. It later became a common choice in imitation-learning vision encoders — Diffusion Policy uses it at the end of its ResNet-18 in place of global average pooling.","example":"Levine and colleagues' policy network follows three convolutional layers with a spatial softmax; the last layer's 32 channels each output one feature-point coordinate, which is concatenated with the robot's joint state and passed through fully connected layers to output motor torques, with the whole network having only about 92,000 parameters.","related":["Convolutional Neural Network","Vision Encoder","Visuomotor Policy","Diffusion Policy","Keypoint Detection","End-to-End Training of Deep Visuomotor Policies"]},{"id":"text-encoder","category":"model","sec":2,"tier":2,"sources":[{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/abs/2405.12213"},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)","url":"https://arxiv.org/abs/2410.07864"}],"as_of":"","related_ids":["language-conditioned-policy","clip","siglip","large-language-model","embedding","feature-wise-linear-modulation"],"name":"Text Encoder","alt":"文本编码器","abbr":"","aliases":["Language Encoder","Instruction Encoder"],"one_liner":"A network that turns a text instruction into a sequence of vectors for the rest of the model to use.","explanation":"A text encoder splits a piece of text (such as 'put the cup in the sink') into tokens and turns them into vectors, used as a conditioning input for a policy or generative model. It's usually a pretrained language model, and it's often kept frozen when training a robot policy: Octo uses a roughly 110-million-parameter T5-base, RDT-1B uses a frozen T5-XXL, and CLIP and SigLIP each come with their own text encoder aligned to their image encoder; text-to-image and text-to-video models rely on one too, to understand the prompt. It brings pretrained language knowledge into the policy, helping the model tell different instructions apart. VLA models built on a large language model backbone, such as OpenVLA and π0, don't have a separate text encoder at all — text goes directly into the language model itself as tokens.","example":"Octo encodes instructions into 16 language tokens using a frozen t5-base, fed into the Transformer alongside image tokens; the paper tried swapping in a larger T5 or fine-tuning it, and neither improved performance.","related":["Language-conditioned Policy","CLIP","SigLIP","Large Language Model","Embedding","Feature-wise Linear Modulation"]},{"id":"multimodal-fusion","category":"model","sec":2,"tier":2,"sources":[{"title":"Multimodal Machine Learning: A Survey and Taxonomy (arXiv:1705.09406)","url":"https://arxiv.org/abs/1705.09406"},{"title":"Wikipedia: Multimodal learning","url":"https://en.wikipedia.org/wiki/Multimodal_learning"},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (arXiv:2212.06817)","url":"https://arxiv.org/html/2212.06817"}],"as_of":"","related_ids":["multimodal-large-language-model","cross-attention","feature-wise-linear-modulation","projector-connector","visuo-tactile-fusion","multi-sensor-fusion"],"name":"Multimodal Fusion","alt":"多模态融合","abbr":"","aliases":["Early / Late Fusion"],"one_liner":"Combining information from different sources — image, language, robot state — into a form a model can use jointly.","explanation":"Multimodal fusion means combining information from different modalities — image, text, depth, touch, joint state, and so on — into a representation the model can use jointly. Baltrušaitis and colleagues' 2017 survey lists it as one of five core challenges in multimodal machine learning, alongside representation, translation, alignment, and co-learning. Fusion approaches are usually grouped by where the combination happens: early fusion concatenates features or tokens from each modality right at the start and processes them together; mid fusion encodes each modality separately first, then merges them partway through the network, often via cross-attention; and late fusion lets each modality produce its own result and combines them only at the final decision. Embodied AI is inherently multimodal — a VLA has to look at camera images, read a language instruction, and sense its own state at the same time — and the fusion approach directly determines whether the model can map an instruction like ‘pick up the red cup’ onto the right object and action.","example":"RT-1 uses FiLM layers to inject the language instruction's embedding into its EfficientNet-B3 image encoder, so visual features carry task information from early in the network; π0, by contrast, puts image, language, state, and action tokens into the same Transformer and lets attention handle the fusion.","related":["Multimodal Large Language Model","Cross-Attention","Feature-wise Linear Modulation","Projector / Connector","Visuo-Tactile Fusion","Multi-Sensor Fusion"]},{"id":"vision-language-model","category":"model","sec":2,"tier":1,"sources":[{"title":"Vision Language Models Explained (Hugging Face blog)","url":"https://huggingface.co/blog/vlms"},{"title":"PaliGemma: A versatile 3B VLM for transfer (arXiv:2407.07726)","url":"https://arxiv.org/abs/2407.07726"},{"title":"RT-2: Vision-Language-Action Models (project page)","url":"https://robotics-transformer2.github.io/"}],"as_of":"2024-07","related_ids":["vision-language-action-model","vision-encoder","projector-connector","large-language-model","paligemma","multimodal-large-language-model"],"name":"Vision-Language Model","alt":"视觉语言模型","abbr":"VLM","aliases":["VLM"],"one_liner":"A model that takes in both images and text and answers in text; the base that most VLAs build on.","explanation":"A vision-language model takes in an image and text and outputs text: it can describe a picture, answer questions about it, or point out where an object is. Three parts are common: a vision encoder turns the image into features, a projection layer aligns those features to the language model's input space, and the language model handles understanding and generating the text. 2023's LLaVA, for instance, is a CLIP encoder plus a projection layer plus the Vicuna language model; Google's 2024 open-source PaliGemma combines a SigLIP encoder with Gemma-2B, at about 3 billion parameters. The objects, common sense, and spatial knowledge a VLM picks up from internet image-text data are exactly what robots lack, which is why most VLAs start from a VLM: RT-2 adds robot data on top of PaLI-X and PaLM-E through joint fine-tuning, and π0 is built on PaliGemma. VLMs are also commonly used on their own as the high-level planner in a hierarchical architecture.","example":"Ask a VLM about a kitchen photo, “what's to the left of the sink,” and it answers in text; swap the output for action tokens and train it further on robot data, and you get a VLA like RT-2.","related":["Vision-Language-Action Model","Vision Encoder","Projector / Connector","Large Language Model","PaliGemma","Multimodal Large Language Model"]},{"id":"multimodal-large-language-model","category":"model","sec":2,"tier":2,"sources":[{"title":"A Survey on Multimodal Large Language Models (arXiv:2306.13549)","url":"https://arxiv.org/html/2306.13549v4"},{"title":"Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv:2310.03744)","url":"https://arxiv.org/abs/2310.03744"}],"as_of":"","related_ids":["vision-language-model","large-language-model","native-multimodal","projector-connector","vision-language-action-model","modality-alignment"],"name":"Multimodal Large Language Model","alt":"多模态大语言模型","abbr":"MLLM","aliases":["MLLM","Multimodal LLM"],"one_liner":"A large language model that can also understand non-text inputs like images, video, and audio.","explanation":"A multimodal large language model uses an LLM as its ‘brain’ but can also take in non-text inputs such as images, video, and audio; GPT-4V brought wide attention to this direction in 2023. Yin and colleagues' survey breaks the typical architecture into three parts: a pretrained modality encoder (such as a CLIP or SigLIP vision encoder), a pretrained LLM, and a modality interface connecting the two — an MLP projector, a Q-Former, or cross-attention layers inserted into the LLM. Training usually proceeds in two stages: modality-alignment pretraining, then instruction tuning. An MLLM that only handles images and text is often called a vision-language model (VLM) instead. Its significance for embodied AI is that it brings internet-scale common sense and visual understanding: most VLAs are built by adding an action output on top of an MLLM or VLM, and embodied-reasoning models and task planners are also often fine-tuned from an MLLM.","example":"LLaVA-1.5 uses a CLIP-ViT-L-336px vision encoder connected to the language model through an MLP projector, then is fine-tuned on visual instruction data; GPT-4o, Gemini, and Qwen2.5-VL are also multimodal large language models.","related":["Vision-Language Model","Large Language Model","Native Multimodal","Projector / Connector","Vision-Language-Action Model","Modality Alignment (Alignment Pretraining Stage)"]},{"id":"projector-connector","category":"model","sec":2,"tier":2,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"},{"title":"Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv 2310.03744)","url":"https://arxiv.org/abs/2310.03744"}],"as_of":"","related_ids":["vision-encoder","multimodal-fusion","querying-transformer","perceiver-resampler","modality-alignment","llava"],"name":"Projector / Connector","alt":"投影层","abbr":"","aliases":["Projector","Connector","Vision-Language Connector"],"one_liner":"A small network that maps features from a vision encoder into the input space of a language model.","explanation":"A projector is a small network in a multimodal model that connects two separately pretrained components — most commonly, it maps the image features a vision encoder produces into vectors a large language model can read directly. It's needed because the two sides were pretrained independently, so their feature dimensions and distributions don't match. The original LLaVA used just a single linear layer; LLaVA-1.5 switched to a two-layer MLP (multilayer perceptron) and got better results. More elaborate designs, such as a Q-Former or a Perceiver Resampler, also compress the number of tokens along the way. When training a VLM, it's common to freeze both pretrained sides and train only the projector, which is the modality-alignment stage. The same idea shows up in VLAs for robot state and action too: a linear layer or MLP projects low-dimensional vectors like joint angles into the model's embedding dimension.","example":"OpenVLA concatenates SigLIP and DINOv2 image features channel-wise, then passes them through a two-layer MLP projector before feeding them into the Llama 2 7B language model.","related":["Vision Encoder","Multimodal Fusion","Querying Transformer","Perceiver Resampler","Modality Alignment (Alignment Pretraining Stage)","LLaVA"]},{"id":"learnable-query","category":"model","sec":2,"tier":3,"sources":[{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs (arXiv:2301.12597)","url":"https://arxiv.org/abs/2301.12597"},{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv:2405.12213)","url":"https://arxiv.org/html/2405.12213"},{"title":"VLA-Adapter: An Effective Paradigm for Tiny-Scale VLA Model (arXiv:2509.09372)","url":"https://arxiv.org/html/2509.09372"}],"as_of":"2025-09","related_ids":["querying-transformer","cross-attention","perceiver-resampler","action-head","octo","vla-adapter"],"name":"Learnable Query","alt":"可学习查询","abbr":"","aliases":["Action Query","Readout Token","Object Query"],"one_liner":"A set of vectors trained as model parameters that use attention to gather task-relevant information out of the input.","explanation":"A learnable query is a set of vectors that are randomly initialized and updated during training, rather than coming from the input; they act as 'queries' in an attention computation that reads the input features, pooling the needed information into a fixed number of outputs. Facebook's 2020 object-detection model DETR uses a set of object queries, each producing one detection; 2023's Q-Former in BLIP-2 uses 32 learnable queries to extract visual features from a frozen image encoder before passing them to a large language model. In robot models, Octo inserts readout tokens into its sequence: they can see the preceding observation and task tokens but aren't seen by those tokens in turn, and the action head generates the action from their output; 2025's VLA-Adapter adds ActionQuery tokens inside a vision-language model (64 of them worked best in its experiments), specifically to gather action-relevant multimodal information for the policy network. Its role is to compress an input of varying length into a fixed-size, task-oriented representation.","example":"In BLIP-2, an image is first turned into hundreds of feature tokens by the vision encoder, and Q-Former's 32 query vectors use cross-attention to pool them into 32 outputs, which are projected and prepended to the text before being fed into the language model.","related":["Querying Transformer","Cross-Attention","Perceiver Resampler","Action Head","Octo","VLA-Adapter"]},{"id":"querying-transformer","category":"model","sec":2,"tier":3,"sources":[{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (arXiv 2301.12597)","url":"https://arxiv.org/abs/2301.12597"},{"title":"BLIP-2 full text (HTML, Section 3.1 Model Architecture)","url":"https://arxiv.org/html/2301.12597"}],"as_of":"","related_ids":["projector-connector","perceiver-resampler","learnable-query","cross-attention","vision-language-model","vision-encoder"],"name":"Querying Transformer","alt":"Q-Former","abbr":"Q-Former","aliases":["Q-Former","BLIP-2 Q-Former"],"one_liner":"The small module in BLIP-2 that uses 32 learnable queries to distill image information before handing it to a large language model.","explanation":"Q-Former is a connector module Salesforce introduced in its 2023 BLIP-2 paper, used to bridge a frozen image encoder and a frozen large language model. It's a small Transformer of about 188 million parameters, initialized from BERT-base, that takes in 32 learnable query vectors (768 dimensions each), reads information out of the image features through cross-attention, and outputs a fixed-size 32×768 feature — far smaller than ViT-L/14's raw 257×1024 features, acting as an information bottleneck. Training has two stages: first learning image-text representations alongside the image encoder, then attaching the language model to learn generation. Because only Q-Former and a few other parameters are trained, BLIP-2 beats Flamingo-80B by 8.7% on zero-shot VQAv2 while using 54x fewer trainable parameters. InstructBLIP and others reuse it; LLaVA and Prismatic instead switch to a simpler MLP projector, and OpenVLA's base model uses an MLP too.","example":"BLIP-2 hands the image features extracted by ViT-g to Q-Former, which compresses them into 32 vectors; after a single linear projection, these serve as a 'soft visual prompt' prepended to the text input of the OPT or Flan-T5 language model.","related":["Projector / Connector","Perceiver Resampler","Learnable Query","Cross-Attention","Vision-Language Model","Vision Encoder"]},{"id":"perceiver-resampler","category":"model","sec":2,"tier":3,"sources":[{"title":"Flamingo: a Visual Language Model for Few-Shot Learning (arXiv:2204.14198)","url":"https://arxiv.org/abs/2204.14198"},{"title":"Perceiver: General Perception with Iterative Attention (arXiv:2103.03206)","url":"https://arxiv.org/abs/2103.03206"},{"title":"Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (GR-1, arXiv:2312.13139)","url":"https://arxiv.org/abs/2312.13139"}],"as_of":"","related_ids":["projector-connector","querying-transformer","learnable-query","cross-attention","visual-token","gr-1"],"name":"Perceiver Resampler","alt":"Perceiver 重采样器","abbr":"","aliases":["Perceiver"],"one_liner":"A module that uses a small set of learnable query vectors to compress a variable number of visual features into a fixed number of tokens.","explanation":"The Perceiver Resampler is a connector module DeepMind introduced in its 2022 Flamingo vision-language model, based on ideas from the 2021 Perceiver model. The number of features a vision encoder outputs varies with image resolution and video frame count, and is often large, making it expensive to feed directly into a language model. The resampler instead defines a small, fixed set of learnable query vectors (64 of them in Flamingo) that 'read' all the visual features through cross-attention, producing a fixed number of visual tokens as output. Flamingo's ablations show it outperforms an ordinary Transformer or MLP as the connector. It's conceptually close to BLIP-2's Q-Former, both compressing visual information with query vectors. In robot models, ByteDance's GR-1 uses it to compress image tokens too.","example":"ByteDance's GR-1 first encodes each frame into a large number of patch tokens with an MAE-pretrained ViT, compresses them down with a Perceiver Resampler, then feeds the result together with language and robot state into a GPT-style Transformer to predict actions and future frames.","related":["Projector / Connector","Querying Transformer","Learnable Query","Cross-Attention","Visual Token","GR-1 (ByteDance)"]},{"id":"tokenlearner","category":"model","sec":2,"tier":3,"sources":[{"title":"TokenLearner: What Can 8 Learned Tokens Do for Images and Videos? (arXiv:2106.11297)","url":"https://arxiv.org/abs/2106.11297"},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (arXiv:2212.06817)","url":"https://arxiv.org/abs/2212.06817"}],"as_of":"","related_ids":["visual-token","visual-token-pruning","rt-1","perceiver-resampler","querying-transformer","vision-transformer"],"name":"TokenLearner","alt":"TokenLearner","abbr":"","aliases":[],"one_liner":"A module that adaptively pools a large number of image tokens down into a handful of key tokens, to save compute.","explanation":"TokenLearner is a module Google's Ryoo and colleagues proposed in 2021 (NeurIPS 2021), with a paper title that literally asks what 8 learned tokens can do. A Vision Transformer usually cuts an image into tens to hundreds of patches, one token each, and attention's compute grows with the square of the token count. TokenLearner instead computes a spatial attention map for each output token based on the input content, uses it to weight and pool the feature map, and compresses a large number of tokens down to around 8; later layers only process these few tokens. The paper reaches competitive results on benchmarks like ImageNet and the Kinetics video-recognition set while using noticeably less compute. Its best-known use in embodied AI is in Google's RT-1, where it helps a large model meet the speed requirements of real-time control.","example":"RT-1 uses TokenLearner to compress each image's 81 visual tokens down to 8, so 6 frames of history become just 48 tokens fed into the Transformer; the paper reports this step gives about a 2.4x inference speedup.","related":["Visual Token","Visual Token Pruning","RT-1","Perceiver Resampler","Querying Transformer","Vision Transformer"]},{"id":"paligemma","category":"model","sec":2,"tier":2,"sources":[{"title":"PaliGemma: A versatile 3B VLM for transfer (arXiv:2407.07726)","url":"https://arxiv.org/abs/2407.07726"},{"title":"Google Developers Blog: Introducing PaliGemma 2","url":"https://developers.googleblog.com/en/introducing-paligemma-2-powerful-vision-language-models-simple-fine-tuning/"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"2024-12","related_ids":["vision-language-model","siglip","pi0","action-expert","vision-language-action-model","open-weight-model"],"name":"PaliGemma","alt":"PaliGemma","abbr":"","aliases":["PaliGemma 2"],"one_liner":"Google's open-weight vision-language model, combining a SigLIP visual encoder with a Gemma language model.","explanation":"PaliGemma is an open-weight vision-language model (VLM) Google released in May 2024, combining a SigLIP-So400m visual encoder with a Gemma-2B language model for about 3B parameters total. It isn't meant to be an out-of-the-box chat model; it's positioned as a base model meant to be transferred and fine-tuned, and the paper validates transfer performance on roughly 40 tasks. PaliGemma 2, released in December 2024, switched to the Gemma 2 language model and comes in three sizes (3B, 10B, 28B) and three input resolutions (224, 448, and 896 pixels). It's well known in embodied AI because Physical Intelligence uses it as the VLM backbone for π0: it's small, fast to run, and carries internet-scale vision-language knowledge, which makes it a good base for attaching an action expert to build a VLA.","example":"π0 uses the 3B-parameter PaliGemma as its backbone, adds a separate, randomly initialized action expert of about 300M parameters, for 3.3B parameters total, and generates continuous actions using flow matching.","related":["Vision-Language Model","SigLIP","π0","Action Expert","Vision-Language-Action Model","Open-weight Model"]},{"id":"qwen-vl","category":"model","sec":2,"tier":2,"sources":[{"title":"QwenLM/Qwen3-VL GitHub (News)","url":"https://github.com/QwenLM/Qwen3-VL"},{"title":"Qwen2.5-VL Technical Report (arXiv 2502.13923)","url":"https://arxiv.org/abs/2502.13923"},{"title":"Xiaomi-Robotics-0 (arXiv 2602.12684)","url":"https://arxiv.org/abs/2602.12684"}],"as_of":"2026-09","related_ids":["vision-language-model","multimodal-large-language-model","paligemma","internvl","vision-language-action-model","dynamic-native-resolution"],"name":"Qwen-VL","alt":"通义千问 Qwen-VL","abbr":"Qwen-VL","aliases":["Qwen2-VL","Qwen2.5-VL","Qwen3-VL"],"one_liner":"Alibaba's Qwen team's open-source vision-language model series, frequently used as a backbone for VLA models.","explanation":"Qwen-VL is the vision-language model series from Alibaba's Qwen team, able to look at images and video and answer in text. The first version was released in August 2023. Qwen2.5-VL came out in early 2025, in sizes from 3B to 72B, supports dynamic-resolution input, and can localize objects with boxes or points. Qwen3-VL began rolling out from September 2025, with dense versions from 2B to 32B and two mixture-of-experts versions, 30B-A3B and 235B-A22B, with native support for a 256K context window. Open weights, small model sizes, and strong grounding ability have made it a popular choice as the vision-language backbone for VLA models. Starting with Qwen3.5 in February 2026, Alibaba's mainline models natively support image input themselves.","example":"Xiaomi's Xiaomi-Robotics-0 uses Qwen3-VL-4B-Instruct as its vision-language backbone, followed by a diffusion Transformer that generates actions; Shanghai AI Lab's InternVLA-M1 uses Qwen2.5-VL-3B as its System 2.","related":["Vision-Language Model","Multimodal Large Language Model","PaliGemma","InternVL","Vision-Language-Action Model","Dynamic / Native Resolution"]},{"id":"dynamic-native-resolution","category":"model","sec":2,"tier":3,"sources":[{"title":"Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution (arXiv:2307.06304)","url":"https://arxiv.org/abs/2307.06304"},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv:2409.12191)","url":"https://arxiv.org/abs/2409.12191"},{"title":"Qwen/Qwen2-VL-7B-Instruct model card","url":"https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct"}],"as_of":"","related_ids":["vision-encoder","vision-transformer","visual-token","qwen-vl","rotary-position-embedding","inference-latency"],"name":"Dynamic / Native Resolution","alt":"动态分辨率（原生分辨率输入）","abbr":"","aliases":["Native Resolution","Naive Dynamic Resolution"],"one_liner":"Letting a vision encoder cut an image at its original size, so bigger images produce more visual tokens instead of being rescaled.","explanation":"Early Vision Transformers and CLIP-style encoders required rescaling or cropping every image to a fixed size, such as 224×224, which loses detail or distorts aspect ratio. Dynamic resolution instead keeps the image's original size and aspect ratio and cuts it directly into fixed-size patches, so a bigger image simply produces more tokens. Google's 2023 NaViT used 'sequence packing' to fit images of different sizes into the same training batch; Alibaba's 2024 Qwen2-VL introduced Naive Dynamic Resolution, using 2D rotary position encoding inside the vision encoder to record each patch's row and column, letting it handle any resolution. The benefit is that small images save compute while large images keep their detail, at the cost of a token count that grows with resolution. VLA backbones using this usually cap the token count within a min/max range to control inference latency.","example":"Qwen2-VL maps roughly every 28×28 pixels to one visual token, with a default range of 4 to 16,384 tokens per image, adjustable via min_pixels and max_pixels to balance speed against memory.","related":["Vision Encoder","Vision Transformer","Visual Token","Qwen-VL","Rotary Position Embedding","Inference Latency"]},{"id":"prismatic-vlms","category":"model","sec":2,"tier":3,"sources":[{"title":"Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models (arXiv 2402.07865)","url":"https://arxiv.org/html/2402.07865"},{"title":"TRI-ML/prismatic-vlms (GitHub)","url":"https://github.com/TRI-ML/prismatic-vlms"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"2024-07","related_ids":["openvla","vision-language-model","dinov2","siglip","projector-connector","llama"],"name":"Prismatic VLMs","alt":"Prismatic VLM","abbr":"","aliases":["Prismatic-7B"],"one_liner":"A Stanford/Toyota Research Institute study and open model that systematically compares VLM design choices; OpenVLA's base model.","explanation":"Prismatic VLM comes from a Stanford University and Toyota Research Institute (TRI) paper by Siddharth Karamcheti and colleagues, published at ICML 2024. The authors built a unified training and evaluation framework (12 benchmarks covering visual question answering, object localization, and challenge sets) to compare vision-language-model design choices one at a time: which vision encoder to use, whether to run a separate alignment-pretraining stage first, and whether the language model should be the base or the instruction-tuned version. The main findings: skipping the separate alignment stage and just doing single-stage training doesn't hurt performance and saves 20-25% of the compute; concatenating DINOv2 and SigLIP features gives a clear boost on localization tasks; and an instruction-tuned language model has no significant advantage. The resulting 7B-13B models beat the contemporary InstructBLIP and LLaVA v1.5, with code and dozens of checkpoints released open-source. Its connection to embodied AI is that OpenVLA is built directly on Prismatic-7B as its base, and the OpenVLA paper credits this fused vision encoder with helping spatial reasoning.","example":"OpenVLA's base model, Prismatic-7B, is made of a roughly 600-million-parameter fused DINOv2 + SigLIP vision encoder, a 2-layer MLP projector, and the Llama 2 7B language model.","related":["OpenVLA","Vision-Language Model","DINOv2","SigLIP","Projector / Connector","Llama"]},{"id":"nvidia-eagle-vlm","category":"model","sec":2,"tier":3,"sources":[{"title":"GitHub: NVlabs/EAGLE","url":"https://github.com/NVlabs/EAGLE"},{"title":"NVIDIA GEAR: GR00T N1.5","url":"https://research.nvidia.com/labs/gear/gr00t-n1_5/"},{"title":"GitHub: NVIDIA/Isaac-GR00T","url":"https://github.com/NVIDIA/Isaac-GR00T"}],"as_of":"2026-09","related_ids":["nvidia-isaac-gr00t-n1","vision-language-model","vision-encoder","dual-system-architecture","nvidia-cosmos-reason","vision-language-action-model"],"name":"NVIDIA Eagle VLM","alt":"Eagle（英伟达 VLM）","abbr":"","aliases":["Eagle 2","Eagle 2.5","NVEagle"],"one_liner":"NVIDIA's open-source vision-language model series, used as the vision-language backbone in GR00T N1 through N1.6.","explanation":"Eagle is an open-source vision-language model (VLM) series from NVIDIA's research team. The original Eagle (August 2024) studied how to combine multiple vision encoders, finding that simply concatenating the visual tokens from several complementary encoders worked well; Eagle 2 (January 2025) published the details of its post-training data strategy, with sizes including 1B, 2B, and 9B; Eagle 2.5 (technical report released April 2025) targets long context, supporting up to 128K tokens and strengthening long-video and high-resolution image understanding. Its main role in embodied AI is as the 'System 2' vision-language backbone for GR00T: GR00T N1 uses Eagle 2, while N1.5 and N1.6 use improved versions from the Eagle 2.5 family. According to the official GR00T repository, N1.7 onward switches to Cosmos-Reason2-2B instead.","example":"GR00T N1.5 starts from Eagle 2.5 and fine-tunes it further for object localization and physical understanding, while keeping this VLM frozen during both pretraining and fine-tuning so its parameters never update.","related":["NVIDIA Isaac GR00T N1","Vision-Language Model","Vision Encoder","Dual-System Architecture (System 1 / System 2)","NVIDIA Cosmos Reason","Vision-Language-Action Model"]},{"id":"native-multimodal","category":"model","sec":2,"tier":2,"sources":[{"title":"Google Blog: Introducing Gemini","url":"https://blog.google/technology/ai/google-gemini-ai/"},{"title":"Chameleon: Mixed-Modal Early-Fusion Foundation Models (arXiv:2405.09818)","url":"https://arxiv.org/abs/2405.09818"},{"title":"Meta AI Blog: The Llama 4 herd","url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"}],"as_of":"2025-04","related_ids":["multimodal-large-language-model","unified-multimodal-model","multimodal-fusion","visual-token","bagel","world-action-model"],"name":"Native Multimodal","alt":"原生多模态","abbr":"","aliases":["Native Multimodal Model","Natively Multimodal"],"one_liner":"Training a model on multiple modalities together from the very start of pretraining, instead of bolting a vision module onto a language model.","explanation":"Native multimodal means a model is trained from the start of pretraining on text, image, audio, and video data together, with every modality processed inside the same backbone network. It's the opposite of the ‘stitched-together’ approach, where a vision encoder and a language model are each trained separately first and then connected with a projector layer for further training. Google emphasized that Gemini was natively multimodal when it launched it in December 2023; Meta's Chameleon (2024) discretized images into tokens too and mixed them with text from the very beginning of training (early fusion); Llama 4 (April 2025) likewise claims to be natively multimodal and uses early fusion. Proponents argue this lets modalities combine more deeply and makes it easier to handle understanding and generation together; the cost is higher training expense and a harder balance to strike in the data mix across modalities. In embodied AI, unified multimodal models and world-action models that train video and action within the same sequence follow the same idea.","example":"Gemini was jointly pretrained on text, image, audio, and video data together from the start, whereas LLaVA attaches a vision encoder in front of an already-trained language model and then does further training — a stitched-together design.","related":["Multimodal Large Language Model","Unified Multimodal Model","Multimodal Fusion","Visual Token","BAGEL","World Action Model"]},{"id":"unified-multimodal-model","category":"model","sec":2,"tier":3,"sources":[{"title":"Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities (arXiv:2505.02567)","url":"https://arxiv.org/abs/2505.02567"},{"title":"BAGEL (ByteDance-Seed GitHub)","url":"https://github.com/ByteDance-Seed/Bagel"},{"title":"WorldVLA: Towards Autoregressive Action World Model (arXiv:2506.21539)","url":"https://arxiv.org/abs/2506.21539"}],"as_of":"2025-06","related_ids":["multimodal-large-language-model","native-multimodal","bagel","hybrid-autoregressive-diffusion-architecture","world-model","worldvla"],"name":"Unified Multimodal Model","alt":"统一多模态模型","abbr":"UMM","aliases":["UMM","Unified Understanding and Generation Model"],"one_liner":"A single model that can both understand images and video by answering questions about them, and generate or edit images and video.","explanation":"A unified multimodal model puts multimodal understanding (answering questions about an image) and visual generation (drawing or editing an image, or generating video, from text) into a single model. These used to be two separate lines of work: understanding mostly used autoregressive multimodal large language models, and generation mostly used diffusion models. GPT-4o's native image generation brought wide attention to this direction, and a leading open-source example is ByteDance Seed's BAGEL, released in May 2025 (a Mixture-of-Transformers design, with 7 billion active parameters and 14 billion total). Roughly speaking, these fall into three categories by generation method: autoregressive, diffusion, and hybrid autoregressive-diffusion. Embodied AI cares about this direction because a single model that can both understand a scene and 'imagine what it will look like after this action' can double as a policy and a world model at once.","example":"Alibaba DAMO Academy's WorldVLA puts a VLA and a world model into the same autoregressive framework: it outputs an action given the current view, and also predicts the next frame given the view and an action; the paper reports the two reinforce each other, outperforming a standalone action model or world model.","related":["Multimodal Large Language Model","Native Multimodal","BAGEL","Hybrid Autoregressive-Diffusion Architecture","World Model","WorldVLA"]},{"id":"bagel","category":"model","sec":2,"tier":3,"sources":[{"title":"BAGEL: Emerging Properties in Unified Multimodal Pretraining (arXiv 2505.14683)","url":"https://arxiv.org/abs/2505.14683"},{"title":"ByteDance-Seed/Bagel GitHub","url":"https://github.com/ByteDance-Seed/Bagel"},{"title":"BAGEL 论文 HTML 版","url":"https://arxiv.org/html/2505.14683"}],"as_of":"2025-05","related_ids":["unified-multimodal-model","mixture-of-transformers","siglip","variational-autoencoder","world-model","bytedance-seed"],"name":"BAGEL","alt":"BAGEL","abbr":"","aliases":["BAGEL-7B-MoT","Emerging Properties in Unified Multimodal Pretraining"],"one_liner":"ByteDance Seed's open-source unified multimodal model that can both understand images and generate or edit them.","explanation":"BAGEL is a unified multimodal foundation model ByteDance's Seed team open-sourced in May 2025, with code and weights released under the Apache 2.0 license; a single model handles image-text understanding, text-to-image generation, and image editing all at once. It's built on the Qwen2.5 large language model and uses a Mixture-of-Transformers (MoT) design: understanding and generation each get their own set of parameters, but every token shares the same self-attention at each layer, giving it 7B active parameters and 14B total. The understanding side uses a SigLIP2 vision encoder; the generation side uses FLUX's VAE to compress images into latent space. The model was pretrained on trillions of tokens of interleaved image-text, video, and web data, and the paper reports that capabilities such as free-form image editing, future-frame prediction, viewpoint rotation, and 'world navigation' emerge as scale increases. It demonstrates putting understanding and generation into one model, which is relevant to the unified-multimodal-model and world-model directions in embodied AI.","example":"Given BAGEL an image and an editing instruction in words, it can output the edited image directly; given navigation commands like 'move forward' or 'turn,' it can generate the view after that viewpoint change.","related":["Unified Multimodal Model","Mixture-of-Transformers","SigLIP","Variational Autoencoder","World Model","ByteDance Seed"]},{"id":"generative-model","category":"model","sec":3,"tier":2,"sources":[{"title":"Google Machine Learning: Background: What is a Generative Model?","url":"https://developers.google.com/machine-learning/gan/generative"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["diffusion-model","flow-matching","variational-autoencoder","generative-adversarial-network","action-multimodality","diffusion-policy"],"name":"Generative Model","alt":"生成模型","abbr":"","aliases":[],"one_liner":"A model that learns the probability distribution behind data and can sample new data from it.","explanation":"A generative model learns the distribution of the data itself, p(x), or a conditional distribution p(x|c), so it can produce new samples. This is in contrast to a discriminative model, which only learns p(y|x) and is used for classification or scoring. Common generative models include autoregressive models (which predict the next token one at a time, like GPT), variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, and flow matching. In embodied AI, generative models show up in three main places. First, for generating actions: demonstration data for the same scene often contains several equally valid ways to solve it (action multimodality), and methods like Diffusion Policy and π0 use diffusion or flow matching to model the whole action distribution instead of regressing to a single average. Second, for generating future frames, which is what world models and video-prediction models do. Third, for generating training data itself, such as using a video generation model to synthesize robot demonstrations.","example":"Suppose half the demonstrations go left around an obstacle and half go right: a model trained with direct regression outputs the average of the two and drives straight into it, while a generative model like Diffusion Policy samples either ‘go left’ or ‘go right.’","related":["Diffusion Model","Flow Matching","Variational Autoencoder","Generative Adversarial Network","Action Multimodality","Diffusion Policy"]},{"id":"autoencoder","category":"model","sec":3,"tier":2,"sources":[{"title":"Deep Learning, Chapter 14: Autoencoders (Goodfellow, Bengio, Courville)","url":"https://www.deeplearningbook.org/contents/autoencoders.html"},{"title":"World Models (Ha & Schmidhuber, interactive article)","url":"https://worldmodels.github.io/"}],"as_of":"","related_ids":["variational-autoencoder","conditional-variational-autoencoder","masked-autoencoder","vector-quantized-variational-autoencoder","latent-space","latent-diffusion-model"],"name":"Autoencoder","alt":"自编码器","abbr":"AE","aliases":["AE"],"one_liner":"A network that compresses input into a short vector and then reconstructs it, learning what matters most in the data.","explanation":"An autoencoder consists of an encoder and a decoder: the encoder compresses input, such as an image, into a low-dimensional vector, and the decoder reconstructs the input from that vector, trained so the reconstruction is as close as possible to the original, with no human labels needed. Because the middle vector is smaller than the input, the network is forced to keep only the most important information, which is why it's commonly used for dimensionality reduction and feature learning. The idea goes back decades in neural networks, to work such as LeCun's in 1987. Common variants include the variational autoencoder (VAE), which makes the middle vector follow a probability distribution so new samples can be generated; the masked autoencoder (MAE), which hides part of the input and reconstructs it, used for visual pretraining; and VQ-VAE, which replaces the vector with a discrete index into a codebook. In embodied AI, latent diffusion models, video tokenizers, and latent action models all use this to compress high-dimensional data into a latent space before processing it.","example":"“World Models” used a convolutional VAE to compress every frame of a racing-game screen into a 32-dimensional vector, so the prediction model downstream only ever operates on those 32 numbers.","related":["Variational Autoencoder","Conditional Variational Autoencoder","Masked Autoencoder","Vector-Quantized Variational Autoencoder","Latent Space","Latent Diffusion Model"]},{"id":"latent-space","category":"model","sec":3,"tier":2,"sources":[{"title":"Wikipedia: Latent space","url":"https://en.wikipedia.org/wiki/Latent_space"},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752)","url":"https://arxiv.org/abs/2112.10752"}],"as_of":"","related_ids":["embedding","variational-autoencoder","latent-diffusion-model","latent-world-model","latent-action","representation-learning"],"name":"Latent Space","alt":"潜在空间","abbr":"","aliases":["Latent Feature Space"],"one_liner":"The internal representation space a model compresses raw data into, where similar things end up close together.","explanation":"Latent space is the vector space a neural network uses internally to represent data. An encoder compresses high-dimensional raw data — images, audio, actions — into vectors with far fewer dimensions, and the space those vectors live in is the latent space; when training goes well, samples with similar meaning end up close together in it. Its value is that it saves compute and captures what matters: generating, predicting, or planning in latent space is much cheaper than working directly with pixels, and it's less easily thrown off by irrelevant detail. Autoencoders and variational autoencoders (VAEs, autoencoders that can sample to generate new data) both rely on it. Common uses in embodied AI include latent diffusion models, which denoise and generate images and video in latent space; latent world models, which predict the future in terms of latent states; and latent action models, which learn an abstract ‘action code’ from video with no action labels.","example":"A latent diffusion model (the basis of Stable Diffusion) first uses a pretrained autoencoder to compress an image into a much smaller latent representation, runs denoising diffusion in that space, and only decodes back to pixels at the end — far cheaper to train and run than diffusing directly on pixels.","related":["Embedding","Variational Autoencoder","Latent Diffusion Model","Latent World Model","Latent Action","Representation Learning"]},{"id":"variational-autoencoder","category":"model","sec":3,"tier":2,"sources":[{"title":"Auto-Encoding Variational Bayes (arXiv 1312.6114)","url":"https://arxiv.org/abs/1312.6114"},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models (arXiv 2112.10752)","url":"https://arxiv.org/abs/2112.10752"}],"as_of":"","related_ids":["autoencoder","latent-space","conditional-variational-autoencoder","vector-quantized-variational-autoencoder","latent-diffusion-model","video-tokenizer"],"name":"Variational Autoencoder","alt":"变分自编码器","abbr":"VAE","aliases":["VAE"],"one_liner":"A model that compresses data into a distribution over latent variables, and can sample from it to reconstruct or generate data.","explanation":"The variational autoencoder was proposed by Kingma and Welling in 2013 and is a type of generative model. Its encoder compresses the input (such as an image) into a low-dimensional latent variable, but outputs a distribution (a mean and variance) rather than a single fixed point; its decoder samples from that distribution and reconstructs data from it. Training optimizes two things at once: reconstructions should look right, and the latent distribution should stay close to a standard normal distribution, which keeps the latent space continuous and smooth, so decoding a randomly sampled point still gives a plausible result. Thanks to the reparameterization trick, a model with this kind of random sampling can still be trained with gradient descent. Its most common use today is as a compressor in front of a diffusion model: Stable Diffusion and Wan both use a VAE to compress pixels into latent space before denoising; in robotics, ACT uses its conditional variant, the CVAE, to capture the variety in demonstration actions.","example":"Stable Diffusion's VAE compresses a 512×512 color image into a 64×64×4 latent variable; the diffusion model does all its denoising in that much smaller space, and only the decoder turns the result back into an image at the end.","related":["Autoencoder","Latent Space","Conditional Variational Autoencoder","Vector-Quantized Variational Autoencoder","Latent Diffusion Model","Video Tokenizer"]},{"id":"conditional-variational-autoencoder","category":"model","sec":3,"tier":2,"sources":[{"title":"Learning Structured Output Representation using Deep Conditional Generative Models (NeurIPS 2015)","url":"https://papers.nips.cc/paper_files/paper/2015/hash/8d55a249e6baa5c06772297520da2051-Abstract.html"},{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv:2304.13705)","url":"https://arxiv.org/html/2304.13705"}],"as_of":"","related_ids":["variational-autoencoder","action-chunking-with-transformers","action-multimodality","action-chunking","generative-model","autoencoder"],"name":"Conditional Variational Autoencoder","alt":"条件变分自编码器","abbr":"CVAE","aliases":["CVAE","Conditional VAE"],"one_liner":"A VAE that learns an output distribution given a condition, so the same input can generate several valid outputs.","explanation":"A conditional variational autoencoder extends the variational autoencoder (VAE): both the encoder and decoder additionally take a condition, such as the current observation, and the model learns the distribution of outputs given that condition, rather than one single answer. It was proposed by Sohn, Lee, and Yan at NeurIPS 2015, originally for structured-output prediction tasks such as image segmentation. During training, the encoder sees the real output, compresses its “style” into a latent variable z, and the decoder reconstructs the output from the condition and z; at inference, z is instead drawn from the prior. Imitation learning uses it to handle the variety in demonstrations: for the same scene, different people might take different paths, and ordinary regression would average several valid approaches into one wrong action. ACT (2023) has the encoder take the joint positions and target action sequence to get z, and at test time drops the encoder and just sets z to the prior mean of 0.","example":"ACT uses a CVAE to predict a 100-step action chunk at once, and with only about 10 minutes of demonstration data, performs fine-motor tasks on a low-cost dual-arm setup with an 80% to 90% success rate.","related":["Variational Autoencoder","Action Chunking with Transformers","Action Multimodality","Action Chunking","Generative Model","Autoencoder"]},{"id":"generative-adversarial-network","category":"model","sec":3,"tier":2,"sources":[{"title":"Generative Adversarial Networks (Goodfellow et al., arXiv 1406.2661)","url":"https://arxiv.org/abs/1406.2661"},{"title":"AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control (arXiv 2104.02180)","url":"https://arxiv.org/abs/2104.02180"}],"as_of":"","related_ids":["generative-model","generative-adversarial-imitation-learning","adversarial-motion-priors","diffusion-model","variational-autoencoder","entropy-collapse-mode-collapse"],"name":"Generative Adversarial Network","alt":"生成对抗网络","abbr":"GAN","aliases":["GAN"],"one_liner":"A generative model where a generator fakes data and a discriminator tries to catch the fakes, trained against each other.","explanation":"Generative adversarial networks (GANs) were introduced by Ian Goodfellow and colleagues in 2014 and are a class of generative model — a model that learns to produce new data rather than just classify it. A GAN has two networks: a generator that turns random noise into samples, and a discriminator that judges whether a sample is real or generator-made. The two train in turns, with the generator trying to fool the discriminator and the discriminator trying to catch it; ideally the generator ends up learning the true data distribution. GANs produce an image in a single forward pass, so they're fast, but training is unstable and prone to mode collapse (generating only a few kinds of samples), and in recent years diffusion models have mostly replaced them for image and video generation. The ‘discriminator as a judge’ idea is still common in robotics: both Generative Adversarial Imitation Learning (GAIL) and Adversarial Motion Priors (AMP) use a discriminator to judge how closely a robot's motion resembles demonstration data, then feed that judgment to reinforcement learning as a reward.","example":"Training a simulated humanoid with AMP: the discriminator compares motion clips produced by the policy to motion-capture data, giving higher reward the more human-like they look, which teaches the character natural walking and running gaits.","related":["Generative Model","Generative Adversarial Imitation Learning","Adversarial Motion Priors","Diffusion Model","Variational Autoencoder","Entropy Collapse / Mode Collapse"]},{"id":"normalizing-flow","category":"model","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Flow-based generative model","url":"https://en.wikipedia.org/wiki/Flow-based_generative_model"},{"title":"Variational Inference with Normalizing Flows (arXiv:1505.05770)","url":"https://arxiv.org/abs/1505.05770"},{"title":"Flow Matching for Generative Modeling (arXiv:2210.02747)","url":"https://arxiv.org/abs/2210.02747"}],"as_of":"","related_ids":["generative-model","flow-matching","variational-autoencoder","diffusion-model","generative-adversarial-network","velocity-field"],"name":"Normalizing Flow","alt":"标准化流","abbr":"","aliases":["Flow-based Generative Model"],"one_liner":"A generative model that turns a simple distribution into a complex one through a chain of invertible transforms, with exact probabilities computable.","explanation":"A normalizing flow is a class of generative model: starting from a simple distribution such as a standard Gaussian, it passes the sample through a chain of invertible neural-network transforms to arrive at a complex data distribution. Because every step is invertible, the change-of-variables formula from probability theory gives the exact likelihood of a sample, so the model can be trained directly by maximum likelihood; to generate, you just sample noise and run it forward through the network once. The cost is that every layer must be invertible, with an easy-to-compute Jacobian determinant (which measures how much a transform expands or shrinks volume), which restricts the network design. Rezende and Mohamed popularized the idea for variational inference in 2015; representative models include NICE, RealNVP, and Glow. The 2018 continuous normalizing flow wrote the transform as an ordinary differential equation, and flow matching is precisely a simulation-free way of training a continuous normalizing flow — the two share the word 'flow' for a reason.","example":"Glow (2018) builds a flow model out of coupling layers and invertible 1×1 convolutions, able both to generate realistic face images and to report the exact log-likelihood of any given image.","related":["Generative Model","Flow Matching","Variational Autoencoder","Diffusion Model","Generative Adversarial Network","Velocity Field"]},{"id":"vector-quantization","category":"model","sec":3,"tier":3,"sources":[{"title":"Wikipedia: Vector quantization","url":"https://en.wikipedia.org/wiki/Vector_quantization"},{"title":"SoundStream: An End-to-End Neural Audio Codec (arXiv:2107.03312)","url":"https://arxiv.org/abs/2107.03312"},{"title":"Behavior Generation with Latent Actions (VQ-BeT, arXiv:2403.03181)","url":"https://arxiv.org/abs/2403.03181"}],"as_of":"","related_ids":["vector-quantized-variational-autoencoder","action-tokenizer","finite-scalar-quantization","tokenizer","behavior-transformer","latent-action"],"name":"Vector Quantization","alt":"向量量化","abbr":"VQ","aliases":["VQ","Codebook","Residual VQ","RVQ"],"one_liner":"Preparing a 'codebook' and replacing a continuous vector with the ID of its nearest codeword, turning it into a discrete token.","explanation":"Vector quantization is a classic compression technique from signal processing, systematically developed in the early 1980s by Robert Gray and colleagues. It works by preparing a set of representative vectors called a codebook, where each vector is called a codeword; for any input vector, you find the nearest codeword and record only its ID. This is essentially the same as k-means clustering, with codewords as the cluster centers. In deep learning, VQ is one of the main tools for turning continuous signals like images, audio, and actions into discrete tokens, which then makes it possible to model them with a language model's 'next-token prediction.' Residual vector quantization (RVQ) uses several codebooks in sequence, each quantizing the error left over from the previous level; Google's 2021 SoundStream audio codec uses it to let a single model switch bitrate anywhere between 3 and 18 kbps.","example":"VQ-BeT (ICML 2024) encodes a robot's continuous actions into discrete tokens using hierarchical residual vector quantization, replacing the k-means clustering of its predecessor BeT, and the paper reports about 5x faster inference than Diffusion Policy.","related":["Vector-Quantized Variational Autoencoder","Action Tokenizer","Finite Scalar Quantization","Tokenizer","Behavior Transformer","Latent Action"]},{"id":"vector-quantized-variational-autoencoder","category":"model","sec":3,"tier":3,"sources":[{"title":"Neural Discrete Representation Learning (arXiv:1711.00937)","url":"https://arxiv.org/abs/1711.00937"},{"title":"Genie: Generative Interactive Environments (arXiv:2402.15391)","url":"https://arxiv.org/abs/2402.15391"}],"as_of":"","related_ids":["variational-autoencoder","vector-quantization","video-tokenizer","latent-action-model","genie","lapa"],"name":"Vector-Quantized Variational Autoencoder","alt":"向量量化变分自编码器","abbr":"VQ-VAE","aliases":["VQ-VAE","VQ-VAE-2"],"one_liner":"An autoencoder whose encoded output is quantized by a codebook into discrete IDs, commonly used to turn images, video, or actions into tokens.","explanation":"VQ-VAE was proposed by DeepMind's van den Oord, Vinyals, and Kavukcuoglu in 2017 (NeurIPS 2017). It differs from a variational autoencoder (VAE, a generative model that compresses data into a continuous latent variable and reconstructs it) in two ways: the encoder's output first goes through vector quantization, replaced with the nearest codeword in a codebook, so the latent variable is discrete; and the prior isn't a fixed Gaussian but a separately trained autoregressive model (such as PixelCNN) instead. The quantization step isn't differentiable, so training copies the gradient from the decoder side straight to the encoder with a straight-through estimator, plus a commitment loss that pulls the encoder's output toward the codeword. It also eases posterior collapse, a common VAE problem. Many later image and video tokenizers, and latent action models, are built on top of it.","example":"Google DeepMind's Genie uses a VQ-VAE to compress video into discrete tokens, and its latent action model uses a VQ-VAE-style objective too, learning discrete latent actions with just 8 codewords from unlabeled game video with no action labels, letting a person 'control' the generated world frame by frame.","related":["Variational Autoencoder","Vector Quantization","Video Tokenizer","Latent Action Model","Genie (Original)","LAPA"]},{"id":"finite-scalar-quantization","category":"model","sec":3,"tier":3,"sources":[{"title":"Finite Scalar Quantization: VQ-VAE Made Simple (Mentzer et al., arXiv 2309.15505)","url":"https://arxiv.org/abs/2309.15505"},{"title":"NVIDIA Cosmos Tokenizer (GitHub)","url":"https://github.com/NVIDIA/Cosmos-Tokenizer"}],"as_of":"","related_ids":["vector-quantization","vector-quantized-variational-autoencoder","action-tokenizer","video-tokenizer","token","nvidia-cosmos"],"name":"Finite Scalar Quantization","alt":"有限标量量化","abbr":"FSQ","aliases":["FSQ"],"one_liner":"Rounding each dimension of a continuous vector directly to a few fixed levels, replacing a VQ codebook for discretization.","explanation":"Finite scalar quantization was proposed in 2023 by Mentzer and colleagues at Google Research, in a paper subtitled 'VQ-VAE Made Simple.' Traditional vector quantization (VQ) maintains a learnable codebook, and training often suffers from 'codebook collapse,' where large numbers of code entries never get used, requiring extra tricks like commitment loss and entropy penalties to fix. FSQ instead projects features down to very few dimensions (usually fewer than 10), clips each dimension to a range, and rounds it directly to a small number of fixed levels; the combination of levels across dimensions forms an implicit codebook. There's no codebook to learn and no collapse to worry about, and it performs on par with VQ on tasks like image generation and depth estimation. It's commonly used to turn images, video, or action sequences into discrete tokens for an autoregressive Transformer to process.","example":"NVIDIA's Cosmos Tokenizer discrete version uses FSQ to compress video into discrete tokens, with an index range of about 64,000 (64K).","related":["Vector Quantization","Vector-Quantized Variational Autoencoder","Action Tokenizer","Video Tokenizer","Token","NVIDIA Cosmos"]},{"id":"diffusion-model","category":"model","sec":4,"tier":1,"sources":[{"title":"Denoising Diffusion Probabilistic Models (arXiv 2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"Diffusion model - Wikipedia","url":"https://en.wikipedia.org/wiki/Diffusion_model"}],"as_of":"","related_ids":["diffusion-policy","denoising-diffusion-probabilistic-model","flow-matching","action-multimodality","denoising-steps","generative-model"],"name":"Diffusion Model","alt":"扩散模型","abbr":"","aliases":["Diffusion"],"one_liner":"A generative model that learns to remove noise step by step, then generates new data by denoising from pure noise.","explanation":"A diffusion model has two processes: a forward process that keeps adding Gaussian noise to real data until it becomes pure noise, and a reverse process where a neural network is trained to predict and remove the noise at each step. Generation runs the reverse process from random noise, denoising repeatedly to produce a new sample. The idea was proposed by Sohl-Dickstein and colleagues in 2015, and Ho and colleagues' 2020 DDPM made it genuinely practical, after which it became the mainstream approach behind image and video generators such as Stable Diffusion. In robotics it is used to generate actions: 2023's Diffusion Policy denoises a segment of future action conditioned on the current observation, and can represent several equally valid ways of doing something in the same scene, called action multimodality, which a single regressed average cannot capture. The cost is that denoising takes multiple steps, so inference tends to be slower.","example":"Diffusion Policy (Chi and colleagues, 2023) reached an average success rate 46.9% higher than prior methods across 12 manipulation tasks in 4 benchmarks; on a T-block-pushing task, faced with two equally valid ways to push, going around the left or the right, it learns both modes instead of averaging them into one.","related":["Diffusion Policy","Denoising Diffusion Probabilistic Model","Flow Matching","Action Multimodality","Denoising Steps","Generative Model"]},{"id":"denoising-diffusion-probabilistic-model","category":"model","sec":4,"tier":2,"sources":[{"title":"Denoising Diffusion Probabilistic Models (arXiv 2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["diffusion-model","denoising-diffusion-implicit-model","noise-schedule","denoising-steps","u-net","diffusion-policy"],"name":"Denoising Diffusion Probabilistic Model","alt":"去噪扩散概率模型","abbr":"DDPM","aliases":["DDPM"],"one_liner":"The 2020 diffusion model that made the approach work well: add noise step by step, then learn to remove it step by step.","explanation":"The denoising diffusion probabilistic model was proposed by Jonathan Ho, Ajay Jain, and Pieter Abbeel in 2020, and is the base version behind today's diffusion models. It defines two processes: a forward process that adds Gaussian noise to a real image a little at a time, following a fixed noise schedule, a timetable of how much noise to add at each step, until it becomes pure noise; and a reverse process, where a neural network is trained to predict, and subtract, the noise added at each step. The training objective is simple, just the mean squared error between predicted and true noise. To generate, start from pure noise and denoise step by step into a new sample. The paper used T=1,000 steps and a U-Net, reaching an FID of 3.17 on CIFAR-10. The downside is that sampling needs many steps, which later methods such as DDIM greatly cut down. Robotics' Diffusion Policy applies this same method to action sequences.","example":"Diffusion Policy trains by adding noise to demonstrated action sequences over a 100-step diffusion process, teaching the network to gradually turn noise back into actions; at deployment it switches to DDIM and samples only 10 steps, taking about 0.1 seconds on an RTX 3080.","related":["Diffusion Model","Denoising Diffusion Implicit Model","Noise Schedule","Denoising Steps","U-Net","Diffusion Policy"]},{"id":"flow-matching","category":"model","sec":4,"tier":1,"sources":[{"title":"Flow Matching for Generative Modeling (arXiv 2210.02747)","url":"https://arxiv.org/abs/2210.02747"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"}],"as_of":"","related_ids":["diffusion-model","velocity-field","rectified-flow","flow-matching-loss","pi0","action-expert"],"name":"Flow Matching","alt":"流匹配","abbr":"FM","aliases":["FM","Conditional Flow Matching","CFM"],"one_liner":"Learning a velocity field that smoothly transports noise into real data, then generating samples by integrating along it.","explanation":"Flow matching was proposed by Meta's Lipman and colleagues in 2022 (published at ICLR 2023), for training continuous normalizing flows, a class of generative model that continuously deforms noise into data. The method defines a path between a noise sample and a real sample, most commonly straight-line interpolation, and trains a network to predict the velocity at every point along that path, the velocity field, with a loss that is just the mean squared error between predicted and target velocity; training needs no simulation of the full trajectory. To generate, start from noise and numerically integrate along the network's predicted velocity; for robot actions, a handful to about ten steps is usually enough (π0 uses 10, GR00T N1 uses 4). It belongs to the same family as diffusion models, both turning noise into data step by step, and diffusion can be viewed as one specific path within flow matching; a straight-line path tends to train faster and need fewer sampling steps, in the same spirit as the contemporaneous rectified-flow idea. In robotics, π0 uses it to generate action chunks in its action expert, and GR00T N1 and others use it too.","example":"At inference, π0 starts from Gaussian noise and integrates 10 steps, each of size 0.1, to turn noise into a 50-step-long action chunk; during training, the loss is just the squared difference between the network's predicted velocity field and the target velocity.","related":["Diffusion Model","Velocity Field","Rectified Flow","Flow Matching Loss","π0","Action Expert"]},{"id":"velocity-field","category":"model","sec":4,"tier":3,"sources":[{"title":"Flow Matching for Generative Modeling (arXiv:2210.02747)","url":"https://arxiv.org/abs/2210.02747"},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow (arXiv:2209.03003)","url":"https://arxiv.org/abs/2209.03003"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["flow-matching","rectified-flow","flow-matching-loss","diffusion-flow-samplers","action-expert","pi0"],"name":"Velocity Field","alt":"速度场","abbr":"","aliases":["Vector Field","Velocity Vector Field"],"one_liner":"The function a flow-matching model learns: it tells each sample which direction to move, and how fast, at every timestep.","explanation":"The velocity field is what flow matching and rectified flow — a family of generative models introduced in 2022 by Lipman et al. and, separately, Liu et al. — actually learn. It's a function v(x, t): given a sample x at time t (0 means pure noise, 1 means real data), it outputs the direction and speed x should move right now. During training, the network regresses the velocity along a predetermined “noise-to-data” path, usually a straight line, so the target velocity is simply data minus noise — there's no need to simulate the full generation process step by step. At inference time, the model starts from random noise and integrates along the velocity field with an ODE solver (such as the Euler method) for a handful of steps to produce a sample. Straighter paths need fewer steps, which is why this approach usually samples faster than traditional diffusion models. In robotics, the action expert in π0 predicts the velocity field for a chunk of future actions.","example":"At inference, π0 first samples a chunk of Gaussian noise to serve as the “action chunk.” The action expert predicts a velocity field conditioned on the current image, instruction, and robot state, then integrates it for 10 steps with the forward Euler method, turning the noise into a continuous 50-step action sequence.","related":["Flow Matching","Rectified Flow","Flow Matching Loss","Diffusion / Flow Samplers","Action Expert","π0"]},{"id":"rectified-flow","category":"model","sec":4,"tier":3,"sources":[{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow (arXiv 2209.03003)","url":"https://arxiv.org/abs/2209.03003"},{"title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (SD3, arXiv 2403.03206)","url":"https://arxiv.org/abs/2403.03206"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["flow-matching","velocity-field","diffusion-model","one-step-generation","consistency-model","pi0"],"name":"Rectified Flow","alt":"整流流","abbr":"RF","aliases":["RF"],"one_liner":"A generative model that moves noise toward data along a straight line — the straighter the path, the fewer sampling steps needed.","explanation":"Rectified flow was proposed by Xingchao Liu, Chengyue Gong, and Qiang Liu in their 2022 paper 'Flow Straight and Fast.' It connects a noise sample and a data sample with a straight line and trains a network to predict the velocity along that line; generation starts from noise and solves an ordinary differential equation (ODE) that follows the predicted velocity to arrive at data. The training objective is a simple least-squares regression, and it's essentially equivalent to the flow matching Lipman and colleagues proposed around the same time when both use straight-line paths. Its other contribution is 'reflow': using a trained model to generate noise-data pairs, then retraining on those pairs, which makes the paths progressively straighter — straight enough that even a single Euler step can produce a decent result, which suits pairing with distillation for few-step or one-step generation. Stability AI's Stable Diffusion 3 trains with rectified flow; in robotics, π0's flow matching uses the same straight-line interpolation path to generate actions.","example":"π0 trains by linearly mixing a real action chunk with Gaussian noise according to a time τ, having the action expert predict the velocity between the two; at deployment, it starts from pure noise and integrates 10 Euler steps to get an executable action chunk.","related":["Flow Matching","Velocity Field","Diffusion Model","One-step Generation","Consistency Model","π0"]},{"id":"u-net","category":"model","sec":4,"tier":2,"sources":[{"title":"U-Net: Convolutional Networks for Biomedical Image Segmentation (arXiv 1505.04597)","url":"https://arxiv.org/abs/1505.04597"},{"title":"Scalable Diffusion Models with Transformers (arXiv 2212.09748)","url":"https://arxiv.org/abs/2212.09748"}],"as_of":"","related_ids":["convolutional-neural-network","diffusion-model","diffusion-transformer","latent-diffusion-model","diffusion-policy","residual-network"],"name":"U-Net","alt":"U-Net","abbr":"","aliases":["UNet"],"one_liner":"A U-shaped convolutional network that downsamples layer by layer, then upsamples back, with matching layers connected directly.","explanation":"U-Net is a convolutional network Ronneberger and colleagues proposed in 2015 for medical image segmentation. The left half is an encoder that downsamples layer by layer, shrinking resolution while extracting more abstract features; the right half is a decoder that upsamples layer by layer back to the original size; matching levels on the left and right are joined by skip connections, which pass the encoder's features directly to the decoder — drawn out, the shape looks like the letter U. This lets the network see global context without losing precise location, which suits tasks that take one image in and produce a same-sized result. It later became the standard backbone for the denoising network in diffusion models — Stable Diffusion and Stable Video Diffusion both use it — though the diffusion Transformer (DiT), introduced in 2022, has begun to replace it. Diffusion Policy in robotics also commonly uses a 1D temporal-convolution U-Net to denoise action sequences.","example":"When Stable Diffusion generates an image, at every step it feeds the noisy latent into the U-Net, which predicts the noise in it; that noise is subtracted before the next step, and repeating this dozens of times produces a clean image.","related":["Convolutional Neural Network","Diffusion Model","Diffusion Transformer","Latent Diffusion Model","Diffusion Policy","Residual Network"]},{"id":"diffusion-transformer","category":"model","sec":4,"tier":2,"sources":[{"title":"Scalable Diffusion Models with Transformers (arXiv 2212.09748)","url":"https://arxiv.org/abs/2212.09748"},{"title":"DiT project page (William Peebles)","url":"https://www.wpeebles.com/DiT"},{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)","url":"https://arxiv.org/abs/2410.07864"}],"as_of":"","related_ids":["diffusion-model","transformer","u-net","adaptive-layer-normalization","latent-diffusion-model","rdt-1b"],"name":"Diffusion Transformer","alt":"扩散 Transformer","abbr":"DiT","aliases":["DiT"],"one_liner":"Using a Transformer instead of a U-Net as a diffusion model's denoising network architecture.","explanation":"The Diffusion Transformer was proposed by William Peebles and Saining Xie in 2022 (published at ICCV 2023). Before this, diffusion models' denoising networks were mostly U-Nets, a convolutional encoder-decoder network; DiT replaces this with a standard Transformer: a VAE-compressed latent image is cut into patches and treated as tokens, and adaptive layer normalization (adaLN) injects conditioning information, such as the denoising timestep and class label, into every layer. The paper found that generation quality kept improving with more compute, following a clean scaling trend; the largest model, DiT-XL/2, reached an FID of 2.27 on 256×256 ImageNet. Video-generation models have since widely adopted this structure. Robotics uses it to generate actions too: RDT-1B (Robotics Diffusion Transformer) uses it for bimanual manipulation, and GR00T N1's action module is likewise a DiT variant.","example":"RDT-1B is a diffusion Transformer of about 1.2 billion parameters that, conditioned on a language instruction and camera images, denoises to generate a segment of upcoming action for a dual-arm robot.","related":["Diffusion Model","Transformer","U-Net","Adaptive Layer Normalization","Latent Diffusion Model","RDT-1B"]},{"id":"feature-wise-linear-modulation","category":"model","sec":4,"tier":3,"sources":[{"title":"FiLM: Visual Reasoning with a General Conditioning Layer (Perez et al., AAAI 2018)","url":"https://arxiv.org/abs/1709.07871"},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (arXiv 2212.06817)","url":"https://arxiv.org/abs/2212.06817"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["adaptive-layer-normalization","language-conditioned-policy","rt-1","diffusion-policy","efficientnet","cross-attention"],"name":"Feature-wise Linear Modulation","alt":"FiLM 特征调制","abbr":"FiLM","aliases":["FiLM","FiLM Layer","FiLM Conditioning"],"one_liner":"Using conditioning information to compute a per-channel scale and shift that modulates a network's intermediate features.","explanation":"FiLM was proposed by Ethan Perez, Aaron Courville, and colleagues (AAAI 2018) as a general-purpose layer for injecting conditioning information into a neural network. The idea is simple: a small network computes a scale γ and a shift β for each feature channel from the condition (such as a language-instruction vector), and the intermediate feature is transformed into γ·x + β. The original paper roughly halved the best error rate at the time on the CLEVR visual-reasoning benchmark. It's common in robotics: RT-1 uses FiLM to inject the language instruction into its pretrained EfficientNet image encoder, zero-initializing the layers that produce γ and β so FiLM starts out as an identity transform and doesn't disturb the pretrained weights; the CNN version of Diffusion Policy also uses FiLM at every convolutional layer to inject observation features.","example":"In RT-1, the instruction 'pick up the coke can' is first turned into a vector by the Universal Sentence Encoder, then modulates EfficientNet's layer features through FiLM, so the same image produces different visual features under different instructions.","related":["Adaptive Layer Normalization","Language-conditioned Policy","RT-1","Diffusion Policy","EfficientNet","Cross-Attention"]},{"id":"adaptive-layer-normalization","category":"model","sec":4,"tier":3,"sources":[{"title":"Scalable Diffusion Models with Transformers (DiT, arXiv 2212.09748)","url":"https://arxiv.org/html/2212.09748"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"},{"title":"openpi pi0.py 源码","url":"https://github.com/Physical-Intelligence/openpi/blob/main/src/openpi/models/pi0.py"}],"as_of":"","related_ids":["feature-wise-linear-modulation","diffusion-transformer","normalization-layers","action-expert","nvidia-isaac-gr00t-n1","cross-attention"],"name":"Adaptive Layer Normalization","alt":"自适应层归一化","abbr":"AdaLN","aliases":["AdaLN","adaLN-Zero","adaRMS"],"one_liner":"Computing layer normalization's scale and shift dynamically from conditioning information, such as a diffusion timestep, instead of fixing them.","explanation":"Layer normalization standardizes each token's features, then multiplies by a scale γ and adds a shift β; ordinarily these two parameters are fixed after training. Adaptive layer normalization instead has a small MLP regress γ and β from a conditioning vector — a diffusion timestep, a class label, a language feature — so that whenever the condition changes, the distribution of the whole layer's features changes with it, 'injecting' the condition into the network; the idea is the same as FiLM. It builds on the adaptive normalization used in GANs and U-Net diffusion models. Peebles and Xie used it systematically in the diffusion Transformer (DiT) in 2022 and introduced adaLN-Zero: an extra gating coefficient is regressed and initialized to zero, so each Transformer block starts out as an identity mapping, which makes training more stable. In embodied AI, NVIDIA's GR00T N1 uses AdaLN in its DiT action module to inject the denoising step, while using cross-attention to receive VLM features.","example":"The openpi implementation of π0.5 encodes the flow-matching timestep with a two-layer MLP into 'adarms_cond,' which modulates the normalization layers on the action tokens using adaptive RMSNorm.","related":["Feature-wise Linear Modulation","Diffusion Transformer","Normalization Layers","Action Expert","NVIDIA Isaac GR00T N1","Cross-Attention"]},{"id":"classifier-free-guidance","category":"model","sec":4,"tier":3,"sources":[{"title":"Classifier-Free Diffusion Guidance (arXiv 2207.12598)","url":"https://arxiv.org/abs/2207.12598"},{"title":"Hugging Face Diffusers: Text-to-image（guidance_scale 说明）","url":"https://huggingface.co/docs/diffusers/using-diffusers/conditional_image_generation"}],"as_of":"","related_ids":["diffusion-model","flow-matching","denoising-steps","text-to-video-image-to-video","generative-model","video-generation-model"],"name":"Classifier-Free Guidance","alt":"无分类器引导","abbr":"CFG","aliases":["CFG","Guidance Scale"],"one_liner":"Computing both a conditional and an unconditional prediction and extrapolating between them so generation follows the condition more closely.","explanation":"Classifier-free guidance was proposed by Jonathan Ho and Tim Salimans (a 2021 NeurIPS workshop paper, with a full arXiv version in 2022), for use with diffusion models, flow matching, and other generative models that denoise step by step. During training, the condition (such as a text prompt) is randomly dropped for some examples, so the same network learns to make both conditional and unconditional predictions; during generation, both are computed at every step, and the final direction is 'unconditional result + w × (conditional result − unconditional result),' where w is called the guidance scale. A larger w follows the condition more closely but reduces diversity, and too large a value introduces artifacts. It replaces the earlier 'classifier guidance,' which needed training a separate classifier, and has become a standard setting in text-to-image and text-to-video models, at the cost of one extra forward pass per step. In embodied AI, diffusion-based video world models and action generation models can also use it to control how closely they follow a language instruction.","example":"Generating an image from text with Stable Diffusion v1.5 in Diffusers, turning guidance_scale up from 2.5 to 10.5 makes the image follow the prompt more and more closely, though artifacts start to appear once it's too high.","related":["Diffusion Model","Flow Matching","Denoising Steps","Text-to-Video / Image-to-Video","Generative Model","Video Generation Model"]},{"id":"latent-diffusion-model","category":"model","sec":4,"tier":3,"sources":[{"title":"High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752)","url":"https://arxiv.org/abs/2112.10752"},{"title":"Cosmos World Foundation Model Platform for Physical AI (arXiv:2501.03575)","url":"https://arxiv.org/html/2501.03575"}],"as_of":"2025-01","related_ids":["diffusion-model","variational-autoencoder","diffusion-transformer","video-tokenizer","nvidia-cosmos","cross-attention"],"name":"Latent Diffusion Model","alt":"潜在扩散模型","abbr":"LDM","aliases":["LDM","Latent Diffusion"],"one_liner":"A model that first compresses data into a low-dimensional latent space with an autoencoder, then runs diffusion generation there.","explanation":"A diffusion model generates data by denoising step by step, but doing this directly on pixels is computationally expensive for high-resolution images and video. In December 2021, Rombach and colleagues at Germany's CompVis group proposed the latent diffusion model (CVPR 2022): an autoencoder is trained first to compress an image into a much smaller latent variable, the diffusion model denoises only in that latent space, and a decoder reconstructs the image at the end; conditions such as text are injected via cross-attention. This dramatically cuts training and inference cost, and the open-source Stable Diffusion is built on this framework. Video generation and world models commonly follow the same idea — for example, the diffusion-based world foundation models in NVIDIA's Cosmos run inside the latent space of the Cosmos video tokenizer (an encoder that compresses video into a latent variable), at a spatial-temporal compression ratio of 8×8×8.","example":"Stable Diffusion first compresses a 512×512 color image into a 64×64×4 latent variable, denoises on this much smaller tensor, and only decodes back to a 512×512 image at the end.","related":["Diffusion Model","Variational Autoencoder","Diffusion Transformer","Video Tokenizer","NVIDIA Cosmos","Cross-Attention"]},{"id":"noise-schedule","category":"model","sec":4,"tier":3,"sources":[{"title":"Improved Denoising Diffusion Probabilistic Models (arXiv:2102.09672)","url":"https://arxiv.org/abs/2102.09672"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["diffusion-model","denoising-diffusion-probabilistic-model","denoising-steps","flow-matching","diffusion-policy","prediction-target-parameterization"],"name":"Noise Schedule","alt":"噪声调度","abbr":"","aliases":["Noise Scheduler"],"one_liner":"The timetable specifying how much noise a diffusion model adds at each step, which affects training and generation quality.","explanation":"The noise schedule is a design choice in a diffusion model: it specifies, at step t of the forward noising process, how much of the original data remains and how much is noise. 2020's DDPM used a linear schedule; in 2021, OpenAI's Nichol and Dhariwal found that on low-resolution images this wastes steps, since the later steps are already almost pure noise, and proposed a cosine schedule instead, which removes information at a steady rate through the middle and changes more gently at both ends. The schedule determines how much training gets allocated to each noise level and noticeably affects generation quality; in flow matching, the corresponding design choice is the sampling distribution over training timesteps. Robot policies need this tuning too: Diffusion Policy uses a cosine schedule, and π0 deliberately oversamples high-noise timesteps when training with flow matching.","example":"The Diffusion Policy paper compares options and settles on the cosine schedule proposed by iDDPM, training with 100 diffusion steps but running only 10 DDIM steps at inference on the real robot, producing an action roughly every 0.1 seconds on an RTX 3080.","related":["Diffusion Model","Denoising Diffusion Probabilistic Model","Denoising Steps","Flow Matching","Diffusion Policy","Prediction Target Parameterization"]},{"id":"prediction-target-parameterization","category":"model","sec":4,"tier":3,"sources":[{"title":"Denoising Diffusion Probabilistic Models (DDPM, arXiv 2006.11239)","url":"https://arxiv.org/abs/2006.11239"},{"title":"Progressive Distillation for Fast Sampling of Diffusion Models (v-prediction, arXiv 2202.00512)","url":"https://arxiv.org/abs/2202.00512"},{"title":"Back to Basics: Let Denoising Generative Models Denoise (JiT, arXiv 2511.13720)","url":"https://arxiv.org/abs/2511.13720"}],"as_of":"2025-11","related_ids":["diffusion-model","flow-matching","rectified-flow","denoising-diffusion-probabilistic-model","velocity-field","denoising-loss"],"name":"Prediction Target Parameterization","alt":"预测目标参数化（ε / v / x₀ 预测）","abbr":"","aliases":["ε-prediction","v-prediction","x0-prediction"],"one_liner":"The choice of whether a diffusion model's network should output noise, velocity, or the clean data itself.","explanation":"When training a diffusion or flow-matching model, clean data x0 is first noised into x_t, and the network, seeing x_t, has to output some quantity to compute the loss against — there are several choices for what that quantity should be. ε-prediction has the network guess the noise that was added, which is what 2020's DDPM does; x0-prediction guesses the clean data directly; v-prediction, proposed by Salimans and Ho in 2022, predicts a weighted combination of noise and data, and is more stable with very few sampling steps; the velocity field predicted in flow matching and rectified flow also falls into this category. Given x_t and the timestep, the three are mathematically interconvertible, but they differ in how hard they are to train, how errors get weighted across noise levels, and how well they work with few sampling steps. A 2025 paper by Tianhong Li and Kaiming He, JiT, argues that predicting the clean data directly in high-dimensional pixel space is easier to learn than predicting noise or velocity. In robotics, Diffusion Policy defaults to ε-prediction, while flow-matching VLAs like π0 predict velocity.","example":"Diffusion Policy's noise-prediction network ε_θ takes a noised action sequence as input and outputs the estimated noise; π0's action expert instead outputs a velocity vector, and at deployment integrates 10 steps starting from pure noise to get an action chunk.","related":["Diffusion Model","Flow Matching","Rectified Flow","Denoising Diffusion Probabilistic Model","Velocity Field","Denoising Loss (Diffusion Loss)"]},{"id":"score-function-score-matching","category":"model","sec":4,"tier":3,"sources":[{"title":"Yang Song: Generative Modeling by Estimating Gradients of the Data Distribution (blog)","url":"https://yang-song.net/blog/2021/score/"},{"title":"Generative Modeling by Estimating Gradients of the Data Distribution (arXiv:1907.05600)","url":"https://arxiv.org/abs/1907.05600"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["diffusion-model","denoising-diffusion-probabilistic-model","energy-based-model","diffusion-policy","flow-matching","diffusion-flow-samplers"],"name":"Score Function / Score Matching","alt":"分数函数 / 分数匹配","abbr":"","aliases":["Score","Denoising Score Matching","Score-based Generative Model"],"one_liner":"The 'score' is the gradient of log-probability with respect to the input; score matching is how a network learns it.","explanation":"The score function is the gradient of the log probability density with respect to the input, ∇ₓ log p(x); it points in the direction where the data's probability increases fastest, and computing it doesn't require the hard-to-compute normalization constant. Score matching was proposed by Hyvärinen in 2005 as a way to train a model to fit this gradient without knowing the true distribution; the later denoising score matching variant instead 'adds noise to the data, then learns how to remove it.' In 2019, Song and Ermon learned the score across many noise levels and generated images with Langevin dynamics (a sampling method that follows the gradient while adding a bit of random noise at every step); the 2021 stochastic differential equation (SDE) framework then showed this is essentially the same class of model as denoising diffusion probabilistic models. What a diffusion model predicts as 'noise' is the score up to a rescaling, so understanding the score is the key to understanding what Diffusion Policy is really doing underneath.","example":"The Diffusion Policy paper describes its own method as learning a score gradient field over the action distribution: at inference, starting from random noise, it takes several noisy, Langevin-like steps along this gradient field and ends up with a robot action.","related":["Diffusion Model","Denoising Diffusion Probabilistic Model","Energy-Based Model","Diffusion Policy","Flow Matching","Diffusion / Flow Samplers"]},{"id":"denoising-steps","category":"model","sec":4,"tier":2,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["denoising-diffusion-probabilistic-model","denoising-diffusion-implicit-model","flow-matching","consistency-model","inference-latency","one-step-generation"],"name":"Denoising Steps","alt":"去噪步数","abbr":"NFE","aliases":["Number of Function Evaluations","NFE","Sampling Steps","Inference Steps"],"one_liner":"How many times a diffusion or flow-matching model calls its network to generate one result; this sets inference speed.","explanation":"Diffusion and flow-matching models don't generate a sample in one shot; they start from noise and call the network repeatedly, each call correcting the result a little. The number of network calls is the number of denoising steps, often written precisely as NFE, the number of function evaluations. More steps usually give a more refined result, but time cost grows roughly proportionally, which matters for robots since a policy has to produce actions in real time inside a closed control loop. Values vary widely by method: the original DDPM used 1,000 steps; Diffusion Policy trains with 100 steps but drops to 10 at inference using DDIM; π0 integrates flow matching over 10 steps; GR00T N1 uses just 4. Consistency models and mean-flow methods push toward generating in 1 or 2 steps. Note that the number of diffusion steps used in training and the number of sampling steps used at inference don't have to match — inference steps are generally adjustable at deployment time.","example":"When π0 predicts a 50-step-long action chunk, it integrates 10 steps, each of size δ=0.1, starting from Gaussian noise, meaning it calls the action expert 10 times to produce one action segment.","related":["Denoising Diffusion Probabilistic Model","Denoising Diffusion Implicit Model","Flow Matching","Consistency Model","Inference Latency","One-step Generation"]},{"id":"denoising-diffusion-implicit-model","category":"model","sec":4,"tier":3,"sources":[{"title":"Denoising Diffusion Implicit Models (arXiv:2010.02502)","url":"https://arxiv.org/abs/2010.02502"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv HTML)","url":"https://arxiv.org/html/2303.04137v5"}],"as_of":"","related_ids":["denoising-diffusion-probabilistic-model","diffusion-model","denoising-steps","diffusion-flow-samplers","diffusion-policy","noise-schedule"],"name":"Denoising Diffusion Implicit Model","alt":"去噪扩散隐式模型","abbr":"DDIM","aliases":["DDIM","DDIM Sampler"],"one_liner":"A diffusion speedup method that keeps DDPM's training but turns sampling into a deterministic process that can skip steps.","explanation":"Denoising Diffusion Implicit Models were proposed in 2020 by Stanford's Jiaming Song, Chenlin Meng, and Stefano Ermon, and published at ICLR 2021. A DDPM (Denoising Diffusion Probabilistic Model) generates one sample by walking a Markov chain of hundreds to thousands of denoising steps, which is slow. DDIM constructs a family of non-Markovian diffusion processes with the same training objective as DDPM, so an already-trained DDPM can switch to DDIM sampling with no retraining: it can skip steps, running only a dozen to a few dozen, which the paper reports is 10 to 50 times faster in wall-clock time; and setting the stochastic term to zero makes sampling deterministic, so the same starting noise always gives the same result. It became the starting point for the many fast diffusion samplers that followed, and Diffusion Policy commonly relies on it to cut steps for real-time control on real robots.","example":"Diffusion Policy's real-robot experiments train with 100 steps but run inference with 10 DDIM steps, giving about 0.1 seconds of inference latency per call on an RTX 3080.","related":["Denoising Diffusion Probabilistic Model","Diffusion Model","Denoising Steps","Diffusion / Flow Samplers","Diffusion Policy","Noise Schedule"]},{"id":"diffusion-flow-samplers","category":"model","sec":4,"tier":3,"sources":[{"title":"Score-Based Generative Modeling through Stochastic Differential Equations (arXiv:2011.13456)","url":"https://arxiv.org/abs/2011.13456"},{"title":"DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps (arXiv:2206.00927)","url":"https://arxiv.org/abs/2206.00927"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv HTML)","url":"https://arxiv.org/html/2410.24164v1"}],"as_of":"","related_ids":["diffusion-model","flow-matching","denoising-diffusion-implicit-model","denoising-steps","velocity-field","one-step-generation"],"name":"Diffusion / Flow Samplers","alt":"扩散 / 流采样器（ODE / SDE 求解器）","abbr":"","aliases":["ODE / SDE Solvers","Euler Sampler","DPM-Solver"],"one_liner":"The numerical algorithm that integrates a trained diffusion or flow model from random noise into a sample, step by step.","explanation":"What a diffusion or flow-matching model learns is 'which direction to move at each noise level'; actually generating a sample requires a numerical method to integrate that process forward, which is what a sampler does. In 2020, Song and colleagues wrote the diffusion process as a stochastic differential equation (SDE) and also derived a deterministic ordinary differential equation with the same marginal distributions, the probability-flow ODE: solving the SDE adds random noise at each step, while solving the ODE gives a deterministic result. Common samplers range from the simplest first-order Euler method and DDIM to specialized high-order solvers such as DPM-Solver, proposed in 2022 by Jun Zhu's team at Tsinghua, which can produce high-quality samples in roughly 10 to 20 network calls. Fewer steps means faster but less accurate, and robot policies have to trade off control frequency against action quality.","example":"π0's flow-matching action expert integrates with the forward Euler method over 10 steps (step size 0.1) at inference, turning a stretch of random noise into an action chunk.","related":["Diffusion Model","Flow Matching","Denoising Diffusion Implicit Model","Denoising Steps","Velocity Field","One-step Generation"]},{"id":"consistency-model","category":"model","sec":4,"tier":3,"sources":[{"title":"Consistency Models (arXiv 2303.01469)","url":"https://arxiv.org/abs/2303.01469"},{"title":"Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation (arXiv 2405.07503)","url":"https://arxiv.org/abs/2405.07503"}],"as_of":"","related_ids":["diffusion-model","diffusion-policy","one-step-generation","knowledge-distillation","meanflow","conrft"],"name":"Consistency Model","alt":"一致性模型","abbr":"","aliases":["Consistency Distillation"],"one_liner":"A generative model that maps noise back to data in one step, used to compress diffusion's many sampling steps into one or two.","explanation":"Consistency models were proposed by OpenAI's Yang Song and colleagues in 2023 (ICML 2023). Generating from a diffusion model means walking step by step along a denoising trajectory, which often takes tens to a hundred steps and is slow. A consistency model instead trains a function so that noisy samples at any point along the same trajectory all map to the same endpoint (the clean data) — a property called 'self-consistency' — so a sample can be produced from pure noise in a single step, or refined further with a few more steps for higher quality. There are two ways to train it: distilling from an already-trained diffusion model (consistency distillation), or training from scratch directly. In robotics, Prasad, Bohg, and colleagues' 2024 Consistency Policy distills Diffusion Policy into a consistency policy, running roughly an order of magnitude faster than even the fastest alternative speedup methods at comparable success rate, which suits robots with limited compute; work such as ConRFT also uses a consistency policy for reinforcement fine-tuning of VLAs.","example":"Consistency Policy was run on a laptop GPU across 6 simulated tasks and 3 real-robot tasks, at roughly 10x the speed of other acceleration methods and a success rate close to the original Diffusion Policy.","related":["Diffusion Model","Diffusion Policy","One-step Generation","Knowledge Distillation","MeanFlow","ConRFT"]},{"id":"one-step-generation","category":"model","sec":4,"tier":3,"sources":[{"title":"Consistency Models (arXiv:2303.01469)","url":"https://arxiv.org/abs/2303.01469"},{"title":"Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation (arXiv:2405.07503)","url":"https://arxiv.org/abs/2405.07503"},{"title":"Mean Flows for One-step Generative Modeling (arXiv:2505.13447)","url":"https://arxiv.org/abs/2505.13447"}],"as_of":"","related_ids":["denoising-steps","consistency-model","meanflow","rectified-flow","consistency-policy","inference-latency"],"name":"One-step Generation","alt":"单步生成","abbr":"","aliases":["1-NFE Generation","One-step Sampling"],"one_liner":"Producing a result directly from noise with just one network forward pass, solving diffusion models' slow sampling.","explanation":"One-step generation means a generative model produces a sample with just a single network forward computation (1 NFE, or number of function evaluations = 1). GANs, VAEs, and normalizing flows are naturally one-step generators; diffusion models and flow matching, by contrast, need tens to thousands of iterative steps starting from noise, giving high quality but at a slow pace. For robots, every extra step in an action head adds control latency, so compressing a diffusion or flow-based policy down to one step is an important direction. There are three common approaches: distillation, training a one-step student from a multi-step teacher, as in consistency distillation; changing the training objective so the model learns to 'jump straight to the endpoint' in one step, as in consistency models and MeanFlow; and straightening the generation path, as in rectified flow. The usual cost is a small quality drop or a more complicated training setup.","example":"Consistency Policy distills Diffusion Policy into a consistency model, running an order of magnitude faster than the fastest existing methods and fitting on a laptop GPU; MP1 uses MeanFlow to generate an action in one step, at about 6.8 ms of inference time.","related":["Denoising Steps","Consistency Model","MeanFlow","Rectified Flow","Consistency Policy","Inference Latency"]},{"id":"meanflow","category":"model","sec":4,"tier":3,"sources":[{"title":"Mean Flows for One-step Generative Modeling (arXiv:2505.13447)","url":"https://arxiv.org/abs/2505.13447"},{"title":"MP1: MeanFlow Tames Policy Learning in 1-step for Robotic Manipulation (arXiv:2507.10543)","url":"https://arxiv.org/abs/2507.10543"}],"as_of":"2025-07","related_ids":["flow-matching","velocity-field","one-step-generation","consistency-model","rectified-flow","denoising-steps"],"name":"MeanFlow","alt":"平均流","abbr":"","aliases":["Mean Flows"],"one_liner":"A generative method that learns the 'average velocity' over a time interval, letting it produce a sample from noise in a single step.","explanation":"MeanFlow was proposed in May 2025 by Zhengyang Geng, J. Zico Kolter, Kaiming He, and colleagues as a one-step generation method. Flow matching trains a network to predict the 'instantaneous velocity' at a given moment, and generation requires integrating along that velocity field over many small steps, which is slow at inference. MeanFlow instead predicts the 'average velocity' over a time interval — the total displacement across that interval divided by its length — and derives an identity relating average velocity to instantaneous velocity to use as the training objective, so it needs no pretrained teacher model or distillation and can sample in one step even trained from scratch. The paper reports an FID (a metric for how close generated images are to real ones, lower is better) of 3.43 for one-step generation on ImageNet 256×256. In robotics, work such as MP1 has already used it as an action-generation head to cut a policy's inference latency.","example":"MP1 uses MeanFlow on a point-cloud-based robot-arm policy, generating a whole action chunk in a single network forward pass; the paper reports about 6.8 ms of inference time, roughly 19x faster than the 3D diffusion policy DP3, with a 10.2-point higher average success rate.","related":["Flow Matching","Velocity Field","One-step Generation","Consistency Model","Rectified Flow","Denoising Steps"]},{"id":"discrete-diffusion","category":"model","sec":4,"tier":3,"sources":[{"title":"Structured Denoising Diffusion Models in Discrete State-Spaces (D3PM, arXiv:2107.03006)","url":"https://arxiv.org/abs/2107.03006"},{"title":"Simple and Effective Masked Diffusion Language Models (MDLM, arXiv:2406.07524)","url":"https://arxiv.org/abs/2406.07524"},{"title":"Discrete Diffusion VLA (arXiv:2508.20072)","url":"https://arxiv.org/abs/2508.20072"}],"as_of":"","related_ids":["diffusion-model","diffusion-language-model","parallel-decoding","discrete-diffusion-vla","denoising-diffusion-probabilistic-model","action-binning"],"name":"Discrete Diffusion","alt":"离散扩散","abbr":"","aliases":["Masked Diffusion"],"one_liner":"A diffusion model that adds and removes noise on discrete symbols, like text tokens, instead of continuous values.","explanation":"An ordinary diffusion model adds Gaussian noise to continuous data — pixels, action values — but data like text or discrete action tokens can't have Gaussian noise added to it directly. Discrete diffusion instead adds noise through 'state transitions': at each step, a transition matrix randomly swaps a token for a different value, or for a special [MASK] symbol. Google's Austin and colleagues systematically formalized this framework in D3PM in 2021, and showed that using an 'absorbing state' for noising (once a token becomes MASK, it stays MASK) connects it closely to masked language models and autoregressive models. Masked diffusion has since become the dominant approach — NeurIPS 2024's MDLM, for instance, simplifies the training objective down to a mixture of masked-language-modeling losses. It's the theoretical foundation behind diffusion language models and discrete-diffusion-style VLAs.","example":"Discrete Diffusion VLA discretizes an action chunk into tokens, masks all of them, then progressively reveals the easiest ones first based on confidence, re-masking and recomputing the uncertain positions, and reports a 96.4% average success rate on LIBERO.","related":["Diffusion Model","Diffusion Language Model","Parallel Decoding","Discrete Diffusion VLA","Denoising Diffusion Probabilistic Model","Action Binning"]},{"id":"diffusion-language-model","category":"model","sec":4,"tier":3,"sources":[{"title":"Large Language Diffusion Models (LLaDA, arXiv:2502.09992)","url":"https://arxiv.org/abs/2502.09992"},{"title":"Gemini Diffusion - Google DeepMind","url":"https://deepmind.google/models/gemini-diffusion/"},{"title":"Mercury: Ultra-Fast Language Models Based on Diffusion (arXiv:2506.17298)","url":"https://arxiv.org/abs/2506.17298"}],"as_of":"2026-09","related_ids":["discrete-diffusion","large-language-model","parallel-decoding","autoregressive-decoding","diffusion-model","discrete-diffusion-vla"],"name":"Diffusion Language Model","alt":"扩散语言模型","abbr":"dLLM","aliases":["dLLM","Diffusion LLM","Masked Diffusion Language Model"],"one_liner":"A language model that denoises a whole span of text in parallel using diffusion, instead of generating it word by word.","explanation":"A diffusion language model applies the diffusion idea to text: generation starts from a span of tokens that are all masked or noised, and predicts them together over several steps, gradually filling in content — each step can fill in several tokens at once, and can also revise earlier guesses. Representative work includes LLaDA (2025, from a Renmin University-led team), whose 8B version matches LLaMA3 8B on in-context learning; Inception Labs' commercial model Mercury; and Google DeepMind's experimental Gemini Diffusion, which the company reports sampling at about 1,479 tokens per second. Its appeal is fast parallel decoding and the ability to use context in both directions at once. Embodied AI has picked up the same idea for decoding actions in parallel too, as in Discrete Diffusion VLA.","example":"When LLaDA answers a question, it first generates a whole span of masked tokens, predicts all the masked positions at each step, keeps the ones it's confident about, re-masks the uncertain ones, and repeats over several rounds until everything is filled in.","related":["Discrete Diffusion","Large Language Model","Parallel Decoding","Autoregressive Decoding","Diffusion Model","Discrete Diffusion VLA"]},{"id":"action-head","category":"model","sec":5,"tier":1,"sources":[{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/abs/2405.12213"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"}],"as_of":"","related_ids":["backbone-network","diffusion-action-head","action-expert","continuous-action-regression","octo","action-multimodality"],"name":"Action Head","alt":"动作头","abbr":"","aliases":["Action Decoder","Policy Head"],"one_liner":"The output module attached after a backbone that turns extracted features into concrete robot actions.","explanation":"The action head follows the same “backbone plus head” split used in vision: the backbone extracts features from images, language, and proprioceptive state, and the head turns those features into whatever output the task needs, which for a robot policy means joint angles, an end-effector pose, or gripper open/close. Three kinds of action heads are common: an MLP head that regresses continuous values directly, trained with mean squared error, which tends to average multiple valid ways of doing something into one; a head that discretizes actions and predicts them as classification over tokens; and a diffusion or flow-matching head, which can express a multimodal action distribution. Octo attaches a lightweight diffusion action head after its Transformer output, and when fine-tuning to a new robot it can simply swap in a new action head to match a different action space; GR00T N1 uses a diffusion Transformer as its action module.","example":"Octo is pretrained with a diffusion action head predicting a segment of future action; when transferred to a new arm with a different action dimensionality, the pretrained Transformer body is kept, a new action head matching the new action space is attached, and the whole model is fine-tuned together on about 100 demonstrations.","related":["Backbone Network","Diffusion Action Head","Action Expert","Continuous Action Regression","Octo","Action Multimodality"]},{"id":"action-representation","category":"model","sec":5,"tier":2,"sources":[{"title":"Universal Manipulation Interface (UMI, arXiv:2402.10329)","url":"https://arxiv.org/abs/2402.10329"},{"title":"On the Continuity of Rotation Representations in Neural Networks (arXiv:1812.07035)","url":"https://arxiv.org/abs/1812.07035"},{"title":"A Survey on Vision-Language-Action Models: An Action Tokenization Perspective (arXiv:2507.01925)","url":"https://arxiv.org/abs/2507.01925"}],"as_of":"","related_ids":["action-space","delta-action-vs-absolute-action","6d-rotation-representation","action-tokenizer","latent-action","end-effector-pose"],"name":"Action Representation","alt":"动作表示","abbr":"","aliases":["Action Parameterization"],"one_liner":"The quantity, reference frame, and format used to describe what action a robot should take.","explanation":"Action representation is how a policy's output action gets written down, one of the first things to settle when building a robot-learning system. It involves several choices: what quantity to control (joint angle, end-effector pose, or velocity or torque instead); what it's relative to (a fixed global frame, an increment relative to the last step, or a whole segment relative to the current pose); how to write rotation (Euler angles, quaternions, or the 6D representation, which is more continuous and easier for a neural network to learn); and the output format (continuous values, discrete tokens, or a more abstract intermediate representation such as a latent action or trajectory points). Swapping one action representation for another, with the same model, can change success rate a great deal, so papers usually spell this choice out explicitly and compare alternatives.","example":"The UMI paper compares these on a cup-pouring task: a whole trajectory relative to the current end-effector pose succeeds 20 out of 20 times, step-by-step increments, whose errors accumulate, succeed 16 out of 20, and a global absolute-coordinate representation succeeds only 5 out of 20.","related":["Action Space","Delta (Relative) Action vs. Absolute Action","6D Rotation Representation","Action Tokenizer","Latent Action","End-Effector Pose"]},{"id":"delta-action-vs-absolute-action","category":"model","sec":5,"tier":2,"sources":[{"title":"Universal Manipulation Interface (UMI, arXiv 2402.10329)","url":"https://arxiv.org/html/2402.10329"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["action-space","action-representation","end-effector-pose","coordinate-frame","universal-manipulation-interface","action-chunking"],"name":"Delta (Relative) Action vs. Absolute Action","alt":"增量动作 / 绝对动作","abbr":"","aliases":["Relative Action","Delta Action","Absolute Action","Relative Trajectory"],"one_liner":"Writing an action as “move by this much” (delta) versus “go to this coordinate” (absolute).","explanation":"This is one of the basic choices in action representation. An absolute action gives a target value directly, such as the pose the end-effector, the gripper or tool at the very tip of the arm, should reach, or a target angle for each joint. A delta action instead gives only the change relative to the current state, such as “move 1 centimeter further along x.” Delta actions don't depend on precise calibration between the robot's base and the world frame, and their value distribution is more consistent across different scenes, but errors accumulate as they add up step by step. Absolute actions don't accumulate error, but they need the coordinate frame to be well calibrated. There's a middle ground too: UMI uses a “relative trajectory,” writing a whole action segment as a transform relative to the end-effector pose at the start of that segment. The Diffusion Policy paper separately found that, for its test tasks, position control that outputs a target position beat velocity control that outputs a target velocity. Which one to use depends on how the dataset defines it, and training and deployment must match, or the robot will drift off course.","example":"If the arm's end effector is currently at x=0.40 meters and needs to move to x=0.42 meters, an absolute action writes 0.42, and a delta action writes +0.02. In one UMI experiment, an absolute-action baseline succeeded only 25% of the time because SLAM coordinates and the robot base's coordinates were hard to align, while the relative-trajectory version succeeded 100% of the time.","related":["Action Space","Action Representation","End-Effector Pose","Coordinate Frame","Universal Manipulation Interface","Action Chunking"]},{"id":"continuous-action-regression","category":"model","sec":5,"tier":2,"sources":[{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)","url":"https://arxiv.org/abs/2502.19645"},{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/html/2405.12213"}],"as_of":"2025-02","related_ids":["action-head","action-multimodality","l1-loss","mean-squared-error","diffusion-action-head","action-binning"],"name":"Continuous Action Regression","alt":"连续动作回归","abbr":"","aliases":["Regression Head","L1 Regression","MSE Regression"],"one_liner":"Having a network output continuous action values directly, trained with L1 or mean-squared error against demonstrations.","explanation":"Continuous action regression is the most direct way for a robot policy to output actions: the network ends in a regression head, usually a few fully connected layers, that directly outputs real numbers such as joint angles or end-effector displacement, trained with mean squared error or L1 loss, the squared or absolute difference between the prediction and the demonstrated action, to match human demonstrations. It is simple to implement, and inference needs just one forward pass, so it's fast. Its weakness shows up when a demonstration set contains several valid ways to do the same thing in the same scene, called action multimodality: regression tends to learn the average of those ways, which can end up being none of them; the Octo paper observed that an MSE regression head produces hesitant, indecisive robot motion. Common alternatives are discretizing actions into tokens, or generating them with diffusion or flow matching. Still, 2025's OpenVLA-OFT paired L1 regression with action chunking and parallel decoding to raise LIBERO's average success rate across four task suites from 76.5% to 97.1%, showing regression is good enough for plenty of tasks.","example":"ACT (Action Chunking with Transformers) trains by directly regressing a future sequence of joint angles with an L1 loss; OpenVLA-OFT replaces OpenVLA's discrete-token output with an L1 regression head, computing a whole action segment in one forward pass and increasing action-generation throughput by about 26x.","related":["Action Head","Action Multimodality","L1 Loss","Mean Squared Error","Diffusion Action Head","Action Binning"]},{"id":"action-chunking","category":"model","sec":5,"tier":1,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"}],"as_of":"","related_ids":["action-horizon","temporal-ensembling","real-time-chunking","compounding-error","action-chunking-with-transformers","diffusion-policy"],"name":"Action Chunking","alt":"动作分块","abbr":"","aliases":["Action Chunk","Chunk","Action Sequence Prediction"],"one_liner":"A policy predicts a short sequence of upcoming actions at once, instead of outputting just one action per step.","explanation":"Action chunking means a policy's inference outputs a sequence of the next k steps of action at once, a “chunk,” executes some or all of it, then re-observes and predicts again. The term entered robot learning through Tony Zhao, Chelsea Finn, and colleagues' 2023 ACT paper, borrowed from a neuroscience idea about packaging a sequence of movements into one unit for execution: ACT controls a dual-arm robot at 50Hz, predicting the next 100 steps of joint targets each time. It solves two problems: it cuts the number of decisions to 1/k of the original, easing the compounding error that behavior cloning suffers from small per-step mistakes; and it makes it easier to learn timing habits in human demonstrations, such as pauses, that a single-step policy struggles to model. Mainstream models today — Diffusion Policy, π0 (50 steps at a time), GR00T N1 (16 steps at a time) — all output action chunks, usually stitched together across chunks with temporal ensembling or real-time chunking to avoid sudden jumps in motion.","example":"ACT performed six fine-motor tasks on the low-cost dual-arm platform ALOHA, each using only 10 to 20 minutes (about 50 demonstrations) of data: opening a translucent condiment-cup lid succeeded 84% of the time, inserting a battery into a remote 96%, and the hardest task, threading a velcro strap, only 20%. It outputs the next 100 steps of joint targets each time, about 2 seconds at 50Hz.","related":["Action Horizon","Temporal Ensembling","Real-Time Chunking","Compounding Error","Action Chunking with Transformers","Diffusion Policy"]},{"id":"action-horizon","category":"model","sec":5,"tier":2,"sources":[{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv:2303.04137)","url":"https://arxiv.org/html/2303.04137v5"},{"title":"diffusion_policy 训练配置 train_diffusion_unet_image_workspace.yaml (GitHub)","url":"https://github.com/real-stanford/diffusion_policy/blob/main/diffusion_policy/config/train_diffusion_unet_image_workspace.yaml"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["action-chunking","temporal-ensembling","diffusion-policy","asynchronous-inference","closed-loop-control","real-time-chunking"],"name":"Action Horizon","alt":"动作视界","abbr":"","aliases":["Prediction Horizon","Execution Horizon","Observation Horizon"],"one_liner":"How many steps of history a policy looks at, how many steps of action it predicts, and how many it actually executes.","explanation":"Action horizon describes how far a policy looks, in time, in both directions. The Diffusion Policy (2023) paper splits it into three quantities: the observation horizon To, how many recent steps of observation are fed in; the prediction horizon Tp, how many steps of action get generated in one shot; and the execution horizon Ta, how many of those actually get sent to the robot before replanning, called receding-horizon control. After executing Ta steps, the policy re-observes and predicts again. A longer Ta means smoother, more coherent motion and fewer inference calls, but a slower reaction to sudden changes; a shorter Ta is more responsive but can jitter. The paper found 8 steps worked best for most tasks, and the open-source code defaults to To=2, Tp=16, Ta=8. What VLAs call chunk size roughly corresponds to the prediction horizon.","example":"π0 predicts 50 steps of action at once: on a 50Hz robot, it executes only the first 25 steps, 0.5 seconds, before re-running inference with new images, while on a 20Hz UR5e it re-infers every 16 steps.","related":["Action Chunking","Temporal Ensembling","Diffusion Policy","Asynchronous Inference","Closed-loop Control","Real-Time Chunking"]},{"id":"temporal-ensembling","category":"model","sec":5,"tier":2,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv 2506.07339)","url":"https://arxiv.org/abs/2506.07339"}],"as_of":"","related_ids":["action-chunking","action-chunking-with-transformers","real-time-chunking","action-smoothing","compounding-error","action-horizon"],"name":"Temporal Ensembling","alt":"时序集成","abbr":"","aliases":["Temporal Ensemble","Action Ensemble"],"one_liner":"Predicting an action chunk at every step, then averaging the overlapping chunks' predictions for the same moment before executing.","explanation":"Temporal ensembling is a way of executing action chunks introduced in the 2023 ACT paper (Tony Zhao and colleagues). Naive action chunking only looks at a new observation every k steps, so the action can jump abruptly when it switches to a new chunk. Temporal ensembling instead calls the policy at every control step, so at any given moment there are several overlapping chunks' predictions for it; these are combined with exponential weights w_i = exp(-m·i), where i=0 is the earliest prediction, and the weighted average is what actually gets executed — a smaller m lets new observations take effect faster. It adds no training cost, only extra inference compute, and in ACT's ablations it improved success rate by about 3.3 percentage points. The cost is that it requires an inference call at every step, which is hard to afford when a large model has high latency; the later Real-Time Chunking (RTC) method uses it as a comparison baseline.","example":"ACT predicts the next k steps of action at every step; at time t, it has already received several predictions for t from the calls made at t, t-1, t-2, and so on, and averages them before sending the result to the arm, which is why ALOHA's bimanual motion looks smoother.","related":["Action Chunking","Action Chunking with Transformers","Real-Time Chunking","Action Smoothing","Compounding Error","Action Horizon"]},{"id":"keyframe-action-prediction","category":"model","sec":5,"tier":3,"sources":[{"title":"Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation (arXiv:2209.05451)","url":"https://arxiv.org/abs/2209.05451"},{"title":"Q-attention: Enabling Efficient Learning for Vision-based Robotic Manipulation (arXiv:2105.14829)","url":"https://arxiv.org/abs/2105.14829"},{"title":"Coarse-to-Fine Q-attention (C2F-ARM, arXiv:2106.12534)","url":"https://arxiv.org/abs/2106.12534"}],"as_of":"2022-11","related_ids":["peract","rvt-2","3d-diffuser-actor","motion-planning","rlbench","action-representation"],"name":"Keyframe Action Prediction","alt":"关键帧动作预测","abbr":"","aliases":["Next-Best-Pose Prediction","Next Best Action","Keypose Prediction"],"one_liner":"Predicting only a few key end-effector poses for a task, and leaving the path between them to a motion planner.","explanation":"A standard visuomotor policy outputs dozens of continuous actions per second; a keyframe method instead compresses a demonstration down to a handful of key end-effector poses (such as before grasping, at the moment of gripping, before releasing), and the policy only learns 'where is the next key pose,' with a motion planner generating the path between two poses. Imperial College's Stephen James and colleagues introduced keyframe discovery in 2021's Q-attention and C2F-ARM, and 2022's PerAct follows the same approach: simple rules, such as joint velocity being near zero, automatically pick out keyframes, typically leaving just 2 to 17 per RLBench demonstration; position is then discretized into voxels and rotation into angle bins, turning prediction into a classification problem. PerAct's ablation shows that choosing frames randomly or at even intervals drops performance to zero. This setup learns fast from few examples, and is also used by 3D manipulation policies like RVT and 3D Diffuser Actor; the cost is dependence on a planner, which struggles with dynamic tasks that need continuous adjustment of force and speed.","example":"To have an arm open a drawer, PerAct only has to predict a few key poses in sequence: a pre-grasp pose in front of the handle, gripping the handle, and pulling out to the target position, with a motion planner filling in the path between each pair.","related":["PerAct","RVT-2","3D Diffuser Actor","Motion Planning","RLBench","Action Representation"]},{"id":"action-binning","category":"model","sec":5,"tier":2,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv:2307.15818)","url":"https://arxiv.org/abs/2307.15818"},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)","url":"https://arxiv.org/abs/2501.09747"}],"as_of":"","related_ids":["action-tokenizer","action-representation","discrete-cosine-transform","rt-2","openvla","pi0-fast"],"name":"Action Binning","alt":"分箱离散化","abbr":"","aliases":["Action Discretization","Bucketing"],"one_liner":"Dividing each continuous action dimension into equal bins and using the bin number as a discrete token.","explanation":"Action binning is the simplest way to turn a continuous robot action into a discrete symbol: for each action dimension, such as end-effector x-displacement or gripper open/close, define a value range, divide it into equal bins, and use the bin a value falls into to represent it. RT-2 splits each dimension into 256 bins and maps each bin number onto a token in the language model's vocabulary, so actions can be predicted the same way text is. OpenVLA keeps 256 bins too, but uses the 1st and 99th percentile of the training data instead of the true minimum and maximum to set the range, so a few outliers don't stretch the bins and hurt precision. The downside is that every timestep and every dimension needs its own token, and at high control frequency, neighboring tokens are highly correlated, giving the model a lot of redundant, hard-to-learn tokens, which is exactly what compressive action tokenizers such as FAST were built to fix.","example":"If x-axis displacement ranges from −2 to 2 centimeters and is split into 256 bins, each bin is about 0.016 centimeters wide, so 0.5 centimeters would fall into bin number 160.","related":["Action Tokenizer","Action Representation","Discrete Cosine Transform","RT-2","OpenVLA","π0-FAST"]},{"id":"action-tokenizer","category":"model","sec":5,"tier":2,"sources":[{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)","url":"https://arxiv.org/html/2501.09747"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"}],"as_of":"2025-01","related_ids":["action-binning","discrete-cosine-transform","byte-pair-encoding","vector-quantization","pi0-fast","action-representation"],"name":"Action Tokenizer","alt":"动作分词器","abbr":"","aliases":["Action Tokenization","Action Token"],"one_liner":"A module that encodes continuous actions into discrete tokens and decodes them back into actions at inference time.","explanation":"An action tokenizer converts a continuous action, usually a whole action chunk, into discrete tokens, and decodes them back into an executable action at inference time, letting a VLA output actions the same way it predicts the next token. The simplest version is per-dimension binning, as in RT-2 and OpenVLA's 256 bins per dimension, but with high-frequency data, actions at neighboring timesteps are nearly identical, so binning produces a lot of redundant tokens the model struggles to learn from. FAST, proposed by Physical Intelligence and others in January 2025, instead applies a discrete cosine transform, DCT, converting a signal into the frequency domain, to each action dimension, quantizes it, and compresses it with byte-pair encoding (BPE), cutting π0's training time to about a fifth. Another line of work learns a discrete action codebook using vector quantization (VQ).","example":"One second of 50Hz action data for folding a T-shirt needs 700 tokens with per-dimension binning, but only 53 tokens after FAST compression.","related":["Action Binning","Discrete Cosine Transform","Byte-Pair Encoding","Vector Quantization","π0-FAST","Action Representation"]},{"id":"discrete-cosine-transform","category":"model","sec":5,"tier":3,"sources":[{"title":"Discrete cosine transform - Wikipedia","url":"https://en.wikipedia.org/wiki/Discrete_cosine_transform"},{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv:2501.09747)","url":"https://arxiv.org/abs/2501.09747"}],"as_of":"","related_ids":["action-tokenizer","pi0-fast","byte-pair-encoding","action-binning","action-representation","action-chunking"],"name":"Discrete Cosine Transform","alt":"离散余弦变换","abbr":"DCT","aliases":["DCT"],"one_liner":"A transform that breaks a signal into a weighted sum of cosine waves at different frequencies, widely used for compression.","explanation":"The discrete cosine transform was originally conceived by Nasir Ahmed, who formally proposed it in a 1974 paper with Natarajan and Rao. It represents a discrete signal as a weighted sum of cosine waves ranging from low to high frequency; those weights are the DCT coefficients. For smooth signals, most of the energy concentrates in a few low-frequency coefficients, so dropping or coarsely quantizing the small high-frequency ones compresses the signal a lot with little loss — this is what JPEG images and MPEG/H.26x video rely on. Embodied AI uses it to represent actions: under high-frequency control, adjacent time steps' actions are very similar, so naively discretizing each step produces a lot of redundant tokens. Physical Intelligence's 2025 FAST tokenizer applies DCT to an action chunk first, then quantizes and compresses it with byte-pair encoding, letting an autoregressive VLA learn high-frequency, dexterous tasks well.","example":"FAST applies DCT dimension by dimension to an action chunk, quantizes it so most high-frequency coefficients become zero while low-frequency ones are kept, then compresses the result into a handful of tokens with byte-pair encoding; the paper reports that, paired with π0 trained on 10,000 hours of data, this matches the diffusion version's performance while cutting training time up to 5x.","related":["Action Tokenizer","π0-FAST","Byte-Pair Encoding","Action Binning","Action Representation","Action Chunking"]},{"id":"gaussian-policy","category":"model","sec":5,"tier":3,"sources":[{"title":"OpenAI Spinning Up: Key Concepts in RL (Diagonal Gaussian Policies)","url":"https://spinningup.openai.com/en/latest/spinningup/rl_intro.html"},{"title":"leggedrobotics/rsl_rl (GitHub)","url":"https://github.com/leggedrobotics/rsl_rl"},{"title":"Soft Actor-Critic (Haarnoja et al., ICML 2018)","url":"https://arxiv.org/abs/1801.01290"}],"as_of":"","related_ids":["policy","deterministic-vs-stochastic-policy","proximal-policy-optimization","soft-actor-critic","entropy-regularization","continuous-action-regression"],"name":"Gaussian Policy","alt":"高斯策略","abbr":"","aliases":["Diagonal Gaussian Policy"],"one_liner":"A stochastic policy where the network outputs an action's mean and standard deviation, then samples from that normal distribution.","explanation":"A Gaussian policy is the most common form of stochastic policy for continuous-action reinforcement learning. A neural network outputs a mean for each action dimension given the observation; the standard deviation is either also output by the network or kept as a separate learnable parameter independent of the observation, usually stored in log form to keep it positive. The action is sampled as 'mean + standard deviation × standard normal noise,' with each dimension usually assumed independent, which is why it's also called a diagonal Gaussian policy. It has a closed-form log-probability, which makes computing the policy gradient easy, and the size of the standard deviation directly controls how much exploration happens. Mainstream algorithms like PPO and SAC both use it, and reinforcement learning for legged locomotion control is almost always built on this kind of policy. At deployment, the mean alone is usually taken as the action. Because it has only one peak, it can't represent several sharply different valid actions, which is why imitation learning often replaces it with a GMM or Diffusion Policy.","example":"Training quadruped walking in Isaac Lab with rsl_rl's PPO: the policy outputs the means of 12 target joint angles, paired with a set of learnable standard deviations for exploration during sampling; on the real robot, only the mean is used.","related":["Policy","Deterministic vs. Stochastic Policy","Proximal Policy Optimization","Soft Actor-Critic","Entropy Regularization","Continuous Action Regression"]},{"id":"gaussian-mixture-model","category":"model","sec":5,"tier":3,"sources":[{"title":"Mixture model - Wikipedia","url":"https://en.wikipedia.org/wiki/Mixture_model"},{"title":"robomimic documentation: Algorithms (BC_GMM / BC_RNN_GMM)","url":"https://robomimic.github.io/docs/modules/algorithms.html"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"}],"as_of":"","related_ids":["mixture-density-network","action-multimodality","gaussian-policy","gaussian-mixture-regression-task-parameterized-gmm","robomimic","diffusion-policy"],"name":"Gaussian Mixture Model","alt":"高斯混合模型","abbr":"GMM","aliases":["GMM","GMM Policy","GMM Action Head"],"one_liner":"A probability model that describes data as a weighted combination of several Gaussian distributions, which can have multiple peaks.","explanation":"A Gaussian mixture model assumes data comes from one of K Gaussian (normal) distributions, each with its own mean, variance (covariance), and weight, with the weights summing to 1. Parameters are usually estimated iteratively with the expectation-maximization (EM) algorithm: first compute each point's probability of belonging to each component, then update the parameters using those probabilities. It's a classic tool for clustering and density estimation. Robot learning has two common uses for it. One is traditional learning from demonstration, which uses a GMM together with Gaussian mixture regression (GMR) to encode trajectories from a handful of demonstrations. The other is as a policy's action head, where the network directly outputs several sets of means, variances, and weights (this is called a mixture density network), representing several valid actions for the same scene. robomimic's BC-RNN-GMM works this way, and it's also one of the main comparison baselines in the Diffusion Policy paper. The number of components has to be fixed in advance, and expressiveness is limited for high-dimensional actions.","example":"Going around an obstacle on a table, passing on the left or the right are both fine. A policy with a single Gaussian averages the two into a straight path into the obstacle; a GMM policy can give the left and right options their own components and sample one of them.","related":["Mixture Density Network","Action Multimodality","Gaussian Policy","Gaussian Mixture Regression / Task-Parameterized GMM","robomimic","Diffusion Policy"]},{"id":"mixture-density-network","category":"model","sec":5,"tier":3,"sources":[{"title":"Bishop, Mixture Density Networks (Aston University technical report, 1994)","url":"https://publications.aston.ac.uk/id/eprint/373/"},{"title":"What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic, arXiv:2108.03298)","url":"https://arxiv.org/abs/2108.03298"}],"as_of":"","related_ids":["gaussian-mixture-model","action-multimodality","gaussian-policy","continuous-action-regression","diffusion-policy","inverse-kinematics"],"name":"Mixture Density Network","alt":"混合密度网络","abbr":"MDN","aliases":["MDN","GMM Policy Head"],"one_liner":"A neural network that outputs the parameters of a Gaussian mixture distribution, instead of a single value directly.","explanation":"The mixture density network was proposed by Christopher Bishop in a 1994 technical report. An ordinary regression network trained with mean squared error learns the average output for a given input; when the same input has several correct answers, that average is often none of them. An MDN instead has the network output the weights, means, and variances of several Gaussian distributions, combined into a Gaussian mixture model, letting it represent a multimodal conditional distribution. One of Bishop's original demonstrations was robot-arm inverse kinematics, where the same end-effector position can correspond to several different sets of joint angles. In robot imitation learning, it's commonly used as a policy's output head to handle action multimodality, as in robomimic's GMM policy; later generative action heads such as Diffusion Policy outperform it on many tasks, but an MDN is cheap to train and run, so it's still often used as a baseline.","example":"Going around an obstacle on a table, some demonstrations go left and some go right. A policy trained with mean squared error averages the two into 'drive straight into it'; an MDN outputs two Gaussian components, one for going left and one for right, and just picks one at execution time.","related":["Gaussian Mixture Model","Action Multimodality","Gaussian Policy","Continuous Action Regression","Diffusion Policy","Inverse Kinematics (IK)"]},{"id":"energy-based-model","category":"model","sec":5,"tier":3,"sources":[{"title":"Implicit Behavioral Cloning (Florence et al., arXiv 2109.00137)","url":"https://arxiv.org/abs/2109.00137"},{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"Energy-based model - Wikipedia","url":"https://en.wikipedia.org/wiki/Energy-based_model"}],"as_of":"","related_ids":["behavior-cloning","action-multimodality","diffusion-policy","mixture-density-network","infonce-loss","generative-model"],"name":"Energy-Based Model","alt":"能量模型","abbr":"EBM","aliases":["EBM","Implicit Behavioral Cloning","IBC"],"one_liner":"A model that scores input-output pairs instead of predicting the output directly; the lowest-scoring output is the answer.","explanation":"An energy-based model doesn't predict its output directly; instead, it learns a scoring function E(x, y) — the better x and y match, the lower the 'energy' — and inference means searching for the y with the lowest energy. The best-known robotics application is Google's 2021 Implicit Behavioral Cloning (IBC): a network assigns an energy score to an 'observation, candidate action' pair, then derivative-free optimization or Langevin sampling (noisy gradient descent) is used to find the lowest-scoring action. This lets a single observation map to several valid actions and lets the policy represent discontinuities, and the paper reached about 1 mm of precision on contact-rich real-robot tasks. The cost is that training needs sampled negative examples to approximate the normalization constant, and the Diffusion Policy paper notes that IBC training is unstable and hard to tune. Since then, modeling multimodal actions has mostly shifted to Diffusion Policy and flow matching instead.","example":"On the same tabletop image, pushing a block from the left or from the right both complete the task. An energy-based model assigns low energy to both actions and lands on one of them at inference; a mean-squared-error regression model instead outputs their average, which is wrong either way.","related":["Behavior Cloning","Action Multimodality","Diffusion Policy","Mixture Density Network","InfoNCE Loss","Generative Model"]},{"id":"diffusion-action-head","category":"model","sec":5,"tier":2,"sources":[{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/html/2405.12213"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/html/2503.14734"},{"title":"CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action (arXiv 2411.19650)","url":"https://arxiv.org/abs/2411.19650"}],"as_of":"","related_ids":["action-head","action-expert","diffusion-policy","diffusion-transformer","action-multimodality","continuous-action-regression"],"name":"Diffusion Action Head","alt":"扩散动作头","abbr":"","aliases":["Diffusion Head"],"one_liner":"An output module attached after a policy's backbone that generates continuous actions through diffusion denoising.","explanation":"A diffusion action head is a type of action head, the part of a policy network that produces the final action. A backbone, such as a Transformer or VLM, first encodes images, language, and robot state into features; the diffusion head then conditions on those features and, starting from random noise, denoises over multiple steps to generate a segment of continuous action. Compared with direct regression, it can express a multimodal action distribution: when a demonstration set shows both “go around the left” and “go around the right” as valid, it doesn't average them into a middle path that runs into the obstacle; compared with discretizing actions into tokens, it keeps the precision of continuous values. The cost is that inference needs multiple network calls. It scales from small to large: Octo uses a 3-layer MLP, while GR00T N1 and CogACT use a dedicated diffusion Transformer as their action module; π0's flow-matching action expert works on a similar principle and is often discussed alongside it.","example":"Octo's Transformer backbone outputs the embedding of a readout token, which is handed to a 3-layer-MLP diffusion head that denoises over 20 steps to generate a segment of action. The paper's comparison found the robot moved hesitantly with an MSE regression head, and imprecisely, often grasping at empty air, with a discrete action head.","related":["Action Head","Action Expert","Diffusion Policy","Diffusion Transformer","Action Multimodality","Continuous Action Regression"]},{"id":"residual-policy","category":"model","sec":5,"tier":3,"sources":[{"title":"Residual Policy Learning (Silver et al., arXiv 1812.06298)","url":"https://arxiv.org/abs/1812.06298"},{"title":"Residual Reinforcement Learning for Robot Control (Johannink et al., arXiv 1812.03201)","url":"https://arxiv.org/abs/1812.03201"},{"title":"From Imitation to Refinement -- Residual RL for Precise Assembly (ResiP, arXiv 2407.16677)","url":"https://arxiv.org/html/2407.16677"}],"as_of":"2024-07","related_ids":["residual-reinforcement-learning","diffusion-policy","reinforcement-fine-tuning","action-chunking","model-predictive-control","teacher-student-distillation"],"name":"Residual Policy","alt":"残差策略","abbr":"","aliases":["Residual Policy Learning","RPL"],"one_liner":"Learning a correction on top of an existing controller or policy's output, and adding the two together as the final action.","explanation":"A residual policy doesn't learn an entire policy from scratch; instead, it keeps an existing base policy (a hand-designed controller, model predictive control, or an already-trained imitation-learning policy) and trains a separate network that only outputs a correction to it, with the final action equal to the base action plus the residual. MIT's Silver and colleagues proposed residual policy learning in 2018, and around the same time Johannink, Levine, and colleagues validated residual reinforcement learning on real-robot block assembly. The benefit is that the base policy can already do the task roughly right, so the residual only has to fill in the parts that are hard to model, such as friction, contact, and calibration error, which means a small exploration space, good sample efficiency, and more safety. A common recent pattern is to freeze a diffusion policy or VLA and train a lightweight, closed-loop, step-by-step residual with reinforcement learning to boost success rate on tasks like fine assembly.","example":"ResiP (Pulkit Agrawal's group, 2024) freezes a diffusion policy that uses action chunking, trains a step-by-step closed-loop residual policy with PPO to correct it, substantially raises success rate on simulated tasks like FurnitureBench furniture assembly, and then transfers it to the real robot via teacher-student distillation.","related":["Residual Reinforcement Learning","Diffusion Policy","Reinforcement Fine-Tuning (RL Fine-Tuning)","Action Chunking","Model Predictive Control","Teacher-Student Distillation"]},{"id":"end-to-end","category":"model","sec":6,"tier":1,"sources":[{"title":"End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)","url":"https://arxiv.org/abs/1504.00702"}],"as_of":"","related_ids":["visuomotor-policy","vision-language-action-model","hierarchical-architecture","sense-plan-act","imitation-learning","end-to-end-training-of-deep-visuomotor-policies"],"name":"End-to-End","alt":"端到端","abbr":"E2E","aliases":["E2E","End-to-End Model","End-to-End Learning"],"one_liner":"Using one model to go straight from raw sensor input to control commands, with no hand-designed intermediate modules.","explanation":"A traditional robot system is a modular pipeline: a perception module identifies objects and poses, a planning module computes a path, and a control module tracks the trajectory, each designed and tuned separately and passing information through human-defined interfaces such as object coordinates. End-to-end instead hands the whole chain to one neural network, mapping raw input, such as camera images and proprioceptive state, directly to joint positions or torque commands, trained all at once toward a single objective. An early landmark in robotics is Levine, Finn, and colleagues' deep visuomotor policy work (2015 preprint, published 2016), which used a convolutional network of only about 92,000 parameters to map raw camera images directly to motor torques. The benefit is less hand-engineering and no error compounding across modules, with performance able to grow with more data; the cost is needing a lot of data, and errors are harder to trace to a cause. Most VLAs today are described as end-to-end models.","example":"Levine and colleagues had a PR2 robot's policy network read raw camera images directly, plus the robot's own joint angles and other state, and output torque commands for every joint, learning tasks such as screwing on a bottle cap and hanging a coat hanger on a rod with no separate object-detection or pose-estimation module at test time.","related":["Visuomotor Policy","Vision-Language-Action Model","Hierarchical Architecture","Sense-Plan-Act","Imitation Learning","End-to-End Training of Deep Visuomotor Policies"]},{"id":"embodied-foundation-model","category":"model","sec":6,"tier":1,"sources":[{"title":"Foundation Models in Robotics: Applications, Challenges, and the Future (arXiv 2312.07843)","url":"https://arxiv.org/abs/2312.07843"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["foundation-model","vision-language-action-model","generalist-policy","embodied-reasoning-model","world-model","pre-training"],"name":"Embodied Foundation Model","alt":"具身大模型","abbr":"","aliases":["Robot Foundation Model","RFM"],"one_liner":"A large model pretrained on data spanning many robots and tasks, adaptable to many different robots and jobs.","explanation":"Embodied foundation model extends the foundation-model idea into robotics. A foundation model is one pretrained on large-scale, diverse data, usually with self-supervised learning, meaning it builds training targets from the data itself rather than from human labels, and can then adapt to many downstream tasks, GPT being an example. Embodied foundation models likewise pretrain first on data spanning many robots and tasks, then post-train on a small amount of data for a specific task or a specific robot body, rather than training a dedicated policy from scratch for every task. The term has no fixed boundary: narrowly it often means a generalist policy that outputs actions directly, such as π0, GR00T N1, or Gemini Robotics; broadly it also includes embodied reasoning models responsible for understanding and planning, and world models used to generate data or evaluate policies. Firoozi and colleagues' 2023 survey identifies the main obstacles as scarce robot data, missing safety guarantees, and real-time requirements.","example":"π0 was pretrained on more than 10,000 hours of robot data, including self-collected data spanning 7 robot configurations and 68 tasks, and then post-trained on curated data for long-horizon tasks such as folding laundry and assembling a box.","related":["Foundation Model","Vision-Language-Action Model","Generalist Policy","Embodied Reasoning Model","World Model","Pre-training"]},{"id":"large-behavior-model","category":"model","sec":6,"tier":2,"sources":[{"title":"TRI LBM 项目页：A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation","url":"https://toyotaresearchinstitute.github.io/lbm1/"},{"title":"TRI News (2023-09-19): Toyota Research Institute Unveils Breakthrough in Teaching Robots New Behaviors","url":"https://www.tri.global/news/toyota-research-institute-unveils-breakthrough-teaching-robots-new-behaviors"},{"title":"TRI News (2025-08-20): AI-Powered Robot by Boston Dynamics and TRI Takes a Key Step Towards General-Purpose Humanoids","url":"https://tri.global/news/ai-powered-robot-boston-dynamics-and-toyota-research-institute-takes-key-step-towards-general"}],"as_of":"2025-08","related_ids":["toyota-research-institute","diffusion-policy","foundation-model","behavior-foundation-model","vision-language-action-model","boston-dynamics-atlas-2"],"name":"Large Behavior Model","alt":"大行为模型","abbr":"LBM","aliases":["LBM","TRI LBM"],"one_liner":"Toyota Research Institute's term for a general-purpose, multi-task robot policy pretrained on large amounts of demonstration data.","explanation":"Large Behavior Model is Toyota Research Institute's (TRI) name, by analogy with large language models, for a general-purpose manipulation policy pretrained on a large amount of multi-task robot data and used to output actions directly. TRI first raised the idea of building an LBM when it published its diffusion-policy skill-learning results in September 2023; a July 2025 paper then gave a concrete implementation: a diffusion Transformer that takes wrist and scene camera images, proprioceptive state, and a language instruction, and predicts 16 steps of action at once. It was trained on about 1,700 hours of data, roughly 468 hours of which came from TRI's own bimanual teleoperation. The paper found that after pretraining, learning a new task needs 3-5x less data, and performance scales smoothly with pretraining data size. TRI has partnered with Boston Dynamics since October 2024, and in August 2025 the two demonstrated a single LBM controlling a full electric Atlas, continuously walking while carrying parts and sorting them into bins.","example":"In the August 2025 demo, a single LBM controls Atlas's arms and legs together: it walks, crouches, picks up a part, and sorts it into a bin; when someone slides a closed box across the floor into its path, it adjusts and keeps working.","related":["Toyota Research Institute","Diffusion Policy","Foundation Model","Behavior Foundation Model","Vision-Language-Action Model","Boston Dynamics Atlas (Electric)"]},{"id":"vision-language-action-model","category":"model","sec":6,"tier":1,"sources":[{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv:2307.15818)","url":"https://arxiv.org/abs/2307.15818"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/abs/2406.09246"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"2024-10","related_ids":["vision-language-model","rt-2","openvla","pi0","action-tokenizer","action-expert"],"name":"Vision-Language-Action Model","alt":"视觉-语言-动作模型","abbr":"VLA","aliases":["VLA","VLA Model"],"one_liner":"A large model that looks at images, understands language instructions, and directly outputs robot actions.","explanation":"A vision-language-action model takes camera images and a natural-language instruction as input and directly outputs robot control actions, usually built by taking a pretrained vision-language model, a large model that can look at images and answer questions, and training it further on robot data. The name comes from Google DeepMind's RT-2 paper in July 2023: it writes actions as text tokens and trains them together with web image-text data, letting the robot draw on internet knowledge to handle objects and instructions it has never seen. Since then there has been the open-source OpenVLA (7B parameters, trained on 970,000 robot demonstrations), and Physical Intelligence's π0, which generates continuous action chunks with flow matching. The main difference between them is how actions get output: discrete tokens, a diffusion or flow-matching action head, or some mix of the two.","example":"Given a tabletop photo and the instruction “put the eggplant in the pot,” OpenVLA outputs 7 action tokens, which get decoded into the arm end-effector's translation, rotation, and gripper open/close.","related":["Vision-Language Model","RT-2","OpenVLA","π0","Action Tokenizer","Action Expert"]},{"id":"parallel-decoding","category":"model","sec":6,"tier":3,"sources":[{"title":"Non-Autoregressive Neural Machine Translation (arXiv:1711.02281)","url":"https://arxiv.org/abs/1711.02281"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv:2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"","related_ids":["autoregressive-decoding","action-chunking","openvla-oft","causal-attention","discrete-diffusion","speculative-decoding"],"name":"Parallel Decoding","alt":"并行解码","abbr":"","aliases":["Non-autoregressive Decoding"],"one_liner":"Producing an entire sequence in one forward pass, instead of generating it token by token like autoregressive decoding.","explanation":"Parallel decoding is defined in contrast to autoregressive decoding: an autoregressive model only predicts the next token each time, so producing N tokens takes N forward passes; parallel decoding instead has the model predict every position in a single forward pass. It was first studied systematically in machine translation as 'non-autoregressive translation' (Gu et al., 2017), cutting latency by roughly an order of magnitude at the cost of weaker modeling of dependencies between positions and somewhat lower quality. The issue is especially visible in VLAs: OpenVLA discretizes each action dimension into a single token and generates them one at a time, so a 7-dimensional action needs 7 forward passes, and action chunking makes it even slower. OpenVLA-OFT instead feeds in a set of empty action placeholder embeddings and replaces the causal attention mask with bidirectional attention, producing the entire action chunk in a single forward pass. Methods like discrete diffusion also fall under parallel or few-step parallel decoding.","example":"OpenVLA-OFT combines parallel decoding with action chunking (outputting 8 steps of 7-dimensional action at once) on LIBERO, raising action throughput from OpenVLA's 4.2 Hz to 108.8 Hz, about 26x faster.","related":["Autoregressive Decoding","Action Chunking","OpenVLA-OFT","Causal Attention","Discrete Diffusion","Speculative Decoding"]},{"id":"action-expert","category":"model","sec":6,"tier":1,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"π0 论文 HTML 全文（动作专家参数量与结构）","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["vision-language-action-model","action-head","flow-matching","mixture-of-experts","pi0","backbone-network"],"name":"Action Expert","alt":"动作专家","abbr":"","aliases":["Action Expert Module"],"one_liner":"A separate set of parameters inside a VLA dedicated to producing continuous actions, apart from the vision-and-language backbone.","explanation":"The term “action expert” became popular through Physical Intelligence's 2024 π0. π0 uses a roughly 3-billion-parameter PaliGemma vision-language model as its backbone, plus a separate, newly initialized Transformer of about 300 million parameters that handles robot state and action tokens; that second set of weights is the action expert. The two share the same attention layers but keep separate weights, similar to a mixture-of-experts with just two experts, and information flows only one way: action tokens can attend to the image and text tokens, but not the reverse, so the backbone doesn't get pulled off its original pretraining. The action expert generates continuous action chunks with flow matching; the paper found that giving state and action tokens their own dedicated weights works better than sharing the backbone's weights. π0.5 and SmolVLA keep this same name, and GR00T N1's action module is a similar design.","example":"At inference, π0's PaliGemma backbone first encodes the camera images and language instruction, and the action expert then integrates 10 steps of flow matching to output a continuous action chunk 50 steps long, with a control frequency of up to 50Hz.","related":["Vision-Language-Action Model","Action Head","Flow Matching","Mixture of Experts","π0","Backbone Network"]},{"id":"mixture-of-transformers","category":"model","sec":6,"tier":3,"sources":[{"title":"Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models (arXiv:2411.04996)","url":"https://arxiv.org/abs/2411.04996"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"Emerging Properties in Unified Multimodal Pretraining (BAGEL, arXiv:2505.14683)","url":"https://arxiv.org/abs/2505.14683"}],"as_of":"2025-05","related_ids":["mixture-of-experts","action-expert","pi0","bagel","unified-multimodal-model","self-attention"],"name":"Mixture-of-Transformers","alt":"混合 Transformer 架构","abbr":"MoT","aliases":["MoT","Mixture-of-Transformer-Experts"],"one_liner":"Giving each modality its own set of Transformer parameters, while every layer's self-attention still lets them all see each other.","explanation":"Mixture-of-Transformers was proposed in November 2024 by Stanford's Weixin Liang and researchers at Meta. It splits the feedforward network, attention projection matrices, and layer normalization parameters by modality: text tokens use one set of weights, image tokens use another, but each layer's self-attention is still computed over the whole sequence, so the modalities interact as usual. Unlike a Mixture of Experts (MoE, where a router picks experts per token), MoT assigns a fixed division of labor by modality and needs no routing to be learned. The paper reports that at 7B scale, on a combined text-and-image generation setup, it matches a dense model's performance using only about 55.8% of the compute. In embodied AI, π0's 'VLM backbone + action expert' design is a similar approach (the paper describes it as resembling an MoE with two members), while ByteDance's BAGEL adopts the MoT structure directly.","example":"π0's roughly 3-billion-parameter backbone comes from PaliGemma, plus a separate action expert of about 300 million parameters: image and text tokens use the backbone's weights, and robot state and noisy action tokens use the action expert's weights, with the two parts visible to each other in every layer's attention.","related":["Mixture of Experts","Action Expert","π0","BAGEL","Unified Multimodal Model","Self-Attention"]},{"id":"hybrid-autoregressive-diffusion-architecture","category":"model","sec":6,"tier":3,"sources":[{"title":"Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model (arXiv:2408.11039)","url":"https://arxiv.org/abs/2408.11039"},{"title":"Autoregressive Image Generation without Vector Quantization (MAR, arXiv:2406.11838)","url":"https://arxiv.org/abs/2406.11838"},{"title":"HybridVLA: Collaborative Diffusion and Autoregression in a Unified VLA Model (arXiv:2503.10631)","url":"https://arxiv.org/abs/2503.10631"}],"as_of":"2025-06","related_ids":["hybridvla","diffusion-model","autoregressive-decoding","unified-multimodal-model","diffusion-action-head","action-tokenizer"],"name":"Hybrid Autoregressive-Diffusion Architecture","alt":"自回归-扩散混合架构","abbr":"","aliases":["AR + Diffusion","Hybrid AR-Diffusion"],"one_liner":"An architecture where discrete content is generated autoregressively and continuous content is generated by diffusion, in the same model.","explanation":"Autoregressive means predicting the next token step by step in sequence, as large language models do, which suits discrete content like text and reasoning; a diffusion model starts from noise and denoises step by step, which suits continuous, high-dimensional data like images and motion. A hybrid architecture puts both inside the same network. Transfusion, proposed by Meta and others in 2024, uses a single Transformer to process a sequence that interleaves text and images, computing a next-token-prediction loss for text and a diffusion loss for images; Kaiming He and colleagues' MAR instead models each continuous-valued token with a diffusion loss inside an autoregressive framework, eliminating the need for vector quantization. In robotics, 2025's HybridVLA has a single large language model do both diffusion denoising and autoregressive action prediction, then adaptively fuses the two resulting actions. The motivation is that forcing continuous actions into discrete tokens loses precision, while a purely diffusion-based action head has more trouble directly using a language model's reasoning ability.","example":"Inside the same language model, HybridVLA both generates a continuous action through diffusion denoising and predicts discretized action tokens autoregressively, then fuses the two results into the final action at execution time.","related":["HybridVLA","Diffusion Model","Autoregressive Decoding","Unified Multimodal Model","Diffusion Action Head","Action Tokenizer"]},{"id":"unified-action-space","category":"model","sec":6,"tier":2,"sources":[{"title":"RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation (arXiv 2410.07864)","url":"https://arxiv.org/html/2410.07864"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"},{"title":"Universal Actions for Enhanced Embodied Foundation Models (arXiv 2501.10105)","url":"https://arxiv.org/abs/2501.10105"}],"as_of":"","related_ids":["cross-embodiment","action-space","embodiment-specific-head","latent-action","cross-embodiment-data","rdt-1b"],"name":"Unified Action Space","alt":"统一动作空间","abbr":"","aliases":["Universal Action Space"],"one_liner":"A single fixed format that maps different robots' actions by physical meaning, so their data can be trained together.","explanation":"Different robots use different action formats: a single arm might be 7 dimensions, a bimanual robot 14; some give joint angles, others give end-effector pose (the gripper's position and orientation in space). To train one model on data from multiple robots, the actions first need to be aligned into a common format — this is a unified action space. The most direct approach is to define a vector long enough for everything, with each position fixed to a physical quantity and missing dimensions padded with zero: Tsinghua's RDT-1B uses 128 dimensions, with fixed slots for the left arm, right arm, and base; π0 zero-pads every robot's action to 18 dimensions. A different approach learns an abstract latent action space instead, as in UniAct, then adds a small per-robot decoder to turn it back into a concrete command. This is a foundational piece of cross-embodiment training.","example":"Fitting data from a 6-DOF single-arm robot into RDT-1B's 128-dimensional vector, the arm is treated as the right arm: its joint angles fill only the first 6 slots of the right-arm block, while the left-arm and base slots are all set to zero.","related":["Cross-Embodiment","Action Space","Embodiment-specific Head","Latent Action","Cross-Embodiment Data","RDT-1B"]},{"id":"embodiment-specific-head","category":"model","sec":6,"tier":3,"sources":[{"title":"Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (HPT, arXiv:2409.20537)","url":"https://arxiv.org/abs/2409.20537"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv HTML)","url":"https://arxiv.org/html/2503.14734v2"}],"as_of":"","related_ids":["cross-embodiment","action-head","unified-action-space","heterogeneous-pre-trained-transformers","nvidia-isaac-gr00t-n1","embodiment"],"name":"Embodiment-specific Head","alt":"本体专属头","abbr":"","aliases":["Embodiment-specific Decoder"],"one_liner":"In a cross-embodiment model, a separate output layer per robot type that turns shared features into that robot's actions.","explanation":"Different robots have different state and action dimensions: a single arm with a gripper might be 7 dimensions, while a bimanual robot with dexterous hands might have several dozen. Cross-embodiment models usually share most of their parameters, but attach a small network per embodiment at the input and output ends; the output-side piece is the embodiment-specific head, typically an MLP or a small decoder that maps the shared backbone's representation into an action with that embodiment's specific dimensions and meaning. MIT and Meta's Lirui Wang, Kaiming He, and colleagues proposed the Heterogeneous Pre-trained Transformer (HPT) in 2024, using an 'embodiment-specific stem + shared trunk + embodiment-specific head' design; NVIDIA's GR00T N1 similarly gives each embodiment an MLP to encode state and action, with a dedicated action decoder for output. A different approach is a unified action space, which pads every embodiment's action into the same large vector instead.","example":"GR00T N1 attaches an embodiment-specific MLP action decoder after its last DiT block, letting the same backbone control everything from tabletop arms to humanoids with dexterous hands.","related":["Cross-Embodiment","Action Head","Unified Action Space","Heterogeneous Pre-trained Transformers","NVIDIA Isaac GR00T N1","Embodiment"]},{"id":"state-proprioception-encoder","category":"model","sec":6,"tier":2,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"},{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/abs/2405.12213"}],"as_of":"","related_ids":["proprioception","projector-connector","embodiment-specific-head","action-state-normalization","causal-confusion","history-encoder"],"name":"State / Proprioception Encoder","alt":"状态编码器","abbr":"","aliases":["Proprioception Encoder","State Encoder","State Projector"],"one_liner":"A small network that turns a robot's own state — joint angles, gripper opening — into a vector the model can use.","explanation":"A state encoder processes a robot's proprioception — information about its own state that doesn't need a camera, such as joint angles, end-effector pose, and gripper opening. These are numeric vectors of a few to a few dozen dimensions, and they first need to be mapped into the model's embedding dimension before they can be fed into a Transformer alongside image and text tokens. The implementation is usually simple: π0 and ACT each use a single linear layer, while GR00T N1 gives each embodiment its own MLP to handle robots with different state dimensions. State input lets the policy know where the arm currently is, but it has a side effect too: the Octo team found that adding proprioceptive state often hurt performance, likely because the model over-relies on the strong correlation between state and action — a form of causal confusion.","example":"The bimanual ALOHA robot's state is 14 joint positions across both arms; ACT projects this 14-dimensional vector to 512 dimensions with a single linear layer and feeds it as one token into the Transformer encoder alongside the image features.","related":["Proprioception","Projector / Connector","Embodiment-specific Head","Action / State Normalization","Causal Confusion (Causal Misidentification)","History Encoder"]},{"id":"history-encoder","category":"model","sec":6,"tier":3,"sources":[{"title":"RMA: Rapid Motor Adaptation for Legged Robots (Kumar et al., RSS 2021)","url":"https://arxiv.org/abs/2107.04034"},{"title":"RMA paper full text (ar5iv)","url":"https://ar5iv.labs.arxiv.org/html/2107.04034"}],"as_of":"","related_ids":["rapid-motor-adaptation","privileged-information","teacher-student-distillation","proprioception","system-identification","domain-randomization"],"name":"History Encoder","alt":"历史编码器","abbr":"","aliases":["Adaptation Module","History Observation Encoder"],"one_liner":"A module that compresses a recent window of observations and actions into a vector, letting the policy infer current conditions.","explanation":"A history encoder is the part of a policy network dedicated to processing information from the past several steps, commonly implemented with a 1D convolution, an RNN, or a Transformer. The most typical use is in reinforcement learning for legged robots: real-world parameters like ground friction, payload, and motor condition can't be measured directly, but they leave traces in the recent sequence of states and actions. RMA (Kumar et al., RSS 2021) trains in two stages: first, a base policy is trained in simulation using an 8-dimensional extrinsics vector compressed from privileged information (ground-truth parameters only available in simulation); then, an 'adaptation module' is trained to regress that same vector just from the last 50 steps (0.5 seconds) of state-and-action history. At deployment, the adaptation module runs at 10 Hz while the base policy runs at 100 Hz. This 'teacher first, then student' approach has since been widely reused across legged and humanoid locomotion work.","example":"Suddenly strapping extra weight onto a Unitree A1's back, RMA's adaptation module estimates the new extrinsics from the joint response over the last half second, and the base policy adjusts its gait accordingly in under a second, with no retraining needed.","related":["Rapid Motor Adaptation","Privileged Information","Teacher-Student Distillation","Proprioception","System Identification","Domain Randomization"]},{"id":"point-cloud-encoder","category":"model","sec":6,"tier":3,"sources":[{"title":"PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv:1612.00593)","url":"https://arxiv.org/abs/1612.00593"},{"title":"3D Diffusion Policy (arXiv:2403.03954)","url":"https://arxiv.org/abs/2403.03954"}],"as_of":"","related_ids":["point-cloud","pointnet-pointnet-plus-plus","point-transformer-v3","3d-diffusion-policy","farthest-point-sampling","vision-encoder"],"name":"Point Cloud Encoder","alt":"点云编码器","abbr":"","aliases":["Point Cloud Backbone"],"one_liner":"A module that converts an unordered set of 3D points into feature vectors a neural network can use.","explanation":"A point cloud encoder turns a point cloud from a depth camera or lidar — an unordered set of 3D points with xyz coordinates, and sometimes color — into features. The difficulty is that points have no fixed order and no fixed count, so the convolutions built for image grids don't directly apply. Stanford's 2016 PointNet extracts per-point features with a shared MLP, then aggregates them with an order-independent operation like max pooling, pioneering the direct-on-point-cloud approach; stronger architectures followed, such as PointNet++ and the Point Transformer series. Any policy that uses 3D input in embodied AI depends on this: 3D Diffusion Policy (DP3) first downsamples the point cloud to 512 or 1024 points with farthest point sampling, then uses a lightweight encoder of three MLP layers plus max pooling to get a 64-dimensional feature, and the paper's ablations show this outperforms more complex encoders like PointNet++.","example":"DP3's point cloud encoder deliberately skips the color channel and uses only geometric coordinates, which the paper says generalizes better to changes in an object's appearance.","related":["Point Cloud","PointNet / PointNet++","Point Transformer V3","3D Diffusion Policy","Farthest Point Sampling","Vision Encoder"]},{"id":"3d-vla","category":"model","sec":6,"tier":3,"sources":[{"title":"3D-VLA: A 3D Vision-Language-Action Generative World Model (arXiv 2403.09631)","url":"https://arxiv.org/abs/2403.09631"},{"title":"SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model (arXiv 2501.15830)","url":"https://arxiv.org/abs/2501.15830"},{"title":"SpatialVLA GitHub","url":"https://github.com/SpatialVLA/SpatialVLA"}],"as_of":"2025-01","related_ids":["vision-language-action-model","spatial-reasoning","point-cloud","depth-estimation","3d-diffusion-policy","world-model"],"name":"3D VLA","alt":"3D VLA","abbr":"","aliases":["Spatial VLA","SpatialVLA","3D-VLA"],"one_liner":"A VLA that explicitly feeds spatial information — depth, point clouds, 3D position — into the model.","explanation":"A standard VLA (vision-language-action model, which looks at an image, hears an instruction, and outputs an action directly) mostly takes only 2D images as input, but a robot's actions happen in 3D space, where distance, height, and occlusion are hard to judge accurately from a 2D image alone. 3D VLA is a general term for VLAs that add depth, point clouds, or 3D positional encoding to the input or an intermediate representation. Two papers are representative. 3D-VLA, from a UMass Amherst team in March 2024, adds interaction tokens on top of a 3D large language model and uses a diffusion model to generate a goal image and goal point cloud, 'imagining' the scene after manipulation. SpatialVLA (RSS 2025), by Qu and colleagues in January 2025, is built on PaliGemma2, uses Ego3D position encoding to inject 3D spatial information into visual features, discretizes continuous actions into spatial tokens with an adaptive action grid, and is pretrained on 1.1 million real-robot trajectories from OXE and RH20T. The main goal of this line of work is more stable spatial judgment when the camera viewpoint or the robot changes.","example":"SpatialVLA-4B encodes each image patch's 3D position into its visual token, needs no camera calibration, and can re-partition its action grid to adapt when moved to a new robot.","related":["Vision-Language-Action Model","Spatial Reasoning","Point Cloud","Depth Estimation","3D Diffusion Policy","World Model"]},{"id":"force-aware-vision-language-action-model","category":"model","sec":6,"tier":3,"sources":[{"title":"ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation (arXiv 2505.22159)","url":"https://arxiv.org/abs/2505.22159"},{"title":"ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness (arXiv 2603.15169)","url":"https://arxiv.org/abs/2603.15169"},{"title":"Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning (arXiv 2608.07558)","url":"https://arxiv.org/abs/2608.07558"}],"as_of":"2026-09","related_ids":["vision-language-action-model","six-axis-force-torque-sensor","contact-rich-manipulation","hybrid-force-position-control","vision-tactile-language-action-model","mixture-of-experts"],"name":"Force-aware Vision-Language-Action Model","alt":"力觉 VLA","abbr":"","aliases":["Force-aware VLA"],"one_liner":"A VLA that takes force/torque signals as input alongside vision and language.","explanation":"A force-aware VLA is a vision-language-action model that adds force signals, such as those from a six-axis force/torque sensor (measuring force along three axes and torque around three axes), as an input modality alongside image and language. An ordinary VLA relies only on cameras, but in contact-rich tasks like plugging in a cord, wiping a table, or assembly, whether contact is properly made and how much force is being applied is often invisible to a camera, leading to getting stuck or pressing too hard. The representative work, ForceVLA (Yu et al., NeurIPS 2025), builds on π0 and uses a force-aware mixture-of-experts module (FVLMoE) to fuse a force token during action decoding, reaching 23.2% higher average performance than the π0 baseline across five contact-rich tasks, while simply concatenating force into the input only gave a small improvement. ForceVLA2, from March 2026, adds hybrid force/position control and reports 48% and 35% improvements over π0 and π0.5 respectively.","example":"When plugging in a cord, as the plug touches the edge of the socket the camera image barely changes, but the force sensor reading jumps; ForceVLA uses this signal to correct the pose, reaching up to 80% success on plugging tasks.","related":["Vision-Language-Action Model","Six-Axis Force/Torque Sensor","Contact-rich Manipulation","Hybrid Force/Position Control","Vision-Tactile-Language-Action Model","Mixture of Experts"]},{"id":"vision-tactile-language-action-model","category":"model","sec":6,"tier":3,"sources":[{"title":"VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation (arXiv:2505.09577)","url":"https://arxiv.org/abs/2505.09577"},{"title":"VT-Bridge: Bridging Pretrained Foundation VLAs to VTLAs via Lightweight Residual Adaptation (arXiv:2609.22606)","url":"https://arxiv.org/abs/2609.22606"}],"as_of":"2026-09","related_ids":["vision-language-action-model","vision-based-tactile-sensor","visuo-tactile-fusion","contact-rich-manipulation","peg-in-hole-insertion","force-aware-vision-language-action-model"],"name":"Vision-Tactile-Language-Action Model","alt":"视觉-触觉-语言-动作模型","abbr":"VTLA","aliases":["VTLA","Tactile VLA"],"one_liner":"A robot model that adds touch sensing to VLA, built for contact-heavy tasks that need a physical feel for the world.","explanation":"VTLA refers to models that feed tactile-sensor signals into a robot policy alongside camera images and language instructions to produce actions — an extension of VLA (vision-language-action models). The name comes from a May 2025 paper by researchers at Samsung Research China, the Beijing Academy of Artificial Intelligence, and the Institute of Automation, Chinese Academy of Sciences. They built on a Qwen2-VL 7B backbone, converted readings from GelStereo visual-tactile sensors on the gripper fingertips into image-like inputs, trained on simulated peg-in-hole insertion data, and used direct preference optimization (DPO) to ease the mismatch between token-based classification and continuous control. Touch matters because in contact-rich tasks like inserting parts, twisting caps, or grasping fragile objects, the contact point is often hidden from the camera by the fingers themselves, so vision alone can't tell how much force is being applied or whether something is slipping. VTLA is now used as a general label for this class of model, and work in 2026 still explores turning off-the-shelf VLAs into VTLAs with little extra data.","example":"In the original VTLA paper's real-robot experiments, the model adjusted insertion position and angle step by step using camera images plus fingertip tactile images, reaching a 95% success rate inserting a square peg into a hole with 0.6mm clearance, and 95–100% on peg shapes it had never seen.","related":["Vision-Language-Action Model","Vision-Based Tactile Sensor","Visuo-Tactile Fusion","Contact-rich Manipulation","Peg-in-Hole Insertion","Force-aware Vision-Language-Action Model"]},{"id":"memory-augmented-vla","category":"model","sec":6,"tier":3,"sources":[{"title":"MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation (arXiv:2508.19236)","url":"https://arxiv.org/abs/2508.19236"},{"title":"Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning (MIKASA, arXiv:2502.10550)","url":"https://arxiv.org/abs/2502.10550"}],"as_of":"2026-03","related_ids":["memory-augmented-vla","vision-language-action-model","embodied-memory","long-horizon-task","partially-observable-markov-decision-process","history-encoder"],"name":"Memory-Augmented VLA","alt":"记忆增强 VLA","abbr":"","aliases":[],"one_liner":"A VLA with an added history-memory module, so it can act based on what happened earlier, not just the current frame.","explanation":"Memory-augmented VLA is an umbrella term for a research direction, not one specific model. Most mainstream VLAs (vision-language-action models) only look at the current frame or the last few frames to produce an action, implicitly assuming the current view holds all the information decisions need. But many manipulation tasks depend on history: a button looks the same whether or not it's already been pressed, and an object hidden behind something else isn't visible in the frame at all. Stuffing a long history of frames straight into the model would blow up the token count and inference time. This line of work instead maintains a separate memory: past observations are compressed into features and stored in a memory bank, and at decision time relevant entries are retrieved and fused with the current features before going to the action head. Representative work includes MemoryVLA, proposed in 2025 by Tsinghua and Yuanli Robotics, among others; the MIKASA-Robo benchmark specifically tests this kind of memory ability across 32 tabletop manipulation tasks.","example":"In MemoryVLA's real-robot task 'press three buttons in a specified color order,' a pressed and an unpressed button look identical, so progress can't be judged from the current frame alone; MemoryVLA completes the task using history stored in its memory bank, and the paper reports it beats the strongest baseline by 26 points on long-horizon tasks.","related":["Memory-Augmented VLA","Vision-Language-Action Model","Embodied Memory","Long-horizon Task","Partially Observable Markov Decision Process","History Encoder"]},{"id":"hierarchical-architecture","category":"model","sec":7,"tier":1,"sources":[{"title":"Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan, arXiv:2204.01691)","url":"https://arxiv.org/abs/2204.01691"},{"title":"Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models (arXiv:2502.19417)","url":"https://arxiv.org/html/2502.19417"},{"title":"Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure)","url":"https://www.figure.ai/news/helix"}],"as_of":"2025-02","related_ids":["dual-system-architecture","braincerebellum-architecture","end-to-end","saycan","hi-robot","long-horizon-task"],"name":"Hierarchical Architecture","alt":"分层架构","abbr":"","aliases":["Hierarchical Policy","Hierarchical Scheme","High-Level Planning + Low-Level Execution","Hierarchical VLA"],"one_liner":"Splitting robot decision-making into a high-level planner and a low-level executor, run by two different models.","explanation":"A hierarchical architecture splits “what to do” from “how to do it”: the high level is usually a large language model or vision-language model that understands the instruction and the scene and breaks a long task into steps, running at low frequency; the low level is a skill library, a VLA policy, or a motion controller that turns each step into motor commands, running at high frequency. The two levels can exchange language instructions, keypoints, or a latent vector. An early example is Google's 2022 SayCan, which has a large language model pick the next step from a robot's existing skill library; 2025's Hi Robot, from Physical Intelligence with Stanford and Berkeley, has a vision-language model output short language instructions for π0 to carry out. It is often contrasted with end-to-end approaches, where one model goes directly from input to action, but the two are not mutually exclusive: China's “brain-cerebellum” framing and dual-system architectures are both hierarchical in spirit, and a dual-system model such as Helix, despite having two layers, is trained end-to-end jointly.","example":"Figure's Helix splits into two layers: a 7-billion-parameter vision-language model understands the scene and the instruction at 7 to 9Hz, and an 80-million-parameter low-level policy turns what it outputs into continuous actions at 200Hz.","related":["Dual-System Architecture (System 1 / System 2)","Brain–Cerebellum Architecture","End-to-End","SayCan","Hi Robot","Long-horizon Task"]},{"id":"embodied-reasoning-model","category":"model","sec":7,"tier":1,"sources":[{"title":"Gemini Robotics 1.5 brings AI agents into the physical world (Google DeepMind)","url":"https://deepmind.google/discover/blog/gemini-robotics-15-brings-ai-agents-into-the-physical-world/"},{"title":"Gemini Robotics 模型页 (Google DeepMind)","url":"https://deepmind.google/models/gemini-robotics/"},{"title":"RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration (arXiv 2505.03673)","url":"https://arxiv.org/abs/2505.03673"}],"as_of":"2026-09","related_ids":["embodied-reasoning","dual-system-architecture","hierarchical-architecture","gemini-robotics-er","robobrain","vision-language-action-model"],"name":"Embodied Reasoning Model","alt":"具身推理模型","abbr":"ER","aliases":["ER","Embodied Brain Model","Embodied Brain","ER Model","Brain"],"one_liner":"A multimodal model specialized in understanding the physical world and planning tasks, often called a robot's ‘brain’.","explanation":"An embodied reasoning model is a multimodal large model with strengthened spatial understanding and task planning; it generally does not output joint actions directly, but instead answers questions like where something is, where to grasp it, what step comes next, and whether a step is done, outputting pointing coordinates, trajectories, or a step-by-step plan for an execution layer to carry out. The representative example is Google DeepMind's Gemini Robotics-ER, released in March 2025 (ER stands for embodied reasoning); the newest version listed on the official site is currently ER 2. In China, the Beijing Academy of Artificial Intelligence (BAAI) has its RoboBrain series. This division of labor is commonly called, in China's robotics industry, a “brain-cerebellum” architecture: the “brain” is the reasoning model responsible for understanding and planning, and the “cerebellum” is the lower-level module responsible for motor control and skill execution. Compared with a dual-system architecture, the brain and cerebellum are more often trained separately and connected through an interface, rather than jointly.","example":"In the Gemini Robotics 1.5 pairing, ER 1.5 acts as the “brain”: it can call Google Search to look up local recycling rules, break the task into steps, and hand each step's natural-language instruction to the VLA model, Gemini Robotics 1.5, to carry out the pick-and-place actions.","related":["Embodied Reasoning","Dual-System Architecture (System 1 / System 2)","Hierarchical Architecture","Gemini Robotics-ER","RoboBrain","Vision-Language-Action Model"]},{"id":"dual-system-architecture","category":"model","sec":7,"tier":1,"sources":[{"title":"Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure)","url":"https://www.figure.ai/news/helix"},{"title":"Helix 02 (Figure)","url":"https://www.figure.ai/news/helix-02"},{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"}],"as_of":"2026-01","related_ids":["embodied-reasoning-model","hierarchical-architecture","vision-language-action-model","figure-helix","nvidia-isaac-gr00t-n1","system-0"],"name":"Dual-System Architecture (System 1 / System 2)","alt":"快慢双系统","abbr":"","aliases":["System 1 / System 2","S1 / S2","Fast-Slow System","Fast System","Slow System","System 0/1/2 Architecture"],"one_liner":"A slow, deliberate large model handles understanding and planning, while a fast, lightweight model handles real-time motor control.","explanation":"The dual-system architecture borrows psychologist Daniel Kahneman's language from Thinking, Fast and Slow: System 2 is slow and deliberate, System 1 is fast and intuitive. Applied to robots, System 2 is usually a vision-language model with a few billion to a few tens of billions of parameters, running a few to a dozen times per second, looking at the scene and understanding the instruction, and outputting a semantic goal or a latent vector; System 1 is a much smaller policy network outputting joint actions at up to hundreds of hertz. This combines a large model's common sense with the low-latency demands of real-time control. Representative systems are Figure's Helix (February 2025) and NVIDIA's GR00T N1 (March 2025). Figure's Helix 02, from January 2026, adds a 1kHz System 0 for balance and contact, becoming a three-system architecture.","example":"Figure Helix's S2 is a 7-billion-parameter open-source VLM running at 7 to 9Hz; its S1 is an 80-million-parameter Transformer controlling 35 degrees of freedom across the wrists, fingers, torso, and head at 200Hz. GR00T N1's vision-language part runs at 10Hz on an L40 GPU, and its action part outputs actions at 120Hz.","related":["Embodied Reasoning Model","Hierarchical Architecture","Vision-Language-Action Model","Figure Helix","NVIDIA Isaac GR00T N1","System 0"]},{"id":"system-0","category":"model","sec":7,"tier":3,"sources":[{"title":"Introducing Helix 02: Full-Body Autonomy (Figure)","url":"https://www.figure.ai/news/helix-02"},{"title":"Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure)","url":"https://www.figure.ai/news/helix"}],"as_of":"2026-01","related_ids":["dual-system-architecture","hierarchical-architecture","whole-body-control","learning-based-whole-body-control","figure-helix-02","braincerebellum-architecture"],"name":"System 0","alt":"System 0（三层系统架构）","abbr":"S0","aliases":["S0/S1/S2 Architecture","Three-System Architecture"],"one_liner":"A kilohertz-rate whole-body control network added below the fast/slow two-system split, dedicated to balance, contact, and coordination.","explanation":"System 0 is a layer Figure AI added when it released Helix 02 on January 27, 2026. The original Helix was a dual-system design: System 2 is a vision-language model that understands the scene and instruction and produces an implicit goal; System 1 is a visuomotor policy that turns that goal into joint commands. Helix 02 adds S0 below S1: a roughly 10-million-parameter neural whole-body controller running at 1kHz, responsible for whole-body balance, contact, and coordination; it's trained on more than 1,000 hours of joint-level human motion data across more than 200,000 parallel simulated environments, and transfers to the real robot via domain randomization. S1 was also upgraded to output whole-body joint targets — from legs and torso to fingers — at 200Hz. This way, the upper layers handle 'what to do,' while the bottom layer guarantees 'stay balanced,' which is what lets a humanoid walk and work at the same time.","example":"Figure showed Helix 02 autonomously completing about 4 minutes of dishwasher loading and unloading throughout an entire kitchen, including 61 combined walking-and-manipulation actions, with no human intervention.","related":["Dual-System Architecture (System 1 / System 2)","Hierarchical Architecture","Whole-Body Control","Learning-Based Whole-Body Control","Figure Helix 02","Brain–Cerebellum Architecture"]},{"id":"behavior-foundation-model","category":"model","sec":7,"tier":3,"sources":[{"title":"A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots (arXiv 2506.20487)","url":"https://arxiv.org/abs/2506.20487"},{"title":"BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised RL (arXiv 2511.04131)","url":"https://arxiv.org/abs/2511.04131"},{"title":"arXiv 检索：behavior foundation model humanoid","url":"https://arxiv.org/search/?query=%22behavior+foundation+model%22+humanoid&searchtype=all"}],"as_of":"2025-11","related_ids":["whole-body-control","motion-tracking","bfm-zero","meta-motivo","foundation-model","unsupervised-skill-discovery"],"name":"Behavior Foundation Model","alt":"行为基础模型","abbr":"BFM","aliases":["BFM","Behavioral Foundation Model"],"one_liner":"A large-pretrained humanoid whole-body control model that can adapt to many motion tasks zero-shot or with little tuning.","explanation":"Behavior foundation model is a term that has emerged over the past two years in humanoid whole-body control (low-level motor control that coordinates the legs, torso, and arms together). The traditional approach trains a separate reinforcement-learning controller for each task; a BFM is instead pretrained first on large amounts of human motion data and a wide variety of tasks, learning reusable base skills and behavioral priors, and then adapts to a new task zero-shot or with minimal tuning through a 'prompt' — a reference motion clip, a target pose, or a reward function — playing a role similar to a foundation model in language. Representative work includes Meta researchers' Meta Motivo in April 2025 (based on unsupervised reinforcement learning with forward-backward representations) and BFM-Zero in November 2025 (deployed on a real Unitree G1, where the same policy can be used for motion tracking, reaching a target pose, and reward optimization). In a hierarchical architecture, it can serve as the lower-level motor controller that takes commands from a higher-level layer.","example":"BFM-Zero runs a single pretrained policy on the Unitree G1, with no retraining, and can switch by prompt between tracking a human motion clip or walking to a specified pose.","related":["Whole-Body Control","Motion Tracking","BFM-Zero","Meta Motivo","Foundation Model","Unsupervised Skill Discovery"]},{"id":"human-motion-generation","category":"model","sec":7,"tier":3,"sources":[{"title":"Human Motion Diffusion Model (Tevet et al., arXiv:2209.14916)","url":"https://arxiv.org/abs/2209.14916"},{"title":"MDM 项目主页（ICLR 2023）","url":"https://guytevet.github.io/mdm-page/"},{"title":"HumanML3D 数据集（GitHub）","url":"https://github.com/EricGuo5513/HumanML3D"}],"as_of":"2023-05","related_ids":["mdm","humanml3d","diffusion-model","motion-retargeting","motion-tracking","text-to-motion"],"name":"Human Motion Generation","alt":"人体动作生成模型（文本生成动作）","abbr":"","aliases":["MDM"],"one_liner":"A model that generates a sequence of 3D human skeletal poses from a condition such as a text description.","explanation":"This kind of model takes a description (such as 'a person walks forward and then sits down') or a motion category as input and outputs a frame-by-frame sequence of 3D human skeletal poses. A common training set is HumanML3D, released in 2022, with 14,616 motion clips and 44,970 text descriptions, with the motions drawn from motion-capture datasets such as AMASS. The representative work is MDM (Human Motion Diffusion Model), proposed by Tevet and colleagues and published at ICLR 2023: it implements a diffusion model with a Transformer, and at each step predicts the clean motion directly instead of the noise, which makes it easy to add geometric constraints like foot contact; the same model can also do motion in-betweening, interpolation, and editing of individual body parts. This kind of model was originally built for animation and games; in embodied AI, the generated human motion can be retargeted (mapping human joints onto robot joints) and handed to a humanoid's motion-tracking controller, making it one route to 'make a humanoid do something with a sentence.'","example":"Given MDM the prompt 'a person walks forward and then sits down,' it outputs a 3D skeletal animation of a few steps forward followed by sitting, which can then be retargeted onto a humanoid robot.","related":["MDM (Motion Diffusion Model)","HumanML3D","Diffusion Model","Motion Retargeting","Motion Tracking","Text-to-Motion"]},{"id":"intermediate-representation","category":"model","sec":7,"tier":3,"sources":[{"title":"RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches (arXiv:2311.01977)","url":"https://arxiv.org/abs/2311.01977"},{"title":"HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation (arXiv:2502.05485)","url":"https://arxiv.org/abs/2502.05485"}],"as_of":"2025-05","related_ids":["hierarchical-architecture","rt-trajectory","affordance","semantic-keypoints","visual-prompting","dual-system-architecture"],"name":"Intermediate Representation","alt":"中间表示","abbr":"","aliases":["Mid-level Representation"],"one_liner":"Transitional information a model predicts between an instruction and low-level action, such as a 2D trajectory or keypoints.","explanation":"In robot learning, an intermediate representation means not going directly from an image and instruction to joint commands, but first producing something human-readable and less tied to a specific robot, which a downstream policy or controller then turns into an action. Common forms include an end-effector trajectory drawn on the image, object keypoints, an affordance heatmap (marking where to grasp or push), a subgoal image, and a language subtask. Google DeepMind and colleagues' 2023 RT-Trajectory uses a rough trajectory sketch as the policy's condition, letting it complete new tasks that language conditioning alone couldn't; 2025's HAMSTER has a high-level vision-language model predict a 2D path, which a lower-level, 3D-aware policy then executes, reaching about 20 percentage points higher average success than OpenVLA in real-robot experiments. The benefit is that the high level can be trained on cheap data such as action-free video and simulation, and it's easier for a person to inspect and correct; the cost is that information gets compressed, and a poorly chosen representation limits precision. Compiler design has a same-named concept (IR) with a different meaning.","example":"A language-conditioned policy trained only on pick-and-place data usually can't learn a new task like folding; RT-Trajectory instead conditions on a trajectory sketch drawn over the image, either hand-drawn or produced by a generative model, letting the policy carry out motions it never saw during training.","related":["Hierarchical Architecture","RT-Trajectory","Affordance","Semantic Keypoints","Visual Prompting","Dual-System Architecture (System 1 / System 2)"]},{"id":"visual-prompting-2","category":"model","sec":7,"tier":3,"sources":[{"title":"Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V (arXiv:2310.11441)","url":"https://arxiv.org/abs/2310.11441"},{"title":"MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv:2403.03174)","url":"https://arxiv.org/abs/2403.03174"}],"as_of":"","related_ids":["vision-language-model","visual-grounding","segment-anything-model","moka","pivot","prompt-prompt-engineering"],"name":"Visual Prompting (Set-of-Mark)","alt":"视觉提示（Set-of-Mark 标记提示）","abbr":"SoM","aliases":["SoM","Set-of-Mark Prompting","Mark-based Visual Prompting"],"one_liner":"Drawing numbers, boxes, or dots on an image so a vision-language model can point at a location just by naming its label.","explanation":"Visual prompting means guiding a vision-language model (VLM) by marking up the input image itself, without changing any model weights. The flagship method is Set-of-Mark (SoM), introduced by Microsoft Research in October 2023: a segmentation model such as SAM or SEEM first cuts the image into regions, each region gets a number, mask, or box overlaid on it, and the marked-up image is then shown to a model like GPT-4V. It's hard for a VLM to output precise pixel coordinates directly, but answering “region 3” is much easier — SoM even beat specially fine-tuned models on zero-shot RefCOCOg referring segmentation. In robotics, this is often used to turn “where to grasp, where to place it” into a multiple-choice question: MOKA (2024, UC Berkeley) marks candidate keypoints and waypoints on the image and has the VLM pick the grasp point and motion path, which then gets converted into arm actions; PIVOT instead draws a batch of candidate actions on the image and has the VLM select and iteratively narrow them down.","example":"A tabletop photo is first segmented with SAM into individual objects labeled 1, 2, 3, and so on. Asking GPT-4V “which one is the red cup” gets back the answer “2,” and the program can then pull the precise mask for region 2 and hand it to the grasping module.","related":["Vision-Language Model","Visual Grounding","Segment Anything Model","MOKA","PIVOT","Prompt / Prompt Engineering"]},{"id":"embodied-chain-of-thought","category":"model","sec":7,"tier":2,"sources":[{"title":"Robotic Control via Embodied Chain-of-Thought Reasoning (arXiv 2407.08693)","url":"https://arxiv.org/abs/2407.08693"},{"title":"Embodied Chain-of-Thought Reasoning 项目页","url":"https://embodied-cot.github.io/"},{"title":"Training Strategies for Efficient Embodied Reasoning (arXiv 2505.08243)","url":"https://arxiv.org/abs/2505.08243"}],"as_of":"2025-05","related_ids":["chain-of-thought","embodied-reasoning","vision-language-action-model","openvla","visual-chain-of-thought","inference-latency"],"name":"Embodied Chain-of-Thought","alt":"具身思维链","abbr":"ECoT","aliases":["ECoT","Embodied Chain-of-Thought Reasoning"],"one_liner":"A method that has a robot model reason step by step — plan, subtask, object locations — before producing an action.","explanation":"Embodied Chain-of-Thought (ECoT) was proposed in 2024 by researchers from UC Berkeley, Stanford, and other institutions. A standard vision-language-action (VLA) model looks at an image, hears an instruction, and outputs an action directly. ECoT instead has the model write out a chain of reasoning first — a restatement of the task, an overall plan, the current subtask, the direction the gripper should move next, the gripper's position, and bounding boxes for objects in the scene — and only then output the action. Unlike chain-of-thought in large language models, this reasoning has to stay grounded in the actual image and robot state. The authors used off-the-shelf foundation models to automatically label the BridgeData V2 dataset with these reasoning traces, then trained OpenVLA on them; on hard generalization tasks, absolute success rate rose by 28%, with no extra robot data. The cost is generating many more tokens, which slows inference down; a 2025 follow-up from the same team introduced a lighter training scheme that runs about 3x faster.","example":"Given the instruction ‘put the mushroom in the pot,’ an ECoT model first writes out a plan (find the mushroom → pick it up → move it above the pot → put it down), the current subtask (‘pick up the mushroom’), a movement direction (‘down and to the left’), and bounding boxes for the mushroom, the pot, and the gripper — only then does it output the 7-dimensional arm action.","related":["Chain-of-Thought","Embodied Reasoning","Vision-Language-Action Model","OpenVLA","Visual Chain-of-Thought","Inference Latency"]},{"id":"visual-chain-of-thought","category":"model","sec":7,"tier":3,"sources":[{"title":"CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models (arXiv:2503.22020)","url":"https://arxiv.org/abs/2503.22020"},{"title":"Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning (arXiv:2403.16999)","url":"https://arxiv.org/abs/2403.16999"}],"as_of":"","related_ids":["chain-of-thought","embodied-chain-of-thought","action-chain-of-thought","cot-vla","vision-language-action-model","video-prediction-model"],"name":"Visual Chain-of-Thought","alt":"视觉思维链","abbr":"Visual CoT","aliases":["Visual CoT","Visual Reasoning Chain"],"one_liner":"Having a model first produce visual intermediate steps, like generated images or marked regions, before giving its final answer or action.","explanation":"Chain-of-thought (CoT) originally meant having a large language model write out its reasoning steps before answering; visual chain-of-thought replaces those intermediate steps with visual ones. In multimodal models, the Visual CoT dataset released by Shao and colleagues in 2024 contains 438,000 question-answer pairs with intermediate bounding boxes, training models to first circle the key region of an image, look closer, and only then answer. In robotics, a more common pattern is to first “imagine” the goal image: CoT-VLA (CVPR 2025), from NVIDIA, Stanford, and others, first autoregressively generates an image of a future subgoal, then generates the action chunk to reach it, beating the previous best VLA by 17% on real robots and 6% in simulation. The benefit is that the target state is drawn out explicitly, which makes it easier to check and helps with long-horizon tasks; the cost is slower inference, since an extra image must be generated. It sits alongside other intermediate representations, such as embodied chain-of-thought, which writes out subtasks and object locations in text, and action chain-of-thought.","example":"After receiving a manipulation instruction, CoT-VLA first generates an image of what the scene should look like several steps ahead — say, the object already grasped and near its target — as a subgoal, then generates a sequence of actions conditioned on that image. After execution, it observes again and imagines the next subgoal image.","related":["Chain-of-Thought","Embodied Chain-of-Thought","Action Chain-of-Thought","CoT-VLA","Vision-Language-Action Model","Video Prediction Model"]},{"id":"action-chain-of-thought","category":"model","sec":7,"tier":3,"sources":[{"title":"ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models (arXiv 2601.11404)","url":"https://arxiv.org/abs/2601.11404"},{"title":"ACoT-VLA 论文 HTML 版","url":"https://arxiv.org/html/2601.11404"}],"as_of":"2026-01","related_ids":["chain-of-thought","embodied-chain-of-thought","visual-chain-of-thought","vision-language-action-model","action-head","flow-matching"],"name":"Action Chain-of-Thought","alt":"动作思维链","abbr":"ACoT","aliases":["ACoT","ACoT-VLA"],"one_liner":"Having a VLA first reason out a coarse trajectory in action space, then generate the fine-grained action from it.","explanation":"Chain-of-thought originally meant having a large model write out intermediate reasoning steps before giving an answer. Existing reasoning approaches in VLAs mostly predict subtask text or generate a goal image as the intermediate step — neither of which is an action itself. In January 2026, a team from Beihang University and AgiBot proposed ACoT-VLA (accepted to CVPR 2026), arguing for reasoning directly in action space: an Explicit Action Reasoner (EAR) uses flow matching to first generate a coarse reference trajectory; an Implicit Action Reasoner (IAR) uses learnable queries to extract a latent action prior from the VLM's internal features; the two are then combined via cross-attention to guide the action head in producing the final action sequence. The paper reports a 98.5% average success rate on LIBERO. It can be seen as another form of reasoning alongside embodied chain-of-thought and visual chain-of-thought. Separately, some 2026 navigation work also uses 'Action-CoT' to refer to step-by-step action reasoning.","example":"ACoT-VLA reports a 66.7% average success rate across three manipulation tasks on a real AgiBot G1 robot, and an 88.0% average success rate on LIBERO-Plus after supervised fine-tuning.","related":["Chain-of-Thought","Embodied Chain-of-Thought","Visual Chain-of-Thought","Vision-Language-Action Model","Action Head","Flow Matching"]},{"id":"latent-reasoning","category":"model","sec":7,"tier":3,"sources":[{"title":"Training Large Language Models to Reason in a Continuous Latent Space (Coconut, arXiv:2412.06769)","url":"https://arxiv.org/abs/2412.06769"},{"title":"A Survey on Latent Reasoning (arXiv:2507.06203)","url":"https://arxiv.org/abs/2507.06203"},{"title":"ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning (arXiv:2507.16815)","url":"https://arxiv.org/abs/2507.16815"}],"as_of":"2025-07","related_ids":["chain-of-thought","embodied-chain-of-thought","reasoning","large-language-model","thinkact","inference-latency"],"name":"Latent Reasoning","alt":"潜在推理","abbr":"","aliases":["Latent CoT","Continuous Thought"],"one_liner":"Having a model carry out multi-step reasoning inside its internal hidden states, instead of writing out a chain of thought in words.","explanation":"Chain-of-thought has a large model write out intermediate reasoning steps in words before giving an answer; it works well, but generating a token at every step is slow and constrained by language. Latent reasoning instead keeps the intermediate steps inside the model's continuous hidden state: Coconut, proposed by Meta and others in December 2024, treats the last layer's hidden state as one step of 'continuous thought,' feeding it directly back in as the next step's input embedding instead of decoding it into words, and outperforms word-based chain-of-thought on logic problems that need a lot of search, while generating fewer tokens too. A July 2025 survey defines this direction as reasoning done entirely in continuous hidden states, with no per-token supervision. Embodied AI needs both reasoning and low latency for a VLA, so it borrows the same idea: for example, ThinkAct compresses the reasoning plan a multimodal large model produces into a single 'visual plan latent,' which then guides a downstream action model. The cost is that the intermediate process becomes unreadable, making it hard to inspect or debug.","example":"Solving a logic problem that needs several steps of deduction, Coconut doesn't write out text like 'because A, therefore B'; instead it feeds a hidden vector back into itself as input over several steps, only outputting the answer at the end.","related":["Chain-of-Thought","Embodied Chain-of-Thought","Reasoning","Large Language Model","ThinkAct","Inference Latency"]},{"id":"world-model","category":"model","sec":8,"tier":1,"sources":[{"title":"World Models (Ha & Schmidhuber, arXiv:1803.10122)","url":"https://arxiv.org/abs/1803.10122"},{"title":"Genie 3: A new frontier for world models (Google DeepMind)","url":"https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/"},{"title":"What Are World Models? (NVIDIA Glossary)","url":"https://www.nvidia.com/en-us/glossary/world-models/"},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv:2506.09985)","url":"https://arxiv.org/abs/2506.09985"}],"as_of":"2025-08","related_ids":["latent-world-model","world-foundation-model","world-action-model","nvidia-cosmos","genie-3","model-based-reinforcement-learning"],"name":"World Model","alt":"世界模型","abbr":"WM","aliases":["WM","World Simulator"],"one_liner":"A model that predicts what the world will look like after a given action is taken.","explanation":"A world model is one that has learned how an environment changes: given the current state or frame and an action, it predicts the next state or frame. Ha and Schmidhuber's 2018 paper “World Models” popularized the term: it compresses game frames into a low-dimensional vector with a variational autoencoder, then uses a recurrent network to predict the next step, letting an agent learn a policy inside the model's generated “dream” and transfer it back to the real game. Today the term covers models that predict purely in latent space, meaning compressed features, such as DreamerV3, which serves reinforcement learning, and Meta's V-JEPA 2, used for robot planning, as well as large models that directly generate video frames, such as Google DeepMind's Genie 3 and NVIDIA's Cosmos. In embodied AI, world models are used to generate synthetic training data, evaluate policies “in imagination,” and plan, all to cut down on real-robot trial and error.","example":"Genie 3, released in August 2025, generates a 720p, 24-frames-per-second scene from a text description; as the user presses direction keys to move around, it generates what comes next in real time, and can remember what it saw roughly a minute earlier.","related":["Latent World Model","World Foundation Model","World Action Model","NVIDIA Cosmos","Genie 3","Model-Based Reinforcement Learning"]},{"id":"forward-dynamics-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Neural Network Dynamics for Model-Based Deep RL with Model-Free Fine-Tuning (arXiv 1708.02596)","url":"https://arxiv.org/abs/1708.02596"},{"title":"Curiosity-driven Exploration by Self-supervised Prediction (arXiv 1705.05363)","url":"https://arxiv.org/abs/1705.05363"}],"as_of":"","related_ids":["inverse-dynamics-model","world-model","model-based-reinforcement-learning","model-predictive-control","forward-dynamics","latent-action-model"],"name":"Forward Dynamics Model","alt":"正向动力学模型","abbr":"FDM","aliases":["FDM","Forward Model","Learned Dynamics Model"],"one_liner":"A model that takes the current state and an action and predicts what the next state will look like.","explanation":"A forward dynamics model learns ‘what the world will look like after this action is taken’: it takes the current state (or observation) s_t and an action a_t as input and outputs the next state s_{t+1}. The idea parallels forward dynamics in mechanics, where torques are used to compute acceleration, but in robot learning it usually isn't derived from physics equations — instead, a neural network learns it from interaction data. With such a model, a robot can try out many candidate action sequences inside the model first and execute whichever leads to the best predicted outcome; this is exactly what model-based reinforcement learning and model predictive control (MPC) do. In 2017, Nagabandi et al. combined a learned neural-network dynamics model with MPC to make simulated legged robots track arbitrary trajectories. Today's world models are essentially large forward dynamics models operating on images or a latent space; a forward dynamics model is often paired with an inverse dynamics model, and the decoder of a latent action model belongs to this family too.","example":"In a block-pushing task, the model is given the block's current position and the action ‘push right 2 cm’ and predicts the block's new position afterward; MPC uses it to compare dozens of candidate pushes and execute whichever gets closest to the goal.","related":["Inverse Dynamics Model","World Model","Model-Based Reinforcement Learning","Model Predictive Control","Forward Dynamics","Latent Action Model"]},{"id":"inverse-dynamics-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (arXiv 2206.11795)","url":"https://arxiv.org/html/2206.11795"},{"title":"Learning Universal Policies via Text-Guided Video Generation (UniPi, arXiv 2302.00111)","url":"https://arxiv.org/abs/2302.00111"}],"as_of":"","related_ids":["forward-dynamics-model","latent-action-model","pseudo-action-labels","action-free-video","vpt","unipi"],"name":"Inverse Dynamics Model","alt":"逆动力学模型","abbr":"IDM","aliases":["IDM","Inverse Model"],"one_liner":"A model that looks at two frames, before and after, and infers what action happened in between.","explanation":"An inverse dynamics model runs in the opposite direction from a forward dynamics model: it takes the current state s_t and the next state s_{t+1} (or a pair of consecutive video frames) and outputs the action a_t that caused the change. In classical mechanics, inverse dynamics works backward from a desired motion to the joint torques that would produce it; in robot learning, an IDM is usually a neural network trained on data instead. Its main use is labeling videos that have no action labels: OpenAI's 2022 VPT project first trained an IDM on about 2,000 hours of Minecraft footage recorded together with keyboard and mouse input, then used it to attach pseudo action labels to about 70,000 hours of online video for pretraining a game-playing agent. Because an IDM can see both the past and future frame, predicting the action is much easier for it to learn than predicting an action from the current frame alone. ‘Generate video, then infer actions’ methods such as UniPi also rely on an IDM to translate generated frames into robot actions.","example":"Given two frames — the first shows an open gripper hovering above a cup, the second shows the gripper closed with the cup lifted a few centimeters — an IDM outputs the action vector for ‘close the gripper and move up.’","related":["Forward Dynamics Model","Latent Action Model","Pseudo Action Labels","Action-free Video","VPT","UniPi"]},{"id":"latent-action","category":"model","sec":8,"tier":2,"sources":[{"title":"Genie: Generative Interactive Environments (arXiv 2402.15391)","url":"https://arxiv.org/html/2402.15391"},{"title":"Latent Action Pretraining from Videos (LAPA, arXiv 2410.11758)","url":"https://arxiv.org/abs/2410.11758"},{"title":"AgiBot World Colosseo (GO-1, arXiv 2503.06669)","url":"https://arxiv.org/abs/2503.06669"}],"as_of":"","related_ids":["latent-action-model","latent-action-pretraining","action-free-video","genie","lapa","agibot-go-1"],"name":"Latent Action","alt":"潜在动作","abbr":"","aliases":[],"one_liner":"An abstract action code learned automatically from how consecutive video frames change, with no real action labels needed.","explanation":"A latent action is an action representation a model infers automatically from the change between adjacent video frames, usually a handful of discrete codes or a single low-dimensional vector. It targets a specific problem: there is a huge amount of human and gameplay video on the internet, but none of it carries real action labels like joint angles or button presses, so it can't be used to train a policy directly. A latent action doesn't correspond to any specific robot's joints; it only describes ‘what kind of change happened on screen,’ so the same coding scheme can be shared across human-hand video and video from very different robots. DeepMind's 2024 Genie learned 8 discrete latent actions from about 30,000 hours of platformer game video and used them to control its generated game worlds; methods such as LAPA, UniVLA, and AgiBot's GO-1 instead pretrain a VLA to predict latent actions first, then fine-tune on a small amount of real robot data to learn how to map them to real actions.","example":"The 8 latent actions Genie learned are numbered consistently across the games it can generate, roughly corresponding to moves like left, right, and jump; users can steer a character in a scene the model has never seen just by pressing the numbered action.","related":["Latent Action Model","Latent Action Pretraining","Action-free Video","Genie (Original)","LAPA","AgiBot GO-1"]},{"id":"latent-action-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Genie: Generative Interactive Environments (arXiv 2402.15391)","url":"https://arxiv.org/html/2402.15391"},{"title":"Latent Action Pretraining from Videos (LAPA, arXiv 2410.11758)","url":"https://arxiv.org/abs/2410.11758"},{"title":"UniVLA: Learning to Act Anywhere with Task-centric Latent Actions (arXiv 2505.06111)","url":"https://arxiv.org/abs/2505.06111"}],"as_of":"","related_ids":["latent-action","inverse-dynamics-model","forward-dynamics-model","vector-quantized-variational-autoencoder","genie","univla"],"name":"Latent Action Model","alt":"潜在动作模型","abbr":"LAM","aliases":["LAM"],"one_liner":"The network that produces latent actions from unlabeled video, trained by having an encoder-decoder pair reconstruct the next frame.","explanation":"A latent action model is the network that produces latent actions. The typical design comes from DeepMind's 2024 Genie: an encoder looks at the current frame and the next frame and outputs a latent action, playing the role of an inverse dynamics model; a decoder then takes only the current frame plus that latent action and reconstructs the next frame, playing the role of a forward dynamics model. In between, vector quantization (the VQ-VAE approach) restricts the latent action to a very small discrete codebook — Genie uses just 8 codes — so the latent action can't encode a whole image, only ‘what changed.’ Once trained, the encoder can attach latent-action labels to huge amounts of video. A common difficulty is that task-irrelevant changes, such as camera shake or a person walking through the background, also get encoded; UniVLA addresses this by learning in DINO feature space instead and using the language instruction to strip out such distractions.","example":"LAPA has three stages: first train a latent-action quantization model on video, then train a VLA to predict latent actions from images and instructions, and finally fine-tune on a small amount of robot data to replace latent actions with real ones.","related":["Latent Action","Inverse Dynamics Model","Forward Dynamics Model","Vector-Quantized Variational Autoencoder","Genie (Original)","UniVLA"]},{"id":"vision-language-latent-action","category":"model","sec":8,"tier":3,"sources":[{"title":"AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv:2503.06669)","url":"https://arxiv.org/abs/2503.06669"},{"title":"OpenDriveLab/AgiBot-World (GitHub, GO-1 / GO-1 Air 开源)","url":"https://github.com/OpenDriveLab/AgiBot-World"}],"as_of":"2025-09","related_ids":["agibot-go-1","latent-action-model","latent-action","action-expert","hierarchical-architecture","latent-action-pretraining"],"name":"Vision-Language-Latent-Action","alt":"ViLLA 架构","abbr":"ViLLA","aliases":["ViLLA","Latent Planner + Action Expert"],"one_liner":"The hierarchical architecture behind AgiBot's GO-1: it first predicts latent action tokens, then decodes them into real robot actions.","explanation":"ViLLA is the framework AgiBot introduced in March 2025 alongside GO-1 (Genie Operator-1), its general-purpose embodied foundation model. A standard VLA (vision-language-action model) maps images and instructions directly to actions; ViLLA inserts an intermediate “latent action” layer with three parts. A latent action model (LAM) trains on human videos (such as Ego4D) and robot trajectories, using inverse and forward dynamics to compress the change between two adjacent frames into discrete latent action tokens. A latent planner, built on an InternVL2.5-2B backbone, predicts these tokens from multi-view images and the instruction. An action expert then decodes them into continuous, low-level action chunks using a diffusion objective. The benefit is that human videos with no action labels can still be used for pretraining, improving data efficiency. GO-1 was open-sourced in September 2025, together with a lighter GO-1 Air variant that drops the latent planner.","example":"When GO-1 carries out a tabletop instruction, the latent planner first looks at three camera views plus the instruction and outputs a handful of latent action tokens — an abstract sketch of “what to do next.” The action expert then denoises conditioned on those tokens to generate the next 30 timesteps of continuous action.","related":["AgiBot GO-1","Latent Action Model","Latent Action","Action Expert","Hierarchical Architecture","Latent Action Pretraining"]},{"id":"video-prediction-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Deep Visual Foresight for Planning Robot Motion (arXiv 1610.00696)","url":"https://arxiv.org/abs/1610.00696"},{"title":"Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control (arXiv 1812.00568)","url":"https://arxiv.org/abs/1812.00568"},{"title":"Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations (arXiv 2412.14803)","url":"https://arxiv.org/abs/2412.14803"}],"as_of":"","related_ids":["video-generation-model","world-model","model-predictive-control","forward-dynamics-model","video-prediction-policy","world-action-model"],"name":"Video Prediction Model","alt":"视频预测模型","abbr":"","aliases":["Action-conditioned Video Prediction"],"one_liner":"A model that predicts upcoming frames from recent ones, often given the action that's about to be taken.","explanation":"A video prediction model takes in the past several frames and outputs future frames; in robotics it's usually also given the action the robot is about to take, answering 'what will the camera see if I do this?' — this variant is called action-conditioned video prediction. In 2016, Finn and Levine proposed Visual Foresight: a robot pushes objects around on its own with no human labeling to collect data, trains a prediction model on that data, and then combines it with model predictive control — trying many candidate actions inside the model at each step and executing whichever one's predicted outcome comes closest to the goal — to complete pushing tasks. The difference from a video generation model is that it focuses on continuing forward from an existing view rather than generating from scratch given text. Recent work has shifted to large-scale video diffusion models instead; for example, Video Prediction Policy (VPP) uses a video model's predictive features to drive an inverse dynamics model that outputs actions.","example":"To push a block on a table to a target position, the robot first tries many candidate action sequences inside a video prediction model, compares which one's predicted frames come closest to the goal image, executes just the first step of that sequence, and repeats.","related":["Video Generation Model","World Model","Model Predictive Control","Forward Dynamics Model","Video Prediction Policy","World Action Model"]},{"id":"video-generation-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Video Diffusion Models (arXiv 2204.03458)","url":"https://arxiv.org/abs/2204.03458"},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models (arXiv 2503.20314)","url":"https://arxiv.org/abs/2503.20314"}],"as_of":"2025-04","related_ids":["text-to-video-image-to-video","diffusion-model","diffusion-transformer","video-prediction-model","world-model","wan"],"name":"Video Generation Model","alt":"视频生成模型","abbr":"","aliases":["Video Diffusion Model"],"one_liner":"A generative model that produces a new, coherent video from text, an image, or an existing clip.","explanation":"A video generation model can produce new video from text, an image, or an existing clip. Google's 2022 paper Video Diffusion Models extended image diffusion models to video; since then, the dominant approach has been to compress video into latent space with a video VAE first, then denoise multiple frames together step by step with a diffusion Transformer or U-Net, so consecutive frames stay coherent. Notable examples include OpenAI's Sora, Google's Veo, and Alibaba's open-source Wan. Because these models learn how objects move and get pushed around from huge amounts of video, they're often treated as a starting point for world models: NVIDIA's Cosmos and other world foundation models, 'generate video then infer actions' policies like UniPi, and world action models like DreamZero are all built on top of video generation models.","example":"Wan is released open-source in two sizes, 1.3B and 14B; the 1.3B version needs only about 8 GB of GPU memory, so text-to-video runs on a consumer graphics card.","related":["Text-to-Video / Image-to-Video","Diffusion Model","Diffusion Transformer","Video Prediction Model","World Model","Wan (Alibaba Video Generation Model)"]},{"id":"text-to-video-image-to-video","category":"model","sec":8,"tier":2,"sources":[{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (arXiv 2311.15127)","url":"https://arxiv.org/abs/2311.15127"},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models (arXiv 2503.20314)","url":"https://arxiv.org/abs/2503.20314"},{"title":"World Action Models are Zero-shot Policies (arXiv 2602.15922)","url":"https://arxiv.org/html/2602.15922"}],"as_of":"2026-02","related_ids":["video-generation-model","diffusion-model","latent-diffusion-model","variational-autoencoder","world-action-model","unipi"],"name":"Text-to-Video / Image-to-Video","alt":"文生视频 / 图生视频","abbr":"T2V / I2V","aliases":["T2V","I2V","Text2Video","Image2Video"],"one_liner":"Generating a video from a text description, or from an image plus text.","explanation":"Text-to-video (T2V) and image-to-video (I2V) are the two most common ways of using a video generation model. In T2V, the model generates video from scratch given only a text description; in I2V, it's given an image (usually treated as the first frame) plus text, and animates the scene forward from there. Stability AI's Stable Video Diffusion and Alibaba's Wan both offer versions of each; the dominant approach is to run diffusion or flow-matching denoising step by step inside a latent space produced by a VAE. I2V is more useful for robotics: treating the robot's current camera view as the first frame and the task instruction as the text turns the generated video into a preview of what should happen next. UniPi uses text-guided video generation for planning, while NVIDIA's DreamZero adds action output directly on top of Wan2.1's 14B image-to-video model.","example":"Given a photo of a robot arm and a cup on a table, plus the instruction 'put the cup in the sink,' an image-to-video model generates a few seconds of video starting from that photo, showing the arm carrying out the task.","related":["Video Generation Model","Diffusion Model","Latent Diffusion Model","Variational Autoencoder","World Action Model","UniPi"]},{"id":"video-tokenizer","category":"model","sec":8,"tier":3,"sources":[{"title":"Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation (MAGVIT-v2, arXiv:2310.05737)","url":"https://arxiv.org/abs/2310.05737"},{"title":"NVIDIA Cosmos Tokenizer (GitHub)","url":"https://github.com/NVIDIA/Cosmos-Tokenizer"},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models (arXiv:2503.20314)","url":"https://arxiv.org/abs/2503.20314"}],"as_of":"","related_ids":["variational-autoencoder","latent-space","spacetime-patches","vector-quantization","latent-diffusion-model","nvidia-cosmos"],"name":"Video Tokenizer","alt":"视频分词器","abbr":"","aliases":["Video VAE","3D Causal VAE","Causal Video VAE"],"one_liner":"A codec network that compresses video into a small number of tokens or latent vectors and can reconstruct the footage from them.","explanation":"A video tokenizer is a front-end module used by video generation models and world models. An encoder compresses a video clip across both time and space at once into either discrete tokens or continuous latent vectors, and a decoder reconstructs pixels from them. Models that output discrete tokens (such as Google's MAGVIT-v2) quantize vectors into “vocabulary indices” that plug neatly into an autoregressive Transformer; models that output continuous vectors are video VAEs (variational autoencoders), which a latent diffusion model then denoises inside. “Causal” means that along the time axis each frame only looks at earlier frames, with the first frame encoded on its own — this lets images and video share one tokenizer and lets arbitrarily long videos be processed in segments. The compression ratio determines how many tokens the downstream large model must process, while reconstruction quality caps how sharp the generated video can look. NVIDIA's Cosmos Tokenizer offers continuous and discrete versions at several compression ratios, including 4×8×8 and 8×16×16.","example":"Alibaba's Wan2.1 uses a 3D causal VAE called Wan-VAE that compresses time 4× and each spatial dimension 8×; with feature caching it can encode and decode 1080p video of any length. The diffusion Transformer only denoises inside this compressed latent space, and the decoder reconstructs the final footage.","related":["Variational Autoencoder","Latent Space","Spacetime Patches","Vector Quantization","Latent Diffusion Model","NVIDIA Cosmos"]},{"id":"spacetime-patches","category":"model","sec":8,"tier":3,"sources":[{"title":"Sora: A Review on Background, Technology, Limitations, and Opportunities (arXiv:2402.17177)","url":"https://arxiv.org/abs/2402.17177"},{"title":"ViViT: A Video Vision Transformer (arXiv:2103.15691)","url":"https://arxiv.org/abs/2103.15691"},{"title":"Wikipedia: Sora (text-to-video model)","url":"https://en.wikipedia.org/wiki/Sora_(text-to-video_model)"}],"as_of":"2024-02","related_ids":["vision-transformer","diffusion-transformer","video-tokenizer","video-generation-model","token","sora"],"name":"Spacetime Patches","alt":"时空块","abbr":"","aliases":["Spacetime Latent Patches","Tubelet"],"one_liner":"Small cubes cut jointly across time and space from a video, each treated as one token for a Transformer.","explanation":"A spacetime patch is the video version of 'cutting an image into patches.' A Vision Transformer cuts an image into fixed-size squares, each treated as a token; video adds a time dimension, so it's instead cut into small cubes spanning several frames and several pixels — Google's ViViT (ICCV 2021) calls this a tubelet, commonly extracted with a 3D convolution. OpenAI popularized the term 'spacetime patches' when it released Sora in February 2024: a video-compression network first compresses the video into latent space, and the latent representation is then cut into spacetime patches and handed to a diffusion Transformer to denoise. The benefit is that videos of different resolutions, durations, and aspect ratios can all become token sequences of varying length and be trained together, with no need to crop or rescale first. Most video generation models and video world models today turn video into tokens in a similar way.","example":"If a compressed latent video is cut into blocks of '2 frames × 2 × 2 cells,' each block is flattened and mapped by a linear layer into one token; landscape and portrait videos simply differ in how many blocks they're cut into and how they're arranged, and can be trained in the same batch.","related":["Vision Transformer","Diffusion Transformer","Video Tokenizer","Video Generation Model","Token","Sora"]},{"id":"wan","category":"model","sec":8,"tier":3,"sources":[{"title":"Wan: Open and Advanced Large-Scale Video Generative Models (arXiv:2503.20314)","url":"https://arxiv.org/abs/2503.20314"},{"title":"Wan-Video/Wan2.2 (GitHub)","url":"https://github.com/Wan-Video/Wan2.2"},{"title":"World Action Models are Zero-shot Policies (DreamZero, arXiv:2602.15922)","url":"https://arxiv.org/abs/2602.15922"}],"as_of":"2026-09","related_ids":["video-generation-model","diffusion-transformer","video-tokenizer","flow-matching","mixture-of-experts","dreamzero"],"name":"Wan (Alibaba Video Generation Model)","alt":"通义万相","abbr":"","aliases":["Tongyi Wanxiang","Wan2.1","Wan2.2"],"one_liner":"Alibaba's Tongyi-team video generation model; Wan2.1 and Wan2.2 are open-weight and widely used as a base for world models.","explanation":"Wan (Tongyi Wanxiang in Chinese) is the video generation model family from Alibaba's Tongyi team. Wan2.1, open-sourced in late February 2025, comes in 1.3B and 14B sizes under an Apache 2.0 license; it's a diffusion Transformer trained with flow matching, ships with its own 3D causal video VAE (Wan-VAE) and a umT5 text encoder, and supports text-to-video, image-to-video, and video editing. Wan2.2, released in July 2025, switched to a mixture-of-experts design with two roughly 14B experts — a high-noise expert that sets overall layout and a low-noise expert that fills in detail, with only one active per step — plus a 5B text-and-image-to-video model; audio-driven (S2V) and character-animation (Animate) variants followed. Because Wan has learned a great deal about how objects move from massive video data, and its weights are open, researchers often use it as a starting point for world models or world-action models. As of September 2026, Wan2.2 remains the main open-weight release.","example":"NVIDIA's DreamZero builds on the Wan2.1-I2V-14B-480P image-to-video model as its backbone, continuing training on robot data so the model generates both future images and the corresponding actions at once — functioning as a world-action model that directly controls the robot.","related":["Video Generation Model","Diffusion Transformer","Video Tokenizer","Flow Matching","Mixture of Experts","DreamZero"]},{"id":"world-foundation-model","category":"model","sec":8,"tier":2,"sources":[{"title":"Cosmos World Foundation Model Platform for Physical AI (arXiv 2501.03575)","url":"https://arxiv.org/abs/2501.03575"},{"title":"NVIDIA Glossary: What Are World Models?","url":"https://www.nvidia.com/en-us/glossary/world-models/"},{"title":"NVIDIA Cosmos","url":"https://www.nvidia.com/en-us/ai/cosmos/"}],"as_of":"2026-09","related_ids":["world-model","foundation-model","nvidia-cosmos","video-generation-model","synthetic-data","physical-ai"],"name":"World Foundation Model","alt":"世界基础模型","abbr":"WFM","aliases":["WFM"],"one_liner":"A general-purpose world model pretrained on huge amounts of real video that can be fine-tuned into specialized world models.","explanation":"The term World Foundation Model was introduced in NVIDIA's January 2025 paper announcing the Cosmos platform, describing a general-purpose world model that can be post-trained into customized world models for specific applications. A world model's job is to predict what the world will look like next, given the current view plus an action or instruction. WFM borrows the large language model playbook — pretrain at massive scale first, then fine-tune per task: Cosmos filtered about 100 million clips out of roughly 20 million hours of raw video and used them to train both diffusion and autoregressive model variants. Its uses include generating synthetic training data for robots and self-driving cars, training and evaluating policies inside the model, and turning simulated footage into photorealistic footage. NVIDIA has since released the Cosmos Predict, Transfer, and Reason series, as well as Cosmos 3, which puts reasoning, video, and action generation into a single model.","example":"The Cosmos paper post-trains a pretrained WFM into a robot-manipulation version: given the current view and a sequence of robot actions, it predicts the video of what happens after executing them.","related":["World Model","Foundation Model","NVIDIA Cosmos","Video Generation Model","Synthetic Data","Physical AI"]},{"id":"autoregressive-video-generation","category":"model","sec":8,"tier":3,"sources":[{"title":"From Slow Bidirectional to Fast Autoregressive Video Diffusion Models (CausVid, arXiv 2412.07772)","url":"https://arxiv.org/abs/2412.07772"},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion (arXiv 2506.08009)","url":"https://arxiv.org/abs/2506.08009"},{"title":"Genie 3: A new frontier for world models (Google DeepMind)","url":"https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/"}],"as_of":"2025-08","related_ids":["diffusion-forcing","self-forcing","interactive-world-model","video-generation-model","exposure-bias","key-value-cache"],"name":"Autoregressive Video Generation","alt":"自回归视频生成","abbr":"","aliases":["Causal Video Generation","Streaming Video Generation"],"one_liner":"Generating video forward in time, one frame or chunk at a time, so it can be played as it's produced.","explanation":"Many video diffusion models generate a whole clip at once, with frames attending to each other in both directions, so they must finish the entire computation before producing any output, and it's hard to feed in new actions partway through. Autoregressive video generation instead writes forward in time: each frame or short chunk is conditioned only on what's already been generated, and combined with causal attention and a KV cache, this lets the model stream output and respond in real time to a user's or robot's actions — exactly what an interactive world model needs. The main difficulty is error accumulation, called exposure bias: during training the model sees real history, but at inference it sees its own generated, imperfect history, and image quality can collapse over time. CausVid (CVPR 2025) distills a bidirectional model into a 4-step causal model, running at 9.4 frames per second on a single GPU; Self Forcing (NeurIPS 2025) uses the model's own outputs as history during training, closing the gap between training and inference.","example":"Google DeepMind's Genie 3 generates frames autoregressively one at a time, with each new frame conditioned on an ever-growing history, and can run interactively in real time at 720p and 24 frames per second.","related":["Diffusion Forcing","Self Forcing","Interactive World Model","Video Generation Model","Exposure Bias","Key-Value Cache"]},{"id":"exposure-bias","category":"model","sec":8,"tier":3,"sources":[{"title":"Sequence Level Training with Recurrent Neural Networks (Ranzato et al., ICLR 2016)","url":"https://arxiv.org/abs/1511.06732"},{"title":"Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks (Bengio et al., 2015)","url":"https://arxiv.org/abs/1506.03099"},{"title":"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion (arXiv 2506.08009)","url":"https://arxiv.org/abs/2506.08009"}],"as_of":"","related_ids":["teacher-forcing","autoregressive-decoding","compounding-error","distribution-shift","self-forcing","diffusion-forcing"],"name":"Exposure Bias","alt":"曝光偏差（自回归误差累积）","abbr":"","aliases":["Train-Test Mismatch"],"one_liner":"The problem where a model sees ground truth during training but its own output at inference, so errors compound.","explanation":"Exposure bias is a classic problem in sequence generation models. Autoregressive models (models that predict the next element step by step) are usually trained with teacher forcing, feeding in the real history at every step; but at inference, the only history available is the model's own previous output. Because the model never saw its own mistakes during training, once it drifts off the training distribution, its errors compound further and further. Bengio and colleagues' scheduled sampling (2015, gradually switching training to use the model's own predictions as input) and Ranzato and colleagues' sequence-level training (2016) were early representative fixes. In embodied AI it shows up in two common forms: a behavior-cloning policy drifting further and further from the demonstrated trajectory once it deviates (often called compounding error or covariate shift), and an autoregressive video world model's image quality collapsing after generating dozens of frames in a row — methods like Self Forcing try to ease this by having the model continue generating from its own generated frames during training too.","example":"Given a video world model one current frame and asked to autoregressively predict the next 100 frames: during training it saw the real previous frame at every step, but at inference it sees its own slightly flawed generated frame, so objects tend to deform and the image tends to blur more and more the further out it goes.","related":["Teacher Forcing","Autoregressive Decoding","Compounding Error","Distribution Shift","Self Forcing","Diffusion Forcing"]},{"id":"diffusion-forcing","category":"model","sec":8,"tier":3,"sources":[{"title":"Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion (arXiv:2407.01392)","url":"https://arxiv.org/abs/2407.01392"},{"title":"Diffusion Forcing project page","url":"https://boyuan.space/diffusion-forcing/"}],"as_of":"","related_ids":["diffusion-model","autoregressive-video-generation","next-token-prediction","teacher-forcing","world-model","self-forcing"],"name":"Diffusion Forcing","alt":"扩散强制","abbr":"","aliases":[],"one_liner":"A causal diffusion sequence model trained by adding an independently random noise level to each token in the sequence.","explanation":"Diffusion Forcing was proposed in 2024 by MIT's Boyuan Chen, Vincent Sitzmann, Russ Tedrake, and colleagues, published at NeurIPS 2024. Sequence generation traditionally comes in two flavors: next-token prediction generates one token at a time with flexible length, but tends to drift off course over long rollouts; full-sequence diffusion denoises an entire span at once, which allows guiding the whole trajectory but fixes its length. Diffusion Forcing trains a causal model in which each token in the sequence gets its own independently random noise level; at inference, it can then generate frame by frame like an autoregressive model, extending past the training length, while also using guidance like a diffusion model to steer the whole trajectory toward a desired goal. It's been used for long video generation, planning, and robot control, and is a training scheme that autoregressive video generation and world-model work often reuses or compares against.","example":"In the paper's real-robot experiment, an arm has to swap two randomly placed pieces of fruit using a third slot as a buffer, which requires remembering the initial positions; Diffusion Forcing completed the task, while an imitation-learning baseline without memory failed.","related":["Diffusion Model","Autoregressive Video Generation","Next-Token Prediction","Teacher Forcing","World Model","Self Forcing"]},{"id":"interactive-world-model","category":"model","sec":8,"tier":3,"sources":[{"title":"Genie: Generative Interactive Environments (arXiv:2402.15391)","url":"https://arxiv.org/abs/2402.15391"},{"title":"Learning Interactive Real-World Simulators (UniSim, arXiv:2310.06114)","url":"https://arxiv.org/abs/2310.06114"},{"title":"Genie 3: A new frontier for world models（Google DeepMind 博客）","url":"https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/"}],"as_of":"2025-08","related_ids":["world-model","genie-3","unisim","latent-action-model","video-generation-model","world-model-based-policy-evaluation"],"name":"Interactive World Model","alt":"可交互世界模型","abbr":"","aliases":["Action-conditioned Video Model","Generative Interactive Environment"],"one_liner":"A world model that takes an action as input at every step and generates the next frame from it, usable as a neural simulator.","explanation":"An ordinary video generation model generates a whole clip at once from a piece of text, with no way to intervene partway through; an interactive world model instead takes an action at every step — a keypress, a robot control command, or a text event — and generates the following frame from it, effectively acting as a simulator implemented with a neural network. 2023's UniSim combines many kinds of data to learn how the world visually responds to human and robot actions, and trains policies inside it that deploy zero-shot to the real world; Google DeepMind's 2024 Genie (11 billion parameters) trains only on unlabeled internet video with no action labels, achieving frame-by-frame control through a latent action model (which automatically infers 'what action was taken' from consecutive frames); Genie 3, released in August 2025, can interact in real time at 720p and 24 frames per second while keeping the scene consistent for several minutes. For embodied AI, it can be used to evaluate policies, generate training data, or let an agent learn inside the generated world.","example":"DeepMind placed its SIMA agent inside a world generated by Genie 3: the agent issues navigation actions like moving forward or turning, Genie 3 generates the corresponding view in real time, and the agent uses it to work toward a given goal.","related":["World Model","Genie 3","UniSim","Latent Action Model","Video Generation Model","World-Model-based Policy Evaluation"]},{"id":"latent-world-model","category":"model","sec":8,"tier":3,"sources":[{"title":"Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv:1811.04551)","url":"https://arxiv.org/abs/1811.04551"},{"title":"DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (arXiv:2411.04983)","url":"https://arxiv.org/abs/2411.04983"},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv:2506.09985)","url":"https://arxiv.org/abs/2506.09985"}],"as_of":"2025-06","related_ids":["world-model","recurrent-state-space-model","dino-wm","v-jepa-2","joint-embedding-predictive-architecture","learning-in-imagination"],"name":"Latent World Model","alt":"隐空间世界模型","abbr":"","aliases":["Latent Dynamics Model"],"one_liner":"A world model that predicts 'what happens if this action is taken' in a compressed latent state instead of pixels.","explanation":"A world model predicts the future given the current state and an action. Predicting future images directly means generating a huge amount of pixel-level detail irrelevant to decision-making, which is slow and hard; a latent world model instead encodes observations into a low-dimensional latent state first, predicts only the next latent state within that latent space, and does planning or policy training there too. Hafner and colleagues' 2018 PlaNet introduced the Recurrent State-Space Model (RSSM), which mixes deterministic and stochastic components, and the later Dreamer series uses it to train policies 'in imagination.' A different line of work skips reconstructing pixels altogether: DINO-WM directly predicts the patch features of the pretrained vision model DINOv2; Meta's 2025 V-JEPA 2-AC trains an action-conditioned predictor on fewer than 62 hours of DROID robot video and achieves zero-shot pick-and-place on Franka arms in two different labs. The difficulty is that the latent state can't be inspected directly, and methods that skip pixel reconstruction also have to guard against representation collapse (all inputs getting encoded into nearly identical vectors).","example":"For pick-and-place, V-JEPA 2-AC encodes the goal image into a latent vector, predicts the outcome of several candidate actions inside latent space, and executes whichever one's predicted result lands closest to the goal — without ever generating an image.","related":["World Model","Recurrent State-Space Model","DINO-WM","V-JEPA 2","Joint-Embedding Predictive Architecture","Learning in Imagination"]},{"id":"recurrent-state-space-model","category":"model","sec":8,"tier":3,"sources":[{"title":"Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv 1811.04551)","url":"https://arxiv.org/abs/1811.04551"},{"title":"Mastering Diverse Domains through World Models (DreamerV3, arXiv 2301.04104)","url":"https://arxiv.org/html/2301.04104"},{"title":"Training Agents Inside of Scalable World Models (Dreamer 4, arXiv 2509.24527)","url":"https://arxiv.org/abs/2509.24527"}],"as_of":"2025-09","related_ids":["world-model","latent-world-model","dreamerv3","dreamer-4","learning-in-imagination","recurrent-neural-network"],"name":"Recurrent State-Space Model","alt":"循环状态空间模型","abbr":"RSSM","aliases":["RSSM"],"one_liner":"The core structure of the Dreamer world models, combining a deterministic recurrent state with a stochastic latent to predict the future.","explanation":"The Recurrent State-Space Model is a latent dynamics model Danijar Hafner and colleagues (at Google Brain, DeepMind, and others) proposed in the 2019 PlaNet paper, and it later became the core of the Dreamer through DreamerV3 world models. Its state has two parts: a deterministic recurrent hidden state h, updated by a recurrent network such as a GRU, responsible for remembering history; and a stochastic latent z, representing the uncertain information at the current moment. Given the previous state and an action, the model first updates h, then predicts the next z; during training, a separate encoder infers z from the real image as a target, and the model is also required to reconstruct the image and predict the reward. PlaNet's comparisons show that a purely deterministic or purely stochastic version underperforms the combination of both. With it, an agent can 'imagine' many steps into the future inside latent space to plan or train a policy, greatly reducing how much it needs to interact with the real environment. DreamerV3 changes z into a discrete categorical distribution; 2025's Dreamer 4 switches to a Transformer-based world model instead.","example":"DreamerV3 generates imagined trajectories with its RSSM world model and trains actor and critic networks on them, and the paper describes it as the first algorithm to mine diamonds from scratch in Minecraft with no human data.","related":["World Model","Latent World Model","DreamerV3","Dreamer 4","Learning in Imagination","Recurrent Neural Network"]},{"id":"joint-embedding-predictive-architecture","category":"model","sec":8,"tier":2,"sources":[{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)","url":"https://arxiv.org/abs/2506.09985"},{"title":"Meta AI: I-JEPA, the first AI model based on Yann LeCun's vision","url":"https://ai.meta.com/blog/yann-lecun-ai-model-i-jepa/"},{"title":"Wikipedia: Yann LeCun","url":"https://en.wikipedia.org/wiki/Yann_LeCun"}],"as_of":"2025-11","related_ids":["v-jepa-2","world-model","self-supervised-learning","latent-world-model","masked-autoencoder","representation-collapse"],"name":"Joint-Embedding Predictive Architecture","alt":"联合嵌入预测架构","abbr":"JEPA","aliases":["JEPA","I-JEPA","V-JEPA"],"one_liner":"A self-supervised architecture that predicts masked or future content in an abstract representation space instead of reconstructing raw pixels.","explanation":"JEPA is an architecture Yann LeCun proposed in his 2022 position paper ‘A Path Towards Autonomous Machine Intelligence’; Meta later built an image version, I-JEPA (2023), and a video version, V-JEPA. It encodes part of the input (the context), then has a predictor guess the representation of another part — a masked region, or a future frame — with the loss computed on the representation rather than on pixels. That means the model doesn't need to reconstruct hard-to-predict details like individual leaf texture, and can instead focus on semantic information such as objects and motion, though extra care is needed to keep the representation from collapsing. V-JEPA 2, released in June 2025, was pretrained on more than 1 million hours of video and then fine-tuned on fewer than 62 hours of robot video to produce the action-conditioned V-JEPA 2-AC, which can plan pick-and-place actions zero-shot on a Franka arm given a goal image. LeCun left Meta in November 2025 and founded AMI Labs to continue working on world models.","example":"For pick-and-place, V-JEPA 2-AC is given a goal image showing the cup already placed on the plate; the model plays out several candidate action sequences in representation space and executes whichever one lands closest to the goal image's predicted representation.","related":["V-JEPA 2","World Model","Self-Supervised Learning","Latent World Model","Masked Autoencoder","Representation Collapse"]},{"id":"representation-collapse","category":"model","sec":8,"tier":3,"sources":[{"title":"VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning (arXiv 2105.04906)","url":"https://arxiv.org/abs/2105.04906"},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)","url":"https://arxiv.org/html/2506.09985"}],"as_of":"2025-06","related_ids":["self-supervised-learning","joint-embedding-predictive-architecture","contrastive-learning","stop-gradient","exponential-moving-average","v-jepa-2"],"name":"Representation Collapse","alt":"表征坍缩","abbr":"","aliases":["Dimensional Collapse"],"one_liner":"When a self-supervised encoder outputs the same or nearly the same vector for every input, so the representation carries no information.","explanation":"Representation collapse is a classic failure mode in self-supervised learning and joint-embedding predictive models. If the training objective only requires the representations of two views of the same sample (or of the current moment and the future) to be close to each other, the easiest solution for the encoder is to output the same constant vector for every input: the loss is tiny, but the representation carries no information at all. A milder form is called dimensional collapse, where the vector only makes use of a handful of dimensions. Anti-collapse methods fall roughly into three categories: contrastive learning uses negative samples to push different images' representations apart; VICReg (Bardes, Ponce, LeCun, 2021) explicitly constrains the variance of each dimension and the covariance between dimensions; and BYOL and the JEPA family use a stop-gradient plus an exponential-moving-average (EMA) target encoder. In embodied AI, training video representations or world models in a JEPA-style latent space, such as Meta's V-JEPA 2, has to deal with the same problem.","example":"When training V-JEPA 2, the target representation for a masked video segment is computed by an EMA copy of the encoder's weights and has stop-gradient applied to it, which the paper explains is specifically meant to prevent representation collapse.","related":["Self-Supervised Learning","Joint-Embedding Predictive Architecture","Contrastive Learning","Stop-Gradient","Exponential Moving Average","V-JEPA 2"]},{"id":"4d-world-model","category":"model","sec":8,"tier":3,"sources":[{"title":"TesserAct: Learning 4D Embodied World Models (arXiv 2504.20995)","url":"https://arxiv.org/abs/2504.20995"},{"title":"TesserAct 论文 HTML 版","url":"https://arxiv.org/html/2504.20995"}],"as_of":"2025-04","related_ids":["world-model","4d-reconstruction","video-generation-model","inverse-dynamics-model","3d-vla","spatial-intelligence"],"name":"4D World Model","alt":"4D 世界模型","abbr":"","aliases":["4D Embodied World Model"],"one_liner":"A world model that predicts how a 3D scene changes over time — 3D space plus time.","explanation":"A world model predicts what the world will look like next, given the current observation and an action. Most video world models only generate 2D frames, with no depth or geometry, making it hard for a robot to read an object's precise position in space from them. A 4D world model instead predicts '3D space plus time': every frame carries geometric information, and the frames can be assembled into a 3D scene that evolves over time. The representative work is TesserAct, released in April 2025 by a UMass Amherst-led team: it fine-tunes the CogVideoX video generation model to jointly predict RGB, depth, and normal-vector video (RGB-DN), then reconstructs it into a temporally consistent 4D scene; actions are then computed by an inverse dynamics model (a network that infers the action from before-and-after states) from point-cloud features, producing 7-DOF robot-arm actions. It overlaps with 4D reconstruction, video generation, and spatial intelligence.","example":"Given the current view and a language instruction, TesserAct generates a future video with depth and normals, converts it into a 4D scene, encodes the resulting point cloud with PointNet, combines it with the instruction, and outputs a 7-DOF action.","related":["World Model","4D Reconstruction","Video Generation Model","Inverse Dynamics Model","3D VLA","Spatial Intelligence"]},{"id":"world-action-model","category":"model","sec":8,"tier":2,"sources":[{"title":"World Action Models are Zero-shot Policies (DreamZero, arXiv 2602.15922)","url":"https://arxiv.org/abs/2602.15922"},{"title":"Fast-WAM: Do World Action Models Need Test-time Future Imagination? (arXiv 2603.16666)","url":"https://arxiv.org/abs/2603.16666"}],"as_of":"2026-03","related_ids":["vision-language-action-model","world-model","video-generation-model","dreamzero","fast-wam","video-prediction-policy"],"name":"World Action Model","alt":"世界动作模型","abbr":"WAM","aliases":["WAM","Video-Action Model","VAM"],"one_liner":"A model that predicts future frames and robot actions together, using its prediction about the world to guide the action.","explanation":"The term World Action Model was formally introduced in NVIDIA's February 2026 DreamZero paper: any model that uses world-modeling ability (predicting future state) to predict actions counts as a WAM, a category that retroactively covers earlier video-and-action joint models like GR-1, UWM, and Cosmos Policy. The typical design starts from a pretrained video generation model as backbone, adds action input and output, and trains it to predict future video and actions together. Compared with a VLA built from a vision-language model, a WAM learns physical dynamics from video; DreamZero reported more than double the generalization to new tasks and environments versus the best VLA at the time. It isn't called a 'video-action model' because the predicted target could someday be touch or force instead of video. One open question is whether inference actually needs to generate future frames: Fast-WAM does the joint video-and-action prediction only during training and skips it at inference, reaching similar performance more than 4x faster.","example":"DreamZero uses Wan2.1's 14B image-to-video model as its backbone; given a camera view and a language instruction, it generates future video and actions at the same time, controlling a real robot in real time at 7Hz.","related":["Vision-Language-Action Model","World Model","Video Generation Model","DreamZero","Fast-WAM","Video Prediction Policy"]},{"id":"inference-latency","category":"model","sec":9,"tier":1,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)","url":"https://arxiv.org/html/2410.24164"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)","url":"https://arxiv.org/abs/2506.07339"}],"as_of":"2025-06","related_ids":["inference","action-chunking","asynchronous-inference","real-time-chunking","control-frequency","post-training-quantization"],"name":"Inference Latency","alt":"推理延迟","abbr":"","aliases":["Inference Delay","Model Latency"],"one_liner":"The time a model takes from receiving input to producing output, which determines how quickly a robot can react.","explanation":"Inference latency is the time a model takes, after receiving one frame of observation — image, joint state, instruction — to finish its forward computation and produce an action, usually measured in milliseconds. Large models are compute-heavy, and their latency often can't keep up with a robot's control frequency of tens to hundreds of hertz: the 7-billion-parameter OpenVLA produces an action at only about 6Hz on an RTX 4090; the π0 paper measured about 73 milliseconds for one full inference with 3 camera views on the same GPU, or about 86 milliseconds when computed on a separate computer and sent over Wi-Fi. Too much latency makes a robot pause between action segments, or fail to keep up with a moving object. Common fixes include action chunking, outputting dozens of steps at once; asynchronous inference, computing the next segment while the current one executes; quantization and inference acceleration; and splitting large and small models into layers that each run at their own frequency.","example":"π0 controls a 50Hz robot by outputting 50 steps of action per inference call, executing 25 of them, 0.5 seconds, before running inference again for the next segment, instead of running the large model at every single step.","related":["Inference","Action Chunking","Asynchronous Inference","Real-Time Chunking","Control Frequency","Post-Training Quantization"]},{"id":"asynchronous-inference","category":"model","sec":9,"tier":2,"sources":[{"title":"Asynchronous Inference (LeRobot 文档)","url":"https://huggingface.co/docs/lerobot/async"},{"title":"SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics (arXiv:2506.01844)","url":"https://arxiv.org/abs/2506.01844"},{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv:2506.07339)","url":"https://arxiv.org/abs/2506.07339"}],"as_of":"2025-06","related_ids":["inference-latency","action-chunking","real-time-chunking","policy-server","smolvla","lerobot"],"name":"Asynchronous Inference","alt":"异步推理","abbr":"","aliases":["Async Inference"],"one_liner":"The robot keeps executing the current action chunk while the model computes the next one, with no pause between.","explanation":"Asynchronous inference splits “the model computing an action” and “the robot executing an action” into two parallel processes. In synchronous inference, the robot has to stop and wait once it finishes executing an action chunk until the model computes the next one, and the pause grows more noticeable the bigger the model. The asynchronous approach instead sends the latest observation to the model once the action queue drops below some threshold, so the next chunk finishes computing before the current one runs out, with the overlapping portion merged according to some rule. Hugging Face introduced this mechanism alongside SmolVLA in June 2025, splitting LeRobot into a PolicyServer that runs the model and a RobotClient that controls the robot. The hard part is stitching chunks together: a new chunk is computed from a slightly earlier observation, so switching directly can cause a jump in motion, which is exactly what Physical Intelligence's real-time chunking (RTC) was designed to handle.","example":"In LeRobot, with each inference outputting 50 steps of action and a threshold of 0.5, once fewer than 25 steps remain in the queue, the client sends a new image to the server, and the overlapping portion is merged with a weighted average.","related":["Inference Latency","Action Chunking","Real-Time Chunking","Policy Server (Remote Inference)","SmolVLA","LeRobot"]},{"id":"real-time-chunking","category":"model","sec":9,"tier":2,"sources":[{"title":"Real-Time Execution of Action Chunking Flow Policies (arXiv 2506.07339)","url":"https://arxiv.org/abs/2506.07339"},{"title":"Training-Time Action Conditioning for Efficient Real-Time Chunking (arXiv 2512.05964)","url":"https://arxiv.org/abs/2512.05964"},{"title":"LeRobot Docs: Real-Time Chunking (RTC)","url":"https://huggingface.co/docs/lerobot/main/en/rtc"}],"as_of":"2026-09","related_ids":["action-chunking","asynchronous-inference","inference-latency","flow-matching","temporal-ensembling","pi0-5"],"name":"Real-Time Chunking","alt":"实时动作分块","abbr":"RTC","aliases":["RTC","Real-Time Execution of Action Chunking Flow Policies"],"one_liner":"A method that computes the next action chunk while still executing the current one, blended smoothly into what's already running.","explanation":"Real-Time Chunking was proposed in 2025 by Kevin Black and colleagues at Physical Intelligence as an inference-time algorithm (NeurIPS 2025). Computing one action chunk with a large model takes on the order of a hundred milliseconds; executing chunks synchronously leaves a pause between them, while executing them asynchronously risks a jump where the new chunk doesn't match the old one. RTC starts computing the next chunk while the current one is still executing: the few steps that will inevitably run during inference are 'frozen,' and the rest is treated as an inpainting problem — a guidance term is added during flow-matching denoising, with a soft mask that gradually relaxes the constraint, so the new chunk blends smoothly with the old one. It needs no retraining; the paper demonstrated on π0.5 that the robot could still strike a match and light a candle with more than 300 ms of latency. A training-time version of RTC followed in December 2025, which simulates latency during training and removes the need for guidance computation at inference time.","example":"Turning on the RTC setting for π0.5 in LeRobot, and setting the delay-step count to match measured inference time, means the robot no longer pauses between action chunks while executing them.","related":["Action Chunking","Asynchronous Inference","Inference Latency","Flow Matching","Temporal Ensembling","π0.5"]},{"id":"floating-point-operations","category":"model","sec":9,"tier":2,"sources":[{"title":"Scaling Laws for Neural Language Models (Kaplan et al., 2020)","url":"https://arxiv.org/abs/2001.08361"},{"title":"Scaling Laws for Neural Language Models（ar5iv 全文，含 C≈6N 推导）","url":"https://ar5iv.labs.arxiv.org/html/2001.08361"}],"as_of":"","related_ids":["parameter-count","scaling-law","flops-tflops","inference-latency","on-device-model","visual-token-pruning"],"name":"Floating-Point Operations (FLOPs)","alt":"浮点运算量（FLOPs）","abbr":"FLOPs","aliases":["FLOPs","FLOP"],"one_liner":"The number of floating-point additions and multiplications a computation takes — a measure of how computationally heavy a model is.","explanation":"FLOPs stands for floating-point operations: the total number of floating-point arithmetic operations a computation requires, used as a measure of a model's computational cost. The lowercase ‘s’ just marks the plural, which is a different thing from the uppercase FLOPS (‘per second’), a measure of hardware compute capacity. There is a widely used Transformer estimate, from OpenAI's 2020 scaling-law paper: processing one token in a forward pass through a model with N parameters takes about 2N floating-point operations, and about 6N once training's backward pass is included; total training compute is also often reported in FLOPs. For robots, FLOPs matters because it drives inference latency directly: on the same onboard chip, a policy with more FLOPs runs slower and forces a lower control frequency, which is why on-device deployment usually relies on shrinking the model and pruning visual tokens to cut FLOPs.","example":"Using the 2N rule of thumb, a 7-billion-parameter VLA processing a single token costs about 14 billion floating-point operations (14 GFLOPs) in its forward pass; feeding in 256 image tokens at once costs roughly 3.6 trillion operations for that part alone (not counting attention).","related":["Parameter Count (Model Size)","Scaling Law","FLOPS / TFLOPS (Floating-Point Operations per Second)","Inference Latency","On-device Model","Visual Token Pruning"]},{"id":"key-value-cache","category":"model","sec":9,"tier":2,"sources":[{"title":"Hugging Face Transformers: Cache strategies","url":"https://huggingface.co/docs/transformers/kv_cache"},{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/html/2410.24164"}],"as_of":"","related_ids":["self-attention","causal-attention","autoregressive-decoding","inference-latency","context-length","pi0"],"name":"Key-Value Cache","alt":"KV 缓存","abbr":"KV Cache","aliases":["KV Cache","KV Caching"],"one_liner":"Storing already-computed attention keys and values so later tokens can reuse them instead of recomputing.","explanation":"The KV cache is the standard trick for speeding up Transformer inference. In self-attention, each token produces a query (Q), key (K), and value (V) vector; during autoregressive generation, only one new token is added at each step, and the K and V vectors for all earlier tokens never change. Without caching, every step would have to recompute the entire preceding sequence; with caching, each step only computes the new token and then attends to the stored K and V vectors, which is much faster. The cost is GPU memory: the cache grows linearly with context length and often becomes a bottleneck for long contexts, which is why techniques like cache quantization and sliding windows exist. The KV cache matters just as much for VLA models: π0 caches the K and V vectors for the prefix formed by the image and language inputs, so its 10-step flow-matching denoising process only has to recompute the action part, making one full inference pass take about 73 milliseconds.","example":"One π0 inference pass: image encoding takes about 14 ms and the prefix forward pass about 32 ms, computed only once; the action expert's 10 denoising steps take about 27 ms total, each step reusing the same cached prefix KV.","related":["Self-Attention","Causal Attention","Autoregressive Decoding","Inference Latency","Context Length","π0"]},{"id":"speculative-decoding","category":"model","sec":9,"tier":3,"sources":[{"title":"Fast Inference from Transformers via Speculative Decoding (arXiv:2211.17192)","url":"https://arxiv.org/abs/2211.17192"},{"title":"Accelerating Large Language Model Decoding with Speculative Sampling (arXiv:2302.01318)","url":"https://arxiv.org/abs/2302.01318"},{"title":"Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance (arXiv:2507.22424)","url":"https://arxiv.org/abs/2507.22424"}],"as_of":"2026-03","related_ids":["autoregressive-decoding","inference-latency","parallel-decoding","key-value-cache","vision-language-action-model","openvla"],"name":"Speculative Decoding","alt":"投机解码","abbr":"","aliases":["Speculative Sampling"],"one_liner":"Having a small model quickly guess a few tokens, then having the large model verify them all in one parallel pass, with no change in output.","explanation":"Speculative decoding is a large-model inference speedup technique, proposed by Google's Leviathan and colleagues in late 2022 (ICML 2023), with DeepMind's Chen and colleagues independently proposing the same idea as 'speculative sampling' around the same time. An autoregressive model has to run the full large model for every single token it generates, which is slow. Speculative decoding instead has a cheap draft model guess several tokens in a row, then has the large model score all of them in one forward pass; a specific acceptance rule keeps the prefix of correct guesses and resamples from the first wrong position onward — so the output distribution exactly matches what the large model would have produced generating alone, with no retraining needed. Google measured a 2-3x speedup on T5-XXL. In embodied AI, VLAs that output discrete action tokens, such as OpenVLA, use it to cut inference latency too, and 2026 saw improved methods that incorporate kinematic information.","example":"Spec-VLA (EMNLP 2025) found that applying standard speculative decoding directly to a VLA gave limited speedup, so it relaxes the acceptance condition using the relative distance between action tokens, raising accepted length by 44% on OpenVLA for a 1.42x speedup with no drop in success rate.","related":["Autoregressive Decoding","Inference Latency","Parallel Decoding","Key-Value Cache","Vision-Language-Action Model","OpenVLA"]},{"id":"post-training-quantization","category":"model","sec":9,"tier":3,"sources":[{"title":"A White Paper on Neural Network Quantization (Nagel et al., arXiv 2106.08295)","url":"https://arxiv.org/abs/2106.08295"},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/html/2406.09246"},{"title":"GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (arXiv 2210.17323)","url":"https://arxiv.org/abs/2210.17323"}],"as_of":"2024-06","related_ids":["quantization-aware-training","pruning","on-device-edge-deployment","inference-latency","nvidia-tensorrt","numerical-precision-formats"],"name":"Post-Training Quantization","alt":"训练后量化","abbr":"PTQ","aliases":["PTQ","Offline Quantization"],"one_liner":"Converting a trained model's weights and activations directly to low-bit numbers after training, with no retraining needed.","explanation":"Post-training quantization is a model compression method: after a model is trained at normal precision (FP32, BF16, etc.), its weights, and sometimes its activations too, are converted to a low-bit representation such as INT8 or INT4. It usually only needs a small batch of calibration data to measure the value ranges, requires no labeled data, and involves no retraining. Qualcomm AI Research's quantization white paper summarizes that most models keep accuracy close to floating point at 8 bits with PTQ; pushing lower tends to hurt accuracy and needs methods like GPTQ, which use second-order information, or a switch to quantization-aware training. For embodied AI, it mainly serves deployment: VLAs routinely have billions of parameters, while a robot's onboard memory and compute are limited, and quantization directly cuts memory use and speeds up inference. Deployment tools such as TensorRT and ONNX Runtime both support PTQ.","example":"The OpenVLA paper runs its 7B model at 4-bit quantized inference, reaching 71.9% success on BridgeData V2 tasks versus 71.3% at BF16, while GPU memory drops from 16.8GB to 7.0GB.","related":["Quantization-Aware Training","Pruning","On-Device / Edge Deployment","Inference Latency","NVIDIA TensorRT","Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)"]},{"id":"quantization-aware-training","category":"model","sec":9,"tier":3,"sources":[{"title":"Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (arXiv 1712.05877)","url":"https://arxiv.org/abs/1712.05877"},{"title":"Gemma 3 QAT Models: Bringing state-of-the-Art AI to consumer GPUs (Google Developers Blog)","url":"https://developers.googleblog.com/en/gemma-3-quantized-aware-trained-state-of-the-art-ai-to-consumer-gpus/"},{"title":"BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation (arXiv 2506.07530)","url":"https://arxiv.org/abs/2506.07530"}],"as_of":"2025-06","related_ids":["post-training-quantization","pruning","knowledge-distillation","on-device-edge-deployment","mixed-precision-training","numerical-precision-formats"],"name":"Quantization-Aware Training","alt":"量化感知训练","abbr":"QAT","aliases":["QAT","Fake-Quantization Training"],"one_liner":"Simulating low-bit error during training itself, so the model adapts in advance and loses less accuracy after quantization.","explanation":"Quantization-aware training inserts 'fake quantization' operations during training or fine-tuning: the forward pass rounds weights and activations to low-bit values (such as INT8 or INT4) before using them in computation, while backpropagation treats the rounding as an identity function (a straight-through estimator) to work around its non-differentiability, letting the model learn to work under quantization error. Google's Jacob and colleagues' 2018 paper on integer-only inference is a representative example of this approach. Compared with post-training quantization, QAT needs training data and extra compute, but keeps accuracy noticeably better at 4 bits and below; a common recipe is to train normally first, then do a short QAT fine-tuning pass. In embodied AI, BitVLA uses a 'quantize then distill' form of quantization-aware training, guided by a full-precision teacher model, to compress its vision encoder to 1.58 bits, and the paper reports 11x less memory than OpenVLA-OFT at comparable performance.","example":"Google's April 2025 QAT release of Gemma 3 runs about 5,000 steps of QAT before quantizing to int4, cutting the 27B model's memory from 54GB at BF16 down to 14.1GB, with Google reporting 54% less perplexity loss from quantization than quantizing directly.","related":["Post-Training Quantization","Pruning","Knowledge Distillation","On-Device / Edge Deployment","Mixed-Precision Training","Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)"]},{"id":"pruning","category":"model","sec":9,"tier":3,"sources":[{"title":"Learning both Weights and Connections for Efficient Neural Networks (Han et al., arXiv 1506.02626)","url":"https://arxiv.org/abs/1506.02626"},{"title":"EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models (arXiv 2506.10100)","url":"https://arxiv.org/abs/2506.10100"}],"as_of":"2025-06","related_ids":["visual-token-pruning","post-training-quantization","quantization-aware-training","knowledge-distillation","on-device-edge-deployment","early-exit"],"name":"Pruning","alt":"剪枝","abbr":"","aliases":["Network Pruning"],"one_liner":"Deleting unimportant weights, channels, or layers from a model to make it smaller and faster.","explanation":"Pruning is a class of model compression methods: it identifies parts of a network that have little effect on the output and removes them. Deleting individual weights is called unstructured pruning; deleting whole channels, attention heads, or layers is called structured pruning. The idea traces back to LeCun and colleagues' 1989 Optimal Brain Damage; in 2015, Song Han and colleagues proposed a three-step 'train, prune small weights, retrain' recipe that cut AlexNet's parameters from 61 million to 6.7 million and compressed VGG-16 13x, with almost no drop in ImageNet accuracy. Unstructured pruning produces a sparse matrix, which needs specialized hardware or libraries to actually speed things up; structured pruning shrinks the matrix directly, making it easier to get a real speedup. In embodied AI, pruning is often combined with quantization and knowledge distillation to fit a large model onto a robot's onboard compute, and VLA-specific work has emerged that prunes redundant language layers or visual tokens.","example":"EfficientVLA, with no additional training, prunes functionally redundant layers from CogACT's language module, filters out unimportant visual tokens, and caches intermediate features in the diffusion action head, giving a 1.93x inference speedup on SIMPLER with only a 0.6-point drop in success rate.","related":["Visual Token Pruning","Post-Training Quantization","Quantization-Aware Training","Knowledge Distillation","On-Device / Edge Deployment","Early Exit"]},{"id":"visual-token-pruning","category":"model","sec":9,"tier":3,"sources":[{"title":"An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models (FastV, arXiv:2403.06764)","url":"https://arxiv.org/abs/2403.06764"},{"title":"EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models (arXiv:2506.10100)","url":"https://arxiv.org/abs/2506.10100"}],"as_of":"2026-09","related_ids":["visual-token","inference-latency","pruning","attention-mechanism","key-value-cache","on-device-edge-deployment"],"name":"Visual Token Pruning","alt":"视觉 token 剪枝","abbr":"","aliases":["Token Pruning","Visual Token Compression"],"one_liner":"Dropping or merging unimportant image tokens at inference time so vision-language and VLA models run faster.","explanation":"VLMs and VLAs cut each image into hundreds of visual tokens — vectors representing small image patches — before feeding them into the language model, and the count multiplies with multiple cameras or multiple frames. Since attention computation scales roughly quadratically with the number of tokens, this is a major source of inference latency. Visual token pruning keeps only a small set of important tokens, ranked by some importance score (commonly how much attention text or action tokens pay to them), and drops or merges the rest; many methods need no retraining and can be applied directly. The landmark method FastV (ECCV 2024) found that a model's attention to image tokens becomes very sparse after the second layer, so it prunes half of them after the shallow layers, cutting LLaVA-1.5-13B's computation by about 45% with almost no drop in performance. In robotics this directly affects control frequency: EfficientVLA combines visual token selection, layer skipping, and caching to speed up CogACT by 1.93×, and 2025–2026 saw a wave of pruning methods designed specifically for VLAs.","example":"In FastV, the image first passes through the language model's first two layers as usual; after that, only the half of image tokens with the highest attention scores continue into later layers. No weights change and no retraining is needed, yet LLaVA-1.5's inference computation drops noticeably.","related":["Visual Token","Inference Latency","Pruning","Attention Mechanism","Key-Value Cache","On-Device / Edge Deployment"]},{"id":"early-exit","category":"model","sec":9,"tier":3,"sources":[{"title":"BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks (arXiv:1709.01686)","url":"https://arxiv.org/abs/1709.01686"},{"title":"DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution (arXiv:2411.02359)","url":"https://arxiv.org/abs/2411.02359"}],"as_of":"","related_ids":["inference-latency","on-device-model","visual-token-pruning","pruning","speculative-decoding","mixture-of-experts"],"name":"Early Exit","alt":"早退机制","abbr":"","aliases":["Multi-exit Network","Dynamic Inference"],"one_liner":"Letting an easy input produce its output at a middle layer of the network, skipping the remaining layers' computation.","explanation":"A deep network normally runs every input through all of its layers, but many easy examples can already be judged correctly from shallow features. Early exit inserts several 'exits' (small classification or output heads) partway through the network; at inference, as soon as one exit is confident enough, the model outputs there and stops, and only hard examples run through the whole network. BranchyNet, from Harvard's Teerapittayanon and colleagues, is an early representative of this idea, later adopted for BERT and large language models. Robot control fits this pattern well, since the action is simple most of the time and only occasionally needs complex reasoning; Tsinghua's Gao Huang and colleagues proposed DeeR-VLA at NeurIPS 2024, turning a multimodal large model into a multi-exit structure that decides when to exit based on a compute, latency, and memory budget.","example":"DeeR-VLA reports on the CALVIN benchmark that it cuts the language model's compute by 5.2-6.5x and GPU memory by 2-6x, with essentially no drop in task performance.","related":["Inference Latency","On-device Model","Visual Token Pruning","Pruning","Speculative Decoding","Mixture of Experts"]},{"id":"on-device-model","category":"model","sec":9,"tier":2,"sources":[{"title":"Google DeepMind: Gemini Robotics On-Device brings AI to local robotic devices","url":"https://deepmind.google/discover/blog/gemini-robotics-on-device-brings-ai-to-local-robotic-devices/"},{"title":"Apple Machine Learning Research: Updates to Apple's On-Device and Server Foundation Language Models","url":"https://machinelearning.apple.com/research/apple-foundation-models-2025-updates"}],"as_of":"2025-07","related_ids":["on-device-edge-deployment","cloud-edge-device-collaboration","post-training-quantization","nvidia-jetson","gemini-robotics-on-device","inference-latency"],"name":"On-device Model","alt":"端侧模型","abbr":"","aliases":["Edge Model","On-device AI"],"one_liner":"A model that runs locally on a device's own chip — a robot, a phone — instead of relying on the cloud.","explanation":"An on-device model is deployed to run locally on an end device — for example, a robot's onboard computer (such as an NVIDIA Jetson), a phone chip, or a car's chip — as opposed to a model that runs inference in a cloud data center. Device compute, memory, and power are all limited, so on-device models are usually smaller and compressed using quantization, pruning, distillation, and similar techniques; Apple's 2025 on-device foundation model, for example, has about 3B parameters, with weights compressed to 2 bits each via quantization-aware training. For robots, running locally has real practical benefits: it still works without a network connection, avoids network round-trip latency, and keeps data from leaving the device. The tradeoff is that on-device models are usually less capable than large cloud models, so a common setup has a large cloud model handle slower planning and reasoning while a small on-device model outputs actions at high frequency — cloud-edge-device collaboration.","example":"Google DeepMind released Gemini Robotics On-Device in June 2025: this VLA can run locally on the robot with no network connection, adapts to new tasks with just 50 to 100 demonstrations, and has been validated on a bimanual Franka FR3 and the Apptronik Apollo humanoid.","related":["On-Device / Edge Deployment","Cloud-Edge-Device Collaboration","Post-Training Quantization","NVIDIA Jetson","Gemini Robotics On-Device","Inference Latency"]},{"id":"uncertainty-estimation","category":"model","sec":9,"tier":3,"sources":[{"title":"Wikipedia: Uncertainty quantification","url":"https://en.wikipedia.org/wiki/Uncertainty_quantification"},{"title":"Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles (arXiv:1612.01474)","url":"https://arxiv.org/abs/1612.01474"},{"title":"Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners (arXiv:2307.01928)","url":"https://arxiv.org/abs/2307.01928"}],"as_of":"","related_ids":["out-of-distribution","robustness","embodied-safety","failure-recovery","human-in-the-loop","knowno"],"name":"Uncertainty Estimation","alt":"不确定性估计","abbr":"","aliases":["Uncertainty Quantification","UQ"],"one_liner":"Having a model report how confident it is alongside its prediction, so it's clear when to stop or ask for help.","explanation":"Uncertainty estimation studies how to quantify how confident a model is in its own output. It's usually split into two kinds: aleatoric uncertainty comes from randomness in the data itself, such as sensor noise, and doesn't go away with more data; epistemic uncertainty comes from the model not having seen or learned something well, and can be reduced with more data. Common methods include deep ensembles (training several models and looking at how much they disagree), Monte Carlo dropout (randomly dropping units multiple times at inference and looking at the spread of results), and conformal prediction (producing a candidate set with a statistical coverage guarantee). For robots, this matters directly for safety: when facing an out-of-distribution scene, high uncertainty can trigger slowing down, stopping, asking a person to take over, or actively collecting more data; model-based reinforcement learning also often treats disagreement across an ensemble of models as an exploration signal.","example":"KnowNo (CoRL 2023) uses conformal prediction to measure a large-language-model planner's uncertainty: when the candidate-action set has more than one option — say, two bowls are on the table and the instruction doesn't say which one to pick up — the robot proactively asks a person, keeping task success guaranteed while asking for help as rarely as possible.","related":["Out-of-Distribution","Robustness","Embodied Safety","Failure Recovery","Human-in-the-Loop","KnowNo"]},{"id":"gaussian-process","category":"model","sec":9,"tier":3,"sources":[{"title":"Gaussian process - Wikipedia","url":"https://en.wikipedia.org/wiki/Gaussian_process"},{"title":"PILCO: A Model-Based and Data-Efficient Approach to Policy Search (Deisenroth & Rasmussen, ICML 2011)","url":"https://icml.cc/2011/papers/323_icmlpaper.pdf"}],"as_of":"","related_ids":["system-identification","model-based-reinforcement-learning","uncertainty-estimation","sample-efficiency","safe-reinforcement-learning","kalman-filter"],"name":"Gaussian Process","alt":"高斯过程","abbr":"GP","aliases":["GP","Gaussian Process Regression","GPR","Kriging"],"one_liner":"A method that puts a probability distribution directly over functions, giving a prediction with uncertainty at every point.","explanation":"A Gaussian process is a stochastic process: for any finite set of input points, the corresponding function values follow a joint (multivariate) normal distribution. It's fully specified by a mean function and a covariance function, also called a kernel, which describes how correlated the function values at two input points are. Used for regression, it lets you compute a predicted mean and variance at any new input given a small number of observations, giving uncertainty for free. The standard reference is Rasmussen and Williams's 2006 textbook 'Gaussian Processes for Machine Learning.' Its drawback is that computation scales as n³ with the number of data points n, so sparse approximations are needed once data gets large. In robotics it's mostly used where data is scarce: learning a dynamics model (such as Deisenroth and Rasmussen's 2011 PILCO, which learns control from scratch in just a handful of trials), or tuning control parameters with Bayesian optimization.","example":"PILCO uses a Gaussian process to learn the dynamics of a cart-pole, folding the model's uncertainty into its long-horizon predictions, and learns to control it using only a handful of real trials, where ordinary reinforcement learning often needs hundreds or thousands.","related":["System Identification","Model-Based Reinforcement Learning","Uncertainty Estimation","Sample Efficiency","Safe Reinforcement Learning","Kalman Filter"]},{"id":"interpretability","category":"model","sec":9,"tier":3,"sources":[{"title":"Mapping the Mind of a Large Language Model（Anthropic, 2024-05）","url":"https://www.anthropic.com/research/mapping-mind-language-model"},{"title":"Mechanistic interpretability for steering vision-language-action models (arXiv:2509.00328)","url":"https://arxiv.org/abs/2509.00328"}],"as_of":"2025-08","related_ids":["vision-language-action-model","large-language-model","neural-network","embodied-safety","uncertainty-estimation"],"name":"Interpretability","alt":"可解释性（机制可解释性）","abbr":"","aliases":["Mechanistic Interpretability","Explainable AI"],"one_liner":"Studying how a neural network's internals produce its output; mechanistic interpretability breaks that computation into human-understandable pieces.","explanation":"A deep network has hundreds of millions of parameters, and the computation from input to output is unreadable to a person. Interpretability research tries to answer 'why did the model produce this output,' with methods ranging from attention heatmaps and saliency maps to mechanistic interpretability: reverse-engineering the network, like a program, to find which concepts (features) it represents internally and how those features wire together into circuits that carry out a computation. One difficulty is that a single neuron often participates in representing several concepts at once; in May 2024, Anthropic used dictionary learning (a sparse decomposition method) to extract millions of features from Claude 3 Sonnet, including one that activates for the text and images of 'Golden Gate Bridge' across many languages. For robots, this matters for debugging and safety: in 2025 a Berkeley team found directions inside π0-FAST and OpenVLA corresponding to concepts like 'fast/slow' and 'high/low,' and adjusting these activations at inference time changes the robot's actions in real time, with no fine-tuning or reward signal needed.","example":"The Berkeley team ran π0-FAST on a UR5 arm carrying a toy: amplifying the internal activation associated with 'fast' made the arm move faster, and amplifying the one associated with 'low' lowered the peak height of the carrying trajectory, all without ever changing the model's weights.","related":["Vision-Language-Action Model","Large Language Model","Neural Network","Embodied Safety","Uncertainty Estimation"]},{"id":"openai-gpt-series","category":"named_model","sec":0,"tier":2,"sources":[{"title":"Products and applications of OpenAI - Text generation（GPT-n 系列发布表，Wikipedia）","url":"https://en.wikipedia.org/wiki/Products_and_applications_of_OpenAI"},{"title":"GPT-4o - Wikipedia","url":"https://en.wikipedia.org/wiki/GPT-4o"},{"title":"GPT-6 - Wikipedia","url":"https://en.wikipedia.org/wiki/GPT-6"}],"as_of":"2026-09","related_ids":["large-language-model","multimodal-large-language-model","openai","transformer","next-token-prediction","rekep"],"name":"OpenAI GPT Series","alt":"GPT 系列（GPT-4o / GPT-5）","abbr":"GPT","aliases":["Generative Pre-trained Transformer","GPT-4o","GPT-5"],"one_liner":"OpenAI's large language model family; from GPT-4o onward it handles images and audio directly, often used by robots for high-level planning.","explanation":"GPT stands for Generative Pre-trained Transformer, and here refers to the large language model family OpenAI has released since 2018: pretrained on massive text with next-token prediction, then fine-tuned and aligned. GPT-1 (2018, 117 million parameters), GPT-2 (2019, 1.5 billion), and GPT-3 (2020, 175 billion) successively showed that capability keeps improving with scale; GPT-4 (March 2023) began accepting image input; GPT-4o (May 2024, the “o” for “omni”) can process and generate text, images, and audio; GPT-5 launched on August 7, 2025, followed by iterations like 5.1 and 5.2; GPT-6 launched in September 2026, first opening an Astra version to paying users on September 4, then adding Sol and Luna versions on September 22. All of these are closed-source, accessible only through ChatGPT or the API. In embodied AI they're commonly used as an external “brain”: looking at images to understand a scene, breaking an instruction into steps, writing control code, or picking action primitives, then handing off to a low-level controller — the model itself doesn't output joint actions directly.","example":"ReKep first uses DINOv2 to mark numbered candidate keypoints on an image, then gives the annotated image and a language instruction to GPT-4o, which writes several Python constraint functions describing how the keypoints should relate to each other at each stage; an optimizer then solves for the end-effector's motion trajectory.","related":["Large Language Model","Multimodal Large Language Model","OpenAI","Transformer","Next-Token Prediction","ReKep"]},{"id":"flamingo","category":"named_model","sec":0,"tier":3,"sources":[{"title":"Flamingo: a Visual Language Model for Few-Shot Learning (arXiv 2204.14198)","url":"https://arxiv.org/abs/2204.14198"},{"title":"Tackling multiple tasks with a single visual language model (Google DeepMind 博客)","url":"https://deepmind.google/discover/blog/tackling-multiple-tasks-with-a-single-visual-language-model/"}],"as_of":"2022-04","related_ids":["vision-language-model","perceiver-resampler","cross-attention","few-shot","in-context-learning","roboflamingo"],"name":"Flamingo","alt":"Flamingo","abbr":"","aliases":["DeepMind Flamingo","a Visual Language Model for Few-Shot Learning"],"one_liner":"A 2022 DeepMind vision-language model that learns a new task from just a few image-text examples, an early landmark VLM.","explanation":"Flamingo is a vision-language model (VLM, a model that can look at images and read and write text at the same time) released by DeepMind in April 2022, published at NeurIPS 2022, with its largest version at 80 billion parameters. It freezes an already-trained vision encoder and language model (DeepMind's 70-billion-parameter Chinchilla) and adds two new kinds of module in between: a Perceiver Resampler that compresses any number of image features down to a fixed number of visual tokens, and gated cross-attention layers inserted into the language model so it can “look at” the image while generating text. It is trained on web data where images and text are interleaved, which lets it do few-shot learning the way large language models do: put a few image-text examples in the prompt, and it completes a new task with no parameter updates. Across 16 benchmarks, it beat prior few-shot methods using just 4 examples per task. The community's open-source reproduction, OpenFlamingo, was later used by RoboFlamingo as the backbone for a robot policy.","example":"Put two example animal photos with captions in the prompt, followed by a new photo, and Flamingo will write a caption for it in the same format — with no additional training required.","related":["Vision-Language Model","Perceiver Resampler","Cross-Attention","Few-shot","In-Context Learning","RoboFlamingo"]},{"id":"llava","category":"named_model","sec":0,"tier":2,"sources":[{"title":"Visual Instruction Tuning (arXiv 2304.08485)","url":"https://arxiv.org/abs/2304.08485"},{"title":"Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, arXiv 2310.03744)","url":"https://arxiv.org/abs/2310.03744"},{"title":"LLaVA 项目主页","url":"https://llava-vl.github.io/"}],"as_of":"2023-10","related_ids":["vision-language-model","multimodal-large-language-model","instruction-tuning","projector-connector","clip","prismatic-vlms"],"name":"LLaVA","alt":"LLaVA（视觉指令微调架构）","abbr":"LLaVA","aliases":["Large Language and Vision Assistant","Visual Instruction Tuning","LLaVA-1.5"],"one_liner":"A 2023 open-source multimodal model that connects a vision encoder to an LLM through a projection layer, then fine-tunes on instructions.","explanation":"LLaVA was proposed in April 2023 by Haotian Liu, Chunyuan Li, and colleagues at the University of Wisconsin-Madison, Microsoft Research, and Columbia University; the paper, “Visual Instruction Tuning,” was an oral presentation at NeurIPS 2023. The architecture is simple: a CLIP ViT-L/14 vision encoder extracts image features, which pass through a projection layer (mapping visual features into the language model's word-embedding space) and into a Vicuna large language model. Training has two stages: first the language model is frozen and only the projection layer is trained, to align the two modalities; then the whole thing is fine-tuned end to end. The key ingredient is data: GPT-4, given only text, generated 158,000 image-instruction pairs covering conversations, detailed descriptions, and complex reasoning. LLaVA-1.5, from October 2023, replaced the projection layer with a two-layer MLP, trained in about a day on a single 8-GPU A100 node using only public data, and set the state of the art across 11 benchmarks. “Vision encoder + projection layer + LLM + instruction tuning” became the common recipe for most later open-source VLMs, and many VLAs add an action output on top of this kind of VLM.","example":"When generating training data, GPT-4 never sees the image itself — only a text description of it and the object bounding-box coordinates — and from that it writes multi-turn questions, answers, and reasoning problems about the image; those Q&A pairs are then paired with the actual image to train LLaVA.","related":["Vision-Language Model","Multimodal Large Language Model","Instruction Tuning","Projector / Connector","CLIP","Prismatic VLMs"]},{"id":"pali-x","category":"named_model","sec":0,"tier":3,"sources":[{"title":"arXiv 2305.18565: PaLI-X","url":"https://arxiv.org/abs/2305.18565"},{"title":"RT-2 项目主页","url":"https://robotics-transformer2.github.io/"}],"as_of":"2023-07","related_ids":["rt-2","vision-language-model","palm-e","vision-transformer","vision-language-action-model","encoder-decoder"],"name":"PaLI-X","alt":"PaLI-X","abbr":"","aliases":["PaLI-X: On Scaling up a Multilingual Vision and Language Model"],"one_liner":"A roughly 55-billion-parameter multilingual vision-language model from Google, one of the two backbones behind RT-2.","explanation":"PaLI-X is a multilingual vision-language model (VLM) that Google Research released in May 2023, a scaled-up version of the earlier PaLI. Its vision encoder is ViT-22B, a 22-billion-parameter Vision Transformer; its language component is a 32-billion-parameter UL2 encoder-decoder; together they total roughly 55 billion parameters. The paper's main finding is that scaling up both the vision and language sides together keeps paying off, and training mixed prefix-completion and masked-token-completion objectives. After fine-tuning, PaLI-X set new state-of-the-art results on more than 15 benchmarks and showed emergent abilities it was never specifically trained for, such as complex object counting and object detection using non-English category names. In embodied AI, PaLI-X is best known as one of the two backbones behind RT-2: RT-2 fine-tuned PaLI-X (55B) and PaLM-E (12B) separately, training each on robot actions represented as text tokens, producing some of the earliest vision-language-action (VLA) models.","example":"RT-2-PaLI-X-55B: PaLI-X was co-fine-tuned on web-scale image-text data together with robot trajectory data, so it could look at an image, read an instruction, and directly output discretized action tokens.","related":["RT-2","Vision-Language Model","PaLM-E","Vision Transformer","Vision-Language-Action Model","Encoder-Decoder"]},{"id":"google-gemini","category":"named_model","sec":0,"tier":2,"sources":[{"title":"Gemini: A Family of Highly Capable Multimodal Models (arXiv 2312.11805)","url":"https://arxiv.org/abs/2312.11805"},{"title":"Gemini (language model) - Wikipedia","url":"https://en.wikipedia.org/wiki/Gemini_(language_model)"},{"title":"Gemini Robotics brings AI into the physical world (Google DeepMind blog)","url":"https://deepmind.google/discover/blog/gemini-robotics-brings-ai-into-the-physical-world/"}],"as_of":"2026-09","related_ids":["multimodal-large-language-model","native-multimodal","gemini-robotics","gemini-robotics-er","google-deepmind","vision-language-model"],"name":"Google Gemini","alt":"Gemini 系列（谷歌多模态大模型）","abbr":"","aliases":["Gemini"],"one_liner":"Google DeepMind's family of natively multimodal large models, also the foundation for robotics models such as Gemini Robotics.","explanation":"Gemini is Google DeepMind's family of multimodal large models, succeeding LaMDA and PaLM 2. The first version, Gemini 1.0, launched in December 2023 in three sizes — Ultra, Pro, and Nano. From the start of training it handles text, images, audio, video, and code together, rather than training a language model first and bolting on a vision module afterward — an approach often called “natively multimodal.” Iteration has been fast: 1.5 in 2024 (very long context), 2.0 and 2.5 from late 2024 into 2025, Gemini 3 in November 2025, and as of September 2026 several versions within the 3.x series. In embodied AI it's used two ways: directly, to look at images, understand a scene, break down tasks, and do high-level planning; and as a backbone for training robot models — Google's Gemini Robotics (a VLA) and Gemini Robotics-ER (embodied reasoning), both released in March 2025, are built on top of Gemini 2.0.","example":"Gemini Robotics adds “physical action” as a new output modality on top of Gemini 2.0: after seeing the camera feed and hearing an instruction, the model directly outputs actions that control the robot.","related":["Multimodal Large Language Model","Native Multimodal","Gemini Robotics","Gemini Robotics-ER","Google DeepMind","Vision-Language Model"]},{"id":"internvl","category":"named_model","sec":0,"tier":3,"sources":[{"title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks (arXiv 2312.14238)","url":"https://arxiv.org/abs/2312.14238"},{"title":"OpenGVLab/InternVL (GitHub)","url":"https://github.com/OpenGVLab/InternVL"}],"as_of":"2026-09","related_ids":["vision-language-model","multimodal-large-language-model","shanghai-artificial-intelligence-laboratory","internvl","internvla","agibot-go-1"],"name":"InternVL","alt":"书生 InternVL","abbr":"InternVL","aliases":["OpenGVLab InternVL","InternVL-Chat"],"one_liner":"Shanghai AI Lab's open-source vision-language model series, starting from a 6-billion-parameter vision encoder and scaling up from there.","explanation":"InternVL is an open-source vision-language model series from Shanghai AI Lab's OpenGVLab team, with the Chinese name “Shusheng Wanxiang.” The original paper was released in December 2023 and selected as an oral presentation at CVPR 2024; the idea was to scale up the vision encoder, InternViT, to 6 billion parameters, then progressively align it with a large language model, evaluated across 32 vision-language benchmarks. Later versions — 1.5, 2, 2.5, 3, and 3.5 — kept the same “vision encoder + MLP projection layer + large language model” structure, ranging in size from 1 billion to 241 billion total parameters, with code open-sourced under the MIT license; as of September 2026, the 3.5 series is the latest. In embodied AI, it's commonly used as the vision-language backbone for VLA (vision-language-action) models.","example":"AgiBot GO-1's latent planner uses InternVL2.5-2B as its backbone: it reads in multiple camera feeds and an instruction, first predicts latent action tokens, then hands them to an action expert to generate continuous actions.","related":["Vision-Language Model","Multimodal Large Language Model","Shanghai Artificial Intelligence Laboratory","InternVL","InternVLA (Shanghai AI Laboratory)","AgiBot GO-1"]},{"id":"molmo","category":"named_model","sec":0,"tier":3,"sources":[{"title":"Molmo (Ai2 blog)","url":"https://allenai.org/blog/molmo"},{"title":"Molmo and PixMo (arXiv 2409.17146)","url":"https://arxiv.org/abs/2409.17146"},{"title":"Molmo 2 (Ai2 blog)","url":"https://allenai.org/blog/molmo2"}],"as_of":"2025-12","related_ids":["vision-language-model","pointing","molmoact","open-weight-model","allen-institute-for-ai","clip"],"name":"Molmo (Ai2)","alt":"Molmo","abbr":"","aliases":["Molmo and PixMo","Molmo 2"],"one_liner":"An Ai2 vision-language model open-sourced with both weights and training data that can answer questions by “pointing” on the image.","explanation":"Molmo is an open vision-language model family released by the Allen Institute for AI (Ai2) in September 2024, in several sizes — MolmoE-1B, Molmo-7B-O, Molmo-7B-D, and Molmo-72B — using CLIP as the vision encoder and OLMo or Qwen2, respectively, for the language component. It has two distinguishing features. First, its training dataset, PixMo, is released alongside it, and the data was not obtained by distilling a closed-source model — PixMo's detailed image captions come from annotators describing images out loud, then transcribed and cleaned up. Second, it can “point”: it can output the 2D coordinates of an object in an image directly, useful for counting and localization. Pointing is practical for robots, since it can tell a robot where to grasp or where to place something, and Ai2's later MolmoAct was built on top of Molmo. Molmo 2, released in December 2025, extended these abilities to video understanding, spatiotemporal grounding, and object tracking.","example":"Ask Molmo “how many cups are in this image,” and it first places a point on each cup, then gives the total count.","related":["Vision-Language Model","Pointing","MolmoAct","Open-weight Model","Allen Institute for AI","CLIP"]},{"id":"mvp","category":"named_model","sec":0,"tier":3,"sources":[{"title":"arXiv 2203.06173: Masked Visual Pre-training for Motor Control","url":"https://arxiv.org/abs/2203.06173"},{"title":"arXiv 2210.03109: Real-World Robot Learning with Masked Visual Pre-training","url":"https://arxiv.org/abs/2210.03109"}],"as_of":"2022-10","related_ids":["pre-trained-visual-representation","masked-autoencoder","self-supervised-learning","r3m","vc-1","vision-transformer"],"name":"MVP","alt":"MVP（掩码视觉预训练）","abbr":"MVP","aliases":["Masked Visual Pre-training","Masked Visual Pre-training for Motor Control","Real-World Robot Learning with Masked Visual Pre-training"],"one_liner":"Uses a masked autoencoder to pretrain a vision encoder on huge amounts of natural images, then freezes it for robot use.","explanation":"MVP (Masked Visual Pre-training) refers to two papers from Jitendra Malik and Trevor Darrell's groups at UC Berkeley: Masked Visual Pre-training for Motor Control (Tete Xiao and colleagues), from March 2022, validated in simulation, and Real-World Robot Learning with Masked Visual Pre-training (Ilija Radosavovic and colleagues, CoRL 2022), from October 2022, extending it to the real robot. The method first self-supervised-pretrains a ViT vision encoder on web images and first-person video using a masked autoencoder (MAE, which hides most small patches of an image and has the network fill them back in), then freezes the encoder and trains only a small control module on top of it. Results show this representation beats CLIP, ImageNet-supervised pretraining, and training from scratch; a 307-million-parameter ViT trained on 4.5 million images keeps improving further still. Together with R3M and VC-1, it helped drive forward the direction of pretrained visual representations.","example":"Freeze an MVP-pretrained ViT to serve as the robot's eyes, and training only a small control head on top of it with a handful of demonstrations is enough to learn grasping on a new task.","related":["Pre-trained Visual Representation","Masked Autoencoder","Self-Supervised Learning","R3M","VC-1","Vision Transformer"]},{"id":"r3m","category":"named_model","sec":0,"tier":3,"sources":[{"title":"R3M (arXiv 2203.12601)","url":"https://arxiv.org/abs/2203.12601"}],"as_of":"2022-03","related_ids":["pre-trained-visual-representation","vc-1","mvp","vip","ego4d","time-contrastive-networks"],"name":"R3M","alt":"R3M","abbr":"R3M","aliases":["R3M: A Universal Visual Representation for Robot Manipulation"],"one_liner":"A general-purpose visual representation for robot manipulation, pretrained on first-person human videos.","explanation":"R3M was released by Suraj Nair, Chelsea Finn, Abhinav Gupta, and colleagues at Stanford and Meta AI in March 2022, published at CoRL 2022. Robot demonstrations are scarce, which makes it hard to train a good visual encoder from scratch. R3M instead pretrains an image encoder on Ego4D, a large dataset of first-person human videos, with a combined objective: time-contrastive learning (frames close together in time should have similar representations), aligning video with its language description, and an L1 penalty that keeps the representation sparse and compact. The encoder is then frozen and used as a perception module, with only a policy trained on top for the downstream task. Across 12 simulated manipulation tasks, this raises success rates more than 20 points over training from scratch, and more than 10 points over CLIP or MoCo representations. R3M represents the 'pretrained visual representation' line of work and is often compared with VC-1, MVP, and VIP.","example":"In a cluttered real apartment, a Franka arm using a frozen R3M encoder as its visual module learned to manipulate objects with just 20 demonstrations per task.","related":["Pre-trained Visual Representation","VC-1","MVP","VIP","Ego4D","Time-Contrastive Networks"]},{"id":"vip","category":"named_model","sec":0,"tier":3,"sources":[{"title":"VIP (arXiv 2210.00030)","url":"https://arxiv.org/abs/2210.00030"},{"title":"VIP project page","url":"https://sites.google.com/view/vip-rl"}],"as_of":"2023-03","related_ids":["r3m","vc-1","pre-trained-visual-representation","time-contrastive-networks","goal-conditioned-reinforcement-learning","ego4d"],"name":"VIP","alt":"VIP（价值隐式预训练）","abbr":"VIP","aliases":["Value-Implicit Pre-Training","VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training"],"one_liner":"A self-supervised pretraining method on human videos that produces both a visual representation and a dense reward at once.","explanation":"VIP was released by Jason Ma, Amy Zhang, and colleagues at Meta AI (FAIR) and the University of Pennsylvania in September 2022, published at ICLR 2023 (Spotlight). Robot reinforcement learning is often stuck on two things at once: no good visual features, and no easy-to-write reward function. VIP frames 'learning a representation from human video' as an offline goal-conditioned reinforcement learning problem, and derives a value-function objective that needs no action labels — essentially a form of implicit time-contrastive learning, where frames closer in time to completing the task sit closer to the goal in feature space. After pretraining on Ego4D first-person videos, it is used frozen: the reward is simply the change in feature-space distance between the current frame and a goal image, which supplies a dense reward for a wide range of simulated and real-robot tasks; on real robots, as few as about 20 trajectories are enough for few-shot offline reinforcement learning.","example":"Given a goal photo showing 'the drawer already closed,' VIP encodes every camera frame and the goal image into features; as the distance shrinks, the robot gets positive reward, learning to push the drawer shut without anyone writing a reward function by hand.","related":["R3M","VC-1","Pre-trained Visual Representation","Time-Contrastive Networks","Goal-Conditioned Reinforcement Learning","Ego4D"]},{"id":"vc-1","category":"named_model","sec":0,"tier":3,"sources":[{"title":"Where are we in the search for an Artificial Visual Cortex (arXiv 2303.18240)","url":"https://arxiv.org/abs/2303.18240"},{"title":"VC-1 project page","url":"https://eai-vc.github.io/"}],"as_of":"2023-03","related_ids":["pre-trained-visual-representation","masked-autoencoder","r3m","vip","ego4d","vision-transformer"],"name":"VC-1","alt":"VC-1","abbr":"VC-1","aliases":["Visual Cortex 1","Artificial Visual Cortex"],"one_liner":"Meta's visual encoder for embodied tasks, pretrained with masked autoencoding on more than 4,000 hours of first-person video.","explanation":"VC-1 is research released by Meta AI (FAIR) in March 2023, and was, at the time, the largest systematic evaluation of pretrained visual representations — off-the-shelf visual encoders meant to serve as a robot's 'eyes.' The authors first built CortexBench, containing 17 tasks spanning locomotion, navigation, dexterous manipulation, and mobile manipulation; they then combined more than 4,000 hours of first-person video from 7 sources with ImageNet and used a masked autoencoder (MAE, which masks out image patches and reconstructs them) to train ViTs of various sizes, the largest being ViT-L, which is VC-1. The conclusion: no single visual representation is best on every task, and scaling up data size and diversity only helps on average; once adapted to a specific task, VC-1 matches or beats the best previously known results across the board. Both the model and code are open-sourced.","example":"When building a robot-arm imitation-learning project, VC-1 can be used directly as a frozen image encoder, turning camera images into feature vectors on top of which a small policy network is trained.","related":["Pre-trained Visual Representation","Masked Autoencoder","R3M","VIP","Ego4D","Vision Transformer"]},{"id":"saycan","category":"named_model","sec":1,"tier":2,"sources":[{"title":"SayCan 项目主页","url":"https://say-can.github.io/"}],"as_of":"2022-08","related_ids":["affordance","language-grounding","llm-based-task-planning","inner-monologue","palm-e","value-function"],"name":"SayCan","alt":"SayCan","abbr":"","aliases":["Do As I Can, Not As I Say: Grounding Language in Robotic Affordances","PaLM-SayCan"],"one_liner":"Combines what a language model says is useful with what a value function says is achievable to pick each action.","explanation":"SayCan is work released by the Google Robotics team and Everyday Robots in April 2022. Large language models have broad commonsense knowledge, but don't know what the robot in front of them can actually do in its current situation; letting an LLM write a plan directly often produces steps that aren't achievable. SayCan instead equips the robot with a set of pretrained skills (such as “pick up the sponge” or “go to the table”); at each step, the language model scores how useful each skill would be toward completing the instruction, and a value function learned through reinforcement learning scores how likely each skill is to succeed from the current state (its affordance); the two scores are multiplied together, and the highest-scoring skill is executed, repeating until the task ends. Using PaLM in place of the original language model, PaLM-SayCan reached an 84% planning success rate and a 74% execution success rate across 101 instructions in a real kitchen. It's one of the founding works on using large models for high-level robot task planning.","example":"When a user says “I spilled my Coke, can you help me get something to clean it up,” SayCan selects, in sequence: find a sponge, pick up the sponge, bring it to you, done.","related":["Affordance","Language Grounding","LLM-based Task Planning","Inner Monologue","PaLM-E","Value Function"]},{"id":"socratic-models","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language (arXiv 2204.00598)","url":"https://arxiv.org/abs/2204.00598"},{"title":"Socratic Models project page","url":"https://socraticmodels.github.io/"}],"as_of":"2022-05","related_ids":["large-language-model","vision-language-model","zero-shot","saycan","code-as-policies","inner-monologue"],"name":"Socratic Models","alt":"Socratic Models（苏格拉底模型）","abbr":"SMs","aliases":["Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language"],"one_liner":"A framework that chains several off-the-shelf large models together zero-shot using natural language as the common interface for multimodal tasks.","explanation":"Socratic Models was released by Andy Zeng, Pete Florence, and colleagues at Google in April 2022. Different foundation models have different strengths: vision-language models (VLMs) understand images, large language models (LLMs) understand commonsense and reasoning, and audio models understand sound. Rather than fine-tuning any of them, Socratic Models uses natural language as a shared interface: one model's output is written into another model's prompt, letting them exchange information as if in conversation and compose new capabilities zero-shot. The paper demonstrates this on first-person video question answering, multimodal assistant dialogue, and robot perception and planning: a vision model first turns tabletop objects into text descriptions, then an LLM breaks the instruction down into a sequence of pick-and-place actions, which are handed to a pretrained language-conditioned policy for execution. Along with SayCan and Code as Policies from the same period, it represents an early approach to using large language models as robot planners.","example":"A user says 'put all the fruit in the bowl'; a vision model first lists an apple, a banana, and a bowl on the table, and a language model writes this out as 'pick up the apple and put it in the bowl; pick up the banana and put it in the bowl,' which a lower-level pick-and-place policy then executes step by step.","related":["Large Language Model","Vision-Language Model","Zero-shot","SayCan","Code as Policies","Inner Monologue"]},{"id":"inner-monologue","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Inner Monologue (arXiv 2207.05608)","url":"https://arxiv.org/abs/2207.05608"},{"title":"Inner Monologue project page","url":"https://innermonologue.github.io/"},{"title":"Inner Monologue (PMLR v205, CoRL 2022)","url":"https://proceedings.mlr.press/v205/huang23c.html"}],"as_of":"2022-12","related_ids":["saycan","llm-based-task-planning","success-detector","closed-loop-control","long-horizon-task","code-as-policies"],"name":"Inner Monologue","alt":"Inner Monologue","abbr":"","aliases":["Embodied Reasoning through Planning with Language Models"],"one_liner":"Feeds environment feedback back into a large language model as text, letting the robot adjust its plan as it goes.","explanation":"Inner Monologue was released in July 2022 by Wenlong Huang, Fei Xia, Brian Ichter, and colleagues at Robotics at Google, published at CoRL 2022. At the time, work such as SayCan already used large language models to break a high-level instruction into a sequence of skills, but the plan was fixed once made, and the model had no way of knowing when something went wrong during execution. Inner Monologue needs no extra training: it writes several kinds of feedback back into the LLM's prompt in natural language — whether a skill succeeded (success detection), what's in the scene (passive or active scene description), and any additional human instructions — forming an “inner monologue” that the LLM uses to decide its next step, retry, or revise the plan. Across three settings — simulated and real tabletop object rearrangement, and long-horizon mobile manipulation in a real kitchen — this closed-loop language feedback clearly raised the instruction-completion rate. It is an early representative example of using an LLM for closed-loop robot planning.","example":"For example, when the robot's grasp fails, a success detector writes back “action failed”; reading this, the LLM schedules a re-grasp attempt instead of just moving on to the next step.","related":["SayCan","LLM-based Task Planning","Success Detector","Closed-loop Control","Long-horizon Task","Code as Policies"]},{"id":"code-as-policies","category":"named_model","sec":1,"tier":2,"sources":[{"title":"Code as Policies (arXiv 2209.07753)","url":"https://arxiv.org/abs/2209.07753"},{"title":"Code as Policies 项目主页","url":"https://code-as-policies.github.io/"}],"as_of":"2022-09","related_ids":["large-language-model","saycan","llm-based-task-planning","voxposer","eureka","skill-primitive"],"name":"Code as Policies","alt":"代码即策略","abbr":"CaP","aliases":["CaP","Code as Policies: Language Model Programs for Embodied Control"],"one_liner":"Has a large language model write Python code directly, calling perception and control APIs to command a robot.","explanation":"Code as Policies was released by Robotics at Google in September 2022. Earlier work using LLMs to control robots, such as SayCan, mostly had the model pick the next step from a fixed list of skills. Code as Policies instead has an LLM — good at writing code — generate a Python program directly from a natural-language instruction: it calls perception APIs like object detection to get positions, uses NumPy to compute coordinates, then calls control primitives such as grasp or move, and can even write loops and conditionals. When it hits an undefined function, it recursively generates that function too, a process called hierarchical code generation. This lets it handle vague instructions that need spatial reasoning or concrete numeric values. The paper demonstrated this on tabletop manipulation, whiteboard drawing, and mobile robots, and the idea of an LLM writing code to drive a robot resurfaces later in VoxPoser and Eureka.","example":"Given the instruction “arrange the blocks in a horizontal line near the top,” the generated code first detects every block's position, computes a row of evenly spaced target points above the table, then calls the pick-and-place function for each one in turn.","related":["Large Language Model","SayCan","LLM-based Task Planning","VoxPoser","Eureka","Skill Primitive"]},{"id":"progprompt","category":"named_model","sec":1,"tier":3,"sources":[{"title":"arXiv 2209.11302: ProgPrompt","url":"https://arxiv.org/abs/2209.11302"},{"title":"ProgPrompt 项目主页","url":"https://progprompt.github.io/"}],"as_of":"2023-05","related_ids":["llm-based-task-planning","saycan","code-as-policies","inner-monologue","virtualhome-simulating-household-activities-via-programs","prompt-prompt-engineering"],"name":"ProgPrompt","alt":"ProgPrompt","abbr":"","aliases":["ProgPrompt: Generating Situated Robot Task Plans using Large Language Models"],"one_liner":"A method that prompts a large language model with Python-like code to generate robot task plans the robot can actually execute.","explanation":"ProgPrompt was released by Ishika Singh, Animesh Garg, and colleagues at the University of Southern California and NVIDIA in September 2022, published at ICRA 2023 with an extended version in Autonomous Robots. Asking a large language model (LLM) to write a plan directly in natural language often produces actions the robot cannot perform or references to objects that are not actually in the scene. ProgPrompt instead writes the prompt as code: it opens with import statements listing the available actions, gives the list of objects present in the scene, and provides a few example tasks written as Python functions (the function name is the task, the body is the steps), then has the LLM continue by writing the function for a new task. Steps in the generated plan are grouped with comments, similar to chain-of-thought reasoning, and assert statements check preconditions and trigger recovery actions when they fail. ProgPrompt achieved the best results of its time in the VirtualHome household simulator and was also validated on tabletop tasks with a real robot arm; it belongs to the same early wave of 'LLM as task planner' work as SayCan and Code as Policies.","example":"The prompt begins with 'from actions import walk, grab, putin, open, close', lists the scene's objects such as salmon and microwave, and gives one example function, 'throw_away_lime()'; the LLM then continues by writing out 'microwave_salmon()' step by step.","related":["LLM-based Task Planning","SayCan","Code as Policies","Inner Monologue","VirtualHome: Simulating Household Activities via Programs","Prompt / Prompt Engineering"]},{"id":"palm-e","category":"named_model","sec":1,"tier":2,"sources":[{"title":"PaLM-E 项目主页","url":"https://palm-e.github.io/"},{"title":"PaLM-E: An Embodied Multimodal Language Model (arXiv:2303.03378)","url":"https://arxiv.org/abs/2303.03378"}],"as_of":"2023-03","related_ids":["saycan","rt-2","multimodal-large-language-model","llm-based-task-planning","vision-transformer","language-grounding"],"name":"PaLM-E","alt":"PaLM-E","abbr":"PaLM-E","aliases":["PaLM-E-562B","PaLM-E: An Embodied Multimodal Language Model"],"one_liner":"Google's 2023 embodied multimodal large language model, which feeds images and robot-state estimates directly into PaLM.","explanation":"PaLM-E is the embodied multimodal language model Google and TU Berlin released in March 2023; its largest version, PaLM-E-562B, has 562 billion parameters and combines the PaLM language model with a ViT vision encoder. It encodes continuous inputs like images and estimated robot state into vectors, interleaves them with text tokens into a single “multimodal sentence,” and feeds that into the language model. Its output is a mid-level plan in the form of text, which is then handed off to low-level skill policies to execute — the model itself doesn't output motor commands directly. The paper found positive transfer from training on web image-text data together with robot data, and the 562B version also achieved state-of-the-art results on OK-VQA visual question answering at the time. It's a landmark example of using a large model for robot task planning, and a predecessor of VLAs like RT-2.","example":"Given a kitchen photo and the instruction “bring me the chips from the drawer,” PaLM-E generates a sequence of substeps — go to the drawer, open the drawer, take out the chips — which a low-level policy then executes one by one.","related":["SayCan","RT-2","Multimodal Large Language Model","LLM-based Task Planning","Vision Transformer","Language Grounding"]},{"id":"embodiedgpt","category":"named_model","sec":1,"tier":3,"sources":[{"title":"EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought (arXiv 2305.15021)","url":"https://arxiv.org/abs/2305.15021"},{"title":"NeurIPS 2023 论文页","url":"https://papers.nips.cc/paper_files/paper/2023/hash/4ec43957eda1126ad4887995d05fae3b-Abstract-Conference.html"}],"as_of":"2023-12","related_ids":["embodied-chain-of-thought","chain-of-thought","egocentric-video","ego4d","llm-based-task-planning","franka-kitchen"],"name":"EmbodiedGPT","alt":"EmbodiedGPT","abbr":"EmbodiedGPT","aliases":["Vision-Language Pre-Training via Embodied Chain of Thought"],"one_liner":"A 2023 embodied multimodal model from HKU and Shanghai AI Lab that uses chain of thought to generate step-by-step plans.","explanation":"EmbodiedGPT was released in May 2023 by the University of Hong Kong, Shanghai AI Lab, and others, published at NeurIPS 2023, and is one of the early landmark works applying large models to embodied planning. The authors selected clips from Ego4D first-person video and annotated them, in chain-of-thought form, as step-by-step sub-goals, building a planning dataset called EgoCOT; they then used prefix tuning to adapt a 7B language model to this kind of data, so it outputs a step-by-step plan after seeing an image. The key design is treating the plan the language model generates as a query, used to extract task-relevant features from the image, which are handed to a low-level policy network — closing the loop between high-level planning and low-level control.","example":"On the Franka Kitchen and Meta-World simulated control tasks, EmbodiedGPT's success rate is 1.6 times and 1.3 times that of a BLIP-2 baseline fine-tuned on Ego4D, respectively.","related":["Embodied Chain-of-Thought","Chain-of-Thought","Egocentric Video","Ego4D","LLM-based Task Planning","Franka Kitchen"]},{"id":"voyager","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Voyager (arXiv 2305.16291)","url":"https://arxiv.org/abs/2305.16291"},{"title":"Voyager 项目主页","url":"https://voyager.minedojo.org/"},{"title":"Voyager (OpenReview, TMLR)","url":"https://openreview.net/forum?id=ehfRiF0R3a"}],"as_of":"2024-03","related_ids":["embodied-agent","large-language-model","minecraft-environments","code-as-policies","curriculum-learning","continual-learning"],"name":"Voyager","alt":"Voyager","abbr":"","aliases":["MineDojo Voyager","Voyager: An Open-Ended Embodied Agent with Large Language Models"],"one_liner":"A GPT-4-driven, lifelong-learning agent in Minecraft that sets its own goals and builds up a library of skills by writing code.","explanation":"Voyager was released in May 2023 by Guanzhi Wang, Linxi Fan, Yuke Zhu, Anima Anandkumar, and colleagues at NVIDIA, Caltech, UT Austin, Stanford, and Arizona State University, later accepted at TMLR, with code open-sourced. It is an agent that runs in Minecraft without fine-tuning any model parameters, working purely by calling GPT-4. It has three core components: an automatic curriculum that proposes the next exploration goal based on the current state; a skill library that stores validated, successful behaviors as executable code, retrievable and composable by description; and iterative prompting, which feeds environment feedback, execution errors, and self-checks back to GPT-4 to revise the code. Because skills accumulate as code, they are easy to reuse and this also mitigates catastrophic forgetting. Compared with prior methods, Voyager obtains 3.3 times more distinct items, travels 2.3 times farther, and unlocks key tech-tree milestones up to 15.3 times faster, making it a representative example of the 'large model plus code skill library' style of embodied agent.","example":"After learning to craft a workbench and a wooden pickaxe, Voyager stores that code in its skill library; dropped into a brand-new Minecraft world, it directly retrieves and composes these skills to tackle new tasks.","related":["Embodied Agent","Large Language Model","Minecraft Environments","Code as Policies","Curriculum Learning","Continual Learning"]},{"id":"knowno","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners (arXiv 2307.01928)","url":"https://arxiv.org/abs/2307.01928"},{"title":"KnowNo 项目页","url":"https://robot-help.github.io/"}],"as_of":"2023-11","related_ids":["large-language-model","llm-based-task-planning","uncertainty-estimation","hallucination","human-in-the-loop","saycan"],"name":"KnowNo","alt":"KnowNo（会求助的机器人）","abbr":"","aliases":["Robots That Ask For Help","Uncertainty Alignment for LLM Planners"],"one_liner":"Has an LLM planner ask a person when uncertain, using conformal prediction to give a statistical success guarantee.","explanation":"KnowNo was released in July 2023 by Princeton University and Google DeepMind, and won the CoRL 2023 best student paper award. When a large language model plans robot tasks, an ambiguous instruction often leads it to confidently produce the wrong plan. KnowNo turns each planning step into a multiple-choice question: it first has the model list several candidate actions, then uses conformal prediction (a method that needs only a small amount of calibration data and gives a statistical coverage guarantee) to narrow those down to a candidate set, guaranteeing the correct option falls within that set with a probability the user sets, such as 80%. If only one option remains in the set, the robot executes it directly; if several remain, it stops and asks a person. It requires no fine-tuning of the model, and was validated on mobile manipulation, tabletop rearrangement, and bimanual manipulation, keeping the number of times it has to ask for help as low as possible while still guaranteeing success.","example":"In a mobile-manipulation setting, the user says “put the chips in the drawer,” but there's more than one bag of chips and more than one drawer present; multiple options remain in the candidate set, so the robot stops to ask for clarification before acting.","related":["Large Language Model","LLM-based Task Planning","Uncertainty Estimation","Hallucination","Human-in-the-Loop","SayCan"]},{"id":"sayplan","category":"named_model","sec":1,"tier":3,"sources":[{"title":"SayPlan (arXiv 2307.06135)","url":"https://arxiv.org/abs/2307.06135"},{"title":"SayPlan project page","url":"https://sayplan.github.io/"}],"as_of":"2023-09","related_ids":["3d-scene-graph","llm-based-task-planning","saycan","long-horizon-task","replanning","mobile-manipulation"],"name":"SayPlan","alt":"SayPlan","abbr":"","aliases":["SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning"],"one_liner":"A method that uses a 3D scene graph to let a large language model do long-horizon task planning across a large, multi-floor space.","explanation":"SayPlan was released in July 2023 by researchers at the QUT Centre for Robotics, the University of Adelaide, and CSIRO's Data61, an oral-presentation paper at CoRL 2023. When using a large language model for robot task planning, writing every room and object in an entire building into the prompt quickly exceeds the context length. SayPlan represents the environment as a hierarchical 3D scene graph (floors, rooms, objects, and their states); the model is first shown only a collapsed, high-level version of this structure and performs a semantic search by 'expanding' and 'collapsing' nodes, keeping only the sub-graph relevant to the task. The actual path planning is left to classical algorithms like Dijkstra's, with the model only responsible for high-level steps. The generated plan is first checked in a scene-graph simulator, and infeasible steps — such as putting something into a cabinet that was never opened — are fed back to the model for revision. The test environments went up to 3 floors, 36 rooms, and 140 objects, with the final plan executed on a real mobile manipulation robot.","example":"For the instruction 'I'm hungry, bring some food to my desk,' SayPlan expands the kitchen in the scene graph to find the fridge and some food, generates a plan of 'walk to the fridge, open the fridge, take out an apple, walk to the desk, put it down,' and confirms in the simulator that every step is feasible.","related":["3D Scene Graph","LLM-based Task Planning","SayCan","Long-horizon Task","Replanning","Mobile Manipulation"]},{"id":"3d-llm","category":"named_model","sec":1,"tier":3,"sources":[{"title":"3D-LLM: Injecting the 3D World into Large Language Models (arXiv 2307.12981)","url":"https://arxiv.org/abs/2307.12981"},{"title":"3D-LLM 代码仓库 (GitHub)","url":"https://github.com/UMass-Foundation-Model/3D-LLM"}],"as_of":"2023-12","related_ids":["3d-vla","multimodal-large-language-model","spatial-reasoning","3d-visual-grounding","leo","scannet"],"name":"3D-LLM","alt":"3D-LLM","abbr":"","aliases":["3D-LLM: Injecting the 3D World into Large Language Models"],"one_liner":"A large language model that takes 3D scene features directly as input to answer spatial questions and break down tasks.","explanation":"3D-LLM was proposed in July 2023 by Yining Hong, Chuang Gan, and colleagues from UCLA, UMass Amherst, MIT, and other institutions, a NeurIPS 2023 spotlight. LLMs and vision-language models only handle text or 2D images, and struggle with inherently three-dimensional concepts like spatial relationships and layout. 3D-LLM renders a 3D scene from multiple viewpoints into images, extracts features with a 2D feature extractor, and maps them back onto 3D points to get semantically rich 3D features, which then feed into an off-the-shelf vision-language model such as BLIP-2 for training; it adds a 3D localization mechanism to help the model understand position. The authors designed three prompting methods and collected more than 300,000 3D-language data points covering scene captioning, 3D question answering, task decomposition, localization, and navigation. It beat the previous best method's BLEU-1 score on ScanQA by 9%. It's one of the earlier works connecting 3D scenes to large models, and later work like 3D-VLA and LEO continues this direction.","example":"Given a 3D scan of a room, 3D-LLM can answer “which side of the sofa is the fridge on,” or break down “make breakfast” into steps like walking to the kitchen, opening the fridge, and taking out eggs.","related":["3D VLA","Multimodal Large Language Model","Spatial Reasoning","3D Visual Grounding","LEO (BIGAI)","ScanNet"]},{"id":"leo","category":"named_model","sec":1,"tier":3,"sources":[{"title":"An Embodied Generalist Agent in 3D World (arXiv 2311.12871)","url":"https://arxiv.org/abs/2311.12871"},{"title":"LEO 项目页","url":"https://embodied-generalist.github.io/"}],"as_of":"2024-07","related_ids":["3d-vla","3d-llm","beijing-institute-for-general-artificial-intelligence","embodied-agent","object-centric-representation","lora"],"name":"LEO (BIGAI)","alt":"LEO（3D 具身通才智能体）","abbr":"LEO","aliases":["An Embodied Generalist Agent in 3D World"],"one_liner":"A 3D embodied generalist model from BIGAI that understands 3D scenes and can answer questions, navigate, and manipulate objects.","explanation":"LEO was proposed by the Beijing Institute for General Artificial Intelligence (BIGAI) together with Peking University, Carnegie Mellon University, and Tsinghua University, released in November 2023 and published at ICML 2024. At the time, most multimodal large models only handled 2D images, and struggled with tasks defined in a 3D scene. LEO concatenates first-person images, object-centric 3D tokens (each object's point cloud encoded by PointNet++, with a spatial Transformer then modeling relationships between objects), and a text instruction into a single sequence, fed into a Vicuna-7B fine-tuned with LoRA, with actions also output as discrete tokens. Training has two stages: first 3D vision-language alignment, then 3D vision-language-action instruction tuning, with the data generated with the help of a large model. It can do 3D captioning, question answering, embodied reasoning, navigation, and manipulation, and is an early representative example of a 3D VLA.","example":"Given a 3D scan of a room, and asked “what's on the table next to the sofa,” LEO answers in text; for object navigation, it looks at first-person images and outputs actions like moving forward or turning left, step by step, to find the target.","related":["3D VLA","3D-LLM","Beijing Institute for General Artificial Intelligence","Embodied Agent","Object-centric Representation","LoRA"]},{"id":"language-to-rewards","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Language to Rewards for Robotic Skill Synthesis (arXiv 2306.08647)","url":"https://arxiv.org/abs/2306.08647"},{"title":"Language to Rewards 项目页","url":"https://language-to-reward.github.io/"}],"as_of":"2023-11","related_ids":["reward-function","model-predictive-control","large-language-model","code-as-policies","eureka","mujoco"],"name":"Language to Rewards","alt":"Language to Rewards（L2R）","abbr":"L2R","aliases":["L2R","Language to Rewards for Robotic Skill Synthesis"],"one_liner":"Has a large language model translate plain instructions into a reward function, handed to a real-time optimizer to generate robot motion.","explanation":"Language to Rewards was released in June 2023 by Google DeepMind, an oral presentation at CoRL 2023. Having a large model output actions directly, or call preset skills, makes it hard to produce low-level motions like “stand up” or “moonwalk.” L2R instead has the large model write a reward function: a reward translator first expands the instruction into a description of the motion, then writes it as reward code (a set of objectives and weights); a motion controller, MuJoCo MPC (a real-time optimizer based on model predictive control), then solves for the action that satisfies that reward, replanning in real time as the user adds corrections. Across 17 tasks on a simulated quadruped and a dexterous robot hand, it completed 90%, versus 50% for a baseline built on preset skill primitives; it also demonstrated pushing an object on a real robot arm. It belongs to the same “have a large model write the reward” line of work as Eureka.","example":"The user tells a simulated quadruped to “walk backward like doing the moonwalk,” watches the result, then adds a few corrections; after a few rounds of this back-and-forth, the robot has learned to moonwalk.","related":["Reward Function","Model Predictive Control","Large Language Model","Code as Policies","Eureka","MuJoCo (Multi-Joint dynamics with Contact)"]},{"id":"eureka","category":"named_model","sec":1,"tier":2,"sources":[{"title":"Eureka (arXiv 2310.12931)","url":"https://arxiv.org/abs/2310.12931"}],"as_of":"2024-04","related_ids":["reward-function","reward-engineering","large-language-model","dreureka","code-as-policies","isaac-gym"],"name":"Eureka","alt":"Eureka","abbr":"","aliases":["Eureka: Human-Level Reward Design via Coding Large Language Models"],"one_liner":"Has GPT-4 write reward-function code and repeatedly refine it, automating the reward design that reinforcement learning depends on.","explanation":"Eureka was released in October 2023 by NVIDIA together with the University of Pennsylvania, Caltech, and UT Austin, and published at ICLR 2024. Reinforcement learning's performance depends heavily on its reward function, and writing one by hand is slow and relies on expert intuition. Eureka gives GPT-4 the environment's source code and a task description and has it write several candidate reward functions in one pass; each is used to train a policy in parallel GPU simulation, and statistics from training on each reward term are fed back to the model (“reward reflection”) so it can rewrite an improved version for the next round, iterating like an evolutionary search. Across 10 robot embodiments and 29 open-source reinforcement-learning environments, its rewards beat human expert-designed ones on 83% of tasks, for an average normalized improvement of 52%. A follow-up, DrEureka, applies the same idea to sim-to-real transfer.","example":"Using a reward Eureka designed, a simulated Shadow dexterous hand learned to spin a pen rapidly and continuously between its fingers.","related":["Reward Function","Reward Engineering","Large Language Model","DrEureka","Code as Policies","Isaac Gym"]},{"id":"dreureka","category":"named_model","sec":1,"tier":3,"sources":[{"title":"DrEureka: Language Model Guided Sim-To-Real Transfer (arXiv 2406.01967)","url":"https://arxiv.org/abs/2406.01967"},{"title":"DrEureka 项目主页","url":"https://eureka-research.github.io/dr-eureka/"},{"title":"DrEureka 代码仓库（GitHub）","url":"https://github.com/eureka-research/DrEureka"}],"as_of":"2024-06","related_ids":["eureka","sim-to-real-transfer","domain-randomization","reward-function","reward-engineering","large-language-model"],"name":"DrEureka","alt":"DrEureka","abbr":"","aliases":["Language Model Guided Sim-To-Real Transfer"],"one_liner":"Uses a large language model to automatically write reward functions and domain-randomization ranges, transferring a simulation-trained policy to a real robot.","explanation":"DrEureka is a method proposed in June 2024 by the University of Pennsylvania, NVIDIA, and UT Austin, published at RSS 2024, extending the same team's earlier Eureka (which used a large language model to write reward functions). Moving a policy trained in simulation onto a real robot usually requires a person to repeatedly hand-tune the reward function and the domain-randomization ranges (randomizing physical parameters like friction and mass during training so the policy can tolerate real-world discrepancies). DrEureka automates this in three steps: first, a large language model generates a reward function with safety constraints and trains a policy in simulation with it; then that policy is tested across a range of physical parameters to find the range over which it still works well (reward-aware physics priors, RAPP); finally, the large language model writes a domain-randomization configuration within that range, a final policy is trained on it, and that policy is deployed directly. The simulator used is Isaac Gym, and the policy takes only proprioceptive input.","example":"On the Unitree Go1 quadruped, the policy DrEureka trained let the robot dog balance and walk on a yoga ball, with no manual, repeated parameter tuning during training; on standard quadruped forward-walking and dexterous-hand cube-rotation tasks, it performed on par with or better than hand-designed solutions.","related":["Eureka","Sim-to-Real Transfer","Domain Randomization","Reward Function","Reward Engineering","Large Language Model"]},{"id":"voxposer","category":"named_model","sec":1,"tier":2,"sources":[{"title":"VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models (arXiv 2307.05973)","url":"https://arxiv.org/abs/2307.05973"},{"title":"VoxPoser 项目页","url":"https://voxposer.github.io/"}],"as_of":"2023-11","related_ids":["llm-based-task-planning","code-as-policies","affordance","intermediate-representation","rekep","motion-planning"],"name":"VoxPoser","alt":"VoxPoser","abbr":"","aliases":["VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models"],"one_liner":"A zero-shot method where a large model writes code to build a 3D value map, which a planner turns into a trajectory.","explanation":"VoxPoser was released in July 2023 by Fei-Fei Li and Jiajun Wu's groups at Stanford, working with Yunzhu Li at UIUC, and received an oral presentation at CoRL 2023, with Wenlong Huang as first author. Large language models know a lot of commonsense, but not where in 3D space an action should land. VoxPoser has an LLM write code, from an instruction, that calls a vision-language model to locate relevant objects, then assembles an “affordance map” (where to go) and a “constraint map” (where to avoid) on a voxel grid (space cut into small cubes); these maps become cost functions handed to a motion planner, which computes the end-effector's trajectory. The whole process needs no task-specific policy training, so it's zero-shot, and it replans in closed loop during execution, handling objects that get moved. It's a landmark example of the “large model plus intermediate representation plus classical planning” approach, and later work like ReKep continues in the same direction.","example":"Given the instruction “open the top drawer, careful of the vase nearby,” the LLM-written code first locates the top drawer's handle and marks the area around it with high value; it then locates the vase and marks the nearby region as high cost; the planner uses this to compute a trajectory to the handle that avoids the vase.","related":["LLM-based Task Planning","Code as Policies","Affordance","Intermediate Representation","ReKep","Motion Planning"]},{"id":"f3rm","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation (arXiv 2308.07931)","url":"https://arxiv.org/abs/2308.07931"},{"title":"F3RM 项目主页","url":"https://f3rm.github.io/"}],"as_of":"2023-11","related_ids":["distilled-feature-fields","neural-radiance-fields","clip","lerf","open-vocabulary","few-shot"],"name":"F3RM","alt":"F3RM","abbr":"F3RM","aliases":["Feature Fields for Robotic Manipulation","Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation"],"one_liner":"A 2023 MIT method that distills 2D features like CLIP into a 3D field, letting a robot grasp objects from language instructions with few demonstrations.","explanation":"F3RM (Feature Fields for Robotic Manipulation) comes from William Shen, Ge Yang, Phillip Isola, and colleagues at MIT CSAIL, published at CoRL 2023. 2D image models like CLIP and DINO understand semantics but have no notion of an object's precise 3D position and shape, and robotic grasping can't do without geometry. F3RM first takes multiple photos of a scene and trains a neural radiance field (NeRF, a method that reconstructs a 3D scene from multi-view photos), while simultaneously distilling features from models like CLIP into that same 3D field, producing a distilled feature field where every point in space carries a semantic feature. A robot then optimizes its gripper pose within this field, learning 6-DOF grasping and placing from just a handful of demonstrations; the target object can be specified with free-form text, and the approach generalizes to objects it has never seen. It shares its core idea with LERF (which embeds language features into a radiance field), and is a representative example of carrying 2D foundation-model knowledge into 3D for manipulation.","example":"Given just two demonstrations per task, such as grasping a mug by its handle or its rim, the robot can pick up mugs of different colors and sizes; typing a short text description lets it pick out the matching object in a scene to grasp.","related":["Distilled Feature Fields","Neural Radiance Fields","CLIP","LERF","Open-vocabulary","Few-shot"]},{"id":"ok-robot","category":"named_model","sec":1,"tier":3,"sources":[{"title":"OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics (arXiv 2401.12202)","url":"https://arxiv.org/abs/2401.12202"},{"title":"OK-Robot 项目页","url":"https://ok-robot.github.io/"},{"title":"ok-robot/ok-robot GitHub 仓库","url":"https://github.com/ok-robot/ok-robot"}],"as_of":"2024-07","related_ids":["open-vocabulary","zero-shot","mobile-manipulation","anygrasp","hello-robot-stretch","semantic-map"],"name":"OK-Robot","alt":"OK-Robot","abbr":"","aliases":["What Really Matters in Integrating Open-Knowledge Models for Robotics"],"one_liner":"A 2024 NYU pick-and-place system built by combining off-the-shelf vision-language models with navigation and grasping modules.","explanation":"OK-Robot was proposed by Lerrel Pinto's group at NYU in collaboration with Meta (Peiqi Liu, Mahi Shafiullah, and others), posted to arXiv in January 2024, published at RSS 2024. “OK” stands for Open Knowledge, meaning models pretrained openly on internet-scale data. It trains no new end-to-end policy; instead, it combines off-the-shelf modules: first, a lidar-equipped iPhone (using the Record3D app) scans the room once to build a voxel map carrying semantic features like CLIP's; on receiving an instruction like “put this object somewhere,” it finds the object in the map, navigates to it, computes a grasp pose with AnyGrasp, then navigates to the target location and places it down. The whole system runs on a Hello Robot Stretch, with no retraining needed when moving to a new home. The paper's focus is summarizing which details actually determine success or failure when assembling a system like this, and it tallies the causes of failure.","example":"Across 171 pick-and-place attempts in 10 real homes in New York, overall success rate was 58.5%, rising to 82% in tidier environments; the most common failures were the semantic map locating the wrong object (9.3%), a difficult grasp pose (8.0%), and hardware issues (7.5%).","related":["Open-vocabulary","Zero-shot","Mobile Manipulation","AnyGrasp","Hello Robot Stretch","Semantic Map"]},{"id":"spatialvlm","category":"named_model","sec":1,"tier":3,"sources":[{"title":"SpatialVLM (arXiv 2401.12168)","url":"https://arxiv.org/abs/2401.12168"},{"title":"SpatialVLM project page","url":"https://spatial-vlm.github.io/"}],"as_of":"2024-06","related_ids":["spatial-reasoning","vision-language-model","visual-question-answering","monocular-depth-estimation","spatial-intelligence","google-deepmind"],"name":"SpatialVLM","alt":"SpatialVLM","abbr":"","aliases":["Spatial VLM","SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities"],"one_liner":"A method that teaches a vision-language model to estimate distance and size using a massive, automatically generated set of 3D spatial question-answer pairs.","explanation":"SpatialVLM was released by Google DeepMind together with MIT and Stanford in January 2024, published at CVPR 2024. The authors found that vision-language models (VLMs) can recognize what's in an image but struggle with quantitative spatial questions like 'how far is the cup from the box' or 'which one is taller,' because their training data lacks 3D spatial knowledge. They built an automated data pipeline: real photos go through object detection, segmentation, and metric depth estimation, lifting 2D images into 3D point clouds with real-world scale, and template-based spatial question-answer pairs are then generated from them — 2 billion question-answer pairs from 10 million images, the first internet-scale dataset for metric spatial reasoning. Models trained on this data show clear gains on both qualitative and quantitative spatial questions, can chain with a large language model for multi-step spatial reasoning, and can use the resulting distance estimates as a dense reward for robot tasks.","example":"Asked 'roughly how many centimeters is the red block from the blue bowl,' an ordinary VLM typically gives only a vague description, while SpatialVLM can give a direct distance estimate with units.","related":["Spatial Reasoning","Vision-Language Model","Visual Question Answering","Monocular Depth Estimation","Spatial Intelligence","Google DeepMind"]},{"id":"pivot","category":"named_model","sec":1,"tier":3,"sources":[{"title":"arXiv 2402.07872: PIVOT","url":"https://arxiv.org/abs/2402.07872"},{"title":"ICML 2024 论文页（PMLR v235）","url":"https://proceedings.mlr.press/v235/nasiriany24a.html"},{"title":"PIVOT 项目主页","url":"https://pivot-prompt.github.io/"}],"as_of":"2024-07","related_ids":["visual-prompting-2","moka","vision-language-model","zero-shot","cross-entropy-method","robopoint"],"name":"PIVOT","alt":"PIVOT（迭代视觉提示）","abbr":"","aliases":["Iterative Visual Prompting","PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs"],"one_liner":"A method that controls robots zero-shot by drawing candidate actions on an image and having a VLM repeatedly pick and narrow them down.","explanation":"PIVOT was led by Google DeepMind (Soroush Nasiriany, Fei Xia, and 21 other co-authors), released in February 2024 and published at ICML 2024. Vision-language models (VLMs) can only output text, but robots need continuous coordinates and actions. PIVOT reframes the problem as iterative visual question answering: it samples a batch of candidates — target points, movement directions, or trajectories — and draws them on the image as numbered arrows or dots; the VLM picks the best few; a new distribution is fit around the selected candidates and resampled, narrowing the range each round, until a final action emerges after a few iterations. The whole process needs no robot training data, and works for real-robot navigation, tabletop manipulation, following instructions in simulation, and image grounding — though the authors acknowledge the success rate is still far from practical. PIVOT represents the visual-prompting line of work, alongside similar approaches like MOKA and Set-of-Mark prompting.","example":"To send a mobile robot to grab a soda can on a table: in round one, several numbered arrows for possible directions are drawn on the image and the VLM picks arrows 3 and 5; the next round redraws arrows only near those two directions and picks again, until the direction is precise enough.","related":["Visual Prompting (Set-of-Mark)","MOKA","Vision-Language Model","Zero-shot","Cross-Entropy Method","RoboPoint"]},{"id":"moka","category":"named_model","sec":1,"tier":3,"sources":[{"title":"MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting (arXiv 2403.03174)","url":"https://arxiv.org/abs/2403.03174"},{"title":"MOKA project page","url":"https://moka-manipulation.github.io/"}],"as_of":"2024-09","related_ids":["visual-prompting-2","affordance","vision-language-model","zero-shot","tool-use","pivot"],"name":"MOKA","alt":"MOKA（标记式视觉提示操作）","abbr":"MOKA","aliases":["Open-World Robotic Manipulation through Mark-Based Visual Prompting"],"one_liner":"A training-free method that marks up an image for GPT-4V to pick keypoints, then converts those into robot-arm actions.","explanation":"MOKA was released in March 2024 by Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine at UC Berkeley, published at RSS 2024. A vision-language model has plenty of common sense, but doesn't directly output robot actions. MOKA rewrites “how to manipulate this” as a picture-answering question: it overlays candidate points, a grid, and text labels onto the camera image as marks (visual prompting), and has GPT-4V hierarchically pick out a grasp point, the point where a tool contacts an object, a target point, and intermediate waypoints, which are then converted into robot-arm motion. It requires no robot-data collection for a new task, and can handle tabletop tasks such as tool use, deformable-object manipulation, and object rearrangement; successful experience gathered during execution can also serve as in-context examples, or be distilled into a policy network. It is a representative example of connecting a VLM to a robot through visual prompting.","example":"Given the instruction “use the brush to sweep the debris aside,” GPT-4V picks, on the marked image, the grasp point on the brush handle, the contact point where the bristles touch the table, and waypoints along the sweeping direction, and the robot executes them in that order.","related":["Visual Prompting (Set-of-Mark)","Affordance","Vision-Language Model","Zero-shot","Tool Use","PIVOT"]},{"id":"copa","category":"named_model","sec":1,"tier":3,"sources":[{"title":"CoPa (arXiv:2403.08248)","url":"https://arxiv.org/abs/2403.08248"},{"title":"CoPa 项目主页","url":"https://copa-2024.github.io/"}],"as_of":"2024-03","related_ids":["rekep","voxposer","omnimanip","visual-prompting-2","task-oriented-grasping","affordance"],"name":"CoPa","alt":"CoPa","abbr":"","aliases":["Spatial Constraints of Parts","CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models"],"one_liner":"A training-free framework where GPT-4V finds relevant object parts and writes spatial constraints to follow open-ended manipulation instructions.","explanation":"CoPa was proposed in March 2024 by Yang Gao's group at Tsinghua University, the Shanghai Qi Zhi Institute, Shanghai Jiao Tong University, and the Shanghai AI Lab. Rather than training a robot policy, it has a foundation model (a general-purpose large model pretrained on massive data) directly supply the geometric information manipulation needs. A manipulation is split into two steps. First, task-oriented grasping: GPT-4V, combined with Set-of-Mark visual prompting, picks where to grasp by narrowing from the whole object down to a specific part, and GraspNet then generates the grasp pose. Second, task-oriented motion planning: a vision-language model identifies the task-relevant parts (such as a hammer's head and a nail), writes out the spatial constraints they should satisfy, solves for the target pose after grasping, and hands it to a motion planner to execute. Like VoxPoser and ReKep, it belongs to the “large model supplies constraints, classical planner executes” approach, and it can handle open-ended instructions and unseen objects.","example":"Given the instruction “hammer in the nail,” CoPa first has GPT-4V select the hammer's handle as the grasp part, then locates the hammer's striking face and the nail, constrains the striking face to align with the nail with the swing direction matching the nail's axis, solves for the hammer's pose right before striking, and executes it.","related":["ReKep","VoxPoser","OmniManip","Visual Prompting (Set-of-Mark)","Task-Oriented Grasping","Affordance"]},{"id":"robopoint","category":"named_model","sec":1,"tier":3,"sources":[{"title":"arXiv 2406.10721: RoboPoint","url":"https://arxiv.org/abs/2406.10721"},{"title":"RoboPoint 项目主页","url":"https://robo-point.github.io/"}],"as_of":"2024-11","related_ids":["affordance","pointing","spatial-reasoning","intermediate-representation","synthetic-data","pivot"],"name":"RoboPoint","alt":"RoboPoint","abbr":"","aliases":["RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics"],"one_liner":"A vision-language model that points to where to place or grasp something on an image, following a language instruction.","explanation":"RoboPoint was released by the University of Washington, NVIDIA, the Allen Institute for AI, and others in June 2024, accepted at CoRL 2024. It focuses on spatial affordance: given an image and an instruction (such as 'put it in the empty spot to the right of the plate'), it marks a set of feasible target locations on the image as 2D points, which are then projected into 3D using a depth map and handed to a motion planner for execution. Its training data is entirely synthesized automatically from procedurally generated 3D scenes, requiring no real-robot data or human demonstrations; the language backbone is Vicuna-13B. The paper reports 21.8% higher spatial-affordance prediction accuracy than GPT-4o and the visual-prompting method PIVOT, and 30.5% higher downstream task success, with applications in manipulation, navigation, and AR assistance. This 'output a point, then hand off to a planner' pattern is a common intermediate representation between VLMs and robots.","example":"For the instruction 'put the cup in the empty spot between the two books,' RoboPoint outputs a set of points on the image falling within the empty spot, and the robot arm uses one of them as the placement target.","related":["Affordance","Pointing","Spatial Reasoning","Intermediate Representation","Synthetic Data","PIVOT"]},{"id":"rekep","category":"named_model","sec":1,"tier":2,"sources":[{"title":"ReKep 项目主页","url":"https://rekep-robot.github.io/"}],"as_of":"2024-09","related_ids":["voxposer","omnimanip","affordance","semantic-keypoints","dinov2","llm-based-task-planning"],"name":"ReKep","alt":"ReKep","abbr":"ReKep","aliases":["Relational Keypoint Constraints","ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation"],"one_liner":"Has a large model write task constraints between keypoints as code, then solves for robot motion with optimization.","explanation":"ReKep (Relational Keypoint Constraints) is a training-free manipulation method proposed in September 2024 by Fei-Fei Li's group at Stanford together with Columbia University. It first uses DINOv2 visual features to find candidate 3D keypoints in an RGB-D image, then gives the numbered, annotated image and a language instruction to GPT-4o, which writes several Python functions: each takes keypoint coordinates as input and outputs a cost, expressing a relation like “the spout must line up with the cup's opening,” with a task optionally broken into multiple stages. A hierarchical optimizer then solves in real time for a sequence of end-effector poses, and replans in closed loop as it tracks the keypoints. This means no task-specific training data is needed to pour tea, fold clothes, or put on shoes with both arms, on either single-arm or bimanual robots — a representative example of the “large model writes constraints, optimizer produces motion” approach.","example":"Given the instruction “pour tea into the cup,” the constraints GPT-4o writes include: during the grasp stage, the gripper must be near a keypoint on the pot's handle; during the pour stage, the spout keypoint must sit directly above the cup's opening keypoint with the pot tilted. The optimizer computes the full motion from these.","related":["VoxPoser","OmniManip","Affordance","Semantic Keypoints","DINOv2","LLM-based Task Planning"]},{"id":"omnimanip","category":"named_model","sec":1,"tier":3,"sources":[{"title":"arXiv 2501.03841: OmniManip","url":"https://arxiv.org/abs/2501.03841"},{"title":"OmniManip 项目主页","url":"https://omnimanip.github.io/"}],"as_of":"2025-06","related_ids":["rekep","voxposer","copa","affordance","intermediate-representation","6d-object-pose-estimation"],"name":"OmniManip","alt":"OmniManip","abbr":"","aliases":["OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints"],"one_liner":"A zero-shot manipulation framework that turns a vision-language model's reasoning into point-and-direction constraints defined in each object's own coordinate frame.","explanation":"OmniManip comes from Hao Dong's lab at Peking University working with AgiBot (the PKU-AgiBot joint lab), released in January 2025 and selected as a CVPR 2025 Highlight paper. Vision-language models (VLMs) have broad commonsense knowledge but cannot reliably output precise 3D positions and orientations. OmniManip first places each object into its own canonical space — a coordinate frame aligned to the object's function rather than just its shape — and defines 'interaction primitives' inside it, such as an interaction point and an interaction direction. These primitives become spatial constraints that the VLM selects and checks. During execution, the system tracks each object's 6D pose (3D position plus 3D orientation) in real time and updates the trajectory accordingly, so both planning and execution run as closed loops. None of this requires fine-tuning the VLM, yet the method generalizes to many manipulation tasks in a zero-shot setting. It belongs to the same 'VLM plus intermediate representation' family as ReKep, VoxPoser, and CoPa; the authors also note it can be used to auto-generate simulation data.","example":"Take 'pour tea into a cup': the VLM first recognizes the teapot and cup, picks the spout point and pouring direction in the teapot's canonical space and the rim point on the cup, and uses these to compute the end-effector pose; during execution it keeps tracking both objects' 6D poses and corrects the trajectory.","related":["ReKep","VoxPoser","CoPa","Affordance","Intermediate Representation","6D Object Pose Estimation"]},{"id":"spatiallm","category":"named_model","sec":1,"tier":3,"sources":[{"title":"SpatialLM: Training Large Language Models for Structured Indoor Modeling (arXiv 2506.07491)","url":"https://arxiv.org/abs/2506.07491"},{"title":"manycore-research/SpatialLM (GitHub)","url":"https://github.com/manycore-research/SpatialLM"}],"as_of":"2025-09","related_ids":["point-cloud","3d-object-detection","scene-understanding","manycore-tech","spatial-intelligence","multimodal-large-language-model"],"name":"SpatialLM","alt":"SpatialLM（群核空间大模型）","abbr":"","aliases":["SpatialLM 1.1","Manycore SpatialLM","SpatialLM: Training Large Language Models for Structured Indoor Modeling"],"one_liner":"Manycore's open-source 3D large model that reads an indoor point cloud and outputs a structured layout of walls, doors, windows, and furniture.","explanation":"SpatialLM was open-sourced by Hangzhou-based Manycore Tech (owner of Kujiale) in March 2025, with a technical report in June, and was selected for NeurIPS 2025. It takes as input a point cloud of an indoor scene — which can come from phone-video reconstruction, an RGB-D camera, or LiDAR — and outputs a structured scene description: the positions of walls, doors, and windows, and 3D boxes for furniture with category and orientation. Rather than designing a separate network for each task, as earlier work did, it reuses a standard multimodal large-model structure: a point-cloud encoder turns geometric information into tokens, and a small open-source language model such as Llama 1B or Qwen 0.5B then writes out the scene item by item as text. Training used point clouds and annotations from 12,328 synthetic indoor scenes (54,778 rooms). Version 1.1, released in June, switched to the Sonata point-cloud encoder and added detection restricted to user-specified categories. SpatialLM can supply structured spatial information for robot navigation and indoor layout understanding.","example":"A phone is used to film a walkthrough of a room, which is reconstructed into a point cloud via MASt3R-SLAM and fed into SpatialLM, yielding each wall's endpoints, the positions of doors and windows, and 3D boxes for the bed and nightstand.","related":["Point Cloud","3D Object Detection","Scene Understanding","Manycore Tech","Spatial Intelligence","Multimodal Large Language Model"]},{"id":"embodied-r1","category":"named_model","sec":1,"tier":3,"sources":[{"title":"Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation (arXiv 2508.13998)","url":"https://arxiv.org/abs/2508.13998"},{"title":"Embodied-R1 GitHub 仓库","url":"https://github.com/pickxiguapi/Embodied-R1"}],"as_of":"2026-03","related_ids":["embodied-reasoning-model","pointing","reinforcement-fine-tuning","group-relative-policy-optimization","intermediate-representation","affordance"],"name":"Embodied-R1","alt":"Embodied-R1","abbr":"","aliases":["Embodied R1","Reinforced Embodied Reasoning for General Robotic Manipulation"],"one_liner":"A 3B embodied-reasoning model from Tianjin University that uses “pointing” as an intermediate representation, trained with reinforcement fine-tuning.","explanation":"Embodied-R1 is work released in August 2025 by a team at Tianjin University, accepted to ICLR 2026. The authors call the gap between “understanding what to do” and “actually doing it” the seeing-to-doing gap, and propose using “pointing” — outputting a point, a region, or a sequence of trajectory points on an image — as an intermediate representation that is independent of any specific robot body, defining four abilities: referring-expression grounding, region grounding, functional-part grounding, and visual-trajectory generation. The model is built on Qwen2.5-VL-3B, and is trained with two-stage reinforcement fine-tuning on a self-built dataset, Embodied-Points-200K, using the GRPO algorithm with automatically scored, task-specific rewards. The points and trajectories the model outputs are then handed off to lower-level modules such as motion planning for execution. The weights and dataset are both open-source.","example":"Without any task-specific fine-tuning, Embodied-R1 reaches a 56.2% success rate in SimplerEnv simulation and 87.5% across 8 real-robot xArm tasks, which the paper reports as a 62% improvement over strong baselines.","related":["Embodied Reasoning Model","Pointing","Reinforcement Fine-Tuning (RL Fine-Tuning)","Group Relative Policy Optimization","Intermediate Representation","Affordance"]},{"id":"end-to-end-training-of-deep-visuomotor-policies","category":"named_model","sec":2,"tier":3,"sources":[{"title":"End-to-End Training of Deep Visuomotor Policies (arXiv 1504.00702)","url":"https://arxiv.org/abs/1504.00702"},{"title":"JMLR 17(39):1-40, 2016","url":"https://www.jmlr.org/papers/v17/15-522.html"}],"as_of":"","related_ids":["visuomotor-policy","end-to-end","spatial-softmax","trajectory-optimization","reinforcement-learning","imitation-learning"],"name":"End-to-End Training of Deep Visuomotor Policies","alt":"端到端视觉运动策略（引导策略搜索）","abbr":"GPS","aliases":["Guided Policy Search","GPS","Levine et al. 2016"],"one_liner":"A landmark 2015 Berkeley paper that used a convolutional network to output robot-arm joint torques directly from camera images.","explanation":"This paper was released in April 2015 by Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel at UC Berkeley, and published in JMLR in 2016. At the time, most robots were designed with perception and control as separate components; the paper set out to test whether jointly training both inside a single network would work better. The policy is a convolutional network with about 92,000 parameters, taking a monocular image and joint state as input and outputting joint torques for a 7-DOF arm at 20Hz. Training uses guided policy search: with object positions known during training, trajectory optimization first finds good actions, and supervised learning then has an image-only policy imitate them — turning reinforcement learning into supervised learning. The spatial softmax layer introduced in this paper was later adopted by many subsequent visuomotor policies.","example":"The PR2 robot learned to hang a coat hanger on a rod, fit blocks into a shape-sorting cube, use the claw of a toy hammer to pull out a nail, and screw on a bottle cap, judging target positions entirely from the camera; each policy took 3 to 4 hours to train in total, of which only about 15 minutes was actual real-robot execution.","related":["Visuomotor Policy","End-to-End","Spatial Softmax","Trajectory Optimization","Reinforcement Learning","Imitation Learning"]},{"id":"google-arm-farm","category":"named_model","sec":2,"tier":3,"sources":[{"title":"Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection (arXiv 1603.02199)","url":"https://arxiv.org/abs/1603.02199"},{"title":"Deep Learning for Robots: Learning from Large-Scale Interaction (Google Research Blog, 2016-03)","url":"https://research.google/blog/deep-learning-for-robots-learning-from-large-scale-interaction/"}],"as_of":"2016-03","related_ids":["google-arm-farm","hand-eye-coordination","visual-servoing","self-supervised-learning","qt-opt","mt-opt"],"name":"Google Arm Farm","alt":"Google 机械臂农场（大规模抓取自监督）","abbr":"","aliases":["Arm Farm","Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection"],"one_liner":"Google's 2016 project where more than a dozen robot arms attempted grasps over 800,000 times to learn grasping from the data.","explanation":"This is work from Google's Sergey Levine, Peter Pastor, Alex Krizhevsky, and Deirdre Quillen, released in March 2016 (an extended version of an ISER 2016 paper), nicknamed the “arm farm” for the sight of a whole row of robot arms working at once. Over about two months, 6 to 14 robot arms autonomously attempted more than 800,000 grasps, with success or failure judged automatically, with no manual labeling. This data was used to train a convolutional neural network: given a monocular camera image and a candidate gripper motion, it predicts whether that motion would lead to a successful grasp; at execution time, the system continuously picks the best motion and adjusts as it watches (visual servoing), with no camera calibration needed. Closed-loop grasping cut the failure rate from 34% under open-loop control down to 18%. This opened up the “large-scale real-robot data plus deep learning” approach that QT-Opt and MT-Opt later followed.","example":"If the gripper drifts off target mid-grasp, the robot corrects course in real time based on what it sees; faced with a pile of objects crowded together, it will even nudge one aside before grasping it.","related":["Google Arm Farm","Hand-Eye Coordination","Visual Servoing","Self-Supervised Learning","QT-Opt","MT-Opt"]},{"id":"dex-net-2-0","category":"named_model","sec":2,"tier":3,"sources":[{"title":"Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics (arXiv:1703.09312)","url":"https://arxiv.org/abs/1703.09312"},{"title":"Dex-Net 项目主页（Berkeley AUTOLAB）","url":"https://berkeleyautomation.github.io/dex-net/"}],"as_of":"2019","related_ids":["grasping","grasp-quality-metric","bin-picking","synthetic-data","convolutional-neural-network","grasp-pose-detection"],"name":"Dex-Net 2.0","alt":"Dex-Net（GQ-CNN 抓取网络）","abbr":"Dex-Net","aliases":["Dexterity Network","GQ-CNN","Grasp Quality CNN"],"one_liner":"Berkeley's grasp-scoring network, trained on simulated data, that judges from a depth image which grasp will hold best.","explanation":"Dex-Net is the grasping research project of Ken Goldberg's group (AUTOLAB) at UC Berkeley; its best-known version, Dex-Net 2.0, was published at RSS 2017. Instead of collecting real-robot data, it labels a huge collection of 3D object models using analytic grasp metrics (formulas from geometry and mechanics that score how stable a grasp is), synthesizing 6.7 million “point cloud + grasp + label” examples, then trains a convolutional network, GQ-CNN: given a depth image and a candidate grasp (planar position, angle, depth), it outputs a success probability, and the robot executes whichever candidate scores highest. It's an early, representative example of “train on synthetic simulated data, deploy directly to the real robot” for grasping; a later version, 3.0, extended this to suction grippers, and 4.0 (Science Robotics, 2019) uses both parallel-jaw grippers and suction together for bin picking.","example":"Dex-Net 2.0 used GQ-CNN to plan two-finger parallel-jaw grasps on an ABB YuMi robot: 93% success on 8 known objects, and about 99% precision for grasps judged “stable” on household objects it had never seen.","related":["Grasping","Grasp Quality Metric","Bin Picking","Synthetic Data","Convolutional Neural Network","Grasp Pose Detection"]},{"id":"qt-opt","category":"named_model","sec":2,"tier":3,"sources":[{"title":"QT-Opt (arXiv 1806.10293)","url":"https://arxiv.org/abs/1806.10293"}],"as_of":"2018-06","related_ids":["q-transformer","mt-opt","real-world-reinforcement-learning","q-function","cross-entropy-method","google-arm-farm"],"name":"QT-Opt","alt":"QT-Opt","abbr":"QT-Opt","aliases":["QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation"],"one_liner":"Google's large-scale reinforcement learning grasping system, trained on more than 580,000 real-robot grasp attempts to learn a visual Q-function.","explanation":"QT-Opt was released by Dmitry Kalashnikov, Sergey Levine, and colleagues at Google Brain and X in June 2018 and published at CoRL 2018. At the time, most grasping systems first chose a grasp point and then executed it open-loop. QT-Opt instead uses deep reinforcement learning to learn a Q-function directly from an overhead RGB camera image, re-deciding how the gripper should move at every step to achieve closed-loop grasping; because it is hard to directly maximize over a continuous action space, it uses the cross-entropy method (an iterative sampling-based optimization) to search for the best action. Training used more than 580,000 real grasp attempts, reaching a 96% success rate on objects never seen before, and the policy spontaneously learned behaviors like re-grasping and nudging an object before picking it up. QT-Opt is a landmark result for large-scale real-robot reinforcement learning, and both MT-Opt and Q-Transformer build on it.","example":"If an object is pushed away by a person mid-grasp, the QT-Opt policy adjusts the gripper's position based on the new camera image and re-attempts the grasp, rather than closing on empty space as originally planned.","related":["Q-Transformer","MT-Opt","Real-World Reinforcement Learning","Q-Function","Cross-Entropy Method","Google Arm Farm"]},{"id":"mt-opt","category":"named_model","sec":2,"tier":3,"sources":[{"title":"MT-Opt (arXiv 2104.08212)","url":"https://arxiv.org/abs/2104.08212"},{"title":"Multi-Task Robotic Reinforcement Learning at Scale (Google Research blog, 2021-04-19)","url":"https://research.google/blog/multi-task-robotic-reinforcement-learning-at-scale/"},{"title":"MT-Opt project page","url":"https://karolhausman.github.io/mt-opt/"}],"as_of":"2021-04","related_ids":["qt-opt","reinforcement-learning","multi-task-learning","success-detector","real-world-reinforcement-learning","offline-reinforcement-learning"],"name":"MT-Opt","alt":"MT-Opt","abbr":"","aliases":["Continuous Multi-Task Robotic Reinforcement Learning at Scale"],"one_liner":"A Google multi-task reinforcement-learning system that used 7 real robots to gather 9,600 hours of data while learning 12 tasks at once.","explanation":"MT-Opt was released in April 2021 by Dmitry Kalashnikov, Chelsea Finn, Sergey Levine, Karol Hausman, and colleagues at Robotics at Google, a multi-task extension of QT-Opt (Google's large-scale Q-learning-based grasping reinforcement-learning system). The team used 7 robots to continuously gather about 9,600 robot-hours across more than 800,000 episodes over 57 days, learning 12 real tasks at once — picking a specific object, placing it in a container, aligning objects, covering, and more. Two design choices are key: a multi-task success detector that automatically judges whether a task is done and assigns reward, and sharing episodes from one task with the others while rebalancing the amount of data, so tasks with little data can borrow experience from more common ones. Average success rate on rare tasks rose from 1% under single-task QT-Opt to 50%, and a new task could be fine-tuned in about a day. It is an important piece of Google's exploration of scaling up real-robot learning before RT-1.","example":"To teach the robot a new task, “cover an object with a towel,” there's no need to train from scratch; collecting about a day's worth of additional data to fine-tune on top of already-learned skills like grasping is enough.","related":["QT-Opt","Reinforcement Learning","Multi-Task Learning","Success Detector","Real-World Reinforcement Learning","Offline Reinforcement Learning"]},{"id":"decision-transformer","category":"named_model","sec":2,"tier":3,"sources":[{"title":"Decision Transformer (arXiv:2106.01345)","url":"https://arxiv.org/abs/2106.01345"},{"title":"Decision Transformer 项目主页","url":"https://sites.google.com/berkeley.edu/decision-transformer"}],"as_of":"2021-06","related_ids":["offline-reinforcement-learning","transformer","return-conditioning","causal-attention","decision-diffuser","gato"],"name":"Decision Transformer","alt":"决策 Transformer","abbr":"DT","aliases":["DT"],"one_liner":"Treats reinforcement learning as sequence modeling: given a target return, a GPT-style model predicts actions one step at a time.","explanation":"Decision Transformer was proposed in June 2021 by Lili Chen, Kevin Lu, and colleagues at Berkeley, Facebook AI Research, and Google Brain, published at NeurIPS 2021. It doesn't fit a value function or compute policy gradients; instead, it writes a trajectory as a sequence of tokens alternating “return-to-go” (how much reward is still wanted from this point to the end), state, and action, and trains a GPT-style Transformer with causal masking, using supervised learning, to predict the next action. At test time, a target return is set at the start, and the reward received at each step is subtracted from it as the episode goes on. Using only offline data, it matched or beat the leading model-free offline reinforcement-learning methods of the time on Atari, OpenAI Gym, and Key-to-Door tasks, helping make “treat decision-making as sequence modeling” one of the mainstream approaches.","example":"In an Atari game, a fairly high target return is set at the start; the model reads in the most recent frames, past actions, and remaining return-to-go, and outputs the next button press; every time points are scored, the remaining return-to-go drops accordingly before the next action is generated.","related":["Offline Reinforcement Learning","Transformer","Return Conditioning","Causal Attention","Decision Diffuser","Gato"]},{"id":"gato","category":"named_model","sec":2,"tier":3,"sources":[{"title":"A Generalist Agent (arXiv 2205.06175)","url":"https://arxiv.org/abs/2205.06175"},{"title":"A Generalist Agent（Google DeepMind 博客）","url":"https://deepmind.google/discover/blog/a-generalist-agent/"}],"as_of":"2022-05","related_ids":["generalist-policy","transformer","token","robocat","multi-task-learning","cross-embodiment"],"name":"Gato","alt":"Gato","abbr":"","aliases":["A Generalist Agent"],"one_liner":"A 2022 DeepMind generalist agent whose single set of weights can play games, caption images, chat, and control a robot arm.","explanation":"Gato is a generalist agent DeepMind released in May 2022, in a paper titled A Generalist Agent. The mainstream approach at the time was one model per task; Gato instead used a single roughly 1.2-billion-parameter Transformer, with one set of weights, to cover 604 tasks: playing Atari games, captioning images, holding conversations, controlling various robots in simulation, and stacking blocks with a real robot arm. The key idea is converting every modality into tokens: text is split by a tokenizer, images are cut into patches, and discrete or continuous values like joint angles and button presses are also encoded as tokens, all concatenated into a single sequence and predicted one at a time like a language model, with the training loss computed only on the text and action tokens. At inference, the model autoregressively samples action tokens one at a time, with a context window of 1,024 tokens. It is an early representative of the “one large model controlling many embodiments” idea, and DeepMind's later RoboCat carried over Gato's architecture.","example":"The same Gato model presses controller buttons in an Atari game one moment, generates a caption for a photo the next, and then controls a real robot arm to stack colored blocks.","related":["Generalist Policy","Transformer","Token","RoboCat","Multi-Task Learning","Cross-Embodiment"]},{"id":"vima","category":"named_model","sec":2,"tier":3,"sources":[{"title":"VIMA (arXiv 2210.03094)","url":"https://arxiv.org/abs/2210.03094"},{"title":"VIMA project page","url":"https://vimalabs.github.io/"}],"as_of":"2023-05","related_ids":["vima-bench","language-conditioned-policy","compositional-generalization","tabletop-manipulation","imitation-learning","transformer"],"name":"VIMA","alt":"VIMA","abbr":"VIMA","aliases":["VIMA: General Robot Manipulation with Multimodal Prompts"],"one_liner":"A Transformer robot agent that uses interleaved text-and-image 'multimodal prompts' to describe manipulation tasks in one unified format.","explanation":"VIMA was released in October 2022 by researchers at Stanford, NVIDIA, Caltech, and other institutions (Yunfan Jiang, Linxi Fan, Yuke Zhu, Fei-Fei Li, and others), published at ICML 2023. Robot tasks can be specified in many different ways — showing a demonstration to imitate, describing it in language, or giving a goal image. VIMA unifies all of these into a single 'text interleaved with images' multimodal prompt, such as 'put [image of an object] into [image of a container],' with a Transformer reading the prompt and autoregressively outputting actions. The authors also built the VIMA-Bench simulation benchmark: 17 task templates that can procedurally generate thousands of tabletop tasks, more than 600,000 expert trajectories, and a four-level evaluation of increasingly hard generalization. The paper reports up to 2.9x higher success than other designs in the hardest zero-shot setting.","example":"A VIMA prompt is a sentence with two small images embedded in it: 'put [image of a red block] onto [image of a green plate]'; VIMA finds the matching objects on a simulated tabletop and completes the pick-and-place, and the same model also handles a prompt that first shows a demonstration image and then says 'do it like this.'","related":["VIMA-Bench","Language-conditioned Policy","Compositional Generalization","Tabletop Manipulation","Imitation Learning","Transformer"]},{"id":"rt-1","category":"named_model","sec":2,"tier":1,"sources":[{"title":"RT-1: Robotics Transformer for Real-World Control at Scale (arXiv 2212.06817)","url":"https://arxiv.org/abs/2212.06817"},{"title":"RT-1 项目主页","url":"https://robotics-transformer1.github.io/"}],"as_of":"2022-12","related_ids":["rt-2","rt-x","rt-1-robot-action-dataset","transformer","efficientnet","tokenlearner"],"name":"RT-1","alt":"RT-1","abbr":"RT-1","aliases":["Robotics Transformer 1","RT-1: Robotics Transformer for Real-World Control at Scale"],"one_liner":"Google's 2022 Transformer policy for real robots, trained on 130,000-plus real-robot demonstrations across 700-plus tasks.","explanation":"RT-1 is a robot control model released by the Google Robotics team and Everyday Robots in December 2022. At the time, most robot learning trained one small model per task; RT-1 set out to test whether the recipe behind language and vision breakthroughs — a large-capacity model trained on large, diverse data — would also work for robots. The team used 13 mobile manipulators over 17 months to collect more than 130,000 demonstrations across more than 700 tasks. The model takes in images and a language instruction: images go through an EfficientNet feature extractor, FiLM layers (which let the language instruction modulate the visual features) fold in the instruction, TokenLearner compresses the number of tokens, and a Transformer outputs discretized actions, controlling the arm and base in closed loop at 3Hz. It reached 97% success on trained instructions and generalized to new tasks, distractors, and backgrounds better than the baselines of the time. Its data later became part of Open X-Embodiment, and both its architecture and data underpin RT-2.","example":"In a Google office kitchen environment, RT-1 followed language instructions to pick and place objects, open and close drawers, and put items away in drawers.","related":["RT-2","RT-X","RT-1 Robot Action Dataset","Transformer","EfficientNet","TokenLearner"]},{"id":"rt-2","category":"named_model","sec":2,"tier":1,"sources":[{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (arXiv 2307.15818)","url":"https://arxiv.org/abs/2307.15818"},{"title":"RT-2 项目主页","url":"https://robotics-transformer2.github.io/"}],"as_of":"2023-07","related_ids":["vision-language-action-model","rt-1","palm-e","co-training","action-binning","openvla"],"name":"RT-2","alt":"RT-2","abbr":"RT-2","aliases":["Robotics Transformer 2","RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control"],"one_liner":"Google DeepMind's 2023 model that outputs robot actions as text tokens — the paper that coined the term VLA.","explanation":"RT-2 is a model Google DeepMind released in July 2023, and its paper was the first to use the term “vision-language-action model” (VLA). The approach: take a vision-language model already pretrained on internet image-text data (PaLI-X or PaLM-E), discretize each dimension of a robot action into 256 bins, write them out as a string of number tokens, and have the model output them just like ordinary text. During training, RT-1's robot data is co-fine-tuned together with web-scale visual question-answering data, so the model doesn't lose its existing knowledge. This lets the robot draw directly on commonsense knowledge learned from the web, producing what the paper calls emergent capabilities: recognizing objects and symbols absent from the robot data, and understanding semantic concepts like “the smallest one” or “something that could be used as a hammer.” The model comes in two sizes, 12 billion parameters (PaLM-E version) and 55 billion parameters (PaLI-X version), and it opened the path that later VLA models such as OpenVLA and π0 followed.","example":"Told to “move the banana to the sum of two plus one,” RT-2 can place the banana on the spot marked 3, even though the robot's own training data contains no such arithmetic task.","related":["Vision-Language-Action Model","RT-1","PaLM-E","Co-training","Action Binning","OpenVLA"]},{"id":"rt-x","category":"named_model","sec":2,"tier":2,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models 项目主页","url":"https://robotics-transformer-x.github.io/"}],"as_of":"2023-10","related_ids":["open-x-embodiment","rt-1","rt-2","cross-embodiment","octo","openvla"],"name":"RT-X","alt":"RT-X","abbr":"RT-X","aliases":["RT-1-X","RT-2-X"],"one_liner":"RT-1 and RT-2 retrained on Open X-Embodiment, a large dataset pooling robot data across 22 different robot embodiments.","explanation":"RT-X is a set of models released in October 2023, led by Google DeepMind in collaboration with 21 institutions and 34 labs, alongside the Open X-Embodiment dataset. That dataset pools 60 existing robot datasets, 22 different robot embodiments, and more than 1 million real-robot trajectories. RT-1-X and RT-2-X keep the architectures of RT-1 (a dedicated robot Transformer) and RT-2 (a VLA fine-tuned from a vision-language model) unchanged, simply swapping in this combined dataset for training. The results show positive transfer from cross-embodiment training: RT-1-X beat the original methods used by individual labs by about 50% on average in low-data settings, and RT-2-X scored roughly 3x higher than RT-2 on emergent-skill evaluations. It drove wider cross-embodiment data sharing, and later models such as Octo and OpenVLA are all trained on this same dataset.","example":"RT-2-X picked up spatial concepts that were only present in other labs' data, learning to tell apart instructions that differ by a single preposition, such as “put the apple on the cloth” versus “put the apple next to the cloth.”","related":["Open X-Embodiment","RT-1","RT-2","Cross-Embodiment","Octo","OpenVLA"]},{"id":"rt-trajectory","category":"named_model","sec":2,"tier":3,"sources":[{"title":"RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches (arXiv 2311.01977)","url":"https://arxiv.org/abs/2311.01977"},{"title":"RT-Trajectory project page","url":"https://rt-trajectory.github.io/"}],"as_of":"2024-01","related_ids":["rt-1","rt-2","intermediate-representation","task-generalization","hindsight-relabeling","goal-conditioned-policy"],"name":"RT-Trajectory","alt":"RT-Trajectory","abbr":"","aliases":["RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches"],"one_liner":"A policy that replaces language instructions with a trajectory sketch drawn on the image, showing the robot how the motion should go.","explanation":"RT-Trajectory was released by Google DeepMind together with UC San Diego, Stanford, and Intrinsic in November 2023, selected as an ICLR 2024 Spotlight. A language instruction only says 'what to do,' which is often not specific enough for tasks the model never saw during training. RT-Trajectory instead conditions on a 'trajectory sketch': the path the end effector should follow is drawn as a curve overlaid on the camera image, with color encoding time and height, and marks showing where the gripper should open or close. No manual annotation is needed for training — the end-effector positions recorded in each demonstration are simply projected onto the image after the fact to generate the sketch (hindsight). The policy backbone reuses RT-1. At test time, sketches can be hand-drawn by a person, extracted from human videos, or generated by a large model. Across 7 tasks unseen during training, it clearly outperforms language-conditioned baselines like RT-1 and RT-2, as well as goal-image-conditioned baselines.","example":"Training data is mostly pick-and-place; at test time, a person draws a curve on the image that first grabs one corner of a cloth and then pulls it toward the opposite side, and the robot follows it to perform 'fold the cloth,' a task absent from training.","related":["RT-1","RT-2","Intermediate Representation","Task Generalization","Hindsight Relabeling","Goal-conditioned Policy"]},{"id":"autort","category":"named_model","sec":2,"tier":3,"sources":[{"title":"AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents (arXiv 2401.12963)","url":"https://arxiv.org/abs/2401.12963"},{"title":"Shaping the future of advanced robotics (Google DeepMind blog, 2024-01-04)","url":"https://deepmind.google/discover/blog/shaping-the-future-of-advanced-robotics/"},{"title":"AutoRT 项目页","url":"https://auto-rt.github.io/"}],"as_of":"2024-01","related_ids":["autonomous-data-collection","asimov-s-three-laws-of-robotics","rt-2","rt-1","saycan","llm-based-task-planning"],"name":"AutoRT","alt":"AutoRT","abbr":"","aliases":["AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents"],"one_liner":"DeepMind's 2024 system that uses a large model to automatically assign tasks to a fleet of robots and collect real data.","explanation":"AutoRT is a system Google DeepMind announced in January 2024, with authors including Karol Hausman, Fei Xia, Chelsea Finn, and Sergey Levine. Training embodied foundation models needs real data, but having one person watch one robot to collect it doesn't scale. AutoRT has an off-the-shelf large model act as dispatcher: robots explore autonomously in an office building, a vision-language model describes the scene and objects, and a large language model proposes tasks based on that. Each task is first filtered by a “robot constitution” (basic rules inspired by Asimov's Three Laws, plus safety rules and embodiment-capability rules), then assigned — depending on available staffing — to teleoperation, a scripted grasping policy, or RT-2 to execute. Over 7 months it orchestrated robots (up to more than 20 at once) across 4 buildings, collecting 77,000 real-robot episodes covering more than 6,650 distinct instructions, with one person able to supervise 3 to 5 robots at once.","example":"If the LLM proposes “open the fridge and get a drink,” the constitution's embodiment rule would filter it out for a single-arm robot, since the task needs two hands; tasks involving people or animals, sharp or fragile objects, or electrical appliances get blocked by the safety rules.","related":["Autonomous Data Collection","Asimov's Three Laws of Robotics","RT-2","RT-1","SayCan","LLM-based Task Planning"]},{"id":"rt-h","category":"named_model","sec":2,"tier":3,"sources":[{"title":"RT-H: Action Hierarchies Using Language (arXiv 2403.01823)","url":"https://arxiv.org/abs/2403.01823"},{"title":"RT-H project page","url":"https://rt-hierarchy.github.io/"}],"as_of":"2024-06","related_ids":["rt-2","hierarchical-architecture","language-corrections","human-in-the-loop","intermediate-representation","vision-language-action-model"],"name":"RT-H","alt":"RT-H","abbr":"RT-H","aliases":["RT-Hierarchy","RT-H: Action Hierarchies Using Language"],"one_liner":"A hierarchical VLA from Google that has the robot first state a 'language motion' like 'move arm forward' before outputting the action.","explanation":"RT-H was released by Google DeepMind and Stanford University in March 2024, with authors including Suneel Belkhale and Dorsa Sadigh. VLAs like RT-2 map directly from a task instruction, such as 'put the soda can in the drawer,' to motor commands, which makes it hard to learn the action structure shared across different tasks. RT-H inserts a middle layer of 'language motions' — fine-grained phrases like 'move arm forward' or 'close gripper': the same vision-language model, co-trained with internet data, first predicts a language motion from the task and the image, then outputs the specific action based on that language motion and the image. This lets semantically different tasks share the same underlying motions, and a person can correct the robot mid-execution simply by speaking, with those correction episodes then reused for further training. The paper reports roughly 15% better performance than RT-2 on multi-task data, and that learning from language interventions works better than learning from teleoperated interventions.","example":"If the robot's hand drifts off-target while opening a drawer, a person can just say 'move your arm to the left'; RT-H treats this as a new language motion and continues execution, and the correction is logged for further training.","related":["RT-2","Hierarchical Architecture","Language Corrections","Human-in-the-Loop","Intermediate Representation","Vision-Language-Action Model"]},{"id":"robocat","category":"named_model","sec":2,"tier":3,"sources":[{"title":"arXiv 2306.11706: RoboCat","url":"https://arxiv.org/abs/2306.11706"},{"title":"Google DeepMind 博客：RoboCat","url":"https://deepmind.google/discover/blog/robocat-a-self-improving-robotic-agent/"}],"as_of":"2023-12","related_ids":["gato","decision-transformer","cross-embodiment","self-improvement","goal-conditioned-policy","google-deepmind"],"name":"RoboCat","alt":"RoboCat","abbr":"","aliases":["RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation"],"one_liner":"DeepMind's cross-embodiment manipulation agent, built on Gato, that generates its own data to keep improving.","explanation":"RoboCat is a robot manipulation agent released by Google DeepMind in June 2023, reusing the architecture of Gato, DeepMind's multimodal generalist model: it is a decision Transformer conditioned on a goal image — given a picture of 'what the task looks like when done,' it outputs actions. It trains on millions of trajectories from a variety of real and simulated robot arms, and can handle robots whose observation and action formats differ from one another. Its defining feature is a 'self-improvement' loop: a task-specific version is first fine-tuned on 100 to 1,000 human demonstrations, then lets the robot practice roughly 10,000 times on its own to generate new data, which is folded back into the training set to retrain the generalist model. According to DeepMind's blog, later versions raised the success rate on new tasks from 36% to 74%. RoboCat is an early landmark demonstrating that cross-embodiment data can speed up learning new skills.","example":"RoboCat learned to operate a new robot arm fitted with a three-fingered gripper within a few hours, reaching an 86% success rate at grasping a gear after just 1,000 demonstrations.","related":["Gato","Decision Transformer","Cross-Embodiment","Self-improvement","Goal-conditioned Policy","Google DeepMind"]},{"id":"roboflamingo","category":"named_model","sec":2,"tier":3,"sources":[{"title":"arXiv 2311.01378: Vision-Language Foundation Models as Effective Robot Imitators","url":"https://arxiv.org/abs/2311.01378"},{"title":"RoboFlamingo 项目主页","url":"https://roboflamingo.github.io/"},{"title":"GitHub: RoboFlamingo","url":"https://github.com/RoboFlamingo/RoboFlamingo"}],"as_of":"2024-02","related_ids":["vision-language-model","vision-language-action-model","calvin-benchmark","imitation-learning","robovlms","flamingo"],"name":"RoboFlamingo","alt":"RoboFlamingo","abbr":"","aliases":["Vision-Language Foundation Models as Effective Robot Imitators"],"one_liner":"An early piece of work that fine-tuned the open-source OpenFlamingo vision-language model directly into a robot manipulation policy.","explanation":"RoboFlamingo was released in November 2023 by ByteDance Research together with Tsinghua University, Shanghai Jiao Tong University, and the National University of Singapore, one of the earlier efforts to turn an open-source vision-language model (VLM) directly into a manipulation policy. It uses OpenFlamingo as its backbone, which at each step understands the current image and the language instruction, followed by an explicit policy head (such as an LSTM) that aggregates history and outputs the robot arm's actions; it is fine-tuned with imitation learning purely on language-annotated demonstration data. This split between 'understanding' and 'decision-making' let it be trained on a single 8-GPU server. On the CALVIN long-horizon benchmark, it completed an average of 4.09 consecutive tasks, clearly ahead of prior methods, though the paper only validated it in simulation. The same authors later expanded this idea into the systematic study RoboVLMs.","example":"Given five consecutive instructions in CALVIN (such as 'open the drawer' and 'push the blue block to the left'), RoboFlamingo could on average complete about 4 of them in a row.","related":["Vision-Language Model","Vision-Language-Action Model","CALVIN Benchmark","Imitation Learning","RoboVLMs","Flamingo"]},{"id":"transporter-networks","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Transporter Networks (arXiv 2010.14406)","url":"https://arxiv.org/abs/2010.14406"},{"title":"Transporter Networks project page","url":"https://transporternets.github.io/"}],"as_of":"2020-10","related_ids":["cliport","pick-and-place","rearrangement","tabletop-manipulation","imitation-learning","equivariant-policy-equivariant-neural-network"],"name":"Transporter Networks","alt":"Transporter Networks","abbr":"","aliases":["Transporter","Transporter Networks: Rearranging the Visual World for Robotic Manipulation"],"one_liner":"A highly sample-efficient manipulation network that treats pick-and-place as 'moving one region of the image to another location.'","explanation":"Transporter Networks was released by Robotics at Google (Andy Zeng, Pete Florence, and others) in October 2020, published at CoRL 2020 and a finalist for the best paper award. It treats tabletop pick-and-place tasks as a sequence of 'spatial displacements': it first predicts where to grasp from a top-down image, then crops the depth features around that grasp point and cross-correlates them against the features of the whole scene (comparing them at every location) to find the best placement position and rotation angle. This design needs no object detection step and naturally exploits translation and rotation symmetry, making it very data-efficient — a small number of demonstrations are enough to learn tasks like stacking blocks, kit assembly, rope manipulation, and pushing objects, and it extends to 6-DoF pick-and-place as well. The paper also open-sourced Ravens, a PyBullet-based simulation task suite. The later CLIPort added CLIP-based language understanding on top of it, becoming a common baseline for language-conditioned manipulation.","example":"For a kit-assembly task, the model looks at a top-down image to find where to grasp each part, which slot in the mold to place it in, and how much to rotate it — learning this from just a handful of demonstrations.","related":["CLIPort","Pick-and-Place","Rearrangement","Tabletop Manipulation","Imitation Learning","Equivariant Policy / Equivariant Neural Network"]},{"id":"cliport","category":"named_model","sec":3,"tier":3,"sources":[{"title":"CLIPort: What and Where Pathways for Robotic Manipulation (arXiv 2109.12098)","url":"https://arxiv.org/abs/2109.12098"},{"title":"CLIPort 项目主页","url":"https://cliport.github.io/"}],"as_of":"2021-09","related_ids":["clip","transporter-networks","language-conditioned-policy","tabletop-manipulation","peract","pick-and-place"],"name":"CLIPort","alt":"CLIPort","abbr":"","aliases":["CLIPort: What and Where Pathways for Robotic Manipulation"],"one_liner":"A language-conditioned manipulation policy that combines CLIP's semantic understanding with Transporter Networks' pixel-level spatial precision.","explanation":"CLIPort is work by Mohit Shridhar, Lucas Manuelli, and Dieter Fox at the University of Washington and NVIDIA, published at CoRL 2021. It borrows the neuroscience idea of separate “what” and “where” visual pathways. The semantic pathway uses a pretrained CLIP (an image-text contrastive learning model) to understand “what to act on” — color, shape, object category; the spatial pathway uses the fully convolutional network from Transporter Networks to process RGB-D images and decide “where to pick up, where to place,” outputting pixel-level heatmaps for grasping and placement. Fusing the two lets it follow language instructions on tabletop tasks like packing a box or folding cloth without needing object poses, segmentation masks, or symbolic state. It's an early, representative example of connecting large-scale pretrained image-text models to robot manipulation, and the same authors' later PerAct extends this idea to 3D voxels.","example":"A single multi-task policy learned to follow language instructions across 9 real tabletop tasks using just 179 paired real image-action examples.","related":["CLIP","Transporter Networks","Language-conditioned Policy","Tabletop Manipulation","PerAct","Pick-and-Place"]},{"id":"neural-descriptor-fields","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation (arXiv 2112.05124)","url":"https://arxiv.org/abs/2112.05124"},{"title":"NDF 项目页","url":"https://yilundu.github.io/ndf/"}],"as_of":"2022-05","related_ids":["object-centric-representation","equivariant-policy-equivariant-neural-network","few-shot","pick-and-place","pose","imitation-learning"],"name":"Neural Descriptor Fields","alt":"神经描述子场","abbr":"NDF","aliases":["NDF","SE(3)-Equivariant Object Representations for Manipulation"],"one_liner":"An MIT 2021 object representation that transfers a manipulation skill to new objects of the same category, in new poses, from just a few demonstrations.","explanation":"Neural Descriptor Fields were proposed by Anthony Simeonov, Yilun Du, Pulkit Agrawal, Vincent Sitzmann, and colleagues at MIT, posted to arXiv in December 2021, published at ICRA 2022. It represents an object as a function: given any 3D point near the object, it outputs a descriptor vector, such that functionally equivalent locations on objects of the same category — the handle of different mugs, say — get similar descriptors. The descriptor comes from an occupancy network self-supervised on a 3D-reconstruction task, with no manual keypoint labeling needed. The network is SE(3)-equivariant (however the object rotates or moves, the descriptors transform the same way), so it can handle the object upright, on its side, or upside down. During a demonstration, the descriptors of a set of points near the gripper are recorded; at test time, optimization finds the gripper pose whose descriptors best match them.","example":"Given just 10 demonstrations of “grasp the mug's rim and hang it on a rack,” a Franka arm can hang up mugs it has never seen, in any orientation; the paper reports an overall success rate above 85%.","related":["Object-centric Representation","Equivariant Policy / Equivariant Neural Network","Few-shot","Pick-and-Place","Pose","Imitation Learning"]},{"id":"peract","category":"named_model","sec":3,"tier":3,"sources":[{"title":"arXiv 2209.05451: Perceiver-Actor","url":"https://arxiv.org/abs/2209.05451"},{"title":"PerAct 项目主页","url":"https://peract.github.io/"}],"as_of":"2022-11","related_ids":["rvt-2","3d-diffuser-actor","rlbench","voxel","keyframe-action-prediction","cliport"],"name":"PerAct","alt":"PerAct","abbr":"PerAct","aliases":["Perceiver-Actor","Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation"],"one_liner":"A multi-task manipulation policy that voxelizes the scene and uses a Perceiver Transformer to predict the next key pose.","explanation":"PerAct was introduced by Mohit Shridhar, Lucas Manuelli, and Dieter Fox at the University of Washington and NVIDIA, presented at CoRL 2022. It converts RGB-D observations into a 100×100×100 voxel grid (a 3D grid of cube-shaped cells) and feeds it into a PerceiverIO Transformer together with the language instruction. Rather than outputting a continuous trajectory, PerAct predicts the 'next best voxel': which cell the end effector should move to, a discretized rotation, whether the gripper should open or close, and whether collision avoidance is needed; a motion planner then carries the arm to this key pose. A single model trains jointly on 18 RLBench tasks (249 variations) and 7 real-world tasks, each with only a handful of demonstrations. PerAct established the 'combine a 3D scene representation with keyframe action prediction' recipe, and later work such as RVT, 3D Diffuser Actor, and BridgeVLA all use it as a baseline.","example":"For the instruction 'open the middle drawer,' PerAct first predicts which voxel cell holds the handle and what orientation to grasp it at; once the gripper closes on the handle, it then predicts the end-effector position that pulls the drawer open.","related":["RVT-2","3D Diffuser Actor","RLBench","Voxel","Keyframe Action Prediction","CLIPort"]},{"id":"motion-policy-networks","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Motion Policy Networks (arXiv 2210.12209)","url":"https://arxiv.org/abs/2210.12209"},{"title":"MπNets project page","url":"https://mpinets.github.io/"},{"title":"NVlabs/motion-policy-networks (GitHub)","url":"https://github.com/NVlabs/motion-policy-networks"}],"as_of":"2022-10","related_ids":["neural-motion-planning","motion-planning","point-cloud","obstacle-avoidance","sampling-based-planning","geometric-fabrics"],"name":"Motion Policy Networks","alt":"运动策略网络","abbr":"MπNets","aliases":["MπNets","MPiNets"],"one_liner":"A neural motion planner that generates collision-free robot-arm motion directly from a single depth camera's point cloud.","explanation":"Motion Policy Networks was released in October 2022 by Adam Fishman, Byron Boots, Dieter Fox, and colleagues at the University of Washington and NVIDIA, published at CoRL 2022. Classical motion planning (such as sampling-based planning like RRT) needs a complete, accurate model of the environment and is slow to compute in cluttered scenes. MπNets replaces this with an end-to-end neural network: given the point cloud a single depth camera sees of the scene plus the robot's current state, it outputs joint motion toward the target pose step by step, which strung together forms a smooth, obstacle-avoiding trajectory. All training data is generated automatically in simulation using classical planning tools (OMPL and Geometric Fabrics), covering more than 500,000 environments and more than 3 million planning problems. The result beats prior neural planners by 46%, is much faster than a global planner, can handle dynamic scenes, and, even trained only on simulation data, transfers to noisy partial point clouds on a real robot. Code, weights, and data are open-source.","example":"To reach into one compartment of a cabinet to grab something, MπNets just looks at the depth camera's point cloud once and gives a real-time joint trajectory that routes around the cabinet's panels, with no need to first build a complete map and then run a planner.","related":["Neural Motion Planning","Motion Planning","Point Cloud","Obstacle Avoidance","Sampling-Based Planning","Geometric Fabrics"]},{"id":"behavior-transformer","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Behavior Transformers: Cloning k modes with one stone (arXiv 2206.11251)","url":"https://arxiv.org/abs/2206.11251"},{"title":"Behavior Generation with Latent Actions (VQ-BeT, arXiv 2403.03181)","url":"https://arxiv.org/abs/2403.03181"},{"title":"VQ-BeT 项目主页","url":"https://sjlee.cc/vq-bet/"}],"as_of":"2024-07","related_ids":["action-multimodality","behavior-cloning","vector-quantization","diffusion-policy","action-tokenizer","franka-kitchen"],"name":"Behavior Transformer","alt":"BeT / VQ-BeT","abbr":"BeT","aliases":["BeT","VQ-BeT","Vector-Quantized Behavior Transformer"],"one_liner":"A Transformer-based imitation-learning method that learns several different valid behaviors at once from multimodal demonstration data.","explanation":"BeT (Behavior Transformer) is a behavior-cloning method proposed in 2022 by Lerrel Pinto's group at NYU. Human demonstrations often show “action multimodality”: several equally valid ways to act in the same situation, and direct regression averages them into one wrong action. BeT first clusters continuous actions into a number of bins with k-means, has a Transformer predict which bin to pick, and then predicts a continuous offset to correct it into a precise action — preserving multiple behavior modes this way. VQ-BeT, from 2024 (NYU and Seoul National University, ICML 2024), replaces k-means with residual vector quantization, which suits high-dimensional actions and long action sequences better, running inference at roughly 5x the speed of a diffusion policy. BeT and diffusion policies are two representative lines of work from the same period addressing the same multimodality problem.","example":"In the Franka Kitchen simulated kitchen, demonstrators complete a set of subtasks in different orders each time, and BeT is used to test whether a model can learn and reproduce these different behavior patterns.","related":["Action Multimodality","Behavior Cloning","Vector Quantization","Diffusion Policy","Action Tokenizer","Franka Kitchen"]},{"id":"diffuser","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Planning with Diffusion for Flexible Behavior Synthesis (arXiv:2205.09991)","url":"https://arxiv.org/abs/2205.09991"},{"title":"Diffuser 项目主页","url":"https://diffusion-planning.github.io/"}],"as_of":"2022","related_ids":["diffusion-model","model-based-reinforcement-learning","trajectory-optimization","decision-diffuser","diffusion-policy","offline-reinforcement-learning"],"name":"Diffuser","alt":"Diffuser（扩散规划器）","abbr":"","aliases":["Planning with Diffusion for Flexible Behavior Synthesis"],"one_liner":"A planning method that generates an entire trajectory as one object to be denoised, one of the first uses of diffusion for decision-making.","explanation":"Diffuser is an ICML 2022 paper by Michael Janner and Sergey Levine at UC Berkeley, and Yilun Du and Joshua Tenenbaum at MIT. Traditional model-based reinforcement learning first learns a dynamics model and then plans with an optimizer, and errors in the model are often amplified by the planner. Diffuser merges the two steps: a diffusion model directly models an entire “state plus action” trajectory, and planning becomes starting from noise and repeatedly denoising it into a trajectory. To get high return, the denoising process is guided with the gradient of a value function; to reach a specific goal, the start and end states are fixed and the model fills in the middle (similar to image inpainting). The same model can switch tasks with no retraining. It's a precursor to later “diffusion model for decision-making” work such as Decision Diffuser and Diffusion Policy.","example":"In the Maze2D task, fixing the start and end states, Diffuser “inpaints” a full feasible path in between by denoising; in a block-stacking task, swapping in a different guidance function lets the same model stack towers under different constraints.","related":["Diffusion Model","Model-Based Reinforcement Learning","Trajectory Optimization","Decision Diffuser","Diffusion Policy","Offline Reinforcement Learning"]},{"id":"decision-diffuser","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Is Conditional Generative Modeling all you need for Decision-Making? (arXiv:2211.15657)","url":"https://arxiv.org/abs/2211.15657"},{"title":"Decision Diffuser 项目主页（ICLR 2023 Oral）","url":"https://anuragajay.github.io/decision-diffuser/"}],"as_of":"2023-07","related_ids":[null,"diffusion-model","classifier-free-guidance","inverse-dynamics-model","offline-reinforcement-learning","return-conditioning"],"name":"Decision Diffuser","alt":"Decision Diffuser（决策扩散器）","abbr":"","aliases":["Is Conditional Generative Modeling all you need for Decision-Making?"],"one_liner":"Generates future state trajectories with a return-conditioned diffusion model, then infers actions from them, skipping dynamic programming.","explanation":"Decision Diffuser was proposed in November 2022 by Pulkit Agrawal's group (Improbable AI Lab) and CSAIL at MIT, an oral presentation at ICLR 2023. Offline reinforcement learning usually needs to learn a value function and do dynamic programming (repeatedly using the Bellman equation to estimate long-term return), which tends to be unstable to train. Decision Diffuser instead treats decision-making as conditional generation: a diffusion model generates only a sequence of future states, conditioned on a desired return, a constraint, or a skill, strengthened with classifier-free guidance; an inverse-dynamics model then infers the action to take from each pair of adjacent states. It beat the leading methods of the time on the D4RL offline reinforcement-learning benchmark, and even though it only ever saw a single constraint or skill at training time, at test time it could combine multiple conditions together. It's a representative follow-up to Diffuser in using diffusion models for decision-making, in the same lineage as later “generate the future, then infer actions with inverse dynamics” approaches like UniPi.","example":"In a Kuka-arm block-stacking experiment, each training trajectory satisfies only a single constraint of the form “A is on top of B,” but at test time several constraints are given as conditions together, and Decision Diffuser generates a stacking plan that satisfies all of them at once.","related":["Diffuser","Diffusion Model","Classifier-Free Guidance","Inverse Dynamics Model","Offline Reinforcement Learning","Return Conditioning"]},{"id":"diffusion-policy","category":"named_model","sec":3,"tier":1,"sources":[{"title":"Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)","url":"https://arxiv.org/abs/2303.04137"},{"title":"Diffusion Policy 项目主页","url":"https://diffusion-policy.cs.columbia.edu/"}],"as_of":"2024-03","related_ids":["diffusion-model","action-multimodality","denoising-diffusion-probabilistic-model","action-chunking","3d-diffusion-policy","push-t"],"name":"Diffusion Policy","alt":"扩散策略","abbr":"DP","aliases":["DP","Diffusion Policy: Visuomotor Policy Learning via Action Diffusion"],"one_liner":"An imitation-learning policy that generates robot action sequences by gradually denoising from random noise, using a diffusion model.","explanation":"Diffusion Policy was introduced by Cheng Chi and colleagues in Shuran Song's group at Columbia University, working with Toyota Research Institute and MIT; it appeared on arXiv in March 2023, was published at RSS 2023, and an extended version appeared in IJRR in 2024. It formulates a robot policy as a conditional denoising diffusion process: conditioned on observations such as camera images, it starts from random noise and denoises step by step into a chunk of future actions. The advantage is that it can represent multimodal actions — when a scene has several equally valid ways to act, it doesn't average them into one wrong action the way direct regression would. The paper also combines this with receding-horizon control (predict a chunk, execute part of it, then re-predict), and beat the best prior methods by an average of 46.9% across 12 tasks in 4 benchmarks. It became an important source for the diffusion and flow-matching action heads later used in models like 3D Diffusion Policy and π0.","example":"Push-T is its signature task: the robot uses the end of a cylindrical rod to push a T-shaped block on a table into a target position and orientation. Since going around from the left or the right both work, and both appear in the demonstrations, Diffusion Policy learns both, but commits to just one of them on each run.","related":["Diffusion Model","Action Multimodality","Denoising Diffusion Probabilistic Model","Action Chunking","3D Diffusion Policy","Push-T"]},{"id":"action-chunking-with-transformers","category":"named_model","sec":3,"tier":1,"sources":[{"title":"Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705)","url":"https://arxiv.org/abs/2304.13705"},{"title":"ALOHA / ACT 项目主页","url":"https://tonyzhaozh.github.io/aloha/"}],"as_of":"2023-04","related_ids":["action-chunking","temporal-ensembling","conditional-variational-autoencoder","aloha","mobile-aloha","behavior-cloning"],"name":"Action Chunking with Transformers","alt":"ACT","abbr":"ACT","aliases":["ACT","Action Chunking Transformer"],"one_liner":"A 2023 Stanford imitation-learning policy that predicts a whole chunk of actions at once, letting low-cost bimanual robots do fine manipulation.","explanation":"ACT is an imitation-learning algorithm introduced in April 2023 by Tony Zhao, Chelsea Finn, and colleagues at Stanford, with collaborators from UC Berkeley and Meta, published alongside the low-cost bimanual platform ALOHA (roughly $20,000 in hardware) at RSS 2023. Behavior cloning — learning actions by copying human demonstrations — tends to accumulate small errors at every step, which is especially damaging for fine-grained tasks. ACT has a Transformer output an entire chunk of future actions at once, so fewer decisions are made and less error accumulates. During training it wraps a conditional variational autoencoder (CVAE, a generative model that can represent multiple valid ways of doing the same thing) around the policy to absorb the randomness in human demonstrations; at execution time, temporal ensembling averages overlapping action chunks with weights to keep motion smooth. Because it's structurally simple and needs little data, ACT went on to become a standard baseline for projects like Mobile ALOHA and LeRobot.","example":"On ALOHA, ACT learned six fine-grained tasks — such as opening a translucent condiment cup and slotting a battery into place — from only about 10 minutes of teleoperated demonstrations per task, reaching 80–90% success rates.","related":["Action Chunking","Temporal Ensembling","Conditional Variational Autoencoder","ALOHA","Mobile ALOHA","Behavior Cloning"]},{"id":"roboagent","category":"named_model","sec":3,"tier":3,"sources":[{"title":"arXiv 2309.01918: RoboAgent","url":"https://arxiv.org/abs/2309.01918"},{"title":"RoboAgent 项目主页（RoboPen）","url":"https://robopen.github.io/"}],"as_of":"2024-05","related_ids":["action-chunking","action-chunking-with-transformers","data-augmentation","segment-anything-model","multi-task-learning","language-conditioned-policy"],"name":"RoboAgent","alt":"RoboAgent（MT-ACT）","abbr":"MT-ACT","aliases":["MT-ACT","Multi-Task Action Chunking Transformer","RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking"],"one_liner":"A multi-skill kitchen manipulation agent from CMU and Meta, trained on just 7,500 demonstrations.","explanation":"RoboAgent was released by Carnegie Mellon University and Meta AI in September 2023 (accepted at ICRA 2024), addressing the question of whether a small amount of real-robot data can still train a robot with many skills, given how expensive that data is. The approach combines two ideas. First, semantic augmentation: the Segment Anything Model (SAM) segments objects in each frame, and their shape, color, and texture are then altered to multiply the existing data. Second, MT-ACT (Multi-Task Action Chunking Transformer), which extends ACT's action chunking (predicting a short block of future actions at once) to a multi-task setting where a language instruction distinguishes between tasks. Using only 7,500 teleoperated trajectories, a single policy learned 12 skills across 38 kitchen tasks, outperforming prior methods by more than 40% in unseen scenarios. The associated data was open-sourced under the name RoboSet.","example":"A single demonstration of 'pull open the drawer' can be semantically augmented into many training samples with different-looking drawers and countertops, all used together to train MT-ACT.","related":["Action Chunking","Action Chunking with Transformers","Data Augmentation","Segment Anything Model","Multi-Task Learning","Language-conditioned Policy"]},{"id":"dobb-e","category":"named_model","sec":3,"tier":3,"sources":[{"title":"On Bringing Robots Home (arXiv 2311.16098)","url":"https://arxiv.org/abs/2311.16098"},{"title":"Dobb·E 项目主页","url":"https://dobb-e.com/"}],"as_of":"2023-11","related_ids":["handheld-gripper-data-collection","universal-manipulation-interface","hello-robot-stretch","household-tasks","pre-trained-visual-representation","robot-utility-models"],"name":"Dobb-E","alt":"Dobb·E","abbr":"Dobb-E","aliases":["On Bringing Robots Home","An Open-Source, General Framework for Learning Household Robotic Manipulation"],"one_liner":"An open-source household-robot framework from NYU and Meta, released in 2023, that teaches a robot a new task from just 5 minutes of demonstration.","explanation":"Dobb-E is an open-source household robot-learning framework released in November 2023 by Lerrel Pinto's group at NYU together with Meta, in a paper titled On Bringing Robots Home. Real homes vary enormously in lighting, furniture, and objects, so policies trained in a lab often fail once moved into a house, while collecting robot data inside people's homes is expensive. The team built a demonstration tool called the Stick: a $25 reacher-grabber fitted with a few 3D-printed parts and an iPhone, which an ordinary person can hold and use to record demonstrations just by going about a chore. Using it, they recorded 13 hours of data across 22 New York households (the HoNY dataset), pretrained a visual representation called HPR (built on ResNet-34) with self-supervised learning, and deployed it on the commercial mobile robot Hello Robot Stretch, where a new task needs only a handful of demonstrations plus brief fine-tuning. Code, data, models, and hardware designs are all open-source. Dobb-E is an early example of the “handheld gripper data collection” approach, and the same team later built Robot Utility Models.","example":"Over about 30 days of testing across 10 households in the New York area, Dobb-E attempted 109 household tasks, using just 5 minutes of demonstration and 15 minutes of model adaptation per new task, reaching an overall success rate of 81%.","related":["Handheld Gripper Data Collection","Universal Manipulation Interface","Hello Robot Stretch","Household Tasks","Pre-trained Visual Representation","Robot Utility Models"]},{"id":"mobile-aloha","category":"named_model","sec":3,"tier":1,"sources":[{"title":"Mobile ALOHA (arXiv 2401.02117)","url":"https://arxiv.org/abs/2401.02117"},{"title":"Mobile ALOHA 项目主页","url":"https://mobile-aloha.github.io/"}],"as_of":"2024-01","related_ids":["aloha","action-chunking-with-transformers","mobile-manipulation","co-training","whole-body-teleoperation","agilex-robotics"],"name":"Mobile ALOHA","alt":"Mobile ALOHA","abbr":"","aliases":["Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation"],"one_liner":"Stanford's 2024 project that puts the ALOHA bimanual robot on a mobile base for low-cost whole-body teleoperation and imitation learning.","explanation":"Mobile ALOHA is work published in January 2024 by Zipeng Fu, Tony Zhao, and Chelsea Finn at Stanford, presented at CoRL 2024. It mounts the previously tabletop-fixed ALOHA bimanual platform onto a Tracer mobile base from AgileX Robotics; the operator's waist is linked to the base so that walking moves the base along, while both hands drive the master arms, teleoperating the arms and base together. The full hardware budget is about $32,000. Actions are 16-dimensional: 14 joint positions for the two arms, plus the base's linear and angular velocity. The problem it targets is data collection and learning for mobile manipulation — using both hands while on the move. Another key finding is co-training: training on the new mobile data together with existing stationary ALOHA data raises success rates by up to 90%, using only about 50 demonstrations per task. The policy itself uses off-the-shelf imitation-learning methods such as ACT and Diffusion Policy.","example":"Mobile ALOHA has autonomously stir-fried shrimp and plated it, opened a two-door cabinet to store a heavy pot, called and boarded an elevator, and rinsed a used pan under a faucet.","related":["ALOHA","Action Chunking with Transformers","Mobile Manipulation","Co-training","Whole-Body Teleoperation","AgileX Robotics"]},{"id":"aloha-unleashed","category":"named_model","sec":3,"tier":3,"sources":[{"title":"ALOHA Unleashed: A Simple Recipe for Robot Dexterity (arXiv 2410.13126)","url":"https://arxiv.org/abs/2410.13126"},{"title":"ALOHA Unleashed 项目页","url":"https://aloha-unleashed.github.io/"},{"title":"Proceedings of The 8th Conference on Robot Learning, PMLR 270","url":"https://proceedings.mlr.press/v270/zhao25b.html"}],"as_of":"2024-10","related_ids":["aloha-2","diffusion-policy","action-chunking-with-transformers","bimanual-manipulation","imitation-learning","mobile-aloha"],"name":"ALOHA Unleashed","alt":"ALOHA Unleashed","abbr":"","aliases":["ALOHA Unleashed: A Simple Recipe for Robot Dexterity"],"one_liner":"DeepMind's 2024 work: massive teleoperated data plus a diffusion policy teach a bimanual robot to tie shoelaces.","explanation":"ALOHA Unleashed is work by Tony Zhao, Ayzaan Wahid, Chelsea Finn, and colleagues at Google DeepMind, published at CoRL 2024, asking how far pure imitation learning alone can push bimanual dexterity. The recipe is deliberately simple — hence the name. On the low-cost bimanual platform ALOHA 2, 35 operators teleoperated under a standardized protocol to collect more than 26,000 demonstrations across 5 real tasks; the policy is a Transformer encoder-decoder with a diffusion loss (that is, a diffusion policy), about 217 million parameters, predicting the next 50 steps of action in one shot. This let the robot autonomously tie shoelaces, hang a shirt on a hanger, and swap a fingertip on another robot. The paper also found that swapping in ACT-style L1 regression, at the same model size, dropped the shirt task's success rate from 70% to 25%.","example":"In the shoelace-tying task, the robot first centers the shoe on the table and straightens the laces, then ties a bow. With the shoe centered and the laces already laid flat, success was 70%; with the shoe rotated up to ±45° and the laces unstraightened, success dropped to 40%.","related":["ALOHA 2","Diffusion Policy","Action Chunking with Transformers","Bimanual Manipulation","Imitation Learning","Mobile ALOHA"]},{"id":"3d-diffuser-actor","category":"named_model","sec":3,"tier":3,"sources":[{"title":"3D Diffuser Actor: Policy Diffusion with 3D Scene Representations (arXiv 2402.10885)","url":"https://arxiv.org/abs/2402.10885"},{"title":"3D Diffuser Actor 项目页","url":"https://3d-diffuser-actor.github.io/"}],"as_of":"2024-07","related_ids":["diffusion-policy","3d-diffusion-policy","peract","rlbench","calvin-benchmark","keyframe-action-prediction"],"name":"3D Diffuser Actor","alt":"3D Diffuser Actor","abbr":"","aliases":["3D Diffuser Actor: Policy Diffusion with 3D Scene Representations"],"one_liner":"An imitation-learning policy that combines diffusion policies with 3D scene representations to generate robot-arm end-effector trajectories.","explanation":"3D Diffuser Actor was proposed in February 2024 by Katerina Fragkiadaki's group at Carnegie Mellon University (Tsung-Wei Ke, Nikolaos Gkanatsios), published at CoRL 2024. Two earlier lines of work each had a strength: diffusion policies can represent multiple valid ways of doing something (action multimodality), while 3D policies fuse multi-view images into 3D features using depth, making them more robust to camera-viewpoint changes. This model combines both: it lifts image features into 3D points using depth, then uses a denoising Transformer with 3D relative-position attention, conditioned on the language instruction and proprioceptive state, to progressively denoise a noised trajectory of end-effector poses. At release it beat the previous best method by 18.1 absolute percentage points in the multi-view RLBench setting and 13.1 points single-view, with a 9% relative improvement on CALVIN; on a real Franka arm, it learned 12 tasks from roughly a dozen demonstrations each.","example":"On RLBench's “open drawer” task, it builds a 3D feature point cloud of the scene from multiple RGB-D cameras, then repeatedly denoises from random noise to produce the arm's next key pose — a position in front of the handle with a specific gripper orientation.","related":["Diffusion Policy","3D Diffusion Policy","PerAct","RLBench","CALVIN Benchmark","Keyframe Action Prediction"]},{"id":"3d-diffusion-policy","category":"named_model","sec":3,"tier":2,"sources":[{"title":"3D Diffusion Policy (arXiv 2403.03954)","url":"https://arxiv.org/abs/2403.03954"},{"title":"3D Diffusion Policy 项目主页","url":"https://3d-diffusion-policy.github.io/"}],"as_of":"2024-09","related_ids":["diffusion-policy","point-cloud","point-cloud-encoder","idp3","imitation-learning","shanghai-qi-zhi-institute"],"name":"3D Diffusion Policy","alt":"3D 扩散策略","abbr":"DP3","aliases":["DP3","3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations"],"one_liner":"A diffusion policy that takes sparse point clouds as input and learns manipulation from very few demonstrations.","explanation":"3D Diffusion Policy (DP3) was proposed in March 2024 by researchers at the Shanghai Qi Zhi Institute and Huazhe Xu's group at Tsinghua University's Institute for Interdisciplinary Information Sciences, together with Shanghai Jiao Tong University and the Shanghai AI Lab (first author Yanjie Ze), and published at RSS 2024. Building on Diffusion Policy, it replaces 2D images with a sparse point cloud from a depth camera (a set of points with 3D coordinates) as input, compresses it with a lightweight point-cloud encoder into a compact 3D feature, and denoises actions conditioned on that feature. Because a 3D representation directly captures an object's position and geometry, it's less sensitive to changes in viewpoint or appearance, so it needs fewer demonstrations. The paper used just 10 demonstrations per task across 72 simulated tasks, improving 24.2% over baselines, and 40 demonstrations per task on 4 real-robot tasks, reaching 85% success with very few safety-violating actions. A later extension, iDP3, scales it up to humanoid robots.","example":"Real-robot experiments included rolling play-dough and folding it into a “dumpling” with an Allegro dexterous hand or a parallel gripper, touching a block with a power drill, and pouring out dried pork floss — each task using 40 demonstrations.","related":["Diffusion Policy","Point Cloud","Point Cloud Encoder","iDP3","Imitation Learning","Shanghai Qi Zhi Institute"]},{"id":"idp3","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Generalizable Humanoid Manipulation with 3D Diffusion Policies (arXiv 2410.10803)","url":"https://arxiv.org/abs/2410.10803"},{"title":"Project page","url":"https://humanoid-manipulation.github.io/"},{"title":"YanjieZe/Improved-3D-Diffusion-Policy (GitHub)","url":"https://github.com/YanjieZe/Improved-3D-Diffusion-Policy"}],"as_of":"2025-09","related_ids":["3d-diffusion-policy","diffusion-policy","point-cloud","humanoid-robot","fourier-gr-1","vr-teleoperation"],"name":"iDP3","alt":"iDP3","abbr":"iDP3","aliases":["Improved 3D Diffusion Policy","Generalizable Humanoid Manipulation with 3D Diffusion Policies"],"one_liner":"An improved 3D Diffusion Policy that lets a humanoid trained on data from a single scene generalize to new scenes.","explanation":"iDP3 was released in October 2024 by Yanjie Ze, Jiajun Wu, and colleagues at Stanford, together with Simon Fraser University, UPenn, UIUC, and CMU, published at IROS 2025, an improved version of 3D Diffusion Policy (DP3). The original DP3 processes point clouds in world coordinates, which requires camera calibration and cropping the point cloud by segmentation — not a good fit for a humanoid whose head camera moves along with its body. iDP3 switches to an “egocentric” 3D point cloud in camera coordinates instead, eliminating the need for calibration and segmentation, while also scaling up the input point-cloud size, improving the vision encoder, and lengthening the action-prediction horizon. The accompanying system includes upper-body teleoperation based on an Apple Vision Pro, and a 25-degree-of-freedom Fourier GR1 humanoid platform (with a RealSense L515 depth camera) mounted on a height-adjustable cart. Using data collected from just a single scene, the robot can complete tasks in many new scenes using only its onboard compute.","example":"A policy trained only on demonstration data collected in one scene can, once deployed, perform the same kind of manipulation in other real-world scenes it has never seen.","related":["3D Diffusion Policy","Diffusion Policy","Point Cloud","Humanoid Robot","Fourier GR-1","VR Teleoperation"]},{"id":"consistency-policy","category":"named_model","sec":3,"tier":3,"sources":[{"title":"Consistency Policy (arXiv:2405.07503)","url":"https://arxiv.org/abs/2405.07503"},{"title":"Consistency Policy 项目主页（RSS 2024）","url":"https://consistency-policy.github.io/"}],"as_of":"2024-06","related_ids":["diffusion-policy","consistency-model","one-step-generation","knowledge-distillation","inference-latency","conrft"],"name":"Consistency Policy","alt":"Consistency Policy（一致性策略）","abbr":"","aliases":["Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation"],"one_liner":"Distills a diffusion policy into a fast visuomotor policy that generates an action in a single step.","explanation":"Consistency Policy was proposed in May 2024 by Jeannette Bohg's group at Stanford (first author Aaditya Prasad) with Jimmy Wu and others at Princeton, published at RSS 2024. Diffusion Policy produces high-quality actions, but generating one requires tens to hundreds of denoising steps, too slow for a laptop GPU or onboard computer. Consistency Policy treats a trained diffusion policy as a “teacher” and distills it using the consistency trajectory model (CTM) objective: the student network is trained so that starting from any point along the denoising trajectory, it can jump straight to the same endpoint (“self-consistency”), so at inference it generates an entire action in one step. Tested across 6 simulated tasks (Robomimic, Push-T, Franka Kitchen) and 3 real-robot tasks, it runs an order of magnitude faster than comparable methods with roughly the same success rate. It's a representative approach for speeding up diffusion policies, and later work such as ConRFT uses it as the policy backbone.","example":"In real-robot experiments, a robot running Consistency Policy on just a laptop-class GPU completed tasks like plugging in a cord and cleaning up trash, using single-step generation in place of Diffusion Policy's many denoising steps.","related":["Diffusion Policy","Consistency Model","One-step Generation","Knowledge Distillation","Inference Latency","ConRFT"]},{"id":"rvt-2","category":"named_model","sec":3,"tier":3,"sources":[{"title":"RVT-2: Learning Precise Manipulation from Few Demonstrations (arXiv 2406.08545)","url":"https://arxiv.org/abs/2406.08545"},{"title":"RVT-2 project page","url":"https://robotic-view-transformer-2.github.io/"}],"as_of":"2024-06","related_ids":["peract","keyframe-action-prediction","multi-view","rlbench","peg-in-hole-insertion","few-shot"],"name":"RVT-2","alt":"RVT-2","abbr":"RVT-2","aliases":["Robotic View Transformer 2","RVT-2: Learning Precise Manipulation from Few Demonstrations"],"one_liner":"NVIDIA's multi-view 3D manipulation policy that learns millimeter-precision insertion from about 10 demonstrations per task.","explanation":"RVT-2 was released by an NVIDIA research team (Ankit Goyal, Dieter Fox, and others) in June 2024, published at RSS 2024, an upgrade of RVT (Robotic View Transformer). Methods in this family first re-render the point cloud from an RGB-D camera into several virtual-viewpoint images, then use a Transformer on these images to predict the next key pose — where the gripper should go, its orientation, and whether to open or close — before handing off to a motion planner for execution. RVT-2 adds coarse-to-fine multi-stage inference: it first finds a rough region in the whole scene, then zooms in on that region for precise prediction; it also conditions rotation prediction on position, and speeds things up with a custom renderer and a more efficient training implementation. The result is 6x faster training and 2x faster inference than RVT, with multi-task success on RLBench rising from 65% to 82%; on a real robot, using just one RGB-D camera and about 10 demonstrations per task, it can perform millimeter-precision tasks like inserting a peg or plugging in a plug.","example":"About 10 demonstrations of 'insert the pin into the hole' are recorded on a real robot arm, and RVT-2 learns to align and insert into a hole with only a very small clearance.","related":["PerAct","Keyframe Action Prediction","Multi-View","RLBench","Peg-in-Hole Insertion","Few-shot"]},{"id":"robot-utility-models","category":"named_model","sec":3,"tier":3,"sources":[{"title":"arXiv 2409.05865: Robot Utility Models","url":"https://arxiv.org/abs/2409.05865"},{"title":"Robot Utility Models 项目主页","url":"https://robotutilitymodels.com/"}],"as_of":"2024-09","related_ids":["zero-shot","scene-generalization","handheld-gripper-data-collection","dobb-e","behavior-transformer","specialist-policy"],"name":"Robot Utility Models","alt":"Robot Utility Models","abbr":"RUM","aliases":["RUM","RUMs","Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments"],"one_liner":"Policies that perform single tasks such as opening a cabinet or a drawer in unfamiliar homes with no fine-tuning at all.","explanation":"Robot Utility Models (RUM) was released by researchers at New York University, Hello Robot, and Meta in September 2024. The problem it addresses: robot policies usually need fresh data collection and fine-tuning every time they move to a new room, whereas language and vision models can be used off the shelf. The authors trained one specialist policy each for five tasks — opening a cabinet, opening a drawer, picking up a tissue, picking up a paper bag, and righting a fallen object — using Stick-v2, a handheld gripper fitted with an iPhone, to collect about 1,000 demonstrations per task across roughly 40 environments, with VQ-BeT as the policy architecture. At deployment, GPT-4o judges whether an attempt succeeded and triggers a retry on failure. The result is about 90% success on unseen environments and objects, and the policies transfer to other robots such as the xArm. The authors' main conclusion is that data diversity matters more than the training algorithm or policy architecture. Code, data, and hardware designs are all open-sourced.","example":"RUM's cabinet-opening policy is mounted on a Hello Robot Stretch and placed in an apartment where no data was ever collected; without any fine-tuning, it goes and opens the kitchen cabinet.","related":["Zero-shot","Scene Generalization","Handheld Gripper Data Collection","Dobb-E","Behavior Transformer","Specialist Policy"]},{"id":"3d-vitac","category":"named_model","sec":3,"tier":3,"sources":[{"title":"3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv 2410.24091)","url":"https://arxiv.org/abs/2410.24091"},{"title":"3D-ViTac 项目页","url":"https://binghao-huang.github.io/3D-ViTac/"}],"as_of":"2024-10","related_ids":["visuo-tactile-fusion","tactile-sensor","diffusion-policy","point-cloud","3d-diffusion-policy","bimanual-manipulation"],"name":"3D-ViTac","alt":"3D-ViTac","abbr":"","aliases":["3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing"],"one_liner":"A system that places low-cost tactile-sensor readings and a visual point cloud in the same 3D space to learn fine manipulation.","explanation":"3D-ViTac was proposed in October 2024 by researchers from Columbia University, UIUC, and the University of Washington (first author Binghao Huang, advised by Yunzhu Li), published at CoRL 2024. Vision alone isn't enough when the gripper itself blocks the camera's view of an object, or when the task requires judging how hard to squeeze. The system covers the soft gripper's fingertips with flexible piezoresistive tactile pads: each pad has a 16×16 grid of 256 sensing cells, each about 3 square millimeters and under 1 millimeter thick, costing about $20, for 1,024 cells total across a bimanual four-finger setup. The key idea is to use robot kinematics to compute each tactile cell's 3D position, turning its reading into a “tactile point,” merged into a single point cloud together with the camera's point cloud, and then trained with imitation learning using a diffusion policy. Across four long-horizon tasks — steaming an egg, preparing grapes, collecting hex keys, and handing over a sandwich — overall success reached 80–90%, versus only 45–60% for a vision-only version.","example":"While gripping an egg, the camera can't tell how much force is being applied, but the tactile points directly show contact location and pressure, letting the policy hold it gently without crushing it; when reorienting a hex key in the gripper, the tactile points tell the policy the key's current orientation.","related":["Visuo-Tactile Fusion","Tactile Sensor","Diffusion Policy","Point Cloud","3D Diffusion Policy","Bimanual Manipulation"]},{"id":"octo","category":"named_model","sec":4,"tier":2,"sources":[{"title":"Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)","url":"https://arxiv.org/abs/2405.12213"},{"title":"Octo 项目主页","url":"https://octo-models.github.io/"},{"title":"RSS 2024 论文页（Robotics: Science and Systems XX, p090）","url":"https://www.roboticsproceedings.org/rss20/p090.html"}],"as_of":"2024-07","related_ids":["open-x-embodiment","generalist-policy","diffusion-action-head","cross-embodiment","openvla","rt-x"],"name":"Octo","alt":"Octo","abbr":"","aliases":["Octo: An Open-Source Generalist Robot Policy","Octo-Small","Octo-Base"],"one_liner":"A 2024 open-source generalist robot policy pretrained on 800,000 cross-embodiment trajectories, quick to fine-tune to new robots.","explanation":"Octo was released in May 2024 by the Octo Model Team, a group of researchers from UC Berkeley, Stanford, CMU, and Google DeepMind, and published at RSS 2024. It's a Transformer-based generalist policy pretrained on about 800,000 trajectories from 25 sub-datasets of Open X-Embodiment (a cross-embodiment robot dataset pooled from many institutions), released in two sizes: Octo-Small (27 million parameters) and Octo-Base (93 million parameters). A task can be specified either with a language instruction or with a single goal image; the Transformer backbone reads in task and observation tokens, and a diffusion action head outputs continuous actions, able to represent multimodal action distributions. Its emphasis is on being open and easy to modify: code, weights, and the training pipeline are all public, the input and output modules are modular, and it can be fine-tuned to a new sensor setup and action space on a consumer GPU in a few hours. It beat the next-best baseline by an average of 52% across 6 fine-tuning evaluation scenarios, and later work such as OpenVLA often uses it as a comparison baseline.","example":"In the “Berkeley Bimanual” task, Octo — pretrained only on single-arm data — was fine-tuned onto an ALOHA bimanual platform made of two ViperX arms: the right arm picks up a marker from the table while the left arm removes its cap, with the action space switched to joint-position control.","related":["Open X-Embodiment","Generalist Policy","Diffusion Action Head","Cross-Embodiment","OpenVLA","RT-X"]},{"id":"openvla","category":"named_model","sec":4,"tier":1,"sources":[{"title":"OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)","url":"https://arxiv.org/abs/2406.09246"},{"title":"OpenVLA 项目主页","url":"https://openvla.github.io/"}],"as_of":"2024-09","related_ids":["vision-language-action-model","rt-2","open-x-embodiment","prismatic-vlms","openvla-oft","lora"],"name":"OpenVLA","alt":"OpenVLA","abbr":"","aliases":["OpenVLA-7B","OpenVLA: An Open-Source Vision-Language-Action Model"],"one_liner":"A 7-billion-parameter vision-language-action model that Stanford, UC Berkeley, and collaborators open-sourced in 2024, code and weights included.","explanation":"OpenVLA was released in June 2024 by researchers from Stanford, UC Berkeley, Toyota Research Institute, Google DeepMind, Physical Intelligence, and MIT, and published at CoRL 2024. It's built on the Prismatic vision-language model: the vision encoder fuses features from SigLIP and DINOv2, the language model is Llama 2 7B, and the total parameter count is about 7 billion. For actions, it follows RT-2's approach of discretizing each action dimension into 256 bins and having the language model predict them as tokens, one at a time. It was trained on 970,000 robot trajectories from the Open X-Embodiment dataset, and beat the 55-billion-parameter closed-source RT-2-X by 16.5 absolute percentage points in success rate across 29 tasks. Just as important, its code and weights are fully open, and it can be fine-tuned on consumer GPUs with LoRA (low-rank adaptation, which trains only a small number of added parameters) — which made it a common starting point and baseline for later VLA research, such as OpenVLA-OFT.","example":"To adapt OpenVLA to a new robot arm and task, fine-tuning only about 1.4% of its parameters with LoRA achieves results comparable to full-parameter fine-tuning.","related":["Vision-Language-Action Model","RT-2","Open X-Embodiment","Prismatic VLMs","OpenVLA-OFT","LoRA"]},{"id":"openvla-oft","category":"named_model","sec":4,"tier":2,"sources":[{"title":"OpenVLA-OFT 项目主页","url":"https://openvla-oft.github.io/"},{"title":"Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (arXiv:2502.19645)","url":"https://arxiv.org/abs/2502.19645"}],"as_of":"2025-04","related_ids":["openvla","vision-language-action-model","action-chunking","parallel-decoding","feature-wise-linear-modulation","libero-benchmark"],"name":"OpenVLA-OFT","alt":"OpenVLA-OFT","abbr":"OFT","aliases":["OFT","OpenVLA-OFT+","Optimized Fine-Tuning (VLA recipe)"],"one_liner":"Stanford's fine-tuning recipe for VLAs that makes OpenVLA generate actions 26 times faster, with a higher success rate.","explanation":"OpenVLA-OFT is a VLA fine-tuning recipe proposed by Moo Jin Kim, Chelsea Finn, and Percy Liang at Stanford in February 2025. The original OpenVLA discretizes actions into tokens and outputs them one at a time, autoregressively, which is slow and hard to use for high-frequency control. OFT changes four things during fine-tuning: parallel decoding (computing all actions in a single forward pass), action chunking (predicting multiple steps at once), a continuous action representation, and an L1 regression loss; OFT+ additionally adds FiLM feature modulation to strengthen how well the model follows language instructions. The result: average success rate on LIBERO rose from 76.5% to 97.1%, action-generation throughput increased 26x, and on a real ALOHA bimanual robot it beat other fine-tuned VLAs, including π0 and RDT-1B. It's now commonly used as a strong baseline for VLA fine-tuning.","example":"When fine-tuning OpenVLA on LIBERO, switching the output from discrete action tokens emitted one at a time to an entire continuous action chunk regressed in a single forward pass sharply raises inference throughput and also improves success rate.","related":["OpenVLA","Vision-Language-Action Model","Action Chunking","Parallel Decoding","Feature-wise Linear Modulation","LIBERO Benchmark"]},{"id":"crossformer","category":"named_model","sec":4,"tier":3,"sources":[{"title":"Scaling Cross-Embodied Learning / CrossFormer (arXiv:2408.11812)","url":"https://arxiv.org/abs/2408.11812"},{"title":"CrossFormer 项目主页（CoRL 2024 Oral）","url":"https://crossformer-model.github.io/"}],"as_of":"2024-08","related_ids":["cross-embodiment","octo","embodiment-specific-head","open-x-embodiment","heterogeneous-pre-trained-transformers","action-chunking"],"name":"CrossFormer","alt":"CrossFormer","abbr":"","aliases":["Scaling Cross-Embodied Learning","CrossFormer: Scaling Cross-Embodied Learning (One Policy for Manipulation, Navigation, Locomotion and Aviation)"],"one_liner":"A single set of weights that controls single arms, bimanual arms, quadrupeds, ground vehicles, and drones with one cross-embodiment Transformer.","explanation":"CrossFormer was proposed in August 2024 by Sergey Levine's group at Berkeley with CMU, an oral presentation at CoRL 2024. Different robots vary in camera count, proprioception, action dimensionality, and control frequency, and earlier cross-embodiment training often had to hand-align observation and action spaces, or drop some inputs. CrossFormer instead cuts multi-camera images, proprioception, and the task (language or a goal image) all into tokens arranged in a sequence, fed into a single decoder-only Transformer shared across all embodiments; readout tokens are inserted into the sequence, then routed to different action heads by embodiment category, each outputting an action chunk of the matching dimensionality — such as a 7D end-effector delta for a single arm, 14D joint positions for a bimanual robot, or a 2D waypoint for navigation. The model, about 130 million parameters, is trained on 900,000 trajectories across 20 embodiments; on real robots it matches embodiment-specific policies and clearly beats earlier cross-embodiment methods.","example":"The same CrossFormer outputs a 4-step, 7-dimensional end-effector action chunk when connected to a single-arm WidowX, a 100-step, 14-dimensional joint-position chunk when connected to a bimanual ALOHA, and a 12-dimensional joint target at every step when connected to a Go1 quadruped.","related":["Cross-Embodiment","Octo","Embodiment-specific Head","Open X-Embodiment","Heterogeneous Pre-trained Transformers","Action Chunking"]},{"id":"heterogeneous-pre-trained-transformers","category":"named_model","sec":4,"tier":3,"sources":[{"title":"Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers (arXiv 2409.20537)","url":"https://arxiv.org/abs/2409.20537"},{"title":"HPT 项目页","url":"https://liruiw.github.io/hpt/"}],"as_of":"2024-12","related_ids":["cross-embodiment","heterogeneous-data","pre-training","embodiment-specific-head","scaling-law","open-x-embodiment"],"name":"Heterogeneous Pre-trained Transformers","alt":"异构预训练 Transformer","abbr":"HPT","aliases":["HPT"],"one_liner":"A 2024 method from Kaiming He's group at MIT that pretrains one shared backbone jointly across many different robots' data.","explanation":"HPT was released in September 2024 by Lirui Wang and Kaiming He at MIT CSAIL together with Xinlei Chen at Meta FAIR, published at NeurIPS 2024 as a Spotlight paper. Robot data is highly “heterogeneous”: different robots vary in number of cameras, joint count, and control scheme, making it hard to train them all together inside one model. HPT splits the network into three parts: a small stem (encoder) per embodiment that converts proprioception and images into a fixed number of tokens; a large, shared Transformer trunk in the middle that learns a representation independent of embodiment and task; and an output-side head that maps that representation into an action for a given task. Pretraining used more than 50 datasets and about 200,000 trajectories, drawn from real-robot teleoperation, simulation, human video, and already-deployed robots.","example":"Attaching the pretrained trunk to a new robot only requires training a new stem and head for it; the authors report more than 20% improvement in fine-tuned policy performance across several simulation benchmarks and unseen real-robot tasks.","related":["Cross-Embodiment","Heterogeneous Data","Pre-training","Embodiment-specific Head","Scaling Law","Open X-Embodiment"]},{"id":"rdt-1b","category":"named_model","sec":4,"tier":2,"sources":[{"title":"RDT-1B 项目主页","url":"https://rdt-robotics.github.io/rdt-robotics/"}],"as_of":"2024-10","related_ids":["diffusion-policy","diffusion-transformer","bimanual-manipulation","unified-action-space","action-multimodality","rdt2"],"name":"RDT-1B","alt":"RDT-1B","abbr":"RDT","aliases":["RDT","Robotics Diffusion Transformer"],"one_liner":"Tsinghua's 2024 diffusion foundation model for bimanual manipulation, with about 1.2 billion parameters and a unified action space.","explanation":"RDT-1B (Robotics Diffusion Transformer) is a robot diffusion foundation model, about 1.2 billion parameters, released by a Tsinghua University team in October 2024. In bimanual manipulation the same situation often has several equally valid ways to act (action multimodality), and direct regression tends to average them into a wrong action; RDT instead uses a diffusion Transformer that denoises step by step to generate an action sequence, which can represent this kind of distribution. To make use of data from many different robots, it designs a “physically interpretable unified action space” that places physical quantities like joint angles and end-effector pose from different robots into fixed positions in one shared vector. The model is first pretrained on more than 1 million trajectories across 46 datasets, then fine-tuned on more than 6,000 ALOHA bimanual demonstrations, after which it can learn a new skill from just 1 to 5 demonstrations. Its successor is RDT2.","example":"Given only 1 to 5 demonstrations, RDT-1B can learn a new bimanual manipulation skill on an ALOHA robot, and it can even execute zero-shot on objects and scenes it has never seen.","related":["Diffusion Policy","Diffusion Transformer","Bimanual Manipulation","Unified Action Space","Action Multimodality","RDT2"]},{"id":"rdt2","category":"named_model","sec":4,"tier":3,"sources":[{"title":"RDT2 project page","url":"https://rdt-robotics.github.io/rdt2/"},{"title":"RDT2 (arXiv 2602.03310)","url":"https://arxiv.org/abs/2602.03310"}],"as_of":"2026-02","related_ids":["rdt-1b","universal-manipulation-interface","handheld-gripper-data-collection","cross-embodiment","flow-matching","vector-quantization"],"name":"RDT2","alt":"RDT2","abbr":"","aliases":["Robotics Diffusion Transformer 2","RDT2-VQ","RDT2-FM","RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization"],"one_liner":"A VLA foundation model from Tsinghua, trained on tens of thousands of hours of UMI data, that deploys zero-shot on new robot arms.","explanation":"RDT2 comes from Jun Zhu's lab at Tsinghua University (the RDT team), open-sourced in September 2025 with the paper following in February 2026; it is the successor to RDT-1B. The problem it addresses: switching to a different robot arm usually forces a full recollection of data and re-fine-tuning. The team improved UMI (Universal Manipulation Interface), a handheld gripper data-collection device, and used it to gather more than 10,000 hours of demonstrations across about 100 locations, mostly real homes; because the same UMI gripper is used for both data collection and deployment, the embodiment gap stays small. The model backbone is a 7-billion-parameter Qwen2.5-VL, trained in three stages: first, residual vector quantization turns actions into discrete tokens (RDT2-VQ); then this is replaced with a roughly 400-million-parameter action expert that outputs continuous actions (RDT2-FM); finally, the model is distilled for speed. The team states it is the first foundation model that can perform simple tasks like pick-and-place zero-shot on an unseen robot embodiment, and it has also been demonstrated playing table tennis.","example":"RDT2-FM is connected directly to a robot arm it never saw during training, fitted with the same UMI-style gripper, and without any fine-tuning it can pick up and place objects following language instructions.","related":["RDT-1B","Universal Manipulation Interface","Handheld Gripper Data Collection","Cross-Embodiment","Flow Matching","Vector Quantization"]},{"id":"robomamba","category":"named_model","sec":4,"tier":3,"sources":[{"title":"arXiv 2406.04339: RoboMamba","url":"https://arxiv.org/abs/2406.04339"},{"title":"RoboMamba 项目主页","url":"https://sites.google.com/view/robomamba-web"}],"as_of":"2024-12","related_ids":["mamba","state-space-model","vision-language-action-model","inference-latency","end-effector-pose","parameter-efficient-fine-tuning"],"name":"RoboMamba","alt":"RoboMamba","abbr":"","aliases":["RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation"],"one_liner":"An efficient VLA that replaces the Transformer language backbone with the Mamba state-space model.","explanation":"RoboMamba was released by Peking University together with Zhiping AI and the Beijing Academy of Artificial Intelligence in June 2024, accepted at NeurIPS 2024. VLAs at the time mostly used Transformer-based large language models as their backbone, which made inference slow and expensive. RoboMamba instead uses Mamba, a selective state-space model whose computation scales linearly with sequence length, as the language model, attached to a vision encoder; it first goes through alignment pretraining and joint training on general and robot-instruction data to gain reasoning ability, then the whole model is frozen and only a very small policy head (about 0.1% of the model's parameters) is added to predict the end effector's SE(3) pose — its position and orientation. The paper reports roughly 3x faster inference than existing VLAs, with competitive pose prediction in both simulation and on a real robot.","example":"RoboMamba can answer a reasoning question like 'which object on the table could be used to hold water,' and also output the position and orientation the gripper should reach after receiving a manipulation instruction.","related":["Mamba","State Space Model","Vision-Language-Action Model","Inference Latency","End-Effector Pose","Parameter-Efficient Fine-Tuning"]},{"id":"tinyvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"TinyVLA (arXiv 2409.12514)","url":"https://arxiv.org/abs/2409.12514"},{"title":"TinyVLA project page","url":"https://tiny-vla.github.io/"}],"as_of":"2025-05","related_ids":["vision-language-action-model","openvla","diffusion-policy","lora","inference-latency","dexvla"],"name":"TinyVLA","alt":"TinyVLA","abbr":"","aliases":["Tiny-VLA","TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation"],"one_liner":"A fast VLA that pairs a small multimodal model with a diffusion policy head and skips robot-data pretraining altogether.","explanation":"TinyVLA was released by Midea Group's AI Lab together with East China Normal University and others (Junjie Wen, Yichen Zhu, and colleagues) in September 2024, published in IEEE RA-L 2025. Models like OpenVLA at the time were slow at inference and also required pretraining on large-scale robot data first. TinyVLA takes a different approach: it first trains a family of small vision-language models, ranging from 70 million to 1.4 billion parameters, using Pythia as the language model and following the LLaVA training recipe, to serve as the policy backbone; when fine-tuning on robot data, the pretrained part is frozen and only about 5% of the parameters are trained via LoRA, with a diffusion policy decoder attached to output continuous actions directly. On real single-arm Franka and dual-arm UR5 robots, the largest variant, TinyVLA-H, beats OpenVLA's success rate by 25.7 points while using 5.5x fewer parameters and roughly 20x lower inference latency. The same team later built DexVLA.","example":"On a dual-arm UR5 task, OpenVLA — which relies on pretraining with single-arm Open X-Embodiment data — struggles, while TinyVLA, which skips robot pretraining and fine-tunes directly, performs better.","related":["Vision-Language-Action Model","OpenVLA","Diffusion Policy","LoRA","Inference Latency","DexVLA"]},{"id":"cogact","category":"named_model","sec":4,"tier":3,"sources":[{"title":"CogACT (arXiv 2411.19650)","url":"https://arxiv.org/abs/2411.19650"},{"title":"CogACT 项目主页","url":"https://cogact.github.io/"}],"as_of":"2024-11","related_ids":["vision-language-action-model","diffusion-action-head","action-expert","openvla","prismatic-vlms","temporal-ensembling"],"name":"CogACT","alt":"CogACT","abbr":"","aliases":["CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation"],"one_liner":"An open-source VLA that attaches a diffusion Transformer action module behind a frozen vision-language backbone.","explanation":"CogACT is a vision-language-action model released in November 2024 by Microsoft Research Asia together with Tsinghua University, the University of Science and Technology of China, and others. At the time, models like OpenVLA had the VLM output discrete action tokens directly, limiting precision and continuity. CogACT separates cognition from action: a VLM such as Prismatic-7B understands the image and instruction and outputs a cognitive feature, and a diffusion Transformer (DiT) action module of up to about 300 million parameters then generates a continuous action sequence conditioned on that feature. At inference, adaptive action ensembling averages only similar action predictions together, avoiding mixing different modes. Trained on about 400,000 trajectories from Open X-Embodiment, it beats OpenVLA by more than 35% in SIMPLER simulation and 55% on real robots, and even outperforms the 55-billion-parameter RT-2-X in simulation.","example":"Researchers tested CogACT on two different real arms, a Realman and a Franka, and it still completed manipulation tasks with objects and backgrounds it had never seen.","related":["Vision-Language-Action Model","Diffusion Action Head","Action Expert","OpenVLA","Prismatic VLMs","Temporal Ensembling"]},{"id":"tracevla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"TraceVLA (arXiv 2412.10345)","url":"https://arxiv.org/abs/2412.10345"},{"title":"TraceVLA project page","url":"https://tracevla.github.io/"}],"as_of":"2025-01","related_ids":["openvla","visual-prompting","tracking-any-point","cotracker","simplerenv","vision-language-action-model"],"name":"TraceVLA","alt":"TraceVLA","abbr":"","aliases":["Visual Trace Prompting","TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies"],"one_liner":"A method that draws the recent motion trace of key points onto the image as a prompt, boosting a VLA's sense of space and time.","explanation":"TraceVLA was released by the University of Maryland and Microsoft Research (Ruijie Zheng, Jianwei Yang, and others) in December 2024, published at ICLR 2025. VLAs usually see only a single current frame and have no sense of how the robot and objects have just been moving. TraceVLA introduces 'visual trace prompting': the point-tracking model CoTracker follows key points in the image over the past several frames, and their trace is drawn directly onto the image, so the model receives both the original frame and the frame with the trace overlaid. The authors used this to fine-tune OpenVLA on 150,000 manipulation trajectories they collected, producing TraceVLA, which beats OpenVLA by about 10 points across 137 configurations in SimplerEnv and reaches 3.5 times OpenVLA's performance on 4 real-robot WidowX tasks. A smaller version based on the 4-billion-parameter Phi-3-Vision comes close to the 7-billion-parameter OpenVLA.","example":"A robot arm is moving a spoon toward a towel; the input image has a colored trace overlaid showing where the gripper and spoon were in the last few steps, letting the model judge the direction and progress of the motion when deciding the next action.","related":["OpenVLA","Visual Prompting","Tracking Any Point","CoTracker","SimplerEnv","Vision-Language-Action Model"]},{"id":"robovlms","category":"named_model","sec":4,"tier":3,"sources":[{"title":"arXiv 2412.14058: RoboVLMs","url":"https://arxiv.org/abs/2412.14058"},{"title":"RoboVLMs 项目主页","url":"https://robovlms.github.io/"}],"as_of":"2026-02","related_ids":["vision-language-action-model","roboflamingo","action-head","cross-embodiment-data","paligemma","ablation-study"],"name":"RoboVLMs","alt":"RoboVLMs","abbr":"","aliases":["Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models","What Matters in Building Vision-Language-Action Models for Generalist Robots"],"one_liner":"A systematic study and open-source framework comparing which VLM backbone, architecture, and training data work best for building a VLA.","explanation":"RoboVLMs was released in December 2024 by Tsinghua University, ByteDance Research, the Institute of Automation at the Chinese Academy of Sciences, Shanghai Jiao Tong University, the National University of Singapore, and others. It is both a systematic experimental study and an open-source framework of the same name that makes it easy to plug a new vision-language model into a VLA. It answers several of the key design choices when building a VLA: which VLM backbone to use, how to organize action output and history information, and when to bring in cross-embodiment data. The authors compared more than 8 VLM backbones and 4 policy architectures across over 600 experiments. The main findings: an independent policy head that outputs continuous actions and takes in multiple frames of history performs best; backbones with thorough vision-language pretraining, such as KosMos and PaliGemma, do noticeably better; and pretraining on cross-embodiment data first helps robustness. The best configuration averaged 4.49 consecutive completed tasks on CALVIN.","example":"To try a new vision-language model as a VLA backbone, one can swap just the backbone inside the RoboVLMs framework while keeping everything else fixed, and compare it directly against the KosMos and PaliGemma versions on CALVIN.","related":["Vision-Language-Action Model","RoboFlamingo","Action Head","Cross-Embodiment Data","PaliGemma","Ablation Study"]},{"id":"uniact","category":"named_model","sec":4,"tier":3,"sources":[{"title":"Universal Actions for Enhanced Embodied Foundation Models (arXiv 2501.10105)","url":"https://arxiv.org/abs/2501.10105"},{"title":"UniAct project page","url":"https://2toinf.github.io/UniAct/"}],"as_of":"2025-03","related_ids":["unified-action-space","cross-embodiment","vector-quantization","embodiment-specific-head","openvla","latent-action"],"name":"UniAct","alt":"UniAct（通用动作空间）","abbr":"UniAct","aliases":["Universal Actions","UniAct: Universal Actions for Enhanced Embodied Foundation Models"],"one_liner":"A method for training an embodied foundation model on a shared, discrete set of 'universal actions' common across different robots.","explanation":"UniAct was released in January 2025 by Tsinghua University's Institute for AI Industry Research (AIR, Xianyuan Zhan's group) together with SenseTime, Peking University, BUPT, and the Shanghai Artificial Intelligence Laboratory, published at CVPR 2025. Different robots have very different action spaces — different numbers of joints, control schemes, and coordinate frames — so training directly on mixed cross-embodiment data causes interference. UniAct instead learns a 'universal action space': a vision-language model maps observations and instructions to a discrete universal action drawn from a vector-quantized codebook, where each code represents an atomic behavior shared across robots; a lightweight, robot-specific decoder head then translates the universal action into that robot's actual control commands. The 0.5B-parameter UniAct matches or beats the 7B OpenVLA across several real-robot and simulation evaluations; adapting to a new robot mainly just requires training a new lightweight decoder head.","example":"After pretraining together on data from several robot arms such as WidowX and Franka, adapting to a new arm only requires training a small decoder head, which then translates the universal actions into that arm's control commands.","related":["Unified Action Space","Cross-Embodiment","Vector Quantization","Embodiment-specific Head","OpenVLA","Latent Action"]},{"id":"hamster","category":"named_model","sec":4,"tier":3,"sources":[{"title":"HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation (arXiv 2502.05485)","url":"https://arxiv.org/abs/2502.05485"},{"title":"HAMSTER 项目页","url":"https://hamster-robot.github.io/"}],"as_of":"2025-05","related_ids":["hierarchical-architecture","intermediate-representation","vision-language-action-model","rvt-2","3d-diffuser-actor","openvla"],"name":"HAMSTER","alt":"HAMSTER（分层动作模型）","abbr":"HAMSTER","aliases":["Hierarchical Action Models for Open-World Robot Manipulation"],"one_liner":"A 2025 hierarchical VLA from NVIDIA and others where a high level sketches a 2D path and a low-level 3D policy follows it.","explanation":"HAMSTER was proposed in February 2025 by researchers at NVIDIA, the University of Washington, and USC, published at ICLR 2025. Ordinary VLA models fine-tune a vision-language model (VLM) directly to output actions, which can only be trained on expensive real-robot data. HAMSTER splits the system into two layers: the high level is a fine-tuned VLM (built on VILA) that looks at an RGB image and a task description and sketches a rough 2D path showing roughly how the end effector should move; the low level is a control policy that handles 3D input (the paper tries RVT-2 and 3D Diffuser Actor), using that path as guidance for precise manipulation. Because the high level only has to output a 2D path, it can be trained on cheap “out-of-domain” data — action-label-free video, hand-drawn sketches, simulation data. The high level doesn't handle fine motor control and the low level doesn't handle task reasoning; each does what it's good at.","example":"Tested on a real robot across 7 generalization axes (such as new objects, new backgrounds, and new instruction semantics), HAMSTER beats OpenVLA by about 20 percentage points in average success rate, a roughly 50% relative improvement.","related":["Hierarchical Architecture","Intermediate Representation","Vision-Language-Action Model","RVT-2","3D Diffuser Actor","OpenVLA"]},{"id":"dexvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"DexVLA (arXiv:2502.05855)","url":"https://arxiv.org/abs/2502.05855"},{"title":"DexVLA 论文 HTML 全文 v3","url":"https://arxiv.org/html/2502.05855v3"}],"as_of":"2025-08","related_ids":["vision-language-action-model","action-expert","diffusion-action-head","cross-embodiment","pi0","qwen-vl"],"name":"DexVLA","alt":"DexVLA","abbr":"","aliases":["DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control"],"one_liner":"A VLA from Midea and others that attaches a roughly billion-parameter diffusion action expert to a vision-language model, adapting to many robots.","explanation":"DexVLA was released in February 2025 by researchers at Midea Group, East China Normal University, and Shanghai University, published at CoRL 2025. It pairs the vision-language model Qwen2-VL (2 billion parameters) with a pluggable diffusion action expert: the VLM handles looking at images and understanding instructions, while the action expert, built on the ScaleDP architecture and scaled up to about 1 billion parameters, uses multiple output heads to adapt to different robots and is dedicated to generating continuous actions. Training follows an “embodiment curriculum” with three stages: first, the action expert is pretrained on its own using about 100 hours of cross-embodiment data; then it's attached to the VLM and aligned to a specific embodiment; finally, it's post-trained on data annotated with substeps for a specific task. It aims to fix earlier VLAs' weak action representation, expensive training, and difficulty switching robots — the same “VLM plus action expert” approach as π0.","example":"The paper tests on a single-arm Franka (gripper or dexterous hand), a bimanual UR5e, and a bimanual AgileX robot; with no task-specific adaptation it folds a shirt (score 0.92), and on the full laundry-folding task it scores 0.4, versus 0.2 for π0 under the same conditions.","related":["Vision-Language-Action Model","Action Expert","Diffusion Action Head","Cross-Embodiment","π0","Qwen-VL"]},{"id":"magma","category":"named_model","sec":4,"tier":3,"sources":[{"title":"arXiv 2502.13130: Magma","url":"https://arxiv.org/abs/2502.13130"},{"title":"GitHub: microsoft/Magma","url":"https://github.com/microsoft/Magma"}],"as_of":"2025-02","related_ids":["visual-prompting-2","vision-language-action-model","multimodal-large-language-model","open-x-embodiment","llama","microsoft-research"],"name":"Magma (Microsoft)","alt":"Magma","abbr":"","aliases":["Magma-8B","A Foundation Model for Multimodal AI Agents"],"one_liner":"A multimodal agent foundation model from Microsoft that can both operate software interfaces and control a robot arm.","explanation":"Magma is a multimodal foundation model released by Microsoft Research in February 2025, published at CVPR 2025, with the open-source version Magma-8B (language backbone Llama-3-8B). It aims to have one model handle agent tasks in both the digital and physical world: clicking through web pages and phone interfaces (UI navigation), and controlling a robot arm to pick and place objects. The key is two kinds of annotation: Set-of-Mark (SoM) numbers the actionable elements in an image, letting the model express where to act by “choosing which mark”; Trace-of-Mark (ToM) labels the future motion trajectory of marked points in a video, teaching the model to predict what happens next. With ToM, it can learn spatiotemporal planning from large amounts of instructional video that has no action labels at all. Training data includes UI-navigation data, Open X-Embodiment robot data, and web video.","example":"The same Magma-8B model can output “click the button marked 3” on a phone screenshot, and also output a robot arm's movement trajectory and gripper action on a robot camera feed.","related":["Visual Prompting (Set-of-Mark)","Vision-Language-Action Model","Multimodal Large Language Model","Open X-Embodiment","Llama","Microsoft Research"]},{"id":"chatvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"ChatVLA (arXiv 2502.14420)","url":"https://arxiv.org/abs/2502.14420"},{"title":"ChatVLA 项目主页","url":"https://chatvla.github.io/"}],"as_of":"2025-11","related_ids":["vision-language-action-model","catastrophic-forgetting","mixture-of-experts","co-training","dexvla","knowledge-insulation"],"name":"ChatVLA","alt":"ChatVLA","abbr":"","aliases":["ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model"],"one_liner":"A unified VLA model that can both answer questions about images and directly control a robot.","explanation":"ChatVLA was proposed in February 2025 by teams at Midea Group and East China Normal University, and was selected for an oral presentation at the EMNLP 2025 main conference. It targets a common problem in VLAs: fine-tuning on robot data tends to wipe out the base VLM's original image question-answering ability (the paper calls this spurious forgetting), while training control and understanding data together makes the two interfere with each other. Its fix has two parts. First, staged alignment training: the model is trained on robot data alone until it can control the robot, then image-text data is mixed back in at a 1:3 ratio to restore understanding. Second, a mixture-of-experts structure: attention layers are shared, but feed-forward layers split into an understanding expert and a control expert, routed by task. Built on Qwen2-VL-2B, it scores 47.2 on MMStar, with multimodal understanding clearly stronger than earlier VLAs, and it also beats OpenVLA and ECoT across 25 real-robot tasks.","example":"The same ChatVLA model can both answer questions about a photo and carry out instructions like grasp, place, push, or hang in scenes such as a bathroom, kitchen, or tabletop.","related":["Vision-Language-Action Model","Catastrophic Forgetting","Mixture of Experts","Co-training","DexVLA","Knowledge Insulation"]},{"id":"dexgraspvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"DexGraspVLA (arXiv:2502.20900)","url":"https://arxiv.org/abs/2502.20900"},{"title":"DexGraspVLA 项目主页","url":"https://dexgraspvla.github.io/"}],"as_of":"2025-11","related_ids":["vision-language-action-model","dexterous-manipulation","hierarchical-architecture","diffusion-policy","dinov2","psibot"],"name":"DexGraspVLA","alt":"DexGraspVLA","abbr":"","aliases":["DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping"],"one_liner":"A hierarchical dexterous-hand grasping framework from Peking University and PsiBot: a large model plans, a diffusion policy acts.","explanation":"DexGraspVLA was released in February 2025 by teams including Peking University's Institute for Artificial Intelligence and the Peking University–PsiBot (灵初智能) joint lab, later accepted as an oral presentation at AAAI 2026. It has two layers. The upper layer uses an off-the-shelf vision-language model (such as Qwen-VL) to understand the instruction and draw a box around the target object in the image. The lower layer is a diffusion controller: SAM and Cutie continuously track the target's mask, a frozen DINOv2 extracts visual features, and a diffusion Transformer outputs arm and dexterous-hand actions. The core idea is to use foundation models first to turn highly variable images and language into a more stable representation, narrowing the gap between training and test scenes, so that imitation learning from only a small number of human demonstrations can still generalize to a huge number of new scenes.","example":"After training on 2,094 cluttered-scene grasping demonstrations collected across 36 household objects, it was tested zero-shot across roughly 1,287 scenes combining 360 new objects, 6 new backgrounds, and 3 new lighting conditions, reaching an overall success rate of 90.8% on a Realman 7-DOF arm paired with PsiBot's 6-DOF G0-R dexterous hand.","related":["Vision-Language-Action Model","Dexterous Manipulation","Hierarchical Architecture","Diffusion Policy","DINOv2","PsiBot"]},{"id":"hybridvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"HybridVLA (arXiv 2503.10631)","url":"https://arxiv.org/abs/2503.10631"},{"title":"HybridVLA project page","url":"https://hybrid-vla.github.io/"},{"title":"PKU-HMI-Lab/Hybrid-VLA (GitHub)","url":"https://github.com/PKU-HMI-Lab/Hybrid-VLA"}],"as_of":"2025-06","related_ids":["vision-language-action-model","hybrid-autoregressive-diffusion-architecture","diffusion-action-head","action-binning","cogact","openvla"],"name":"HybridVLA","alt":"HybridVLA","abbr":"","aliases":["Hybrid-VLA","Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model"],"one_liner":"A VLA that predicts actions both by diffusion and autoregressively, inside the same large language model.","explanation":"HybridVLA was released in March 2025 by Shanghang Zhang's group at Peking University, together with the Beijing Academy of Artificial Intelligence, CUHK, and Fudan University. Autoregressive VLA models discretize actions into tokens, which lets them inherit a vision-language model's reasoning ability, but discretization breaks the continuity of actions and hurts fine motor control; diffusion VLA models instead attach a separate diffusion head to output continuous actions, but only use the features a VLM extracts, without benefiting from token-by-token generative reasoning. HybridVLA embeds diffusion denoising directly into a large language model's next-token-prediction process, so one model produces both a diffusion action and an autoregressive action at once, which are then adaptively fused through “collaborative action ensembling.” The model is built on Prismatic 7B (a LLaMA-2 backbone); the paper reports average success rates 14% and 19% higher than the previous best methods on simulation and real-robot tasks, respectively.","example":"At every step, the same HybridVLA model produces two action predictions, one diffusion-based and one autoregressive, which are fused before being handed to the robot arm to execute.","related":["Vision-Language-Action Model","Hybrid Autoregressive-Diffusion Architecture","Diffusion Action Head","Action Binning","CogACT","OpenVLA"]},{"id":"dita","category":"named_model","sec":4,"tier":3,"sources":[{"title":"Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy (arXiv 2503.19757)","url":"https://arxiv.org/abs/2503.19757"},{"title":"Dita 项目主页","url":"https://robodita.github.io/"}],"as_of":"2025-09","related_ids":["diffusion-transformer","vision-language-action-model","diffusion-action-head","open-x-embodiment","causal-attention","libero-benchmark"],"name":"Dita","alt":"Dita","abbr":"","aliases":["RoboDita","Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy"],"one_liner":"An open-source, roughly 330-million-parameter generalist VLA policy that denoises a whole action sequence directly with a single Transformer.","explanation":"Dita is a generalist robot policy released in March 2025 by researchers from Shanghai AI Lab, Zhejiang University, SenseTime, CUHK, Peking University, Tsinghua University, and other institutions, published at ICCV 2025. Many diffusion-based VLA models attach a small diffusion action head — a small network dedicated to turning noise into actions — after the main model, so the action generation only ever sees compressed features. Dita instead concatenates language tokens, visual tokens from historical images, and noisy action tokens into a single causal Transformer and denoises them together, a form of in-context conditioning that lets the action align directly with raw visual detail and better capture action deltas and differences across environments. The model is pretrained on cross-embodiment data from Open X-Embodiment, with about 334 million total parameters (about 221 million trainable), and matches or approaches the best contemporary results on simulation benchmarks including SimplerEnv, LIBERO, CALVIN, and ManiSkill2. The code is open-source, and the project positions itself as a lightweight, easily reproducible diffusion VLA baseline.","example":"When moved to a new physical robot and a new scene, Dita can be fine-tuned with just 10 demonstrations per task from a single third-person camera, and still complete multi-step, long-horizon tasks while handling changes in background, clutter placement, and lighting.","related":["Diffusion Transformer","Vision-Language-Action Model","Diffusion Action Head","Open X-Embodiment","Causal Attention","LIBERO Benchmark"]},{"id":"cot-vla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"CoT-VLA (arXiv:2503.22020, CVPR 2025)","url":"https://arxiv.org/abs/2503.22020"},{"title":"CoT-VLA 项目主页","url":"https://cot-vla.github.io/"}],"as_of":"2025-03","related_ids":["visual-chain-of-thought","vision-language-action-model","chain-of-thought","action-chunking","action-free-video","susie"],"name":"CoT-VLA","alt":"CoT-VLA","abbr":"","aliases":["CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models","Visual Chain-of-Thought VLA"],"one_liner":"A 7-billion-parameter VLA that first generates a future subgoal image as its “thought,” then outputs an action to reach it.","explanation":"CoT-VLA was proposed in March 2025 by NVIDIA, Stanford, MIT, and others, published at CVPR 2025. Earlier VLAs mapped images and instructions directly to actions with no reasoning step in between. CoT-VLA replaces chain-of-thought with a visual form: the model first autoregressively generates a subgoal frame several steps into the future — what the scene should look like once part of the task is done — then generates an action chunk targeting that image, executing it in closed loop. Its backbone is VILA-U, a multimodal model that can both understand and generate images, at 7B parameters total; image generation uses causal attention, while action decoding uses full attention. Because predicting a subgoal image needs no action labels, action-label-free video like EPIC-KITCHENS can be used for training too. The paper reports beating the strongest VLA baseline of the time by 17% on real-robot tasks and 6% on simulation benchmarks.","example":"Given “put the bowl in the drawer,” CoT-VLA first draws what the scene should look like a few steps ahead — the bowl already picked up and near the drawer — then outputs a chunk of arm actions to reach that image; after executing, it looks at the new image and repeats.","related":["Visual Chain-of-Thought","Vision-Language-Action Model","Chain-of-Thought","Action Chunking","Action-free Video","SuSIE"]},{"id":"smolvla","category":"named_model","sec":4,"tier":2,"sources":[{"title":"SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data (Hugging Face Blog)","url":"https://huggingface.co/blog/smolvla"}],"as_of":"2025-06","related_ids":["lerobot","vision-language-action-model","action-expert","flow-matching","asynchronous-inference","so-100-so-101-arm"],"name":"SmolVLA","alt":"SmolVLA","abbr":"","aliases":["SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics"],"one_liner":"Hugging Face's open-source, roughly 450-million-parameter small VLA that trains and runs on ordinary consumer hardware, even a laptop.","explanation":"SmolVLA is an open-source vision-language-action model, about 450 million parameters, released in June 2025 by Hugging Face's LeRobot team. It's built on the small vision-language model SmolVLM2, followed by an action expert of about 100 million parameters trained with flow matching; to save compute, it keeps only 64 visual tokens per frame and skips some network layers. Pretraining uses only 487 community-uploaded LeRobot datasets, about 10 million frames — an order of magnitude less data than typical VLAs. It also supports asynchronous inference: the robot starts computing the next action chunk while still executing the current one, which the team says speeds up task completion by about 30%. On LIBERO, Meta-World, and real SO-100/SO-101 robots, it performs close to or better than larger models, and it can run on consumer GPUs or even a MacBook.","example":"Collect a batch of demonstrations with LeRobot of an SO-101 arm picking up blocks, fine-tune SmolVLA on a single consumer GPU, then deploy it back onto the same arm.","related":["LeRobot","Vision-Language-Action Model","Action Expert","Flow Matching","Asynchronous Inference","SO-100 / SO-101 Arm"]},{"id":"beast","category":"named_model","sec":4,"tier":3,"sources":[{"title":"BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning (arXiv 2506.06072)","url":"https://arxiv.org/abs/2506.06072"}],"as_of":"2025-10","related_ids":["action-tokenizer","action-chunking","pi0-fast","parallel-decoding","florence-2","action-chunking-with-transformers"],"name":"BEAST","alt":"BEAST","abbr":"BEAST","aliases":["B-spline Encoded Action Sequence Tokenizer","BEAST-F","BEAST-ACT"],"one_liner":"An action tokenizer that compresses a chunk of actions into a fixed number of tokens using B-spline control points.","explanation":"BEAST is an action tokenizer proposed in June 2025 by a team at the Karlsruhe Institute of Technology (KIT) in Germany, published at NeurIPS 2025. The idea is to fit an action sequence with a B-spline curve (a smooth curve whose shape is set by a small number of “control points”), then treat the control points as tokens: they can be quantized into discrete tokens for a language model, or kept as continuous values. Compared with tokenizers like FAST, which need separate training and produce variable-length output, BEAST needs no training and always produces a fixed number of tokens per chunk, so decoding can happen in parallel in a single forward pass; the B-spline also guarantees smooth transitions between the start and end of adjacent action chunks, reducing jitter. The paper plugs it into small models like Florence-2 (BEAST-F) and into architectures like ACT, gets competitive results on CALVIN and LIBERO, and reaches roughly twice π0's inference throughput.","example":"BEAST-ACT turns the 100-step action sequence ACT would normally predict step by step into just 15 control points, which a B-spline then reconstructs into a smooth trajectory.","related":["Action Tokenizer","Action Chunking","π0-FAST","Parallel Decoding","Florence-2","Action Chunking with Transformers"]},{"id":"bridgevla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"BridgeVLA (arXiv 2506.07961)","url":"https://arxiv.org/abs/2506.07961"},{"title":"BridgeVLA GitHub","url":"https://github.com/BridgeVLA/BridgeVLA"}],"as_of":"2026-08","related_ids":["3d-vla","paligemma","rvt-2","rlbench","the-colosseum-a-benchmark-for-evaluating-generalization-for","keyframe-action-prediction"],"name":"BridgeVLA","alt":"BridgeVLA","abbr":"","aliases":["BridgeVLA++","BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models"],"one_liner":"A 3D-manipulation VLA that projects point clouds into 2D images and reads off actions as heatmaps on them.","explanation":"BridgeVLA was proposed in June 2025 by teams at the Institute of Automation, Chinese Academy of Sciences, and ByteDance Seed, published at NeurIPS 2025. The problem: VLMs are pretrained on 2D images and text, while 3D manipulation needs to handle point clouds and output 3D poses — the two are misaligned, so knowledge transfers poorly between them. BridgeVLA renders a point cloud into several multi-view 2D images and feeds them into PaliGemma, having the model output a 2D heatmap on each image (each pixel representing the likelihood of “act here”), then combines the multi-view heatmaps to determine the 3D position the end effector should go to; before formal training, it's also pretrained to output heatmaps using object-detection data. It raised success rate on RLBench from 81.4% to 88.2%. In August 2026 the team also open-sourced BridgeVLA++.","example":"Tested on a real Franka Research 3 arm across more than 10 tasks, given only 3 demonstrations per task, it reached an average success rate of 96.8%.","related":["3D VLA","PaliGemma","RVT-2","RLBench","The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation","Keyframe Action Prediction"]},{"id":"dreamvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"DreamVLA (arXiv 2507.04447)","url":"https://arxiv.org/abs/2507.04447"},{"title":"DreamVLA 项目主页","url":"https://zhangwenyao1.github.io/DreamVLA/"},{"title":"DreamVLA 代码仓库（GitHub）","url":"https://github.com/Zhangwenyao1/DreamVLA"}],"as_of":"2025-09","related_ids":["vision-language-action-model","world-model","inverse-dynamics-model","diffusion-transformer","attention-mask","calvin-benchmark"],"name":"DreamVLA","alt":"DreamVLA","abbr":"","aliases":["A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge"],"one_liner":"A VLA model that first predicts future dynamic regions, depth, and semantic features, then generates actions conditioned on them.","explanation":"DreamVLA is a vision-language-action (VLA) model proposed in July 2025 by researchers from Shanghai Jiao Tong University, the Eastern Institute of Technology in Ningbo, Tsinghua University, Peking University, Galbot, and other institutions, published at NeurIPS 2025. Some VLA models first “imagine” a complete next frame before producing an action, but most pixels in a full image are irrelevant to the task. DreamVLA instead predicts only three compact kinds of “world knowledge”: where things will move (dynamic regions), monocular depth, and high-level semantics (using features from DINOv2 and SAM); conditioned on these, a diffusion Transformer then generates a segment of future actions — effectively predicting the outcome first and working backward to figure out what to do, an inverse-dynamics-style approach. To keep the three kinds of information from interfering with each other inside attention, it separates them with a block-structured attention mask. It is a representative example of the “VLA combined with world-model-style prediction” direction.","example":"On the CALVIN ABC-D long-horizon benchmark, it completes an average of 4.44 consecutive sub-tasks, and reaches a 76.7% success rate on real-robot manipulation tasks.","related":["Vision-Language-Action Model","World Model","Inverse Dynamics Model","Diffusion Transformer","Attention Mask","CALVIN Benchmark"]},{"id":"thinkact","category":"named_model","sec":4,"tier":3,"sources":[{"title":"ThinkAct (arXiv 2507.16815)","url":"https://arxiv.org/abs/2507.16815"},{"title":"ThinkAct project page","url":"https://jasper0314-huang.github.io/thinkact-vla/"},{"title":"Fast-ThinkAct (arXiv 2601.09708)","url":"https://arxiv.org/abs/2601.09708"}],"as_of":"2026-01","related_ids":["embodied-reasoning","dual-system-architecture","chain-of-thought","group-relative-policy-optimization","latent-reasoning","libero-benchmark"],"name":"ThinkAct","alt":"ThinkAct","abbr":"","aliases":["Fast-ThinkAct (follow-up version)","ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning"],"one_liner":"A reasoning-based VLA that first has a multimodal large model work out a visual plan, then hands it to an action model to execute.","explanation":"ThinkAct was released by NVIDIA and National Taiwan University (Chi-Pin Huang, Fu-En Yang, and others) in July 2025, published at NeurIPS 2025. Most VLAs map directly from an image and instruction to an action with no explicit reasoning step, which makes multi-step, long-horizon tasks difficult. ThinkAct uses a dual-system design: the 'thinking' part is a multimodal large model initialized from Qwen2.5-VL 7B, first given a supervised fine-tuning cold start, then trained with GRPO reinforcement learning to generate embodied reasoning plans, with rewards coming from visual signals — whether the endpoint of the planned trajectory matches the demonstration, and whether the whole trajectory is consistent with it. The reasoning output is compressed into a visual planning latent, which conditions a downstream action model that produces the specific actions. On manipulation benchmarks such as SimplerEnv and LIBERO, and on several embodied-reasoning benchmarks, it shows few-shot adaptation, long-horizon planning, and self-correction. A follow-up, Fast-ThinkAct, was published at CVPR 2026.","example":"Given 'put the carrot on the plate,' ThinkAct first reasons that it needs to grasp the carrot and then move it above the plate, and plans the gripper's motion trajectory, which the action model then executes; if the grasp fails partway through, it can replan.","related":["Embodied Reasoning","Dual-System Architecture (System 1 / System 2)","Chain-of-Thought","Group Relative Policy Optimization","Latent Reasoning","LIBERO Benchmark"]},{"id":"molmoact","category":"named_model","sec":4,"tier":3,"sources":[{"title":"MolmoAct (Ai2 blog)","url":"https://allenai.org/blog/molmoact"},{"title":"MolmoAct: Action Reasoning Models that can Reason in Space (arXiv 2508.07917)","url":"https://arxiv.org/abs/2508.07917"},{"title":"MolmoAct 2 (Ai2 blog)","url":"https://allenai.org/blog/molmoact2"}],"as_of":"2026-05","related_ids":["vision-language-action-model","molmo","embodied-reasoning","action-chain-of-thought","flow-matching","molmoact2-bimanualyam"],"name":"MolmoAct","alt":"MolmoAct","abbr":"","aliases":["MolmoAct2","Action Reasoning Model (ARM)","Action Reasoning Models that can Reason in Space"],"one_liner":"Ai2's open-source “action reasoning model,” a VLA that reasons about depth and sketches a trajectory before outputting an action.","explanation":"MolmoAct was released by Ai2 in August 2025, built on the company's own Molmo vision-language model, with 7 billion parameters; the model, code, data, and evaluation scripts are all open-source. It splits VLA decision-making into three steps: first it outputs perceptual tokens carrying depth information to understand 3D space, then it draws a waypoint trajectory for the end effector directly on the image, and finally it decodes that trajectory into specific arm and gripper actions. Because the planned trajectory is overlaid directly on the image, a person can see what it intends to do before execution, and can also draw a path on a phone or tablet to guide it. At release, it reported a 72.1% success rate on SimplerEnv out-of-distribution tasks and 86.6% on LIBERO. MolmoAct2, from May 2026, switched to the embodied-reasoning backbone Molmo2-ER and attached a flow-matching action expert, cutting single-action-inference time from about 6.7 seconds down to 180 milliseconds, and released a companion dataset of more than 720 hours of bimanual YAM data.","example":"Given the instruction “put the pillow on the sofa,” MolmoAct first draws a trajectory from pillow to sofa on the camera view; after the user confirms it or redraws it on a tablet, it converts that into robot-arm actions to execute.","related":["Vision-Language-Action Model","Molmo (Ai2)","Embodied Reasoning","Action Chain-of-Thought","Flow Matching","MolmoAct2-BimanualYAM"]},{"id":"memoryvla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"arXiv 2508.19236: MemoryVLA","url":"https://arxiv.org/abs/2508.19236"},{"title":"MemoryVLA 项目主页","url":"https://shihao1895.github.io/MemoryVLA/"}],"as_of":"2026-01","related_ids":["memory-augmented-vla","embodied-memory","vision-language-action-model","long-horizon-task","diffusion-action-head","dexmal"],"name":"MemoryVLA","alt":"MemoryVLA","abbr":"","aliases":["Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation"],"one_liner":"Adds working memory and a memory bank to a VLA, so the robot remembers what it has already seen and done.","explanation":"MemoryVLA is work from Gao Huang's group at Tsinghua University, in collaboration with Dexmal (原力灵机), Megvii, and others, released in August 2025 and published at ICLR 2026. Most VLA models decide on an action by looking only at the current frame, but for many manipulation tasks a single current frame isn't enough — the footage right before and right after pressing a button can look nearly identical, and without any memory of history there's no way to tell whether it's already been pressed. MemoryVLA borrows the cognitive-science concepts of working memory and episodic memory: a 7B vision-language model encodes observations into perceptual tokens and cognitive tokens, serving as working memory; a separate memory bank stores past low-level detail and high-level semantics, retrieved as needed and fused with current information; a diffusion action expert then outputs a segment of action. The paper reports a 71.9% success rate on SimplerEnv-Bridge, 96.5% on LIBERO, and 84.0% on real-robot tasks, with especially large gains on long-horizon tasks.","example":"For a task like “press three buttons in sequence,” the footage barely changes after each press, so only the memory bank remembering which ones are already pressed can prevent the robot from pressing the same one twice or skipping one.","related":["Memory-Augmented VLA","Embodied Memory","Vision-Language-Action Model","Long-horizon Task","Diffusion Action Head","Dexmal"]},{"id":"discrete-diffusion-vla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"Discrete Diffusion VLA (arXiv:2508.20072)","url":"https://arxiv.org/abs/2508.20072"},{"title":"Discrete Diffusion VLA 论文 HTML 全文 v4","url":"https://arxiv.org/html/2508.20072v4"}],"as_of":"2026-05","related_ids":["discrete-diffusion","vision-language-action-model","openvla","action-binning","parallel-decoding","autoregressive-decoding"],"name":"Discrete Diffusion VLA","alt":"Discrete Diffusion VLA（离散扩散 VLA）","abbr":"","aliases":["Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies"],"one_liner":"Decodes VLA actions with discrete diffusion: action tokens are generated in parallel, confident ones fixed first, the rest filled in over several rounds.","explanation":"Discrete Diffusion VLA was released in August 2025 by researchers at the University of Hong Kong, Shanghai Jiao Tong University, and other institutions (corresponding authors Yao Mu and Ping Luo), and has been accepted at ICML 2026. Discrete VLAs like OpenVLA bin actions into tokens and generate them one at a time, autoregressively, which is slow and can't revise early mistakes; continuous diffusion-head VLAs like π0 need a separate action module bolted on. This model instead changes the action-chunk tokens in OpenVLA (built on Prismatic-7B, a Llama 2 backbone) to use bidirectional attention, and decodes them with discrete diffusion (masked prediction): all tokens start masked, each round predicts in parallel, high-confidence tokens are locked in first, the rest keep iterating, and already-filled tokens that remain uncertain can be re-masked and corrected. Action decoding shares a single Transformer and a cross-entropy objective with the language model, which also better preserves the original VLM's abilities.","example":"On simulation benchmarks it reaches 96.4% average success on LIBERO, 71.2% on SimplerEnv-Fractal visual matching, and 54.2% on SimplerEnv-Bridge; the authors also validated it on a real AgileX Cobot Magic platform.","related":["Discrete Diffusion","Vision-Language-Action Model","OpenVLA","Action Binning","Parallel Decoding","Autoregressive Decoding"]},{"id":"vla-adapter","category":"named_model","sec":4,"tier":3,"sources":[{"title":"VLA-Adapter (arXiv 2509.09372)","url":"https://arxiv.org/abs/2509.09372"},{"title":"VLA-Adapter 项目主页","url":"https://vla-adapter.github.io/"},{"title":"OpenHelix-Team/VLA-Adapter (GitHub)","url":"https://github.com/OpenHelix-Team/VLA-Adapter"}],"as_of":"2025-09","related_ids":["vision-language-action-model","learnable-query","adapter","openvla-oft","libero-benchmark","vla-rft"],"name":"VLA-Adapter","alt":"VLA-Adapter","abbr":"","aliases":["VLA-Adapter-Pro","VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model"],"one_liner":"A VLA that reaches top-tier performance without any robot pretraining, using a 0.5B small model plus a lightweight policy module.","explanation":"VLA-Adapter was released in September 2025 by Beijing University of Posts and Telecommunications, Westlake University, Zhejiang University, HKUST (Guangzhou), the OpenHelix team, and others, with code and weights open-sourced. Typical VLAs rely on large vision-language models (VLMs) and pretraining on massive robot datasets, which is expensive. The authors systematically compared which layers and features of a VLM are best suited to condition action generation, and designed a roughly 97-million-parameter policy module accordingly: Bridge Attention injects raw vision-language features from various VLM layers, along with features from a set of learnable queries (ActionQuery), into the action space, with the amount injected controlled by learnable parameters. Using only Qwen2.5-0.5B as the backbone and no robot-data pretraining, it reaches 97.3% average success on LIBERO (98.5% for the Pro version) and an average completed length of 4.50 on CALVIN ABC→D (Pro version). This sharply lowers the barrier to training and deploying a VLA, and later work such as VLA-RFT builds on it.","example":"The paper reports a usable model can be trained on a single consumer GPU in about 8 hours; inference throughput reaches 219.2 Hz, compared with 71.4 Hz for OpenVLA-OFT under the same conditions.","related":["Vision-Language-Action Model","Learnable Query","Adapter","OpenVLA-OFT","LIBERO Benchmark","VLA-RFT"]},{"id":"x-vla","category":"named_model","sec":4,"tier":3,"sources":[{"title":"X-VLA (arXiv:2510.10274)","url":"https://arxiv.org/abs/2510.10274"},{"title":"2toinf/X-VLA (GitHub)","url":"https://github.com/2toinf/X-VLA"}],"as_of":"2026-09","related_ids":["cross-embodiment","prompt-tuning-soft-prompt","flow-matching","vision-language-action-model","lerobot","cross-embodiment-data"],"name":"X-VLA","alt":"X-VLA","abbr":"","aliases":["X-VLA-0.9B","X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"],"one_liner":"A 0.9B cross-embodiment VLA that unifies many robots by giving each one its own learnable 'soft prompt.'","explanation":"X-VLA was proposed in October 2025 by Tsinghua University's Institute for AI Industry Research (AIR) together with the Shanghai Artificial Intelligence Laboratory and others, and has been accepted at ICLR 2026. The difficulty in cross-embodiment training is that different robots have very different cameras and action spaces, so training on them mixed together causes interference. X-VLA instead gives each data source its own set of learnable embedding vectors as a 'soft prompt' that tells the model which hardware it's currently dealing with, while sharing the rest of the backbone entirely; actions are generated with flow matching, and the backbone is a standard Transformer encoder. The 0.9-billion-parameter version is pretrained on 290,000 trajectories combined from DROID, RoboMIND, and AgiBot, and leads on 6 simulation benchmarks — including LIBERO, SimplerEnv, and CALVIN — and 3 real-robot platforms. The code is open-sourced under Apache-2.0 and has been integrated into LeRobot.","example":"On a real-robot clothes-folding task (Soft-Fold), X-VLA-0.9B is reported to reach 100% success, folding 33 garments per hour.","related":["Cross-Embodiment","Prompt Tuning / Soft Prompt","Flow Matching","Vision-Language-Action Model","LeRobot","Cross-Embodiment Data"]},{"id":"vla-0","category":"named_model","sec":4,"tier":3,"sources":[{"title":"VLA-0: Building State-of-the-Art VLAs with Zero Modification (arXiv 2510.13054)","url":"https://arxiv.org/abs/2510.13054"},{"title":"VLA-0 项目主页","url":"https://vla0.github.io/"}],"as_of":"2025-10","related_ids":["vision-language-action-model","action-tokenizer","action-binning","qwen-vl","libero-benchmark","smolvla"],"name":"VLA-0","alt":"VLA-0","abbr":"VLA-0","aliases":["VLA-0 (Action as Text)","VLA-0: Building State-of-the-Art VLAs with Zero Modification"],"one_liner":"NVIDIA's minimalist VLA that makes no architecture changes at all, having the VLM write out actions directly as plain digit text.","explanation":"VLA-0 is work by Ankit Goyal and colleagues at NVIDIA, released in October 2025. Typical VLAs (vision-language-action models) either add special action tokens to the vision-language model's vocabulary or attach a separate action head; VLA-0 changes nothing at all, having Qwen2.5-VL-3B output actions in plain text: each continuous action dimension is normalized into an integer between 0 and 1000, written out as digits, and decoded back afterward. The paper points out that making this work well needs a couple of tricks: during training, characters in the target digit string are randomly masked to force the model to actually look at the image rather than just continuing the text; at inference, predictions from adjacent timesteps are ensembled and averaged. The result is 94.7% average success on LIBERO, beating OpenVLA-OFT and SmolVLA when trained only on that benchmark's own data, and also beating models pretrained on large-scale robot data such as π0 and GR00T N1; on a real SO-100 robot it beats SmolVLA by 12.5 points. It is a reminder to get the simple baseline solid before reaching for something fancier.","example":"Given a camera image and an instruction, the model answers like a question, outputting a string of integers between 0 and 1000 as text, with each number corresponding to one action dimension; these are un-normalized and handed to the robot arm for execution.","related":["Vision-Language-Action Model","Action Tokenizer","Action Binning","Qwen-VL","LIBERO Benchmark","SmolVLA"]},{"id":"pi0","category":"named_model","sec":5,"tier":1,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)","url":"https://arxiv.org/abs/2410.24164"},{"title":"Physical-Intelligence/openpi GitHub 仓库","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"2025-09","related_ids":["flow-matching","action-expert","paligemma","pi0-5","pi0-fast","physical-intelligence"],"name":"π0","alt":"π0","abbr":"π0","aliases":["pi-zero","pi0","π-zero","π0: A Vision-Language-Action Flow Model for General Robot Control"],"one_liner":"Physical Intelligence's 2024 VLA that generates continuous robot actions with flow matching; code and weights are open-sourced.","explanation":"π0 is a general-purpose robot foundation model released on October 31, 2024, by the US embodied-AI company Physical Intelligence (PI). It's built on Google's 3-billion-parameter PaliGemma vision-language model, plus an “action expert” of about 300 million parameters that uses flow matching — a method similar to diffusion models that generates continuous values by gradually denoising — to produce the next 50 steps of action in one shot, controlling at up to 50Hz. Unlike RT-2 or OpenVLA, which discretize actions into tokens, this continuous generation suits high-frequency, dexterous tasks like folding laundry better. It was pretrained on more than 10,000 hours of robot data spanning 7 robot embodiments and 68 tasks, mixed with open datasets like OXE and DROID; after pretraining it can follow language instructions directly, and can also be fine-tuned on a small amount of data to learn new skills. PI later open-sourced the code and weights in the openpi repository, and it can run inference and LoRA fine-tuning on a single RTX 4090.","example":"In the paper's demos, a fine-tuned π0 takes clothes out of a dryer, carries them to a table, and folds each one; it can also assemble cardboard boxes and clear a table.","related":["Flow Matching","Action Expert","PaliGemma","π0.5","π0-FAST","Physical Intelligence"]},{"id":"pi0-fast","category":"named_model","sec":5,"tier":2,"sources":[{"title":"FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)","url":"https://arxiv.org/abs/2501.09747"},{"title":"FAST: Efficient Robot Action Tokenization (Physical Intelligence)","url":"https://www.pi.website/research/fast"},{"title":"openpi (GitHub, Physical Intelligence)","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"2025-01","related_ids":["pi0","action-tokenizer","discrete-cosine-transform","byte-pair-encoding","action-binning","autoregressive-decoding"],"name":"π0-FAST","alt":"π0-FAST","abbr":"","aliases":["pi0-FAST","pi0_fast","π0 + FAST"],"one_liner":"An autoregressive VLA that tokenizes π0's actions into discrete tokens with the FAST tokenizer and predicts them one by one.","explanation":"π0-FAST is a model Physical Intelligence released in January 2025 alongside the FAST action tokenizer; it shares π0's backbone network and training data, differing only in how actions are output. π0 generates continuous actions with flow matching, while π0-FAST turns actions into discrete tokens and predicts them one at a time, like a language model. Traditional discretization — binning each dimension at each timestep separately — barely learns anything on high-frequency, dexterous tasks. FAST instead applies a discrete cosine transform (DCT, the same transform used in JPEG compression) to a chunk of actions, quantizes it, and compresses it further with byte-pair encoding (BPE), so a chunk of actions typically needs only 30–60 tokens. This trains up to 5x faster with performance comparable to the flow-matching version, and produced the first generalist policy trained on the DROID dataset able to follow instructions zero-shot in new environments. The cost is that autoregressive decoding is noticeably slower than flow matching. Weights are open-sourced in openpi.","example":"A bimanual robot controlled at 50Hz with 14 dimensions per step means 700 numbers per second of action. The old per-dimension binning approach would need 700 tokens; after FAST compression, only a few dozen remain, letting the model learn to fold clothes with next-token prediction, then convert tokens back into continuous actions at inference time.","related":["π0","Action Tokenizer","Discrete Cosine Transform","Byte-Pair Encoding","Action Binning","Autoregressive Decoding"]},{"id":"pi0-5","category":"named_model","sec":5,"tier":1,"sources":[{"title":"π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054)","url":"https://arxiv.org/abs/2504.16054"},{"title":"Physical-Intelligence/openpi GitHub 仓库","url":"https://github.com/Physical-Intelligence/openpi"}],"as_of":"2025-09","related_ids":["pi0","co-training","pi0-fast","open-world","long-horizon-task","pi-star-0-6"],"name":"π0.5","alt":"π0.5","abbr":"π0.5","aliases":["pi0.5","pi05","π0.5: a Vision-Language-Action Model with Open-World Generalization"],"one_liner":"PI's 2025 VLA that handles long-horizon tasks like tidying a kitchen in real homes it has never seen.","explanation":"π0.5 is the upgraded version of π0 that Physical Intelligence released on April 22, 2025, focused on open-world generalization — getting a robot to work in homes it never visited during training. The approach co-trains on a mix of data sources: roughly 400 hours of mobile-manipulator data collected across about 100 homes makes up only about 2.4% of the pretraining examples, with the rest coming from other robots, lab data, web image-text question answering, and high-level semantic annotations. At inference, the model first predicts the next subtask in words (such as “pick up the plate”), then generates low-level actions conditioned on that — similar to chain-of-thought. During training, the pretraining stage compresses actions into discrete FAST tokens for efficiency, and post-training attaches a flow-matching action expert that outputs continuous actions. The paper tested it in 3 real homes excluded from training, where it completed multistep tasks like tidying a kitchen or straightening a bedroom. The weights were open-sourced in openpi in September 2025.","example":"Given the instruction “clean the kitchen” in a home it has never visited, π0.5 first states a subtask like “put the plate in the sink,” then executes the corresponding actions, working through the task step by step.","related":["π0","Co-training","π0-FAST","Open-world","Long-horizon Task","π*0.6"]},{"id":"pi-star-0-6","category":"named_model","sec":5,"tier":2,"sources":[{"title":"π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)","url":"https://arxiv.org/abs/2511.14759"},{"title":"π*0.6: a VLA that Learns from Experience (Physical Intelligence blog)","url":"https://www.pi.website/blog/pistar06"}],"as_of":"2025-11","related_ids":["recap","vision-language-action-model","reinforcement-fine-tuning","advantage-conditioning","value-function","physical-intelligence"],"name":"π*0.6","alt":"π*0.6","abbr":"π*0.6","aliases":["π0.6","pi0.6","pi*0.6","pi-star-0.6"],"one_liner":"π0.6 after reinforcement learning with RECAP, able to keep improving from demonstrations, corrections, and its own experience.","explanation":"π*0.6 is a vision-language-action model Physical Intelligence released in November 2025. Its base, π0.6, builds on π0.5 but switches to a Gemma 3 4B vision-language backbone and grows the action expert to 860 million parameters; the asterisk marks that it was further trained with reinforcement learning using the RECAP method. A model trained purely on human demonstrations tends to drift further off course after a small mistake (compounding error), making it hard to succeed reliably. RECAP trains on three kinds of data at once: human demonstrations, corrections made when an expert teleoperates in to take over, and the robot's own successes and failures from autonomous runs. It first trains a value function that estimates how many steps remain to completion, uses that to judge whether a stretch of actions made things better or worse (its advantage), then trains the model together with an “Advantage: positive / negative” text condition, and at deployment only asks it for the “positive” behavior. On the hardest tasks, this more than doubled throughput and roughly halved the failure rate.","example":"After RECAP training, π*0.6 made espresso drinks continuously for 13 hours, folded unfamiliar laundry without interruption for more than two hours in a new home, and assembled real shipping boxes on a factory floor.","related":["RECAP","Vision-Language-Action Model","Reinforcement Fine-Tuning (RL Fine-Tuning)","Advantage Conditioning","Value Function","Physical Intelligence"]},{"id":"pi0-7","category":"named_model","sec":5,"tier":2,"sources":[{"title":"π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities (arXiv 2604.15483)","url":"https://arxiv.org/abs/2604.15483"},{"title":"π0.7: a Steerable Model with Emergent Capabilities (Physical Intelligence blog)","url":"https://www.pi.website/blog/pi07"}],"as_of":"2026-04","related_ids":["steerability","pi-star-0-6","pi0-5","compositional-generalization","cross-embodiment","hi-robot"],"name":"π0.7","alt":"π0.7","abbr":"","aliases":["pi0.7","pi07"],"one_liner":"PI's 2026 generalist robot model that can be steered with rich prompts and shows early signs of compositional generalization.","explanation":"π0.7 is a robot foundation model Physical Intelligence released in April 2026. Simply mixing data from different robots, human videos, and autonomous runs (including failures) together in training doesn't work well. π0.7's approach is to attach a richer prompt to every piece of data: language describing the task and its substeps, metadata like speed and quality, a label for whether joint-space or end-effector control was used, and subgoal images showing what things look like after a substep is done (which, at inference, can be generated by a lightweight world model). This lets data of different quality and different ways of doing the same task all be used, and at inference the prompt specifies “how to do it” — steerability. PI reports that without fine-tuning it matches the level of task-specific π*0.6 models at folding laundry, making coffee, and folding paper boxes, and shows the first signs of compositional generalization: it learns to use an air fryer it has never seen through step-by-step language guidance, and folds laundry on a bimanual UR5e with no laundry-folding data of its own.","example":"Asked to put a sweet potato in an air fryer, the robot only partially completes the task on a few tries when given just one instruction; guided step by step in language by a person, it finishes the task; fine-tuning a high-level policy on that guidance data then lets it generate its own substeps and complete the task fully autonomously.","related":["Steerability","π*0.6","π0.5","Compositional Generalization","Cross-Embodiment","Hi Robot"]},{"id":"hi-robot","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models (arXiv 2502.19417)","url":"https://arxiv.org/abs/2502.19417"},{"title":"Hi Robot 论文 HTML 全文","url":"https://arxiv.org/html/2502.19417v2"}],"as_of":"2025-07","related_ids":["hierarchical-architecture","dual-system-architecture","pi0","instruction-following","language-corrections","physical-intelligence"],"name":"Hi Robot","alt":"Hi Robot","abbr":"","aliases":["Hierarchical Interactive Robot","Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models"],"one_liner":"A 2025 Physical Intelligence hierarchical system where a high-level VLM breaks down complex instructions and low-level π0 executes them.","explanation":"Hi Robot was released in February 2025 by Physical Intelligence together with researchers at Stanford and UC Berkeley, published at ICML 2025. A single VLA is good at executing simple instructions like “pick up the cup,” but struggles with complex requests that carry conditions, and with a user interrupting mid-execution to correct it. Hi Robot splits the system into two layers: the high level is a vision-language model built on PaliGemma-3B that looks at the current view and what the user says, reasons out the simple instruction to execute next, and can also talk back to the user; the low level is π0, which turns that simple instruction into continuous actions. To train the high level, the team cut teleoperation demonstrations into short skill segments, then had a large VLM work backward to infer “what the user might have said at that moment, and how the robot should respond,” synthesizing conversational training data and skipping manual annotation. The system was tested on a single-arm UR5e, a bimanual ARX, and a mobile bimanual ARX.","example":"When the user says “make me a vegetarian sandwich,” the high level breaks it down into step-by-step sub-instructions for getting each ingredient, skipping any meat; while clearing a table, if the user says “that's not trash,” the robot stops and adjusts what it's doing.","related":["Hierarchical Architecture","Dual-System Architecture (System 1 / System 2)","π0","Instruction Following","Language Corrections","Physical Intelligence"]},{"id":"recap","category":"named_model","sec":5,"tier":2,"sources":[{"title":"π*0.6: a VLA That Learns From Experience (arXiv:2511.14759)","url":"https://arxiv.org/abs/2511.14759"}],"as_of":"2025-11","related_ids":["pi-star-0-6","advantage-conditioning","value-function","offline-reinforcement-learning","human-in-the-loop","classifier-free-guidance"],"name":"RECAP","alt":"RECAP","abbr":"RECAP","aliases":["RL with Experience and Corrections via Advantage-conditioned Policies"],"one_liner":"Physical Intelligence's method for letting a VLA keep improving from its own experience plus human corrections.","explanation":"RECAP is a training method Physical Intelligence released alongside π*0.6 in November 2025. Pure imitation learning can only ever match the level of its demonstrations, while a deployed robot generates a large stream of successes, failures, and human takeovers. RECAP combines all three kinds of data — offline demonstrations, trajectories from the robot's own autonomous runs, and corrections from an expert who teleoperates in to intervene during a run. It first trains a value function that predicts how many steps remain until task completion, uses that to judge whether each action was better or worse than average (its advantage), then trains the policy together with an “Advantage: positive / negative” text token as input. At inference, conditioning on “positive” pushes the model to output better actions. Deployment, retraining the value function, and fine-tuning the policy can be repeated in a loop. On tasks like folding laundry, assembling boxes, and making coffee, this more than doubled throughput and roughly halved the failure rate.","example":"After training with RECAP, PI reports that π*0.6 can make espresso drinks continuously for 13 hours, and fold unfamiliar laundry in unfamiliar homes for more than two hours at a stretch.","related":["π*0.6","Advantage Conditioning","Value Function","Offline Reinforcement Learning","Human-in-the-Loop","Classifier-Free Guidance"]},{"id":"q-transformer","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Q-Transformer (arXiv 2309.10150)","url":"https://arxiv.org/abs/2309.10150"},{"title":"Q-Transformer project page","url":"https://qtransformer.github.io/"}],"as_of":"2023-09","related_ids":["qt-opt","offline-reinforcement-learning","q-function","rt-1","conservative-q-learning","decision-transformer"],"name":"Q-Transformer","alt":"Q-Transformer","abbr":"","aliases":["Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions"],"one_liner":"A large-scale offline reinforcement learning method that uses a Transformer to estimate Q-values one action dimension at a time.","explanation":"Q-Transformer was released by Yevgen Chebotar, Sergey Levine, and colleagues at Google DeepMind in September 2023 and published at CoRL 2023. The problem it targets: robots often have both human demonstrations and a much larger pool of autonomously collected data, including failures, and offline reinforcement learning (training only from existing data, without further interaction) offers a way to use both. The method represents the Q-function (a value function that scores a state-action pair) with a Transformer: each action dimension is discretized into a number of bins, and the Q-value for each dimension is predicted one at a time, similar to generating tokens. Training is stabilized with conservative regularization (pushing the value of unseen actions down to a minimum) and Monte Carlo return estimates. On RT-1's real-robot multi-task benchmark, Q-Transformer outperforms RT-1, IQL, and Decision Transformer, and is especially good at making use of failure data.","example":"On RT-1's real-robot multi-task data, Q-Transformer trains on both human demonstrations and failure episodes from the robot's own autonomous attempts, outperforming RT-1, which only imitates successful demonstrations.","related":["QT-Opt","Offline Reinforcement Learning","Q-Function","RT-1","Conservative Q-Learning","Decision Transformer"]},{"id":"serl","category":"named_model","sec":5,"tier":3,"sources":[{"title":"SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (arXiv 2401.16013)","url":"https://arxiv.org/abs/2401.16013"},{"title":"SERL project page","url":"https://serl-robot.github.io/"}],"as_of":"2024-05","related_ids":["hil-serl","real-world-reinforcement-learning","reinforcement-learning-with-prior-data","success-detector","impedance-control","sample-efficiency"],"name":"SERL","alt":"SERL","abbr":"SERL","aliases":["Sample-Efficient Robotic Reinforcement Learning","SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning"],"one_liner":"An open-source real-world reinforcement learning toolkit from Berkeley and others that can train a policy on a real robot arm in tens of minutes.","explanation":"SERL was released by researchers at UC Berkeley, Stanford, the University of Washington, and Intrinsic (Jianlan Luo, Sergey Levine, and others) in January 2024, published at ICRA 2024. Getting reinforcement learning to work on real robots is hard not just because of the algorithms, but because of engineering problems: how to define the reward, how to reset after each episode, and whether the low-level controller is safe. SERL packages solutions to these as open-source tools: RLPD, a sample-efficient off-policy algorithm that can also make use of a small number of human demonstrations; an image classifier that judges task success and serves as the reward; two alternating 'forward' and 'backward' policies that reset the environment automatically, removing the need for manual resets; and an impedance controller for the Franka arm that keeps contact safe and compliant. In the paper's experiments, tasks like PCB component insertion, cable routing, and object relocation reach close to 100% success after just 25 to 50 minutes of training. The follow-up work HIL-SERL adds live human corrections on top of this.","example":"For PCB component insertion, a small number of human demonstrations are recorded first, then an image classifier that judges 'is the component seated correctly' is trained as the reward; the arm then practices on its own and reaches reliable insertion in under an hour.","related":["HIL-SERL","Real-World Reinforcement Learning","Reinforcement Learning with Prior Data","Success Detector","Impedance Control","Sample Efficiency"]},{"id":"hil-serl","category":"named_model","sec":5,"tier":2,"sources":[{"title":"Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (arXiv 2410.21845)","url":"https://arxiv.org/abs/2410.21845"},{"title":"HIL-SERL 项目主页","url":"https://hil-serl.github.io/"},{"title":"Jianlan Luo 个人主页（HIL-SERL 发表于 Science Robotics 2025）","url":"https://people.eecs.berkeley.edu/~jianlanluo/"}],"as_of":"2025-08","related_ids":["serl","human-in-the-loop","real-world-reinforcement-learning","human-gated-dagger","reward-model","reinforcement-fine-tuning"],"name":"HIL-SERL","alt":"HIL-SERL","abbr":"HIL-SERL","aliases":["Human-in-the-Loop Sample-Efficient Robotic Reinforcement Learning","Precise and Dexterous Robotic Manipulation via Human-in-the-Loop RL"],"one_liner":"Berkeley's real-robot RL system where a human can take over anytime, learning fine manipulation in 1 to 2.5 hours.","explanation":"HIL-SERL is a real-robot visual reinforcement-learning system from Jianlan Luo, Charles Xu, Jeffrey Wu, and Sergey Levine at UC Berkeley, posted to arXiv in October 2024 and featured on the cover of Science Robotics in August 2025. It builds on the same group's open-source SERL software stack. First, teleoperation collects a handful of successful and failed examples to train a binary classifier as a sparse reward signal; a small number of demonstrations are then loaded into a replay buffer; online training uses an off-policy algorithm based on RLPD, which can reuse old data repeatedly, and an operator can take over and correct the robot at any time with a device like a SpaceMouse, with that correction data also fed back into learning. Real-robot RL is usually too slow and unstable to be practical; HIL-SERL uses a pretrained vision backbone, demonstrations plus human corrections, and a safe low-level controller to cut training down to 1 to 2.5 hours, reaching near-100% success on most tasks — roughly double the success rate and 1.8x the execution speed of imitation-learning baselines.","example":"Tasks in the paper include Jenga extraction — the arm whips a block out of a tower — flipping objects in a pan, and precision assembly tasks like a circuit board, an IKEA shelf, a car dashboard, and a timing belt.","related":["SERL","Human-in-the-Loop","Real-World Reinforcement Learning","Human-Gated DAgger","Reward Model","Reinforcement Fine-Tuning (RL Fine-Tuning)"]},{"id":"diffusion-policy-policy-optimization","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Diffusion Policy Policy Optimization (arXiv:2409.00588)","url":"https://arxiv.org/abs/2409.00588"},{"title":"DPPO 项目主页","url":"https://diffusion-ppo.github.io/"},{"title":"Allen Z. Ren 个人主页（论文列表，标注 ICLR 2025）","url":"https://allenzren.github.io/"}],"as_of":"2025","related_ids":["diffusion-policy","proximal-policy-optimization","reinforcement-fine-tuning","policy-gradient","sim-to-real-transfer","reinflow"],"name":"Diffusion Policy Policy Optimization","alt":"DPPO（扩散策略策略优化）","abbr":"DPPO","aliases":["DPPO"],"one_liner":"Fine-tunes a diffusion policy directly with PPO policy-gradient reinforcement learning by treating each denoising step as a decision.","explanation":"DPPO was proposed in September 2024 by Allen Z. Ren and colleagues at Princeton, together with researchers at MIT, Toyota Research Institute, Carnegie Mellon University, and Harvard, published at ICLR 2025. A diffusion policy is trained with imitation learning, so its performance is capped by the quality of the demonstrations; policy-gradient reinforcement learning methods like PPO need to compute action probabilities, and multi-step diffusion denoising doesn't lend itself to that directly, so this kind of fine-tuning was widely assumed to be inefficient. DPPO splits the process into two nested Markov decision processes: the outer one is interacting with the environment, and the inner one treats each denoising step as a Gaussian-sampling “action,” whose probability can be computed, making it possible to fine-tune end to end with PPO. Experiments show it explores close to the demonstration data's distribution, trains stably, and produces more robust policies after fine-tuning, making it a common baseline for reinforcement-learning fine-tuning of diffusion and flow-matching policies.","example":"On the Furniture-Bench furniture-assembly simulation tasks, DPPO raised a pretrained diffusion policy's success rate from 57% to 97% on One-Leg and from 12% to 87% on Lamp, and the assembly policy also transferred zero-shot to a real robot.","related":["Diffusion Policy","Proximal Policy Optimization","Reinforcement Fine-Tuning (RL Fine-Tuning)","Policy Gradient","Sim-to-Real Transfer","ReinFlow"]},{"id":"v-gps","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Steering Your Generalists (arXiv 2410.13816)","url":"https://arxiv.org/abs/2410.13816"},{"title":"V-GPS project page","url":"https://nakamotoo.github.io/V-GPS/"}],"as_of":"2024-10","related_ids":["value-guided-sampling","calibrated-q-learning","offline-reinforcement-learning","q-function","inference-time-compute","octo"],"name":"V-GPS","alt":"V-GPS","abbr":"V-GPS","aliases":["Value-Guided Policy Steering","Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance"],"one_liner":"A method that re-ranks a generalist policy's candidate actions at deployment time using a value function learned with offline reinforcement learning.","explanation":"V-GPS was released by Mitsuhiko Nakamoto, Sergey Levine, and colleagues at UC Berkeley and CMU in October 2024, published at CoRL 2024. Generalist robot policies are trained on demonstration data of wildly varying quality, and the larger the dataset, the harder it is to clean up. V-GPS leaves the original policy untouched: it first trains a Q-function (which scores a state-action pair) with offline reinforcement learning (using only existing data, mainly with Cal-QL) on BridgeData V2 and RT-1 data; at deployment, the generalist policy samples K candidate actions at once (the paper tries 10 and 50), and the Q-function picks the highest-scoring one to execute. It requires no fine-tuning and no access to the policy's weights, and the same value function improves five different policies — Octo, RT-1-X, OpenVLA, and others — making it an early example of trading extra inference-time compute for better performance in robotics.","example":"On a real robot, having Octo pick up sushi and put it in a bowl: Octo first generates several candidate actions at each step, V-GPS's Q-function scores each one, and only the highest-scoring action is executed; the project page reports a clear success-rate increase on this kind of task.","related":["Value-Guided Sampling","Calibrated Q-Learning","Offline Reinforcement Learning","Q-Function","Inference-Time Compute","Octo"]},{"id":"generative-value-learning","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Vision Language Models are In-Context Value Learners (arXiv:2411.04549)","url":"https://arxiv.org/abs/2411.04549"},{"title":"Generative Value Learning 项目主页","url":"https://generative-value-learning.github.io/"}],"as_of":"2025-01","related_ids":["progress-reward-model","vlm-as-reward","value-function","success-detector","in-context-learning","data-curation"],"name":"Generative Value Learning (GVL)","alt":"GVL（生成式价值学习）","abbr":"GVL","aliases":["Vision Language Models are In-Context Value Learners"],"one_liner":"Has a vision-language model estimate task progress from shuffled video frames, turning it into a general-purpose value function.","explanation":"GVL was proposed in November 2024 by researchers at Google DeepMind, together with the University of Pennsylvania and Stanford, published at ICLR 2025. “Value” here means what percentage of the task is complete, which can be used to judge success or failure, filter data, or serve as a reward for reinforcement learning. Directly having a vision-language model (VLM) score each frame in chronological order works poorly, because adjacent frames are highly correlated and the model tends to just output a monotonically increasing score based on order alone. GVL instead shuffles the video frames first, then has the model (the paper uses Gemini-1.5-Pro) estimate completion percentage frame by frame, forcing it to actually look at what's in each frame. It needs no task-specific training, works zero-shot or few-shot across more than 300 real tasks, and can even learn by putting human videos or other robots' demonstrations into its context.","example":"Use GVL to score a batch of demonstration videos for progress; clips where the predicted progress correlates poorly with the true chronological order (usually failed or low-quality demonstrations) can be filtered out, and the remaining data used to train an imitation-learning policy.","related":["Progress Reward Model","VLM-as-Reward","Value Function","Success Detector","In-Context Learning","Data Curation"]},{"id":"grape","category":"named_model","sec":5,"tier":3,"sources":[{"title":"GRAPE: Generalizing Robot Policy via Preference Alignment (arXiv 2411.19309)","url":"https://arxiv.org/abs/2411.19309"},{"title":"GRAPE 项目页","url":"https://grape-vla.github.io/"}],"as_of":"2025-02","related_ids":["direct-preference-optimization","vision-language-action-model","openvla","behavior-cloning","generalization","reinforcement-fine-tuning"],"name":"GRAPE","alt":"GRAPE","abbr":"GRAPE","aliases":["Generalizing Robot Policy via Preference Alignment"],"one_liner":"A method that uses preference alignment on successful and failed trajectories to improve a VLA's generalization to new tasks.","explanation":"GRAPE was proposed in November 2024 by researchers at UNC Chapel Hill, the University of Washington, and the University of Chicago, with experiments built on OpenVLA as the base model. Most VLA models only do behavior cloning (learning by copying demonstrations) on successful demonstrations; having never seen failure, they don't know which behaviors to avoid, and tend to make mistakes when generalizing to new tasks. GRAPE borrows the idea of preference alignment from large language models: it has the model execute a task many times, ranks the resulting trajectories from best to worst by score, and then uses trajectory-wise preference optimization (TPO) to bias the model toward the better trajectories — so even failed trajectories provide useful information. For scoring, a complex task is first broken into stages, and a vision-language model generates spatiotemporal constraints for each stage to score against; swapping in a different set of constraints lets alignment target different goals, such as “safer” or “more efficient.”","example":"The authors report that GRAPE improves success rate by 51.79% on in-distribution tasks and 58.20% on unseen tasks; when aligned toward safety and efficiency goals, collision rate drops 37.44% and the number of execution steps falls 11.15%.","related":["Direct Preference Optimization","Vision-Language-Action Model","OpenVLA","Behavior Cloning","Generalization","Reinforcement Fine-Tuning (RL Fine-Tuning)"]},{"id":"conrft","category":"named_model","sec":5,"tier":3,"sources":[{"title":"ConRFT (arXiv 2502.05450)","url":"https://arxiv.org/abs/2502.05450"},{"title":"ConRFT 项目主页","url":"https://cccedric.github.io/conrft/"}],"as_of":"2025-04","related_ids":["reinforcement-fine-tuning","consistency-policy","hil-serl","human-in-the-loop","octo","real-world-reinforcement-learning"],"name":"ConRFT","alt":"ConRFT","abbr":"ConRFT","aliases":["Cal-ConRFT","HIL-ConRFT"],"one_liner":"A method that fine-tunes a VLA with reinforcement learning using a consistency policy, offline first, then online on the real robot.","explanation":"ConRFT was proposed in February 2025 by Dongbin Zhao's group at the Institute of Automation, Chinese Academy of Sciences, published at RSS 2025. Supervised fine-tuning of a VLA on a small number of demonstrations often isn't reliable enough for contact-rich real-robot tasks, while running reinforcement learning directly on the real robot is slow and unsafe. ConRFT works in two stages. The offline stage (Cal-ConRFT) combines behavior cloning with Q-learning to learn a reasonably stable policy and value estimate from a small number of demonstrations. The online stage (HIL-ConRFT) continues fine-tuning with reinforcement learning on the real robot, letting a person take over and correct it at any time to keep exploration safe. The action head uses a consistency policy (a diffusion-style model that can generate an action in one or a few steps), which fits well with reinforcement learning. Built on Octo-small, after 45–90 minutes of online fine-tuning across 8 real tasks, average success rate reached 96.3%, 144% higher than pure supervised fine-tuning.","example":"Tasks included picking up a banana, opening a drawer, putting bread in a toaster, mounting a car wheel, and hanging a Chinese knot; during online training, the operator could take over whenever the robot was about to make a mistake.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Consistency Policy","HIL-SERL","Human-in-the-Loop","Octo","Real-World Reinforcement Learning"]},{"id":"ript-vla","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Interactive Post-Training for Vision-Language-Action Models (arXiv 2505.17016)","url":"https://arxiv.org/abs/2505.17016"}],"as_of":"2025-05","related_ids":["reinforcement-fine-tuning","post-training","openvla-oft","libero-benchmark","proximal-policy-optimization","simplevla-rl"],"name":"RIPT-VLA","alt":"RIPT-VLA（交互式后训练）","abbr":"","aliases":["Reinforcement Interactive Post-Training","RIPT-VLA: Interactive Post-Training for Vision-Language-Action Models"],"one_liner":"A method that post-trains a pretrained VLA with reinforcement learning using only a binary success/failure reward.","explanation":"RIPT-VLA was released by Shuhan Tan, Philipp Krähenbühl, and colleagues at UT Austin in May 2025. VLAs are normally pretrained and then supervised-fine-tuned on expert demonstrations, but this works poorly when demonstrations are scarce. RIPT-VLA adds a third stage, 'interactive post-training': the model repeatedly attempts the task in the environment and is trained with reinforcement learning using only a binary reward for task success or failure. It samples the same initial state multiple times and estimates the advantage using the average score of the other attempts in the same group as a baseline (a leave-one-out estimator), then updates the policy with PPO; groups where every attempt succeeds or every attempt fails carry no learning signal and are discarded and resampled. It improves the QueST model by 21.2% and pushes the 7B OpenVLA-OFT to 97.5% on LIBERO; given just a single demonstration, it can take a model with 4% success up to 97% within 15 iterations.","example":"A VLA fine-tuned on just one demonstration starts at only 4% success; letting it repeatedly attempt the task in simulation and scoring each attempt as success or failure raises its success rate to 97% after 15 iterations.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Post-training","OpenVLA-OFT","LIBERO Benchmark","Proximal Policy Optimization","SimpleVLA-RL"]},{"id":"vla-rl","category":"named_model","sec":5,"tier":3,"sources":[{"title":"VLA-RL (arXiv 2505.18719)","url":"https://arxiv.org/abs/2505.18719"},{"title":"GuanxingLu/vlarl (GitHub)","url":"https://github.com/GuanxingLu/vlarl"}],"as_of":"2025-05","related_ids":["reinforcement-fine-tuning","proximal-policy-optimization","reward-model","openvla","libero-benchmark","simplevla-rl"],"name":"VLA-RL","alt":"VLA-RL","abbr":"","aliases":["VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning"],"one_liner":"An early framework that further improves an autoregressive VLA, such as OpenVLA, using online reinforcement learning.","explanation":"VLA-RL was proposed in May 2025 by researchers at Tsinghua Shenzhen International Graduate School and Nanyang Technological University, with code open-sourced. A VLA trained purely by imitating demonstrations has only seen a limited set of states, and tends to fail as soon as it drifts out of distribution; VLA-RL instead lets a pretrained autoregressive VLA keep improving online by trying things itself in the environment. It treats one robot manipulation trajectory as a multi-turn multimodal conversation and applies trajectory-level reinforcement learning with PPO (Proximal Policy Optimization); to ease the sparse-reward problem, it fine-tunes a vision-language model into a robotic process reward model, with training labels coming from automatically segmented task stages (using keyframes such as moments when the gripper becomes stable). On the engineering side, it also uses curriculum-based task selection, GPU-load-balanced parallel environments, batched decoding, and value-network warm-up. VLA-RL is one of the earlier works to systematically demonstrate that 'VLA plus online RL' is workable, and the authors also observed that increasing test-time optimization keeps improving results further.","example":"Across 40 manipulation tasks in LIBERO, VLA-RL raised OpenVLA-7B's average success rate from 76.5% to 81.0%, on par with π0-FAST.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","Proximal Policy Optimization","Reward Model","OpenVLA","LIBERO Benchmark","SimpleVLA-RL"]},{"id":"reinflow","category":"named_model","sec":5,"tier":3,"sources":[{"title":"ReinFlow (arXiv 2505.22094)","url":"https://arxiv.org/abs/2505.22094"},{"title":"ReinFlow project page","url":"https://reinflow.github.io/"}],"as_of":"2025-05","related_ids":["flow-matching","diffusion-policy-policy-optimization","pirl","reinforcement-fine-tuning","rectified-flow","policy-gradient"],"name":"ReinFlow","alt":"ReinFlow","abbr":"","aliases":["ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning"],"one_liner":"A method that injects learnable noise into a flow-matching policy so it can be fine-tuned with online reinforcement learning.","explanation":"ReinFlow was released by Tonghe Zhang, Chao Yu, Yu Wang, and colleagues at Carnegie Mellon University, Tsinghua University, and others in May 2025, published at NeurIPS 2025. Flow-matching policies (generative policies that turn noise into actions step by step using a velocity field) hit an obstacle when researchers try to improve them further with reinforcement learning: the generation process is deterministic, so it has no well-defined action probability and little exploration. ReinFlow adds a learnable Gaussian noise term at each denoising step, turning generation into a discrete-time Markov process, which makes the likelihood exactly computable and lets the policy be fine-tuned with policy gradients; the noise network is used only during training and dropped at inference time. On legged-locomotion tasks, this raises the episode reward of rectified-flow policies by about 135% on average, with roughly 83% less compute than DPPO, a diffusion-policy fine-tuning method. ReinFlow belongs to the same line of work as DPPO and πRL: using reinforcement learning to fine-tune generative policies.","example":"A Shortcut flow policy that needs very few denoising steps is first trained on demonstrations, then fine-tuned with ReinFlow using online reinforcement learning on robomimic simulated manipulation tasks, producing a clear jump in success rate.","related":["Flow Matching","Diffusion Policy Policy Optimization","πRL","Reinforcement Fine-Tuning (RL Fine-Tuning)","Rectified Flow","Policy Gradient"]},{"id":"diffusion-steering-via-reinforcement-learning","category":"named_model","sec":5,"tier":3,"sources":[{"title":"Steering Your Diffusion Policy with Latent Space Reinforcement Learning (arXiv:2506.15799)","url":"https://arxiv.org/abs/2506.15799"},{"title":"DSRL 项目主页","url":"https://diffusion-steering.github.io/"}],"as_of":"2025-06","related_ids":["noise-space-policy-steering","diffusion-policy","pi0","real-world-reinforcement-learning","residual-reinforcement-learning","diffusion-policy-policy-optimization"],"name":"Diffusion Steering via Reinforcement Learning","alt":"DSRL（扩散策略噪声空间强化学习）","abbr":"DSRL","aliases":["DSRL","Steering Your Diffusion Policy with Latent Space RL"],"one_liner":"Leaves a diffusion policy's weights untouched and uses reinforcement learning only to pick its input noise, steering it toward better actions.","explanation":"DSRL was proposed in June 2025 by Andrew Wagenmaker and colleagues in Sergey Levine's group at UC Berkeley, together with researchers at the University of Washington and Amazon, published at CoRL 2025. A diffusion policy generates an action by first sampling random noise and then denoising it, and for the same observation, different initial noise leads to different actions. DSRL treats this initial noise as a new “action space,” training a small reinforcement-learning policy (an MLP) that outputs noise given the observation, which is then handed to the frozen diffusion policy to denoise. The base policy's weights are never touched — it's only called as a black box — and because the noise always maps to some reasonable action from the demonstration data, exploration is better directed and needs less real-robot interaction. It's well suited to quickly and autonomously improving a behavior-cloned policy in a new environment, and it's a representative example of steering a policy through its noise space.","example":"The authors treated a public π0 checkpoint (trained on DROID) as a black box and ran reinforcement learning only in its noise space, improving performance on real-robot manipulation tasks without fine-tuning π0 itself.","related":["Noise-Space Policy Steering","Diffusion Policy","π0","Real-World Reinforcement Learning","Residual Reinforcement Learning","Diffusion Policy Policy Optimization"]},{"id":"robomonkey","category":"named_model","sec":5,"tier":3,"sources":[{"title":"arXiv 2506.17811: RoboMonkey","url":"https://arxiv.org/abs/2506.17811"},{"title":"RoboMonkey 项目主页","url":"https://robomonkey-vla.github.io/"}],"as_of":"2025-07","related_ids":["inference-time-compute","best-of-n-sampling","value-guided-sampling","vision-language-action-model","openvla","reward-model"],"name":"RoboMonkey","alt":"RoboMonkey","abbr":"","aliases":["RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models"],"one_liner":"A method that samples several candidate action sets at deployment time and uses a VLM-based verifier to pick the best one, improving VLA robustness.","explanation":"RoboMonkey was released by researchers at Stanford, UC Berkeley, and NVIDIA in June 2025, accepted at CoRL 2025, bringing the 'test-time scaling' idea from large language models over to VLAs. It leaves the original policy untouched: at deployment time, it samples several actions from the VLA, adds Gaussian perturbations, and uses majority voting to build a candidate set; a vision-language-model-based action verifier then scores the candidates and picks the best one to execute. The verifier is trained on automatically synthesized preference data (ranked by each candidate action's distance to the ground-truth action); the authors found that verification accuracy keeps improving with more synthetic data, and action error follows a roughly power-law relationship with the number of samples. Paired with models like OpenVLA, it delivers a 25-point absolute improvement on out-of-distribution tasks and 9 points on in-distribution tasks; an optimized serving engine can sample and verify 16 candidate actions in about 650 milliseconds.","example":"When OpenVLA encounters an object it has never seen, RoboMonkey has it sample several sets of candidate grasping actions at once, and the verifier picks out the set most likely to succeed for execution.","related":["Inference-Time Compute","Best-of-N Sampling","Value-Guided Sampling","Vision-Language-Action Model","OpenVLA","Reward Model"]},{"id":"rac","category":"named_model","sec":5,"tier":3,"sources":[{"title":"RaC (arXiv 2509.07953)","url":"https://arxiv.org/abs/2509.07953"},{"title":"RaC project page","url":"https://rac-scaling-robot.github.io/"}],"as_of":"2025-09","related_ids":["recovery-and-correction-data","human-in-the-loop","human-intervention-data","long-horizon-task","compounding-error","hil-serl"],"name":"RaC","alt":"RaC","abbr":"RaC","aliases":["Recovery and Correction","RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction"],"one_liner":"A method that scales long-horizon task data by having a human take over right before failure, first recovering and then correcting.","explanation":"RaC was released by Zheyuan Hu, Zackory Erickson, Aviral Kumar, and colleagues at Carnegie Mellon University in September 2025. The authors found that on contact-rich, long-horizon tasks involving deformable objects, simply piling on more expert demonstrations makes imitation-learning success rates plateau, because expert data almost never shows what to do after something goes wrong. RaC adds a human-in-the-loop fine-tuning stage after imitation-learning pretraining: while the policy is running, an operator takes over just before it fails, first steering the robot back to a familiar, in-distribution state (recovery), then demonstrating how to complete the current sub-task (correction), and ending the episode there. Across three real-robot bimanual tasks — hanging a shirt, sealing a food-storage container lid, and packing a takeout box — plus one simulated assembly task, RaC beats the previous best methods using roughly a tenth of the collection time and samples, and success rate scales roughly linearly with the number of recovery interventions performed during rollouts.","example":"While hanging a shirt, the policy is about to hang it crookedly on the hanger; the operator takes over, first moves the arm back to a normal hanger-holding pose, then demonstrates hanging it correctly, and this intervention data is added to the fine-tuning set.","related":["Recovery and Correction Data","Human-in-the-Loop","Human Intervention Data","Long-horizon Task","Compounding Error","HIL-SERL"]},{"id":"simplevla-rl","category":"named_model","sec":5,"tier":3,"sources":[{"title":"SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (arXiv 2509.09674)","url":"https://arxiv.org/abs/2509.09674"},{"title":"PRIME-RL/SimpleVLA-RL (GitHub)","url":"https://github.com/PRIME-RL/SimpleVLA-RL"}],"as_of":"2026-01","related_ids":["reinforcement-fine-tuning","openvla-oft","verl","reinforcement-learning-with-verifiable-rewards","robotwin","libero-benchmark"],"name":"SimpleVLA-RL","alt":"SimpleVLA-RL","abbr":"","aliases":["SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning"],"one_liner":"An open-source framework that runs large-scale online reinforcement learning on a VLA using only a success/failure reward.","explanation":"SimpleVLA-RL was released and open-sourced in September 2025 by Tsinghua University, the Shanghai Artificial Intelligence Laboratory, and other institutions, accepted at ICLR 2026. VLAs are usually trained with supervised fine-tuning (SFT) to imitate human demonstrations, but high-quality real-robot demonstrations are expensive, and policies trained this way tend to fail when they hit out-of-distribution situations. SimpleVLA-RL borrows the idea of using reinforcement learning to improve reasoning in large language models, adapting veRL, an RL framework built for large models, for VLAs — adding interactive trajectory sampling, parallel rendering across many environments, and distributed training; the reward looks only at whether the task ultimately succeeded (0 or 1), with no hand-designed reward shaping. Starting from OpenVLA-OFT as the base model, it reaches state-of-the-art results on LIBERO at the time and beats π0 on RoboTwin 1.0 and 2.0; when SFT is done with just one demonstration per task, LIBERO-Long success can go from 17.3% to 91.7%. The authors also observed a 'pushcut' phenomenon, where RL training discovers behaviors absent from the demonstrations, such as pushing an object into place instead of picking it up as demonstrated.","example":"In LIBERO simulation, OpenVLA-OFT is first fine-tuned on one demonstration per task until it occasionally succeeds, then lets it repeatedly try in a large number of parallel environments, rewarded only by final success or failure, which sharply raises its success rate.","related":["Reinforcement Fine-Tuning (RL Fine-Tuning)","OpenVLA-OFT","veRL (Volcano Engine Reinforcement Learning)","Reinforcement Learning with Verifiable Rewards","RoboTwin","LIBERO Benchmark"]},{"id":"vla-rft","category":"named_model","sec":5,"tier":3,"sources":[{"title":"VLA-RFT (arXiv 2510.00406)","url":"https://arxiv.org/abs/2510.00406"},{"title":"VLA-RFT 项目主页","url":"https://vla-rft.github.io/"}],"as_of":"2025-10","related_ids":["world-model","reinforcement-fine-tuning","group-relative-policy-optimization","vla-adapter","wmpo","learning-in-imagination"],"name":"VLA-RFT","alt":"VLA-RFT","abbr":"","aliases":["VLA Reinforcement Fine-Tuning in World Simulators","VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators"],"one_liner":"A method that runs reinforcement learning inside a learned world model to make a VLA more robust in just a few hundred steps.","explanation":"VLA-RFT was proposed in October 2025 by Westlake University, Zhejiang University, the OpenHelix team, and others. A VLA trained with imitation learning alone tends to accumulate errors and fail under perturbations; reinforcement learning can help, but real-robot interaction is expensive and traditional simulators still have a sim-to-real gap. The method first trains a roughly 138-million-parameter autoregressive world model on real interaction data, predicting future frames from the current image and action, and uses it as a controllable simulator; the policy then rolls out whole trajectories inside it, with reward computed from the pixel-level (L1) and perceptual (LPIPS) error between the predicted frames and the expert reference trajectory's frames — a verifiable reward — and the policy is updated with GRPO (Group Relative Policy Optimization). The base policy is VLA-Adapter. Fine-tuning for under 400 steps already beats the supervised fine-tuning baseline, and the result is more robust to perturbations in object position and initial state, showing that a world model can serve as a practical environment for VLA post-training.","example":"On LIBERO, the VLA-Adapter supervised fine-tuning baseline reaches 86.6% average success; after 400 steps of reinforcement fine-tuning inside the world model, this rises to 91.1%.","related":["World Model","Reinforcement Fine-Tuning (RL Fine-Tuning)","Group Relative Policy Optimization","VLA-Adapter","WMPO","Learning in Imagination"]},{"id":"rl-100","category":"named_model","sec":5,"tier":3,"sources":[{"title":"RL-100 (arXiv 2510.14830)","url":"https://arxiv.org/abs/2510.14830"},{"title":"RL-100 project page","url":"https://lei-kun.github.io/RL-100/"},{"title":"Tech Xplore: RL-100 framework helps robots refine learned tasks","url":"https://techxplore.com/news/2026-08-rl-framework-robots-refine-tasks.html"}],"as_of":"2026-08","related_ids":["real-world-reinforcement-learning","diffusion-policy","consistency-model","proximal-policy-optimization","offline-to-online-reinforcement-learning","hil-serl"],"name":"RL-100","alt":"RL-100","abbr":"","aliases":["RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning"],"one_liner":"A diffusion-policy-based real-world reinforcement learning framework that reached 1,000-for-1,000 success across eight manipulation tasks.","explanation":"RL-100 was released in October 2025 by teams from Shanghai Jiao Tong University, Shanghai Qi Zhi Institute, and Tsinghua University (Huazhe Xu and colleagues), published in Science Robotics in 2026. It targets the near-expert-level reliability that home and factory deployment require. The framework is built on a diffusion visuomotor policy and proceeds in three stages — imitation learning from human demonstrations, offline reinforcement learning, and real-world online reinforcement learning — all three sharing a single clipped PPO objective applied over the denoising process, which keeps improvements conservative and stable; consistency distillation then compresses the multi-step denoising into a single step to meet the demands of high-frequency control. Across eight real-robot tasks — pushing objects, bowling, pouring water, folding cloth, screwing in a bolt, juicing, and folding paper boxes, among others — it achieved 1,000 successes out of 1,000 trials, with completion speed matching or exceeding expert teleoperators.","example":"A juicing robot deployed in a shopping mall served a continuous stream of random customers zero-shot for about seven hours without a single failure.","related":["Real-World Reinforcement Learning","Diffusion Policy","Consistency Model","Proximal Policy Optimization","Offline-to-Online Reinforcement Learning","HIL-SERL"]},{"id":"pirl","category":"named_model","sec":5,"tier":3,"sources":[{"title":"πRL (arXiv:2510.25889)","url":"https://arxiv.org/abs/2510.25889"},{"title":"πRL 论文 HTML 版","url":"https://arxiv.org/html/2510.25889"}],"as_of":"2026-01","related_ids":["pi0","pi0-5","rlinf","reinforcement-fine-tuning","flow-matching","proximal-policy-optimization"],"name":"πRL","alt":"πRL","abbr":"","aliases":["piRL","pi_RL","πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models"],"one_liner":"An open-source framework for online reinforcement-learning fine-tuning of flow-matching VLAs like π0 and π0.5.","explanation":"πRL was released in October 2025 by teams from Tsinghua University, Peking University, the Institute of Automation at the Chinese Academy of Sciences, and others, built on RLinf, an open-source reinforcement learning framework. The difficulty is that a flow-matching VLA generates actions through multi-step denoising, which has no tractable log-probability for an action, making it hard to directly apply policy-gradient algorithms like PPO. πRL offers two solutions: Flow-Noise models the denoising process as a discrete-time MDP with a learnable noise network that makes the log-likelihood exactly computable, and Flow-SDE turns the deterministic ODE sampling into stochastic SDE sampling, forming a two-level MDP out of denoising and environment interaction that makes exploration easier. Both substantially improve a policy already fine-tuned with a small amount of supervised data, on both LIBERO and ManiSkill, and it has also been validated on GR00T N1.5.","example":"A π0 policy fine-tuned with supervised learning on a small number of demonstrations reaches only 57.6% success on LIBERO; after reinforcement learning with πRL, this rises to 97.6%, and π0.5 goes from 77.1% to 98.3%.","related":["π0","π0.5","RLinf","Reinforcement Fine-Tuning (RL Fine-Tuning)","Flow Matching","Proximal Policy Optimization"]},{"id":"wmpo","category":"named_model","sec":5,"tier":3,"sources":[{"title":"WMPO (arXiv:2511.09515)","url":"https://arxiv.org/abs/2511.09515"},{"title":"WMPO 项目主页","url":"https://wm-po.github.io"}],"as_of":"2025-11","related_ids":["world-model","reinforcement-fine-tuning","group-relative-policy-optimization","learning-in-imagination","openvla-oft","vla-rft"],"name":"WMPO","alt":"WMPO（基于世界模型的策略优化）","abbr":"WMPO","aliases":["World-Model-based Policy Optimization","WMPO: World Model-based Policy Optimization for Vision-Language-Action Models"],"one_liner":"A method that lets a VLA do reinforcement learning on trajectories 'imagined' by a video world model, with no real-robot interaction.","explanation":"WMPO was proposed in November 2025 by researchers at HKUST and ByteDance Seed. A VLA trained purely on expert demonstrations never learns to correct itself after a failure, while doing reinforcement learning directly on a real robot is too sample-expensive. WMPO first trains a video world model that predicts pixel-level frames, lets the VLA repeatedly practice inside trajectories 'imagined' by this model, and updates the policy with on-policy GRPO (Group Relative Policy Optimization: sampling several trajectories for the same task and scoring them against each other), all without ever interacting with the real environment. Working in pixel space rather than a latent space keeps the imagined frames aligned with the visual features the VLA already learned from pretraining on internet images. The experiments use OpenVLA-OFT as the base policy, and WMPO outperforms GRPO and DPO baselines on both MimicGen simulation tasks and the real-robot Mobile ALOHA, with self-correcting behavior emerging along the way.","example":"On a real-robot 'insert the block onto the peg' task with a 5-millimeter clearance, the base policy succeeds 53% of the time, DPO reaches 60%, and WMPO reaches 70% (30 trials each).","related":["World Model","Reinforcement Fine-Tuning (RL Fine-Tuning)","Group Relative Policy Optimization","Learning in Imagination","OpenVLA-OFT","VLA-RFT"]},{"id":"vla-opd","category":"named_model","sec":5,"tier":3,"sources":[{"title":"VLA-OPD (arXiv 2603.26666)","url":"https://arxiv.org/abs/2603.26666"},{"title":"VLA-OPD 项目主页","url":"https://irpn-lab.github.io/VLA-OPD/"}],"as_of":"2026-03","related_ids":["on-policy-distillation","supervised-fine-tuning","reinforcement-fine-tuning","catastrophic-forgetting","kullback-leibler-divergence","simplevla-rl"],"name":"VLA-OPD","alt":"VLA-OPD","abbr":"VLA-OPD","aliases":["On-Policy VLA Distillation","VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation"],"one_liner":"A VLA post-training method where a strong teacher corrects the student token by token on trajectories the student generated itself.","explanation":"VLA-OPD is a VLA post-training framework proposed by a team at HKUST (Guangzhou) in March 2026. Offline supervised fine-tuning (SFT) learns only from demonstration data and has no way to handle states the model drifts into on its own, and it is also prone to catastrophic forgetting of pretrained abilities; online reinforcement learning, meanwhile, is held back by sparse rewards and poor sample efficiency. VLA-OPD takes a middle path — on-policy distillation: the student policy runs in the environment on its own, and a stronger teacher policy provides dense, token-by-token supervision on exactly the states the student itself reaches, with no dependence on environment reward. The loss uses reverse KL divergence, which the authors argue biases the model toward 'committing to one mode,' making it more stable than forward KL (prone to exploding entropy) or hard cross-entropy (prone to premature entropy collapse). In experiments, the teacher is an expert trained by SimpleVLA-RL and the student is OpenVLA-OFT; on LIBERO and RoboTwin 2.0, VLA-OPD is more sample-efficient than reinforcement learning, more robust than SFT, and forgets less.","example":"On LIBERO, a student fine-tuned on just one demonstration per task starts at 48.9% average success; distilling it with VLA-OPD raises this to 87.4%, and following up with GRPO reinforcement learning brings it to 93.4%, close to the teacher's 93.9%.","related":["On-Policy Distillation","Supervised Fine-Tuning","Reinforcement Fine-Tuning (RL Fine-Tuning)","Catastrophic Forgetting","Kullback-Leibler Divergence","SimpleVLA-RL"]},{"id":"robometer","category":"named_model","sec":5,"tier":3,"sources":[{"title":"arXiv 2603.02115: Robometer","url":"https://arxiv.org/abs/2603.02115"},{"title":"Robometer 项目主页","url":"https://robometer.github.io/"}],"as_of":"2026-05","related_ids":["reward-model","progress-reward-model","vlm-as-reward","dense-reward","success-detector","failure-data"],"name":"Robometer","alt":"Robometer","abbr":"","aliases":["RBM-1M","Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons"],"one_liner":"A general-purpose robot reward model trained on both task-progress labels and pairwise trajectory comparisons.","explanation":"Robometer was released in March 2026 by researchers from the University of Southern California, the University of Washington, MIT, the Allen Institute for AI, NVIDIA, and other institutions, accepted at RSS 2026. Both reinforcement learning and data curation need a reward model that can judge how well a robot is doing, but earlier approaches relied mainly on frame-by-frame progress labels on expert demonstrations, leaving large amounts of failed and suboptimal trajectories unused. Robometer combines two supervision signals: a frame-level progress loss that anchors the reward's scale using expert data, and a preference loss from pairwise trajectory comparisons that learns which of two trajectories is better, which lets it also learn from failure data. The authors built RBM-1M, a dataset of more than one million trajectories spanning 21 robot embodiments, and the model is based on Qwen3-VL (main version 4B). The resulting reward can be used for online and offline reinforcement learning, failure detection, and retrieving data for imitation learning.","example":"Given a video of a robot performing 'put the cup in the drawer,' Robometer outputs a frame-by-frame task-progress score; this curve can be used directly as a dense reward for reinforcement learning, or to judge whether the attempt failed.","related":["Reward Model","Progress Reward Model","VLM-as-Reward","Dense Reward","Success Detector","Failure Data"]},{"id":"nvidia-isaac-gr00t-n1","category":"named_model","sec":6,"tier":1,"sources":[{"title":"GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)","url":"https://arxiv.org/abs/2503.14734"},{"title":"NVIDIA/Isaac-GR00T GitHub 仓库","url":"https://github.com/NVIDIA/Isaac-GR00T"},{"title":"GR00T N1.6 研究页（NVIDIA GEAR）","url":"https://research.nvidia.com/labs/gear/gr00t-n1_6/"}],"as_of":"2026-04","related_ids":["dual-system-architecture","data-pyramid","flow-matching","diffusion-transformer","nvidia-generalist-embodied-agent-research-lab","egoscale"],"name":"NVIDIA Isaac GR00T N1","alt":"GR00T N1 系列","abbr":"GR00T","aliases":["GR00T","GR00T N1 Series","GR00T N1.5","GR00T N1.6","GR00T N1.7","Isaac GR00T","Isaac-GR00T Repository"],"one_liner":"NVIDIA's open humanoid-robot foundation model family, built on a fast-slow architecture and continually updated since its 2025 launch.","explanation":"GR00T N1 is an open humanoid-robot foundation model (about 2.2 billion parameters) that NVIDIA released at GTC on March 18, 2025, with code under Apache 2.0 and weights under NVIDIA's open model license, both usable commercially. It uses a fast-slow, two-system design: System 2 is a vision-language model (Eagle-2 in N1) that looks at images and understands instructions; System 1 is a diffusion Transformer trained with flow matching that turns this into continuous action chunks, with each robot using its own state and action encoder/decoder. Data is organized as a “data pyramid”: web and human videos at the base, synthetic data from simulation and video models in the middle, and real-robot data at the top. Later releases followed: N1.5 (June 2025) froze the VLM and added a FLARE loss that lets it learn from human videos; N1.6 (December 2025) switched to a Cosmos-family VLM, doubled the action head, and moved to predicting relative actions; N1.7 (from April 2026), at about 3 billion parameters, switched its VLM to Cosmos-Reason2-2B and added pretraining on about 20,000 hours of EgoScale first-person human video.","example":"NVIDIA provides GR00T-N1.7-3B on Hugging Face along with versions fine-tuned on LIBERO, DROID, and SimplerEnv; the README recommends a GPU with at least 16GB of memory for inference and 40GB or more for fine-tuning.","related":["Dual-System Architecture (System 1 / System 2)","Data Pyramid","Flow Matching","Diffusion Transformer","NVIDIA Generalist Embodied Agent Research Lab","EgoScale"]},{"id":"nvidia-isaac-gr00t-n2","category":"named_model","sec":6,"tier":2,"sources":[{"title":"NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (NVIDIA Newsroom, 2026-03-16)","url":"https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world"},{"title":"World Action Models are Zero-shot Policies (DreamZero, arXiv 2602.15922)","url":"https://arxiv.org/abs/2602.15922"},{"title":"NVIDIA/Isaac-GR00T GitHub 仓库","url":"https://github.com/NVIDIA/Isaac-GR00T"}],"as_of":"2026-09","related_ids":["nvidia-isaac-gr00t-n1","dreamzero","world-action-model","vision-language-action-model","roboarena","nvidia"],"name":"NVIDIA Isaac GR00T N2","alt":"GR00T N2","abbr":"","aliases":["Isaac GR00T N2","GR00T N2"],"one_liner":"NVIDIA's next-generation robot foundation model, previewed for 2026, moving from a VLA design to a world-action-model architecture.","explanation":"GR00T N2 is the next-generation general-purpose robot foundation model NVIDIA previewed at GTC 2026 on March 16, 2026, succeeding the GR00T N1 series (N1, N1.5, N1.6, N1.7). The N1 series followed a VLA design: a vision-language model understands the scene and instruction, and a diffusion-style action head outputs actions. According to NVIDIA's announcement, N2 is “based on DreamZero research” and switches to a world-action model (WAM) architecture: built on a pretrained video generation model, it predicts future frames and robot actions at the same time, learning physical dynamics from video rather than just semantics. NVIDIA says it succeeds at new tasks in new environments more than twice as often as leading VLAs, and at launch it ranked first on the MolmoSpaces and RoboArena general-policy benchmarks. NVIDIA planned to release it before the end of 2026; as of September 2026, the newest GR00T weights NVIDIA had published on Hugging Face were still from the N1.7 series.","example":"Its research foundation, DreamZero, is built on a 14-billion-parameter autoregressive video diffusion model that, once optimized, can control a robot in real time closed-loop at 7Hz; adapting to a new robot takes only about 30 minutes of play data.","related":["NVIDIA Isaac GR00T N1","DreamZero","World Action Model","Vision-Language-Action Model","RoboArena","NVIDIA"]},{"id":"nvidia-cosmos","category":"named_model","sec":6,"tier":2,"sources":[{"title":"Cosmos World Foundation Model Platform for Physical AI (arXiv 2501.03575)","url":"https://arxiv.org/abs/2501.03575"},{"title":"NVIDIA Launches Cosmos World Foundation Model Platform to Accelerate Physical AI Development (NVIDIA Newsroom, 2025-01-06)","url":"https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-world-foundation-model-platform-to-accelerate-physical-ai-development"},{"title":"Cosmos 3: Omnimodal World Models for Physical AI (arXiv 2606.02800)","url":"https://arxiv.org/abs/2606.02800"}],"as_of":"2026-06","related_ids":["world-foundation-model","nvidia-cosmos-predict","nvidia-cosmos-transfer","nvidia-cosmos-reason","cosmos-3","nvidia"],"name":"NVIDIA Cosmos","alt":"Cosmos","abbr":"","aliases":["Cosmos World Foundation Models","Cosmos WFM"],"one_liner":"NVIDIA's open world-foundation-model platform, using video generation to produce training data and rehearsal environments for robots and self-driving cars.","explanation":"Cosmos is the world-foundation-model platform NVIDIA released at CES on January 6, 2025. A world foundation model (WFM) is a general-purpose model first pretrained on massive amounts of video to learn “what happens next in the world,” then post-trained for a specific robotics or self-driving scenario. The platform includes a video-curation pipeline (cutting about 20 million hours of raw video down to roughly 100 million clips), the Cosmos Tokenizer video tokenizer, both diffusion-based and autoregressive pretrained models, and post-training examples, all released with open weights. Several product lines followed: Cosmos Predict, which predicts future frames; Cosmos Transfer, which generates photorealistic video conditioned on structured inputs like segmentation or depth maps; and Cosmos Reason, a multimodal large model for physical commonsense and embodied reasoning. Cosmos 3, from June 2026, folds language, image, video, audio, and action into a single mixture Transformer model, offered in Edge, Nano (16B), and Super (64B) sizes, released under the Linux Foundation's OpenMDW-1.1 license.","example":"One use of Cosmos Transfer is sim-to-real: taking segmentation maps, depth maps, and edge maps rendered by a simulator as conditioning, it generates video that looks like real footage, used to expand training data for robots and self-driving cars.","related":["World Foundation Model","NVIDIA Cosmos Predict","NVIDIA Cosmos Transfer","NVIDIA Cosmos Reason","Cosmos 3","NVIDIA"]},{"id":"nvidia-cosmos-predict","category":"named_model","sec":6,"tier":3,"sources":[{"title":"nvidia-cosmos/cosmos-predict2.5 GitHub 仓库","url":"https://github.com/nvidia-cosmos/cosmos-predict2.5"},{"title":"nvidia-cosmos/cosmos-predict2 GitHub 仓库","url":"https://github.com/nvidia-cosmos/cosmos-predict2"},{"title":"Cosmos World Foundation Model Platform for Physical AI (arXiv 2501.03575)","url":"https://arxiv.org/abs/2501.03575"}],"as_of":"2026-06","related_ids":["nvidia-cosmos","world-foundation-model","video-generation-model","cosmos-policy","dreamgen","cosmos-3"],"name":"NVIDIA Cosmos Predict","alt":"Cosmos Predict","abbr":"","aliases":["Cosmos-Predict1","Cosmos-Predict2","Cosmos-Predict2.5"],"one_liner":"The video world model branch of NVIDIA's Cosmos family, generating what comes next from text, an image, or video.","explanation":"Cosmos Predict is the branch of NVIDIA's Cosmos world foundation model family responsible for predicting future frames. Predict1 launched with the Cosmos platform in January 2025, with 7B/14B diffusion models and several autoregressive model sizes; Predict2 opened up in June 2025, with 2B and 14B video models; Predict2.5, from October 2025, switched to flow-based generation, uses Cosmos-Reason1 as its text encoder, and merges text-to-video, image-to-video, and video continuation into one model. It's positioned as a fine-tunable base for robotics and autonomous driving: post-trained on robot data, it can become an action-conditioned simulator, a multi-view generator, or be used directly as a policy (as in Cosmos Policy). Starting in 2026, NVIDIA has shifted its main focus to the unified Cosmos 3, and the Predict repository is no longer under active development.","example":"NVIDIA provides a version of Cosmos-Predict2 post-trained with action conditioning on the Bridge robot dataset: given the current frame and a sequence of arm actions, it generates video of what happens after executing them; GR00T Dreams also uses a post-trained version of it to generate robot training video.","related":["NVIDIA Cosmos","World Foundation Model","Video Generation Model","Cosmos Policy","DreamGen","Cosmos 3"]},{"id":"nvidia-cosmos-transfer","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control (arXiv 2503.14492)","url":"https://arxiv.org/abs/2503.14492"},{"title":"nvidia-cosmos/cosmos-transfer2.5 GitHub 仓库","url":"https://github.com/nvidia-cosmos/cosmos-transfer2.5"}],"as_of":"2026-02","related_ids":["nvidia-cosmos","sim-to-real-transfer","sim-to-real-gap","generative-data-augmentation","synthetic-data","nvidia-cosmos-predict"],"name":"NVIDIA Cosmos Transfer","alt":"Cosmos Transfer","abbr":"","aliases":["Cosmos-Transfer1","Cosmos-Transfer2.5"],"one_liner":"NVIDIA's controllable video-generation model that produces photorealistic video conditioned on structure maps like depth, segmentation, and edges.","explanation":"Cosmos Transfer is the conditional-generation branch of NVIDIA's Cosmos family, using a multi-branch ControlNet structure (feeding extra structure maps in as generation conditions). Given spatial control signals such as a segmentation map, depth map, edge map, blurred footage, high-definition map, or lidar, it outputs photorealistic video that keeps the same layout and motion, but can change lighting, weather, and materials. Cosmos-Transfer1, from March 2025, can assign different weights to different control signals at different locations in the frame; Cosmos-Transfer2.5, from October 2025, has 2B parameters and is built on Predict2.5, with a low-latency distilled version following in February 2026. Its main use is data augmentation: turning simulator-rendered robot or driving footage into photorealistic video to narrow the sim-to-real gap, or turning a piece of real footage into rainy or nighttime versions to expand training data.","example":"Render a robot-arm grasping video in Isaac Sim, extract its depth and segmentation maps, and feed them to Cosmos Transfer; swapping in different text prompts yields multiple photorealistic versions with different backgrounds and lighting, used to train a policy more robust to visual distractions.","related":["NVIDIA Cosmos","Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Generative Data Augmentation","Synthetic Data","NVIDIA Cosmos Predict"]},{"id":"nvidia-cosmos-reason","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning (arXiv 2503.15558)","url":"https://arxiv.org/abs/2503.15558"},{"title":"nvidia-cosmos/cosmos-reason2 GitHub 仓库","url":"https://github.com/nvidia-cosmos/cosmos-reason2"},{"title":"nvidia-cosmos/cosmos-reason1 GitHub 仓库","url":"https://github.com/nvidia-cosmos/cosmos-reason1"}],"as_of":"2026-04","related_ids":["vision-language-model","embodied-reasoning-model","chain-of-thought","nvidia-cosmos","qwen-vl","nvidia-alpamayo"],"name":"NVIDIA Cosmos Reason","alt":"Cosmos Reason","abbr":"","aliases":["Cosmos-Reason1","Cosmos-Reason2"],"one_liner":"NVIDIA's reasoning vision-language model for physical AI, which watches a video, reasons through it in writing, then answers.","explanation":"Cosmos Reason is the vision-language model in NVIDIA's Cosmos family responsible for understanding and reasoning. Cosmos-Reason1, from March 2025, trains in two steps: supervised fine-tuning on physical-AI data first, then reinforcement learning, so the model writes out a chain of thought before answering; the paper covers 7B and 56B sizes, with the open-sourced version being the 7B built on Qwen2.5-VL. It focuses on physical common sense (space, time, basic physics) and embodied reasoning (what a robot should do next). Cosmos-Reason2, from December 2025, switched to Qwen3-VL, offered at 2B and 8B, with a 32B added in April 2026. Common uses include flagging physical errors in generated video, filtering and annotating training data, and serving as a robot's planning module; Cosmos-Predict2.5 uses it as a text encoder, and the driving model Alpamayo also uses it as its backbone.","example":"Given an AI-generated robot manipulation video and asked “does this footage obey physics,” it first outputs its reasoning process (such as whether an object moves without cause, or clips through another object), then gives its judgment, which can be used to automatically filter out unqualified synthetic data.","related":["Vision-Language Model","Embodied Reasoning Model","Chain-of-Thought","NVIDIA Cosmos","Qwen-VL","NVIDIA Alpamayo"]},{"id":"cosmos-policy","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Cosmos Policy (arXiv:2601.16163)","url":"https://arxiv.org/abs/2601.16163"},{"title":"Cosmos Policy 项目主页（NVIDIA Research）","url":"https://research.nvidia.com/labs/cosmos-lab/cosmos-policy/"}],"as_of":"2026-01","related_ids":["nvidia-cosmos-predict","video-generation-model","world-action-model","value-function","openvla-oft","libero-benchmark"],"name":"Cosmos Policy","alt":"Cosmos Policy","abbr":"","aliases":["Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning"],"one_liner":"Fine-tunes a video generation model directly into a robot policy that predicts actions, future frames, and success confidence together.","explanation":"Cosmos Policy was proposed by NVIDIA and Stanford in January 2026, first author Moo Jin Kim (also first author of OpenVLA-OFT), and accepted at ICLR 2026. It starts from the video generation model Cosmos-Predict2-2B and, without changing its architecture or adding an action head, post-trains it just once on robot demonstration data: proprioceptive state, action chunks, and value (expected return) are all encoded as “latent frames” and inserted into the video diffusion model's latent sequence, denoised together with the image frames. This way, one pass gives three things at once: the action to execute, the future frames that will result, and how confident the model is that this step succeeds. At inference, it can execute the action directly, or sample several candidates and use the predicted future and value to pick the best one (test-time planning). It reaches 98.5% average success on LIBERO and 67.1% on RoboCasa (50 demonstrations per task), representative of the “fine-tune a video model directly into a policy” approach.","example":"On real ALOHA bimanual tasks such as putting candy into a sealed bag, in planning mode Cosmos Policy first samples 8 candidate action chunks, then scores them using its own predicted future frames and value, and executes whichever one it's most confident in.","related":["NVIDIA Cosmos Predict","Video Generation Model","World Action Model","Value Function","OpenVLA-OFT","LIBERO Benchmark"]},{"id":"cosmos-3","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Cosmos 3: Omnimodal World Models for Physical AI (arXiv:2606.02800)","url":"https://arxiv.org/abs/2606.02800"},{"title":"NVIDIA Cosmos GitHub（Cosmos 3 模型家族与发布记录）","url":"https://github.com/nvidia/cosmos"}],"as_of":"2026-07","related_ids":["nvidia-cosmos","world-foundation-model","mixture-of-transformers","unified-multimodal-model","world-action-model","cosmos-policy"],"name":"Cosmos 3","alt":"Cosmos 3","abbr":"","aliases":["Cosmos3","NVIDIA Cosmos 3","Cosmos 3: Omnimodal World Models for Physical AI"],"one_liner":"NVIDIA's third-generation omnimodal world model: a single model that understands video, generates it, and also outputs actions.","explanation":"This is the third generation of NVIDIA's Cosmos world foundation model family, released in May 2026, with the technical report posted to arXiv in June 2026. The first two generations split understanding (Cosmos Reason) and video generation (Cosmos Predict) into separate models; Cosmos 3 merges them into one using a mixture-of-transformers (MoT) architecture: an autoregressive language tower reads text and reasons, a diffusion generation tower outputs images, video, audio, and action sequences, and the two towers share attention. The same set of weights can act as a vision-language model, a text-to-video or image-to-video model, a world simulator that predicts future frames from actions, and can also be post-trained into a robot policy. It's released in Super (64B) and Nano (16B) sizes, with an Edge (4B) version for Jetson edge devices added in July 2026; code and weights use the OpenMDW-1.1 license.","example":"According to the technical report, Cosmos3-Nano-Policy, post-trained on DROID data, ranked first on the RoboArena real-robot leaderboard (a snapshot from May 30, 2026) with an Elo of 1870.","related":["NVIDIA Cosmos","World Foundation Model","Mixture-of-Transformers","Unified Multimodal Model","World Action Model","Cosmos Policy"]},{"id":"dreamgen","category":"named_model","sec":6,"tier":3,"sources":[{"title":"DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv 2505.12705)","url":"https://arxiv.org/abs/2505.12705"},{"title":"DreamGen 项目主页（NVIDIA GEAR）","url":"https://research.nvidia.com/labs/gear/dreamgen/"},{"title":"NVIDIA Computex 2025 新闻稿：GR00T N1.5 与 GR00T-Dreams","url":"https://nvidianews.nvidia.com/news/nvidia-powers-humanoid-robot-industry-with-cloud-to-robot-computing-platforms-for-physical-ai"}],"as_of":"2025-05","related_ids":["neural-trajectories","video-generation-model","inverse-dynamics-model","latent-action-model","nvidia-isaac-gr00t-n1","nvidia-cosmos-predict"],"name":"DreamGen","alt":"DreamGen（GR00T Dreams）","abbr":"","aliases":["GR00T Dreams","GR00T-Dreams","Isaac GR00T-Dreams","Unlocking Generalization in Robot Learning through Video World Models"],"one_liner":"An NVIDIA 2025 method that uses a video world model to generate robot videos and infer actions from them to synthesize training data.","explanation":"DreamGen is a synthetic-data pipeline proposed by NVIDIA's GEAR Lab and others in May 2025; its open-source implementation is called the Isaac GR00T-Dreams blueprint, announced that same month at Computex in Taipei. Teaching a robot a new skill usually requires a human to teleoperate it to collect large amounts of data. DreamGen works in four steps: first fine-tune an image-to-video generation model on data from the target robot (the open-source implementation uses Cosmos-Predict2); then, given a single starting frame and a language instruction, have the model generate a realistic video of the robot performing the new action in a new environment; next, use an inverse dynamics model or a latent action model to infer pseudo action labels from that video; and finally train a visuomotor policy on these “neural trajectories” together with real data. The paper also proposes DreamGen Bench, finding that higher video-generation quality leads to a better-trained policy. NVIDIA says that using GR00T-Dreams, it generated training data for GR00T N1.5 in 36 hours, versus roughly three months for manual collection.","example":"Using only teleoperation data from a single pick-and-place task in one environment, plus videos generated by DreamGen, the GR1 humanoid robot learned 22 new behaviors and could perform them in 10 environments it had never seen.","related":["Neural Trajectories","Video Generation Model","Inverse Dynamics Model","Latent Action Model","NVIDIA Isaac GR00T N1","NVIDIA Cosmos Predict"]},{"id":"dreamdojo","category":"named_model","sec":6,"tier":3,"sources":[{"title":"DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos (arXiv 2602.06949)","url":"https://arxiv.org/abs/2602.06949"},{"title":"DreamDojo 项目主页","url":"https://dreamdojo-world.github.io/"}],"as_of":"2026-02","related_ids":["world-model","latent-action","egocentric-video","egodex","nvidia-cosmos-predict","world-model-based-policy-evaluation"],"name":"DreamDojo","alt":"DreamDojo","abbr":"","aliases":["A Generalist Robot World Model from Large-Scale Human Videos"],"one_liner":"A generalist robot world model NVIDIA released in 2026, pretrained on about 44,000 hours of human video.","explanation":"DreamDojo is a generalist robot world model — a model that predicts future video frames conditioned on actions — released in February 2026, led by NVIDIA together with UC Berkeley, HKUST (Hong Kong University of Science and Technology), Stanford, and others. Robot data is scarce and expensive to collect, while videos of humans doing everyday things are abundant but carry no action labels. DreamDojo assembles about 44,700 hours of first-person human video (mainly its own crowdsourced dataset, DreamDojo-HV, about 43,800 hours, plus EgoDex and a small amount of lab data), which the paper describes as the largest video collection used to pretrain a world model at the time. It uses continuous latent actions learned from the video itself as a unified “proxy action.” The model first learns how objects behave under interaction from unlabeled video, and is then post-trained on robot data to turn it into a simulator that accepts real action commands. The model is built on Cosmos-Predict2.5, comes in 2B and 14B versions, and after distillation generates in real time at about 10.8 frames per second on an H100, making it usable for policy evaluation, real-time teleoperation, and model-based planning.","example":"After post-training, it was adapted to robots including Fourier's GR-1, Unitree's G1, AgiBot, and the YAM arm: given a current frame and a sequence of actions, the model generates in real time what the video would look like after executing them, letting a policy be evaluated without running it on a real robot.","related":["World Model","Latent Action","Egocentric Video","EgoDex","NVIDIA Cosmos Predict","World-Model-based Policy Evaluation"]},{"id":"dreamzero","category":"named_model","sec":6,"tier":3,"sources":[{"title":"World Action Models are Zero-shot Policies (arXiv 2602.15922)","url":"https://arxiv.org/abs/2602.15922"},{"title":"DreamZero 项目主页","url":"https://dreamzero0.github.io/"}],"as_of":"2026-02","related_ids":["world-action-model","video-generation-model","wan","zero-shot","cross-embodiment","fast-wam"],"name":"DreamZero","alt":"DreamZero","abbr":"","aliases":["World Action Models are Zero-shot Policies"],"one_liner":"NVIDIA's 2026 world action model, built on a 14B video diffusion model that predicts frames and actions together.","explanation":"DreamZero is a world action model (WAM, a model that jointly predicts future video frames and actions) released in February 2026 by an NVIDIA research team including Jim Fan and Yuke Zhu, in a paper titled World Action Models are Zero-shot Policies. Mainstream VLA models use a vision-language model as their backbone, which is good at understanding semantics but has limited grasp of physical dynamics. DreamZero instead uses a pretrained video diffusion model as its backbone (Tongyi Wanxiang's Wan2.1 image-to-video model, 14B parameters), autoregressively generating future video and robot actions together, treating video as dense supervision for “how the world changes” — which lets it learn from diverse, non-repetitive, heterogeneous robot data. The paper reports that its generalization to new tasks and new environments is more than double that of the best VLA models at the time; after model and system optimization, a single inference takes about 150 milliseconds on a GB200, enabling 7Hz closed-loop control. The team says it will open-source the model weights and inference code.","example":"Trained mainly on about 500 hours of teleoperation data from the AgiBot G1, it reaches an average task-progress score of 62.2% on tasks seen during training, versus 27.4% for the best VLA baseline; adding just 10–20 minutes of video from other robots or humans lifts performance on unseen tasks by more than 42% relative, and just 30 minutes of play data is enough to adapt it to the new YAM robot.","related":["World Action Model","Video Generation Model","Wan (Alibaba Video Generation Model)","Zero-shot","Cross-Embodiment","Fast-WAM"]},{"id":"egoscale","category":"named_model","sec":6,"tier":3,"sources":[{"title":"EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (arXiv 2602.16710)","url":"https://arxiv.org/abs/2602.16710"},{"title":"EgoScale 项目页（NVIDIA GEAR）","url":"https://research.nvidia.com/labs/gear/egoscale/"}],"as_of":"2026-02","related_ids":["egocentric-video","pretraining-on-human-videos","dexterous-manipulation","scaling-law","mid-training","nvidia-isaac-gr00t-n1"],"name":"EgoScale","alt":"EgoScale","abbr":"","aliases":["Scaling Dexterous Manipulation with Diverse Egocentric Human Data"],"one_liner":"An NVIDIA 2026 project that pretrains a dexterous-hand VLA on 20,000 hours of first-person human video.","explanation":"EgoScale is a paper released in February 2026 by NVIDIA's GEAR Lab together with UC Berkeley and the University of Maryland, aimed at answering whether human video can teach a robot to work with a multi-fingered dexterous hand at scale. The method has three steps: first, extract wrist motion from 20,854 hours of action-annotated, first-person human video, and retarget human hand poses into the joint space of a 22-degree-of-freedom Sharpa dexterous hand, to pretrain a flow-matching VLA structurally similar to GR00T N1; next, mid-train on a small amount of aligned data where a human and a robot perform the same action in the same scene; finally, post-train on specific tasks. The paper finds a log-linear scaling relationship between the amount of human data and validation loss, with average success rate improving 54% compared to skipping human pretraining.","example":"On a Galaxea R1 Pro fitted with the 22-DOF Sharpa dexterous hand, given just 1 robot demonstration, average success on a shirt-folding task reached as high as 88%; switching to a three-fingered Unitree G1 hand, human pretraining still delivered more than a 30-percentage-point absolute improvement.","related":["Egocentric Video","Pretraining on Human Videos","Dexterous Manipulation","Scaling Law","Mid-training","NVIDIA Isaac GR00T N1"]},{"id":"figure-helix","category":"named_model","sec":6,"tier":1,"sources":[{"title":"Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure AI)","url":"https://www.figure.ai/news/helix"}],"as_of":"2025-02","related_ids":["dual-system-architecture","vision-language-action-model","figure-ai","figure-helix-02","figure-02"],"name":"Figure Helix","alt":"Helix","abbr":"","aliases":["Helix","Helix VLA"],"one_liner":"Figure AI's 2025 vision-language-action model for its humanoid robot, using a fast-slow two-system architecture to control the whole upper body.","explanation":"Helix is a vision-language-action model (VLA — a model that looks at images, listens to instructions, and outputs actions directly) released on February 20, 2025, by the US humanoid-robot company Figure AI for use on its Figure humanoid robots. It uses a fast-slow, two-system design: System 2 is a 7-billion-parameter vision-language model that understands the scene and the language instruction at 7–9Hz and compresses the intent into a latent vector; System 1 is an 80-million-parameter Transformer that turns that latent vector, together with real-time perception, into actions at 200Hz. It controls the entire upper body at once — 35 degrees of freedom, including the wrists, every finger, the torso, and head orientation. Figure says training used only about 500 hours of teleoperation data, that a single set of weights covers many tasks without task-specific fine-tuning, and that the two systems each run on one of two low-power embedded GPUs onboard the robot. Later versions include Helix 02.","example":"In the launch demo, two Figure robots collaborated on putting away a bag of groceries neither had seen before, following spoken instructions and sorting items into the fridge and drawers.","related":["Dual-System Architecture (System 1 / System 2)","Vision-Language-Action Model","Figure AI","Figure Helix 02","Figure 02"]},{"id":"figure-helix-02","category":"named_model","sec":6,"tier":2,"sources":[{"title":"Helix 02: Full-Body Autonomy (Figure AI)","url":"https://www.figure.ai/news/helix-02"}],"as_of":"2026-01","related_ids":["figure-helix","figure-ai","dual-system-architecture","system-0","loco-manipulation","whole-body-control"],"name":"Figure Helix 02","alt":"Helix 02","abbr":"","aliases":["Helix-02","Helix 02 (Full-Body Autonomy)"],"one_liner":"Figure's 2026 full-body humanoid control model, mapping raw pixels directly to whole-body control for long-horizon tasks.","explanation":"Figure AI released this in the US on January 27, 2026, as the successor to February 2025's Helix. The original Helix used a fast-slow two-system design to control the upper body; Helix 02 extends this to the whole robot, coordinating walking and manipulation together, and restructures it into three layers. System 2 understands the scene and language and sets behavioral goals; System 1 runs at 200Hz, mapping all sensor input directly into whole-body joint commands; a new System 0 runs at 1kHz with about 10 million parameters, trained on more than 1,000 hours of human motion data, and handles balance, contact, and whole-body coordination. On the hardware side, it adds palm cameras and fingertip touch sensing that can detect about 3 grams of force. In an official demo, it performed a 4-minute, fully autonomous, end-to-end task around a dishwasher made up of 61 separate loco-manipulation actions.","example":"Taking a pill out of a blister pack, precisely dispensing 5 milliliters of liquid with a syringe, and unscrewing a bottle cap are all dexterous tasks Figure showed Helix 02 performing.","related":["Figure Helix","Figure AI","Dual-System Architecture (System 1 / System 2)","System 0","Loco-manipulation","Whole-Body Control"]},{"id":"helix-2-5","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Helix 2.5: Zero-Shot 30-Home Generalization (Figure AI, 2026-09-17)","url":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"},{"title":"Figure AI 新闻列表","url":"https://www.figure.ai/news"}],"as_of":"2026-09","related_ids":["figure-ai","figure-helix","figure-helix-02","figure-03","human-video-data","zero-shot"],"name":"Helix 2.5","alt":"Helix 2.5","abbr":"","aliases":["Figure Helix 2.5"],"one_liner":"A humanoid robot model Figure AI released in September 2026 that does household chores zero-shot in 30 unfamiliar homes.","explanation":"Helix 2.5 is a robot model released on September 17, 2026 by the US humanoid robot company Figure AI, running on the Figure 03 humanoid, and the successor to Helix and Helix 02. Unlike its predecessors, which started from a vision-language model, the company says Helix 2.5 is pretrained entirely on Index, Figure's self-built, global-scale dataset of human behavior, which it says gains about 35 minutes of new human experience every second. The release's headline result is zero-shot generalization: across 30 San Francisco Bay Area homes the robot had never visited, with no data collection or fine-tuning and a single fixed set of model weights per task, it performed three household chores that require moving and manipulating at the same time — tidying a living room, folding towels, and making a bed. In the company's own comparison, a policy trained from random initialization reached 9% success, while Helix 2.5 starting from Index pretraining reached 56%; and matching Helix 02's success rate required only half as much task-specific data.","example":"Figure 03 walks into an unfamiliar home and, using the same model weights, puts scattered items in the living room back in place, folds towels, and makes the bed — with no data ever collected in that home beforehand.","related":["Figure AI","Figure Helix","Figure Helix 02","Figure 03","Human Video Data","Zero-shot"]},{"id":"1x-world-model","category":"named_model","sec":6,"tier":3,"sources":[{"title":"1X World Model (1X, 2024-09)","url":"https://www.1x.tech/discover/1x-world-model"},{"title":"1X World Model: evaluating Redwood AI (1X, 2025-06)","url":"https://www.1x.tech/discover/redwood-ai-world-model"},{"title":"1X World Model | From Video to Action: A New Way Robots Learn (1X, 2026-01)","url":"https://www.1x.tech/discover/world-model-self-learning"}],"as_of":"2026-01","related_ids":["world-model","world-model-based-policy-evaluation","inverse-dynamics-model","video-generation-model","1x-neo","redwood"],"name":"1X World Model","alt":"1X 世界模型","abbr":"1XWM","aliases":["1X World Model Challenge"],"one_liner":"Humanoid company 1X's video world model, first used to evaluate policies, later used directly to control NEO.","explanation":"The 1X World Model is a video-generation world model developed by the humanoid robotics company 1X. It was first announced in September 2024, trained on thousands of hours of video and action data collected by EVE robots in homes and offices: given the current image and a sequence of actions, it predicts the resulting video, and can simulate falling objects, cloth, doors, and drawers; 1X also released more than 100 hours of data and launched the 1X World Model Challenge. From June 2025 it was used to evaluate NEO's Redwood policy — running rollouts inside the model to compare policies and checkpoints while reducing real-robot testing. In January 2026, 1X connected it directly to NEO as a policy: a 14-billion-parameter text-conditioned video diffusion model first generates a video of the task being completed, and an inverse-dynamics model then infers actions from that footage; training used web video, 900 hours of first-person human video, and 70 hours of NEO data. Generating a 5-second video currently takes about 11 seconds, and dexterous tasks like pouring water remain difficult.","example":"Given the phrase “pull out a tissue” and the current image, 1XWM generates a video of NEO completing the action, and the inverse-dynamics model derives the actions for NEO to execute from it; generating several candidate videos in parallel and picking the best one raised this task's success rate from 30% to 45%.","related":["World Model","World-Model-based Policy Evaluation","Inverse Dynamics Model","Video Generation Model","1X NEO","Redwood"]},{"id":"redwood","category":"named_model","sec":6,"tier":3,"sources":[{"title":"1X: Redwood AI","url":"https://www.1x.tech/discover/redwood-ai"}],"as_of":"2025-06","related_ids":["1x-world-model","1x-neo","vision-language-action-model","mobile-manipulation","on-device-edge-deployment","whole-body-control"],"name":"Redwood","alt":"1X Redwood","abbr":"","aliases":["Redwood AI","1X Redwood"],"one_liner":"1X's vision-language-action model for its home humanoid NEO, running entirely on the robot's onboard GPU.","explanation":"Redwood was released by humanoid robot company 1X Technologies in June 2025. It is a roughly 160-million-parameter vision-language-action (VLA) model — it looks at the camera feed, listens to instructions, and outputs actions directly — running entirely on the onboard embedded GPU of the NEO Gamma robot, producing actions at about 5 Hz, and continuing to work even with a poor network connection. It targets mobile manipulation around the home, such as picking things up while walking or opening doors, and can coordinate walking, arm movement, and pelvis posture commands at the same time, enabling behaviors like bending down to pick something up or bracing against a wall with the free hand while pulling open a heavy door. Training data comes from teleoperated and autonomous episodes recorded on 1X's two robot generations, EVE and NEO, in offices and employees' homes, with failure episodes included in training as well. Users can give instructions through voice conversation.","example":"A user tells NEO, 'bring me the cup on the kitchen counter,' and Redwood controls it to walk to the kitchen, pick up the cup, and bring it back.","related":["1X World Model","1X NEO","Vision-Language-Action Model","Mobile Manipulation","On-Device / Edge Deployment","Whole-Body Control"]},{"id":"gemini-robotics","category":"named_model","sec":6,"tier":1,"sources":[{"title":"Gemini Robotics brings AI into the physical world (Google DeepMind blog)","url":"https://deepmind.google/discover/blog/gemini-robotics-brings-ai-into-the-physical-world/"},{"title":"Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)","url":"https://arxiv.org/abs/2503.20020"}],"as_of":"2025-03","related_ids":["gemini-robotics-er","gemini-robotics-1-5","gemini-robotics-on-device","vision-language-action-model","aloha-2","google-deepmind"],"name":"Gemini Robotics","alt":"Gemini Robotics","abbr":"","aliases":["Gemini Robotics 1.0"],"one_liner":"Google DeepMind's 2025 VLA built on Gemini 2.0, able to directly control robots for dexterous manipulation.","explanation":"Gemini Robotics is a vision-language-action model released by Google DeepMind on March 12, 2025, built on top of the Gemini 2.0 multimodal model. Released alongside it, Gemini Robotics-ER focuses on spatial understanding and embodied reasoning, such as object detection and predicting trajectories and grasps. The goal is to bring the general understanding of large models into the physical world, and Google highlighted three properties: generalization (handling new objects, scenes, and instructions, scoring more than twice as well as other leading VLAs at the time on a general-generalization benchmark), interactivity (understanding conversational instructions and adjusting in real time as the environment or instructions change), and dexterity (such as folding origami or packing a snack into a zip-lock bag). It was trained mainly on the bimanual ALOHA 2 platform, but also adapts to Franka arms and Apptronik's Apollo humanoid robot; the technical report says a new task can be fine-tuned with around 100 demonstrations. Later versions include On-Device and 1.5.","example":"In official demos, an ALOHA 2 bimanual robot running Gemini Robotics folds origami and packs a snack into a zip-lock bag on spoken command; if someone moves the target object mid-task, the robot adjusts its motion and keeps going.","related":["Gemini Robotics-ER","Gemini Robotics 1.5","Gemini Robotics On-Device","Vision-Language-Action Model","ALOHA 2","Google DeepMind"]},{"id":"gemini-robotics-er","category":"named_model","sec":6,"tier":2,"sources":[{"title":"Gemini Robotics: Bringing AI into the Physical World (arXiv 2503.20020)","url":"https://arxiv.org/abs/2503.20020"},{"title":"Gemini Robotics ER - Google DeepMind","url":"https://deepmind.google/models/gemini-robotics/gemini-robotics-er/"},{"title":"Gemini Robotics ER 2 (Google blog)","url":"https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/"}],"as_of":"2026-07","related_ids":["embodied-reasoning","embodied-reasoning-model","gemini-robotics","gemini-robotics-1-5","erqa","braincerebellum-architecture"],"name":"Gemini Robotics-ER","alt":"Gemini Robotics-ER","abbr":"","aliases":["Gemini Robotics-ER 1.5","Gemini Robotics-ER 1.6","Gemini Robotics ER 2"],"one_liner":"Google DeepMind's embodied-reasoning model, which understands spatial layouts, makes plans, and judges whether a robot task is done.","explanation":"ER stands for Embodied Reasoning. Google DeepMind introduced it in March 2025 alongside the original Gemini Robotics: a vision-language model built on Gemini with strengthened 3D perception and spatial understanding, able to point out object locations and predict grasp points and trajectories. It generally doesn't output joint-level actions directly; instead it answers questions like where something is, which step comes first, and whether a step is complete, then hands the plan to a VLA or robot interface to execute — functioning as the “slow” system in a fast-slow dual-system setup. ER 1.5, from September 2025, is available through the Gemini API, later updated to 1.6; ER 2, from July 2026, adds video understanding and progress tracking, and can orchestrate action models through multistep tasks, coordinate multiple robots, and call tools such as Google Search.","example":"ER 2 watches a video of a robot at work, judges whether the current step is finished, has the robot retry if not, and moves on to the next step once it is.","related":["Embodied Reasoning","Embodied Reasoning Model","Gemini Robotics","Gemini Robotics 1.5","ERQA","Brain–Cerebellum Architecture"]},{"id":"gemini-robotics-on-device","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Gemini Robotics On-Device brings AI to local robotic devices (Google DeepMind 博客, 2025-06-24)","url":"https://deepmind.google/discover/blog/gemini-robotics-on-device-brings-ai-to-local-robotic-devices/"},{"title":"Gemini Robotics On-Device 2 模型页","url":"https://deepmind.google/models/gemini-robotics/gemini-robotics-on-device/"},{"title":"Gemini Robotics 模型总览","url":"https://deepmind.google/models/gemini-robotics/"}],"as_of":"2026-09","related_ids":["gemini-robotics","gemini-robotics-2","on-device-model","gemini-robotics-sdk","aloha","apptronik-apollo"],"name":"Gemini Robotics On-Device","alt":"Gemini Robotics On-Device","abbr":"","aliases":["Gemini Robotics On-Device 2"],"one_liner":"Google DeepMind's lightweight Gemini Robotics VLA that runs directly on the robot itself, with no internet connection required.","explanation":"Gemini Robotics On-Device is a vision-language-action model (VLA, a model that takes in images and spoken instructions and outputs robot actions directly) released by Google DeepMind on June 24, 2025, a stripped-down version of Gemini Robotics specifically optimized to run on a robot's own onboard compute with no cloud connection needed — suited to poor-connectivity or latency-sensitive settings. It targets bimanual manipulation mainly, capable of dexterous tasks like unzipping a zipper or folding clothes, and supports fine-tuning to a new task with just 50 to 100 demonstrations. The model was trained on the ALOHA bimanual platform, and official demos showed it transferred to a dual-arm Franka FR3 setup and to Apptronik's Apollo humanoid. It was made available, together with the Gemini Robotics SDK, through a trusted-tester program. DeepMind later released Gemini Robotics On-Device 2, which it says can adapt to a new robot embodiment with fewer than 200 examples, also currently available through the trusted-tester program.","example":"A developer collects 50–100 demonstrations for a new task, such as folding a new type of garment, and can fine-tune the On-Device model to that task; inference then runs entirely on the robot itself, so it keeps working even with no network connection.","related":["Gemini Robotics","Gemini Robotics 2","On-device Model","Gemini Robotics SDK","ALOHA","Apptronik Apollo"]},{"id":"gemini-robotics-1-5","category":"named_model","sec":6,"tier":2,"sources":[{"title":"Gemini Robotics 1.5 brings AI agents into the physical world (Google DeepMind)","url":"https://deepmind.google/discover/blog/gemini-robotics-15-brings-ai-agents-into-the-physical-world/"}],"as_of":"2025-09","related_ids":["gemini-robotics","gemini-robotics-er","vision-language-action-model","cross-embodiment","embodied-chain-of-thought","apptronik-apollo"],"name":"Gemini Robotics 1.5","alt":"Gemini Robotics 1.5","abbr":"","aliases":[],"one_liner":"Google DeepMind's VLA that thinks in words before acting, with skills that transfer across different kinds of robots.","explanation":"Google DeepMind released this on September 25, 2025, as an upgraded vision-language-action model (VLA) building on the original Gemini Robotics from March of that year. It adds two new capabilities. First, think-before-acting: before producing actions, it generates an internal chain of reasoning in natural language, breaking a semantically complex instruction down into simple steps. Second, Motion Transfer: a task trained only on the bimanual ALOHA 2 platform can run directly on the Apptronik Apollo humanoid or a dual-arm Franka setup, and vice versa — reusing skills across different robot embodiments. It works together with Gemini Robotics-ER 1.5, an embodied-reasoning model released at the same time: ER 1.5 handles high-level planning, directing 1.5 step by step in natural language. At launch, 1.5 was available only to select partners.","example":"A task taught only on the ALOHA 2 bimanual platform can be carried out on the Apollo humanoid robot with no additional training.","related":["Gemini Robotics","Gemini Robotics-ER","Vision-Language-Action Model","Cross-Embodiment","Embodied Chain-of-Thought","Apptronik Apollo"]},{"id":"gemini-robotics-2","category":"named_model","sec":6,"tier":2,"sources":[{"title":"Gemini Robotics 2 brings whole body intelligence to robots (Google DeepMind)","url":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"},{"title":"Gemini Robotics 模型页 (Google DeepMind)","url":"https://deepmind.google/models/gemini-robotics/"}],"as_of":"2026-07","related_ids":["gemini-robotics-1-5","gemini-robotics-er","gemini-robotics-on-device","whole-body-control","apptronik-apollo","sharpawave"],"name":"Gemini Robotics 2","alt":"Gemini Robotics 2","abbr":"","aliases":[],"one_liner":"Google DeepMind's 2026 robotics model generation, with the new ability to control a humanoid's whole body.","explanation":"Google DeepMind released this on July 30, 2026, as the generation following 1.5, made up of three models: the VLA model Gemini Robotics 2, the embodied-reasoning model Gemini Robotics ER 2, and an On-Device 2 version that runs locally on the robot. The main new capability is full-body control: it can drive an entire humanoid robot while it walks, crouches, and manipulates objects at the same time, and Google demonstrated multistep tasks such as retrieving items from a shelf on the Apptronik Apollo 2. It can also control the 22-degree-of-freedom, five-fingered SharpaWave dexterous hand, as well as ordinary two-finger grippers; adapting On-Device 2 to a new dual-arm embodiment takes only a few hours. The VLA and On-Device versions are currently available only to early partners, while ER 2 is accessible through Google AI Studio.","example":"Under Gemini Robotics 2's control, an Apptronik Apollo 2 humanoid walks to a shelf, crouches, and retrieves an item, completing a multistep task.","related":["Gemini Robotics 1.5","Gemini Robotics-ER","Gemini Robotics On-Device","Whole-Body Control","Apptronik Apollo","SharpaWave"]},{"id":"veo-world-simulator","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Evaluating Gemini Robotics Policies in a Veo World Simulator (arXiv 2512.10675)","url":"https://arxiv.org/abs/2512.10675"},{"title":"论文 HTML 版（Veo 2、ALOHA 2 与相关性结果）","url":"https://arxiv.org/html/2512.10675v2"}],"as_of":"2026-01","related_ids":["world-model-based-policy-evaluation","gemini-robotics","video-generation-model","interactive-world-model","sim-to-real-correlation","mean-maximum-rank-violation"],"name":"Veo World Simulator","alt":"Veo 世界模拟器","abbr":"","aliases":["Evaluating Gemini Robotics Policies in a Veo World Simulator"],"one_liner":"A Google DeepMind system that uses the Veo video model to 'imagine' a robot's execution in order to evaluate policies.","explanation":"The Veo World Simulator is a technical report released by Google DeepMind's Gemini Robotics team in December 2025. Real-robot evaluation is slow, expensive, and hard to extend to new objects or backgrounds not seen before. The team adapted the Veo 2 video generation model into a world simulator: given the current image and a sequence of future robot poses, it generates the resulting frames from all four camera views of an ALOHA 2 bimanual platform, and uses image editing plus multi-view inpainting to alter the real scene with new objects, backgrounds, and distractors. This makes it possible to run policies inside 'generated video,' predict the relative strengths and weaknesses of different policy versions, compare which kinds of generalization are harder, and even conduct red-teaming to surface unsafe behaviors. The authors validated the agreement between these predictions and real-robot results using 8 versions of Gemini Robotics policies, 5 tasks, and more than 1,600 real-robot trials.","example":"In one red-team test, the simulator generated a scene where the instruction 'quick, grab the red block!' exposed a policy bumping into a nearby hand; in another, a policy closed a laptop before moving a pair of scissors out of the way, risking damage to the screen.","related":["World-Model-based Policy Evaluation","Gemini Robotics","Video Generation Model","Interactive World Model","Sim-to-Real Correlation","Mean Maximum Rank Violation"]},{"id":"rho-alpha","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Microsoft Research: Advancing AI for the physical world","url":"https://www.microsoft.com/en-us/research/story/advancing-ai-for-the-physical-world/"}],"as_of":"2026-01","related_ids":["vision-language-action-model","vision-tactile-language-action-model","tactile-sensor","bimanual-manipulation","human-in-the-loop","microsoft-research"],"name":"Rho-alpha","alt":"微软 Rho-alpha","abbr":"ρα","aliases":["ρα","Rho-Alpha","Microsoft Research robotics model derived from Phi"],"one_liner":"Microsoft's first robotics model, built from its Phi vision-language models and adding touch sensing.","explanation":"Rho-alpha (written ρα) was released by Microsoft Research on January 21, 2026, Microsoft's first robotics model, derived from its Phi family of vision-language models. Microsoft calls it 'VLA+': on top of the usual vision-language-action model setup — looking at the scene, listening to an instruction, and outputting actions directly — it adds touch to perception, with force sensing still under development, and is exploring how a robot can keep adapting during use based on a person's corrective feedback. It converts natural-language instructions into control signals for two-arm manipulation. Training data includes real-robot demonstrations, simulated tasks, and web-scale visual question-answering data. At launch it was evaluated on dual UR5e arms fitted with tactile sensors and on a humanoid robot, using the BusyBox benchmark. It launched as an early-access research program, with Microsoft saying it would later become available on Microsoft Foundry.","example":"An operator gives a single natural-language instruction, and a pair of UR5e arms fitted with tactile sensors carries out a two-handed manipulation task, using both vision and touch to judge whether the action has succeeded.","related":["Vision-Language-Action Model","Vision-Tactile-Language-Action Model","Tactile Sensor","Bimanual Manipulation","Human-in-the-Loop","Microsoft Research"]},{"id":"rfm-1","category":"named_model","sec":6,"tier":3,"sources":[{"title":"IEEE Spectrum: Covariant Announces a Universal AI Platform for Robots","url":"https://spectrum.ieee.org/covariant-foundation-model"},{"title":"TechCrunch: Covariant is building ChatGPT for robots","url":"https://techcrunch.com/2024/03/11/covariant-is-building-chatgpt-for-robots/"},{"title":"Wikipedia: Covariant (company)","url":"https://en.wikipedia.org/wiki/Covariant_(company)"}],"as_of":"2024-08","related_ids":["foundation-model","world-model","bin-picking","covariant","vacuum-suction-cup","amazon-frontier-ai-and-robotics"],"name":"RFM-1","alt":"Covariant RFM-1","abbr":"RFM-1","aliases":["Robotics Foundation Model 1","Covariant RFM-1"],"one_liner":"Covariant's roughly 8-billion-parameter robot foundation model, trained on warehouse picking data.","explanation":"RFM-1 is a robot foundation model released by warehouse-robotics AI company Covariant on March 11, 2024, with about 8 billion parameters. Its training data comes from Covariant's own picking arms deployed in warehouses across 15 countries, reportedly amounting to tens of millions of trajectories, including images, video, joint angles, and readings such as force and suction-cup vacuum level, plus text. Both its inputs and outputs can be text, images, video, or robot actions: it can understand natural-language instructions, and it can also predict what the scene will look like after a given action, effectively acting as a learned physics simulator to anticipate the consequences of an action. Its limitation is that it mainly covers vacuum-suction warehouse picking, with limited generalization to entirely new objects and situations. In August 2024, Amazon obtained a non-exclusive license to Covariant's technology and brought on founders Pieter Abbeel, Peter Chen, and Rocky Duan.","example":"An operator can tell a picking arm in natural language to 'pick up the red one'; RFM-1 can also generate a predicted image of the scene after executing a given grasp, letting the outcome be anticipated in advance.","related":["Foundation Model","World Model","Bin Picking","Covariant","Vacuum Suction Cup","Amazon Frontier AI & Robotics"]},{"id":"dyna-1","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Dynamism v1 (DYNA-1) Model: A Breakthrough in Performance and Production-Ready Embodied AI (Dyna Robotics, 2025-06)","url":"https://www.dyna.co/research/dyna-1"},{"title":"DYNA-1 Pre-Training: Zero-Shot Dexterity Is Here (Dyna Robotics, 2025-11)","url":"https://www.dyna.co/research/pre-training"}],"as_of":"2025-11","related_ids":["dyna-robotics","dyna-2","reward-model","real-world-reinforcement-learning","data-flywheel","embodied-foundation-model"],"name":"DYNA-1","alt":"DYNA-1","abbr":"","aliases":["Dyna-1","Dynamism v1","DYNA-1 Foundation Model"],"one_liner":"Dyna Robotics' first-generation robot foundation model, released in 2025, best known for folding napkins autonomously for hours at a stretch.","explanation":"DYNA-1 is the first embodied foundation model from the US startup Dyna Robotics, released in June 2025, with the official full name Dynamism v1. It is built for continuous operation in commercial settings: on a restaurant napkin-folding task, the company reports 24 hours of unattended autonomous operation with a 99.4% success rate, folding more than 850 napkins at about 60% of human speed. Its core approach is to put a self-developed reward model — a model that scores the robot's own performance — inside the training loop, letting the robot explore and self-correct, and continuing to generate new data during deployment. In November 2025, the company also said its pretrained base model, without any further fine-tuning, could fold clothes and sort packages in environments it had never seen, and could learn a new task from about an hour of demonstration. Architecture details have not been disclosed.","example":"According to the company's blog, in week 1 the robot could only fold napkins continuously for 5 minutes; by week 6 it could run non-stop for a full 24 hours. DYNA-1 was then deployed folding napkins at a paying customer's site.","related":["Dyna Robotics","DYNA-2","Reward Model","Real-World Reinforcement Learning","Data Flywheel","Embodied Foundation Model"]},{"id":"dyna-2","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models (Dyna Robotics, 2026-08)","url":"https://www.dyna.co/dyna-2"},{"title":"Not Just a Model, But a Product (Dyna Robotics, 2026-08)","url":"https://www.dyna.co/research/scaling-customer-deployments"},{"title":"Dyna Robotics 官网","url":"https://www.dyna.co/"}],"as_of":"2026-08","related_ids":["dyna-1","world-action-model","egocentric-video","pretraining-on-human-videos","scaling-law","dyna-robotics"],"name":"DYNA-2","alt":"Dyna DYNA-2","abbr":"","aliases":["Dyna-2"],"one_liner":"Dyna Robotics' 2026 world action model, pretrained on a million hours of first-person human video.","explanation":"DYNA-2 is Dyna Robotics' second-generation model, released in August 2026, a world action model (one that jointly predicts future video frames and future actions). It uses a video diffusion model as its backbone, with a hybrid Transformer that processes the video stream and the action stream separately, trained with flow matching so it can denoise future video and actions either jointly or separately. Its pretraining data is more than a million hours of head-mounted, first-person human video, covering everyday activities like cooking, tidying up, folding clothes, and assembly. The company reports observing a scaling law from human data to robot performance: as human pretraining data increases in stages from 1,000 hours to 1 million hours, the average normalized performance (relative to the achievable ceiling) across 14 real-robot tasks rises in step, to 20%, 28%, 45%, and 53%. In the company's layered system, DYNA-2 serves as System 1, taking direction from DYNA-VLM, which handles reasoning.","example":"According to the company, when folding napkins, DYNA-1 manages about 35 per hour at a 75% acceptable-quality rate, while DYNA-2 reaches 95 per hour at 93%; customer Din Tai Fung is now expanding Dyna's robots from a pilot restaurant to more of its locations and central kitchens.","related":["DYNA-1","World Action Model","Egocentric Video","Pretraining on Human Videos","Scaling Law","Dyna Robotics"]},{"id":"skild-brain","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Building the general-purpose robotic brain (Skild AI blog, 2025-07)","url":"https://www.skild.ai/blogs/building-the-general-purpose-robotic-brain"},{"title":"The case for an omni-bodied robot brain (Skild AI blog)","url":"https://www.skild.ai/blogs/omni-bodied"},{"title":"Announcing Series C (Skild AI blog, 2026-01)","url":"https://www.skild.ai/blogs/series-c"}],"as_of":"2026-04","related_ids":["skild-ai","cross-embodiment","embodiment-agnostic","skild-s1","hierarchical-architecture","sim-to-real-transfer"],"name":"Skild Brain","alt":"Skild Brain","abbr":"","aliases":["Omni-bodied Brain"],"one_liner":"Skild AI's general-purpose robot brain, with a single model able to control quadrupeds, humanoids, arms, and other embodiments.","explanation":"Skild Brain is a robot foundation model that American robotics company Skild AI made public in July 2025; the company was founded in 2023 by Carnegie Mellon University's Deepak Pathak and Abhinav Gupta. Skild Brain's central idea is 'omni-bodied': a single model can control quadrupeds, humanoids, tabletop arms, and mobile manipulation robots, without needing a separate design for each robot type. Structurally it has two layers: a low-frequency, high-level policy handles manipulation and navigation decisions, while a high-frequency, low-level policy converts these into joint angles and motor torques. Training relies mainly on large-scale simulation and internet video, followed by targeted real-robot data for post-training. The company states it trained on roughly 100,000 different simulated robot morphologies, so the model cannot simply memorize a solution for one specific body; in demonstrations, when a robot's lower leg is broken, its knee joint is locked, or a wheel gets stuck, it adjusts its gait within a few seconds and keeps moving. In January 2026, Skild AI raised a $1.4 billion Series C (led by SoftBank, valuing the company at more than $14 billion); in March it announced a partnership with ABB Robotics and Universal Robots to deploy Skild Brain in industrial settings, and in April it acquired Zebra Technologies' robotics business.","example":"With one leg of a quadruped robot locked at the knee — effectively leaving it only three usable legs — Skild Brain shifts its center of mass within two or three seconds and continues walking with a new gait.","related":["Skild AI","Cross-Embodiment","Embodiment-agnostic","Skild S1","Hierarchical Architecture","Sim-to-Real Transfer"]},{"id":"skild-s1","category":"named_model","sec":6,"tier":3,"sources":[{"title":"Introducing S1: In-Context Learning for Robotics (Skild AI blog)","url":"https://www.skild.ai/blogs/s1"},{"title":"Skild AI blog index","url":"https://www.skild.ai/blogs"}],"as_of":"2026-08","related_ids":["in-context-learning","skild-ai","skild-brain","human-video-data","universal-manipulation-interface","long-horizon-task"],"name":"Skild S1","alt":"Skild S1","abbr":"S1","aliases":["Skild AI S1"],"one_liner":"Skild AI's robot foundation model that can perform a task it never trained on after watching a single demonstration video.","explanation":"Skild S1 was released in August 2026 by American robotics company Skild AI, described by the company as its flagship robot foundation model. Its core idea is 'in-context learning': at deployment, the robot is shown a demonstration video, and the model infers the demonstrator's intent, the correspondence between objects, and task progress from it, then executes the task directly without updating any weights, even if that exact task never appeared during pretraining. Traditionally, a VLA facing a new task would need fresh teleoperated data and fine-tuning; S1 aims to skip that step. Its pretraining data mixes robot teleoperation, UMI handheld data collection, first-person human video, and simulation data. According to the company's blog, at a pretraining scale of 100,000 hours, it reaches 96% success on tasks seen during training and 66% on unseen tasks (versus 9% for a language-instruction baseline); one demonstration is said to be worth roughly 380 post-training data points, and it can handle long-horizon tasks lasting up to about 10 minutes. S1 is already used by commercial customers and is offered externally through early access.","example":"A staff member demonstrates repotting a plant once, and after watching the video, S1 goes on to complete this never-trained-on task; the company states that from setting up the scene to autonomous execution took just 11 minutes.","related":["In-Context Learning","Skild AI","Skild Brain","Human Video Data","Universal Manipulation Interface","Long-horizon Task"]},{"id":"field-foundation-models","category":"named_model","sec":6,"tier":3,"sources":[{"title":"FieldAI 官网","url":"https://www.fieldai.com/"},{"title":"Boston Dynamics and FieldAI Partner to Bring Robots Into Uncharted and Dynamic Environments","url":"https://www.fieldai.com/news/boston-dynamics-and-fieldai-partner-to-bring-robots-into-uncharted-and-dynamic-environments"},{"title":"Caterpillar and FieldAI Advance AI-Powered Industrial Innovation (PR Newswire, 2026-09-02)","url":"https://www.prnewswire.com/news-releases/caterpillar-and-fieldai-advance-ai-powered-industrial-innovation-302866862.html"}],"as_of":"2026-09","related_ids":["field-ai","foundation-model","world-model","uncertainty-estimation","boston-dynamics-spot","on-device-edge-deployment"],"name":"Field Foundation Models (Field AI)","alt":"Field AI FFM（场景基础模型）","abbr":"FFM","aliases":["FFM"],"one_liner":"Field AI's robot foundation model, built to operate autonomously in unmapped construction sites and industrial facilities.","explanation":"FFM (Field Foundation Models) is the core model line of Field AI, a US robot-software company headquartered in Irvine, California. Its founder and CEO is Ali Agha, and the team includes people from NASA's Jet Propulsion Laboratory and Google DeepMind, with experience on Mars exploration and DARPA programs. The company describes FFM as a “physics-first” embodied foundation model: it combines data-driven neural networks with physics-based reasoning and uncertainty estimation, built around a risk-assessing “belief world model” that lets a robot move and work autonomously in unstructured environments with no pre-built map, no GPS, and no preset path — running entirely on onboard compute. The same model can be installed on robots of different shapes from different manufacturers, which the company describes as “one brain for any machine.” According to its press releases, Field AI has raised more than $400 million in total funding, from investors including Bezos Expeditions, Khosla Ventures, and NVIDIA's venture arm NVentures.","example":"In March 2026, Field AI announced a partnership with Boston Dynamics to put FFM on the Spot quadruped, letting it move and work autonomously on construction sites that have no map and change day to day; that September, the company also announced a partnership with Caterpillar.","related":["Field AI","Foundation Model","World Model","Uncertainty Estimation","Boston Dynamics Spot","On-Device / Edge Deployment"]},{"id":"gen-0","category":"named_model","sec":6,"tier":2,"sources":[{"title":"GEN-0: Embodied Foundation Models That Scale with Physical Interaction (Generalist AI)","url":"https://generalistai.com/blog/gen-0"},{"title":"Generalist AI Blog","url":"https://generalistai.com/blog"}],"as_of":"2025-11","related_ids":["gen-1","generalist-ai","scaling-law","ossification","embodied-foundation-model","real-robot-data"],"name":"GEN-0","alt":"GEN-0","abbr":"","aliases":["GEN-0: Embodied Foundation Models That Scale with Physical Interaction","Generalist AI GEN-0"],"one_liner":"Generalist AI's embodied foundation model, trained on more than 270,000 hours of real-world interaction data from homes and workplaces.","explanation":"GEN-0 is an embodied foundation model that the US robotics company Generalist AI released on November 4, 2025, trained on more than 270,000 hours of real manipulation trajectories collected from thousands of homes, warehouses, and workplaces worldwide, with the company saying it was adding roughly 10,000 more hours each week. Its main finding is a scaling law: downstream performance improves as a power law with more pretraining data. Model size shows a “phase transition” around 7 billion parameters — a 1-billion-parameter model “ossifies” when faced with massive data, losing the ability to learn new things, while larger models keep improving. Architecturally it uses what the company calls Harmonic Reasoning, letting perception and action tokens flow in parallel, asynchronous, continuous time rather than relying on a fast-slow dual-system design. It was followed by GEN-1 in April 2026 and GEN-1.5 in August 2026.","example":"In the company's experiments, a 1-billion-parameter model plateaus when trained on large-scale data, while models of 7 billion parameters or more keep improving as more data is added.","related":["GEN-1","Generalist AI","Scaling Law","Ossification","Embodied Foundation Model","Real-Robot Data"]},{"id":"gen-1","category":"named_model","sec":6,"tier":3,"sources":[{"title":"GEN-1: Scaling Embodied Foundation Models to Mastery (Generalist AI Blog)","url":"https://generalistai.com/blog/gen-1"},{"title":"GEN-1.5: Embodied Foundation Models are One-Shot Learners (Generalist AI Blog)","url":"https://generalistai.com/blog/gen-1.5"}],"as_of":"2026-08","related_ids":["gen-0","generalist-ai","embodied-foundation-model","in-context-learning","scaling-law","real-robot-data"],"name":"GEN-1","alt":"GEN-1","abbr":"","aliases":["GEN-1.5","Scaling Embodied Foundation Models to Mastery"],"one_liner":"Generalist AI's second-generation embodied foundation model, aiming for reliable, fast “mastery” of simple tasks rather than just getting them done.","explanation":"GEN-1 is an embodied foundation model released in April 2026 by the US robotics company Generalist AI, the successor to GEN-0. It was pretrained on more than 500,000 hours of real physical-interaction data, mostly collected from people wearing low-cost wearable devices while doing everyday activities. The company defines its goal as “mastery”: reliable, fast, and able to improvise. According to the company, fine-tuning on just about 1 hour of robot data per task brought success on tasks where the previous model scored around 64% up to 99%, folding a paper box in about 12.1 seconds — roughly 3 times faster than the previous best. The August 2026 version, GEN-1.5, gained in-context learning ability (learning from a demonstration without changing its weights): it can perform a new task after watching just a 3-to-12-second demonstration, with an average single-shot success rate of about 59%.","example":"In an official demo, GEN-1 folded paper boxes 200 times in a row and assembled blocks 1,800 times in a row, to demonstrate its reliability over long, continuous operation.","related":["GEN-0","Generalist AI","Embodied Foundation Model","In-Context Learning","Scaling Law","Real-Robot Data"]},{"id":"sunday-robotics-act-1","category":"named_model","sec":6,"tier":3,"sources":[{"title":"ACT-1: A Robot Foundation Model Trained on Zero Robot Data (Sunday blog, 2025-11-19)","url":"https://www.sunday.ai/blog/no-robot-data"},{"title":"Sunday Robotics 公司页（创始人介绍）","url":"https://www.sunday.ai/company"},{"title":"ACT-2 Preview: Generalizing Reliability (Sunday blog, 2026-07-17)","url":"https://www.sunday.ai/blog/act-2-preview"}],"as_of":"2026-07","related_ids":["skill-capture-glove","sunday-robotics-memo","sunday-robotics","robot-free-data-collection","mobile-manipulation","long-horizon-task"],"name":"Sunday Robotics ACT-1","alt":"Sunday ACT-1","abbr":"ACT-1","aliases":["ACT-1: A Robot Foundation Model Trained on Zero Robot Data","ACT-1 (act one)"],"one_liner":"Sunday Robotics' 2025 household robot model, trained with not a single teleoperated robot demonstration in its data.","explanation":"ACT-1 is the first foundation model announced in November 2025 by Sunday Robotics, a US home-robotics startup founded by ALOHA co-creator Tony Zhao and Diffusion Policy co-creator Cheng Chi. Its training data contains no teleoperated trajectories at all: a person wears a “skill-capture glove” matching the robot hand's shape and sensor layout, does household chores while wearing it, and Skill Transform then strips out the human-specific parts of the footage and motion, converting it into robot data (the company reports roughly a 90% conversion success rate) — sidestepping the fact that teleoperation is slow, expensive, and hard to scale. ACT-1 drives the household robot Memo, using a single end-to-end model for both long-horizon manipulation and navigation via a 3D map; demos include clearing a table into a dishwasher, folding socks, and making coffee, plus tidying a table in an Airbnb it had never visited. In July 2026 the company previewed ACT-2.","example":"In the “table to dishwasher” task, ACT-1 has Memo clear wine glasses, ceramic plates, and metal utensils from a table, scrape off leftovers, load the dishwasher, and start it — 33 kinds of dexterous interactions, 68 in total, across 21 different objects, over more than 130 feet (about 40 meters) of travel.","related":["Skill Capture Glove","Sunday Robotics Memo","Sunday Robotics","Robot-free (Embodiment-free) Data Collection","Mobile Manipulation","Long-horizon Task"]},{"id":"gene-26-5","category":"named_model","sec":6,"tier":3,"sources":[{"title":"GENE-26.5: Advancing Robotic Manipulation to Human Level (Genesis AI Blog)","url":"https://www.genesis.ai/blog/gene-26-5-advancing-robotic-manipulation-to-human-level"}],"as_of":"2026-05","related_ids":["genesis-ai","genesis","dexterous-hand","tactile-glove","flow-matching","embodiment-gap"],"name":"GENE-26.5","alt":"Genesis AI GENE-26.5","abbr":"","aliases":["GENE"],"one_liner":"Genesis AI's first robotics foundation model system, released in May 2026, focused on dexterous manipulation.","explanation":"GENE-26.5 was released on May 7, 2026 by the robotics company Genesis AI, the first public version of its GENE model family — the “26.5” in the name refers to May 2026. The company frames manipulation as a full-stack systems problem: using a high-DOF dexterous hand close to a human hand to narrow the embodiment gap; collecting human data with a tactile-equipped data glove plus first-person and third-person video (more than 200,000 hours total with partners); modeling the joint distribution of language, vision, proprioception, touch, and action with flow matching; running large-scale closed-loop evaluation in the company's own high-fidelity simulator, Genesis World; and building its own low-latency control stack. The company says most difficult skills need less than an hour of task-specific robot data.","example":"In an official demo, the same model autonomously completed a roughly 4-minute cooking routine of more than 20 sub-tasks at normal speed, including cracking an egg one-handed and holding a tomato steady with one hand while cutting it with the other.","related":["Genesis AI","Genesis","Dexterous Hand","Tactile Glove","Flow Matching","Embodiment Gap"]},{"id":"isaac-0-5","category":"named_model","sec":6,"tier":3,"sources":[{"title":"PerceptronAI/Isaac-0.5 模型卡 (Hugging Face)","url":"https://huggingface.co/PerceptronAI/Isaac-0.5"},{"title":"PerceptronAI/Isaac-0.1 模型卡 (Hugging Face)","url":"https://huggingface.co/PerceptronAI/Isaac-0.1"}],"as_of":"2026-08","related_ids":["vision-language-action-model","embodied-foundation-model","open-weight-model","flow-matching","scaling-law","pi0-fast"],"name":"Isaac 0.5 (Perceptron)","alt":"Perceptron Isaac 0.5","abbr":"","aliases":["Perceptron Isaac","Isaac 0.1","Isaac 0.2"],"one_liner":"A 36-billion-parameter open-weight robot foundation model from Perceptron, released in 2026, unifying seeing, thinking, and acting.","explanation":"Isaac 0.5 is a robot foundation model the US startup Perceptron released with open weights under the Apache 2.0 license in August 2026; the company was founded by a team that had worked on Meta's Chameleon multimodal model, and had previously released the 1B–2B-parameter perception-language models Isaac 0.1 and 0.2. It has no relation to NVIDIA's Isaac platform. The model has 36 billion parameters total, built on the Qwen family of vision-language models as its backbone with added sparse mixture-of-experts layers; it takes in images, video, instructions, robot state, and action history, and can output text, coordinates, task progress, or actions — continuous actions are generated by a flow-matching expert, and discrete actions use FAST action tokens. The company says its training data includes 100,000 hours of experience from more than 35 robot types plus 1 million hours of ordinary video.","example":"The company's own fitted data-mixing law: to reach the same action-prediction loss, with only 1,000 hours of ordinary video you need about 5,900 hours of teleoperation data, but once ordinary video reaches 1 million hours, only about 28 hours are needed.","related":["Vision-Language-Action Model","Embodied Foundation Model","Open-weight Model","Flow Matching","Scaling Law","π0-FAST"]},{"id":"gr-1","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (arXiv 2312.13139)","url":"https://arxiv.org/abs/2312.13139"},{"title":"GR-1 项目页","url":"https://gr1-manipulation.github.io/"},{"title":"bytedance/GR-1 GitHub 仓库（ICLR 2024）","url":"https://github.com/bytedance/GR-1"}],"as_of":"2024-01","related_ids":["video-prediction-model","gr-2","seed-gr-3","calvin-benchmark","language-conditioned-policy","pre-training"],"name":"GR-1 (ByteDance)","alt":"字节 GR-1","abbr":"GR-1","aliases":["Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation"],"one_liner":"A late-2023 ByteDance manipulation model that first learns to predict future video frames, then learns to output actions.","explanation":"GR-1 was released by ByteDance Research in December 2023, published at ICLR 2024. It is a GPT-style Transformer: it takes in a language instruction, a history of images, and robot state, and outputs both a future frame and a robot action at the same time. Training happens in two steps: first, video-prediction pretraining on large-scale video, learning only “what will the next frame look like,” which needs no action labels; then fine-tuning on robot data, so the model outputs actions while still predicting frames. The authors' reasoning is that predicting video teaches the model how objects get pushed and how they change, knowledge that is useful for manipulation. On the CALVIN benchmark, it raised the success rate from 88.9% to 94.9%, and zero-shot generalization success on unseen scenes from 53.3% to 85.4%. It is the first generation of ByteDance's GR series, followed by GR-2 and GR-3. Note that this is not the same GR-1 as Fourier's humanoid robot.","example":"In the CALVIN simulated kitchen tabletop, GR-1 follows language instructions to complete a sequence of sub-tasks in a row, such as “open the drawer,” “push the blue block to the left,” and “turn on the light bulb.”","related":["Video Prediction Model","GR-2 (ByteDance)","Seed GR-3","CALVIN Benchmark","Language-conditioned Policy","Pre-training"]},{"id":"gr-2","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation (arXiv 2410.06158)","url":"https://arxiv.org/abs/2410.06158"},{"title":"GR-2 论文 HTML 全文","url":"https://arxiv.org/html/2410.06158v1"}],"as_of":"2024-10","related_ids":["gr-1","seed-gr-3","video-generation-model","pretraining-on-human-videos","conditional-variational-autoencoder","bin-picking"],"name":"GR-2 (ByteDance)","alt":"字节 GR-2","abbr":"GR-2","aliases":["A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation"],"one_liner":"ByteDance's second-generation manipulation model, released in 2024, pretrained first on 38 million web videos before learning to act.","explanation":"GR-2 was released by ByteDance Research in October 2024, an upgrade over GR-1 that the authors call a “video-language-action model.” The first stage does video-generation pretraining on 38 million internet videos (more than 50 billion tokens), drawn from everyday human-activity video sources like HowTo100M, Ego4D, Something-Something V2, and EPIC-KITCHENS, teaching the model how footage will unfold next; the second stage fine-tunes on robot trajectories, predicting future frames and an action trajectory at the same time. Images are cut into discrete tokens with VQGAN, and the action trajectory is generated by a conditional variational autoencoder (CVAE). The paper reports an average 97.7% success rate across more than 100 tasks, along with strong generalization to new backgrounds, environments, objects, and tasks. It was deployed on a real Kinova Gen3 arm, with trajectory optimization and a real-time-tracking whole-body control algorithm carrying out the actions.","example":"In an industrial bin-picking experiment requiring picking a specified item out of a cluttered bin, across 122 object types (67 of them unseen during training), GR-2 reached an average 79.0% success rate, versus just 33.3% for GR-1.","related":["GR-1 (ByteDance)","Seed GR-3","Video Generation Model","Pretraining on Human Videos","Conditional Variational Autoencoder","Bin Picking"]},{"id":"seed-gr-3","category":"named_model","sec":7,"tier":2,"sources":[{"title":"GR-3 Technical Report (arXiv:2507.15493)","url":"https://arxiv.org/abs/2507.15493"}],"as_of":"2025-07","related_ids":["gr-2","gr-1","bytedance-seed","vision-language-action-model","flow-matching","co-training"],"name":"Seed GR-3","alt":"字节 GR-3","abbr":"GR-3","aliases":["ByteDance Seed GR-3","Generalist Robot Model 3"],"one_liner":"ByteDance Seed's 2025 general-purpose robot VLA, about 4 billion parameters, controlling its own ByteMini bimanual robot.","explanation":"GR-3 is a vision-language-action model the ByteDance Seed team released in July 2025, following GR-1 and GR-2. It has about 4 billion parameters, using Qwen2.5-VL-3B as its vision-language backbone, followed by a diffusion Transformer action head trained with flow matching that generates one action chunk at a time. Training draws on three kinds of data: robot trajectories for imitation learning; web image-text data co-trained in to preserve understanding of new objects and abstract instructions; and human trajectories captured with a VR device for few-shot adaptation — the paper reports that adding just 10 demonstrations per unseen object raised pick-and-place success from 57.8% to 86.7%. Its companion robot, ByteMini, is a 22-degree-of-freedom bimanual mobile platform, and GR-3 outperforms π0 on long-horizon tasks such as clearing a table and hanging up clothes.","example":"In the clothes-hanging task, GR-3 coordinates ByteMini's two arms to thread a hanger into a garment and then hang it on a rack — an example of bimanual deformable-object manipulation.","related":["GR-2 (ByteDance)","GR-1 (ByteDance)","ByteDance Seed","Vision-Language-Action Model","Flow Matching","Co-training"]},{"id":"gr-dexter","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GR-Dexter Technical Report (arXiv 2512.24210)","url":"https://arxiv.org/abs/2512.24210"},{"title":"GR-Dexter 技术报告 HTML 全文","url":"https://arxiv.org/html/2512.24210v2"}],"as_of":"2025-12","related_ids":["dexterous-manipulation","bimanual-manipulation","dexterous-hand","seed-gr-3","vision-language-action-model","data-glove"],"name":"GR-Dexter","alt":"字节 GR-Dexter","abbr":"GR-Dexter","aliases":["ByteDance Seed VLA for bimanual high-DoF dexterous-hand robots"],"one_liner":"A ByteDance Seed VLA from late 2025 for bimanual dexterous hands, built together with its own hand hardware and teleoperation system.","explanation":"GR-Dexter was released as a technical report by ByteDance's Seed team in late December 2025. Most VLA (vision-language-action) models only control a two-finger gripper; switching to a pair of high-DOF dexterous hands enlarges the action space, the hands frequently occlude the object, and real-robot data becomes more expensive too. GR-Dexter tackles hardware, data, and model together: a self-developed, 21-degree-of-freedom, linkage-driven ByteDexter V2 dexterous hand with piezoresistive tactile sensors in the fingertips, mounted on two Franka Research 3 arms; data collected through bimanual teleoperation using a Meta Quest headset plus Manus data gloves; and a model that carries over GR-3's hybrid Transformer structure at about 4 billion parameters, trained on a mix of teleoperation trajectories, vision-language data, curated cross-embodiment data, and more than 800 hours of human hand-trajectory data.","example":"On the long-horizon “organize cosmetics” task, success rate is 0.97 under standard layouts and still 0.89 across 5 unseen layouts; on a general pick-and-place task with 23 unseen objects, success rate is 0.85.","related":["Dexterous Manipulation","Bimanual Manipulation","Dexterous Hand","Seed GR-3","Vision-Language-Action Model","Data Glove"]},{"id":"gr-rl","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation (arXiv 2512.01801)","url":"https://arxiv.org/abs/2512.01801"},{"title":"GR-RL 论文 HTML 全文","url":"https://arxiv.org/html/2512.01801v3"}],"as_of":"2025-12","related_ids":["seed-gr-3","reinforcement-fine-tuning","offline-reinforcement-learning","noise-space-policy-steering","data-curation","long-horizon-task"],"name":"GR-RL","alt":"字节 GR-RL","abbr":"","aliases":["Going Dexterous and Precise for Long-Horizon Robotic Manipulation"],"one_liner":"ByteDance Seed uses reinforcement learning to turn a generalist VLA into a specialist, the first learned policy to lace a shoe autonomously.","explanation":"GR-RL was released by ByteDance's Seed team in December 2025, starting from the generalist model GR-3. Its premise is that ordinary VLA training assumes human demonstrations are optimal, but in fine-grained, long-horizon dexterous tasks, demonstrations often contain jitter and unnecessary motion. GR-RL works in three steps: first, offline reinforcement learning with sparse rewards, using the learned Q-value as an estimate of “task progress” to cut out segments that don't contribute to progress; then morphological symmetry augmentation — flipping images left-right, swapping the left and right wrist cameras, and rewriting left/right words in the instruction to match — to expand the data; and finally online reinforcement learning, which trains a latent-space noise predictor that steers a flow-matching policy's output toward higher reward, aligning behavior between training and actual deployment. The experimental platform is ByteMini-v2, a wheeled robot with two 7-degree-of-freedom arms.","example":"Shoe-lacing: threading a lace through a shoe's eyelets in sequence requires long-horizon planning, millimeter-level precision, and compliant handling of a soft lace. GR-RL reaches an 83.3% success rate, which the authors say is the first learned policy able to complete this task autonomously.","related":["Seed GR-3","Reinforcement Fine-Tuning (RL Fine-Tuning)","Offline Reinforcement Learning","Noise-Space Policy Steering","Data Curation","Long-horizon Task"]},{"id":"robix","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Robix (arXiv 2509.01106)","url":"https://arxiv.org/abs/2509.01106"},{"title":"Robix project page","url":"https://robix-seed.github.io/robix/"}],"as_of":"2025-09","related_ids":["seed-gr-3","hierarchical-architecture","dual-system-architecture","embodied-reasoning","llm-based-task-planning","bytedance-seed"],"name":"Robix","alt":"字节 Robix","abbr":"Robix","aliases":["Robix-7B","Robix-32B","Robix: A Unified Model for Robot Interaction, Reasoning and Planning"],"one_liner":"ByteDance Seed's high-level robot 'brain' model that unifies conversation, reasoning, and task planning.","explanation":"Robix was released by the ByteDance Seed team in September 2025, in 7B and 32B parameter sizes. It serves as the high-level cognitive layer in a hierarchical robot system: given the camera feed and what the user says, a single vision-language model handles reasoning, long-horizon task planning, and natural-language interaction all at once, producing both atomic instructions for a lower-level controller and things to say back to the user; the lower level can be ByteDance's own GR-3 VLA model. Robix can proactively ask clarifying questions about ambiguous instructions, revise its plan in real time when interrupted mid-execution, and apply commonsense reasoning, such as filtering food choices by dietary restrictions. Training has three stages: continued pretraining to strengthen embodied-reasoning abilities such as spatial understanding, supervised fine-tuning that unifies interaction and planning into reasoning-action sequences, and reinforcement learning to improve consistency between reasoning and action. The authors report it outperforms commercial large-model baselines on real tasks such as clearing a dinner table and grocery shopping.","example":"A user asks the robot to clear the table, then mid-task adds, 'leave the cup for now' — Robix immediately revises its plan, confirms verbally, and sends the lower-level VLA a new atomic instruction.","related":["Seed GR-3","Hierarchical Architecture","Dual-System Architecture (System 1 / System 2)","Embodied Reasoning","LLM-based Task Planning","ByteDance Seed"]},{"id":"era-42","category":"named_model","sec":7,"tier":3,"sources":[{"title":"星动纪元官网 · ERA-42 模型页","url":"https://www.robotera.com/model"},{"title":"星动纪元官网 · 关于我们","url":"https://www.robotera.com/about/us"}],"as_of":"2025-11","related_ids":["robotera","vision-language-action-model","video-prediction-policy","ctrl-world","robotera-xhand1","robotera-l7"],"name":"ERA-42","alt":"星动纪元 ERA-42","abbr":"","aliases":["ERA42","RobotEra end-to-end native robot model"],"one_liner":"RobotEra's end-to-end VLA embodied model, released in late 2024, driving the company's dexterous hands and humanoid robots.","explanation":"ERA-42 is the embodied large model from Beijing-based RobotEra, described on the company's website as an “end-to-end VLA embodied model.” RobotEra was founded in August 2023; its founder, Jianyu Chen, is an assistant professor at Tsinghua University's Institute for Interdisciplinary Information Sciences (IIIS), and Tsinghua holds a stake in the company. According to the company's own timeline, ERA-42 was released in December 2024, with the first version demonstrated on a single arm plus a five-fingered dexterous hand; it moved to dual arms in March 2025, and by July was driving the same model across all 55 degrees of freedom of the full-size humanoid robot L7. The technical work the company's website lists as groundwork includes papers on the video prediction policy VPP, UP-VLA, the online reinforcement-learning method iRe-VLA, and the controllable world model Ctrl-World. The company says it now has 100 dexterous manipulation skills deployed in logistics, manufacturing, and commercial services; the architecture and parameter count have not been fully disclosed.","example":"In RobotEra's warehousing solution, a robot running on ERA-42 completes the full order-outbound-picking-scanning-packing sequence, using the 12-DOF XHAND1 dexterous hand to grasp irregularly shaped medicine boxes and flip them to scan the barcode.","related":["RobotEra","Vision-Language-Action Model","Video Prediction Policy","Ctrl-World","RobotEra XHAND1","RobotEra L7"]},{"id":"agibot-go-1","category":"named_model","sec":7,"tier":2,"sources":[{"title":"AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems (arXiv 2503.06669)","url":"https://arxiv.org/abs/2503.06669"},{"title":"OpenDriveLab/AgiBot-World（GitHub，含 GO-1 / GO-1 Air 开源说明）","url":"https://github.com/OpenDriveLab/AgiBot-World"}],"as_of":"2025-09","related_ids":["vision-language-latent-action","latent-action","agibot-world","agibot-go-2","action-expert","agibot"],"name":"AgiBot GO-1","alt":"智元 GO-1（启元大模型）","abbr":"GO-1","aliases":["Genie Operator-1","GO-1 Air"],"one_liner":"AgiBot's 2025 general-purpose embodied foundation model, using latent actions so human video can help train it too.","explanation":"GO-1 is the general-purpose manipulation policy AgiBot released on March 10, 2025, alongside the technical report for its AgiBot World dataset. It proposes the ViLLA (vision-language-latent-action) framework in three parts: a latent action model compresses the change between adjacent frames into discrete latent action tokens, and can be trained on human videos with no action labels, such as Ego4D; a latent planner, built on the InternVL2.5-2B vision-language model, predicts these tokens from images and instructions; an action expert then uses a diffusion objective to output continuous, high-frequency actions. The paper reports that pretraining on AgiBot World — over 1 million trajectories collected from 100 robots — improves average performance 30% over using Open X-Embodiment; on complex real-robot tasks GO-1 reaches over 60% success, 32 points higher than RDT. The model was open-sourced in September 2025, together with a lighter GO-1 Air variant that drops the latent planner; its successor is 2026's GO-2.","example":"The open-sourced GO-1 weights are hosted on Hugging Face; the official documentation says inference needs about 7GB of GPU memory, and full-parameter fine-tuning with batch size 16 needs about 70GB.","related":["Vision-Language-Latent-Action","Latent Action","AgiBot World","AgiBot GO-2","Action Expert","AgiBot"]},{"id":"agibot-go-2","category":"named_model","sec":7,"tier":2,"sources":[{"title":"AGIBOT Unveils Genie Operator-2 (GO-2) Next-Gen Embodied Foundation Model","url":"https://www.agibot.com/article/231/detail/56.html"},{"title":"AGIBOT Unveils New Generation of Embodied AI Robots and Models (APC 2026, 2026-04-17)","url":"https://www.agibot.com/article/231/detail/63.html"},{"title":"ACoT-VLA (AgibotTech GitHub)","url":"https://github.com/AgibotTech/ACoT-VLA"}],"as_of":"2026-04","related_ids":["agibot-go-1","vision-language-latent-action","action-chain-of-thought","dual-system-architecture","agibot","libero-benchmark"],"name":"AgiBot GO-2","alt":"智元 GO-2（Genie Operator-2）","abbr":"GO-2","aliases":["Genie Operator-2","GO-2 (ViLLA)"],"one_liner":"AgiBot's 2026 second-generation embodied foundation model: it first plans a coarse action sequence, then executes it.","explanation":"GO-2 is the embodied foundation model AgiBot released on April 17, 2026, at its “2026 Partner Conference” in Shanghai — the successor to 2025's GO-1, which the company describes as a ViLLA-style embodied foundation model. Two ideas are central. First, “action chain-of-thought”: instead of first writing text or generating an image and only then deriving actions, the model directly generates a sequence of coarse-grained action intentions in action space as a high-level plan; the related paper, ACoT-VLA, was accepted at CVPR 2026. Second, an asynchronous two-system design: a low-frequency semantic planning module produces the plan, and a high-frequency action-following module combines it with real-time observations to output control signals — this part was accepted at ACL 2026. AgiBot reports an average LIBERO success rate of 98.5% and 86.6% zero-shot on LIBERO-Plus. Released alongside it was GE-2, a world-action model used to test policies in virtual environments.","example":"AgiBot reports that GO-2, trained only on simulated data, reached an 82.9% success rate in real-robot testing.","related":["AgiBot GO-1","Vision-Language-Latent-Action","Action Chain-of-Thought","Dual-System Architecture (System 1 / System 2)","AgiBot","LIBERO Benchmark"]},{"id":"enerverse","category":"named_model","sec":7,"tier":3,"sources":[{"title":"EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation (arXiv 2501.01895)","url":"https://arxiv.org/abs/2501.01895"},{"title":"EnerVerse 项目页","url":"https://sites.google.com/view/enerverse"}],"as_of":"2025-11","related_ids":["video-prediction-model","world-model","genie-envisioner","3d-gaussian-splatting","multi-view","video-prediction-policy"],"name":"EnerVerse","alt":"EnerVerse","abbr":"","aliases":["EnerVerse-A","EnerVerse-D","Envisioning Embodied Future Space for Robotics Manipulation"],"one_liner":"A generative robot model from AgiBot and others that first predicts future multi-view frames with video diffusion, then produces actions.","explanation":"EnerVerse was released in January 2025 by AgiBot, Shanghai AI Lab, and other teams, and was later accepted to NeurIPS 2025. Its approach is to predict the future first and decide on actions second: an autoregressive video diffusion model generates future “embodied space” frames segment by segment based on an instruction, using sparse context memory to support long-horizon tasks. To represent 3D scenes, the authors propose Free Anchor Views, a multi-view video representation whose viewpoints are chosen flexibly per task. EnerVerse-A is a policy head attached after the generative model that converts the predicted 4D representation into robot actions; EnerVerse-D is a data engine that combines the generative model with 4D Gaussian splatting, used to automatically synthesize data and narrow the gap between simulation and reality.","example":"Given a manipulation instruction, the model first generates future video clips, from several viewpoints, of the robot completing the task, and the policy head then outputs actions based on them; the paper reports producing an 8-step action chunk in about 280 milliseconds.","related":["Video Prediction Model","World Model","Genie Envisioner (AgiBot)","3D Gaussian Splatting","Multi-View","Video Prediction Policy"]},{"id":"genie-envisioner","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Genie Envisioner (arXiv:2508.05635)","url":"https://arxiv.org/abs/2508.05635"},{"title":"AgibotTech/Genie-Envisioner (GitHub)","url":"https://github.com/AgibotTech/Genie-Envisioner"},{"title":"GE-Act 2.0 项目主页","url":"https://ge-act-v2.github.io/"}],"as_of":"2026-09","related_ids":["agibot","world-action-model","agibot-world","ewmbench","neural-simulator","agibot-go-1"],"name":"Genie Envisioner (AgiBot)","alt":"智元 Genie Envisioner 世界模型","abbr":"GE","aliases":["GE-Base","GE-Act","GE-Sim","GE-Act 2.0","A Unified World Foundation Platform for Robotic Manipulation"],"one_liner":"AgiBot's platform that unifies a video world model, a policy, a neural simulator, and evaluation into one system.","explanation":"Genie Envisioner is a robot-manipulation platform released in August 2025 by the AgiBot team. Its core, GE-Base, is a diffusion model that generates video from a language instruction, trained on about 1 million clips (2,967 hours total) of dual-arm real-robot data from AgiBot World, learning to predict how the next frames will unfold. GE-Act uses a roughly 160-million-parameter flow-matching decoder to convert GE-Base's latent representation into actions; GE-Sim takes in an action and generates future frames, serving as a neural simulator for closed-loop evaluation; and EWMBench is used to measure the quality of this kind of world model. Code and weights are open-source. The September 2026 GE-Act 2.0 switched to training entirely from scratch on manipulation data, chaining together an autoencoder, a single-step visual planner, and an inverse dynamics model.","example":"In the paper, adapting to a new task with just about 1 hour of data lets GE-Act transfer to other bimanual platforms, such as AgileX's Cobot Magic, to carry out manipulation tasks.","related":["AgiBot","World Action Model","AgiBot World","EWMBench","Neural Simulator","AgiBot GO-1"]},{"id":"robobrain","category":"named_model","sec":7,"tier":3,"sources":[{"title":"arXiv 2502.21257: RoboBrain (CVPR 2025)","url":"https://arxiv.org/abs/2502.21257"},{"title":"arXiv 2507.02029: RoboBrain 2.0 Technical Report","url":"https://arxiv.org/abs/2507.02029"},{"title":"GitHub: FlagOpen/RoboBrain2.5","url":"https://github.com/FlagOpen/RoboBrain2.5"}],"as_of":"2026-03","related_ids":["embodied-foundation-model","braincerebellum-architecture","affordance","spatial-reasoning","roboos","beijing-academy-of-artificial-intelligence"],"name":"RoboBrain","alt":"智源 RoboBrain（具身大脑）","abbr":"","aliases":["RoboBrain 1.0","RoboBrain 2.0","RoboBrain 2.5","BAAI embodied brain model"],"one_liner":"An open-source series of embodied 'brain' models from BAAI in Beijing, responsible for task planning and spatial understanding.","explanation":"RoboBrain is a series of open-source embodied 'brain' models from the Beijing Academy of Artificial Intelligence (BAAI). Given an image and an instruction, it outputs intermediate results such as task decomposition, affordances (where an object can be grasped or pressed), and end-effector trajectories, which are then handed to a lower-level controller for execution; it does not directly output motor commands itself. Version 1.0 was released in February 2025 and accepted at CVPR 2025, together with the ShareRobot annotated dataset; version 2.0's technical report followed in July 2025, with 7B and 32B sizes, adding capabilities such as spatial referring and long-horizon planning for multi-robot collaboration; version 2.5 began releasing 8B and 4B weights from January 2026 onward, upgrading from predicting 2D image coordinates to predicting 3D spatial trajectories with depth, and adding a general-purpose reward model that estimates task progress. RoboBrain is one of the more complete open-source 'brain' models in the domestic 'brain-cerebellum' hierarchical approach.","example":"Given a kitchen photo and the instruction 'put the bowl in the microwave,' RoboBrain first breaks the task into sub-steps, then marks on the image where the bowl should be grasped and the hand's movement trajectory, for a lower-level controller to execute.","related":["Embodied Foundation Model","Brain–Cerebellum Architecture","Affordance","Spatial Reasoning","RoboOS","Beijing Academy of Artificial Intelligence"]},{"id":"govla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"智平方发布全新一代智能机器人AlphaBot 2，开启AGI终端新时代！（智平方官网）","url":"https://ai2robotics.com/%E6%99%BA%E5%B9%B3%E6%96%B9%E5%8F%91%E5%B8%83%E5%85%A8%E6%96%B0%E4%B8%80%E4%BB%A3%E6%99%BA%E8%83%BD%E6%9C%BA%E5%99%A8%E4%BA%BAalphabot-2%E5%BC%80%E5%90%AFagi%E7%BB%88%E7%AB%AF%E6%96%B0/"}],"as_of":"2025-04","related_ids":["ai2-robotics","ai2-robotics-alphabot","dual-system-architecture","vision-language-action-model","robomamba","mobile-manipulation"],"name":"GOVLA (Global & Omni-body Vision-Language-Action)","alt":"智平方 GOVLA（AlphaBrain）","abbr":"GOVLA","aliases":["Alpha Brain","AlphaBrain","AI2R Brain"],"one_liner":"AI² Robotics' large VLA model that outputs both full-body actions and a movement trajectory at once, across arbitrary embodiments and spaces.","explanation":"GOVLA was introduced by AI² Robotics in April 2025 alongside its AlphaBot 2 robot, at the same time the company upgraded its earlier embodied-model brand, AI2R Brain, to Alpha Brain, whose core is GOVLA (Global & Omni-body Vision-Language-Action model). It has three parts: a spatial-interaction foundation model, a slow system, and a fast system. The slow system (System 2) handles complex reasoning, task decomposition, and language interaction, while the fast system (System 1) outputs whole-body control actions and a movement trajectory. The company's framing is that ordinary VLA models only output robot-arm actions and can only operate within tabletop range, whereas GOVLA targets tasks ranging from the tabletop to open environments and from single-arm to whole-body coordination, and incorporates technology from DeepSeek in its construction to strengthen long-horizon reasoning.","example":"As the company's own example puts it: an ordinary VLA robot needs a person to first put the ingredients on the table before it can work, whereas a robot running GOVLA can go get ingredients from the refrigerator itself, make breakfast, and bring it to the table.","related":["AI² Robotics","AI² Robotics AlphaBot","Dual-System Architecture (System 1 / System 2)","Vision-Language-Action Model","RoboMamba","Mobile Manipulation"]},{"id":"graspvla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data (arXiv 2505.03233)","url":"https://arxiv.org/abs/2505.03233"},{"title":"GraspVLA 项目页","url":"https://pku-epic.github.io/GraspVLA-web/"}],"as_of":"2025-08","related_ids":["syngrasp-1b","synthetic-data","sim-to-real-transfer","domain-randomization","flow-matching","grasping"],"name":"GraspVLA","alt":"银河通用 GraspVLA","abbr":"","aliases":["a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data"],"one_liner":"A 2025 grasping foundation model from Galbot and others, pretrained mainly on a billion frames of simulated synthetic data.","explanation":"GraspVLA was proposed by Galbot together with He Wang's group at Peking University, the University of Hong Kong, and the Beijing Academy of Artificial Intelligence, released in May 2025 and published at CoRL 2025. Real-robot data is expensive and hard to scale up, so the team instead generated data at large scale in simulation: the SynGrasp-1B dataset has about 1 billion frames, spanning 240 categories and more than 10,000 object models, with heavy randomization (domain randomization) of initial pose, placement, background, lighting, and material, and photorealistic rendering to narrow the sim-to-real gap. The model uses “progressive action generation”: it autoregressively predicts a 2D detection box for the target object, then predicts a grasp pose, and finally generates an action chunk with flow matching; training mixes this synthetic data with internet image-text data, letting it grasp unseen object categories from open-vocabulary instructions. The dataset and model weights are open-source.","example":"After pretraining on synthetic data alone, the model can grasp unseen objects on a real tabletop zero-shot; if a particular scene calls for a specific grasping style, a small number of real-robot demonstrations is enough for few-shot fine-tuning.","related":["SynGrasp-1B","Synthetic Data","Sim-to-Real Transfer","Domain Randomization","Flow Matching","Grasping"]},{"id":"astrabrain","category":"named_model","sec":7,"tier":3,"sources":[{"title":"银河通用官网（银河星脑 AstraBrain 架构介绍）","url":"https://www.galbot.com/"},{"title":"中国唯一具身智能大模型重点实验室，银河通用牵头获批！（银河通用公众号，2026-08-10）","url":"https://mp.weixin.qq.com/s/6IdDqb0yxYQU7eQKAsolDw"},{"title":"颠覆后训练范式！无需人类动作标签，银河通用WAM-TTT让机器人换厨房不丢手艺（腾讯科技，2026-07-16）","url":"https://mp.weixin.qq.com/s/LMtt3FGB6zyB7Boex4S6kg"}],"as_of":"2026-08","related_ids":["braincerebellum-architecture","world-action-model","learning-based-whole-body-control","scaling-law","graspvla","trackvla"],"name":"AstraBrain","alt":"银河星脑 AstraBrain","abbr":"","aliases":["Galbot AstraBrain","AstraBrain-WBC","AstraBrain-WAM","AstraBrain-Dex"],"one_liner":"Galbot's embodied model family, connecting a “brain” that plans with a “cerebellum” that handles real-time whole-body control.","explanation":"AstraBrain is Galbot's umbrella name for its embodied model family; the Beijing-based humanoid company describes it as an end-to-end “brain–cerebellum–neural control” architecture. AstraBrain-WAM is the “brain,” a world-action model (predicting future images and generating actions at the same time); Galbot's website lists it at 3 to 13 billion parameters, outputting high-level commands at 5–10Hz and handling task understanding and long-horizon planning, and the company says it can draw on simulated and real-robot data, and human and robot data, with or without action labels, all in one framework. AstraBrain-WBC is the “cerebellum,” an 80-million-to-300-million-parameter Transformer that does real-time whole-body and hand control at 100–250Hz, trained on large-scale human motion data. AstraBrain-Dex is a world model aimed at dexterous hands. WAM 0.5, WBC 0.5, and a test-time training framework called WAM-TTT, all disclosed in 2026, belong to this same family. Most of these figures and claims come from the company itself.","example":"In a 2026 demo, Galbot showed a humanoid robot autonomously playing tennis — seeing the incoming ball, predicting its trajectory, moving into position to hit it, and keeping its balance — which the company describes as whole-body control achieved jointly by the brain and cerebellum models.","related":["Brain–Cerebellum Architecture","World Action Model","Learning-Based Whole-Body Control","Scaling Law","GraspVLA","TrackVLA"]},{"id":"worldvla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"WorldVLA (arXiv:2506.21539)","url":"https://arxiv.org/abs/2506.21539"},{"title":"alibaba-damo-academy/WorldVLA (GitHub)","url":"https://github.com/alibaba-damo-academy/WorldVLA"}],"as_of":"2025-11","related_ids":["world-action-model","vision-language-action-model","autoregressive-decoding","attention-mask","rynnvla-002","alibaba-damo-academy"],"name":"WorldVLA","alt":"WorldVLA","abbr":"","aliases":["WorldVLA: Towards Autoregressive Action World Model"],"one_liner":"Alibaba DAMO Academy's model that merges a VLA and a world model into one autoregressive system, outputting actions and predicting the next frame.","explanation":"WorldVLA was released by Alibaba's DAMO Academy in June 2025, built on an autoregressive image-and-text model in the Chameleon family, converting images, text, and actions all into tokens fed into the same Transformer. The single model plays two roles at once: as an action model, it generates actions from an image and an instruction; as a world model, it predicts the next frame from the current image and an action. The authors found that training the two jointly improves both. They also found that when actions are generated autoregressively one after another, errors in earlier actions propagate to later ones, so they introduce an attention mask that, when generating the current action, hides previous actions and looks only at the image and instruction, which noticeably improves multi-step action generation on LIBERO. The project was upgraded to RynnVLA-002 in November 2025.","example":"On a LIBERO simulation task, the same model can both output the robot arm's next segment of action and generate the resulting frame given a specified action.","related":["World Action Model","Vision-Language-Action Model","Autoregressive Decoding","Attention Mask","RynnVLA-002","Alibaba DAMO Academy"]},{"id":"rynnvla-002","category":"named_model","sec":7,"tier":3,"sources":[{"title":"RynnVLA-002: A Unified Vision-Language-Action and World Model (arXiv 2511.17502)","url":"https://arxiv.org/abs/2511.17502"},{"title":"alibaba-damo-academy/RynnVLA-002 (GitHub)","url":"https://github.com/alibaba-damo-academy/RynnVLA-002"},{"title":"alibaba-damo-academy/RynnVLA-001 (GitHub)","url":"https://github.com/alibaba-damo-academy/RynnVLA-001"}],"as_of":"2026-05","related_ids":["world-action-model","world-model","vision-language-action-model","worldvla","libero-benchmark","rynnbrain"],"name":"RynnVLA-002","alt":"达摩院 RynnVLA-002","abbr":"","aliases":["RynnVLA","RynnVLA-002: A Unified Vision-Language-Action and World Model"],"one_liner":"An open-source VLA from Alibaba DAMO Academy that merges an action model and a world model into one autoregressive network.","explanation":"RynnVLA-002 was released and open-sourced by Alibaba's DAMO Academy in November 2025. Its predecessor, RynnVLA-001 (August 2025), was a 7B-parameter VLA that first went through video-generation pretraining on first-person human videos, then transferred the manipulation skills it learned onto a robot arm. RynnVLA-002 merges two kinds of models into one: the VLA part outputs actions from images and language instructions, and the world-model part predicts the next frame from the current image and an action. Both share a single autoregressive backbone based on Chameleon, treating images, text, and actions all as tokens, plus a separate Action Transformer that outputs continuous actions, with support for wrist-camera and robot-state input. The authors' view is that learning to predict 'how the image changes as a result of an action' helps the model understand physical dynamics, which in turn improves action quality. The paper reports 97.4% success on the LIBERO simulation benchmark without pretraining, and roughly a 50% success-rate gain from adding the world model on real-robot LeRobot tasks.","example":"The same model can output the robot arm's next segment of motion for 'put the block in the box,' and can also generate the wrist-camera image that would result from executing a given action.","related":["World Action Model","World Model","Vision-Language-Action Model","WorldVLA","LIBERO Benchmark","RynnBrain"]},{"id":"rynnbrain","category":"named_model","sec":7,"tier":3,"sources":[{"title":"alibaba-damo-academy/RynnBrain (GitHub)","url":"https://github.com/alibaba-damo-academy/RynnBrain"},{"title":"RynnBrain: Open Embodied Foundation Models (arXiv 2602.14979)","url":"https://arxiv.org/abs/2602.14979"},{"title":"RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model (arXiv 2607.17977)","url":"https://arxiv.org/abs/2607.17977"}],"as_of":"2026-07","related_ids":["alibaba-damo-academy","embodied-foundation-model","embodied-reasoning-model","braincerebellum-architecture","rynnvla-002","mixture-of-experts"],"name":"RynnBrain","alt":"达摩院 RynnBrain","abbr":"","aliases":["RynnBrain 1.0","RynnBrain 1.1","Alibaba DAMO Academy Open Embodied Foundation Models"],"one_liner":"Alibaba DAMO Academy's open-source embodied 'brain' model, for understanding first-person video, localizing objects, and planning tasks.","explanation":"RynnBrain was open-sourced by Alibaba's DAMO Academy in February 2026, positioned as a spatiotemporal foundation model that serves as a robot's 'brain' rather than a policy that outputs motor commands directly. It combines four capabilities in one model: understanding first-person video, spatiotemporal grounding (locating objects and operable parts in video), reasoning grounded in physical commonsense, and task planning. Version 1.0 came in 2B, 8B, and 30B-A3B sizes (the last a mixture-of-experts model with 30B total and about 3B active parameters), and was further post-trained into RynnBrain-Nav (navigation), -Plan (planning), -VLA (action), and a spatial-reasoning-focused -CoP variant, alongside the RynnBrain-Bench evaluation suite. Version 1.1, released in July 2026, is built on Qwen3.5 and scales up to 2B, 9B, and 122B-A10B (the largest also a mixture-of-experts model), adding contact-point prediction, with the 2B and 9B versions also supporting native 3D grounding; the company states its largest model beats the closed- and open-source models it evaluated against on three spatial-reasoning benchmarks — VSI-Bench, MMSI, and RefSpatial-Bench. Code and weights are released under Apache 2.0.","example":"Given 'put the cup on the table into the sink,' RynnBrain first marks the location of the cup and the sink in the first-person image, then breaks the task into steps — walk to the table, pick up the cup, walk to the sink, put it down — for a lower-level action model to execute.","related":["Alibaba DAMO Academy","Embodied Foundation Model","Embodied Reasoning Model","Brain–Cerebellum Architecture","RynnVLA-002","Mixture of Experts"]},{"id":"being-h0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (arXiv 2507.15597)","url":"https://arxiv.org/abs/2507.15597"},{"title":"Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization (arXiv 2601.12993)","url":"https://arxiv.org/abs/2601.12993"},{"title":"BeingBeyond/Being-H0 GitHub","url":"https://github.com/BeingBeyond/Being-H0"}],"as_of":"2026-05","related_ids":["pretraining-on-human-videos","unihand","vision-language-action-model","dexterous-manipulation","cross-embodiment","beingbeyond"],"name":"Being-H0","alt":"智在无界 Being-H0","abbr":"","aliases":["Being-H0.5","Being-H0.7","Being-H Series"],"one_liner":"BeingBeyond's series of dexterous-manipulation VLA models, pretrained on large-scale video of human hands doing manipulation.","explanation":"Being-H0 is a vision-language-action model released in July 2025 by Peking University, Renmin University of China, and the embodied-AI company BeingBeyond (智在无界), later accepted at ICML 2026. Its idea is to treat the human hand as a universal manipulator: it first pretrains on human hand-manipulation data, using a part-specific action tokenizer to compress wrist and finger motion into tokens, then does 3D spatial alignment, and finally post-trains on a small amount of robot data to transfer onto a dexterous robot hand. Its companion dataset, UniHand, pools motion capture, VR, and ordinary video, totaling about 1,100 hours and 165 million instruction samples. Being-H0.5, from January 2026, expands the data to more than 35,000 hours across 30 embodiments, with an emphasis on cross-embodiment transfer; Being-H0.7, from April, switches to a latent-space world-action model.","example":"In real-robot experiments, the researchers post-trained Being-H0 onto a Franka arm fitted with an Inspire six-degree-of-freedom dexterous hand for dexterous manipulation tasks.","related":["Pretraining on Human Videos","UniHand","Vision-Language-Action Model","Dexterous Manipulation","Cross-Embodiment","BeingBeyond"]},{"id":"galaxea-g0-dual-system-vla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Galaxea Open-World Dataset and G0 Dual-System VLA Model (arXiv 2509.00576)","url":"https://arxiv.org/abs/2509.00576"},{"title":"OpenGalaxea/GalaxeaVLA（GitHub）","url":"https://github.com/OpenGalaxea/GalaxeaVLA"}],"as_of":"2026-08","related_ids":["galaxea-ai","dual-system-architecture","galaxea-open-world-dataset","galaxea-r1","paligemma","vision-language-action-model"],"name":"Galaxea G0 Dual-System VLA","alt":"星海图 G0","abbr":"","aliases":["G0","G0-VLA","Galaxea G0"],"one_liner":"Galaxea's 2025 open-source fast-slow dual-system robot model, where a VLM breaks down the task and a VLA executes the actions.","explanation":"G0 is a dual-system robot foundation model from the embodied-AI company Galaxea (Galaxea AI), released in late August 2025 and open-sourced in September, launched alongside the Galaxea Open-World Dataset. It splits into a fast-slow dual system: the slow system, G0-VLM, is fine-tuned from Qwen2.5-VL and handles understanding instructions and breaking a long task into sub-tasks; the fast system, G0-VLA, has a vision-language backbone initialized from PaliGemma and outputs actions based on the current view, the robot's own state, and the sub-task instruction. Training has three stages: cross-embodiment pretraining on data from multiple robots including OXE, then pretraining on the company's own single-embodiment data, and finally post-training on specific tasks. The paper found the second stage — pretraining on real data from the same type of robot — to be the most important. The dataset was collected with the Galaxea R1 Lite across 50 real-world scenes spanning homes, retail, dining, and offices, totaling about 500 hours and 100,000 trajectories. Galaxea later released G0Plus (January 2026) and G0.5 (June 2026).","example":"When the user says “help me make the bed,” G0-VLM breaks it into a sequence of sub-task instructions, and G0-VLA drives the R1 Lite's two arms and mobile base step by step according to those instructions, until the bed is made.","related":["Galaxea AI","Dual-System Architecture (System 1 / System 2)","Galaxea Open-World Dataset","Galaxea R1","PaliGemma","Vision-Language-Action Model"]},{"id":"eo-1","category":"named_model","sec":7,"tier":3,"sources":[{"title":"EO-1: An Open Unified Embodied Foundation Model for General Robot Control (arXiv 2508.21112)","url":"https://arxiv.org/abs/2508.21112"},{"title":"EO-1 GitHub 仓库","url":"https://github.com/EO-Robotics/EO1"},{"title":"EO-1 项目页","url":"https://eo-robotics.ai/eo-1"}],"as_of":"2026-02","related_ids":["vision-language-action-model","hybrid-autoregressive-diffusion-architecture","flow-matching","embodied-reasoning","shanghai-artificial-intelligence-laboratory","lerobot"],"name":"EO-1","alt":"EO-1（EmbodiedOneVision）","abbr":"EO-1","aliases":["EmbodiedOneVision","EO1","EO-Robotics","An Open Unified Embodied Foundation Model for General Robot Control"],"one_liner":"A 3B open-source unified embodied model from Shanghai AI Lab where a single network both reasons in language and outputs actions.","explanation":"EO-1 was released in August 2025, led by Shanghai AI Lab with real-robot support from AgiBot, with weights, training code, and data all open-source. It is built on Qwen2.5-VL-3B as a decoder-only Transformer: text is generated autoregressively, token by token, while actions are generated as continuous values by denoising through flow matching, with the two trained jointly inside the same model. Its companion dataset, EO-Data1.5M, contains 1.5 million interleaved “image-text-action” samples covering physical common sense, task reasoning, spatial understanding, and manipulation trajectories, letting the model reason and act within the same stretch of context. The company reports that it surpassed the then-current open-source models on ERQA, LIBERO, SimplerEnv, and its own EO-Bench.","example":"Real-robot testing covered four robots — Franka, WidowX 250, AgiBot G-1, and the LeRobot SO-100 — on tasks including long-horizon dexterous manipulation and tasks that require reasoning before acting.","related":["Vision-Language-Action Model","Hybrid Autoregressive-Diffusion Architecture","Flow Matching","Embodied Reasoning","Shanghai Artificial Intelligence Laboratory","LeRobot"]},{"id":"internvla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy (arXiv 2510.13778)","url":"https://arxiv.org/abs/2510.13778"},{"title":"InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation (arXiv 2601.02456)","url":"https://arxiv.org/abs/2601.02456"},{"title":"InternVLA-N1 项目主页","url":"https://internrobotics.github.io/internvla-n1.github.io/"}],"as_of":"2026-07","related_ids":["vision-language-action-model","dual-system-architecture","interndata-a1","vision-and-language-navigation","shanghai-artificial-intelligence-laboratory","world-action-model"],"name":"InternVLA (Shanghai AI Laboratory)","alt":"上海AI实验室 InternVLA 系列","abbr":"","aliases":["InternVLA-M1","InternVLA-A1","InternVLA-A1.5","InternVLA-N1"],"one_liner":"A family of embodied models from Shanghai AI Lab: M1 and A1 for manipulation, N1 for navigation.","explanation":"InternVLA is a family of VLA (vision-language-action) models open-sourced by Shanghai AI Lab's embodied-AI team (GitHub organization InternRobotics). InternVLA-M1 (October 2025) is built on Qwen2.5-VL, first pretrained for spatial grounding on 2.3 million box, point, and trajectory annotations, then guided by spatial prompts to a diffusion-Transformer action expert that produces actions. InternVLA-A1 (January 2026, in 2B and 3B sizes) uses a Mixture-of-Transformers architecture that puts three experts — understanding, future-frame prediction, and action — into one model, pretrained on real-robot data, synthetic simulation data (such as InternData-A1), and human video, totaling 692 million frames. A1.5 (July 2026) switches to Qwen3.5-2B, and during training has a frozen Wan2.2 video model supervise “foresight tokens,” while inference generates no video at all. InternVLA-N1 is for navigation: a slow system marks waypoints on the image, and a fast system uses a diffusion policy to output trajectories at more than 30Hz.","example":"When InternVLA-A1 sorts items off a moving conveyor belt, it has to predict where an object will be an instant later before grasping it; the current version of the paper reports it beats π0.5 by 26.7% on this kind of dynamic task.","related":["Vision-Language-Action Model","Dual-System Architecture (System 1 / System 2)","InternData-A1","Vision-and-Language Navigation","Shanghai Artificial Intelligence Laboratory","World Action Model"]},{"id":"wall-a","category":"named_model","sec":7,"tier":3,"sources":[{"title":"自变量机器人官网","url":"https://www.x2robot.com/"},{"title":"字节、阿里、美团首次在具身智能「同框」，十亿级融资背后，自变量到底凭什么？（智东西，2026-01-12）","url":"https://zhidx.com/p/528318.html"},{"title":"Igniting VLMs toward the Embodied Space (WALL-OSS, arXiv 2509.11766)","url":"https://arxiv.org/abs/2509.11766"}],"as_of":"2026-09","related_ids":["x-square-robot","wall-oss","vision-language-action-model","world-model","embodied-foundation-model","mixture-of-experts"],"name":"WALL-A","alt":"自变量 WALL-A","abbr":"","aliases":["WALL-A Manipulation Foundation Model","WALL-A series"],"one_liner":"X Square Robot's in-house, closed-source embodied manipulation model series, running end-to-end from perception to motor control.","explanation":"WALL-A is a series of in-house manipulation foundation models built by Shenzhen-based embodied AI company X Square Robot (自变量机器人); it is not open-sourced. The company's website describes it as achieving 'unified intelligence across the full pipeline, from perception and understanding to action control' — in other words, an end-to-end vision-language-action approach. According to a January 2026 report from Zhidongxi, the WALL-A series combines a VLA with a world model, using the world model to predict how the environment's state changes over time and reasoning jointly with vision to decide on actions. The WALL-OSS paper, open-sourced in September 2025, describes WALL-A's architecture as a tightly coupled design with shared self-attention and separate feed-forward layers per modality, and WALL-OSS follows the same idea, making it viewable as an open-source version of the WALL-A approach. In April 2026 the company released a further 'world unified model' called WALL-B, and its later public demonstrations — housework, logistics sorting — have mostly been driven by WALL-B. WALL-A's initial release date and parameter count could not be verified from public sources for this entry.","example":"According to Zhidongxi's report, X Square Robot's 'Quantum No. 1' used its models to demonstrate mobile manipulation across both outdoor and indoor settings: breaking down and recycling a takeout box outdoors, then making its own way through a building door and elevator to deliver it indoors.","related":["X Square Robot","WALL-OSS","Vision-Language-Action Model","World Model","Embodied Foundation Model","Mixture of Experts"]},{"id":"wall-oss","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Igniting VLMs toward the Embodied Space (arXiv:2509.11766)","url":"https://arxiv.org/abs/2509.11766"},{"title":"X-Square-Robot/wall-x (GitHub)","url":"https://github.com/X-Square-Robot/wall-x"},{"title":"x-square-robot/wall-oss-0.5 (Hugging Face)","url":"https://huggingface.co/x-square-robot/wall-oss-0.5"}],"as_of":"2026-06","related_ids":["x-square-robot","wall-a","vision-language-action-model","embodied-chain-of-thought","flow-matching","qwen-vl"],"name":"WALL-OSS","alt":"自变量 WALL-OSS","abbr":"","aliases":["WALL-OSS 0.5","WALL-OSS-FLOW","WALL-OSS-FAST"],"one_liner":"X Square Robot's open-source embodied foundation model that turns a vision-language model into a VLA that outputs actions directly.","explanation":"WALL-OSS is an embodied foundation model open-sourced by X Square Robot (自变量机器人) in September 2025, described in the paper 'Igniting VLMs toward the Embodied Space.' It uses Qwen2.5-VL-3B as its backbone, with a tightly coupled mixture-of-experts structure (different feed-forward layers assigned to different objectives depending on the training task) that connects 'instruction to reasoning to sub-task planning to continuous action' inside one differentiable model. Training first grounds the model with embodied question-answering and FAST discrete-action data, then teaches continuous actions using flow matching, on more than 10,000 hours of data. Two versions were released, FLOW and FAST, under the Apache-2.0 license. In May 2026 the company released WALL-OSS 0.5 (about 4 billion parameters), built to deploy directly from its pretrained weights with no fine-tuning required.","example":"Loading the WALL-OSS-FLOW weights through the official wall-x codebase, a researcher fine-tunes it on their own LeRobot-format data for a tabletop tidying task.","related":["X Square Robot","WALL-A","Vision-Language-Action Model","Embodied Chain-of-Thought","Flow Matching","Qwen-VL"]},{"id":"unifolm","category":"named_model","sec":7,"tier":3,"sources":[{"title":"unitreerobotics/unifolm-world-model-action (GitHub)","url":"https://github.com/unitreerobotics/unifolm-world-model-action"},{"title":"unitreerobotics/unifolm-vla (GitHub)","url":"https://github.com/unitreerobotics/unifolm-vla"},{"title":"unitreerobotics/unifolm-wla (GitHub)","url":"https://github.com/unitreerobotics/unifolm-wla"}],"as_of":"2026-09","related_ids":["unitree-robotics","unitree-g1","world-action-model","vision-language-action-model","unitree-z1-robotic-arm","open-weight-model"],"name":"UnifoLM","alt":"宇树 UnifoLM 系列","abbr":"","aliases":["UnifoLM-WMA-0","UnifoLM-VLA-0","UnifoLM-WLA-1.0"],"one_liner":"Unitree's open-source series of robot foundation models, spanning a world model, a VLA, and a general-purpose humanoid model.","explanation":"UnifoLM is a series of open-source robot foundation models from Unitree Robotics, with three main releases so far. UnifoLM-WMA-0 (September 2025) uses a world-model-plus-action-head architecture: the world model predicts the future images that result from the robot interacting with the environment, which can either assist an action head with decision-making or serve as an interactive simulator to generate training data; it supports the Z1 arm and the G1 humanoid. UnifoLM-VLA-0 (January 2026) is based on Qwen2.5-VL, continues pretraining on manipulation data, and uses a single policy to perform 12 categories of manipulation tasks on the G1. UnifoLM-WLA-1.0 (September 2026) is a 6-billion-parameter general-purpose humanoid foundation model that the company says was trained on about 2,500 hours of real-robot data, covering 64 tasks across tabletop and whole-body manipulation, supporting both two-finger grippers and several five-fingered dexterous hands, open-sourced under Apache 2.0. This series reflects Unitree's expansion from selling hardware into open-source models as well.","example":"UnifoLM-WMA-0's world model can predict the upcoming images as a Unitree Z1 arm stacks boxes, which can both help choose the next action and generate synthetic training data.","related":["Unitree Robotics","Unitree G1","World Action Model","Vision-Language-Action Model","Unitree Z1 Robotic Arm","Open-weight Model"]},{"id":"wow","category":"named_model","sec":7,"tier":3,"sources":[{"title":"WoW (arXiv:2509.22642)","url":"https://arxiv.org/abs/2509.22642"},{"title":"WoW 项目主页","url":"https://wow-world-model.github.io/"}],"as_of":"2025-10","related_ids":["world-model","video-generation-model","inverse-dynamics-model","diffusion-transformer","beijing-humanoid-robot-innovation-center","xr-1"],"name":"WoW","alt":"WoW 具身世界模型","abbr":"WoW","aliases":["WoW-DiT","WoW: Towards a World omniscient World model Through Embodied Interaction"],"one_liner":"A 14-billion-parameter embodied video world model trained on 2 million real-robot interaction trajectories.","explanation":"WoW was released in September 2025 by the Beijing Humanoid Robot Innovation Center together with Peking University and HKUST. The authors argue that watching passive video alone cannot teach real physical intuition, which instead requires learning from large amounts of causally connected interaction; they trained a 14-billion-parameter diffusion Transformer video generation model on 2 million real robot trajectories spanning 12 robot platforms. Because generated video often contains physical errors, the authors use a framework called SOPHIA, which has a vision-language model act as a reviewer, repeatedly checking the output and rewriting the prompt to steer generation toward something more physically plausible; an inverse dynamics model (which infers the action from consecutive frames) then translates the predicted video into robot-arm actions. The team also released WoWBench, a benchmark for evaluating physical consistency and causal reasoning.","example":"Given a tabletop image and a manipulation instruction, WoW first generates a video of the action being completed, and then an inverse dynamics module translates that video into the end effector's 7-degree-of-freedom actions for execution.","related":["World Model","Video Generation Model","Inverse Dynamics Model","Diffusion Transformer","Beijing Humanoid Robot Innovation Center","XR-1"]},{"id":"pelican-vl","category":"named_model","sec":7,"tier":3,"sources":[{"title":"arXiv 2511.00108: Pelican-VL 1.0","url":"https://arxiv.org/abs/2511.00108"},{"title":"Hugging Face: X-Humanoid/Pelican1.0-VL-7B","url":"https://huggingface.co/X-Humanoid/Pelican1.0-VL-7B"},{"title":"Hugging Face: X-Humanoid 组织页","url":"https://huggingface.co/X-Humanoid"}],"as_of":"2026-09","related_ids":["embodied-reasoning-model","vision-language-model","qwen-vl","group-relative-policy-optimization","braincerebellum-architecture","beijing-humanoid-robot-innovation-center"],"name":"Pelican-VL","alt":"北京人形 Pelican-VL","abbr":"","aliases":["Pelican-VL 1.0","Pelican1.0-VL","Pelican-VL: A Foundation Brain Model for Embodied Intelligence"],"one_liner":"An open-source vision-language 'brain' model for embodied robots, released by China's X-Humanoid center.","explanation":"Pelican-VL is an embodied 'brain' model from the Beijing Humanoid Robot Innovation Center (X-Humanoid). The technical report came out in late October 2025, with weights open-sourced in November 2025; it is built on Qwen2.5-VL, and the first release shipped 7B and 72B parameter versions under the Apache 2.0 license. Rather than outputting joint-level actions directly, Pelican-VL handles spatial understanding, affordance reasoning, task planning, and function calling, and issues instructions to a lower-level controller or a vision-language-action (VLA) model. Its training method, DPPO (Deliberate Practice Policy Optimization), first uses GRPO (Group Relative Policy Optimization) reinforcement learning to find cases the model handles poorly, turns those hard cases into supervised fine-tuning data, and alternates between RL and SFT in a loop. The report claims a 20.3% average improvement over the base model, and leads open-source models of similar size on embodied benchmarks such as Where2Place and RefSpatialBench. Its Hugging Face page later added 3B and 235B-A22B versions as well.","example":"In one real-robot experiment from the report, the model predicts and continuously corrects grip force from the camera feed to complete a contact-rich grasp; in another, it acts as a unified 'brain' coordinating several different robots on a long-horizon task.","related":["Embodied Reasoning Model","Vision-Language Model","Qwen-VL","Group Relative Policy Optimization","Brain–Cerebellum Architecture","Beijing Humanoid Robot Innovation Center"]},{"id":"xr-1","category":"named_model","sec":7,"tier":3,"sources":[{"title":"XR-1 (arXiv:2511.02776)","url":"https://arxiv.org/abs/2511.02776"},{"title":"XR-1 论文 HTML 版","url":"https://arxiv.org/html/2511.02776"}],"as_of":"2026-09","related_ids":["cross-embodiment","vector-quantized-variational-autoencoder","latent-action","beijing-humanoid-robot-innovation-center","tiangong","pi0"],"name":"XR-1","alt":"北京人形 XR-1","abbr":"XR-1","aliases":["X-Humanoid XR-1","XR-1-Light"],"one_liner":"A VLA foundation model from X-Humanoid that pretrains across robot embodiments using a 'unified vision-motion encoding.'","explanation":"XR-1 was released by the Beijing Humanoid Robot Innovation Center (X-Humanoid) in November 2025 (some authors also hold positions at Peking University and Beihang University), with the paper accepted as an oral presentation at ICML 2026. Its core idea is UVMC (Unified Vision-Motion Coding): a dual-branch vector-quantized variational autoencoder encodes 'how the image changes' and 'how the robot moves' into the same discrete codebook, with an alignment loss keeping the two consistent. Training has three stages: self-supervised learning of UVMC; pretraining on about 164 million frames of cross-embodiment data (Open-X, the team's own XR-D, RoboMIND, and Ego4D human video); and finally task-specific fine-tuning. The main model reuses the π0 architecture, with a lighter version, XR-1-Light, based on Florence-2. Code and weights are open-sourced on GitHub, Hugging Face, and ModelScope.","example":"The authors ran more than 14,000 real-robot trials across six embodiments — Tiangong 1.0/2.0, single- and dual-arm UR5e, dual-arm Franka, and AgileX Cobot Magic — covering more than 120 manipulation tasks.","related":["Cross-Embodiment","Vector-Quantized Variational Autoencoder","Latent Action","Beijing Humanoid Robot Innovation Center","Tiangong","π0"]},{"id":"gigabrain-0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GigaBrain-0 (arXiv:2510.19430)","url":"https://arxiv.org/abs/2510.19430"},{"title":"GigaBrain-0.5M* (arXiv:2602.12099)","url":"https://arxiv.org/abs/2602.12099"},{"title":"GigaBrain-0.7 (arXiv:2608.15875)","url":"https://arxiv.org/abs/2608.15875"}],"as_of":"2026-08","related_ids":["gigaai","gigaworld-0","vision-language-action-model","embodied-chain-of-thought","knowledge-insulation","synthetic-data"],"name":"GigaBrain-0","alt":"极佳 GigaBrain-0","abbr":"","aliases":["GigaBrain","GigaBrain-0.5","GigaBrain-0.5M*","GigaBrain-0.7","A World Model-Powered Vision-Language-Action Model"],"one_liner":"A VLA model series from GigaAI trained mainly on data generated by a world model.","explanation":"GigaBrain-0 is a vision-language-action (VLA) model released by GigaAI in October 2025. To address the fact that real-robot data is expensive and scarce, it replaces a large share of training data with data generated by a world model, including video generation, appearance transfer, human-video transfer, viewpoint transfer, and sim-to-real transfer. The model takes RGB-D images as input to strengthen spatial awareness, first produces an embodied chain of thought (writing out intermediate reasoning in language) before outputting an action, and uses knowledge insulation to keep the reasoning training from interfering with the action output. There is also a GigaBrain-0-Small variant that can run on a Jetson AGX Orin. Later versions followed: GigaBrain-0.5, trained on more than 10,000 hours of real-robot data; 0.5M*, which adds world-model-based reinforcement learning; and 0.7, from August 2026, which uses more than 37,000 hours of data and a three-system architecture.","example":"The paper gradually raised the share of world-model-generated data in the training mix from 0% to 90%, and the model's generalization to changes in object appearance, placement, and camera viewpoint improved markedly as a result.","related":["GigaAI","GigaWorld-0","Vision-Language-Action Model","Embodied Chain-of-Thought","Knowledge Insulation","Synthetic Data"]},{"id":"gigaworld-0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"GigaWorld-0 (arXiv:2511.19861)","url":"https://arxiv.org/abs/2511.19861"},{"title":"GigaWorld-0 项目主页","url":"https://giga-world-0.github.io/"},{"title":"GigaWorld-1 (arXiv:2607.02642)","url":"https://arxiv.org/abs/2607.02642"}],"as_of":"2026-07","related_ids":["gigabrain-0","gigaai","world-model","synthetic-data","3d-gaussian-splatting","world-model-based-policy-evaluation"],"name":"GigaWorld-0","alt":"极佳 GigaWorld-0","abbr":"","aliases":["GigaWorld-0-Video","GigaWorld-0-3D","World Models as Data Engine to Empower Embodied AI"],"one_liner":"GigaAI's framework that uses a world model as a “data engine” to mass-produce robot training data.","explanation":"GigaWorld-0 was released by GigaAI in November 2025, positioned as a world-model framework for producing training data for vision-language-action (VLA) models. It has two parts: GigaWorld-0-Video uses a large-scale video-generation model to produce texture-rich robot manipulation videos, with fine-grained control over appearance, camera viewpoint, and action semantics; GigaWorld-0-3D combines 3D generation, 3D Gaussian splatting reconstruction, differentiable physics system identification, and motion planning to keep results geometrically consistent and physically plausible. Its companion training framework, GigaTrain, uses FP8 precision and sparse attention to save compute. GigaBrain-0, trained on this synthetic data, performs well on real robots. Code and models are open-source; in July 2026 the team also released GigaWorld-1, aimed at policy evaluation.","example":"Take a real-robot manipulation video and change the tabletop texture, lighting, or camera viewpoint, generating multiple versions of training data that look different but share the same action, used to improve a VLA's generalization to appearance changes.","related":["GigaBrain-0","GigaAI","World Model","Synthetic Data","3D Gaussian Splatting","World-Model-based Policy Evaluation"]},{"id":"mimo-embodied","category":"named_model","sec":7,"tier":3,"sources":[{"title":"MiMo-Embodied: X-Embodied Foundation Model Technical Report (arXiv 2511.16518)","url":"https://arxiv.org/abs/2511.16518"},{"title":"XiaomiMiMo/MiMo-Embodied (GitHub)","url":"https://github.com/XiaomiMiMo/MiMo-Embodied"}],"as_of":"2026-04","related_ids":["vision-language-model","embodied-reasoning","affordance","spatial-reasoning","autonomous-driving","cross-embodiment"],"name":"MiMo-Embodied (Xiaomi)","alt":"小米 MiMo-Embodied","abbr":"","aliases":["MiMo-Embodied-7B","X-Embodied Foundation Model"],"one_liner":"An open-source 7B vision-language model from Xiaomi covering both autonomous-driving and embodied-AI understanding and planning.","explanation":"MiMo-Embodied is a vision-language model (VLM) whose technical report was released, and weights open-sourced, by Xiaomi's embodied-AI team in November 2025; it has about 7 billion parameters, with weights public on Hugging Face. The premise is that autonomous driving and indoor robots both need spatial understanding and planning, yet are usually trained entirely separately. MiMo-Embodied puts both kinds of data into a single model, trained through multiple stages with curated data plus chain-of-thought and reinforcement-learning fine-tuning, covering affordance prediction, task planning, and spatial understanding on the embodied side, and environment perception, state prediction, and driving planning on the driving side. The report says it matches or beats comparable open- and closed-source models across 17 embodied-AI benchmarks and 12 autonomous-driving benchmarks, and observes positive transfer between the two domains, each reinforcing the other. It evaluates understanding, reasoning, and planning ability; it does not directly output a robot's joint actions itself.","example":"Show it a kitchen photo and ask “where should the cup be grasped,” and it can point out the graspable location on the image; show it driving footage, and it can describe the state of surrounding vehicles and give a next driving plan.","related":["Vision-Language Model","Embodied Reasoning","Affordance","Spatial Reasoning","Autonomous Driving","Cross-Embodiment"]},{"id":"xiaomi-robotics-0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Xiaomi-Robotics-0 (arXiv:2602.12684)","url":"https://arxiv.org/abs/2602.12684"},{"title":"Xiaomi-Robotics-0 技术报告 HTML 版","url":"https://arxiv.org/html/2602.12684"},{"title":"Xiaomi-Robotics-1 (arXiv:2607.15330)","url":"https://arxiv.org/abs/2607.15330"}],"as_of":"2026-07","related_ids":["vision-language-action-model","asynchronous-inference","real-time-chunking","action-expert","xiaomi","mimo-embodied"],"name":"Xiaomi-Robotics-0","alt":"小米 Xiaomi-Robotics-0","abbr":"","aliases":["Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution"],"one_liner":"Xiaomi's open-source, 4.7-billion-parameter VLA, built to execute actions in real time and smoothly on a consumer GPU.","explanation":"Xiaomi-Robotics-0 is a vision-language-action model released and open-sourced by Xiaomi in February 2026, with 4.7 billion parameters in total, using Qwen3-VL-4B as its vision-language backbone and attaching a diffusion Transformer action expert that generates actions via flow matching. Pretraining used roughly 200 million timesteps of cross-embodiment robot trajectories (from DROID, MolmoAct data, and Xiaomi's own collected data), plus more than 80 million vision-language samples. Its main focus is fixing the action stutter caused by slow inference: it uses asynchronous execution, inferring the next action segment while still executing the previous one, and feeds already-committed actions back into the model as a prefix, paired with a Λ-shaped attention mask to keep the transition smooth; on an RTX 4090, inference latency is about 80 milliseconds. In July 2026, Xiaomi followed up with Xiaomi-Robotics-1, trained on more than 100,000 hours of real-robot data.","example":"The company reports 98.7% average success on LIBERO and an average of 4.75 consecutively completed tasks on CALVIN ABC→D; on a real robot, it demonstrated two bimanual tasks — disassembling LEGO and folding a towel.","related":["Vision-Language-Action Model","Asynchronous Inference","Real-Time Chunking","Action Expert","Xiaomi","MiMo-Embodied (Xiaomi)"]},{"id":"kairos","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI (arXiv 2606.16533，v1 题为 A Native World Model Stack for Physical AI)","url":"https://arxiv.org/abs/2606.16533"},{"title":"kairos-agi/kairos (GitHub)","url":"https://github.com/kairos-agi/kairos"},{"title":"大晓机器人官网","url":"https://www.acerobotics.com"}],"as_of":"2026-07","related_ids":["world-model","world-action-model","embodied-foundation-model","video-generation-model","ace-robotics","robotwin"],"name":"Kairos (ACE Robotics)","alt":"大晓机器人 开悟世界模型","abbr":"","aliases":["A Native World Model Stack for Physical AI","A Regret-Aware Native World-Action Model Stack for Physical AI"],"one_liner":"A 4-billion-parameter embodied world model from ACE Robotics that does understanding, video generation, and action prediction all at once.","explanation":"Kairos is the world model of Shanghai-based embodied-AI company ACE Robotics (Daxiao Robotics); its technical report lists Xiaogang Wang and Dacheng Tao among the authors, and the company is reportedly led by Xiaogang Wang, a co-founder of SenseTime. Kairos 3.0 was released in December 2025 with 4-billion-parameter pretrained weights open-sourced; a technical report followed in June 2026, and version 3.1 plus world-action-model inference code was open-sourced in July. Its central claim is that a world model doesn't need to render every pixel photorealistically — what matters is retaining the information useful for control, such as object state, contact, task progress, and the consequences of an action. Concretely, it orders training data in stages — ordinary video first, then human behavior, then robot interaction — uses a unified architecture for understanding, generation, and prediction all at once, and uses hybrid linear temporal attention to cut the cost of long-horizon inference.","example":"The open-sourced kairos-4B-robot-RoboTwin2.0 weights jointly predict future frames and actions across more than 50 bimanual tasks on RoboTwin 2.0, achieving what the company says is the best result yet on that benchmark.","related":["World Model","World Action Model","Embodied Foundation Model","Video Generation Model","ACE Robotics","RoboTwin"]},{"id":"ubtech-thinker","category":"named_model","sec":7,"tier":3,"sources":[{"title":"UBTECH-Robot/Thinker (GitHub)","url":"https://github.com/UBTECH-Robot/Thinker"},{"title":"Thinker: A vision-language foundation model for embodied intelligence (arXiv 2601.21199)","url":"https://arxiv.org/abs/2601.21199"},{"title":"Embodied Intelligence 2026: Farewell to Narrative Hype, Practical Deployment Reigns Supreme (36Kr)","url":"https://eu.36kr.com/en/p/3953394550537606"}],"as_of":"2026-08","related_ids":["ubtech-robotics","ubtech-walker-s2","vision-language-model","world-model","vision-language-action-model","qwen-vl"],"name":"UBTech Thinker","alt":"优必选 Thinker","abbr":"","aliases":["Thinker-4B","Thinker-WM","Thinker-VLA"],"one_liner":"UBTech's in-house embodied model stack, made up of a foundation model, a world model, and an action model.","explanation":"Thinker is the in-house embodied model system UBTech has built for its humanoid robots, reportedly organized into three layers: a foundation model called Thinker, a world model called Thinker-WM, and an action model called Thinker-VLA. Thinker itself is a vision-language model for embodied intelligence; UBTech open-sourced Thinker-4B (based on the Qwen3-VL architecture, 4B parameters, non-commercial license) along with a paper in January 2026. It targets problems that arise when a general vision-language model is applied to robots, such as viewpoint confusion and weak temporal understanding, and is trained on first-person video, visual grounding, spatial understanding, and chain-of-thought data, focused on task planning, visual grounding, and spatial understanding; the company states it leads on 7 embodied benchmarks. According to reports, Thinker-WM ranks first on the LIBERO benchmark, and Thinker-VLA raises inference efficiency by 176% in industrial settings. This model stack powers UBTech's industrial humanoid robots, including the Walker S series.","example":"Given a first-person image from a robot and the instruction 'put the part in the bin,' Thinker-4B can output a step-by-step task plan and also draw a box around the target part's location in the image.","related":["UBTech Robotics","UBTech Walker S2","Vision-Language Model","World Model","Vision-Language-Action Model","Qwen-VL"]},{"id":"spirit-v1-5","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Spirit-v1.5: Clean Data Is the Enemy of Great Robot Foundation Models (Spirit AI Blog)","url":"https://www.spirit-ai.com/en/blog/spirit-v1-5"},{"title":"Spirit-AI-Team/spirit-v1.5 (GitHub)","url":"https://github.com/Spirit-AI-Team/spirit-v1.5"},{"title":"Spirit-AI-robotics/Spirit-v1.5 (Hugging Face)","url":"https://huggingface.co/Spirit-AI-robotics/Spirit-v1.5"}],"as_of":"2026-04","related_ids":["spirit-ai","vision-language-action-model","robochallenge","qwen-vl","diffusion-transformer","data-diversity"],"name":"Spirit v1.5","alt":"千寻 Spirit v1.5","abbr":"","aliases":["Spirit-v1.5","Spirit AI Spirit v1.5"],"one_liner":"An open-source VLA foundation model from Spirit AI, built on the premise that messy, unscripted data makes a better pretraining set.","explanation":"Spirit v1.5 is a vision-language-action (VLA) model released and open-sourced by Chinese startup Spirit AI (千寻智能) in January 2026, with the inference code (MIT license), base weights, and one fine-tuned checkpoint (Apache 2.0 license) made public, followed by fine-tuning code in April. Its architecture is the common 'VLM plus action head' pattern: Qwen3-VL-4B as the vision-language backbone, followed by a diffusion Transformer (DiT) action head that generates continuous actions. Its central claim, stated in the title of its technical blog post, is that 'clean data is the enemy of a good robot foundation model': rather than giving data collectors a fixed script or staging objects carefully, they are simply given a goal and left to complete a chain of real tasks freely, so the resulting data naturally contains failed retries and task switching. The team states this raised effective per-collector data-collection time by 200%. At release, it ranked first on RoboChallenge's Table30 real-robot leaderboard.","example":"A data collector sets themselves a goal, such as 'use the robot to mix a drink today,' and the whole session — preparing ingredients, noticing a wrong ratio, adjusting, adding or removing ingredients — is recorded continuously and used for pretraining.","related":["Spirit AI","Vision-Language-Action Model","RoboChallenge","Qwen-VL","Diffusion Transformer","Data Diversity"]},{"id":"lingbot-vla","category":"named_model","sec":7,"tier":3,"sources":[{"title":"arXiv 2601.18692: A Pragmatic VLA Foundation Model","url":"https://arxiv.org/abs/2601.18692"},{"title":"Robbyant 官网：LingBot-VLA","url":"https://technology.robbyant.com/lingbot-vla"},{"title":"arXiv 2607.06403: From Foundation to Application (LingBot-VLA 2.0)","url":"https://arxiv.org/abs/2607.06403"}],"as_of":"2026-07","related_ids":["vision-language-action-model","qwen-vl","cross-embodiment","post-training","pi0","robbyant"],"name":"LingBot-VLA (Robbyant)","alt":"蚂蚁灵波 LingBot-VLA","abbr":"","aliases":["LingBot-VLA 2.0","A Pragmatic VLA Foundation Model"],"one_liner":"Robbyant's open-source VLA foundation model, pretrained on about 20,000 hours of bimanual real-robot data across 9 arm configurations.","explanation":"LingBot-VLA is a vision-language-action (VLA) foundation model open-sourced by Robbyant in January 2026, in a paper titled A Pragmatic VLA Foundation Model, emphasizing practicality: strong generalization and low data and compute cost when adapting to a new platform. It uses Qwen2.5-VL-3B as its vision-language backbone (also supporting PaliGemma), pretrained on about 20,000 hours of real-robot data across 9 mainstream bimanual configurations, and systematically evaluated with 100 tasks each (GM-100) on 4 platforms including AgiBot G1, AgileX, and Galaxea R1 Pro. Its companion training code reaches a throughput of 261 samples per second on 8 GPUs, 1.5 to 2.8 times faster than existing VLA codebases. The July 2026 version 2.0 expanded the data to about 60,000 hours (including 10,000 hours of human first-person video), and extended the action space to the head, waist, base, and dexterous hands.","example":"In the GM-100 evaluation, each task gets only 130 post-training demonstrations, and success rate is then compared against other VLA models on the same real-robot platform.","related":["Vision-Language-Action Model","Qwen-VL","Cross-Embodiment","Post-training","π0","Robbyant"]},{"id":"lingbot-va","category":"named_model","sec":7,"tier":3,"sources":[{"title":"arXiv 2601.21998: Causal World Modeling for Robot Control","url":"https://arxiv.org/abs/2601.21998"},{"title":"GitHub: Robbyant/lingbot-va","url":"https://github.com/Robbyant/lingbot-va"},{"title":"arXiv 2607.08639: Native Video-Action Pretraining for Generalizable Robot Control (LingBot-VA 2.0)","url":"https://arxiv.org/abs/2607.08639"}],"as_of":"2026-07","related_ids":["world-action-model","world-model","mixture-of-transformers","asynchronous-inference","autoregressive-video-generation","robbyant"],"name":"LingBot-VA (Robbyant)","alt":"蚂蚁灵波 LingBot-VA","abbr":"","aliases":["LingBot-VA 2.0","Causal World Modeling for Robot Control"],"one_liner":"Robbyant's open-source video-action world model that predicts future frames and outputs robot actions at the same time.","explanation":"LingBot-VA is a robot control model open-sourced in January 2026 by Robbyant, the embodied-AI company under Ant Group, in a paper titled Causal World Modeling for Robot Control, accepted to RSS 2026. It follows the “world action model” approach: it generates segment by segment with autoregressive diffusion, alternately predicting future video frames and actions within the same sequence; visual and action tokens share a latent space, processed by a Mixture-of-Transformers (MoT). At execution time, each segment is corrected against real observations (closed-loop rollout), and action prediction runs asynchronously in parallel with motor execution to cut latency. Its video encoding reuses the VAE from Tongyi Wanxiang's Wan2.2. Version 2.0, released in July 2026, switched to causal pretraining from scratch with a sparse MoE backbone.","example":"The company reports success rates of 92.9% (easy) and 91.6% (hard) across 50 bimanual simulation tasks on RoboTwin 2.0, and an average of 98.5% on LIBERO.","related":["World Action Model","World Model","Mixture-of-Transformers","Asynchronous Inference","Autoregressive Video Generation","Robbyant"]},{"id":"lingbot-world","category":"named_model","sec":7,"tier":3,"sources":[{"title":"arXiv 2601.20540: Advancing Open-source World Models","url":"https://arxiv.org/abs/2601.20540"},{"title":"GitHub: Robbyant/lingbot-world","url":"https://github.com/Robbyant/lingbot-world"},{"title":"arXiv 2607.07534: Infinite Worlds with Versatile Interactions (LingBot-World 2.0)","url":"https://arxiv.org/abs/2607.07534"}],"as_of":"2026-07","related_ids":["world-model","interactive-world-model","genie-3","wan","autoregressive-video-generation","robbyant"],"name":"LingBot-World (Robbyant)","alt":"蚂蚁灵波 LingBot-World","abbr":"","aliases":["LingBot-World 2.0","LingBot-World-Infinity","Advancing Open-source World Models"],"one_liner":"Robbyant's open-source interactive world model that generates the next frame in real time from keyboard and camera commands.","explanation":"LingBot-World is a world model Robbyant open-sourced in January 2026, adapted from the Tongyi Wanxiang Wan2.2 video-generation model. Given a starting image or text description, the user controls it step by step with keyboard actions or camera pose, and the model generates the corresponding footage — the paper calls this a world simulator. The company reports it keeps a scene consistent over minutes-long durations (called long-term memory), and interacts in real time at 16 frames per second with under 1 second of latency; code and weights are released under the Apache 2.0 license, aimed at narrowing the gap between open- and closed-source world models, for content creation, games, and robot learning. Version 2.0 (LingBot-World-Infinity), from July 2026, supports interaction of unlimited duration, with a distilled real-time version driving 720p, 60fps video, offered at 14B and 1.3B sizes.","example":"Given an indoor photo as a starting point, pressing W moves forward and A/D turns, and the model generates the following footage from the corresponding viewpoint, segment by segment.","related":["World Model","Interactive World Model","Genie 3","Wan (Alibaba Video Generation Model)","Autoregressive Video Generation","Robbyant"]},{"id":"dm0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI (arXiv 2602.14974)","url":"https://arxiv.org/abs/2602.14974"},{"title":"Dexmal/DM05 模型卡（Hugging Face）","url":"https://huggingface.co/Dexmal/DM05"},{"title":"Dexbotic 代码仓库（GitHub）","url":"https://github.com/Dexmal/dexbotic"}],"as_of":"2026-09","related_ids":["vision-language-action-model","action-expert","flow-matching","mid-training","robochallenge","dexbotic"],"name":"DM0","alt":"原力灵机 DM0","abbr":"","aliases":["Dexmal DM0","DM0.5","An Embodied-Native Vision-Language-Action Model towards Physical AI"],"one_liner":"An “embodiment-native” VLA model open-sourced by Dexmal in 2026, pretrained from the start on a mix of driving and robot data.","explanation":"DM0 is a vision-language-action (VLA) model released and open-sourced in February 2026 by Dexmal (原力灵机) together with StepFun (阶跃星辰). Most VLA models start from a vision-language model that has only ever seen internet images and text, then fine-tune it on robot data; DM0 instead follows an “embodiment-native” philosophy, training from the start on a mix of web text, autonomous-driving scenes, and robot interaction logs — about 1.2 trillion tokens in all. Training runs in three stages: pretraining a unified VLM (its language component is based on Qwen3-1.7B), mid-training that adds a flow-matching action expert (a sub-network dedicated to producing continuous actions) on top of it, and post-training that fine-tunes with a mixed strategy. When training on embodied data, gradients from the action expert are not backpropagated into the VLM, to preserve its general understanding ability; the model also uses a spatial chain of thought, reasoning about spatial relationships before producing an action. It has about 2 billion parameters and supports both manipulation and navigation in one model. In July 2026, Dexmal released a roughly 6-billion-parameter successor, DM0.5, along with an accompanying open-source framework called OpenDM.","example":"According to the paper, on the RoboChallenge Table30 real-robot benchmark, DM0 reached an average success rate of 62.0% in the task-specific setting and 37.3% in the general setting, ranking first on both at the time of release.","related":["Vision-Language-Action Model","Action Expert","Flow Matching","Mid-training","RoboChallenge","Dexbotic (Dexmal VLA toolbox)"]},{"id":"holobrain-0","category":"named_model","sec":7,"tier":3,"sources":[{"title":"HoloBrain-0 Technical Report (arXiv 2602.12062)","url":"https://arxiv.org/abs/2602.12062"},{"title":"HoloBrain-0 技术报告 HTML 全文","url":"https://arxiv.org/html/2602.12062v1"}],"as_of":"2026-02","related_ids":["vision-language-action-model","cross-embodiment","unified-robot-description-format","robotwin","on-device-edge-deployment","horizon-robotics"],"name":"HoloBrain-0","alt":"地平线 HoloBrain-0","abbr":"HoloBrain","aliases":["Horizon Robotics HoloBrain","HoloBrain-0 Technical Report"],"one_liner":"A VLA framework Horizon Robotics open-sourced in early 2026 that feeds camera parameters and robot structure into the model as priors.","explanation":"HoloBrain-0 was released as a technical report and open-sourced by the Horizon Robotics team in February 2026. Its core design explicitly feeds robot-embodiment priors into the VLA: the parameters of multiple cameras, and a URDF file describing the robot's joint structure, used to strengthen 3D spatial reasoning and adapt across different embodiments. Training follows a “pretrain then post-train” route: pretraining data comes from real robots — a bimanual Piper, AgiBot G1, and Franka — plus a bimanual UR5 and bimanual ARX in simulation, and the human-hand video dataset EgoDex, followed by post-training on specific tasks. The report says it achieves state-of-the-art results on the RoboTwin 2.0, LIBERO, and GenieSim simulation benchmarks, and also performs well on long-horizon real-robot tasks. What's open-sourced includes the pretrained model, post-training checkpoints for various simulation suites and real-robot tasks, and RoboOrchard, a full-stack toolchain covering data curation, training, and deployment.","example":"It has a lightweight, roughly 200-million-parameter (0.2B) version based on GroundingDINO-Tiny that performs comparably to much larger baselines with low inference latency, suited to deploying directly on a robot's onboard chip; there's also a roughly 1.1-billion-parameter version based on Qwen2.5-VL-3B.","related":["Vision-Language-Action Model","Cross-Embodiment","Unified Robot Description Format","RoboTwin","On-Device / Edge Deployment","Horizon Robotics"]},{"id":"tars-robotics-awe","category":"named_model","sec":7,"tier":3,"sources":[{"title":"丁文超博士发布通用具身大模型 AWE3.0（中国日报网，推广信息）","url":"https://cn.chinadaily.com.cn/a/202603/17/WS69b8fa3ba310942cc49a396b.html"},{"title":"「能干活」的通用具身大模型 AWE3.0 亮相（新华网客户端）","url":"https://app.xinhuanet.com/news/article.html?articleId=3b05aef1a3ea11cbc6b2f7ab79e329ce"},{"title":"它石智航@WAIC 2026：具身原生基座模型 AWE 摘得 SAIL 之星（知乎专栏）","url":"https://zhuanlan.zhihu.com/p/2063322866691657972"}],"as_of":"2026-07","related_ids":["tars-robotics","world-in-your-hands","human-video-data","latent-action","embodied-foundation-model","visuo-tactile-fusion"],"name":"TARS Robotics AWE","alt":"它石智航 AWE","abbr":"AWE","aliases":["AWE (AI World Engine)","AWE 3.0","AWE 3.5"],"one_liner":"TARS Robotics' general-purpose embodied foundation model, whose main selling point is training on large-scale human manipulation data.","explanation":"AWE is the general-purpose embodied foundation model from TARS Robotics (它石智航), a Shanghai embodied-AI company reportedly founded in February 2025; the company describes it as built on a proprietary “AI World Engine” architecture. The company released AWE 3.0 on March 14, 2026, with three main selling points: training on its own WIYH human-manipulation dataset (which the company says exceeds 1 million hours); generating actions in “latent space” (a compressed latent representation), which it says cuts jitter by more than 45%; and added high-density tactile sensing. At WAIC in July 2026, AWE won a SAIL Award, with the company's booth showing a robot doing wire-harness assembly driven by AWE 3.5. It represents a domestic line of work centered on human data as the primary training source, but there is currently no public paper, and all the figures above come from the company's own announcements.","example":"According to the company's launch demo, a TARS robot running AWE 3.0 completed more than a hundred sub-millimeter-precision wire-harness assemblies in one hour.","related":["TARS Robotics","World In Your Hands (WIYH)","Human Video Data","Latent Action","Embodied Foundation Model","Visuo-Tactile Fusion"]},{"id":"psi-r2","category":"named_model","sec":7,"tier":3,"sources":[{"title":"灵初智能：新一代具身模型发布，全球最大人类手部操作数据集开源","url":"https://www.psibot.ai/%e7%81%b5%e5%88%9d%e6%99%ba%e8%83%bd%e6%96%b0%e4%b8%80%e4%bb%a3%e5%85%b7%e8%ba%ab%e6%a8%a1%e5%9e%8b%e5%8f%91%e5%b8%83%ef%bc%8c%e5%85%a8%e7%90%83%e6%9c%80%e5%a4%a7%e4%ba%ba%e7%b1%bb%e6%89%8b%e9%83%a8/"},{"title":"灵初智能：Psi-R2.5 正式发布","url":"https://www.psibot.ai/%e7%81%b5%e5%88%9d%e6%99%ba%e8%83%bdpsi-r2-5%e6%ad%a3%e5%bc%8f%e5%8f%91%e5%b8%83%ef%bc%81%e6%a8%a1%e5%9e%8b%e8%83%bd%e5%8a%9b%e6%8c%81%e7%bb%ad%e5%bc%ba%e5%8c%96%ef%bc%8c%e7%a0%b4%e8%a7%a3/"}],"as_of":"2026-09","related_ids":["world-action-model","psibot","human-video-data","data-glove","world-model","dreamzero"],"name":"Psi-R2","alt":"灵初 Psi-R2","abbr":"","aliases":["Psi R2","PsiR2","PsiBot Psi-R2"],"one_liner":"A world-action model from the Chinese startup PsiBot, pretrained on roughly 100,000 hours of human manipulation data.","explanation":"Psi-R2 is an embodied model released on April 10, 2026, by the Chinese startup PsiBot (灵初智能). The company describes it as the first 'world-action model' pretrained at the scale of 100,000-plus hours of human manipulation data: a model that takes an image and a language instruction and jointly predicts future video frames and the robot's actions. The data comes from human demonstrations captured with the company's own exoskeleton gloves, covering scenes such as industrial assembly, everyday chores, and object grasping; PsiBot also open-sourced an initial batch of 1,000 hours of human hand-operation data, including vision, language, joint angles, and touch. The company states that fine-tuning on fewer than 100 real-robot trajectories is enough for the model to perform long-horizon, fine-grained tasks such as assembling a phone, industrial packaging, or folding paper boxes. Alongside Psi-R2, PsiBot released an action-conditioned world model called Psi-W0, used to convert human data into robot-usable data and to run reinforcement learning on the policy inside the model. A follow-up, Psi-R2.5, released on September 28, 2026, switched to a two-level structure with a high-level VLM breaking tasks down and a lower-level video model producing the actions.","example":"In an official demo, after fine-tuning on fewer than 100 real-robot trajectories, the robot completes a multi-step, fine-grained task such as assembling a phone.","related":["World Action Model","PsiBot","Human Video Data","Data Glove","World Model","DreamZero"]},{"id":"tencent-hy-embodied","category":"named_model","sec":7,"tier":3,"sources":[{"title":"Tencent-Hunyuan/HY-Embodied (GitHub)","url":"https://github.com/Tencent-Hunyuan/HY-Embodied"},{"title":"HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents (arXiv 2604.07430)","url":"https://arxiv.org/abs/2604.07430"},{"title":"Tencent-Hunyuan/Hy-Embodied-0.5-VLA (GitHub)","url":"https://github.com/Tencent-Hunyuan/Hy-Embodied-0.5-VLA"}],"as_of":"2026-07","related_ids":["embodied-foundation-model","vision-language-model","vision-language-action-model","mixture-of-transformers","universal-manipulation-interface","tencent-tairos-embodied-ai-open-platform"],"name":"Tencent HY-Embodied (Hunyuan Embodied)","alt":"腾讯 HY-Embodied（混元具身）","abbr":"HY-Embodied","aliases":["HY-Embodied-0.5","HY-Embodied-0.5-X","Hy-Embodied-0.5-VLA","HY-VLA-0.5","Hy-Embodied-VLM-1.0"],"one_liner":"An open-source embodied foundation model series from Tencent Robotics X and the Hunyuan team, including both an embodied VLM and a VLA.","explanation":"This is an open-source embodied foundation model family from Tencent's Robotics X Lab and its Hunyuan vision team. HY-Embodied-0.5, from April 2026, is a vision-language model built for robots, strengthening spatial and temporal perception and embodied reasoning, using a Mixture-of-Transformers (MoT) architecture, with a 2B-active (4B total) on-device version and a 32B version — the MoT-2B was open-sourced. The same month's 0.5-X continues post-training on top of it, focused on task planning and risk judgment. June's Hy-Embodied-0.5-VLA adds a flow-matching action expert on this backbone, trained on more than 10,000 hours of bimanual data collected with the company's own fingertip UMI device, with over 2,000 hours of that open-sourced; in July, a MoE-architecture VLM-1.0 followed (about 3B active / 30B total parameters). The project page sits on Tencent's Tairos embodied-AI open platform.","example":"Hy-Embodied-0.5-VLA reports success rates of 90.9% (Clean) and 90.1% (Randomized) on the RoboTwin 2.0 simulation benchmark.","related":["Embodied Foundation Model","Vision-Language Model","Vision-Language-Action Model","Mixture-of-Transformers","Universal Manipulation Interface","Tencent Tairos Embodied AI Open Platform"]},{"id":"qwen-robot-series","category":"named_model","sec":7,"tier":2,"sources":[{"title":"Qwen-RobotManip Technical Report (arXiv:2606.17846)","url":"https://arxiv.org/abs/2606.17846"},{"title":"QwenLM/Qwen-RobotNav GitHub 仓库","url":"https://github.com/QwenLM/Qwen-RobotNav"},{"title":"Qwen-RobotWorld Technical Report (arXiv:2606.17030)","url":"https://arxiv.org/abs/2606.17030"}],"as_of":"2026-09","related_ids":["qwen-vl","vision-language-action-model","world-model","vision-and-language-navigation","flow-matching","cross-embodiment"],"name":"Qwen-Robot Series","alt":"千问 Qwen-Robot 系列","abbr":"","aliases":["Qwen-RobotManip","Qwen-RobotNav","Qwen-RobotWorld"],"one_liner":"Alibaba's Qwen team released this trio of 2026 embodied models, one each for manipulation, navigation, and world modeling.","explanation":"Qwen-Robot is a group of embodied models Alibaba's Qwen team released together in June 2026, all built on the Qwen vision-language model, with each of the three members handling a different job. Qwen-RobotManip is a manipulation VLA: a Qwen-VL backbone followed by a flow-matching DiT action head, trained only on open-source robot data plus robot trajectories synthesized from human-hand videos, totaling about 38,100 hours of pretraining data, with an emphasis on aligning data across different robot embodiments before scaling up. Qwen-RobotNav uses a unified waypoint-prediction interface to handle vision-language navigation, object search, target tracking, and autonomous driving all at once. Qwen-RobotWorld is a language-conditioned video world model used to synthesize training data and evaluate policies. The official repository states there is currently no plan to release weights for Manip or Nav.","example":"Qwen-RobotNav was deployed zero-shot on a Unitree Go2 quadruped, running inference on a Jetson Thor at about 5Hz to navigate unfamiliar environments by language instruction.","related":["Qwen-VL","Vision-Language-Action Model","World Model","Vision-and-Language Navigation","Flow Matching","Cross-Embodiment"]},{"id":"minicpm-robot-series","category":"named_model","sec":7,"tier":3,"sources":[{"title":"OpenBMB/MiniCPM-Robot (GitHub)","url":"https://github.com/OpenBMB/MiniCPM-Robot"}],"as_of":"2026-07","related_ids":["vision-language-action-model","on-device-model","embodied-visual-tracking","visual-token-pruning","unitree-go2","pi0-5"],"name":"MiniCPM-Robot series (ModelBest / OpenBMB)","alt":"面壁 MiniCPM-Robot 系列","abbr":"","aliases":["MiniCPM-RobotManip","MiniCPM-RobotTrack"],"one_liner":"MiniCPM's embodied model family, focused on small, on-device models for manipulation and for following a target on command.","explanation":"This is an embodied model family open-sourced in July 2026 by the OpenBMB open-source community, led by ModelBest (面壁智能), with two models in the first release. MiniCPM-RobotManip is a 1.5-billion-parameter generalist VLA (vision-language-action model), with one set of weights covering multiple tasks; it uses streaming inference to keep past frames in context, retaining up to about 1 minute of visual memory, and reuses MiniCPM-V 4.6's visual-token compression, cutting each frame from 256 tokens down to 64. The company reports results on benchmarks like LIBERO, CALVIN, and RoboTwin 2.0 that come close to or beat the much larger π0.5. MiniCPM-RobotTrack is built on MiniCPM4-0.5B, about 900 million parameters, dedicated to following a target on a language instruction; it runs purely on vision on a Unitree Go2 robot dog's onboard compute, at more than 5 frames per second with about 180 milliseconds of end-to-end latency, which the company says is the best open-source result on EVT-Bench.","example":"Say “follow that person” on a Unitree Go2 EDU, and RobotTrack follows along using only its onboard camera and local compute; official demos include riding an elevator and passing through an underground parking garage.","related":["Vision-Language-Action Model","On-device Model","Embodied Visual Tracking","Visual Token Pruning","Unitree Go2","π0.5"]},{"id":"world-models","category":"named_model","sec":8,"tier":3,"sources":[{"title":"World Models (arXiv:1803.10122)","url":"https://arxiv.org/abs/1803.10122"},{"title":"World Models 交互式论文页","url":"https://worldmodels.github.io/"}],"as_of":"2018-12","related_ids":["world-model","learning-in-imagination","variational-autoencoder","recurrent-neural-network","mixture-density-network","dreamerv3"],"name":"World Models","alt":"World Models 论文（Ha & Schmidhuber）","abbr":"","aliases":["Recurrent World Models Facilitate Policy Evolution","World Models (Ha & Schmidhuber, 2018)"],"one_liner":"A classic 2018 world-model paper that trained an agent's policy entirely inside a 'dream' the model learned on its own.","explanation":"World Models is a paper released in March 2018 by David Ha (Google Brain) and Jürgen Schmidhuber (NNAISENSE), published the same year at NeurIPS under the title 'Recurrent World Models Facilitate Policy Evolution.' The agent has three parts: V, a variational autoencoder that compresses a 64×64 image into a vector of a few dozen dimensions; M, a mixture density network combined with a recurrent neural network that predicts the probability distribution of the next vector; and C, a linear controller with only about a thousand parameters, trained with the evolutionary strategy CMA-ES. The key result is that the controller can be trained entirely inside a 'dream' generated by M, and still work when placed back into the real VizDoom game. This paper made the term 'world model' popular again, and later work such as the Dreamer series continues down this path.","example":"On the CarRacing-v0 driving task, the agent scored 906±21, ahead of prior methods' 591–838; on the VizDoom fireball-dodging task, a controller trained only inside the dream scored 1092±556 once transferred back to the real game.","related":["World Model","Learning in Imagination","Variational Autoencoder","Recurrent Neural Network","Mixture Density Network","DreamerV3"]},{"id":"planet","category":"named_model","sec":8,"tier":3,"sources":[{"title":"arXiv 1811.04551: Learning Latent Dynamics for Planning from Pixels","url":"https://arxiv.org/abs/1811.04551"},{"title":"ICML 2019 论文页（PMLR v97）","url":"https://proceedings.mlr.press/v97/hafner19a.html"},{"title":"Google Research Blog: Introducing PlaNet","url":"https://research.google/blog/introducing-planet-a-deep-planning-network-for-reinforcement-learning/"}],"as_of":"2019-06","related_ids":["recurrent-state-space-model","dreamerv3","world-model","model-based-reinforcement-learning","model-predictive-control","cross-entropy-method"],"name":"PlaNet","alt":"PlaNet","abbr":"","aliases":["Deep Planning Network","PlaNet: Learning Latent Dynamics for Planning from Pixels"],"one_liner":"A model-based reinforcement learning agent that learns a latent-space world model from pixels and plans actions by imagining outcomes.","explanation":"PlaNet was released by Danijar Hafner and colleagues at Google (a Google Brain and DeepMind collaboration) in November 2018 and published at ICML 2019. It works from images alone: it first learns a world model that compresses each frame into a latent state and predicts how that latent state changes under different actions, and how much reward it would earn. The model's core is the Recurrent State-Space Model (RSSM), which combines a deterministic path and a stochastic path so multi-step predictions stay stable; the paper also introduces a multi-step training objective called latent overshooting. At decision time, PlaNet does not train a separate policy network — instead, it imagines many candidate action sequences inside the latent space and picks the one with the highest predicted return, executing only its first step before replanning (online planning). On the DeepMind Control continuous-control benchmark, it needs far fewer interaction episodes than model-free methods; Google's blog reported roughly 50 times better data efficiency on average. The later Dreamer series reuses the RSSM but learns a policy inside imagination instead.","example":"On the 'cheetah run' task, PlaNet works from camera images alone, comparing tens of thousands of imagined action sequences in latent space at every step and executing just the first action of the sequence with the highest predicted reward.","related":["Recurrent State-Space Model","DreamerV3","World Model","Model-Based Reinforcement Learning","Model Predictive Control","Cross-Entropy Method"]},{"id":"muzero","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (arXiv 1911.08265)","url":"https://arxiv.org/abs/1911.08265"},{"title":"MuZero: Mastering Go, chess, shogi and Atari without rules (Google DeepMind blog)","url":"https://deepmind.google/discover/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules/"}],"as_of":"2020-12","related_ids":["model-based-reinforcement-learning","world-model","monte-carlo-tree-search","latent-world-model","value-function","dreamerv3"],"name":"MuZero","alt":"MuZero","abbr":"MuZero","aliases":["Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model"],"one_liner":"A DeepMind algorithm that masters Go and Atari through tree search inside a learned internal model, with no rules given.","explanation":"MuZero was posted as a preprint by Julian Schrittwieser, David Silver, and colleagues at DeepMind in November 2019, and published in Nature in December 2020, the successor to AlphaGo and AlphaZero. AlphaZero needs to know the rules of a game to simulate it out “in its head”; MuZero no longer needs the rules at all: it learns a latent-space model that predicts only the three things most useful for decision-making — reward, policy, and value — without reconstructing the full image, then uses Monte Carlo tree search to plan ahead and choose an action inside that learned model. It reached AlphaZero's level at Go, chess, and shogi, and beat every algorithm at the time across 57 Atari games. It is a representative example of model-based reinforcement learning, embodying the idea that a world model only needs to serve planning, and is often discussed alongside the Dreamer series.","example":"Playing an Atari game, MuZero is never told the rules; seeing only the screen and the score, it simulates the consequences of several different moves inside its own learned internal model and picks whichever has the highest expected return.","related":["Model-Based Reinforcement Learning","World Model","Monte Carlo Tree Search","Latent World Model","Value Function","DreamerV3"]},{"id":"daydreamer","category":"named_model","sec":8,"tier":3,"sources":[{"title":"DayDreamer (arXiv:2206.14176)","url":"https://arxiv.org/abs/2206.14176"},{"title":"DayDreamer 项目主页（CoRL 2022）","url":"https://danijar.com/project/daydreamer/"}],"as_of":"2022-06","related_ids":["world-model","learning-in-imagination","model-based-reinforcement-learning","real-world-reinforcement-learning","recurrent-state-space-model","dreamerv3"],"name":"DayDreamer","alt":"DayDreamer","abbr":"","aliases":["DayDreamer: World Models for Physical Robot Learning"],"one_liner":"Runs the Dreamer world model directly on real robots; a quadruped learns to walk from scratch in an hour.","explanation":"DayDreamer was proposed in June 2022 by Pieter Abbeel and Ken Goldberg's groups at Berkeley, with authors including Danijar Hafner of the Dreamer series, published at CoRL 2022. The hard part of real-robot reinforcement learning is that trial and error is expensive, so most work trains in simulation first and transfers afterward. DayDreamer instead moves the Dreamer algorithm directly onto the real robot: as the robot interacts, it uses the collected data to learn a world model (predicting “what I'll see and what reward I'll get if I take this action”), and the policy learns mostly from imagined rollouts inside that world model, sharply reducing real-world trial and error. The same hyperparameters worked across four robots: an A1 quadruped learned to roll over, stand up, and walk from scratch in about 1 hour, and learned to resist being pushed in about 10 minutes; UR5 and xArm arms learned pick-and-place from camera images; a Sphero wheeled robot learned to navigate to a goal. It demonstrated the sample efficiency of world-model-based reinforcement learning in the real physical world.","example":"With no simulation pretraining and no manual resets, an A1 quadruped robot learned to roll over, stand up, and walk forward from scratch after about 1 hour of learning on real ground.","related":["World Model","Learning in Imagination","Model-Based Reinforcement Learning","Real-World Reinforcement Learning","Recurrent State-Space Model","DreamerV3"]},{"id":"dreamerv3","category":"named_model","sec":8,"tier":2,"sources":[{"title":"DreamerV3 (arXiv 2301.04104)","url":"https://arxiv.org/abs/2301.04104"},{"title":"DreamerV3 项目页（Danijar Hafner）","url":"https://danijar.com/project/dreamerv3/"},{"title":"DreamerV2: Mastering Atari with Discrete World Models (arXiv 2010.02193)","url":"https://arxiv.org/abs/2010.02193"}],"as_of":"2025","related_ids":["world-model","recurrent-state-space-model","learning-in-imagination","model-based-reinforcement-learning","daydreamer","dreamer-4"],"name":"DreamerV3","alt":"DreamerV3","abbr":"","aliases":["Dreamer Series","DreamerV1","DreamerV2"],"one_liner":"A world-model reinforcement-learning algorithm that trains a policy by imagining rollouts, using one fixed set of hyperparameters.","explanation":"The Dreamer series is led by Danijar Hafner. The original 2019 Dreamer proposed learning behavior by “imagining” the future inside a latent space; 2020's DreamerV2 switched to discrete latent variables and was the first world-model-based agent to reach human-level performance on Atari, beating top single-GPU model-free methods like Rainbow and IQN; DreamerV3 was posted publicly in January 2023 and published in Nature in 2025. The approach: first learn a world model from interaction experience that predicts the next latent state and reward, then train the policy on imagined trajectories generated by that model, using the real environment mainly to collect data. Through robustness tricks like normalization, balancing, and transformations, it solves more than 150 different tasks with the same fixed hyperparameters, and it became the first algorithm to mine diamonds in Minecraft from scratch without human data or a curriculum.","example":"In Minecraft, DreamerV3 learned from scratch, without human demonstrations, to chop trees, craft tools, and eventually mine a diamond.","related":["World Model","Recurrent State-Space Model","Learning in Imagination","Model-Based Reinforcement Learning","DayDreamer","Dreamer 4"]},{"id":"dreamer-4","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Training Agents Inside of Scalable World Models (arXiv 2509.24527)","url":"https://arxiv.org/abs/2509.24527"},{"title":"Dreamer 4 项目页（danijar.com）","url":"https://danijar.com/project/dreamer4/"}],"as_of":"2025-09","related_ids":["world-model","learning-in-imagination","dreamerv3","model-based-reinforcement-learning","vpt","offline-reinforcement-learning"],"name":"Dreamer 4","alt":"Dreamer 4","abbr":"","aliases":["Training Agents Inside of Scalable World Models"],"one_liner":"A 2025 Google DeepMind world-model agent that mined diamonds in Minecraft using only offline data.","explanation":"Dreamer 4 is an agent released in September 2025 by Danijar Hafner, Wilson Yan, and Timothy Lillicrap at Google DeepMind, the newest generation in the Dreamer line (following DreamerV3 and earlier versions). The core idea behind the Dreamer approach is “learning in imagination”: first learn a world model, then train the policy with reinforcement learning entirely on experience generated by that model, without having to repeatedly interact with the real environment. Dreamer 4 scales the world model to about 2 billion parameters (a 400-million-parameter video tokenizer plus a 1.6-billion-parameter dynamics model), using a training objective called “shortcut forcing” together with an efficient Transformer to achieve real-time interactive inference on a single GPU. It learns most of its knowledge from video with no action labels, needing only a small amount of action-labeled data to learn action conditioning. It is the first agent to obtain a diamond in Minecraft using purely offline data. Because trial and error is slow and unsafe on real robots, this style of training inside a world model is seen as highly relevant to robotics.","example":"Mining a diamond in Minecraft requires more than 20,000 consecutive mouse and keyboard actions from raw pixels; trained on just 2,541 hours of offline player recordings, Dreamer 4 obtained a diamond in 0.7% of episodes, outperforming OpenAI's offline VPT agent while using about 100 times less data.","related":["World Model","Learning in Imagination","DreamerV3","Model-Based Reinforcement Learning","VPT","Offline Reinforcement Learning"]},{"id":"td-mpc2","category":"named_model","sec":8,"tier":3,"sources":[{"title":"TD-MPC2: Scalable, Robust World Models for Continuous Control (arXiv 2310.16828)","url":"https://arxiv.org/abs/2310.16828"},{"title":"TD-MPC2 project page","url":"https://www.tdmpc2.com/"}],"as_of":"2024-01","related_ids":["model-based-reinforcement-learning","world-model","model-predictive-control","latent-world-model","temporal-difference-learning","dreamerv3"],"name":"TD-MPC2","alt":"TD-MPC2","abbr":"TD-MPC2","aliases":["TD-MPC 2","TD-MPC2: Scalable, Robust World Models for Continuous Control"],"one_liner":"A reinforcement learning algorithm that plans inside a learned latent-space world model, using one set of hyperparameters across hundreds of control tasks.","explanation":"TD-MPC2 was released by Nicklas Hansen, Hao Su, and Xiaolong Wang at UC San Diego in October 2023, an ICLR 2024 Spotlight paper and an improvement on the same team's earlier TD-MPC. It is a model-based reinforcement learning method: it first learns a world model that predicts the next state, reward, and value purely in a latent space, with no image reconstruction (that is, no decoder); at decision time, it performs model predictive control (MPC) within that latent space, relying on short-horizon rollouts from the model while longer-horizon returns are estimated by a value function learned via temporal difference (TD) learning, which is where the name comes from. TD-MPC2 performs reliably with the same set of hyperparameters across 104 continuous-control tasks spanning four domains — DMControl, Meta-World, ManiSkill2, and MyoSuite — and the team also trained a 317-million-parameter multi-task agent that handles 80 tasks, showing that capability keeps growing with model and data scale.","example":"The same TD-MPC2 multi-task model can control the cheetah-run task in DMControl and also open a drawer with a robot arm in Meta-World, without needing separate hyperparameter tuning for each task.","related":["Model-Based Reinforcement Learning","World Model","Model Predictive Control","Latent World Model","Temporal-Difference Learning","DreamerV3"]},{"id":"dino-wm","category":"named_model","sec":8,"tier":3,"sources":[{"title":"DINO-WM (arXiv:2411.04983)","url":"https://arxiv.org/abs/2411.04983"},{"title":"DINO-WM 项目主页","url":"https://dino-wm.github.io/"},{"title":"Gaoyue Zhou 个人主页（标注 ICML 2025）","url":"https://gaoyuezhou.github.io/"}],"as_of":"2025","related_ids":["world-model","latent-world-model","dinov2","model-predictive-control","joint-embedding-predictive-architecture","v-jepa-2"],"name":"DINO-WM","alt":"DINO-WM","abbr":"DINO-WM","aliases":["DINO World Model","DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning"],"one_liner":"A world model that predicts the future in DINOv2's image-feature space, able to plan zero-shot toward new goals after training.","explanation":"DINO-WM was proposed by Gaoyue Zhou and Hengkai Pan at NYU with Yann LeCun and Lerrel Pinto (the former also at Meta FAIR), released in November 2024 and published at ICML 2025. Many world models need to reconstruct pixels, or are tied to a specific task reward. DINO-WM instead uses a pretrained, frozen DINOv2 to extract patch features, and trains only a predictor: given the current features and an action, it predicts the next step's features, with no image reconstruction at all. Training uses only offline collected trajectories, needing no expert demonstrations, reward model, or inverse-dynamics model. At test time, given a goal image, it uses model-predictive control (MPC) to search for a sequence of actions whose predicted future features come as close as possible to the goal's features — zero-shot planning. It represents the approach of building a world model inside a pretrained representation's latent space, in a similar spirit to JEPA and V-JEPA 2.","example":"In the Push-T task, given a goal image of the desired arrangement, DINO-WM rolls out different pushing actions in feature space and picks the sequence that brings the T-shaped block closest to the goal pose; the same method is also used for maze navigation and manipulating rope and granular materials.","related":["World Model","Latent World Model","DINOv2","Model Predictive Control","Joint-Embedding Predictive Architecture","V-JEPA 2"]},{"id":"v-jepa-2","category":"named_model","sec":8,"tier":2,"sources":[{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)","url":"https://arxiv.org/abs/2506.09985"},{"title":"Introducing V-JEPA 2 (Meta AI blog)","url":"https://ai.meta.com/blog/v-jepa-2-world-model-benchmarks/"},{"title":"V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (arXiv 2603.14482)","url":"https://arxiv.org/abs/2603.14482"}],"as_of":"2026-06","related_ids":["joint-embedding-predictive-architecture","world-model","latent-world-model","self-supervised-learning","model-predictive-control","droid"],"name":"V-JEPA 2","alt":"V-JEPA 2","abbr":"V-JEPA 2","aliases":["V-JEPA 2-AC","V-JEPA","V-JEPA 2.1"],"one_liner":"Meta's self-supervised video model that predicts video in feature space, usable as a world model for planning robot actions.","explanation":"V-JEPA 2 is an open-source video model, about 1.2 billion parameters, released by Meta FAIR (Yann LeCun's team) in June 2025. It follows the JEPA (Joint Embedding Predictive Architecture) approach: instead of generating pixels, it masks part of a video and has the model predict the masked content in an abstract feature space, learning motion and physical regularities this way. The first stage does self-supervised pretraining on more than 1 million hours of web video and 1 million images; the second stage freezes the encoder and trains an action-conditioned predictor, called V-JEPA 2-AC, on fewer than 62 hours of unlabeled robot video from the DROID dataset. In use, given a goal image, the model searches in feature space for the action whose predicted outcome comes closest to the goal (model-predictive control). On Franka arms in two different labs, it grasped and placed unseen objects zero-shot with 65–80% success. In 2026 Meta also released V-JEPA 2.1, with improved dense features.","example":"In a new lab where no data was ever collected, given a photo showing an object already placed at its target position, V-JEPA 2-AC imagines the outcome of several candidate actions in feature space at each step, executes whichever one lands closest to the goal image, and moves the object into place step by step.","related":["Joint-Embedding Predictive Architecture","World Model","Latent World Model","Self-Supervised Learning","Model Predictive Control","DROID (Distributed Robot Interaction Dataset)"]},{"id":"stable-video-diffusion","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (arXiv 2311.15127)","url":"https://arxiv.org/abs/2311.15127"},{"title":"Introducing Stable Video Diffusion (Stability AI)","url":"https://stability.ai/news-updates/stable-video-diffusion-open-ai-video-model"},{"title":"Video Prediction Policy (arXiv 2412.14803)","url":"https://arxiv.org/abs/2412.14803"}],"as_of":"2023-11","related_ids":["video-generation-model","latent-diffusion-model","video-prediction-policy","text-to-video-image-to-video","world-model","diffusion-model"],"name":"Stable Video Diffusion","alt":"Stable Video Diffusion","abbr":"SVD","aliases":["SVD-XT","Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets"],"one_liner":"Stability AI's open-source image-to-video diffusion model, commonly used in robotics research as a backbone for video prediction.","explanation":"Stable Video Diffusion is an open-source video generation model released by Stability AI on November 21, 2023, a latent diffusion model (which compresses images into a low-dimensional latent space and denoises step by step within it). It adds temporal layers on top of the Stable Diffusion image model and summarizes training into three stages — text-to-image pretraining, large-scale video pretraining, and high-quality video fine-tuning — emphasizing how much video-data curation matters. The initial release included two image-to-video models generating 14 and 25 frames respectively (the latter called SVD-XT), with a settable frame rate of 3 to 30 fps; it launched as a research preview, not for commercial use. Thanks to its open weights and modest size of about 1.5 billion parameters, SVD became a commonly used video backbone in embodied AI — for example, Video Prediction Policy (VPP) adds language conditioning on top of SVD, fine-tunes it on robot video, and extracts predictive representations of the future from it to output actions.","example":"Video Prediction Policy (VPP) fine-tunes SVD into a manipulation-video prediction model: given the current frame and the instruction 'open the drawer,' it first predicts features of the upcoming frames and then uses them to output the robot arm's actions.","related":["Video Generation Model","Latent Diffusion Model","Video Prediction Policy","Text-to-Video / Image-to-Video","World Model","Diffusion Model"]},{"id":"sora","category":"named_model","sec":8,"tier":2,"sources":[{"title":"Sora (text-to-video model) - Wikipedia","url":"https://en.wikipedia.org/wiki/Sora_(text-to-video_model)"},{"title":"Sora (人工智能模型) - 维基百科","url":"https://zh.wikipedia.org/wiki/Sora_(人工智能模型)"}],"as_of":"2026-09","related_ids":["video-generation-model","world-model","diffusion-transformer","spacetime-patches","world-foundation-model","intuitive-physics"],"name":"Sora","alt":"Sora（视频生成即世界模拟器）","abbr":"","aliases":["Sora 2","Video Generation Models as World Simulators"],"one_liner":"OpenAI's text-to-video model, introduced with a technical report arguing that video generation models can double as world simulators.","explanation":"Sora is OpenAI's video generation model, first shown in February 2024 and opened to paying users that December, alongside a technical report titled “Video generation models as world simulators.” It's a diffusion Transformer: video is compressed into a latent space, cut into spacetime patches used as tokens, and then generated by progressive denoising. The report argues that scaling up video generation models naturally produces abilities like 3D consistency and object permanence, and may be a path toward a simulator of the physical world. This connected the video-generation and world-model lines of research, and also sparked debate over whether it truly understands physics — the report itself acknowledges that interactions like glass shattering aren't simulated accurately. Sora 2, with a companion social app, launched on September 30, 2025; OpenAI announced it was discontinuing Sora in March 2026, the app was taken down on April 26, and the API was shut off on September 24.","example":"In the technical report, Sora generated Minecraft footage while simultaneously controlling the player character and rendering the surrounding world — cited as an example of it “simulating a digital world.”","related":["Video Generation Model","World Model","Diffusion Transformer","Spacetime Patches","World Foundation Model","Intuitive Physics"]},{"id":"genie","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Genie: Generative Interactive Environments (arXiv:2402.15391)","url":"https://arxiv.org/abs/2402.15391"},{"title":"Genie 项目主页","url":"https://sites.google.com/view/genie-2024/home"}],"as_of":"2024-02","related_ids":["genie-2","genie-3","latent-action-model","video-tokenizer","interactive-world-model","lapa"],"name":"Genie (Original)","alt":"Genie（初代）","abbr":"","aliases":["Genie 1","Generative Interactive Environments"],"one_liner":"Google DeepMind's model that learns a “playable world” from unlabeled video, a generative interactive environment.","explanation":"Genie was released by Google DeepMind in February 2024, with 11 billion parameters, and is the first generative interactive environment trained in an unsupervised way from nothing but unlabeled internet video. It has three components: a spatiotemporal video tokenizer that compresses video into discrete tokens; a latent action model that infers, from a pair of consecutive frames, which action occurred, out of a small set of discrete action codes (8 in the paper); and an autoregressive dynamics model that predicts the next frame from the current frame and the action. The key point is that no action labels are needed during training, yet a person can still control the generated world frame by frame. It was trained mainly on 2D platform-jumping game videos, and was also validated on RT-1 robot video; the latent-action idea was later carried over by robot-pretraining work such as LAPA.","example":"Give Genie a hand-drawn sketch, and it will turn it into a 2D platform game you can play frame by frame with your actions.","related":["Genie 2","Genie 3","Latent Action Model","Video Tokenizer","Interactive World Model","LAPA"]},{"id":"genie-2","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Genie 2: A large-scale foundation world model (Google DeepMind Blog)","url":"https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/"}],"as_of":"2024-12","related_ids":["genie","genie-3","world-model","interactive-world-model","latent-diffusion-model","sima-2"],"name":"Genie 2","alt":"Genie 2","abbr":"","aliases":[],"one_liner":"Google DeepMind's large-scale world model that generates a controllable 3D world from a single image.","explanation":"Genie 2 is a foundation world model released by Google DeepMind on December 4, 2024, the successor to the original Genie. Given a single image — either a real photo or output from the Imagen 3 text-to-image model — it generates a 3D environment that can be controlled with a keyboard and mouse, staying visually consistent for up to about a minute, though most examples hold for 10–20 seconds. Its architecture is an autoregressive latent diffusion model trained on large-scale video data: it denoises and generates frame by frame inside a compressed latent space. It can simulate effects like gravity, water, and smoke, along with object interactions, and remembers content that has moved out of view. Its main use is providing diverse training and evaluation environments for agents such as SIMA; its successor is Genie 3.","example":"Given a real-world photo, Genie 2 can turn it into an interactive 3D scene that you can walk through in first person using the keyboard.","related":["Genie (Original)","Genie 3","World Model","Interactive World Model","Latent Diffusion Model","SIMA 2"]},{"id":"genie-3","category":"named_model","sec":8,"tier":2,"sources":[{"title":"Genie 3: A new frontier for world models (Google DeepMind blog)","url":"https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/"},{"title":"Genie (world model) - Wikipedia","url":"https://en.wikipedia.org/wiki/Genie_(world_model)"}],"as_of":"2026-01","related_ids":["world-model","interactive-world-model","genie","genie-2","waymo-world-model","autoregressive-video-generation"],"name":"Genie 3","alt":"Genie 3","abbr":"","aliases":["Genie3"],"one_liner":"Google DeepMind's 2025 real-time interactive world model: one text prompt generates a virtual world you can walk through.","explanation":"Genie 3 is the general-purpose world model Google DeepMind released on August 5, 2025. Given a text description, it generates a dynamic world that can be explored and controlled in real time: 720p at 24 frames per second, holding consistency for several minutes, with visual memory reaching back about 1 minute. Frames are generated autoregressively one at a time, each conditioned on a growing history of past frames and the user's actions; it also supports “promptable world events,” letting text change the weather or add new objects or characters. This lets it serve as an environment for training and evaluating agents without hand-building a simulator for every task. Its limits are a restricted set of actions an agent can take and sessions of only a few minutes. At launch it was available only to a small number of researchers and creators; on January 29, 2026, Google opened it to AI Ultra subscribers as Project Genie (capped at 60 seconds per session), and Waymo has also built a self-driving world model on top of it.","example":"DeepMind placed its SIMA agent inside a world generated by Genie 3 and gave it a goal. The agent only sends navigation actions like move forward or turn; Genie 3 has no idea what the goal is and simply simulates the next frames from those actions, and as long as consistency holds long enough, the agent can carry out longer action sequences.","related":["World Model","Interactive World Model","Genie (Original)","Genie 2","Waymo World Model","Autoregressive Video Generation"]},{"id":"sima-2","category":"named_model","sec":8,"tier":3,"sources":[{"title":"SIMA 2: An agent that plays, reasons, and learns with you in virtual 3D worlds (Google DeepMind blog)","url":"https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/"},{"title":"SIMA 2: A Generalist Embodied Agent for Virtual Worlds (arXiv 2512.04797)","url":"https://arxiv.org/abs/2512.04797"}],"as_of":"2025-12","related_ids":["genie-3","embodied-agent","self-improvement","google-gemini","google-deepmind","instruction-following"],"name":"SIMA 2","alt":"SIMA 2","abbr":"SIMA","aliases":["Scalable Instructable Multiworld Agent 2"],"one_liner":"A Gemini-based Google DeepMind agent that follows instructions, reasons, and self-improves across many different 3D game worlds.","explanation":"SIMA 2 was released by Google DeepMind in November 2025 as a limited research preview, with a technical report published in December. Its predecessor, SIMA (Scalable Instructable Multiworld Agent, released 2024), could carry out more than 600 simple skill instructions, like 'turn left' or 'climb the ladder,' across a range of commercial 3D games. SIMA 2 is built around a Gemini model and goes beyond following instructions: it can understand a high-level goal, hold a conversation with the user, explain what it plans to do, and accept instructions in the form of images, sketches, or even emoji. It can also self-improve: Gemini sets it tasks and scores its performance, and the agent trains its next version on the experience it generates by playing, no longer relying on human demonstrations. It can also orient itself and act on instructions in entirely new worlds generated in real time by Genie 3. DeepMind views SIMA 2 as a step toward a general embodied agent that will eventually be used on real robots.","example":"Told to 'go to the house that's the color of a ripe tomato,' SIMA 2 first reasons that a ripe tomato is red, then goes looking for the red house.","related":["Genie 3","Embodied Agent","Self-improvement","Google Gemini","Google DeepMind","Instruction Following"]},{"id":"diamond","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Diffusion for World Modeling: Visual Details Matter in Atari (arXiv:2405.12399)","url":"https://arxiv.org/abs/2405.12399"},{"title":"DIAMOND 项目主页","url":"https://diamond-wm.github.io/"}],"as_of":"2024-10","related_ids":["world-model","diffusion-model","learning-in-imagination","gamengen","dreamerv3","arcade-learning-environment-atari-100k"],"name":"DIAMOND","alt":"DIAMOND（扩散世界模型）","abbr":"DIAMOND","aliases":["Diffusion for World Modeling","DIAMOND: Visual Details Matter in Atari"],"one_liner":"Generates frames one at a time with a diffusion model as a world model, letting an RL agent train inside a generated game.","explanation":"DIAMOND was released in May 2024 by researchers at the University of Geneva, the University of Edinburgh, and Microsoft Research, a NeurIPS 2024 Spotlight paper. Many earlier world models (such as IRIS and DreamerV3) compress frames into discrete tokens or a latent variable before predicting, which easily loses small but important visual details. DIAMOND instead uses a diffusion model directly to generate the next frame from recent frames and the action, and works out the key design choices needed to keep this stable over long rollouts. An agent is trained with reinforcement learning entirely inside this generated environment, then tested in the real game: on the Atari 100k benchmark it reaches an average human-normalized score of 1.46, the best result at the time for an agent trained purely with a world model. It also shows that a diffusion world model can serve as a playable, interactive game engine, in a similar spirit to GameNGen and Genie.","example":"The authors trained a 381-million-parameter diffusion world model on 87 hours of human Counter-Strike: Global Offensive (CS:GO) match data; it runs at about 10 frames per second on an RTX 3090, and a person can “play” inside it in real time with a keyboard and mouse.","related":["World Model","Diffusion Model","Learning in Imagination","GameNGen","DreamerV3","Arcade Learning Environment (ALE) / Atari 100k"]},{"id":"gamengen","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Diffusion Models Are Real-Time Game Engines (arXiv 2408.14837)","url":"https://arxiv.org/abs/2408.14837"},{"title":"GameNGen 项目主页","url":"https://gamengen.github.io/"}],"as_of":"2024-08","related_ids":["interactive-world-model","diffusion-model","autoregressive-video-generation","diamond","genie-2","world-model"],"name":"GameNGen","alt":"GameNGen（神经游戏引擎）","abbr":"GameNGen","aliases":["Neural Game Engine","Diffusion Models Are Real-Time Game Engines"],"one_liner":"A 2024 Google project that uses a diffusion model to generate playable DOOM footage in real time, replacing the game engine.","explanation":"GameNGen is a paper released in August 2024 by researchers at Google Research and Google DeepMind, demonstrating that a neural network can simulate a complex game and be played in real time. The game chosen was the classic first-person shooter DOOM. The method has two steps: first, train a reinforcement-learning agent to play the game, recording footage and key presses; then adapt Stable Diffusion 1.4 into a diffusion model that predicts the next frame conditioned on a number of past frames and the player's actions. Autoregressive generation — using self-generated frames to generate the next ones — tends to get blurrier and blurrier over time, so the authors added Gaussian noise to history frames during training, teaching the model to correct itself, which lets it run stably for several minutes; they also separately fine-tuned the image decoder to reduce compression artifacts. It reaches about 20 frames per second on a single TPU. This work drew wide attention to interactive world models, and is often discussed alongside DIAMOND and Genie 2.","example":"The player presses keys to shoot, open doors, and pick up health packs, and the diffusion model generates the footage frame by frame, including the on-screen health and ammo counts; the generated frames reach a PSNR of 29.4, comparable to lossy JPEG compression, and human raters could barely do better than chance at telling short real and generated clips apart.","related":["Interactive World Model","Diffusion Model","Autoregressive Video Generation","DIAMOND","Genie 2","World Model"]},{"id":"matrix-game","category":"named_model","sec":8,"tier":3,"sources":[{"title":"arXiv 2508.13009: Matrix-Game 2.0","url":"https://arxiv.org/abs/2508.13009"},{"title":"GitHub: SkyworkAI/Matrix-Game","url":"https://github.com/SkyworkAI/Matrix-Game"},{"title":"arXiv 2604.08995: Matrix-Game 3.0","url":"https://arxiv.org/abs/2604.08995"}],"as_of":"2026-04","related_ids":["interactive-world-model","world-model","genie-3","autoregressive-video-generation","self-forcing","diffusion-step-distillation"],"name":"Matrix-Game (Skywork)","alt":"昆仑万维 Matrix-Game","abbr":"","aliases":["Matrix-Game 2.0","Matrix-Game 3.0"],"one_liner":"Skywork's open-source real-time interactive world model that generates game footage frame by frame from keyboard and mouse input.","explanation":"Matrix-Game is a series of interactive world models open-sourced by Skywork AI, under Kunlun Wanwei; version 1.0 was released in May 2025, and 2.0 in August 2025. The model reads the current frame plus the user's keyboard and mouse input at each frame, then generates the next segment of video, producing a game world generated as you play. 2.0's main ideas are three: automatically collecting about 1,200 hours of interaction-annotated video using Unreal Engine and the GTA5 environment; injecting keyboard-and-mouse actions into the model frame by frame; and distilling the diffusion model into a few-step causal autoregressive generator that produces minutes-long video in real time at 25 FPS, with weights and code both open-sourced. Version 3.0, released in late March 2026, added camera-position-based long-term memory, reaching 720p at 40 FPS with a 5B model. This family of models follows the same line as Genie 3, and is seen as a potential tool for robot simulation and data generation.","example":"In a GTA-style driving scene, holding W and moving the mouse right makes the model generate, in real time, footage of the car driving forward and turning right.","related":["Interactive World Model","World Model","Genie 3","Autoregressive Video Generation","Self Forcing","Diffusion Step Distillation"]},{"id":"hunyuanworld","category":"named_model","sec":8,"tier":3,"sources":[{"title":"HunyuanWorld 1.0 technical report (arXiv 2507.21809)","url":"https://arxiv.org/abs/2507.21809"},{"title":"Tencent-Hunyuan/HunyuanWorld-1.0 (GitHub, news timeline)","url":"https://github.com/Tencent-Hunyuan/HunyuanWorld-1.0"},{"title":"Voyager: Long-Range and World-Consistent Video Diffusion (arXiv 2506.04225)","url":"https://arxiv.org/abs/2506.04225"}],"as_of":"2026-05","related_ids":["world-model","interactive-world-model","generative-simulation","marble","genie-3","spatial-intelligence"],"name":"HunyuanWorld (Tencent)","alt":"腾讯混元世界模型","abbr":"HunyuanWorld","aliases":["HunyuanWorld 1.0","HunyuanWorld-Voyager","HY-World","WorldPlay"],"one_liner":"Tencent Hunyuan's open-source 3D world-generation series that turns text or an image into an explorable 3D scene.","explanation":"HunyuanWorld is an open-source world-generation model series from Tencent's Hunyuan team. HunyuanWorld 1.0, from July 2025, uses a 360° panorama as an intermediate representation, layering the scene semantically and reconstructing it into an exportable 3D mesh, generating an explorable 3D world from a single sentence or image; Voyager, from September 2025, is an RGB-D video diffusion model that generates geometrically consistent color and depth video along a camera path the user provides, yielding a 3D point cloud directly. Later releases followed: 1.1 (WorldMirror, reconstruction from video or multi-view images), 1.5 (WorldPlay, December 2025, a real-time interactive world model), and HY-World 2.0 (April 2026). For embodied AI, it's used mainly to mass-produce simulation scenes and assets — it leans toward “generating a 3D environment,” a different focus from a world model that predicts the consequences of a robot's actions.","example":"Given the prompt “a seaside town of wooden cabins,” HunyuanWorld 1.0 generates an explorable 360° 3D scene with an exportable mesh, ready to import into a game engine or simulator.","related":["World Model","Interactive World Model","Generative Simulation","Marble (World Labs)","Genie 3","Spatial Intelligence"]},{"id":"marble","category":"named_model","sec":8,"tier":3,"sources":[{"title":"World Labs: Marble: A Multimodal World Model","url":"https://www.worldlabs.ai/blog/marble-world-model"},{"title":"World Labs: Atlas: A World Model for Spatial Intelligence","url":"https://www.worldlabs.ai/blog/atlas"},{"title":"World Labs is Joining AMD","url":"https://www.worldlabs.ai/blog/amd-announcement"}],"as_of":"2026-09","related_ids":["spatial-intelligence","world-model","3d-gaussian-splatting","generative-simulation","world-labs","real-to-sim-to-real"],"name":"Marble (World Labs)","alt":"Marble（World Labs）","abbr":"","aliases":[],"one_liner":"World Labs' world-generation product that turns text, images, or video into a persistent, exportable 3D scene.","explanation":"Marble is the generative world model product that World Labs, co-founded by Fei-Fei Li, formally launched on November 12, 2025. A user provides text, one or more images, or video, or first blocks out a rough 3D layout with the companion tool Chisel, and Marble generates a persistent 3D scene that can be freely walked through and viewed, exportable as a Gaussian splat (representing the scene with a large number of semi-transparent ellipsoids), a triangle mesh (including a simplified mesh for collision), or a video rendered along a specified camera path. Unlike a video world model that generates frames one at a time, it outputs explicit 3D assets that can be dropped into a game engine or simulator. In July 2026, World Labs demonstrated using generated scenes for real-to-sim-to-real robot training; in September, its underlying model, Atlas, was released and will power later versions of Marble; on September 28, AMD announced it would acquire World Labs, with the deal expected to close before the end of 2026.","example":"Upload a photo of a living room, and Marble generates a 3D living room you can look around in; export the collision mesh and drop it into a simulator to let a robot practice navigating inside it.","related":["Spatial Intelligence","World Model","3D Gaussian Splatting","Generative Simulation","World Labs","Real-to-Sim-to-Real"]},{"id":"mano","category":"named_model","sec":8,"tier":3,"sources":[{"title":"MANO 官网（MPI-IS）","url":"https://mano.is.tue.mpg.de/"},{"title":"arXiv 2201.02610: Embodied Hands: Modeling and Capturing Hands and Bodies Together","url":"https://arxiv.org/abs/2201.02610"},{"title":"GitHub: hassony2/manopth（MANO 的 PyTorch 实现）","url":"https://github.com/hassony2/manopth"}],"as_of":"2017-11","related_ids":["smpl","hamer","hand-pose-estimation","motion-retargeting","dexterous-hand","human-video-data"],"name":"MANO","alt":"MANO 手部模型","abbr":"MANO","aliases":["hand Model with Articulated and Non-rigid defOrmations"],"one_liner":"The most widely used parametric 3D hand model, describing a hand with a small number of shape and pose parameters.","explanation":"MANO, short for hand Model with Articulated and Non-rigid defOrmations, was proposed by Javier Romero and Dimitrios Tzionas in Michael Black's group at the Max Planck Institute for Intelligent Systems in Germany, published at SIGGRAPH Asia 2017. It was learned from about 1,000 high-precision 3D hand scans of 31 people: given shape parameters (commonly 10-dimensional, describing the hand's size and thickness) and pose parameters (the rotation of each finger joint, which can be compressed to fewer dimensions with principal component analysis), it outputs a 3D hand mesh complete with finger bending and skin deformation. It can also attach to the SMPL body model to form SMPL+H. In embodied AI, MANO is a common representation for learning dexterous manipulation from human-hand video: hand-reconstruction methods like HaMeR output MANO parameters, which are then mapped, through motion retargeting, onto a robot dexterous hand's joints.","example":"MANO parameters are estimated frame by frame from a first-person video of a person picking up a cup, yielding a trajectory of finger joint angles, which is then retargeted into a demonstration for a robot dexterous hand.","related":["SMPL","HaMeR","Hand Pose Estimation","Motion Retargeting","Dexterous Hand","Human Video Data"]},{"id":"vpt","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Video PreTraining (VPT) (arXiv 2206.11795)","url":"https://arxiv.org/abs/2206.11795"},{"title":"VPT 论文 HTML 版（数据规模细节）","url":"https://arxiv.org/html/2206.11795"}],"as_of":"2022-06","related_ids":["inverse-dynamics-model","pseudo-action-labels","action-free-video","behavior-cloning","minecraft-environments","latent-action-pretraining"],"name":"VPT","alt":"VPT（视频预训练）","abbr":"VPT","aliases":["Video PreTraining","Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos"],"one_liner":"OpenAI's approach of first training an inverse dynamics model to label online video with actions, then learning to play Minecraft by imitation.","explanation":"VPT was released by OpenAI in June 2022, published at NeurIPS 2022. The difficulty with decision-making tasks is that while there's plenty of video online, none of it comes labeled with 'which key was pressed at each moment.' VPT first pays human contractors to play Minecraft while recording their keyboard and mouse input (about 1,962 hours), and uses this to train an inverse dynamics model (IDM), which infers the action taken between two frames; the IDM is then used to automatically pseudo-label about 70,000 hours of online gameplay video with actions, and behavior cloning trains a base policy on this data, which operates the game directly at 20 Hz through the same keyboard-and-mouse interface a human would use. This base policy already shows some zero-shot ability, and after further imitation learning and reinforcement-learning fine-tuning, it became the first system to synthesize a diamond tool in Minecraft from scratch (a task that takes a skilled human player about 20 minutes and 24,000 actions) — something reinforcement learning starting from nothing could essentially never achieve. The 'train an IDM on a little labeled data, then pseudo-label a lot of video' idea has since been widely borrowed in robotics.","example":"After reinforcement-learning fine-tuning, the VPT base model can start from an empty inventory in Minecraft and work all the way through gathering materials and crafting tools to finally make a diamond tool, something essentially impossible for reinforcement learning trained from scratch.","related":["Inverse Dynamics Model","Pseudo Action Labels","Action-free Video","Behavior Cloning","Minecraft Environments","Latent Action Pretraining"]},{"id":"mimicplay","category":"named_model","sec":8,"tier":3,"sources":[{"title":"MimicPlay: Long-Horizon Imitation Learning by Watching Human Play (arXiv 2302.12422)","url":"https://arxiv.org/abs/2302.12422"},{"title":"MimicPlay project page","url":"https://mimic-play.github.io/"}],"as_of":"2023-10","related_ids":["imitation-learning","long-horizon-task","play-data","human-video-data","hierarchical-architecture","visuomotor-policy"],"name":"MimicPlay","alt":"MimicPlay","abbr":"","aliases":["Long-Horizon Imitation Learning by Watching Human Play"],"one_liner":"An imitation-learning method that learns high-level plans from video of human hands playing freely, then low-level actions from a little teleoperation data.","explanation":"MimicPlay was released in February 2023 by Chen Wang, Linxi Fan, Fei-Fei Li, Yuke Zhu, and colleagues at Stanford, NVIDIA, and other institutions, an oral presentation at CoRL 2023. Collecting data purely through teleoperation is expensive for long-horizon tasks (tasks requiring many steps done in sequence). MimicPlay uses a hierarchical structure: the high level learns a latent plan from “human play data” — video of a person freely moving objects around a scene by hand — representing what the 3D trajectory of a human hand should look like given a goal image; the low level trains a visuomotor policy on only a small number of teleoperation demonstrations, outputting robot actions that follow this plan. Human-hand video is cheap and works across embodiments, filling exactly the gap left by scarce robot data. Across 14 real long-horizon manipulation tasks, it beat the contemporary baselines on success rate, generalization, and robustness to disturbance.","example":"For a multi-step task like “open the microwave, put the bowl inside, then close the door,” the high level predicts the 3D trajectory a human hand should follow given the goal image, and the low-level policy drives the robot arm step by step along that trajectory to complete the task.","related":["Imitation Learning","Long-horizon Task","Play Data","Human Video Data","Hierarchical Architecture","Visuomotor Policy"]},{"id":"vrb","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Affordances from Human Videos as a Versatile Representation for Robotics (arXiv 2304.08488)","url":"https://arxiv.org/abs/2304.08488"},{"title":"VRB 项目主页","url":"https://robo-affordances.github.io/"}],"as_of":"2023-06","related_ids":["affordance","human-video-data","egocentric-video","epic-kitchens","imitation-from-observation","affordance-detection"],"name":"VRB","alt":"VRB（从人类视频学可供性）","abbr":"VRB","aliases":["Vision-Robotics Bridge","VRB: Affordances from Human Videos as a Versatile Representation for Robotics"],"one_liner":"A method that learns 'where to grasp and which way to move afterward' from human video and hands that affordance directly to a robot.","explanation":"VRB (Vision-Robotics Bridge) is work by Shikhar Bahl, Russell Mendonca, and colleagues in Deepak Pathak's group at Carnegie Mellon University, together with Meta AI, posted to arXiv in April 2023 and published at CVPR 2023. Affordance refers to what an object 'can be used for' — a drawer handle, for instance, affords pulling. VRB trains a visual affordance model on large amounts of first-person human video from datasets like EPIC-KITCHENS and Ego4D: given a scene image, it outputs two things — a contact heatmap (where a person is most likely to reach) and the wrist's post-contact motion trajectory (which direction the hand moves after making contact). Because this representation does not depend on any particular robot's body, it can plug into many different robot-learning setups: offline imitation learning, exploration, goal-conditioned learning, and as an action parameterization for reinforcement learning. The paper validated it across 4 real-world environments, more than 10 tasks, and 2 robot platforms, an early example of using human video to make up for scarce robot data.","example":"Given a kitchen photo, the model marks a high contact probability at the cabinet handle and predicts a trajectory of 'grasp, then pull outward'; the robot reaches for that spot and follows the trajectory to attempt opening the cabinet.","related":["Affordance","Human Video Data","Egocentric Video","EPIC-KITCHENS","Imitation from Observation","Affordance Detection"]},{"id":"atm","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Any-point Trajectory Modeling for Policy Learning (arXiv 2401.00025)","url":"https://arxiv.org/abs/2401.00025"},{"title":"ATM 项目页","url":"https://xingyu-lin.github.io/atm/"},{"title":"Robotics: Science and Systems XX (RSS 2024) 论文集","url":"https://www.roboticsproceedings.org/rss20/index.html"}],"as_of":"2024-07","related_ids":["tracking-any-point","cotracker","action-free-video","pretraining-on-human-videos","libero-benchmark","intermediate-representation"],"name":"ATM","alt":"ATM（任意点轨迹建模）","abbr":"ATM","aliases":["Any-point Trajectory Modeling","Any-point Trajectory Modeling for Policy Learning"],"one_liner":"Learns how any point in a scene will move from video, then uses that predicted motion to guide a robot policy.","explanation":"ATM was proposed by researchers at UC Berkeley, Tsinghua's Institute for Interdisciplinary Information Sciences, and collaborators (including Yang Gao and Pieter Abbeel), published at RSS 2024. Action-labeled robot data is expensive; action-free video is abundant. Earlier video pretraining mostly predicted future frames pixel by pixel, which is computationally heavy and full of irrelevant detail. ATM instead predicts point motion: it first uses the point tracker CoTracker to label 2D trajectories for points in a video, then trains a Transformer to predict, from the current image, a language instruction, and a set of point positions, where those points will go in the future; a policy is then trained to output actions using the predicted trajectories as a subgoal, needing only a small number of action-labeled demonstrations. Across more than 130 tasks including LIBERO, it beat video-pretraining baselines by about 80% on average, and can also transfer skills from human videos.","example":"Given the instruction “open the middle drawer of the cabinet,” the model first draws, on the current image, how points such as the drawer handle will move as it's pulled open, and the policy then outputs arm actions guided by these predicted trajectories.","related":["Tracking Any Point","CoTracker","Action-free Video","Pretraining on Human Videos","LIBERO Benchmark","Intermediate Representation"]},{"id":"lapa","category":"named_model","sec":8,"tier":2,"sources":[{"title":"Latent Action Pretraining from Videos (arXiv 2410.11758)","url":"https://arxiv.org/abs/2410.11758"},{"title":"LAPA 项目主页","url":"https://latentactionpretraining.github.io/"}],"as_of":"2025-04","related_ids":["latent-action","latent-action-pretraining","latent-action-model","vector-quantized-variational-autoencoder","action-free-video","openvla"],"name":"LAPA","alt":"LAPA","abbr":"LAPA","aliases":["LAPA: Latent Action Pretraining from Videos","Latent Action Pretraining for General Action Models"],"one_liner":"Pretrains a VLA on action-label-free video by first learning latent actions, then mapping them to real actions with little robot data.","explanation":"LAPA was proposed in October 2024 by researchers from KAIST, the University of Washington, Microsoft Research, NVIDIA, and the Allen Institute for AI, published at ICLR 2025. Pretraining a VLA normally requires action labels collected through teleoperation, which limits both data sources and scale; LAPA instead aims to use the huge amount of action-label-free video already on the web. It works in three steps. First, a VQ-VAE-style objective trains a quantization model that encodes the change between two consecutive frames into a discrete latent action. Second, a vision-language model (a 7B LWM) is pretrained to predict these latent actions from images and a task description. Finally, it's fine-tuned on a small amount of real-robot data that maps the latent actions to real robot actions. The paper reports that on real-robot tasks requiring language understanding and generalization, it beats OpenVLA — which was trained with real action labels — by 6.22%, while being more than 30x more pretraining-efficient.","example":"Pretraining latent actions on only about 220,000 everyday human manipulation videos from Something-Something V2, then fine-tuning on a small number of robot trajectories, already outperforms OpenVLA pretrained on the Bridge robot dataset.","related":["Latent Action","Latent Action Pretraining","Latent Action Model","Vector-Quantized Variational Autoencoder","Action-free Video","OpenVLA"]},{"id":"phantom","category":"named_model","sec":8,"tier":3,"sources":[{"title":"arXiv 2503.00779: Phantom","url":"https://arxiv.org/abs/2503.00779"},{"title":"Phantom 项目主页","url":"https://phantom-human-videos.github.io/"}],"as_of":"2025-09","related_ids":["robotizing-human-videos-human-to-robot-video-translation","human-video-data","robot-free-data-collection","hand-pose-estimation","egomimic","embodiment-gap"],"name":"Phantom","alt":"Phantom（无机器人训练）","abbr":"","aliases":["Phantom: Training Robots Without Robots Using Only Human Videos"],"one_liner":"A method that trains robot policies purely from human demonstration videos by digitally replacing the human hand with a rendered robot arm.","explanation":"Phantom comes from Jeannette Bohg's lab at Stanford University (Marion Lepert and colleagues), released in March 2025 and presented at CoRL 2025. Teleoperated data collection is expensive, while human videos are cheap — but a human hand looks nothing like a robot gripper, and the videos carry no robot action labels. Phantom's solution: estimate hand pose in every frame and convert it into robot end-effector actions, then use image inpainting to erase the human hand from each frame and render a virtual robotic arm in its place, so the training images look like a robot performing the task. This lets a policy be trained with zero robot data and deployed zero-shot on a Franka or a Kinova Gen3 arm, completing tasks such as pick-and-place, cup stacking, tying rope, sweeping, and insertion, with a reported success rate as high as 92%. It belongs to the 'human-video-to-robot' family of methods.","example":"A researcher demonstrates 'stack the cups' with their own hand on a tabletop; Phantom automatically replaces the hand in the video with a rendered gripper and generates action labels, and the resulting policy runs directly on a Franka arm.","related":["Robotizing Human Videos / Human-to-Robot Video Translation","Human Video Data","Robot-free (Embodiment-free) Data Collection","Hand Pose Estimation","EgoMimic","Embodiment Gap"]},{"id":"univla","category":"named_model","sec":8,"tier":3,"sources":[{"title":"UniVLA (arXiv 2505.06111)","url":"https://arxiv.org/abs/2505.06111"},{"title":"OpenDriveLab/UniVLA GitHub","url":"https://github.com/OpenDriveLab/UniVLA"}],"as_of":"2025-05","related_ids":["latent-action-model","latent-action","vision-language-action-model","openvla","lapa","cross-embodiment"],"name":"UniVLA","alt":"UniVLA","abbr":"","aliases":["UniVLA: Learning to Act Anywhere with Task-centric Latent Actions"],"one_liner":"A framework that trains a cross-embodiment VLA by learning 'task-relevant latent actions' from video.","explanation":"UniVLA was released by the University of Hong Kong's OpenDriveLab and AgiBot in May 2025, published at RSS 2025. Most VLAs (vision-language-action models) rely on large amounts of action-labeled robot data and are tied to a single robot. UniVLA first trains a latent action model: looking at two consecutive frames in DINOv2 feature space, together with the language instruction, it separates task-relevant changes from irrelevant ones like camera shake, and quantizes the relevant changes into discrete latent action tokens, which lets videos without action labels — including human videos — be used for pretraining too. It then uses Prismatic-7B as the backbone to predict these latent actions, and when deploying to a specific robot, only adds a small decoder head of about 12 million parameters to translate them into real actions. The paper reports that with under 1/20th of OpenVLA's pretraining compute and only 1/10th of its downstream data, it outperforms OpenVLA on benchmarks including LIBERO, CALVIN, and R2R.","example":"Robot arm data, navigation data, and human manipulation videos are all fed together into UniVLA's latent action model; the same learned latent actions, paired with different small decoder heads, can then drive a robot arm in LIBERO simulation and a navigation agent in R2R.","related":["Latent Action Model","Latent Action","Vision-Language-Action Model","OpenVLA","LAPA","Cross-Embodiment"]},{"id":"egovla","category":"named_model","sec":8,"tier":3,"sources":[{"title":"EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos (arXiv 2507.12440)","url":"https://arxiv.org/abs/2507.12440"},{"title":"EgoVLA 项目页","url":"https://rchalyang.github.io/EgoVLA/"}],"as_of":"2025-07","related_ids":["egocentric-video","pretraining-on-human-videos","mano","motion-retargeting","vision-language-action-model","egoscale"],"name":"EgoVLA","alt":"EgoVLA","abbr":"EgoVLA","aliases":["Learning Vision-Language-Action Models from Egocentric Human Videos"],"one_liner":"A VLA pretrained on first-person human video that converts human hand motion into humanoid robot actions.","explanation":"EgoVLA is a paper released in July 2025 by UC San Diego together with UIUC, MIT, NVIDIA, and others. The motivation is that real-robot data is tied to specific robot hardware and limited in scale, while first-person human video is abundant and covers a much wider range of scenes. It uses the NVILA-2B vision-language model as its backbone, first learning to predict human wrist pose and MANO hand parameters (a parametric model of the human hand) from about 500,000 human-video image-action pairs; the robot hand is then converted into that same MANO action space, and the model is fine-tuned on a small number of robot demonstrations. At deployment, wrist pose is converted into arm joint angles through inverse kinematics, and finger motion is mapped to dexterous-hand joints by a small MLP.","example":"On the authors' Ego Humanoid Manipulation Benchmark, built on Isaac Lab (a simulated Unitree H1 fitted with an Inspire dexterous hand, across 12 tasks), EgoVLA pretrained on human video clearly outperforms baselines trained only on robot data, with especially large gains on long-horizon and fine-grained manipulation tasks.","related":["Egocentric Video","Pretraining on Human Videos","MANO","Motion Retargeting","Vision-Language-Action Model","EgoScale"]},{"id":"unipi","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Learning Universal Policies via Text-Guided Video Generation (arXiv 2302.00111)","url":"https://arxiv.org/abs/2302.00111"},{"title":"UniPi project page","url":"https://universal-policy.github.io/unipi/"}],"as_of":"2023-11","related_ids":["unisim","video-generation-model","inverse-dynamics-model","video-prediction-policy","susie","world-model"],"name":"UniPi","alt":"UniPi","abbr":"","aliases":["Universal Policies via Text-Guided Video Generation","UniPi: Learning Universal Policies via Text-Guided Video Generation"],"one_liner":"A method that first generates a video of a task being completed from text, then infers the robot's actions from that video.","explanation":"UniPi was released by MIT, Google Brain, UC Berkeley, and the University of Alberta (Yilun Du, Pieter Abbeel, and others) in January 2023, published at NeurIPS 2023. It reframes decision-making as 'text-conditioned video generation': given a text goal and the current image, a video diffusion model generates a future video of the task being completed, which serves as the plan, and an inverse dynamics model (which computes the action between two adjacent frames) then derives the robot action the robot should take from each pair of consecutive frames. Because different robots and environments are all unified into images, the model can share knowledge across tasks; text goals can be freely combined for compositional generalization; and pretraining on internet image-text-video data also improves generalization to new instructions. UniPi is an early representative of the 'video generation model as policy' approach, and later work such as UniSim and Video Prediction Policy follows a similar idea.","example":"Given 'put the red block in the blue bowl,' UniPi first generates a short video of a robot completing this task, then an inverse dynamics model derives the actions frame by frame from that video for execution.","related":["UniSim","Video Generation Model","Inverse Dynamics Model","Video Prediction Policy","SuSIE","World Model"]},{"id":"unisim","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Learning Interactive Real-World Simulators (arXiv 2310.06114)","url":"https://arxiv.org/abs/2310.06114"},{"title":"UniSim project page","url":"https://universal-simulator.github.io/unisim/"}],"as_of":"2024-05","related_ids":["unipi","interactive-world-model","neural-simulator","world-model","genie","video-generation-model"],"name":"UniSim","alt":"UniSim","abbr":"","aliases":["Universal Simulator","Learning Interactive Real-World Simulators"],"one_liner":"A learned 'real-world simulator,' built from a video generation model, that responds to actions.","explanation":"UniSim was released by UC Berkeley, Google DeepMind, and MIT (Sherry Yang, Yilun Du, Pieter Abbeel, and others) in October 2023, winning the ICLR 2024 Outstanding Paper Award. It uses a video generation model to learn a real-world simulator: given the current image and an action, it generates the resulting image. The action can either be a high-level language instruction like 'open the drawer' or a low-level control command like 'move to a given coordinate.' No single dataset covers all of this information, so the authors combined image data (rich in object variety), robot data (rich in actions), and navigation data (rich in movement), among others, for training. Once trained, vision-language policies and reinforcement learning policies can be trained inside UniSim and then transferred zero-shot to real robots; the data it generates can also help train other models, such as video captioning. UniSim is an early landmark for interactive world models and neural simulators.","example":"Given a kitchen image and the instruction 'open the drawer,' UniSim generates the subsequent video of the drawer being pulled open; a reinforcement learning policy can then practice through trial and error inside these generated images.","related":["UniPi","Interactive World Model","Neural Simulator","World Model","Genie (Original)","Video Generation Model"]},{"id":"avdc","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Learning to Act from Actionless Videos through Dense Correspondences (arXiv 2310.08576)","url":"https://arxiv.org/abs/2310.08576"},{"title":"OpenReview: ICLR 2024 spotlight","url":"https://openreview.net/forum?id=Mhb5fpA1T0"},{"title":"AVDC 项目页","url":"https://flow-diffusion.github.io/"}],"as_of":"2024-01","related_ids":["action-free-video","unipi","video-generation-model","optical-flow","imitation-from-observation","human-video-data"],"name":"AVDC","alt":"AVDC（从无动作视频学动作）","abbr":"AVDC","aliases":["Actionless Video through Dense Correspondences","Learning to Act from Actionless Videos through Dense Correspondences"],"one_liner":"First generates a video of a robot doing the task, then uses optical flow to derive the actions — no action labels needed.","explanation":"AVDC was proposed by researchers at National Taiwan University working with MIT (including Yilun Du and Joshua Tenenbaum), published at ICLR 2024 (Spotlight). What makes robot data expensive is the action labels; video with no action labels at all is far more plentiful. Given a current image and a text instruction, AVDC first uses a text-conditioned diffusion video-generation model to “imagine” a video of the task being completed; it then estimates optical flow between adjacent frames (which way each pixel moves), treats this as a dense correspondence, combines it with the first frame's depth to compute the rigid-body pose change of the object, and finally converts that into the arm's grasping and moving actions. The policy is trained using only RGB video, and was validated on Meta-World manipulation, iTHOR navigation, and a real Franka arm; the authors also open-sourced a video-model framework that trains in a day on 4 GPUs. It belongs to the same “use video generation as the policy” line of work as UniPi.","example":"Training a video model on just 198 videos of a human hand pushing objects — with no robot actions of any kind — and then using it, with no fine-tuning, to control a simulated robot arm on a pushing task reached 90% success over 40 trials, an example of transfer from human video to a robot with a different embodiment.","related":["Action-free Video","UniPi","Video Generation Model","Optical Flow","Imitation from Observation","Human Video Data"]},{"id":"susie","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models (arXiv 2310.10639)","url":"https://arxiv.org/abs/2310.10639"},{"title":"SuSIE project page","url":"https://rail-berkeley.github.io/susie/"}],"as_of":"2023-10","related_ids":["goal-conditioned-policy","diffusion-model","hierarchical-architecture","calvin-benchmark","unipi","zero-shot"],"name":"SuSIE","alt":"SuSIE","abbr":"SuSIE","aliases":["Subgoal Synthesis via Image Editing","Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models"],"one_liner":"A method that uses an image-editing diffusion model as a high-level planner, sketching a subgoal image for a low-level policy to reach.","explanation":"SuSIE was released by Kevin Black, Sergey Levine, Chelsea Finn, and colleagues at UC Berkeley and Stanford in October 2023. The method has two levels. The high level is InstructPix2Pix, an open-source image-editing model, fine-tuned on human video and robot data: given the current camera image and a language instruction, it outputs a subgoal image showing 'what it should look like' at some point soon. The low level is a goal-conditioned policy that looks only at the target image, not the language, and is responsible for getting the robot from the current image to that subgoal image. The two levels alternate until the task is complete, letting the high level draw on internet-scale image pretraining to handle new objects and instructions, while the low level focuses purely on precise control. It achieved state-of-the-art results at the time in the zero-shot setting on the CALVIN benchmark, and outperformed RT-2-X in real-robot experiments too. SuSIE belongs to the same 'generate an image, then derive the action' family as UniPi.","example":"Given the instruction 'put the yellow block in the drawer,' SuSIE first generates an image of the gripper already holding the yellow block as a subgoal; once the low-level policy reaches that state, it generates the next subgoal image showing the block inside the drawer.","related":["Goal-conditioned Policy","Diffusion Model","Hierarchical Architecture","CALVIN Benchmark","UniPi","Zero-shot"]},{"id":"gen2act","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Gen2Act (arXiv:2409.16283)","url":"https://arxiv.org/abs/2409.16283"},{"title":"Gen2Act 项目主页","url":"https://homangab.github.io/gen2act/"}],"as_of":"2024-09","related_ids":["video-generation-model","human-video-data","tracking-any-point","behavior-cloning","unipi","dreamgen"],"name":"Gen2Act","alt":"Gen2Act","abbr":"","aliases":["Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation"],"one_liner":"A manipulation method that first generates a video of a human performing the task, then has the robot execute it by following that video.","explanation":"Gen2Act was proposed in September 2024 by researchers at Google DeepMind, Carnegie Mellon University, and Stanford. It splits “follow a language instruction to manipulate an object” into two steps: first, a video-generation model trained on web video generates, zero-shot with no fine-tuning, a video of a human completing the task in the current scene; then a robot policy watches this video and executes it, with a point-trajectory prediction loss added during training (predicting how points in the image move) so the policy learns to read motion information out of the video. Generalizing across objects and actions is handed off to the generative model, which has seen a vast amount of web video, while the robot data only has to teach “translating the action in the video into a robot action” — which lets it handle object categories and actions that never appear in the robot data itself.","example":"Action types absent from the robot data, such as stirring in a circle or dragging an object in a new direction, can be carried out with the help of a generated human video; chaining several such steps together can also accomplish long-horizon tasks like “clear the table” or “make coffee.”","related":["Video Generation Model","Human Video Data","Tracking Any Point","Behavior Cloning","UniPi","DreamGen"]},{"id":"video-prediction-policy","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Video Prediction Policy (arXiv 2412.14803)","url":"https://arxiv.org/abs/2412.14803"},{"title":"VPP project page","url":"https://video-prediction-policy.github.io/"}],"as_of":"2025-05","related_ids":["video-prediction-model","inverse-dynamics-model","diffusion-policy","calvin-benchmark","vidar","robotera"],"name":"Video Prediction Policy","alt":"视频预测策略","abbr":"VPP","aliases":["Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations"],"one_liner":"A robot policy that guides action generation using the internal 'prediction of the future' features inside a video diffusion model.","explanation":"Video Prediction Policy was released in December 2024 by Jianyu Chen's lab at Tsinghua University together with RobotEra, UC Berkeley, the Shanghai Artificial Intelligence Laboratory, and others, selected as an ICML 2025 Spotlight. A typical visual encoder only looks at a single image, or compares two images, and cannot capture 'what's about to happen next.' VPP's hypothesis is that a video diffusion model's internal features, when predicting an upcoming frame, already encode both the current image and a prediction of future dynamics. The approach first fine-tunes a pretrained video model on robot data and internet videos of human manipulation, then takes its intermediate predictive representation as a conditioning signal to train an implicit inverse dynamics model (which infers the action from the 'predicted future') that outputs the action. The paper reports an 18.6% relative improvement over the previous best method on the CALVIN ABC-D generalization benchmark, and a 31.6% success-rate improvement on complex real-robot dexterous-hand manipulation tasks.","example":"On a Franka arm and an XHand dexterous hand, VPP forms a predictive representation of the next few frames internally from a language instruction and outputs the action from that, without ever having to fully render out the whole future video.","related":["Video Prediction Model","Inverse Dynamics Model","Diffusion Policy","CALVIN Benchmark","Vidar","RobotEra"]},{"id":"seer","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation (arXiv 2412.15109)","url":"https://arxiv.org/abs/2412.15109"},{"title":"OpenRobotLab/Seer (GitHub)","url":"https://github.com/OpenRobotLab/Seer"}],"as_of":"2025-01","related_ids":["inverse-dynamics-model","video-prediction-policy","droid","calvin-benchmark","libero-benchmark","shanghai-artificial-intelligence-laboratory"],"name":"Seer","alt":"Seer（预测式逆动力学模型）","abbr":"PIDM","aliases":["Predictive Inverse Dynamics Model","Seer: Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation"],"one_liner":"An end-to-end manipulation policy that first predicts the upcoming frame, then infers the action from it using inverse dynamics.","explanation":"Seer was released by the Shanghai Artificial Intelligence Laboratory and other institutions in December 2024, selected for an oral presentation at ICLR 2025. Common approaches either imitate actions directly from the current frame (behavior cloning), or first generate future frames with a video model and separately infer actions afterward, training the two parts apart. Seer proposes the 'Predictive Inverse Dynamics Model' (PIDM): inside one end-to-end trained Transformer, it first predicts the frame the robot is about to see, then uses an inverse dynamics model (which infers what action was taken from a current state and a target state) to compute the action from 'now' and the 'predicted future,' with visual prediction and action prediction optimized jointly. It can be pretrained on large-scale robot data such as DROID, then fine-tuned on a small amount of downstream data. The paper reports a 13% improvement over the previous best method on LIBERO-LONG, 21% on CALVIN ABC-D, and 43% on real-robot tasks.","example":"While performing 'open the drawer,' Seer first predicts the frame showing the gripper approaching the handle a moment later, then computes how the arm should move from the current frame and this predicted frame.","related":["Inverse Dynamics Model","Video Prediction Policy","DROID (Distributed Robot Interaction Dataset)","CALVIN Benchmark","LIBERO Benchmark","Shanghai Artificial Intelligence Laboratory"]},{"id":"uva","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Unified Video Action Model (arXiv 2503.00200)","url":"https://arxiv.org/abs/2503.00200"},{"title":"UVA project page","url":"https://unified-video-action-model.github.io/"}],"as_of":"2025-04","related_ids":["video-prediction-policy","world-action-model","forward-dynamics-model","inverse-dynamics-model","diffusion-policy","universal-manipulation-interface"],"name":"UVA","alt":"UVA","abbr":"UVA","aliases":["Unified Video Action Model"],"one_liner":"A robot model in which video and action share one latent representation, letting inference skip video generation for speed.","explanation":"UVA was released by Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song at Stanford University in February 2025, published at RSS 2025. Using video generation to build a robot policy lets the model learn how the environment changes, but action precision and inference speed have typically lagged behind policies that output actions directly. UVA has video and action share one joint latent representation, learned together during training but decoded separately: two lightweight diffusion heads decode the future frame and the action independently, and at inference time video generation can simply be skipped and only the action decoded, which keeps it fast. Combined with masked training, the same model can then serve as a policy, a forward dynamics model, an inverse dynamics model, or a video prediction model. Across real-robot multi-task experiments like folding a towel and arranging cups, it beats a diffusion-policy baseline trained on UMI data when facing unseen environments and objects.","example":"While a robot folds a towel, UVA decodes only the next action and skips frame generation to stay real-time; offline, that same model can also predict the resulting image given a specified action.","related":["Video Prediction Policy","World Action Model","Forward Dynamics Model","Inverse Dynamics Model","Diffusion Policy","Universal Manipulation Interface"]},{"id":"uwm","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Unified World Models (arXiv 2504.02792)","url":"https://arxiv.org/abs/2504.02792"},{"title":"UWM project page","url":"https://weirdlabuw.github.io/uwm/"}],"as_of":"2025-05","related_ids":["world-model","diffusion-model","action-free-video","inverse-dynamics-model","droid","uva"],"name":"UWM","alt":"UWM（统一世界模型）","abbr":"UWM","aliases":["Unified World Models","Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets"],"one_liner":"A robot framework that merges action diffusion and video diffusion into one model and can pretrain on action-free video.","explanation":"UWM was released in April 2025 by the University of Washington (Abhishek Gupta's group) and Toyota Research Institute, published at RSS 2025. Imitation learning needs robot data labeled with actions, while the huge amount of video without action labels is hard to use directly. UWM runs action diffusion and video diffusion inside one Transformer at the same time, giving each modality its own independent denoising timestep: setting a modality's timestep to pure noise is equivalent to not looking at it at all. This lets the same model act as a policy, a forward dynamics model, an inverse dynamics model, or a video predictor, just by combining timesteps differently; for action-free video, the missing action is simply treated as fully noised during training. Pretrained on the large-scale DROID dataset and then fine-tuned, it generalizes better than pretraining with plain behavior cloning, and adding action-free video on top improves results further.","example":"Robot trajectories from DROID are pretrained together with a batch of manipulation videos that have no action labels to train UWM, which is then fine-tuned into a task-specific policy using a small number of demonstrations.","related":["World Model","Diffusion Model","Action-free Video","Inverse Dynamics Model","DROID (Distributed Robot Interaction Dataset)","UVA"]},{"id":"vidar","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Vidar (arXiv 2507.12898)","url":"https://arxiv.org/abs/2507.12898"},{"title":"Vidar 论文 HTML v4（机构与 Vidu 2.0 基座）","url":"https://arxiv.org/html/2507.12898v4"},{"title":"Vidar & AnyPos project page","url":"https://embodiedfoundation.github.io/vidar_anypos"}],"as_of":"2025-12","related_ids":["video-generation-model","inverse-dynamics-model","video-prediction-policy","cross-embodiment","world-action-model","shengshu-technology"],"name":"Vidar","alt":"生数 Vidar","abbr":"Vidar","aliases":["Embodied Video Diffusion Model for Generalist Manipulation","Vidar: VIdeo Diffusion for Action Reasoning"],"one_liner":"A robot manipulation model that first predicts future frames with a video diffusion model, then infers the action from them.","explanation":"Vidar was released in July 2025 by Jun Zhu's lab (TSAIL) at Tsinghua University; its name stands for VIdeo Diffusion for Action Reasoning. Real-robot experiments are built on Vidu 2.0, the video model from ShengShu Technology, and the paper notes that part of the work was done at ShengShu, with reports describing it as a joint release between ShengShu and Tsinghua. The approach has two steps: a video diffusion model first 'imagines' an upcoming video from the instruction and current image; then a masked inverse dynamics model (MIDM), which infers actions from consecutive frames, translates that video into robot actions, and automatically learns to focus only on action-relevant pixels such as the robot arm, filtering out background clutter. The video model first goes through continued embodied-domain pretraining on 750,000 multi-view trajectories across three robot platforms. On a robot it has never seen, the paper reports that only about 20 minutes of human demonstration is needed to beat baseline methods, and it generalizes to new tasks, backgrounds, and camera layouts.","example":"A new Aloha bimanual robot is adapted with only about 20 minutes of human demonstration; Vidar then generates a video of the task being completed from a new language instruction, and an inverse dynamics model translates that video segment by segment into joint actions for execution.","related":["Video Generation Model","Inverse Dynamics Model","Video Prediction Policy","Cross-Embodiment","World Action Model","ShengShu Technology"]},{"id":"ctrl-world","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Ctrl-World (arXiv:2510.10125)","url":"https://arxiv.org/abs/2510.10125"},{"title":"Ctrl-World 项目主页（ICLR 2026）","url":"https://ctrl-world.github.io/"}],"as_of":"2026-03","related_ids":["world-model","world-model-based-policy-evaluation","droid","pi0-5","stable-video-diffusion","synthetic-data"],"name":"Ctrl-World","alt":"Ctrl-World","abbr":"","aliases":["Ctrl-World: A Controllable Generative World Model for Robot Manipulation"],"one_liner":"A manipulation world model that generates multi-view future frames from actions, used to evaluate and improve policies “in imagination.”","explanation":"Ctrl-World was proposed in October 2025 by Chelsea Finn's group at Stanford with Jianyu Chen's group at Tsinghua, accepted at ICLR 2026. Evaluating and improving a general-purpose robot policy usually needs a large number of real-robot trials, which is slow and expensive. Ctrl-World is initialized from the 1.5-billion-parameter video diffusion model Stable Video Diffusion, and trained on the DROID dataset (about 95,000 trajectories across 564 scenes) into a world model that generates future frames given an action, with three modifications: jointly predicting two third-person views plus a wrist camera; sparsely sampling history frames and embedding the arm's pose so the model can retrieve relevant past frames by pose and stay consistent over long horizons; and injecting actions frame by frame for centimeter-level control precision. A policy can interact inside it continuously for more than 20 seconds; it's used to rank π0, π0-FAST, and π0.5, matching real-robot results, and fine-tuning π0.5 on successful trajectories picked from imagined rollouts raised its success rate on unfamiliar instructions and objects from 38.7% to 83.4%.","example":"Given π0.5 an instruction it hasn't practiced, running 400 imagined trials in Ctrl-World with reworded instructions and randomized arm resets, then manually picking 25–50 successful trajectories and fine-tuning the policy on them for 2,000 steps, is enough to meaningfully improve it.","related":["World Model","World-Model-based Policy Evaluation","DROID (Distributed Robot Interaction Dataset)","π0.5","Stable Video Diffusion","Synthetic Data"]},{"id":"motus","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Motus: A Unified Latent Action World Model (arXiv 2512.13030)","url":"https://arxiv.org/abs/2512.13030"},{"title":"thu-ml/Motus (GitHub)","url":"https://github.com/thu-ml/Motus"},{"title":"Motus project page","url":"https://motus-robotics.github.io/motus"}],"as_of":"2025-12","related_ids":["world-action-model","latent-action","mixture-of-transformers","inverse-dynamics-model","data-pyramid","robotwin"],"name":"Motus","alt":"Motus","abbr":"","aliases":["A Unified Latent Action World Model"],"one_liner":"A latent-action world model that combines understanding, video generation, and action into three experts inside one model.","explanation":"Motus was released in December 2025 by Jun Zhu's group at Tsinghua University together with Peking University and Horizon Robotics, with code and weights open-sourced under Apache 2.0. Earlier approaches often split understanding, a world model (predicting future frames), and control into separate models, making it hard to jointly exploit large-scale heterogeneous data. Motus uses a Mixture-of-Transformers (MoT) architecture to connect an understanding expert, a video-generation expert, and an action expert together — the video part is built on Wan2.2-5B and the understanding part on Qwen3-VL-2B, about 8 billion parameters in total — and uses UniDiffuser-like scheduling to switch among modes such as world model, VLA, inverse dynamics model, and video generation. It learns latent actions from optical flow, letting video with no action labels also join pretraining, paired with three-stage training and a six-tier data pyramid. Across 50 tasks on RoboTwin 2.0, it reaches an average success rate of about 87%, which the paper reports as 15% higher than X-VLA and 45% higher than π0.5.","example":"The same Motus model: given the current frame and an instruction, it outputs an action directly (VLA mode); given a frame and an action, it predicts the following video (world-model mode); given a before-and-after pair of frames, it infers what action happened in between (inverse-dynamics mode).","related":["World Action Model","Latent Action","Mixture-of-Transformers","Inverse Dynamics Model","Data Pyramid","RoboTwin"]},{"id":"fast-wam","category":"named_model","sec":8,"tier":3,"sources":[{"title":"Fast-WAM: Do World Action Models Need Test-time Future Imagination? (arXiv 2603.16666)","url":"https://arxiv.org/abs/2603.16666"}],"as_of":"2026-03","related_ids":["world-action-model","video-generation-model","action-expert","inference-latency","lingbot-va","motus"],"name":"Fast-WAM","alt":"Fast-WAM","abbr":"","aliases":["Do World Action Models Need Test-time Future Imagination?"],"one_liner":"A 2026 Tsinghua and Galaxea world action model that learns video prediction during training but skips imagining the future at inference, making it faster.","explanation":"Fast-WAM is a paper released in March 2026 by Tianyuan Yuan, Hang Zhao, and colleagues at Tsinghua University's Institute for Interdisciplinary Information Sciences together with Galaxea. Most world action models (WAMs, models that jointly predict future video and robot actions) follow an “imagine, then act” approach: a video diffusion model generates a few future frames, and actions are produced from those, with the repeated denoising making inference slow. The authors separated two questions to test independently: whether to jointly learn video prediction during training, and whether to actually generate future frames at inference. They found that dropping the imagination step at inference barely hurts performance, while dropping joint video training during training hurts it substantially — showing that video prediction's real value lies in training a better world representation, not in generating images at test time. The model uses the Wan2.2-5B video diffusion Transformer as its backbone plus a roughly 1-billion-parameter action expert, about 6 billion parameters in total; inference latency is 190 milliseconds, more than 4 times faster than comparable “imagine, then act” models.","example":"With no robot-data pretraining at all, Fast-WAM reaches a 91.8% success rate on RoboTwin 2.0 and an average of 97.6% on LIBERO, and completed a long-horizon real-robot towel-folding task on the Galaxea R1 Lite.","related":["World Action Model","Video Generation Model","Action Expert","Inference Latency","LingBot-VA (Robbyant)","Motus"]},{"id":"flux-3-action","category":"named_model","sec":8,"tier":3,"sources":[{"title":"FLUX 3 Action 模型页（Black Forest Labs）","url":"https://bfl.ai/models/flux-3-action"},{"title":"Black Forest Labs 官网","url":"https://bfl.ai/"}],"as_of":"2026-09","related_ids":["world-action-model","open-weight-model","cosmos-3","fast-wam","droid","so-100-so-101-arm"],"name":"FLUX 3 Action","alt":"FLUX 3 Action（Black Forest Labs）","abbr":"","aliases":["Black Forest Labs open-weight world action model"],"one_liner":"A 7B open-weight world action model from Germany's Black Forest Labs, released September 2026, predicting frames and actions together.","explanation":"FLUX 3 Action is a robot model released in September 2026 by Black Forest Labs, the German company known for its FLUX line of image-generation models, with 7 billion parameters and weights open on Hugging Face. It is a world action model (WAM, a model that jointly predicts future video frames and robot actions), adapted from the company's multimodal FLUX 3 backbone, which was pretrained on large amounts of video, image, and audio data, mostly video. The model takes in camera images, joint state, and a text instruction, and outputs both a robot action and a predicted future frame at the same time, aiming to carry the physical common sense learned from video pretraining over into control. The company says it supports multiple action spaces, including end-effector pose and joint angles, and demonstrated it on a Franka arm and the SO-101. It is an example of an image- and video-generation company moving into embodied AI.","example":"The company reports that a guidance-distilled version reaches a 42.2% success rate on the RoboLab-120 simulation leaderboard, ahead of NVIDIA Cosmos 3 Nano's 36.8%, and a 93.3% success rate across 10 real-robot DROID manipulation tasks.","related":["World Action Model","Open-weight Model","Cosmos 3","Fast-WAM","DROID (Distributed Robot Interaction Dataset)","SO-100 / SO-101 Arm"]},{"id":"anymal-rl-locomotion-series","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Learning agile and dynamic motor skills for legged robots (Hwangbo et al., Science Robotics 2019, arXiv 1901.08652)","url":"https://arxiv.org/abs/1901.08652"},{"title":"Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020, arXiv 2010.11251)","url":"https://arxiv.org/abs/2010.11251"},{"title":"Learning robust perceptive locomotion for quadrupedal robots in the wild (Miki et al., Science Robotics 2022, arXiv 2201.08117)","url":"https://arxiv.org/abs/2201.08117"}],"as_of":"2022-01","related_ids":["rl-based-locomotion-control","actuator-modeling","teacher-student-distillation","privileged-information","terrain-curriculum","anybotics-anymal"],"name":"ANYmal RL Locomotion Series","alt":"ANYmal 强化学习运控系列（执行器网络 / 教师-学生盲走 / 感知行走）","abbr":"","aliases":["Learning Agile and Dynamic Motor Skills for Legged Robots","Learning Quadrupedal Locomotion over Challenging Terrain","Learning Robust Perceptive Locomotion for Quadrupedal Robots in the Wild"],"one_liner":"Three Science Robotics papers (2019, 2020, 2022) from ETH Zurich that took simulation-trained RL locomotion out into the real wild on the ANYmal quadruped.","explanation":"This refers to three papers published in Science Robotics by ETH Zurich's Robotic Systems Lab (Marco Hutter's group), all using the ANYmal quadruped as their platform. In 2019, Hwangbo and colleagues introduced actuator networks: a neural network trained on real-robot data to simulate motor response, plugged into the simulator to shrink the sim-to-real gap, letting a policy trained purely in simulation run directly on the real robot — 25% faster than the previous record, and able to get back up after falling. In 2020, Lee and colleagues used teacher-student distillation: a teacher is trained with access to privileged information such as ground-truth terrain, and a student that uses only proprioception learns to imitate it, combined with a terrain curriculum, so the robot can walk over mud, snow, and rubble without seeing the terrain at all (blind locomotion); this controller was used in the DARPA Subterranean Challenge. In 2022, Miki and colleagues used an attention-based recurrent encoder to fuse a height map with proprioception, automatically leaning more on proprioception whenever vision becomes unreliable.","example":"The 2022 perceptive-locomotion controller let an ANYmal complete a 2.2-kilometer hiking route rated “difficult,” with 120 meters of elevation gain, on Switzerland's Mount Etzel in 78 minutes — almost exactly the 76 minutes a hiking-planning tool would recommend for a person — stopping only once to repair a dislodged foot cover and swap batteries.","related":["RL-based Locomotion Control","Actuator Modeling (Actuator Network)","Teacher-Student Distillation","Privileged Information","Terrain Curriculum","ANYbotics ANYmal"]},{"id":"rapid-motor-adaptation","category":"named_model","sec":9,"tier":2,"sources":[{"title":"RMA: Rapid Motor Adaptation for Legged Robots 项目主页","url":"https://ashish-kmr.github.io/rma-legged-robots/"}],"as_of":"2021-07","related_ids":["sim-to-real-transfer","privileged-information","teacher-student-distillation","rl-based-locomotion-control","domain-randomization","hora"],"name":"Rapid Motor Adaptation","alt":"快速运动适应","abbr":"RMA","aliases":["RMA"],"one_liner":"A method that lets quadruped robots sense terrain or load changes and adjust gait within a fraction of a second.","explanation":"RMA was proposed by Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik at UC Berkeley and Carnegie Mellon University, published at RSS 2021. Real ground varies in softness, friction, load, and motor wear, and a policy trained only in simulation often falls on the real robot. RMA has two parts. A base policy, trained in simulation, has access to these environment parameters (privileged information) and compresses them into a low-dimensional “extrinsics” vector that modulates its actions. An adaptation module sees only a recent window of joint state and action history, and learns to infer that same vector. Deployed on a Unitree A1 with no real-robot fine-tuning, it adapts within a fraction of a second to sand, foam mats, stairs, and slippery ground. This approach — using privileged information during training and estimating it from history at deployment — later became a standard recipe for sim-to-real transfer in legged robots.","example":"A sudden load is added to a Unitree A1's back; the adaptation module estimates the new extrinsics vector from the recent change in joint states, and the policy immediately adjusts its gait and keeps walking, with no retraining needed.","related":["Sim-to-Real Transfer","Privileged Information","Teacher-Student Distillation","RL-based Locomotion Control","Domain Randomization","HORA"]},{"id":"walk-these-ways","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Walk These Ways (arXiv 2212.03238)","url":"https://arxiv.org/abs/2212.03238"},{"title":"Walk These Ways 项目主页","url":"https://gmargo11.github.io/walk-these-ways/"},{"title":"Improbable-AI/walk-these-ways (GitHub)","url":"https://github.com/Improbable-AI/walk-these-ways"}],"as_of":"2022-12","related_ids":["legged-locomotion","gait","rl-based-locomotion-control","isaac-gym","unitree-go1","sim-to-real-transfer"],"name":"Walk These Ways","alt":"Walk These Ways","abbr":"","aliases":["Multiplicity of Behavior (MoB) Controller","Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior"],"one_liner":"A quadruped locomotion controller that learns many gaits in one policy and can be retuned live at deployment to handle new terrain.","explanation":"Walk These Ways is a quadruped locomotion method from Gabriel Margolis and Pulkit Agrawal at MIT's Improbable AI Lab, posted to arXiv in December 2022 and an oral-presentation paper at CoRL 2022. When a reinforcement-learning-trained walking policy fails outside its training distribution, the usual fix is to go back, tweak the reward or environment, and retrain — a slow cycle. The authors instead propose 'Multiplicity of Behavior' (MoB): train a single policy that takes both a velocity command and a set of behavior parameters, including footfall timing (which can switch between trotting, bounding, pacing, and other gaits), stepping frequency, body height, body pitch, stance width, and foot-lift height. Different parameter combinations complete the same task in different ways with different generalization properties, and a person or a script can switch between them live at deployment without retraining. The policy is trained in Isaac Gym and deployed on the Unitree Go1 at a 50 Hz control rate; the code is open-sourced and has often served as a starting point for quadruped locomotion work since.","example":"The paper shows that raising the stepping frequency helps when sprinting on slippery ground; climbing stairs works better with a low stepping frequency and high foot lift; and when pushed by a person, lowering the foot lift and widening the stance improves stability.","related":["Legged Locomotion","Gait","RL-based Locomotion Control","Isaac Gym","Unitree Go1","Sim-to-Real Transfer"]},{"id":"dreamwaq","category":"named_model","sec":9,"tier":3,"sources":[{"title":"DreamWaQ (arXiv 2301.10602)","url":"https://arxiv.org/abs/2301.10602"}],"as_of":"2023-05","related_ids":["blind-locomotion","proprioception","asymmetric-actor-critic","privileged-information","variational-autoencoder","unitree-a1"],"name":"DreamWaQ","alt":"DreamWaQ","abbr":"","aliases":["Learning Robust Quadrupedal Locomotion With Implicit Terrain Imagination via Deep Reinforcement Learning"],"one_liner":"A 2023 KAIST reinforcement-learning method for blind quadruped locomotion that “imagines” the terrain underfoot using only proprioception.","explanation":"DreamWaQ is a locomotion control method for quadruped robots from Hyun Myung's group at KAIST (Korea Advanced Institute of Science and Technology), published at ICRA 2023. Many approaches to walking over complex terrain rely on a camera or lidar, but these sensors can fail in bad weather or poor lighting. DreamWaQ instead walks blind, using only proprioception (the robot's own internal signals, such as joint angles, angular velocities, and body orientation): it trains a context-aided estimator network (CENet) that estimates the body's velocity from the last few steps of observation, and uses a variational autoencoder to infer a latent variable representing the terrain — effectively an “implicit imagination” of what's underfoot; the policy then combines these estimates to output joint actions. Training is done in Isaac Gym with 4,096 domain-randomized parallel environments, using an asymmetric actor-critic — where the critic can see privileged information available only in simulation, while the policy sees only what a real robot could observe — before being transferred zero-shot to the real robot.","example":"On the Unitree A1 quadruped, the DreamWaQ policy, without any external perception, was able to cross complex terrain such as steps during a single long-distance continuous outdoor walk.","related":["Blind Locomotion","Proprioception","Asymmetric Actor-Critic","Privileged Information","Variational Autoencoder","Unitree A1"]},{"id":"barkour","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Barkour: Benchmarking Animal-level Agility with Quadruped Robots (arXiv 2305.14654)","url":"https://arxiv.org/abs/2305.14654"},{"title":"Barkour: Benchmarking animal-level agility with quadruped robots (Google Research Blog)","url":"https://research.google/blog/barkour-benchmarking-animal-level-agility-with-quadruped-robots/"}],"as_of":"2023-05","related_ids":["quadruped-robot","legged-locomotion","parkour","benchmark","teacher-student-distillation","rl-based-locomotion-control"],"name":"Barkour","alt":"Barkour","abbr":"","aliases":["Barkour: Benchmarking Animal-level Agility with Quadruped Robots"],"one_liner":"A quadruped-robot agility benchmark from Google, modeled on dog agility trials, for measuring animal-level agility in a repeatable way.","explanation":"Barkour is a quadruped-robot agility benchmark released by the Google Research team (later folded into Google DeepMind) in May 2023, its name blending “bark” and “parkour.” Modeled on canine agility trials, it sets up weave poles, an A-frame ramp, a 0.5-meter broad jump, and a finish-line table in a 5-meter-by-5-meter course, scoring 0 to 1 based on whether each obstacle is cleared and whether the time comes close to that of a small dog (about 10 seconds). The team first trained specialized skills — walking, climbing, jumping — separately with reinforcement learning, then used teacher-student distillation to combine them into a single Transformer-based generalist locomotion policy, which completed the course on the team's own quadruped robot in about 20 seconds, roughly half a small dog's speed. Its contribution is giving “animal-level agility” a quantifiable, reproducible way to compare systems.","example":"The team's own quadruped robot typically completes the Barkour course in about 20 seconds, versus about 10 seconds for a small dog.","related":["Quadruped Robot","Legged Locomotion","Parkour","Benchmark","Teacher-Student Distillation","RL-based Locomotion Control"]},{"id":"robot-parkour-learning","category":"named_model","sec":9,"tier":3,"sources":[{"title":"arXiv 2309.05665: Robot Parkour Learning","url":"https://arxiv.org/abs/2309.05665"},{"title":"Robot Parkour Learning 项目主页","url":"https://robot-parkour.github.io/"}],"as_of":"2023-11","related_ids":["parkour","perceptive-locomotion","teacher-student-distillation","dagger","sim-to-real-transfer","extreme-parkour"],"name":"Robot Parkour Learning","alt":"Robot Parkour Learning","abbr":"","aliases":["Quadruped Robot Parkour Learning"],"one_liner":"A parkour policy that lets a low-cost quadruped robot dog climb, jump, crawl, and squeeze through obstacles using only a depth camera.","explanation":"Robot Parkour Learning was released by the Shanghai Qi Zhi Institute, Tsinghua University, Stanford, CMU, ShanghaiTech, and other institutions in September 2023, receiving an oral presentation at CoRL 2023 and reaching the finals for the Best Systems Paper Award. It teaches low-cost quadruped robots like Unitree's A1 and Go1 to climb high obstacles, leap wide gaps, crawl under low barriers, squeeze through narrow gaps, and run. The method has three steps: first, reinforcement learning pretraining in simulation, where the robot is allowed to 'pass through' obstacles, treating physics violations as only a soft penalty to encourage exploration; then fine-tuning each individual skill with full physical constraints restored; finally, using DAgger to distill the several specialized skills into a single vision-based policy that relies only on an onboard depth camera. Once deployed on a real robot, it can choose the appropriate action for the obstacle in front of it on its own, without relying on animal motion references.","example":"A robot dog encounters a low horizontal bar; based on its depth image, the policy automatically chooses to lower its body and crawl under it rather than trying to jump over it.","related":["Parkour","Perceptive Locomotion","Teacher-Student Distillation","DAgger","Sim-to-Real Transfer","Extreme Parkour"]},{"id":"extreme-parkour","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Extreme Parkour with Legged Robots (arXiv 2309.14341)","url":"https://arxiv.org/abs/2309.14341"},{"title":"Extreme Parkour 项目主页","url":"https://extreme-parkour.github.io/"}],"as_of":"2023-09","related_ids":["parkour","robot-parkour-learning","perceptive-locomotion","teacher-student-distillation","unitree-a1","sim-to-real-transfer"],"name":"Extreme Parkour","alt":"Extreme Parkour","abbr":"","aliases":["Extreme Parkour with Legged Robots"],"one_liner":"A 2023 CMU project where a low-cost quadruped does parkour — jumping high and far — using just one depth camera and a single network.","explanation":"Extreme Parkour is work released in September 2023 by Xuxin Cheng, Deepak Pathak, and colleagues at Carnegie Mellon University, published at ICRA 2024. They used a low-cost Unitree A1 quadruped, whose motor control isn't especially precise, with just one Intel RealSense D435 depth camera on its head producing low-frequency, jittery, noisy footage. Where traditional approaches carefully design perception, mapping, planning, and control as separate modules, this work goes end-to-end: a neural network is trained with large-scale reinforcement learning in simulation to output joint actions directly from the depth image. Training happens in two stages — first using precise terrain information only available in simulation, then distilled into a policy that works from the depth image alone and can even decide which direction to jump; the reward is mostly unified into a single term, the dot product between the velocity direction and the goal direction, with no need to hand-design a reward for each type of obstacle. It pushed the difficulty ceiling for legged-robot parkour up substantially, and is often cited alongside the contemporaneous Robot Parkour Learning.","example":"The A1 robot dog can jump onto a 0.5-meter-tall box (about twice its own height), leap across a 0.8-meter-wide gap (about twice its body length), and even walk upside down on just its two front legs.","related":["Parkour","Robot Parkour Learning","Perceptive Locomotion","Teacher-Student Distillation","Unitree A1","Sim-to-Real Transfer"]},{"id":"him","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Hybrid Internal Model (arXiv 2312.11460)","url":"https://arxiv.org/abs/2312.11460"},{"title":"OpenRobotLab/HIMLoco (GitHub)","url":"https://github.com/OpenRobotLab/HIMLoco"}],"as_of":"2024-01","related_ids":["legged-locomotion","proprioception","blind-locomotion","contrastive-learning","rapid-motor-adaptation","shanghai-artificial-intelligence-laboratory"],"name":"HIM","alt":"HIM（混合内部模型）","abbr":"HIM","aliases":["Hybrid Internal Model","HIMLoco","Learning Agile Legged Locomotion with Simulated Robot Response"],"one_liner":"A legged-locomotion method that uses only proprioception, inferring terrain and disturbances from the robot's own response.","explanation":"HIM was released in December 2023 by Shanghai AI Lab's OpenRobotLab (Jiangmiao Pang's group), published at ICLR 2024, with its code repository named HIMLoco. A legged robot's sensors only ever give incomplete, noisy observations, and external state such as ground friction and terrain height is hard to estimate directly. HIM borrows the idea of internal model control from classical control theory, treating this external state as a disturbance to be inferred from the robot's own response: a “hybrid internal embedding” represents both explicit body velocity and implicit stability information at once, and contrastive learning is used to pull it close to the robot's actual next-moment state. It needs only proprioception from joint encoders and the IMU, skipping the usual two-stage teacher-student imitation process; the paper reports that training for about an hour on a single RTX 4090 is enough to let a quadruped handle a variety of terrains and external-force disturbances.","example":"A quadruped relying only on joint encoders and the IMU can keep walking under terrain and external shoves it never saw during training.","related":["Legged Locomotion","Proprioception","Blind Locomotion","Contrastive Learning","Rapid Motor Adaptation","Shanghai Artificial Intelligence Laboratory"]},{"id":"abs","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion (arXiv 2401.17583)","url":"https://arxiv.org/abs/2401.17583"},{"title":"ABS 项目页","url":"https://agile-but-safe.github.io/"},{"title":"Robotics: Science and Systems XX (RSS 2024) 论文集","url":"https://www.roboticsproceedings.org/rss20/index.html"}],"as_of":"2024-05","related_ids":["legged-locomotion","rl-based-locomotion-control","obstacle-avoidance","safe-reinforcement-learning","hamilton-jacobi-reachability-analysis","unitree-go1"],"name":"ABS","alt":"ABS（敏捷且安全的足式运动）","abbr":"ABS","aliases":["Agile But Safe","Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion"],"one_liner":"CMU's 2024 quadruped framework: sprint at high speed, switching to a learned collision-avoidance policy whenever a learned safety value says to.","explanation":"ABS was developed by Guanya Shi and Changliu Liu's groups at Carnegie Mellon University with ETH Zurich, published at RSS 2024. Earlier quadruped obstacle-avoidance controllers, in the name of safety, usually stayed under 1 m/s; controllers built purely for agility didn't worry about collisions at all. ABS uses two policies: an agile policy for high-speed running between obstacles, and a recovery policy for emergency avoidance; a learned reach-avoid value network estimates in real time whether the current state is safe, decides when to switch between the two, and supplies the optimization target for the recovery policy. Obstacle information is compressed into distances along 11 rays, predicted by a network from a depth image. Each module is trained in Isaac Gym and deployed to a real Unitree Go1, running entirely on onboard sensing and compute, reaching a peak speed of 3.1 m/s. It's a representative example of combining control-theoretic safety guarantees with reinforcement-learning-based locomotion.","example":"A Unitree Go1 sprints down a dim hallway at an average of 1.5 m/s and a peak of 2.5 m/s; when someone walks toward it or suddenly sticks out a leg to block it, a drop in the safety value triggers a switch to the recovery policy to dodge, then switches back to the agile policy to keep running once it's safe again.","related":["Legged Locomotion","RL-based Locomotion Control","Obstacle Avoidance","Safe Reinforcement Learning","Hamilton-Jacobi Reachability Analysis","Unitree Go1"]},{"id":"dial-mpc","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing (arXiv:2409.15610)","url":"https://arxiv.org/abs/2409.15610"},{"title":"DIAL-MPC 项目主页（LeCAR Lab）","url":"https://lecar-lab.github.io/dial-mpc/"}],"as_of":"2025","related_ids":["sampling-based-mpc","model-predictive-path-integral-control","model-predictive-control","legged-locomotion","diffusion-model","unitree-go2"],"name":"DIAL-MPC","alt":"DIAL-MPC（扩散式退火足式 MPC）","abbr":"DIAL-MPC","aliases":["Diffusion-Inspired Annealing for Legged MPC","Full-Order Sampling-Based MPC for Torque-Level Locomotion Control via Diffusion-Style Annealing"],"one_liner":"A training-free, sampling-based MPC for legged robots that borrows diffusion-style annealing to optimize whole-body motion in real time.","explanation":"DIAL-MPC was proposed by Guanya Shi's group (LeCAR Lab) at Carnegie Mellon University in September 2024, and was a best-paper finalist at ICRA 2025. Real-time optimal control for legged robots usually requires a simplified model (such as a single rigid body) or a pre-specified contact schedule, because the full dynamics are high-dimensional and non-convex. Sampling-based MPC methods (such as MPPI) sample many candidate action sequences every control cycle, roll them out in a model, and average them by weight — but a single round of sampling is often noisy or gets stuck in a local solution. The authors identify a connection between MPPI and single-step diffusion denoising, and change it to iterate over multiple rounds the way a diffusion model does, annealing the noise level step by step: searching broadly at first, then converging to fine detail. It needs no training and no model simplification, doing real-time, torque-level control directly on the full-order dynamics.","example":"On a Unitree Go2 quadruped, DIAL-MPC performed precise, loaded jumps and trajectory tracking in real time; the paper reports 13.4x lower tracking error than standard MPPI and about 50% better performance than a reinforcement-learning policy on a climbing task, with no training at all.","related":["Sampling-based MPC","Model Predictive Path Integral Control","Model Predictive Control","Legged Locomotion","Diffusion Model","Unitree Go2"]},{"id":"dactyl","category":"named_model","sec":9,"tier":3,"sources":[{"title":"arXiv 1808.00177: Learning Dexterous In-Hand Manipulation","url":"https://arxiv.org/abs/1808.00177"},{"title":"arXiv 1910.07113: Solving Rubik's Cube with a Robot Hand","url":"https://arxiv.org/abs/1910.07113"}],"as_of":"2019-10","related_ids":["automatic-domain-randomization","domain-randomization","sim-to-real-transfer","in-hand-manipulation","shadow-dexterous-hand","proximal-policy-optimization"],"name":"Dactyl","alt":"Dactyl（OpenAI 魔方灵巧手）","abbr":"","aliases":["OpenAI Dactyl","Learning Dexterous In-Hand Manipulation","Solving Rubik's Cube with a Robot Hand"],"one_liner":"OpenAI's project that trained a five-fingered robot hand entirely in simulation with reinforcement learning, then transferred it directly to the real hand.","explanation":"Dactyl is OpenAI's dexterous manipulation project, built on the Shadow Dexterous Hand, a five-fingered robotic hand. The 2018 paper 'Learning Dexterous In-Hand Manipulation' had the hand reorient a block to a target pose in its palm: the policy was trained entirely in simulation with reinforcement learning and transferred straight to the real hand using domain randomization (randomizing physical parameters such as friction and object appearance during training), with no human demonstrations at all; human-like behaviors such as finger gaiting emerged on their own. The 2019 follow-up, 'Solving Rubik's Cube with a Robot Hand,' introduced automatic domain randomization (ADR), which automatically widens the randomization ranges as training progresses. A Kociemba solver computes the cube-solving move sequence; the neural network only handles the physical manipulation. The paper reports about 60% success at 15 face rotations and about 20% at the hardest 26-rotation sequences. Dactyl is a landmark result for both sim-to-real transfer and dexterous manipulation.","example":"In the Rubik's Cube experiments, the cube's pose was estimated from three camera views with a CNN, individual face angles were read from a sensor-equipped 'Giiker' smart cube, and fingertip positions came from a motion-capture system — all fed into a recurrent neural network policy that controlled the fingers.","related":["Automatic Domain Randomization","Domain Randomization","Sim-to-Real Transfer","In-hand Manipulation","Shadow Dexterous Hand","Proximal Policy Optimization"]},{"id":"hora","category":"named_model","sec":9,"tier":3,"sources":[{"title":"In-Hand Object Rotation via Rapid Motor Adaptation (arXiv 2210.04887)","url":"https://arxiv.org/abs/2210.04887"},{"title":"HORA project page","url":"https://haozhi.io/hora/"}],"as_of":"2022-10","related_ids":["rapid-motor-adaptation","in-hand-manipulation","dexterous-manipulation","privileged-information","sim-to-real-transfer","allegro-hand"],"name":"HORA","alt":"HORA（手内物体旋转 + RMA）","abbr":"HORA","aliases":["In-Hand Object Rotation via Rapid Motor Adaptation"],"one_liner":"A reinforcement-learning method that uses only fingertip and joint proprioception to keep rotating a variety of objects in-hand.","explanation":"HORA was released in October 2022 by Haozhi Qi, Jitendra Malik, and colleagues at UC Berkeley and Meta AI, published at CoRL 2022. The goal is to have the Allegro four-fingered dexterous hand rotate an object continuously about one axis using only its fingertips. The method follows Rapid Motor Adaptation (RMA): first, a base policy is trained with reinforcement learning in simulation, with access to privileged information such as the object's size, mass, and friction; then an adaptation module is trained to infer those same object properties using only a recent history of joint proprioception. The policy is trained in simulation on cylinders alone, yet deploys to the real robot with no fine-tuning, rotating objects that vary in size, shape, and weight; a stable “finger gait” emerges naturally during training. It is a representative example of sim-to-real transfer for in-hand manipulation.","example":"A policy that only ever saw cylinders in simulation, once deployed on a real Allegro hand, can rotate everyday objects of varying shapes and weights without using a camera at all.","related":["Rapid Motor Adaptation","In-hand Manipulation","Dexterous Manipulation","Privileged Information","Sim-to-Real Transfer","Allegro Hand"]},{"id":"visual-dexterity","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes (arXiv 2211.11744)","url":"https://arxiv.org/abs/2211.11744"},{"title":"Visual Dexterity 项目主页","url":"https://taochenshh.github.io/projects/visual-dexterity"}],"as_of":"2023-11","related_ids":["in-hand-manipulation","dexterous-manipulation","teacher-student-distillation","sim-to-real-transfer","dactyl","hora"],"name":"Visual Dexterity","alt":"Visual Dexterity（视觉手内重定向）","abbr":"","aliases":["Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes"],"one_liner":"MIT's system that uses a single depth camera to let a low-cost dexterous hand reorient unseen objects in the air in real time.","explanation":"Visual Dexterity is work by Tao Chen, Pulkit Agrawal, and colleagues at MIT CSAIL's Improbable AI Lab, posted to arXiv in November 2022 and published in Science Robotics in 2023. In-hand reorientation means turning an object to any target orientation using only the fingers, with no tabletop to rest on — a hard problem because of complex contact and self-occlusion by the hand. The method first trains a 'teacher' in simulation with reinforcement learning that can read the object's true state, then distills it into a 'student' that works only from depth-camera point clouds (processed with a sparse convolutional network, running at about 12 Hz in real time). Training used about 150 objects, on hardware built from the open-source D'Claw three-fingered (9-DoF) or four-fingered hand, for a total system cost under $5,000; it can even reorient objects while palm-down, fighting gravity, with a median completion time of about 7 seconds. This shows that simulation training combined with visual input can generalize to novel object shapes, making it a landmark in in-hand manipulation after Dactyl.","example":"With its palm facing down, the robot hand holds a plastic duck it never saw during training in midair and reorients it to a target pose using only depth-camera point clouds; the duck was dropped in 56% of trials, and when it wasn't dropped, about 75% of trials landed within 23 degrees of the target.","related":["In-hand Manipulation","Dexterous Manipulation","Teacher-Student Distillation","Sim-to-Real Transfer","Dactyl","HORA"]},{"id":"dextrah-g","category":"named_model","sec":9,"tier":3,"sources":[{"title":"DextrAH-G (arXiv:2407.02274)","url":"https://arxiv.org/abs/2407.02274"},{"title":"DextrAH-G 项目主页","url":"https://sites.google.com/view/dextrah-g"}],"as_of":"2024-10","related_ids":["geometric-fabrics","dexterous-manipulation","teacher-student-distillation","privileged-information","sim-to-real-transfer","isaac-gym"],"name":"DextrAH-G","alt":"DextrAH-G","abbr":"","aliases":["DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric Fabrics"],"one_liner":"NVIDIA's arm-and-hand dexterous grasping system, trained entirely in simulation, grasping and carrying objects continuously from depth images alone.","explanation":"DextrAH-G is a dexterous-grasping method released by NVIDIA together with researchers at Stanford, the University of Utah, and Berkeley in July 2024, published at CoRL 2024. It controls a KUKA arm plus an Allegro dexterous hand — 23 motors in total — and the challenge is the high action dimensionality, the sim-to-real gap, and the need to avoid collisions and respect joint limits. The approach has three steps: train a teacher policy in Isaac Gym with reinforcement learning that can see privileged information such as the object's true pose; distill it into a student policy that sees only a depth image; then deploy zero-shot to the real robot. The key detail is that the policy's output doesn't go straight to the motors — it's handed to Geometric Fabrics, NVIDIA's reactive motion-generation method, which handles obstacle avoidance, joint limits, and posture, both protecting the hardware and making reinforcement-learning exploration easier.","example":"In a bin-picking test, DextrAH-G, using only a single RealSense depth camera mounted at the edge of the table, continuously grasped and carried objects with an 87% success rate over 256 attempts, at about 10.7 seconds per pick-and-place cycle.","related":["Geometric Fabrics","Dexterous Manipulation","Teacher-Student Distillation","Privileged Information","Sim-to-Real Transfer","Isaac Gym"]},{"id":"dexteritygen","category":"named_model","sec":9,"tier":3,"sources":[{"title":"DexterityGen: Foundation Controller for Unprecedented Dexterity (arXiv:2502.04307)","url":"https://arxiv.org/abs/2502.04307"},{"title":"DexGen 项目主页","url":"https://zhaohengyin.github.io/dexteritygen/"}],"as_of":"2025-02","related_ids":["dexterous-manipulation","in-hand-manipulation","reinforcement-learning","diffusion-model","teleoperation","meta-fundamental-ai-research"],"name":"DexterityGen","alt":"DexterityGen","abbr":"DexGen","aliases":["DexGen"],"one_liner":"Meta and Berkeley's low-level dexterous-hand controller that turns a human's rough teleoperation intent into precise finger motion.","explanation":"DexterityGen (DexGen for short) is a dexterous-hand control method released in February 2025 by teams at Meta FAIR and UC Berkeley. A dexterous hand has many degrees of freedom, and direct human teleoperation struggles to reliably perform fine motions like spinning a pen in-hand or turning a screw; pure reinforcement learning, on the other hand, has trouble learning long, complex tasks. Its division of labor: first, reinforcement learning in simulation trains a huge library of in-hand motion primitives — rotation, translation, and so on — collecting about 10^10 state transitions; that data then trains a diffusion model as a “foundation controller,” which generates fingertip keypoint motion from the current observation, converted into joint commands by an inverse-dynamics model. At deployment, a higher-level source (such as human teleoperation) only needs to give a rough intent, and DexGen refines it into stable, dexterous motion, while also being able to reject dangerous actions.","example":"On a Franka arm fitted with an Allegro dexterous hand, an operator wearing a Manus data glove teleoperates with DexGen's help to pick up and use a pen, a syringe, and a screwdriver; the paper reports the object stays gripped without dropping 10 to 100 times longer than baselines.","related":["Dexterous Manipulation","In-hand Manipulation","Reinforcement Learning","Diffusion Model","Teleoperation","Meta Fundamental AI Research"]},{"id":"google-deepmind-table-tennis-robot","category":"named_model","sec":9,"tier":3,"sources":[{"title":"Achieving Human Level Competitive Robot Table Tennis (arXiv 2408.03906)","url":"https://arxiv.org/abs/2408.03906"},{"title":"Competitive Robot Table Tennis 项目页","url":"https://sites.google.com/view/competitive-robot-table-tennis/home"}],"as_of":"2024-08","related_ids":["sim-to-real-transfer","hierarchical-architecture","reinforcement-learning","dynamic-manipulation","mujoco","google-deepmind"],"name":"Google DeepMind Table Tennis Robot","alt":"DeepMind 乒乓球机器人","abbr":"","aliases":["Achieving Human Level Competitive Robot Table Tennis","Robot Table Tennis"],"one_liner":"DeepMind's 2024 table-tennis robot, the first learning-based robot to reach amateur human level in real matches.","explanation":"This work was released by Google DeepMind in August 2024, and the paper claims it's the first learning-based robot to reach amateur human level in competitive table tennis. The hardware is a 6-degree-of-freedom ABB IRB 1100 arm mounted on two linear rails that let it move forward, back, left, and right, with two 125Hz cameras tracking the ball. Control has two layers: low-level specialized skills (such as forehand and backhand), each carrying a “skill descriptor” recording what it's good and bad at; and a high-level controller that picks a skill based on the incoming ball and statistics about the opponent, adapting to that opponent in real time during a match. The skills are trained with reinforcement learning in MuJoCo simulation and deployed zero-shot to the real robot; the distribution of incoming balls observed in real matches is then fed back into simulation, forming an iterative training curriculum. Playing 29 human opponents it had never faced before, the robot won 45% of matches overall: it won every match against beginners, 55% against intermediate players, and lost every match against advanced players.","example":"Against an opponent it has never played before, the high-level controller keeps tallying each low-level skill's return success rate during the match, and increasingly picks whichever skills are working against that particular opponent.","related":["Sim-to-Real Transfer","Hierarchical Architecture","Reinforcement Learning","Dynamic Manipulation","MuJoCo (Multi-Joint dynamics with Contact)","Google DeepMind"]},{"id":"deepmimic","category":"named_model","sec":10,"tier":2,"sources":[{"title":"DeepMimic (arXiv 1804.02717)","url":"https://arxiv.org/abs/1804.02717"}],"as_of":"2018-04","related_ids":["motion-tracking","reinforcement-learning","reference-state-initialization","early-termination","adversarial-motion-priors","motion-capture"],"name":"DeepMimic","alt":"DeepMimic","abbr":"","aliases":["DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills"],"one_liner":"Uses reinforcement learning to make a simulated character imitate motion-capture clips, learning physically realistic flips and martial arts.","explanation":"DeepMimic was published at SIGGRAPH 2018 by Xue Bin Peng and colleagues at UC Berkeley and the University of British Columbia. Simulated characters trained with reinforcement learning alone often move stiffly and unnaturally. DeepMimic adds an imitation reward: at every step in physics simulation, the character scores higher the closer its pose is to a reference motion clip (such as motion-capture data), and this can be combined with an additional task reward. Two of its training tricks later became standard practice: reference state initialization, which starts each episode from a random frame of the reference motion, and early termination, which ends an episode as soon as the character falls. It taught humanoid characters, the Atlas robot, and even a dinosaur to walk, flip, and perform martial-arts moves. Today's motion-tracking approaches for humanoid robots, such as ASAP and BeyondMimic, continue this same line of thinking.","example":"Given a motion-capture clip of a human backflip, a simulated humanoid character trained this way learns to perform the backflip in physics simulation and land on its feet.","related":["Motion Tracking","Reinforcement Learning","Reference State Initialization","Early Termination","Adversarial Motion Priors","Motion Capture"]},{"id":"ase","category":"named_model","sec":10,"tier":3,"sources":[{"title":"ASE: Large-Scale Reusable Adversarial Skill Embeddings for Physically Simulated Characters (arXiv 2205.01906)","url":"https://arxiv.org/abs/2205.01906"},{"title":"ASE 项目页（Xue Bin Peng）","url":"https://xbpeng.github.io/projects/ASE/index.html"}],"as_of":"2022-07","related_ids":["adversarial-motion-priors","generative-adversarial-imitation-learning","unsupervised-skill-discovery","latent-space","behavior-foundation-model","isaac-gym"],"name":"ASE","alt":"ASE（对抗技能嵌入）","abbr":"ASE","aliases":["Adversarial Skill Embeddings","ASE: Large-Scale Reusable Adversarial Skill Embeddings for Physically Simulated Characters"],"one_liner":"Berkeley and NVIDIA's 2022 method for learning a reusable latent space of skills from motion-capture clips.","explanation":"ASE was proposed by Xue Bin Peng, Sergey Levine, Sanja Fidler, and colleagues at UC Berkeley and NVIDIA, published at SIGGRAPH 2022. In simulated character animation, every new task usually means training a policy from scratch, relearning basic motions like walking and running again and again. ASE works in two layers. During pretraining, on a large batch of unlabeled, unsegmented motion clips, it trains a low-level policy conditioned on a latent variable z, using adversarial imitation learning (a discriminator judges whether a motion looks human, as in the data) to keep motions natural, and unsupervised skill discovery so that different values of z correspond to different skills. For a downstream task, the low-level policy is frozen and only a high-level policy that outputs z is trained, which needs only a simple reward. Using parallel simulation in Isaac Gym, the low-level policy was trained on more than 10 billion samples. It's a follow-up to AMP (Adversarial Motion Priors).","example":"A simulated humanoid character carrying a sword and shield first pretrains its skill latent space on about 30 minutes of motion data across 187 clips; afterward, given only a reward for “run to the target and knock it down,” the high-level policy can compose running, sword swings, and a strike into a coherent sequence.","related":["Adversarial Motion Priors","Generative Adversarial Imitation Learning","Unsupervised Skill Discovery","Latent Space","Behavior Foundation Model","Isaac Gym"]},{"id":"mdm","category":"named_model","sec":10,"tier":3,"sources":[{"title":"arXiv 2209.14916: Human Motion Diffusion Model","url":"https://arxiv.org/abs/2209.14916"},{"title":"MDM 项目主页","url":"https://guytevet.github.io/mdm-page/"}],"as_of":"2022-09","related_ids":["diffusion-model","human-motion-generation","humanml3d","classifier-free-guidance","motion-retargeting","motion-tracking"],"name":"MDM (Motion Diffusion Model)","alt":"MDM（人体动作扩散模型）","abbr":"MDM","aliases":["Human Motion Diffusion Model"],"one_liner":"A landmark method that generates a 3D human motion sequence from text or an action category using a diffusion model.","explanation":"MDM (Human Motion Diffusion Model) is work released in September 2022 by Guy Tevet, Amit Bermano, and colleagues at Tel Aviv University in Israel, published at ICLR 2023. It applies the diffusion models used in image generation to human motion: using a Transformer as the network, it progressively denoises a sequence of joint motion from noise, with classifier-free guidance letting generation be controlled by text or an action category. A key design choice is predicting the clean motion directly at every step rather than the noise, which makes it possible to add geometric losses on joint position, velocity, and foot contact, producing more natural motion. It achieved the best results at the time on text-to-motion benchmarks such as HumanML3D and KIT, and became the baseline for a large amount of follow-up motion-generation work. In embodied AI, the human motion this kind of model generates can, after motion retargeting, be handed to a humanoid robot's motion-tracking controller for execution.","example":"Given the prompt “a person walks forward a few steps and then sits down,” MDM generates a corresponding 3D human skeletal motion sequence.","related":["Diffusion Model","Human Motion Generation","HumanML3D","Classifier-Free Guidance","Motion Retargeting","Motion Tracking"]},{"id":"perpetual-humanoid-control","category":"named_model","sec":10,"tier":3,"sources":[{"title":"arXiv 2305.06456: Perpetual Humanoid Control for Real-time Simulated Avatars","url":"https://arxiv.org/abs/2305.06456"},{"title":"GitHub: ZhengyiLuo/PHC","url":"https://github.com/ZhengyiLuo/PHC"}],"as_of":"2026-09","related_ids":["motion-tracking","deepmimic","maskedmimic","amass","h2o","catastrophic-forgetting"],"name":"Perpetual Humanoid Control","alt":"PHC","abbr":"PHC","aliases":["PHC","Perpetual Humanoid Control for Real-time Simulated Avatars"],"one_liner":"A physics-based humanoid controller that tracks huge libraries of human motion in simulation and gets back up on its own after falling.","explanation":"PHC (Perpetual Humanoid Control) is a physics-based motion-imitation controller from Zhengyi Luo and colleagues at Carnegie Mellon University and Meta Reality Labs, presented at ICCV 2023 and trained with reinforcement learning in Isaac Gym. The goal is for a simulated humanoid to track noisy reference motions in real time — motions coming from video-based pose estimation or text-to-motion generation — without any external stabilizing forces. The key idea is Progressive Multiplicative Control Policy (PMCP): whenever the controller hits a motion it cannot learn, a new sub-network is added to handle it, which lets it learn nearly 10,000 motion clips from the AMASS motion-capture library without catastrophic forgetting (losing previously learned skills while acquiring new ones). The paper reports a 98.9% success rate on the training set and 96.4% on the held-out test set. PHC also learned to recover naturally after falling and resume tracking. Its open-source code was later extended to Unitree's H1 and G1 humanoid models, and follow-up work such as PHC+ and PULSE builds on it.","example":"A simulated character is driven in real time by poses estimated from a monocular video; if it gets tripped partway through, PHC lets it stand back up on its own and keep tracking the reference motion from where it left off.","related":["Motion Tracking","DeepMimic","MaskedMimic","AMASS (Archive of Motion Capture as Surface Shapes)","H2O","Catastrophic Forgetting"]},{"id":"maskedmimic","category":"named_model","sec":10,"tier":3,"sources":[{"title":"arXiv 2409.14393: MaskedMimic","url":"https://arxiv.org/abs/2409.14393"},{"title":"NVIDIA Research: MaskedMimic 项目页","url":"https://research.nvidia.com/labs/par/maskedmimic/"}],"as_of":"2024-09","related_ids":["motion-tracking","perpetual-humanoid-control","deepmimic","conditional-variational-autoencoder","dagger","protomotions"],"name":"MaskedMimic","alt":"MaskedMimic","abbr":"","aliases":["Unified Physics-Based Character Control Through Masked Motion Inpainting"],"one_liner":"An NVIDIA unified physics-based humanoid controller that treats every control mode as filling in missing motion.","explanation":"MaskedMimic was released in September 2024 by Chen Tessler, Xue Bin Peng, and colleagues at NVIDIA Research, published at SIGGRAPH Asia 2024 (ACM TOG). Previously, controlling a humanoid character in physics simulation — tracking a motion, walking, reaching for an object, acting out text — usually meant training a separate policy with its own reward for each. MaskedMimic unifies all of these as a motion-inpainting problem: given only partial constraints (target positions for a few joints, some keyframes, a line of text, or an object to interact with), with the rest masked out, the model generates complete, physically plausible whole-body motion. Training happens in two steps: first, reinforcement learning trains a teacher controller on AMASS motion-capture data that fully tracks a reference motion; then DAgger-style behavior cloning distills it into a conditional-VAE student policy that can handle randomly masked input. The code is included in NVIDIA's open-source framework, ProtoMotions.","example":"Given only the head and hand target positions corresponding to a VR headset and its two controllers, with every other joint masked out, MaskedMimic generates complete whole-body motion for standing, walking, and reaching.","related":["Motion Tracking","Perpetual Humanoid Control","DeepMimic","Conditional Variational Autoencoder","DAgger","ProtoMotions (NVIDIA GPU-accelerated humanoid simulation & learning framework)"]},{"id":"meta-motivo","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models (arXiv 2504.11054)","url":"https://arxiv.org/abs/2504.11054"},{"title":"Meta FAIR: Sharing new research, models, and datasets (2024-12-12)","url":"https://ai.meta.com/blog/meta-fair-updates-agents-robustness-safety-architecture/"},{"title":"facebookresearch/metamotivo (GitHub)","url":"https://github.com/facebookresearch/metamotivo"}],"as_of":"2025-04","related_ids":["behavior-foundation-model","unsupervised-skill-discovery","motion-tracking","zero-shot","bfm-zero","adversarial-motion-priors"],"name":"Meta Motivo","alt":"Meta Motivo（人形行为基础模型）","abbr":"","aliases":["FB-CPR","Forward-Backward Representations with Conditional-Policy Regularization","Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models"],"one_liner":"Meta's simulated-humanoid behavior foundation model that does motion tracking, reaching a target pose, and reward-driven behavior with no further training.","explanation":"Meta Motivo was released by Meta FAIR in December 2024, with the paper published at ICLR 2025; the company calls it the first behavior foundation model for humanoids. It controls a virtual humanoid character in physics simulation (the HumEnv environment), not a real robot. Its core algorithm, FB-CPR, builds on forward-backward representations, an unsupervised reinforcement-learning method that encodes state, reward, and policy into the same latent space; a discriminator is added on top, constraining the policy to stay close to unlabeled motion-capture data, so the learned motion looks human while still generalizing. Once pretrained, giving it a reference motion, a target pose, or a reward function directly yields the corresponding policy, with no additional training or planning needed; it is also somewhat robust to changes like gravity, wind, and external shoves. Code and several models, ranging from 24.5 million to 288 million parameters, are all open-source.","example":"Feed the same pretrained model a motion-capture reference clip, a target standing pose, or a “spin in place” reward function, and it can directly control the simulated humanoid to do each one, with no retraining needed per task.","related":["Behavior Foundation Model","Unsupervised Skill Discovery","Motion Tracking","Zero-shot","BFM-Zero","Adversarial Motion Priors"]},{"id":"real-world-humanoid-locomotion-with-reinforcement-learning","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Real-World Humanoid Locomotion with Reinforcement Learning (arXiv 2303.03381)","url":"https://arxiv.org/abs/2303.03381"},{"title":"Project page","url":"https://learning-humanoid-locomotion.github.io/"}],"as_of":"2024","related_ids":["rl-based-locomotion-control","sim-to-real-transfer","teacher-student-distillation","domain-randomization","agility-robotics-digit","humanoid-locomotion-as-next-token-prediction"],"name":"Real-World Humanoid Locomotion with Reinforcement Learning","alt":"真实世界人形强化学习行走","abbr":"","aliases":["Real-World Humanoid Locomotion with RL","Radosavovic et al., Science Robotics 2024"],"one_liner":"A Transformer walking controller trained purely with reinforcement learning in simulation, deployed zero-shot to make a humanoid walk outdoors.","explanation":"This work was released by Ilija Radosavovic, Jitendra Malik, and colleagues at UC Berkeley in March 2023 and published in Science Robotics in 2024. Humanoid walking had mostly relied on classical controllers, which struggle to adapt to new environments. This paper uses a causal Transformer as the controller: it takes in a window of past proprioceptive observations and actions and outputs the next action, adapting to terrain 'in context' from its history without updating any parameters. Training happens entirely in Isaac Gym simulation: a teacher policy is first trained with privileged information (state information unavailable on the real robot), then a student policy is trained by combining imitation of the teacher with reinforcement learning, together with domain randomization, before being deployed zero-shot to Agility Robotics' Digit humanoid. The resulting controller walks across a range of outdoor terrains and resists pushes, making this an early landmark for learning-based humanoid locomotion control.","example":"Digit walks on outdoor grass and rubberized running tracks it never saw during training, and stays stable even when pushed from the side by a person.","related":["RL-based Locomotion Control","Sim-to-Real Transfer","Teacher-Student Distillation","Domain Randomization","Agility Robotics Digit","Humanoid Locomotion as Next Token Prediction"]},{"id":"op3-soccer","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning (arXiv 2304.13653)","url":"https://arxiv.org/abs/2304.13653"},{"title":"OP3 Soccer 项目页","url":"https://sites.google.com/view/op3-soccer"}],"as_of":"2024-04","related_ids":["reinforcement-learning","sim-to-real-transfer","domain-randomization","self-play","policy-distillation","google-deepmind"],"name":"OP3 Soccer (DeepMind)","alt":"OP3 足球（DeepMind 双足踢球）","abbr":"","aliases":["Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning"],"one_liner":"DeepMind uses deep reinforcement learning to teach the small humanoid robot OP3 to play one-on-one soccer.","explanation":"This is work from Google DeepMind, posted to arXiv in April 2023 and published in Science Robotics in April 2024. The robot is Robotis's OP3, a low-cost small humanoid with 20 driven joints, and the task is simplified one-on-one soccer. Training happens entirely in MuJoCo simulation: skills like getting up and shooting are trained separately first, then distilled into a single policy, which keeps improving by playing against its own past versions (self-play). A relatively high control frequency, targeted dynamics randomization, and disturbances applied during training let the policy transfer to the real robot zero-shot. Compared to a scripted controller, it walks 181% faster, turns 302% faster, gets up 63% quicker, and kicks 34% faster. It's an early representative example of deep reinforcement learning producing agile, whole-body motion on a bipedal robot.","example":"During a match, the robot anticipates where the ball will go and turns sideways to block an opponent's shot; if knocked down, it gets back up quickly and keeps chasing the ball.","related":["Reinforcement Learning","Sim-to-Real Transfer","Domain Randomization","Self-Play","Policy Distillation","Google DeepMind"]},{"id":"humanoid-locomotion-as-next-token-prediction","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Humanoid Locomotion as Next Token Prediction (arXiv 2402.19469)","url":"https://arxiv.org/abs/2402.19469"},{"title":"Project page","url":"https://humanoid-next-token-prediction.github.io/"}],"as_of":"2024-02","related_ids":["next-token-prediction","autoregressive-decoding","action-free-video","bipedal-locomotion","agility-robotics-digit","real-world-humanoid-locomotion-with-reinforcement-learning"],"name":"Humanoid Locomotion as Next Token Prediction","alt":"人形行走即下一个 token 预测","abbr":"","aliases":[],"one_liner":"Treats a humanoid robot's walking control as a language-model-style “predict the next token” problem, learned with a causal Transformer.","explanation":"This was released in February 2024 by Ilija Radosavovic, Koushil Sreenath, Jitendra Malik, and colleagues at UC Berkeley. It frames real humanoid robot control as a problem similar to language modeling: a causal Transformer autoregressively predicts a sensorimotor trajectory made of observations and actions, where each token predicts the next token of the same modality. This lets data lacking action labels — such as action trajectories extracted from human video — also participate in training. The data comes from simulated trajectories generated by existing neural-network policies and model-based controllers, human motion-capture data, and YouTube videos of humans. The model was deployed on Agility Robotics' full-size humanoid Digit, walking zero-shot on the streets of San Francisco; using just 27 hours of walking data is also enough to transfer to the real robot, and it generalizes to a backward-walking command that was never in the training data.","example":"There is no backward-walking command in the training data, yet once deployed, the model can still make Digit walk backward on command.","related":["Next-Token Prediction","Autoregressive Decoding","Action-free Video","Bipedal Locomotion","Agility Robotics Digit","Real-World Humanoid Locomotion with Reinforcement Learning"]},{"id":"humanoid-parkour-learning","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Humanoid Parkour Learning (arXiv 2406.10759)","url":"https://arxiv.org/abs/2406.10759"},{"title":"Humanoid Parkour Learning project page","url":"https://humanoid4parkour.github.io/"}],"as_of":"2024-09","related_ids":["parkour","perceptive-locomotion","teacher-student-distillation","dagger","unitree-h1","robot-parkour-learning"],"name":"Humanoid Parkour Learning","alt":"人形跑酷学习","abbr":"","aliases":[],"one_liner":"An end-to-end parkour policy that lets a humanoid jump onto platforms, cross gaps, and clear hurdles using just its head depth camera.","explanation":"This was released in June 2024 by Ziwen Zhuang, Shenzhe Yao, and Hang Zhao at the Shanghai Qi Zhi Institute, ShanghaiTech University, and Tsinghua University, published at CoRL 2024, and is the humanoid counterpart of the same group's quadruped work, Robot Parkour Learning. It needs no human motion reference at all, training a vision-based, end-to-end whole-body control policy with reinforcement learning: first, flat-ground walking is trained on terrain with fractal noise (the noise naturally encourages the robot to lift its feet), then a “privileged” parkour policy that reads ground-truth terrain directly is trained across 10 obstacle types, and finally that's distilled with DAgger into a student policy that only sees images from a head-mounted depth camera. Deployed on a Unitree H1, it can jump onto a 0.42 m platform, cross a 0.8 m gap, run outdoors at 1.8 m/s, and autonomously choose which parkour move to use when given only a steering command.","example":"The operator only uses a joystick to steer; when the H1 sees a platform ahead, it decides on its own to jump up onto it.","related":["Parkour","Perceptive Locomotion","Teacher-Student Distillation","DAgger","Unitree H1","Robot Parkour Learning"]},{"id":"denoising-world-model-learning","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning (arXiv:2408.14472)","url":"https://arxiv.org/abs/2408.14472"}],"as_of":"2024-08","related_ids":["asymmetric-actor-critic","privileged-information","sim-to-real-transfer","domain-randomization","blind-locomotion","robotera"],"name":"Denoising World Model Learning","alt":"DWL（去噪世界模型学习）","abbr":"DWL","aliases":["DWL","Advancing Humanoid Locomotion: Mastering Challenging Terrains with DWL"],"one_liner":"An end-to-end reinforcement-learning walking framework using only proprioception, letting a humanoid cross snow, slopes, and stairs.","explanation":"DWL was proposed in 2024 by RobotEra together with Jianyu Chen's group at Tsinghua University and the Shanghai Qi Zhi Institute, and was a best-paper finalist at RSS 2024. When a humanoid moves from simulation to a real robot, it faces environmental disturbances, inaccurate dynamics modeling, sensor noise, and quantities like linear velocity or contact force that simply can't be measured directly. DWL treats all of this as noise added to the true state: during simulated training, noise is added to observations and unmeasurable quantities are masked out, a GRU encoder extracts a latent state from a window of history, and a decoder reconstructs the full state — including privileged information like friction coefficient, external force, and terrain height — a kind of “denoising” world model; the policy is trained on the latent state with PPO, while the critic sees the full state directly (asymmetric actor-critic). With no camera or lidar, the policy transfers zero-shot to the 1.2-meter XBot-S and the 1.65-meter XBot-L, walking over snow, slopes, stairs, and rough ground with the same set of parameters.","example":"In indoor tests, DWL reached 100% success climbing and descending 10-centimeter-high steps, on slopes, and on uneven ground; PPO with the denoising loss removed managed only 20% success climbing stairs.","related":["Asymmetric Actor-Critic","Privileged Information","Sim-to-Real Transfer","Domain Randomization","Blind Locomotion","RobotEra"]},{"id":"hugwbc","category":"named_model","sec":10,"tier":3,"sources":[{"title":"HugWBC (arXiv 2502.03206)","url":"https://arxiv.org/abs/2502.03206"},{"title":"HugWBC project page","url":"https://hugwbc.github.io/"},{"title":"apexrl/HugWBC (GitHub)","url":"https://github.com/apexrl/HugWBC"}],"as_of":"2025-04","related_ids":["whole-body-control","learning-based-whole-body-control","gait","loco-manipulation","unitree-h1","hover"],"name":"HugWBC","alt":"HugWBC","abbr":"","aliases":["A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion"],"one_liner":"A whole-body controller where one policy lets a humanoid walk, run, jump, and hop, with adjustable step frequency and foot-lift height.","explanation":"HugWBC was released in February 2025 by Shanghai Jiao Tong University (Weinan Zhang's group) together with Shanghai AI Lab (Jiangmiao Pang) and others, published at RSS 2025. Most humanoid walking controllers only know one fixed gait, with no adjustable parameters and little room to extend. HugWBC designs a unified command space: besides velocity, it can also specify gait (walk/run, jumping, standing, hopping on one foot), step frequency, foot-lift height, body height, waist rotation, and body pitch. Training adds a symmetry loss, and uses “intervention training” to let the upper body be taken over by an external signal, such as teleoperation, which supports walking and manipulating at the same time. The policy is trained in Isaac Gym and deployed on a Unitree H1, with code open-sourced.","example":"A remote control makes the H1 walk at a specified step frequency and foot-lift height, while the operator simultaneously takes over its arms through teleoperation to carry a box.","related":["Whole-Body Control","Learning-Based Whole-Body Control","Gait","Loco-manipulation","Unitree H1","HOVER"]},{"id":"host","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Learning Humanoid Standing-up Control across Diverse Postures (arXiv 2502.08378)","url":"https://arxiv.org/abs/2502.08378"},{"title":"HoST project page","url":"https://taohuang13.github.io/humanoid-standingup.github.io/"}],"as_of":"2025-04","related_ids":["fall-recovery","rl-based-locomotion-control","curriculum-learning","sim-to-real-transfer","unitree-g1","shanghai-artificial-intelligence-laboratory"],"name":"HoST (Humanoid Standing-up)","alt":"HoST（人形起身）","abbr":"HoST","aliases":["Learning Humanoid Standing-up Control across Diverse Postures"],"one_liner":"A control framework that uses reinforcement learning to learn, from scratch, how a humanoid can stand up from many different postures.","explanation":"HoST was released in February 2025 by Shanghai AI Lab (Jiangmiao Pang's group) together with Shanghai Jiao Tong University, the University of Hong Kong, Zhejiang University, and CUHK, published at RSS 2025 and nominated for best systems paper. It addresses how a humanoid robot gets itself back up after falling: earlier methods either ran only in simulation and ignored real motor limits, or relied on standing-up trajectories hand-designed for a specific ground surface. HoST learns from scratch with reinforcement learning across a variety of simulated terrains, using no reference motion at all; it uses multiple value networks (critics), each evaluating a different group of rewards, paired with curriculum learning, plus smoothness regularization and an implicit velocity cap to prevent jitter and overly forceful motion on the real robot. Once trained, it was deployed directly to a Unitree G1, able to stand up from postures like lying down or leaning against a wall in a variety of indoor and outdoor settings, and is a representative work in fall-recovery research.","example":"Whether a Unitree G1 is lying on the ground or leaning against a wall, the HoST policy can control it to stand up on its own, with no human support or overhead rig needed.","related":["Fall Recovery","RL-based Locomotion Control","Curriculum Learning","Sim-to-Real Transfer","Unitree G1","Shanghai Artificial Intelligence Laboratory"]},{"id":"exbody","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Expressive Whole-Body Control for Humanoid Robots (arXiv 2402.16796)","url":"https://arxiv.org/abs/2402.16796"},{"title":"ExBody 项目页","url":"https://expressive-humanoid.github.io/"}],"as_of":"2024-07","related_ids":["exbody2","motion-tracking","whole-body-control","motion-retargeting","unitree-h1","massively-parallel-reinforcement-learning"],"name":"ExBody","alt":"ExBody","abbr":"ExBody","aliases":["Expressive Whole-Body Control for Humanoid Robots"],"one_liner":"A 2024 UC San Diego method where a humanoid's upper body imitates human motion while the legs just track velocity steadily.","explanation":"ExBody was released in February 2024 by Xiaolong Wang's group at UC San Diego, published at RSS 2024, and is one of the earlier works to apply large-scale human motion-capture data to whole-body control on a real humanoid robot. The difficulty is that humans and robots differ greatly in degrees of freedom and strength, so having the whole robot imitate a person frame by frame tends to make it fall over. ExBody's solution is to split the work: the upper body imitates the reference motion's joint angles and keypoints frame by frame, while the legs aren't required to match frame by frame at all, only to robustly follow the overall velocity and heading given by the reference motion. The data comes from about 780 clips in the CMU motion-capture database, retargeted onto the 19-degree-of-freedom Unitree H1, trained with massively parallel reinforcement learning in Isaac Gym, and then transferred to the real robot. Later work such as ExBody2 often uses it as a comparison baseline.","example":"On the real robot, the H1 can walk while performing motions in different styles at the same time — such as a zombie walk, high-fiving and shaking hands with a person, or even dancing together with someone.","related":["ExBody2","Motion Tracking","Whole-Body Control","Motion Retargeting","Unitree H1","Massively Parallel Reinforcement Learning"]},{"id":"exbody2","category":"named_model","sec":10,"tier":3,"sources":[{"title":"ExBody2: Advanced Expressive Humanoid Whole-Body Control (arXiv 2412.13196)","url":"https://arxiv.org/abs/2412.13196"},{"title":"ExBody2 项目页","url":"https://exbody2.github.io/"}],"as_of":"2025-03","related_ids":["exbody","motion-tracking","whole-body-control","teacher-student-distillation","privileged-information","omnih2o"],"name":"ExBody2","alt":"ExBody2","abbr":"ExBody2","aliases":["Advanced Expressive Humanoid Whole-Body Control"],"one_liner":"A humanoid whole-body motion-tracking method from UC San Diego and others that lets the Unitree G1 walk, squat, and dance.","explanation":"ExBody2 was released in December 2024 by Xiaolong Wang's group at UC San Diego together with UC Berkeley and MIT, an upgrade over the original ExBody. ExBody only had the upper body imitate a reference motion; ExBody2 switches to full-body tracking, mainly through three changes: first, automatic data filtering — a base policy is run over the CMU motion-capture data, and motions the robot can't achieve are filtered out by tracking error, balancing feasibility against diversity; second, teacher-student distillation, where a teacher policy uses privileged information (state unavailable on a real robot) in simulation, and is then distilled into a student policy that uses only onboard observations; and third, tracking keypoints in a local coordinate frame, decoupled from overall velocity tracking. The authors also found that training a general policy first and then fine-tuning it on a specific category of motion gives higher accuracy.","example":"On the Unitree G1, ExBody2 can dance the cha-cha following a reference motion; a specialized policy fine-tuned on dance data has noticeably lower joint-tracking error than either the general policy or one trained from scratch.","related":["ExBody","Motion Tracking","Whole-Body Control","Teacher-Student Distillation","Privileged Information","OmniH2O"]},{"id":"h2o","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation (arXiv 2403.04436)","url":"https://arxiv.org/abs/2403.04436"},{"title":"H2O 项目页","url":"https://human2humanoid.com/"}],"as_of":"2024-10","related_ids":["whole-body-teleoperation","motion-retargeting","motion-tracking","omnih2o","amass","unitree-h1"],"name":"H2O","alt":"H2O（人到人形）","abbr":"H2O","aliases":["Human to Humanoid","Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation"],"one_liner":"A 2024 CMU humanoid teleoperation framework that uses a single RGB camera to make a robot mimic a person's whole-body motion in real time.","explanation":"H2O was released in March 2024 by teams under Guanya Shi, Changliu Liu, and Kris Kitani at Carnegie Mellon University, published at IROS 2024 as an oral presentation, with Tairan He and Zhengyi Luo as co-first authors. It uses nothing but an ordinary RGB camera: the pose estimator HybrIK computes human body pose from the footage in real time, and a whole-body tracking policy trained with reinforcement learning drives a Unitree H1 to follow along. The key step is “sim-to-data”: about 10,000 motion clips from the AMASS motion-capture dataset are first retargeted onto the H1, then tried out in simulation by an imitation policy with access to privileged information, filtering out motions the robot can't achieve, leaving about 8,500 clips for training; the trained policy is deployed to the real robot with no extra tuning. The authors call this the first learning-based real-time whole-body humanoid teleoperation. Follow-up work, OmniH2O, extended it to VR teleoperation and autonomous learning from teleoperation data.","example":"The operator stands in front of the camera and walks, kicks, turns, waves, and throws punches, and the H1 follows along with the same motions in real time; a backflip was also demonstrated.","related":["Whole-Body Teleoperation","Motion Retargeting","Motion Tracking","OmniH2O","AMASS (Archive of Motion Capture as Surface Shapes)","Unitree H1"]},{"id":"omnih2o","category":"named_model","sec":10,"tier":2,"sources":[{"title":"OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning (arXiv 2406.08858)","url":"https://arxiv.org/abs/2406.08858"},{"title":"OmniH2O 项目主页","url":"https://omni.human2humanoid.com/"}],"as_of":"2024-11","related_ids":["h2o","whole-body-teleoperation","teacher-student-distillation","privileged-information","motion-retargeting","humanplus"],"name":"OmniH2O","alt":"OmniH2O","abbr":"","aliases":["OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning","Omni Human-to-Humanoid"],"one_liner":"CMU's 2024 humanoid whole-body teleoperation system: VR, cameras, voice, or GPT-4o can all drive the same robot.","explanation":"OmniH2O is a humanoid whole-body teleoperation and learning system from Tairan He, Zhengyi Luo, Guanya Shi, and colleagues at Carnegie Mellon University and Shanghai Jiao Tong University, proposed in June 2024 and published at CoRL 2024 as an upgrade to the same group's H2O. Its core idea is to use “kinematic pose” — the position and orientation of key body parts — as a unified control interface: whether a command comes from a VR headset, an RGB camera, voice, or GPT-4o, it's first converted into a target pose, which the same whole-body control policy then tracks. That policy is trained in simulation with reinforcement learning: after large-scale retargeting and augmentation of AMASS human motion data, a teacher policy is trained using privileged information (full state only available in simulation), then distilled into a student policy that uses only sparse sensor inputs and can run on the real robot. The platform is a Unitree H1 fitted with dexterous hands. The team also released OmniH2O-6, the first whole-body-control dataset for humanoids covering 6 everyday tasks, and used it to train autonomous skills with Diffusion Policy.","example":"Autonomous shadow-boxing: the robot's head-camera feed is sent to GPT-4o with a prompt saying to throw a left jab at a blue target, a right jab at a red target, and stay still if there's no target; GPT-4o answers only A, B, or C each time, and OmniH2O's whole-body policy carries out the chosen action.","related":["H2O","Whole-Body Teleoperation","Teacher-Student Distillation","Privileged Information","Motion Retargeting","HumanPlus"]},{"id":"humanplus","category":"named_model","sec":10,"tier":2,"sources":[{"title":"HumanPlus: Humanoid Shadowing and Imitation from Humans (arXiv 2406.10454)","url":"https://arxiv.org/abs/2406.10454"},{"title":"HumanPlus 项目主页","url":"https://humanoid-ai.github.io/"}],"as_of":"2024-11","related_ids":["humanoid-robot","whole-body-teleoperation","unitree-h1","action-chunking-with-transformers","amass","omnih2o"],"name":"HumanPlus","alt":"HumanPlus","abbr":"","aliases":["HumanPlus: Humanoid Shadowing and Imitation from Humans"],"one_liner":"Stanford's 2024 humanoid system: one RGB camera lets a robot shadow a person in real time, then learn skills from it.","explanation":"HumanPlus is a humanoid data-collection and learning system proposed in June 2024 by Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, and Chelsea Finn at Stanford, published at CoRL 2024 and a finalist for best paper. Its hardware is a Unitree H1 fitted with two 6-DOF Inspire dexterous hands, for 33 degrees of freedom total. It works in two steps. First, a low-level controller called the Humanoid Shadowing Transformer is trained in simulation with reinforcement learning on about 40 hours of human motion data from AMASS; once deployed, a single RGB camera is enough to estimate the operator's body and hand pose, letting the robot “shadow” them in real time. This teleoperation method is then used to collect demonstrations, training a Humanoid Imitation Transformer, built on ACT, that acts autonomously using first-person video from a stereo RGB camera on the robot's head. The work shows that whole-body teleoperation doesn't necessarily require expensive motion-capture equipment.","example":"Standing up and walking after putting on shoes: using at most 40 shadowing-collected demonstrations per task, the trained robot completed it autonomously with 60% success; unloading a warehouse shelf reached 90%, and typing reached 80%.","related":["Humanoid Robot","Whole-Body Teleoperation","Unitree H1","Action Chunking with Transformers","AMASS (Archive of Motion Capture as Surface Shapes)","OmniH2O"]},{"id":"okami","category":"named_model","sec":10,"tier":3,"sources":[{"title":"OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation (arXiv 2410.11792)","url":"https://arxiv.org/abs/2410.11792"},{"title":"OKAMI 项目页","url":"https://ut-austin-rpl.github.io/OKAMI/"}],"as_of":"2024-11","related_ids":["imitation-from-observation","motion-retargeting","human-video-data","one-shot-imitation-learning","humanoid-robot","fourier-gr-1"],"name":"OKAMI","alt":"OKAMI","abbr":"OKAMI","aliases":["Teaching Humanoid Robots Manipulation Skills through Single Video Imitation"],"one_liner":"A 2024 UT Austin and NVIDIA method that teaches a humanoid robot to manipulate objects from watching a single human video.","explanation":"OKAMI was proposed by Yuke Zhu's group at UT Austin together with NVIDIA Research, posted to arXiv in October 2024, an oral presentation at CoRL 2024. The goal is to have a humanoid robot learn a manipulation task from watching just a single RGB-D video of a human demonstration, with no teleoperation data collection needed. The first step analyzes the video: GPT-4V identifies task-relevant objects, Grounded-SAM segments and Cutie tracks them, and the person's body and hand motion is reconstructed (using the SMPL-H model) to get a reference plan. The second step is object-aware motion retargeting: it first locates where the objects are in the current scene, then adjusts the human arm trajectory to that object position before mapping it onto the robot, with finger motion transferred along with it. The experimental platform is a Fourier GR1 humanoid fitted with two 6-degree-of-freedom Inspire dexterous hands. Successfully executed trajectories can also serve as data for training a closed-loop visuomotor policy.","example":"Across 6 tasks — bagging groceries, sprinkling salt, putting a snack on a plate, closing a laptop, and others — OKAMI's average success rate is 71.7%, 58.3 percentage points above the baseline ORION; a visuomotor policy trained on the trajectories it produces reaches an average success rate of 79.2%.","related":["Imitation from Observation","Motion Retargeting","Human Video Data","One-shot Imitation Learning","Humanoid Robot","Fourier GR-1"]},{"id":"hover","category":"named_model","sec":10,"tier":3,"sources":[{"title":"HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots (arXiv 2410.21229)","url":"https://arxiv.org/abs/2410.21229"},{"title":"HOVER project page","url":"https://hover-versatile-humanoid.github.io/"},{"title":"NVlabs/HOVER (GitHub)","url":"https://github.com/NVlabs/HOVER"}],"as_of":"2025-03","related_ids":["whole-body-control","learning-based-whole-body-control","policy-distillation","motion-tracking","omnih2o","nvidia-generalist-embodied-agent-research-lab"],"name":"HOVER","alt":"HOVER","abbr":"HOVER","aliases":["Versatile Neural Whole-Body Controller for Humanoid Robots"],"one_liner":"A versatile neural whole-body controller from NVIDIA and others where one policy works across many different command modes.","explanation":"HOVER was released in October 2024 by NVIDIA (Linxi Fan, Yuke Zhu, and others) together with CMU, UC Berkeley, UT Austin, and UCSD, published at ICRA 2025. Navigation, loco-manipulation, and tabletop manipulation each need a different control interface for a humanoid: navigation cares about torso (root) velocity, while tabletop manipulation cares about upper-body joint positions. Previously, each interface was usually trained as its own separate, non-interchangeable policy. HOVER first trains a teacher whole-body motion-tracking policy that imitates AMASS human motion data in simulation, then distills it into a student policy; during training, mode masks and sparse masks are applied separately to upper- and lower-body commands, letting the same policy switch among the command modes used by H2O, OmniH2O, ExBody, and HumanPlus with no retraining needed. The code is open-sourced on Isaac Lab, and it was deployed on a real Unitree H1.","example":"The same HOVER policy can either take just torso-velocity commands to make the H1 walk, or switch to a VR teleoperation mode that tracks only the head and hand positions.","related":["Whole-Body Control","Learning-Based Whole-Body Control","Policy Distillation","Motion Tracking","OmniH2O","NVIDIA Generalist Embodied Agent Research Lab"]},{"id":"uh-1","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Learning from Massive Human Videos for Universal Humanoid Pose Control (arXiv 2412.14172)","url":"https://arxiv.org/abs/2412.14172"},{"title":"UH-1 project page (PSI Lab)","url":"https://psi-lab.ai/UH-1/"}],"as_of":"2024-12","related_ids":["humanoid-x","human-video-data","motion-retargeting","text-to-motion","action-tokenizer","humanoid-robot"],"name":"UH-1","alt":"UH-1 / Humanoid-X","abbr":"UH-1","aliases":["Universal Humanoid Pose Control"],"one_liner":"A large model that learns from massive internet human videos and generates humanoid robot motion from text instructions.","explanation":"UH-1 was released in December 2024 by the University of Southern California (Yue Wang's group), UC Berkeley, and Toyota Research Institute, later given an oral presentation at Humanoids 2025. Humanoid robot data mostly comes from reinforcement learning and teleoperation, which is hard to scale up; this work instead learns from internet human videos. The team first built the Humanoid-X dataset: 3D human poses are extracted from videos and automatically captioned, then retargeted into humanoid keypoints and actions, yielding about 164,000 motion clips and more than 20 million robot poses. UH-1 is then trained on this: humanoid motions are first discretized into tokens, and a Transformer autoregressively generates motion tokens conditioned on a text instruction. The output can either be keypoints, which are handed to a goal-conditioned policy for tracking, or robot actions directly, executed open-loop. UH-1 demonstrates a route for scaling up humanoid data by turning human video plus text into humanoid motion.","example":"Given the text 'wave hello,' UH-1 generates a corresponding sequence of humanoid motion, which is handed to a lower-level controller to drive the humanoid robot's wave.","related":["Humanoid-X","Human Video Data","Motion Retargeting","Text-to-Motion","Action Tokenizer","Humanoid Robot"]},{"id":"asap","category":"named_model","sec":10,"tier":2,"sources":[{"title":"ASAP (arXiv 2502.01143)","url":"https://arxiv.org/abs/2502.01143"},{"title":"ASAP 项目主页","url":"https://agile.human2humanoid.com/"}],"as_of":"2025-04","related_ids":["sim-to-real-transfer","sim-to-real-gap","residual-policy","motion-tracking","unitree-g1","domain-randomization"],"name":"ASAP","alt":"ASAP","abbr":"ASAP","aliases":["ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills"],"one_liner":"Uses real-robot data to train a correction model that closes the sim-to-real gap for agile humanoid moves.","explanation":"ASAP is a whole-body humanoid control method released by Carnegie Mellon University and NVIDIA in February 2025, published at RSS 2025. Motions trained in simulation often look distorted on the real robot, because simulated physics doesn't match real motors and contact dynamics — the sim-to-real gap. ASAP works in two stages. First, it trains a motion-tracking policy in simulation on retargeted human motion data. Then it deploys that policy on the real robot to collect trajectories, and trains a delta (residual) action model that learns how much extra action the simulator needs to add to match real-robot behavior; this residual is plugged back into the simulator to fine-tune the original policy. On the Unitree G1, it outperforms common approaches like system identification and domain randomization, and the code is open-sourced.","example":"With ASAP, a Unitree G1 reproduced signature celebration moves from athletes like Cristiano Ronaldo, Kobe Bryant, and LeBron James, as well as forward and side jumps over a meter high.","related":["Sim-to-Real Transfer","Sim-to-Real Gap (Reality Gap)","Residual Policy","Motion Tracking","Unitree G1","Domain Randomization"]},{"id":"homie","category":"named_model","sec":10,"tier":3,"sources":[{"title":"HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit (arXiv 2502.13013)","url":"https://arxiv.org/abs/2502.13013"},{"title":"HOMIE 项目页","url":"https://homietele.github.io/"}],"as_of":"2025-06","related_ids":["exoskeleton-teleoperation","loco-manipulation","whole-body-teleoperation","data-glove","unitree-g1","shanghai-artificial-intelligence-laboratory"],"name":"HOMIE","alt":"HOMIE","abbr":"HOMIE","aliases":["OpenHomie","Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit"],"one_liner":"A 2025 Shanghai AI Lab humanoid teleoperation system that controls the whole body with an exoskeleton, gloves, and foot pedals.","explanation":"HOMIE was released in February 2025 by Shanghai AI Lab together with the Chinese University of Hong Kong (Jiangmiao Pang and Dahua Lin's groups), published at RSS 2025, and fully open-sourced as OpenHomie. A humanoid needs to walk and work with its hands at the same time (loco-manipulation), which makes it hard to control both halves of the body at once during teleoperation. HOMIE is a semi-autonomous solution: the lower body runs a reinforcement-learning policy that walks, turns, and crouches to a commanded height on instruction, while adapting to whatever pose the upper body is in (trained with an upper-body pose curriculum, a height-tracking reward, and left-right symmetry); the operator gives movement commands with foot pedals, wears an exoskeleton isomorphic to the robot's arms (joints correspond one to one, so the read-out angles serve directly as commands, with no inverse kinematics needed), and controls the dexterous hands with Hall-sensor motion-sensing gloves. The whole hardware setup costs about $500; policies were trained in Isaac Gym for both the Unitree G1 and the Fourier GR-1, with real-robot experiments done on the G1.","example":"For example, the operator presses pedals to walk the robot to a table and crouch down, while using the exoskeleton arms and gloves to control both hands to pick something up; the authors report task completion time about half that of previous systems.","related":["Exoskeleton Teleoperation","Loco-manipulation","Whole-Body Teleoperation","Data Glove","Unitree G1","Shanghai Artificial Intelligence Laboratory"]},{"id":"twist-2","category":"named_model","sec":10,"tier":2,"sources":[{"title":"TWIST: Teleoperated Whole-Body Imitation System (arXiv 2505.02833)","url":"https://arxiv.org/abs/2505.02833"},{"title":"TWIST 项目页","url":"https://yanjieze.com/projects/TWIST/"},{"title":"TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System (arXiv 2511.02832)","url":"https://arxiv.org/abs/2511.02832"}],"as_of":"2025-11","related_ids":["whole-body-teleoperation","motion-retargeting","motion-tracking","teacher-student-distillation","twist2","unitree-g1"],"name":"TWIST","alt":"TWIST","abbr":"TWIST","aliases":["TWIST: Teleoperated Whole-Body Imitation System"],"one_liner":"Stanford's 2025 whole-body teleoperation system for humanoids: the robot mirrors the operator's entire body in real time.","explanation":"TWIST was released in May 2025 by Jiajun Wu and C. Karen Liu's teams at Stanford, working with Xue Bin Peng at Simon Fraser University, and published at CoRL 2025 with Yanjie Ze as first author. Earlier humanoid teleoperation usually controlled the upper and lower body separately, which couldn't produce coordinated whole-body moves like squatting to lift a box or kicking a ball. TWIST captures a person's full-body motion with OptiTrack optical motion capture, retargets it in real time (converting human motion into robot joint angles) onto a 29-degree-of-freedom Unitree G1, and tracks it with a single unified neural-network controller. That controller is trained in Isaac Gym simulation on about 42 hours of human motion-capture data: first a teacher policy that can see 2 seconds of future reference motion is trained, then it's distilled — using reinforcement learning plus behavior cloning — into a student policy that sees only the current frame, cutting latency. A later version, TWIST2, switches to a PICO VR headset and no longer needs motion capture, making it easier to collect data at scale.","example":"An operator on the motion-capture stage bends down to pick up a box from the floor, and the G1 mirrors the same whole-body motion in sync; when the operator kicks a ball or dances a waltz step, the robot follows along in real time.","related":["Whole-Body Teleoperation","Motion Retargeting","Motion Tracking","Teacher-Student Distillation","TWIST2","Unitree G1"]},{"id":"twist2","category":"named_model","sec":10,"tier":3,"sources":[{"title":"TWIST2 (arXiv 2511.02832)","url":"https://arxiv.org/abs/2511.02832"},{"title":"TWIST2 project page","url":"https://yanjieze.com/projects/TWIST2/"}],"as_of":"2025-11","related_ids":["twist","whole-body-teleoperation","vr-teleoperation","unitree-g1","hierarchical-architecture","motion-tracking"],"name":"TWIST2","alt":"TWIST2","abbr":"","aliases":["TWIST 2","TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System"],"one_liner":"A portable teleoperation system that collects whole-body humanoid data with a VR headset, with no motion-capture studio needed.","explanation":"TWIST2 was released in November 2025 by a joint team from Stanford, Amazon FAR, USC, UC Berkeley, and CMU (authors including Yanjie Ze, Pieter Abbeel, Guanya Shi, Jiajun Wu, and C. Karen Liu), an upgrade to the earlier full-body teleoperation system TWIST. Whole-body humanoid data is hard to collect, and its predecessor relied on optical motion capture tied to a fixed studio. TWIST2 instead captures the operator's full-body motion with a PICO 4 Ultra headset plus two PICO body trackers, and adds a roughly $250, 2-degree-of-freedom active neck to the Unitree G1 so the robot has its own first-person view; the whole setup can be carried into any environment. The paper reports that 15 minutes of use can collect about 100 demonstrations at close to 100% success. On top of this, a hierarchical policy is trained: a low-level whole-body motion tracker learned through simulated reinforcement learning, and a high-level visuomotor policy trained with imitation learning on the collected data. The system, hardware design, and dataset are all open-sourced.","example":"An operator wears the PICO headset and straps on the trackers to perform pick-and-place motions, and the Unitree G1 reproduces the whole-body motion in real time, recording over a hundred bimanual pick-and-place demonstrations in just 15 minutes.","related":["Twist","Whole-Body Teleoperation","VR Teleoperation","Unitree G1","Hierarchical Architecture","Motion Tracking"]},{"id":"videomimic","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Visual Imitation Enables Contextual Humanoid Control (arXiv 2505.03729)","url":"https://arxiv.org/abs/2505.03729"},{"title":"VideoMimic project page","url":"https://www.videomimic.net/"}],"as_of":"2025-09","related_ids":["real-to-sim-to-real","4d-reconstruction","motion-retargeting","motion-tracking","human-video-data","unitree-g1"],"name":"VideoMimic","alt":"VideoMimic","abbr":"","aliases":["Visual Imitation Enables Contextual Humanoid Control"],"one_liner":"A humanoid robot method that learns skills like climbing stairs and sitting down in a chair from ordinary phone-shot human videos.","explanation":"VideoMimic was released by Angjoo Kanazawa, Jitendra Malik, Pieter Abbeel, and colleagues at UC Berkeley in May 2025, winning the CoRL 2025 Best Student Paper Award. Teaching a humanoid robot 'environment-aware' actions, such as climbing stairs or sitting down in a chair, requires knowing both how a human moves and what the surrounding terrain looks like. VideoMimic is a real-to-sim-to-real pipeline: from a monocular video, it simultaneously reconstructs a metrically accurate 4D human trajectory and the scene's geometry; the human motion is retargeted onto the robot, and the scene is converted into a simulator mesh; a policy that tracks these motions is then trained in simulation with reinforcement learning, and distilled into a single policy that relies only on proprioception, a height map of the terrain around the body, and a goal direction. The final policy is deployed on a 23-degree-of-freedom Unitree G1.","example":"A phone video of a person walking up some steps and then sitting down on a bench is fed through VideoMimic, and the Unitree G1 can then autonomously climb stairs and sit down and stand up in a real environment, with the same policy working for different staircases and chairs.","related":["Real-to-Sim-to-Real","4D Reconstruction","Motion Retargeting","Motion Tracking","Human Video Data","Unitree G1"]},{"id":"amo","category":"named_model","sec":10,"tier":3,"sources":[{"title":"AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control (arXiv 2505.03738)","url":"https://arxiv.org/abs/2505.03738"},{"title":"AMO 项目页（RSS 2025）","url":"https://amo-humanoid.github.io/"}],"as_of":"2025-05","related_ids":["whole-body-control","learning-based-whole-body-control","trajectory-optimization","teacher-student-distillation","unitree-g1","whole-body-teleoperation"],"name":"AMO","alt":"AMO","abbr":"AMO","aliases":["Adaptive Motion Optimization","AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control"],"one_liner":"UCSD's 2025 humanoid whole-body control method, using trajectory optimization to help an RL policy bend and reach far.","explanation":"AMO was proposed by Xiaolong Wang's group at UC San Diego, published at RSS 2025. Picking something off the floor or reaching a high shelf requires a humanoid to coordinate bending at the waist, twisting the torso, and flexing the legs — but reinforcement-learning policies trained by imitating human motion-capture data rarely see this kind of extreme torso posture, and become unstable on out-of-distribution commands. AMO first uses trajectory optimization (solving for joint trajectories under dynamics constraints) to batch-generate data on “how the legs should move for a given torso orientation and height,” training a small MLP module that supplies a reference pose to the lower-body policy in real time; that lower-body policy is trained in Isaac Gym with teacher-student distillation. The system is deployed on a 29-degree-of-freedom Unitree G1, can be teleoperated with VR, and teleoperated data can also train a Transformer policy to act autonomously.","example":"In the paper's demos, a G1 moves cans between surfaces at different heights, taking a bottle from a tall shelf on the left and setting it on a low table on the right, and can also straighten both legs to place a bottle on a high shelf; during teleoperation, 3 poses from the VR device are converted into control commands.","related":["Whole-Body Control","Learning-Based Whole-Body Control","Trajectory Optimization","Teacher-Student Distillation","Unitree G1","Whole-Body Teleoperation"]},{"id":"falcon","category":"named_model","sec":10,"tier":3,"sources":[{"title":"FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation (arXiv 2505.06776)","url":"https://arxiv.org/abs/2505.06776"},{"title":"FALCON 项目主页（CMU LeCAR Lab）","url":"https://lecar-lab.github.io/falcon-humanoid/"}],"as_of":"2025-11","related_ids":["loco-manipulation","whole-body-control","curriculum-learning","unitree-g1","booster-robotics-t1","field-ai"],"name":"FALCON","alt":"FALCON","abbr":"","aliases":["Learning Force-Adaptive Humanoid Loco-Manipulation"],"one_liner":"A 2025 CMU humanoid training framework that lets a robot push, pull, and carry heavy loads steadily while walking.","explanation":"FALCON is a whole-body control method for humanoid robots released in May 2025 by CMU's LeCAR Lab (Guanya Shi's group) together with Field AI and others, accepted as an oral presentation at L4DC 2026. When a humanoid walks and works at the same time (loco-manipulation), a large force on its hands — from pulling a cart, opening a door, or carrying something heavy — can easily throw the lower body off balance, while the upper body also struggles to reach its target accurately. FALCON uses dual-agent reinforcement learning: a lower-body policy keeps walking stable under force disturbances, and an upper-body policy sends the hand to its target while implicitly compensating for the external force, both trained jointly and sharing proprioception (joint angles, velocities, and other self-state); this is paired with a 3D force curriculum that gradually increases the force applied to the hands during training, while staying under each joint's torque limit. Upper-body tracking accuracy is about twice that of baselines, and the same pipeline, with no reward changes, works on both the Unitree G1 and the Booster Robotics T1.","example":"In real-robot experiments, the humanoid walked while pulling a cart under 0–100 N of force, opened a door with both arms against 0–40 N of resistance, and could carry a load while walking, squatting, and turning.","related":["Loco-manipulation","Whole-Body Control","Curriculum Learning","Unitree G1","Booster Robotics T1","Field AI"]},{"id":"clone","category":"named_model","sec":10,"tier":3,"sources":[{"title":"CLONE (arXiv 2506.08931)","url":"https://arxiv.org/abs/2506.08931"},{"title":"CLONE 论文 HTML 版 (arXiv 2506.08931v2)","url":"https://arxiv.org/html/2506.08931v2"}],"as_of":"2025-08","related_ids":["whole-body-teleoperation","mixture-of-experts","unitree-g1","apple-vision-pro","motion-tracking","omnih2o"],"name":"CLONE","alt":"CLONE","abbr":"","aliases":["CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks"],"one_liner":"A humanoid whole-body teleoperation system that tracks only the head and hands, correcting drift in closed loop.","explanation":"CLONE was proposed in June 2025 by teams at the Beijing Institute for General Artificial Intelligence (BIGAI), Peking University, and Beijing Institute of Technology. Earlier humanoid teleoperation often controlled the upper and lower body separately to stay stable, producing uncoordinated motion; robots also drift further from the operator's actual position the longer they walk (accumulated drift). CLONE captures only the operator's head and hand poses with an Apple Vision Pro headset, and a mixture-of-experts (MoE) whole-body policy generates whole-body joint actions for a Unitree G1, with the robot's actual global position fed back to the policy in real time for closed-loop drift correction. The paper reports an average position error of about 5.1 centimeters over 8.9 meters of straight-line walking, and it can perform long-horizon tasks requiring whole-body coordination, such as squatting to pick something up off the floor. Systems like this are also a basic tool for collecting demonstration data on humanoid robots.","example":"An operator wearing an Apple Vision Pro walks around a room and bends down; the G1 walks over and squats to pick an object up off the floor, following along.","related":["Whole-Body Teleoperation","Mixture of Experts","Unitree G1","Apple Vision Pro","Motion Tracking","OmniH2O"]},{"id":"kungfubot","category":"named_model","sec":10,"tier":3,"sources":[{"title":"KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills (arXiv 2506.12851)","url":"https://arxiv.org/abs/2506.12851"},{"title":"KungfuBot 项目页","url":"https://kungfubot.github.io"},{"title":"KungfuBot2: Learning Versatile Motion Skills for Humanoid Whole-Body Control (arXiv 2509.16638)","url":"https://arxiv.org/abs/2509.16638"}],"as_of":"2025-10","related_ids":["motion-tracking","motion-retargeting","unitree-g1","smpl","asymmetric-actor-critic","learning-based-whole-body-control"],"name":"KungfuBot","alt":"KungfuBot","abbr":"PBHC","aliases":["PBHC","KungfuBot2","VMS"],"one_liner":"A humanoid whole-body motion-imitation framework that teaches a Unitree G1 highly dynamic moves like kung fu and dance.","explanation":"KungfuBot was released in June 2025 by China Telecom's AI Research Institute (TeleAI) together with Shanghai Jiao Tong University, East China University of Science and Technology, Harbin Institute of Technology, and ShanghaiTech University, accepted to NeurIPS 2025, with the method itself named PBHC and its code open-sourced. Earlier humanoid motion-tracking work could mostly only imitate slow, gentle motions. PBHC first processes the data: it extracts SMPL human poses from video, uses physical metrics to filter out motions the robot can't achieve and correct foot contact, then retargets the result onto a Unitree G1; it then trains a tracking policy with reinforcement learning, turning the tolerance for tracking error into a bi-level optimization that adapts automatically to the error, effectively an automatic curriculum. A September 2025 follow-up, KungfuBot2, uses an orthogonal mixture of experts to let a single policy master multiple skills, and can stably track motions lasting up to several minutes.","example":"The G1 performs a jumping kick, a spinning kick, a horse-stance punch, and tai chi movements, all copied from human video; the project page showcases 15 such highly dynamic skills in total.","related":["Motion Tracking","Motion Retargeting","Unitree G1","SMPL","Asymmetric Actor-Critic","Learning-Based Whole-Body Control"]},{"id":"leverb","category":"named_model","sec":10,"tier":3,"sources":[{"title":"LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction (arXiv 2506.13751)","url":"https://arxiv.org/abs/2506.13751"}],"as_of":"2025-09","related_ids":["learning-based-whole-body-control","dual-system-architecture","latent-action","conditional-variational-autoencoder","unitree-g1","dagger"],"name":"LeVERB","alt":"LeVERB","abbr":"","aliases":["Latent Vision-Language-Encoded Robot Behavior","LeVERB-Bench"],"one_liner":"A dual-system framework that links a vision-language model to a humanoid whole-body controller through a “latent verb.”","explanation":"LeVERB was released in June 2025, led by UC Berkeley, with collaborators including Xue Bin Peng, Trevor Darrell, and Koushil Sreenath. Existing VLA models mostly assume the low-level controller only accepts manually defined commands like end-effector pose or base velocity, limiting them to quasi-static tasks. LeVERB splits into two layers: the high-level LeVERB-VL (System 2, 10Hz) encodes first- and third-person images and the instruction with SigLIP, learning a “latent verb” space through a conditional variational autoencoder; the low-level LeVERB-A (System 1, 50Hz) is a whole-body controller, first trained with PPO as a teacher that tracks reference motions, then distilled with DAgger into a student conditioned on the latent verb. The authors used IsaacSim to render motion-capture playback and built LeVERB-Bench, covering more than 150 tasks. Deployed zero-shot on a Unitree G1, it reaches 80% success on simple visual navigation and 58.5% overall, 7.8 times that of a naive hierarchical baseline.","example":"Tell the robot “walk over to the red chair and sit down,” and the high-level model looks at the camera feed and outputs a latent verb, which the low-level controller uses to make the G1 walk over, turn, and sit.","related":["Learning-Based Whole-Body Control","Dual-System Architecture (System 1 / System 2)","Latent Action","Conditional Variational Autoencoder","Unitree G1","DAgger"]},{"id":"gmt","category":"named_model","sec":10,"tier":3,"sources":[{"title":"GMT (arXiv:2506.14770)","url":"https://arxiv.org/abs/2506.14770"},{"title":"GMT 项目主页","url":"https://gmt-humanoid.github.io/"}],"as_of":"2025-09","related_ids":["motion-tracking","mixture-of-experts","unitree-g1","teacher-student-distillation","amass","beyondmimic"],"name":"GMT","alt":"GMT","abbr":"GMT","aliases":["General Motion Tracking"],"one_liner":"A whole-body control method that uses one unified policy to make a humanoid track many different human motions.","explanation":"GMT was proposed in June 2025 by Xiaolong Wang's group at UC San Diego together with Xue Bin Peng at Simon Fraser University. Motion tracking means having a robot imitate a reference motion — such as human motion-capture data — in real time. Earlier methods often needed one policy per motion category, or per-category fine-tuning; GMT instead uses a single policy that covers walking, kicking, soccer kicks, dancing, and more. Two ideas are key: adaptive sampling, which automatically practices harder clips more during training, and a motion mixture-of-experts (MoE), which lets different parts of the network specialize in different types of motion. Training first uses PPO to produce a teacher policy with access to privileged information, then distills it with DAgger into a student policy that sees only proprioception. The data comes from about 33 hours of motion drawn from AMASS and LAFAN1, deployed on a real 23-degree-of-freedom Unitree G1.","example":"The same GMT policy on a Unitree G1 can perform martial-arts kicks and soccer kicks, and also imitate a drunken walk, a crouching walk, and dancing, all without training separately for each motion category.","related":["Motion Tracking","Mixture of Experts","Unitree G1","Teacher-Student Distillation","AMASS (Archive of Motion Capture as Surface Shapes)","BeyondMimic"]},{"id":"unitracker","category":"named_model","sec":10,"tier":3,"sources":[{"title":"UniTracker (arXiv 2507.07356)","url":"https://arxiv.org/abs/2507.07356"},{"title":"UniTracker 论文 HTML 版（含机构与实验设置）","url":"https://arxiv.org/html/2507.07356v3"}],"as_of":"2025-09","related_ids":["motion-tracking","conditional-variational-autoencoder","teacher-student-distillation","privileged-information","amass","unitree-g1"],"name":"UniTracker","alt":"UniTracker","abbr":"","aliases":["UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots"],"one_liner":"A three-stage whole-body motion tracking framework that lets a single policy make a humanoid robot track a wide range of human motions.","explanation":"UniTracker was released in July 2025 by Shanghai Jiao Tong University, the Shanghai Artificial Intelligence Laboratory, and other institutions. Motion tracking means having a robot imitate a reference human motion in real time; the difficulty is that one policy has to cover thousands of different motions, and a real robot cannot access the complete state information available in simulation. UniTracker has three stages: first, a teacher policy is trained in simulation using privileged information (full state information unavailable on the real robot); then it is distilled into a real-robot-ready student policy, which learns a global latent variable for each motion via a conditional variational autoencoder (CVAE) to reduce drift in global quantities like orientation when only partial observations are available; finally, a fast-adaptation module fine-tunes individually or in batches on motions that are hard to track. The training data is 8,179 human motion clips filtered from AMASS, validated both in simulation and on a real Unitree G1.","example":"Given a dance motion clip taken from AMASS, the same UniTracker policy can make a real Unitree G1 perform it; for a particularly hard motion sequence the robot cannot track directly, the third-stage fast-adaptation module fine-tunes on it individually or in a batch.","related":["Motion Tracking","Conditional Variational Autoencoder","Teacher-Student Distillation","Privileged Information","AMASS (Archive of Motion Capture as Surface Shapes)","Unitree G1"]},{"id":"beyondmimic","category":"named_model","sec":10,"tier":2,"sources":[{"title":"BeyondMimic (arXiv 2508.08241)","url":"https://arxiv.org/abs/2508.08241"},{"title":"BeyondMimic 项目主页","url":"https://beyondmimic.github.io/"}],"as_of":"2025-11","related_ids":["motion-tracking","diffusion-model","deepmimic","asap","unitree-g1","lafan1"],"name":"BeyondMimic","alt":"BeyondMimic","abbr":"","aliases":["BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion"],"one_liner":"Trains a humanoid to reproduce human motion, then distills that into a guided diffusion model for new tasks.","explanation":"BeyondMimic is a whole-body humanoid control framework released in August 2025 by Koushil Sreenath's group at UC Berkeley and C. Karen Liu's group at Stanford, built in two stages. The first stage is motion tracking: a simple, unified reinforcement-learning recipe learns in simulation to reproduce motions from the LAFAN1 motion-capture dataset, then deploys zero-shot to a real Unitree G1, performing highly dynamic moves like aerial cartwheels, spinning kicks, and sprinting. The second stage distills a chosen set of tracking policies into a single latent-space diffusion model; at inference, sampling is guided with a simple cost function, letting it handle new tasks — waypoint navigation, joystick teleoperation, obstacle avoidance — with no retraining. It pushes “imitating motion” forward into “using motion as a prior for general-purpose control,” and the motion-tracking code is open-sourced.","example":"The same diffusion model can switch from “follow the joystick” to “walk around obstacles to a target point” just by swapping in a different cost function.","related":["Motion Tracking","Diffusion Model","DeepMimic","ASAP","Unitree G1","LAFAN1"]},{"id":"any2track","category":"named_model","sec":10,"tier":3,"sources":[{"title":"Track Any Motions under Any Disturbances (arXiv 2509.13833)","url":"https://arxiv.org/abs/2509.13833"},{"title":"Any2Track 项目页","url":"https://zzk273.github.io/Any2Track/"}],"as_of":"2025-09","related_ids":["motion-tracking","rapid-motor-adaptation","sim-to-real-transfer","unitree-g1","amass","gmt"],"name":"Any2Track","alt":"Any2Track","abbr":"","aliases":["Track Any Motions under Any Disturbances","AnyTracker","AnyAdapter"],"one_liner":"A humanoid motion tracker from Tsinghua, Peking University, and Galbot that keeps tracking a motion even when pushed, pulled, or loaded down.","explanation":"Any2Track was proposed jointly by Tsinghua University, Peking University, Galbot, and the Shanghai Qi Zhi Institute (corresponding author Li Yi), released in September 2025. Motion tracking means having a humanoid reproduce a reference motion in real time, a core capability underlying whole-body control and teleoperation; most existing trackers had only been validated on flat ground with no external forces. Any2Track uses two-stage reinforcement learning. First, a general tracker called AnyTracker is trained on LAFAN1 and AMASS motion-capture data, learning highly dynamic, multi-contact motions of many kinds with a single policy. It's then frozen, and an adaptation module called AnyAdapter is added, extracting dynamics features from a recent window of state-action history (trained with an auxiliary task of predicting future states, like a world model) to compensate online for changes in terrain, external force, and load — without degrading the original tracking ability. The policy is trained in simulation and deployed zero-shot to a real Unitree G1.","example":"In real-robot tests, a loaded-down G1 tracks a reference motion over uneven ground while someone pulls it with a rope and pushes it with a foot, and the robot stays stable and keeps completing the motion.","related":["Motion Tracking","Rapid Motor Adaptation","Sim-to-Real Transfer","Unitree G1","AMASS (Archive of Motion Capture as Surface Shapes)","GMT"]},{"id":"bfm-zero","category":"named_model","sec":10,"tier":3,"sources":[{"title":"BFM-Zero (arXiv 2511.04131)","url":"https://arxiv.org/abs/2511.04131"},{"title":"BFM-Zero 项目主页","url":"https://lecar-lab.github.io/BFM-Zero/"}],"as_of":"2025-11","related_ids":["behavior-foundation-model","unsupervised-skill-discovery","motion-tracking","unitree-g1","meta-motivo","learning-based-whole-body-control"],"name":"BFM-Zero","alt":"BFM-Zero","abbr":"","aliases":["BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised RL"],"one_liner":"A humanoid whole-body control foundation model trained with no task reward at all, switching tasks through a simple “prompt.”","explanation":"BFM-Zero is a humanoid behavioral foundation model released in November 2025 by Guanya Shi's group at Carnegie Mellon University together with Meta and other institutions. It's trained with unsupervised reinforcement learning: no specific task reward is given during training; instead, a forward-backward representation (a method that maps states and tasks into the same latent space) learns a shared latent space, into which reference motions, target poses, and reward functions can all be encoded as vectors. At deployment, giving the corresponding vector — the “prompt” — lets the same policy perform motion tracking, reaching a target pose, or optimizing a given reward, all zero-shot, and it also supports few-shot adaptation. The team demonstrated push recovery and getting back up after a fall on a real Unitree G1, calling it the first behavioral foundation model that can switch tasks on a real humanoid through prompts.","example":"The same BFM-Zero policy deployed on a Unitree G1 does motion tracking when given a reference motion, poses itself to match a given target pose, or acts to optimize a given reward function — all without retraining in between.","related":["Behavior Foundation Model","Unsupervised Skill Discovery","Motion Tracking","Unitree G1","Meta Motivo","Learning-Based Whole-Body Control"]},{"id":"sonic","category":"named_model","sec":10,"tier":3,"sources":[{"title":"SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control (arXiv 2511.07820)","url":"https://arxiv.org/abs/2511.07820"},{"title":"GEAR-SONIC project page","url":"https://nvlabs.github.io/GEAR-SONIC/"}],"as_of":"2026-09","related_ids":["motion-tracking","learning-based-whole-body-control","unitree-g1","nvidia-isaac-gr00t-n1","nvidia-generalist-embodied-agent-research-lab","gr00t-wholebodycontrol"],"name":"SONIC","alt":"SONIC","abbr":"SONIC","aliases":["GEAR-SONIC","SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control"],"one_liner":"NVIDIA's general-purpose humanoid whole-body control foundation model, built by scaling up motion tracking training.","explanation":"SONIC was released by NVIDIA's GEAR lab (Zhengyi Luo, Linxi Fan, Yuke Zhu, and others) in November 2025, later published in Science Robotics (2026). The premise is that humanoid controllers can also get better simply by scaling up, and motion tracking — having a robot reproduce a piece of human motion in real time — is well suited to that, since motion-capture data already comes with dense supervision and needs no hand-designed reward. The team scaled the network from 1.2 million to 42 million parameters, trained on about 700 hours of motion-capture data (over 100 million frames) using roughly 21,000 GPU-hours, and deployed it on the Unitree G1. Through a unified token interface, the same policy can take motion commands from a joystick plus a real-time motion planner, VR teleoperation, video imitation, or text- and music-driven motion generation, and it can also sit behind a VLA such as GR00T N1.5 to perform whole-body mobile manipulation. The code is open-sourced in the GR00T-WholeBodyControl repository.","example":"An operator wearing VR gear tracks only their head and hands; SONIC uses a motion planner to fill in the lower body, controlling the G1 to complete a task like pushing a lawnmower while walking.","related":["Motion Tracking","Learning-Based Whole-Body Control","Unitree G1","NVIDIA Isaac GR00T N1","NVIDIA Generalist Embodied Agent Research Lab","GR00T-WholeBodyControl"]},{"id":"clip-on-wheels","category":"named_model","sec":11,"tier":3,"sources":[{"title":"CoWs on Pasture (arXiv:2203.10421)","url":"https://arxiv.org/abs/2203.10421"},{"title":"CoWs on Pasture 项目主页（CVPR 2023）","url":"https://cow.cs.columbia.edu/"}],"as_of":"2023-06","related_ids":["object-goal-navigation","clip","open-vocabulary","zero-shot","frontier-based-exploration","vlfm"],"name":"CLIP on Wheels","alt":"CoW（CLIP on Wheels）","abbr":"CoW","aliases":["CoW","CoWs on Pasture"],"one_liner":"A baseline that bolts an open-vocabulary model like CLIP onto a mobile robot to find objects from text, with no navigation training.","explanation":"CoW was proposed by Shuran Song's group at Columbia University with the University of Washington in 2022, published at CVPR 2023. It studies “language-driven zero-shot object navigation”: a robot must find a target described in a sentence (such as “the toy airplane under the bed”) in an unfamiliar house, with no navigation training on those specific objects or scenes. CoW's approach is deliberately simple: while it's not yet confident it recognizes the target, it wanders using a classical exploration strategy; once an open-vocabulary model like CLIP (which can recognize objects from arbitrary text) locates the target in view with enough confidence, it plans a path there. The authors evaluated 21 CoW variants and proposed the Pasture benchmark, testing rare objects, objects described by appearance or spatial relation, and occluded objects. The best CoW beat the previous best method by 15.6 percentage points on a RoboTHOR object subset.","example":"For a goal like “tie-dye surfboard” — a category absent from navigation datasets — CoW first explores the room, and once CLIP judges some view a good enough match for that phrase, it plans a path to that location as the target.","related":["Object-Goal Navigation","CLIP","Open-vocabulary","Zero-shot","Frontier-based Exploration","VLFM"]},{"id":"lm-nav","category":"named_model","sec":11,"tier":3,"sources":[{"title":"arXiv 2207.04429: LM-Nav","url":"https://arxiv.org/abs/2207.04429"},{"title":"LM-Nav 项目主页","url":"https://sites.google.com/view/lmnav"}],"as_of":"2022-07","related_ids":["vision-and-language-navigation","clip","topological-map","llm-based-task-planning","gnm","vint"],"name":"LM-Nav","alt":"LM-Nav","abbr":"LM-Nav","aliases":["Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action"],"one_liner":"Combines GPT-3, CLIP, and a visual navigation model so a robot can navigate by following natural-language instructions.","explanation":"LM-Nav was released in July 2022 by Dhruv Shah, Brian Ichter, Sergey Levine, and colleagues, published at CoRL 2022. Directing a robot's navigation with language usually needs a lot of trajectory data with text descriptions, which is expensive to annotate. LM-Nav does no fine-tuning at all and uses no language-labeled robot data; instead, it combines three off-the-shelf pretrained models: the large language model GPT-3 breaks the instruction down into a sequence of landmarks; the image-text model CLIP judges which landmark corresponds to what the robot's camera sees; and the visual navigation model ViNG builds a topological map of the environment from previously collected images and runs a go-to-point policy. The system then searches for the shortest route that passes through those landmarks in order. It completed long-distance navigation in real outdoor environments, and is an early representative example of assembling a robot system out of foundation models.","example":"For example, the user says “go past the stop sign and head to the white building”; GPT-3 extracts the two landmarks “stop sign” and “white building,” CLIP locates the corresponding positions in the topological map, and ViNG drives to them in sequence.","related":["Vision-and-Language Navigation","CLIP","Topological Map","LLM-based Task Planning","GNM","ViNT"]},{"id":"vlmaps","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Visual Language Maps for Robot Navigation (arXiv 2210.05714)","url":"https://arxiv.org/abs/2210.05714"},{"title":"VLMaps 项目主页","url":"https://vlmaps.github.io/"},{"title":"vlmaps/vlmaps (GitHub)","url":"https://github.com/vlmaps/vlmaps"}],"as_of":"2023-03","related_ids":["semantic-map","vision-and-language-navigation","open-vocabulary","code-as-policies","clip","vlfm"],"name":"VLMaps","alt":"VLMaps","abbr":"VLMaps","aliases":["Visual Language Maps","Visual Language Maps for Robot Navigation"],"one_liner":"A system that writes vision-language features into a 3D map, letting a robot navigate using phrases like 'three meters to the right of the chair.'","explanation":"VLMaps was proposed by Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard at the University of Freiburg, Google Research, and the Nuremberg Institute of Technology, posted to arXiv in October 2022 and published at ICRA 2023. As it moves through an environment, the robot builds a map from RGB-D video: an open-vocabulary segmentation model, LSeg, computes a vision-language feature for every pixel, which is projected onto 3D surfaces using depth and pose and then compressed into a top-down grid map. At query time, words like 'sofa' or 'fridge' are encoded with a text encoder and compared against the map's features by similarity, letting the system locate any object — an open-vocabulary semantic map. Complex instructions are first handed to a large language model (the GPT-3 family), which writes code that calls the map's interface, so the system can handle spatial relations like 'between two objects' or 'three meters to the right.' The same map can also generate a separate obstacle map for each different robot. VLMaps is an early representative of connecting foundation models to navigation maps.","example":"Told to 'move to the spot three meters to the right of the chair,' the language model turns the instruction into code that first looks up the chair's position on the map, then computes the point three meters to its right for the robot to head toward.","related":["Semantic Map","Vision-and-Language Navigation","Open-vocabulary","Code as Policies","CLIP","VLFM"]},{"id":"gnm","category":"named_model","sec":11,"tier":3,"sources":[{"title":"GNM: A General Navigation Model to Drive Any Robot (arXiv:2210.03370)","url":"https://arxiv.org/abs/2210.03370"},{"title":"GNM 项目主页","url":"https://sites.google.com/view/drive-any-robot"}],"as_of":"2023-05","related_ids":["vint","nomad","cross-embodiment","image-goal-navigation","topological-map","berkeley-artificial-intelligence-research"],"name":"GNM","alt":"GNM（通用导航模型）","abbr":"GNM","aliases":["General Navigation Model"],"one_liner":"A general visual-navigation model trained on mixed data from 6 robot types that can drive different robots.","explanation":"GNM was proposed in October 2022 by Sergey Levine's group at UC Berkeley (Dhruv Shah and others), published at ICRA 2023. It pools about 60 hours of navigation data from 6 robot types — TurtleBot2, Jackal, Spot, an RC car, an all-terrain vehicle, and more — to train a single image-goal navigation policy: given the current image, a few past frames (used to infer embodiment context, i.e., “which kind of robot am I”), and a goal image, it outputs the time-distance to the goal and the next 5 normalized waypoints. A normalized action space is what lets it work across different robots. The finding is that a single policy trained on this heterogeneous data outperforms any policy trained on a single dataset, and it can even deploy directly on robots absent from the training set, such as a quadrotor drone. Follow-up work includes ViNT and NoMaD.","example":"At deployment, a topological map made of images along a route is built first; GNM estimates how far the current view is from each node, and Dijkstra's algorithm then plans a sequence of sub-goals for the robot to navigate to one by one.","related":["ViNT","NoMaD","Cross-Embodiment","Image-Goal Navigation","Topological Map","Berkeley Artificial Intelligence Research"]},{"id":"vint","category":"named_model","sec":11,"tier":3,"sources":[{"title":"ViNT (arXiv 2306.14846)","url":"https://arxiv.org/abs/2306.14846"},{"title":"ViNT project page","url":"https://general-navigation-models.github.io/vint/index.html"}],"as_of":"2023-10","related_ids":["gnm","nomad","image-goal-navigation","navigation","foundation-model","prompt-tuning-soft-prompt"],"name":"ViNT","alt":"ViNT","abbr":"ViNT","aliases":["Visual Navigation Transformer","ViNT: A Foundation Model for Visual Navigation"],"one_liner":"A visual navigation foundation model from Berkeley, trained on navigation data from many kinds of robots, that finds a goal from an image.","explanation":"ViNT was released by Dhruv Shah, Sergey Levine, and colleagues at UC Berkeley in June 2023, an oral-presentation paper at CoRL 2023. Earlier navigation models were mostly trained on data from a single robot in a single kind of environment. ViNT encodes images with EfficientNet followed by a Transformer, taking in the current frame plus several recent past frames along with a goal image, and predicting how far away the goal is and what to do next; its training data comes from multiple robot platforms, totaling hundreds of hours of navigation data. Paired with a diffusion model that generates candidate subgoal images, it can explore unfamiliar environments, and combined with long-range heuristics like GPS it can perform kilometer-scale navigation; it can also be adapted through prompt tuning to take GPS waypoints or turn-by-turn instructions as the goal instead. ViNT builds on the earlier GNM, and was itself followed by NoMaD.","example":"A robot first drives along a route taking a series of photos; later, given just one of those photos as a goal image, ViNT can navigate back to the spot where that photo was taken in the same environment, and the same model can be deployed on different mobile robots.","related":["GNM","NoMaD","Image-Goal Navigation","Navigation","Foundation Model","Prompt Tuning / Soft Prompt"]},{"id":"nomad","category":"named_model","sec":11,"tier":3,"sources":[{"title":"NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration (arXiv 2310.07896)","url":"https://arxiv.org/abs/2310.07896"},{"title":"NoMaD 项目页","url":"https://general-navigation-models.github.io/nomad/"}],"as_of":"2024-05","related_ids":["diffusion-policy","vint","gnm","image-goal-navigation","active-exploration","action-multimodality"],"name":"NoMaD","alt":"NoMaD","abbr":"NoMaD","aliases":["Goal Masked Diffusion Policies for Navigation and Exploration"],"one_liner":"A 2023 Berkeley navigation diffusion policy where one model can both explore freely and head toward a goal image.","explanation":"NoMaD was proposed in October 2023 by Sergey Levine's group at UC Berkeley (Ajay Sridhar, Dhruv Shah, and others), published at ICRA 2024, where the project page says it won that year's best paper award. Robot navigation commonly needs two things: goal-free exploration in an unfamiliar environment, and heading toward a goal once a photo of it is seen. Previously this took two separate models; NoMaD merges them with “goal masking”: half of training samples have the goal image masked out, and at inference, masking it produces exploration while revealing it produces goal-reaching. It uses a ViNT-style Transformer to encode recent frames and the goal image, then a diffusion model to generate a future sequence of waypoints, able to express multiple viable options at a fork in the path (action multimodality). The model has about 19 million parameters, and was trained on more than 100 hours of real, multi-robot data from GNM, SACSoN, and others.","example":"Exploring real environments on a LoCoBot, NoMaD reaches a 98% success rate with an average of 0.2 collisions; a compared baseline that first generates a sub-goal image with diffusion and then navigates gets 77% and 1.7 collisions, while also having about 15 times more parameters.","related":["Diffusion Policy","ViNT","GNM","Image-Goal Navigation","Active Exploration","Action Multimodality"]},{"id":"vlfm","category":"named_model","sec":11,"tier":3,"sources":[{"title":"VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation (arXiv 2312.03275)","url":"https://arxiv.org/abs/2312.03275"},{"title":"VLFM 项目页","url":"https://naoki.io/portfolio/vlfm"}],"as_of":"2024-05","related_ids":["object-goal-navigation","frontier-based-exploration","zero-shot","success-weighted-by-path-length","occupancy-grid-map","vlmaps"],"name":"VLFM","alt":"VLFM（视觉语言前沿地图）","abbr":"VLFM","aliases":["Vision-Language Frontier Maps","VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation"],"one_liner":"A method that scores exploration frontiers with a vision-language model to find a specified object in an unfamiliar environment, zero-shot.","explanation":"VLFM was proposed by Naoki Yokoyama, Dhruv Batra, Bernadette Bucher, and colleagues at Georgia Tech and the Boston Dynamics AI Institute, posted to arXiv in December 2023 and published at ICRA 2024, where it won best paper in cognitive robotics. The task is object-goal navigation: finding an object of a given category in an environment the robot has never visited. It builds an occupancy map from depth data and identifies the frontier between known and unknown regions; at the same time, it uses BLIP-2 to compute the similarity between the current image and text prompts like 'the target is likely to be nearby,' recording this in a language value map, and then chooses the highest-value frontier to explore. Once the target is spotted, YOLOv7 or Grounding DINO detects it and Mobile-SAM segments it out for the robot to approach. The whole pipeline needs no navigation-specific training and achieved state-of-the-art results at the time by SPL (success weighted by path length) on Gibson, HM3D, and MP3D, and was deployed directly on a Boston Dynamics Spot.","example":"In an office building with no pre-built map, once Spot is given the category of object to find, it preferentially heads in the direction the vision-language model judges more likely to contain that object, and stops next to it once found.","related":["Object-Goal Navigation","Frontier-based Exploration","Zero-shot","Success weighted by Path Length","Occupancy Grid Map","VLMaps"]},{"id":"navid","category":"named_model","sec":11,"tier":3,"sources":[{"title":"NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation (arXiv 2402.15852)","url":"https://arxiv.org/abs/2402.15852"},{"title":"NaVid 项目页","url":"https://pku-epic.github.io/NaVid/"},{"title":"Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks (arXiv 2412.06224)","url":"https://arxiv.org/abs/2412.06224"}],"as_of":"2025-02","related_ids":["vision-and-language-navigation","room-to-room","vision-language-model","navila","navfom","sim-to-real-transfer"],"name":"NaVid","alt":"NaVid","abbr":"","aliases":["Uni-NaVid","Video-based VLM Plans the Next Step for Vision-and-Language Navigation"],"one_liner":"A 2024 video-LLM navigation method from Peking University and others that decides the next step using only monocular video.","explanation":"NaVid was proposed in February 2024 by He Wang's group at Peking University together with the Beijing Academy of Artificial Intelligence, the University of Adelaide, Galbot, and others, published at RSS 2024, addressing vision-and-language navigation (reaching a destination by following a one-sentence instruction). Earlier methods mostly relied on maps, odometry, or depth maps; NaVid uses only monocular RGB video: the current frame is compressed to 64 tokens and each history frame to 4, fed into a video large language model built on Vicuna-7B, which outputs the next step directly in text, such as how far to move forward, how much to turn, or to stop. Using only RGB, it reaches a 37.4% success rate on R2R-CE, and transfers to a real robot. The follow-up Uni-NaVid (RSS 2025) merges four task types — instruction navigation, object finding, embodied question answering, and person following — into a single model, trained on 3.6 million samples.","example":"Tested across 4 real indoor scenes with 200 instructions total, NaVid running on a Turtlebot4 achieved about 66% success on simple instructions and about 48% on multi-step compound instructions.","related":["Vision-and-Language Navigation","Room-to-Room","Vision-Language Model","NaVILA","NavFoM (Galbot)","Sim-to-Real Transfer"]},{"id":"poliformer","category":"named_model","sec":11,"tier":3,"sources":[{"title":"arXiv 2406.20083: PoliFormer","url":"https://arxiv.org/abs/2406.20083"},{"title":"GitHub: allenai/poliformer","url":"https://github.com/allenai/poliformer"}],"as_of":"2024-11","related_ids":["object-goal-navigation","on-policy","procthor","ai2-thor","allen-institute-for-ai","sim-to-real-transfer"],"name":"PoliFormer","alt":"PoliFormer","abbr":"","aliases":["PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators"],"one_liner":"A Transformer navigation policy trained purely with large-scale on-policy reinforcement learning in simulation, then deployed straight to real robots.","explanation":"PoliFormer was released by the Allen Institute for AI (Ai2) in June 2024 and published at CoRL 2024. It uses RGB images only: a Vision Transformer encoder (DINOv2 in the code) encodes each frame, followed by a causal Transformer decoder that aggregates a fairly long history and outputs navigation actions. Training happens entirely in simulation: on-policy reinforcement learning across a huge number of procedurally generated houses from ProcTHOR, run in parallel across many machines for hundreds of millions of interactions. It reaches 85.5% success on the CHORES-S object-goal navigation benchmark, a 28.5-percentage-point absolute improvement over the previous best method, and deploys without any additional tuning to two different real robots, the LoCoBot and the Stretch RE-1; it also transfers directly to downstream tasks like object tracking and open-vocabulary navigation. PoliFormer shows that reinforcement learning combined with Transformers and large-scale simulation also benefits from scale.","example":"Told to 'find the apple in the kitchen,' a Stretch robot running PoliFormer decides whether to move forward or turn based only on its head-camera feed, step by step, and stops next to the apple once it spots it.","related":["Object-Goal Navigation","On-Policy","ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","AI2-THOR","Allen Institute for AI","Sim-to-Real Transfer"]},{"id":"mobility-vla","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Mobility VLA (arXiv 2407.07775)","url":"https://arxiv.org/abs/2407.07775"},{"title":"Mobility VLA (arXiv HTML full text)","url":"https://arxiv.org/html/2407.07775"}],"as_of":"2024-07","related_ids":["vision-and-language-navigation","topological-map","hierarchical-architecture","google-gemini","context-length","structure-from-motion"],"name":"Mobility VLA","alt":"Mobility VLA","abbr":"","aliases":["MINT (Multimodal Instruction Navigation with demonstration Tours)","Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs"],"one_liner":"A hierarchical navigation system that has a long-context VLM watch a tour video, then follows an image-and-text instruction to find the destination.","explanation":"This was released by Google DeepMind in July 2024. The task it targets is called MINT (Multimodal Instruction Navigation with demonstration Tours): someone first walks through the environment with a camera to record a tour video, and afterward the user can give instructions with text plus an image — for example, holding an object and asking “where does this go back.” The system has two layers: the high level uses Gemini 1.5 Pro, with a context length of up to 1 million tokens, to read through the entire tour video and the instruction and locate the frame where the target is; the low level uses COLMAP (a tool that recovers camera pose from images) to build a topological map from the video — a map where locations are nodes and passable connections are edges — and generates waypoint actions from it for the base to execute. In a real, occupied 836-square-meter office, end-to-end success rates for instructions requiring reasoning and for multimodal instructions were 86% and 90%, respectively.","example":"The user holds up a charger and asks “where should this go back,” and the robot first locates, within the tour video, the desk where the charger belongs, then drives there along the topological map.","related":["Vision-and-Language Navigation","Topological Map","Hierarchical Architecture","Google Gemini","Context Length","Structure from Motion"]},{"id":"navigation-world-models","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Navigation World Models (arXiv 2412.03572)","url":"https://arxiv.org/abs/2412.03572"},{"title":"Navigation World Models 项目页","url":"https://www.amirbar.net/nwm/"}],"as_of":"2025-06","related_ids":["world-model","video-prediction-model","diffusion-transformer","nomad","image-goal-navigation","meta-fundamental-ai-research"],"name":"Navigation World Models (Meta)","alt":"导航世界模型","abbr":"NWM","aliases":["NWM"],"one_liner":"A late-2024 Meta video world model for navigation that imagines what you'd see after taking a given action.","explanation":"Navigation World Models was proposed by Amir Bar, Yann LeCun, and colleagues at Meta FAIR, together with NYU and Berkeley, posted to arXiv in December 2024, and received a CVPR 2025 best-paper honorable mention. It is a controllable video-generation model: given past frames and a navigation action (which way to go, how much to turn), it predicts the footage you'd see next. The model is a 1-billion-parameter conditional diffusion Transformer (CDiT), trained on first-person video from both humans and robots. With it, a robot can simulate several candidate routes inside the model first and pick whichever reaches the goal; it can also score and rank trajectories sampled by an existing policy like NoMaD; and it can imagine walking through an unfamiliar environment given just a single photo. The authors also note that generating for a long time in an unfamiliar environment gradually drifts the footage back toward the training data.","example":"Given a starting photo and a goal photo, NWM generates the along-the-way footage for a batch of candidate action sequences one by one, and picks whichever ends up closest to the goal photo to execute.","related":["World Model","Video Prediction Model","Diffusion Transformer","NoMaD","Image-Goal Navigation","Meta Fundamental AI Research"]},{"id":"navila","category":"named_model","sec":11,"tier":3,"sources":[{"title":"NaVILA: Legged Robot Vision-Language-Action Model for Navigation (arXiv 2412.04453)","url":"https://arxiv.org/abs/2412.04453"},{"title":"NaVILA 项目页","url":"https://navila-bot.github.io/"}],"as_of":"2025-06","related_ids":["vision-and-language-navigation","vision-language-action-model","legged-locomotion","hierarchical-architecture","rl-based-locomotion-control","navid"],"name":"NaVILA","alt":"NaVILA","abbr":"","aliases":["Legged Robot Vision-Language-Action Model for Navigation"],"one_liner":"A legged-robot navigation VLA where a large model gives mid-level actions in text and an RL locomotion controller does the walking.","explanation":"NaVILA was proposed by Xiaolong Wang's group at UC San Diego together with NVIDIA and USC, posted to arXiv in December 2024, published at RSS 2025. A vision-language model is good at understanding images and instructions, but outputting leg joint commands directly is hard, so NaVILA splits into two layers: the high level is a VLA fine-tuned from NVIDIA's VILA (8B), which looks at the camera view and the instruction and outputs, in text, a mid-level action with distance and angle, such as “move forward 75 centimeters”; the low level is a locomotion policy trained with reinforcement learning that reads a lidar-generated height map and turns that mid-level action into commands for a quadruped's 12 joints. Besides simulated navigation data, training data also includes about 2,000 first-person YouTube tour videos. It reaches a 54% success rate on R2R-CE, and the paper also released VLN-CE-Isaac, a benchmark built on Isaac Lab.","example":"In real-robot tests across 25 instructions spanning office, home, and outdoor settings, NaVILA running on a Unitree Go2 quadruped reached 88% success, and 75% on complex multi-room instructions; the same model also works on a Booster T1 humanoid with no retraining.","related":["Vision-and-Language Navigation","Vision-Language-Action Model","Legged Locomotion","Hierarchical Architecture","RL-based Locomotion Control","NaVid"]},{"id":"trackvla","category":"named_model","sec":11,"tier":3,"sources":[{"title":"TrackVLA: Embodied Visual Tracking in the Wild (arXiv 2505.23189)","url":"https://arxiv.org/abs/2505.23189"},{"title":"TrackVLA project page","url":"https://pku-epic.github.io/TrackVLA-web/"}],"as_of":"2025-05","related_ids":["embodied-visual-tracking","vision-language-action-model","navfom","navigation","diffusion-model","unitree-go2"],"name":"TrackVLA","alt":"银河通用 TrackVLA","abbr":"","aliases":["Galbot TrackVLA","TrackVLA: Embodied Visual Tracking in the Wild"],"one_liner":"A Galbot VLA for embodied visual tracking that simultaneously recognizes and plans a path to follow a target, using only its own first-person camera.","explanation":"TrackVLA was released in May 2025 by Peking University's EPIC Lab (He Wang's group) together with Galbot and other institutions, published at CoRL 2025, a vision-language-action (VLA) model. The task is embodied visual tracking: using only its own first-person camera, the robot must recognize a specified target in a dynamic environment and keep following it. Earlier approaches split 'recognizing the target' and 'planning where to go' into two separate modules, which lets errors compound; TrackVLA instead has both share a single large-language-model backbone (Vicuna-7B), with a language head handling recognition and an anchor-based diffusion model outputting the movement trajectory. Training uses about 1.7 million samples, half of them tracking examples from the team's own EVT-Bench simulation benchmark and half video question-answering recognition samples. It has been deployed on a real Unitree Go2 quadruped, with the model running on a remote RTX 4090 server at roughly 10 frames per second, and it is reasonably robust to occlusion and fast target movement.","example":"Told to 'follow the person in the blue shirt,' a Unitree Go2 running only a single head-mounted RGB camera recognizes the target in a crowd and keeps following it, and can pick the target back up after it is briefly blocked from view.","related":["Embodied Visual Tracking","Vision-Language-Action Model","NavFoM (Galbot)","Navigation","Diffusion Model","Unitree Go2"]},{"id":"navfom","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Embodied Navigation Foundation Model (arXiv 2509.12129)","url":"https://arxiv.org/abs/2509.12129"},{"title":"NavFoM 项目页","url":"https://pku-epic.github.io/NavFoM-Web/"},{"title":"银河通用发布全球首个跨本体全域环视导航大模型NavFoM（网易，2025-11-05）","url":"https://c.m.163.com/news/a/KDJOHKOP0519D45U.html"}],"as_of":"2025-11","related_ids":["cross-embodiment","vision-and-language-navigation","embodied-visual-tracking","trackvla","navid","astrabrain"],"name":"NavFoM (Galbot)","alt":"银河通用 NavFoM","abbr":"NavFoM","aliases":["Embodied Navigation Foundation Model"],"one_liner":"A 2025 navigation foundation model from Galbot and Peking University, with one set of weights adapting to many embodiments and navigation tasks.","explanation":"NavFoM was proposed by Galbot together with Peking University, the University of Adelaide, Zhejiang University, and others; the paper went up on arXiv in September 2025, and the company formally launched it in November, promoted as “the world's first cross-embodiment, all-around navigation foundation model.” Earlier navigation models were mostly trained separately, one robot and one task at a time; NavFoM instead uses 8.02 million navigation samples (quadruped, drone, wheeled robot, car) plus 4.76 million image-text and video question-answer samples to simultaneously learn instruction navigation, object finding, target tracking, and autonomous driving. It uses Qwen2-7B as its language backbone, concatenating DINOv2 and SigLIP visual features, with special tokens marking which camera and which moment each frame comes from, supporting 1 to 8 camera feeds; an MLP finally outputs waypoints, handed off to the embodiment's local planner for execution.","example":"The same weights, with no task-specific fine-tuning, are evaluated on benchmarks including VLN-CE instruction navigation, HM3D-OVON open-vocabulary object finding, EVT-Bench target tracking, and NAVSIM autonomous driving; Peking University and Galbot's later UrbanVLA was also trained on top of it.","related":["Cross-Embodiment","Vision-and-Language Navigation","Embodied Visual Tracking","TrackVLA","NaVid","AstraBrain"]},{"id":"alvinn","category":"named_model","sec":11,"tier":3,"sources":[{"title":"ALVINN: An Autonomous Land Vehicle in a Neural Network (NIPS 1988 proceedings)","url":"https://proceedings.neurips.cc/paper/1988/hash/812b4ba287f5ee0bc9d43bbf5bbe87fb-Abstract.html"},{"title":"End to End Learning for Self-Driving Cars (NVIDIA, arXiv 1604.07316)","url":"https://arxiv.org/abs/1604.07316"}],"as_of":"2016-04","related_ids":["end-to-end","behavior-cloning","autonomous-driving","imitation-learning","multilayer-perceptron","tesla-fsd-v12"],"name":"ALVINN","alt":"ALVINN","abbr":"ALVINN","aliases":["Autonomous Land Vehicle In a Neural Network"],"one_liner":"Carnegie Mellon's 1988 driving neural network, which maps camera images of the road directly to a steering command.","explanation":"ALVINN is work by Dean Pomerleau at Carnegie Mellon University, published at NIPS 1988. Self-driving at the time relied on hand-designed image-processing rules that broke easily when lighting or road conditions changed. ALVINN used a backpropagation network with just one hidden layer of 29 units instead: it took in a 30×32 camera image and an 8×32 laser range-finder image, and produced 45 output units representing the steering curvature to follow, with no hand-written road-detection rules in between. The network was first trained on 1,200 synthetically generated road images, then installed on CMU's NAVLAB test vehicle, which it drove autonomously at 0.5 m/s along a 400-meter wooded path on campus. It's often regarded as one of the earliest examples of behavior cloning and end-to-end driving, and NVIDIA's 2016 end-to-end driving system, DAVE-2, cites it as an inspiration in its paper.","example":"For each frame of road image it reads in, the network's 45 steering output units light up most strongly at the position that determines direction: the center unit means drive straight, and units further toward either side mean turn more sharply left or right.","related":["End-to-End","Behavior Cloning","Autonomous Driving","Imitation Learning","Multilayer Perceptron","Tesla FSD v12"]},{"id":"uniad","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Planning-oriented Autonomous Driving (arXiv 2212.10156)","url":"https://arxiv.org/abs/2212.10156"},{"title":"OpenDriveLab/UniAD (GitHub)","url":"https://github.com/OpenDriveLab/UniAD"},{"title":"CVPR 2023 Awards","url":"https://cvpr2023.thecvf.com/Conferences/2023/Awards"}],"as_of":"2025-10","related_ids":["autonomous-driving","end-to-end","tesla-fsd-v12","drivevlm","birds-eye-view","occupancy-network"],"name":"UniAD","alt":"UniAD（规划导向的端到端自动驾驶）","abbr":"UniAD","aliases":["Unified Autonomous Driving","Planning-oriented Autonomous Driving"],"one_liner":"An end-to-end autonomous-driving framework that chains perception, prediction, and planning into one network, with everything serving the final plan.","explanation":"UniAD was released by the Shanghai Artificial Intelligence Laboratory's OpenDriveLab (Hongyang Li's group) and other institutions in December 2022, winning the CVPR 2023 Best Paper Award. Traditional autonomous driving splits perception, prediction, and planning into separate modules or parallel multi-task heads optimized independently, which loses information and compounds errors between modules. UniAD's principle is 'planning-oriented': a single network sequentially performs object tracking, online mapping, motion prediction, occupancy prediction, and trajectory planning, with information passed between modules through shared query vectors, and the whole system optimized toward the final planning outcome. It outperformed prior methods across every metric on the nuScenes dataset, becoming a landmark for end-to-end autonomous driving that is also frequently cited when discussing end-to-end design in embodied AI more broadly. The team released a UniAD 2.0 code version in October 2025.","example":"UniAD takes in images from multiple surround-view cameras on a car, tracks nearby vehicles, draws lane lines, and predicts their movement over the next few seconds, all within the same network, and directly outputs the ego vehicle's upcoming driving trajectory.","related":["Autonomous Driving","End-to-End","Tesla FSD v12","DriveVLM","Bird’s-Eye View","Occupancy Network"]},{"id":"tesla-fsd-v12","category":"named_model","sec":11,"tier":2,"sources":[{"title":"Tesla Software Release 2024.3 release notes (FSD Supervised v12)","url":"https://tesla-info.com/release/2024.3"},{"title":"Tesla pushes end-to-end neural networks for highway driving, but only for newer vehicles (Electrek, 2024-11)","url":"https://electrek.co/2024/11/14/tesla-pushes-end-to-end-neural-networks-for-highway-driving-but-only-for-newer-vehicles/"},{"title":"Elon Musk demonstrates Tesla FSD 12 in a live stream (Tesla Oracle, 2023-08)","url":"https://www.teslaoracle.com/2023/08/27/elon-musk-demonstrates-tesla-fsd-12-no-code-autopilot-ai"}],"as_of":"2024-11","related_ids":["end-to-end","autonomous-driving","imitation-learning","tesla","data-flywheel","uniad"],"name":"Tesla FSD v12","alt":"特斯拉 FSD V12（端到端自动驾驶）","abbr":"FSD","aliases":["FSD Beta v12","Full Self-Driving (Supervised) v12","End-to-End Neural Network Driving"],"one_liner":"Tesla's 2024 driver-assist release that replaced its city-street driving stack with a single end-to-end neural network.","explanation":"FSD (Full Self-Driving) is Tesla's driver-assistance software; its official name includes “Supervised,” meaning a driver must still monitor it. V12 changed the underlying architecture. Elon Musk livestreamed a test drive in August 2023, it rolled out to users in early 2024, and v12.3 expanded the rollout with a free one-month trial for US owners in March. According to Tesla, the city-driving stack was replaced with a single end-to-end neural network trained on millions of video clips, replacing more than 300,000 lines of hand-written C++ rules: camera footage goes in, driving controls come out, with no hand-designed rules in between. In November 2024, v12.5.6.3 extended end-to-end driving to highways too, but only for cars with HW4 hardware. V12 turned “end-to-end” into an industry buzzword, and both the self-driving and embodied-AI communities often cite it as evidence that a data-driven approach can work; Tesla has not published model details.","example":"During Musk's roughly 45-minute livestreamed test drive of V12 in August 2023, he took over the wheel only once, when the car misread a signal at a busy intersection and was about to run a red light. He said there was no code specifically written to handle traffic lights — this kind of behavior was learned entirely from fleet video.","related":["End-to-End","Autonomous Driving","Imitation Learning","Tesla","Data Flywheel","UniAD"]},{"id":"drivevlm","category":"named_model","sec":11,"tier":3,"sources":[{"title":"DriveVLM (arXiv 2402.12289)","url":"https://arxiv.org/abs/2402.12289"},{"title":"DriveVLM 项目主页","url":"https://tsinghua-mars-lab.github.io/DriveVLM/"}],"as_of":"2024-11","related_ids":["dual-system-architecture","autonomous-driving","vision-language-model","chain-of-thought","long-tail-problem","qwen-vl"],"name":"DriveVLM","alt":"DriveVLM（快慢双系统智驾）","abbr":"","aliases":["DriveVLM-Dual","The Convergence of Autonomous Driving and Large Vision-Language Models"],"one_liner":"A 2024 Tsinghua and Li Auto autonomous-driving system where a vision-language model thinks slowly and a traditional planner executes fast.","explanation":"DriveVLM is an autonomous-driving method proposed in February 2024 by Hang Zhao's group at Tsinghua University's Institute for Interdisciplinary Information Sciences together with Li Auto, published at CoRL 2024. Traditional autonomous-driving pipelines have limited ability to handle rare, complex long-tail scenarios. DriveVLM uses a vision-language model (VLM, built on Qwen-VL) to work through a chain of thought in three steps — scene description, scene analysis, and hierarchical planning — before finally outputting a driving trajectory. But a VLM isn't precise enough at spatial localization and is slow to run inference, so the paper also proposes DriveVLM-Dual: the VLM outputs a rough reference trajectory at low frequency (the slow system), while a traditional perception-and-planning module refines it into the actual trajectory at high frequency (the fast system), the two working together asynchronously. The paper reports the system has been deployed in production vehicles, with average inference of about 410 milliseconds on an onboard platform with two Orin X chips. This division of labor matches the fast-slow dual-system architecture also seen in robotics, such as Figure's Helix.","example":"When it encounters an unusual road situation, the VLM first describes the scene, identifies key objects that could affect the vehicle and analyzes their impact, then gives a high-level decision such as “slow down and go around” along with a rough trajectory, which the traditional planner refines in real time into an executable trajectory.","related":["Dual-System Architecture (System 1 / System 2)","Autonomous Driving","Vision-Language Model","Chain-of-Thought","Long-tail Problem","Qwen-VL"]},{"id":"gaia-2","category":"named_model","sec":11,"tier":3,"sources":[{"title":"GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving (arXiv 2503.20523)","url":"https://arxiv.org/abs/2503.20523"},{"title":"GAIA-2（Wayve 官方博客）","url":"https://wayve.ai/thinking/gaia-2/"}],"as_of":"2025-03","related_ids":["world-model","autonomous-driving","latent-diffusion-model","synthetic-data","waymo-world-model","nvidia-cosmos"],"name":"GAIA-2 (Wayve)","alt":"Wayve GAIA-2","abbr":"GAIA","aliases":["A Controllable Multi-View Generative World Model for Autonomous Driving"],"one_liner":"A 2025 driving world model from UK self-driving company Wayve that generates controllable, multi-camera driving video.","explanation":"GAIA-2 is a generative world model released by the UK self-driving company Wayve on March 26, 2025, the successor to GAIA-1. “World model” here means a video-generation model that can “imagine” future driving-scene footage conditioned on given inputs. GAIA-2 moves away from GAIA-1's autoregressive token generation to a video-tokenizer-plus-latent-diffusion architecture (compressing video into a latent space first, then denoising and generating within that latent space), natively supporting several cameras generated at once that stay consistent with each other in both time and space, with data covering the UK, the US, and Germany. Generation can be controlled along several axes: the ego vehicle's speed and steering, the behavior of other vehicles and pedestrians, weather and time of day, and road structure such as lanes and intersections. It is used to mass-produce synthetic data, create variations on real driving logs, and generate rare, dangerous scenarios to test driving models. It is one of the earlier examples of world models being put to practical use in physical AI.","example":"Take a real multi-camera driving log, change the weather to heavy rain, and add a car suddenly cutting in, to generate a new set of surround-view video used to check how a driving model reacts to this rare situation.","related":["World Model","Autonomous Driving","Latent Diffusion Model","Synthetic Data","Waymo World Model","NVIDIA Cosmos"]},{"id":"nvidia-alpamayo","category":"named_model","sec":11,"tier":3,"sources":[{"title":"Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail (arXiv 2511.00088)","url":"https://arxiv.org/abs/2511.00088"},{"title":"NVlabs/alpamayo GitHub 仓库（Alpamayo 1）","url":"https://github.com/NVlabs/alpamayo"},{"title":"NVIDIA Announces Alpamayo Family of Open-Source AI Models（NVIDIA Newsroom，2026-01-05）","url":"https://nvidianews.nvidia.com/news/alpamayo-autonomous-vehicle-development"}],"as_of":"2026-01","related_ids":["vision-language-action-model","autonomous-driving","chain-of-thought","nvidia-cosmos-reason","action-expert","long-tail-problem"],"name":"NVIDIA Alpamayo","alt":"英伟达 Alpamayo（驾驶推理 VLA）","abbr":"AR1","aliases":["Alpamayo-R1","Alpamayo 1","Alpamayo-R1-10B","NVIDIA DRIVE Alpamayo-R1"],"one_liner":"NVIDIA's open autonomous-driving reasoning VLA, which writes out in text why it's driving this way before outputting a trajectory.","explanation":"Alpamayo-R1 is NVIDIA's autonomous-driving vision-language-action model; the paper went up on arXiv in late October 2025, weights opened on Hugging Face in December, and it was renamed Alpamayo 1 at CES in January 2026, released together with the AlpaSim simulation framework and more than 1,700 hours of open driving data. It targets rare, long-tail road situations: imitating trajectories alone tends to fail in unseen scenarios, so the model first writes out a “causal chain” of reasoning (what it sees, and therefore what to do), then generates the trajectory. The open version uses the Cosmos-Reason vision-language model as its backbone (8.2B), plus a 2.3B diffusion-based action expert, taking in 4 camera feeds and the ego vehicle's motion history, and outputting a 6.4-second future trajectory. The paper reports up to 12% better planning accuracy on hard scenarios, a 35% drop in close-call near-misses in closed-loop simulation, and about 99 milliseconds of onboard latency.","example":"When the lane ahead is blocked by construction cones, the model first outputs an explanation — such as “lane ahead is closed, need to slow down and change lanes to go around” — then gives the corresponding 6.4-second driving trajectory, letting an engineer check the text against the trajectory to see why it drove that way.","related":["Vision-Language-Action Model","Autonomous Driving","Chain-of-Thought","NVIDIA Cosmos Reason","Action Expert","Long-tail Problem"]},{"id":"waymo-world-model","category":"named_model","sec":11,"tier":3,"sources":[{"title":"The Waymo World Model: A New Frontier for Autonomous Driving Simulation (Waymo Blog)","url":"https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-frontier-for-autonomous-driving-simulation"}],"as_of":"2026-02","related_ids":["genie-3","world-model","autonomous-driving","interactive-world-model","long-tail-problem","lidar"],"name":"Waymo World Model","alt":"Waymo 世界模型","abbr":"","aliases":["The Waymo World Model (built on Genie 3)"],"one_liner":"Waymo's Genie-3-based generative driving simulator, able to generate camera and lidar data together.","explanation":"The Waymo World Model is a generative world model that Waymo announced on February 6, 2026, for large-scale autonomous-driving simulation. It is built on top of Genie 3, Google DeepMind's interactive world model, leveraging Genie 3's pretraining on huge amounts of video to generate photorealistic, interactive 3D driving environments, and outputting both camera and LiDAR sensor data at once. It mainly targets autonomous driving's long-tail problem: dangerous situations like extreme weather, natural disasters, or rare objects are seldom encountered during real-world test driving, yet must still be tested for in advance. It supports three kinds of control: changing the self-driving car's own actions to run 'what if we had driven differently' counterfactuals; controlling road layout and traffic conditions; and using language to change the time of day or weather, or even generate an entirely synthetic scene. It illustrates how general-purpose world models are starting to be used as both a test environment and a data source for embodied systems.","example":"A single sentence can turn a real, sunny daytime test-drive scene into a rainy night, and the self-driving car can be made to change lanes differently, to observe how the autonomous-driving system reacts on the same stretch of road.","related":["Genie 3","World Model","Autonomous Driving","Interactive World Model","Long-tail Problem","LiDAR"]},{"id":"google-deepmind","category":"company","sec":0,"tier":1,"sources":[{"title":"Google DeepMind - Wikipedia","url":"https://en.wikipedia.org/wiki/Google_DeepMind"},{"title":"Gemini Robotics 2 brings whole-body intelligence to robots","url":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"}],"as_of":"2026-07","related_ids":["rt-2","gemini-robotics","gemini-robotics-2","open-x-embodiment","genie-3","vision-language-action-model"],"name":"Google DeepMind","alt":"谷歌 DeepMind","abbr":"GDM","aliases":["DeepMind","Robotics at Google","Google Brain robotics team"],"one_liner":"Google's AI research division, source of the RT series and Gemini Robotics.","explanation":"DeepMind was founded in London in 2010 by Demis Hassabis, Shane Legg, and Mustafa Suleyman, acquired by Google in January 2014, and merged with Google Brain in April 2023 to form Google DeepMind, still headquartered in London. Around the time of the merger (with Google Brain's own robotics team, Robotics at Google, folded in), its landmark robotics work included RT-1, RT-2 (one of the first VLAs), the cross-embodiment dataset Open X-Embodiment, RoboCat, and ALOHA Unleashed. Starting in 2025, the focus shifted to the Gemini Robotics series: the first version in March, 1.5 in September, and Gemini Robotics 2, which supports whole-body humanoid control, on July 30, 2026. On the hardware side, it partners with Apptronik and Boston Dynamics.","example":"RT-2 directly fine-tuned a vision-language model trained on web image-text data to output robot action tokens, helping popularize the term 'VLA.'","related":["RT-2","Gemini Robotics","Gemini Robotics 2","Open X-Embodiment","Genie 3","Vision-Language-Action Model"]},{"id":"nvidia","category":"company","sec":0,"tier":1,"sources":[{"title":"NVIDIA and Global Robotics Leaders Take Physical AI to the Real World (GTC 2026)","url":"https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world"},{"title":"Jetson Thor","url":"https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/"},{"title":"Hugging Face - Wikipedia（英伟达收购报道）","url":"https://en.wikipedia.org/wiki/Hugging_Face"}],"as_of":"2026-09","related_ids":["nvidia-three-computer-solution","nvidia-isaac-lab","nvidia-isaac-gr00t-n1","nvidia-cosmos","nvidia-jetson-thor","nvidia-generalist-embodied-agent-research-lab"],"name":"NVIDIA","alt":"英伟达","abbr":"","aliases":["Nvidia"],"one_liner":"The GPU giant, supplying a full stack of compute and software for training, simulating, and deploying robots.","explanation":"NVIDIA was founded in 1993 by Jensen Huang and colleagues, headquartered in Santa Clara, California, and became the leading supplier of AI training compute through its GPUs. In embodied AI it promotes a 'three computers' framing: DGX clusters for training models, Omniverse and Isaac Sim/Isaac Lab for simulation and synthetic data, and Jetson Thor for running models onboard the robot. It also builds its own models: the GR00T N1 series of humanoid foundation models (starting March 2025) and the Cosmos world foundation model (with the third generation, Cosmos 3, announced at GTC in March 2026); at the same event it also previewed GR00T N2, shifting toward a 'world action model' approach, planned to open up before year's end. Research is led by teams including the GEAR lab. On September 3, 2026, NVIDIA announced it would acquire Hugging Face, pending regulatory approval.","example":"Many quadruped and humanoid reinforcement-learning locomotion policies are trained in Isaac Lab across thousands of parallel simulated environments, then deployed onto Jetson hardware.","related":["NVIDIA Three-Computer Solution","NVIDIA Isaac Lab","NVIDIA Isaac GR00T N1","NVIDIA Cosmos","NVIDIA Jetson Thor","NVIDIA Generalist Embodied Agent Research Lab"]},{"id":"tesla","category":"company","sec":0,"tier":1,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"}],"as_of":"2026-09","related_ids":["tesla-optimus","tesla-optimus-v3","tesla-supply-chain","tesla-fsd-v12","tesla-ai-day","mass-production"],"name":"Tesla","alt":"特斯拉","abbr":"","aliases":["Tesla Optimus team"],"one_liner":"The electric-car company developing the Optimus humanoid robot.","explanation":"Tesla was founded in 2003 and is headquartered in Austin, Texas, known for its electric vehicles and FSD driver-assistance system. Its embodied AI project is the Optimus humanoid robot: unveiled at AI Day in August 2021, with a second generation shown in December 2023, built around reusing the vision neural networks, batteries, and motor supply chain from Tesla's self-driving effort. In June 2025, project lead Milan Kovac departed, with Autopilot lead Ashok Elluswamy taking over. Elon Musk has said the target price is around $30,000; the Fremont factory is installing a line meant for a million units a year, with an even larger line planned at the Texas factory. As of September 2026, a third generation intended for mass production had not yet been unveiled. Optimus's production plans have driven a wave of domestic Chinese parts suppliers, referred to as the 'Tesla supply chain' (T链).","example":"","related":["Tesla Optimus","Tesla Optimus V3 (Optimus Gen 3)","Tesla (Optimus) Supply Chain","Tesla FSD v12","Tesla AI Day","Mass Production"]},{"id":"hugging-face","category":"company","sec":0,"tier":1,"sources":[{"title":"Hugging Face - Wikipedia","url":"https://en.wikipedia.org/wiki/Hugging_Face"},{"title":"SmolVLA - Hugging Face Blog","url":"https://huggingface.co/blog/smolvla"}],"as_of":"2026-09","related_ids":["lerobot","smolvla","so-100-so-101-arm","pollen-robotics-reachy-mini","pollen-robotics","lerobotdataset"],"name":"Hugging Face","alt":"Hugging Face","abbr":"HF","aliases":["抱抱脸 (colloquial Chinese nickname)","Hugging Face Hub"],"one_liner":"An open-source AI model and dataset hosting platform, whose robotics effort is LeRobot.","explanation":"Hugging Face was founded in 2016 by French entrepreneurs Clément Delangue, Julien Chaumond, and Thomas Wolf, and is headquartered in New York. It is best known for the Hub, its model and dataset hosting platform, and the Transformers library — most researchers who download open-source weights or upload datasets pass through it. In robotics, it built the open-source framework LeRobot (a unified dataset format plus training and deployment code), paired with the SO-100/SO-101, an open-source robot arm whose parts cost a bit over $100; in April 2025 it acquired French humanoid robot company Pollen Robotics (maker of Reachy 2), and in June it released the small VLA SmolVLA. After a funding round in August 2023 it was valued at $4.5 billion. On September 3, 2026, NVIDIA announced it would acquire Hugging Face for about $12.9 billion, a deal expected to close in the first half of 2027 pending regulatory approval; NVIDIA said Hugging Face would keep its brand and continue operating as an open platform.","example":"A newcomer teleoperates an SO-101 with LeRobot to record a few dozen demonstrations, uploads them to the Hub, and then trains ACT or SmolVLA with a single command.","related":["LeRobot","SmolVLA","SO-100 / SO-101 Arm","Pollen Robotics Reachy Mini","Pollen Robotics","LeRobotDataset"]},{"id":"nvidia-generalist-embodied-agent-research-lab","category":"company","sec":0,"tier":2,"sources":[{"title":"NVIDIA GEAR Lab","url":"https://research.nvidia.com/labs/gear/"},{"title":"GR00T N1 paper","url":"https://arxiv.org/abs/2503.14734"}],"as_of":"2026-04","related_ids":["nvidia-isaac-gr00t-n1",null,"eureka","voyager","dreamgen","sonic"],"name":"NVIDIA Generalist Embodied Agent Research Lab","alt":"英伟达 GEAR 实验室","abbr":"GEAR","aliases":["NVIDIA GEAR Lab","GEAR Lab","GEAR"],"one_liner":"NVIDIA Research's group for general-purpose embodied agents, the team behind the GR00T humanoid model.","explanation":"GEAR is the NVIDIA Research team working on general-purpose embodied agents, co-led by Jim Fan (a Stanford PhD) and Yuke Zhu (also a professor at UT Austin); it reportedly became a formal team in early 2024, with the goal of building foundation models for agents in both virtual and physical worlds. Before GEAR, its members had worked on the Minecraft agent Voyager, MineDojo, and Eureka, which uses a large language model to write reward functions. Since forming, the team's flagship project has been the GR00T humanoid foundation model, from N1 (open-sourced in March 2025) through N1.7 (April 2026), alongside related work such as the DreamGen synthetic-data pipeline, the SONIC whole-body controller, EgoScale human-video pretraining, and the DreamZero world-action model.","example":"Developers can download the GR00T N1.x weights and fine-tune them on a few dozen demonstrations from their own robot to get a policy that follows instructions.","related":["NVIDIA Isaac GR00T N1","NVIDIA","Eureka","Voyager","DreamGen","SONIC"]},{"id":"openai","category":"company","sec":0,"tier":2,"sources":[{"title":"OpenAI - Wikipedia","url":"https://en.wikipedia.org/wiki/OpenAI"},{"title":"Caitlin Kalinowski - Wikipedia","url":"https://en.wikipedia.org/wiki/Caitlin_Kalinowski"},{"title":"Solving Rubik's Cube with a Robot Hand","url":"https://arxiv.org/abs/1910.07113"}],"as_of":"2026-04","related_ids":["dactyl","clip","openai-gpt-series","sora",null],"name":"OpenAI","alt":"OpenAI","abbr":"","aliases":[],"one_liner":"The maker of ChatGPT, which in its early years also built a robot hand that solved a Rubik's Cube.","explanation":"OpenAI is an AI company founded in San Francisco in December 2015 by Sam Altman, Elon Musk, Ilya Sutskever, Greg Brockman, and others. It is best known for the GPT model series and ChatGPT, and it also created CLIP and Sora. Its earlier connection to embodied AI was robotics: from 2018 to 2019, its Dactyl project combined reinforcement learning with domain randomization to train a Shadow dexterous hand in simulation, then transferred that policy to solve a Rubik's Cube on the real hand; the robotics team was reportedly disbanded in 2021. In November 2024, OpenAI hired Caitlin Kalinowski to lead robotics and consumer hardware, and she resigned in March 2026. The company converted to a public-benefit corporation structure in October 2025, and was valued at about $852 billion post-money in April 2026.","example":"Many robot-learning models use OpenAI's CLIP as a vision or text encoder, placing images and language into the same embedding space.","related":["Dactyl","CLIP","OpenAI GPT Series","Sora","Automatic Domain Randomization"]},{"id":"meta","category":"company","sec":0,"tier":2,"sources":[{"title":"Meta buys robotics startup to bolster its humanoid AI ambitions (TechCrunch, 2026-05)","url":"https://techcrunch.com/2026/05/01/meta-buys-robotics-startup-to-bolster-its-humanoid-ai-ambitions/"},{"title":"Meta wants to become the Android of robotics (Engadget, 2025-09)","url":"https://www.engadget.com/big-tech/meta-wants-to-become-the-android-of-robotics-220701800.html"},{"title":"Meta Platforms accelerates robotics efforts with leadership changes (DIGITIMES, 2025-11)","url":"https://www.digitimes.com/news/a20251119PD205/meta-robotics-humanoid-robot-smart-glasses.html"}],"as_of":"2026-05","related_ids":["meta-fundamental-ai-research","assured-robot-intelligence","v-jepa-2","digit-360","meta-motivo","ami-labs"],"name":"Meta","alt":"Meta","abbr":"","aliases":["Meta Platforms","Facebook"],"one_liner":"Facebook's parent company, whose FAIR lab researches world models while a separate team builds humanoid robots.","explanation":"Meta began as Facebook, founded by Mark Zuckerberg in 2004, and was renamed in 2021; it is headquartered in Menlo Park, California. Two threads connect it to embodied AI. First is its foundational-research division, FAIR, which has produced the video world model V-JEPA 2, the vision-and-touch sensor Digit 360, and the simulated-humanoid-control model Meta Motivo, most of them released open source. Second is a humanoid-robotics team formed within Reality Labs in February 2025, led by former Cruise CEO Marc Whitten, working internally under the project name Metabot; CTO Andrew Bosworth said in September 2025 that Meta wants to build a software platform it can license to other robot makers, similar to Android. In May 2026, Meta acquired Assured Robot Intelligence, a startup working on full-body humanoid control, folding the team into its Superintelligence Labs (MSL).","example":"V-JEPA 2-AC pairs an action predictor with Meta's open-sourced video model to serve as a world model, letting a Franka arm perform zero-shot pick-and-place.","related":["Meta Fundamental AI Research","Assured Robot Intelligence","V-JEPA 2","Digit 360","Meta Motivo","AMI Labs"]},{"id":"meta-fundamental-ai-research","category":"company","sec":0,"tier":2,"sources":[{"title":"Meta AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Meta_AI"},{"title":"Yann LeCun - Wikipedia","url":"https://en.wikipedia.org/wiki/Yann_LeCun"}],"as_of":"2026-03","related_ids":["v-jepa-2",null,"dinov2","habitat",null,"joint-embedding-predictive-architecture"],"name":"Meta Fundamental AI Research","alt":"Meta FAIR","abbr":"FAIR","aliases":["FAIR","Facebook AI Research","Meta AI"],"one_liner":"Meta's foundational AI research lab, known for open models like Segment Anything, DINO, V-JEPA, and the Habitat simulator.","explanation":"FAIR is Meta's (formerly Facebook's) fundamental AI research division, founded in 2013 as Facebook AI Research with Yann LeCun as its first lab director. Its original lab was in New York, and it later grew teams in Menlo Park, London, and Paris, among other locations. It has produced a long list of results relevant to embodied AI: the Segment Anything Model (SAM) family, the self-supervised vision models DINOv2 and DINOv3, the video world model V-JEPA 2, the indoor-navigation simulator Habitat, and the DIGIT and Digit 360 visuotactile sensors. Most of these are open-sourced and widely used as building blocks in robotics research. In 2025, Meta folded FAIR and its large-model teams into Meta Superintelligence Labs (MSL); LeCun left the company in November 2025 to found AMI Labs, a world-model startup.","example":"Many manipulation policies use DINOv2 as their vision encoder and SAM to cut the target object's mask out of an image.","related":["V-JEPA 2","Segment Anything Model","DINOv2","Habitat","DIGIT Visuotactile Sensor","Joint-Embedding Predictive Architecture"]},{"id":"amazon-robotics","category":"company","sec":0,"tier":2,"sources":[{"title":"Amazon Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Amazon_Robotics"},{"title":"Amazon launches a new AI foundation model to power its robotic fleet and deploys its 1 millionth robot","url":"https://www.aboutamazon.com/news/operations/amazon-million-robots-ai-foundation-model"}],"as_of":"2025-07","related_ids":["amazon-vulcan","autonomous-mobile-robot","amazon-frontier-ai-and-robotics","covariant","agility-robotics","fleet-management-system"],"name":"Amazon Robotics","alt":"亚马逊机器人","abbr":"","aliases":["Kiva Systems"],"one_liner":"Amazon's warehouse robotics division, evolved from Kiva Systems, now running a fleet of more than a million robots.","explanation":"Amazon Robotics grew out of Kiva Systems, founded in 2003 by Mick Mountz, Peter Wurman, and Raffaello D'Andrea. Kiva's orange carts drove underneath shelving units and carried the whole shelf to a picker, flipping 'person goes to the goods' into 'goods come to the person.' Amazon acquired Kiva for $775 million in 2012 and renamed it Amazon Robotics in 2015; it is headquartered in Massachusetts. Amazon has since rolled out Proteus, a mobile robot that can navigate autonomously around people, Hercules, which carries shelving units, and Vulcan, a picking robot with a sense of touch, among others. By 2025 Amazon had deployed more than a million robots across more than 300 facilities, and also released DeepFleet, a generative foundation model that routes the whole fleet's movement, which the company says improved travel efficiency by about 10%. Amazon Robotics represents specialized warehouse automation already deployed at massive scale, a useful point of comparison against the general-purpose humanoid robot approach.","example":"Vulcan uses force sensing to push aside other items inside a shelf bin crammed with goods, completing both stowing and picking.","related":["Amazon Vulcan","Autonomous Mobile Robot","Amazon Frontier AI & Robotics","Covariant","Agility Robotics","Fleet Management System (e.g. Open-RMF)"]},{"id":"softbank-group","category":"company","sec":0,"tier":2,"sources":[{"title":"SoftBank Group - Wikipedia","url":"https://en.wikipedia.org/wiki/SoftBank_Group"},{"title":"Acquisition of ABB Ltd's Robotics Business（SoftBank Group 新闻稿，2025-10-08）","url":"https://group.softbank/en/news/press/20251008"},{"title":"Skild AI Series C","url":"https://www.skild.ai/blogs/series-c"}],"as_of":"2026-01","related_ids":["skild-ai","boston-dynamics","abb-robotics","softbank-robotics-pepper","softbank-robotics-nao","agile-robots"],"name":"SoftBank Group","alt":"软银集团","abbr":"SBG","aliases":["SBG","SoftBank"],"one_liner":"A Japanese tech-investment conglomerate making large recent bets on robotics and physical AI.","explanation":"SoftBank Group was founded by Masayoshi Son in 1981 and is headquartered in Tokyo; it is primarily a tech-investment holding company, running a Vision Fund with over $100 billion under management, holding roughly 90% of chip designer Arm, and ranking among OpenAI's major shareholders. Its ties to robotics run deep: in 2012 it took control of Aldebaran, the French developer of NAO, and launched Pepper with it in 2014 (selling Aldebaran back off in 2022); in 2017 it bought Boston Dynamics from Google, then sold 80% of it to Hyundai Motor Group in 2020. Since 2025, SoftBank has treated robotics as part of a broader physical-AI strategy: in October it announced a $5.375 billion acquisition of ABB's robotics business (originally expected to close in mid-to-late 2026), and in January 2026 it led a $1.4 billion round in Skild AI, having earlier also invested in Agile Robots. Its name turns up often behind the large funding rounds in embodied-AI news.","example":"In January 2026, Skild AI raised $1.4 billion at a valuation above $14 billion, in a round led by SoftBank.","related":["Skild AI","Boston Dynamics","ABB Robotics","SoftBank Robotics Pepper","SoftBank Robotics NAO","Agile Robots"]},{"id":"huawei","category":"company","sec":0,"tier":2,"sources":[{"title":"Huawei - Wikipedia","url":"https://en.wikipedia.org/wiki/Huawei"},{"title":"华为云 CloudRobo 具身智能平台","url":"https://www.huaweicloud.com/product/cloudrobo.html"}],"as_of":"2026-09","related_ids":["huawei-cloud-cloudrobo-embodied-ai-platform","compute-architecture-for-neural-networks","mindspore","cloud-edge-device-collaboration","robot-brain-company","m-robots-os"],"name":"Huawei","alt":"华为","abbr":"","aliases":["Huawei Cloud","Huawei Technologies Co., Ltd."],"one_liner":"A Shenzhen telecom and tech giant entering embodied AI through chips, cloud platforms, and large models.","explanation":"Huawei was founded in Shenzhen in September 1987; its founder, Ren Zhengfei, previously served in the People's Liberation Army's engineering corps, and its main businesses are telecom equipment, smartphones, chips, and cloud computing. Huawei does not mass-produce humanoid robots; instead it builds the underlying compute, platforms, and models: its Ascend AI chips and CANN software stack, its MindSpore deep-learning framework, its Pangu large model, and Huawei Cloud's embodied-AI development platform, CloudRobo. CloudRobo is currently in open beta, providing manipulation and navigation data, simulation assets, model training, and combined simulation-plus-real-robot evaluation, and it opens up an R2C protocol that the company says can shrink the time to adapt a new robot body from weeks to hours. Huawei has reportedly also set up an embodied-AI industry innovation center in Shenzhen, partnering with multiple robot-body makers.","example":"A robot-body maker connects its own robot through CloudRobo's R2C protocol, trains a model on the platform's data in the cloud, and then runs it through simulation and real-robot evaluation.","related":["Huawei Cloud CloudRobo Embodied AI Platform","Compute Architecture for Neural Networks (Huawei Ascend, CANN)","MindSpore","Cloud-Edge-Device Collaboration","Robot-Brain (Model-Only) Company","M-Robots OS (OpenHarmony-based Robot OS)"]},{"id":"alibaba-group","category":"company","sec":0,"tier":2,"sources":[{"title":"阿里千问发布 Qwen-Robot 具身模型系列（量子位，2026-06）","url":"https://www.qbitai.com/2026/06/435873.html"},{"title":"阿里入场 具身智能迎来超级玩家（财联社）","url":"https://www.cls.cn/detail/2164819"},{"title":"首度布局人形机器人 阿里投了这家深圳企业（财联社，2024-05）","url":"https://www.cls.cn/detail/1681264"}],"as_of":"2026-06","related_ids":["qwen-robot-series","alibaba-damo-academy","rynnbrain","x-square-robot","limx-dynamics","pragmatik-labs"],"name":"Alibaba Group","alt":"阿里巴巴","abbr":"","aliases":["Alibaba"],"one_liner":"A Hangzhou internet giant self-developing embodied models under the Qwen brand while investing in multiple robot-body makers.","explanation":"Alibaba Group was founded in Hangzhou in 1999 by Jack Ma and others, with core businesses in e-commerce and Alibaba Cloud. In embodied AI it follows a “build the brain in-house, invest in the body” strategy. On models, its DAMO Academy open-sourced the embodied-brain model RynnBrain in February 2026, and its Qwen team released the Qwen-Robot family on June 16 — the manipulation model RobotManip, the navigation model RobotNav, and the world model RobotWorld. On investment, it took a stake in LimX Dynamics in May 2024, and in September 2025 Alibaba Cloud led a nearly RMB 1 billion Series A+ round in X Square Robot; it has also invested in RobotEra and Noematrix. In October 2025, Junyang Lin, then head of Qwen's technical team, said the team had already formed a robotics and embodied-AI group internally; he left in March 2026 and later founded Pragmatik Labs.","example":"When news says “Alibaba is entering embodied AI,” it usually means moves like Qwen releasing the Qwen-Robot models or Alibaba Cloud leading a round in X Square Robot — complete robot bodies mostly come through its investments in body makers.","related":["Qwen-Robot Series","Alibaba DAMO Academy","RynnBrain","X Square Robot","LimX Dynamics","Pragmatik Labs"]},{"id":"robbyant","category":"company","sec":0,"tier":2,"sources":[{"title":"Robbyant 官网","url":"https://www.robbyant.com/"},{"title":"Robbyant Technology","url":"https://technology.robbyant.com/"},{"title":"Robbyant | LinkedIn","url":"https://www.linkedin.com/company/robbyant"}],"as_of":"2026-09","related_ids":["lingbot-vla","lingbot-va","lingbot-world","lingbot-depth",null,null],"name":"Robbyant","alt":"蚂蚁灵波","abbr":"","aliases":["Lingbo Technology","Ant Group Embodied Intelligence"],"one_liner":"An embodied-AI company under Ant Group that open-sources the LingBot family of robot models.","explanation":"Robbyant is the embodied-AI company under Ant Group (the fintech affiliate of Alibaba), with the slogan “a brain for every robot,” targeting service scenarios such as eldercare, medical assistance, and housework, and building both models and complete robots. Starting in January 2026, it began open-sourcing its LingBot model family in quick succession: LingBot-Depth (depth completion for transparent and reflective objects), LingBot-VLA (a vision-language-action model pretrained on about 20,000 hours of real-robot data), LingBot-World (an interactive world model adapted from Alibaba's Tongyi Wanxiang video model), and LingBot-VA (a world-action model); LingBot-VLA, LingBot-World, and LingBot-VA were all later upgraded to version 2.0. According to its LinkedIn page, its GitHub repository passed 30,000 stars within 200 days. Hardware listed on its website includes the Robbyant R2 robot with swappable dexterous hands and grippers.","example":"Researchers can take LingBot-VLA's open-source weights and fine-tune them on a small amount of data from their own dual-arm robot to do tabletop manipulation.","related":["LingBot-VLA (Robbyant)","LingBot-VA (Robbyant)","LingBot-World (Robbyant)","LingBot-Depth","World Action Model","Open-weight Model"]},{"id":"bytedance-seed","category":"company","sec":0,"tier":2,"sources":[{"title":"ByteDance Seed 官网","url":"https://seed.bytedance.com/en/"},{"title":"ByteDance - Wikipedia","url":"https://en.wikipedia.org/wiki/ByteDance"}],"as_of":"2026-08","related_ids":["seed-gr-3","gr-rl","robix","gr-1","bagel","vision-language-action-model"],"name":"ByteDance Seed","alt":"字节跳动 Seed","abbr":"","aliases":["字节 Seed (short Chinese form)","Seed robotics team","Seed"],"one_liner":"ByteDance's foundation-model research team, whose robotics effort is the GR series of VLAs.","explanation":"ByteDance Seed is the foundation-model research team ByteDance formed in 2023; the Doubao large language model, the image generator Seedream, and the video generator Seedance all come out of this team. Its robotics GR series began within ByteDance's research group: GR-1 (late 2023) and GR-2 (October 2024) first went through generative pretraining on large-scale video, then were fine-tuned on robot data. Starting in 2025, releases have come under the Seed name: in July, the roughly 4-billion-parameter vision-language-action model GR-3 and the bimanual mobile robot ByteMini; in September, Robix, responsible for high-level reasoning and human-robot dialogue; and in December, GR-Dexter for dexterous hands and GR-RL, which incorporates reinforcement learning. The team has also open-sourced BAGEL, a unified multimodal model. The overall approach combines video generation, vision-language models, and real-robot data to build a general-purpose robot capable of long-horizon, fine-grained tasks.","example":"GR-3 drives ByteMini to clear a dinner table and hang clothes on a hanger, following language instructions.","related":["Seed GR-3","GR-RL","Robix","GR-1 (ByteDance)","BAGEL","Vision-Language-Action Model"]},{"id":"tencent-robotics-x-lab","category":"company","sec":0,"tier":2,"sources":[{"title":"极客公园：具身智能，腾讯「低调入局」","url":"https://www.geekpark.net/news/351962"},{"title":"华尔街见闻：对话腾讯首席科学家张正友","url":"https://wallstreetcn.com/articles/3777373"}],"as_of":"2026-07","related_ids":["tencent-tairos-embodied-ai-open-platform","tencent-hy-embodied","quadruped-robot","wheel-legged-robot","embodied-foundation-model","one-brain-multiple-robots"],"name":"Tencent Robotics X Lab","alt":"腾讯 Robotics X 实验室","abbr":"","aliases":["Tencent Robotics X"],"one_liner":"Tencent's robotics research lab, founded in 2018, now focused on an embodied-AI software platform.","explanation":"Tencent set up this frontier robotics research unit in 2018, led by Tencent chief scientist Zhengyou Zhang, who previously worked at Microsoft Research and devised the widely used “Zhang's method” for camera calibration. At the time, China had almost no mature robot bodies to partner with, so the lab built its own hardware and software from scratch, developing the multimodal quadruped Max, the wheel-legged robot Ollie, and a household-robot prototype nicknamed “Xiaowu.” In early 2025 Tencent said explicitly that it wanted to be a partner to robot makers rather than sell its own hardware; that July, at the World Artificial Intelligence Conference, it launched the embodied-AI open platform Tairos, offering large models, developer tools, and data services in a modular way, already working with makers such as Unitree, AgiBot, Dobot, and LimX Dynamics. It later worked with the Hunyuan team to release the Hy-Embodied family of embodied models.","example":"A robot maker can skip building its own “brain” and instead call the perception and planning models on the Tairos platform, connecting them to its own hardware for guided tours or delivery.","related":["Tencent Tairos Embodied AI Open Platform","Tencent HY-Embodied (Hunyuan Embodied)","Quadruped Robot","Wheel-legged Robot","Embodied Foundation Model","One Brain, Multiple Robots"]},{"id":"jd-com","category":"company","sec":0,"tier":2,"sources":[{"title":"证券时报：京东投资三家具身智能机器人企业","url":"https://stcn.com/article/detail/2666033.html"},{"title":"钛媒体：京东，具身智能大玩家？","url":"https://www.tmtpost.com/8152876.html"},{"title":"新京报：资本疯狂投喂机器人赛道","url":"https://www.bjnews.com.cn/detail/1782822930129028.html"}],"as_of":"2026-09","related_ids":["agibot","limx-dynamics","spirit-ai","engineai","meituan","data-flywheel"],"name":"JD.com","alt":"京东（具身智能投资）","abbr":"","aliases":["JD Group","Jingdong"],"one_liner":"An e-commerce and logistics giant that has invested heavily in embodied AI since 2025 while building its own data and testing grounds.","explanation":"JD.com is an e-commerce and logistics group founded by Richard Liu (Liu Qiangdong) in 1998, headquartered in Beijing. It reportedly set up a group-level embodied-AI unit in March 2025, led by Hui Shen, a former SenseTime vice president; its JD Explore Academy open-sourced the dual-arm mobile-manipulation dataset JD ManiData. In 2025 it invested in AgiBot, then on July 21 led rounds in Spirit AI, LimX Dynamics, and EngineAI on the same day, followed by investments in RoboScience and PaXini Tech, spanning robot bodies, large models, touch sensing, and dexterous hands. In June 2026, an affiliated fund co-led an angel round of more than $200 million for the world-model company Wujie Dynamics. It has also connected its conversational agent JoyInside to robots, is building an embodied-data collection center, has open-sourced the human-perspective dataset EgoLive, and is turning its own malls and warehouses into robot training grounds.","example":"On July 21, 2025, Spirit AI announced a Pre-A+ funding round of nearly RMB 600 million, led by JD.com.","related":["AgiBot","LimX Dynamics","Spirit AI","EngineAI","Meituan","Data Flywheel"]},{"id":"meituan","category":"company","sec":0,"tier":2,"sources":[{"title":"美团，重注具身智能（投中网）","url":"https://www.chinaventure.com.cn/news/78-20250710-387094.html"},{"title":"美团另一面：投资中国硬科技8年，宇树科技的早期支持者（经济观察网）","url":"http://m.eeo.com.cn/2026/0604/901917.shtml"},{"title":"外卖巨头决定深耕科技投资（观点网）","url":"https://www.guandian.cn/article/20260509/559873.html"}],"as_of":"2026-06","related_ids":["unitree-robotics","galaxea-ai","x-square-robot","tars-robotics","mech-mind-robotics","jd-com"],"name":"Meituan","alt":"美团（具身智能投资）","abbr":"","aliases":["Meituan Strategic Investment","Meituan Longzhu"],"one_liner":"A Chinese local-services platform that is also one of the most active investors in embodied AI.","explanation":"Meituan is a food-delivery and local-services platform founded by Wang Xing in 2010, headquartered in Beijing. It invests in robotics through two vehicles, Meituan Strategic Investment and Meituan Longzhu: it invested in Unitree Robotics twice in 2024 and, after multiple rounds, became its second-largest shareholder; it was also an angel investor in Galbot, and has invested in X Square Robot, Galaxea AI, Mech-Mind Robotics, Pudu Robotics, and Flexiv Robotics, leading a round in TARS Robotics for the first time in July 2025. It reportedly had invested in at least 16 embodied-AI companies as of June 2026. Meituan has its own instant-delivery use case, set up the Meituan Robotics Institute in July 2022, and also works on drones and unmanned delivery vehicles.","example":"According to Tianyancha data from July 2025, Meituan-affiliated entities together held about 10.46% of Unitree Robotics, second only to founder Wang Xingxing, making it the second-largest shareholder.","related":["Unitree Robotics","Galaxea AI","X Square Robot","TARS Robotics","Mech-Mind Robotics","JD.com"]},{"id":"intrinsic","category":"company","sec":0,"tier":3,"sources":[{"title":"Intrinsic 官网","url":"https://www.intrinsic.ai/"},{"title":"Intrinsic Blog","url":"https://www.intrinsic.ai/blog"}],"as_of":"2026-09","related_ids":["everyday-robots","google-deepmind","robot-operating-system","open-robotics","fanuc","motion-planning"],"name":"Intrinsic (Alphabet)","alt":"Intrinsic","abbr":"","aliases":[],"one_liner":"An industrial-robot software company incubated at Alphabet's X, folded into Google in 2026.","explanation":"Intrinsic is a robotics software company incubated at X, Alphabet's “moonshot factory,” spun out as an independent company in 2021. It doesn't build robots; instead it builds software that makes industrial robots easier to program: its development environment, Flowstate, integrates vision-based pose estimation, collision-free motion planning, and force control; it released the Intrinsic vision model in October 2025 and open-sourced Intrinsic Core, ROS-compatible, in September 2026. In December 2022 it took on the commercial division (OSRC) of Open Robotics, the organization that maintains ROS. Its partners include NVIDIA, FANUC, and TRUMPF; in November 2025 it announced a joint venture with Foxconn to build AI-driven manufacturing solutions. In February 2026, Intrinsic announced it would be folded into Google.","example":"An engineer combines pose estimation, motion planning, and force-control modules in Flowstate to build a workstation workflow, validates it in simulation, then deploys it to a real robot arm.","related":["Everyday Robots (Alphabet X)","Google DeepMind","Robot Operating System","Open Robotics","FANUC","Motion Planning"]},{"id":"pollen-robotics","category":"company","sec":0,"tier":3,"sources":[{"title":"Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition","url":"https://huggingface.co/blog/hugging-face-pollen-robotics-acquisition"},{"title":"Reachy Mini in the Wild | Pollen Robotics","url":"https://pollen-robotics.com/reachy-mini/community"}],"as_of":"2026-05","related_ids":["hugging-face","pollen-robotics-reachy-2","pollen-robotics-reachy-mini","lerobot","open-source-hardware","vr-teleoperation"],"name":"Pollen Robotics","alt":"Pollen Robotics","abbr":"","aliases":["Pollen"],"one_liner":"A French open-source robotics company behind Reachy, acquired by Hugging Face in 2025.","explanation":"Pollen Robotics was founded in 2016 in Bordeaux, France, by Matthieu Lapeyre and Pierre Rouanet, who came out of the Flowers team at the French national research institute Inria. It builds open hardware-and-software robots: Reachy 2 is a torso-only humanoid with an omnidirectional-wheel base, two 7-degree-of-freedom arms, and support for VR teleoperation, priced around $70,000; Reachy Mini is a small, Python-programmable desktop robot priced at $299 or $449. On April 14, 2025, Hugging Face acquired the company, folding its technology into Hugging Face's open-source robotics library, LeRobot. In May 2026, Hugging Face launched an open-source app store for Reachy Mini with more than 200 apps.","example":"Researchers use LeRobot to collect teleoperated demonstrations on a Reachy 2, then train an imitation-learning policy on the data.","related":["Hugging Face","Pollen Robotics Reachy 2","Pollen Robotics Reachy Mini","LeRobot","Open-Source Hardware (OSHW)","VR Teleoperation"]},{"id":"microsoft-research","category":"company","sec":0,"tier":3,"sources":[{"title":"Rho-Alpha - Microsoft Foundry Labs","url":"https://labs.ai.azure.com/projects/rho-alpha"},{"title":"Microsoft Research Unveils Rho-Alpha Robotics Model (Pulse 2.0)","url":"https://pulse2.com/microsoft-rho-alpha-robotics-model"}],"as_of":"2026-01","related_ids":["rho-alpha","magma","florence-2","vision-language-action-model","tactile-sensor","teleoperation"],"name":"Microsoft Research","alt":"微软（微软研究院）","abbr":"MSR","aliases":["MSR"],"one_liner":"Microsoft's fundamental research arm, which released the Phi-derived robotics model Rho-alpha in 2026.","explanation":"Microsoft Research was founded in 1991, headquartered in Redmond, Washington, with labs around the world, including Microsoft Research Asia in Beijing. It has a long track record in fundamental research on computer vision and multimodal models, producing work such as the vision model Florence-2 and the multimodal agent model Magma. In January 2026 it released Rho-alpha (ρα), Microsoft's first robotics model derived from its Phi family of vision-language models (its earlier Magma model could already do some robot manipulation); Rho-alpha turns natural-language instructions into control signals for dual-arm manipulation, takes touch as an input alongside vision, and can keep learning from corrections made during human teleoperation while deployed. It is one example of a big tech company extending a general-purpose multimodal model into robot control.","example":"Rho-alpha: derived from the Phi vision-language model, it takes images, touch, and language instructions as input and outputs dual-arm manipulation actions.","related":["Rho-alpha","Magma (Microsoft)","Florence-2","Vision-Language-Action Model","Tactile Sensor","Teleoperation"]},{"id":"assured-robot-intelligence","category":"company","sec":0,"tier":3,"sources":[{"title":"Meta Accelerates Push Into Robotics Intelligence With New Acquisition (Business Insider)","url":"https://www.businessinsider.com/meta-acquires-assured-robot-intelligence-humanoid-robotics-2026-5"},{"title":"Lerrel Pinto homepage","url":"https://lerrelpinto.com"}],"as_of":"2026-05","related_ids":["whole-body-control","humanoid-robot","embodied-foundation-model","meta-fundamental-ai-research","learning-based-whole-body-control"],"name":"Assured Robot Intelligence","alt":"Assured Robot Intelligence","abbr":"ARI","aliases":["ARI"],"one_liner":"A startup building whole-body control models for humanoid robots, acquired by Meta in 2026.","explanation":"Assured Robot Intelligence (ARI) was a roughly 20-person robotics AI startup headquartered in San Diego, reportedly founded around 2025. Co-founder and CEO Lerrel Pinto is a professor at New York University, and co-founder and CTO Xiaolong Wang is an associate professor at UC San Diego who had previously done research at NVIDIA. The company's goal was to build foundation models for humanoid robots, particularly whole-body control — a single model coordinating legs, arms, and torso together. Its seed round was funded by AIX Ventures. In May 2026, Meta announced it was acquiring ARI for an undisclosed amount, folding the team into Meta Superintelligence Labs (MSL), with Pinto leading frontier robotics models and Wang serving as research director.","example":"","related":["Whole-Body Control","Humanoid Robot","Embodied Foundation Model","Meta Fundamental AI Research","Learning-Based Whole-Body Control"]},{"id":"amazon-frontier-ai-and-robotics","category":"company","sec":0,"tier":3,"sources":[{"title":"Amazon hires from AI robotics startup Covariant, licenses technology","url":"https://www.aboutamazon.com/news/company-news/amazon-covariant-ai-robots"},{"title":"Pieter Abbeel - Amazon Science","url":"https://www.amazon.science/author/pieter-abbeel"}],"as_of":"2026-09","related_ids":["covariant","amazon-robotics","holosoma","omniretarget","whole-body-control","humanoid-robot"],"name":"Amazon Frontier AI & Robotics","alt":"亚马逊前沿 AI 与机器人团队","abbr":"FAR","aliases":["Amazon FAR"],"one_liner":"Amazon's robot foundation-model team, built around the Covariant team it absorbed.","explanation":"Amazon FAR is Amazon's internal frontier robotics research team. In August 2024, Amazon licensed the technology of the AI-robotics company Covariant and hired its founders — Pieter Abbeel (a UC Berkeley professor), Peter Chen, and Rocky Duan — along with roughly a quarter of its staff; this group became the core of FAR, which researches robot foundation models with a focus on whole-body control and data generation for humanoid robots. Its results include Holosoma, a framework for training humanoid reinforcement learning in simulation and deploying it to the real robot, and OmniRetarget, a motion-retargeting method. Abbeel reportedly later moved to lead frontier-model research within Amazon's AGI organization. FAR is a separate effort from Amazon's warehouse and logistics robotics business, Amazon Robotics.","example":"Holosoma packages a humanoid robot's simulation training and real-robot deployment into one open-source pipeline.","related":["Covariant","Amazon Robotics","Holosoma (Amazon FAR humanoid RL framework)","OmniRetarget","Whole-Body Control","Humanoid Robot"]},{"id":"covariant","category":"company","sec":0,"tier":3,"sources":[{"title":"Covariant (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Covariant_(company)"}],"as_of":"2025","related_ids":["rfm-1","berkeley-artificial-intelligence-research","amazon-frontier-ai-and-robotics","bin-picking","order-picking","foundation-model"],"name":"Covariant","alt":"Covariant","abbr":"","aliases":[],"one_liner":"A Berkeley-linked warehouse-picking AI company that released RFM-1, with its core team absorbed by Amazon in 2024.","explanation":"Covariant was founded in Emeryville, California in 2017 by Pieter Abbeel, Peter Chen, Rocky Duan, and Tianhao Zhang. Abbeel heads Berkeley's Robot Learning Lab, and the other three were his students; Abbeel, Chen, and Duan had also previously done research together at OpenAI. Its product, Covariant Brain, is AI software for warehouse picking arms, using imitation learning and reinforcement learning to let a robot arm grasp a wide variety of goods out of cluttered bins. In March 2024, it used data accumulated from deployments to release the robot foundation model RFM-1. The company raised about $222 million in total funding, valued at around $625 million in 2023. In August 2024, Amazon obtained a non-exclusive license to its technology and brought on founders Abbeel, Chen, and Duan along with roughly a quarter of its staff, a so-called 'reverse acquihire'; reports say the company has seen little activity since.","example":"A robot arm running Covariant Brain picks items one by one out of mixed bins in an e-commerce warehouse and places them on a conveyor belt.","related":["RFM-1","Berkeley Artificial Intelligence Research","Amazon Frontier AI & Robotics","Bin Picking","Order Picking","Foundation Model"]},{"id":"alibaba-damo-academy","category":"company","sec":0,"tier":3,"sources":[{"title":"RynnBrain: Open Embodied Foundation Models (arXiv 2602.14979)","url":"https://arxiv.org/abs/2602.14979"},{"title":"阿里达摩院与国家人工智能应用中试基地（具身智能）达成战略合作","url":"https://damo.alibaba.com/events/32026051817790876259766230?language=zh"}],"as_of":"2026-05","related_ids":["rynnbrain","rynnvla-002","embodied-foundation-model","vision-language-action-model","qwen-vl"],"name":"Alibaba DAMO Academy","alt":"阿里达摩院","abbr":"DAMO","aliases":["DAMO Academy","DAMO"],"one_liner":"Alibaba's frontier research institute, source of the open-source RynnBrain and other embodied-AI models.","explanation":"DAMO Academy is a research institute Alibaba founded in 2017, headquartered in Hangzhou, spanning areas including chips, computer vision, and multimodal AI. In embodied AI, its main output is the Rynn family of open-source models: the RynnVLA series are vision-language-action models that turn an image and a spoken instruction directly into robot-arm actions, while RynnBrain, open-sourced in early 2026, is an embodied “brain” foundation model that emphasizes spatiotemporal memory and spatial reasoning — it understands a scene and plans tasks rather than directly driving motors. In 2026 it also formed a strategic partnership with China's National Pilot Base for Artificial Intelligence Applications (embodied intelligence). Newcomers reading embodied-AI papers will often see it listed among baselines or open-source models.","example":"RynnBrain can watch a first-person video from a robot, answer where an object was placed earlier, and then plan the next action.","related":["RynnBrain","RynnVLA-002","Embodied Foundation Model","Vision-Language-Action Model","Qwen-VL"]},{"id":"physical-intelligence","category":"company","sec":1,"tier":1,"sources":[{"title":"Physical Intelligence Inc. - Wikipedia","url":"https://en.wikipedia.org/wiki/Physical_Intelligence_Inc."},{"title":"π0.7 - Physical Intelligence","url":"https://www.pi.website/blog/pi07"},{"title":"π*0.6 - Physical Intelligence","url":"https://www.pi.website/blog/pistar06"}],"as_of":"2026-04","related_ids":["pi0","pi0-5","pi-star-0-6","pi0-7","openpi","robot-brain-company"],"name":"Physical Intelligence","alt":"Physical Intelligence","abbr":"PI","aliases":["π","Physical Intelligence (physical-AI company)"],"one_liner":"An American startup that builds only robot 'brains,' maker of the π series of VLAs.","explanation":"Physical Intelligence was founded in San Francisco in 2024. Co-founders include former Google robotics researchers Karol Hausman and Brian Ichter, Berkeley professor Sergey Levine, Stanford professor Chelsea Finn, and investor Lachy Groom, among others. The company builds no robot hardware of its own, only general-purpose models that can control many different robots: π0 (October 2024, using a flow-matching action expert), π0-FAST, π0.5 (open-world generalization), π*0.6 (using RECAP reinforcement learning), and π0.7 in April 2026, with some weights and code open-sourced through openpi. Funding: $400 million in 2024 (valuing the company at $2.4 billion), $600 million in 2025 ($5.6 billion valuation), and $1 billion in 2026 ($11 billion valuation).","example":"Many labs take the π0 weights from openpi and fine-tune them on their own robot arm using just tens of hours of data.","related":["π0","π0.5","π*0.6","π0.7","openpi (Physical Intelligence)","Robot-Brain (Model-Only) Company"]},{"id":"skild-ai","category":"company","sec":1,"tier":1,"sources":[{"title":"Skild AI Series C","url":"https://www.skild.ai/blogs/series-c"},{"title":"Skild AI | LinkedIn","url":"https://www.linkedin.com/company/skild-ai"},{"title":"Skild S1","url":"https://www.skild.ai/blogs/s1"}],"as_of":"2026-08","related_ids":["skild-brain","skild-s1",null,null,null,null],"name":"Skild AI","alt":"Skild AI","abbr":"","aliases":["Skild"],"one_liner":"A US company building a cross-robot general foundation model, Skild Brain, without making its own robot hardware.","explanation":"Skild AI is a US company founded in 2023, headquartered in Pittsburgh with an additional team in the San Francisco Bay Area. Its founders, Deepak Pathak and Abhinav Gupta, are both robotics professors at Carnegie Mellon University. The company builds only the “brain,” not the robot body: it introduced Skild Brain in July 2025, a single model that can control quadrupeds, humanoids, robot arms, and mobile manipulators, trained mainly on large-scale simulation and internet video; in August 2026 it released S1, which can perform a new task after watching just one demonstration video. In January 2026, it raised $1.4 billion in funding led by SoftBank with participation from NVIDIA and others, at a valuation of more than $14 billion; the company says it generated about $30 million in revenue over a few months in 2025, already deployed in security, warehousing, and factory-assembly settings.","example":"In a Skild Brain demo, a quadruped robot that lost part of one leg adjusted its gait within seconds and kept walking.","related":["Skild Brain","Skild S1","Cross-Embodiment","Robot-Brain (Model-only) Company","One Brain, Multiple Robots","In-Context Learning"]},{"id":"generalist-ai","category":"company","sec":1,"tier":2,"sources":[{"title":"Generalist AI Blog","url":"https://generalistai.com/blog"},{"title":"Generalist AI - About","url":"https://generalistai.com/about"},{"title":"Generalist | LinkedIn","url":"https://www.linkedin.com/company/generalistai"}],"as_of":"2026-09","related_ids":["gen-0","gen-1",null,"scaling-law","ossification","real-robot-data"],"name":"Generalist AI","alt":"Generalist AI","abbr":"","aliases":["Generalist"],"one_liner":"A US embodied-AI company that trains its GEN-series foundation models on huge amounts of real-robot interaction data.","explanation":"Generalist AI is a US robotics AI company founded in 2024, based in the San Francisco Bay Area (San Mateo, California), with a second team in Boston. CEO Pete Florence previously worked on PaLM-E and RT-2 at Google DeepMind, and CTO Andrew Barry came from Boston Dynamics. Its approach is to collect large amounts of real-world physical-interaction data itself and use it to train embodied foundation models. GEN-0, released in November 2025, was trained on more than 270,000 hours of real-robot data and was the company's first reported scaling-law result on a robot. GEN-1 (April 2026) focused on performing simple tasks skillfully and quickly. GEN-1.5 (August 2026) can learn a new task after watching a demonstration only about a dozen seconds long. According to its LinkedIn page, its most recent funding round raised $400 million, bringing its total funding to more than $500 million.","example":"Fine-tuned on only about an hour of robot data, GEN-1 reportedly reached a 99% success rate on tasks like folding a paper box.","related":["GEN-0","GEN-1","Embodied Foundation Model","Scaling Law","Ossification","Real-Robot Data"]},{"id":"genesis-ai","category":"company","sec":1,"tier":2,"sources":[{"title":"Khosla-backed robotics startup Genesis AI has gone full stack, demo shows (TechCrunch, 2026-05-06)","url":"https://techcrunch.com/2026/05/06/khosla-backed-robotics-startup-genesis-ai-has-gone-full-stack-demo-shows/"},{"title":"Genesis AI 官网","url":"https://www.genesis.ai/"}],"as_of":"2026-05","related_ids":[null,null,null,null,null,null],"name":"Genesis AI","alt":"Genesis AI","abbr":"","aliases":["Genesis AI SAS"],"one_liner":"A robotics company that grew out of the Genesis simulator, building dexterous manipulation with matched hardware and software.","explanation":"Genesis AI came out of stealth in July 2025 with a $105 million seed round led by Eclipse and Khosla Ventures, with offices in Paris and California, later expanding to London. Its CEO, Xian Zhou, is the lead originator of the open-source physics simulator Genesis; its president, Théophile Gervet, was previously a researcher at Mistral AI. The company takes a full-stack approach: it partnered with Wuji Robotics to build a dexterous hand close to the size of a human hand, paired with sensor-equipped gloves to capture human manipulation data, and built its own Genesis World simulator for training and evaluation. In May 2026 it released the robotics foundation model GENE-26.5, demonstrating dexterous manipulation such as cooking, playing piano, and solving a Rubik's Cube. Its website is currently collecting a waitlist for a home/commercial robot called Eno, with targeted customer deployments planned to begin in late 2026.","example":"In the GENE-26.5 release demo, a dexterous hand trained on human glove data plus a small amount of robot data solved a Rubik's Cube.","related":["Genesis Simulator","GENE-26.5 (Genesis AI robotics foundation model)","Wuji Hand","Dexterous Manipulation","Data Glove","Hardware-Software Integration"]},{"id":"dyna-robotics","category":"company","sec":1,"tier":3,"sources":[{"title":"Yahoo Finance: Dyna Robotics Raises $23.5 Million","url":"https://finance.yahoo.com/news/dyna-robotics-raises-23-5-130000239.html"},{"title":"Bloomberg: Dyna Robotics Raises $120 Million in Funding From Nvidia, Amazon","url":"https://www.bloomberg.com/news/articles/2025-09-15/dyna-robotics-raises-120-million-in-funding-from-nvidia-amazon"},{"title":"Dyna: DYNA-2","url":"https://www.dyna.co/dyna-2"}],"as_of":"2026-08","related_ids":["dyna-1","dyna-2","physical-intelligence","world-action-model","reward-model","data-flywheel"],"name":"Dyna Robotics","alt":"Dyna Robotics","abbr":"","aliases":["Dyna","DYNA Robotics"],"one_liner":"A US embodied-foundation-model startup that gets robots running long, reliable shifts in real commercial settings.","explanation":"Dyna Robotics was founded in 2024, headquartered in Redwood City, California. Co-founders Lindon Gao and York Yang had previously started Caper AI, a cashier-free smart shopping-cart company that sold to Instacart for $350 million; another co-founder, Jason Ma, was previously a research scientist at DeepMind. Its approach is to first get a robot to do one thing very reliably in a real commercial setting, such as folding towels or prepping food, then keep collecting data while deployed and use it to iterate on a more general model. It raised a $23.5 million seed round in March 2025 and a $120 million Series A in September, with participation from NVIDIA and Amazon; Bloomberg reported a valuation above $600 million. On models, it released DYNA-1 in June 2025, followed in August 2026 by DYNA-2, a world-action model pretrained on more than a million hours of first-person human video.","example":"The company reports that DYNA-1 ran a restaurant napkin-folding task continuously for 24 hours with no human intervention, at a 99.4% success rate.","related":["DYNA-1","DYNA-2","Physical Intelligence","World Action Model","Reward Model","Data Flywheel"]},{"id":"field-ai","category":"company","sec":1,"tier":3,"sources":[{"title":"FieldAI News","url":"https://www.fieldai.com/news"}],"as_of":"2026-09","related_ids":[null,null,null,null,null],"name":"Field AI","alt":"Field AI","abbr":"","aliases":["FieldAI"],"one_liner":"A US startup building a general-purpose “brain” for robots working in the field and on job sites.","explanation":"Field AI is a US robotics software company headquartered in Irvine, California; its founder and CEO, Ali Agha, previously led a team at NASA's Jet Propulsion Laboratory competing in the DARPA Subterranean Challenge. It doesn't build robot bodies; instead it builds autonomy software that can be installed on robots of different shapes, centered on FFM (Field Foundation Models), which combine neural networks with physics-based reasoning and uncertainty estimation, letting a robot move and work autonomously on job sites, mines, and other field environments without pre-built maps or GPS, running entirely on onboard compute. The company reportedly raised about $400 million in total during 2025, with investors including Bezos Expeditions, Khosla Ventures, and NVIDIA. In 2026 it announced partnerships with Boston Dynamics, NVIDIA, and Caterpillar in succession.","example":"Field AI's software can be installed on a quadruped or wheeled base to autonomously patrol a construction site and log progress.","related":["Field Foundation Models (FieldAI)","Boston Dynamics","One Brain, Multiple Robots","Robot-Brain (Model-only) Company","Navigation"]},{"id":"flexion-robotics","category":"company","sec":1,"tier":3,"sources":[{"title":"Flexion Robotics - About","url":"https://flexion.ai/about"},{"title":"Flexion Robotics 官网","url":"https://flexion.ai/"}],"as_of":"2026-07","related_ids":[null,"legged-gym",null,null,null],"name":"Flexion Robotics","alt":"Flexion","abbr":"","aliases":["Flexion Robotics AG","Flexion AI"],"one_liner":"An ETH Zurich spin-off building an autonomy software stack for humanoid robots.","explanation":"Flexion Robotics is a humanoid-robot software company headquartered in Zurich with an additional office in San Francisco. Co-founder and CEO Nikita Rudin did his PhD at ETH Zurich and later worked as an NVIDIA researcher, and is the lead author of legged_gym; CTO David Höller also came out of ETH and NVIDIA; and ETH Robotic Systems Lab professor Marco Hutter is a co-founder and advisor. The company doesn't build its own robot bodies; instead it aims to build a hardware-agnostic “autonomy stack” for humanoids, covering everything from understanding instructions and planning down to the low-level control for manipulation and walking, with an emphasis on training with reinforcement learning in simulation and transferring to the real robot. It released Reflect v0 in November 2025 and Reflect v1.0, aimed at long-horizon autonomous work, in June 2026; in July it announced sim-to-real research partnerships with Niantic Spatial and NVIDIA.","example":"","related":["ETH Zurich Robotic Systems Lab","legged_gym","Sim-to-Real Transfer","Robot-Brain (Model-only) Company","Humanoid Robot"]},{"id":"rhoda-ai","category":"company","sec":1,"tier":3,"sources":[{"title":"Rhoda AI 官网","url":"https://www.rhoda.ai/"}],"as_of":"2026-09","related_ids":["video-generation-model","world-action-model","vision-language-action-model","internet-video-data","post-training"],"name":"Rhoda AI","alt":"Rhoda AI","abbr":"","aliases":[],"one_liner":"A US embodied-AI startup that pre-trains robot policies on internet video, aimed at industrial settings.","explanation":"Rhoda AI is a US embodied-AI startup building robot foundation models for industrial sites. Its website describes a core capability called FutureVision and a “Direct Video Action” model: it first pre-trains on large-scale internet video so the model learns to predict how a scene will change, then fine-tunes on just 1 to 10 hours of robot trajectory data to get a policy that outputs actions directly — which the company says generalizes better than a conventional vision-language-action (VLA) pipeline. It also builds its own wheeled robot platform, with demonstrations of tasks such as returns processing, bearing kitting, and unboxing across automotive, manufacturing, logistics, and e-commerce settings. Investors listed on its website include Khosla Ventures, John Doerr, Temasek, and Samsung NEXT; the founding team, founding year, and funding amount are not disclosed on the site and are omitted here.","example":"","related":["Video Generation Model","World Action Model","Vision-Language-Action Model","Internet Video Data","Post-training"]},{"id":"sunday-robotics","category":"company","sec":1,"tier":3,"sources":[{"title":"Sunday Raises $165M to Launch First Autonomous Robots by Thanksgiving（GlobeNewswire，2026-03-12）","url":"https://www.globenewswire.com/news-release/2026/03/12/3254877/0/en/sunday-raises-165m-to-launch-first-autonomous-robots-by-thanksgiving.html"},{"title":"Sunday Robotics 官网","url":"https://www.sunday.ai/"},{"title":"ACT-2 Preview: Generalizing Reliability（Sunday blog，2026-07-17）","url":"https://www.sunday.ai/blog/act-2-preview"}],"as_of":"2026-07","related_ids":["sunday-robotics-memo","sunday-robotics-act-1","skill-capture-glove","robot-free-data-collection","action-chunking-with-transformers","diffusion-policy"],"name":"Sunday Robotics","alt":"Sunday Robotics","abbr":"","aliases":["Sunday"],"one_liner":"A US home-robotics company founded by the creators of ALOHA and Diffusion Policy, building the Memo robot.","explanation":"Sunday Robotics, usually known simply as Sunday, is headquartered in Mountain View, California, and was founded by two Stanford PhDs: CEO Tony Zhao, an author of ACT and ALOHA, and CTO Cheng Chi, an author of Diffusion Policy and UMI. The company came out of stealth in November 2025, simultaneously unveiling the household robot Memo, the foundation model ACT-1, and a “skill-capture glove.” Its core approach skips robot teleoperation for data collection: instead, a person wears a glove structurally matched to the robot's hand and does real household chores, and that data is converted into robot-usable training data. In March 2026 it closed a $165 million Series B led by Coatue at a $1.15 billion valuation; it aims to deliver its first test units to 50 households before Thanksgiving 2026, and previewed ACT-2 that July.","example":"Memo uses a single ACT-1 model to complete long-horizon tasks such as clearing an entire dinner table and loading the dishwasher, including in homes it has never visited before.","related":["Sunday Robotics Memo","Sunday Robotics ACT-1","Skill Capture Glove","Robot-free (Embodiment-free) Data Collection","Action Chunking with Transformers","Diffusion Policy"]},{"id":"the-bot-company","category":"company","sec":1,"tier":3,"sources":[{"title":"Reuters: Former Cruise CEO Vogt's robotics startup valued at $2 billion","url":"https://www.reuters.com/technology/former-cruise-ceo-vogts-robotics-startup-valued-2-billion-new-funding-sources-2025-03-21"},{"title":"Bloomberg: Cruise Founder Kyle Vogt's Robotics Startup Eyes $4 Billion Valuation","url":"https://www.bloomberg.com/news/articles/2025-10-28/cruise-founder-kyle-vogt-s-robotics-startup-eyes-4-billion-valuation"}],"as_of":"2025-10","related_ids":[null,"household-tasks","mobile-manipulation","autonomous-driving-talent-moving-into-embodied-ai","sunday-robotics",null],"name":"The Bot Company","alt":"The Bot Company","abbr":"","aliases":["Bot Company"],"one_liner":"A US startup founded by Cruise's Kyle Vogt to build household robots.","explanation":"The Bot Company was founded in San Francisco in 2024 by Kyle Vogt, who had earlier co-founded the livestreaming platform Justin.tv (the precursor to Twitch) and the self-driving company Cruise, where he served as CEO; co-founders Paril Jain and Luke Holoubek came from Tesla and Cruise. The company's goal is a robot that can tidy up and do everyday household chores, combining a mobile base, robot arms, and large models — a classic case of a team moving from self-driving into embodied AI. Product details have stayed largely under wraps, but fundraising has moved fast: a $150 million round led by Greenoaks in March 2025 at a $2 billion valuation, and, per an October 2025 Bloomberg report, a further $250 million raised at a valuation above $4 billion.","example":"","related":["Kyle Vogt","Household Tasks","Mobile Manipulation","Autonomous-Driving Talent Moving into Embodied AI","Sunday Robotics","1X (1X Technologies)"]},{"id":"mind-robotics","category":"company","sec":1,"tier":3,"sources":[{"title":"Rivian spinout Mind Robotics valued at $3.4 billion in new funding round (Reuters)","url":"https://www.reuters.com/legal/transactional/rivian-spinout-mind-robotics-valued-34-billion-new-funding-round-2026-05-13"},{"title":"Mind Robotics raises Series A to develop AI-driven industrial automation (The Robot Report)","url":"https://www.therobotreport.com/mind-robotics-raises-series-a-develop-ai-driven-industrial-automation"},{"title":"Rivian spinoff Mind Robotics raises another $400M (TechCrunch)","url":"https://techcrunch.com/2026/05/13/rivian-spinoff-mind-robotics-raises-another-400m"}],"as_of":"2026-05","related_ids":["industrial-robot","physical-ai","automakers-entering-humanoid-robotics","brownfield-deployment","funding-rounds-and-valuation"],"name":"Mind Robotics","alt":"Mind Robotics","abbr":"","aliases":[],"one_liner":"A US factory-automation AI robotics startup spun out of the EV maker Rivian.","explanation":"Mind Robotics is an independent company Rivian, the US electric-vehicle maker, spun out in November 2025; it began as an internal project called “Project Synapse” and was founded by Rivian founder and CEO RJ Scaringe. Its goal is to develop AI-driven industrial robots that make factories run more efficiently. Its funding has moved quickly: a $115 million seed round led by Eclipse in November 2025; a $500 million Series A co-led by Accel and a16z in March 2026 at a $2 billion valuation; and a further $400 million led by Kleiner Perkins in May 2026, reportedly at a $3.4 billion valuation, for more than $1 billion raised in total. It is a representative example of an automaker extending its manufacturing expertise into embodied AI.","example":"In May 2026, Kleiner Perkins led a $400 million round, reportedly pushing the company's valuation up to $3.4 billion.","related":["Industrial Robot","Physical AI","Automakers Entering Humanoid Robotics","Brownfield Deployment","Funding Rounds & Valuation"]},{"id":"project-prometheus","category":"company","sec":1,"tier":3,"sources":[{"title":"Project Prometheus (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Project_Prometheus_(company)"}],"as_of":"2026-08","related_ids":[null,null,"world-model","generalist-ai","skild-ai"],"name":"Project Prometheus","alt":"普罗米修斯计划","abbr":"","aliases":["Prometheus","Prometheus Industries"],"one_liner":"A Bezos-backed AI company building physical AI for engineering and manufacturing.","explanation":"Project Prometheus became public in November 2025 as a US AI startup co-led by Amazon founder Jeff Bezos and Vik Bajaj, a chemist and physicist formerly of Google X, who serve as co-CEOs; startup funding was $6.2 billion, partly from Bezos himself. It is headquartered in San Francisco, with offices in London and Zurich. The company wants AI to learn through trial and error in the real physical world rather than relying only on digital data, applying this to engineering and manufacturing in computing, aerospace, and automotive industries. In November 2025 it acquired the agent startup General Agents, and by December 2025 it had more than 120 employees. In 2026 it renamed itself Prometheus, leased a former steel warehouse in West Oakland that August, and, according to reports, is also raising money for a holding company that acquires traditional industrial businesses.","example":"","related":["Physical AI","Embodied AI","World Model","Generalist AI","Skild AI"]},{"id":"bedrock-robotics","category":"company","sec":1,"tier":3,"sources":[{"title":"Bedrock Robotics brings in $80M for construction retrofit kits (The Robot Report)","url":"https://www.therobotreport.com/bedrock-robotics-brings-in-80m-for-construction-retrofit-kits"},{"title":"Bedrock Robotics raises $270M for autonomous machine solution","url":"https://theconstructionbroadsheet.com/bedrock-robotics-raises-m-for-autonomous-machine-solution-p2178-176.htm"}],"as_of":"2026-02","related_ids":["autonomous-driving","lidar","physical-ai","autonomous-driving-talent-moving-into-embodied-ai","real-world-deployment"],"name":"Bedrock Robotics","alt":"Bedrock Robotics","abbr":"","aliases":[],"one_liner":"A US company retrofitting excavators and other heavy machinery with self-driving kits.","explanation":"Bedrock Robotics was founded in San Francisco in May 2024 by three former Waymo leaders and a former Segment engineering lead; CEO Boris Sofman previously led Waymo's self-driving trucks program and had earlier co-founded the consumer robotics company Anki. Rather than build new machines, it retrofits existing excavators and loaders with a kit of lidar, 360-degree cameras, and a compute unit, letting the heavy equipment operate autonomously and easing labor shortages in construction. It came out of stealth in July 2025 with $80 million in funding, then closed a $270 million Series B in February 2026, co-led by CapitalG and the Valor Atreides AI Fund. It is a representative example of autonomous-driving talent moving into physical AI.","example":"A retrofitted excavator autonomously digs a building's foundation on a job site while a person only supervises remotely.","related":["Autonomous Driving","LiDAR","Physical AI","Autonomous-Driving Talent Moving into Embodied AI","Real-world Deployment"]},{"id":"wayve","category":"company","sec":1,"tier":3,"sources":[{"title":"Wayve - Wikipedia","url":"https://en.wikipedia.org/wiki/Wayve"},{"title":"Wayve raises $1.2B at $8.6B valuation to scale embodied AI for autonomous driving（Tech.eu）","url":"https://tech.eu/2026/02/25/wayve-raises-12b-at-86b-valuation-to-scale-embodied-ai-for-autonomous-driving/"},{"title":"Wayve and Uber Launch First-Ever Autonomous Rides in the UK（Wayve）","url":"https://wayve.ai/press/wayve-uber-launch-autonomous-rides/"}],"as_of":"2026-09","related_ids":["gaia-2","autonomous-driving","end-to-end","world-model","vision-language-action-model","tesla-fsd-v12"],"name":"Wayve","alt":"Wayve","abbr":"","aliases":["Wayve Technologies"],"one_liner":"A London end-to-end autonomous-driving company that calls its own work embodied AI and now runs robotaxis with Uber.","explanation":"Wayve was founded in Cambridge in 2017 by Cambridge machine-learning PhD student Alex Kendall, now CEO, and Amar Shah, and is now headquartered in London. It follows an end-to-end approach it calls AV2.0: a single neural network outputs driving actions directly from camera footage, without relying on high-definition maps or hand-written rules, and the company describes its own work as embodied AI. Notable models include LINGO-2, which explains its driving decisions in language as it drives, and the GAIA family of driving-video world models. In May 2024 it closed a $1.05 billion Series C led by SoftBank; in February 2026 it closed a $1.2 billion Series D at an $8.6 billion post-money valuation, with NVIDIA, Microsoft, Uber, Mercedes-Benz, Nissan, and Stellantis among the investors. On September 3, 2026, it launched a safety-driver-monitored robotaxi service with Uber in London.","example":"While driving, LINGO-2 narrates its reasoning in real time, saying things like “there's a pedestrian crossing ahead, I'm slowing down.”","related":["GAIA-2 (Wayve)","Autonomous Driving","End-to-End","World Model","Vision-Language-Action Model","Tesla FSD v12"]},{"id":"figure-ai","category":"company","sec":2,"tier":1,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"Figure: Helix 2.5","url":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"}],"as_of":"2026-09","related_ids":["figure-03","figure-helix","figure-helix-02","helix-2-5","botq","humanoid-robot"],"name":"Figure AI","alt":"Figure AI","abbr":"","aliases":["Figure"],"one_liner":"An American humanoid robot company, maker of the Figure robot series and the Helix model.","explanation":"Figure AI was founded in 2022 by Brett Adcock, who previously founded electric-aircraft company Archer Aviation and hiring platform Vettery; the company is headquartered in San Jose, California. It builds both humanoid robot hardware (Figure 01, 02, 03) and the robot 'brain': its in-house vision-language-action model Helix (February 2025), followed by the whole-body-control Helix 02 and, in September 2026, Helix 2.5; it also built its own BotQ factory for mass production. In January 2024 it began an in-plant pilot with BMW; a model partnership with OpenAI ended after about a year. On funding, it raised a $675 million Series B in February 2024 and a Series C of more than $1 billion in September 2025 at a post-money valuation of $39 billion. According to reports, by April 2026 it had delivered more than 350 units of Figure 03.","example":"Helix 2.5, running on Figure 03, completed zero-shot tasks such as tidying the living room, folding towels, and making the bed in 30 homes it had never visited before.","related":["Figure 03","Figure Helix","Figure Helix 02","Helix 2.5","BotQ","Humanoid Robot"]},{"id":"boston-dynamics","category":"company","sec":2,"tier":1,"sources":[{"title":"Boston Dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Boston_Dynamics"},{"title":"CES 2026: Boston Dynamics unveils new Atlas","url":"https://www.robotics247.com/article/ces-2026-boston-dynamics-unveils-new-atlas-humanoid-robot/technologies"}],"as_of":"2026-02","related_ids":["boston-dynamics-atlas-2","boston-dynamics-atlas","boston-dynamics-spot","boston-dynamics-stretch","hyundai-motor-group","google-deepmind"],"name":"Boston Dynamics","alt":"波士顿动力","abbr":"BD","aliases":["波动 (colloquial Chinese nickname)"],"one_liner":"A veteran legged-robot company, maker of the Spot robot dog and the Atlas humanoid.","explanation":"Boston Dynamics was founded in 1992 by Marc Raibert, growing out of his Leg Laboratory at MIT, and is headquartered in Waltham, Massachusetts. Google acquired it in 2013, it moved to SoftBank in 2017, and in December 2020 Hyundai Motor Group agreed to buy an 80% stake for about $880 million, a deal that closed in June 2021. The company has long relied on model-based control to achieve highly dynamic legged locomotion, with landmark robots including BigDog, the quadruped Spot (commercialized starting 2019, sold to the public from 2020), the warehouse unloading robot Stretch, and the humanoid Atlas. The hydraulic version of Atlas was retired in April 2024 and replaced by an electric version; a production version of Atlas was unveiled at CES in January 2026, alongside a partnership with Google DeepMind to train foundation models for it. In February 2026, CEO Robert Playter retired and CFO Amanda McMaster became interim CEO.","example":"Spot, commonly used for factory inspection, carries a thermal camera and a robotic arm on its back, and was one of the first quadruped robots to be commercialized at scale.","related":["Boston Dynamics Atlas (Electric)","Boston Dynamics Atlas (Hydraulic)","Boston Dynamics Spot","Boston Dynamics Stretch","Hyundai Motor Group","Google DeepMind"]},{"id":"1x-technologies","category":"company","sec":2,"tier":1,"sources":[{"title":"1X Technologies - Wikipedia","url":"https://en.wikipedia.org/wiki/1X_Technologies"}],"as_of":"2025-10","related_ids":["1x-neo","1x-eve","redwood","1x-world-model","remote-teleoperation-takeover","openai"],"name":"1X Technologies","alt":"1X","abbr":"","aliases":["Halodi Robotics (former name)"],"one_liner":"A Norway-founded humanoid robot company, best known for its home humanoid NEO.","explanation":"1X was founded in Norway in 2014 by Bernt Øivind Børnich under the name Halodi Robotics, renamed 1X in 2022; it is now headquartered in Palo Alto, California, while still keeping a factory in Moss, Norway. Its first product, EVE, paired a wheeled base with a humanoid upper body, used for security and logistics; the company then shifted to the bipedal home humanoid NEO, unveiling the Beta in August 2024 and the Gamma in February 2025, and opening preorders on October 28, 2025, priced at a $20,000 outright purchase or $499 a month, with deliveries planned for 2026. It builds its own VLA model, Redwood, and a video-generation-based 1X World Model. Funding: a $23.5 million Series A2 in March 2023 (led by the OpenAI Startup Fund) and a $100 million Series B in January 2024; reports say it sought up to $1 billion in new funding in September 2025.","example":"Household tasks NEO cannot handle can be taken over remotely by a 1X employee through teleoperation, which has also raised privacy concerns.","related":["1X NEO","1X EVE","Redwood","1X World Model","Remote Teleoperation Takeover","OpenAI"]},{"id":"agility-robotics","category":"company","sec":2,"tier":1,"sources":[{"title":"Agility Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Agility_Robotics"},{"title":"Agility Robotics plans to go public via SPAC in a $2.5B deal (TechCrunch)","url":"https://techcrunch.com/2026/06/24/agility-robotics-plans-to-go-public-via-spac-in-a-2-5b-deal/"}],"as_of":"2026-09","related_ids":["agility-robotics-digit","agility-robotics-cassie","robofab","robot-as-a-service","special-purpose-acquisition-company","humanoid-robot"],"name":"Agility Robotics","alt":"Agility Robotics","abbr":"","aliases":["Agility"],"one_liner":"An American bipedal humanoid robot company whose Digit robot is already commercially deployed in logistics warehouses.","explanation":"Agility Robotics was spun out of Oregon State University's Dynamic Robotics Laboratory in 2015 and is headquartered in Salem, Oregon; co-founder Jonathan Hurst founded that lab and has long studied dynamic bipedal walking, and the current CEO is Peggy Johnson. Its early product was Cassie, a legs-only research platform; a torso and arms were later added to create the humanoid Digit, mainly used to move totes in warehouses, billed under a 'robot-as-a-service' model to customers including GXO, Schaeffler, a Toyota plant in Canada, and Mercado Libre. In 2023 the company announced its own humanoid factory, RoboFab, in Salem. In March 2026 the brand was shortened to Agility; in June it announced a merger with Churchill Capital Corp XI to go public via a SPAC at a deal valuation of about $2.5 billion, planning to trade under the ticker AGLT; in September it released Digit 5, its fifth generation, built around cage-free collaborative safety.","example":"In 2024, GXO signed the first robot-as-a-service contract with Agility, using Digit to move totes at a Spanx warehouse in Georgia.","related":["Agility Robotics Digit","Agility Robotics Cassie","RoboFab","Robot-as-a-Service","Special Purpose Acquisition Company","Humanoid Robot"]},{"id":"apptronik","category":"company","sec":2,"tier":2,"sources":[{"title":"Humanoid robot startup Apptronik has now raised $935M at a $5B+ valuation (TechCrunch)","url":"https://techcrunch.com/2026/02/11/humanoid-robot-startup-apptronik-has-now-raised-935m-at-a-5b-valuation/"},{"title":"Apptronik 官网","url":"https://apptronik.com/"}],"as_of":"2026-06","related_ids":["apptronik-apollo","google-deepmind","gemini-robotics","nasa-valkyrie","darpa-robotics-challenge","humanoid-robot"],"name":"Apptronik","alt":"Apptronik","abbr":"","aliases":[],"one_liner":"An Austin, Texas humanoid robot company, maker of Apollo, partnered with Google DeepMind.","explanation":"Apptronik was founded in Austin, Texas in 2016. Founders Jeff Cardenas and Nick Paine came out of the University of Texas at Austin's Human Centered Robotics Lab, whose members had competed in the DARPA Robotics Challenge using NASA's Valkyrie humanoid. The company builds its own electric actuators and unveiled its humanoid, Apollo, in August 2023, targeting repetitive physical work in factories and warehouses such as moving boxes, machine tending, and material feeding; Mercedes-Benz has been piloting it since 2024, alongside partners GXO, Jabil, and NVIDIA. Apptronik partners with Google DeepMind, running Gemini Robotics models on Apollo. On funding, it raised a $350 million Series A in February 2025 (with Google participating), followed by several top-ups; by February 2026 the Series A total reached $935 million at a roughly $5.3 billion post-money valuation. In June 2026 it launched Apollo 2, available with either a bipedal or wheeled base.","example":"In a Mercedes-Benz factory pilot, Apollo delivers parts boxes to the assembly line.","related":["Apptronik Apollo","Google DeepMind","Gemini Robotics","NASA Valkyrie (R5)","DARPA Robotics Challenge","Humanoid Robot"]},{"id":"neura-robotics","category":"company","sec":2,"tier":2,"sources":[{"title":"NEURA Robotics Announces Record Series C of up to $1.4B","url":"https://neura-robotics.com/record-series-c"},{"title":"Humanoid robotics company Neura Robotics backed by Amazon, Nvidia (CNBC)","url":"https://www.cnbc.com/2026/06/10/neura-robotics-funding-ai-humanoid-robots.html"},{"title":"NEURA Robotics Secures €120 Million in Series B Funding","url":"https://neura-robotics.com/neura-robotics-secures-euro-120-million-series-b"}],"as_of":"2026-06","related_ids":["neura-robotics-4ne1","collaborative-robot","humanoid-robot","physical-ai","nvidia"],"name":"NEURA Robotics","alt":"NEURA Robotics","abbr":"","aliases":["NEURA"],"one_liner":"A German “cognitive robotics” company making collaborative arms and the 4NE1 humanoid robot.","explanation":"NEURA Robotics was founded in Metzingen, Germany in 2019; its founder and CEO is David Reger. It focuses on “cognitive robots” — robots with built-in vision, force sensing, and AI capability — with products including the MAiRA and LARA collaborative arms and the 4NE1 humanoid robot, alongside a software and data platform that lets customers train and share robot skills. It previously closed a €120 million Series B led by Lingotto; in 2026 it announced a Series C of up to $1.4 billion, with investors including Tether, Qualcomm, Amazon, and NVIDIA, and CNBC reported in June 2026 that it was valued at about $7 billion, with part of the funding tied to performance milestones.","example":"The full-size humanoid robot 4NE1.","related":["NEURA Robotics 4NE1","Collaborative Robot","Humanoid Robot","Physical AI","NVIDIA"]},{"id":"sanctuary-ai","category":"company","sec":2,"tier":3,"sources":[{"title":"Sanctuary AI Unveils Phoenix","url":"https://sanctuary.ai/news/sanctuary-ai-unveils-phoenix-a-humanoid-general-purpose-robot-designed-for-work"},{"title":"Sanctuary AI Phoenix 2026 review (RoboZaps)","url":"https://blog.robozaps.com/b/sanctuary-ai-phoenix-review"}],"as_of":"2026-06","related_ids":["sanctuary-ai-phoenix","humanoid-robot","dexterous-hand","hydraulic-actuation","teleoperation"],"name":"Sanctuary AI","alt":"Sanctuary AI","abbr":"","aliases":[],"one_liner":"A Canadian humanoid-robot company building the Phoenix humanoid and its Carbon control system.","explanation":"Sanctuary AI was founded in Vancouver, Canada, in 2018 by co-founders including Geordie Rose and Suzanne Gildert, who also founded the quantum-computing company D-Wave. Its goal is a general-purpose humanoid robot; its product line, Phoenix, released its 6th generation in May 2023 and its 7th and 8th generations in 2024, distinguished by hydraulically driven, highly articulated dexterous hands, paired with an AI control system called Carbon. Early on, the company relied mainly on teleoperation to collect human demonstrations for training. Investors have included Bell, Magna, and the Canadian government, with cumulative funding over CA$140 million by 2024. In November 2024, Rose was replaced by the board amid layoffs; the company reportedly pivoted in 2026 toward providing physical-AI software for existing industrial robots rather than focusing on complete humanoids.","example":"The 8th-generation Phoenix (December 2024) switched to a wheeled base, used mainly to collect manipulation data.","related":["Sanctuary AI Phoenix","Humanoid Robot","Dexterous Hand","Hydraulic Actuation","Teleoperation"]},{"id":"foundation-robotics","category":"company","sec":2,"tier":3,"sources":[{"title":"Foundation 官网","url":"https://foundation.bot/"},{"title":"Foundation Emerges With Phantom Humanoid - Humanoids Daily","url":"https://www.humanoidsdaily.com/news/foundation-emerges-with-phantom-humanoid-betting-on-novel-actuators-and-hybrid-ai"}],"as_of":"2026-07","related_ids":[null,null,null,null,null],"name":"Foundation Robotics","alt":"Foundation","abbr":"","aliases":["Foundation Robotics Labs","Foundation Future Industries"],"one_liner":"A US humanoid-robot startup focused on defense and industrial use, maker of the Phantom robot.","explanation":"This is a humanoid-robot company based in San Francisco, operating under the legal name Foundation Future Industries; its CEO, Sankaet Pathak, previously founded the fintech company Synapse. It was reportedly founded in 2023 and also has a site in Munich. Its product is the full-size humanoid robot Phantom, built with in-house actuators including a rolling-contact reducer; it relies only on cameras for perception, with no lidar, and runs an AI system called Cortex. Unlike most humanoid companies, it openly emphasizes defense as a focus, reportedly winning a Pentagon research contract worth about $24 million; its website showed demos in 2026 of load-carrying and mortar operation, while also piloting the robot in consumer-goods, beverage, and glass-manufacturing factories. The company has stated a goal of deploying tens of thousands of units in 2026, though that so far remains only a plan.","example":"","related":["Foundation Robotics Phantom","Humanoid Robot","Full-size Humanoid Robot","Vision-Only Approach","Robot-as-a-Service"]},{"id":"persona-ai","category":"company","sec":2,"tier":3,"sources":[{"title":"Persona AI Raises $27M Oversubscribed Pre-Seed","url":"https://finance.yahoo.com/news/persona-ai-raises-27m-oversubscribed-131500748.html"},{"title":"HD Hyundai and Persona AI Sign Agreement to Deploy Humanoid Welding Robots","url":"https://www.prnewswire.com/news-releases/hd-hyundai-and-persona-ai-sign-agreement-to-deploy-humanoid-welding-robots-for-shipbuilding-automation-302449258.html"},{"title":"Persona AI 官网","url":"https://persona.ai"}],"as_of":"2026-09","related_ids":["humanoid-robot",null,"robot-as-a-service","dirty-dull-and-dangerous-jobs","florida-institute-for-human-and-machine-cognition"],"name":"Persona AI","alt":"Persona AI","abbr":"","aliases":[],"one_liner":"A Houston humanoid-robot company targeting heavy industry, especially shipyard welding.","explanation":"Persona AI was founded in Houston, Texas, in 2024. Its leadership includes CEO Nic Radford, CTO Jerry Pratt — a bipedal-robot researcher — and co-founder and COO Jide Akinyode. The company builds humanoid robots for heavy industry, with shipyard welding as its main focus alongside listed use cases such as mining, inspection, and steel-structure fabrication; commercially it plans to run on a robots-as-a-service model, charging by usage so customers don't need to buy the hardware outright. In May 2025 it announced an oversubscribed pre-seed round of $27 million. It has signed an agreement with South Korea's HD Hyundai group to develop humanoid welding robots for shipbuilding, with prototype delivery planned for late 2026 and on-site testing and commercial deployment starting in 2027.","example":"Developing a bipedal humanoid robot that can walk into a ship's hull structure and weld.","related":["Humanoid Robot","HD Hyundai","Robot-as-a-Service","Dirty, Dull, and Dangerous (3D) Jobs","Florida Institute for Human and Machine Cognition"]},{"id":"humanoid","category":"company","sec":2,"tier":3,"sources":[{"title":"Humanoid 官网","url":"https://thehumanoid.ai/"},{"title":"Humanoid News","url":"https://thehumanoid.ai/news/"},{"title":"Humanoid Secures Landmark Deal with Schaeffler","url":"https://thehumanoid.ai/news/humanoid-secures-landmark-deal-with-schaeffler-to-deploy-thousands-of-humanoid-robots/"}],"as_of":"2026-07","related_ids":["humanoid-hmnd-01","humanoid-robot","wheeled-humanoid-robot","schaeffler","figure-ai","apptronik"],"name":"Humanoid","alt":"Humanoid（英国人形机器人公司）","abbr":"","aliases":["SKL Robotics"],"one_liner":"A London humanoid-robot company building the HMND 01 in wheeled and bipedal versions.","explanation":"Humanoid is a humanoid-robot company founded in London in 2024, legally registered as SKL Robotics; its founder is reportedly investor Artem Sokolov. Its product is the HMND 01 Alpha, aimed at factories and logistics, available in a wheeled-base version and a bipedal version. Its companion robot-AI framework, KinetIQ, handles end-to-end scheduling for fleets of robots, and in June 2026 it introduced KinetIQ Ascend, focused on manipulation reliability and speed. In 2026 it announced partnerships in succession with Siemens, and then with NVIDIA, Schaeffler, and Bosch; its website says Schaeffler plans to deploy several thousand of the wheeled robots across its global factories by 2032. In July 2026 it closed a $152 million Series A at a post-money valuation of $1.35 billion, which it says makes it Europe's first pure-play humanoid-robot unicorn.","example":"In May 2026, Schaeffler signed an agreement with Humanoid to deploy HMND 01 humanoid robots in its own factories.","related":["Humanoid HMND 01","Humanoid Robot","Wheeled Humanoid Robot","Schaeffler","Figure AI","Apptronik"]},{"id":"rainbow-robotics","category":"company","sec":2,"tier":3,"sources":[{"title":"Rainbow Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Rainbow_Robotics"}],"as_of":"2025-03","related_ids":["rainbow-robotics-rb-y1","samsung-electronics","darpa-robotics-challenge","collaborative-robot","wheeled-humanoid-robot","humanoid-robot"],"name":"Rainbow Robotics","alt":"Rainbow Robotics","abbr":"","aliases":[],"one_liner":"A Korean company founded by KAIST's humanoid HUBO team, under Samsung's control since 2025.","explanation":"Rainbow Robotics was founded in February 2011 and is headquartered in Daejeon, South Korea. Founder Jun-ho Oh is a professor at KAIST's Humanoid Robot Research Center and led development of the bipedal humanoid HUBO; its DRC-HUBO won the 2015 DARPA Robotics Challenge. The company later commercialized its joint and control technology: the RB series of collaborative robots (arms that can work alongside people in the same space), quadruped robots, and, released in 2024, the wheeled dual-arm humanoid RB-Y1, which overseas labs often use as a platform for collecting demonstration data and training VLA models. The company is listed on the Korea Exchange (ticker 277810); Samsung Electronics reportedly announced in late 2024 that it would raise its stake to about 35%, becoming the largest shareholder after regulatory approval in March 2025.","example":"RB-Y1 uses a wheeled base plus two 7-degree-of-freedom arms, sidestepping bipedal balance issues, which makes it well suited to mobile-manipulation research.","related":["Rainbow Robotics RB-Y1","Samsung Electronics","DARPA Robotics Challenge","Collaborative Robot","Wheeled Humanoid Robot","Humanoid Robot"]},{"id":"clone-robotics","category":"company","sec":2,"tier":3,"sources":[{"title":"Clone Robotics 官网","url":"https://www.clonerobotics.com/"},{"title":"Protoclone 报道（Interesting Engineering）","url":"https://interestingengineering.com/innovation/video-worlds-first-humanoid-lifelike-muscles"}],"as_of":"2026-09","related_ids":["clone-robotics-protoclone",null,null,null,null],"name":"Clone Robotics","alt":"Clone Robotics","abbr":"","aliases":["Clone"],"one_liner":"A company building humanoid robots driven by artificial muscles that copy human anatomy.","explanation":"Clone Robotics got its start in 2021, reportedly headquartered in Wrocław, Poland, founded by Dhanush Radhakrishnan and Łukasz Koźlik. Instead of the usual “motor plus gearbox” approach, it models the human skeleton and muscles directly: fluid-driven artificial muscles it calls Myofiber pull on a polymer skeleton. The company started with robotic hands — the Clone Hand reportedly cost about $2,800 to build as of 2023 — and in February 2025 it unveiled Protoclone, a musculoskeletal humanoid prototype that was still suspended and had not demonstrated autonomous walking. Its website now lists a bipedal humanoid for individuals and businesses, Clone (Alpha), with a next-generation Neoclone in development. If this approach works, the resulting robots should move more smoothly and look more human, but at the cost of much harder control and reliability problems.","example":"Protoclone is driven by more than 1,000 Myofiber artificial muscles and is claimed to have more than 200 degrees of freedom.","related":["Clone Robotics Protoclone","Artificial Muscle","Bio-inspired Robot","Tendon-Driven Actuation","Hyper-realistic (Bionic) Humanoid Robot / Android"]},{"id":"engineered-arts","category":"company","sec":2,"tier":3,"sources":[{"title":"Engineered Arts - Wikipedia","url":"https://en.wikipedia.org/wiki/Engineered_Arts"},{"title":"Engineered Arts restructures in US and secures $10M to scale up humanoid robots（SiliconANGLE）","url":"https://siliconangle.com/2024/12/17/engineered-arts-restructures-us-secures-10m-scale-humanoid-robots/"},{"title":"Ameca (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Ameca_(robot)"},{"title":"About - Engineered Arts（官网）","url":"https://engineeredarts.com/us/about"}],"as_of":"2026-09","related_ids":["engineered-arts-ameca","hyper-realistic-humanoid-robot","uncanny-valley","human-robot-interaction","sophia","guided-tours-and-reception"],"name":"Engineered Arts","alt":"Engineered Arts","abbr":"","aliases":[],"one_liner":"A British humanoid-robot company known for Ameca, prized for its lifelike facial expressions.","explanation":"Engineered Arts was founded in Falmouth, Cornwall, UK, in October 2004 by Will Jackson. While building exhibits for London's Science Museum in the 1990s, Jackson wanted a machine that could explain things to visitors over and over — the idea that started the company; its 2005 mechanical theater piece for the Eden Project led to its first humanoid robot, RoboThespian. Ameca, which debuted at CES in January 2022, became well known for its nuanced expressions, driven by the company's own Tritium system, and can connect to a large language model for conversation, though it cannot walk; later variants include the head-and-shoulders-only Ameca Desktop and tabletop units named Ami and Azi. In December 2024 the company reorganized under a US entity based in Redwood City, California, closing a $10 million Series A led by Helium-3 Ventures. It says it has deployed more than 250 robots, including RoboThespian and other past products, across more than 30 countries.","example":"Connected to a large language model, Ameca can hold a conversation with a visitor while blinking, frowning, and looking surprised in response.","related":["Engineered Arts Ameca","Hyper-Realistic Humanoid Robot","Uncanny Valley","Human-Robot Interaction","Sophia (Hanson Robotics)","Guided Tours & Reception"]},{"id":"anybotics","category":"company","sec":2,"tier":3,"sources":[{"title":"ANYbotics earns strategic investment from Climate Investment (The Robot Report)","url":"https://www.therobotreport.com/anybotics-earns-strategic-investment-from-climate-investment"},{"title":"ANYbotics raises additional USD 20 million (Startupticker)","url":"https://www.startupticker.ch/en/news/anybotics-raises-additional-usd-20-million"}],"as_of":"2026-08","related_ids":["anybotics-anymal","eth-zurich-robotic-systems-lab","quadruped-robot","inspection-robot","rl-based-locomotion-control"],"name":"ANYbotics","alt":"ANYbotics","abbr":"","aliases":[],"one_liner":"A spin-out of ETH Zurich building the ANYmal industrial-inspection quadruped robot.","explanation":"ANYbotics was founded in 2016 in Zurich, Switzerland, as a spin-out of the Robotic Systems Lab at ETH Zurich (Marco Hutter's group). Its product is the ANYmal quadruped robot, used for autonomous inspection of industrial sites in chemicals, energy, and mining: it patrols set routes carrying cameras, thermal imagers, and gas sensors, taking over work in hazardous areas that would otherwise require a person. ANYmal is also a classic research platform for legged reinforcement learning, having been used to validate landmark work such as actuator networks, blind locomotion, and perceptive locomotion. In September 2025 it raised a further $20 million, bringing reported total funding to more than $150 million; its first explosion-proof-certified quadruped, ANYmal X, is planned to begin deliveries in 2026.","example":"At a petrochemical plant, ANYmal follows a fixed route to read gauges and detect gas leaks.","related":["ANYbotics ANYmal","ETH Zurich Robotic Systems Lab","Quadruped Robot","Inspection Robot","RL-based Locomotion Control"]},{"id":"hello-robot","category":"company","sec":2,"tier":3,"sources":[{"title":"Hello Robot - About","url":"https://hello-robot.com/about"},{"title":"Hello Robot 官网（Stretch 4）","url":"https://hello-robot.com/"}],"as_of":"2026-09","related_ids":["hello-robot-stretch","mobile-manipulation","mobile-manipulator","ok-robot","dobb-e","household-tasks"],"name":"Hello Robot","alt":"Hello Robot","abbr":"","aliases":["Hello Robot Inc."],"one_liner":"A US mobile-manipulation robotics company, maker of the lightweight, open-source Stretch robot.","explanation":"Hello Robot is a US robotics company founded in 2017, with offices in Atlanta, Georgia and Martinez, California (which handles manufacturing). CEO Aaron Edsinger previously served as Google's director of robotics and had earlier founded Meka Robotics and other companies; CTO Charlie Kemp was formerly an associate professor at Georgia Tech, researching assistive mobile manipulation. The company has a single product line, Stretch: a differential-drive base with a vertical lift column and a telescoping arm, with few degrees of freedom, light weight, and open-source software, aimed at research, in-home assistance (such as helping someone with limited mobility fetch objects), and enterprise use. It's a common platform for academic work on household mobile manipulation — both OK-Robot and Dobb-E ran their experiments on Stretch. As of September 2026, its website is selling the fourth-generation Stretch 4 for $29,950.","example":"NYU's OK-Robot combines open-vocabulary detection with a grasping model on a Stretch robot to complete “bring me object X” tasks in real homes.","related":["Hello Robot Stretch","Mobile Manipulation","Mobile Manipulator","OK-Robot","Dobb-E","Household Tasks"]},{"id":"collaborative-robotics","category":"company","sec":2,"tier":3,"sources":[{"title":"Cobot 团队页","url":"https://www.co.bot/team"},{"title":"Cobot 官网","url":"https://www.co.bot/"}],"as_of":"2026-09","related_ids":[null,null,null,null,"scale-ai"],"name":"Collaborative Robotics","alt":"Collaborative Robotics","abbr":"Cobot","aliases":["Cobot"],"one_liner":"A US company founded by a former Amazon Robotics VP, building the Proxie mobile robot that works alongside people.","explanation":"Collaborative Robotics (often shortened to Cobot) was founded in 2022, with offices in Seattle and Santa Clara. Its founder and CEO, Brad Porter, studied computer science at MIT, served as vice president of Amazon Robotics where he oversaw the deployment of more than 500,000 robots, and later became CTO of Scale AI. Its flagship product is Proxie, a mobile manipulation robot designed to work safely right next to people, aimed at labor-short settings such as hospitals and logistics, doing repetitive physical work like material transport. Its Series A was funded by Sequoia, Khosla Ventures, and Mayo Clinic, and its Series B, reportedly $100 million in 2024, was led by General Catalyst; it later formed a foundation-model team. Note that the company is a distinct entity from the generic term “cobot” for a collaborative robot.","example":"","related":["Collaborative Robot","Mobile Manipulator","Human-Robot Collaboration","Labor Shortage (Aging Workforce)","Scale AI"]},{"id":"unitree-robotics","category":"company","sec":3,"tier":1,"sources":[{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"},{"title":"Unitree G1","url":"https://www.unitree.com/g1/"},{"title":"unitreerobotics/unifolm-wla - GitHub","url":"https://github.com/unitreerobotics/unifolm-wla"}],"as_of":"2026-08","related_ids":["unitree-g1","unitree-go2","unitree-h2","unifolm","hangzhou-s-six-little-dragons","star-market"],"name":"Unitree Robotics","alt":"宇树科技","abbr":"","aliases":["Unitree","宇树 (colloquial Chinese short form)"],"one_liner":"A Hangzhou maker of quadruped and humanoid robots, known for low-cost, high-performance hardware.","explanation":"Unitree Robotics was founded in Hangzhou in August 2016 by Xingxing Wang, who built a quadruped prototype called XDog at Shanghai University during his graduate studies. Unitree drove down the price of legged robots sharply through in-house motors and control software: its quadrupeds include the Go1, Go2, and B2, and its humanoids include the H1, the G1 (released 2024, starting at RMB 99,000), and the H2 (October 2025). Thanks to their low price and open SDK, the G1 and Go2 have become among the most commonly used real robots in university locomotion-control and humanoid-VLA papers; an H1 performing the dance 'Yangbot' at China's 2025 CCTV Spring Festival Gala (China's most-watched TV broadcast) went viral. The company also open-sources its UnifoLM model series. It began IPO coaching in July 2025 and listed on the Shanghai Stock Exchange's STAR Market in August 2026 (688836).","example":"Many humanoid reinforcement-learning papers use unitree_rl_gym to train a walking policy in simulation, then deploy it onto a real Unitree G1.","related":["Unitree G1","Unitree Go2","Unitree H2","UnifoLM","Hangzhou's Six Little Dragons","STAR Market"]},{"id":"agibot","category":"company","sec":3,"tier":1,"sources":[{"title":"智元机器人 - 维基百科","url":"https://zh.wikipedia.org/wiki/智元机器人"},{"title":"AgiBot - Wikipedia","url":"https://en.wikipedia.org/wiki/AgiBot"},{"title":"AgiBot GO-2 发布","url":"https://www.agibot.com/article/231/detail/56.html"}],"as_of":"2026-04","related_ids":["agibot-go-1","agibot-go-2","agibot-world","agibot-a3","swancor-advanced-materials","agibot-g1g5-embodied-ai-roadmap"],"name":"AgiBot","alt":"智元机器人","abbr":"","aliases":["Zhiyuan Robotics","AgiBot Innovation (智元新创)"],"one_liner":"A Shanghai humanoid robot company that builds its own hardware, models, and datasets in-house.","explanation":"AgiBot (智元机器人) was founded in Shanghai in February 2023. Co-founder Zhihui Peng (彭志辉, known online as 稚晖君) previously joined Huawei under its 'Genius Youth' program; the CEO is Tenghua Deng. Its product lines span the Yuanzheng (远征) A2 humanoid series, the wheeled dual-arm Lingjing (精灵) G1/G2, the tabletop humanoid Lingxi (灵犀) X1/X2, the quadruped D1, and the OmniHand dexterous hand. AgiBot builds both models and data in-house: in March 2025 it released GO-1 (a ViLLA-architecture model) alongside the AgiBot World dataset of a million trajectories, followed by GO-2 in April 2026. Mass production began in December 2024, with the 1,000th unit shipped in January 2025. According to reports, in July 2025 AgiBot announced it had taken a controlling stake in STAR Market-listed Swancor Advanced Materials, and it plans to list in Hong Kong in 2026.","example":"AgiBot records AgiBot World using its own factory-scale data collection operation, running around a hundred robots, and then trains GO-1 on that data — a strategy that advances hardware, data, and models together.","related":["AgiBot GO-1","AgiBot GO-2","AgiBot World","AgiBot A3","Swancor Advanced Materials","AgiBot G1–G5 Embodied-AI Roadmap"]},{"id":"ubtech-robotics","category":"company","sec":3,"tier":1,"sources":[{"title":"UBtech Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/UBtech_Robotics"},{"title":"Walker S2 - UBTECH","url":"https://www.ubtrobot.com/en/humanoid/products/walker-s2"},{"title":"UBTECH-Robot/Thinker - GitHub","url":"https://github.com/UBTECH-Robot/Thinker"}],"as_of":"2026-02","related_ids":["ubtech-walker-s2","ubtech-thinker","first-listed-humanoid-robot-stock","hkex-chapter-18c","factory-pilot-deployment","industrial-robot"],"name":"UBTech Robotics","alt":"优必选","abbr":"","aliases":["UBTECH","UBTech Technology (优必选科技)"],"one_liner":"A Shenzhen humanoid robot company, listed in Hong Kong, best known for its industrial Walker S humanoid.","explanation":"UBTech was founded in Shenzhen in March 2012 by Jian Zhou. It started out with the small Alpha series of humanoid robots and educational robots, later moving into the larger Walker humanoid line. It listed on the Hong Kong Stock Exchange (9880) in December 2023, earning the nickname 'the first humanoid robot stock.' Its current flagship product is the industrial Walker S series, piloted in car factories for material handling and sorting; the Walker S2, released in July 2025, can swap its own battery using both arms, enabling continuous operation, and reports say it landed several orders worth over RMB 100 million each in 2025. On the model side, UBTech has built its own Thinker model stack in-house, open-sourcing the Thinker-4B vision-language model in early 2026.","example":"The Walker S2 autonomously swaps its own battery in a car factory in about 3 minutes, with no need to stop for charging.","related":["UBTech Walker S2","UBTech Thinker","First Listed Humanoid Robot Stock","HKEX Chapter 18C","Factory Pilot Deployment","Industrial Robot"]},{"id":"galbot","category":"company","sec":3,"tier":1,"sources":[{"title":"银河通用：成立34个月、6轮融资70亿（36氪）","url":"https://www.36kr.com/p/3978672823417862"},{"title":"Galbot Raises $362 Million in Fresh Funding, Eyes Hong Kong IPO (Caixin Global, 2026-03)","url":"https://www.caixinglobal.com/2026-03-03/galbot-raises-362-million-in-fresh-funding-eyes-hong-kong-ipo-102418742.html"},{"title":"银河通用加码零售场景应用，机器人在北京海淀开店（21世纪经济报道，2025-08）","url":"https://www.21jingji.com/article/20250808/herald/0c3a2e4ba2568d4a86a4647721408a51.html"}],"as_of":"2026-08","related_ids":["galbot-g1","galbot-s1","astrabrain","graspvla","navfom","synthetic-data"],"name":"Galbot","alt":"银河通用","abbr":"","aliases":["Beijing Galbot Robotics"],"one_liner":"A Beijing embodied-AI company founded by PKU's He Wang, building wheeled humanoid robots and embodied foundation models.","explanation":"Galbot was founded in Beijing in May 2023; founder He Wang holds a Stanford PhD and teaches at Peking University's School of Computer Science, with a long research background in robot grasping and manipulation. Its main products are the wheeled dual-arm humanoid Galbot G1 and the industrial heavy-payload variant S1, with a bipedal robot, ET1, added in August 2026. On models, it has released the grasping model GraspVLA and the navigation model NavFoM, and refers to its overall embodied-model stack as AstraBrain, notable for training heavily on simulated synthetic data. Its robots are already deployed at a CATL factory, in pharmacies, and in unstaffed retail stores, and it appeared on China Central Television's 2026 Spring Festival Gala. In March 2026 it raised RMB 2.5 billion, with investors including the National Artificial Intelligence Industry Investment Fund; it has reportedly raised about RMB 7 billion cumulatively at a valuation above RMB 20 billion and is preparing a Hong Kong listing.","example":"At a smart pharmacy run with Meituan Grocery in Beijing, a single Galbot manages more than 5,000 medicines across roughly 40 square meters of floor space, pulling items off the shelf by order with, the company says, no teleoperation involved.","related":["Galbot G1","Galbot S1","AstraBrain","GraspVLA","NavFoM (Galbot)","Synthetic Data"]},{"id":"galaxea-ai","category":"company","sec":3,"tier":1,"sources":[{"title":"星海图 关于我们","url":"https://www.galaxea-ai.com/about"},{"title":"星海图官网","url":"https://galaxea-ai.com/"}],"as_of":"2026-06","related_ids":["galaxea-g0-dual-system-vla","galaxea-r1","galaxea-open-world-dataset","dual-system-architecture","mobile-manipulation","tsinghua-university-institute-for-interdisciplinary-informat"],"name":"Galaxea AI","alt":"星海图","abbr":"","aliases":["Galaxea","Galaxea (Beijing) Artificial Intelligence Technology Co."],"one_liner":"A Beijing embodied AI company that builds both the wheeled bimanual robot R1 and the open-source G0 model.","explanation":"Galaxea AI was founded in Beijing in September 2023, following a 'hardware plus intelligence' strategy of building both robot bodies and foundation models in-house. Its website says the core team has experience shipping autonomous-driving products at scale; the CEO is reportedly Jiyang Gao, and co-founders include scholars from Tsinghua University's Institute for Interdisciplinary Information Sciences. Its hardware includes the R1 series of wheeled bimanual robots (R1, R1 Pro, R1 Lite) and the A1 robot arm, and its website says its customers include Stanford, Physical Intelligence, and nearly a hundred other institutions. On the model side, it released and open-sourced G0, a dual-system (fast-and-slow) model, at the end of August 2025, alongside the open-sourced Galaxea Open-World Dataset of about 500 hours and 100,000 trajectories; it later released G0Plus (January 2026) and G0.5 (June 2026). The company's website discloses a Series A of nearly RMB 300 million in February 2025, led by Ant Group.","example":"A researcher downloads the Galaxea Open-World Dataset and fine-tunes G0 on an R1 Lite for a task like tidying a tabletop.","related":["Galaxea G0 Dual-System VLA","Galaxea R1","Galaxea Open-World Dataset","Dual-System Architecture (System 1 / System 2)","Mobile Manipulation","Tsinghua University Institute for Interdisciplinary Information Sciences"]},{"id":"fourier","category":"company","sec":3,"tier":2,"sources":[{"title":"Fourier (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Fourier_Intelligence"},{"title":"傅利叶智能 - 维基百科","url":"https://zh.wikipedia.org/wiki/傅利叶智能"}],"as_of":"2025","related_ids":["fourier-gr-1","fourier-gr-2","fourier-gr-3","fourier-n1","fourier-actionnet","exoskeleton"],"name":"Fourier","alt":"傅利叶","abbr":"","aliases":["Fourier Intelligence"],"one_liner":"A Shanghai humanoid robot company that started in rehabilitation robotics, maker of the GR series.","explanation":"Fourier was founded in Shanghai in 2015 by Jie Gu, a graduate of Shanghai Jiao Tong University who had previously worked in sales management at National Instruments (NI); the company is named after the mathematician Fourier. It started out making rehabilitation robots and exoskeletons for hospital rehab departments, and began developing humanoid robots in 2019, using its own in-house FSA integrated actuators for the joints. It released the GR-1 in 2023, one of China's earlier humanoid platforms delivered at volume to universities and research institutions; the GR-2 followed in September 2024, then the GR-3, aimed at care and companionship settings, alongside the open-sourced real-robot dataset ActionNet. Investors include IDG Capital and Saudi Aramco, and reports say it closed a Series E of nearly RMB 800 million in early 2025. Its distinguishing feature is entering from rehabilitation medicine and pushing humanoid robots toward elder care and caregiving applications.","example":"iDP3, the humanoid version of 3D Diffusion Policy, was validated on a real Fourier GR-1.","related":["Fourier GR-1","Fourier GR-2","Fourier GR-3","Fourier N1","Fourier ActionNet","Exoskeleton"]},{"id":"limx-dynamics","category":"company","sec":3,"tier":2,"sources":[{"title":"LimX Dynamics - About","url":"https://www.limxdynamics.com/en/about"},{"title":"逐际动力完成A轮融资，人形机器人赛道融资升温（第一财经）","url":"https://www.yicai.com/news/102191399.html"},{"title":"张巍（百度百科）","url":"https://baike.baidu.com/item/%E5%BC%A0%E5%B7%8D/63753106"}],"as_of":"2026-07","related_ids":["limx-dynamics-tron-1","limx-dynamics-tron-2","limx-dynamics-oli","legged-robot","rl-based-locomotion-control","humanoid-robot"],"name":"LimX Dynamics","alt":"逐际动力","abbr":"","aliases":["LimX"],"one_liner":"A Shenzhen legged- and humanoid-robot company behind the TRON series and the Oli humanoid.","explanation":"LimX Dynamics was founded in Shenzhen in January 2022; its founder, Wei Zhang, is a professor at the Southern University of Science and Technology who previously taught at Ohio State University and has long researched legged-robot control. The company started with leg locomotion control, and its product line includes the TRON 1 biped, whose foot can be swapped between point, flat, and wheeled configurations; the TRON 2, which can switch between dual-arm, bipedal, and wheeled-leg forms; the full-size humanoid Oli; and Luna, released in May 2026. On the software side it has COSA, a humanoid “brain” system, and it open-sourced the VLA engineering framework FluxVLA Engine in April 2026. Alibaba led its Series A in July 2025; it closed a $200 million Series B in February 2026, and a further Pre-IPO round of about $200 million that July.","example":"A university lab buys a TRON 1, fits it with point feet, trains a walking policy with reinforcement learning in simulation, then deploys it to the real robot to verify sim-to-real transfer.","related":["LimX Dynamics TRON 1","LimX Dynamics TRON 2","LimX Dynamics Oli","Legged Robot","RL-based Locomotion Control","Humanoid Robot"]},{"id":"engineai","category":"company","sec":3,"tier":2,"sources":[{"title":"众擎机器人官网：完成2亿美元B轮融资，估值破百亿","url":"https://www.engineai.com.cn/about-news-media/55.html"},{"title":"21经济网：众擎机器人创始人赵同阳：下一代产品定位打工机器人","url":"https://www.21jingji.com/article/20251128/herald/42125069ac6bccd01f5b88914a62a867.html"},{"title":"观察者网风闻：众擎机器人据报已秘密递表港交所","url":"https://user.guancha.cn/main/content?id=1670242"}],"as_of":"2026-08","related_ids":["engineai-t800","engineai-pm01","engineai-se01","straight-knee-walking","humanoid-robot","rl-based-locomotion-control"],"name":"EngineAI","alt":"众擎机器人","abbr":"","aliases":["Zhongqing Robotics","Engine AI"],"one_liner":"A Shenzhen humanoid-robot company known for straight-knee walking and dynamic moves like front flips.","explanation":"EngineAI was founded in Shenzhen in October 2023; its founder, Tongyang Zhao, previously led the robotics team at Peng Xing Intelligent, XPeng's robotics affiliate. The company prioritizes “physical fitness first” — making the robot body and its motion control (what the industry calls the “cerebellum”) strong before anything else. In October 2024, its SE01 achieved straight-knee walking (unlike the bent-knee gait of most humanoids), its PM01 went viral on video doing a front flip, and in December 2025 it released the full-size humanoid T800. Funding has moved quickly: JD.com led its Series A1 in 2025; in April 2026 it closed a $200 million Series B co-led by Henan Investment Group's Huirong Fund and Luxshare Precision at a valuation above RMB 10 billion, followed by a Series B+, and it has reportedly confidentially filed for a Hong Kong listing. The first batch of T800 units rolled off the line in May 2026, and in August, at the World Robot Conference, it unveiled its embodied-AI engine, EngineAI Awaken.","example":"After T800's release, its movements looked so smooth that some accused the footage of being CGI; the company responded with a video of a T800 kicking down its own CEO, Tongyang Zhao.","related":["EngineAI T800","EngineAI PM01","EngineAI SE01","Straight-Knee Walking","Humanoid Robot","RL-based Locomotion Control"]},{"id":"booster-robotics","category":"company","sec":3,"tier":2,"sources":[{"title":"加速进化官网","url":"https://www.booster.tech/zh/"},{"title":"Booster Robotics 官网（英文）","url":"https://www.booster.tech/"}],"as_of":"2026-09","related_ids":["booster-robotics-t1","booster-robotics-k1","robocup",null,null,null],"name":"Booster Robotics","alt":"加速进化","abbr":"","aliases":["Booster"],"one_liner":"A Beijing humanoid-robot company making small and mid-size humanoids for developers and robot soccer.","explanation":"Booster Robotics was founded in Beijing in 2023; its founder and CEO is reportedly Hao Cheng. It positions itself as an “embodied development platform,” selling to universities, competition teams, and developers with an open SDK that makes it easy to build reinforcement-learning locomotion control and higher-level algorithms. Its products include the roughly 1.2-meter Booster T1, the roughly 95-centimeter entry-level K1 (reportedly starting around $4,999), and the flagship development platform T2. It is best known in robot soccer: in 2025, Tsinghua's Hulk team won the RoboCup humanoid AdultSize championship on a T1; the company's website says it swept all three titles in the bipedal humanoid divisions at RoboCup 2026, where 38 of the 59 competing teams used its robots. Its website states it has raised nearly RMB 1 billion in total funding.","example":"At RoboCup 2025, Tsinghua's Hulk team won the humanoid AdultSize championship on a Booster T1, while Germany's HTWK team won the KidSize division on a K1.","related":["Booster Robotics T1","Booster Robotics K1","RoboCup","Small-size Humanoid Robot","RL-based Locomotion Control","Research & Education Market"]},{"id":"robotera","category":"company","sec":3,"tier":2,"sources":[{"title":"星动纪元官网：关于我们","url":"https://www.robotera.com/about/us"}],"as_of":"2026-06","related_ids":["era-42","robotera-star1","robotera-l7","robotera-xhand1","video-prediction-policy","tsinghua-university-institute-for-interdisciplinary-informat"],"name":"RobotEra","alt":"星动纪元","abbr":"","aliases":["Beijing RobotEra"],"one_liner":"A Tsinghua-incubated humanoid-robot company building its own hardware, dexterous hands, and models in-house.","explanation":"RobotEra was founded in Beijing in August 2023; founder Jianyu Chen is an assistant professor at Tsinghua University's Institute for Interdisciplinary Information Sciences (IIIS), and the company was incubated with Tsinghua as a shareholder, describing itself as fully self-developed across both hardware and software. On hardware it released the humanoid STAR1 in August 2024, the full-size humanoid L7 in July 2025 (55 degrees of freedom across the body), the 12-degree-of-freedom dexterous hand XHAND1, and an upgraded XHAND 1 PRO in June 2026. On models it has released the end-to-end VLA model ERA-42, and the team has also published papers on the video-prediction policy VPP and the controllable world model Ctrl-World. It is one of the few Chinese companies building its own robot body, dexterous hand, and large model all in-house, with products aimed at logistics, manufacturing, and commercial services.","example":"The same ERA-42 model drives the L7, whether it's showing off street-dance-level dynamic moves or fine manipulation like tightening a screw or sorting items.","related":["ERA-42","RobotEra STAR1","RobotEra L7","RobotEra XHAND1","Video Prediction Policy","Tsinghua University Institute for Interdisciplinary Information Sciences"]},{"id":"noetix-robotics","category":"company","sec":3,"tier":2,"sources":[{"title":"松延动力 - 人形机器人公司资料、融资、产品与估值 | RobotHub","url":"https://www.robothub.app/zh/companies/noetix-robotics"},{"title":"仿生机器人、万元级人形机器人「出圈」（21经济网）","url":"https://www.21jingji.com/article/20260407/herald/532f01a3612369c961d706065fd9bc76.html"},{"title":"95后姜哲源创建松延动力2年后，春晚再现机器人蔡明（瑞财经）","url":"https://m.rccaijing.com/news-7429165880347128834.html"}],"as_of":"2026-09","related_ids":["noetix-n2","noetix-bumi","10-000-yuan-class-humanoid-robot","hyper-realistic-humanoid-robot","small-size-humanoid-robot","research-and-education-market"],"name":"Noetix Robotics","alt":"松延动力","abbr":"","aliases":["Noetix"],"one_liner":"A Beijing humanoid-robot company known for the low-cost N2, the sub-RMB-10,000 Bumi, and bionic faces.","explanation":"Noetix Robotics was founded in Changping, Beijing in September 2023; its founder, Zheyuan Jiang, is a Tsinghua-trained member of China's post-1995 generation. The company pursues a standardized, low-cost strategy: small humanoids such as the N2, aimed at education, research, exhibitions, and commercial performances, reportedly starting at RMB 39,900; in October 2025 it released Bumi, priced at RMB 9,998, one of the earliest humanoid robots priced under RMB 10,000. It also makes the bionic-face robot Hobbs; the “robot Cai Ming” that appeared on CCTV's 2026 Spring Festival Gala, China's most-watched TV broadcast, came from Noetix. According to public data platforms, the company has raised about RMB 1.5 billion in total, including about RMB 1 billion in a March 2026 round; media reports in September 2026 said it was pursuing an IPO.","example":"The Bumi is priced at RMB 9,998, aimed at home companionship and education.","related":["Noetix N2","Noetix Bumi","10,000-Yuan-Class Humanoid Robot","Hyper-Realistic Humanoid Robot","Small-size Humanoid Robot","Research & Education Market"]},{"id":"leju-robotics","category":"company","sec":3,"tier":2,"sources":[{"title":"深圳新闻网：乐聚智能IPO获受理","url":"https://www.sznews.com/news/content/2026-05/20/content_32057558.htm"},{"title":"南都：乐聚机器人开启上市辅导","url":"https://m.mp.oeeee.com/a/BAAFRD0000202510311166071.html"},{"title":"新浪财经：乐聚机器人完成IPO辅导验收","url":"https://finance.sina.com.cn/wm/2026-04-22/doc-inhvkqha6091928.shtml"}],"as_of":"2026-05","related_ids":["leju-kuavo","humanoid-robot","full-size-humanoid-robot","small-size-humanoid-robot","m-robots-os","research-and-education-market"],"name":"Leju Robotics","alt":"乐聚机器人","abbr":"","aliases":["Leju","Leju Intelligent"],"one_liner":"A Shenzhen humanoid-robot company behind the Kuavo robot, now pursuing a ChiNext IPO in 2026.","explanation":"Leju Intelligent (Shenzhen) Co., Ltd. was founded in 2016; its three founders, Xiaokun Leng (chairman and CTO), Lin Chang (CEO), and Ziwei An (COO), all hold PhDs from Harbin Institute of Technology and together control about 33.72% of the company. It started out with small, 30-70-centimeter educational humanoid robots, which reached nearly 5,000 primary, secondary, and vocational schools nationwide; in December 2023 it released the full-size bipedal humanoid Kuavo, running the open-source OpenHarmony operating system, which has been piloted in factories including BAIC. Its investors include Tencent and Shenzhen Capital Group, and it closed a Pre-IPO round of nearly RMB 1.5 billion in October 2025. On May 19, 2026, the Shenzhen Stock Exchange accepted its ChiNext IPO application, making it the first company to qualify under ChiNext's fourth listing standard; IDC data puts its 2025 humanoid-robot shipments third worldwide.","example":"A Kuavo robot carried the torch during the Shenzhen leg of the 15th National Games.","related":["Leju Kuavo","Humanoid Robot","Full-size Humanoid Robot","Small-size Humanoid Robot","M-Robots OS (OpenHarmony-based Robot OS)","Research & Education Market"]},{"id":"magiclab","category":"company","sec":3,"tier":2,"sources":[{"title":"管理团队调整不到一周，魔法原子完成新一轮5亿元融资（界面新闻）","url":"https://m.jiemian.com/article/14090359.html"},{"title":"登春晚、筹划上市，创始人兼CEO突然离职（21经济网）","url":"https://www.21jingji.com/article/20260306/herald/dcf1a111991ff644ef5b79393144c1fd.html"},{"title":"魔法原子完成1.5亿元天使轮融资（魔法原子官网）","url":"https://www.magiclab.top/news/29"}],"as_of":"2026-03","related_ids":["magiclab-magicbot","humanoid-robot","quadruped-robot","xiaomi-cyberdog","lumos-robotics","factory-pilot-deployment"],"name":"MagicLab","alt":"魔法原子","abbr":"","aliases":[],"one_liner":"A humanoid-robot company incubated by Dreame Technology, maker of MagicBot and MagicDog.","explanation":"MagicLab launched publicly in January 2024, incubated with participation from Dreame Technology; its core team came from Xiaomi's CyberDog (“Tiedan”) robot-dog project, and its main entity is based in Suzhou, Jiangsu, with an additional office in Wuxi. Its products include the full-size humanoid MagicBot Gen1, the small, highly dynamic bipedal MagicBot Z1, the quadruped MagicDog, and dexterous hands, starting out in factory scenarios like material handling and quality inspection. It raised an RMB 150 million angel round in December 2024, performed as a robotics partner on CCTV's 2026 Spring Festival Gala, China's most-watched TV broadcast, and that March founder Changzheng Wu stepped down, with CTO Chunyu Chen taking over as legal representative, after which the company closed a new RMB 500 million round and says it is accelerating toward an IPO.","example":"MagicLab has put multiple MagicBot units on a factory line to pilot quality inspection, material handling, and parts placement, testing whether humanoid robots can work in a real factory.","related":["MagicLab MagicBot","Humanoid Robot","Quadruped Robot","Xiaomi CyberDog","Lumos Robotics","Factory Pilot Deployment"]},{"id":"astribot","category":"company","sec":3,"tier":2,"sources":[{"title":"深圳，又跑出一家百亿具身智能独角兽（智东西）","url":"https://m.zhidx.com/p/562710.html"},{"title":"星尘智能官网","url":"https://www.astribot.com/index.html"}],"as_of":"2026-09","related_ids":["astribot-s1","tendon-driven-actuation","wheeled-humanoid-robot","dual-arm-robot","teleoperation"],"name":"Astribot","alt":"星尘智能","abbr":"","aliases":[],"one_liner":"A Shenzhen company building the tendon-driven AI robot Astribot S1.","explanation":"Astribot was founded in Shenzhen in December 2022. Its founder and CEO, Jie Lai, previously led Baidu's Xiaodu robot team and was an early core member of Tencent's Robotics X lab. Its flagship product, Astribot S1, is a wheeled, dual-arm humanoid robot whose joints use tendon-driven actuation — cables that transmit motor force to the joints — which keeps the arms light, lets them move quickly, makes contact with people safer, and lowers cost. The company develops its hardware, teleoperation, and models all in-house, and demonstrated the S1 performing autonomous household tasks such as opening a bottle, peeling a cucumber, and writing calligraphy. Its investors include Ant Group and Jinqiu Fund; in June 2026 it announced it had closed three Series B rounds within three months, raising more than RMB 1 billion in total at a valuation above RMB 10 billion.","example":"In Astribot S1's release video, the robot autonomously folds clothes, pours a drink, and tosses a wok while stir-frying.","related":["Astribot S1","Tendon-Driven Actuation","Wheeled Humanoid Robot","Dual-arm Robot","Teleoperation"]},{"id":"deep-robotics","category":"company","sec":3,"tier":2,"sources":[{"title":"上交所：杭州云深处科技股份有限公司招股说明书（申报稿）","url":"https://static.sse.com.cn/stock/disclosure/announcement/c/202605/002190_20260518_G44F.pdf"},{"title":"21经济网：云深处完成上市辅导","url":"https://www.21jingji.com/article/20260502/herald/2737465f8b9147a684f6b64e3683dcad.html"},{"title":"新浪财经：云深处科创板IPO更新财务资料","url":"https://finance.sina.com.cn/jjxw/2026-09-28/doc-initkxux9642412.shtml"}],"as_of":"2026-09","related_ids":["deep-robotics-jueying-x30","deep-robotics-lite3","deep-robotics-lynx","deep-robotics-dr02","quadruped-robot","inspection-robot"],"name":"DEEP Robotics","alt":"云深处科技","abbr":"","aliases":["Hangzhou DEEP Robotics Co., Ltd."],"one_liner":"A Hangzhou quadruped-robot maker focused on power-line inspection and other industrial uses, now pursuing a STAR Market IPO.","explanation":"DEEP Robotics was founded in Hangzhou in 2017. Its founder and CEO, Qiuguo Zhu, is an associate professor at Zhejiang University's College of Control Science and Engineering who has long studied humanoid and bio-inspired robots; he started the company with fellow Zhejiang University PhD Chao Li (now CTO), and it is one of the so-called “Six Little Dragons of Hangzhou.” It is one of the earliest companies in China to commercialize quadruped robots, with products including the Jueying series of quadrupeds (X30, Lite3), the Lynx series of wheeled-leg robots (wheels mounted at the leg tips), and the DR series of humanoids, sold mainly to customers in power-line inspection, emergency response and firefighting, industrial inspection, and security. Its prospectus shows RMB 337 million in revenue in 2025, its first profitable year. Its STAR Market IPO application was accepted on May 18, 2026, it responded to a first round of inquiries in August, updated its financial materials on September 28, and reported a net loss of about RMB 8.8 million for the first half of 2026; the review is still ongoing.","example":"DEEP Robotics' DR02, released in 2025, is billed as the first humanoid robot with full-body IP66 water and dust protection, carrying over the sealing know-how it built for its outdoor-patrolling robot dogs.","related":["DEEP Robotics Jueying X30","DEEP Robotics Lite3","DEEP Robotics Lynx","DEEP Robotics DR02","Quadruped Robot","Inspection Robot"]},{"id":"swancor-advanced-materials","category":"company","sec":3,"tier":3,"sources":[{"title":"智元机器人拟收购上纬新材63.62%股份（华尔街见闻）","url":"https://wallstreetcn.com/articles/3750667"},{"title":"从「智元资本壳」到启元机器人，上纬新材惊险一跃（界面新闻）","url":"https://www.jiemian.com/article/14955818.html"},{"title":"启元机器人开启预订 个人机器人叩开家庭新消费大门（科技日报，2026-08-23）","url":"https://www.stdaily.com/web/gdxw/2026-08/23/content_568520.html"}],"as_of":"2026-08","related_ids":["agibot","backdoor-listing","star-market","humanoid-robot-concept-stocks","consumer-grade-robot"],"name":"Swancor Advanced Materials","alt":"上纬新材","abbr":"","aliases":["Swancor Qiyuan"],"one_liner":"A STAR Market materials company now controlled by AgiBot, which has pivoted it into consumer robots.","explanation":"Swancor Advanced Materials traces back to Swancor Fine Chemical, founded in 2000 and headquartered in Shanghai; it listed on the Shanghai STAR Market in 2020 (ticker 688585), originally making environmentally friendly, new-energy materials such as resins for wind-turbine blades. In July 2025, AgiBot announced it was taking control through a combination of a negotiated share transfer and a tender offer, ultimately holding about 63.62% of the company; AgiBot co-founder Zhihui Peng (known online as “Zhihuijun”) became chairman. The market initially read this as AgiBot pursuing a backdoor listing, though AgiBot said it had no such plan within three years. Since then, Swancor has independently built consumer- and home-facing robots under the “Swancor Qiyuan” brand, kept separate from AgiBot's commercial and industrial robots. Per its 2026 interim report, it had launched two products, the Qiyuan Q1 and T1, with RMB 210 million in advance customer payments; first-half revenue was RMB 803 million, with a net loss attributable to the parent of RMB 167 million due to robotics R&D spending.","example":"In August 2026, Qiyuan Robotics opened pre-orders, reportedly priced around RMB 20,000, aimed at home users.","related":["AgiBot","Backdoor Listing","STAR Market","Humanoid Robot Concept Stocks","Consumer-Grade Robot"]},{"id":"zhejiang-fenglong-electric","category":"company","sec":3,"tier":3,"sources":[{"title":"优必选正式入主锋龙股份（证券时报网）","url":"https://stcn.com/article/detail/3674855.html"},{"title":"方达助力优必选科技收购锋龙股份控制性股权","url":"http://fangdalaw.com/content/print34_9321.html"},{"title":"锋龙股份再回应：优必选三年内不会借壳上市（新京报）","url":"https://m.bjnews.com.cn/detail/1766920585129742.html"}],"as_of":"2026-04","related_ids":["ubtech-robotics","backdoor-listing","humanoid-robot-concept-stocks","first-listed-humanoid-robot-stock"],"name":"Zhejiang Fenglong Electric","alt":"锋龙股份","abbr":"","aliases":[],"one_liner":"A Zhejiang precision-parts listed company that has come under UBTech's control since 2026.","explanation":"Zhejiang Fenglong Electric (002931.SZ) listed on the Shenzhen Stock Exchange in 2018, with core businesses in three lines of precision components — for garden machinery, automobiles, and hydraulics — supplying complete-equipment makers including STIHL and Honda. It entered embodied-AI news because of a change in control: on December 24, 2025, Hong Kong-listed humanoid-robot company UBTech announced it would acquire about 43% of Fenglong for roughly RMB 1.665 billion through a combination of a negotiated share transfer and a partial tender offer; the share transfer closed in March 2026, making UBTech founder Jian Zhou the actual controller, and the partial tender offer completed on April 24, bringing UBTech's total stake to about 43.01%. The market speculated this might be UBTech's route back to an A-share listing, though Fenglong said UBTech would not pursue a backdoor listing within three years.","example":"","related":["UBTech Robotics","Backdoor Listing","Humanoid Robot Concept Stocks","First Listed Humanoid Robot Stock"]},{"id":"kepler-robotics","category":"company","sec":3,"tier":3,"sources":[{"title":"证券时报：上市公司入主在即，开普勒机器人发声回应","url":"https://www.stcn.com/article/detail/3921288.html"},{"title":"新京报：CEO出走后，开普勒机器人被卖给上市公司","url":"https://m.bjnews.com.cn/detail/1779279357129638.html"},{"title":"财联社：杭州柯林拟3亿元控股开普勒机器人","url":"https://www.cls.cn/detail/2375901"}],"as_of":"2026-05","related_ids":["kepler-forerunner-k2","humanoid-robot","planetary-roller-screw","full-size-humanoid-robot","mass-production","factory-pilot-deployment"],"name":"Kepler Robotics","alt":"开普勒机器人","abbr":"","aliases":["Shanghai Kepler Robotics"],"one_liner":"A Shanghai industrial-humanoid company, maker of the Forerunner K2 “Bumblebee” robot.","explanation":"Shanghai Kepler Robotics Co., Ltd. was founded in August 2023; its founder and controlling shareholder, Hua Yang, had previously founded Chunmi Technology, a company in Xiaomi's ecosystem chain. It focuses on “blue-collar” humanoid robots that can do heavy work in factories: it released the Forerunner K1 in 2023 and the Forerunner K2 in October 2024, with the commercial K2 “Bumblebee” version reportedly entering mass production in September 2025 for jobs such as material handling and high-altitude welding. Former CEO Debo Hu left in February 2026 and went on to found a separate embodied-brain company. Kepler has not yet turned a profit: in 2025 it reported revenue of about RMB 4.34 million and a net loss of about RMB 66.94 million. In May 2026, the STAR Market-listed company Hangzhou Kolion announced plans to acquire an additional 41.57% stake for up to RMB 300 million, which would give it a controlling 51% stake.","example":"The dual arms on the Forerunner K2 can carry 25-30 kg, and the company says one hour of charging gives it 8 hours of runtime.","related":["Kepler Forerunner K2","Humanoid Robot","Planetary Roller Screw","Full-size Humanoid Robot","Mass Production","Factory Pilot Deployment"]},{"id":"casbot","category":"company","sec":3,"tier":3,"sources":[{"title":"灵宝 CASBOT 关于我们","url":"https://www.casbot.tech/about"},{"title":"灵宝 CASBOT 官网","url":"https://www.casbot.tech/"}],"as_of":"2026-02","related_ids":[null,null,null,null,null],"name":"CASBOT (Beijing Zhongke Huiling Robot Technology)","alt":"中科慧灵（灵宝 CASBOT）","abbr":"","aliases":["Lingbao CASBOT","Zhongke Huiling"],"one_liner":"A Beijing humanoid-robot company whose products include the bipedal CASBOT 01/02 and the wheeled W1.","explanation":"Beijing Zhongke Huiling Robot Technology Co., Ltd. was founded in August 2023 and is headquartered at the Dongsheng Science Park in Zhongguancun, Haidian, Beijing; it markets its products under the brand “Lingbao CASBOT,” launched in October 2024. It builds full-size humanoid and embodied-AI robots: its first full-size bipedal humanoid, CASBOT 01, launched in November 2024; a second generation, CASBOT 02, followed in June 2025; and in August 2025 it released the height-adjustable wheeled robot CASBOT W1 and the lightweight dexterous hand Handle-L1. Target scenarios include commercial, cultural, and tourism venues, research and education, industrial manufacturing, and specialized field work. Its website states its angel round reached more than RMB 100 million by February 2025, with a further angel+ round of about RMB 100 million that same year, and in February 2026 it was working with partners to scale up production.","example":"","related":["Humanoid Robot","Full-size Humanoid Robot","Wheeled Humanoid Robot","Dexterous Hand","Hundred-Robot War (crowded humanoid market)"]},{"id":"kunlunxing-robotics","category":"company","sec":3,"tier":3,"sources":[{"title":"公司注册10天，估值逾10亿美元！理想智驾大牛创业（量子位，2026-03）","url":"https://www.qbitai.com/2026/03/394149.html"},{"title":"90天融了3轮，昆仑行完成数十亿元融资（投资界，2026-06）","url":"https://news.pedaily.cn/202606/565436.shtml"},{"title":"产业生态「强磁场」吸引昆仑行机器人极速落地（北京经济技术开发区，2026-06）","url":"https://kfqgw.beijing.gov.cn/ywdt/gdcyfzgd/zdxm/202606/t20260610_4694343.html"}],"as_of":"2026-06","related_ids":["li-auto-inc","autonomous-driving-talent-moving-into-embodied-ai","humanoid-robot","world-model","simplexity-robotics","funding-rounds-and-valuation"],"name":"Kunlunxing Robotics","alt":"昆仑行机器人","abbr":"","aliases":["Kunlunxing"],"one_liner":"A Beijing humanoid-robot company founded in 2026 by former Alibaba Cloud and Li Auto executives.","explanation":"Kunlunxing was registered in Beijing on March 16, 2026, based in the Beijing Economic-Technological Development Area (E-Town). Founder and CEO Geng Ren previously served as Huawei's CEO for an overseas country operation and as Alibaba's vice president and president of Alibaba Cloud's China region, later becoming president of ENN Group; co-founder and CTO Xianpeng Lang is Li Auto's former head of autonomous driving, who left after being reassigned in January 2026 to lead the company's humanoid-robot hardware effort. The company builds general-purpose humanoid robots, developing the robot body and an embodied foundation model together, positioning itself against Tesla's Optimus, and promoting technology it calls the Kunlun world model and a two-system architecture. In June 2026 it announced that in under 90 days since registration it had closed three funding rounds totaling several billion RMB, at a valuation above $1 billion, with investors including Hillhouse Ventures, Banyan Capital, Casstar, and Sinovation Ventures; as of June 2026 it had no public product yet.","example":"","related":["Li Auto Inc.","Autonomous-Driving Talent Moving into Embodied AI","Humanoid Robot","World Model","Simplexity Robotics","Funding Rounds & Valuation"]},{"id":"unix-ai","category":"company","sec":3,"tier":3,"sources":[{"title":"UniX AI 官网：关于优理奇","url":"https://www.unix-group.ai/cn/About-UNIX-AI"},{"title":"科技日报：优理奇发布新一代高性能具身智能机器人","url":"https://www.stdaily.com/web/gdxw/2026-02/15/content_474669.html"}],"as_of":"2026-02","related_ids":["wheeled-humanoid-robot","mobile-manipulation","dual-arm-robot","vision-based-tactile-sensor","real-world-deployment"],"name":"UniX AI","alt":"优理奇","abbr":"","aliases":[],"one_liner":"A Suzhou embodied-AI company known for its wheeled dual-arm humanoid, Wanda.","explanation":"UniX AI, formally UniX AI Technology (Suzhou) Co., Ltd., was founded in 2024 and is headquartered in Suzhou, Jiangsu. Founder and CEO Fengyu Yang earned his undergraduate degree at the University of Michigan and researched robot visuo-tactile perception during his PhD at Yale; Hesheng Wang, a professor at Shanghai Jiao Tong University, serves as chief scientist. The company builds its full stack in-house, developing its own core components, with products including the Wanda series of wheeled dual-arm humanoids, the bipedal humanoid Martian, and, released in February 2026, the Wanda Panther series, which pairs an 8-degree-of-freedom arm with a four-wheel, four-wheel-steer chassis. Its target settings include security, retail warehousing, hotels, guided tours, and eldercare. On funding, after several angel rounds, in December 2025 it announced two combined “angel++” and “angel+++” rounds totaling RMB 300 million.","example":"Wanda uses its two arms to tidy items and restock shelves in a hotel or micro-fulfillment warehouse, with its base carrying it between areas.","related":["Wheeled Humanoid Robot","Mobile Manipulation","Dual-arm Robot","Vision-Based Tactile Sensor","Real-world Deployment"]},{"id":"lumos-robotics","category":"company","sec":3,"tier":3,"sources":[{"title":"鹿明机器人完成数亿元A1及A2轮融资（雷峰网，2026-05-11）","url":"https://m.leiphone.com/category/industrynews/VqCT68PHdtdfvHXy.html"},{"title":"半年三轮，鹿明机器人完成天使++轮融资（投资界）","url":"https://news.pedaily.cn/202505/550373.shtml"},{"title":"清华系具身智能公司获数亿Pre-A轮融资（36氪）","url":"https://eu.36kr.com/zh/p/3582338244230272"}],"as_of":"2026-05","related_ids":["universal-manipulation-interface","fastumi","robot-free-data-collection","humanoid-robot","magiclab","joint-actuator-module"],"name":"Lumos Robotics","alt":"鹿明机器人","abbr":"","aliases":["Lumos","Shenzhen Lumos Robotics Technology"],"one_liner":"A Shenzhen humanoid-robot company that also makes FastUMI, an embodiment-free data-collection device.","explanation":"Lumos Robotics was founded in Bao'an, Shenzhen in 2024; its founder and CEO, Chao Yu, graduated from Tsinghua and previously led Dreame Technology's embodied-robotics business, and had worked on Xiaomi's CyberDog (“Tiedan”). Its products include the full-size LUS humanoid series, the heavy-payload, wheeled-arm MOS series, and components including joint actuator modules and visuotactile modules; on the data side, it is best known for FastUMI, a handheld data-collection device that follows the Universal Manipulation Interface (UMI) approach and needs no real robot, with a product family released in March 2026. It closed several angel and Pre-A rounds in 2025, then Series A1 and A2 rounds led by Mitsubishi Electric in May 2026, for close to RMB 1 billion in total funding.","example":"A data collector holds a FastUMI gripper and performs actions like opening a drawer or folding clothes in a real home; the recorded video and pose trajectory are used directly to train a robot policy.","related":["Universal Manipulation Interface","FastUMI","Robot-free (Embodiment-free) Data Collection","Humanoid Robot","MagicLab","Joint Actuator Module"]},{"id":"aheadform","category":"company","sec":3,"tier":3,"sources":[{"title":"首形科技获得新一轮数亿元A1轮融资（腾讯新闻）","url":"https://news.qq.com/rain/a/20260407A05Y5300"},{"title":"给机器人做「脸」，28岁哥大博士收获百万粉丝（科学网）","url":"https://news.sciencenet.cn/htmlnews/2025/8/550104.shtm"}],"as_of":"2026-07","related_ids":["hyper-realistic-humanoid-robot","uncanny-valley","human-robot-interaction","companion-robot","engineered-arts-ameca"],"name":"AheadForm","alt":"首形科技","abbr":"","aliases":[],"one_liner":"A Chinese startup building hyper-realistic robots with expressive, humanlike faces.","explanation":"AheadForm was founded in June 2024. Its founder, Yuhang Hu, earned his PhD at Hod Lipson's Creative Machines Lab at Columbia University, where he built the face robot Emo, which could predict and mirror a person's smile about 0.84 seconds before it happened; that work was published in Science Robotics. The company focuses on bionic faces: its own silicone skin, about 30 motors driving facial movement, plus emotion- and expression-generation models that let its robots blink, raise an eyebrow, and lip-sync, aimed at emotional interaction, tourism, and custom character IP. Its investors include Shunwei Capital and Wuyuan Capital; in April 2026 it announced a Series A1 round worth several hundred million RMB, led by Huakong Fund.","example":"In late 2025, working with NetEase's online game Nishuihan (Justice), AheadForm built a 1:1-scale bionic robot of an in-game character that could make eye contact with onlookers.","related":["Hyper-Realistic Humanoid Robot","Uncanny Valley","Human-Robot Interaction","Companion Robot","Engineered Arts Ameca"]},{"id":"vbot","category":"company","sec":3,"tier":3,"sources":[{"title":"证券时报：维他动力完成近5亿元融资","url":"https://www.stcn.com/article/detail/3902691.html"},{"title":"极客公园：对话维他动力余轶南","url":"https://www.geekpark.net/news/364058"}],"as_of":"2026-05","related_ids":["quadruped-robot","consumer-grade-robot","companion-robot","autonomous-driving-talent-moving-into-embodied-ai","mass-production","vbot-super-robot-dog"],"name":"Vbot","alt":"维他动力","abbr":"Vbot","aliases":[],"one_liner":"A consumer robot-dog company founded by former Horizon Robotics VP Yinan Yu.","explanation":"Vbot was founded in December 2024 by former Horizon Robotics vice president Yinan Yu, together with Wei Song, Horizon's former chief software-platform architect, and Zhelun Zhao, Li Auto's former director of intelligent-driving products — another representative case of talent moving from autonomous driving into embodied AI. The company set out from day one to build for home consumers; its first product, the Vbot Super Robot Dog (nicknamed “Big-Head BoBo”), is built around not needing a remote control, with autonomous following and companionship as its selling points. It was unveiled and opened for pre-order in late 2025, and after several rounds of pilot production began deliveries on May 8, 2026. That same month it closed a nearly RMB 500 million Pre-A round co-led by Oriental Fortune Capital, Huatai Zijin, and Fosun RZ Capital, bringing cumulative funding above RMB 700 million, earmarked for mass production, its sales network, and a next-generation humanoid robot.","example":"Walking Vbot in the park, it recognizes its owner on its own and follows along, no handheld remote required.","related":["Quadruped Robot","Consumer-Grade Robot","Companion Robot","Autonomous-Driving Talent Moving into Embodied AI","Mass Production","Vbot Super Robot Dog"]},{"id":"x-square-robot","category":"company","sec":4,"tier":1,"sources":[{"title":"自变量机器人官网","url":"https://www.x2robot.com/"},{"title":"智东西：自变量机器人报道","url":"https://zhidx.com/p/528318.html"}],"as_of":"2026-09","related_ids":["wall-a","wall-oss","vision-language-action-model","world-model","end-to-end"],"name":"X Square Robot","alt":"自变量机器人","abbr":"","aliases":["X Square","X-Square Robot"],"one_liner":"A Shenzhen embodied-AI company building its own end-to-end WALL model family and wheeled dual-arm robots.","explanation":"X Square Robot was founded in December 2023 and is headquartered in Shenzhen, focused on end-to-end embodied foundation models for general-purpose robots. According to the tech-news outlet Zhidongxi, founder and CEO Qian Wang holds a bachelor's and a master's degree from Tsinghua University and did robotics research during his PhD at the University of Southern California; co-founder and CTO Hao Wang previously led the large-model team's algorithms at the International Digital Economy Academy (IDEA) in the Guangdong-Hong Kong-Macao Greater Bay Area. Its models include the closed-source WALL-A series, WALL-OSS (open-sourced in September 2025), and a unified world model, WALL-B, released in 2026; its hardware is the wheeled, dual-arm “Quantum” series. It reportedly closed a RMB 1 billion Series A++ round in early 2026, with investors including ByteDance, Sequoia China, and Shenzhen Capital Group, after earlier investment from Alibaba and Meituan.","example":"X Square Robot's open-source WALL-OSS, built on the Qwen2.5-VL backbone, can turn a single instruction into reasoning, subtasks, and continuous actions in one pass.","related":["WALL-A","WALL-OSS","Vision-Language-Action Model","World Model","End-to-End"]},{"id":"spirit-ai","category":"company","sec":4,"tier":1,"sources":[{"title":"千寻智能官网新闻（发展历程）","url":"https://spirit-ai.com/news/8"},{"title":"首发 | 3个月近50亿，千寻打破具身融资纪录（投资界，2026-06）","url":"https://news.pedaily.cn/202606/564786.shtml"},{"title":"千寻智能韩峰涛：我们坚持使用「脏数据」（证券时报网，2026-07-20）","url":"https://stcn.com/article/detail/4029005.html"},{"title":"千寻智能再获10亿元融资，顺为资本和云锋基金联合领投（澎湃新闻，2026-04-07）","url":"https://m.thepaper.cn/newsDetail_forward_32915291"},{"title":"千寻智能完成15亿元A+轮融资（中国基金报，2026-06-03）","url":"https://www.chnfund.com/article/AR0ab851f9-7f8c-386e-c6c5-3a219ef0e2c5"}],"as_of":"2026-07","related_ids":["spirit-v1-5","vision-language-action-model","humanoid-robot","robochallenge","university-big-tech-autonomous-driving-founder-lineage"],"name":"Spirit AI","alt":"千寻智能","abbr":"","aliases":["Qianxun"],"one_liner":"A Hangzhou embodied-AI company building the Moz1 humanoid and the open-source Spirit VLA model.","explanation":"Spirit AI was registered in Hangzhou in January 2024. Founder and CEO Fengtao Han previously co-founded the industrial-robot maker ROKAE as its CTO; co-founder and chief scientist Yang Gao is an assistant professor at Tsinghua's IIIS with a PhD from UC Berkeley; co-founder and COO Lingyin Zheng leads commercialization. It builds both hardware and models: its humanoid, Moz1, emphasizes whole-body force control, while its VLA model, Spirit v1.5, open-sourced in January 2026, ranked first on the RoboChallenge Table30 real-robot leaderboard at release. Its data philosophy favors “messy data,” including failures and retries, over heavy cleaning. Fundraising has been intense: two rounds in February 2026 totaling nearly RMB 2 billion, another RMB 1 billion in April co-led by Shunwei Capital and Yunfeng Capital, and a RMB 1.5 billion Series A+ round in June — four rounds in roughly three months. Media often cite a cumulative “nearly RMB 5 billion,” though the disclosed amounts add up to about RMB 4.5 billion; valuation had already passed RMB 20 billion by April. It reportedly already does connector-insertion work on a CATL production line.","example":"At WAIC 2026, after being told to “tidy up the living room,” Moz1 put a can of cola in the fridge and loaded dirty bowls into the dishwasher on its own, with the screen showing how it broke the task down in real time.","related":["Spirit v1.5","Vision-Language-Action Model","Humanoid Robot","RoboChallenge","University / Big-Tech / Autonomous-Driving Founder Lineage"]},{"id":"ai2-robotics","category":"company","sec":4,"tier":1,"sources":[{"title":"具身公司智平方融资估值超200亿元（财新）","url":"https://m.caixin.com/m/2026-06-29/102458633.html"},{"title":"智平方成粤港澳大湾区首个估值200亿具身智能独角兽（电子工程专辑）","url":"https://www.eet-china.com/news/202606304967.html"}],"as_of":"2026-06","related_ids":["govla","ai2-robotics-alphabot","vision-language-action-model","wheeled-humanoid-robot","real-world-deployment"],"name":"AI² Robotics","alt":"智平方","abbr":"","aliases":["AI Squared Robotics"],"one_liner":"A Shenzhen general-purpose embodied-robot company behind the GOVLA model and the AlphaBot humanoid.","explanation":"AI² Robotics was founded in Shenzhen in early 2023. Its founder, Yandong Guo, holds a PhD in electrical and computer engineering from Purdue University, worked as a Microsoft researcher, and later served as chief scientist at XPeng Motors and OPPO. The company builds its own embodied foundation model, GOVLA (Global & Omni-body Vision-Language-Action, also called AlphaBrain), and its hardware is the AlphaBot (“Aibao”) line of wheeled dual-arm humanoid robots, aimed mainly at factory settings in semiconductors, automotive, and display-panel manufacturing; it reportedly signed a three-year, 1,000-unit order with display maker HKC. In February 2026 it closed a Series B round of more than RMB 1 billion at a valuation above RMB 10 billion, and on June 29 it announced a further series of rounds totaling nearly RMB 5 billion, at a valuation above RMB 20 billion.","example":"In a display-panel factory, AlphaBot moves material trays and loads and unloads parts, driven by the GOVLA model.","related":["GOVLA (Global & Omni-body Vision-Language-Action)","AI² Robotics AlphaBot","Vision-Language-Action Model","Wheeled Humanoid Robot","Real-world Deployment"]},{"id":"tars-robotics","category":"company","sec":4,"tier":2,"sources":[{"title":"它石智航融资1.2亿美元，创下今年最大天使轮纪录（界面新闻，2025-03-26）","url":"https://www.jiemian.com/article/12523148.html"},{"title":"6个月内15家智能家居创企估值突破100亿（36氪）","url":"https://eu.36kr.com/zh/p/3874254534710276"},{"title":"arXiv 2512.24310: World In Your Hands","url":"https://arxiv.org/abs/2512.24310"}],"as_of":"2026-07","related_ids":["tars-robotics-awe","world-in-your-hands","human-video-data","wearable-data-collection","autonomous-driving-talent-moving-into-embodied-ai"],"name":"TARS Robotics","alt":"它石智航","abbr":"TARS","aliases":["TARS"],"one_liner":"A Shanghai embodied-AI company founded by former Huawei and Baidu autonomous-driving executives, training its AWE model mainly on human data.","explanation":"TARS Robotics was founded in Shanghai in February 2025. CEO Yilun Chen previously served as CTO of Huawei's vehicle-BU autonomous-driving unit and as DJI's chief machine-vision engineer; chairman Zhenyu Li previously led Baidu's Intelligent Driving Group, overseeing Apollo and the Apollo Go robotaxi service; chief scientist Wenchao Ding was formerly a Huawei “Genius Youth” hire and a Fudan University researcher. The company's approach centers on human operation data: data collectors wear its custom capture rig while working in factories, supermarkets, and hotels, producing the WIYH dataset, used to train its general-purpose embodied foundation model, AWE (AWE 3.0 shipped in March 2026); its hardware includes the industrial A-series robots, the general-purpose T-series robots, and the TARS DexHand dexterous hand. In March 2025 it closed a $120 million angel round led by Lanchi Ventures and Qiming Venture Partners; it reportedly closed a $455 million Pre-A round in April 2026 co-led by Hillhouse and Sequoia China at a valuation of about RMB 13 billion.","example":"At WAIC 2026, TARS's booth used AWE 3.5 to drive a robot through automotive wiring-harness assembly.","related":["TARS Robotics AWE","World In Your Hands (WIYH)","Human Video Data","Wearable Data Collection","Autonomous-Driving Talent Moving into Embodied AI"]},{"id":"psibot","category":"company","sec":4,"tier":2,"sources":[{"title":"About Us - 灵初智能","url":"https://www.psibot.ai/about-us_zh"},{"title":"灵初智能完成A轮超亿美元融资（财联社）","url":"https://www.cls.cn/detail/2466860"},{"title":"灵初智能完成过亿美元融资（投资界）","url":"https://www.sohu.com/a/1068190102_122014422"}],"as_of":"2026-08","related_ids":["psi-r2","dexterous-manipulation","data-glove","human-video-data","world-model","tuopu-group"],"name":"PsiBot","alt":"灵初智能","abbr":"","aliases":["Lingchu Intelligence"],"one_liner":"A Shanghai embodied-AI company focused on dexterous-manipulation models and low-cost human data collection.","explanation":"PsiBot was founded in 2024 and is based in Xuhui, Shanghai. Founder Qibin Wang spent nearly 20 years in the phone, smart-speaker, and robotics industries; the company also runs a joint lab with Peking University, whose chief scientist, Yaodong Yang, works on reinforcement learning. Its focus is dexterous manipulation — a multi-fingered hand performing fine motor tasks — built on a two-system architecture: the policy model Psi-R2 breaks long-horizon tasks into steps and generates continuous motion, while Psi-W0 is an action-conditioned world model. On data, it follows a “human data” approach, using its Psi-SynEngine collection pipeline and SynGlove data glove to cheaply gather real-world data, already validated on sorting tasks in logistics warehouses. In March 2026 it announced angel and Pre-A rounds totaling RMB 2 billion; by August 2026 it reportedly closed a new round of over $100 million, with investors including Tuopu Group and a fund under Chery.","example":"DexKnot, from the PKU–PsiBot joint lab, taught a dexterous hand to tie a knot in a bag, aimed at packing tasks at supermarkets.","related":["Psi-R2","Dexterous Manipulation","Data Glove","Human Video Data","World Model","Tuopu Group"]},{"id":"dexmal","category":"company","sec":4,"tier":3,"sources":[{"title":"新京报：原力灵机完成2亿元天使轮融资","url":"https://m.bjnews.com.cn/detail/1742906053129898.html"},{"title":"新浪科技：Dexmal原力灵机融资近10亿元，阿里巴巴、蔚来资本分别领投","url":"https://finance.sina.com.cn/tech/digi/2025-11-14/doc-infxiwez1707401.shtml"},{"title":"瑞财经：Dexmal原力灵机两轮融资近10亿元","url":"https://m.rccaijing.com/news-7395014056673473687.html"}],"as_of":"2026-07","related_ids":["dm0","dexbotic","robochallenge","vision-language-action-model","hugging-face","open-source-hardware"],"name":"Dexmal","alt":"原力灵机","abbr":"","aliases":["Dexmal Yuanli Lingji","Yuanli Lingji (Chongqing) Intelligent Technology"],"one_liner":"An embodied-AI company founded by Megvii co-founder Wenbin Tang, building VLA models, open tools, and benchmarks.","explanation":"Dexmal was founded in March 2025, registered in Chongqing. Its founder and CEO, Wenbin Tang, came out of Tsinghua's elite Yao Class and co-founded and served as CTO of the computer-vision company Megvii; core team members Haoqiang Fan, Erjin Zhou, and Tiancai Wang also came from Megvii. It closed a RMB 200 million angel round the same month it was founded, and in November 2025 announced a Series A (led by NIO Capital) and Series A+ (exclusively from Alibaba) totaling nearly RMB 1 billion. It has built several pieces of industry infrastructure: the open-source VLA toolbox Dexbotic (VLA meaning vision-language-action model), the open-source hardware platform DOS-W1, and, together with Hugging Face, the real-robot benchmark platform RoboChallenge. In February 2026 it released the open-source model DM0 jointly with StepFun, followed by DM0.5 in July. Commercially, it started with logistics and warehousing.","example":"RoboChallenge has every team deploy its model onto the same real robots to run the same set of tasks, comparing VLA models by real-robot success rate rather than simulation scores.","related":["DM0","Dexbotic (Dexmal VLA toolbox)","RoboChallenge","Vision-Language-Action Model","Hugging Face","Open-Source Hardware (OSHW)"]},{"id":"ace-robotics","category":"company","sec":4,"tier":3,"sources":[{"title":"中国日报网：首创ACE具身研发范式，大晓机器人构建具身智能开放新生态","url":"https://cn.chinadaily.com.cn/a/202512/19/WS69450572a310942cc49978f4.html"},{"title":"人民网上海：大晓机器人完成天使+轮融资","url":"http://sh.people.com.cn/n2/2026/0616/c176738-41611926.html"},{"title":"百度百科：大晓机器人","url":"https://baike.baidu.com/item/%E5%A4%A7%E6%99%93%E6%9C%BA%E5%99%A8%E4%BA%BA/67054231"}],"as_of":"2026-07","related_ids":["kairos","world-model","robot-brain-company","one-brain-multiple-robots","on-device-model","agibot"],"name":"ACE Robotics","alt":"大晓机器人","abbr":"","aliases":["SenseTime Daxiao","Daxiao Robotics"],"one_liner":"A Shanghai robot-brain company, led by SenseTime co-founder Xiaogang Wang, built around a world model.","explanation":"ACE Robotics (Daxiao Robotics) was founded in Shanghai in December 2025, chaired by Xiaogang Wang, a co-founder and executive director of SenseTime, with Dacheng Tao, a fellow of the Australian Academy of Science, as chief scientist; a SenseTime-affiliated fund is among its existing investors. Rather than mainly building complete robots, it focuses on the “brain”: its core product is the Kairos world model, which first learns to predict how the world will change after an action, then uses that prediction to decide what to do. It open-sourced version 3.0 the same month it was founded, and also released A1, an embodied-brain module that can be installed on other companies' robot dogs and humanoids. In June 2026 it announced an angel+ round, reportedly raising several hundred million dollars in total in the first half of the year; in July it said it planned to deploy in 1,000 retail stores within a year, partnering with companies including Galbot and AgiBot.","example":"During a media visit, a humanoid robot running the 4-billion-parameter on-device Kairos-4B model, with no cloud connection, autonomously watered plants and poured cereal from the fridge into a bowl.","related":["Kairos (ACE Robotics)","World Model","Robot-Brain (Model-Only) Company","One Brain, Multiple Robots","On-device Model","AgiBot"]},{"id":"noematrix","category":"company","sec":4,"tier":3,"sources":[{"title":"具身智能领域再掀波澜！穹彻智能完成Pre-A轮融资 | 穹彻智能","url":"https://www.noematrix.ai/news/noematrix_pre-a"},{"title":"穹彻智能完成A轮融资 | 穹彻智能","url":"https://www.noematrix.ai/news/Noematrix-A++"},{"title":"上海穹彻智能科技有限公司_百度百科","url":"https://baike.baidu.com/item/%E4%B8%8A%E6%B5%B7%E7%A9%B9%E5%BD%BB%E6%99%BA%E8%83%BD%E7%A7%91%E6%8A%80%E6%9C%89%E9%99%90%E5%85%AC%E5%8F%B8/64871743"}],"as_of":"2026-08","related_ids":["flexiv-robotics","sjtu-mvig-lab","embodied-foundation-model","world-action-model","robot-free-data-collection","hybrid-force-position-control"],"name":"Noematrix","alt":"穹彻智能","abbr":"","aliases":["Shanghai Noematrix Technology"],"one_liner":"An embodied-AI “brain” company incubated by Flexiv Robotics and co-founded by Cewu Lu.","explanation":"Noematrix was founded in Shanghai in November 2023, strategically incubated by the force-controlled robotics company Flexiv Robotics; its co-founders include Shanghai Jiao Tong University professor Cewu Lu (who did a postdoc at Stanford's AI Lab) and Flexiv founder Shiquan Wang. Its core product is the general-purpose embodied brain Noematrix Brain (version 2.0 released in July 2025), along with a toolchain that includes CoMiner, a companion data-collection system, emphasizing pretraining on real-world scene data followed by post-training with hybrid force/position control; its robots are already deployed at scale in pharmacies. In February 2026 it closed a Series A of several hundred million RMB led by C Capital, followed in June by strategic investment from institutions including SJTU's AI Future Fund, and in August it released a technical preview of its embodied world-action model, Noe-0.","example":"A pharmacy medicine-picking robot: it needs no changes to the existing shelving, fits in about 2.5 square meters, and connects directly to the store's order system.","related":["Flexiv Robotics","SJTU MVIG Lab","Embodied Foundation Model","World Action Model","Robot-free (Embodiment-free) Data Collection","Hybrid Force/Position Control"]},{"id":"beingbeyond","category":"company","sec":4,"tier":3,"sources":[{"title":"BeingBeyond 官网","url":"https://www.beingbeyond.com/"},{"title":"BeingBeyond GitHub","url":"https://github.com/BeingBeyond"},{"title":"Being-H0 (arXiv 2507.15597)","url":"https://arxiv.org/abs/2507.15597"}],"as_of":"2026-03","related_ids":["being-h0",null,null,null,"dexumi"],"name":"BeingBeyond","alt":"智在无界","abbr":"","aliases":[],"one_liner":"A Beijing embodied-foundation-model startup whose Being-H series learns dexterous manipulation from videos of human hands.","explanation":"BeingBeyond is an embodied-AI foundation model company based in Haidian, Beijing, founded by a team from Peking University led by Zongqing Lu (reportedly in 2025), who is also the corresponding author on most of its papers. Its core idea is “human-centric”: pretraining on large amounts of human-hand video and motion-capture data, then transferring that to robot dexterous hands and humanoids. Its model lineup includes the Being-H series for dexterous-hand manipulation (Being-H0 was accepted to ICML 2026, followed by H0.5 and H0.7), Being-M for whole-body motion generation, the multimodal model Being-VL, and the humanoid whole-body control framework BumbleBee (a NeurIPS 2025 Spotlight). On the hardware side it has the D1 desktop dexterous arm and the Being-Actor humanoid whole-body teleoperation system, and in March 2026 it released the data-collection device U1.","example":"Being-H0 is first pretrained on about 1,100 hours of human-hand data called UniHand, then fine-tuned on a small amount of real-robot data to transfer to a dexterous hand for grasping and manipulation.","related":["Being-H0","Pretraining on Human Videos","Dexterous Manipulation","Cross-Embodiment","DexUMI"]},{"id":"xingyuanzhi-robotics","category":"company","sec":4,"tier":3,"sources":[{"title":"星源智机器人完成Pre-A轮融资，10个月累计融资10亿元（新京报）","url":"https://m.bjnews.com.cn/detail/1780455817129694.html"},{"title":"成立10个月累计融资10亿 智源系星源智（科创板日报）","url":"https://www.cls.cn/detail/2391282"}],"as_of":"2026-06","related_ids":["beijing-academy-of-artificial-intelligence","robobrain","robot-brain-company","on-device-edge-deployment","world-model"],"name":"Xingyuanzhi Robotics","alt":"星源智","abbr":"","aliases":["Xingyuanzhi"],"one_liner":"A BAAI-incubated embodied-brain company in Beijing pursuing a hardware-software-integrated, on-device deployment approach.","explanation":"Xingyuanzhi Robotics was founded in Beijing on August 1, 2025, incubated by the Beijing Academy of Artificial Intelligence (BAAI). Founder and CEO Dong Liu previously served as general manager of JD's autonomous-driving unit; co-founder Yadong Mu is a researcher at Peking University and a BAAI scholar. The company does not build robot bodies; instead it makes an “embodied brain” — an upper-level model that understands a scene and plans a task — that can be installed on different robots, following a hardware-software-integrated, on-device deployment approach that runs the large model directly on the robot itself rather than depending on the cloud. Within a month of founding it closed a RMB 200 million angel round, with investors including Casstar, Hillhouse, and Yuanhe Origin Capital; on June 3, 2026, it closed a Pre-A round, bringing cumulative funding to about RMB 1 billion within ten months of founding, aimed at embodied-brain and world-model R&D and mass production.","example":"","related":["Beijing Academy of Artificial Intelligence","RoboBrain","Robot-Brain (Model-Only) Company","On-Device / Edge Deployment","World Model"]},{"id":"five-ages","category":"company","sec":4,"tier":3,"sources":[{"title":"BridgeVLA (arXiv 2506.07961, affiliations incl. FiveAges)","url":"https://arxiv.org/html/2506.07961"},{"title":"GitHub: fiveages-sim/open-deploy-ws","url":"https://github.com/fiveages-sim/open-deploy-ws"}],"as_of":"2025-07","related_ids":[null,null,null,null,null],"name":"Five Ages","alt":"中科第五纪","abbr":"","aliases":["FiveAges"],"one_liner":"A Chinese embodied-AI startup reportedly closely connected to the Institute of Automation, Chinese Academy of Sciences.","explanation":"Five Ages is a Beijing-based embodied-AI startup. It reportedly has close ties to the Institute of Automation, Chinese Academy of Sciences (CASIA): the June 2025 paper on the 3D manipulation model BridgeVLA lists FiveAges as an author affiliation alongside CASIA and ByteDance Seed. On GitHub, under the name fiveages-sim, the company has open-sourced a set of engineering code, including ROS 2 deployment workspaces for dual-arm robots (such as the Dobot CR5 and ARX robot arms), Isaac Sim scene assets, and reinforcement-learning examples, with a focus leaning toward robot-arm manipulation models and sim-to-real deployment. Public information about its founding date and funding is limited, so this entry does not go into further detail.","example":"","related":["BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models","Robot-Brain (Model-only) Company","NVIDIA Isaac Sim","Robot Operating System 2","Dobot"]},{"id":"wujie-dynamics","category":"company","sec":4,"tier":3,"sources":[{"title":"36氪：无界动力完成超2亿美元天使轮融资","url":"https://m.36kr.com/p/3869370059035913"},{"title":"新浪财经：无界动力获3亿元天使融资","url":"https://finance.sina.com.cn/stock/hkstock/hkzmt/2025-11-11/doc-infwzrrv1941527.shtml"}],"as_of":"2026-07","related_ids":["latent-world-model","reinforcement-learning","world-model","manipulation","autonomous-driving-talent-moving-into-embodied-ai","braincerebellum-architecture"],"name":"Wujie Dynamics","alt":"无界动力","abbr":"","aliases":[],"one_liner":"A Beijing company founded by former Horizon Robotics VP Yufeng Zhang, building a general-purpose robot brain.","explanation":"Wujie Dynamics was founded in Beijing in March 2025. Founder and CEO Yufeng Zhang previously worked in R&D management at Sony and ARM before joining Horizon Robotics in 2017, where he rose to vice president and president of its smart-vehicle business unit. The company focuses on a general-purpose “brain” and “manipulation intelligence” for robots — getting a robot's hands to reliably grasp, assemble, and complete other tasks — using a technical approach that combines a latent-space world model, which predicts how the environment will change within a compressed feature space, with reinforcement learning; its self-developed robot body, the K15, has already entered mass production, starting with combined hardware-and-software solutions for industrial and commercial settings. The company states its global order book totals nearly $100 million. Fundraising has moved fast: in June 2026 it announced an angel round of more than $200 million, with a Pre-A round of nearly $200 million close to completion, backed by investors including Sequoia China, Hillhouse Ventures, and JD.com.","example":"","related":["Latent World Model","Reinforcement Learning","World Model","Manipulation","Autonomous-Driving Talent Moving into Embodied AI","Brain–Cerebellum Architecture"]},{"id":"simplexity-robotics","category":"company","sec":4,"tier":3,"sources":[{"title":"腾讯和阿里同时押注具身智能创企，至简动力5轮累计融资20亿元（21经济网，2026-03-09）","url":"https://www.21jingji.com/article/20260309/herald/4faae25f2c3c91bb0a84c3942cc445b2.html"},{"title":"至简动力半年完成5轮融资累计20亿元（每日商报，2026-03-11）","url":"https://mdaily.hangzhou.com.cn/mrsb/2026/03/11/article_detail_3_20260311A078.html"},{"title":"别人还在叠衣服，至简动力的100台机器人已下车间（极客公园，2026-07-11）","url":"https://www.geekpark.net/news/367186"}],"as_of":"2026-07","related_ids":["autonomous-driving-talent-moving-into-embodied-ai","university-big-tech-autonomous-driving-founder-lineage","funding-rounds-and-valuation","machine-tending","real-world-deployment"],"name":"Simplexity Robotics","alt":"至简动力","abbr":"","aliases":[],"one_liner":"An embodied-AI company founded by former Li Auto autonomous-driving executives, building both models and robot bodies.","explanation":"Simplexity Robotics was founded in late July 2025, reportedly headquartered in Hangzhou with additional offices in Beijing, Shanghai, and Suzhou. Its three founders all came from Li Auto: chairman Kai Wang is Li Auto's former CTO, CEO Peng Jia is its former head of autonomous-driving technology, and COO Jiajia Wang is its former head of autonomous-driving mass production — a typical example of a team moving from autonomous driving into embodied AI. The company works on foundation models, a data loop, and robot bodies together, aiming first at enclosed settings such as factory floors, supermarkets, and logistics. As of March 2026 it had reportedly closed five funding rounds in under six months, totaling about RMB 2 billion, at a valuation above $1 billion, with investors including Sequoia China, Lanchi Ventures, Tencent, and Alibaba. On July 6, 2026, it announced in Suzhou that its first full-scenario robot, the i7 Pro, had completed its first batch delivery of 100 units.","example":"In July 2026, Simplexity Robotics held a delivery ceremony in Suzhou announcing that 100 i7 Pro units were entering factory floors, and it built a production line where an embodied robot loads and unloads CNC machine tools.","related":["Autonomous-Driving Talent Moving into Embodied AI","University / Big-Tech / Autonomous-Driving Founder Lineage","Funding Rounds & Valuation","Machine Tending","Real-world Deployment"]},{"id":"morphi","category":"company","sec":4,"tier":3,"sources":[{"title":"成立半年估值超70亿，墨奇智能刷新国内具身智能首轮融资规模纪录（凤凰网财经）","url":"https://finance.ifeng.com/c/8uZNoCWQZF8"},{"title":"成立半年估值超70亿，墨奇智能创国内具身智能天使轮融资纪录（搜狐/财闻）","url":"https://www.sohu.com/a/1046673669_122014422"}],"as_of":"2026-07","related_ids":["autonomous-driving-talent-moving-into-embodied-ai","university-big-tech-autonomous-driving-founder-lineage","data-flywheel","household-tasks","service-robot","funding-rounds-and-valuation"],"name":"Morphi","alt":"墨奇智能","abbr":"","aliases":["Moqi Intelligence"],"one_liner":"An embodied-AI company co-founded by a former Huawei autonomous-driving lead, aiming at a general home robot.","explanation":"Morphi (Moqi Intelligence) was founded in late 2025 by CTO Qingqiu Huang and CEO Wenli Gao. Huang has an undergraduate degree from Tsinghua's automation department and a PhD from the Chinese University of Hong Kong's Multimedia Lab, and previously led the AI division of Huawei's autonomous-driving vehicle unit; Gao co-founded the cross-border logistics company iMile. The company wants to bring the data loops, on-device model optimization, and mass-production quality control proven in autonomous driving over to robotics, planning to first deploy in commercial settings such as hotel service and last-mile delivery to gather real-world data, before moving into the home. In July 2026 it announced a Series A round of angel-stage funding exceeding RMB 1 billion, led by Alibaba and Tencent, at a reported post-money valuation above RMB 7 billion.","example":"In July 2026, Morphi's angel round of more than RMB 1 billion, led by Alibaba and Tencent, set a record for China's largest publicly disclosed first funding round in embodied AI.","related":["Autonomous-Driving Talent Moving into Embodied AI","University / Big-Tech / Autonomous-Driving Founder Lineage","Data Flywheel","Household Tasks","Service Robot","Funding Rounds & Valuation"]},{"id":"pragmatik-labs","category":"company","sec":4,"tier":3,"sources":[{"title":"刚刚，林俊旸官宣创业公司Pragmatik Labs（新浪科技，2026-08）","url":"https://finance.sina.com.cn/tech/roll/2026-08-12/doc-inimzcpz4117680.shtml"},{"title":"上海抢走林俊旸：首轮融资阵容出炉（投资界，2026-08）","url":"https://news.pedaily.cn/202608/567608.shtml"},{"title":"林俊旸新公司卜拉格工商信息与首轮估值（量子位，2026-06）","url":"https://www.qbitai.com/2026/06/436138.html"}],"as_of":"2026-08","related_ids":["alibaba-group","qwen-vl","qwen-robot-series","embodied-agent","world-model"],"name":"Pragmatik Labs","alt":"语用科技","abbr":"p7k","aliases":["p7k"],"one_liner":"An AI-agent company founded by former Qwen tech lead Junyang Lin, with embodied AI among its directions.","explanation":"Pragmatik Labs was founded by Junyang Lin, the former technical lead of Alibaba's Qwen large-model team. Born in 1993, Lin holds a master's degree from Peking University, became Qwen's technical lead in late 2022 — Alibaba's youngest P10-level employee — and left in March 2026. Between May and June he registered several entities in Shanghai, including Pragmatik (Shanghai) Technology and Shanghai Pragmatik Technology, which is why early reports referred to the company as “Pragmatik”; it officially launched on August 12. The company researches next-generation agents spanning both the digital and physical worlds: digital agents for knowledge work and business operations, and physical agents — that is, embodied AI — for completing long-horizon tasks in the real world. Reports say its first funding round totaled several hundred million dollars, with Banyan Capital and Sequoia China each leading roughly $100 million and Tencent contributing about $20 million, at a post-money valuation of about $2 billion; no model or product had been released at the time of the announcement.","example":"","related":["Alibaba Group","Qwen-VL","Qwen-Robot Series","Embodied Agent","World Model"]},{"id":"xiaomi","category":"company","sec":5,"tier":1,"sources":[{"title":"Xiaomi - Wikipedia","url":"https://en.wikipedia.org/wiki/Xiaomi"},{"title":"Xiaomi-Robotics-0 (arXiv 2602.12684)","url":"https://arxiv.org/abs/2602.12684"}],"as_of":"2026-07","related_ids":["xiaomi-robotics-0","mimo-embodied","xiaomi-cyberdog","xiaomi-cyberone","xiaomi-cybergear-micro-motor","automakers-entering-humanoid-robotics"],"name":"Xiaomi","alt":"小米","abbr":"","aliases":["Xiaomi Robotics","Xiaomi Corporation"],"one_liner":"A Chinese consumer-electronics company making phones and EVs that also builds robot dogs, a humanoid, and open VLA models.","explanation":"Xiaomi was founded in Beijing in April 2010 by Lei Jun and others, listed on the Hong Kong Stock Exchange in 2018, and has expanded from smartphones into home appliances and electric vehicles. On the robotics side, it released the open-source quadruped robot CyberDog in 2021, showed the humanoid CyberOne in 2022, and launched the CyberGear integrated joint motor, priced at RMB 499, in 2023. Starting in 2025, its embodied-AI team began releasing models: in November 2025 it open-sourced MiMo-Embodied, a vision-language model built for both robotics and autonomous driving; in February 2026 it open-sourced Xiaomi-Robotics-0, a 4.7-billion-parameter VLA model designed to reason while acting so its motion stays smooth; and in July 2026 it released Xiaomi-Robotics-1, trained on more than 100,000 hours of real-robot data.","example":"Xiaomi-Robotics-0 runs with roughly 80 ms of inference latency on an RTX 4090, using asynchronous execution to keep the robot arm's motion smooth.","related":["Xiaomi-Robotics-0","MiMo-Embodied (Xiaomi)","Xiaomi CyberDog","Xiaomi CyberOne","Xiaomi CyberGear Micro-Motor","Automakers Entering Humanoid Robotics"]},{"id":"xpeng","category":"company","sec":5,"tier":1,"sources":[{"title":"XPeng - Wikipedia","url":"https://en.wikipedia.org/wiki/XPeng"},{"title":"CnEVPost: XPeng unveils next-gen IRON humanoid robot","url":"https://cnevpost.com/2025/11/05/xpeng-unveils-next-gen-iron-humanoid-robot/"}],"as_of":"2026-06","related_ids":["xpeng-iron","automakers-entering-humanoid-robotics","humanoid-robot","vision-language-action-model","autonomous-driving-talent-moving-into-embodied-ai","tesla"],"name":"XPeng","alt":"小鹏汽车","abbr":"","aliases":["XPeng Motors","XPeng Robotics"],"one_liner":"A Guangzhou smart-EV maker building the IRON humanoid robot, a leading example of automakers entering humanoid robotics.","explanation":"XPeng is a smart electric-vehicle company founded in Guangzhou in 2014; its chairman, He Xiaopeng, previously founded the browser company UCWeb and later served as a senior Alibaba executive. It listed on the New York Stock Exchange in 2020 and took a secondary primary listing on the Hong Kong Stock Exchange in 2021. XPeng is extending the vision systems, chips, and large models it built for autonomous driving into robotics: it first unveiled the humanoid robot IRON in 2024, then a new generation in November 2025 with 82 degrees of freedom across the body, three of XPeng's own Turing AI chips, and XPeng's own VLA model. It is targeting mass production by the end of 2026, initially for commercial roles such as store greeting and guided tours, and is also working with Baosteel to pilot industrial inspection. He Xiaopeng reportedly announced in June 2026 that he would personally lead the robotics business.","example":"The new-generation IRON stands about 1.78 meters tall with 22 degrees of freedom in each hand, and is planned to first work as a greeter in XPeng's own stores.","related":["XPeng IRON","Automakers Entering Humanoid Robotics","Humanoid Robot","Vision-Language-Action Model","Autonomous-Driving Talent Moving into Embodied AI","Tesla"]},{"id":"byd-company-limited","category":"company","sec":5,"tier":2,"sources":[{"title":"比亚迪：「人形机器人代号尧舜禹」等消息均不属实（IT之家，2026-06）","url":"https://www.ithome.com/0/960/816.htm"},{"title":"BYD unveils 5.2-feet-tall Xiao Di humanoid robot on showroom floors in China (Interesting Engineering, 2026-08)","url":"https://interestingengineering.com/ai-robotics/byd-xiao-di-humanoid-robot-china"},{"title":"比亚迪全球招聘具身智能人才（量子位，2024-12）","url":"https://www.qbitai.com/2024/12/238878.html"}],"as_of":"2026-08","related_ids":["automakers-entering-humanoid-robotics","humanoid-robot","xpeng","li-auto-inc","tesla","autonomous-driving-talent-moving-into-embodied-ai"],"name":"BYD Company Limited","alt":"比亚迪","abbr":"BYD","aliases":["BYD"],"one_liner":"The world's top-selling EV maker, which unveiled the humanoid robot “Xiaodi” in 2026 for use in showrooms.","explanation":"BYD was founded in Shenzhen in 1995 by Chuanfu Wang and others, starting in batteries before entering the auto industry in 2003, and is now the world's top-selling new-energy vehicle maker. According to its December 2024 campus-recruiting materials, its embodied-AI research team was formed in 2022 and had already built process, collaborative, and mobile robots for its own factories, and was hiring to work on humanoid, bipedal, and quadruped robots. In May 2026, executive vice president Ke Li publicly confirmed for the first time that BYD was developing a humanoid robot, saying she hoped every dealership could have two or three for greeting customers and explaining models; in June, online reports of a project codenamed “Yao Shun Yu” and plans for 20,000 units of internal use within the year were denied by BYD. In August, its first humanoid robot, “Xiaodi,” was unveiled publicly at the “Di Space” experience center in Zhengzhou, standing 1.61 meters tall with 31 degrees of freedom across its body, initially used for greeting customers, explaining models, and demonstrating in-car features.","example":"As Ke Li described it, an in-store humanoid robot would handle greeting, model explanations, and demonstrating car-infotainment features alongside sales staff, rather than replacing them.","related":["Automakers Entering Humanoid Robotics","Humanoid Robot","XPeng","Li Auto Inc.","Tesla","Autonomous-Driving Talent Moving into Embodied AI"]},{"id":"hyundai-motor-group","category":"company","sec":5,"tier":2,"sources":[{"title":"Hyundai Motor Group - Wikipedia","url":"https://en.wikipedia.org/wiki/Hyundai_Motor_Group"},{"title":"Atlas (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Atlas_(robot)"},{"title":"Boston Dynamics News","url":"https://bostondynamics.com/news/"}],"as_of":"2026-01","related_ids":["boston-dynamics","boston-dynamics-atlas-2","automakers-entering-humanoid-robotics","gemini-robotics","softbank-group","humanoid-robot"],"name":"Hyundai Motor Group","alt":"现代汽车集团","abbr":"","aliases":["Hyundai"],"one_liner":"A South Korean automaker group, Boston Dynamics' controlling shareholder and first major customer.","explanation":"Hyundai Motor Group took its current form in 1998 when Hyundai Motor acquired Kia, headquartered in Seoul, South Korea, and today includes Hyundai, Kia, Genesis, and Hyundai Mobis; by production volume it was the world's third-largest automaker group in 2023. In June 2021 it completed the purchase of about 80% of Boston Dynamics from SoftBank, bringing it into the robotics business. At CES 2026, the group unveiled its AI robotics strategy, and Boston Dynamics showed a production version of the electric Atlas built for automotive assembly, planned for deployment in 2028 at the Hyundai Motor Group Metaplant America (HMGMA) in Georgia, with Atlas also to be paired with Google's Gemini Robotics model. It is a textbook case of an automaker entering humanoid robotics: both a shareholder and the first major customer.","example":"Boston Dynamics opened an applications center at the HMGMA campus, letting Atlas practice sorting automotive parts in a real factory environment.","related":["Boston Dynamics","Boston Dynamics Atlas (Electric)","Automakers Entering Humanoid Robotics","Gemini Robotics","SoftBank Group","Humanoid Robot"]},{"id":"honor","category":"company","sec":5,"tier":2,"sources":[{"title":"荣耀手机公司入局人形机器人，百万年薪抢人（36氪）","url":"https://eu.36kr.com/zh/p/3697024553873025"},{"title":"HONOR Advances Its AI Vision at MWC 2026 with Robot Phone, Humanoid Robot","url":"https://www.honor.com/global/news/honor-mwc2026-launch/"},{"title":"荣耀机器人包揽亦庄「半马」前三名（北京日报，2026-04）","url":"https://news.bjd.com.cn/2026/04/19/11697522.shtml"},{"title":"Ratified: world records for Kiplimo, Tharp and Wanyonyi（World Athletics，2026-09-03）","url":"https://worldathletics.org/news/press-releases/ratified-world-records-kiplimo-tharp-wanyonyi"}],"as_of":"2026-07","related_ids":["honor-lightning-humanoid-robot","humanoid-robot-half-marathon","world-humanoid-robot-games","joint-actuator-module","liquid-cooled-joint-actuators"],"name":"HONOR","alt":"荣耀","abbr":"","aliases":["Honor Device"],"one_liner":"A Shenzhen phone maker that has bet on robotics since 2025, with its self-developed humanoid “Lightning” winning a half marathon.","explanation":"HONOR was founded in 2013 as a phone sub-brand under Huawei, becoming independent in November 2020, and is headquartered in Shenzhen. In March 2025 it announced its “Alpha Strategy,” pledging $10 billion over five years to transform into an AI-device ecosystem company, with robotics as one pillar; that April it formed a New Industry Incubation department covering labs for embodied AI, embodied data, powertrains, and bionic robot bodies. On March 1, 2026, at Mobile World Congress it showed a “robot phone” with a mechanical gimbal and unveiled its first self-developed humanoid robot; on April 19, its self-developed humanoid “Lightning” won the Beijing E-Town Humanoid Robot Half Marathon in 50 minutes 26 seconds, with HONOR's three entries sweeping the top three places. The company is pursuing an A-share listing, completing listing-counseling registration in June 2025.","example":"At the E-Town half marathon, “Lightning” ran the roughly 21-kilometer course in fully autonomous navigation mode, finishing in a net time of 50 minutes 26 seconds — faster than the human men's half-marathon world record at the time, 57 minutes 20 seconds, set by Ugandan runner Jacob Kiplimo in Lisbon in March 2026.","related":["Honor Lightning Humanoid Robot","Humanoid Robot Half Marathon (Beijing E-Town)","World Humanoid Robot Games","Joint Actuator Module","Liquid-Cooled Joint Actuators (Active Thermal Management)"]},{"id":"dji","category":"company","sec":5,"tier":2,"sources":[{"title":"DJI - Wikipedia","url":"https://en.wikipedia.org/wiki/DJI"}],"as_of":"2026-02","related_ids":["robomaster-robocon-university-robotics-competitions","livox","livox-mid-360","unmanned-aerial-vehicle","robot-vacuum-cleaner","lidar"],"name":"DJI","alt":"大疆创新","abbr":"DJI","aliases":["大疆 (colloquial Chinese short form)","Shenzhen DJI Sciences and Technologies Ltd."],"one_liner":"The Shenzhen consumer-drone leader, whose portfolio also includes the RoboMaster competition and Livox lidar.","explanation":"DJI was founded in Shenzhen in 2006 by Frank Wang, who began prototyping drone flight controllers in his dorm room while studying at HKUST; the company is headquartered in Shenzhen's Nanshan district. DJI built its business on consumer drones, holding more than 90% of the global consumer drone market as of June 2024, and its products also include the Osmo handheld camera line, Ronin stabilizers, agricultural drones, and the RoboMaster educational robot. Its overlap with embodied AI comes mainly through three things: it hosts the RoboMaster robotics competition for university students, which has trained a large number of robotics engineers; it incubated the lidar company Livox, whose Mid-360 sensor is used by many humanoid and quadruped robots; and in 2025 it launched the robot vacuum Romo, entering home robotics. On the U.S. side, the FCC banned the import and sale of new DJI drones in December 2025, and DJI sued the U.S. government starting in February 2026.","example":"The Livox Mid-360 lidar used on the Unitree G1 humanoid's head is the same sensor DJI's spin-off makes.","related":["RoboMaster / ROBOCON University Robotics Competitions","Livox","Livox Mid-360","Unmanned Aerial Vehicle (UAV)","Robot Vacuum Cleaner","LiDAR"]},{"id":"samsung-electronics","category":"company","sec":5,"tier":3,"sources":[{"title":"Rainbow Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Rainbow_Robotics"},{"title":"Samsung advances in-house humanoid robot development with cost edge","url":"https://interestingengineering.com/ai-robotics/samsungs-humanoid-robot-with-lower-costs"}],"as_of":"2026-07","related_ids":["rainbow-robotics-rb-y1","rainbow-robotics","humanoid-robot","wheeled-humanoid-robot","actuator"],"name":"Samsung Electronics","alt":"三星电子（机器人业务）","abbr":"","aliases":["Samsung"],"one_liner":"South Korea's electronics giant, which took control of Rainbow Robotics in 2025 and is investing in humanoids and physical AI.","explanation":"Samsung Electronics is South Korea's largest electronics company, with core businesses in memory chips, phones, and home appliances. Its key move in robotics has been investing in Rainbow Robotics, founded by KAIST's humanoid HUBO team: it announced in late 2024 that it would raise its stake to about 35%, becoming the largest shareholder, completed the takeover after regulatory approval in March 2025, and set up a Future Robotics Business Team reporting directly to the CEO, led by Rainbow founder Jun-ho Oh. Rainbow's wheeled dual-arm robot RB-Y1 has reportedly been tested in logistics settings; Samsung is also developing humanoid robots internally and plans to use its home-appliance motor technology in joint actuators to cut costs. Samsung reportedly formed an RX division in July 2026 to coordinate its robotics business.","example":"Rainbow Robotics' RB-Y1 wheeled dual-arm robot has reportedly been tested in Samsung's logistics operations.","related":["Rainbow Robotics RB-Y1","Rainbow Robotics","Humanoid Robot","Wheeled Humanoid Robot","Actuator"]},{"id":"lg-electronics","category":"company","sec":5,"tier":3,"sources":[{"title":"LG Acquires Majority Stake in Bear Robotics to Bolster Robotics Capabilities（LG 官网）","url":"https://www.lg.com/global/newsroom/news/corporate/lg-acquires-majority-stake-in-bear-robotics-to-bolster-robotics-capabilities/"},{"title":"LG's charge into humanoid robotics focuses on joints, eyes（The Korea Herald）","url":"https://www.koreaherald.com/article/10701560"},{"title":"LG moves deeper into humanoid supply chain（The Korea Herald）","url":"https://www.koreaherald.com/article/10808594"},{"title":"LG humanoid set for 2027 debut with Nvidia AI brain（The Korea Herald）","url":"https://www.koreaherald.com/article/10841656"}],"as_of":"2026-08","related_ids":["lg-cloid","joint-actuator-module","household-tasks","service-robot","samsung-electronics","consumer-electronics-show"],"name":"LG Electronics","alt":"LG 电子","abbr":"LGE","aliases":["LGE","LG"],"one_liner":"A South Korean appliance giant building the home humanoid CLoiD and its AXIUM joint-actuator brand.","explanation":"LG Electronics traces back to GoldStar, founded by In-hwoi Koo in 1958, is headquartered in Seoul, and is known for appliances and TVs. Its robotics business started years ago under the CLOi brand, for commercial robots such as guides and delivery units; in March 2024 it invested $60 million for a stake in Silicon Valley delivery-robot company Bear Robotics, raising its stake to a controlling 51% by January 2025 and folding CLOi's commercial-robot business into it. In January 2026 at CES it unveiled the wheeled home humanoid CLoiD and a joint-actuator brand called AXIUM, integrating a motor, driver, and reducer into one unit; that March its CEO told shareholders that 2026 would be “year one” of its humanoid-robot business, with plans to build out AXIUM's mass-production system within the year, first for use in CLoiD, and to supply it externally starting in 2027; reports say CLoiD will first be used for material handling in LG's own factories. In August, LG Group signed an agreement with NVIDIA, planning to launch a bipedal humanoid built on the Jetson Thor chip in the first quarter of 2027.","example":"At CES 2026, CLoiD was shown taking milk out of the fridge, putting a croissant in the oven, and folding laundry.","related":["LG CLOiD","Joint Actuator Module","Household Tasks","Service Robot","Samsung Electronics","Consumer Electronics Show"]},{"id":"li-auto-inc","category":"company","sec":5,"tier":3,"sources":[{"title":"理想汽车内部会曝光：必做人形机器人（36氪，2026-01）","url":"https://www.36kr.com/p/3658883629212293"},{"title":"21独家｜理想汽车将在今年年内发布一款双轮机器人（21世纪经济报道，2026-03）","url":"https://www.21jingji.com/article/20260305/herald/c18a1aac9354fceb783a07969c1e9164.html"},{"title":"具身智能「上下半场」：李想在理想 L9 发布会上的表述（新浪财经，2026-05）","url":"https://finance.sina.com.cn/roll/2026-05-15/doc-inhxyimn6529650.shtml"}],"as_of":"2026-07","related_ids":["automakers-entering-humanoid-robotics","kunlunxing-robotics","simplexity-robotics","autonomous-driving-talent-moving-into-embodied-ai","xpeng","autonomous-driving"],"name":"Li Auto Inc.","alt":"理想汽车","abbr":"","aliases":["Li Auto"],"one_liner":"A Beijing EV maker that announced in 2026 it will build humanoid robots, calling autonomous driving embodied AI's first act.","explanation":"Li Auto was founded in Beijing in 2015 by Xiang Li, known for extended-range and pure-electric SUVs; it listed on Nasdaq in July 2020 and on the Hong Kong Stock Exchange in August 2021. At an all-staff meeting on January 26, 2026, Li said Li Auto “will definitely build humanoid robots, and will bring one to market as fast as possible”; in May he said “autonomous driving is the first act of embodied AI, and general-purpose humanoid robots are the second act.” Reports describe its robotics project under the codename Nexus, planning a factory-oriented two-wheeled robot and a bipedal humanoid, with the two-wheeled model due within 2026. Former autonomous-driving head Xianpeng Lang was reassigned to lead robot hardware early in the year before leaving to found Kunlunxing Robotics; CTO Kai Wang, who left even earlier, co-founded Simplexity Robotics.","example":"Li broke down humanoid-robot generalization into three stages: reaching the level of a 6-year-old child by 2030–2035, a 12-year-old by 2035–2040, and something close to an 18-year-old around 2040 — a process he says will take 15 to 20 years.","related":["Automakers Entering Humanoid Robotics","Kunlunxing Robotics","Simplexity Robotics","Autonomous-Driving Talent Moving into Embodied AI","XPeng","Autonomous Driving"]},{"id":"gac-group","category":"company","sec":5,"tier":3,"sources":[{"title":"广汽集团发布第三代具身智能人形机器人 GoMate（IT之家）","url":"https://www.ithome.com/0/820/290.htm"},{"title":"广汽集团内部孵化具身智能机器人公司慧仑科技（广汽集团官网）","url":"https://www.gacgroup.com/cn/news/detail?baseid=19104"},{"title":"广汽旗下人形机器人公司慧仑科技完成亿元融资（广汽集团官网）","url":"https://www.gacgroup.com/cn/news/detail?baseid=19182"}],"as_of":"2026-08","related_ids":["automakers-entering-humanoid-robotics","wheel-legged-robot","inspection-robot","joint-actuator-module","xpeng"],"name":"GAC Group","alt":"广汽集团","abbr":"GAC","aliases":["GAC","Guangzhou Automobile Group"],"one_liner":"A Guangzhou state-owned automaker that incubated Huilun Technology to build its wheel-legged GoMate humanoid robots.","explanation":"GAC Group was formed in 1997, is headquartered in Guangzhou, and is a Guangzhou municipal state-owned automaker, dual-listed in Hong Kong and on the A-share market. It began developing humanoid robots in early 2022, and in December 2024 released its third-generation GoMate: a wheel-legged design that can switch between a four-wheel and two-wheel stance, with self-developed core components. In October 2025 it followed with the fourth-generation GoMate Mini, mainly for security patrol work. In February 2026, GAC incubated Guangdong Huilun Technology to independently handle robot R&D and sales; that August, Huilun closed a round of more than RMB 100 million (with CRRC Guochuang Fund and CMB International, among others), saying nearly 50 GoMate Mini units were already deployed with orders worth close to RMB 10 million, and that it plans large-scale mass production by 2027.","example":"GoMate Mini patrols along a set route in a business park or auto plant, using its cameras to spot anomalies and report them; Huilun says its longest-running deployment has operated continuously for 12 months.","related":["Automakers Entering Humanoid Robotics","Wheel-legged Robot","Inspection Robot","Joint Actuator Module","XPeng"]},{"id":"aimoga-robotics","category":"company","sec":5,"tier":3,"sources":[{"title":"售价超28万元！奇瑞墨甲机器人上线京东（新浪财经）","url":"https://finance.sina.com.cn/roll/2026-04-14/doc-inhunttq4621338.shtml"},{"title":"背靠奇瑞求上市，墨甲机器人何时「独立」行走？（36氪）","url":"https://www.36kr.com/p/3957680691903616"},{"title":"Chery's robot unit Aimoga prepares for IPO, targets 10,000 deliveries（CnEVPost）","url":"https://cnevpost.com/2026/08/19/chery-aimoga-prepares-ipo/"}],"as_of":"2026-08","related_ids":["automakers-entering-humanoid-robotics","humanoid-robot","guided-tours-and-reception","quadruped-robot","agibot"],"name":"AiMOGA Robotics","alt":"墨甲机器人","abbr":"","aliases":["AiMOGA"],"one_liner":"A Chery-incubated robotics company building reception humanoids, police robots, and robot dogs, now preparing an IPO.","explanation":"AiMOGA Robotics grew out of preliminary humanoid-robot research that Chery Automobile started in 2022 (its first prototype rolled off the line in late 2023); it was registered in Wuhu, Anhui, in January 2025, with Chery directly holding 76.8% and general manager Guibing Zhang serving as Chery's executive vice president and head of international business. Its products include the reception humanoid “Mornine,” used for tasks such as dealership explanations and mall guidance; a traffic-directing “smart police” robot; quadruped robot dogs; and a home companion robot. In January 2026 it closed an angel round of more than RMB 100 million at a post-money valuation of RMB 2.5 billion, with AgiBot and IDG Capital among the investors; in April the bipedal humanoid Mornine M1 went on sale on JD.com priced at RMB 285,800. In August, Zhang told Reuters the company was preparing for an IPO, with cumulative deliveries of various robots exceeding 3,000 units, about 2,000 overseas, targeting 10,000 deliveries by 2027.","example":"In a dealership, Mornine explains car models and answers customer questions in multiple languages.","related":["Automakers Entering Humanoid Robotics","Humanoid Robot","Guided Tours & Reception","Quadruped Robot","AgiBot"]},{"id":"mobileye","category":"company","sec":5,"tier":3,"sources":[{"title":"Mobileye To Acquire Mentee Robotics to Accelerate Physical AI Leadership","url":"https://www.mobileye.com/news/mobileye-to-acquire-mentee-robotics-to-accelerate-physical-ai-leadership"},{"title":"Mobileye to acquire humanoid robotics startup Mentee for $900 million (Reuters)","url":"https://www.reuters.com/world/asia-pacific/mobileye-acquire-humanoid-robotics-startup-mentee-900-million-2026-01-06"}],"as_of":"2026-01","related_ids":["mentee-robotics","mentee-robotics-menteebot","autonomous-driving","autonomous-driving-talent-moving-into-embodied-ai","humanoid-robot","physical-ai"],"name":"Mobileye","alt":"Mobileye","abbr":"","aliases":[],"one_liner":"An Israeli autonomous-driving chip company that announced the acquisition of humanoid-robot maker Mentee in 2026.","explanation":"Mobileye was founded in Jerusalem, Israel in 1999; co-founder Amnon Shashua is a computer-vision scholar. It built its business on the EyeQ line of vision chips and driver-assistance systems, was acquired by Intel in 2017, and re-listed on the Nasdaq in 2022, with Intel remaining its largest shareholder. On January 6, 2026, Mobileye announced it was acquiring the humanoid-robot company Mentee Robotics for about $900 million (about $612 million in cash, with the rest in up to 26.2 million Mobileye shares), calling it “Mobileye 3.0” — bringing its autonomous-driving perception, decision-making, and mass-production experience to humanoid robots, with customer proof-of-concept deployments planned for 2026 and mass-market production for 2028. It is a textbook case of autonomous-driving talent moving into embodied AI.","example":"Mobileye acquired Mentee Robotics for about $900 million, bringing the humanoid robot MenteeBot under its umbrella.","related":["Mentee Robotics","Mentee Robotics MenteeBot","Autonomous Driving","Autonomous-Driving Talent Moving into Embodied AI","Humanoid Robot","Physical AI"]},{"id":"mentee-robotics","category":"company","sec":5,"tier":3,"sources":[{"title":"Mobileye To Acquire Mentee Robotics（Mobileye News）","url":"https://www.mobileye.com/news/mobileye-to-acquire-mentee-robotics-to-accelerate-physical-ai-leadership/"},{"title":"Mobileye acquires humanoid robot startup Mentee Robotics for $900M（TechCrunch）","url":"https://techcrunch.com/2026/01/06/mobileye-acquires-humanoid-robot-startup-mentee-robotics-for-900m"}],"as_of":"2026-01","related_ids":["mentee-robotics-menteebot","mobileye","humanoid-robot","sim-to-real-transfer","vision-only-approach","autonomous-driving-talent-moving-into-embodied-ai"],"name":"Mentee Robotics","alt":"Mentee Robotics","abbr":"","aliases":["Mentee"],"one_liner":"An Israeli humanoid-robot company founded by Mobileye's founder, acquired by Mobileye in 2026.","explanation":"Mentee Robotics was founded in Israel in 2022; its co-founders include Mobileye founder Amnon Shashua and CEO Lior Wolf. Its product is the general-purpose humanoid robot MenteeBot: the reported third generation stands about 175 cm tall with a payload of about 25 kg, relies on cameras alone for perception, uses in-house actuators and a hot-swappable battery, and leans heavily on sim-to-real transfer for training, emphasizing autonomous task completion from natural-language instructions rather than teleoperation. On January 6, 2026, Mobileye announced it was acquiring Mentee for about $900 million, planning customer proof-of-concept trials in 2026 and mass production in 2028, extending its autonomous-driving perception and chip capabilities into humanoid robots.","example":"In a Mentee demo, a user tells MenteeBot in one sentence to fetch a box from a shelf, and the robot understands the scene, plans a path, and completes the task on its own.","related":["Mentee Robotics MenteeBot","Mobileye","Humanoid Robot","Sim-to-Real Transfer","Vision-Only Approach","Autonomous-Driving Talent Moving into Embodied AI"]},{"id":"franka-robotics","category":"company","sec":6,"tier":2,"sources":[{"title":"Franka @ CES 2026: Powering the Future of Embodied AI","url":"https://franka.de/news/franka-ces-2026-powering-the-future-of-embodied-ai"},{"title":"Agile Robots acquires Franka Emika (Munich Startup, 2023-11)","url":"https://www.munich-startup.de/en/95730/agile-robots-takes-over-franka-emika/"},{"title":"Deutscher Zukunftspreis 2017, Team 2 (Franka Emika)","url":"https://www.deutscher-zukunftspreis.de/en/team-2-2017"}],"as_of":"2026-01","related_ids":["franka-emika-panda-franka-research-3","agile-robots","franka-hand","libfranka-franka-control-interface","droid","nvidia-isaac-gr00t-n1"],"name":"Franka Robotics","alt":"Franka","abbr":"","aliases":["Franka Emika"],"one_liner":"A Munich force-controlled robot-arm maker whose Panda / FR3 arms are among the most common in robotics research.","explanation":"Franka's predecessor, Franka Emika, was founded in Munich, Germany, in 2016 by Sami Haddadin and others; Haddadin had previously done robotics research at the German Aerospace Center, and the team won Germany's Future Prize in 2017 for the 7-axis force-controlled collaborative arm Panda. Panda and its research-oriented successor, FR3, have a torque sensor at every joint and expose a 1 kHz real-time control interface, making them the most common robot arm seen in robot-learning papers and datasets such as DROID. The company filed for insolvency in August 2023 after a shareholder dispute and was acquired by Agile Robots that November, taking the name Franka Robotics. It has since moved toward embodied AI: in 2025 it launched the dual-arm prototype FR3 Duo, and in January 2026 at CES it used that prototype for the first public on-device, end-to-end demo of NVIDIA's GR00T N1.6, while also offering a robot-data-collection service.","example":"The DROID dataset was collected by 13 institutions using the same setup — a Franka Panda plus cameras — so a policy trained on DROID is often tested directly on a real Franka arm.","related":["Franka Emika Panda / Franka Research 3","Agile Robots","Franka Hand","libfranka / Franka Control Interface (FCI)","DROID (Distributed Robot Interaction Dataset)","NVIDIA Isaac GR00T N1"]},{"id":"universal-robots","category":"company","sec":6,"tier":2,"sources":[{"title":"Universal Robots - Wikipedia","url":"https://en.wikipedia.org/wiki/Universal_Robots"}],"as_of":"2026-09","related_ids":["collaborative-robot","universal-robots-ur5e","rtde","streaming-servo-control","6-axis-robot-arm"],"name":"Universal Robots","alt":"优傲机器人","abbr":"UR","aliases":["UR"],"one_liner":"A Danish maker of collaborative robot arms that shipped the first commercially viable cobot; its UR5e is a lab staple.","explanation":"Universal Robots was founded in 2005 in Odense, Denmark, by researchers Esben Østergaard, Kasper Støy, and Kristian Kassow. In 2008 it shipped what is widely regarded as the first commercially viable collaborative robot, or “cobot” — an arm that can safely work in the same space as a person — and in 2015 it was acquired by the US test-equipment company Teradyne for $285 million. Its product line has grown from the UR3, UR5, and UR10 through the 2018 “e-Series” (including the UR5e) to the UR20 (2022) and UR30 (2024), with cumulative sales of more than 100,000 units. Its open interfaces, including RTDE (Real-Time Data Exchange), which reads and writes joint state in real time, have made it a common robot-arm brand in robot-learning labs and real-robot datasets.","example":"Many real-robot datasets and VLA papers collect data on a UR5e, streaming joint or end-effector commands over RTDE at several hundred hertz.","related":["Collaborative Robot","Universal Robots UR5e","RTDE","Streaming Servo Control","6-Axis Robot Arm"]},{"id":"agilex-robotics","category":"company","sec":6,"tier":2,"sources":[{"title":"AgileX Robotics - About Us","url":"https://global.agilex.ai/pages/about-us"}],"as_of":"2026-09","related_ids":["agilex-piper","agilex-cobot-magic","agilex-pika","mobile-aloha","mobile-base","leader-follower-teleoperation"],"name":"AgileX Robotics","alt":"松灵机器人","abbr":"","aliases":["松灵 (colloquial Chinese short form)","AgileX"],"one_liner":"A maker of mobile robot chassis and low-cost robot arms, a common hardware supplier for embodied AI research.","explanation":"AgileX Robotics was founded in 2016 and is headquartered in Shenzhen, starting out with mobile robot chassis and self-driving solutions; its products include the Scout, Hunter, Tracer, and Ranger chassis lines and the LIMO education platform. As embodied AI took off, it launched the low-cost six-axis robot arms PiPER and NERO, the Cobot Magic bimanual mobile platform (a recreation of Stanford's Mobile ALOHA), and the Pika handheld gripper kit, which can collect data without a real robot. Thanks to its low prices and complete open-source ROS drivers, many Chinese universities and companies use its bimanual platforms for imitation learning and data collection. The company states it has worked with more than 1,000 companies and over 50 universities.","example":"A lab buys a Cobot Magic setup, uses leader-follower teleoperation to record bimanual clothes-folding data, and then trains ACT or a diffusion policy on it.","related":["AgileX PiPER","AgileX Cobot Magic","AgileX Pika","Mobile ALOHA","Mobile Base (Chassis)","Leader-Follower Teleoperation"]},{"id":"dobot","category":"company","sec":6,"tier":2,"sources":[{"title":"财联社：机器人冲刺A股又添一例（越疆启动A股上市）","url":"https://www.cls.cn/detail/2245078"},{"title":"21财经：大湾区H回A首单，越疆科技创业板IPO即将上会","url":"https://m.21jingji.com/article/20260715/herald/c3f19a7d8572e488f9e91e10f9c3f6e5.html"},{"title":"21经济网：终止、等待、收购，机器人企业资本化路径分化","url":"https://www.21jingji.com/article/20260927/herald/a5a861b4b1cacffd91fba28f215f2b8b.html"}],"as_of":"2026-09","related_ids":["collaborative-robot","6-axis-robot-arm","dobot-magician","dobot-atom","desktop-robot-arm","research-and-education-market"],"name":"Dobot","alt":"越疆科技","abbr":"","aliases":["Yuejiang Technology","Shenzhen Dobot"],"one_liner":"A Shenzhen collaborative-arm maker, Hong Kong's first listed “cobot” company, now also building a humanoid.","explanation":"Dobot was founded in Nanshan, Shenzhen, in 2015; its founder, Peichao Liu, serves as chairman and general manager. It started with the desktop educational robot arm Dobot Magician, and its main product line is now 6-axis collaborative robots (light arms that can safely work in the same space as a person); this segment made up about 60% of total revenue in the first half of 2025, and the company says it shipped more collaborative robots than any other maker worldwide in 2025, with more than 100,000 units installed cumulatively. It listed on the Hong Kong Stock Exchange on December 23, 2024, widely called Hong Kong's “first collaborative-robot stock.” In recent years it has expanded into embodied AI: it unveiled the humanoid robot Atom in March 2025 and is also working on multi-legged robots. In 2026 it began a return to A-shares; its ChiNext IPO application was accepted on April 27, passed listing-committee review on July 22, and as of September had not yet completed registration.","example":"Universities and schools commonly use the Dobot Magician to teach robot-arm kinematics, while factories use its CR and Nova collaborative arms for loading and assembly.","related":["Collaborative Robot","6-Axis Robot Arm","Dobot Magician","Dobot Atom","Desktop Robot Arm","Research & Education Market"]},{"id":"agile-robots","category":"company","sec":6,"tier":3,"sources":[{"title":"Agile Robots - Wikipedia","url":"https://en.wikipedia.org/wiki/Agile_Robots"},{"title":"Humanoid Agile ONE embodies Physical AI at Hannover Messe 2026","url":"https://www.agile-robots.com/en/news/detail/humanoid-agile-one-embodies-physical-ai-at-hannover-messe-2026"}],"as_of":"2026-04","related_ids":["franka-emika-panda-franka-research-3","dlr-institute-of-robotics-and-mechatronics","collaborative-robot","7-dof-robot-arm","force-control","humanoid-robot"],"name":"Agile Robots","alt":"思灵机器人","abbr":"","aliases":["Agile Robots SE"],"one_liner":"A Munich robot-arm and automation company spun out of the German Aerospace Center, now Franka Emika's owner.","explanation":"Agile Robots (legally Agile Robots SE) was founded in Munich, Germany in 2018 by Zhaopeng Chen and Peter Meusel, both of whom came from the DLR Institute of Robotics and Mechatronics at the German Aerospace Center. Its main products are force-controlled 7-axis collaborative robot arms (such as the Diana 7), mobile robots, and full automation packages for manufacturing and logistics. In 2021, a Series C round led by SoftBank's Vision Fund 2 made it a unicorn; in November 2023 it acquired the bankrupt Franka Emika (maker of the Franka robot arm), and in September 2025 it took full ownership of idealworks, a robotics venture incubated by BMW. In April 2026, at the Hannover Messe trade fair, it unveiled its first industrial humanoid robot, Agile ONE. For newcomers, its most direct connection is to the Franka arm widely used in research.","example":"The Franka Research 3 arm commonly used in labs is made by Franka Emika, which Agile Robots acquired in 2023.","related":["Franka Emika Panda / Franka Research 3","DLR Institute of Robotics and Mechatronics","Collaborative Robot","7-DoF Robot Arm","Force Control","Humanoid Robot"]},{"id":"trossen-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"Trossen Robotics: Aloha Robot, A Low-Cost Bimanual Platform","url":"https://www.trossenrobotics.com/post/aloha-robot-low-cost-bimanual-platform"},{"title":"LinkedIn: Matt Trossen","url":"https://www.linkedin.com/in/matttrossen"}],"as_of":"2026-09","related_ids":["trossen-robotics-widowx-250","trossen-robotics-viperx-300","aloha","leader-follower-teleoperation","bridgedata-v2","mobile-aloha"],"name":"Trossen Robotics","alt":"Trossen Robotics","abbr":"","aliases":["Interbotix"],"one_liner":"A US research robot-arm maker selling WidowX, ViperX, and complete ALOHA kits.","explanation":"Trossen Robotics is a US robotics hardware company founded in 2004 by Matt Trossen, headquartered in Downers Grove, Illinois, and long focused on selling robot kits for research and education. It became well known in embodied-AI circles because its low-cost Interbotix arms are used in a huge number of academic projects: the BridgeData V2 dataset used a WidowX 250, and Stanford's ALOHA dual-arm teleoperation platform uses a ViperX 300 as the follower arm and a WidowX as the leader. The company now sells complete ALOHA hardware kits directly, in both stationary and mobile versions, including the leader and follower arms, multiple cameras, and data-collection software — letting a lab start collecting imitation-learning data within hours of unboxing.","example":"Many labs reproducing ACT or Mobile ALOHA buy a complete ALOHA dual-arm kit straight from Trossen to collect their demonstration data.","related":["Trossen Robotics WidowX 250","Trossen Robotics ViperX 300","ALOHA","Leader-Follower Teleoperation","BridgeData V2","Mobile ALOHA"]},{"id":"kinova","category":"company","sec":6,"tier":3,"sources":[{"title":"Jaco - ROBOTS: Your Guide to the World of Robotics（IEEE）","url":"https://robotsguide.com/robots/jaco"},{"title":"Kinova raises $60 million in new financing（Newswire.ca）","url":"https://www.newswire.ca/news-releases/kinova-raises-60-million-in-new-financing-the-company-s-expansion-into-the-industrial-automation-market-continues-826413359.html"},{"title":"Kinova Celebrates 20 Years of Innovation with the Launch of KIMA（Newswire.ca）","url":"https://www.newswire.ca/news-releases/kinova-celebrates-20-years-of-innovation-with-the-launch-of-kima-its-medical-robotic-arm-875537715.html"}],"as_of":"2026-06","related_ids":["kinova-gen3","7-dof-robot-arm","collaborative-robot","franka-emika-panda-franka-research-3","universal-robots-ur5e","surgical-robot"],"name":"Kinova","alt":"Kinova","abbr":"","aliases":["Kinova Robotics"],"one_liner":"A Canadian lightweight robot-arm maker that started with wheelchair-mounted assistive arms; its Gen3 is a common research arm.","explanation":"Kinova was founded in 2006 by Charles Deguire, now CEO, and Louis-Joseph L'Écuyer, and is headquartered in Boisbriand, Quebec, near Montreal. Several of Deguire's uncles have muscular dystrophy and rely on power wheelchairs, and one of them had built his own robot arm; the company's 2009 wheelchair-mounted assistive arm, JACO, was named after him. It later developed a lightweight research product line, with the Gen3 (available in 6- and 7-degree-of-freedom versions) becoming common in university labs. In February 2022 it raised CA$60 million — CA$40 million led by Graham Partners and CA$20 million from Canada's Strategic Innovation Fund — to expand into industrial automation, launching the industrial collaborative arm Link 6 that same year. In June 2026, marking its 20th anniversary, it released the surgical and endoscopic medical arm KIMA, with a 3-kilogram payload and under 13 kilograms of its own weight.","example":"Mounted on a power wheelchair, JACO helps a user with limited upper-body strength eat and open doors on their own.","related":["Kinova Gen3","7-DoF Robot Arm","Collaborative Robot","Franka Emika Panda / Franka Research 3","Universal Robots UR5e","Surgical Robot"]},{"id":"ufactory","category":"company","sec":6,"tier":3,"sources":[{"title":"UArm机器人 项目信息（36氪创投平台）","url":"https://pitchhub.36kr.com/project/1678224239801347"},{"title":"Cheetah Mobile to Acquire Controlling Stake in UFACTORY（猎豹移动投资者关系）","url":"https://ir.cmcm.com/2025-07-28-Cheetah-Mobile-to-Acquire-Controlling-Stake-in-UFACTORY-to-Accelerate-Its-Robotics-Commercialization-Strategy"},{"title":"Cheetah Mobile Announces Second Quarter 2026 Unaudited Consolidated Financial Results（PR Newswire）","url":"http://www.prnewswire.com/news-releases/cheetah-mobile-announces-second-quarter-2026-unaudited-consolidated-financial-results-302875945.html"}],"as_of":"2026-09","related_ids":["ufactory-xarm","collaborative-robot","desktop-robot-arm","7-dof-robot-arm","gello","open-source-hardware"],"name":"UFACTORY","alt":"UFACTORY","abbr":"","aliases":["uFactory"],"one_liner":"A Shenzhen lightweight robot-arm maker behind the xArm and uArm, taken over by Cheetah Mobile in 2025.","explanation":"UFACTORY was founded in Shenzhen in December 2013, formally Shenzhen UFactory Technology Co., Ltd., by Shitao Deng, a maker born in 1989. It got its start with the Kickstarter-funded desktop 4-axis arm uArm in early 2014, then moved into lightweight collaborative arms: the xArm 5/6/7 (700 mm reach, 3–5 kg payload), the entry-level Lite 6, and the UFACTORY 850, all cheaper than a Franka or a UR arm, with a Python SDK and ROS support, which has made it a common arm for university labs doing imitation learning and teleoperation data collection. Reports say its products sell to more than 80 countries and regions, mainly overseas, and the business is profitable. In July 2025, Cheetah Mobile acquired 60.8% of the company for about RMB 99.5 million, bringing its total stake to roughly 80% and making it the controlling shareholder; Cheetah Mobile said its “robotics and other” revenue grew 72.5% year-on-year in the second quarter of 2026, partly from UFACTORY's consolidated revenue.","example":"Open-source teleoperation projects such as GELLO offer an xArm version, letting researchers record demonstration data on it and use it to train imitation-learning policies.","related":["UFACTORY xArm","Collaborative Robot","Desktop Robot Arm","7-DoF Robot Arm","GELLO","Open-Source Hardware (OSHW)"]},{"id":"realman-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"RealMan Robotics 官网","url":"https://www.realman-robotics.com/"}],"as_of":"2026-09","related_ids":["realman-rm-series-arm","collaborative-robot","wheeled-humanoid-robot","payload-to-weight-ratio","teleoperation","embodied-ai-data-service-provider"],"name":"RealMan Robotics","alt":"睿尔曼智能","abbr":"","aliases":["RealMan"],"one_liner":"A Beijing lightweight robot-arm maker whose RM-series arms are common on domestic dual-arm and wheeled humanoid platforms.","explanation":"RealMan Robotics is a Beijing-based robotics company whose flagship product is what it calls an ultra-lightweight humanoid-style robot arm — the RM series, including the 6-axis RM65 and the 7-axis RM75. The controller is built into the arm itself, it runs on 24V DC power, and it has a high payload-to-weight ratio, making it easy to mount on a mobile base or dual-arm platform, which is why many Chinese embodied-AI teams and universities use it as their arm of choice. In recent years the product line has expanded to include the WHJ joint modules, the RealBot series of wheeled humanoids, and teleoperation networks and data-collection platforms; the company's website positions it as system-level infrastructure for physical AI and also offers ODM/OEM customization. Its founding year and funding figures are not listed on its official site and are omitted here.","example":"Some compound robots mount two RM75 arms on a lifting column and mobile base, using teleoperation to collect dual-arm data for training policies.","related":["RealMan RM Series Arm","Collaborative Robot","Wheeled Humanoid Robot","Payload-to-Weight Ratio","Teleoperation","Embodied AI Data Service Provider"]},{"id":"flexiv-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"Flexiv - About","url":"https://www.flexiv.com/about"},{"title":"Flexiv - News","url":"https://www.flexiv.com/news"}],"as_of":"2026-03","related_ids":[null,null,null,null,null],"name":"Flexiv Robotics","alt":"非夕科技","abbr":"","aliases":["Flexiv"],"one_liner":"A US-China robotics company building force-controlled “adaptive robots,” with its Rizon arm as the flagship product.","explanation":"Flexiv was founded in Santa Clara, California in 2016, with a core team that came out of Stanford's robotics and AI research lab; its founder and CEO is Shiquan Wang. It now has business hubs in Silicon Valley, Shanghai, Beijing, Munich, and Singapore. Its flagship product is the Rizon, a 7-axis “adaptive” robot arm: every joint carries a torque sensor, enabling high-precision force control and hybrid force/position control, making it suited to contact-heavy work like polishing, assembly, and connector insertion that requires a sense of “touch,” and it is also a common experimental platform for researchers studying contact-rich manipulation. According to the company's own news page, it announced a robotics-simulation partnership with NVIDIA in June 2025 and received an investment from a long-term investment institution in March 2026 to expand its global deployment.","example":"Using a Rizon arm to polish a car part: the arm follows the curved surface at a set pressure, without needing to know the workpiece's exact shape in advance.","related":["Flexiv Rizon","Force Control","Hybrid Force/Position Control","Contact-rich Manipulation","Collaborative Robot"]},{"id":"jaka-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"21世纪经济报道：9岁上海机器人，孙正义3亿投它","url":"https://www.21jingji.com/article/20230519/herald/e78d1c32fbe40cc6bc522856bbbc215f.html"},{"title":"新浪财经：节卡机器人IPO终止","url":"https://finance.sina.com.cn/stock/s/2025-12-22/doc-inhcsexf4575422.shtml"},{"title":"上海交大机动学院：节卡共建通用智能机器人联合研究中心","url":"https://me.sjtu.edu.cn/hzdt/77443.html"}],"as_of":"2026-01","related_ids":["collaborative-robot","6-axis-robot-arm","kinesthetic-teaching","universal-robots","dobot","flexiv-robotics"],"name":"JAKA Robotics","alt":"节卡机器人","abbr":"","aliases":["JAKA"],"one_liner":"A Shanghai collaborative-arm maker founded in 2014 whose STAR Market IPO was later terminated.","explanation":"JAKA Robotics was founded in Shanghai in July 2014; its founder and chairman, Mingyang Li, graduated from Shanghai Jiao Tong University and previously worked in sales at Tetra Pak, and the company's early R&D was done in partnership with Shanghai Jiao Tong University's robotics institute. It started out integrating packaging lines for dairy gift boxes, launched its first Zu-series collaborative robot in 2017, and has since expanded into the Zu, Pro, and C series, with payloads of roughly 1-20 kg, supporting kinesthetic teaching and work in the same space as people. Its investors include SoftBank Vision Fund and Temasek. It filed for a STAR Market IPO in May 2023, which was terminated on December 19, 2025. In January 2026 it set up a joint research center for general-purpose intelligent robots with Shanghai Jiao Tong University, beginning to move into embodied AI.","example":"A factory uses a JAKA Zu-series collaborative arm for machine tending; a worker drags the arm through the motion once to teach it.","related":["Collaborative Robot","6-Axis Robot Arm","Kinesthetic Teaching","Universal Robots","Dobot","Flexiv Robotics"]},{"id":"rokae","category":"company","sec":6,"tier":3,"sources":[{"title":"珞石机器人官网","url":"https://www.rokae.com/"},{"title":"珞石机器人投资者关系","url":"https://ircn.rokae.com/"},{"title":"珞石 2026 中期报告（港交所披露易）","url":"https://www1.hkexnews.hk/listedco/listconews/sehk/2026/0923/2026092300881_c.pdf"}],"as_of":"2026-09","related_ids":["collaborative-robot","industrial-robot","force-control","7-dof-robot-arm","jaka-robotics","dobot"],"name":"ROKAE","alt":"珞石机器人","abbr":"","aliases":["ROKAE Robotics"],"one_liner":"A Beijing industrial and collaborative-robot maker known for its xMate force-controlled cobot arms, now also building embodied-AI robots.","explanation":"ROKAE Robotics, formally ROKAE (Beijing) Robot Co., Ltd., is headquartered in Beijing and works on the R&D, manufacturing, and commercialization of intelligent robots. Its products fall into three groups: the xMate line of flexible collaborative robots — the CR, SR, and ZR series, with payloads from 3 to 45 kilograms — which emphasize joint torque sensing and force control for tasks that need a delicate touch, such as polishing and assembly; the NB/XB industrial robots and SCARA arms; and, more recently, embodied-AI products such as force-controlled humanoid arms and wheeled humanoids. Its website states more than 1,000 customers across 40-plus countries and regions and positions the company as a core infrastructure builder toward physical AI. It is listed on the Hong Kong Stock Exchange, and disclosure documents such as its 2026 interim report are available on its investor-relations page.","example":"","related":["Collaborative Robot","Industrial Robot","Force Control","7-DoF Robot Arm","JAKA Robotics","Dobot"]},{"id":"abb-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"ABB Group - Wikipedia","url":"https://en.wikipedia.org/wiki/ABB_Group"}],"as_of":"2025-10","related_ids":["big-four-of-industrial-robotics","abb-yumi","industrial-robot","softbank-group","fanuc","kuka"],"name":"ABB Robotics","alt":"ABB","abbr":"","aliases":["ABB Group","ABB Ltd"],"one_liner":"A Swiss electrification and automation group, one of industrial robotics' “Big Four,” now selling its robotics unit to SoftBank.","explanation":"ABB is an electrification and automation group headquartered in Zurich, Switzerland, formed in 1988 through the merger of Sweden's ASEA and Switzerland's BBC (Brown, Boveri & Cie); together with FANUC, KUKA, and Yaskawa, it is considered one of industrial robotics' “Big Four.” Its predecessor ASEA was the first to build a microprocessor into an industrial robot. Its product line includes the IRB series of industrial arms used for welding, painting, and material handling, as well as YuMi, a dual-arm collaborative robot designed to work alongside people on the same line. Its robotics division had about $2.3 billion in revenue and roughly 7,000 employees in 2024. In October 2025, ABB announced it would sell its robotics business to SoftBank Group for about $5.4 billion, replacing an earlier plan to spin the unit off in an IPO in 2026; the deal still requires regulatory approval.","example":"An electronics factory uses ABB's dual-arm YuMi cobot for small-parts assembly, working alongside human workers with no safety fencing.","related":["Big Four of Industrial Robotics","ABB YuMi","Industrial Robot","SoftBank Group","FANUC","KUKA"]},{"id":"kuka","category":"company","sec":6,"tier":3,"sources":[{"title":"Wikipedia: KUKA","url":"https://en.wikipedia.org/wiki/KUKA"},{"title":"KUKA: LBR iiwa","url":"https://www.kuka.com/en-us/products/robotics-systems/industrial-robots/lbr-iiwa"}],"as_of":"2025","related_ids":["big-four-of-industrial-robotics","kuka-lbr-iiwa","industrial-robot","fanuc","abb-robotics","joint-torque-sensor"],"name":"KUKA","alt":"库卡","abbr":"","aliases":["Midea KUKA","KUKA AG"],"one_liner":"A German industrial-robot maker, one of the “Big Four,” owned by China's Midea Group since 2017.","explanation":"KUKA was founded in 1898 by Johann Josef Keller and Jakob Knappich in Augsburg, Germany — the name KUKA is an abbreviation of the company's early German name — and it is still headquartered in Augsburg today. Grouped with FANUC, ABB, and Yaskawa as one of the “Big Four” of industrial robotics, it mainly builds industrial robot arms for uses such as automotive welding, along with logistics and medical automation. In 2016, China's Midea Group launched a tender offer worth about €4.5 billion; the deal closed in January 2017 with Midea holding about 94.55%, and Midea bought out the remaining shares in 2022, delisting the company, which in China is often called “Midea KUKA.” In embodied-AI research, its 7-axis collaborative arm, the LBR iiwa, is a common sight — every joint carries a torque sensor, and it is often used for force control and contact-rich manipulation experiments.","example":"Dexterous-grasping research such as DextrAH-G runs real-robot experiments using a KUKA arm paired with an Allegro dexterous hand.","related":["Big Four of Industrial Robotics","KUKA LBR iiwa","Industrial Robot","FANUC","ABB Robotics","Joint Torque Sensor"]},{"id":"fanuc","category":"company","sec":6,"tier":3,"sources":[{"title":"Wikipedia: FANUC","url":"https://en.wikipedia.org/wiki/FANUC"},{"title":"FANUC Strengthens Collaboration with NVIDIA (2026-05-15)","url":"https://www.fanuc.co.jp/en/profile/pr/newsrelease/2026/notice20260515.html"}],"as_of":"2026-09","related_ids":[null,null,null,null,null,null],"name":"FANUC","alt":"发那科","abbr":"","aliases":["FANUC Corporation"],"one_liner":"A Japanese giant in industrial robots and CNC systems, one of the “Big Four” of industrial robotics.","explanation":"FANUC is a Japanese company headquartered in Oshino, Yamanashi Prefecture, which grew out of Fujitsu's numerical-control division in the 1950s and became independent in 1972 under founder Seiuemon Inaba. Its main business is CNC (computer numerical control) systems that control machine tools, industrial robots, and factory automation equipment; its signature yellow robot arms are found throughout auto and electronics factories, and it is grouped with ABB, KUKA, and Yaskawa as one of the “Big Four” of industrial robotics. It is listed on the Tokyo Stock Exchange. It represents the traditional “teach-and-playback” camp of industrial robots, but has recently begun adopting embodied-AI technology: in May 2026 it announced a partnership with Google to have AI agents operate its robots, deepened its collaboration with NVIDIA by integrating Isaac Sim into its own offline simulation software ROBOGUIDE, and demonstrated dual-arm clothes-folding via imitation learning using the GR00T N model; that September it released an AI agent for welding.","example":"Most of the rows of yellow six-axis welding robots on an auto-body assembly line are FANUC machines.","related":["Industrial Robot","Big Four of Industrial Robotics (FANUC, ABB, KUKA, Yaskawa)","Teach-and-Playback Programming","NVIDIA","NVIDIA Isaac Sim","NVIDIA Isaac GR00T N1 / N1.5 / N1.6 / N1.7"]},{"id":"yaskawa-electric-corporation","category":"company","sec":6,"tier":3,"sources":[{"title":"Yaskawa Electric Corporation - Wikipedia","url":"https://en.wikipedia.org/wiki/Yaskawa_Electric_Corporation"}],"as_of":"2026-07","related_ids":["big-four-of-industrial-robotics","industrial-robot","fanuc","kuka","servo-motor"],"name":"Yaskawa Electric Corporation","alt":"安川电机","abbr":"","aliases":["Yaskawa"],"one_liner":"A century-old Japanese industrial-controls company, a major maker of servo motors and MOTOMAN industrial robots.","explanation":"Yaskawa Electric was founded in 1915 and is headquartered in Kitakyushu, Fukuoka Prefecture, Japan; it is listed on the Tokyo Stock Exchange and is a Nikkei 225 constituent. Its core businesses are servo motors, motion controllers, inverters, and the MOTOMAN line of industrial robots, a common sight on welding, handling, and assembly lines, and it is counted with FANUC, ABB, and KUKA as one of the “Big Four” of industrial robotics. In 1969 it trademarked the term “Mechatronics,” which later became a generic industry term. In embodied AI, Yaskawa represents the traditional industrial-robotics path: taught-and-programmed motion, high precision, and high reliability. According to reports, in July 2026 Yaskawa joined FANUC and Kawasaki Heavy Industries in a Fujitsu-led project to build an NVIDIA-based physical-AI coordination platform for factories, hospitals, and homes.","example":"Rows of MOTOMAN welding robots on an automotive body shop's line repeat a taught trajectory over and over.","related":["Big Four of Industrial Robotics","Industrial Robot","FANUC","KUKA","Servo Motor"]},{"id":"estun-automation","category":"company","sec":6,"tier":3,"sources":[{"title":"埃斯顿官网：埃斯顿港股上市，A+H双资本平台战略加速国际化布局","url":"https://www.estun.com/market/718.html"},{"title":"证券时报：埃斯顿总裁吴侃专访","url":"https://stcn.com/article/detail/3678067.html"},{"title":"钛媒体：埃斯顿的中国机器人故事，要用真金白银撑起来","url":"https://www.tmtpost.com/7894283.html"}],"as_of":"2026-03","related_ids":["industrial-robot","big-four-of-industrial-robotics","domestic-substitution","servo-motor","6-axis-robot-arm","palletizing-depalletizing"],"name":"Estun Automation","alt":"埃斯顿","abbr":"","aliases":["Nanjing Estun Automation Co., Ltd.","Estun"],"one_liner":"A Nanjing industrial-robot leader reportedly now China's top-shipping domestic maker of industrial robots.","explanation":"Estun was founded in Nanjing in 1993 by Bo Wu, and is now led by his son, Kan Wu, as vice chairman and president. It started out making motion-control components such as CNC systems and servo motors, then moved downstream into building complete industrial robots, giving it an in-house chain from core components to finished machines; its Nanjing factory runs a line where “robots build robots.” It listed on the Shenzhen Stock Exchange in 2015 (002747). Reportedly, in 2025 its industrial-robot shipments in the Chinese market surpassed the “Big Four” — FANUC, ABB, KUKA, and Yaskawa — for the first time, making it the top seller domestically. On March 9, 2026, it listed on the Hong Kong Stock Exchange main board (02715), becoming China's first industrial-robot company with both A-share and H-share listings; the proceeds are earmarked mainly for overseas capacity, acquisitions, and R&D, and the company is also moving into embodied AI for industrial settings.","example":"Among domestic brands, Estun ships more of the 6-axis industrial arms commonly seen on welding, palletizing, automotive, and battery production lines than almost any other Chinese maker.","related":["Industrial Robot","Big Four of Industrial Robotics","Domestic Substitution","Servo Motor","6-Axis Robot Arm","Palletizing / Depalletizing"]},{"id":"siasun-robot-and-automation","category":"company","sec":6,"tier":3,"sources":[{"title":"新松机器人投资者关系活动记录表（2026-08）","url":"https://www.siasun.com/uploads/file/20260826/3d506a6db414377d37a734d09201c3f4.pdf"},{"title":"新松参展2025世界智能制造大会","url":"https://www.siasun.com/news-detail954.html"}],"as_of":"2026-08","related_ids":["industrial-robot","collaborative-robot","autonomous-mobile-robot","humanoid-robot","domestic-substitution"],"name":"SIASUN Robot & Automation","alt":"新松机器人","abbr":"","aliases":["SIASUN"],"one_liner":"A long-established Chinese industrial-robot maker controlled by the Chinese Academy of Sciences' Shenyang Institute of Automation.","explanation":"SIASUN Robot & Automation was founded in Shenyang in 2000; its controlling shareholder is the Chinese Academy of Sciences' Shenyang Institute of Automation, and the company's name honors Xinsong Jiang, often called “the father of Chinese robotics.” It listed on the Shenzhen Stock Exchange's ChiNext board in 2009 under the stock abbreviation “Robot” (ticker 300024). Its business spans industrial robots, mobile robots, clean-room and semiconductor equipment, specialty robots, and automated production lines; in 2025 automated assembly and inspection lines were its largest revenue source, with industrial-robot revenue around RMB 1.1 billion. It is an early representative of domestic industrial robots and AGVs, and in recent years has also developed humanoid robots and embodied AI, including its Ruike series of dual-arm humanoids, though these have not yet produced significant profit at scale.","example":"SIASUN showed its Ruike MR73A humanoid robot at the 2025 World Manufacturing Convention, demonstrating carrying, inspection, and guided tours.","related":["Industrial Robot","Collaborative Robot","Autonomous Mobile Robot","Humanoid Robot","Domestic Substitution"]},{"id":"inovance-technology","category":"company","sec":6,"tier":3,"sources":[{"title":"Inovance - Wikipedia","url":"https://en.wikipedia.org/wiki/Inovance_Technology"}],"as_of":"2024-01","related_ids":["servo-motor","programmable-logic-controller","selective-compliance-assembly-robot-arm","domestic-substitution","industrial-robot","estun-automation"],"name":"Inovance Technology","alt":"汇川技术","abbr":"","aliases":["Inovance"],"one_liner":"A Shenzhen industrial-automation leader making servo systems, PLCs, and industrial robots.","explanation":"Inovance Technology was founded in Shenzhen in April 2003 by Xingming Zhu and a group of former Huawei engineers, sometimes called “Little Huawei” within the industry. It listed on the Shenzhen Stock Exchange's ChiNext board in September 2010 (300124). Its main businesses are frequency converters, servo systems, PLCs (programmable logic controllers), and industrial robots, and through subsidiaries it also makes electric drive and control systems for new-energy vehicles. Wikipedia describes it as China's largest industrial-automation company and the country's second-largest maker of industrial robots. Servo motors, drivers, and controllers are exactly the core components of a robot's joints, so in discussions of the embodied-AI supply chain, Inovance is often cited as a representative company for component supply and domestic substitution.","example":"","related":["Servo Motor","Programmable Logic Controller (PLC)","Selective Compliance Assembly Robot Arm","Domestic Substitution","Industrial Robot","Estun Automation"]},{"id":"tianji-intelligence","category":"company","sec":6,"tier":3,"sources":[{"title":"证券时报：天机智能完成10亿元融资","url":"https://www.stcn.com/article/detail/3925507.html"},{"title":"科创板日报：天机智能B++轮落地","url":"https://www.chinastarmarket.cn/detail/2477835"}],"as_of":"2026-09","related_ids":["industrial-robot","6-axis-robot-arm","selective-compliance-assembly-robot-arm","joint-torque-sensor","machines-replacing-humans","yaskawa-electric-corporation"],"name":"Tianji Intelligence","alt":"天机智能","abbr":"","aliases":["Tianji Robotics"],"one_liner":"A Dongguan industrial-robot maker under Apple supplier Luxshare Precision, valued near RMB 10 billion in 2026.","explanation":"Tianji Intelligence, formally Guangdong Tianji Intelligent Systems Co., Ltd., was founded in 2015 and is headquartered in Songshan Lake, Dongguan, Guangdong; its largest shareholder is Apple supply-chain company Luxshare Precision (holding about 27%), with chairman and general manager Xi Chen. In 2017 it formed a joint venture with Yaskawa China called Tianji Robotics. Its products cover small- and medium-payload 6-axis industrial robots (the TR series) and SCARA robots — a horizontal-jointed arm suited to fast, flat-plane handling (the SR series) — running its self-developed Tianji Fusion control system, and it says it has developed its own MEMS joint-torque sensors for robot joints. In recent years it has moved from traditional automation toward embodied AI: in May 2026 it closed combined Series B and B+ rounds of RMB 1 billion, led by Hillhouse Ventures and Meituan's strategic investment arm with Tencent following on, at a post-money valuation near RMB 10 billion; that September it raised a further round involving Ant Group and Tencent.","example":"Shoe, apparel, and food-service factories use Tianji's 6-axis arms and SCARA robots in place of manual labor for machine tending and sorting.","related":["Industrial Robot","6-Axis Robot Arm","Selective Compliance Assembly Robot Arm","Joint Torque Sensor","Machines Replacing Humans","Yaskawa Electric Corporation"]},{"id":"pudu-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"关于我们 - 普渡科技","url":"https://www.pudurobotics.com/zh-HK/company"},{"title":"普渡机器人创始人张涛（深圳市发改委）","url":"https://fgw.sz.gov.cn/ztzl/qtztzl/szscjmyjjfzzhfwpt/mqfc/myqyjdxsj/content/post_12623955.html"},{"title":"这家百亿机器人独角兽要IPO了（腾讯新闻）","url":"https://view.inews.qq.com/a/20260613A02A6F00"}],"as_of":"2026-06","related_ids":["service-robot","autonomous-mobile-robot","humanoid-robot","one-brain-multiple-robots","going-global","hkex-chapter-18c"],"name":"Pudu Robotics","alt":"普渡机器人","abbr":"","aliases":["PUDU"],"one_liner":"A Shenzhen commercial service-robot company with high shipment volumes in restaurant delivery and cleaning robots.","explanation":"Pudu Robotics (Shenzhen Pudu Technology Co., Ltd.) was founded in Shenzhen in January 2016. Founder Tao Zhang studied mechanical and electronic engineering as an undergraduate, did graduate research in software algorithms at the Hong Kong University of Science and Technology, and had also co-founded the tech media outlet Leiphone. The company started with restaurant food-delivery robots and now has four product lines — service delivery, commercial cleaning, industrial delivery, and general-purpose embodied AI — using a “one brain, multiple forms” architecture to build specialized, semi-humanoid, and humanoid robots such as the PUDU D9. Its website states cumulative global shipments of more than 130,000 units across 85-plus countries and regions; industrial-delivery revenue reportedly doubled year-on-year in the first quarter of 2026. In June 2026 it reportedly began preparing for a Hong Kong IPO.","example":"The delivery robots carrying dishes to tables in many hot-pot restaurants are Pudu products.","related":["Service Robot","Autonomous Mobile Robot","Humanoid Robot","One Brain, Multiple Robots","Going Global","HKEX Chapter 18C"]},{"id":"keenon-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"KEENON Robotics Showcases Humanoid Robot at CES 2026 for First Time（PR Newswire）","url":"https://www.prnewswire.co.uk/news-releases/keenon-robotics-showcases-humanoid-robot-at-ces-2026-for-first-time-and-unveils-first-robotic-lawn-mower-expanding-its-robotic-services-into-new-realms-302654186.html"},{"title":"擎朗智能发布人形具身服务机器人 XMAN-R1（IT之家）","url":"https://www.ithome.com/0/841/931.htm"},{"title":"擎朗智能考虑今年赴港上市（新浪财经·新股消息）","url":"https://finance.sina.com.cn/stock/hkstock/ggscyd/2026-01-19/doc-inhhvmxx6921579.shtml"}],"as_of":"2026-07","related_ids":["service-robot","wheeled-humanoid-robot","pudu-robotics","autonomous-mobile-robot","robot-rental"],"name":"KEENON Robotics","alt":"擎朗智能","abbr":"","aliases":["KEENON"],"one_liner":"A Shanghai commercial service-robot company with high delivery-robot shipment volumes, now also building humanoids.","explanation":"KEENON Robotics was founded in Shanghai in 2010; founder and CEO Tong Li previously worked at Microsoft Research Asia's engineering institute on the Microsoft Robotics Studio development platform. It builds commercial service robots for restaurant delivery, hotel delivery, cleaning, and medical transport, and offers a pay-monthly “robot-as-a-hire” model overseas. The company says it has shipped more than 100,000 units cumulatively, citing IDC data that it ranks first worldwide in commercial service-robot shipments. In March 2025 it released the wheeled humanoid service robot XMAN-R1, able to handle a continuous sequence from taking an order to plating, delivering, and clearing dishes, followed by the bipedal humanoid XMAN-F1. It closed a $200 million Series D led by the SoftBank Vision Fund in 2021; it reportedly was considering a Hong Kong listing within 2026, aiming to raise about $200 million, as of January that year.","example":"In a restaurant, XMAN-R1 pours drinks and arranges trays, then hands off to a delivery robot to carry the dishes to the table.","related":["Service Robot","Wheeled Humanoid Robot","Pudu Robotics","Autonomous Mobile Robot","Robot Rental"]},{"id":"geek-plus","category":"company","sec":6,"tier":3,"sources":[{"title":"市值超210亿，机器人超级独角兽登陆港交所（澎湃新闻）","url":"https://m.thepaper.cn/newsDetail_forward_31139514"},{"title":"全球首款仓储通用人形机器人：极智嘉发布 Gino 1（IT之家）","url":"https://www.ithome.com/0/921/115.htm"},{"title":"极智嘉(2590.HK)亮相2026 WAIC：「一核双引擎」战略落地（网易）","url":"https://www.163.com/dy/article/L22DTDJS05198ETO.html"}],"as_of":"2026-07","related_ids":["autonomous-mobile-robot","order-picking","fleet-management-system","wheeled-humanoid-robot","4d-world-model","amazon-robotics"],"name":"Geek+","alt":"极智嘉","abbr":"","aliases":["Geekplus"],"one_liner":"A Beijing warehouse-logistics robotics company building picking and transport AMRs, now also moving into embodied AI after its Hong Kong listing.","explanation":"Geek+ was founded in Beijing in 2015; founder and CEO Yong Zheng holds a bachelor's and master's in industrial engineering from Tsinghua, previously worked in production operations at ABB and Saint-Gobain, and later moved into robotics-industry investing. The company builds autonomous mobile robots (AMRs) for warehouse logistics: goods-to-person systems that bring shelves or totes to a picker, plus sorting and transport robots, and it cites Interact Analysis data ranking it first worldwide in AMR shipments for seven consecutive years. It listed on the Hong Kong Stock Exchange on July 9, 2025 (2590.HK), raising about HK$2.7 billion. In February 2026 it released the wheeled humanoid Gino 1 for warehouse settings — dual arms with a three-fingered dexterous hand for picking, tote handling, and packing — and that July, at WAIC, it launched the embodied-AI framework Gravity and its core model Gravity 4D, which predicts future frames, 3D structure, and motion together; its embodied-AI subsidiary also began raising its first independent funding round.","example":"On the WAIC 2026 show floor, Gino 1 picked items one at a time out of a tote — chips, bread, a small rubber bowl — working alongside a transport robot to complete the picking process.","related":["Autonomous Mobile Robot","Order Picking","Fleet Management System (e.g. Open-RMF)","Wheeled Humanoid Robot","4D World Model","Amazon Robotics"]},{"id":"youibot-robotics","category":"company","sec":6,"tier":3,"sources":[{"title":"优艾智合递交IPO招股书，拟赴香港上市（新浪财经）","url":"http://finance.sina.com.cn/wm/2026-04-02/doc-inhtcnfs0898056.shtml"},{"title":"合肥优艾智合机器人股份有限公司公告（港交所披露易）","url":"https://www1.hkexnews.hk/app/sehk/2026/108376/documents/sehk26033101825_c.pdf"},{"title":"优艾智合冲刺港股IPO（中国基金报）","url":"https://www.chnfund.com/article/AR7324d0cc-b4de-1dd7-6649-3a1ca2dd5835"}],"as_of":"2026-03","related_ids":["mobile-manipulator","mobile-manipulation","hkex-chapter-18c","one-brain-multiple-robots","machine-tending"],"name":"Youibot Robotics","alt":"优艾智合","abbr":"","aliases":["Youibot"],"one_liner":"A Chinese mobile-manipulation robot company mainly serving semiconductor, energy, and other factory settings.","explanation":"Youibot Robotics was founded in 2017 by Zhaohui Zhang, a Xi'an Jiao Tong University PhD, and others, starting out in Shenzhen; its listing entity is Hefei Youibot Robotics Co., Ltd. It builds mobile manipulators — a mobile base combined with a robot arm — plus fleet-scheduling software and models, describing its approach as “one brain, multiple forms.” Its main customers are semiconductor wafer fabs and the power, energy and chemicals, battery, and consumer-electronics industries, for tasks such as material handling, machine tending, and inspection. Shareholders include SIG Asia Investments, Lanchi Ventures, and SBVA (formerly SoftBank Ventures Asia). It first filed for a Hong Kong IPO in March 2025; after that filing lapsed on September 26, its refiling under Chapter 18C took effect on March 31, 2026, sponsored solely by CICC. Its prospectus disclosed 2025 revenue of about RMB 340 million and a net loss of about RMB 384 million.","example":"In a wafer fab, Youibot's mobile manipulators carry wafer cassettes between tools and handle machine loading and unloading.","related":["Mobile Manipulator","Mobile Manipulation","HKEX Chapter 18C","One Brain, Multiple Robots","Machine Tending"]},{"id":"dexterity","category":"company","sec":6,"tier":3,"sources":[{"title":"Dexterity: About Us","url":"https://dexterity.ai/about"},{"title":"Dexterity: Meet the Mech","url":"https://dexterity.ai/blog/meet-the-mech"},{"title":"Yahoo Finance: Dexterity secures $95m, reaching $1.65bn valuation","url":"https://finance.yahoo.com/news/dexterity-secures-95m-reaching-1-110002439.html"}],"as_of":"2026-09","related_ids":["tote-handling","palletizing-depalletizing","sorting","dual-arm-robot","mobile-manipulator","covariant"],"name":"Dexterity","alt":"Dexterity","abbr":"","aliases":["Dexterity AI","Dexterity, Inc."],"one_liner":"A US logistics-robotics company using AI dual-arm robots to load trucks and sort packages.","explanation":"Dexterity was founded in 2017, headquartered in Redwood City, California; its founder and CEO, Samir Menon, holds a computer science PhD from Stanford, and much of the founding team also came out of Stanford's robotics research community. It focuses on the heaviest, messiest jobs in logistics warehouses, such as loading and unloading trucks, depalletizing, and parcel sorting. Rather than betting on a single large model, it combines multiple smaller, task-specific models with force control and perception. In March 2025 it launched Mech, a mobile base carrying two large arms with a roughly 16-foot reach that the company says can move boxes weighing more than 130 pounds, aimed at truck loading; that same month it raised $95 million at a $1.65 billion valuation. The company says it logged more than 100 million autonomous actions in live production in 2025, and in 2026 FedEx named it a key technology partner at its investor day.","example":"At a warehouse dock, Mech takes boxes off a conveyor belt one at a time and packs them into a truck trailer — a physically demanding job with high worker turnover.","related":["Tote Handling","Palletizing / Depalletizing","Sorting","Dual-arm Robot","Mobile Manipulator","Covariant"]},{"id":"sereact","category":"company","sec":6,"tier":3,"sources":[{"title":"Zalando joins Sereact's $116M Series B (MassRobotics)","url":"https://www.massrobotics.org/zalando-joins-sereacts-116m-series-b-to-accelerate-ai-powered-warehouse-automation"},{"title":"AI Startup Sereact Raises $110 Million (Bloomberg)","url":"https://www.bloomberg.com/news/articles/2026-04-27/ai-startup-sereact-raises-110-million-for-robots-that-predict-consequences"}],"as_of":"2026-07","related_ids":["bin-picking","order-picking","world-model","dual-arm-robot","tote-handling"],"name":"Sereact","alt":"Sereact","abbr":"","aliases":[],"one_liner":"A German warehouse-robotics AI company that runs picking robots on its own Cortex model.","explanation":"Sereact was founded in Stuttgart, Germany, in 2021, later adding a Boston office; founders Ralf Gulde and Marc Tuscher both came out of the University of Stuttgart. It builds software for robots in warehouse logistics: its own robot foundation model, Cortex, and a 3D perception system drive robot arms through tasks such as order picking and returns processing, with customers including BMW, Mercedes-Benz, and IKEA, and more than 200 systems deployed. Its typical setups are single-arm picking stations and dual-arm returns-processing stations, and it is also developing a wheeled-base humanoid robot. It closed a €25 million Series A in January 2025; in April 2026 it raised a $110 million Series B led by Headline, which grew to $116 million after Zalando joined, with the funds going toward a model that lets robots predict the consequences of an action before taking it.","example":"In an e-commerce warehouse, Sereact's picking workstation uses a robot arm to grab differently shaped items out of a tote and place them into an order box.","related":["Bin Picking","Order Picking","World Model","Dual-arm Robot","Tote Handling"]},{"id":"world-labs","category":"company","sec":7,"tier":2,"sources":[{"title":"AMD to Acquire World Labs to Advance the Future of AI Compute","url":"https://ir.amd.com/news-events/press-releases/detail/1299/amd-to-acquire-world-labs-to-advance-the-future-of-ai-compute"},{"title":"World Labs: joining AMD","url":"https://www.worldlabs.ai/blog/amd-announcement"},{"title":"World Labs: Marble","url":"https://www.worldlabs.ai/blog/marble-world-model"}],"as_of":"2026-09","related_ids":["spatial-intelligence","world-model","marble","advanced-micro-devices","gaussian-splatting-based-simulation"],"name":"World Labs","alt":"World Labs","abbr":"","aliases":["Fei-Fei Li's World Labs"],"one_liner":"Fei-Fei Li's spatial-intelligence startup, which builds 3D world generation; AMD announced its acquisition in September 2026.","explanation":"World Labs was founded in 2024 by Fei-Fei Li together with Justin Johnson, Ben Mildenhall, and others, headquartered in San Francisco and focused on spatial intelligence — getting models to understand and generate three-dimensional worlds. When it came out of stealth in September 2024, it had raised $230 million in total funding at a $1 billion valuation. In November 2025 it launched Marble, which turns text, images, or video into a 3D scene a user can freely walk through and view, exportable as a Gaussian splat or a mesh; it raised a further $1 billion round in February 2026 and released its underlying Atlas model in September 2026. On September 28, 2026, AMD announced an all-stock acquisition of World Labs for about $8.2 billion, with Li becoming AMD's chief scientist; the deal is expected to close before the end of the year. The 3D scenes it generates can also serve as robot simulation environments.","example":"Marble can turn a single photo of a kitchen into a 3D scene, which can be exported as a mesh and dropped into a simulator as a robot training environment.","related":["Spatial Intelligence","World Model","Marble (World Labs)","Advanced Micro Devices","Gaussian Splatting-based Simulation"]},{"id":"ami-labs","category":"company","sec":7,"tier":2,"sources":[{"title":"Yann LeCun's AMI Labs raises $1.03B to build world models (TechCrunch, 2026-03)","url":"https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/"},{"title":"Why AMI Labs’ Alexandre LeBrun won’t call his AI AGI or superintelligence (TechCrunch, 2026-07)","url":"https://techcrunch.com/2026/07/16/why-ami-labs-alexandre-lebrun-wont-call-his-ai-agi-or-superintelligence/"},{"title":"Advanced Machine Intelligence Labs - Wikipedia","url":"https://en.wikipedia.org/wiki/Advanced_Machine_Intelligence_Labs"}],"as_of":"2026-07","related_ids":["joint-embedding-predictive-architecture","world-model","latent-world-model","v-jepa-2","meta-fundamental-ai-research","world-labs"],"name":"AMI Labs","alt":"AMI Labs","abbr":"AMI","aliases":["Advanced Machine Intelligence Labs","AMI"],"one_liner":"A Paris world-model company founded by Yann LeCun after leaving Meta, built around his JEPA approach.","explanation":"AMI Labs is a company founded by Turing Award winner Yann LeCun after leaving Meta, established in Paris in December 2025, with additional offices in New York, Montreal, and Singapore. LeCun serves as executive chairman; CEO Alexandre LeBrun was previously CEO of the medical-AI company Nabla, and became Nabla's chairman and chief AI scientist to take the role, with Nabla also becoming AMI's first partner; Saining Xie serves as chief scientist. The company builds world models around LeCun's proposed JEPA (Joint-Embedding Predictive Architecture): predicting what happens next in an abstract representation space, rather than generating output token by token the way large language models do. In March 2026 it closed a $1.03 billion seed round at a $3.5 billion pre-money valuation, with investors including NVIDIA, Samsung, and Toyota. Target industries include robotics, manufacturing, wearables, and healthcare; as of July 2026 it had no public product yet, saying it would release papers and code.","example":"","related":["Joint-Embedding Predictive Architecture","World Model","Latent World Model","V-JEPA 2","Meta Fundamental AI Research","World Labs"]},{"id":"gigaai","category":"company","sec":7,"tier":2,"sources":[{"title":"极佳科技官网","url":"https://gigaai.cc/"},{"title":"GigaAI GitHub organization (open-gigaai)","url":"https://github.com/open-gigaai"},{"title":"GigaBrain-0 (arXiv 2510.19430)","url":"https://arxiv.org/abs/2510.19430"}],"as_of":"2026-03","related_ids":["gigabrain-0","gigaworld-0","world-model","vision-language-action-model","synthetic-data","robochallenge"],"name":"GigaAI","alt":"极佳科技","abbr":"","aliases":["GigaAI Vision"],"one_liner":"A Chinese AI startup that uses world models to manufacture training data for embodied foundation models.","explanation":"GigaAI (also known as GigaAI Vision) is a Chinese AI company reportedly founded by Guan Huang. It started out building world models for autonomous driving, the DriveDreamer series, before pivoting to embodied AI. Its core approach is to use world models — models that can generate future video frames — to mass-produce training data, then use that data to train vision-language-action (VLA) models: GigaWorld-0 generates robot manipulation videos and 3D scenes, and the GigaBrain-0 series are the resulting VLA models, with code open-sourced under the open-gigaai organization on GitHub. The company has also released Maker H01, a robot with dual arms on a mobile base. Its website states it closed a RMB 1 billion Pre-B round in March 2026. It is one of the representative companies pursuing the “world model as data engine” approach.","example":"GigaBrain-0 is trained on real-robot data supplemented with video generated by GigaWorld-0; the company's website says it topped the RoboChallenge leaderboard in February 2026.","related":["GigaBrain-0","GigaWorld-0","World Model","Vision-Language-Action Model","Synthetic Data","RoboChallenge"]},{"id":"lightwheel","category":"company","sec":7,"tier":2,"sources":[{"title":"界面新闻：光轮智能完成10亿元融资","url":"https://www.jiemian.com/article/14098498.html"},{"title":"Lightwheel: RoboFinals","url":"https://lightwheel.ai/robofinals"},{"title":"GitHub: LightwheelAI/leisaac","url":"https://github.com/LightwheelAI/leisaac"}],"as_of":"2026-03","related_ids":["simulation-data","synthetic-data","simready-assets","nvidia-isaac-lab-arena","robofinals","leisaac"],"name":"Lightwheel","alt":"光轮智能","abbr":"","aliases":["Lightwheel AI"],"one_liner":"A company building simulation assets, synthetic data, and evaluation tools for robot training, a 2026 unicorn.","explanation":"Lightwheel was founded in January 2023, registered in Beijing; its founder and CEO, Chen Xie, graduated from Peking University and Columbia University and previously led autonomous-driving simulation at NVIDIA, Cruise, and NIO. It supplies simulation assets (3D objects and scenes with physical properties, ready to drop directly into a simulator), synthetic data, and simulation-based evaluation for robot training: it worked with NVIDIA to develop the Isaac Lab-Arena evaluation framework, open-sourced LeIsaac, a simulated-teleoperation tool for the SO-101 arm, and released the industrial-grade evaluation platform RoboFinals in December 2025. In March 2026 it closed combined Series A++ and A+++ funding of RMB 1 billion, with investors including New Hope Group; the company says this made it the first unicorn in the embodied-data space.","example":"A researcher can use LeIsaac to teleoperate a simulated SO-101 in Isaac Lab with a real leader arm, then fine-tune a GR00T policy on the collected data.","related":["Simulation Data","Synthetic Data","SimReady Assets","NVIDIA Isaac Lab-Arena","RoboFinals (Lightwheel industrial-grade simulation evaluation platform)","LeIsaac"]},{"id":"general-intuition","category":"company","sec":7,"tier":3,"sources":[{"title":"General Intuition 官网","url":"https://www.generalintuition.com/"},{"title":"Valor, Point72 back General Intuition at $6B valuation (TechCrunch, 2026-08-24)","url":"https://techcrunch.com/2026/08/24/valor-point72-back-general-intuition-at-6b-valuation-as-ai-startup-pushes-into-robotics/"}],"as_of":"2026-08","related_ids":[null,null,null,null,null],"name":"General Intuition","alt":"General Intuition","abbr":"","aliases":[],"one_liner":"An AI lab training world models and agents on massive amounts of gameplay footage.","explanation":"General Intuition is an AI research company headquartered in New York, spun out in October 2025 from Medal, a platform for sharing gameplay clips; its CEO is Medal founder Pim de Witte. Its premise is that Medal holds enormous volumes of player gameplay recordings, complete with logged keyboard and mouse input — effectively a huge set of observation-action pairs — that can be used to train world models and agents that understand space and time and can predict the consequences of an action, which can then be transferred to physical settings like robotics. It has released MIRA, a multiplayer world model built with Kyutai and Epic Games, trained on Rocket League footage and running in real time at 20 frames per second. On funding, it raised a seed round of about $134 million in October 2025, then $320 million at a $2.3 billion valuation in June 2026; TechCrunch reported that by August it was in talks for a new round at a $6 billion pre-money valuation, with robotics as a key focus.","example":"","related":["World Model","Observation-Action Pair","Internet Video Data","Minecraft Environments (MineDojo / MineRL)","Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos"]},{"id":"runway","category":"company","sec":7,"tier":3,"sources":[{"title":"Runway (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Runway_(company)"},{"title":"Introducing Runway GWM-1","url":"https://runway.com/research/introducing-runway-gwm-1"}],"as_of":"2026-02","related_ids":["video-generation-model","world-model","interactive-world-model","neural-simulator","world-model-based-policy-evaluation","synthetic-data"],"name":"Runway","alt":"Runway","abbr":"","aliases":["Runway AI","RunwayML"],"one_liner":"A New York video-generation company that followed its Gen series with GWM-1, a world model with a robotics variant.","explanation":"Runway was founded in New York in 2018 by three co-founders — Cristóbal Valenzuela, Alejandro Matamala, and Anastasis Germanidis — who met at NYU's ITP program for art and technology. It became known for its Gen-series video-generation models (Gen-4.5 shipped in late 2025), and in 2022 it co-released Stable Diffusion together with the CompVis group at LMU Munich. On December 11, 2025, building on Gen-4.5, it released the general-purpose world model GWM-1: an autoregressive model that generates video frame by frame in real time and can be steered with inputs such as camera movement or robot actions; its GWM Robotics variant provides a Python SDK that generates multi-view future frames conditioned on robot actions, used to synthesize training data and evaluate policies in a virtual environment. On funding, it raised $315 million in February 2026 at a $5.3 billion valuation.","example":"","related":["Video Generation Model","World Model","Interactive World Model","Neural Simulator","World-Model-based Policy Evaluation","Synthetic Data"]},{"id":"shengshu-technology","category":"company","sec":7,"tier":3,"sources":[{"title":"生数科技完成近20亿元B轮融资（量子位）","url":"https://www.qbitai.com/2026/04/398772.html"},{"title":"京企生数科技完成超6亿元A+轮融资（国家科技传播中心）","url":"https://www.ncsti.gov.cn/kjdt/xwjj/202602/t20260208_237672.html"}],"as_of":"2026-04","related_ids":["vidar","motus","video-generation-model","world-action-model","world-model","diffusion-transformer"],"name":"ShengShu Technology","alt":"生数科技","abbr":"","aliases":["ShengShu"],"one_liner":"A Tsinghua-linked multimodal generation company behind the Vidu video model, now extending into embodied world models.","explanation":"ShengShu Technology was founded in Beijing in March 2023, with a core team from Tsinghua University; co-founder and chief scientist Jun Zhu is a Tsinghua computer science professor, and Jiayu Tang is CEO. In September 2022 the team proposed the U-ViT architecture, one of the earlier works to use a Transformer as the backbone of a diffusion model. Its signature product is the video-generation model Vidu, released in April 2024 and launched globally that July. It has extended video generation into robotics: working with Tsinghua on the embodied model Vidar, which predicts future frames with a video diffusion model and then decodes them into actions, and separately building the world-action model Motus. In April 2026 it closed a Series B of nearly RMB 2 billion led by Alibaba Cloud, positioning itself as building a “general-purpose world model.”","example":"ShengShu's Vidar first generates a video of the robot completing a task, then uses an inverse-dynamics model to infer the action at each step from the video.","related":["Vidar","Motus","Video Generation Model","World Action Model","World Model","Diffusion Transformer"]},{"id":"hillbot","category":"company","sec":7,"tier":3,"sources":[{"title":"Hillbot 官网","url":"https://www.hillbot.ai/"},{"title":"ManiSkill (haosulab) GitHub","url":"https://github.com/haosulab/ManiSkill"}],"as_of":"2026-09","related_ids":["maniskill","sapien","synthetic-data","sim-to-real-transfer","real-robot-data-camp-vs-sim-data-camp","gpu-accelerated-parallel-simulation"],"name":"Hillbot","alt":"Hillbot","abbr":"","aliases":[],"one_liner":"A US embodied-AI startup training robot skills with a mix of simulated and real-robot data.","explanation":"Hillbot is a US embodied-AI startup whose slogan is “building general-purpose robots one skill at a time.” Its approach combines real-world data with large volumes of synthetic data generated in a simulator to train manipulation skills that generalize; its website shows demos across multiple robot bodies, including an arm opening a cabinet door, a quadruped, and a dual-arm system. Its website lists the open-source simulation platform ManiSkill and the simulation framework SAPIEN as its own products; both projects came out of Hao Su's lab at UC San Diego, and Su is reportedly a co-founder of Hillbot. Its founding date and funding are not disclosed on its website. For newcomers, it's a representative example of the “simulation data camp” among robotics startups.","example":"Large numbers of cabinet-opening and grasping demonstrations are first generated in parallel on GPUs inside ManiSkill, then used together with a small amount of real-robot data to train a policy.","related":["ManiSkill","SAPIEN (SimulAted Part-based Interactive ENvironment)","Synthetic Data","Sim-to-Real Transfer","Real-Robot-Data Camp vs. Sim-Data Camp","GPU-Accelerated Parallel Simulation"]},{"id":"sudo-technology","category":"company","sec":7,"tier":3,"sources":[{"title":"20亿美金苏度科技具身首秀：0真机数据，zero-shot，98%首次抓取成功率（量子位 / 腾讯新闻，2026-04-20）","url":"https://news.qq.com/rain/a/20260420A056AL00"},{"title":"上海，跑出一家百亿独角兽！（东方财富财富号，2026-04-23）","url":"https://caifuhao.eastmoney.com/news/20260423145145282740910"}],"as_of":"2026-04","related_ids":["sim-to-real-transfer","reinforcement-learning","zero-shot","real-robot-data-camp-vs-sim-data-camp","hillbot","3d-vision"],"name":"Sudo Technology","alt":"苏度科技","abbr":"","aliases":["Sudo"],"one_liner":"A Shanghai embodied-AI company betting on simulation-only training, launching the Sudo R1 robot system in 2026.","explanation":"Sudo Technology was founded in Shanghai in May 2025. Co-founder and CEO Zheng Han is a serial entrepreneur who helped start the smart-hardware company ZEPP; chief technology advisor Hao Su was a core contributor to ImageNet, ShapeNet, and PointNet, formerly an associate professor at UC San Diego, and became a professor at Fudan University in 2026; much of the core team comes from Hillbot, which Su also helped found. The company follows a sim-to-real path, training mainly through simulation and reinforcement learning while trying to avoid real-robot teleoperated data. In April 2026 it released its first fully self-developed hardware-and-software robot system, #Sudo R1, described as using an integrated “3D world model plus reinforcement learning” design, with a company-reported 98% zero-shot first-attempt grasp success rate. That same month it reportedly closed a $500 million Pre-A round at a roughly $2 billion valuation.","example":"In the #Sudo R1 launch demo, the company said no real-robot data was used in training at all — the model went straight from training to grasping unfamiliar objects on a real robot.","related":["Sim-to-Real Transfer","Reinforcement Learning","Zero-shot","Real-Robot-Data Camp vs. Sim-Data Camp","Hillbot","3D Vision"]},{"id":"dexforce","category":"company","sec":7,"tier":3,"sources":[{"title":"跨维智能官网：公司介绍","url":"https://www.dexforce.com/about.html"},{"title":"跨维智能完成10亿元B轮融资 - DexForce","url":"https://dexforce.com/news_detail.php?id=1261"},{"title":"证券时报：跨维智能完成10亿元B轮融资将迎IPO","url":"https://www.stcn.com/article/detail/3988923.html"}],"as_of":"2026-07","related_ids":["generative-simulation","synthetic-data","simulation-data","sim-to-real-transfer","world-model","real-robot-data-camp-vs-sim-data-camp"],"name":"DexForce","alt":"跨维智能","abbr":"","aliases":["DexForce Technology"],"one_liner":"A Shenzhen embodied-AI company that trains robots mainly on synthetic data from generative simulation.","explanation":"DexForce was founded in Shenzhen in June 2021. Its founder, Kui Jia, is a professor at the Chinese University of Hong Kong, Shenzhen, with a long research background in deep learning and 3D geometric vision. Where most companies collect data by having a person teleoperate a real robot, DexForce takes a “generative simulation” approach: its own DexVerse embodied-AI engine generates virtual scenes and synthetic data in bulk to train models. Jia's reasoning is that physical variation — object positions, materials, environments — can all be generated in simulation, while “what to do,” semantic knowledge, still depends more on real data. Its other products include the DexWorldModel world model, the DexForce W1 Pro humanoid robot, and the DexSense vision-only sensor; the company says it is already deployed across more than 50 industry verticals. On June 30, 2026, it announced a Series B round of more than RMB 1 billion at a post-money valuation above RMB 10 billion, and is reportedly preparing an IPO.","example":"Jia has estimated that a single teleoperator can only collect 100-150 demonstrations a day, which is why DexForce chose to generate large volumes of synthetic data in simulation to cover physical-level generalization instead.","related":["Generative Simulation","Synthetic Data","Simulation Data","Sim-to-Real Transfer","World Model","Real-Robot-Data Camp vs. Sim-Data Camp"]},{"id":"songying-technology","category":"company","sec":7,"tier":3,"sources":[{"title":"松应科技新闻中心（官网）","url":"https://www.orca3d.cn/news.html"},{"title":"实时物理AI仿真平台松应科技完成天使轮融资 中科创星领投（品玩，2025-03-20）","url":"https://www.pingwest.com/a/303218"},{"title":"NIE 2025 | 松应科技聂凯旋：物理AI仿真系统（弗若斯特沙利文）","url":"https://www.frostchina.com/content/activity/detail/68bfe473f3453aa11bf09fbf"}],"as_of":"2026-08","related_ids":["simulator","nvidia-isaac-sim","simulation-data","sim-to-real-gap","domestic-substitution","embodied-ai-training-ground"],"name":"Songying Technology","alt":"松应科技","abbr":"","aliases":["ORCA"],"one_liner":"A Shanghai maker of the ORCA physical-AI simulation platform, positioned as a domestic alternative to NVIDIA Isaac Sim.","explanation":"Songying Technology was founded in Shanghai in November 2021; founder and CEO Kaixuan Nie previously served as deputy general manager of Huawei Cloud's Kunpeng solutions unit and as head of its cloud-gaming business. Its main product is the physical-AI simulation platform ORCA (now called ORCA OS), which provides physically based rendering, parallel simulation, sensor simulation, and synthetic-data generation, letting robots train and collect data in a virtual environment before transferring to real hardware; it is positioned as a domestic alternative to NVIDIA's Isaac Sim / Omniverse and is adapted for Chinese-made GPUs. ORCA 1.0 was commercialized in late 2024, with customers including humanoid-robot body makers and national- and provincial-level humanoid robot innovation centers. On funding, it closed an angel round led by Casstar in March 2025; in August 2026 it announced consecutive Series A and A1 rounds totaling several hundred million RMB, led by CICC Capital and Wuhan Hongshan Capital.","example":"When the National and Local Co-built Humanoid Robotics Innovation Center's “Qilin” training ground launched, its virtual training environment was built by Songying's ORCA, used to generate simulated training data for humanoid robots.","related":["Simulator","NVIDIA Isaac Sim","Simulation Data","Sim-to-Real Gap (Reality Gap)","Domestic Substitution","Embodied AI Training Ground (Robot Data Collection Center)"]},{"id":"51world","category":"company","sec":7,"tier":3,"sources":[{"title":"51Aes 企业简介","url":"https://www.51aes.com/about/company"}],"as_of":"2026-09","related_ids":["digital-twin","synthetic-data","real-to-sim-to-real","physical-ai","simulator","autonomous-driving"],"name":"51WORLD","alt":"五一视界","abbr":"","aliases":["Beijing 51WORLD Digital Twin Technology Co., Ltd.","51Aes"],"one_liner":"A Beijing digital-twin company, listed in Hong Kong, that supplies simulation scenes and synthetic data for embodied AI.","explanation":"51WORLD, formally Beijing 51WORLD Digital Twin Technology Co., Ltd., was founded in Beijing in February 2015 and is listed on the Hong Kong Stock Exchange (ticker 6651.HK). Its core business is digital twins: using computer graphics and AI to rebuild real-world places such as cities, industrial parks, factories, and transportation infrastructure as real-time-rendered 3D scenes, through products including its AES digital-twin platform and the low-code development platform WDP, along with the synthetic-data and simulation platform 51Sim and the digital-earth platform 51Earth. In recent years it has positioned itself as “physical AI infrastructure,” applying its 3D-scene and simulation capabilities to embodied AI by providing training data and test environments for a real-to-sim-to-real training loop.","example":"A real warehouse can be rebuilt as a digital-twin scene, then used to generate large batches of labeled synthetic images for training a robot's perception model.","related":["Digital Twin","Synthetic Data","Real-to-Sim-to-Real","Physical AI","Simulator","Autonomous Driving"]},{"id":"manycore-tech","category":"company","sec":7,"tier":3,"sources":[{"title":"群核科技（维基百科）","url":"https://zh.wikipedia.org/wiki/群核科技"},{"title":"SpatialLM: Training Large Language Models for Structured Indoor Modeling（arXiv 2506.07491）","url":"https://arxiv.org/abs/2506.07491"}],"as_of":"2026-04","related_ids":["spatiallm","spatial-intelligence","synthetic-data","simulation-assets","hangzhou-s-six-little-dragons","digital-twin"],"name":"Manycore Tech","alt":"群核科技","abbr":"","aliases":["Kujiale","Coohom"],"one_liner":"A Hangzhou spatial-design software company behind Kujiale, now also building spatial intelligence and simulation data.","explanation":"Manycore Tech was founded in Hangzhou in 2011; its three founders, Xiaohuang Huang, Hang Chen, and Hao Zhu, all did graduate studies at the University of Illinois Urbana-Champaign. Its main business is the cloud-based interior-design software Kujiale (launched in 2013) and its international version, Coohom, which together have amassed a huge library of renderable indoor 3D scenes. In November 2024 it launched the spatial-intelligence platform SpatialVerse, using those scenes to generate simulation training environments and synthetic data for robots; in 2025 it open-sourced SpatialLM, which converts point clouds into structured floor-plan descriptions. It is one of the so-called “Six Little Dragons of Hangzhou,” and on April 17, 2026 it listed on the Hong Kong Stock Exchange, the first of the six to go public.","example":"A robotics team pulls complete indoor scenes from SpatialVerse into a simulator to generate training data for household tasks in bulk.","related":["SpatialLM","Spatial Intelligence","Synthetic Data","Simulation Assets","Hangzhou's Six Little Dragons","Digital Twin"]},{"id":"scale-ai","category":"company","sec":7,"tier":3,"sources":[{"title":"Scale AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Scale_AI"}],"as_of":"2026-03","related_ids":["data-annotation","embodied-ai-data-service-provider","reinforcement-learning-from-human-feedback","autonomous-driving","openai"],"name":"Scale AI","alt":"Scale AI","abbr":"","aliases":[],"one_liner":"A US AI data-labeling company that supplies training data for large models and autonomous driving.","explanation":"Scale AI was founded in San Francisco in 2016 by Alexandr Wang and Lucy Guo. Its core business is data annotation and model evaluation: it organizes large teams of human labelers to prepare training data for autonomous driving, computer vision, and large language models, including the human-preference data needed for reinforcement learning from human feedback (RLHF). In June 2025, Meta paid over $14 billion for a 49% non-voting stake, and Wang joined Meta to lead its superintelligence team, with Jason Droege taking over as CEO; in March 2026 the company launched a research division called Scale Labs. In embodied AI, Scale represents the “data-service provider” role, and it is reported to be expanding into robotics and physical-AI data as well.","example":"Scale's early customers were mostly autonomous-driving companies, which used it to label vehicles and pedestrians in LiDAR point clouds and camera images.","related":["Data Annotation","Embodied AI Data Service Provider","Reinforcement Learning from Human Feedback","Autonomous Driving","OpenAI"]},{"id":"build-ai","category":"company","sec":7,"tier":3,"sources":[{"title":"Build AI 官网","url":"https://www.build.ai/"},{"title":"Egocentric-10K on Hugging Face","url":"https://huggingface.co/datasets/builddotai/Egocentric-10K"}],"as_of":"2026-03","related_ids":["egocentric-10k",null,null,null,"embodied-ai-data-service-provider"],"name":"Build AI","alt":"Build AI","abbr":"","aliases":[],"one_liner":"A US embodied-data company that pays factory workers to wear head cameras and record first-person video.","explanation":"Build AI is an embodied-data startup headquartered in San Francisco with an additional site in Shenzhen, founded by Eddy Xu, Patrick Rim, Jonathan Jia, and Zikang Jiang, and registered as a public-benefit corporation. It describes itself as a “hyperscale data supplier for physical AI,” handling everything in-house from capture hardware and production to logistics, data collection, and model training; its core product is first-person video of workers performing real jobs in actual factories, recorded with head-mounted cameras. In August 2025 it open-sourced Egocentric-10K (about 10,000 hours from 85 factories), followed by Egocentric-100K (December 2025) and Egocentric-1M (March 2026). Its website states it has raised more than $30 million and reached an annualized revenue run rate of $100 million in 2026.","example":"Egocentric-10K is open-sourced on Hugging Face under Apache 2.0, and can be used for pretraining on human videos or for latent-action pretraining.","related":["Egocentric-10K","Egocentric Video","Pretraining on Human Videos","Robot-free (Embodiment-free) Data Collection","Embodied AI Data Service Provider"]},{"id":"inspire-robots","category":"company","sec":8,"tier":2,"sources":[{"title":"INSPIRE ROBOTS - About Us","url":"https://en.inspire-robots.com/about-us/"},{"title":"因时机器人官网","url":"https://www.inspire-robots.com/"}],"as_of":"2026-08","related_ids":["inspire-rh56-dexterous-hand",null,"linear-actuator",null,null,null],"name":"Inspire Robots","alt":"因时机器人","abbr":"","aliases":["Inspire","INSPIRE-ROBOTS"],"one_liner":"A Beijing maker of dexterous hands and miniature electric linear actuators; its RH56 hand appears on many humanoid robots.","explanation":"Inspire Robots was founded in 2016 in Beijing's Shijingshan district and specializes in miniature precision motion components. Its core product is a small linear electric actuator, which it then builds into five-fingered humanoid dexterous hands and electric parallel grippers. In its RH56 series of dexterous hands, each finger is driven by one small linear actuator through a linkage, giving a mid-range price and reliability that has made it a popular choice: it ships on humanoid robots such as Unitree's H1 and G1 and Fourier's GR1, and shows up frequently as the hand used in embodied-AI papers. The company says it has more than 2,000 customers and, as of 2025, had shipped over 10,000 dexterous hands; it raised a pre-A round in 2017 and a B+ round in 2023. At the World Robot Conference in August 2026, it launched the RH524J1 series, a 24-degree-of-freedom dexterous hand.","example":"Humanoid-robot papers such as HumanPlus and OKAMI both used Inspire's 6-DoF dexterous hand as their experimental hardware.","related":["Inspire RH56 Dexterous Hand","Dexterous Hand","Linear Actuator (Electric Cylinder)","Linkage Transmission","Unitree H1","Core Components"]},{"id":"linkerbot","category":"company","sec":8,"tier":2,"sources":[{"title":"灵心巧手完成近15亿元B轮融资（证券时报网，2026-02-12）","url":"https://www.stcn.com/article/detail/3641823.html"},{"title":"超亿元，机器人灵巧手最大种子轮诞生！（证券时报网）","url":"https://stcn.com/article/detail/1647503.html"},{"title":"灵心巧手官网","url":"https://www.linkerbot.cn"}],"as_of":"2026-05","related_ids":["linker-hand","dexterous-hand","data-glove","tendon-driven-actuation","end-effector","humanoid-robot"],"name":"Linkerbot","alt":"灵心巧手","abbr":"","aliases":["LinkerBot"],"one_liner":"A Beijing dexterous-hand company, maker of the high-DoF Linker Hand series of robot hands.","explanation":"Linkerbot was founded in 2019, headquartered in Beijing; its founder and CEO, Yong Zhou, has more than a decade of experience in internet products and robotics. It specializes in dexterous hands (multi-fingered, multi-degree-of-freedom end effectors modeled on the human hand), launching the Linker Hand series in 2024 with degrees of freedom ranging from 6 to 42 and covering linkage, tendon-driven, and direct-drive transmission, plus data gloves and teleoperation arms for data capture. In April 2025 it raised a seed round of more than RMB 100 million, and in February 2026 it closed a Series B of nearly RMB 1.5 billion at a valuation above RMB 10 billion, targeting deliveries of 50,000 to 100,000 dexterous hands that year; it was reportedly seeking a valuation of about $6 billion in its next round as of May.","example":"A humanoid-robot company mounts a Linker Hand L20 at the end of an arm; an operator wearing a data glove teleoperates it to collect grasping and pinching data.","related":["Linker Hand","Dexterous Hand","Data Glove","Tendon-Driven Actuation","End Effector","Humanoid Robot"]},{"id":"shenzhen-zhaowei-machinery-and-electronics","category":"company","sec":8,"tier":3,"sources":[{"title":"兆威机电港股上市，「灵巧手」龙头开启全球化新征程（证券时报网）","url":"https://www.stcn.com/article/detail/3667158.html"},{"title":"资讯：灵巧手微型驱动龙头兆威机电港股上市（艾邦机器人）","url":"https://www.aibangbots.com/a/7704"}],"as_of":"2026-09","related_ids":["dexterous-hand","coreless-motor","micro-lead-screw","speed-reducer-gearbox","core-components"],"name":"Shenzhen Zhaowei Machinery & Electronics","alt":"兆威机电","abbr":"","aliases":["Zhaowei"],"one_liner":"A Shenzhen micro-transmission and micro-drive maker building dexterous hands and micro gearboxes, dual-listed A+H.","explanation":"Zhaowei Machinery & Electronics was founded in Shenzhen in 2001, making micro transmission systems, micro-motors, and electronic control systems used in smart cars, consumer electronics and medical devices, and industrial equipment; it says it can mass-produce transmission systems under 6 millimeters in diameter, down to as small as 3.4 millimeters. Since a dexterous hand's fingers have very little internal space, fitting a motor and reduction mechanism into a finger joint calls for exactly this kind of micro transmission, which is why Zhaowei built its ZWHAND series of dexterous hands and is described as the first Chinese company to commercialize a high-degree-of-freedom dexterous hand. It listed on the Shenzhen Stock Exchange in December 2020 (ticker 003021) and on the Hong Kong Stock Exchange's Main Board on March 9, 2026 (02692.HK), becoming a dual A+H company, with the proceeds going mainly toward its embodied-robotics business; it released a next-generation dexterous hand in September 2026.","example":"A humanoid-robot maker might buy Zhaowei's complete dexterous hand, or just its micro gearboxes to fit into its own finger-joint design.","related":["Dexterous Hand","Coreless Motor","Micro Lead Screw","Speed Reducer / Gearbox","Core Components"]},{"id":"brainco","category":"company","sec":8,"tier":3,"sources":[{"title":"Wikipedia: BrainCo","url":"https://en.wikipedia.org/wiki/BrainCo"},{"title":"BrainCo Revo 2 参数文档","url":"https://www.brainco-hz.com/docs/revolimb-hand/en/revo2/parameters.html"}],"as_of":"2026-01","related_ids":["brainco-revo-hand",null,"hangzhou-s-six-little-dragons",null,null],"name":"BrainCo","alt":"强脑科技","abbr":"","aliases":[],"one_liner":"A Hangzhou brain-computer-interface company that extended its smart-prosthetics work into the Revo robot hand.","explanation":"BrainCo was founded in 2015 by Bicheng Han, a Harvard PhD student in brain science, and is headquartered in Hangzhou with an additional office in Somerville, Massachusetts; it is one of the so-called “Six Little Dragons of Hangzhou.” Its main business is non-invasive brain-computer interfaces: headbands that read EEG (electroencephalogram) brain signals for focus- and sleep-related products, plus smart prosthetic hands and legs controlled by muscle (EMG) signals. Its connection to embodied AI comes from turning that prosthetics technology into the Revo series of robot dexterous hands, which are often mounted on humanoid robots. On funding, it was valued at more than $1.3 billion in August 2025, and reportedly raised about RMB 2 billion in January 2026 while confidentially filing for a Hong Kong Stock Exchange listing sponsored by CICC and UBS.","example":"The Revo 2 dexterous hand weighs about 383 grams, uses 6 motors to drive 11 degrees of freedom, and can be mounted on a humanoid robot over CAN FD or EtherCAT.","related":["BrainCo Revo Hand","Dexterous Hand","Hangzhou's Six Little Dragons","Underactuation","Humanoid Robot"]},{"id":"agilink","category":"company","sec":8,"tier":3,"sources":[{"title":"头部互联网大厂领投，「临界点」再获数亿元融资丨智能涌现独家（36氪）","url":"https://m.36kr.com/p/3678179258425990"},{"title":"智元拆出一家「灵巧手」独角兽，成立不足5月估值破10亿美元（澎湃/搜狐）","url":"https://m.sohu.com/a/1024972075_260616"}],"as_of":"2026-06","related_ids":["agibot","dexterous-hand","fully-direct-drive-dexterous-hand","agibot-omnihand","core-components","dexterous-manipulation"],"name":"AGILINK","alt":"临界点","abbr":"AGILINK","aliases":["Shanghai AGILINK Innovative Intelligent Technology Co., Ltd."],"one_liner":"A dexterous-hand company spun out of AgiBot's own dexterous-hand division.","explanation":"AGILINK was registered in Shanghai in January 2026, spun out of AgiBot's dexterous-hand division and held by AgiBot-affiliated companies. Its founder, Kun Xiong, graduated from the Hong Kong University of Science and Technology's Robotics Institute, was an early member of Tencent's Robotics X lab, later worked at the International Digital Economy Academy (IDEA) and Inovance Technology, and joined AgiBot in November 2024 to lead its dexterous-hand work. The company builds multi-degree-of-freedom five-fingered dexterous hands and grippers, used to give robots fine-grained grasping and manipulation abilities. It raised funding in quick succession after founding, and by May-June 2026 was reportedly valued at over $1 billion, becoming a unicorn in the dexterous-hand space. In April 2026 it launched the OmniHand 3 series, whose flagship Ultra-T model uses a hybrid of tendon-driven and direct-drive actuation. It is one example of AgiBot's strategy of spinning off core components into independently funded companies.","example":"At CES 2026, an LG robot fitted with AGILINK's dexterous hand reportedly demonstrated fine-grained manipulation.","related":["AgiBot","Dexterous Hand","Fully Direct-Drive Dexterous Hand","AgiBot OmniHand (OmniHand 2025 / OmniHand Pro 2025)","Core Components","Dexterous Manipulation"]},{"id":"shadow-robot-company","category":"company","sec":8,"tier":3,"sources":[{"title":"Shadow Robot unveils DEX-EE, developed with Google DeepMind","url":"https://shadowrobot.com/shadow-robot-unveils-the-worlds-most-robust-dexterous-robot-hand-developed-in-partnership-with-google-deepmind"},{"title":"Shadow Robot Dexterous Hand & DEX-EE (Robots International)","url":"https://www.robotsinternational.com/Shadow-Robot.htm"}],"as_of":"2026-09","related_ids":["shadow-dexterous-hand","dex-ee","dexterous-hand","dactyl","google-deepmind","tendon-driven-actuation"],"name":"Shadow Robot Company","alt":"Shadow Robot Company","abbr":"","aliases":["Shadow Robot"],"one_liner":"A long-running British dexterous-hand company behind the Shadow Hand and the DEX-EE.","explanation":"The Shadow Robot Company was founded in London in 1987, making it one of the oldest robotics companies still operating, now led by Rich Walker. Its best-known product is the Shadow Dexterous Hand: a tendon-driven, five-fingered hand roughly the size of a human hand with more than 20 degrees of freedom, long considered the standard hardware for dexterous-manipulation research. OpenAI's Dactyl project used it to train a policy in simulation and transfer it to the real hand, achieving in-hand object reorientation and single-handed Rubik's-cube solving. The company later partnered with Google DeepMind to develop the three-fingered DEX-EE hand, built for durability so it can withstand the long, intensive real-robot training that reinforcement learning requires while collecting large amounts of tactile and other sensor data.","example":"OpenAI's 2019 demonstration of a Shadow Hand solving a Rubik's Cube one-handed is a classic example of sim-to-real transfer.","related":["Shadow Dexterous Hand","DEX-EE (Shadow Robot × Google DeepMind)","Dexterous Hand","Dactyl","Google DeepMind","Tendon-Driven Actuation"]},{"id":"robotiq","category":"company","sec":8,"tier":3,"sources":[{"title":"Robotiq 官网","url":"https://robotiq.com/"},{"title":"Robotiq | LinkedIn","url":"https://www.linkedin.com/company/robotiq"},{"title":"2F-85 & 2F-140 Adaptive Gripper","url":"https://robotiq.com/products/2f85-140-adaptive-robot-gripper"}],"as_of":"2026-09","related_ids":["robotiq-2f-85-gripper",null,null,null,"droid",null],"name":"Robotiq","alt":"Robotiq","abbr":"","aliases":[],"one_liner":"A Canadian maker of grippers and plug-and-play accessories for collaborative robot arms.","explanation":"Robotiq is a Canadian automation company founded in 2008 and headquartered in Lévis, Quebec, with additional offices in North America and Europe. It mainly builds plug-and-play accessories and workstations for collaborative robot arms: adaptive two-finger grippers (2F-85, 2F-140), the parallel Hand-E gripper, vacuum grippers, the FT 300 six-axis force/torque sensor, wrist cameras, tactile fingertips, and complete workstations for palletizing, machine tending, and screwdriving. In embodied-AI research, the 2F-85 has become something close to a default end effector on Franka and UR arms, and large-scale datasets such as DROID were collected using it. In recent years, Robotiq has open-sourced ROS 2 drivers, a C++ SDK, and Isaac Sim simulation assets, and has begun positioning itself around physical AI.","example":"The DROID dataset was collected on a Franka Panda arm fitted with a Robotiq 2F-85 gripper.","related":["Robotiq 2F-85 Gripper","Adaptive / Underactuated Gripper","Collaborative Robot","Six-Axis Force/Torque Sensor","DROID (Distributed Robot Interaction Dataset)","Palletizing / Depalletizing"]},{"id":"dh-robotics","category":"company","sec":8,"tier":3,"sources":[{"title":"大寰机器人官网","url":"https://www.dh-robotics.com"},{"title":"2026世界机器人大会展商信息：大寰机器人","url":"https://www.worldrobotconference.com/expo/company/706.html"},{"title":"机器人大讲堂：核心团队师承戴建生院士，大寰机器人完成C+轮融资","url":"https://www.leaderobot.com/news/3743"}],"as_of":"2026-08","related_ids":["gripper","electric-gripper","dexterous-hand","end-effector","force-control","core-components"],"name":"DH-Robotics","alt":"大寰机器人","abbr":"","aliases":["Dahuan","Shenzhen DH-Robotics Technology Co., Ltd."],"one_liner":"A Shenzhen maker of electric grippers, also building dexterous hands and precision force-controlled actuators.","explanation":"DH-Robotics was founded in Shenzhen in December 2016. Its founder, Jie Sun, earned an undergraduate degree in mechanical engineering from Tsinghua University and a PhD from King's College London, and its core team studied under mechanism scholar Jian S. Dai. It builds the “hand” at the end of a robot arm and related parts: electric grippers (driven by a motor rather than a pneumatic cylinder, giving precise control over opening position and grip force), dexterous hands, voice-coil actuators, servo electric cylinders, and drivers, with precision force control and direct-drive force feedback as its core technology. Its products are used mainly on production lines for consumer electronics, auto parts, and medical devices, and also pair with collaborative robot arms and humanoid robots. ByteDance took a strategic stake in 2022, and in February 2024 the company closed a Series C+ round with participation from Qiming Venture Partners and others. It is designated a national-level “Little Giant” enterprise for specialized, sophisticated innovation.","example":"On a phone camera-module production line, a robot arm fitted with a DH-Robotics electric gripper picks up fragile lens components using a small, precisely set grip force.","related":["Gripper","Electric Gripper","Dexterous Hand","End Effector","Force Control","Core Components"]},{"id":"leaderdrive","category":"company","sec":8,"tier":3,"sources":[{"title":"财联社：绿的谐波筹划发行H股并在香港联交所上市","url":"https://www.cls.cn/detail/2465192"},{"title":"钛媒体：卖铲人绿的谐波","url":"https://www.tmtpost.com/8108147.html"},{"title":"新浪财经：绿的谐波科创板招股说明书","url":"https://vip.stock.finance.sina.com.cn/corp/view/vISSUE_RaiseExplanationDetail.php?stockid=688017&id=6536262"}],"as_of":"2026-08","related_ids":["strain-wave-gear","laifual","speed-reducer-gearbox","joint-actuator-module","domestic-substitution","tesla-supply-chain"],"name":"Leaderdrive","alt":"绿的谐波","abbr":"","aliases":["Suzhou Leaderdrive Transmission Technology"],"one_liner":"China's leading harmonic-reducer maker, listed on the STAR Market and now planning a Hong Kong listing.","explanation":"Suzhou Leaderdrive Transmission Technology Co., Ltd. was founded in 2011; its chairman, Yuyu Zuo, began developing harmonic reducers independently in 2003, and he controls the company together with his younger brother, Jing Zuo. It broke Japan's Harmonic Drive Systems' long-standing monopoly on harmonic reducers, and listed on the STAR Market in August 2020 (688017), often called China's “first harmonic-reducer stock.” A harmonic reducer is compact and has very little backlash, making it a standard component in the arm joints of collaborative robots and humanoid robots, which has made Leaderdrive a sought-after supplier in the humanoid supply chain. In 2025 it reported RMB 571 million in revenue and RMB 124 million in net profit attributable to shareholders, with a domestic market share of about 27.5%. On August 26, 2026, its board approved a plan to issue H-shares and list on the Hong Kong Stock Exchange main board.","example":"The forearm and wrist joints of an industrial robot commonly use a harmonic reducer to convert a motor's high rotation speed into high torque.","related":["Strain Wave Gear (Harmonic Drive)","Laifual","Speed Reducer / Gearbox","Joint Actuator Module","Domestic Substitution","Tesla (Optimus) Supply Chain"]},{"id":"laifual","category":"company","sec":8,"tier":3,"sources":[{"title":"财联社：估值15亿的来福谐波冲刺港股IPO","url":"https://www.cls.cn/detail/2387467"},{"title":"香港经济日报：来福谐波（03952）IPO","url":"https://knowledge.hket.com/article/4151158/"},{"title":"经济通：来福谐波 03952 招股资讯","url":"https://www.etnet.com.hk/www/tc/stocks/ipo-info.php?code=03952"}],"as_of":"2026-06","related_ids":["strain-wave-gear","leaderdrive","joint-actuator-module","hkex-chapter-18c","domestic-substitution","core-components"],"name":"Laifual","alt":"来福谐波","abbr":"","aliases":["Laifual Drive","Laifual Transmission"],"one_liner":"A Zhejiang harmonic-reducer maker that listed in Hong Kong in June 2026.","explanation":"Zhejiang Laifual Drive Co., Ltd. was founded in 2013; its chairman, Jie Zhang, born in the 1990s and a graduate of the New Jersey Institute of Technology, has run the company since 2017. Its main business is harmonic reducers, extended into joint actuator modules, robot arms, and automation workstations, sold mainly for humanoid and industrial robots. Citing data from consultancy CIC, its prospectus states that by 2025 shipment volume it ranked second among China's robot harmonic-reducer suppliers, with a 21.4% share, behind only Leaderdrive; it had about RMB 261 million in revenue in 2025 and was still unprofitable. Having twice failed to complete a STAR Market listing, it listed on the Hong Kong Stock Exchange main board on June 30, 2026, under ticker 03952.","example":"A humanoid robot arm's rotary joint often pairs a Laifual or Leaderdrive harmonic reducer with a frameless motor to form an integrated joint actuator module.","related":["Strain Wave Gear (Harmonic Drive)","Leaderdrive","Joint Actuator Module","HKEX Chapter 18C","Domestic Substitution","Core Components"]},{"id":"nabtesco","category":"company","sec":8,"tier":3,"sources":[{"title":"Precision Reduction Gears | Nabtesco Corporation","url":"https://www.nabtesco.com/en/products/robot"},{"title":"Nabtesco 2026 Company Profile | PitchBook","url":"https://pitchbook.com/profiles/company/60420-79"}],"as_of":"2026-09","related_ids":["rotate-vector-reducer","cycloidal-reducer","speed-reducer-gearbox","strain-wave-gear","big-four-of-industrial-robotics","domestic-substitution"],"name":"Nabtesco","alt":"纳博特斯克","abbr":"","aliases":["Nabtesco Corporation"],"one_liner":"A Japanese precision-reducer giant, the main supplier of RV reducers for industrial robots.","explanation":"Nabtesco was formed in 2003 through the merger of Teijin Seiki and Nabco, headquartered in Tokyo. It is best known for its precision-reducer business: its RV (rotate vector) reducers combine a cycloidal stage with a planetary-gear stage for high stiffness and shock resistance, suited to heavily loaded robot joints, with flagship series including RV-N, RV-C, and RV-E. The company estimates its own global share of precision reducers for the joints of medium and large industrial robots at about 60%. For embodied AI, it represents how core components like reducers have long been dominated by Japanese suppliers, and it is the main benchmark Chinese reducer makers compare themselves against in pursuing domestic substitution.","example":"The RV-N series precision reducer is used in the heavily loaded joints of medium and large industrial robots.","related":["Rotate Vector (RV) Reducer","Cycloidal Reducer","Speed Reducer / Gearbox","Strain Wave Gear (Harmonic Drive)","Big Four of Industrial Robotics","Domestic Substitution"]},{"id":"zhejiang-shuanghuan-driveline","category":"company","sec":8,"tier":3,"sources":[{"title":"关于双环｜浙江双环传动机械股份有限公司","url":"https://www.gearsnet.com/about.html"},{"title":"机器人公司IPO收紧？环动科技终止IPO（新浪财经）","url":"https://finance.sina.com.cn/roll/2026-09-22/doc-inissitm5657890.shtml"}],"as_of":"2026-09","related_ids":["rotate-vector-reducer","speed-reducer-gearbox","nabtesco","domestic-substitution","industrial-robot"],"name":"Zhejiang Shuanghuan Driveline","alt":"双环传动","abbr":"","aliases":[],"one_liner":"A Hangzhou gear-making leader whose subsidiary Huandong Technology is a major Chinese RV-reducer maker.","explanation":"Zhejiang Shuanghuan Driveline (Zhejiang Shuanghuan Driveline Machinery Co., Ltd.) was established in 1980, with its management headquarters in Hangzhou, and listed on the Shenzhen Stock Exchange in September 2010 (ticker 002472); the company describes itself as China's first listed company specializing in gear manufacturing, with customers including ZF, Cummins, and Caterpillar. Its core business is gears for automobiles and new-energy vehicles; in 2020 it established the subsidiary Huandong Technology to make RV reducers — a heavy-duty, high-stiffness precision reducer commonly used at a large industrial robot's joints, such as the waist and upper arm. The RV-reducer market has long been dominated by Japan's Nabtesco, and Huandong Technology is one of the leading Chinese makers pursuing domestic substitution, supplying Chinese industrial robots at scale. Shuanghuan Driveline holds about 61% of Huandong Technology; Huandong Technology's application for a spinoff listing on the STAR Market was accepted in November 2024 but voluntarily withdrawn in September 2026, with its review status now listed as “terminated.”","example":"The RV reducer fitted into the base and upper-arm joints of a Chinese-made 6-axis industrial robot may well come from Huandong Technology.","related":["Rotate Vector (RV) Reducer","Speed Reducer / Gearbox","Nabtesco","Domestic Substitution","Industrial Robot"]},{"id":"ningbo-zhongda-leader-intelligent-transmission","category":"company","sec":8,"tier":3,"sources":[{"title":"中大力德：机器人运动执行部件制造的多面选手（券商研报）","url":"https://pdf.dfcfw.com/pdf/H3_AP202212201581217130_1.pdf"},{"title":"人形机器人一体化关节 - 宁波中大力德智能传动股份有限公司","url":"https://www.zd-motor.com/products_detail/499.html"}],"as_of":"2026-03","related_ids":["planetary-gearbox","rotate-vector-reducer","strain-wave-gear","joint-actuator-module","core-components","domestic-substitution"],"name":"Ningbo ZhongDa Leader Intelligent Transmission","alt":"中大力德","abbr":"","aliases":["ZD Leader","ZhongDa Leader"],"one_liner":"A Ningbo reducer and motor maker that builds planetary, RV, and harmonic reducers alike.","explanation":"ZhongDa Leader (Ningbo ZhongDa Leader Intelligent Transmission Co., Ltd.) was founded in 2006, growing out of the ZhongDa Motor Factory established in 1998, headquartered in Ningbo, Zhejiang, and listed on the Shenzhen Stock Exchange in 2017 (002896). It started out as a geared-motor maker and is one of the few Chinese companies that manufactures precision planetary reducers, RV reducers, and harmonic reducers all at once, together with matching servo drivers and various motors, with products used in industrial robots, smart logistics, and new-energy equipment. With the rise of humanoid robots, it has launched an integrated humanoid joint actuator module that combines the motor, reducer, and driver in one unit, which is why it is commonly listed among humanoid-robot core-component and domestic-substitution concept companies.","example":"An integrated humanoid-robot joint: the motor, reducer, and driver combined into a single module.","related":["Planetary Gearbox","Rotate Vector (RV) Reducer","Strain Wave Gear (Harmonic Drive)","Joint Actuator Module","Core Components","Domestic Substitution"]},{"id":"rollvis","category":"company","sec":8,"tier":3,"sources":[{"title":"Rollvis SA 官网","url":"https://www.rollvis.com/en/"}],"as_of":"2026-09","related_ids":["planetary-roller-screw","inverted-planetary-roller-screw","ball-screw","linear-actuator","tesla-optimus","domestic-substitution"],"name":"Rollvis","alt":"Rollvis（行星滚柱丝杠）","abbr":"","aliases":["Rollvis SA"],"one_liner":"A long-established Swiss maker of planetary roller screws, a key supplier for humanoid robots' linear joints.","explanation":"Rollvis SA was founded in 1970 and is based in Plan-les-Ouates, near Geneva, Switzerland, specializing in planetary roller screws: several threaded rollers replace the steel balls used in a ball screw to turn a motor's rotation into linear push-pull motion, giving more contact lines and therefore higher load capacity, stiffness, and shock resistance than a ball screw. Its product line includes the standard RV/HRV, the inverted RVI, the recirculating RVR, and the differential RVD types, and its traditional customers are in aerospace, medical devices, research, and industry. Humanoid robots commonly use linear actuators — a motor paired with a planetary roller screw — at joints such as the knee, elbow, and ankle; this kind of high-precision screw has long been dominated by a handful of European makers, and Rollvis is one frequently mentioned, as well as a benchmark for domestic substitution in China.","example":"According to supply-chain teardown analyses, Tesla's Optimus uses an inverted planetary roller screw for its linear joints — exactly the type of product Rollvis's RVI line makes.","related":["Planetary Roller Screw","Inverted Planetary Roller Screw","Ball Screw","Linear Actuator (Electric Cylinder)","Tesla Optimus","Domestic Substitution"]},{"id":"shanghai-beite-technology","category":"company","sec":8,"tier":3,"sources":[{"title":"人形机器人丝杠成最贵零部件，绑定核心客户的2家公司（东方财富）","url":"https://caifuhao.eastmoney.com/news/20260530193720434562320"},{"title":"人形机器人催化丝杠国产化（券商研报）","url":"https://pdf.dfcfw.com/pdf/H3_AP202506271698187795_1.pdf?1751008182000.pdf="}],"as_of":"2026-05","related_ids":["planetary-roller-screw","linear-actuator","inverted-planetary-roller-screw","tesla-supply-chain","core-components"],"name":"Shanghai Beite Technology","alt":"北特科技","abbr":"","aliases":["Beite"],"one_liner":"A Shanghai auto-chassis parts maker moving into planetary roller screws for humanoid robots.","explanation":"Shanghai Beite Technology is a Shanghai automotive-parts company listed on the Shanghai Stock Exchange's Main Board (ticker 603009), with core businesses in chassis parts, air-conditioning compressors, and lightweight aluminum components for cars; chassis parts accounted for more than 60% of its roughly RMB 2.3 billion revenue in 2025. Humanoid robots' linear joints commonly use planetary roller screws to convert a motor's rotation into high-force linear motion, a component that demands high machining precision and has long relied on imports. Beite is applying the precision-manufacturing capability it built making auto parts to enter this segment; it reportedly invested about RMB 1.85 billion in 2024 to build screw-manufacturing capacity in Kunshan, Jiangsu, and, according to brokerage research reports, its planetary roller screws have already gone through small-batch trial production and sample shipments with several Chinese humanoid-robot body makers.","example":"A humanoid robot's knee joint commonly uses a linear actuator — a motor plus a planetary roller screw — to drive the lower leg; the screw itself is the part Beite is targeting.","related":["Planetary Roller Screw","Linear Actuator (Electric Cylinder)","Inverted Planetary Roller Screw","Tesla (Optimus) Supply Chain","Core Components"]},{"id":"tuopu-group","category":"company","sec":8,"tier":3,"sources":[{"title":"证券之星：拓普集团港股IPO","url":"https://4g.stockstar.com/detail/IG2026041300021443"},{"title":"铸造头条：拓普集团赴港IPO","url":"https://zhuzaotoutiao.com/xw/html/22845.shtml"}],"as_of":"2026-04","related_ids":["actuator","linear-actuator","planetary-roller-screw","tesla-supply-chain","tier-1-supplier","humanoid-robot-concept-stocks"],"name":"Tuopu Group","alt":"拓普集团","abbr":"","aliases":["Ningbo Tuopu Group"],"one_liner":"A Ningbo auto-parts giant transforming itself into a humanoid-robot actuator supplier.","explanation":"Tuopu Group traces back to 1983, when founder Jianshu Wu started making auto parts in Ningbo; its predecessor, Ningbo Tuopu Brake Systems, was set up in 2004, and the company listed on the Shanghai Stock Exchange in 2015 (ticker 601689). Its core business covers damping systems, interior components, chassis, automotive electronics, and thermal management, and it is a major supplier to new-energy automakers including Tesla. It is applying its automotive supply-chain capability to enter humanoid robotics, starting with linear actuators — components that turn a motor's rotation into linear push-pull motion, commonly used in robot arms and legs — and expanding into rotary actuators, dexterous-hand motors, structural parts, and sensors, with a planned Ningbo robot-core-component base carrying roughly RMB 5 billion in total investment. The market widely views it as an actuator supplier for Tesla's Optimus, though the company has not confirmed this item by item. In December 2025 it announced plans for a Hong Kong listing, filing with the Hong Kong Stock Exchange in April 2026 to pursue a dual A+H listing.","example":"A humanoid robot needs dozens of actuators per unit, and automotive-supply-chain vendors like Tuopu bring car-grade mass-production lines to bear on cutting the cost per unit.","related":["Actuator","Linear Actuator (Electric Cylinder)","Planetary Roller Screw","Tesla (Optimus) Supply Chain","Tier-1 Supplier","Humanoid Robot Concept Stocks"]},{"id":"sanhua-intelligent-controls","category":"company","sec":8,"tier":3,"sources":[{"title":"天册助力三花智控于香港联交所主板成功上市","url":"https://www.tclawfirm.com/content-1655.html"},{"title":"股价因「机器人」飙涨 三花智控2025年净利至多增50%（观点网）","url":"https://www.guandian.cn/m/show/533596"}],"as_of":"2025-10","related_ids":["actuator","joint-actuator-module","tesla-optimus","tesla-supply-chain","core-components","humanoid-robot-concept-stocks"],"name":"Sanhua Intelligent Controls","alt":"三花智控","abbr":"","aliases":["Zhejiang Sanhua Intelligent Controls"],"one_liner":"A leading maker of refrigeration control components that has moved into electromechanical actuators for humanoid robots.","explanation":"Sanhua Intelligent Controls (Zhejiang Sanhua Intelligent Controls Co., Ltd.) was founded in 1994 under the Sanhua Group, listed on the Shenzhen Stock Exchange in 2005, and listed again on the Hong Kong Stock Exchange's Main Board on June 23, 2025, making it a dual A+H listed company. Its core business is refrigeration control components such as valves for air conditioners and refrigerators, plus thermal-management parts for new-energy vehicles. It entered embodied AI by developing electromechanical actuators for humanoid robots — joint-drive units that integrate a motor with a gearbox or screw and sensors. It is widely seen as a representative company in Tesla's Optimus supply chain, and its stock rose sharply in 2025 on humanoid-robot sentiment.","example":"Between September and October 2025, Sanhua's A-shares rose from around RMB 30 to a record high above RMB 53, with the market mainly betting on its robot-actuator business.","related":["Actuator","Joint Actuator Module","Tesla Optimus","Tesla (Optimus) Supply Chain","Core Components","Humanoid Robot Concept Stocks"]},{"id":"schaeffler","category":"company","sec":8,"tier":3,"sources":[{"title":"Schaeffler and Neura Robotics launch future-oriented technology partnership","url":"https://www.schaeffler.com/en/media/press-releases/press-releases-detail.jsp?id=88136515"},{"title":"Schaeffler enters into strategic partnership with Hexagon Robotics","url":"https://markets.businessinsider.com/news/stocks/eqs-news-humanoid-robotics-schaeffler-enters-into-strategic-partnership-with-hexagon-robotics-1036047911"}],"as_of":"2026-04","related_ids":["neura-robotics","actuator","planetary-gearbox","strain-wave-gear","selling-shovels"],"name":"Schaeffler","alt":"舍弗勒","abbr":"","aliases":["Schaeffler Group"],"one_liner":"A German bearings and auto-parts giant now supplying joint actuators for humanoid robots.","explanation":"Schaeffler was founded in 1946 and is headquartered in Herzogenaurach, Germany, known for rolling bearings and automotive drivetrain components under brands such as INA and FAG. In recent years it has extended its precision-mechanics expertise into humanoid robots, launching an actuator platform built on planetary and harmonic (strain-wave) reducers. In November 2025 it formed a technology partnership with Germany's NEURA Robotics to supply actuators for its humanoids, with plans to deploy thousands of humanoid robots across its own global factories by 2035; in April 2026 it separately signed agreements with Switzerland's Hexagon Robotics and Vietnam's VinDynamics to supply reducers and actuators. It represents a broader pattern of traditional automotive-parts suppliers “selling shovels” to the humanoid-robot boom.","example":"The planetary-reduction actuators Schaeffler supplies to NEURA are used at rotating joints such as the shoulder, elbow, knee, and wrist, with a rated torque up to about 250 Nm.","related":["NEURA Robotics","Actuator","Planetary Gearbox","Strain Wave Gear (Harmonic Drive)","Selling Shovels"]},{"id":"maxon","category":"company","sec":8,"tier":3,"sources":[{"title":"Maxon Group（Wikipedia）","url":"https://en.wikipedia.org/wiki/Maxon_Motor"}],"as_of":"2025-06","related_ids":["coreless-motor","brushless-dc-motor","faulhaber","dexterous-hand","speed-reducer-gearbox","hexagon-aeon"],"name":"maxon","alt":"Maxon","abbr":"","aliases":["Maxon Group","maxon motor"],"one_liner":"A Swiss precision-motor maker known for coreless motors, which have driven Mars rovers.","explanation":"Maxon was founded in Sachseln, Switzerland in 1961 as Interelectric AG, and was renamed from maxon motor to Maxon in 2019; it has about 3,000 employees. It makes precision DC motors, brushless motors, gearboxes, encoders, and motor controllers, and is especially known for coreless motors — small motors whose rotor has no iron core, giving low inertia and smooth operation at low speed. NASA's Sojourner, Spirit, Opportunity, and Perseverance Mars rovers all used its motors. In embodied AI, this type of small motor is commonly used in dexterous-hand fingers and small joints; Maxon is also reportedly the actuator partner for Hexagon's humanoid robot, AEON.","example":"Many dexterous hands fit a coreless motor and a miniature lead screw into each finger; maxon and FAULHABER are the most common overseas suppliers of this type of motor.","related":["Coreless Motor","Brushless DC Motor","FAULHABER","Dexterous Hand","Speed Reducer / Gearbox","Hexagon AEON"]},{"id":"faulhaber","category":"company","sec":8,"tier":3,"sources":[{"title":"Wikipedia (de): Faulhaber (Unternehmen)","url":"https://de.wikipedia.org/wiki/Faulhaber_(Unternehmen)"},{"title":"FAULHABER - About us","url":"https://www.faulhaber.com/en/about-us/"}],"as_of":"2025","related_ids":[null,null,null,null,null],"name":"FAULHABER","alt":"冯哈伯","abbr":"","aliases":["Dr. Fritz Faulhaber GmbH & Co. KG"],"one_liner":"A German micro-motor maker, one of the original inventors and top suppliers of coreless motors.","explanation":"FAULHABER is a German family business founded in 1947 by Fritz Faulhaber, headquartered in Schönaich, Baden-Württemberg. Its best-known technology is the skew-wound, ironless rotor coil it patented in 1965 — what is commonly called a coreless motor today, with no iron core in the rotor, giving it small size, low inertia, and smooth rotation. Its products include miniature DC motors, brushless motors, stepper motors, linear motors, and matching precision gearheads, encoders, and drivers, used in medical devices, aerospace, optics, and robotics. Because a dexterous hand's fingers have very little room, they commonly use a coreless motor paired with a miniature lead screw or gearhead to drive each joint, and FAULHABER and Switzerland's maxon are the two leading international suppliers of this type of motor.","example":"Some dexterous hands fit a small-diameter coreless motor into each finger, paired with a gearhead to drive the finger joint.","related":["Coreless Motor","maxon","Dexterous Hand","Micro Lead Screw","Core Components"]},{"id":"moons-electric","category":"company","sec":8,"tier":3,"sources":[{"title":"空心杯电机是机器人灵巧手的核心驱动技术 - 鸣志官网","url":"https://www.moons.com.cn/article/cn-techschool-stepmotor-00106"},{"title":"上海鸣志电器股份有限公司-展商信息-2026世界机器人大会","url":"https://www.worldrobotconference.com/expo/company/520.html"},{"title":"上海鸣志电器股份有限公司2024年年度报告","url":"https://static.cninfo.com.cn/finalpage/2025-04-26/1223327773.PDF"}],"as_of":"2026-08","related_ids":["coreless-motor","stepper-motor","dexterous-hand","core-components","domestic-substitution","tesla-supply-chain"],"name":"MOONS' Electric","alt":"鸣志电器","abbr":"","aliases":["MOONS'"],"one_liner":"A Shanghai control-motor maker whose coreless motors are seen as key dexterous-hand components.","explanation":"MOONS' Electric (Shanghai MOONS' Electric Co., Ltd.) was founded in 1994, headquartered in Shanghai, and listed on the Shanghai Stock Exchange in 2017 (603728). It mainly makes control motors and drive systems, with stepper motors as its traditional strength, used in semiconductor, lithium-battery, and photovoltaic equipment, medical devices, and industrial automation. It entered the embodied-AI conversation mainly through its brushless coreless motors: this type of motor, whose rotor has no iron core, is small and has low inertia, making it well suited to driving the fingers of a dexterous hand, which has gotten it grouped with humanoid-robot supply-chain concept stocks. At the 2026 World Robot Conference, it exhibited a coreless motor paired with a planetary gearbox, along with miniature frameless motors.","example":"The brushless coreless motor ECH11026 weighs about 18.5 grams and can be used to drive a dexterous hand's fingers.","related":["Coreless Motor","Stepper Motor","Dexterous Hand","Core Components","Domestic Substitution","Tesla (Optimus) Supply Chain"]},{"id":"kollmorgen","category":"company","sec":8,"tier":3,"sources":[{"title":"Kollmorgen: History","url":"https://www.kollmorgen.com/en-us/company/history"},{"title":"Regal Rexnord and Altra receipt of regulatory approvals for merger","url":"https://investors.regalrexnord.com/investors/ir-news/press-release-details/2023/REGAL-REXNORD-AND-ALTRA-ANNOUNCE-RECEIPT-OF-ALL-REQUIRED-REGULATORY-APPROVALS-FOR-MERGER/default.aspx"},{"title":"艾邦机器人：无框力矩电机全球23家供应商","url":"https://www.aibangbots.com/a/1764"}],"as_of":"2026","related_ids":["frameless-torque-motor","servo-motor","joint-actuator-module","strain-wave-gear","maxon","core-components"],"name":"Kollmorgen","alt":"科尔摩根","abbr":"","aliases":[],"one_liner":"A long-established US motion-control maker whose frameless torque motors are common in robot joints.","explanation":"Kollmorgen traces back to periscope designs by the German optics expert Friedrich Kollmorgen, and was incorporated in New York in 1916; it is now headquartered in Radford, Virginia, and marked its 110th anniversary in 2026. It has changed hands several times: owned by Danaher, then spun off in 2016 into Fortive (itself a Danaher spinoff), folded into Altra Industrial Motion in 2018, and became a brand under Regal Rexnord after it acquired Altra in 2023. Its main products are servo motors, drivers, and frameless torque motors, and it also makes navigation control systems for AGVs. A frameless motor is sold as just a stator and rotor, which a robot maker mounts directly inside its own joint housing — well suited to the compact joints of collaborative arms and humanoid robots; its TBM2G series is designed specifically for humanoid robot joints.","example":"A humanoid robot's rotary joints commonly use a frameless motor like Kollmorgen's TBM2G, paired with a harmonic reducer and dual encoders to form a joint actuator module.","related":["Frameless Torque Motor","Servo Motor","Joint Actuator Module","Strain Wave Gear (Harmonic Drive)","maxon","Core Components"]},{"id":"zeroerr","category":"company","sec":8,"tier":3,"sources":[{"title":"头部人形机器人关节公司半年再获新融资，同创伟业领投数亿元（36氪）","url":"https://m.36kr.com/p/3885232033378308"},{"title":"零差云控完成数亿元C++轮融资（新浪科技）","url":"https://finance.sina.com.cn/tech/digi/2026-07-13/doc-inihrsvc4207356.shtml"},{"title":"零差云控官网","url":"https://zeroerr.cn"}],"as_of":"2026-07","related_ids":["joint-actuator-module","actuator","collaborative-robot","humanoid-robot","integrated-drive-and-control"],"name":"ZeroErr","alt":"零差云控","abbr":"","aliases":[],"one_liner":"A Shenzhen maker of robot joint modules, focused on standardized, integrated eRob joints.","explanation":"ZeroErr was founded in Shenzhen in 2016 by Xiqing Jia. It builds robot joint modules — integrated units combining a motor, gearbox, encoder, and driver — with its eRob series used in collaborative robots, medical robots, automation equipment, and humanoid robots. Its founder's philosophy is “fake customization, real standardization”: most of what customers call custom requirements really just means lighter, smaller, stronger, or cheaper, needs that can be met by iterating on a standard product. As humanoid robots move toward mass production, joint-supply capacity has become a bottleneck, drawing investor attention to upstream component makers like this one. In July 2026 it closed a Series C++ round of several hundred million RMB, led by Vertex Ventures and joined by Guotai Junan Innovation Investment, with existing investor Huakong Fund adding to its stake; the funds go toward expanding capacity and overseas markets.","example":"A small collaborative-arm team can buy six eRob joint modules and assemble them directly into a 6-axis robot arm, without designing its own motors and gearboxes.","related":["Joint Actuator Module","Actuator","Collaborative Robot","Humanoid Robot","Integrated Drive and Control"]},{"id":"damiao-technology","category":"company","sec":8,"tier":3,"sources":[{"title":"达妙科技官网","url":"https://www.mdmbot.com/"},{"title":"DM-J4310-2EC 产品页","url":"https://www.mdmbot.com/index.php?c=show&id=84"}],"as_of":"2025","related_ids":["damiao-dm-j4310-2ec-joint-motor",null,null,null,null],"name":"DAMIAO Technology","alt":"达妙科技","abbr":"","aliases":["DAMIAO"],"one_liner":"A Shenzhen joint-motor maker whose DM-J4310 is a common small joint in student and open-source robots.","explanation":"Shenzhen DAMIAO Technology Co., Ltd. was founded in 2019 and is based at the Nanshan University Town entrepreneurship park in Shenzhen. It makes integrated joint motors, hub motors, and gimbal motors for robots, along with the DM40-DM100 series of motor drivers and development boards. Its representative product, the DM-J4310-2EC, integrates a brushless motor, gearbox, driver, and dual encoders in one unit, communicates over CAN bus, and supports “MIT mode,” which sends position, velocity, stiffness, damping, and feed-forward torque all at once. Because it is inexpensive and simple to wire up, it is common in robot dogs, robot arms, bio-inspired robots, and robotics competitions, and it is a frequent choice for lighter-load joints in many open-source humanoids and open-source robot arms.","example":"Many open-source dual-arm robots and small humanoids use the DM-J4310 directly for wrist and head joints, controlling it with MIT-mode commands sent over CAN.","related":["DAMIAO DM-J4310-2EC Joint Motor","Joint Actuator Module","MIT Mode (MIT Cheetah-style Joint Motor Command)","Controller Area Network / CAN with Flexible Data-Rate","Open-source Hardware"]},{"id":"robstride","category":"company","sec":8,"tier":3,"sources":[{"title":"RobStride 产品页","url":"https://www.robstride.com/products"}],"as_of":"2026-09","related_ids":["joint-actuator-module","quasi-direct-drive","mit-mode","xiaomi-cybergear-micro-motor","planetary-gearbox","controller-area-network"],"name":"RobStride","alt":"灵足时代","abbr":"","aliases":["RobStride Dynamics"],"one_liner":"A Chinese maker of robot joint actuators whose RS-series integrated joints are common in quadrupeds and small humanoids.","explanation":"RobStride Dynamics is a Chinese maker of robot joint modules, best known for its RS00–RS06 line of integrated joints: a brushless motor, planetary gearbox, driver, and encoder packed into a single disc-shaped module that takes commands over a CAN bus and supports position, velocity, torque, and MIT-mode control, which sends target position, velocity, stiffness, damping, and feedforward torque all at once. These “quasi-direct-drive” joints have a low gear ratio and can be back-driven, which suits reinforcement-learning-based locomotion control that outputs joint targets directly, making them a common power source for quadrupeds, small humanoids, and desktop robot arms. Its core team is reported to come from Xiaomi's CyberGear micro-motor project. No reliable source was found for its founding year or funding, so those are omitted here.","example":"","related":["Joint Actuator Module","Quasi-Direct Drive","MIT Mode","Xiaomi CyberGear Micro-Motor","Planetary Gearbox","Controller Area Network (CAN)"]},{"id":"robotis","category":"company","sec":8,"tier":3,"sources":[{"title":"로보티즈 - 위키백과","url":"https://ko.wikipedia.org/wiki/로보티즈"}],"as_of":"2025-08","related_ids":["robotis-dynamixel-servo","servo","turtlebot","koch-v1-1-arm","trossen-robotics-widowx-250","robot-operating-system-2"],"name":"ROBOTIS","alt":"ROBOTIS","abbr":"","aliases":[],"one_liner":"A South Korean robotics company behind DYNAMIXEL servos and the TurtleBot3, a staple in research and education.","explanation":"ROBOTIS was founded in March 1999 by Byung-soo Kim, is headquartered in Seoul, South Korea, and listed on KOSDAQ in October 2018. Its best-known product is the DYNAMIXEL smart servo: a motor, gearbox, driver, and encoder integrated into a single module, with multiple units daisy-chained on one bus and controlled together, which is why low-cost arms and teleoperation leader arms such as Koch, GELLO, and WidowX all use it. The company also co-developed the TurtleBot3 teaching platform with Open Robotics, and makes the OpenMANIPULATOR robot arm, the GAEMI delivery robot, and the AI Worker humanoid, among other products. According to the Korean-language Wikipedia, the company spun off its autonomous-delivery-robot business in June 2025 and announced a capital increase of about 100 billion won that August.","example":"The Koch v1.1 arm featured in LeRobot's early tutorials is built out of Dynamixel XL430 and XL330 servos.","related":["ROBOTIS Dynamixel Servo","Servo (Smart Serial Bus Servo)","TurtleBot","Koch v1.1 Arm","Trossen Robotics WidowX 250","Robot Operating System 2"]},{"id":"seer-robotics","category":"company","sec":8,"tier":3,"sources":[{"title":"「机器人大脑」第一股诞生！仙工智能正式登陆港交所","url":"https://seer-robotics.ai/zh/media/333"},{"title":"仙工智能 06106 招股資訊（經濟通）","url":"https://www.etnet.com.hk/www/tc/stocks/ipo-info.php?code=06106"}],"as_of":"2026-06","related_ids":["autonomous-mobile-robot","robot-controller","mobile-manipulator","hkex-chapter-18c","automated-guided-vehicle"],"name":"SEER Robotics","alt":"仙工智能","abbr":"","aliases":["SEER"],"one_liner":"A Shanghai maker of mobile-robot controllers, the world's top seller by volume, listed in Hong Kong in 2026.","explanation":"SEER Robotics is headquartered in Shanghai; founder and CEO is Yue Zhao. Its core product is the robot controller — the “brain” of a mobile robot, integrating localization and navigation, motion control, and fleet-scheduling software — sold to integrators and body makers for use in autonomous mobile robots (AMRs), unmanned forklifts, and mobile manipulators; the company also builds complete robots and embodied-AI systems. According to its IPO prospectus, by 2025 shipment volume its robot controllers held roughly a 24.8% global and 45.2% China market share, both ranked first. It listed under Chapter 18C of the Hong Kong Stock Exchange's rules (for specialist technology companies) on the Main Board on June 24, 2026 (ticker 06106), earning the nickname “the first robot-brain stock.”","example":"A factory AMR can use a SEER controller to handle mapping, localization, and path planning, so the body maker only needs to build the chassis hardware.","related":["Autonomous Mobile Robot","Robot Controller","Mobile Manipulator","HKEX Chapter 18C","Automated Guided Vehicle"]},{"id":"realsense","category":"company","sec":9,"tier":2,"sources":[{"title":"Intel let RealSense go. Now Cognex is paying $600 million to buy it (Calcalist, 2026-09)","url":"https://www.calcalistech.com/ctechnews/article/s1tqir19fl"},{"title":"Cognex to Acquire RealSense, Expanding Machine Vision Leadership into High-Growth Robotic Perception Market (Cognex, 2026-09)","url":"https://investor.cognex.com/news/news-details/2026/Cognex-to-Acquire-RealSense-Expanding-Machine-Vision-Leadership-into-High-Growth-Robotic-Perception-Market/default.aspx"},{"title":"RealSense Completes Spin Out from Intel, Raises $50 Million (Intel Capital, 2025-07)","url":"https://www.intelcapital.com/realsense-completes-spin-out-from-intel-raises-50-million-to-accelerate-ai-powered-vision-for-robotics-and-biometrics/"}],"as_of":"2026-09","related_ids":["realsense-depth-camera","intel-realsense-sdk-2-0","depth-camera","active-stereo","wrist-camera","orbbec"],"name":"RealSense","alt":"RealSense","abbr":"","aliases":["Intel RealSense"],"one_liner":"A depth-camera company spun out of Intel; its D435 and D405 are staples in robotics labs.","explanation":"RealSense grew out of a 3D-camera project Intel launched in 2014, and the D400-series stereo depth cameras it released in 2018 — the D415 and D435, later joined by the D435i, D405, D455, and others — became the most common RGB-D camera, outputting both a color image and a depth map, in robotics labs. On July 11, 2025, it spun off from Intel as an independent company, led by former Intel executive Nadav Orbach as CEO, closing a $50 million Series A with Intel Capital and MediaTek's venture arm participating; it is headquartered in California, though most employees are reportedly in Haifa, Israel. The company says 60% of the world's autonomous mobile robots and humanoid robots use its depth cameras. On September 22, 2026, machine-vision company Cognex announced a $500 million cash acquisition, plus about $100 million for employee retention and stock incentives, expected to close in the fourth quarter.","example":"In tabletop-manipulation experiments, it's common to mount one D435 at the side of the table as a third-person view and one D405 next to the gripper as a wrist camera, feeding both RGB-D streams into the policy.","related":["RealSense Depth Camera (D435i / D405)","Intel RealSense SDK 2.0 (librealsense)","Depth Camera","Active Stereo","Wrist Camera","Orbbec"]},{"id":"orbbec","category":"company","sec":9,"tier":2,"sources":[{"title":"Orbbec - Wikipedia","url":"https://en.wikipedia.org/wiki/Orbbec"},{"title":"Orbbec unveils Gemini 330 series","url":"https://www.orbbec.com/news/orbbec-unveils-gemini-330-series-of-stereo-vision-3d-cameras-powered-by-latest-asic-for-outdoor-and-indoor-performance/"}],"as_of":"2025-11","related_ids":["orbbec-gemini-330-series","orbbec-femto-bolt",null,null,null,"realsense-depth-camera"],"name":"Orbbec","alt":"奥比中光","abbr":"","aliases":[],"one_liner":"A Shenzhen maker of 3D vision sensors, and a common supplier of the depth cameras used on robots.","explanation":"Orbbec was founded in Shenzhen in 2013 by Howard Huang, who came from a postdoctoral research background; it also has a team in Michigan, and it trades on the Shanghai Stock Exchange's STAR Market (ticker 688322). It builds a full range of 3D vision sensors, covering structured light, indirect time-of-flight (iToF), stereo, and lidar, and it designs its own depth-computation chips. In robotics, its most widely used products are the Gemini 330 series of stereo depth cameras and the Femto Bolt, made in partnership with Microsoft to be compatible with Azure Kinect; the Gemini 330 series has been integrated into NVIDIA Isaac Perceptor. According to Wikipedia, its vision systems were used on the Tiangong Ultra robot that won the 100-meter race at the 2025 World Humanoid Robot Games, and the company was profitable through the first three quarters of that year.","example":"Labs often mount a Femto Bolt at the edge of a table as a third-person camera to capture point clouds for training manipulation policies.","related":["Orbbec Gemini 330 Series","Orbbec Femto Bolt","Depth Camera (RGB-D Camera)","Structured Light","Time of Flight","RealSense Depth Camera (D435i / D405)"]},{"id":"paxini-tech","category":"company","sec":9,"tier":2,"sources":[{"title":"帕西尼获超10亿元B轮融资（证券时报）","url":"https://www.stcn.com/article/detail/3670368.html"},{"title":"独家对话帕西尼许晋诚（钛媒体）","url":"https://www.tmtpost.com/7966743.html"}],"as_of":"2026-03","related_ids":["tactile-sensor","paxini-tora-one","paxini-px-6ax","vision-tactile-language-action-model","tactile-data","dexterous-hand"],"name":"PaXini Tech","alt":"帕西尼感知","abbr":"","aliases":["PaXini"],"one_liner":"A Shenzhen tactile-sensor company that has expanded into dexterous hands, humanoids, and embodied data.","explanation":"PaXini Tech was founded in 2021 and is headquartered in Shenzhen. Founder and CEO Jincheng Xu comes out of Professor Shigeki Sugano's lab at Waseda University in Japan and has long worked on humanoid robots and tactile sensors. Its core product is a multi-dimensional tactile sensor — the PX-6AX series, which measures pressure and shear force at once — on top of which it has built the dexterous hand DexH13 and the humanoid robot TORA-ONE, set up data-collection factories that produce real-robot data with touch included, and developed a vision-tactile-language-action model called OmniVTLA and a world model called HyperCosmos. Fundraising has been intense since 2025, with investors including JD.com and BYD; in March 2026 it closed a Series B of more than RMB 1 billion at a valuation above RMB 10 billion.","example":"The hands on the TORA-ONE humanoid robot use PaXini's own multi-dimensional tactile sensors.","related":["Tactile Sensor","PaXini TORA-ONE","PaXini PX-6AX","Vision-Tactile-Language-Action Model","Tactile Data","Dexterous Hand"]},{"id":"tashan-technology","category":"company","sec":9,"tier":3,"sources":[{"title":"他山科技完成数亿元B轮融资，全链路布局触觉感知产业（新浪财经，2026-07-10）","url":"https://finance.sina.com.cn/wm/2026-07-10/doc-inihispu3758074.shtml"},{"title":"清华、北航校友造触觉，横扫中国机器人市场半壁江山（智东西）","url":"https://m.zhidx.com/p/482791.html"},{"title":"他山科技一个季度内连续完成两轮融资（猎云网 / 东方财富）","url":"https://caifuhao.eastmoney.com/news/20251128170229122318000"}],"as_of":"2026-07","related_ids":["tactile-sensor","capacitive-tactile-sensing","electronic-skin","dexterous-hand","tactile-data","robomind"],"name":"Tashan Technology","alt":"他山科技","abbr":"","aliases":["TASHAN"],"one_liner":"A Beijing maker of tactile-sensing chips and sensors, a leading Chinese supplier of touch for humanoid robots.","explanation":"Tashan Technology, formally Beijing Tashan Technology Co., Ltd., was founded in 2017 by CEO Yang Ma, Tengchen Sun, and CTO Wuqiang Yang — a Tsinghua PhD in precision instruments and a professor at the University of Manchester — with a team drawing on Tsinghua and Beihang backgrounds. The company started with tactile-sensing chips, packing capacitive touch sensing small enough to fit in a fingertip into a dedicated chip (a “touch MCU”), then building up tactile sensors, tactile data, and algorithms on top to give dexterous hands and humanoid robots “fingertip touch.” It reportedly holds more than 80% of the Chinese market for humanoid-robot tactile sensors. On funding, it announced combined Series A3 and A4 rounds of several hundred million RMB in November 2025, and a further Series B of several hundred million RMB in July 2026, with investors including Joyson Electronics and Taiping Innovation; the company says its order volume in the first half of 2026 already exceeded four times the whole of the previous year.","example":"In the tactile-augmented clips of the RoboMIND 2.0 dataset, the normal and shear forces were measured with Tashan's tactile sensors.","related":["Tactile Sensor","Capacitive Tactile Sensing","Electronic Skin","Dexterous Hand","Tactile Data","RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation)"]},{"id":"daimon-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"戴盟机器人 关于我们","url":"https://www.dmrobot.com/about.html"},{"title":"戴盟机器人官网","url":"https://www.dmrobot.com/"}],"as_of":"2026-06","related_ids":["daimon-dm-tac-visuotactile-sensor","daimon-infinity",null,null,"gelsight"],"name":"Daimon Robotics","alt":"戴盟机器人","abbr":"","aliases":["Daimon"],"one_liner":"A Shenzhen company, incubated at HKUST, making visuotactile sensors and touch-inclusive embodied datasets.","explanation":"Daimon Robotics was incubated at the Hong Kong University of Science and Technology; its core team began researching visuotactile sensing in 2017 and began formal operations in 2023, headquartered in Bao'an, Shenzhen, with an R&D center in Hong Kong; reported co-founders include HKUST professor Yu Wang. A visuotactile sensor places a small camera behind an elastic surface and photographs how that surface deforms on contact, from which it infers contact shape and force. Its products are the DM-Tac series of sensors (the general-purpose W2, the pointed X, the fingertip F, and the gripper-mounted G), whose first generation launched in April 2025, plus the DM-Flux edge-compute platform and a teleoperation data-collection system. In April 2026, together with several institutions, it released Daimon-Infinity, an all-modality dataset that includes touch, and in June it released a unified evaluation framework.","example":"The DM-Tac W2 has a resolution of 384×288 and a 120 Hz sampling rate, and can output contact geometry, 3D force distribution, and six-axis resultant force.","related":["Daimon DM-Tac Visuotactile Sensor","Daimon-Infinity","Vision-Based Tactile Sensor","Tactile Data","GelSight"]},{"id":"weitai-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"新京报：纬钛机器人完成近亿元天使轮及天使+轮融资","url":"https://m.bjnews.com.cn/detail/1744725950168547.html"},{"title":"观察者网：小米领投的纬钛机器人完成亿元融资","url":"https://www.guancha.cn/economy/2025_04_16_772456.shtml"}],"as_of":"2025-04","related_ids":["vision-based-tactile-sensor","gelsight","tactile-sensor","dexterous-hand","contact-rich-manipulation","robotic-assembly"],"name":"Weitai Robotics","alt":"纬钛科技","abbr":"","aliases":[],"one_liner":"A Shanghai vision-based tactile-sensor company founded by a co-developer of the GelSight fingertip sensor.","explanation":"Usually known as Weitai Robotics, the company was founded in Shanghai in January 2024. Founder Rui Li studied for his PhD at MIT under Edward Adelson, the inventor of GelSight, and helped develop the GelSight fingertip sensor; its core team comes out of MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). Vision-based tactile sensing works by placing a small camera beneath a soft rubber surface and photographing how that surface deforms on contact, computing contact shape and force from the image. Weitai builds vision-based tactile sensors, bionic fingertips, and dexterous hands around this approach, with a focus on tasks such as precision assembly that need coordinated hand-eye feedback. In April 2025 it announced combined angel and angel+ rounds of nearly RMB 100 million, led by Xiaomi's strategic investment arm.","example":"When a robot arm plugs in a connector, the vision-based tactile sensor on its fingertip can tell whether the plug is misaligned and adjust its position accordingly.","related":["Vision-Based Tactile Sensor","GelSight","Tactile Sensor","Dexterous Hand","Contact-rich Manipulation","Robotic Assembly"]},{"id":"xense-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"千觉机器人宣布连续完成两轮数亿元战略融资（新浪科技）","url":"https://finance.sina.com.cn/tech/roll/2026-09-20/doc-inismrfk6918460.shtml"},{"title":"千觉机器人连续完成两轮数亿元战略融资（腾讯新闻）","url":"https://news.qq.com/rain/a/20260921A08QNP00"},{"title":"千觉机器人官网：完成天使轮融资","url":"https://xenserobotics.com/article/387/detail/6"}],"as_of":"2026-09","related_ids":["vision-based-tactile-sensor","tactile-sensor","tactile-simulation","tactile-data","dexterous-manipulation"],"name":"Xense Robotics","alt":"千觉机器人","abbr":"","aliases":["Xense"],"one_liner":"A Shanghai embodied-touch company building vision-based tactile sensors, tactile data, and tactile models.","explanation":"Xense Robotics was founded in Shanghai in May 2024; founder Daolin Ma is a professor at Shanghai Jiao Tong University and an ICRA 2021 Best Paper winner who earned his PhD at MIT. The company's main product is vision-based tactile sensing — using a camera to capture how a soft rubber surface deforms in order to measure contact force and texture — offered as fingertip and gripper sensors and wearable tactile-data-collection gloves, alongside the Xense Sim tactile simulator, the TacVerse dataset, and the X-TouchMind V1 tactile model, aiming to fill the sense of touch that fine robot manipulation still lacks. In October 2025 it closed a Pre-A round in the tens of millions of RMB led by Futeng Capital with Li Auto participating; in July 2026 it raised a further round in the tens of millions, and that September it announced two consecutive strategic rounds of several hundred million RMB, with investors including Lanchi Ventures and a fund under CICC Capital.","example":"Mounting a fingertip-type vision-based tactile sensor on a gripper lets it judge from the tactile image whether an egg is slipping while grasping it, and adjust its grip force accordingly.","related":["Vision-Based Tactile Sensor","Tactile Sensor","Tactile Simulation","Tactile Data","Dexterous Manipulation"]},{"id":"xela-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"XELA Robotics: About","url":"https://xelarobotics.com/about"},{"title":"TOKYO UPDATES: Giving Robots a Human Sense of Touch","url":"https://www.tokyoupdates.metro.tokyo.lg.jp/en/post-1786"}],"as_of":"2026-09","related_ids":["uskin","magnetic-tactile-sensing","tactile-sensor","electronic-skin","slip-detection","allegro-hand"],"name":"XELA Robotics","alt":"XELA Robotics","abbr":"","aliases":["XELA"],"one_liner":"A Japanese tactile-sensor company whose uSkin product measures both pressure and shear force.","explanation":"XELA Robotics was founded in Tokyo in August 2018 as a spinout from Waseda University; CEO Alexander Schmitz has researched tactile sensors for more than a decade. Its flagship product, uSkin, is a thin, sheet-like magnetic tactile sensor: each sensing unit contains a small magnet and a magnetic-field sensor, and when force is applied the magnet shifts position, letting the unit measure both normal (pressing) force and shear (sideways) force at once, for a full three-axis reading. It can be attached to a dexterous hand's fingertips, palm, or a gripper's surface to give a robot contact and slip information, with companion software called uAi to process the tactile data. In research it is commonly seen retrofitted onto Allegro Hands, used for tasks such as grasping fragile objects and detecting slip.","example":"Covering an Allegro Hand's fingertips and palm with uSkin lets the robot feel when an egg is about to slip and add just enough grip force in response.","related":["uSkin","Magnetic Tactile Sensing","Tactile Sensor","Electronic Skin","Slip Detection","Allegro Hand"]},{"id":"ati-industrial-automation","category":"company","sec":9,"tier":3,"sources":[{"title":"Novanta acquiring ATI Industrial Automation for $172M (The Robot Report)","url":"https://www.therobotreport.com/novanta-acquiring-ati-industrial-automation-for-172m"},{"title":"Novanta Announces Agreement to Acquire ATI","url":"https://investors.novanta.com/news/news-details/2021/Novanta-Announces-Agreement-to-Acquire-ATI/default.aspx"}],"as_of":"2021-12","related_ids":["six-axis-force-torque-sensor","tool-changer","force-control","contact-rich-manipulation","end-effector"],"name":"ATI Industrial Automation","alt":"ATI 工业自动化","abbr":"ATI","aliases":["ATI"],"one_liner":"A long-established US maker of six-axis force sensors and tool-changers for robot arms.","explanation":"ATI Industrial Automation was founded in 1989 and is headquartered in Apex, North Carolina. It makes components mounted at the end of a robot arm: six-axis force/torque sensors (measuring force and torque along three axes each at once), tool changers, and crash protectors, serving both industrial and surgical robots. Its Nano and Mini series of six-axis force sensors are common in research, and many papers on contact-rich manipulation, force control, and assembly use them to measure the forces at the end effector. In 2021, the US precision-technology company Novanta announced it was acquiring ATI for about $172 million.","example":"In a peg-in-hole assembly experiment, an ATI six-axis force sensor is mounted between the arm's flange and the gripper to read contact forces in real time during insertion.","related":["Six-Axis Force/Torque Sensor","Tool Changer","Force Control","Contact-rich Manipulation","End Effector"]},{"id":"sunrise-instruments","category":"company","sec":9,"tier":3,"sources":[{"title":"关于我们 - 宇立仪器","url":"https://www.srisensor.com.cn/about.html"},{"title":"宇立仪器官网首页（产品）","url":"https://www.srisensor.com.cn"}],"as_of":"2026-08","related_ids":["six-axis-force-torque-sensor","joint-torque-sensor","force-control","kunwei-technology","ati-industrial-automation","domestic-substitution"],"name":"Sunrise Instruments","alt":"宇立仪器","abbr":"SRI","aliases":["SRI"],"one_liner":"A long-established Chinese six-axis force-sensor maker that also builds force-controlled polishing and automotive-test equipment.","explanation":"Sunrise Instruments (SRI) was founded in 2007, with bases in Nanning and Shanghai and additional offices in Michigan (US), Taiwan, and South Korea. Founder Dr. York Huang graduated from Wayne State University and previously worked at Ford and the crash-test-dummy maker Humanetics. The company builds three product lines around “force measurement and force control”: multi-axis force sensors (six-axis force/torque sensors, joint torque sensors, and others), the iGrinder line of intelligent floating force-controlled polishing equipment, and automotive crash-test sensors. A six-axis force/torque sensor, mounted at a robot arm's wrist or a humanoid's wrist or ankle, tells the controller how much force and torque the end effector is experiencing, forming the basis for force control, assembly, and polishing. As humanoid robots have taken off, the company has released ultra-thin models, and industry reports commonly list it as one of the leading Chinese makers of six-axis force sensors.","example":"Sunrise's M35XX-series six-axis force sensor is only 9.2 millimeters thick, which the company says makes it suitable for fitting inside a humanoid robot's wrist.","related":["Six-Axis Force/Torque Sensor","Joint Torque Sensor","Force Control","Kunwei Technology","ATI Industrial Automation","Domestic Substitution"]},{"id":"kunwei-technology","category":"company","sec":9,"tier":3,"sources":[{"title":"人民网：坤维科技B++超亿元轮融资落地","url":"http://finance.people.com.cn/n1/2026/0609/c1004-40736552.html"},{"title":"新浪财经：小米、高瓴联手，坤维科技完成B轮融资","url":"https://finance.sina.com.cn/jjxw/2025-02-10/doc-ineizeny3466088.shtml"},{"title":"坤维科技官网","url":"https://www.kunweitech.com"}],"as_of":"2026-06","related_ids":["six-axis-force-torque-sensor","joint-torque-sensor","keli-sensing-technology","sunrise-instruments","ati-industrial-automation","domestic-substitution"],"name":"Kunwei Technology","alt":"坤维科技","abbr":"","aliases":["Changzhou Kunwei Sensing Technology"],"one_liner":"A Changzhou six-axis force sensor maker whose core team came from aerospace research institutes.","explanation":"Kunwei Technology, formally Changzhou Kunwei Sensing Technology Co., Ltd., was founded in 2018 by Lin Xiong, with a core team drawn from aerospace research institutes. It specializes in six-axis force sensors for robots, alongside joint torque sensors, single-axis force sensors, and data-acquisition modules, and was a lead drafter of the Chinese national standard GB/T 43199-2023 for testing multi-axis force/torque sensors on robots. For humanoid robots it has launched the HRS series, built for tight, high-impact locations such as the wrist and ankle. In February 2025 it closed a Series B with participation from Xiaomi and Hillhouse, and in June 2026 a Series B++ round of more than RMB 100 million led by Huatai Zijin and Jinrong Street Capital. Citing data from the market researcher MIR, the company says it holds 53% of the six-axis force sensor market for humanoid and collaborative robots.","example":"A six-axis force sensor in a humanoid robot's ankle can measure ground reaction force and center of pressure, helping with balance control.","related":["Six-Axis Force/Torque Sensor","Joint Torque Sensor","Keli Sensing Technology","Sunrise Instruments","ATI Industrial Automation","Domestic Substitution"]},{"id":"keli-sensing-technology","category":"company","sec":9,"tier":3,"sources":[{"title":"2026世界机器人大会展商信息：宁波柯力传感","url":"https://www.worldrobotconference.com/expo/company/528.html"},{"title":"中国仪器仪表行业协会：宁波柯力传感首次公开发行A股股票上市","url":"http://www.cima.org.cn/nnews.asp?vid=24912"}],"as_of":"2026-08","related_ids":["six-axis-force-torque-sensor","strain-gauge","kunwei-technology","sunrise-instruments","core-components","humanoid-robot-concept-stocks"],"name":"Keli Sensing Technology","alt":"柯力传感","abbr":"","aliases":["Ningbo Keli Sensing Technology"],"one_liner":"A Ningbo leader in weighing sensors now moving into six-axis force sensors for robots.","explanation":"Ningbo Keli Sensing Technology Co., Ltd. was founded in 1995; its chairman is Jiandong Ke, and it listed on the Shanghai Stock Exchange main board in August 2019 (603662). It has long made strain-gauge force sensors and instruments for weighing, and describes itself as one of the world's largest makers of steel load cells, while also building industrial-IoT systems. With the rise of humanoid robots, the company has made robot sensors a strategic priority, centered on six-axis force sensors while also building out tactile sensing and IMUs; its chairman says it has mastered key techniques including structural and algorithmic decoupling. It is one of the stocks commonly discussed as a “humanoid robot concept stock” on the secondary market, and it exhibited at the 2026 World Robot Conference.","example":"Mounting a six-axis force sensor on a robot arm's wrist lets it sense the forces at its end effector in real time during polishing or peg insertion.","related":["Six-Axis Force/Torque Sensor","Strain Gauge","Kunwei Technology","Sunrise Instruments","Core Components","Humanoid Robot Concept Stocks"]},{"id":"stereolabs","category":"company","sec":9,"tier":3,"sources":[{"title":"Ouster Acquires StereoLabs（Ouster 投资者新闻，2026-02-09）","url":"https://investors.ouster.com/news-releases/news-release-details/ouster-acquires-stereolabs-creating-world-leading-physical-ai"},{"title":"Lidar-maker Ouster buys vision company StereoLabs（TechCrunch，2026-02-09）","url":"https://techcrunch.com/2026/02/09/lidar-maker-ouster-buys-vision-company-stereolabs-as-sensor-consolidation-continues"},{"title":"About Us | Stereolabs","url":"https://www.stereolabs.com/about"}],"as_of":"2026-02","related_ids":["stereolabs-zed","stereo-camera","depth-camera","stereo-matching","droid","open-television"],"name":"Stereolabs","alt":"Stereolabs","abbr":"","aliases":["ZED"],"one_liner":"A French stereo-camera company behind the ZED line, acquired by Ouster in 2026.","explanation":"Stereolabs is a French 3D-vision company founded in 2010 by co-founders Cecile Schmollgruber, Edwin Azzam, and Olivier Braun. Its main product is the ZED line of stereo cameras: two lenses shoot the same scene, and depth is computed through stereo matching rather than by actively projecting light, so it also works outdoors in sunlight; the companion ZED SDK provides depth maps, point clouds, localization, and object detection, with a ROS 2 driver available. The company states it has shipped more than 90,000 ZED cameras to over 10,000 customers. In embodied AI, ZED is commonly used as a head-mounted or third-person camera, and also for streaming stereo video back during teleoperation. In February 2026, LiDAR maker Ouster acquired Stereolabs for $35 million in cash plus 1.8 million shares, making it a wholly owned subsidiary while the original founders stayed on to lead the team.","example":"Every capture rig in the DROID dataset pairs two ZED 2 cameras as third-person views, with a ZED Mini mounted on the wrist.","related":["Stereolabs ZED","Stereo Camera","Depth Camera","Stereo Matching","DROID (Distributed Robot Interaction Dataset)","Open-TeleVision"]},{"id":"percipio","category":"company","sec":9,"tier":3,"sources":[{"title":"关于图漾 - 图漾科技官网","url":"http://www.percipio.xyz/about/about-percipio"},{"title":"图漾科技 项目信息 - 36氪","url":"https://pitchhub.36kr.com/project/1678358612997129"}],"as_of":"2023-07","related_ids":["depth-camera","structured-light","time-of-flight","orbbec","mech-mind-robotics","palletizing-depalletizing"],"name":"Percipio","alt":"图漾科技","abbr":"","aliases":["Percipio.XYZ","Shanghai Percipio Technology"],"one_liner":"A Shanghai maker of 3D industrial cameras that supplies depth cameras to robots and logistics equipment.","explanation":"Percipio (Shanghai Percipio Technology Co., Ltd.) was founded in June 2015 and is headquartered in Shanghai; founder Zheping Fei graduated from the electronic engineering department at Fudan University. The company makes 3D industrial cameras — cameras that output a depth map — and supporting software, with product lines covering time-of-flight ranging, speckle structured light, and fringe structured light, used in industrial automation, logistics, and mobile robots. It follows a components-only model, selling cameras rather than complete robots. On funding, its 2021 Series B+ round included Yunfeng Capital, followed by a Series C and C+ in 2023. Within embodied AI it is an upstream perception-component supplier, in the same space as Orbbec and Mech-Mind Robotics.","example":"A depalletizing robot in a warehouse uses a Percipio structured-light camera to capture a depth image of a cardboard box before computing where to grasp it.","related":["Depth Camera","Structured Light","Time of Flight","Orbbec","Mech-Mind Robotics","Palletizing / Depalletizing"]},{"id":"mech-mind-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"Mech-Mind Robotics（Wikipedia）","url":"https://en.wikipedia.org/wiki/Mech-Mind_Robotics"},{"title":"梅卡曼德機器人 09615 招股資訊（經濟通）","url":"https://www.etnet.com.hk/www/tc/stocks/ipo-info.php?code=09615"},{"title":"機器人腦企業梅卡曼德登港股（BigGo 財經）","url":"https://finance.biggo.com.tw/news/2f65c7bc-350e-4140-934a-affd91f87866"}],"as_of":"2026-09","related_ids":["3d-vision-guided-robotics","bin-picking","palletizing-depalletizing","structured-light","hkex-chapter-18c","machine-vision"],"name":"Mech-Mind Robotics","alt":"梅卡曼德","abbr":"","aliases":["Mech-Mind"],"one_liner":"A Beijing industrial 3D-vision and robot-intelligence company, known for its Mech-Eye camera.","explanation":"Mech-Mind was founded in Beijing in 2016; its founder and CEO is Tianlan Shao. It doesn't build robot bodies; instead it supplies the “eyes, brain, and hand” components for robot arms: the Mech-Eye industrial 3D camera handles perception, software including Mech-Vision and Mech-Viz plus the multimodal large model Mech-GPT handle recognition and planning, and it also makes the Mech-Hand dexterous hand. Its typical applications are 3D vision-guided bin picking, mixed-case depalletizing, and machine tending, with customers mostly in automotive, manufacturing, and logistics; its products sell into about 50 countries and regions. Its investors include IDG, Qiming Venture Partners, Intel Capital, and Meituan; it listed on the Hong Kong Stock Exchange on September 1, 2026, at an IPO price of HK$101.7.","example":"A logistics center uses a Mech-Eye camera to photograph a pallet of mixed boxes; the software identifies each box's position and guides a robot arm with a suction gripper to depalletize them one by one.","related":["3D Vision-Guided Robotics","Bin Picking","Palletizing / Depalletizing","Structured Light","HKEX Chapter 18C","Machine Vision"]},{"id":"lyte","category":"company","sec":9,"tier":3,"sources":[{"title":"Lyte raises $165M to help robots better sense their surroundings（The Robot Report）","url":"https://www.therobotreport.com/lyte-raises-165m-help-robots-better-sense-their-surroundings"},{"title":"Lyte Emerges from Stealth with $107M（The AI Insider, 2026-01-06）","url":"https://theaiinsider.tech/2026/01/06/lyte-emerges-from-stealth-with-107m-to-build-the-perception-foundation-for-physical-ai"},{"title":"Lyte, founded by Apple's Face ID engineers, raises $165M at $1.6B（Tech Funding News）","url":"https://techfundingnews.com/lyte-founded-by-apples-face-id-engineers-raises-165m-at-1-6b-to-build-robot-perception"}],"as_of":"2026-09","related_ids":["multi-sensor-fusion","depth-camera","lidar","exteroception","physical-ai"],"name":"Lyte","alt":"Lyte","abbr":"","aliases":["Lyte AI"],"one_liner":"A US robot-perception company founded by veterans of Apple's Face ID team.","explanation":"Lyte was founded in Silicon Valley, California in 2021; founders Alexander Shpunt, Arman Hajati, and Yuval Gerson had worked on Apple's Face ID, and the team also has roots in PrimeSense, the company behind Microsoft's Kinect depth camera. It builds a perception foundation for robots: combining its own chips, multimodal sensors, and spatial-understanding software into a single platform, LyteVision, so robot makers don't have to piece together cameras, radar, and algorithms themselves. It came out of stealth in January 2026 with $107 million raised across Series A and B, winning a CES 2026 Best Innovation award in robotics; in September it closed a $165 million Series C led by Maverick Silicon at a $1.6 billion post-money valuation, and it is already shipping to customers in inspection, logistics, and manufacturing.","example":"An inspection-robot company installs a LyteVision module directly and gets fused depth and obstacle information for navigation, skipping the work of calibrating and fusing multiple sensors itself.","related":["Multi-Sensor Fusion","Depth Camera","LiDAR","Exteroception","Physical AI"]},{"id":"hesai-technology","category":"company","sec":9,"tier":3,"sources":[{"title":"Hesai Group - Wikipedia","url":"https://en.wikipedia.org/wiki/Hesai_Group"},{"title":"禾赛 JT128 产品页","url":"https://www.hesaitech.com/product/jt128/"}],"as_of":"2025-11","related_ids":["hesai-jt128","lidar","robosense","livox","point-cloud","autonomous-driving"],"name":"Hesai Technology","alt":"禾赛科技","abbr":"","aliases":["Hesai","Hesai Group"],"one_liner":"A Shanghai lidar company whose sensors are used in both cars and robots.","explanation":"Hesai Technology was founded in Shanghai in October 2014 by Yifan Li (CEO), Kai Sun, and Shaoqing Xiang, building lidar — sensors that measure distance with lasers and output 3D point clouds. Its main customers are automakers and self-driving companies; Wikipedia cites figures putting its share of the automotive lidar market at around 30%. It listed on the Nasdaq in February 2023 and on the Hong Kong Stock Exchange in September 2025. On robotics, it has launched small lidars with a very wide vertical field of view, such as the JT series, used on embodied-AI robots, delivery and cleaning robots, and AGVs/AMRs (automated guided vehicles and autonomous mobile robots) for mapping, localization, and obstacle avoidance. It reportedly began construction of its first overseas factory in Bangkok, Thailand, in November 2025, and plans to raise annual capacity from 2 million to 4 million units in 2026.","example":"The JT128 has a 189° vertical field of view, so when mounted on a robot it can see both the floor at its feet and the space overhead at once — a much smaller blind spot than a traditional spinning lidar.","related":["Hesai JT128","LiDAR","RoboSense","Livox","Point Cloud","Autonomous Driving"]},{"id":"robosense","category":"company","sec":9,"tier":3,"sources":[{"title":"RoboSense - Wikipedia","url":"https://en.wikipedia.org/wiki/RoboSense"}],"as_of":"2026-09","related_ids":["lidar","robosense-airy","robosense-ac1","solid-state-lidar","hesai-technology","autonomous-driving"],"name":"RoboSense","alt":"速腾聚创","abbr":"","aliases":["RoboSense Technology"],"one_liner":"A Shenzhen LiDAR company, listed in Hong Kong, that has extended its automotive LiDAR technology into robot sensing.","explanation":"RoboSense was founded in Shenzhen in August 2014 by chairman Chunxin Qiu together with Xiaorui Zhu and Letian Liu. LiDAR (light detection and ranging) is a sensor that measures distance with laser light to directly produce a 3D point cloud of the surroundings; RoboSense started with mechanical spinning LiDAR before mass-producing automotive-grade units such as its M series and solid-state units, and by early 2024 was described as the world's largest automotive LiDAR supplier, with shareholders including BYD, Xiaomi, and Cainiao. The company listed on the Hong Kong Stock Exchange on January 5, 2024 (ticker 2498). Over the past two years it has extended into robotics, launching the hemisphere-field-of-view Airy LiDAR and the AC1 active camera, which combines depth and RGB, and open-sourcing drivers and perception algorithms for quadrupeds, humanoids, and mobile robots.","example":"Mounting one Airy unit on a quadruped or humanoid's head lets it see both its surroundings and the ground under its feet at once, useful for mapping and obstacle avoidance.","related":["LiDAR","RoboSense Airy","RoboSense AC1","Solid-State LiDAR","Hesai Technology","Autonomous Driving"]},{"id":"livox","category":"company","sec":9,"tier":3,"sources":[{"title":"追觅旗下可庭科技 x 览沃达成战略合作（Livox 新闻，含公司简介）","url":"https://www.livoxtech.com/cn/news/dreame-and-livox-form-strategic-partnership"},{"title":"激光雷达在智能驾驶场景的破局之路（Livox 新闻）","url":"https://www.livoxtech.com/cn/news/10"},{"title":"Livox 览沃科技官网","url":"https://www.livoxtech.com/cn"}],"as_of":"2026-09","related_ids":["livox-mid-360","lidar","fast-lio2","simultaneous-localization-and-mapping","solid-state-lidar","dji"],"name":"Livox","alt":"览沃科技","abbr":"","aliases":["Shenzhen Livox Technology"],"one_liner":"A lidar company incubated inside DJI; its Mid-360 is a common choice on robots.","explanation":"Livox was founded in 2016, headquartered in Shenzhen; the company describes itself as an independent company created through DJI's internal incubation mechanism, focused on 3D lidar (sensors that emit laser pulses to measure distance and produce a surrounding point cloud). It keeps costs down using a rotating-mirror hybrid solid-state design with a non-repetitive scan pattern, launched its Mid series in January 2019, and says it has served more than 1,500 customers. Its most important product for the robotics community is the Mid-360: a 360° horizontal field of view, a built-in IMU, and a weight of about 265 grams, commonly mounted on quadrupeds, humanoids, and drones to run SLAM and lidar-inertial odometry methods such as FAST-LIO. Newer products include the kilometer-range Avia 2 and a lower-cost Mid-360S.","example":"In one of Livox's own case studies, a drone equipped with a Mid-360 navigates autonomously, flying fast through complex environments like forests while avoiding obstacles.","related":["Livox Mid-360","LiDAR","FAST-LIO2","Simultaneous Localization and Mapping","Solid-State LiDAR","DJI"]},{"id":"adaps-photonics","category":"company","sec":9,"tier":3,"sources":[{"title":"灵明光子 Adaps Photonics 官网","url":"https://www.adapsphotonics.com/"}],"as_of":"2026-09","related_ids":["single-photon-avalanche-diode","direct-time-of-flight","lidar","solid-state-lidar","depth-camera","vertical-cavity-surface-emitting-laser"],"name":"Adaps Photonics","alt":"灵明光子","abbr":"","aliases":["Adaps Photonics Technology"],"one_liner":"A Chinese maker of SPAD single-photon sensor chips, used for dToF ranging in lidar and phones.","explanation":"Adaps Photonics is a Chinese 3D-sensing chip company founded in 2018 by a team of returnee PhDs, with offices in Shenzhen, Hangzhou, Shanghai's Zhangjiang district, and Deqing County in Zhejiang province. Its core product is the SPAD (single-photon avalanche diode), a photosensor that can respond to individual photons, and the dToF (direct time-of-flight, which measures distance from the round-trip time of a light pulse) chips built around it. These include area-array dToF chips, multi-zone spot depth sensors, and SiPM (silicon photomultiplier) devices, built with backside-illuminated 3D stacking that layers the photosensitive layer directly on top of the computing circuitry. Its chips are used for autofocus and depth sensing in automotive lidar, smartphones, and XR headsets, and it also targets robot obstacle avoidance and SLAM.","example":"A solid-state lidar's receiver uses an Adaps Photonics area-array SPAD dToF chip, paired with a VCSEL transmitter, to measure distance.","related":["Single-Photon Avalanche Diode","Direct Time of Flight","LiDAR","Solid-State LiDAR","Depth Camera","Vertical-Cavity Surface-Emitting Laser"]},{"id":"optitrack","category":"company","sec":9,"tier":3,"sources":[{"title":"About | OptiTrack","url":"https://www.optitrack.com/about"},{"title":"NaturalPoint 官网","url":"https://www.naturalpoint.com"}],"as_of":"","related_ids":["optical-motion-capture","motion-capture","vicon","ground-truth","motion-retargeting","xsens"],"name":"OptiTrack","alt":"OptiTrack","abbr":"","aliases":["NaturalPoint"],"one_liner":"An American optical motion-capture brand that robotics labs use to measure ground truth.","explanation":"OptiTrack is an optical motion-capture brand owned by the American company NaturalPoint, headquartered in Corvallis, Oregon. The company's early products were head-tracking devices such as SmartNav and TrackIR, before it developed multi-camera motion-capture systems: infrared cameras mounted around a capture volume (such as its PrimeX series) track reflective markers attached to a person or object, and the Motive software computes sub-millimeter position and orientation in real time. In embodied AI it is mainly used to provide high-precision ground-truth pose — a measurement treated as the correct answer — for robots, drones, or objects, and also to record human motion that is later retargeted onto humanoid robots. Along with Vicon, it is one of the two optical motion-capture systems most commonly found in robotics labs.","example":"In quadruped or drone experiments, researchers treat the body pose measured by OptiTrack as ground truth to evaluate a state-estimation algorithm's error.","related":["Optical Motion Capture","Motion Capture","Vicon","Ground Truth","Motion Retargeting","Xsens"]},{"id":"vicon","category":"company","sec":9,"tier":3,"sources":[{"title":"Vicon: About Us","url":"https://www.vicon.com/about-us"},{"title":"Vicon: What Is Motion Capture","url":"https://www.vicon.com/about-us/what-is-motion-capture"}],"as_of":"2026-09","related_ids":["optical-motion-capture","motion-capture","optitrack","motion-retargeting","state-estimation","ground-truth"],"name":"Vicon","alt":"Vicon","abbr":"","aliases":["Vicon Motion Systems"],"one_liner":"A British optical motion-capture company, a common source of ground-truth pose in robotics labs.","explanation":"Vicon is a British optical motion-capture company based in Oxford, part of the listed company Oxford Metrics. The first Vicon product launched in 1979 out of a subsidiary of Oxford Instruments, becoming independent after a 1984 management buyout. Its systems surround a capture area with multiple infrared cameras that track reflective markers attached to a person or object, computing millimeter-accurate 3D position and orientation in real time. Beyond film animation and sports biomechanics, robotics labs commonly use it to supply external ground-truth pose for quadrupeds, drones, and robot arms — for evaluating state-estimation error — or to record human motion that is later retargeted onto humanoid robots. It competes directly with OptiTrack and is generally priced higher.","example":"When training a quadruped to run, researchers record the body's true position with Vicon and compare it against the robot's own estimated odometry to see how much it has drifted.","related":["Optical Motion Capture","Motion Capture","OptiTrack","Motion Retargeting","State Estimation","Ground Truth"]},{"id":"xsens","category":"company","sec":9,"tier":3,"sources":[{"title":"Xsens - Wikipedia","url":"https://en.wikipedia.org/wiki/Xsens"}],"as_of":"2026-09","related_ids":["inertial-motion-capture","motion-capture","inertial-measurement-unit","motion-retargeting","whole-body-teleoperation","optitrack"],"name":"Xsens","alt":"Xsens","abbr":"","aliases":["Movella"],"one_liner":"A Dutch maker of inertial motion-capture suits and IMUs, whose MVN suit is widely used to record human motion.","explanation":"Xsens was founded in Enschede, Netherlands, in 2000 by University of Twente graduates Casper Peeters and Per Slycke, building sensors and motion-capture systems around MEMS inertial measurement units — small chips that measure acceleration and angular velocity; its flagship products are the MTi-series IMU and the MVN inertial motion-capture suit. The company has changed hands several times: it came under mCube in 2017, its parent renamed itself Movella in 2021, Movella listed on Nasdaq in 2023 and was delisted in 2024, and according to Wikipedia it reverted to the Xsens brand name in 2026. Unlike optical motion capture, inertial motion capture needs no external cameras — a wearer can record full-body joint motion in any location — so it is commonly used to capture human motion that is then retargeted onto humanoid robots for teleoperation or training data.","example":"An operator wearing an MVN suit performs a motion, and every joint angle is retargeted onto a humanoid robot in real time for whole-body teleoperation.","related":["Inertial Motion Capture","Motion Capture","Inertial Measurement Unit","Motion Retargeting","Whole-Body Teleoperation","OptiTrack"]},{"id":"noitom","category":"company","sec":9,"tier":3,"sources":[{"title":"诺亦腾 Noitom 官网","url":"https://www.noitom.com.cn"},{"title":"启明星｜诺亦腾机器人完成Pre-A+轮融资，启明创投领投","url":"https://www.qimingvc.com/cn/news/%E5%90%AF%E6%98%8E%E6%98%9F%EF%BD%9C%E8%AF%BA%E4%BA%A6%E8%85%BE%E6%9C%BA%E5%99%A8%E4%BA%BA%E5%AE%8C%E6%88%90pre-a%E8%BD%AE%E8%9E%8D%E8%B5%84%EF%BC%8C%E5%90%AF%E6%98%8E%E5%88%9B%E6%8A%95%E9%A2%86%E6%8A%95"}],"as_of":"2025-12","related_ids":["inertial-motion-capture","motion-capture","motion-retargeting","whole-body-teleoperation","embodied-ai-data-service-provider","xsens"],"name":"Noitom","alt":"诺亦腾","abbr":"","aliases":["Noitom Ltd.","Noitom Technology"],"one_liner":"A Beijing motion-capture company whose gear is used for robot teleoperation and human-motion data.","explanation":"Noitom was founded in Beijing in 2012; its core technology is inertial motion capture — sensors with inertial measurement units worn on the body compute full-body pose in real time without relying on external cameras — alongside a hybrid optical-and-inertial approach. Its products include the Perception Neuron (PN) series and PN Hybrid, originally used in film and games, VR, sports, and medicine. With the rise of embodied AI, its motion-capture gear is now used to teleoperate humanoid robots, or to record human motion and retarget it onto a robot. Co-founder Ruoli Dai separately founded Noitom Robotics, which calls itself “a robotics company that doesn't build robots,” focused entirely on embodied training data; it closed a Pre-A+ round led by Qiming Venture Partners in December 2025, with several hundred million RMB raised in total.","example":"A data collector wears a PN inertial motion-capture suit to teleoperate a humanoid robot in real time while recording motion data.","related":["Inertial Motion Capture","Motion Capture","Motion Retargeting","Whole-Body Teleoperation","Embodied AI Data Service Provider","Xsens"]},{"id":"manus","category":"company","sec":9,"tier":3,"sources":[{"title":"MANUS - About us","url":"https://www.manus-meta.com/about-us"},{"title":"MANUS To Attend IROS 2026（Humanoid Robotics Technology）","url":"https://humanoidroboticstechnology.com/industry-news/manus-to-attend-iros-2026"}],"as_of":"2026-09","related_ids":["data-glove","teleoperation","motion-retargeting","dexterous-hand","motion-capture","haptic-glove"],"name":"MANUS","alt":"MANUS","abbr":"","aliases":["Manus VR","MANUS Technology Group"],"one_liner":"A Dutch data-glove maker whose Metagloves are widely used for dexterous-manipulation data capture and teleoperation.","explanation":"MANUS is a Dutch company founded in 2014, headquartered in Eindhoven; it started out making VR gloves (formerly known as Manus VR) before shifting to high-precision hand motion capture. Its main products include the Quantum Metagloves (2022), which use electromagnetic tracking on the fingers, the Metagloves Pro (2024), the haptic-feedback Metagloves Pro Haptic (2025), and the modular MANUS Nexus glove system. The company says its products are used by more than 2,000 robotics companies, AI labs, and motion-capture studios. In embodied AI, it is commonly paired with a VR headset for teleoperating a dexterous hand, retargeting the angles of a human hand's joints into robot commands to collect data. Note that it is a different product from the AI agent product also named Manus.","example":"ByteDance's GR-Dexter used a Meta Quest headset together with MANUS data gloves for two-handed teleoperation, collecting training data for a dexterous hand.","related":["Data Glove","Teleoperation","Motion Retargeting","Dexterous Hand","Motion Capture","Haptic Glove"]},{"id":"qualcomm","category":"company","sec":9,"tier":3,"sources":[{"title":"Qualcomm - Wikipedia","url":"https://en.wikipedia.org/wiki/Qualcomm"},{"title":"Qualcomm Introduces a Full Suite of Robotics Technologies（Edge AI and Vision Alliance）","url":"https://www.edge-ai-vision.com/2026/01/qualcomm-introduces-a-full-suite-of-robotics-technologies-powering-physical-ai-from-household-robots-up-to-full-size-humanoids/"},{"title":"Qualcomm to Acquire PickNik to Advance the Future of Open Robotics and Physical AI（高通官网）","url":"https://www.qualcomm.com/news/releases/2026/09/qualcomm-to-acquire-picknik-to-advance-the-future-of-open-roboti"}],"as_of":"2026-09","related_ids":["qualcomm-dragonwing-iq10","nvidia-jetson-thor","nvidia-jetson","arduino-esp32-microcontroller-boards","moveit-motion-planning-framework","onboard-compute-platform"],"name":"Qualcomm","alt":"高通","abbr":"","aliases":["Qualcomm Technologies"],"one_liner":"The US chip giant that launched its Dragonwing robotics chips and acquired Arduino and PickNik.","explanation":"Qualcomm was founded in 1985 and is headquartered in San Diego, California, started by Irwin Jacobs and six other former employees of the communications company Linkabit, and is known for CDMA technology and Snapdragon phone chips. On robotics, it launched the RB5 robotics development platform in 2020 and later grouped its industrial and embedded chips under the Dragonwing brand. In October 2025 it acquired the open-hardware company Arduino, launching the UNO Q development board built on its own chip; in January 2026 at CES it released the Dragonwing IQ10 processor and a full robotics hardware-and-software solution for full-size humanoids and industrial AMRs, saying it was co-designing the compute platform for Figure's next-generation humanoid, positioning it as a rival to NVIDIA's Jetson Thor. In September 2026 it also announced the acquisition of PickNik, the maintainer of MoveIt, pledging that MoveIt would remain open source.","example":"The Arduino UNO Q runs a Linux application and a real-time control program on the same board at once, letting a small robot handle both vision recognition and motor control.","related":["Qualcomm Dragonwing IQ10","NVIDIA Jetson Thor","NVIDIA Jetson","Arduino / ESP32 Microcontroller Boards","MoveIt Motion Planning Framework","Onboard Compute Platform"]},{"id":"advanced-micro-devices","category":"company","sec":9,"tier":3,"sources":[{"title":"AMD to Acquire World Labs to Advance the Future of AI Compute","url":"https://ir.amd.com/news-events/press-releases/detail/1299/amd-to-acquire-world-labs-to-advance-the-future-of-ai-compute"},{"title":"AMD - Wikipedia","url":"https://en.wikipedia.org/wiki/AMD"}],"as_of":"2026-09","related_ids":["world-labs","nvidia","cuda","spatial-intelligence","world-model"],"name":"Advanced Micro Devices","alt":"AMD","abbr":"AMD","aliases":["AMD"],"one_liner":"A US chipmaker known for CPUs, GPUs, and AI accelerators, which announced the acquisition of World Labs in 2026.","explanation":"AMD is a US chip company founded in 1969 by Jerry Sanders and others, headquartered in Santa Clara, California, and currently led by chair and CEO Lisa Su. Its products include the Ryzen and EPYC processors, Radeon graphics cards, the Instinct line of AI accelerators, and the open-source ROCm software stack, positioned as a rival to NVIDIA's CUDA; in 2022 it acquired the FPGA maker Xilinx for about $50 billion, gaining embedded vision-computing platforms such as Kria. Its most direct link to embodied AI: on September 28, 2026, AMD announced an all-stock acquisition of Fei-Fei Li's World Labs for about $8.2 billion, with Li becoming AMD's executive vice president and chief scientist. The deal is expected to close before the end of 2026 and is AMD's second-largest acquisition after Xilinx.","example":"In announcing the World Labs acquisition, AMD said the goal was to connect its hardware, software, and open models into one end-to-end AI ecosystem.","related":["World Labs","NVIDIA","CUDA","Spatial Intelligence","World Model"]},{"id":"horizon-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"Horizon Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Horizon_Robotics"},{"title":"HoloBrain-0 (arXiv 2602.12062)","url":"https://arxiv.org/abs/2602.12062"}],"as_of":"2026-02","related_ids":["holobrain-0","d-robotics","d-robotics-rdk-s100","horizon-robotics-openexplorer","autonomous-driving-talent-moving-into-embodied-ai","autonomous-driving"],"name":"Horizon Robotics","alt":"地平线","abbr":"","aliases":[],"one_liner":"A Beijing smart-driving chip company also building robot chips and VLA models.","explanation":"Horizon Robotics was founded in Beijing in July 2015; its founder, Kai Yu, previously led autonomous-driving and other research at Baidu. Its main business is its Journey series of smart-driving chips and accompanying algorithms, reportedly holding nearly half of China's autonomous-driving chip market in 2023; it listed on the Hong Kong Stock Exchange in October 2024 (ticker 9660). It connects to embodied AI in two ways: first, its robot computing-platform business has spun off as the independent company D-Robotics, which makes robot development boards such as the RDK S100; second, its own team open-sourced the VLA model HoloBrain-0 in February 2026, along with RoboOrchard, a toolchain covering data, training, and deployment. It is a representative example of autonomous-driving talent moving into embodied AI.","example":"HoloBrain-0 feeds camera parameters and a robot's URDF structural description into the model together, and is pretrained on data from multiple robot arms plus human-hand video.","related":["HoloBrain-0","D-Robotics","D-Robotics RDK S100","Horizon Robotics OpenExplorer","Autonomous-Driving Talent Moving into Embodied AI","Autonomous Driving"]},{"id":"d-robotics","category":"company","sec":9,"tier":3,"sources":[{"title":"地瓜机器人开发者社区","url":"https://developer.d-robotics.cc/"},{"title":"D-Robotics 官网（英文）","url":"https://en.d-robotics.cc/"},{"title":"量子位：地瓜机器人发布 RDK S100","url":"https://www.qbitai.com/2025/06/292932.html"}],"as_of":"2026-09","related_ids":[null,"d-robotics-rdk-s100","togetheros-bot",null,null,null],"name":"D-Robotics","alt":"地瓜机器人","abbr":"","aliases":[],"one_liner":"A robot-development-platform company spun out of Horizon Robotics, making RDK boards and robot software.","explanation":"D-Robotics was spun out of Horizon Robotics' robotics business; its CEO is Cong Wang. It builds “robot development infrastructure”: its RDK (Robot Development Kit) line of boards ranges from the 5-TOPS RDK X3 and 10-TOPS X5 up to the 80-128 TOPS S100 and the 560-TOPS S600, alongside software including TogetheROS.Bot (compatible with ROS 2), RDK OS, the RDK Studio development environment, and the NodeHub app center. The RDK S100, released in June 2025, puts the BPU that runs models and the MCU that handles real-time motor control on the same chip, positioning it as a rival to NVIDIA's Jetson. It is widely used by students, competition teams, and Chinese humanoid-robot makers, and it took part in ROSCon 2026 in September 2026.","example":"Run YOLO object detection on an RDK X5, then send the results to a ROS 2 navigation node through TogetheROS.Bot.","related":["Horizon Robotics","D-Robotics RDK S100","TogetheROS.Bot (D-Robotics)","NVIDIA Jetson","On-Device / Edge Deployment","Robot Operating System 2"]},{"id":"stanford-artificial-intelligence-laboratory","category":"company","sec":10,"tier":1,"sources":[{"title":"About SAIL - Stanford AI Lab","url":"https://ai.stanford.edu/about/"}],"as_of":"2026-09","related_ids":["berkeley-artificial-intelligence-research","mobile-aloha","openvla","voxposer","behavior-1k","toyota-research-institute"],"name":"Stanford Artificial Intelligence Laboratory","alt":"斯坦福人工智能实验室","abbr":"SAIL","aliases":["Stanford AI Lab","SAIL"],"one_liner":"Stanford University's AI lab, founded in 1963, and a major center for robot learning research.","explanation":"SAIL is the artificial intelligence lab in Stanford University's computer science department, founded in 1963 by John McCarthy, who coined the term “artificial intelligence.” Its current director is Chris Manning, a natural-language-processing researcher, and past and present members include Fei-Fei Li and Andrew Ng. Its research spans machine learning, computer vision, natural language processing, and robotics. Several widely cited results in embodied AI have come out of SAIL-affiliated groups, including the low-cost mobile bimanual platform Mobile ALOHA, the open-source vision-language-action model OpenVLA (developed with Berkeley, Toyota Research Institute, and others), VoxPoser, which uses a large language model to generate 3D value maps, and the household benchmark BEHAVIOR-1K.","example":"Mobile ALOHA (2024), released by a Stanford team, used a mobile bimanual platform costing about $32,000 to learn household chores like stir-frying shrimp and pushing in chairs from just a few dozen demonstrations.","related":["Berkeley Artificial Intelligence Research","Mobile ALOHA","OpenVLA","VoxPoser","BEHAVIOR-1K (BEHAVIOR Challenge)","Toyota Research Institute"]},{"id":"berkeley-artificial-intelligence-research","category":"company","sec":10,"tier":2,"sources":[{"title":"Welcome to the BAIR Blog","url":"https://bair.berkeley.edu/blog/2017/06/20/welcome/"},{"title":"Covariant (company) - Wikipedia","url":"https://en.wikipedia.org/wiki/Covariant_(company)"}],"as_of":"","related_ids":["physical-intelligence","covariant","octo","bridgedata-v2","hil-serl","dex-net-2-0"],"name":"Berkeley Artificial Intelligence Research","alt":"伯克利人工智能研究实验室","abbr":"BAIR","aliases":["Berkeley AI Research"],"one_liner":"A joint AI lab at UC Berkeley, and a major hub for robot learning research.","explanation":"BAIR brings together UC Berkeley research groups working on computer vision, machine learning, natural language processing, planning, and robotics, and launched the BAIR blog in 2017 to publish its results. It is one of the most prolific academic institutions in robot learning: Sergey Levine's group has produced real-world deep reinforcement learning, BridgeData V2, Octo, SERL, and HIL-SERL; Pieter Abbeel's group has been highly influential in imitation learning and reinforcement learning; Ken Goldberg's AUTOLAB built the grasping network Dex-Net; and Jitendra Malik's group developed Rapid Motor Adaptation (RMA) for legged locomotion control. A number of embodied AI companies have spun out of BAIR — Covariant, for instance, was founded by Abbeel and his students, and Levine is a co-founder of Physical Intelligence.","example":"The open-source generalist robot policy Octo was led by a BAIR team and trained on Open X-Embodiment data.","related":["Physical Intelligence","Covariant","Octo","BridgeData V2","HIL-SERL","Dex-Net 2.0"]},{"id":"cmu-robotics-institute","category":"company","sec":10,"tier":2,"sources":[{"title":"Robotics Institute - Wikipedia","url":"https://en.wikipedia.org/wiki/Robotics_Institute"}],"as_of":"","related_ids":["skild-ai","moravec-s-paradox","extreme-parkour","leap-hand","asap","autonomous-driving"],"name":"CMU Robotics Institute","alt":"卡内基梅隆大学机器人研究所","abbr":"CMU RI","aliases":["Carnegie Mellon Robotics Institute"],"one_liner":"The world's first academic robotics department, founded in 1979 and based in Pittsburgh.","explanation":"The CMU Robotics Institute was founded in 1979 by Raj Reddy and Angel Jordan with a $3 million grant from Westinghouse Electric, the world's first academic department dedicated to robotics, and it launched the world's first robotics PhD program in 1988. In its early years it was known for autonomous driving: the Navlab series of self-driving vehicles, and Boss, which won the 2007 DARPA Urban Challenge; Hans Moravec, who proposed 'Moravec's paradox,' also came out of the institute. In recent years it has been equally active in embodied AI — for example, Deepak Pathak's group produced Extreme Parkour and the LEAP Hand, and Guanya Shi's group contributed to the humanoid whole-body control work ASAP. Pathak and Abhinav Gupta also co-founded the robot foundation-model company Skild AI.","example":"CMU's self-driving car Boss won the 2007 DARPA Urban Challenge, completing the 55-mile course in about 4 hours and 20 minutes.","related":["Skild AI","Moravec's Paradox","Extreme Parkour","LEAP Hand","ASAP","Autonomous Driving"]},{"id":"mit-computer-science-and-artificial-intelligence-laboratory","category":"company","sec":10,"tier":2,"sources":[{"title":"MIT Computer Science and Artificial Intelligence Laboratory - Wikipedia","url":"https://en.wikipedia.org/wiki/MIT_Computer_Science_and_Artificial_Intelligence_Laboratory"}],"as_of":"","related_ids":["drake","subsumption-architecture",null,"f3rm","heterogeneous-pre-trained-transformers","visual-dexterity"],"name":"MIT Computer Science and Artificial Intelligence Laboratory","alt":"MIT 计算机科学与人工智能实验室","abbr":"CSAIL","aliases":["MIT CSAIL","CSAIL"],"one_liner":"MIT's largest computer science and AI lab, and a long-time center of robotics research.","explanation":"CSAIL is MIT's largest on-campus lab, formed on July 1, 2003 by merging the Laboratory for Computer Science (LCS) and the AI Lab. It is housed in the Stata Center in Cambridge, Massachusetts, and has been directed by Daniela Rus since 2012. Robotics has long been one of its strengths: former director Rodney Brooks proposed the subsumption architecture and co-founded iRobot, and Boston Dynamics founder Marc Raibert also came out of CSAIL. Today's embodied-AI researchers regularly run into its output, including the Drake robotics toolbox from Russ Tedrake's group, Visual Dexterity and the Decision Diffuser from Pulkit Agrawal's group, and projects such as F3RM and the Heterogeneous Pre-trained Transformer (HPT).","example":"HPT was proposed by CSAIL's Lirui Wang, Kaiming He, and collaborators at Meta FAIR, for joint pretraining across data from many different robots.","related":["Drake","Subsumption Architecture","Boston Dynamics","F3RM","Heterogeneous Pre-trained Transformers","Visual Dexterity"]},{"id":"eth-zurich-robotic-systems-lab","category":"company","sec":10,"tier":2,"sources":[{"title":"ETH Zurich Robotic Systems Lab - People","url":"https://rsl.ethz.ch/the-lab/people.html"},{"title":"Wikipedia: ANYbotics","url":"https://en.wikipedia.org/wiki/ANYbotics"}],"as_of":"2026-09","related_ids":["anybotics",null,"legged-gym",null,null,null],"name":"ETH Zurich Robotic Systems Lab","alt":"苏黎世联邦理工机器人系统实验室","abbr":"RSL","aliases":["ETH RSL","Legged Robotics (ETH)"],"one_liner":"ETH Zurich's center for legged robots and reinforcement-learning locomotion control, birthplace of ANYmal.","explanation":"The Robotic Systems Lab is part of ETH Zurich's Department of Mechanical and Process Engineering, led by Professor Marco Hutter, with a GitHub presence under the name leggedrobotics. It is best known for legged robots: the ANYmal quadruped was developed here, spinning out the company ANYbotics in 2016. The lab is one of the pioneers of the “train with reinforcement learning in simulation, then transfer to the real robot” approach to locomotion control, and has published a series of ANYmal-related results in Science Robotics, including actuator networks, teacher-student blind locomotion, and perceptive locomotion; it has also open-sourced the training codebases legged_gym and rsl_rl, now widely reused across quadruped and humanoid projects. Hutter also co-leads the Zurich branch of the RAI Institute (formerly the Boston Dynamics AI Institute), and lab alumni have gone on to found companies such as Flexion.","example":"Many humanoid and quadruped reinforcement-learning projects fork legged_gym directly and use the PPO implementation in rsl_rl to train a walking policy.","related":["ANYbotics","ANYbotics ANYmal","legged_gym","RSL RL","RL-based Locomotion Control","Flexion Robotics"]},{"id":"toyota-research-institute","category":"company","sec":10,"tier":2,"sources":[{"title":"Toyota Research Institute - Wikipedia","url":"https://en.wikipedia.org/wiki/Toyota_Research_Institute"},{"title":"TRI: AI-Powered Robot by Boston Dynamics and Toyota Research Institute","url":"https://tri.global/news/ai-powered-robot-boston-dynamics-and-toyota-research-institute-takes-key-step-towards-general"}],"as_of":"2025-08","related_ids":["large-behavior-model","diffusion-policy","boston-dynamics-atlas-2","boston-dynamics","openvla","prismatic-vlms"],"name":"Toyota Research Institute","alt":"丰田研究院","abbr":"TRI","aliases":["TRI"],"one_liner":"Toyota's US-based AI and robotics research arm, an early driver of diffusion policies and large behavior models.","explanation":"Toyota Research Institute is a research organization Toyota Motor Corporation established in the US in 2016, headquartered in Los Altos, California, with an additional site in Cambridge, Massachusetts. Its founding CEO was Gill Pratt, a roboticist and former DARPA program manager, and Toyota initially committed $1 billion over five years. Its focus areas include autonomous driving, materials science, and robotics. On the robotics side, TRI helped develop diffusion policy (2023, with Columbia University and others), which made generating actions with a diffusion model a mainstream technique; it later put more emphasis on large behavior models, general-purpose manipulation policies trained on large amounts of multi-task demonstrations, and in October 2024 worked with Boston Dynamics to apply one to the electric Atlas. TRI also contributed to OpenVLA and its backbone, Prismatic VLM.","example":"In 2025, TRI and Boston Dynamics showed a language-conditioned large behavior model controlling the electric Atlas through a multi-step tidying task.","related":["Large Behavior Model","Diffusion Policy","Boston Dynamics Atlas (Electric)","Boston Dynamics","OpenVLA","Prismatic VLMs"]},{"id":"beijing-academy-of-artificial-intelligence","category":"company","sec":10,"tier":2,"sources":[{"title":"Beijing Academy of Artificial Intelligence - Wikipedia","url":"https://en.wikipedia.org/wiki/Beijing_Academy_of_Artificial_Intelligence"}],"as_of":"2026-01","related_ids":["robobrain","roboos","braincerebellum-architecture","xingyuanzhi-robotics","embodied-foundation-model","open-weight-model"],"name":"Beijing Academy of Artificial Intelligence","alt":"北京智源人工智能研究院","abbr":"BAAI","aliases":["智源研究院 (short Chinese form)","智源 (colloquial Chinese short form)"],"one_liner":"A nonprofit AI research institute in Beijing, known for large models and the open-source embodied brain model RoboBrain.","explanation":"BAAI is a nonprofit AI research institute founded in Beijing in November 2018; its founding board chair, Hongjiang Zhang, previously served as CEO of Kingsoft. BAAI first became known for large models: Wudao 2.0, with 1.75 trillion parameters, the multimodal model Emu3, and the widely used BGE embedding model for retrieval, most of them open-sourced. In embodied AI, it released the embodied brain model RoboBrain (version 1.0 in February 2025, version 2.5 open-sourced in early 2026) and RoboOS, a cross-embodiment multi-robot collaboration framework; the approach has the brain model handle task decomposition, affordance, and trajectory prediction, then hands off to a lower-level controller or a VLA (vision-language-action model) for execution. BAAI has also incubated the embodied-AI company Xingyuanzhi Robotics. In March 2025, it was added to the U.S. Commerce Department's Entity List.","example":"Given a tabletop photo and the instruction 'put the cup to the left of the plate,' RoboBrain outputs where on the cup to grasp it and a movement trajectory point.","related":["RoboBrain","RoboOS","Brain–Cerebellum Architecture","Xingyuanzhi Robotics","Embodied Foundation Model","Open-weight Model"]},{"id":"shanghai-artificial-intelligence-laboratory","category":"company","sec":10,"tier":2,"sources":[{"title":"上海人工智能实验室官网","url":"https://www.shlab.org.cn/"},{"title":"上海人工智能实验室 - 百度百科","url":"https://baike.baidu.com/item/上海人工智能实验室"}],"as_of":"2026-07","related_ids":["internvla","internvl","internutopia","interndata-a1","genmanip"],"name":"Shanghai Artificial Intelligence Laboratory","alt":"上海人工智能实验室","abbr":"Shanghai AI Lab","aliases":["Shanghai AI Lab"],"one_liner":"A Shanghai AI research institute known for its open-source Intern model family and the InternVLA robotics models.","explanation":"Shanghai AI Lab is a new type of research institute that launched at the World Artificial Intelligence Conference in July 2020; scholars including Xiaoou Tang helped establish it early on, and its current director and chief scientist is Bowen Zhou. It is known for open-sourcing its work: the Intern family of large models (the InternLM language model and the InternVL multimodal model), the OpenMMLab computer-vision toolbox, the OpenCompass evaluation platform, and the MinerU document-parsing tool. Its embodied-AI work is led by the InternRobotics team, which has open-sourced the InternVLA-M1/A1/N1 vision-language-action models, the InternUtopia simulation platform, and the InternData-A1 synthetic dataset. In July 2026, at the World Artificial Intelligence Conference, it unveiled “Intern-Duanyan” (duanyan being a renowned type of Chinese inkstone), a platform for scientific discovery.","example":"InternVLA-A1 is pretrained on the synthetic InternData-A1 dataset and on AgiBot World, and can be downloaded and fine-tuned directly.","related":["InternVLA (Shanghai AI Laboratory)","InternVL","InternUtopia","InternData-A1","GenManip"]},{"id":"tsinghua-university-institute-for-interdisciplinary-informat","category":"company","sec":10,"tier":2,"sources":[{"title":"IIIS Introduction - Tsinghua University","url":"https://iiis.tsinghua.edu.cn/en/About/Introduction.htm"}],"as_of":"2025-12","related_ids":["shanghai-qi-zhi-institute","robotera","galaxea-ai","spirit-ai","3d-diffusion-policy"],"name":"Tsinghua University Institute for Interdisciplinary Information Sciences","alt":"清华大学交叉信息研究院（清华叉院）","abbr":"IIIS","aliases":["Tsinghua IIIS","IIIS"],"one_liner":"A Tsinghua University institute led by Andrew Yao, and one of China's key sources of embodied-AI talent and startups.","explanation":"The Institute for Interdisciplinary Information Sciences (IIIS) was founded at Tsinghua University in 2011 and is led by Turing Award winner Andrew Yao; the “Yao Class,” an elite computer-science undergraduate program he started in 2005, is also part of the institute. Its research and teaching span computer science, quantum information, and artificial intelligence. In recent years it has become a major source of embodied-AI work in China: several of its younger faculty work on robot learning, reinforcement learning, and autonomous driving, and have published results such as 3D Diffusion Policy (DP3); the founding teams of embodied-AI companies including RobotEra, Galaxea AI, and Spirit AI all include IIIS faculty. When a paper lists “Tsinghua IIIS” as an affiliation, this is the institute it means.","example":"3D Diffusion Policy (DP3, 2024) was proposed by Huazhe Xu's group at Tsinghua IIIS together with the Shanghai Qi Zhi Institute and others; its diffusion policy takes point-cloud input and can learn manipulation from only a small number of demonstrations.","related":["Shanghai Qi Zhi Institute","RobotEra","Galaxea AI","Spirit AI","3D Diffusion Policy"]},{"id":"beijing-humanoid-robot-innovation-center","category":"company","sec":10,"tier":2,"sources":[{"title":"北京人形机器人创新中心 关于我们","url":"https://www.x-humanoid.com/about.html"},{"title":"北京人形机器人创新中心官网","url":"https://www.x-humanoid.com/"}],"as_of":"2026-09","related_ids":["tiangong","huisi-kaiwu","pelican-vl","xr-1","robomind","humanoid-robot-half-marathon"],"name":"Beijing Humanoid Robot Innovation Center","alt":"北京人形机器人创新中心","abbr":"X-Humanoid","aliases":["北京人形 (colloquial Chinese short form)","National-Local Co-Built Embodied Intelligent Robot Innovation Center","Beijing Innovation Center of Humanoid Robotics"],"one_liner":"A Beijing humanoid robot public platform company, maker of the Tiangong robot series and open-source embodied models.","explanation":"The Beijing Humanoid Robot Innovation Center was established as a corporate innovation center in November 2023 in Beijing's Yizhuang tech zone (Beijing Economic-Technological Development Area); reports say it was jointly funded by UBTech, Xiaomi, Jingcheng Electromechanical, and Yizhuang's state-owned capital, among others. In October 2024, China's Ministry of Industry and Information Technology and the Beijing municipal government jointly designated it the 'National-Local Co-Built Embodied Intelligent Robot Innovation Center,' and it describes itself as China's first full-stack embodied AI hardware-and-software company. It positions itself as a shared industry platform: on hardware, its openly extensible Tiangong humanoid series, with Tiangong Ultra winning the Beijing Yizhuang Humanoid Robot Half Marathon in April 2025 and the line now at Tiangong 3.0; on software, the general-purpose embodied AI platform Huisi Kaiwu, the real-robot dataset RoboMIND, and, released around November 2025, the open-source embodied brain model Pelican-VL and the cross-embodiment VLA model XR-1.","example":"Tiangong Ultra finished the 2025 Beijing Yizhuang Humanoid Robot Half Marathon in 2 hours, 40 minutes, and 42 seconds, taking first place.","related":["Tiangong","Huisi Kaiwu (X-Humanoid general embodied AI platform)","Pelican-VL","XR-1","RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation)","Humanoid Robot Half Marathon (Beijing E-Town)"]},{"id":"shanghai-qi-zhi-institute","category":"company","sec":10,"tier":3,"sources":[{"title":"人形机器人「小星」问世，期智研究院瞄准具身通用人工智能（解放日报）","url":"https://www.jfdaily.com/wx/detail.do?id=638686"},{"title":"上海期智研究院瞄准全球前五AI高地（上海交大新闻网转载）","url":"https://news.sjtu.edu.cn/mtjj/20210108/139743.html"},{"title":"姚期智建的4个研究院，成了VC疯抢的项目库（投中网）","url":"https://m.chinaventure.com.cn/news/80-20260914-393251.html"}],"as_of":"2026-09","related_ids":["robotera","tsinghua-university-institute-for-interdisciplinary-informat","embodied-agi","humanoid-robot","world-artificial-intelligence-conference"],"name":"Shanghai Qi Zhi Institute","alt":"上海期智研究院","abbr":"","aliases":["Qi Zhi Institute"],"one_liner":"A Shanghai research institute led by Andrew Yao that has incubated embodied-AI companies such as RobotEra.","explanation":"The Shanghai Qi Zhi Institute was founded in 2020, led by Turing Award winner Andrew Chi-Chih Yao, who also directs Tsinghua University's Institute for Interdisciplinary Information Sciences. It is based at AI Tower in the West Bund Smart Valley in Xuhui, Shanghai, and is a new-style research institute backed by the Shanghai municipal government, covering frontier fields including AI and robotics. In embodied AI, it showed its own humanoid robot, nicknamed Xiaoxing, at the 2023 World Artificial Intelligence Conference, where Yao proposed that AI's next major challenge would be “embodied AGI.” The institute co-incubated the humanoid-robot company RobotEra (founded August 2023) together with Tsinghua's IIIS; RobotEra's founder, Jianyu Chen, is an IIIS assistant professor and a principal investigator at the Qi Zhi Institute.","example":"RobotEra, co-incubated by Tsinghua's IIIS and the Shanghai Qi Zhi Institute, is the institute's best-known embodied-AI project.","related":["RobotEra","Tsinghua University Institute for Interdisciplinary Information Sciences","Embodied AGI","Humanoid Robot","World Artificial Intelligence Conference"]},{"id":"national-and-local-co-built-humanoid-robotics-innovation-cen","category":"company","sec":10,"tier":3,"sources":[{"title":"OpenLoong 开源社区","url":"https://www.openloong.org.cn/cn"},{"title":"国家地方共建人形机器人创新中心发布青龙","url":"https://www.leaderobot.com/news/4397"},{"title":"白虎-VTouch 发布报道","url":"https://news.qq.com/rain/a/20260126A04W2D00"}],"as_of":"2026-01","related_ids":[null,"baihu-vtouch-visuo-tactile-dataset","beijing-humanoid-robot-innovation-center",null,null],"name":"National and Local Co-built Humanoid Robotics Innovation Center (Shanghai)","alt":"国家地方共建人形机器人创新中心（上海人形机器人创新中心）","abbr":"","aliases":["Shanghai Humanoid Robot Innovation Center","Humanoid Robot (Shanghai) Co., Ltd."],"one_liner":"A national-level humanoid-robotics platform in Shanghai, China, that built the open-source humanoid robot Qinglong.","explanation":"This is a humanoid-robotics innovation platform backed jointly by China's central and Shanghai municipal governments, reportedly formed in 2024 and operated through Humanoid Robot (Shanghai) Co., Ltd., with Lei Jiang as chief scientist. It is meant to serve as a public platform for the industry rather than sell its own products. On July 6, 2024, at the World Artificial Intelligence Conference (WAIC) in Shanghai, it unveiled the full-size open-source humanoid robot Qinglong (185 cm tall, with 43 active degrees of freedom). It also runs the OpenLoong open-source community, incubated by the OpenAtom Foundation, which publishes Qinglong's full hardware drawings, control framework, and whole-body dynamics software; the center has since added the smaller Qinglong Mini, the “Gewu” simulation platform, and the “Baihu” dataset. In January 2026, together with Vtouch Technology, it released the Baihu-VTouch visuotactile dataset.","example":"A university lab can download Qinglong's hardware designs and control code from the OpenLoong community to reproduce or modify its own humanoid robot.","related":["Qinglong (OpenLoong)","Baihu-VTouch Visuo-Tactile Dataset","Beijing Humanoid Robot Innovation Center","Open-source Hardware","Humanoid Robot"]},{"id":"zhejiang-humanoid-robot-innovation-center","category":"company","sec":10,"tier":3,"sources":[{"title":"公司简介｜浙江人形机器人创新中心有限公司","url":"https://www.zj-humanoid.com/about"},{"title":"浙江人形机器人创新中心获4.5亿元Pre-A轮融资（经济参考网）","url":"http://jjckb.xinhuanet.com/20260123/db7e1207902344a0800570141d46eeb2/c.html"},{"title":"浙江人形机器人创新中心在宁波启动（科技部）","url":"https://www.most.gov.cn/dfkj/zj/zxdt/202404/t20240418_190373.html"}],"as_of":"2026-03","related_ids":["beijing-humanoid-robot-innovation-center","national-and-local-co-built-humanoid-robotics-innovation-cen","braincerebellum-architecture","wheeled-humanoid-robot","humanoid-robot"],"name":"Zhejiang Humanoid Robot Innovation Center","alt":"浙江人形机器人创新中心","abbr":"","aliases":["Zhejiang Humanoid Center"],"one_liner":"A Ningbo-based humanoid R&D company co-founded by the Ningbo city government and Zhejiang University's Rong Xiong team.","explanation":"The Zhejiang Humanoid Robot Innovation Center was founded in late 2023 in Haishu District, Ningbo, co-built by the Ningbo municipal government and Professor Rong Xiong's team at Zhejiang University's Institute of Cyber-Systems and Control, formally launching on March 27, 2024. Its focus is a humanoid robot's “brain and cerebellum” — high-level understanding and planning paired with lower-level motion control — and complete robot bodies: in March 2024 it released the full-size, 39-degree-of-freedom prototype Lingchang Zhe No.1, followed by the 41-degree-of-freedom Lingchang Zhe No.2 that August, and its current lineup includes the NAVIAI WA2 wheeled-arm humanoid and the WA1. Like similar centers in Beijing and Shanghai, it is a government-backed industry platform that also raises money as a company: in January 2026 it closed a RMB 450 million Pre-A round with investors including Supcon Technology and Legend Capital, and in March it partnered with Germany's KION Group to showcase logistics solutions at LogiMAT 2026.","example":"The NAVIAI WA2 wheeled-arm humanoid gave continuous tote-grasping demonstrations at the logistics trade show.","related":["Beijing Humanoid Robot Innovation Center","National and Local Co-built Humanoid Robotics Innovation Center (Shanghai)","Brain–Cerebellum Architecture","Wheeled Humanoid Robot","Humanoid Robot"]},{"id":"beijing-institute-for-general-artificial-intelligence","category":"company","sec":10,"tier":3,"sources":[{"title":"BIGAI 官网 About","url":"https://www.bigai.ai/about"},{"title":"Wikipedia: Beijing Institute for General Artificial Intelligence","url":"https://en.wikipedia.org/wiki/Beijing_Institute_for_General_Artificial_Intelligence"},{"title":"BIGAI 官网","url":"https://www.bigai.ai/"}],"as_of":"2026-07","related_ids":["leo",null,"beijing-academy-of-artificial-intelligence","pku-epic-lab","world-humanoid-robot-games"],"name":"Beijing Institute for General Artificial Intelligence","alt":"北京通用人工智能研究院","abbr":"BIGAI","aliases":["BIGAI"],"one_liner":"A Beijing research institute led by Song-Chun Zhu, focused on general and embodied artificial intelligence.","explanation":"The Beijing Institute for General Artificial Intelligence was founded in 2020, backed by the Beijing municipal government and the Ministry of Science and Technology, in partnership with Peking University and Tsinghua University; its director is Song-Chun Zhu, who moved to the US in 1992, spent 18 years on the UCLA faculty, and returned to China in 2020. It champions a “small data, big tasks” approach, emphasizing reasoning and value-driven behavior rather than simply scaling up large models. Its results include the virtual agent Tong Tong (released January 2024, upgraded to version 2.0 that April), a General Intelligence Test (通智测试) benchmark, the 3D embodied generalist model LEO (ICML 2024), the simulation-training platform TongSIM (open-sourced December 2025), and the TongAgents agent framework. It also works on humanoid robot locomotion control, and has won awards at the World Humanoid Robot Games.","example":"LEO feeds first-person images, object-level 3D point clouds, and text instructions into a large language model together, letting it answer questions, navigate, and manipulate objects within a 3D scene.","related":["LEO (BIGAI)","Embodied AI","Beijing Academy of Artificial Intelligence","PKU EPIC Lab","World Humanoid Robot Games"]},{"id":"pku-epic-lab","category":"company","sec":10,"tier":3,"sources":[{"title":"PKU EPIC Lab 主页","url":"https://pku-epic.github.io/"}],"as_of":"2026-09","related_ids":["graspvla","navid","dexgraspnet","galbot-g1","syngrasp-1b","university-big-tech-autonomous-driving-founder-lineage"],"name":"PKU EPIC Lab","alt":"北京大学 EPIC 实验室（王鹤组）","abbr":"EPIC Lab","aliases":["PKU Embodied Perception and InteraCtion Lab","He Wang's Lab"],"one_liner":"He Wang's lab at Peking University for embodied perception and interaction, closely tied to Galbot.","explanation":"EPIC Lab sits within Peking University's Center on Frontiers of Computing Studies (CFCS) and is led by He Wang. It studies how robots perceive and interact with complex 3D environments, covering dexterous grasping, vision-language-action (VLA) models, embodied navigation, and 3D object perception. Notable work includes the large-scale dexterous-grasping dataset DexGraspNet, the grasping method UniDexGrasp, the cross-category part-level manipulation benchmark GAPartNet, the video-based navigation model NaVid, and GraspVLA, a grasping foundation model built jointly with Galbot. He Wang is also the founder of Galbot, and the lab's results are often deployed on Galbot's robots, making it a typical example of an academic-origin embodied AI team.","example":"GraspVLA is first pre-trained on the billion-scale synthetic grasping dataset SynGrasp-1B, then deployed on Galbot's robots.","related":["GraspVLA","NaVid","DexGraspNet","Galbot G1","SynGrasp-1B","University / Big-Tech / Autonomous-Driving Founder Lineage"]},{"id":"sjtu-mvig-lab","category":"company","sec":10,"tier":3,"sources":[{"title":"MVIG 实验室主页","url":"https://www.mvig.org/"},{"title":"GraspNet 项目主页","url":"https://graspnet.net/"},{"title":"上海交大新跑出一家具身智能公司「穹彻智能」（雷峰网）","url":"https://m.leiphone.com/category/ai/IPEN8fseWTn7UvjV.html"}],"as_of":"2026-09","related_ids":["graspnet-1billion","anygrasp","rh20t","airexo","oakink","noematrix"],"name":"SJTU MVIG Lab","alt":"上海交通大学 MVIG 实验室（卢策吾组）","abbr":"MVIG","aliases":["SJTU Machine Vision and Intelligence Group","Cewu Lu's Lab"],"one_liner":"Cewu Lu's lab at Shanghai Jiao Tong University, one of China's earlier academic teams working on robot manipulation data and grasping.","explanation":"MVIG is the Machine Vision and Intelligence Group at Shanghai Jiao Tong University, led by Professor Cewu Lu. It started out working on human pose and activity understanding, open-sourcing the pose-estimation tool AlphaPose, before turning to embodied AI and robot manipulation. Its guiding idea is to have robots learn general-purpose behavior from large amounts of human activity video and demonstrations. The lab's embodied-AI work is mostly datasets and foundational tools: the grasping benchmark GraspNet-1Billion and the grasp-perception system AnyGrasp, the real-robot manipulation dataset RH20T, the hand-object interaction dataset OakInk, the dual-arm exoskeleton AirExo, and force-aware policies such as FoAR and RDP. In 2023, Cewu Lu founded the embodied-AI company Noematrix to commercialize part of the lab's work. Readers of embodied-AI papers will often see these datasets and tools cited as benchmarks or data-collection methods.","example":"Many grasping papers call the AnyGrasp SDK directly: feed in a depth camera's point cloud, and it returns a set of scored gripper grasp poses for the scene.","related":["GraspNet-1Billion","AnyGrasp","RH20T","AirExo","OakInk","Noematrix"]},{"id":"opendrivelab","category":"company","sec":10,"tier":3,"sources":[{"title":"OpenDriveLab 官网","url":"https://opendrivelab.com/"}],"as_of":"2026-07","related_ids":["uniad","agibot-world","agibot","end-to-end","autonomous-driving-talent-moving-into-embodied-ai"],"name":"OpenDriveLab","alt":"香港大学 OpenDriveLab","abbr":"","aliases":["OpenDriveLab (HKU)"],"one_liner":"A University of Hong Kong lab led by Hongyang Li, working on end-to-end autonomous driving and embodied AI.","explanation":"OpenDriveLab was founded in 2021 and is now based at the University of Hong Kong under Hongyang Li, with teams in both Hong Kong and Shanghai. It first became known for autonomous driving: UniAD folds perception, prediction, and planning into a single end-to-end network — going straight from sensor input to a driving plan — and won the CVPR 2023 Best Paper award. The lab later moved into embodied AI, partnering with AgiBot to release the large-scale real-robot manipulation dataset AgiBot World (a Best Paper finalist at IROS 2025), and it also works on humanoid robot control and long-horizon manipulation. In February 2026 it announced partnerships with Unitree, Noitom, and BrainCo. For newcomers, it stands as one of the representative academic teams that moved from autonomous driving into embodied AI.","example":"UniAD (CVPR 2023 Best Paper) and the AgiBot World dataset both came out of this lab.","related":["UniAD","AgiBot World","AgiBot","End-to-End","Autonomous-Driving Talent Moving into Embodied AI"]},{"id":"allen-institute-for-ai","category":"company","sec":10,"tier":3,"sources":[{"title":"MolmoAct 2: An open foundation for robots that work in the real world (Ai2)","url":"https://allenai.org/blog/molmoact2"},{"title":"Ai2 releases MolmoAct 2 (SiliconANGLE)","url":"https://siliconangle.com/2026/05/05/ai2-releases-molmoact-2-enhancing-robot-intelligence-real-world"}],"as_of":"2026-05","related_ids":["molmoact","molmo","ai2-thor","procthor","molmospaces","open-weight-model"],"name":"Allen Institute for AI","alt":"艾伦人工智能研究所","abbr":"Ai2","aliases":["AI2","Ai2"],"one_liner":"A US nonprofit AI research institute known for fully open models and embodied simulation environments.","explanation":"Ai2 was founded in 2014 by Microsoft co-founder Paul Allen, is headquartered in Seattle, and operates as a nonprofit research institute; its current CEO is Ali Farhadi. Its hallmark is releasing weights, training data, and code all together as open source, with flagship projects including the language model OLMo and the vision-language model Molmo. On the embodied side, it built indoor simulation environments including AI2-THOR, ProcTHOR, and Holodeck; in 2025 it released MolmoAct, an action model that can reason in space, followed by MolmoAct 2 in May 2026, releasing its code and integrating with LeRobot. Students getting started with open-source VLA models often end up using its models and simulators.","example":"Within a few weeks of release, MolmoAct 2 had been downloaded more than 400,000 times and shipped with LeRobot integration.","related":["MolmoAct","Molmo (Ai2)","AI2-THOR","ProcTHOR (Large-Scale Embodied AI Using Procedural Generation)","MolmoSpaces","Open-weight Model"]},{"id":"robotics-and-ai-institute","category":"company","sec":10,"tier":3,"sources":[{"title":"RAI Institute - About","url":"https://rai-inst.com/about/"},{"title":"RAI Institute 官网首页","url":"https://rai-inst.com/"}],"as_of":"2026-05","related_ids":["boston-dynamics","hyundai-motor-group","rl-based-locomotion-control","dexterous-manipulation","boston-dynamics-spot","eth-zurich-robotic-systems-lab"],"name":"Robotics and AI Institute","alt":"RAI 研究所（机器人与人工智能研究所）","abbr":"RAI Institute","aliases":["RAI Institute","Boston Dynamics AI Institute","The AI Institute"],"one_liner":"A robotics basic-research institute led by Boston Dynamics founder Marc Raibert and funded by Hyundai.","explanation":"The RAI Institute began in 2022 as the Boston Dynamics AI Institute, launched by Hyundai Motor Group together with Boston Dynamics with a reported initial investment of more than $400 million, before later taking its current name. Executive Director Marc Raibert, the founder of Boston Dynamics, leads the institute from its main campus in Cambridge, Massachusetts, with an additional office in Zurich, Switzerland. Rather than selling products, it focuses on long-term basic research in robotics and AI; its website lists dexterous manipulation, learning methods for control, data-driven AI models, navigation in complex environments, and robot ethics as its research areas. It collaborates with Boston Dynamics on reinforcement-learning-based locomotion control. In May 2026 it published a research demo of a robot using onboard vision and a multi-fingered hand to juggle by throwing and catching.","example":"","related":["Boston Dynamics","Hyundai Motor Group","RL-based Locomotion Control","Dexterous Manipulation","Boston Dynamics Spot","ETH Zurich Robotic Systems Lab"]},{"id":"dlr-institute-of-robotics-and-mechatronics","category":"company","sec":10,"tier":3,"sources":[{"title":"DLR: History of the institute","url":"https://www.dlr.de/en/rm/about-us/institute/history"},{"title":"ROBOTS Guide: Rollin' Justin","url":"https://robotsguide.com/robots/justin"}],"as_of":"2022","related_ids":["kuka-lbr-iiwa","impedance-control","joint-torque-sensor","collaborative-robot","dexterous-hand","kuka"],"name":"DLR Institute of Robotics and Mechatronics","alt":"德国宇航中心机器人与机电研究所","abbr":"DLR RMC","aliases":["DLR-RM","Robotics and Mechatronics Center"],"one_liner":"The German Aerospace Center's robotics institute, creator of lightweight robot arms and the Justin humanoid.","explanation":"This is a research institute under the German Aerospace Center (DLR), located in Oberpfaffenhofen near Munich; Alin Albu-Schäffer has been its director since 2012, and it is one of Europe's most influential robotics research institutions. It began with space teleoperation — ROTEX, in 1993, was Germany's first space-robotics experiment. Its biggest contribution is its lightweight robot arms, which carry torque sensors and can do impedance control (letting a joint behave like a spring, so it moves compliantly): it launched the first-generation LBR in 1995, then licensed the LBR III to KUKA in 2004, which later grew into the commercial LBR iiwa. It has also developed the DLR dexterous hand and the dual-arm humanoid Justin (2006, and again in 2008 with a wheeled base as Rollin' Justin), and in recent years has worked on satellite-servicing robots and lunar-surface exploration trials.","example":"The KUKA LBR iiwa, a collaborative arm widely used in industry, traces its technology back to the institute's third-generation lightweight arm, the LBR III.","related":["KUKA LBR iiwa","Impedance Control","Joint Torque Sensor","Collaborative Robot","Dexterous Hand","KUKA"]},{"id":"florida-institute-for-human-and-machine-cognition","category":"company","sec":10,"tier":3,"sources":[{"title":"Wikipedia: Florida Institute for Human and Machine Cognition","url":"https://en.wikipedia.org/wiki/Florida_Institute_for_Human_and_Machine_Cognition"},{"title":"IHMC Robotics","url":"https://robots.ihmc.us/"}],"as_of":"2016","related_ids":[null,null,null,null,null],"name":"Florida Institute for Human and Machine Cognition","alt":"IHMC（美国人机认知研究所）","abbr":"IHMC","aliases":["IHMC"],"one_liner":"A nonprofit Florida research institute and a leading team in humanoid-robot walking control.","explanation":"IHMC was founded in 1990 by Kenneth Ford and others on the campus of the University of West Florida, headquartered in Pensacola, Florida, as a nonprofit institute within the Florida State University System, researching artificial intelligence, robotics, exoskeletons, and human performance. Its best-known robotics work is bipedal walking and whole-body control: in the DARPA Robotics Challenge, IHMC used Boston Dynamics' hydraulic Atlas to win first place in the virtual round, then took second place at the 2015 final. The team has long open-sourced its humanoid whole-body control and walking software, used on platforms including Atlas and NASA's Valkyrie, and it also researches walking-assistance exoskeletons for people with paraplegia, competing in the first Cybathlon in 2016.","example":"At the DARPA Robotics Challenge final, IHMC's Atlas fell and was damaged but still completed its tasks, taking second place.","related":["DARPA Robotics Challenge (DRC)","Boston Dynamics Atlas (Hydraulic)","NASA Valkyrie (R5)","Whole-Body Control","Exoskeleton"]},{"id":"disney-research","category":"company","sec":10,"tier":3,"sources":[{"title":"Nvidia and Google DeepMind will help power Disney's cute robots（TechCrunch）","url":"https://techcrunch.com/2025/03/18/nvidia-and-google-deepmind-will-help-power-disneys-cute-robots/"},{"title":"NVIDIA GTC: Walt Disney Imagineering's Olaf Robotic Character（Disney Experiences）","url":"https://disneyexperiences.com/nvidia-gtc-olaf-robotic-character/"},{"title":"Olaf: Bringing an Animated Character to Life in the Physical World（Disney Research）","url":"https://la.disneyresearch.com/publication/olaf-bringing-an-animated-character-to-life-in-the-physical-world/"}],"as_of":"2026-03","related_ids":["disney-research-bdx-droid","newton-physics-engine","rl-based-locomotion-control","sim-to-real-transfer","nvidia-gtc"],"name":"Disney Research","alt":"迪士尼研究院","abbr":"","aliases":["DisneyResearch|Studios"],"one_liner":"The Walt Disney Company's R&D arm, which uses reinforcement learning to turn animated characters into physical robots.","explanation":"Disney Research was founded by The Walt Disney Company in 2008 to work on graphics, vision, machine learning, and robotics. It now splits into two parts: DisneyResearch|Studios in Zurich, led by ETH graphics professor Markus Gross, serves film production; the robotics side sits within Walt Disney Imagineering, with teams in Los Angeles and Zurich. Its robotics group, led by Moritz Bächer in Zurich, specializes in “character robots”: animators first design personality-filled motions, and reinforcement learning trains a control policy in simulation so the physical robot moves like the character without falling over. Its best-known work includes the Star Wars-style bipedal BDX droid, which appeared in NVIDIA's GTC keynote in March 2025, and Olaf from Frozen, who shared the stage with NVIDIA CEO Jensen Huang at GTC in March 2026 and has performed on a boat ride at Disneyland Paris since March 29. Disney Research also co-launched the open-source physics engine Newton together with NVIDIA and Google DeepMind.","example":"Olaf hides his two legs under a foam “skirt”; when training its control policy, the team even fed motor temperature in as an input, to keep the small motors in its skinny neck from overheating.","related":["Disney Research BDX Droid","Newton Physics Engine","RL-based Locomotion Control","Sim-to-Real Transfer","NVIDIA GTC"]},{"id":"honda-motor-co-ltd","category":"company","sec":11,"tier":2,"sources":[{"title":"Honda P2 Humanoid Bipedal Robot Recognized as IEEE Milestone (Honda, 2026-04)","url":"https://global.honda/en/topics/2026/c_2026-04-28aeng.html"},{"title":"Honda Avatar Robot | Honda Technology","url":"https://global.honda/en/tech/Avatar_robot/"},{"title":"Honda Robotics Returns: The Dexterous Hand After ASIMO (Intelligent Living, 2026-08)","url":"https://www.intelligentliving.co/honda-robotics-dexterous-hand-asimo/"}],"as_of":"2026-08","related_ids":["honda-asimo","zero-moment-point","bipedal-locomotion","hrp-humanoid-robot-series","dexterous-hand","teleoperation"],"name":"Honda Motor Co., Ltd.","alt":"本田","abbr":"","aliases":["Honda"],"one_liner":"The Japanese automaker behind ASIMO, a pioneer of bipedal humanoid robot research.","explanation":"Honda Motor was founded by Soichiro Honda in 1948, is headquartered in Tokyo, and is best known for motorcycles and cars. It began researching bipedal robots in 1986; its 1996 prototype, P2, carried its own power source and could walk stably on slopes and stairs, earning an IEEE Milestone designation in April 2026. Its year-2000 robot, ASIMO, walked and ran using model-based balance control such as the zero moment point, and became the benchmark humanoid robot of the era before deep learning took off. In 2003, Honda set up Honda Research Institutes (HRI) in Japan, the US, and Germany for frontier AI and robotics research. Around ASIMO's retirement in March 2022, Honda shifted toward more practical directions: it unveiled a teleoperated “avatar robot” in 2021, and in May 2026 showed a 16-degree-of-freedom multi-fingered dexterous hand in Tokyo, reportedly planned for factory assembly use in the early 2030s.","example":"Honda's avatar robot is remotely controlled by a person, while its AI infers what the operator is trying to grasp and automatically fine-tunes the hand's posture and grip force — aimed at work like disaster response or equipment repair where sending a person is impractical.","related":["Honda ASIMO","Zero Moment Point","Bipedal Locomotion","HRP Humanoid Robot Series","Dexterous Hand","Teleoperation"]},{"id":"willow-garage","category":"company","sec":11,"tier":3,"sources":[{"title":"ROBOTS Guide: PR2","url":"https://robotsguide.com/robots/pr2"},{"title":"Clearpath Robotics: Clearpath Welcomes PR2 to the Family","url":"https://clearpathrobotics.com/blog/2014/01/clearpath-welcomes-pr2"}],"as_of":"2014-01","related_ids":["robot-operating-system","willow-garage-pr2","open-robotics","opencv","point-cloud-library","mobile-manipulation"],"name":"Willow Garage","alt":"Willow Garage","abbr":"","aliases":[],"one_liner":"The US robotics research company that incubated ROS and the PR2, shut down in 2014.","explanation":"Willow Garage was founded in late 2006 by early Google engineer Scott Hassan, based in Menlo Park, California. He funded a project by Stanford researchers Keenan Wyrobek and Eric Berger called PR1, and building on it the company developed the open-source robotics framework ROS and the dual-arm mobile robot PR2 (Personal Robot 2), giving away PR2 units to multiple universities so researchers could share code on a common platform. Willow Garage also funded the development of OpenCV and the Point Cloud Library (PCL). The Open Source Robotics Foundation (OSRF) was founded in 2012 and took over ROS maintenance in 2013; the company itself shut down in 2014, handing PR2 support off to Clearpath Robotics. ROS remains the most widely used software framework in robotics today, which is its biggest legacy.","example":"","related":["Robot Operating System","Willow Garage PR2 (Personal Robot 2)","Open Robotics","OpenCV (Open Source Computer Vision Library)","Point Cloud Library (PCL)","Mobile Manipulation"]},{"id":"open-robotics","category":"company","sec":11,"tier":3,"sources":[{"title":"Open Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Open_Robotics"}],"as_of":"2024-04","related_ids":["robot-operating-system","robot-operating-system-2","gazebo","willow-garage","intrinsic","fleet-management-system"],"name":"Open Robotics","alt":"Open Robotics / 开源机器人基金会","abbr":"OSRF","aliases":["Open Source Robotics Foundation","OSRF"],"one_liner":"The US nonprofit that maintains ROS, Gazebo, and other open-source robotics software.","explanation":"Founded in 2012 as a spinout from the robotics lab Willow Garage, Open Robotics is headquartered in Mountain View, California, and is the public-facing name of the Open Source Robotics Foundation (OSRF). It maintains the Robot Operating System (ROS, the most widely used communication and tooling framework for robot software), the Gazebo simulator, and the multi-robot fleet framework Open-RMF. In December 2022, Intrinsic — a robotics company under Google's parent Alphabet — acquired its for-profit subsidiary OSRC, while the foundation itself kept operating independently. In April 2024 it launched the Open Source Robotics Alliance (OSRA), a membership program to fund ROS and related projects long-term. For newcomers, installing ROS 2, reading its docs, or checking a release roadmap will eventually lead back to a site or repository that Open Robotics maintains.","example":"The release schedule and official docs for every ROS 2 distribution (such as Humble or Jazzy) are maintained under OSRF's lead.","related":["Robot Operating System","Robot Operating System 2","Gazebo","Willow Garage","Intrinsic (Alphabet)","Fleet Management System (e.g. Open-RMF)"]},{"id":"rethink-robotics","category":"company","sec":11,"tier":3,"sources":[{"title":"Rethink Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Rethink_Robotics"}],"as_of":"2025","related_ids":["rethink-robotics-baxter","rethink-robotics-sawyer","collaborative-robot","kinesthetic-teaching","series-elastic-actuator"],"name":"Rethink Robotics","alt":"Rethink Robotics","abbr":"","aliases":["Heartland Robotics"],"one_liner":"The collaborative-robot pioneer behind Baxter and Sawyer, now defunct after several shutdowns and restarts.","explanation":"Rethink Robotics was founded in Boston in 2008 as Heartland Robotics by Rodney Brooks — an iRobot co-founder and MIT professor — and Ann Whittaker. It launched the dual-arm robot Baxter in 2012 and the single-arm Sawyer in 2015, both built around low cost, hand-guided teaching, and a “soft” response when they bump into a person; the research edition became a staple platform in robot-learning labs through the 2010s. Sales fell short of expectations, and the company shut down and sold its assets in October 2018; Germany's HAHN Group bought the patents, trademarks, and Intera software and kept selling Sawyer. HAHN's subsidiary Rethink Robotics GmbH closed in August 2024; a new company relaunched with new products that September, but it shut down again in 2025.","example":"The robot arm modeled in the Meta-World simulation benchmark is a Sawyer.","related":["Rethink Robotics Baxter","Rethink Robotics Sawyer","Collaborative Robot","Kinesthetic Teaching","Series Elastic Actuator (SEA)"]},{"id":"everyday-robots","category":"company","sec":11,"tier":3,"sources":[{"title":"X, the moonshot factory: Everyday Robots","url":"https://x.company/projects/everyday-robots/"}],"as_of":"2023-01","related_ids":[null,null,"rt-1","saycan",null],"name":"Everyday Robots (Alphabet X)","alt":"Everyday Robots","abbr":"","aliases":["Everyday Robot Project","Google X Everyday Robots Project"],"one_liner":"A general-purpose robot project from Alphabet's X lab, shut down in 2023.","explanation":"This was a general-purpose robotics project incubated at X, Alphabet's “moonshot factory,” aiming to get robots to learn a wide range of tasks in unstructured environments like homes and offices, rather than being programmed task by task. The team built its own wheeled, single-arm mobile manipulator, deployed in fleets around Google offices to wipe down tables, sort trash, and open doors, trained at large scale using cloud-based simulation. Its biggest impact was academic: work done with Google Research, including SayCan and RT-1, used these robots to collect data and run experiments, and the RT-1 data was later folded into Open X-Embodiment. The project was shut down in early 2023 during Alphabet layoffs, with some people and technology absorbed into Google's robotics research, which later became part of Google DeepMind.","example":"The RT-1 paper's 13 robots and roughly 130,000 demonstrations were collected on Everyday Robots' mobile manipulators.","related":["Everyday Robots Mobile Manipulator","Google DeepMind","RT-1","SayCan","Open X-Embodiment"]},{"id":"hanson-robotics","category":"company","sec":11,"tier":3,"sources":[{"title":"Hanson Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Hanson_Robotics"},{"title":"David Hanson (robotics designer) - Wikipedia","url":"https://en.wikipedia.org/wiki/David_Hanson_(robotics_designer)"},{"title":"Meet Grace, the healthcare robot COVID-19 created（China Daily HK）","url":"https://www.chinadailyhk.com/hk/article/222857"},{"title":"Awakening Health Launches Humanoid Robot Healthcare Assistant Named Grace（Voicebot.ai）","url":"https://voicebot.ai/2020/11/12/awakening-health-launches-humanoid-robot-healthcare-assistant-named-grace/"}],"as_of":"2021-06","related_ids":["sophia","hyper-realistic-humanoid-robot","uncanny-valley","human-robot-interaction","engineered-arts-ameca","companion-robot"],"name":"Hanson Robotics","alt":"汉森机器人","abbr":"","aliases":[],"one_liner":"A Hong Kong maker of realistic-faced humanoid robots, creator of Sophia.","explanation":"Hanson Robotics was founded by American robot designer David Hanson in Dallas (sources give both 2003 and 2007 as the founding year), relocating to Hong Kong Science Park in 2013, where it remains headquartered. Hanson studied film animation as an undergraduate, earned his PhD at the University of Texas at Dallas, and had worked as a sculptor and materials researcher at Disney Imagineering. The company's specialty is realistic faces: the head is covered in a proprietary elastic skin-like material called Frubber, with motors underneath pulling it into expressions. Its best-known creation is Sophia, unveiled in 2016, along with Albert HUBO, an Einstein-faced robot built with KAIST's HUBO humanoid. It also formed a joint venture, Awakening Health, with Singularity Studio, a spinoff of SingularityNET, releasing the Sophia-platform care robot Grace in November 2020. Its robots emphasize expression and conversation demos, with very limited walking and manipulation ability, and the company has often been criticized for overstating its AI capabilities.","example":"The care robot Grace has a thermal-imaging camera in its chest to take an elderly patient's temperature, and can converse in English, Mandarin, and Cantonese.","related":["Sophia (Hanson Robotics)","Hyper-Realistic Humanoid Robot","Uncanny Valley","Human-Robot Interaction","Engineered Arts Ameca","Companion Robot"]},{"id":"cloudminds","category":"company","sec":11,"tier":3,"sources":[{"title":"Wikipedia: CloudMinds","url":"https://en.wikipedia.org/wiki/CloudMinds"},{"title":"维基百科：达闼科技","url":"https://zh.wikipedia.org/wiki/达闼科技"},{"title":"维基百科：黄晓庆","url":"https://zh.wikipedia.org/wiki/黄晓庆"}],"as_of":"2020-07","related_ids":[null,null,null,null,null],"name":"CloudMinds","alt":"达闼科技","abbr":"","aliases":["CloudMinds Technology"],"one_liner":"An early Chinese proponent of “cloud robots,” which puts a robot's brain in the cloud.","explanation":"CloudMinds was founded in March 2015 by Xiaoqing Huang (also known as Bill Huang), formerly president of the China Mobile Research Institute and senior vice president and CTO of UTStarcom; the company is dual-headquartered in Beijing and California. Its core idea is a “cloud brain”: the robot body only handles sensing and physical execution, while recognition and decision-making run in the cloud, connected over a secure private network, with a human able to step in remotely when the AI is unsure. Its investors included SoftBank and Foxconn. In 2019 it filed with the SEC for a New York Stock Exchange IPO, was added to the US Entity List in May 2020, and withdrew its IPO filing that July, reportedly losing about three-quarters of its orders as a result. Before China's current embodied-AI boom, it was one of the representative companies in cloud robotics and service robots.","example":"","related":["Cloud-Edge-Device Collaboration","Service Robot","Remote Teleoperation Takeover (Human Fallback)","Humanoid Robot","SoftBank Group"]},{"id":"k-scale-labs","category":"company","sec":11,"tier":3,"sources":[{"title":"Humanoids Daily: K-Scale Labs Cancels K-Bot Orders, Open-Sources All IP","url":"https://www.humanoidsdaily.com/news/k-scale-labs-cancels-k-bot-orders-open-sources-all-ip-after-funding-fails"},{"title":"Mike Kalil: K-Scale Labs Disrupts Silicon Valley with Open-Source Humanoids","url":"https://mikekalil.com/blog/kscale-labs"},{"title":"The Robot Report: 6 lessons I learned watching a robotics startup die","url":"https://www.therobotreport.com/6-lessons-learned-watching-a-robotics-startup-die-from-the-inside"}],"as_of":"2025-11","related_ids":["humanoid-robot","open-source-hardware","small-size-humanoid-robot","mujoco","jax","mass-production"],"name":"K-Scale Labs","alt":"K-Scale Labs","abbr":"","aliases":["K-Scale"],"one_liner":"A US startup building affordable, open-source humanoid robots, which shut down in November 2025.","explanation":"K-Scale Labs was founded in Palo Alto, California in 2024 by Benjamin Bolte, backed by Y Combinator; Bolte had previously worked at Meta FAIR and on Tesla's Autopilot team. The company set out to build open-source humanoid robots that ordinary developers could afford and modify themselves: the roughly 1.4-meter K-Bot was priced at about $9,000 for developers, the 46-centimeter desktop humanoid Z-Bot was about $999, and it open-sourced the MuJoCo-and-JAX-based reinforcement-learning training library ksim and the inference-export tool kinfer. On November 4, 2025, Bolte wrote to customers saying fundraising had fallen through, canceled all K-Bot preorders with refunds, laid off most of the staff, and released all its designs under an open-source license. It is often cited in discussions of the open-source-hardware approach and the funding barrier to mass production.","example":"K-Bot's parts were designed to be manufactured on a 3D printer with a 256×256 mm print bed, with a reported materials cost under $10,000.","related":["Humanoid Robot","Open-Source Hardware (OSHW)","Small-size Humanoid Robot","MuJoCo (Multi-Joint dynamics with Contact)","JAX","Mass Production"]},{"id":"defense-advanced-research-projects-agency","category":"company","sec":11,"tier":3,"sources":[{"title":"Wikipedia: DARPA Robotics Challenge","url":"https://en.wikipedia.org/wiki/DARPA_Robotics_Challenge"},{"title":"NASA JPL: NASA Robots Compete in DARPA's Subterranean Challenge Final","url":"https://www.jpl.nasa.gov/news/nasa-robots-compete-in-darpas-subterranean-challenge-final"},{"title":"Open Robotics: SubT Part 1 Introduction","url":"https://www.openrobotics.org/blog/2022/2/3/open-robotics-and-the-darpa-subterranean-challenge"}],"as_of":"2026-09","related_ids":["darpa-robotics-challenge","boston-dynamics-atlas","florida-institute-for-human-and-machine-cognition","autonomous-driving","boston-dynamics","anybotics-anymal"],"name":"Defense Advanced Research Projects Agency","alt":"DARPA（美国国防高级研究计划局）","abbr":"DARPA","aliases":["DARPA","ARPA"],"one_liner":"The US Department of Defense's frontier-research funding agency, which has run several challenges that pushed robotics and self-driving cars forward.","explanation":"DARPA is a research-funding agency under the US Department of Defense, founded in 1958 under the name ARPA and headquartered in Arlington, Virginia. It does not do research itself; instead, it poses problems and funds universities and companies to solve them — the internet's precursor, ARPANET, came out of this model. In robotics, it is best known for a series of prize challenges: the 2004-2007 Grand Challenges for self-driving cars directly helped give rise to today's autonomous-driving industry; the 2012-2015 DARPA Robotics Challenge (DRC) required humanoid robots to drive a vehicle, open doors, and turn valves at a simulated disaster site, and DARPA funded Boston Dynamics to build the hydraulic Atlas for competing teams to use; the 2018-2021 Subterranean Challenge (SubT) tested multi-robot autonomous exploration in GPS-denied tunnels and caves. Its Triage Challenge, on robots assisting with casualty triage, reportedly has its final round scheduled for November 2026.","example":"At the 2015 DRC final, South Korea's KAIST won with DRC-HUBO, IHMC took second place piloting an Atlas, and widespread footage of humanoid robots falling over showed the public the limits of the technology at the time.","related":["DARPA Robotics Challenge","Boston Dynamics Atlas (Hydraulic)","Florida Institute for Human and Machine Cognition","Autonomous Driving","Boston Dynamics","ANYbotics ANYmal"]},{"id":"national-ai-industry-investment-fund","category":"company","sec":11,"tier":3,"sources":[{"title":"600亿国家人工智能基金将开展投资布局（证券时报）","url":"https://stcn.com/article/detail/1653447.html"},{"title":"600亿，国家级AI基金登场（36氪）","url":"https://m.36kr.com/p/3249409418879235"}],"as_of":"2025-04","related_ids":["patient-capital","ai-plus-initiative","new-quality-productive-forces","funding-rounds-and-valuation","embodied-ai-bubble"],"name":"National AI Industry Investment Fund","alt":"国家人工智能产业投资基金","abbr":"","aliases":["National AI Fund"],"one_liner":"A RMB 60.06 billion state-level AI industry fund set up by China's industry and finance ministries.","explanation":"The National AI Industry Investment Fund was established on January 17, 2025, led by China's Ministry of Industry and Information Technology and Ministry of Finance, with total committed capital of RMB 60.06 billion and a 13-year term, registered in Xuhui, Shanghai. According to business-registry filings, its main backers are Phase III of the National Integrated Circuit Industry Investment Fund (commonly called the “Big Fund” Phase III) and Guozhitou (Shanghai) Private Fund Management Co., Ltd. It makes equity investments across the whole AI industry chain — compute, algorithms, data, and enabling applications — following a stated principle of investing moderately early, small, and at the frontier. In April 2025, representatives of the fund said at a Shenzhen Stock Exchange forum on embodied AI that they took the field very seriously and would invest as the industry develops. It is a representative example of state-backed “patient capital” entering the embodied-AI race.","example":"At an April 2025 Shenzhen Stock Exchange forum on embodied-AI industrialization, the fund's preparatory team said it would invest in the field.","related":["Patient Capital","AI+ Initiative","New Quality Productive Forces","Funding Rounds & Valuation","Embodied AI Bubble"]},{"id":"ieee-robotics-and-automation-society","category":"company","sec":11,"tier":3,"sources":[{"title":"About RAS - IEEE Robotics and Automation Society","url":"https://www.ieee-ras.org/about-ras"}],"as_of":"","related_ids":["ieee-international-conference-on-robotics-and-automation","ieee-rsj-international-conference-on-intelligent-robots-and","ieee-t-ro-ijrr-ra-l","ieee-ras-international-conference-on-humanoid-robots","top-tier-conferences-and-journals"],"name":"IEEE Robotics and Automation Society","alt":"IEEE 机器人与自动化学会","abbr":"IEEE RAS","aliases":["RAS","IEEE RAS"],"one_liner":"The IEEE's robotics academic society, organizer of ICRA and other conferences and journals like T-RO.","explanation":"The IEEE Robotics and Automation Society (IEEE RAS) is the professional society within the Institute of Electrical and Electronics Engineers (IEEE) devoted to robotics and automation. It grew out of the IEEE Robotics and Automation Council, founded in 1984, which became a full society in 1987. It organizes ICRA, robotics' largest conference (first held in Atlanta in 1984), co-organizes IROS, and also runs conferences including CASE and Humanoids; its journals include IEEE T-RO, RA-L, T-ASE, and RA Magazine. Most of the IEEE robotics conferences and journals a newcomer encounters while submitting papers or searching the literature are organized under it.","example":"Many embodied-AI papers are first submitted to the journal RA-L, then, once accepted, presented at either ICRA or IROS.","related":["IEEE International Conference on Robotics and Automation","IEEE/RSJ International Conference on Intelligent Robots and Systems","IEEE T-RO / IJRR / RA-L","IEEE-RAS International Conference on Humanoid Robots","Top-tier Conferences & Journals (CCF-A Venues)"]},{"id":"international-federation-of-robotics","category":"company","sec":11,"tier":3,"sources":[{"title":"IFR International Federation of Robotics","url":"https://ifr.org/"},{"title":"IFR Press Releases","url":"https://ifr.org/ifr-press-releases"}],"as_of":"2026-09","related_ids":["robot-density","industrial-robot","service-robot","big-four-of-industrial-robotics","ieee-robotics-and-automation-society"],"name":"International Federation of Robotics","alt":"国际机器人联合会","abbr":"IFR","aliases":["IFR"],"one_liner":"The global robotics industry body that publishes the annual “World Robotics” statistical report.","explanation":"The International Federation of Robotics (IFR) was founded in 1987 and is headquartered in Frankfurt, Germany, with members including national robotics associations, research institutions, and robot manufacturers. It is best known for its annual “World Robotics” report, which tracks yearly installations and the operating stock of industrial and service robots worldwide, along with robot density (the number of robots per 10,000 manufacturing workers). Most of the global or national installation figures cited in media coverage and research reports trace back to this report. “World Robotics 2026,” released on September 24, 2026, put the number of industrial robots operating in factories worldwide at about 5 million, with more than 600,000 newly installed in 2025, up 11% year over year.","example":"When a research report ranks countries by robot density, the source is usually IFR's annual statistics.","related":["Robot Density","Industrial Robot","Service Robot","Big Four of Industrial Robotics","IEEE Robotics and Automation Society"]},{"id":"upstream-midstream-downstream-of-the-industry-chain","category":"industry","sec":0,"tier":1,"sources":[{"title":"新华网：一文了解人形机器人产业链（2025-11-21）","url":"http://www.news.cn/finance/20251121/1a2e6771a8154e358805eb857ce4b8b7/c.html"},{"title":"Supply chain - Wikipedia","url":"https://en.wikipedia.org/wiki/Supply_chain"}],"as_of":"","related_ids":["core-components","robot-body-maker","scenario-owner","system-integrator","domestic-substitution","per-unit-content-value"],"name":"Upstream / Midstream / Downstream of the Industry Chain","alt":"产业链上中下游","abbr":"","aliases":["Upstream Components / Midstream Bodies / Downstream Applications"],"one_liner":"A way of splitting an industry into components, complete products, and end applications, following the order of production.","explanation":"In embodied AI and humanoid robotics, “upstream” refers to core components and foundational hardware and software — reducers, motors, screws, sensors, chips, dexterous hands; “midstream” refers to body makers, which integrate those components into a complete robot along with control software and models; and “downstream” refers to application scenarios and customers, such as auto factories, logistics, commercial services, or research and education. Brokerage research reports, policy documents, and investment news commonly use this three-part split to analyze where the profit sits and where the technology is bottlenecked, or “choked.” For a newcomer, this framework is a quick way to place a company: a joint-module maker sits upstream, a complete-robot maker sits midstream, and a company doing scenario integration and operations for factory customers sits downstream.","example":"Leaderdrive makes harmonic reducers and sits upstream; UBTech builds complete Walker S-series humanoid robots and sits midstream; an automaker that buys humanoid robots for its factory floor sits downstream.","related":["Core Components","Robot Body Maker","Scenario Owner (End Customer)","System Integrator","Domestic Substitution","Per-Unit Content Value"]},{"id":"robot-body-maker","category":"industry","sec":0,"tier":1,"sources":[{"title":"Original equipment manufacturer - Wikipedia","url":"https://en.wikipedia.org/wiki/Original_equipment_manufacturer"},{"title":"Unitree Robotics 官网","url":"https://www.unitree.com/"}],"as_of":"","related_ids":["embodiment","robot-brain-company","hardware-software-integration","upstream-midstream-downstream-of-the-industry-chain","unitree-robotics","core-components"],"name":"Robot Body Maker","alt":"本体厂商","abbr":"","aliases":["Robot OEM","Body Maker"],"one_liner":"A company that designs and manufactures a robot's complete physical hardware, such as a humanoid or quadruped maker.","explanation":"“Embodiment” refers to a robot's physical body — its mechanical structure, joint motors, sensors, and onboard compute — and a “body maker” is a company that builds and sells the complete robot, often called a robot OEM in English (note that in Chinese, “OEM” usually implies contract manufacturing under someone else's brand, a different meaning). In the industry's supply chain, body makers sit in the middle, with upstream suppliers of reducers, motors, sensors, and other components below them and scenario owners and system integrators above. Unitree, UBTech, AgiBot, and Figure are all body makers in this sense. The contrasting category is a “brain company,” which builds only models and no hardware; some companies do both hardware and software in-house. When reading industry news or picking a research platform, it often helps to first work out whether a company is a body maker or a brain company.","example":"Unitree Robotics sells complete robots such as the Go2 quadruped and the G1 humanoid, making it a classic robot body maker.","related":["Embodiment","Robot-Brain (Model-Only) Company","Hardware-Software Integration","Upstream / Midstream / Downstream of the Industry Chain","Unitree Robotics","Core Components"]},{"id":"robot-brain-company","category":"industry","sec":0,"tier":2,"sources":[{"title":"Physical Intelligence 官网","url":"https://www.physicalintelligence.company/"}],"as_of":"","related_ids":["robot-body-maker","hardware-software-integration","hardware-software-decoupling","one-brain-multiple-robots","physical-intelligence","skild-ai"],"name":"Robot-Brain (Model-Only) Company","alt":"大脑公司","abbr":"","aliases":["“Brain” Company","Embodied-Model Company"],"one_liner":"A company that builds general-purpose robot models but doesn't manufacture the robot hardware itself.","explanation":"This refers to companies that focus on developing embodied foundation models — such as VLA or world models — while largely not mass-producing robot hardware themselves, as distinct from “embodiment makers,” who build the physical robots. Their idea is to build a general-purpose “brain” that can be installed on different robots, using cross-embodiment generalization to cover many kinds of hardware; Physical Intelligence (the π0 series) and Skild AI are representative examples. Supporters argue the model is the real moat and hardware will gradually standardize; skeptics argue that without owning the hardware, it's hard to get enough real-robot data or to co-optimize software and hardware together. This is part of the broader “hardware–software integration vs. decoupling” debate in the embodied-AI industry.","example":"Physical Intelligence trains π0 on data collected from robot arms made by several different manufacturers, without selling any robots of its own.","related":["Robot Body Maker","Hardware-Software Integration","Hardware-Software Decoupling","One Brain, Multiple Robots","Physical Intelligence","Skild AI"]},{"id":"three-pillars-of-embodied-ai-data-model-embodiment","category":"industry","sec":0,"tier":2,"sources":[{"title":"Open X-Embodiment: Robotic Learning Datasets and RT-X Models","url":"https://robotics-transformer-x.github.io/"}],"as_of":"","related_ids":["embodiment","vision-language-action-model","data-scarcity","robot-body-maker","robot-brain-company","embodied-ai-data-service-provider"],"name":"Three Pillars of Embodied AI: Data, Model, Embodiment","alt":"具身智能三要素（数据 / 模型 / 本体）","abbr":"","aliases":["Data / Model / Embodiment Framework"],"one_liner":"An industry shorthand for breaking embodied AI down into three parts: data, model, and embodiment.","explanation":"This is a breakdown commonly used in industry reports and talks, not a formal framework from any single paper. Data refers to the demonstrations, simulations, and human videos used to train a robot; model refers to the policy that turns observations into actions, such as a VLA model; embodiment refers to the robot hardware itself — its joints, dexterous hands, and sensors. The three constrain each other: the embodiment determines what data can be collected and what actions are possible, the amount of data determines how well the model can learn, and the model's capability in turn determines how valuable the embodiment is. Companies also tend to position themselves along these same three lines, as embodiment makers, “brain” companies, or data service providers.","example":"","related":["Embodiment","Vision-Language-Action Model","Data Scarcity","Robot Body Maker","Robot-Brain (Model-Only) Company","Embodied AI Data Service Provider"]},{"id":"embodied-ai-data-service-provider","category":"industry","sec":0,"tier":2,"sources":[{"title":"数据决定上限：25家国内具身智能数据采集厂商盘点（艾邦机器人）","url":"https://www.aibangbots.com/a/11921"},{"title":"钛媒体：机器人还没学会做家务，卖数据的已经先赚到了钱","url":"https://www.tmtpost.com/8062934.html"},{"title":"Wikipedia: Scale AI","url":"https://en.wikipedia.org/wiki/Scale_AI"}],"as_of":"2026-09","related_ids":["data-collector","embodied-ai-training-ground","data-annotation","data-quality-control","scale-ai","valid-data"],"name":"Embodied AI Data Service Provider","alt":"具身数据服务商（数据采集服务商）","abbr":"","aliases":["Data Collection Service Provider"],"one_liner":"A company whose main business is collecting, labeling, and selling training data for robot companies.","explanation":"This describes a company whose main business isn't selling robots or models, but supplying training data to embodied-AI companies. Its work includes building data-collection sites, hiring data collectors to gather demonstrations using teleoperation, motion capture, data gloves, or handheld grippers; cleaning the data, breaking it into subtasks, adding language labels, and running quality checks; and delivering by the hour or by the clip, or selling ready-made datasets or collection equipment outright. This role exists because imitation learning needs large volumes of real-robot data, and building an in-house collection team is expensive and slow for a robotics company. It's the robotics-era counterpart to a labeling company like Scale AI in the large-language-model era. When evaluating a provider, the main things to check are its data format, whether its collection hardware matches your own robot, and what share of the data is actually usable.","example":"A robotics company that wants a clothes-folding skill outsources it to a data provider, which uses a matching robot arm to teleoperate several hundred hours of demonstrations at its collection site, labels the subtasks, and delivers the data in LeRobot format.","related":["Data Collector (Teleoperator)","Embodied AI Training Ground (Robot Data Collection Center)","Data Annotation","Data Quality Control","Scale AI","Valid (Usable) Data"]},{"id":"selling-shovels","category":"industry","sec":0,"tier":2,"sources":[{"title":"Investopedia: Pick-and-Shovel Play","url":"https://www.investopedia.com/terms/p/pick-and-shovel-play.asp"}],"as_of":"","related_ids":["core-components","embodied-ai-data-service-provider",null,"upstream-midstream-downstream-of-the-industry-chain","tesla-supply-chain"],"name":"Selling Shovels","alt":"卖铲子","abbr":"","aliases":["Picks-and-Shovels Play","Selling Water (Gold-Rush Supplier)"],"one_liner":"Not betting on who builds the best robot, but supplying everyone who's trying to build one.","explanation":"The phrase comes from the 19th-century American gold rush: most prospectors didn't strike it rich, but the merchants selling shovels and water made steady money regardless. Applied to embodied AI, it describes businesses that don't build complete robots or general-purpose models themselves, but instead supply the whole industry with what it needs — core components (reducers, lead screws, dexterous hands, sensors), simulation and training platforms, data-collection services, or compute chips. Whichever integrator eventually wins, these suppliers benefit, spreading their risk more broadly. NVIDIA is often cited as the archetypal example: it supplies GPUs, Jetson controllers, and the Isaac simulation platform to nearly every robotics company.","example":"A data-collection service provider supplies teleoperated data to several embodied-AI companies without training a general-purpose model of its own.","related":["Core Components","Embodied AI Data Service Provider","NVIDIA","Upstream / Midstream / Downstream of the Industry Chain","Tesla (Optimus) Supply Chain"]},{"id":"nvidia-three-computer-solution","category":"industry","sec":0,"tier":3,"sources":[{"title":"Physical AI Accelerated by Three NVIDIA Computers for Robot Training, Simulation and Inference (NVIDIA Blog)","url":"https://blogs.nvidia.com/blog/three-computers-robotics/"}],"as_of":"2025-08","related_ids":["nvidia","physical-ai","nvidia-jetson-thor","nvidia-isaac-sim","nvidia-cosmos","selling-shovels"],"name":"NVIDIA Three-Computer Solution","alt":"英伟达三台计算机","abbr":"","aliases":["Train / Simulate / Deploy Framework"],"one_liner":"NVIDIA's framework splitting robot development into three “computers”: for training, simulation, and deployment.","explanation":"“Three computers” is the framework NVIDIA uses to promote physical AI, first laid out in an official blog post in October 2024 and updated in August 2025. The first computer is DGX, used to train robot foundation models — for example, post-training GR00T or Cosmos. The second is Omniverse and Cosmos running on RTX PRO servers, paired with Isaac Sim and Isaac Lab, used for simulation, generating synthetic data, and testing policies. The third is a Jetson AGX Thor installed on the robot itself, handling real-time inference and control. The framework strings together “train, simulate, deploy” into one chain, with each link matched to NVIDIA hardware and software; understanding it clarifies NVIDIA's position in embodied AI — it doesn't build complete robots, but supplies the toolchain that everyone else uses.","example":"A humanoid-robotics team trains a policy on DGX, runs reinforcement learning and testing in Isaac Lab, and then deploys the model to the robot's onboard Jetson Thor — exactly this division of labor.","related":["NVIDIA","Physical AI","NVIDIA Jetson Thor","NVIDIA Isaac Sim","NVIDIA Cosmos","Selling Shovels"]},{"id":"big-four-of-industrial-robotics","category":"industry","sec":0,"tier":2,"sources":[{"title":"Industrial robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Industrial_robot"},{"title":"KUKA - Wikipedia","url":"https://en.wikipedia.org/wiki/KUKA"}],"as_of":"2025-10","related_ids":["industrial-robot","fanuc","abb-robotics","kuka","yaskawa-electric-corporation","domestic-substitution"],"name":"Big Four of Industrial Robotics","alt":"工业机器人四大家族","abbr":"","aliases":["FANUC, ABB, KUKA, Yaskawa"],"one_liner":"The industry's collective name for four long-established industrial-robot giants: FANUC, ABB, KUKA, and Yaskawa.","explanation":"This is the industry's collective term for four established makers of industrial robots: Japan's FANUC, Switzerland's ABB, Germany's KUKA (majority-owned by China's Midea Group since 2017), and Japan's Yaskawa Electric. They have long held the bulk of the market for multi-joint industrial arms used in car welding, handling, and painting lines, and control the core technology in controllers, servo drives, and reducers. When discussing China's industrial-robot sector, the “Big Four” is commonly used as the benchmark that domestic makers are trying to catch up to and displace through domestic substitution. According to reports, ABB announced in 2025 that it would sell its robotics business to SoftBank Group.","example":"The rows of orange KUKA and yellow FANUC 6-axis arms on an automotive body shop's line are a textbook example of the Big Four at work.","related":["Industrial Robot","FANUC","ABB Robotics","KUKA","Yaskawa Electric Corporation","Domestic Substitution"]},{"id":"system-integrator","category":"industry","sec":0,"tier":2,"sources":[{"title":"System integrator - Wikipedia","url":"https://en.wikipedia.org/wiki/System_integrator"}],"as_of":"","related_ids":["upstream-midstream-downstream-of-the-industry-chain","non-standard-automation","industrial-robot","scenario-owner","to-business","project-based-delivery-vs-productization"],"name":"System Integrator","alt":"系统集成商","abbr":"SI","aliases":["SI"],"one_liner":"A company that assembles robots, tooling, and software into a complete production-line solution for a customer.","explanation":"A system integrator sits downstream in the robotics supply chain. It typically doesn't manufacture robots itself; instead it buys hardware — robot arms, AGVs (automated guided vehicles), and so on — and adds fixtures, sensors, a PLC (programmable logic controller), and control software, designing, installing, and commissioning it all into a working line tailored to a customer factory's specific process. Deploying industrial robots almost always goes through an integrator, because every factory's parts and cycle time differ. When embodied-AI companies discuss enterprise (ToB) deployment, they often have to decide whether to build integration capability in-house or partner with an integrator.","example":"On an auto-body welding line, an integrator combines multiple industrial robots, welding guns, fixtures, and a PLC into a complete automated welding line delivered to the automaker.","related":["Upstream / Midstream / Downstream of the Industry Chain","Non-Standard (Custom) Automation","Industrial Robot","Scenario Owner (End Customer)","To Business (B2B)","Project-Based Delivery vs. Productization"]},{"id":"scenario-owner","category":"industry","sec":0,"tier":3,"sources":[{"title":"End user - Wikipedia","url":"https://en.wikipedia.org/wiki/End_user"}],"as_of":"","related_ids":["robot-body-maker","robot-brain-company","system-integrator","real-world-deployment","return-on-investment-payback-period","factory-pilot-deployment"],"name":"Scenario Owner (End Customer)","alt":"场景方","abbr":"","aliases":[],"one_liner":"The party with a real-world use case who actually uses the robot and pays for it — the factory, warehouse, or store.","explanation":"“Scenario owner” is how the embodied-AI industry refers to the end user: a company or institution that has a real working environment, such as a car factory, logistics warehouse, retail store, or hospital. The supply chain also typically includes embodiment makers (who build the robots), “brain” companies (who build the models), and system integrators (who wire the equipment into the line), but it's the scenario owner who decides what job the robot needs to do and how much efficiency it needs to hit before they'll pay for it. Whether a robot can move from demo to real deployment largely depends on whether the scenario owner is willing to open up its site, share data, and run the return-on-investment math.","example":"An automaker, acting as the scenario owner, lets a humanoid robot run a pilot moving parts bins in its final-assembly workshop.","related":["Robot Body Maker","Robot-Brain (Model-Only) Company","System Integrator","Real-world Deployment","Return on Investment / Payback Period","Factory Pilot Deployment"]},{"id":"automakers-entering-humanoid-robotics","category":"industry","sec":0,"tier":2,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"title":"Boston Dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Boston_Dynamics"}],"as_of":"2025-12","related_ids":["tesla-optimus","xpeng-iron","hyundai-motor-group","autonomous-driving-talent-moving-into-embodied-ai","tesla-supply-chain","factory-pilot-deployment"],"name":"Automakers Entering Humanoid Robotics","alt":"车企造人形（车企入局）","abbr":"","aliases":[],"one_liner":"Car companies using their supply chains and self-driving know-how to build humanoid robots.","explanation":"This describes automakers developing or investing in humanoid robots in-house. Examples include Tesla's Optimus, announced in 2021; XPeng's IRON, unveiled in 2024; Xiaomi's CyberOne from 2022; and Hyundai Motor Group, which took a controlling stake in Boston Dynamics in 2021. The logic behind automakers entering the field is that technology and supply chains carry over: motors, batteries, sensors, and chips are already things they procure, the perception and end-to-end modeling experience from self-driving teams transfers to robotics, and the mass-production management skills built for car plants apply too; at the same time, their own factories double as a ready-made testing ground, known as in-factory training. This trend is often discussed alongside the “T-chain” and the broader shift of self-driving talent into embodied AI.","example":"XPeng unveiled its humanoid robot IRON in 2024, saying it shares an AI chip and some autonomous-driving capability with its smart cars.","related":["Tesla Optimus","XPeng IRON","Hyundai Motor Group","Autonomous-Driving Talent Moving into Embodied AI","Tesla (Optimus) Supply Chain","Factory Pilot Deployment"]},{"id":"autonomous-driving-talent-moving-into-embodied-ai","category":"industry","sec":0,"tier":3,"sources":[{"title":"它石智航融资1.2亿美元，创下今年最大天使轮纪录（界面新闻，2025-03-26）","url":"https://www.jiemian.com/article/12523148.html"}],"as_of":"2025-03","related_ids":["university-big-tech-autonomous-driving-founder-lineage","autonomous-driving","end-to-end","data-flywheel","tars-robotics","automakers-entering-humanoid-robotics"],"name":"Autonomous-Driving Talent Moving into Embodied AI","alt":"智驾转具身","abbr":"","aliases":["AD-to-Robotics Talent Migration"],"one_liner":"The trend, since 2023, of autonomous-driving engineers and founders switching into embodied-AI robotics.","explanation":"This refers to the phenomenon, starting around 2023, of a wave of autonomous-driving (“smart driving”) professionals moving into robotics — including technical leaders at automakers and autonomous-driving companies leaving to found startups, as well as large numbers of perception, planning, and data engineers changing jobs. The reason is that the two technology stacks overlap heavily: multi-sensor perception, end-to-end models, the data flywheel (collecting more data after deployment to retrain the model), simulation testing, and automotive-grade mass-production experience all transfer; at the same time, competition in autonomous driving has been narrowing while embodied-AI fundraising has been hot. The industry accordingly sorts founding teams into three lineages: academic, big-tech, and autonomous-driving. Teams with an autonomous-driving background typically bring stronger engineering and mass-production discipline, but still have to learn arm manipulation, contact, and force control — problems that don't come up in cars.","example":"TARS Robotics's CEO, Yilun Chen, previously served as CTO of Huawei's autonomous-driving unit, and its chairman, Zhenyu Li, previously headed Baidu's Intelligent Driving Group; the two founded their embodied-AI company in 2025.","related":["University / Big-Tech / Autonomous-Driving Founder Lineage","Autonomous Driving","End-to-End","Data Flywheel","TARS Robotics","Automakers Entering Humanoid Robotics"]},{"id":"university-big-tech-autonomous-driving-founder-lineage","category":"industry","sec":0,"tier":3,"sources":[{"title":"36氪（具身智能创业报道）","url":"https://www.36kr.com/"}],"as_of":"","related_ids":["autonomous-driving-talent-moving-into-embodied-ai","robot-body-maker","robot-brain-company","funding-rounds-and-valuation","galbot","agibot"],"name":"University / Big-Tech / Autonomous-Driving Founder Lineage","alt":"高校系 / 大厂系 / 智驾系","abbr":"","aliases":["Academic / Big-Tech / AD Lineage"],"one_liner":"An informal three-way grouping of embodied-AI startups by their founders' professional background.","explanation":"This is an informal classification used by Chinese investors and media to group embodied-AI startups by where their founders came from. The “academic” lineage is founded by university professors or PhD teams, with technical strength built on papers and algorithms; the “big-tech” lineage comes from internet or tech giants such as Huawei, ByteDance, or Alibaba, and tends to be strong in engineering and productization; the “autonomous-driving” lineage comes from self-driving companies or automakers' autonomous-driving divisions, bringing experience with data flywheels, mass production, and automotive-grade engineering into robotics. The grouping helps in reading a company's technical approach and likely weaknesses — an academic-lineage company is often questioned on mass-production ability, while big-tech- and autonomous-driving-lineage companies are often questioned on how much real embodiment experience they have. It's a rough grouping, and many teams are a mix of more than one lineage.","example":"As reported, Galbot founder He Wang is an assistant professor at Peking University and is often classified as academic lineage; AgiBot co-founder Zhihui Peng came from Huawei and is often classified as big-tech lineage.","related":["Autonomous-Driving Talent Moving into Embodied AI","Robot Body Maker","Robot-Brain (Model-Only) Company","Funding Rounds & Valuation","Galbot","AgiBot"]},{"id":"technical-route-debate","category":"industry","sec":1,"tier":2,"sources":[{"title":"Vision-language-action model - Wikipedia","url":"https://en.wikipedia.org/wiki/Vision-language-action_model"}],"as_of":"","related_ids":["real-robot-data-camp-vs-sim-data-camp","form-factor-debate","end-to-end","braincerebellum-architecture","world-model","consensus-non-consensus"],"name":"Technical Route Debate","alt":"路线之争","abbr":"","aliases":["Technology Roadmap Debate"],"one_liner":"The industry's unresolved disagreement over which technical approach embodied AI should take.","explanation":"Embodied AI hasn't converged on an agreed-upon approach the way large language models have, and the industry often labels several competing pairs of options a “route debate”: end-to-end VLA (vision-language-action) models versus a layered brain–cerebellum architecture; real-robot data versus simulation or human-video data as the primary training source; VLA versus world models; and humanoid versus wheeled or other non-humanoid embodiments. Understanding these divides helps in reading how different companies make technical trade-offs and frame their marketing, which often directly reflects the bet they're making.","example":"Some companies advocate collecting large volumes of teleoperated real-robot data, while others advocate pretraining on synthetic data generated in simulation — this is the data route debate.","related":["Real-Robot-Data Camp vs. Sim-Data Camp","Form-Factor Debate","End-to-End","Brain–Cerebellum Architecture","World Model","Consensus / Non-Consensus"]},{"id":"form-factor-debate","category":"industry","sec":1,"tier":2,"sources":[{"title":"Wikipedia: Humanoid robot","url":"https://en.wikipedia.org/wiki/Humanoid_robot"},{"title":"逐际动力 TRON 2 具身机器人发布：可变化三种形态（IT之家）","url":"https://www.ithome.com/0/906/067.htm"}],"as_of":"2026-09","related_ids":["humanoid-robot","wheeled-humanoid-robot","wheel-legged-robot","legged-robot","technical-route-debate","embodiment"],"name":"Form-Factor Debate","alt":"形态之争","abbr":"","aliases":["Humanoid vs. Non-humanoid","Wheeled vs. Legged"],"one_liner":"The ongoing debate over what shape a general-purpose robot should take, and whether it should move on wheels or legs.","explanation":"This is a long-running debate in embodied AI over what a robot's body should look like, made up of two separate arguments. The first is humanoid versus non-humanoid: those in favor of humanoid form argue that human environments, tools, and data — human videos, motion capture — are all built around the human body, so a humanoid shape can reuse them directly; opponents argue that a biped is unstable and expensive, and that an arm plus a mobile base is enough for most tasks. The second is wheeled versus legged: wheels are faster, more energy-efficient, and more stable on flat ground, while legs can climb stairs and cross obstacles. Compromise designs include wheeled humanoid robots, with a wheeled lower body and dual arms up top, and wheel-legged robots. Underneath the debate are trade-offs in cost, reliability, and target use case, with no settled answer yet, and many companies hedge by releasing more than one form factor at once.","example":"LimX Dynamics' TRON 2 can switch between dual-arm, legged, and wheel-legged configurations, an example of a maker hedging its bets on form factor.","related":["Humanoid Robot","Wheeled Humanoid Robot","Wheel-legged Robot","Legged Robot","Technical Route Debate","Embodiment"]},{"id":"real-robot-data-camp-vs-sim-data-camp","category":"industry","sec":1,"tier":2,"sources":[{"title":"AgiBot World (GitHub)","url":"https://github.com/OpenDriveLab/AgiBot-World"},{"title":"NVIDIA Isaac Lab","url":"https://developer.nvidia.com/isaac/lab"}],"as_of":"2025","related_ids":[null,null,null,null,null,"technical-route-debate"],"name":"Real-Robot-Data Camp vs. Sim-Data Camp","alt":"真机派 / 仿真派","abbr":"","aliases":["Real-Data Camp / Sim-Data Camp","Real-World vs. Simulation Data Debate"],"one_liner":"The debate over whether robot training data should mainly come from real robots or from simulation.","explanation":"This is how China's robotics industry refers to its ongoing debate over the data pipeline for embodied AI. The real-data camp argues that the fine details of real physical interaction are too hard for simulation to reproduce faithfully, and favors collecting large volumes of real-robot data via teleoperation — Physical Intelligence and AgiBot's AgiBot World dataset are representative examples. The sim-data camp argues real-robot collection is too slow and expensive, and favors generating synthetic data at scale in simulators using domain randomization and procedural generation, then applying sim-to-real transfer — Galbot's use of synthetic data to train grasping models, and NVIDIA's simulation toolchain, are representative examples. In practice most teams blend both approaches and are increasingly adding human-video data too; the “data pyramid” concept is exactly this idea of layering different data types by how much of each is used.","example":"Galbot's GraspVLA is pretrained mainly on roughly a billion synthetic grasping examples, while AgiBot World is a large-scale dataset collected from real robots at a dedicated data-collection facility.","related":["Real-Robot Data","Simulation Data","Synthetic Data","Sim-to-Real Transfer","Data Pyramid","Technical Route Debate"]},{"id":"full-stack-in-house-development","category":"industry","sec":1,"tier":2,"sources":[{"title":"星动纪元：关于我们","url":"https://www.robotera.com/about/us"},{"title":"北京人形机器人创新中心：关于我们","url":"https://www.x-humanoid.com/about.html"}],"as_of":"2026-09","related_ids":["hardware-software-integration","robot-body-maker","robot-brain-company","joint-actuator-module","hardware-software-decoupling","robotera"],"name":"Full-Stack In-House Development","alt":"全栈自研","abbr":"","aliases":["Vertical Integration"],"one_liner":"Developing the robot body, its core components, control, and the large model all in-house rather than buying them in.","explanation":"This means a robotics company handles nearly every major layer itself — mechanical structure, joint modules, dexterous hands, electronics, motion control, perception, and the upper-level large model — rather than buying someone else's robot body or calling on someone else's model. The upside is that hardware and software can be iterated together, problems are easier to trace, data formats stay consistent, and long-term costs are more controllable; the downside is that it takes heavy investment and a long timeline, and the team has to be capable across the board. The contrasting models are a body maker that builds only hardware, a brain company that builds only models, and an approach that buys in core components and integrates them. The phrase “full-stack in-house” gets used loosely in marketing, so it's worth checking whether the key components and models really were built in-house.","example":"RobotEra's website describes itself as “fully self-developed across hardware and software,” building the humanoid robot L7, the dexterous hand XHAND1, and the end-to-end model ERA-42 all at once.","related":["Hardware-Software Integration","Robot Body Maker","Robot-Brain (Model-Only) Company","Joint Actuator Module","Hardware-Software Decoupling","RobotEra"]},{"id":"hardware-software-integration","category":"industry","sec":1,"tier":2,"sources":[{"title":"新京报：星源智完成 Pre-A 轮融资","url":"https://m.bjnews.com.cn/detail/1780455817129694.html"},{"title":"36氪：无界动力","url":"https://m.36kr.com/p/3869370059035913"}],"as_of":"2026-06","related_ids":["full-stack-in-house-development","hardware-software-decoupling","robot-body-maker","robot-brain-company","on-device-edge-deployment","one-brain-multiple-robots"],"name":"Hardware-Software Integration","alt":"软硬一体","abbr":"","aliases":[],"one_liner":"Hardware and algorithms designed and tuned together by the same party, delivered as one combined solution.","explanation":"This means a robot's hardware — its body, sensors, and compute platform — and its software — control, perception, and models — are designed, tuned, and delivered together as a single package. It differs from full-stack in-house development in emphasis: full-stack in-house stresses that every layer is built internally, while hardware-software integration stresses that the hardware and software are well matched and work together out of the box for the customer, with some of the hardware possibly bought in from outside. The reasoning is that an embodied model's performance depends heavily on the specific robot body it runs on — where the camera sits, how much delay a joint has, how fast the control loop runs — all of which affect a policy's performance, and building hardware and software separately makes them easy to mismatch. The opposite concept is hardware-software decoupling, where a model tries not to be tied to one specific robot body, so a single brain can adapt to many different robots.","example":"Reportedly, Xingyuanzhi Robotics doesn't build its own robot body, but it delivers its embodied-brain model together with an on-device compute platform as one package, describing its approach as hardware-software integration with on-device deployment.","related":["Full-Stack In-House Development","Hardware-Software Decoupling","Robot Body Maker","Robot-Brain (Model-Only) Company","On-Device / Edge Deployment","One Brain, Multiple Robots"]},{"id":"hardware-software-decoupling","category":"industry","sec":1,"tier":3,"sources":[{"title":"π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv)","url":"https://arxiv.org/abs/2410.24164"}],"as_of":"","related_ids":["robot-brain-company","hardware-software-integration","one-brain-multiple-robots","cross-embodiment","embodiment-agnostic"],"name":"Hardware-Software Decoupling","alt":"软硬解耦（模型与本体解耦）","abbr":"","aliases":["Model–Embodiment Decoupling"],"one_liner":"The robot's intelligence isn't tied to one specific hardware platform, and can run on different robots.","explanation":"Hardware-software decoupling means developing a robot's “brain” (the perception and decision-making model and software) separately from its “embodiment” (mechanical structure, motors, sensors), with the model adapting to multiple robots through a unified interface, while the hardware can also be paired with models from different vendors. This is the business premise behind “brain” companies: building only the model, not a complete robot, and training a general-purpose policy on cross-embodiment data. The opposite is hardware-software integration, where one company designs both the hardware and the model together and deeply co-optimizes them. Proponents of decoupling value scale and reuse; skeptics argue control performance can't be separated from deep knowledge of the specific hardware.","example":"Physical Intelligence doesn't build robots itself; its π0 model is trained on data from robots with many different configurations, and can control single-arm, dual-arm, and mobile robots alike.","related":["Robot-Brain (Model-Only) Company","Hardware-Software Integration","One Brain, Multiple Robots","Cross-Embodiment","Embodiment-agnostic"]},{"id":"one-brain-multiple-robots","category":"industry","sec":1,"tier":2,"sources":[{"title":"Skild AI","url":"https://www.skild.ai/"},{"title":"π0: Our First Generalist Policy - Physical Intelligence","url":"https://www.physicalintelligence.company/blog/pi0"}],"as_of":"2025","related_ids":[null,null,null,null,"robot-brain-company","hardware-software-decoupling"],"name":"One Brain, Multiple Robots","alt":"一脑多机","abbr":"","aliases":["One Brain, Multiple Embodiments","One Brain, Multiple Bodies"],"one_liner":"Using a single embodied-AI model to control many different kinds of robot hardware.","explanation":"“One brain, multiple robots” is a phrase used in China's robotics industry for training a single general-purpose robot “brain” model that can drive robot arms, dual-arm rigs, wheeled platforms, humanoids, and quadrupeds alike, instead of training a separate model for each type of machine. Technically, it relies on cross-embodiment training: data from many different robots is pooled for pretraining, and a unified action space or embodiment-specific action heads then handle differences in joint count and control scheme. The appeal is that data can be reused across robots and a model company isn't locked into one hardware platform; the difficulty is that embodiments differ enough that negative transfer — where training on one robot hurts performance on another — is common. Both “brain” companies and the model–hardware decoupling business model rest on this premise.","example":"Skild AI says its Skild Brain can control multiple robot embodiments; Physical Intelligence's π0 is jointly trained on data from several different robots.","related":["Cross-Embodiment","Embodiment Gap","Unified Action Space","Embodiment-Specific Head","Robot-Brain (Model-Only) Company","Hardware-Software Decoupling"]},{"id":"specialized-first-general-later","category":"industry","sec":1,"tier":2,"sources":[{"title":"Generalist robot policy 概念（Octo 项目页）","url":"https://octo-models.github.io/"}],"as_of":"","related_ids":["general-purpose-robot","specialist-policy","real-world-deployment","technical-route-debate","generalist-policy","data-flywheel"],"name":"Specialized First, General Later","alt":"先专后通","abbr":"","aliases":["Narrow-to-General Strategy"],"one_liner":"Making a robot work well in one narrow scenario first, then gradually broadening it toward generality.","explanation":"This is one school of thought in the embodied-AI industry about how to reach real-world deployment: pick one concrete scenario first — factory loading, logistics sorting, retail shelf restocking — and push the robot's success rate up to something profitable there, then use the revenue and real data from that scenario to improve the model and gradually expand to more tasks, eventually working toward a general-purpose robot. The opposing view is “general first, specialized later”: train a general foundation model first, then fine-tune it for specific scenarios. Advocates of specialized-first argue general models won't reach commercial reliability anytime soon; opponents worry that the deeper a specialized solution goes, the harder it becomes to transfer elsewhere. In practice, most companies hedge by pursuing both approaches at once.","example":"A company first limits its wheeled dual-arm robot to picking items off pharmacy shelves, then expands to other retail settings once it has accumulated enough data.","related":["General-purpose Robot","Specialist Policy","Real-world Deployment","Technical Route Debate","Generalist Policy","Data Flywheel"]},{"id":"agibot-g1g5-embodied-ai-roadmap","category":"industry","sec":1,"tier":3,"sources":[{"title":"AgiBot - Wikipedia","url":"https://en.wikipedia.org/wiki/AgiBot"},{"title":"智元机器人官网","url":"https://www.agibot.com/"}],"as_of":"2024-08","related_ids":["agibot","humanoid-robot-intelligence-level-grading","skill-primitive","end-to-end","agibot-go-1","levels-of-autonomy"],"name":"AgiBot G1–G5 Embodied-AI Roadmap","alt":"智元 G1–G5 技术路线","abbr":"","aliases":["G1–G5 Embodied-AI Evolution Path"],"one_liner":"AgiBot's five-level roadmap for how a robot's “brain” evolves from hand-coded rules to a fully end-to-end general model.","explanation":"This is a technical tiering AgiBot proposed at its product launch event in August 2024 to explain how a robot's “brain” becomes more general step by step. As reported, it runs roughly: G1, basic automation, relying on hand-designed rules and simple visual feedback that must be rebuilt for each new setting; G2, general-purpose atomic skills, where reusable primitives like grasping and placing (“atomic skills”) are sequenced into tasks; G3, end-to-end, where atomic skills are instead learned by a data-driven end-to-end model; G4, a general-purpose manipulation foundation model that generalizes across tasks and scenes; and G5, where perception through execution is handled entirely by one unified foundation model. This is one company's own roadmap, not an industry standard, and it's a different thing from the China Institute of Electronics' humanoid-robot intelligence-level grading. AgiBot's later models, such as GO-1, are positioned as products moving toward the higher levels of this roadmap.","example":"","related":["AgiBot","Humanoid Robot Intelligence Level Grading","Skill Primitive","End-to-End","AgiBot GO-1","Levels of Autonomy"]},{"id":"consensus-non-consensus","category":"industry","sec":1,"tier":3,"sources":[{"title":"Contrarian investing - Wikipedia","url":"https://en.wikipedia.org/wiki/Contrarian_investing"}],"as_of":"","related_ids":["technical-route-debate","form-factor-debate","real-robot-data-camp-vs-sim-data-camp","embodied-ai-bubble","funding-rounds-and-valuation"],"name":"Consensus / Non-Consensus","alt":"共识 / 非共识","abbr":"","aliases":[],"one_liner":"A judgment most people already agree with is “consensus”; one only a few hold is “non-consensus.”","explanation":"This is common language in investing and startup circles. Consensus refers to a judgment most people in an industry already accept, such as “data is the bottleneck for embodied AI”; non-consensus refers to a judgment only a minority believes, not yet widely accepted. Investors often say that only a “non-consensus and correct” view can produce outsized returns, because consensus views are already reflected in prices and valuations. Many debates in embodied AI can be described with this pair of terms: humanoid versus non-humanoid, real-robot data versus simulation data, end-to-end versus layered architectures. A given judgment tends to shift from non-consensus to consensus as evidence accumulates, at which point anyone entering later generally can't capture the early-mover returns anymore. When reading financing news and interviews, this term is often used to explain why a founder chose a particular technical route.","example":"A few years ago, “control multiple types of robots with one large model” was non-consensus; after models such as π0 and GR00T were released, it gradually became industry consensus.","related":["Technical Route Debate","Form-Factor Debate","Real-Robot-Data Camp vs. Sim-Data Camp","Embodied AI Bubble","Funding Rounds & Valuation"]},{"id":"demo","category":"industry","sec":2,"tier":1,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"}],"as_of":"","related_ids":["teleoperated-demo","fully-autonomous","playback-speed-label","edited-demo","cherry-picking","one-take-video"],"name":"Demo (Demonstration Video)","alt":"Demo（演示视频）","abbr":"","aliases":["Demo"],"one_liner":"A video companies or teams use to show off a robot's abilities, usually a selected highlight clip.","explanation":"Nearly every embodied-AI company or team accompanies a model or product launch with a video of a robot folding laundry, cooking, or moving boxes — known in the field simply as a demo. It's the most direct material a newcomer has for gauging a company's progress, but it carries limited information: how many attempts it took to succeed, whether it was teleoperated, whether it was sped up, and whether the scene was staged in advance often aren't visible in the footage. When watching a demo, first look for on-screen labels such as “1x” or “autonomous,” then check the paper or technical report for success rates and test conditions. Related jargon includes teleoperated demo, edited demo, cherry-picking, and one-take video.","example":"At Tesla's We, Robot event in October 2024, Optimus chatted and interacted with guests on stage; it later came out that these interactions were mostly driven by a person operating it remotely, and Tesla was criticized for not disclosing that at the event.","related":["Teleoperated Demo","Fully Autonomous","Playback Speed Label","Edited Demo","Cherry-picking","One-Take (Uncut) Video"]},{"id":"fully-autonomous","category":"industry","sec":2,"tier":1,"sources":[{"title":"Introducing Helix 02: Full-Body Autonomy (Figure, 2026-01-27)","url":"https://www.figure.ai/news/helix-02"}],"as_of":"2026-01","related_ids":["teleoperated-demo","demo","playback-speed-label","remote-teleoperation-takeover","intervention-rate","levels-of-autonomy"],"name":"Fully Autonomous","alt":"全自主","abbr":"","aliases":["Autonomous Label","Non-teleoperated"],"one_liner":"The robot carries out a task under its own model's control the whole way through, with no person operating it from behind the scenes.","explanation":"This term is the counterpart to teleoperation, in which a person wearing a VR headset or holding a leader arm controls the robot in real time; “fully autonomous” means the policy model — a network that maps observations to actions — decides every step on its own. When releasing a demo, companies often label the footage “Autonomous” to signal that no one is operating it, sometimes alongside a “1x speed” label to show it hasn't been sped up. Note that “fully autonomous” only tells you no one was remote-controlling it during execution — it doesn't mean the results weren't cherry-picked, that the scene wasn't staged in advance, or that the task was something the model had never seen before; it's also worth checking whether the whole run was autonomous or whether a person stepped in partway through (see intervention rate).","example":"When Figure released Helix 02 in January 2026, it noted that the demo videos on the page were “all fully autonomous, not teleoperated,” and said the robot loaded and unloaded a dishwasher throughout an entire kitchen in a 4-minute task with no resets and no human intervention.","related":["Teleoperated Demo","Demo (Demonstration Video)","Playback Speed Label","Remote Teleoperation Takeover","Intervention Rate","Levels of Autonomy"]},{"id":"teleoperated-demo","category":"industry","sec":2,"tier":1,"sources":[{"title":"Mobile ALOHA 项目主页（Autonomous Skills / Teleoperation 分栏）","url":"https://mobile-aloha.github.io/"},{"title":"The Decoder: Watch a low-cost, all-purpose robot fry shrimp autonomously（Fu 称居家视频中的机器人「目前」是遥操作）","url":"https://the-decoder.com/watch-an-autonomous-robot-cook-and-then-build-one-yourself/"},{"title":"Teleoperation - Wikipedia","url":"https://en.wikipedia.org/wiki/Teleoperation"}],"as_of":"","related_ids":["teleoperation","fully-autonomous","demo","remote-teleoperation-takeover","mobile-aloha","pre-programmed-motion"],"name":"Teleoperated Demo","alt":"遥操作演示","abbr":"","aliases":["Wizard-of-Oz Demo"],"one_liner":"A video in which the robot is actually being controlled in real time by a person offstage, not making decisions on its own.","explanation":"Teleoperation is when a person controls a robot in real time using a VR headset, a leader-follower arm setup, a motion-capture suit, or similar equipment. It's a legitimate technique in its own right, and the main way training data gets collected — the problem is presenting teleoperated footage as a demonstration of autonomous ability without saying so, which the field calls a “puppet-show demo,” borrowing the English phrase Wizard-of-Oz, meaning someone pulling the strings behind the curtain. Teleoperation can prove the hardware is capable, but it doesn't prove the model has learned anything. To spot it, check whether the footage is labeled teleoperated or autonomous, whether the robot's responses look as instant and varied as a human's, and whether the publisher has given a success rate for autonomous operation.","example":"When Stanford's Mobile ALOHA launched in January 2024, its project page split “autonomous skills” and “teleoperated” videos into separate sections; the household-chore videos that went viral afterward were later described by the authors on social media as teleoperated recordings.","related":["Teleoperation","Fully Autonomous","Demo (Demonstration Video)","Remote Teleoperation Takeover","Mobile ALOHA","Pre-Programmed (Choreographed) Motion"]},{"id":"remote-teleoperation-takeover","category":"industry","sec":2,"tier":2,"sources":[{"title":"1X NEO 官网","url":"https://www.1x.tech/neo"}],"as_of":"2025-10","related_ids":["teleoperation","fully-autonomous","teleoperated-demo","human-intervention-data","deployment-data-backflow","human-in-the-loop"],"name":"Remote Teleoperation Takeover","alt":"远程接管（人工兜底）","abbr":"","aliases":["Human Fallback","Expert Mode (1X NEO)"],"one_liner":"When a robot fails or hits something unfamiliar, a remote human operator takes over to finish the task.","explanation":"This refers to a remote operator taking real-time control of a robot — usually over the network via a VR headset or handheld controller — whenever the robot's autonomous policy fails or meets a situation it hasn't seen before, finishing the task in its place. It's common practice in embodied AI today: models aren't yet reliable enough for commercial use on their own, so human fallback keeps the service running without interruption, and the data recorded during the takeover can be fed back into training (the data feedback loop). When 1X opened preorders for its NEO robot in 2025, it publicly described an “Expert Mode,” where a 1X employee can remotely operate the robot with the user's consent. The practice is also controversial: whether the actions shown in a demo are fully autonomous or secretly teleoperated from offstage depends on whether the maker discloses it clearly.","example":"When the 1X NEO home robot encounters a chore it can't handle, a 1X remote operator wearing a VR headset can take over, with the user's authorization, to finish it.","related":["Teleoperation","Fully Autonomous","Teleoperated Demo","Human Intervention Data","Deployment Data Backflow","Human-in-the-Loop"]},{"id":"pre-programmed-motion","category":"industry","sec":2,"tier":2,"sources":[{"title":"Atlas (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Atlas_(robot)"},{"title":"Boston Dynamics - Wikipedia","url":"https://en.wikipedia.org/wiki/Boston_Dynamics"}],"as_of":"","related_ids":["fully-autonomous",null,null,"demo","commercial-robot-performances","teleoperated-demo"],"name":"Pre-Programmed (Choreographed) Motion","alt":"预编程动作（动作编排）","abbr":"","aliases":["Choreographed Motion","Scripted Motion"],"one_liner":"A fixed motion sequence designed in advance and replayed by the robot, not decided on the spot.","explanation":"Pre-programmed motion means engineers design a sequence of movements ahead of time — a dance, a backflip, a martial-arts routine — and the robot replays or tracks that fixed script. Today these are commonly built with motion capture plus reinforcement learning to train a tracking policy, with balance and force output handled by the controller in real time, which is not technically trivial. But what to do and when is fixed in advance, and doesn't depend on the robot understanding its surroundings. This differs from full autonomy, which requires the robot to decide its next move based on what it actually perceives. When watching a robot dance or box, it's worth separating “this is strong motor control” from “this robot can autonomously complete a task.”","example":"Boston Dynamics' 2020 dance video “Do You Love Me” and the humanoid group dances on CCTV's Spring Festival Gala, China's most-watched TV broadcast, are both examples of pre-choreographed motion.","related":["Fully Autonomous","Motion Tracking","Motion Capture","Demo (Demonstration Video)","Commercial Robot Performances","Teleoperated Demo"]},{"id":"playback-speed-label","category":"industry","sec":2,"tier":2,"sources":[{"title":"Mobile ALOHA project page","url":"https://mobile-aloha.github.io/"},{"title":"π0: Our First Generalist Policy - Physical Intelligence","url":"https://www.physicalintelligence.company/blog/pi0"}],"as_of":"","related_ids":["demo","fully-autonomous","teleoperated-demo","one-take-video","edited-demo","cherry-picking"],"name":"Playback Speed Label","alt":"倍速 / 原速标注","abbr":"","aliases":["1x / Sped-Up Label","Playback Speed Disclosure"],"one_liner":"The “1x,” “2x,” or “10x” marker in the corner of a robot demo video showing its true speed.","explanation":"Robot demo videos often carry a small label in the corner — “1x” for real time, or “2x,” “4x,” “10x” for sped-up playback. Real-robot policies are often slow and pause mid-task, so publishers speed up the footage to keep it watchable and the runtime short. That's not inherently deceptive, but if the label is missing or too small to notice, viewers will overestimate how fast and fluid the robot really is — and speed is directly tied to whether a robot can meet a factory's cycle time and become commercially viable. So when watching a demo, it helps to check for the speed label, then check whether it's fully autonomous, filmed in one continuous take, and whether only successful attempts were kept — judging a robot's real ability means weighing all of these together.","example":"A laundry-folding video labeled “4x” means folding one item in reality likely took about four times as long as shown.","related":["Demo (Demonstration Video)","Fully Autonomous","Teleoperated Demo","One-Take (Uncut) Video","Edited Demo","Cherry-picking"]},{"id":"edited-demo","category":"industry","sec":2,"tier":2,"sources":[{"title":"Gemini (language model) - Wikipedia","url":"https://en.wikipedia.org/wiki/Gemini_(language_model)"}],"as_of":"","related_ids":["demo","cherry-picking","one-take-video","playback-speed-label","teleoperated-demo","fully-autonomous"],"name":"Edited Demo","alt":"剪辑 Demo","abbr":"","aliases":["Staged Video"],"one_liner":"A robot demonstration video that has been cut, spliced together, or sped up.","explanation":"This refers to a demo video that has been edited: several attempts stitched into what looks like one, failures and stalls cut out, playback sped up, or even outright staged. Editing isn't necessarily fraudulent on its own, but it can lead viewers to overestimate a robot's speed, success rate, and level of autonomy. When watching a demo, it's worth checking whether it's a single unbroken take, whether it's labeled as sped up, whether it states full autonomy versus teleoperation, and whether a success rate is disclosed. Rigorous teams will publish an unedited, real-time, full-length video to back up their claims.","example":"Reports say Google's 2023 Gemini demo video drew criticism for being edited in a way that shortened its actual response time.","related":["Demo (Demonstration Video)","Cherry-picking","One-Take (Uncut) Video","Playback Speed Label","Teleoperated Demo","Fully Autonomous"]},{"id":"one-take-video","category":"industry","sec":2,"tier":3,"sources":[{"title":"Long take - Wikipedia","url":"https://en.wikipedia.org/wiki/Long_take"}],"as_of":"","related_ids":["demo","edited-demo","playback-speed-label","fully-autonomous","teleoperated-demo","cherry-picking"],"name":"One-Take (Uncut) Video","alt":"一镜到底","abbr":"","aliases":["Continuous Single-Shot Footage"],"one_liner":"A robot demo filmed from start to finish in a single continuous shot, with no cuts.","explanation":"“One-take” (一镜到底) is originally a filmmaking term for a long take: a single continuous shot with no cuts. In embodied-AI circles, it's used to argue for a demo's credibility: a video stitched together from multiple clips can cut out failed attempts and keep only the successful ones, while a one-take video at least shows that this particular stretch of motion was completed continuously. It doesn't equal full autonomy — someone could still be teleoperating off-camera, and the clip could be the best of many, many takes. So when watching a demo, it's worth checking one-take footage, real-time playback, and full autonomy separately, as three distinct claims.","example":"When a company releases a video of a humanoid robot folding laundry with a corner label reading “1x real speed, one continuous take, fully autonomous,” it's addressing editing, speed, and teleoperation concerns all at once.","related":["Demo (Demonstration Video)","Edited Demo","Playback Speed Label","Fully Autonomous","Teleoperated Demo","Cherry-picking"]},{"id":"cherry-picking","category":"industry","sec":2,"tier":2,"sources":[{"title":"Cherry picking - Wikipedia","url":"https://en.wikipedia.org/wiki/Cherry_picking"}],"as_of":"","related_ids":["edited-demo","demo","success-rate","one-take-video","leaderboard-chasing"],"name":"Cherry-picking","alt":"cherry-pick（挑结果）","abbr":"","aliases":["Cherry-pick"],"one_liner":"Showing only the best-looking results out of many attempts while hiding failures and the average performance.","explanation":"The term originally means picking only the best cherries, and in research and marketing it means showing only the successful or most impressive runs out of many attempts. A robot policy's performance can vary a lot from one trial to the next — the same task might succeed 3 times out of 10 — so releasing only the successful clips can make people overestimate its ability. To spot this, check whether the total number of trials and the success rate are reported, whether the test objects and settings differ from training, and whether the video is an unbroken, single continuous take. Rigorous papers report success rates, confidence intervals, and failure cases.","example":"A paper's project page shows only 3 successful clips of a robot folding clothes, with no mention in the text of how many total attempts were made or what the overall success rate was.","related":["Edited Demo","Demo (Demonstration Video)","Success Rate","One-Take (Uncut) Video","Leaderboard Chasing"]},{"id":"leaderboard-chasing","category":"industry","sec":2,"tier":2,"sources":[{"title":"Goodhart's law - Wikipedia","url":"https://en.wikipedia.org/wiki/Goodhart%27s_law"},{"title":"LIBERO Benchmark","url":"https://libero-project.github.io/"}],"as_of":"","related_ids":["benchmark","benchmark-saturation","libero-benchmark","libero-plus","real-world-evaluation","cherry-picking"],"name":"Leaderboard Chasing","alt":"刷榜","abbr":"","aliases":["Benchmark Hacking","Gaming the Benchmark"],"one_liner":"Optimizing hard for one benchmark's score in a way that inflates it without necessarily improving real ability.","explanation":"“Leaderboard chasing” is AI-community slang for researchers or companies repeatedly tuning hyperparameters, cherry-picking settings, or even designing specifically around a public benchmark's quirks to push their score to the top. Chasing a benchmark in moderation can drive real progress, but overdoing it decouples the score from real capability — an instance of Goodhart's Law, which holds that once a measure becomes a target, it stops being a good measure. In embodied AI, a common version of this is simulation-benchmark saturation: several VLA (vision-language-action) models report near-perfect success rates on LIBERO, yet changing the camera angle or adding minor disturbances causes scores to drop sharply. That's why it's worth checking a paper's generalization tests and real-robot evaluations alongside its headline benchmark numbers.","example":"Many VLA models report above 95% success on LIBERO, while LIBERO-Plus and LIBERO-PRO, which add perturbations, show noticeably lower scores for the same models.","related":["Benchmark","Benchmark Saturation","LIBERO Benchmark","LIBERO-Plus","Real-World Evaluation","Cherry-picking"]},{"id":"arxiv-preprint","category":"industry","sec":3,"tier":1,"sources":[{"title":"About arXiv","url":"https://info.arxiv.org/about/index.html"},{"title":"arXiv spin-out FAQ（2026 年 7 月 1 日起为独立非营利机构 arXiv, Inc.）","url":"https://info.arxiv.org/about/spinout_faq.html"},{"title":"arXiv Help: Availability of submissions（编号按首次公开的月份分配）","url":"https://info.arxiv.org/help/availability.html"}],"as_of":"2026-07","related_ids":["top-tier-conferences-and-journals","conference-on-robot-learning","ieee-international-conference-on-robotics-and-automation","state-of-the-art","pi0"],"name":"arXiv Preprint","alt":"arXiv 预印本","abbr":"","aliases":["arXiv","Preprint","Posting to arXiv"],"one_liner":"A paper version posted publicly before formal publication; most AI-field papers post to arXiv first.","explanation":"arXiv is a free preprint platform created by physicist Paul Ginsparg in 1991, hosted by Cornell University for most of its history, and since July 2026 an independent nonprofit. Authors upload a paper and it goes public after light screening, without peer review. AI and embodied-AI research moves fast, and most new papers are posted to arXiv around the time they are submitted, without waiting for formal acceptance, so “posting to arXiv” is essentially how a paper first appears. Each paper gets an ID number whose first four digits give the year and month it first went public. When reading one, keep in mind that a preprint hasn't been peer-reviewed, so its conclusions and experiments may still change across versions (v1, v2, and so on); it's worth checking whether a formally published version exists before citing it.","example":"The π0 paper is arXiv:2410.24164 — the first four digits, 2410, mean it went public in October 2024.","related":["Top-tier Conferences & Journals (CCF-A Venues)","Conference on Robot Learning","IEEE International Conference on Robotics and Automation","State of the Art (SOTA)","π0"]},{"id":"top-tier-conferences-and-journals","category":"industry","sec":3,"tier":1,"sources":[{"title":"中国计算机学会推荐国际学术会议和期刊目录","url":"https://www.ccf.org.cn/Academic_Evaluation/By_category/"},{"title":"第七版中国计算机学会推荐国际学术会议和期刊目录（CCF 2026）附件页（中国科学院大学人工智能学院）","url":"https://ai.ucas.ac.cn/index.php/zh-cn/jxjy/fzpy1/7675-ccf-2026"}],"as_of":"2026-03","related_ids":["arxiv-preprint","cvpr-iccv-eccv","neurips-icml-iclr","ieee-international-conference-on-robotics-and-automation","conference-on-robot-learning","robotics-science-and-systems"],"name":"Top-tier Conferences & Journals (CCF-A Venues)","alt":"顶会 / 顶刊（CCF-A）","abbr":"","aliases":["Top-tier Venues","CCF-A"],"one_liner":"The conferences and journals widely seen as most influential in a field, commonly benchmarked in China against the CCF's A-tier list.","explanation":"CCF is the China Computer Federation, which publishes a “Recommended List of International Academic Conferences and Journals” that sorts computer-science venues in each subfield into tiers A, B, and C; the A tier is what Chinese universities commonly mean by “top-tier,” and the list is revised every few years, with the seventh edition published in March 2026. Embodied-AI-relevant CCF-A venues include conferences such as CVPR, ICCV, NeurIPS, ICML, and ICLR, and journals such as TPAMI and IJCV. Robotics-specific venues such as ICRA, IROS, CoRL, and RSS are considered top venues within robotics circles, but none of them are CCF-A: ICRA is tier B, IROS is tier C, and CoRL and RSS aren't listed at all, so “top venue” and “CCF-A” aren't the same thing. Computer science overall leans on conferences rather than journals, and papers are commonly posted to arXiv around the time they're submitted.","example":"ECCV is tier B in the CCF list, but within the vision community it's considered one of the “big three” alongside CVPR and ICCV.","related":["arXiv Preprint","CVPR / ICCV / ECCV","NeurIPS / ICML / ICLR","IEEE International Conference on Robotics and Automation","Conference on Robot Learning","Robotics: Science and Systems"]},{"id":"neurips-icml-iclr","category":"industry","sec":3,"tier":2,"sources":[{"title":"NeurIPS","url":"https://neurips.cc/"},{"title":"ICML","url":"https://icml.cc/"},{"title":"ICLR","url":"https://iclr.cc/"}],"as_of":"","related_ids":["top-tier-conferences-and-journals","conference-on-robot-learning","robotics-science-and-systems","ieee-international-conference-on-robotics-and-automation","cvpr-iccv-eccv","arxiv-preprint"],"name":"NeurIPS / ICML / ICLR","alt":"NeurIPS / ICML / ICLR（机器学习三大会）","abbr":"","aliases":["Machine Learning's “Big Three” Conferences"],"one_liner":"The three most influential annual academic conferences in machine learning.","explanation":"NeurIPS (Conference on Neural Information Processing Systems, called NIPS until 2018, usually held in December), ICML (International Conference on Machine Learning, usually held in summer), and ICLR (International Conference on Learning Representations, launched in 2013 by Yoshua Bengio, Yann LeCun, and others, with public review on OpenReview) are together known as machine learning's “big three” conferences. In computer science, conference papers are the primary form of publication, and these three have low acceptance rates and high impact; China's CCF ranking classifies them as top-tier “CCF-A” venues. Methods-heavy embodied-AI work — VLA models, world models, reinforcement-learning algorithms — is often submitted to these three, while more systems- and real-robot-focused work tends to go to CoRL, RSS, or ICRA instead.","example":"The Decision Transformer paper was published at NeurIPS 2021; many VLA and world-model papers are submitted to ICLR.","related":["Top-tier Conferences & Journals (CCF-A Venues)","Conference on Robot Learning","Robotics: Science and Systems","IEEE International Conference on Robotics and Automation","CVPR / ICCV / ECCV","arXiv Preprint"]},{"id":"cvpr-iccv-eccv","category":"industry","sec":3,"tier":2,"sources":[{"title":"Conference on Computer Vision and Pattern Recognition - Wikipedia","url":"https://en.wikipedia.org/wiki/Conference_on_Computer_Vision_and_Pattern_Recognition"},{"title":"CVPR official site","url":"https://cvpr.thecvf.com/"}],"as_of":"","related_ids":["top-tier-conferences-and-journals","neurips-icml-iclr","conference-on-robot-learning","computer-vision","arxiv-preprint"],"name":"CVPR / ICCV / ECCV","alt":"CVPR / ICCV / ECCV（计算机视觉三大会）","abbr":"CVPR / ICCV / ECCV","aliases":["The Big Three of Computer Vision"],"one_liner":"The three most influential international academic conferences in computer vision.","explanation":"CVPR is held every year; ICCV is held in odd-numbered years; and ECCV is held in even-numbered years, originally always in Europe. All three are top venues in vision and are CCF-A tier. Embodied AI leans heavily on vision: work on 3D reconstruction, depth estimation, segmentation, video generation, and world models is heavily published at these venues, and increasingly more VLA and embodied-navigation papers are submitted to them as well, with CVPR often running dedicated embodied-AI workshops.","example":"3D vision foundation models such as DUSt3R and VGGT were both published at CVPR and later widely adopted for robot perception.","related":["Top-tier Conferences & Journals (CCF-A Venues)","NeurIPS / ICML / ICLR","Conference on Robot Learning","Computer Vision","arXiv Preprint"]},{"id":"ieee-international-conference-on-robotics-and-automation","category":"industry","sec":3,"tier":2,"sources":[{"title":"Wikipedia: International Conference on Robotics and Automation","url":"https://en.wikipedia.org/wiki/International_Conference_on_Robotics_and_Automation"}],"as_of":"2025-05","related_ids":["ieee-rsj-international-conference-on-intelligent-robots-and","conference-on-robot-learning","robotics-science-and-systems","ieee-t-ro-ijrr-ra-l","ieee-robotics-and-automation-society","top-tier-conferences-and-journals"],"name":"IEEE International Conference on Robotics and Automation","alt":"ICRA","abbr":"ICRA","aliases":["ICRA"],"one_liner":"The flagship annual robotics conference hosted by the IEEE Robotics and Automation Society.","explanation":"ICRA is hosted by the IEEE Robotics and Automation Society (RAS) and has run annually since 1984, usually in May or June; it's one of the largest academic robotics conferences, drawing thousands of paper submissions each year. Its coverage spans nearly every area of robotics — motion planning, control, perception, SLAM, manipulation, legged locomotion, and human-robot interaction — and robot-learning and VLA papers have grown quickly in recent years. It's considered, alongside IROS, one of robotics' two major general-purpose conferences, and compared with the more focused CoRL and RSS, it covers more ground at a larger scale. It's also linked with the journal RA-L, whose accepted papers can be presented at ICRA. ICRA 2025 was held in Atlanta, USA.","example":"Many classic algorithms debuted at ICRA — Kajita and colleagues' 2003 ZMP preview control, for instance, and Mellinger and Kumar's 2011 minimum-snap trajectory generation for quadrotors.","related":["IEEE/RSJ International Conference on Intelligent Robots and Systems","Conference on Robot Learning","Robotics: Science and Systems","IEEE T-RO / IJRR / RA-L","IEEE Robotics and Automation Society","Top-tier Conferences & Journals (CCF-A Venues)"]},{"id":"ieee-rsj-international-conference-on-intelligent-robots-and","category":"industry","sec":3,"tier":2,"sources":[{"title":"Wikipedia: International Conference on Intelligent Robots and Systems","url":"https://en.wikipedia.org/wiki/International_Conference_on_Intelligent_Robots_and_Systems"}],"as_of":"2026-09","related_ids":["ieee-international-conference-on-robotics-and-automation","conference-on-robot-learning","robotics-science-and-systems","ieee-ras-international-conference-on-humanoid-robots","ieee-t-ro-ijrr-ra-l","top-tier-conferences-and-journals"],"name":"IEEE/RSJ International Conference on Intelligent Robots and Systems","alt":"IROS","abbr":"IROS","aliases":["IROS"],"one_liner":"A major annual robotics conference on par with ICRA, held each fall.","explanation":"IROS, in full the IEEE/RSJ International Conference on Intelligent Robots and Systems, was first held in Tokyo in 1988 and is jointly organized by the IEEE and the Robotics Society of Japan (RSJ), among others; it's usually held in the fall and draws more than 2,000 paper submissions a year. Its research scope overlaps heavily with ICRA's, spanning motion control, perception, SLAM, and robot learning, and the field commonly refers to the two together as robotics' two major conferences. It rotates between Asia, Europe, and the Americas: Abu Dhabi in 2024, Hangzhou in 2025, Pittsburgh, USA, in 2026, and Florence, Italy, set for 2027. For newcomers submitting papers, IROS and ICRA are usually the first choices in robotics.","example":"LIO-SAM, a widely used method in LiDAR SLAM, was published at IROS 2020.","related":["IEEE International Conference on Robotics and Automation","Conference on Robot Learning","Robotics: Science and Systems","IEEE-RAS International Conference on Humanoid Robots","IEEE T-RO / IJRR / RA-L","Top-tier Conferences & Journals (CCF-A Venues)"]},{"id":"conference-on-robot-learning","category":"industry","sec":3,"tier":2,"sources":[{"title":"Conference on Robot Learning (CoRL)","url":"https://www.corl.org/"}],"as_of":"","related_ids":["robotics-science-and-systems","ieee-international-conference-on-robotics-and-automation","ieee-rsj-international-conference-on-intelligent-robots-and","robot-learning","top-tier-conferences-and-journals"],"name":"Conference on Robot Learning","alt":"CoRL","abbr":"CoRL","aliases":["CoRL"],"one_liner":"An annual international academic conference focused on the intersection of robotics and machine learning.","explanation":"CoRL was established in 2017 and is held once a year, focused on robot learning — using machine-learning methods for a robot's perception, decision-making, and control. Compared with the broader, more general ICRA and IROS conferences, it is smaller and more tightly focused, and it's a major publication venue for embodied AI, VLA (vision-language-action) models, imitation learning, and reinforcement learning; its proceedings are collected in PMLR. When tracking the embodied-AI frontier, CoRL is often watched together with RSS and ICRA.","example":"Landmark VLA papers such as RT-2 and OpenVLA were both published at CoRL.","related":["Robotics: Science and Systems","IEEE International Conference on Robotics and Automation","IEEE/RSJ International Conference on Intelligent Robots and Systems","Robot Learning","Top-tier Conferences & Journals (CCF-A Venues)"]},{"id":"robotics-science-and-systems","category":"industry","sec":3,"tier":2,"sources":[{"title":"Robotics: Science and Systems 官网","url":"https://roboticsconference.org/"}],"as_of":"","related_ids":["ieee-international-conference-on-robotics-and-automation","ieee-rsj-international-conference-on-intelligent-robots-and","conference-on-robot-learning","top-tier-conferences-and-journals","science-robotics"],"name":"Robotics: Science and Systems","alt":"RSS","abbr":"RSS","aliases":["RSS Conference"],"one_liner":"A single-track top academic robotics conference held once a year.","explanation":"RSS is one of the most influential academic conferences in robotics, held annually since 2005 as a single-track event — every paper is presented in sequence in the same room — with a far lower acceptance count than ICRA or IROS, and a reputation for strict review and high paper quality. Important work in robot learning, manipulation, and motion planning is often published here, and together with CoRL it's considered one of the two most closely watched robot-learning venues. Newcomers reading papers can treat an “RSS 20XX” citation as a signal of high-quality work in that area; submissions are usually due early in the year, with the conference itself held in summer.","example":"The Diffusion Policy paper was published at RSS 2023.","related":["IEEE International Conference on Robotics and Automation","IEEE/RSJ International Conference on Intelligent Robots and Systems","Conference on Robot Learning","Top-tier Conferences & Journals (CCF-A Venues)","Science Robotics"]},{"id":"science-robotics","category":"industry","sec":3,"tier":2,"sources":[{"title":"Science Robotics 期刊主页","url":"https://www.science.org/journal/scirobotics"}],"as_of":"","related_ids":["ieee-t-ro-ijrr-ra-l","robotics-science-and-systems","top-tier-conferences-and-journals","anymal-rl-locomotion-series"],"name":"Science Robotics","alt":"Science Robotics","abbr":"","aliases":[],"one_liner":"The top robotics-focused journal published under the Science family of journals.","explanation":"Science Robotics is a sister journal of Science, published by the American Association for the Advancement of Science (AAAS). Launched in 2016 and issued monthly, it covers robot hardware, control, learning, medical robotics, bio-inspired design, and related areas. It favors work with complete systems and real-world validation, publishes relatively few papers, and carries substantial influence, making it one of the most respected journals in robotics. Unlike conference papers, journal articles go through a longer review cycle and are written more completely, often accompanied by extensive video material. Multiple landmark papers on reinforcement-learning-based legged locomotion have appeared here.","example":"The paper on perceptive locomotion for the ANYmal quadruped over complex terrain (Miki et al., 2022) was published in Science Robotics.","related":["IEEE T-RO / IJRR / RA-L","Robotics: Science and Systems","Top-tier Conferences & Journals (CCF-A Venues)","ANYmal RL Locomotion Series"]},{"id":"ieee-t-ro-ijrr-ra-l","category":"industry","sec":3,"tier":3,"sources":[{"title":"IEEE Transactions on Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/IEEE_Transactions_on_Robotics"},{"title":"The International Journal of Robotics Research - Wikipedia","url":"https://en.wikipedia.org/wiki/The_International_Journal_of_Robotics_Research"}],"as_of":"","related_ids":["ieee-international-conference-on-robotics-and-automation","ieee-rsj-international-conference-on-intelligent-robots-and","conference-on-robot-learning","robotics-science-and-systems","science-robotics","top-tier-conferences-and-journals"],"name":"IEEE T-RO / IJRR / RA-L","alt":"T-RO / IJRR / RA-L（机器人期刊）","abbr":"","aliases":["IEEE Transactions on Robotics","The International Journal of Robotics Research","IEEE Robotics and Automation Letters"],"one_liner":"The three journals most often cited in robotics: T-RO, IJRR, and RA-L.","explanation":"These three are the most frequently cited journals in robotics. IEEE Transactions on Robotics (T-RO) is run by the IEEE Robotics and Automation Society (RAS), and traces back to a predecessor launched in 1985; it favors complete theoretical and systems work. The International Journal of Robotics Research (IJRR), published by SAGE and launched in 1982, is one of the oldest robotics journals, similarly known for long articles and a high bar for acceptance. IEEE Robotics and Automation Letters (RA-L) is RAS's rapid-publication journal — short articles, fast review — and an accepted RA-L paper can also be presented at a conference like ICRA or IROS. Much work in robot learning is published first at conferences such as CoRL or RSS, with journals more often used to publish extended, complete versions; either way, a journal publication is generally weighted heavily in job applications and evaluations.","example":"An ICRA submission can also go through the “submit to RA-L, present at ICRA” track, which after acceptance counts as both a journal paper and a conference presentation.","related":["IEEE International Conference on Robotics and Automation","IEEE/RSJ International Conference on Intelligent Robots and Systems","Conference on Robot Learning","Robotics: Science and Systems","Science Robotics","Top-tier Conferences & Journals (CCF-A Venues)"]},{"id":"ieee-ras-international-conference-on-humanoid-robots","category":"industry","sec":3,"tier":3,"sources":[{"title":"About IEEE RAS","url":"https://www.ieee-ras.org/about-ras"}],"as_of":"","related_ids":["ieee-international-conference-on-robotics-and-automation","ieee-rsj-international-conference-on-intelligent-robots-and","conference-on-robot-learning","robotics-science-and-systems","humanoid-robot","whole-body-control"],"name":"IEEE-RAS International Conference on Humanoid Robots","alt":"Humanoids 会议","abbr":"Humanoids","aliases":["Humanoids","IEEE-RAS Humanoids"],"one_liner":"An international conference run by the IEEE Robotics and Automation Society, dedicated specifically to humanoid robots.","explanation":"Humanoids is an international conference organized by the IEEE Robotics and Automation Society (RAS), first held in 2000 and roughly annual since, dedicated specifically to humanoid robots. Common topics include bipedal walking and balance control, whole-body control, dexterous humanoid hands, motion imitation and teleoperation, and human-robot interaction. It's smaller in scale than broad robotics conferences like ICRA and IROS, but more focused on its niche — research groups working on humanoid locomotion and whole-body control frequently submit work here, and new humanoid platforms are often demonstrated at the event. Newcomers reading humanoid-related papers will regularly see “Humanoids 20xx” in citation lists.","example":"Papers on whole-body humanoid teleoperation, fall recovery, and bipedal gait are commonly submitted to Humanoids.","related":["IEEE International Conference on Robotics and Automation","IEEE/RSJ International Conference on Intelligent Robots and Systems","Conference on Robot Learning","Robotics: Science and Systems","Humanoid Robot","Whole-Body Control"]},{"id":"china-embodied-ai-conference","category":"industry","sec":3,"tier":3,"sources":[{"title":"中国人工智能学会官网","url":"http://www.caai.cn/"}],"as_of":"","related_ids":["world-robot-conference","world-artificial-intelligence-conference","conference-on-robot-learning","ieee-international-conference-on-robotics-and-automation","embodied-ai"],"name":"China Embodied AI Conference","alt":"中国具身智能大会（CEAI）","abbr":"CEAI","aliases":["CEAI"],"one_liner":"China's domestic annual academic conference focused specifically on embodied AI.","explanation":"The China Embodied AI Conference (CEAI) is a domestic academic conference dedicated to embodied AI, reportedly organized by the Chinese Association for Artificial Intelligence and other bodies, held annually starting in 2024. It typically includes invited talks, themed forums, paper presentations, and robot demonstrations, covering embodied foundation models, robot manipulation and locomotion control, data collection, simulation, and evaluation. Its role is to bring together Chinese researchers and companies working on vision, machine learning, and robot control, letting newcomers use its talks and forum agendas to learn which domestic teams exist and what each is working on. It differs from international venues like ICRA or CoRL in emphasizing exchange over rigorous, formal paper publication.","example":"","related":["World Robot Conference","World Artificial Intelligence Conference","Conference on Robot Learning","IEEE International Conference on Robotics and Automation","Embodied AI"]},{"id":"world-artificial-intelligence-conference","category":"industry","sec":4,"tier":2,"sources":[{"title":"世界人工智能大会官网","url":"https://www.worldaic.com.cn/"},{"title":"World Artificial Intelligence Conference - Wikipedia","url":"https://en.wikipedia.org/wiki/World_Artificial_Intelligence_Conference"}],"as_of":"2025","related_ids":["world-robot-conference","world-humanoid-robot-games","nvidia-gtc","consumer-electronics-show","humanoid-robot"],"name":"World Artificial Intelligence Conference","alt":"世界人工智能大会","abbr":"WAIC","aliases":["WAIC"],"one_liner":"A comprehensive annual AI conference and expo held every year in Shanghai.","explanation":"The World Artificial Intelligence Conference has been held annually in Shanghai since 2018, usually in summer, co-organized by the Shanghai municipal government together with relevant national ministries, and includes a main forum, sub-forums, and a large exhibition. It covers the whole AI field — large language models, chips, autonomous driving — and in recent years the humanoid-robot and embodied-AI exhibition areas have become a major focus, with many Chinese companies choosing WAIC as the venue to unveil new robots or new models, making it an important window for tracking China's embodied-AI progress.","example":"","related":["World Robot Conference","World Humanoid Robot Games","NVIDIA GTC","Consumer Electronics Show","Humanoid Robot"]},{"id":"world-robot-conference","category":"industry","sec":4,"tier":2,"sources":[{"title":"世界机器人大会官网","url":"https://www.worldrobotconference.com/"}],"as_of":"2025","related_ids":["world-artificial-intelligence-conference","world-humanoid-robot-games","humanoid-robot","upstream-midstream-downstream-of-the-industry-chain","beijing-humanoid-robot-innovation-center"],"name":"World Robot Conference","alt":"世界机器人大会","abbr":"WRC","aliases":["WRC"],"one_liner":"An annual robotics-focused conference held in Beijing, combining forums, an expo, and competitions.","explanation":"The World Robot Conference has been held in Beijing since 2015, usually in August, jointly organized by the Beijing municipal government, the Ministry of Industry and Information Technology, the China Association for Science and Technology, and others, with the China Institute of Electronics as organizer, and it consists of a forum, an expo, and a robot competition. Compared with the broader WAIC, it focuses more specifically on robotics itself: industrial robots, components, and humanoid and service robot makers exhibit heavily, and in recent years the number of humanoid robots on display and new product launches — both complete robots and components — have been the main draw.","example":"","related":["World Artificial Intelligence Conference","World Humanoid Robot Games","Humanoid Robot","Upstream / Midstream / Downstream of the Industry Chain","Beijing Humanoid Robot Innovation Center"]},{"id":"nvidia-gtc","category":"industry","sec":4,"tier":2,"sources":[{"title":"NVIDIA GTC","url":"https://www.nvidia.com/gtc/"},{"title":"NVIDIA Isaac GR00T","url":"https://developer.nvidia.com/isaac/gr00t"}],"as_of":"2025-03","related_ids":[null,"nvidia-isaac-gr00t-n1","nvidia-cosmos","newton-physics-engine","nvidia-three-computer-solution","physical-ai"],"name":"NVIDIA GTC","alt":"英伟达 GTC 大会","abbr":"GTC","aliases":["GPU Technology Conference","GTC"],"one_liner":"NVIDIA's annual AI and GPU technology conference, where Jensen Huang unveils new products.","explanation":"GTC, short for GPU Technology Conference, is NVIDIA's flagship technical conference, with its main event usually held each spring in San Jose, California, plus regional editions elsewhere. Founder and CEO Jensen Huang's keynote is where NVIDIA unveils its next-generation GPUs, software platforms, and partner news, making it a bellwether for AI compute and NVIDIA's overall strategy. In recent years robotics and physical AI have been a major focus: at GTC 2024, NVIDIA announced Project GR00T, a foundation-model initiative for humanoid robots; at GTC 2025, it open-sourced GR00T N1 and unveiled the Newton physics engine, developed with Google DeepMind and Disney. Each GTC tends to trigger a wave of coordinated partnership announcements from robotics companies.","example":"In the GTC 2025 keynote, Jensen Huang announced the open-sourcing of the Isaac GR00T N1 humanoid-robot foundation model.","related":["NVIDIA","NVIDIA Isaac GR00T N1","NVIDIA Cosmos","Newton Physics Engine","NVIDIA Three-Computer Solution","Physical AI"]},{"id":"consumer-electronics-show","category":"industry","sec":4,"tier":3,"sources":[{"title":"Consumer Electronics Show - Wikipedia","url":"https://en.wikipedia.org/wiki/Consumer_Electronics_Show"},{"title":"CES 官网","url":"https://www.ces.tech/"}],"as_of":"2026-01","related_ids":["nvidia-gtc","nvidia-cosmos","boston-dynamics-atlas-2","lg-cloid","physical-ai","world-robot-conference"],"name":"Consumer Electronics Show","alt":"CES 国际消费电子展","abbr":"CES","aliases":["CES"],"one_liner":"The world's largest consumer-electronics trade show, held every January in Las Vegas.","explanation":"CES is organized by the US Consumer Technology Association (CTA), first held in 1967, and now takes place every January in Las Vegas, Nevada — a major venue where TV, phone, and automotive-electronics makers, among others, unveil new products. In recent years it has also become an important launch platform for physical AI and humanoid robots: at CES in January 2025, NVIDIA unveiled its Cosmos world foundation model platform, with CEO Jensen Huang saying robotics' “ChatGPT moment” was approaching; at CES in January 2026, Boston Dynamics unveiled the production version of its electric Atlas, LG showed off its home robot CLOiD, and several Chinese humanoid-robot companies demonstrated their products on the show floor. For tracking industry news, January's CES and March's NVIDIA GTC are the two months when announcements cluster most heavily.","example":"At CES in January 2026, Hyundai Motor Group and Boston Dynamics showed off the production version of Atlas and announced plans to deploy it in car factories.","related":["NVIDIA GTC","NVIDIA Cosmos","Boston Dynamics Atlas (Electric)","LG CLOiD","Physical AI","World Robot Conference"]},{"id":"tesla-ai-day","category":"industry","sec":4,"tier":3,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"}],"as_of":"2022-09","related_ids":["tesla-optimus","tesla","tesla-we-robot-event","tesla-supply-chain","humanoid-robot"],"name":"Tesla AI Day","alt":"特斯拉 AI Day","abbr":"","aliases":[],"one_liner":"Tesla's 2021 and 2022 technical showcase events, where the Optimus humanoid robot first appeared.","explanation":"AI Day was a Tesla event aimed at engineers and recruiting, held twice. The first, in August 2021, focused mainly on autonomous driving and the Dojo supercomputer, and ended by unveiling the concept of a humanoid robot called Tesla Bot — at the time, just a person in a costume walking on stage. The second, in September 2022, showed an actual Optimus prototype walking on stage, with Elon Musk saying the target price would be under $20,000. These two events are widely credited with fueling a wave of humanoid-robot startups and investment both in China and abroad.","example":"At AI Day 2022, an early Optimus prototype walked onto the stage without its outer shell and waved to the audience.","related":["Tesla Optimus","Tesla","Tesla We, Robot Event (2024)","Tesla (Optimus) Supply Chain","Humanoid Robot"]},{"id":"tesla-we-robot-event","category":"industry","sec":4,"tier":3,"sources":[{"title":"Optimus (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"title":"Tesla AI & Robotics","url":"https://www.tesla.com/AI"}],"as_of":"2024-10","related_ids":["tesla-optimus","tesla","tesla-ai-day","teleoperated-demo","fully-autonomous","remote-teleoperation-takeover"],"name":"Tesla We, Robot Event (2024)","alt":"特斯拉 We, Robot 发布会","abbr":"","aliases":["We, Robot"],"one_liner":"Tesla's October 2024 event unveiling the Cybercab, with Optimus robots interacting live with attendees.","explanation":"We, Robot was a Tesla event held in October 2024 at the Warner Bros. studio lot in California, mainly unveiling the driverless Cybercab and the Robovan. Several Optimus humanoid robots on-site poured drinks for guests, chatted, and danced, and Musk said Optimus would eventually be priced around $20,000 to $30,000. Multiple outlets later reported that the robots' conversation and some of their movements were remotely operated by humans rather than fully autonomous. The event is commonly cited as a reminder to separate teleoperation from full autonomy when watching robot demonstrations.","example":"At the event, an Optimus poured drinks for guests at the bar, which was reportedly being remotely operated behind the scenes.","related":["Tesla Optimus","Tesla","Tesla AI Day","Teleoperated Demo","Fully Autonomous","Remote Teleoperation Takeover"]},{"id":"yangbot","category":"industry","sec":4,"tier":2,"sources":[{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"}],"as_of":"2025-01","related_ids":["wubot","unitree-h1","unitree-robotics","pre-programmed-motion","commercial-robot-performances"],"name":"Yangbot","alt":"《秧BOT》","abbr":"","aliases":["Unitree H1 Yangge Dance (2025 CCTV Spring Festival Gala)"],"one_liner":"A 2025 CCTV Spring Festival Gala segment in which Unitree H1 humanoid robots performed a traditional yangge dance.","explanation":"Yangbot (《秧BOT》) was a segment on China Media Group's 2025 CCTV Spring Festival Gala — China's most-watched annual TV broadcast — in which Unitree Robotics' H1 humanoid robots performed the yangge, a traditional folk dance, spinning handkerchiefs alongside dancers from the Xinjiang Arts Institute. The segment was the first time humanoid robots appeared as a group on China's biggest television broadcast, and it's widely seen as the moment humanoid robots entered mainstream public awareness; public attention to Unitree, and to the humanoid-robot sector as a whole, rose sharply afterward. The dance moves were pre-choreographed, and the performance doesn't represent the robots having general autonomous ability.","example":"","related":["Wubot","Unitree H1","Unitree Robotics","Pre-Programmed (Choreographed) Motion","Commercial Robot Performances"]},{"id":"wubot","category":"industry","sec":4,"tier":3,"sources":[{"title":"宇树科技官网","url":"https://www.unitree.com/"}],"as_of":"2026-02","related_ids":["yangbot","unitree-robotics","pre-programmed-motion","motion-tracking","whole-body-control","commercial-robot-performances"],"name":"Wubot","alt":"《武BOT》","abbr":"","aliases":["Unitree Kung Fu Performance (2026 CCTV Spring Festival Gala)"],"one_liner":"A 2026 CCTV Spring Festival Gala segment featuring a Unitree humanoid robot martial-arts performance.","explanation":"Wubot (《武BOT》) was a robot martial-arts segment on China Media Group's 2026 CCTV Spring Festival Gala. Humanoid robots from Unitree Robotics took part in the performance, reportedly appearing alongside martial-arts performers to complete fist forms, weapons routines, and other movements. It was Unitree's second appearance on the Spring Festival Gala, following the 2025 segment Yangbot, in which Unitree's H1 performed the yangge folk dance. As with that earlier segment, the movements here were pre-choreographed, trained into a motion-tracking policy using motion-capture data and then executed on the robot; the performance showcases whole-body motor control and balance, and doesn't mean the robot can autonomously understand or respond to arbitrary situations. It has also become another important window through which the Chinese public has come to know humanoid robots.","example":"","related":["Yangbot","Unitree Robotics","Pre-Programmed (Choreographed) Motion","Motion Tracking","Whole-Body Control","Commercial Robot Performances"]},{"id":"hangzhou-s-six-little-dragons","category":"industry","sec":4,"tier":3,"sources":[{"title":"杭州六小龙 - 维基百科","url":"https://zh.wikipedia.org/wiki/%E6%9D%AD%E5%B7%9E%E5%85%AD%E5%B0%8F%E9%BE%99"}],"as_of":"2025-02","related_ids":["unitree-robotics","deep-robotics","brainco","manycore-tech","yangbot"],"name":"Hangzhou's Six Little Dragons","alt":"杭州六小龙","abbr":"","aliases":[],"one_liner":"A 2025 nickname for six Hangzhou tech companies that rose to fame together, two of them robotics makers.","explanation":"“Hangzhou's Six Little Dragons” is a label that became popular in Chinese media and online discussion in early 2025 for six Hangzhou-headquartered tech companies: Game Science, DeepSeek, Unitree Robotics, DEEP Robotics, BrainCo, and Manycore Tech. It caught on after Black Myth: Wukong, the DeepSeek models, and Unitree's robots appearing at the CCTV Spring Festival Gala all went viral in succession. Of the six, Unitree and DEEP Robotics make quadruped and humanoid robots, BrainCo makes brain-computer interfaces and bionic hands, and Manycore Tech makes spatial-design software and spatial-intelligence models, which is why this term often appears in embodied-AI coverage. It's a media label, not an official list.","example":"After Unitree's humanoid robots performed “Yangbot” at the 2025 Spring Festival Gala, the phrase “Hangzhou's Six Little Dragons” spread quickly.","related":["Unitree Robotics","DEEP Robotics","BrainCo","Manycore Tech","Yangbot"]},{"id":"world-humanoid-robot-games","category":"industry","sec":4,"tier":2,"sources":[{"title":"World Humanoid Robot Games - Wikipedia","url":"https://en.wikipedia.org/wiki/World_Humanoid_Robot_Games"}],"as_of":"2025-08","related_ids":["humanoid-robot-half-marathon","world-robot-conference","robocup","humanoid-robot","fall-recovery"],"name":"World Humanoid Robot Games","alt":"世界人形机器人运动会","abbr":"WHRG","aliases":["WHRG","Robot Olympics"],"one_liner":"A multi-sport games held for the first time in Beijing in August 2025, exclusively for humanoid robots.","explanation":"The first World Humanoid Robot Games was held in Beijing in August 2025, featuring competitive events such as track and field, soccer, and combat sports, plus task-based events like material handling and medicine sorting, with teams from Chinese and international companies and universities. Media commonly call it the “robot Olympics.” Its significance is putting locomotion abilities — walking, balance, fall recovery, and multi-robot coordination — on public display for comparison, which also exposed current shortcomings in humanoid robots' battery life and stability.","example":"Unitree's H1 reportedly won the 1,500-meter race at the first Games.","related":["Humanoid Robot Half Marathon (Beijing E-Town)","World Robot Conference","RoboCup","Humanoid Robot","Fall Recovery"]},{"id":"humanoid-robot-half-marathon","category":"industry","sec":4,"tier":2,"sources":[{"title":"北京人形：天工自主跑完北京亦庄半马","url":"https://x-humanoid.com/news-view-164.html"},{"title":"百余台机器人同跑半马 「闪电」超越人类纪录（新华网）","url":"https://www.news.cn/sports/20260419/0834250af22d4432ac708322aa8f7123/c.html"},{"title":"马拉松亚军人形机器人「松延动力 N2」被拍卖（IT之家）","url":"https://www.ithome.com/0/847/782.htm"},{"title":"50分26秒！2026北京亦庄人形机器人半马冠军出炉（澎湃新闻）","url":"https://m.thepaper.cn/newsDetail_forward_33004042"},{"title":"一年提速近两小时、从遥控到自主、跑赢人类！人形机器人「半马」刷新纪录（每日经济新闻，2026-04-19）","url":"https://www.nbd.com.cn/articles/2026-04-19/4345990.html"},{"title":"Kiplimo breaks world half marathon record with 57:20 on Lisbon return（World Athletics，2026-03-08）","url":"https://worldathletics.org/competitions/world-athletics-label-road-races/news/jacob-kiplimo-half-marathon-world-record-lisbon"}],"as_of":"2026-04","related_ids":["world-humanoid-robot-games","tiangong","honor-lightning-humanoid-robot","noetix-n2","beijing-humanoid-robot-innovation-center","battery-runtime"],"name":"Humanoid Robot Half Marathon (Beijing E-Town)","alt":"人形机器人半程马拉松","abbr":"","aliases":["Beijing E-Town Half Marathon"],"one_liner":"A roughly 21-kilometer half marathon for humanoid robots held in Beijing's E-Town district, first run in April 2025.","explanation":"This is a humanoid robot half marathon held in the Beijing Economic-Technological Development Area (E-Town), where robots and human runners race the same day on separate lanes over about 21 kilometers. The first edition, in April 2025, was billed as the world's first: Tiangong Ultra, from the Beijing Humanoid Robot Innovation Center, won in 2:40:42, with Noetix's N2 second. The second edition, on April 19, 2026, drew more than 100 robots; remote-controlled entries had their times multiplied by 1.2 so they could be ranked together with autonomously navigating ones. HONOR's autonomous “Lightning” won with a net time of 50:26, and HONOR swept the top three; a remote-controlled “Lightning” actually crossed the finish line first, in 48:19 net, but after the 1.2× conversion it no longer ranked first. The winning time beat the human men's half-marathon world record then standing, 57:20, set by Ugandan runner Jacob Kiplimo in Lisbon in March 2026 — a widely cited 56:42 was actually his 2025 Barcelona time, which was never ratified by World Athletics. The race tests stability under sustained exertion, motor thermal management, and battery life, not manipulation skill — worth remembering when reading the results.","example":"In the 2026 race, the winning autonomous “Lightning” covered the full 21.0975-kilometer course, swapping its battery once, at the 10.6-kilometer mark.","related":["World Humanoid Robot Games","Tiangong","Honor Lightning Humanoid Robot","Noetix N2","Beijing Humanoid Robot Innovation Center","Battery Runtime"]},{"id":"cmg-world-robot-competition-mecha-fighting-series","category":"industry","sec":4,"tier":3,"sources":[{"title":"Robot combat - Wikipedia","url":"https://en.wikipedia.org/wiki/Robot_combat"}],"as_of":"2025","related_ids":["unitree-g1","unitree-robotics","ultimate-robot-knock-out-legend","teleoperation","fall-recovery","commercial-robot-performances"],"name":"CMG World Robot Competition — Mecha Fighting Series","alt":"CMG 机甲格斗擂台赛","abbr":"","aliases":["CMG Robot Combat Competition","Mecha Fighting Arena"],"one_liner":"A CCTV-organized humanoid-robot combat competition where Unitree G1 units are remotely operated to fight.","explanation":"CMG stands for China Media Group, the English name of China's central broadcaster. The Mecha Fighting Arena is one event within CMG's “World Robot Competition” series, held in Hangzhou in 2025 and broadcast on television; competitors used Unitree G1 humanoid robots, remotely controlling them to throw punches and kicks against each other. What the match mainly showcases is the robot's balance, impact resistance, and ability to autonomously get back up after falling — not autonomous decision-making, since every move is directed by a human controller. It brought humanoid robots into mainstream public view and turned “robot combat” into its own genre of paid performance and competition, later followed by leagues such as URKL. When watching this kind of event, it's important to separate what's teleoperated from what's autonomous.","example":"When a robot is knocked down in a match, it gets back up on its own using a fall-recovery policy — an ability trained with reinforcement learning in simulation.","related":["Unitree G1","Unitree Robotics","Ultimate Robot Knock-out Legend","Teleoperation","Fall Recovery","Commercial Robot Performances"]},{"id":"ultimate-robot-knock-out-legend","category":"industry","sec":4,"tier":3,"sources":[{"title":"众擎机器人官网","url":"https://www.engineai.com.cn/"}],"as_of":"2025-12","related_ids":["cmg-world-robot-competition-mecha-fighting-series","engineai","engineai-t800","teleoperation","commercial-robot-performances","world-humanoid-robot-games"],"name":"Ultimate Robot Knock-out Legend","alt":"URKL 格斗联赛","abbr":"URKL","aliases":["URKL"],"one_liner":"A humanoid-robot combat league launched by EngineAI, with competitors piloting robots into a fighting ring.","explanation":"URKL is a humanoid-robot combat league launched by Shenzhen-based EngineAI: competitors control humanoid robots that fight in a ring, with the winner decided by knockdown or points. This kind of event is similar to the CMG Mecha Fighting Arena, mainly aimed at showcasing a robot's fall resistance, balance, and quick movement, while also driving brand exposure and businesses like paid appearances and rental. It's worth noting that robots in the ring are typically directed by a human through teleoperation or a handheld controller, with a lower-level control policy handling the movement itself — the robot isn't deciding how to fight on its own — so when watching match footage, it's important to separate what the human is directing from what the robot's own ability provides.","example":"","related":["CMG World Robot Competition — Mecha Fighting Series","EngineAI","EngineAI T800","Teleoperation","Commercial Robot Performances","World Humanoid Robot Games"]},{"id":"roboleague","category":"industry","sec":4,"tier":3,"sources":[{"title":"RoboCup 官网（机器人足球赛事对照）","url":"https://www.robocup.org/"}],"as_of":"2025-06","related_ids":["robocup","world-humanoid-robot-games","booster-robotics-t1","multi-robot-collaboration","fall-recovery"],"name":"RoboLeague","alt":"RoboLeague 机器人足球联赛","abbr":"","aliases":["RoboLeague Robot Football League"],"one_liner":"A Chinese fully-autonomous humanoid robot soccer league where the robots see, position, and shoot on their own.","explanation":"RoboLeague is a humanoid-robot soccer league held in China. As reported, a 3-on-3 match took place in June 2025 in Beijing's Yizhuang tech zone, with several university teams using the same small humanoid-robot platform; throughout the match the robots made all decisions autonomously, with no human remote control, and the event also served as a warm-up for the World Humanoid Robot Games. It tests perception (finding the ball, recognizing teammates), locomotion control (running, getting back up after falling), and multi-robot coordination, and the real competition is mainly between each team's own algorithms. It's worth comparing with the much older RoboCup.","example":"As reported, Tsinghua University's Huoshen team won the June 2025 RoboLeague match in Beijing's Yizhuang tech zone.","related":["RoboCup","World Humanoid Robot Games","Booster Robotics T1","Multi-robot Collaboration","Fall Recovery"]},{"id":"robocup","category":"industry","sec":4,"tier":3,"sources":[{"title":"RoboCup - Wikipedia","url":"https://en.wikipedia.org/wiki/RoboCup"},{"title":"RoboCup Federation official site","url":"https://www.robocup.org/"}],"as_of":"2026-09","related_ids":["roboleague","booster-robotics-t1","softbank-robotics-nao","op3-soccer","multi-robot-collaboration","robomaster-robocon-university-robotics-competitions"],"name":"RoboCup","alt":"RoboCup 机器人世界杯","abbr":"","aliases":[],"one_liner":"An annual international robotics competition centered on robot soccer, held since 1997.","explanation":"RoboCup is an annual international robotics competition held since 1997, organized by the RoboCup Federation, with its first edition in Nagoya, Japan. Its long-term goal is that by the middle of this century, a team of fully autonomous humanoid robots will beat that year's human World Cup champions under official FIFA rules. The competition centers on robot soccer, divided into humanoid, standard-platform, small-size, mid-size, and simulation leagues, among others, with additional rescue, home-service (@Home), industrial, and youth events. It's a public proving ground for bipedal walking, multi-robot coordination, and real-time perception and decision-making, and a platform many locomotion-control and multi-agent research groups build on.","example":"At RoboCup 2025 in Brazil, Tsinghua University's Huoshen team won the humanoid adult league using the Booster T1 platform from Booster Robotics; in the standard-platform league, every team competes using the same NAO robot, so only the software differs.","related":["RoboLeague","Booster Robotics T1","SoftBank Robotics NAO","OP3 Soccer (DeepMind)","Multi-robot Collaboration","RoboMaster / ROBOCON University Robotics Competitions"]},{"id":"robomaster-robocon-university-robotics-competitions","category":"industry","sec":4,"tier":3,"sources":[{"title":"RoboMaster 官网","url":"https://www.robomaster.com/"},{"title":"ABU Robocon - Wikipedia","url":"https://en.wikipedia.org/wiki/ABU_Robocon"}],"as_of":"","related_ids":["dji","robocup","embedded-software-development","stmicroelectronics-stm32-mcu-family","darpa-robotics-challenge"],"name":"RoboMaster / ROBOCON University Robotics Competitions","alt":"RoboMaster / ROBOCON 大学生机器人竞赛","abbr":"","aliases":["Robomaster Robotics Competition","ABU Robocon"],"one_liner":"China's two best-known university robotics competitions, where many robotics engineers get their start.","explanation":"RoboMaster is organized by drone maker DJI: university teams design and build their own robots and compete in shooting-based combat, involving mechanical design, electronic control, vision-based aiming, and embedded development. ROBOCON traces back to ABU Robocon, run by the Asia-Pacific Broadcasting Union, with a different task each year — throwing balls, carrying objects, and so on — and China's national qualifying rounds co-organized by China Central Television and others. Both require students to build a working robot completely from scratch, exercising full-system engineering skills. Embodied-AI companies hiring engineers often treat this kind of competition experience as a reference for hands-on ability.","example":"Many core engineers at Chinese robotics startups took part in RoboMaster teams as undergraduates.","related":["DJI","RoboCup","Embedded Software Development","STMicroelectronics STM32 MCU Family","DARPA Robotics Challenge"]},{"id":"darpa-robotics-challenge","category":"industry","sec":4,"tier":3,"sources":[{"title":"DARPA Robotics Challenge - Wikipedia","url":"https://en.wikipedia.org/wiki/DARPA_Robotics_Challenge"}],"as_of":"","related_ids":["defense-advanced-research-projects-agency","boston-dynamics-atlas","florida-institute-for-human-and-machine-cognition","humanoid-robot","special-purpose-robot"],"name":"DARPA Robotics Challenge","alt":"DARPA 机器人挑战赛","abbr":"DRC","aliases":["DRC"],"one_liner":"A 2012–2015 US DARPA competition for disaster-response humanoid robots.","explanation":"The DARPA Robotics Challenge was launched by the US Defense Advanced Research Projects Agency after the 2011 Fukushima nuclear accident, starting in 2012, requiring robots to complete eight tasks simulating a disaster site: driving and exiting a vehicle, opening a door, turning a valve, cutting a hole in a wall with a drill, and crossing rubble and climbing stairs, among others. DARPA funded Boston Dynamics to build several hydraulic Atlas units, which some teams were given to compete with. The final, held in June 2015 in Pomona, California, saw 25 teams compete for a $3.5 million prize; South Korea's KAIST won with DRC-HUBO, followed by IHMC in second place and Carnegie Mellon University's Tartan Rescue in third. The robots moved slowly and fell frequently during the finals, exposing that era's limits in perception, balance, and autonomy, and the event trained a generation of researchers who went on to shape later humanoid robotics.","example":"Footage of multiple robots falling while exiting vehicles or opening doors during the finals was widely circulated as compilation clips.","related":["Defense Advanced Research Projects Agency","Boston Dynamics Atlas (Hydraulic)","Florida Institute for Human and Machine Cognition","Humanoid Robot","Special-purpose Robot"]},{"id":"amazon-picking-challenge","category":"industry","sec":4,"tier":3,"sources":[{"title":"Analysis and Observations from the First Amazon Picking Challenge (arXiv 1601.05484)","url":"https://arxiv.org/abs/1601.05484"}],"as_of":"2017","related_ids":["order-picking","bin-picking","vacuum-suction-cup","amazon-robotics","benchmark","darpa-robotics-challenge"],"name":"Amazon Picking Challenge","alt":"亚马逊拣选挑战赛","abbr":"APC / ARC","aliases":["Amazon Robotics Challenge","APC","ARC"],"one_liner":"A 2015–2017 Amazon-run competition for robots that pick items off warehouse shelves.","explanation":"Amazon launched this competition to address a step in e-commerce warehouses — pulling various items off shelves — that was still done by hand. The first edition was held at ICRA 2015 with 26 teams. Robots had to autonomously identify and retrieve specified items from shelf bins within a time limit, with items varying widely in shape and material and often obscured or wrapped in transparent packaging. A 2016 edition added the task of placing items onto shelves, and after being renamed the Amazon Robotics Challenge (ARC) in 2017, the competition ended. It exposed the difficulty of perception, grasp planning, and end-effector design, with suction cups combined with deep-learning recognition becoming a common approach, and it's seen as an important driver of later research into warehouse picking and unstructured (bin) grasping.","example":"After the first 2015 event, organizers surveyed the 26 competing teams and published a paper analyzing how mechanical design, perception, and planning choices related to performance.","related":["Order Picking","Bin Picking","Vacuum Suction Cup","Amazon Robotics","Benchmark","DARPA Robotics Challenge"]},{"id":"mass-production","category":"industry","sec":5,"tier":1,"sources":[{"title":"BotQ: A High-Volume Manufacturing Facility for Humanoid Robots (Figure)","url":"https://www.figure.ai/news/botq"},{"title":"Mass production - Wikipedia","url":"https://en.wikipedia.org/wiki/Mass_production"}],"as_of":"2025-03","related_ids":["year-one-of-mass-production","bill-of-materials-cost","shipment-volume","yield-rate-manufacturing-consistency","start-of-production","botq"],"name":"Mass Production","alt":"量产","abbr":"","aliases":["Thousand-Unit Mass Production","Scaled Mass Production"],"one_liner":"Moving a product from a handful of hand-built prototypes to large-batch, consistent manufacturing on a fixed process.","explanation":"This means a product's design is finalized and it is then built in large batches on a standardized line with a stable supply chain, as opposed to hand-assembled prototypes or small runs of a few dozen units. Humanoid-robot companies often publicize milestones like “reaching thousand-unit mass production” or “year one of mass production.” Genuine mass production requires solving for consistency and yield — every unit performing the same way, with few defects — and BOM cost, the total cost of a unit's materials, plus a steady delivery cadence, not just proving the product can be built at all. When evaluating a claim, look at actual shipment volume and who it went to, rather than planned capacity or framework orders; a company's own use of the word “mass production” sometimes just means small-batch deliveries have begun.","example":"In March 2025, Figure announced its self-built humanoid-robot factory, BotQ, saying its first-generation line had a maximum annual capacity of 12,000 units — that's a capacity ceiling, not the same as actual units shipped.","related":["Year One of Mass Production","Bill of Materials Cost","Shipment Volume","Yield Rate / Manufacturing Consistency","Start of Production","BotQ"]},{"id":"year-one-of-mass-production","category":"industry","sec":5,"tier":2,"sources":[{"title":"Humanoid robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Humanoid_robot"}],"as_of":"2025","related_ids":["mass-production","shipment-volume","framework-order-letter-of-intent-order","pilot-small-batch-delivery","start-of-production","yield-rate-manufacturing-consistency"],"name":"Year One of Mass Production","alt":"量产元年","abbr":"","aliases":["Humanoid Robot's “Year One of Mass Production”","Year One of Commercialization"],"one_liner":"A media and industry label for “the year batch production and delivery began,” not an official benchmark.","explanation":"“Year one of mass production” is a phrase commonly used by securities analysts, companies, and the media, not an official statistical category. 2025 is often called humanoid robotics' year one of mass production, because several companies announced production and delivery plans on the order of hundreds to thousands of units. Newcomers reading this term should note that different companies set different bars for what counts as “mass production” — it might mean a production line has been built, a small batch has been delivered, or a framework order has been signed — so the real shipment figures depend on actual delivery data and customers, not the slogan.","example":"","related":["Mass Production","Shipment Volume","Framework Order / Letter-of-Intent Order","Pilot / Small-Batch Delivery","Start of Production","Yield Rate / Manufacturing Consistency"]},{"id":"shipment-volume","category":"industry","sec":5,"tier":2,"sources":[{"title":"IFR World Robotics","url":"https://ifr.org/worldrobotics/"}],"as_of":"","related_ids":["mass-production","framework-order-letter-of-intent-order","scaled-deployment","research-and-education-market","international-federation-of-robotics"],"name":"Shipment Volume","alt":"出货量","abbr":"","aliases":["Units Delivered"],"one_liner":"The number of robots a manufacturer actually ships to customers over a given period.","explanation":"Shipment volume is the quantity of product a manufacturer actually ships out during a given period, usually reported annually or quarterly, and it's one of the most direct measures of how far a robotics company's commercialization has progressed. A few things are worth watching when reading this number: shipping a unit doesn't mean a customer is actually using it (it may be sitting in channel inventory); robots sold for research and education or paid appearances mean something different from robots doing real factory work; and framework orders or letters of intent are only agreements, not shipments. Industry reports, from the IFR and various consultancies, compile shipment figures across manufacturers, but the definitions used for humanoid robots aren't standardized, so it's worth checking the source closely before citing a number.","example":"","related":["Mass Production","Framework Order / Letter-of-Intent Order","Scaled Deployment","Research & Education Market","International Federation of Robotics"]},{"id":"framework-order-letter-of-intent-order","category":"industry","sec":5,"tier":3,"sources":[{"title":"Letter of intent - Wikipedia","url":"https://en.wikipedia.org/wiki/Letter_of_intent"},{"title":"Framework agreement - Wikipedia","url":"https://en.wikipedia.org/wiki/Framework_agreement"}],"as_of":"","related_ids":["shipment-volume","mass-production","public-tender-winning-bid-orders","closed-commercial-loop","pilot-small-batch-delivery","embodied-ai-bubble"],"name":"Framework Order / Letter-of-Intent Order","alt":"框架订单 / 意向订单","abbr":"","aliases":["Framework Agreement","Letter of Intent (LOI) Order"],"one_liner":"An order laying out a partnership framework or purchase intent, not the same as actual delivery and revenue.","explanation":"A framework order is when buyer and seller first sign a framework agreement setting out price, specifications, the partnership period, and an expected total volume, with actual quantities and delivery timing determined by formal orders placed later in batches; a letter-of-intent order comes from a letter of intent (LOI), expressing a wish to buy, usually with little or no legal binding force. Humanoid-robot companies often announce, at launch events or fundraising rounds, that they've “won a framework order worth several hundred million RMB” or “thousands of units in intent orders” — these numbers show customers are interested, but don't mean the robots have actually been delivered, accepted, or booked as revenue. When reading this kind of news, it's worth distinguishing between the framework or intent amount, a formal purchase contract, actual units delivered, and revenue confirmed in financial statements — only the latter reflect real commercial progress.","example":"Suppose a company announces an RMB 100 million framework agreement, but the first formal purchase order covers only a few dozen units, with the rest contingent on how the trial goes.","related":["Shipment Volume","Mass Production","Public Tender / Winning-Bid Orders","Closed Commercial Loop","Pilot / Small-Batch Delivery","Embodied AI Bubble"]},{"id":"engineering-productionization","category":"industry","sec":5,"tier":2,"sources":[{"title":"Wikipedia: Technology readiness level","url":"https://en.wikipedia.org/wiki/Technology_readiness_level"},{"title":"Embodied Intelligence 2026: Farewell to Narrative Hype, Practical Deployment Reigns Supreme (36Kr)","url":"https://eu.36kr.com/en/p/3953394550537606"}],"as_of":"","related_ids":["mass-production","yield-rate-manufacturing-consistency","real-world-deployment","number-of-nines","mean-time-between-failures","on-device-edge-deployment"],"name":"Engineering / Productionization","alt":"工程化","abbr":"","aliases":["Productization"],"one_liner":"Turning a technology that works in the lab into a reliable, mass-producible, maintainable product.","explanation":"In the industry there's a common saying that a gap called “engineering” sits between a paper and a product. A policy that hits an 80% success rate in a lab demo can get published; getting it into a factory might require 99%-plus reliability, hours of continuous operation without error, working the same way on another unit of the same robot model, and being quickly repairable when something breaks. Productionization is the work that closes that gap: improving reliability and consistency, handling the long tail of edge-case failures, doing inference optimization and on-device deployment so the model runs directly on the robot's own chip, building out calibration, monitoring, remote maintenance, and OTA updates, and controlling cost and yield along the way. This is often where embodied-AI companies actually separate from one another, which is why “engineering capability” shows up so often in funding pitches and job postings.","example":"A clothes-folding policy only needs to succeed once in a lab demo; putting it into a laundry plant means it has to run 24 hours a day, retry automatically after a failure, and be deployed and updated across dozens of robots at once — all of that counts as productionization.","related":["Mass Production","Yield Rate / Manufacturing Consistency","Real-world Deployment","Number of Nines (Reliability)","Mean Time Between Failures","On-Device / Edge Deployment"]},{"id":"yield-rate-manufacturing-consistency","category":"industry","sec":5,"tier":2,"sources":[{"title":"First pass yield - Wikipedia","url":"https://en.wikipedia.org/wiki/First_pass_yield"}],"as_of":"","related_ids":["mass-production","bill-of-materials-cost","engineering-productionization","joint-actuator-module","joint-zero-position-calibration","mean-time-between-failures"],"name":"Yield Rate / Manufacturing Consistency","alt":"良率 / 一致性","abbr":"","aliases":["Yield / Consistency"],"one_liner":"Yield is the share of units that pass inspection; consistency is how little individual units differ from each other.","explanation":"Yield rate is the proportion of manufactured units that pass inspection on the first try; manufacturing consistency is how much individual units of the same model differ from each other in dimensions, motor parameters, sensor readings, and joint zero positions, among other things. Both are key thresholds on the path from prototype to mass production: low yield drives up cost, while poor consistency means a policy tuned to work on one robot may fail on another unit of the exact same model, and also raises calibration and after-sales costs. The consistency of core components — joint modules, dexterous hands, and reducers — draws particular attention.","example":"If torque constants vary too much between motors in the same batch of joint modules, control parameters tuned on one robot can produce distorted motion when moved to another.","related":["Mass Production","Bill of Materials Cost","Engineering / Productionization","Joint Actuator Module","Joint Zero-Position Calibration (Homing / Offset Calibration)","Mean Time Between Failures"]},{"id":"bill-of-materials-cost","category":"industry","sec":5,"tier":2,"sources":[{"title":"Bill of materials - Wikipedia","url":"https://en.wikipedia.org/wiki/Bill_of_materials"}],"as_of":"","related_ids":["mass-production","core-components","joint-actuator-module","per-unit-content-value","domestic-substitution","10-000-yuan-class-humanoid-robot"],"name":"Bill of Materials Cost","alt":"BOM 成本","abbr":"BOM","aliases":["BOM Cost","BOM"],"one_liner":"The total cost of every component and raw material needed to build one unit of a product.","explanation":"A bill of materials (BOM) lists every part that goes into a product, along with quantity and unit price, and adding it all up gives the BOM cost. It excludes R&D, assembly labor, distribution, and after-sales expenses, but it's typically the largest share of hardware cost. In a humanoid robot's BOM, joint actuators (motors, reducers, screws), dexterous hands, sensors, the main compute chip, and the battery account for a large portion. BOM cost determines how cheaply a complete robot can be sold and whether it can reach mass production and the consumer market, which is why “cutting the BOM” is a central topic for both body makers and the domestic-substitution push in components.","example":"Tearing down a humanoid robot and pricing out its dozens of joint modules, its two dexterous hands, its cameras, and its Jetson compute module one by one is exactly what a BOM analysis does.","related":["Mass Production","Core Components","Joint Actuator Module","Per-Unit Content Value","Domestic Substitution","10,000-Yuan-Class Humanoid Robot"]},{"id":"domestic-substitution","category":"industry","sec":5,"tier":2,"sources":[{"title":"Import substitution industrialization - Wikipedia","url":"https://en.wikipedia.org/wiki/Import_substitution_industrialization"}],"as_of":"","related_ids":["chokepoint-technology","core-components","bill-of-materials-cost","big-four-of-industrial-robotics","leaderdrive","upstream-midstream-downstream-of-the-industry-chain"],"name":"Domestic Substitution","alt":"国产替代","abbr":"","aliases":["Localization"],"one_liner":"Replacing components or equipment that used to depend on imports with products from Chinese makers.","explanation":"This refers to switching from foreign-sourced supply to domestic products, with the “localization rate” measuring how much of a given product is now domestic. It's driven by two things: supply-chain security, and cost — domestic parts are typically cheaper and have shorter lead times. In robotics, harmonic reducers, servo motors, controllers, sensors, main-control chips, and complete industrial-robot bodies are the main targets, with examples like Leaderdrive's harmonic reducers and Inovance's servo systems. Domestic substitution and chokepoint technology are two sides of the same coin, and it's also a key route to cutting a humanoid robot's BOM cost.","example":"A Chinese humanoid-robot maker swapping an imported harmonic reducer in a joint for a domestic one is a case of domestic substitution.","related":["Chokepoint Technology","Core Components","Bill of Materials Cost","Big Four of Industrial Robotics","Leaderdrive","Upstream / Midstream / Downstream of the Industry Chain"]},{"id":"chokepoint-technology","category":"industry","sec":5,"tier":2,"sources":[{"title":"Made in China 2025 - Wikipedia","url":"https://en.wikipedia.org/wiki/Made_in_China_2025"}],"as_of":"","related_ids":["domestic-substitution","core-components","strain-wave-gear","nvidia-jetson","upstream-midstream-downstream-of-the-industry-chain"],"name":"Chokepoint Technology","alt":"卡脖子","abbr":"","aliases":[],"one_liner":"A key technology or component that depends on imports and is hard to replace if supply is ever cut off.","explanation":"In Chinese usage, this phrase describes a link in the supply chain that depends heavily on foreign sources, where a restriction from the supplying country could effectively strangle the industry — the image is literally a hand “choking someone's neck.” In robotics, commonly cited chokepoints include high-end AI chips and GPUs, high-precision reducers, high-end servo motors and encoders, high-precision force sensors, and industrial software such as simulators and CAD tools. These are areas with high technical barriers and long validation cycles, so it takes time for Chinese makers to catch up. Identifying chokepoint technologies is the starting point for government support, investment strategy, and domestic substitution.","example":"Robot compute platforms commonly use NVIDIA's Jetson-series chips; if export restrictions applied, makers would need to find a domestic chip alternative — a classic chokepoint discussion.","related":["Domestic Substitution","Core Components","Strain Wave Gear (Harmonic Drive)","NVIDIA Jetson","Upstream / Midstream / Downstream of the Industry Chain"]},{"id":"tier-1-supplier","category":"industry","sec":5,"tier":3,"sources":[{"title":"Wikipedia: Automotive industry supply chain / Tier 1","url":"https://en.wikipedia.org/wiki/Original_equipment_manufacturer"}],"as_of":"","related_ids":["tesla-supply-chain","upstream-midstream-downstream-of-the-industry-chain","robot-body-maker","sampling-supplier-nomination","core-components","tuopu-group"],"name":"Tier-1 Supplier","alt":"一级供应商（Tier 1）","abbr":"Tier 1","aliases":["Tier 1"],"one_liner":"A component or module maker that supplies a complete-robot maker directly, one tier up from raw parts.","explanation":"Tier 1 is supply-chain grading carried over from the auto industry: a supplier that delivers a subassembly or module directly to an OEM (a company that designs, assembles, and sells the complete product) is a Tier-1 supplier, and whoever supplies that Tier-1 supplier is Tier 2, and so on down the chain. Humanoid robotics has adopted the same language: the embodiment maker plays the role of the OEM, companies supplying it with joint modules, dexterous hands, or complete actuator assemblies are Tier 1, and companies supplying those Tier-1 makers with individual parts like reducers, lead screws, and motors are Tier 2. It matters because it determines who wins the order and who bears quality and delivery responsibility; when the secondary market discusses the “Tesla chain” or a “nomination,” it's often really asking whether a company has become a given embodiment maker's Tier-1 supplier.","example":"As reported, Tuopu Group supplies complete actuator assemblies for Tesla's Optimus, and is accordingly often described as a Tier-1 candidate for Optimus.","related":["Tesla (Optimus) Supply Chain","Upstream / Midstream / Downstream of the Industry Chain","Robot Body Maker","Sampling / Supplier Nomination","Core Components","Tuopu Group"]},{"id":"sampling-supplier-nomination","category":"industry","sec":5,"tier":3,"sources":[{"title":"Production part approval process - Wikipedia","url":"https://en.wikipedia.org/wiki/Production_part_approval_process"}],"as_of":"","related_ids":["start-of-production","tesla-supply-chain","tier-1-supplier","per-unit-content-value","engineering-design-production-validation-test","humanoid-robot-concept-stocks"],"name":"Sampling / Supplier Nomination","alt":"送样 / 定点","abbr":"","aliases":["Design-In"],"one_liner":"A supplier sends samples for testing, and once approved, is formally named as the supplier for a project.","explanation":"This pair of terms comes from the automotive supply chain. Sampling is when a component maker delivers samples to an automaker or a complete-robot maker for testing and validation; nomination (or “design-in”) is when the customer formally notifies a supplier that it has been selected to supply a specific part for a specific project, after which the two sides co-develop the part, and once mass production begins, batch supply follows. The humanoid-robot industry has adopted the same language: listed component makers often say in announcements or investor briefings that they've “sampled to a certain customer” or “received a nomination.” Sampling doesn't guarantee an order, and a nomination doesn't guarantee the eventual production volume, so it's worth reading news carefully to see exactly which stage a company is at.","example":"A lead-screw maker announces that it has sampled its product to several humanoid-robot customers, but has not yet received a nomination.","related":["Start of Production","Tesla (Optimus) Supply Chain","Tier-1 Supplier","Per-Unit Content Value","Engineering / Design / Production Validation Test","Humanoid Robot Concept Stocks"]},{"id":"engineering-design-production-validation-test","category":"industry","sec":5,"tier":3,"sources":[{"title":"Engineering validation test - Wikipedia","url":"https://en.wikipedia.org/wiki/Engineering_validation_test"}],"as_of":"","related_ids":["mass-production","start-of-production","yield-rate-manufacturing-consistency","sampling-supplier-nomination","bill-of-materials-cost","engineering-productionization"],"name":"Engineering / Design / Production Validation Test","alt":"EVT / DVT / PVT（样机验证阶段）","abbr":"EVT / DVT / PVT","aliases":["EVT / DVT / PVT"],"one_liner":"The three hardware-validation gates from prototype to mass production: engineering, design, and production.","explanation":"EVT, DVT, and PVT are hardware development stages standard in the consumer-electronics industry, and robotics companies have adopted them too. EVT (engineering validation test) uses hand-built or small-batch prototypes to confirm that a function or performance target can actually be achieved; DVT (design validation test) runs reliability, environmental, drop, lifespan, and certification testing, and once it passes, the design is essentially frozen; PVT (production validation test) uses the actual production line and tooling to run a small trial batch, checking yield, cycle time, and consistency. Only after this does mass production (MP, also called SOP) begin. Any stage can send the design back for rework if problems turn up. When a humanoid-robot company says it has “entered DVT” or “completed PVT,” that's a more informative signal of how close it is to real mass production than a “launch” or “unveiling” announcement.","example":"A humanoid-robot company doing joint-lifespan testing during DVT discovers its reducer wears out too quickly, and has to switch to a different model before re-testing.","related":["Mass Production","Start of Production","Yield Rate / Manufacturing Consistency","Sampling / Supplier Nomination","Bill of Materials Cost","Engineering / Productionization"]},{"id":"start-of-production","category":"industry","sec":5,"tier":3,"sources":[{"title":"Advanced product quality planning - Wikipedia","url":"https://en.wikipedia.org/wiki/Advanced_product_quality_planning"}],"as_of":"","related_ids":["mass-production","engineering-design-production-validation-test","sampling-supplier-nomination","yield-rate-manufacturing-consistency","data-collection-sop"],"name":"Start of Production","alt":"SOP（量产启动）","abbr":"SOP","aliases":["SOP"],"one_liner":"The point when a product formally begins mass manufacturing, a term borrowed from the auto industry.","explanation":"In the auto industry, SOP marks the point when a given vehicle model or component formally begins batch manufacturing, preceded by prototype validation stages (EVT/DVT/PVT) and component approvals, after which volume ramps up gradually. The humanoid-robot industry has adopted the same language, and a manufacturer announcing “SOP in [month]” means its line has started shipping at mass-production standards. Note that the same abbreviation is also used in quality management for “standard operating procedure,” a different thing — for example, a “collection SOP” in data-collection work refers to the latter, so context matters for telling them apart.","example":"A component maker says its humanoid-robot joint module is scheduled to reach SOP next year.","related":["Mass Production","Engineering / Design / Production Validation Test","Sampling / Supplier Nomination","Yield Rate / Manufacturing Consistency","Data Collection SOP"]},{"id":"botq","category":"industry","sec":5,"tier":3,"sources":[{"title":"Figure: BotQ, a new humanoid robot manufacturing facility","url":"https://www.figure.ai/news/botq"}],"as_of":"2025-03","related_ids":["figure-ai","figure-03","mass-production","robofab","figure-helix","yield-rate-manufacturing-consistency"],"name":"BotQ","alt":"Figure BotQ 工厂","abbr":"","aliases":["Figure AI Manufacturing Facility"],"one_liner":"Figure AI's own humanoid-robot manufacturing plant, with a first line designed for 12,000 units a year.","explanation":"BotQ is the in-house manufacturing plant that US humanoid-robot company Figure AI announced on March 15, 2025. The company says its first-generation line can produce up to 12,000 humanoid robots a year, and that it's building a supply chain able to support 100,000 robots (or 3 million actuators) within four years. Figure also plans to put its own humanoid robots to work on the line itself doing parts assembly and material handling — “robots building robots” — with that share expected to grow over time. Its significance lies in showing humanoid robots moving from hand-assembled prototypes toward factory-scale production, where cost, yield, and consistency have to be engineered, not hand-tuned. Figure's third-generation robot, Figure 03, was designed specifically to be mass-manufactured at BotQ.","example":"","related":["Figure AI","Figure 03","Mass Production","RoboFab","Figure Helix","Yield Rate / Manufacturing Consistency"]},{"id":"robofab","category":"industry","sec":5,"tier":3,"sources":[{"title":"Agility Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Agility_Robotics"}],"as_of":"2026-09","related_ids":["agility-robotics","agility-robotics-digit","botq","mass-production","robot-as-a-service","shipment-volume"],"name":"RoboFab","alt":"Agility RoboFab 工厂","abbr":"","aliases":["Agility Robotics Factory"],"one_liner":"Agility Robotics' dedicated factory in Oregon for building its Digit humanoid robot.","explanation":"RoboFab is the factory US humanoid-robot company Agility Robotics built and owns in Salem, Oregon, announced in 2023, dedicated to producing its bipedal humanoid robot, Digit, and often cited as one of the earliest purpose-built humanoid-robot manufacturing plants. Its planned capacity is reported at up to 10,000 units a year, though actual early output has been well below that. Its significance is in showing humanoid robots starting to move from hand-assembled lab units toward factory-scale production; together with Figure's BotQ and Tesla's Optimus line, it's often cited in discussions of how humanoid-robot mass production is progressing.","example":"Digit units built at RoboFab are deployed to warehouses at logistics companies such as GXO on a robot-as-a-service basis to move parts bins.","related":["Agility Robotics","Agility Robotics Digit","BotQ","Mass Production","Robot-as-a-Service","Shipment Volume"]},{"id":"real-world-deployment","category":"industry","sec":6,"tier":1,"sources":[{"title":"Agility Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Agility_Robotics"}],"as_of":"2024-06","related_ids":["proof-of-concept","pilot-small-batch-delivery","scaled-deployment","factory-pilot-deployment","closed-commercial-loop","killer-app"],"name":"Real-world Deployment","alt":"场景落地","abbr":"","aliases":["Application Scenario","Scenario Deployment"],"one_liner":"A robot moves into an actual factory, warehouse, store, or similar setting and does real work that creates value.","explanation":"“Deployment” means a technology moves out of the lab and into real business operations; “scenario” means the specific use case, such as moving totes at a car factory, picking in a warehouse, or giving tours in a mall. Embodied-AI companies' fundraising and marketing often center on which scenarios they've deployed into, because being able to work reliably in a real environment over time, with a customer willing to pay for it, is what actually demonstrates the technology works. Deployment usually passes through a few stages — proof of concept, a small-scale demonstration that something is feasible, pilot or small-batch delivery, and scaled deployment — with the hard parts being reliability, cycle time (how fast one unit of work gets done), and return on investment. When reading news about this, it's worth distinguishing a showcase-style in-factory pilot from paid, ongoing operation.","example":"Agility Robotics' Digit humanoid robot is reportedly deployed in logistics company GXO's warehouses under a robots-as-a-service (RaaS) model, where the customer pays to use it by subscription rather than buying it outright.","related":["Proof of Concept","Pilot / Small-Batch Delivery","Scaled Deployment","Factory Pilot Deployment","Closed Commercial Loop","Killer App"]},{"id":"proof-of-concept","category":"industry","sec":6,"tier":3,"sources":[{"title":"Proof of concept - Wikipedia","url":"https://en.wikipedia.org/wiki/Proof_of_concept"}],"as_of":"","related_ids":["pilot-small-batch-delivery","scaled-deployment","real-world-deployment","factory-pilot-deployment","return-on-investment-payback-period","cycle-time-units-per-hour"],"name":"Proof of Concept","alt":"概念验证","abbr":"PoC","aliases":["PoC"],"one_liner":"A small-scale trial proving a solution can actually work in a customer's real scenario.","explanation":"Proof of concept (PoC) is the process of using a small-scale trial to demonstrate that a technical approach is workable, done before a formal purchase or investment decision. In robot commercialization, a common approach is for a manufacturer to bring one or two robots to a customer's site and run them on a specific step — loading, sorting, and so on — for a few days to a few weeks, verifying whether success rate, cycle time, and safety meet the bar. Only after passing a PoC does a project move to a pilot or small-batch delivery, and later toward scaled deployment. Many companies' public claims of “partnering with an automaker” are, in reality, still at the PoC stage, some distance from a genuine order.","example":"A logistics company might first let a humanoid robot try moving parts bins in one warehouse for two weeks, tracking success rate and units moved per hour, and only sign a pilot contract once it clears the bar.","related":["Pilot / Small-Batch Delivery","Scaled Deployment","Real-world Deployment","Factory Pilot Deployment","Return on Investment / Payback Period","Cycle Time / Units Per Hour"]},{"id":"pilot-small-batch-delivery","category":"industry","sec":6,"tier":3,"sources":[{"title":"Pilot experiment - Wikipedia","url":"https://en.wikipedia.org/wiki/Pilot_experiment"}],"as_of":"","related_ids":["proof-of-concept","scaled-deployment","factory-pilot-deployment","framework-order-letter-of-intent-order","mass-production","real-world-deployment"],"name":"Pilot / Small-Batch Delivery","alt":"试点 / 小批量交付","abbr":"","aliases":[],"one_liner":"Delivering a few to a few dozen units to select customers first, to validate before scaling up.","explanation":"Pilot and small-batch delivery is the intermediate stage between prototype and scaled deployment: a manufacturer places a handful to a few dozen robots into a customer's factory, warehouse, or store, running them under real working conditions for a period to check success rate, cycle time, failure rate, and return on investment. It comes after proof of concept (PoC) and before scaled deployment. Most of what humanoid-robot companies publicly announce as “delivered” or “entered a factory” is at this stage, and the quantity and real usage time are often undisclosed, so delivery figures in the news should be read carefully to see whether they describe a pilot or a genuine batch order.","example":"A humanoid-robot company announcing it has delivered its first few units to an automaker's factory for material-handling training is a pilot delivery; only once the customer validates it does a batch order typically follow.","related":["Proof of Concept","Scaled Deployment","Factory Pilot Deployment","Framework Order / Letter-of-Intent Order","Mass Production","Real-world Deployment"]},{"id":"factory-pilot-deployment","category":"industry","sec":6,"tier":2,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"UBTech Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/UBTech_Robotics"}],"as_of":"2025","related_ids":["proof-of-concept","pilot-small-batch-delivery","scaled-deployment","real-world-deployment","deployment-data-backflow","cycle-time-units-per-hour"],"name":"Factory Pilot Deployment","alt":"进厂实训","abbr":"","aliases":["In-Factory Trial Run","Shop-Floor Pilot"],"one_liner":"A trial placement of robots on a real factory line to test the job and collect data before wider rollout.","explanation":"“In-factory training” (进厂实训) is a term used in Chinese robotics media and industry for humanoid robots and other embodied-AI products being placed on real production lines — at automakers, electronics assemblers, and similar companies — to pilot specific jobs such as material handling, part feeding, quality inspection, or labeling. It usually sits between proof of concept and scaled deployment: on one hand it tests whether the robot can meet the line's cycle time (the seconds allowed per unit of work), reliability, and safety requirements; on the other, it gathers real-world data to feed back into training. Most of these placements are small-scale trials, not settled orders or genuine worker replacement, so readers should distinguish a pilot from a letter of intent, and both from an actual paid deployment.","example":"Since 2024, Figure has reportedly run line pilots at a BMW plant in the US, and UBTECH's Walker S series has done similar factory trials at several Chinese automakers.","related":["Proof of Concept","Pilot / Small-Batch Delivery","Scaled Deployment","Real-world Deployment","Deployment Data Backflow","Cycle Time / Units Per Hour"]},{"id":"scaled-deployment","category":"industry","sec":6,"tier":2,"sources":[{"title":"IFR World Robotics","url":"https://ifr.org/worldrobotics/"}],"as_of":"","related_ids":["proof-of-concept","pilot-small-batch-delivery","mass-production","return-on-investment-payback-period","data-flywheel","real-world-deployment"],"name":"Scaled Deployment","alt":"规模化部署","abbr":"","aliases":["Large-Scale Deployment"],"one_liner":"Robots going from a handful of pilot units to hundreds or thousands working long-term in real settings.","explanation":"Scaled deployment means robots have moved past proof of concept and small-batch pilots to running in large numbers, long-term and reliably, at a customer's site and actually generating value. It's the key threshold for commercializing embodied AI: a few prototypes succeeding in a demo isn't hard, but hundreds or thousands of units working continuously in real environments require a high enough success rate and reliability, a manageable cost and payback period, and mature operations and remote-monitoring systems. The industry often talks about it alongside “mass production,” but mass production emphasizes being able to build the units, while scaled deployment emphasizes them actually being used, and kept in use. Data generated during deployment feeding back into training is, for many companies, the imagined starting point of a data flywheel.","example":"","related":["Proof of Concept","Pilot / Small-Batch Delivery","Mass Production","Return on Investment / Payback Period","Data Flywheel","Real-world Deployment"]},{"id":"cycle-time-units-per-hour","category":"industry","sec":6,"tier":3,"sources":[{"title":"Takt time - Wikipedia","url":"https://en.wikipedia.org/wiki/Takt_time"},{"title":"Throughput (business) - Wikipedia","url":"https://en.wikipedia.org/wiki/Throughput_(business)"}],"as_of":"","related_ids":["return-on-investment-payback-period","mean-time-between-failures","factory-pilot-deployment","machines-replacing-humans","inference-latency","sorting"],"name":"Cycle Time / Units Per Hour","alt":"节拍 / UPH","abbr":"UPH","aliases":["UPH"],"one_liner":"Cycle time is how many seconds one unit of work takes; UPH is how many units get done per hour.","explanation":"Cycle time is how long it takes a station on a production line to complete one operation, or for an entire line to output one finished unit; UPH (Units Per Hour) is the number of units produced per hour, roughly equal to 3,600 divided by the cycle time in seconds. Manufacturing also uses “takt time,” the target cycle time derived by working backward from customer demand. This pair of metrics is the most direct test a robot faces entering a factory: what the factory cares about isn't whether a robot can do the job at all, but whether it can do it reliably within the required cycle time. Many current humanoid robots and VLA policies move relatively slowly and don't reach 100% success, so their cycle time often lags a skilled worker's — a main reason they're still hard to substitute for humans on the line, and a real-world driver behind work on inference acceleration and action chunking.","example":"If a sorting station has a 6-second cycle time, UPH is 600; if a robot needs 12 seconds per item, it would take two robots to match one worker's output.","related":["Return on Investment / Payback Period","Mean Time Between Failures","Factory Pilot Deployment","Machines Replacing Humans","Inference Latency","Sorting"]},{"id":"number-of-nines","category":"industry","sec":6,"tier":3,"sources":[{"title":"High availability - Wikipedia（「nines」表示法）","url":"https://en.wikipedia.org/wiki/High_availability"}],"as_of":"","related_ids":["success-rate","long-tail-problem","mean-time-between-failures","robustness","mean-time-between-interventions"],"name":"Number of Nines (Reliability)","alt":"几个 9（可靠性）","abbr":"","aliases":["Three Nines / Four Nines"],"one_liner":"Counting the 9s in a percentage to describe reliability or success rate, e.g. 99.9% is “three nines.”","explanation":"“Number of nines” is a shorthand for describing reliability by counting consecutive 9s in a percentage: 99.9% is “three nines,” 99.99% is “four nines.” It became popular first for describing IT system uptime — three nines corresponds to roughly 8.76 hours of downtime a year, four nines to about 53 minutes. In embodied AI it's often used for task success rate: hitting 90% in a lab demo isn't especially hard, but a factory line often demands 99.9% or higher. Each additional nine cuts the failure count by roughly an order of magnitude, and the failures that remain tend to be rare, long-tail cases, which is why each additional nine gets harder to reach than the last — one reason people often say going “from demo to deployment” is so difficult.","example":"If a sorting line picks 10,000 times a day, a 99% success rate means about 100 failures a day requiring human handling, while 99.99% means roughly 1.","related":["Success Rate","Long-tail Problem","Mean Time Between Failures","Robustness","Mean Time Between Interventions"]},{"id":"mean-time-between-failures","category":"industry","sec":6,"tier":3,"sources":[{"title":"Mean time between failures - Wikipedia","url":"https://en.wikipedia.org/wiki/Mean_time_between_failures"}],"as_of":"","related_ids":["mean-time-between-interventions","number-of-nines","mass-production","yield-rate-manufacturing-consistency","return-on-investment-payback-period"],"name":"Mean Time Between Failures","alt":"平均无故障时间","abbr":"MTBF","aliases":["MTBF"],"one_liner":"The average time a repairable machine runs normally between one failure and the next.","explanation":"Mean time between failures (MTBF) is a standard reliability-engineering metric: the average time a repairable piece of equipment runs normally between successive failures, usually estimated as total operating time divided by number of failures, in hours. It's commonly considered alongside mean time to repair (MTTR); together the two determine what fraction of the time equipment is actually usable. For humanoid and other embodied robots, parts like joint modules, reducers, and dexterous-hand tendons are prone to problems, and a low MTBF means frequent downtime for repairs — one of the metrics factory customers care about most when purchasing and deploying robots. Robot-learning papers more commonly report mean intervention interval (how often a human has to take over), a similar idea but tracking a different statistic.","example":"If a robot runs 2,000 hours total and breaks down 4 times, its MTBF is roughly 500 hours.","related":["Mean Time Between Interventions","Number of Nines (Reliability)","Mass Production","Yield Rate / Manufacturing Consistency","Return on Investment / Payback Period"]},{"id":"return-on-investment-payback-period","category":"industry","sec":6,"tier":2,"sources":[{"title":"Investopedia: Payback Period","url":"https://www.investopedia.com/terms/p/paybackperiod.asp"}],"as_of":"","related_ids":["bill-of-materials-cost","cycle-time-units-per-hour","scaled-deployment","machines-replacing-humans","mean-time-between-failures"],"name":"Return on Investment / Payback Period","alt":"投资回报 / 回本周期","abbr":"ROI","aliases":["ROI","Payback Period"],"one_liner":"How long it takes a robot buyer to earn back the purchase cost through the money it saves.","explanation":"Return on investment (ROI) is the ratio of gain to cost; the payback period is how long it takes cumulative gains to catch up with the initial cost. When a company evaluates whether to adopt robots, it weighs total cost — purchase price, deployment and retrofitting, maintenance, electricity — against the wages of the workers it replaces and any productivity gain. Manufacturers buying industrial robots commonly treat the payback period as a key threshold; humanoid robots today are expensive and still fall short of a skilled worker's efficiency and reliability, so the math often doesn't work out yet, which is why the industry frequently asks whether the ROI “pencils out” as a precondition for scaled deployment. Related figures that affect the return include cycle time (UPH) and mean time between failures.","example":"If a robot costs RMB 200,000 in total and replaces one position worth RMB 100,000 a year in wages, the payback period is roughly two years (a simplified estimate that ignores maintenance).","related":["Bill of Materials Cost","Cycle Time / Units Per Hour","Scaled Deployment","Machines Replacing Humans","Mean Time Between Failures"]},{"id":"labor-shortage","category":"industry","sec":6,"tier":2,"sources":[{"title":"Population ageing - Wikipedia","url":"https://en.wikipedia.org/wiki/Population_ageing"},{"title":"International Federation of Robotics","url":"https://ifr.org/"}],"as_of":"","related_ids":["machines-replacing-humans","dirty-dull-and-dangerous-jobs","return-on-investment-payback-period","lights-out-factory","robot-density"],"name":"Labor Shortage","alt":"用工荒 / 劳动力短缺","abbr":"","aliases":["Labor Shortage (Aging Workforce)","Worker Shortage","Hiring Difficulty"],"one_liner":"Factories and similar employers can't find or keep enough workers, often cited as the core case for robots.","explanation":"A labor shortage means significantly more workers are wanted for a job than are willing to take it, and it's common on manufacturing lines and in logistics and caregiving — physically demanding, repetitive work. Causes include an aging population, a shrinking working-age population, and younger workers avoiding factory jobs; China, Japan, South Korea, and Western countries all face versions of this problem. It's the demand argument most often cited in robotics fundraising pitches and government policy documents: as workers become scarcer and more expensive, the payback period for replacing them with robots gets shorter. When reading coverage of this topic, it helps to distinguish a long-term structural shortage from a temporary hiring crunch in one industry, and to trace any headline shortage numbers back to their original source.","example":"Manufacturers cite hiring difficulty and high worker turnover as reasons for bringing in collaborative robots or humanoid robots to handle loading and material transport.","related":["Machines Replacing Humans","Dirty, Dull, and Dangerous (3D) Jobs","Return on Investment / Payback Period","Lights-Out Factory","Robot Density"]},{"id":"machines-replacing-humans","category":"industry","sec":6,"tier":2,"sources":[{"title":"Automation - Wikipedia","url":"https://en.wikipedia.org/wiki/Automation"},{"title":"Technological unemployment - Wikipedia","url":"https://en.wikipedia.org/wiki/Technological_unemployment"}],"as_of":"","related_ids":["labor-shortage","industrial-robot","lights-out-factory","dirty-dull-and-dangerous-jobs","human-robot-mixed-workforce","return-on-investment-payback-period"],"name":"Machines Replacing Humans","alt":"机器换人","abbr":"","aliases":["“Machine for Worker” Substitution","Machines Replacing Human Labor"],"one_liner":"Using robots and automation to replace human workers on production jobs.","explanation":"“Machines replacing humans” (机器换人) is a slogan widely used in China's manufacturing upgrade drive, pushed hard by policy in manufacturing-heavy provinces like Zhejiang and Guangdong starting in the 2010s. The core idea is to use industrial robots and automated lines to replace repetitive, dangerous, hard-to-staff jobs, improving efficiency and quality while reducing dependence on cheap labor. In the past this relied mainly on fixed-program industrial robots, suited only to structured, high-volume work; embodied AI targets the jobs that change setup often or need flexible judgment, letting robots move into work that conventional automation couldn't reach. The phrase also regularly triggers broader public debate about job displacement.","example":"Welding shops at car factories are now run almost entirely by welding robots — an early example of machines replacing humans; some companies are now trying humanoid robots for moving parts bins.","related":["Labor Shortage","Industrial Robot","Lights-Out Factory","Dirty, Dull, and Dangerous (3D) Jobs","Human-Robot Mixed Workforce","Return on Investment / Payback Period"]},{"id":"dirty-dull-and-dangerous-jobs","category":"industry","sec":6,"tier":3,"sources":[{"title":"Dirty, dangerous and demeaning - Wikipedia","url":"https://en.wikipedia.org/wiki/Dirty,_dangerous_and_demeaning"}],"as_of":"","related_ids":["machines-replacing-humans","labor-shortage","special-purpose-robot","inspection-robot","real-world-deployment"],"name":"Dirty, Dull, and Dangerous (3D) Jobs","alt":"3D 工作（脏、累、险）","abbr":"3D","aliases":["3D Jobs"],"one_liner":"Jobs that are dirty, tedious, or dangerous — often seen as where robots will be deployed first.","explanation":"The term traces back to Japan's “3K” (kitanai/dirty, kiken/dangerous, kitsui/grueling), and in English was first framed as dirty, dangerous, and demeaning (or demanding) — describing jobs that are hard to staff and often fall to migrant workers. The robotics industry commonly restates it as dirty, dull, and dangerous, folding in “tedious and repetitive” as well. It's frequently used to argue where robots should go first: these jobs are hard to hire for and have high turnover, replacing them with machines faces less social resistance, and customers are more willing to pay for safety and consistency. In news coverage, “replacing 3D jobs” generally refers to scenarios like foundry grinding and polishing, spray painting, heavy-load transport, chemical- and power-plant inspection, and hazardous-material handling.","example":"Substation inspection, nuclear-plant maintenance, and mining work are commonly listed as 3D scenarios well suited to robot replacement.","related":["Machines Replacing Humans","Labor Shortage","Special-purpose Robot","Inspection Robot","Real-world Deployment"]},{"id":"robot-density","category":"industry","sec":6,"tier":3,"sources":[{"title":"IFR World Robotics","url":"https://ifr.org/worldrobotics/"}],"as_of":"2024-11","related_ids":["international-federation-of-robotics","industrial-robot","machines-replacing-humans","labor-shortage","lights-out-factory"],"name":"Robot Density","alt":"机器人密度","abbr":"","aliases":["Industrial Robot Density"],"one_liner":"The number of industrial robots per 10,000 manufacturing workers, a measure of a country's factory automation.","explanation":"Robot density is a metric used by the International Federation of Robotics (IFR) in its annual World Robotics report, calculated as the number of industrial robots in operation divided by the number of manufacturing workers, multiplied by 10,000. It cancels out the effect of a country's overall size, making it easier to compare factory automation levels across countries. According to IFR data published in 2024, the global average was about 162 units in 2023, with South Korea the highest and China at about 470, already ahead of Germany and Japan. This figure is often cited in discussions of machines replacing humans, labor shortages, and how much room humanoid robots have to enter factories.","example":"IFR data show South Korea's robot density exceeded 1,000 per 10,000 workers in 2023, the highest in the world.","related":["International Federation of Robotics","Industrial Robot","Machines Replacing Humans","Labor Shortage","Lights-Out Factory"]},{"id":"lights-out-factory","category":"industry","sec":6,"tier":3,"sources":[{"title":"Lights out (manufacturing) - Wikipedia","url":"https://en.wikipedia.org/wiki/Lights_out_(manufacturing)"}],"as_of":"","related_ids":["machines-replacing-humans","flexible-manufacturing","industrial-robot","non-standard-automation","big-four-of-industrial-robotics"],"name":"Lights-Out Factory","alt":"黑灯工厂","abbr":"","aliases":["Unmanned Factory","Lights-Out Manufacturing"],"one_liner":"A factory so highly automated it can run without people on-site, even with the lights off.","explanation":"A lights-out factory is one so highly automated it doesn't need people present on the floor, and so can, in principle, keep running with the lights off — hence the term “lights-out manufacturing.” A commonly cited example is Japan's FANUC, reported to have used robots to manufacture robots since 2001, able to run unattended for long stretches. In reality most so-called lights-out factories still need people for maintenance, restocking materials, and handling exceptions; it's more common for only part of a process or workshop to be fully unmanned. The traditional approach relies on fixed-station industrial robots and dedicated equipment, suited to large-volume standardized products; one of the arguments made by embodied-AI companies is that general-purpose robots can cover the flexible steps that dedicated equipment can't handle.","example":"FANUC's robot factory is often cited as a representative lights-out factory: robots assembling robots, unattended overnight.","related":["Machines Replacing Humans","Flexible Manufacturing","Industrial Robot","Non-Standard (Custom) Automation","Big Four of Industrial Robotics"]},{"id":"human-robot-mixed-workforce","category":"industry","sec":6,"tier":3,"sources":[{"title":"UBTECH Walker S Industrial Humanoid Robot","url":"https://www.ubtrobot.com/en/humanoid/products/walker-s"}],"as_of":"","related_ids":["human-robot-collaboration","factory-pilot-deployment","machines-replacing-humans","lights-out-factory","brownfield-deployment"],"name":"Human-Robot Mixed Workforce","alt":"人机混编","abbr":"","aliases":["Human-Robot Mixed Teaming"],"one_liner":"People and robots working side by side on the same line or team, each handling different tasks.","explanation":"A human-robot mixed workforce means robots and human workers are assigned to the same production line or team in a factory or warehouse, each taking on the tasks suited to them, rather than converting an entire line to full automation all at once. It's common in the early stage of humanoid robots entering factories: robots first take over repetitive steps like material handling, loading, and quality inspection, while people handle complex assembly, exception handling, and oversight. This approach requires minimal changes to an existing line, and lets a company gather real-world data and verify reliability, but it also demands safety guarding, human-robot collaboration procedures, and matching cycle times. It sits partway along the same spectrum as “machines replacing humans” and the “lights-out factory.”","example":"When UBTECH's Walker S series of industrial humanoid robots does factory-pilot work, it operates on the same line as human workers, handling steps such as material handling and quality inspection.","related":["Human-Robot Collaboration","Factory Pilot Deployment","Machines Replacing Humans","Lights-Out Factory","Brownfield Deployment"]},{"id":"brownfield-deployment","category":"industry","sec":6,"tier":3,"sources":[{"title":"Figure AI - Wikipedia","url":"https://en.wikipedia.org/wiki/Figure_AI"},{"title":"Greenfield project - Wikipedia","url":"https://en.wikipedia.org/wiki/Greenfield_project"}],"as_of":"","related_ids":["machines-replacing-humans","lights-out-factory","factory-pilot-deployment","humanoid-robot","real-world-deployment","human-robot-mixed-workforce"],"name":"Brownfield Deployment","alt":"棕地部署（存量工厂）","abbr":"","aliases":["Brownfield (Existing-Factory) Deployment"],"one_liner":"Putting robots into an already-built factory designed around people, instead of a purpose-built new plant.","explanation":"“Brownfield” and “greenfield” are a pair of engineering terms: greenfield means building entirely from scratch, where the building and workflow can be designed around automation from day one; brownfield means retrofitting an existing building, production line, or warehouse, where space, workstations, and shelf heights were all designed for people and can't be substantially changed. Most factories are brownfield, and retrofitting them with traditional automation usually means stopping the line for expensive reconstruction. One of the main selling points of humanoid robots and general-purpose mobile manipulators is that they can adapt directly to human-built environments, taking over part of a job without redesigning the line — so whether a company's robot can run reliably in a brownfield setting is an important test of its real deployment ability.","example":"Figure AI's 2024 partnership with BMW, placing humanoid robots into BMW's existing US factory for testing, is an example of brownfield deployment.","related":["Machines Replacing Humans","Lights-Out Factory","Factory Pilot Deployment","Humanoid Robot","Real-world Deployment","Human-Robot Mixed Workforce"]},{"id":"flexible-manufacturing","category":"industry","sec":6,"tier":3,"sources":[{"title":"Flexible manufacturing system - Wikipedia","url":"https://en.wikipedia.org/wiki/Flexible_manufacturing_system"},{"title":"柔性制造系统 - 维基百科","url":"https://zh.wikipedia.org/wiki/柔性制造系统"}],"as_of":"","related_ids":["non-standard-automation","industrial-robot","teach-and-playback-programming","lights-out-factory","general-purpose-robot","machines-replacing-humans"],"name":"Flexible Manufacturing","alt":"柔性制造","abbr":"FMS","aliases":["FMS","Flexible Manufacturing System"],"one_liner":"A production line's ability to switch quickly between products, suited to small batches of many varieties.","explanation":"Flexible manufacturing means a production line can switch quickly between different product models without major equipment changes, suiting it to small-batch, many-variety orders. Flexible manufacturing systems (FMS), which emerged in the 1960s–70s, achieved this with CNC machine tools, automated material handling, and computer scheduling. Traditional industrial robots rely on taught, fixed programs and dedicated fixtures at fixed stations, so switching products means reprogramming and re-tooling, which only pays off at large production volumes. One of embodied AI's core selling points is flexibility: the hope is that robots can adapt to new parts and new procedures using vision and learned policies, with little or no reprogramming, making it possible to enter fast-changing lines such as consumer-electronics assembly and automotive final assembly.","example":"A phone contract manufacturer switches between several models each year; if a robot could learn a new model's assembly motions from just a few demonstrations, changeover time could shrink dramatically.","related":["Non-Standard (Custom) Automation","Industrial Robot","Teach-and-Playback Programming","Lights-Out Factory","General-purpose Robot","Machines Replacing Humans"]},{"id":"non-standard-automation","category":"industry","sec":6,"tier":3,"sources":[{"title":"Automation - Wikipedia","url":"https://en.wikipedia.org/wiki/Automation"}],"as_of":"","related_ids":["system-integrator","project-based-delivery-vs-productization","flexible-manufacturing","industrial-robot","lights-out-factory","general-purpose-robot"],"name":"Non-Standard (Custom) Automation","alt":"非标自动化","abbr":"","aliases":["Non-Standard Automation Equipment"],"one_liner":"Automation equipment custom-designed for the specific needs of one production line.","explanation":"Non-standard automation refers to automation equipment custom-designed and built for the specific needs of a single customer's single production line, as distinct from “standard” products like general-purpose industrial robots or standardized conveyor lines. This kind of project is usually taken on by system integrators or non-standard equipment makers; each project requires fresh mechanical design, electrical work, and programming and commissioning, with long delivery times and experience that's hard to reuse across projects, so the business tends to be project-based. Its connection to embodied AI is this: today, a great many factory processes are still handled by non-standard equipment or by hand, and one of the claims embodied-AI companies make is that general-purpose robots paired with transferable models could replace “a custom equipment set for every line,” turning project-based work into a standardized product.","example":"A machine custom-designed for one battery factory that combines cell loading with appearance inspection is a typical piece of non-standard automation equipment.","related":["System Integrator","Project-Based Delivery vs. Productization","Flexible Manufacturing","Industrial Robot","Lights-Out Factory","General-purpose Robot"]},{"id":"project-based-delivery-vs-productization","category":"industry","sec":6,"tier":3,"sources":[{"title":"Productization - Wikipedia","url":"https://en.wikipedia.org/wiki/Productization"}],"as_of":"","related_ids":["non-standard-automation","system-integrator","general-purpose-robot","scaled-deployment","closed-commercial-loop","engineering-productionization"],"name":"Project-Based Delivery vs. Productization","alt":"项目制 / 产品化","abbr":"","aliases":[],"one_liner":"Custom-building a solution for each customer versus turning it into a standard product sold at scale.","explanation":"Project-based delivery means designing a custom solution for each customer's specific scenario, sending engineers on-site to commission and deliver it, with revenue settled project by project; productization means turning a capability into a standard hardware and software product that a customer can buy and get running with only light configuration, allowing it to be replicated at scale. Project-based work ramps up quickly and individual deals can be large, but it's labor-intensive, hard to scale, and margins are limited; productization requires heavy upfront investment but has low marginal cost. Non-standard automation and system integration in industrial automation mostly fall into the project-based category. What embodied-AI companies are chasing with general-purpose robots is, at heart, using generalization to turn project-based work into a product.","example":"Custom-designing fixtures and vision software for one production line and spending three months on-site commissioning it is project-based work; a robot that, once unboxed, can do many kinds of tasks from natural-language instructions is a productized one.","related":["Non-Standard (Custom) Automation","System Integrator","General-purpose Robot","Scaled Deployment","Closed Commercial Loop","Engineering / Productionization"]},{"id":"to-business","category":"industry","sec":7,"tier":2,"sources":[{"title":"Business-to-business - Wikipedia","url":"https://en.wikipedia.org/wiki/Business-to-business"}],"as_of":"","related_ids":["to-consumer","task-oriented-grasping","real-world-deployment","system-integrator","return-on-investment-payback-period","factory-pilot-deployment"],"name":"To Business (B2B)","alt":"ToB","abbr":"ToB","aliases":["ToB"],"one_liner":"Selling a product to companies, factories, or institutions rather than individual consumers.","explanation":"ToB refers to a business model aimed at enterprise customers. In embodied AI, ToB customers include automakers, logistics warehouses, electronics factories, and universities and research institutes; deals here tend to be large, are evaluated on return on investment and reliability, usually require customization and on-site deployment, and involve a long sales cycle. Most robotics companies today earn most of their actual revenue from ToB work — industrial, logistics, and research and education — because factory settings are relatively structured and easier to deploy into first, compared with homes.","example":"A humanoid-robot company signing a factory-pilot agreement with an automaker to move parts bins on an assembly line is a ToB deal.","related":["To Consumer (B2C)","Task-Oriented Grasping","Real-world Deployment","System Integrator","Return on Investment / Payback Period","Factory Pilot Deployment"]},{"id":"to-consumer","category":"industry","sec":7,"tier":2,"sources":[{"title":"Retail / Business-to-consumer - Wikipedia","url":"https://en.wikipedia.org/wiki/Business-to-consumer"}],"as_of":"","related_ids":["to-business","consumer-grade-robot","companion-robot","household-tasks","robot-vacuum-cleaner"],"name":"To Consumer (B2C)","alt":"ToC","abbr":"ToC","aliases":["ToC"],"one_liner":"Selling a product directly to individual consumers, such as home robots.","explanation":"ToC refers to a business model aimed at individual consumers. In robotics, robot vacuums, robot dogs, and companion robots all fall under ToC. ToC demands a low price, out-of-the-box usability, adequate safety, and a solid after-sales network, and the home environment is highly unstructured — every household's layout and belongings differ — so a general-purpose household humanoid robot is widely considered harder to bring into homes than ToB deployment, even though it's also seen as one of the largest potential markets.","example":"A consumer buying a quadruped robot dog on an e-commerce platform to play with at home is a ToC purchase.","related":["To Business (B2B)","Consumer-Grade Robot","Companion Robot","Household Tasks","Robot Vacuum Cleaner"]},{"id":"to-government","category":"industry","sec":7,"tier":3,"sources":[{"title":"Wikipedia: Business-to-government","url":"https://en.wikipedia.org/wiki/Business-to-government"}],"as_of":"","related_ids":["to-business","to-consumer","public-tender-winning-bid-orders","embodied-ai-training-ground","application-scenario-list-scenario-opening","research-and-education-market"],"name":"To Government (B2G)","alt":"ToG","abbr":"ToG","aliases":["ToG"],"one_liner":"A business model selling to government bodies, state-owned enterprises, and public institutions.","explanation":"ToG refers to a company selling products or services to government agencies and the entities under them, including state-owned enterprises, public universities, research institutes, and local platform companies — standing alongside ToB (selling to businesses) and ToC (selling to individual consumers). In embodied AI, a common form of ToG order is a local government procuring robots to build an embodied-AI training ground, an innovation center, or a pilot-scale testing base, along with research and teaching purchases by universities and vocational schools, usually done through public tender. It's characterized by large individual deal sizes and long payment cycles, and it's often used to support early-stage revenue; the industry also frequently notes that government procurement doesn't mean a robot has actually replaced human labor in a real production setting.","example":"A local government building an embodied-AI training ground procures dozens to over a hundred humanoid robots in one public tender, for use in data collection.","related":["To Business (B2B)","To Consumer (B2C)","Public Tender / Winning-Bid Orders","Embodied AI Training Ground (Robot Data Collection Center)","Application Scenario List / Scenario Opening","Research & Education Market"]},{"id":"public-tender-winning-bid-orders","category":"industry","sec":7,"tier":3,"sources":[{"title":"Call for bids - Wikipedia","url":"https://en.wikipedia.org/wiki/Call_for_bids"}],"as_of":"2026-09","related_ids":["task-oriented-grasping","framework-order-letter-of-intent-order","shipment-volume","research-and-education-market","embodied-ai-training-ground","scenario-owner"],"name":"Public Tender / Winning-Bid Orders","alt":"招投标 / 中标订单（集采）","abbr":"","aliases":["Centralized Procurement"],"one_liner":"State firms, governments, or universities publicly bid out robot purchases, and manufacturers compete for the contract.","explanation":"Public tendering is the legally mandated procurement process used by government bodies, state-owned enterprises, and universities: the buyer publishes a tender notice with its requirements, manufacturers submit bids to compete, and after evaluation, the winning bidder and contract amount are made public. Centralized procurement means several units' needs are combined into one larger tender. Because winning-bid results are published on public resource trading platforms, this is one of the most reliable ways to verify a robot's real order volume from the outside. Many of the largest early humanoid-robot orders have come from tenders by telecom carriers, state-owned enterprises, and research institutions, often for showroom use, training deployments, or data collection.","example":"As reported, in 2025 a subsidiary of China Mobile ran a tender to procure humanoid robots worth about RMB 124 million, won by AgiBot and Unitree, described by media at the time as one of the largest single humanoid-robot orders to date.","related":["Task-Oriented Grasping","Framework Order / Letter-of-Intent Order","Shipment Volume","Research & Education Market","Embodied AI Training Ground (Robot Data Collection Center)","Scenario Owner (End Customer)"]},{"id":"research-and-education-market","category":"industry","sec":7,"tier":2,"sources":[{"title":"Unitree G1 产品页","url":"https://www.unitree.com/g1"}],"as_of":"","related_ids":["edu-edition","secondary-development","shipment-volume","real-world-deployment","closed-commercial-loop"],"name":"Research & Education Market","alt":"科研教育市场","abbr":"","aliases":["Research and Education Sales"],"one_liner":"The market for selling robots to universities, research institutes, and schools for study and teaching.","explanation":"This refers to the segment of the robotics market whose customers are university labs, research institutes, vocational colleges, and primary and secondary schools. These customers don't buy robots to replace labor; they use them to run algorithms, collect data, publish papers, or teach classes, so they place a high value on developability — an open SDK and low-level interfaces — and a low value on reliably getting real work done. At a stage when humanoid and quadruped robots still can't be deployed into factories at scale, the research and education market has been many manufacturers' earliest and most stable source of revenue; companies like Unitree also release dedicated EDU editions aimed at further development. Its drawback is limited scale, so it's generally seen as a transitional source of revenue rather than the destination.","example":"University labs buy the Unitree G1 EDU edition to do reinforcement-learning locomotion control and VLA research.","related":["EDU Edition","Secondary Development (custom development on SDK)","Shipment Volume","Real-world Deployment","Closed Commercial Loop"]},{"id":"edu-edition","category":"industry","sec":7,"tier":2,"sources":[{"title":"Unitree Go2 产品页（规格表）","url":"https://www.unitree.com/go2"},{"title":"Unitree G1 产品页（规格表）","url":"https://www.unitree.com/g1"},{"title":"宇树开发者文档中心","url":"https://support.unitree.com/home/zh/developer"}],"as_of":"2026-09","related_ids":["secondary-development","software-development-kit","research-and-education-market","unitree-go2","unitree-g1","robot-body-maker"],"name":"EDU Edition","alt":"EDU 版（科研教育版）","abbr":"EDU","aliases":["Research & Education Version","Developer Edition"],"one_liner":"A maker's version of a robot aimed at universities and developers, with open low-level access and extra sensors.","explanation":"Chinese robot makers often sell the same robot in several versions, and the one aimed at universities, research institutions, and developers is called the EDU edition. Compared with a consumer or base edition, an EDU edition typically opens up its low-level SDK (software development kit) and joint-level control interfaces, and adds onboard compute, such as an NVIDIA Jetson module, extra sensors, or a dexterous hand, at a noticeably higher price. Research work requires being able to read joint states and send torque or position commands directly to deploy a self-trained policy on real hardware, which is why the real robots used in papers are usually EDU editions. Different makers open up different layers of access, so it's worth checking a maker's spec sheet carefully before buying.","example":"According to Unitree's website, only the EDU edition of the Go2 comes with foot-mounted force sensors; the base edition of the G1 has 23 joints, while the EDU edition, with a dexterous hand and other add-ons, can have up to 43.","related":["Secondary Development (custom development on SDK)","Software Development Kit","Research & Education Market","Unitree Go2","Unitree G1","Robot Body Maker"]},{"id":"consumer-grade-robot","category":"industry","sec":7,"tier":2,"sources":[{"title":"Domestic robot - Wikipedia","url":"https://en.wikipedia.org/wiki/Domestic_robot"}],"as_of":"2025-10","related_ids":["to-consumer","robot-vacuum-cleaner","10-000-yuan-class-humanoid-robot","companion-robot","1x-neo","household-tasks"],"name":"Consumer-Grade Robot","alt":"消费级机器人","abbr":"","aliases":[],"one_liner":"A robot product sold directly to individuals and households, rather than to factories or businesses.","explanation":"This refers to robots sold to and used by ordinary consumers, at home or in personal life, as opposed to industrial-grade or commercial-grade robots sold to factories and businesses. Robot vacuum cleaners are already a mature example, and quadruped robot dogs, desktop robot arms, and companion robots also fall into this category. It demands far more than a research or industrial setting does on price, safety, ease of use, and being maintenance-free, which is why it's simultaneously the hardest and most eagerly anticipated market for humanoid robots. 1X reportedly opened pre-orders for its home humanoid robot, NEO, in 2025.","example":"Robot vacuum cleaners are currently the most widespread consumer-grade robot.","related":["To Consumer (B2C)","Robot Vacuum Cleaner","10,000-Yuan-Class Humanoid Robot","Companion Robot","1X NEO","Household Tasks"]},{"id":"10-000-yuan-class-humanoid-robot","category":"industry","sec":7,"tier":2,"sources":[{"title":"Unitree R1 产品页","url":"https://www.unitree.com/R1"}],"as_of":"2025-10","related_ids":["consumer-grade-robot","small-size-humanoid-robot","unitree-r1","noetix-bumi","price-war","research-and-education-market"],"name":"10,000-Yuan-Class Humanoid Robot","alt":"万元级人形机器人","abbr":"","aliases":[],"one_liner":"A small humanoid robot priced from around RMB 10,000 to several tens of thousands, aimed mainly at individuals and education.","explanation":"Starting in 2025, Chinese makers began pushing humanoid-robot prices down to the ten-thousand-RMB range; these are mostly small units around a meter tall or shorter, with weaker joint torque and payload than full-size models, sold mainly to developers, university teaching programs, science-outreach exhibits, and home entertainment rather than for factory work. Reported examples include Unitree's R1, starting around RMB 39,900, and Noetix's Bumi, about RMB 9,998. These let individuals afford a walking, running, hackable humanoid platform for the first time, though battery life, onboard compute, and dexterity still lag well behind full-size machines. Prices change quickly, so go by a maker's latest listed price.","example":"Noetix's Bumi reportedly sells for RMB 9,998, stands about 94 centimeters tall, and targets education and home users.","related":["Consumer-Grade Robot","Small-size Humanoid Robot","Unitree R1","Noetix Bumi","Price War (Humanoid Robots)","Research & Education Market"]},{"id":"price-war","category":"industry","sec":7,"tier":3,"sources":[{"title":"Price war - Wikipedia","url":"https://en.wikipedia.org/wiki/Price_war"}],"as_of":"2026-09","related_ids":["10-000-yuan-class-humanoid-robot","bill-of-materials-cost","hundred-robot-war","consumer-grade-robot","research-and-education-market","domestic-substitution"],"name":"Price War (Humanoid Robots)","alt":"价格战","abbr":"","aliases":[],"one_liner":"Manufacturers repeatedly undercutting each other's prices to win share of the humanoid-robot market.","explanation":"A price war is when makers of similar products repeatedly cut prices to compete for market share. The humanoid-robot price war has been clearly visible since roughly 2024: Unitree's G1 launched starting at RMB 99,000, and afterward Unitree's R1 and Songyan Power's Bumi, among others, pushed entry-level humanoid prices down to the tens of thousands of RMB, some even under RMB 10,000. Prices have come down partly through domestic supply-chain sourcing, replacing metal with plastic, and smaller designs to cut bill-of-materials cost, and partly as a bid to capture early markets like research/education and paid appearances. Lower-priced models generally have fewer degrees of freedom, less payload, and less onboard compute, and are suited to teaching and development rather than a direct price comparison with full-size factory-grade models.","example":"As reported, Songyan Power's Bumi, launched in October 2025, was priced at RMB 9,998, seen as the moment the humanoid-robot price war dropped below RMB 10,000.","related":["10,000-Yuan-Class Humanoid Robot","Bill of Materials Cost","Hundred-Robot War","Consumer-Grade Robot","Research & Education Market","Domestic Substitution"]},{"id":"going-global","category":"industry","sec":7,"tier":3,"sources":[{"title":"Unitree Robotics 官网（英文）","url":"https://www.unitree.com/"}],"as_of":"","related_ids":["unitree-robotics","consumer-grade-robot","research-and-education-market","consumer-electronics-show","price-war"],"name":"Going Global","alt":"出海","abbr":"","aliases":["Overseas Expansion"],"one_liner":"The common Chinese business term for a company selling products or setting up operations abroad.","explanation":"“Going global” (出海, literally “going out to sea”) is common language in Chinese business, referring to a company expanding its products, services, or manufacturing capacity into overseas markets — exporting complete products, setting up overseas subsidiaries and channels, exhibiting at foreign trade shows, or building factories abroad. In robotics, Chinese manufacturers leverage their supply-chain and price advantages to sell quadruped robots, humanoid robots, and robot arms to overseas university labs, developers, and enterprise customers, which has been an important early revenue source for a number of companies. Going global also means dealing with certification standards, data compliance, after-sales service, and trade-policy hurdles. When reading a company's prospectus or financial reports, its share of overseas revenue is often treated as a measure of how well its global expansion is going.","example":"Unitree Robotics' English-language website sells quadruped and humanoid robots worldwide, and its products are relatively common at overseas research institutions.","related":["Unitree Robotics","Consumer-Grade Robot","Research & Education Market","Consumer Electronics Show","Price War (Humanoid Robots)"]},{"id":"commercial-robot-performances","category":"industry","sec":7,"tier":2,"sources":[{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"}],"as_of":"2025-02","related_ids":["robot-rental","yangbot","pre-programmed-motion","guided-tours-and-reception","unitree-g1"],"name":"Commercial Robot Performances","alt":"商演","abbr":"","aliases":["Robot Shows"],"one_liner":"Renting out humanoid or quadruped robots to perform at commercial events for a fee.","explanation":"This refers to robots performing at mall openings, trade shows, corporate galas, and similar commercial events — dancing, sparring, or interacting with the crowd — with the event organizer paying a fee per show or per day. Demand rose noticeably after a Unitree humanoid performed a yangge folk dance on China Central Television's Spring Festival Gala in 2025. These performances usually run on pre-programmed, choreographed motion, which keeps the technical bar relatively low, yet they remain one of the few ways many humanoid robots currently generate direct revenue, and they've also fueled a robot-rental business. The trend is often used as a talking point for asking when humanoid robots will move beyond performing and start doing real work.","example":"A mall rents two Unitree G1 units to dance at its entrance for an anniversary event, billed by the half-day.","related":["Robot Rental","Yangbot","Pre-Programmed (Choreographed) Motion","Guided Tours & Reception","Unitree G1"]},{"id":"guided-tours-and-reception","category":"industry","sec":7,"tier":3,"sources":[{"title":"Pepper (robot) - Wikipedia","url":"https://en.wikipedia.org/wiki/Pepper_(robot)"}],"as_of":"","related_ids":["service-robot","real-world-deployment","commercial-robot-performances","softbank-robotics-pepper","human-robot-interaction"],"name":"Guided Tours & Reception","alt":"导览接待","abbr":"","aliases":[],"one_liner":"Robots greeting, explaining, and leading the way at exhibition halls, malls, and lobbies.","explanation":"Guided tours and reception refers to robots providing greeting, explanation, Q&A, and wayfinding at exhibition halls, museums, government service centers, shopping malls, and bank branches. It demands little manipulation ability, relying mainly on voice interaction, navigation, and visual presence, which makes it one of the earliest commercialized scenarios for service and humanoid robots. An early representative was SoftBank's Pepper, and in recent years a number of Chinese humanoid robots have also been delivered for showroom explanation and event greeting. Orders in this scenario come relatively quickly, but the real labor value a single unit creates is limited, so the industry generally discusses it separately from “real work” scenarios like factory material handling and sorting when talking about a viable business model.","example":"SoftBank's Pepper, launched in 2015, was deployed in large numbers at stores and banks for greeting and inquiries.","related":["Service Robot","Real-world Deployment","Commercial Robot Performances","SoftBank Robotics Pepper","Human-Robot Interaction"]},{"id":"robot-rental","category":"industry","sec":7,"tier":2,"sources":[{"title":"Robot-as-a-Service 概述（Wikipedia）","url":"https://en.wikipedia.org/wiki/Robot_as_a_service"}],"as_of":"","related_ids":["robot-as-a-service","commercial-robot-performances","guided-tours-and-reception","subscription-model","closed-commercial-loop"],"name":"Robot Rental","alt":"机器人租赁","abbr":"","aliases":["Humanoid Robot Rental"],"one_liner":"Renting a robot out by the day or event instead of selling it outright.","explanation":"Robot rental means the owner leases equipment to a user for a set time or event, common for paid appearances, trade shows, store-opening ceremonies, and guided-tour receptions. Since humanoid robots broke into Chinese public awareness starting in 2025, renting one by the day for events has become a small business of its own, with prices reported to swing widely with a robot's current buzz. Rental lowers the barrier for customers to try a robot, and lets manufacturers and channel partners earn revenue before the robot can handle real physical work. It's close to robot-as-a-service (RaaS), but rental typically provides just the hardware itself, while RaaS more often bundles outcome- or subscription-based service and maintenance.","example":"A shopping mall rents a humanoid robot for a day to dance and interact with customers at its opening event.","related":["Robot-as-a-Service","Commercial Robot Performances","Guided Tours & Reception","Subscription Model","Closed Commercial Loop"]},{"id":"robot-as-a-service","category":"industry","sec":7,"tier":2,"sources":[{"title":"Robot as a service（Wikipedia）","url":"https://en.wikipedia.org/wiki/Robot_as_a_service"}],"as_of":"","related_ids":["robot-rental","subscription-model","return-on-investment-payback-period","scaled-deployment","deployment-data-backflow"],"name":"Robot-as-a-Service","alt":"机器人即服务","abbr":"RaaS","aliases":["RaaS"],"one_liner":"Customers don't buy the robot; they pay by the month or by the amount of work it does.","explanation":"Robot-as-a-service (RaaS) is modeled on the software industry's SaaS (software-as-a-service): the manufacturer keeps ownership of the robot, and the customer pays monthly, hourly, or by the amount of work completed, while the manufacturer handles maintenance, upgrades, and remote monitoring. It converts a customer's large one-time purchase into a predictable operating expense, lowering the barrier to adoption, while also giving the manufacturer continuous access to deployment data to improve its models. Warehouse and logistics robots adopted this model relatively early, and humanoid-robot companies such as Agility Robotics now also supply robots to factories and warehouses on a RaaS basis. The drawback is that the manufacturer must front the hardware cost, which demands strong cash flow.","example":"A warehouse uses a fleet of material-handling robots for a fixed monthly fee, and the manufacturer sends someone to fix any that break down.","related":["Robot Rental","Subscription Model","Return on Investment / Payback Period","Scaled Deployment","Deployment Data Backflow"]},{"id":"subscription-model","category":"industry","sec":7,"tier":3,"sources":[{"title":"1X Technologies 官网","url":"https://www.1x.tech/"}],"as_of":"2025-10","related_ids":["robot-as-a-service","robot-rental","consumer-grade-robot","1x-neo","over-the-air-update","to-consumer"],"name":"Subscription Model","alt":"订阅制","abbr":"","aliases":["Monthly Subscription"],"one_liner":"Paying monthly or yearly to use a robot or its software features, instead of buying it outright.","explanation":"The subscription model is a way of charging where users pay monthly or yearly for the right to use a robot, or for some software capability, while the manufacturer handles maintenance, upgrades, and part of the after-sales support. For users, it lowers the barrier of a one-time purchase; for manufacturers, it produces steadier revenue and lets them keep improving the product through continuous software updates and a data feedback loop. It's close in spirit to robot-as-a-service (RaaS) for enterprise customers, except the subscription model is more often used for consumer-facing products, and also commonly used to charge separately for individual software features.","example":"When 1X opened preorders for its home humanoid robot NEO in 2025, buyers could either pay outright or choose a $499-a-month subscription.","related":["Robot-as-a-Service","Robot Rental","Consumer-Grade Robot","1X NEO","Over-the-Air (OTA) Update","To Consumer (B2C)"]},{"id":"robot-4s-store","category":"industry","sec":7,"tier":3,"sources":[{"title":"4S店 - 百度百科","url":"https://baike.baidu.com/item/4S店"}],"as_of":"2025-08","related_ids":["consumer-grade-robot","robot-rental","to-consumer","real-world-deployment","beijing-humanoid-robot-innovation-center"],"name":"Robot 4S Store","alt":"机器人 4S 店","abbr":"","aliases":["Sales, Service, Spare Parts, Survey Store"],"one_liner":"A robot dealership modeled on car dealerships, combining sales, repair, parts, and feedback in one place.","explanation":"4S was originally auto-dealership terminology for a bundle of four functions: Sale (vehicle sales), Service (after-sales service), Spare parts, and Survey (customer feedback). As reported, China's first “embodied-AI robot 4S store” opened in Beijing's Yizhuang tech zone in August 2025, displaying humanoid robots, robot dogs, and other products from multiple manufacturers and offering purchase, rental, and repair services. It reflects the industry's push to sell robots to ordinary consumers and small and mid-sized customers, for whom after-sales repair and spare-parts supply are exactly the concerns this kind of store is meant to address.","example":"","related":["Consumer-Grade Robot","Robot Rental","To Consumer (B2C)","Real-world Deployment","Beijing Humanoid Robot Innovation Center"]},{"id":"embodied-ai-robot-insurance","category":"industry","sec":7,"tier":3,"sources":[{"title":"Liability insurance - Wikipedia","url":"https://en.wikipedia.org/wiki/Liability_insurance"},{"title":"Product liability - Wikipedia","url":"https://en.wikipedia.org/wiki/Product_liability"}],"as_of":"2025","related_ids":["robot-rental","commercial-robot-performances","embodied-safety","functional-safety","scaled-deployment"],"name":"Embodied-AI Robot Insurance","alt":"具身智能机器人保险","abbr":"","aliases":["Humanoid Robot Insurance"],"one_liner":"Insurance designed for humanoid or embodied robots, covering damage to the robot and harm it causes others.","explanation":"Embodied-AI robot insurance refers to insurance products that insurers design specifically for humanoid robots, quadruped robots, and other embodied products. As reported, Chinese property insurers began offering such products starting in 2025, typically covering two things: property insurance for damage to or loss of the robot itself, and liability insurance for injuries or property damage the robot causes to others while working, performing, or being rented out. It addresses the question of who pays after something goes wrong: once robots start appearing in paid performances, showroom tours, factories, and homes, a single fall or loss of control could cause real damage, and it's often unclear whether the manufacturer, the operator, or the user bears responsibility. Insurance shifts this risk elsewhere, lowering the barrier to rental and deployment, and also pushes manufacturers to provide fault logs and safety data.","example":"","related":["Robot Rental","Commercial Robot Performances","Embodied Safety","Functional Safety","Scaled Deployment"]},{"id":"closed-commercial-loop","category":"industry","sec":7,"tier":2,"sources":[{"title":"Business model - Wikipedia","url":"https://en.wikipedia.org/wiki/Business_model"}],"as_of":"","related_ids":["real-world-deployment","product-market-fit","data-flywheel","return-on-investment-payback-period","scaled-deployment","framework-order-letter-of-intent-order"],"name":"Closed Commercial Loop","alt":"商业闭环","abbr":"","aliases":["Closing the Loop"],"one_liner":"A business model where a product keeps selling, revenue covers its costs, and the cycle sustains itself.","explanation":"This describes a full loop from product to paying customer to revenue flowing back in: customers are willing to pay for real value delivered, that revenue covers R&D, manufacturing, and operating costs, and there's enough left over to keep funding further iteration. Embodied-AI companies are often asked whether they've “closed the loop,” meaning whether they have real orders that hold up without relying on fresh funding, rather than just demos and letters of intent. In robotics specifically, a closed loop is often tied to a data flywheel: deployment brings in revenue, and it also brings in real-world data that improves the model.","example":"If a company's robot does picking work in a warehouse, and the customer pays by the unit and keeps renewing, with the rental fees covering hardware and maintenance costs, that scenario has closed its commercial loop.","related":["Real-world Deployment","Product-Market Fit","Data Flywheel","Return on Investment / Payback Period","Scaled Deployment","Framework Order / Letter-of-Intent Order"]},{"id":"product-market-fit","category":"industry","sec":7,"tier":2,"sources":[{"title":"Product-market fit - Wikipedia","url":"https://en.wikipedia.org/wiki/Product-market_fit"},{"title":"The only thing that matters - Marc Andreessen","url":"https://pmarchive.com/guide_to_startups_part4.html"}],"as_of":"","related_ids":["killer-app","closed-commercial-loop","real-world-deployment","proof-of-concept","scaled-deployment","return-on-investment-payback-period"],"name":"Product-Market Fit","alt":"PMF（产品市场匹配）","abbr":"PMF","aliases":["PMF"],"one_liner":"When a product satisfies a large enough market that customers actively choose to buy it.","explanation":"Product-market fit (PMF) is startup and investing terminology for a product finding a market that genuinely needs it and is large enough to matter, evidenced by customers buying it on their own initiative, repurchasing, and spreading it by word of mouth. The concept originates with venture capitalist Andy Rachleff, and became widely known after Marc Andreessen's 2007 blog post calling it “the only thing that matters” for a startup. In embodied AI, PMF is the question investors ask most often: robot demos can be stunning, but which scenario will pay an acceptable price for a robot that's stable and efficient enough to replace the current solution is still unsettled. Until a company finds PMF, much of its revenue tends to come from research and education sales, exhibitions, and pilot programs.","example":"An investor asking a humanoid robot company, “Besides research/education sales and paid appearances, which use case has customers repurchasing at volume?” is really asking about PMF.","related":["Killer App","Closed Commercial Loop","Real-world Deployment","Proof of Concept","Scaled Deployment","Return on Investment / Payback Period"]},{"id":"killer-app","category":"industry","sec":7,"tier":2,"sources":[{"title":"Killer application - Wikipedia","url":"https://en.wikipedia.org/wiki/Killer_application"},{"title":"VisiCalc - Wikipedia","url":"https://en.wikipedia.org/wiki/VisiCalc"}],"as_of":"","related_ids":["product-market-fit","closed-commercial-loop","real-world-deployment","chatgpt-moment-for-robotics","iphone-moment"],"name":"Killer App","alt":"杀手级应用","abbr":"","aliases":["Killer Use Case","Killer Application"],"one_liner":"The one use case so valuable that customers buy an entire product just to get it.","explanation":"“Killer app” originally comes from the software industry: an application so valuable that it alone justifies buying a whole class of hardware or platform. The classic example is VisiCalc, an electronic spreadsheet released in 1979 that drove many businesses to buy an Apple II computer just to run it. In embodied AI, the debate is which single use case will make humanoid or general-purpose robots something customers actively want to pay for, at a scale large enough to matter. The industry consensus right now is that no such killer app has been found yet: candidates people are betting on include factory material handling, warehouse order picking, household chores, guided tours, and research and education, but none has proven itself. This question sits at the center of discussions about product-market fit and finding a workable business model.","example":"VisiCalc was the killer app for the Apple II; some believe a household robot that could reliably fold laundry and tidy a room would become embodied AI's killer app.","related":["Product-Market Fit","Closed Commercial Loop","Real-world Deployment","ChatGPT Moment for Robotics","iPhone Moment"]},{"id":"chatgpt-moment-for-robotics","category":"industry","sec":8,"tier":1,"sources":[{"title":"NVIDIA Launches Cosmos World Foundation Model Platform to Accelerate Physical AI Development","url":"https://nvidianews.nvidia.com/news/nvidia-launches-cosmos-world-foundation-model-platform-to-accelerate-physical-ai-development"},{"title":"MIT Technology Review: Is robotics about to have its own ChatGPT moment?（2024-04-11）","url":"https://www.technologyreview.com/2024/04/11/1090718/household-robots-ai-data-robotics/"}],"as_of":"2025-01","related_ids":["iphone-moment","imagenet-moment","deepseek-moment-for-embodied-ai","embodied-ai-bubble","general-purpose-robot","nvidia"],"name":"ChatGPT Moment for Robotics","alt":"ChatGPT 时刻","abbr":"","aliases":["GPT Moment"],"one_liner":"Industry shorthand for the hoped-for turning point when general-purpose robot ability suddenly becomes usable and visible to the public.","explanation":"The phrase borrows from how quickly large language models broke into the mainstream after ChatGPT launched in late 2022, and refers to a hoped-for equivalent turning point for robots: a single general-purpose model that lets a robot handle many tasks in unfamiliar environments, at which point ordinary people first feel that robots “actually work.” Media and researchers were already discussing this in 2024, and after NVIDIA CEO Jensen Huang said at CES 2025, while unveiling Cosmos, that “the ChatGPT moment for robotics is coming,” the phrase became a fixture at product launches and in funding news. It isn't a technical metric and has no agreed-upon threshold, so when you hear it, it's worth asking: across how many tasks, how many new environments, and at what success rate does the claim actually hold. Similar phrases include the “iPhone moment” and the “ImageNet moment.”","example":"When NVIDIA unveiled its Cosmos world foundation model at CES 2025, Jensen Huang said the ChatGPT moment for robotics was coming.","related":["iPhone Moment","ImageNet Moment","“DeepSeek Moment” for Embodied AI","Embodied AI Bubble","General-purpose Robot","NVIDIA"]},{"id":"iphone-moment","category":"industry","sec":8,"tier":3,"sources":[{"title":"iPhone (1st generation) - Wikipedia","url":"https://en.wikipedia.org/wiki/IPhone_(1st_generation)"}],"as_of":"","related_ids":["chatgpt-moment-for-robotics","imagenet-moment","consumer-grade-robot","killer-app","embodied-ai-bubble"],"name":"iPhone Moment","alt":"iPhone 时刻","abbr":"","aliases":[],"one_liner":"The turning point when a technology gets its defining product and starts being adopted by the mainstream.","explanation":"The “iPhone moment” borrows the story of Apple's first iPhone in 2007, which drove smartphones into mass adoption, to describe the point when a technology produces a landmark product that the public and industry broadly notice, after which adoption accelerates. As reported, Jensen Huang said at NVIDIA's GTC in March 2023 that AI's “iPhone moment” had already arrived, referring to ChatGPT. In embodied AI, the phrase is often used to ask when humanoid or home robots will get their own iPhone moment — that is, a product priced acceptably that can genuinely do useful work in a home or factory. It's close in meaning to the “ChatGPT moment,” but leans toward product adoption, while “ChatGPT moment” leans toward a breakthrough in capability.","example":"“The iPhone moment for home humanoid robots hasn't arrived yet” usually means: there are plenty of impressive demos, but no product yet that can be sold into ordinary homes and do reliable work.","related":["ChatGPT Moment for Robotics","ImageNet Moment","Consumer-Grade Robot","Killer App","Embodied AI Bubble"]},{"id":"imagenet-moment","category":"industry","sec":8,"tier":3,"sources":[{"title":"ImageNet - Wikipedia","url":"https://en.wikipedia.org/wiki/ImageNet"},{"title":"NLP's ImageNet moment has arrived (The Gradient)","url":"https://thegradient.pub/nlp-imagenet/"}],"as_of":"","related_ids":["data-scarcity","open-x-embodiment","agibot-world","benchmark","chatgpt-moment-for-robotics","iphone-moment"],"name":"ImageNet Moment","alt":"ImageNet 时刻","abbr":"","aliases":[],"one_liner":"The turning point when a large public dataset and benchmark lets one method break out and reshape a field.","explanation":"ImageNet is a large-scale labeled image dataset released in 2009 by a team led by Fei-Fei Li. In 2012, AlexNet used a deep convolutional network to win the ImageNet challenge by a wide margin over other methods, an event widely regarded as the starting point of the deep-learning boom. “ImageNet moment” has since come to describe, more generally, a turning point where a field gets a large enough public dataset and a unified benchmark, one method clearly pulls ahead, and the whole field shifts direction because of it. In 2018, Sebastian Ruder used “NLP's ImageNet moment” to describe the rise of pretrained language models. In embodied AI, people often say the field “still lacks its own ImageNet,” meaning it lacks a robot dataset and evaluation standard that's large enough, uniformly formatted, and widely adopted; datasets like Open X-Embodiment and AgiBot World are discussed in exactly this context.","example":"A common line when discussing embodied data: robot learning is still waiting for its ImageNet moment — the bottleneck is too little real-robot data, in too many incompatible formats.","related":["Data Scarcity","Open X-Embodiment","AgiBot World","Benchmark","ChatGPT Moment for Robotics","iPhone Moment"]},{"id":"deepseek-moment-for-embodied-ai","category":"industry","sec":8,"tier":3,"sources":[{"title":"DeepSeek - Wikipedia","url":"https://en.wikipedia.org/wiki/DeepSeek"}],"as_of":"","related_ids":["chatgpt-moment-for-robotics","iphone-moment","imagenet-moment","open-weight-model","embodied-ai-bubble","consensus-non-consensus"],"name":"“DeepSeek Moment” for Embodied AI","alt":"DeepSeek 时刻","abbr":"","aliases":["DeepSeek Moment"],"one_liner":"The hoped-for turning point when an open, cheap model drastically lowers the barrier to embodied AI, echoing DeepSeek-R1.","explanation":"The phrase comes from DeepSeek's January 2025 release of DeepSeek-R1, an open-weight reasoning model whose performance came close to the top closed-source models of the time at a much lower publicly stated training cost, briefly triggering a sharp sell-off in US AI stocks. The embodied-AI community borrows the phrase to describe an anticipated turning point: an open-source, cheap, and good-enough robot model appears, letting small and mid-sized teams build usable products on top of it and accelerating the whole industry. It differs in emphasis from the “ChatGPT moment,” which stresses capability first making ordinary people feel “this actually works”; the DeepSeek moment instead stresses openness and cost. It's an industry catchphrase with no defined criteria, so when you hear it, it's worth asking exactly what was open-sourced, how much cost actually dropped, and on which tasks the claim holds.","example":"","related":["ChatGPT Moment for Robotics","iPhone Moment","ImageNet Moment","Open-weight Model","Embodied AI Bubble","Consensus / Non-Consensus"]},{"id":"gartner-hype-cycle","category":"industry","sec":8,"tier":3,"sources":[{"title":"Gartner hype cycle - Wikipedia","url":"https://en.wikipedia.org/wiki/Gartner_hype_cycle"}],"as_of":"","related_ids":["embodied-ai-bubble","consensus-non-consensus","chatgpt-moment-for-robotics","year-one-of-mass-production","real-world-deployment"],"name":"Gartner Hype Cycle","alt":"技术成熟度曲线（Gartner 炒作周期）","abbr":"","aliases":["Technology Maturity Curve"],"one_liner":"A five-stage curve describing how a new technology moves from hype to disillusionment to real adoption.","explanation":"The hype cycle was introduced by Gartner analyst Jackie Fenn in 1995, and Gartner now publishes many versions annually across different technology fields. The curve divides public expectations of a new technology into five stages: the innovation trigger, the peak of inflated expectations, the trough of disillusionment, the slope of enlightenment, and the plateau of productivity. Its point is that media buzz and real-world usability often move out of sync — hype is frequently at its highest well before large-scale deployment is anywhere close. Discussions of whether embodied AI is a bubble, or where humanoid robots currently sit, often borrow this chart. It's an empirical framework, not a quantitative prediction, and many technologies never complete the whole curve.","example":"According to reports, Gartner's 2025 AI hype cycle placed generative AI as having already slid from the peak into the trough of disillusionment.","related":["Embodied AI Bubble","Consensus / Non-Consensus","ChatGPT Moment for Robotics","Year One of Mass Production","Real-world Deployment"]},{"id":"embodied-ai-bubble","category":"industry","sec":8,"tier":2,"sources":[{"title":"Embodied Intelligence 2026: Farewell to Narrative Hype, Practical Deployment Reigns Supreme (36Kr)","url":"https://eu.36kr.com/en/p/3953394550537606"},{"title":"Wikipedia: Gartner hype cycle","url":"https://en.wikipedia.org/wiki/Gartner_hype_cycle"}],"as_of":"2025-11","related_ids":["gartner-hype-cycle","hundred-robot-war","homogenization-redundant-construction","framework-order-letter-of-intent-order","demo","closed-commercial-loop"],"name":"Embodied AI Bubble","alt":"具身智能泡沫","abbr":"","aliases":["Humanoid Robot Bubble"],"one_liner":"The worry that funding and valuations in embodied AI are running ahead of the technology's actual maturity and commercial payoff.","explanation":"This is skepticism that the embodied-AI sector, and humanoid robots in particular, has overheated: the number of companies, the amount raised, and valuations have all climbed fast, while scenarios that can work reliably and generate ongoing revenue remain scarce, a good share of announced orders turn out to be framework or letter-of-intent orders, and demo videos often involve teleoperation or editing. People who hold this view worry the sector could follow other overheated technologies into a period of hype followed by a shakeout once the excitement fades, and a large number of companies exiting. The counterargument is that models and hardware are advancing quickly, and heavy early-stage investment is normal for a young industry. Three things are worth checking to judge for yourself: whether demos are fully autonomous, whether orders are genuinely being delivered, and whether a single unit can pay for itself.","example":"At a press briefing in November 2025, China's National Development and Reform Commission reportedly noted that the country already had more than 150 humanoid-robot companies, and flagged the risk of a crowd of highly similar products and squeezed room for R&D.","related":["Gartner Hype Cycle","Hundred-Robot War","Homogenization / Redundant Construction","Framework Order / Letter-of-Intent Order","Demo (Demonstration Video)","Closed Commercial Loop"]},{"id":"hundred-robot-war","category":"industry","sec":8,"tier":3,"sources":[{"title":"国家发改委：防范重复度高的人形机器人产品「扎堆」上市（中国政府网，2025-11）","url":"https://www.gov.cn/yaowen/liebiao/202511/content_7049858.htm"}],"as_of":"2025-11","related_ids":["homogenization-redundant-construction","embodied-ai-bubble","price-war","mass-production","humanoid-robot"],"name":"Hundred-Robot War","alt":"百机大战","abbr":"","aliases":["Crowded Humanoid-Robot Market"],"one_liner":"A phrase describing the crowd of well over a hundred Chinese companies all building humanoid robots at once.","explanation":"“Hundred-robot war” is a phrase used by Chinese media and investors, modeled on the “hundred-model war” used to describe China's 2023 large-language-model boom, describing the sudden surge in the number of humanoid-robot makers and the convergence of their product shapes and specifications. At a press briefing in November 2025, the National Development and Reform Commission said China already had more than 150 humanoid-robot companies, over half of them startups or companies that had pivoted in from other industries, and it warned against a crowd of highly similar products rushing to list, which risks squeezing out room for real R&D. For newcomers, this phrase is a cue that once complete-robot designs start looking more and more alike, the real competition shifts to mass-production cost, real orders from real scenarios, and model capability — and the industry may be headed for consolidation.","example":"When yet another new humanoid robot launches, industry commentary often asks: in the middle of the hundred-robot war, what's actually different about it compared with existing products, and does it have real orders?","related":["Homogenization / Redundant Construction","Embodied AI Bubble","Price War (Humanoid Robots)","Mass Production","Humanoid Robot"]},{"id":"homogenization-redundant-construction","category":"industry","sec":8,"tier":3,"sources":[{"title":"工业和信息化部等七部门关于推动未来产业创新发展的实施意见（中国政府网）","url":"https://www.gov.cn/zhengce/zhengceku/202401/content_6929021.htm"}],"as_of":"2025-11","related_ids":["hundred-robot-war","embodied-ai-bubble","price-war","form-factor-debate","future-industries"],"name":"Homogenization / Redundant Construction","alt":"同质化 / 重复建设","abbr":"","aliases":[],"one_liner":"Many companies and regions building nearly identical products and projects, wasting resources.","explanation":"Homogenization means different companies' products end up looking and functioning almost identically, hard to tell apart; redundant construction is an older term in Chinese policy language, describing different regions repeatedly launching similar projects, parks, and production capacity in the same sector. During the humanoid-robot boom, a large number of companies released complete robots with similar looks and specs, regions competed to build innovation centers and training grounds, and a batch of companies rushed to list around the same time, drawing exactly this kind of criticism. As reported, in November 2025 the National Development and Reform Commission publicly flagged bubble risk in the humanoid-robot sector, saying the number of complete-robot makers had already exceeded 150 and warning against a crowd of highly similar products flooding the market. The positive-framed counterpart in policy documents is “differentiated development.”","example":"The future-industries implementation opinions from the Ministry of Industry and Information Technology and six other departments call on localities to “plan rationally, cultivate precisely, and develop in a differentiated way” based on their own industrial foundations.","related":["Hundred-Robot War","Embodied AI Bubble","Price War (Humanoid Robots)","Form-Factor Debate","Future Industries"]},{"id":"funding-rounds-and-valuation","category":"industry","sec":8,"tier":3,"sources":[{"title":"Unicorn (finance) - Wikipedia","url":"https://en.wikipedia.org/wiki/Unicorn_(finance)"},{"title":"Venture round - Wikipedia","url":"https://en.wikipedia.org/wiki/Venture_round"},{"title":"Figure Exceeds $1B in Series C Funding at $39B Post-Money Valuation","url":"https://www.figure.ai/news/series-c"}],"as_of":"2025-09","related_ids":["pre-ipo-tutoring","star-market","hkex-chapter-18c","embodied-ai-bubble","patient-capital"],"name":"Funding Rounds & Valuation","alt":"融资轮次与估值（天使轮 / A 轮 / Pre-IPO / 独角兽）","abbr":"","aliases":["Angel / Series A–C / Pre-IPO / Unicorn"],"one_liner":"The staged way startups raise money, and the price investors agree the company is worth at each stage.","explanation":"Startups typically raise money in stages: at the seed and angel rounds, often only a team and an idea exist; Series A, B, and C correspond to product validation, scaling, and expansion; Pre-IPO is the last round before going public. In each round, investors exchange money for shares at an agreed valuation, and “post-money valuation” refers to the company's total value right after that round's cash comes in. A still-private startup valued above $1 billion is called a “unicorn,” a term coined by investor Aileen Lee in 2013. Embodied-AI headlines about “completing an A+ round worth hundreds of millions of RMB” or “valuation doubling” all use this vocabulary. A funding round only tells you about fundraising pace, not technology or revenue level, so it should be read alongside information like mass production and order volume.","example":"Figure AI announced a Series C round of over $1 billion in September 2025, at a post-money valuation of $39 billion.","related":["Pre-IPO Tutoring","STAR Market","HKEX Chapter 18C","Embodied AI Bubble","Patient Capital"]},{"id":"total-addressable-market","category":"industry","sec":8,"tier":3,"sources":[{"title":"Wikipedia: Total addressable market","url":"https://en.wikipedia.org/wiki/Total_addressable_market"}],"as_of":"","related_ids":["closed-commercial-loop","labor-shortage","funding-rounds-and-valuation","product-market-fit","real-world-deployment","mass-production"],"name":"Total Addressable Market","alt":"市场空间（TAM）","abbr":"TAM","aliases":["TAM"],"one_liner":"The full market size a product could theoretically capture if it won every target customer.","explanation":"TAM is a standard way venture capital and consulting firms estimate market size: the total annual revenue a product would generate if it captured its entire target customer base. It's commonly used together with SAM (the serviceable market that's actually reachable once geography, channels, and other constraints are considered) and SOM (the share realistically obtainable in the near term), with the three narrowing step by step. Humanoid-robot fundraising materials and research reports frequently estimate TAM, typically by multiplying “number of jobs that could be replaced” by “unit price or labor cost saved per robot.” When reading these numbers, it's important to note the underlying assumptions: TAM describes a ceiling, not how many units can realistically be sold in the near term.","example":"Research reports estimating humanoid-robot TAM commonly start by counting jobs in manufacturing, logistics, and household chores, then multiply by an assumed replacement ratio and unit price.","related":["Closed Commercial Loop","Labor Shortage","Funding Rounds & Valuation","Product-Market Fit","Real-world Deployment","Mass Production"]},{"id":"patient-capital","category":"industry","sec":8,"tier":3,"sources":[{"title":"Patient capital - Wikipedia","url":"https://en.wikipedia.org/wiki/Patient_capital"}],"as_of":"","related_ids":["national-ai-industry-investment-fund","funding-rounds-and-valuation","future-industries","new-quality-productive-forces","homogenization-redundant-construction","embodied-ai-bubble"],"name":"Patient Capital","alt":"耐心资本（国资长线投资）","abbr":"","aliases":["State-Backed Long-Term Investment"],"one_liner":"Investment willing to be held long-term without demanding a quick return, often state-backed in China.","explanation":"Patient capital is an economics concept for capital willing to be held long-term, without pursuing a quick exit or return. Since 2024 it has appeared frequently in Chinese policy documents; the decision from the Third Plenary Session of the 20th Party Central Committee called for “developing patient capital,” encouraging investment that's early, small, long-term, and directed at hard tech. Embodied-AI R&D cycles are long, and mass production and profitability are both still far off, so the typical five-to-seven-year exit window of ordinary venture capital often can't wait that long — making national- and local-level government guidance funds and state-owned platforms important sources of capital in this sector. It also has a side effect: regions competing to fund local projects can encourage homogenization and redundant construction.","example":"The National Artificial Intelligence Industry Investment Fund and various local government guidance funds have participated in the funding rounds of several humanoid-robot companies, moves the media often describes as state-backed patient capital entering the sector.","related":["National AI Industry Investment Fund","Funding Rounds & Valuation","Future Industries","New Quality Productive Forces","Homogenization / Redundant Construction","Embodied AI Bubble"]},{"id":"star-market","category":"industry","sec":8,"tier":2,"sources":[{"title":"STAR Market - Wikipedia","url":"https://en.wikipedia.org/wiki/STAR_Market"},{"title":"上海证券交易所科创板","url":"https://star.sse.com.cn/"}],"as_of":"2025","related_ids":["pre-ipo-tutoring","first-listed-humanoid-robot-stock","hkex-chapter-18c","humanoid-robot-concept-stocks","backdoor-listing"],"name":"STAR Market","alt":"科创板","abbr":"","aliases":["Shanghai Sci-Tech Innovation Board","Shanghai Stock Exchange STAR Market"],"one_liner":"A Shanghai Stock Exchange board for hard-tech companies, piloting China's IPO registration system.","explanation":"The STAR Market (科创板, Sci-Tech Innovation Board) is an independent board of the Shanghai Stock Exchange, opened in July 2019 as the first venue in China to pilot a registration-based IPO system rather than the older approval-based one. It targets “hard tech” companies in fields like semiconductors, advanced equipment, and artificial intelligence, and keeps a listing path open even for companies that aren't yet profitable. Robotics and embodied-AI companies typically carry heavy R&D spending and face long paths to profitability, making the STAR Market one of their main routes to a domestic listing; news phrases like “first humanoid-robot stock” or “beginning listing guidance” are frequently tied to it.","example":"Unitree Robotics reportedly began listing guidance in 2025, aiming to list on the STAR Market.","related":["Pre-IPO Tutoring","First Listed Humanoid Robot Stock","HKEX Chapter 18C","Humanoid Robot Concept Stocks","Backdoor Listing"]},{"id":"hkex-chapter-18c","category":"industry","sec":8,"tier":3,"sources":[{"title":"Black Sesame Technologies - Wikipedia","url":"https://en.wikipedia.org/wiki/Black_Sesame_Technologies"}],"as_of":"2024-08","related_ids":["star-market","pre-ipo-tutoring","first-listed-humanoid-robot-stock","funding-rounds-and-valuation","backdoor-listing"],"name":"HKEX Chapter 18C","alt":"港股 18C 章","abbr":"18C","aliases":["Specialist Technology Companies Listing Regime"],"one_liner":"A Hong Kong Stock Exchange listing channel for hard-tech companies with little or no revenue yet.","explanation":"Chapter 18C is the section of the Hong Kong Stock Exchange's Listing Rules covering “specialist technology companies,” effective since late March 2023. It allows companies in fields such as next-generation information technology, advanced hardware and software, advanced materials, new energy, and new food and agricultural technology to list even when revenue is minimal or the company hasn't yet commercialized, in exchange for a higher required market capitalization and participation by sophisticated independent investors. The market-cap thresholds set at launch were HK$6 billion for commercialized companies and HK$10 billion for pre-commercialization companies, reportedly lowered temporarily at times since. Robotics companies and autonomous-driving chipmakers, which burn cash heavily and reach profitability late, often consider this route.","example":"Black Sesame Technologies listed on the Hong Kong Stock Exchange under Chapter 18C in August 2024, the second company to list under this rule.","related":["STAR Market","Pre-IPO Tutoring","First Listed Humanoid Robot Stock","Funding Rounds & Valuation","Backdoor Listing"]},{"id":"pre-ipo-tutoring","category":"industry","sec":8,"tier":3,"sources":[{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"},{"title":"Initial public offering - Wikipedia","url":"https://en.wikipedia.org/wiki/Initial_public_offering"}],"as_of":"2026-09","related_ids":["star-market","hkex-chapter-18c","first-listed-humanoid-robot-stock","backdoor-listing","special-purpose-acquisition-company","funding-rounds-and-valuation"],"name":"Pre-IPO Tutoring","alt":"上市辅导","abbr":"","aliases":["Listing Guidance","IPO Counseling"],"one_liner":"The mandatory process where a securities firm prepares an A-share company's governance before it files to list.","explanation":"Pre-IPO tutoring is a mandatory step before an initial public offering (IPO) on China's A-share market: a company first signs a tutoring agreement with a securities firm, files for tutoring registration with the local securities regulatory bureau, and the securities firm then works to bring the company's governance, finance, and internal controls up to standard, while training its directors, supervisors, and executives on securities regulations; only after passing an acceptance review can the company submit its listing application to an exchange. Because the registration filing is published on the securities regulator's website, “Company X has begun listing tutoring” is often treated as the first public signal that it's preparing to go public. Since 2025, several humanoid-robot companies have entered tutoring one after another.","example":"According to Wikipedia, Unitree Robotics began listing tutoring with CITIC Securities in July 2025 and listed on the Shanghai Stock Exchange in August 2026.","related":["STAR Market","HKEX Chapter 18C","First Listed Humanoid Robot Stock","Backdoor Listing","Special Purpose Acquisition Company","Funding Rounds & Valuation"]},{"id":"backdoor-listing","category":"industry","sec":8,"tier":3,"sources":[{"title":"智元机器人拟收购上纬新材63.62%股份（华尔街见闻）","url":"https://wallstreetcn.com/articles/3750667"},{"title":"锋龙股份再回应：优必选三年内不会借壳上市（新京报）","url":"https://m.bjnews.com.cn/detail/1766920585129742.html"},{"title":"Reverse takeover - Wikipedia","url":"https://en.wikipedia.org/wiki/Reverse_takeover"}],"as_of":"2026-04","related_ids":["swancor-advanced-materials","zhejiang-fenglong-electric","agibot","star-market","hkex-chapter-18c","special-purpose-acquisition-company"],"name":"Backdoor Listing","alt":"借壳上市","abbr":"RTO","aliases":["Reverse Takeover (RTO)","Shell Listing"],"one_liner":"An unlisted company buys control of a listed “shell” company and injects its own business to go public.","explanation":"A backdoor listing means an unlisted company first acquires control of an already-listed company (the “shell”), then injects its own assets and business into it, achieving a public listing while bypassing the queue for a traditional initial public offering (IPO). On China's A-share market, once a change of control is followed by an asset injection above a certain size, regulators treat it as a “restructuring listing,” reviewed under standards close to those for an IPO. Embodied-AI companies tend to carry high valuations and heavy losses, and a traditional IPO takes a long time, so when one takes control of a listed company, the market often reads it as a possible backdoor-listing signal: in July 2025, AgiBot announced it had acquired about 63.62% of STAR Market-listed Shangwei New Materials, and in 2026 UBTECH completed taking control of Fenglong Co. Both companies publicly stated they had no backdoor-listing plans within three years, so coverage should be read carefully to distinguish “taking a controlling stake” from “actually injecting assets.”","example":"After AgiBot took control of Shangwei New Materials, Shangwei continued operating its consumer-robotics business independently under the “Swancor Qiyuan” brand, rather than immediately folding in all of AgiBot's business.","related":["Swancor Advanced Materials","Zhejiang Fenglong Electric","AgiBot","STAR Market","HKEX Chapter 18C","Special Purpose Acquisition Company"]},{"id":"special-purpose-acquisition-company","category":"industry","sec":8,"tier":3,"sources":[{"title":"Special-purpose acquisition company - Wikipedia","url":"https://en.wikipedia.org/wiki/Special-purpose_acquisition_company"}],"as_of":"2022-01","related_ids":["backdoor-listing","star-market","hkex-chapter-18c","pre-ipo-tutoring","funding-rounds-and-valuation"],"name":"Special Purpose Acquisition Company","alt":"SPAC 上市（特殊目的收购公司）","abbr":"SPAC","aliases":["SPAC","Blank-Check Company"],"one_liner":"A shell company lists first to raise money, then merges with a target to take it public indirectly.","explanation":"A SPAC is a listed company with no real business of its own. It first raises money through an IPO and places it in a trust account, then, within a set window (usually one to two years), looks for an unlisted company to merge with; once the merger completes, the target company effectively becomes the listed entity. Compared with a traditional IPO, this route is faster and lets valuation be negotiated in advance, but it's also frequently criticized for weak disclosure to retail investors. The number of US SPACs surged in 2020–2021, and a batch of robotics companies went public this way; the Hong Kong exchange also introduced a SPAC mechanism starting in 2022.","example":"Exoskeleton and robotics company Sarcos and warehouse-robotics company Berkshire Grey both went public in the US via SPAC in 2021.","related":["Backdoor Listing","STAR Market","HKEX Chapter 18C","Pre-IPO Tutoring","Funding Rounds & Valuation"]},{"id":"first-listed-humanoid-robot-stock","category":"industry","sec":8,"tier":3,"sources":[{"title":"UBtech Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/UBtech_Robotics"},{"title":"Unitree Robotics - Wikipedia","url":"https://en.wikipedia.org/wiki/Unitree_Robotics"}],"as_of":"2026-08","related_ids":["ubtech-robotics","unitree-robotics","star-market","hkex-chapter-18c","pre-ipo-tutoring","humanoid-robot-concept-stocks"],"name":"First Listed Humanoid Robot Stock","alt":"人形机器人第一股","abbr":"","aliases":["“Humanoid Robot's First Stock”"],"one_liner":"A media label for the first humanoid-robot company to list on a given stock market.","explanation":"“First stock” is a habitual label used by the media and securities analysts for the first company to list in a given sector — it's not an official designation. In Hong Kong, the “first humanoid-robot stock” usually refers to UBTECH, which listed on the Hong Kong Stock Exchange's Main Board (ticker 9880) in December 2023. On the A-share side, Unitree Robotics began listing guidance in July 2025 and listed on the Shanghai Stock Exchange's STAR Market (688836) in August 2026, an event media treated as a landmark for A-share humanoid-robot listings. Because sub-sector definitions vary, there are also labels like “first collaborative-robot stock” and “first robot-brain stock.” When you see this kind of title, it's worth going back to the company's actual revenue breakdown to check whether it truly sells humanoid robots, education products, or something else.","example":"At the time of its listing, UBTECH's humanoid-robot revenue share was small; most of its revenue came from education robots and consumer products.","related":["UBTech Robotics","Unitree Robotics","STAR Market","HKEX Chapter 18C","Pre-IPO Tutoring","Humanoid Robot Concept Stocks"]},{"id":"humanoid-robot-concept-stocks","category":"industry","sec":8,"tier":3,"sources":[{"title":"概念股 - 维基百科","url":"https://zh.wikipedia.org/wiki/%E6%A6%82%E5%BF%B5%E8%82%A1"}],"as_of":"","related_ids":["tesla-supply-chain","first-listed-humanoid-robot-stock","per-unit-content-value","upstream-midstream-downstream-of-the-industry-chain","embodied-ai-bubble"],"name":"Humanoid Robot Concept Stocks","alt":"人形机器人概念股","abbr":"","aliases":[],"one_liner":"Stocks the market groups together as “related to humanoid robots,” rising and falling on industry news.","explanation":"“Concept stocks” is common A-share and Hong Kong-market language for stocks the market groups together and trades as a theme because their business, or even just rumors about it, is tied to a hot topic. Humanoid-robot concept stocks include complete-robot makers as well as suppliers of components like reducers, lead screws, motors, and sensors, plus companies believed to have entered Tesla Optimus's supply chain (the “Tesla chain”). News like a Tesla event, a Spring Festival Gala performance, or a major fundraising round from a leading company often moves the whole sector up or down together. It's worth noting that many concept stocks derive only a small share of revenue from robotics, so their share price mostly reflects expectations, not results already delivered.","example":"Auto-parts suppliers such as Sanhua Intelligent Controls and Tuopu Group are often classified as humanoid-robot concept stocks because they're believed to supply actuators for Optimus.","related":["Tesla (Optimus) Supply Chain","First Listed Humanoid Robot Stock","Per-Unit Content Value","Upstream / Midstream / Downstream of the Industry Chain","Embodied AI Bubble"]},{"id":"tesla-supply-chain","category":"industry","sec":8,"tier":3,"sources":[{"title":"Tesla AI & Robotics","url":"https://www.tesla.com/AI"}],"as_of":"","related_ids":["tesla-optimus","tesla","humanoid-robot-concept-stocks","sampling-supplier-nomination","planetary-roller-screw","per-unit-content-value"],"name":"Tesla (Optimus) Supply Chain","alt":"T 链（特斯拉链）","abbr":"T链","aliases":["Tesla Chain","T-Chain"],"one_liner":"Component makers believed to supply, or possibly supply, Tesla's Optimus robot, mostly an A-share market label.","explanation":"The “Tesla chain” is a term from China's capital markets, originally referring to suppliers of Tesla's cars, later extended mainly to companies that might supply Tesla's Optimus humanoid robot, covering actuators, planetary roller screws, harmonic reducers, coreless motors, sensors, and similar parts. Tesla has never published a complete list of Optimus suppliers, so being “part of the Tesla chain” usually comes from a company's own statements, investor-relations notes, or media reports, and a good number of these are still only at the sampling stage. Share prices of these companies often swing sharply on any Optimus news, so it's worth distinguishing a confirmed nomination from mere rumor when reading coverage.","example":"Sanhua Intelligent Controls and Tuopu Group are commonly classified as part of the Tesla chain, though much of the reported partnership detail is described as unconfirmed.","related":["Tesla Optimus","Tesla","Humanoid Robot Concept Stocks","Sampling / Supplier Nomination","Planetary Roller Screw","Per-Unit Content Value"]},{"id":"per-unit-content-value","category":"industry","sec":8,"tier":3,"sources":[{"title":"Bill of materials - Wikipedia","url":"https://en.wikipedia.org/wiki/Bill_of_materials"}],"as_of":"","related_ids":["bill-of-materials-cost","shipment-volume","total-addressable-market","upstream-midstream-downstream-of-the-industry-chain","core-components","tesla-supply-chain"],"name":"Per-Unit Content Value","alt":"单机价值量","abbr":"","aliases":["Component Value per Robot"],"one_liner":"How much money one category of component adds up to inside a single robot, used to size a market.","explanation":"Per-unit content value is a term common in securities-analyst and industry reports: how much a given category of component (or a given supplier's products) is worth in total inside one robot. A humanoid robot, for example, uses dozens of joints, each paired with a reducer, a motor, and a lead screw, and multiplying that out gives the per-unit content value for that category of part. It's related to BOM cost (a complete robot's total bill-of-materials cost), but focuses on one category of component. Analysts estimate a component market's size as “per-unit content value × projected shipment volume,” and use that to judge which upstream companies stand to benefit most — which is why the term often appears alongside the “Tesla chain” and concept stocks.","example":"Research reports estimate that planetary roller screws, harmonic reducers, and frameless torque motors carry relatively high per-unit content value in a humanoid robot, and recommend related upstream companies on that basis.","related":["Bill of Materials Cost","Shipment Volume","Total Addressable Market","Upstream / Midstream / Downstream of the Industry Chain","Core Components","Tesla (Optimus) Supply Chain"]},{"id":"embodied-ai-first-named-in-the-government-work-report","category":"industry","sec":9,"tier":2,"sources":[{"title":"政府工作报告（2025 年 3 月 5 日，中国政府网）","url":"https://www.gov.cn/yaowen/liebiao/202503/content_7013163.htm"}],"as_of":"2025-03","related_ids":["future-industries","new-quality-productive-forces","15th-five-year-plan","ai-plus-initiative","guiding-opinions-on-the-innovative-development-of-humanoid-r","embodied-ai"],"name":"Embodied AI First Named in the Government Work Report","alt":"政府工作报告首次写入具身智能","abbr":"","aliases":["Embodied AI in the 2025 Government Work Report"],"one_liner":"China's 2025 annual Government Work Report named embodied AI for the first time as an industry to cultivate.","explanation":"On March 5, 2025, Premier Qiang Li delivered the Government Work Report at the Third Session of the 14th National People's Congress. In the section on emerging and future industries within the 2025 work plan, it stated: cultivate future industries such as bio-manufacturing, quantum technology, embodied AI, and 6G. This was the first time the term “embodied AI” appeared in a Government Work Report. The report is the State Council's annual policy statement to the National People's Congress, delivered each year at the “Two Sessions”; being named in it signals that embodied AI has been formally designated a national-level industrial priority. Since then, various localities have rolled out dedicated policies, set up industry funds, and built data-collection training grounds, and the industry commonly treats this moment as the point at which embodied AI was formally sanctioned at the national level in China.","example":"The report's exact wording: “cultivate future industries such as bio-manufacturing, quantum technology, embodied AI, and 6G.”","related":["Future Industries","New Quality Productive Forces","15th Five-Year Plan (2026–2030)","AI+ Initiative","Guiding Opinions on the Innovative Development of Humanoid Robots","Embodied AI"]},{"id":"future-industries","category":"industry","sec":9,"tier":3,"sources":[{"title":"工业和信息化部等七部门关于推动未来产业创新发展的实施意见（中国政府网）","url":"https://www.gov.cn/zhengce/zhengceku/202401/content_6929021.htm"}],"as_of":"2024-01","related_ids":["new-quality-productive-forces","guiding-opinions-on-the-innovative-development-of-humanoid-r","15th-five-year-plan","embodied-ai-first-named-in-the-government-work-report","patient-capital"],"name":"Future Industries","alt":"未来产业","abbr":"","aliases":[],"one_liner":"Chinese policy term for industries still at an early technical stage that could become tomorrow's pillar sectors.","explanation":"Future industries is a category in Chinese industrial policy: fields driven by frontier technology that are currently still incubating or just emerging, but could grow into large-scale industries later, as distinct from the already-formed “strategic emerging industries.” In January 2024, the Ministry of Industry and Information Technology and six other departments issued the “Implementation Opinions on Promoting Innovation and Development of Future Industries,” proposing six priority directions — future manufacturing, information, materials, energy, space, and health — and listing humanoid robots as a high-end equipment product to be developed. Embodied AI has since also been written into the future-industries list in the Government Work Report. Understanding this term helps explain why local governments set up dedicated robotics funds, innovation centers, and pilot-scale testing bases.","example":"The “Implementation Opinions on Promoting Innovation and Development of Future Industries” calls for breakthroughs in high-end equipment products such as humanoid robots, quantum computers, and ultra-high-speed trains.","related":["New Quality Productive Forces","Guiding Opinions on the Innovative Development of Humanoid Robots","15th Five-Year Plan (2026–2030)","Embodied AI First Named in the Government Work Report","Patient Capital"]},{"id":"new-quality-productive-forces","category":"industry","sec":9,"tier":3,"sources":[{"title":"New quality productive forces - Wikipedia","url":"https://en.wikipedia.org/wiki/New_quality_productive_forces"}],"as_of":"2024-03","related_ids":["future-industries","15th-five-year-plan","ai-plus-initiative","embodied-ai-first-named-in-the-government-work-report","patient-capital"],"name":"New Quality Productive Forces","alt":"新质生产力","abbr":"","aliases":[],"one_liner":"Chinese policy term for advanced productive capacity led by scientific and technological innovation.","explanation":"New quality productive forces is a Chinese policy term, first proposed by Xi Jinping during a September 2023 inspection tour of Heilongjiang province, referring to advanced productive capacity generated by revolutionary technological breakthroughs and innovative allocation of production factors, with scientific and technological innovation playing the leading role. China's 2024 Government Work Report listed “accelerating the development of new quality productive forces” as a key task for the year, and the phrase has since become a high-frequency term in industrial policy. Humanoid robots and embodied AI, classified under “future industries,” are frequently promoted by local governments and companies as representative examples of new quality productive forces. For newcomers, understanding this term helps make sense of why policy documents, local subsidies, industrial funds, and investment-promotion news so often mention embodied AI.","example":"","related":["Future Industries","15th Five-Year Plan (2026–2030)","AI+ Initiative","Embodied AI First Named in the Government Work Report","Patient Capital"]},{"id":"15th-five-year-plan","category":"industry","sec":9,"tier":3,"sources":[{"title":"Five-year plans of China - Wikipedia","url":"https://en.wikipedia.org/wiki/Five-year_plans_of_China"}],"as_of":"2026-03","related_ids":["future-industries","new-quality-productive-forces","ai-plus-initiative","embodied-ai-first-named-in-the-government-work-report","embodied-ai"],"name":"15th Five-Year Plan (2026–2030)","alt":"十五五规划","abbr":"","aliases":["China's 15th Five-Year Plan"],"one_liner":"China's national economic and social development plan for 2026–2030, in which embodied AI is named a future industry.","explanation":"The 15th Five-Year Plan is China's fifteenth five-year plan, covering 2026 through 2030. The process runs in stages: the Central Committee of the Communist Party first issues recommendations for the plan (approved at the Fourth Plenary Session of the 20th Party Central Committee in October 2025), and the State Council then drafts the full plan, which is approved at the National People's Congress session (the “Two Sessions”) in 2026. For embodied AI, the key point is that the plan's recommendations list embodied intelligence alongside quantum technology, biomanufacturing, brain-computer interfaces, and 6G as future industries to be cultivated as new sources of economic growth. A five-year plan shapes the ministry-level policies, local industrial plans, and state-capital investment of the following years, so financing news and local policy announcements often cite “implementing the 15th Five-Year Plan” as their backdrop.","example":"","related":["Future Industries","New Quality Productive Forces","AI+ Initiative","Embodied AI First Named in the Government Work Report","Embodied AI"]},{"id":"ai-plus-initiative","category":"industry","sec":9,"tier":3,"sources":[{"title":"中国政府网 政策文件库","url":"https://www.gov.cn/zhengce/"}],"as_of":"2025-08","related_ids":["15th-five-year-plan","new-quality-productive-forces","embodied-ai-first-named-in-the-government-work-report","embodied-ai","future-industries"],"name":"AI+ Initiative","alt":"人工智能+ 行动","abbr":"","aliases":["Artificial Intelligence Plus Initiative"],"one_liner":"China's national push to deeply integrate AI into every industry, with smart robots as one focus area.","explanation":"The “AI+” initiative was first proposed in China's 2024 Government Work Report, following the pattern of the earlier “Internet+” initiative — the idea being to connect artificial intelligence into manufacturing, healthcare, transportation, government services, and other industries. In August 2025, the State Council issued the “Opinions on Deeply Implementing the ‘AI+’ Initiative,” proposing that adoption of next-generation smart terminals and AI agents exceed 70% by 2027 and 90% by 2030, and listing smart robots as one category of next-generation smart terminal. For the embodied-AI industry, it forms part of the policy backdrop together with the 15th Five-Year Plan and future-industries policy, and local subsidies, scenario openings, and state-capital investment are often justified by reference to it.","example":"","related":["15th Five-Year Plan (2026–2030)","New Quality Productive Forces","Embodied AI First Named in the Government Work Report","Embodied AI","Future Industries"]},{"id":"robot-plus-application-action-implementation-plan","category":"industry","sec":9,"tier":3,"sources":[{"title":"工业和信息化部官网","url":"https://www.miit.gov.cn/"}],"as_of":"2023-01","related_ids":["guiding-opinions-on-the-innovative-development-of-humanoid-r","robot-density","machines-replacing-humans","application-scenario-list-scenario-opening","service-robot","special-purpose-robot"],"name":"“Robot+” Application Action Implementation Plan","alt":"《“机器人+”应用行动实施方案》","abbr":"","aliases":["Robot+ Initiative (MIIT, 2023)"],"one_liner":"A 2023 policy from 17 Chinese ministries promoting the use of robots across many industries beyond manufacturing.","explanation":"This is a policy document jointly issued in January 2023 by the Ministry of Industry and Information Technology, leading a total of seventeen government departments. Its approach is “robot plus a given industry”: pushing robots to expand beyond traditional factories like automaking and electronics into many more sectors. The plan calls for China's manufacturing robot density (robots per 10,000 workers) to roughly double by 2025 compared with 2020, and lists priority areas including manufacturing, agriculture, construction, energy, commerce and logistics, healthcare, elder care, education, community services, and public safety and hazardous environments, directing local governments and industry to identify and promote model use cases in each. It's a foundational document that predates the later humanoid-robot-specific policy, and is often cited by subsequent local scenario lists and demonstration projects.","example":"","related":["Guiding Opinions on the Innovative Development of Humanoid Robots","Robot Density","Machines Replacing Humans","Application Scenario List / Scenario Opening","Service Robot","Special-purpose Robot"]},{"id":"guiding-opinions-on-the-innovative-development-of-humanoid-r","category":"industry","sec":9,"tier":3,"sources":[{"title":"工业和信息化部等七部门关于推动未来产业创新发展的实施意见（中国政府网，相关后续政策）","url":"https://www.gov.cn/zhengce/zhengceku/202401/content_6929021.htm"}],"as_of":"2023-10","related_ids":["future-industries","braincerebellum-architecture","humanoid-robot","beijing-humanoid-robot-innovation-center","mass-production"],"name":"Guiding Opinions on the Innovative Development of Humanoid Robots","alt":"人形机器人创新发展指导意见","abbr":"","aliases":["MIIT Humanoid Robot Guiding Opinions (2023)"],"one_liner":"A 2023 MIIT policy document, China's first industrial policy specifically for humanoid robots.","explanation":"The “Guiding Opinions on the Innovative Development of Humanoid Robots” was issued by the Ministry of Industry and Information Technology in October 2023, China's first industrial-policy document dedicated specifically to humanoid robots. It calls for establishing an initial innovation system by 2025, achieving breakthroughs in key technologies described as “brain, cerebellum, and limbs,” ensuring safe and reliable supply of core components, and reaching batch production of complete robots, with the goal by 2027 of building a safe and reliable supply chain and industrial chain. The “brain–cerebellum” division of labor commonly used in the industry became further popularized because of this document. Since then, many local humanoid-robot policies and innovation-center construction projects have cited it as their basis.","example":"The document calls for using key technologies such as “brain, cerebellum, and limbs” as breakthrough points to drive batch production of complete robots.","related":["Future Industries","Brain–Cerebellum Architecture","Humanoid Robot","Beijing Humanoid Robot Innovation Center","Mass Production"]},{"id":"unveiling-the-list-competition-mechanism","category":"industry","sec":9,"tier":3,"sources":[{"title":"中国政府网 政策文件库","url":"https://www.gov.cn/zhengce/"}],"as_of":"","related_ids":["application-scenario-list-scenario-opening","chokepoint-technology","domestic-substitution","future-industries","guiding-opinions-on-the-innovative-development-of-humanoid-r"],"name":"“Unveiling the List” Competition Mechanism","alt":"揭榜挂帅","abbr":"","aliases":["Jiebang Guashuai","Open Call for Technical Champions"],"one_liner":"A Chinese government method where officials post key tech problems publicly and let any capable team compete to solve them.","explanation":"“Unveiling the list and taking command” (揭榜挂帅, jiebang guashuai) is a way Chinese science and technology programs are organized: the party posing the challenge — a government agency or a leading company — publishes a “list” of key technical problems that need solving, with clear targets and deadlines, and whoever can meet them may “unveil” the challenge and lead the effort to tackle it, regardless of seniority or affiliation. Successful teams receive funding or orders. The phrase entered central-government speeches around 2016 and was later written into the recommendations for the 14th Five-Year Plan. It's meant to fix a problem with traditional project approval, where funding decisions were based on seniority and track record rather than the actual need. In robotics, tasks related to future industries and humanoid robots from the Ministry of Industry and Information Technology and other agencies are reported to use this mechanism too; news that “a company was selected as a list-unveiling unit” usually means it won a national-level technical challenge.","example":"","related":["Application Scenario List / Scenario Opening","Chokepoint Technology","Domestic Substitution","Future Industries","Guiding Opinions on the Innovative Development of Humanoid Robots"]},{"id":"application-scenario-list-scenario-opening","category":"industry","sec":9,"tier":3,"sources":[{"title":"中华人民共和国科学技术部官网","url":"https://www.most.gov.cn/"}],"as_of":"","related_ids":["unveiling-the-list-competition-mechanism","real-world-deployment","pilot-scale-testing-base","proof-of-concept","embodied-ai-training-ground","task-oriented-grasping"],"name":"Application Scenario List / Scenario Opening","alt":"应用场景清单 / 场景开放","abbr":"","aliases":["Scenario List","Scenario Innovation"],"one_liner":"Governments or state firms publish real business scenarios where new technology can be trialed, and invite companies in.","explanation":"This is a common industrial-policy tool used by local governments, state-owned enterprises, and industrial parks in China: they compile their own real operations — a government-service hall needing a receptionist-guide, a factory needing inspection, a warehouse needing sorting, an elder-care facility, and so on — into a public “scenario list,” inviting companies to apply, with selected companies allowed to trial their technology on-site, sometimes with accompanying procurement or subsidies. At the national level, the Ministry of Science and Technology and other agencies issued a notice in 2022 promoting scenario innovation as a way to drive AI adoption. For robotics companies, the biggest challenge is often finding a customer willing to let a robot onto its site and generate real data and orders, which scenario opening is designed to solve — which is why it appears frequently in local embodied-AI support policies.","example":"","related":["“Unveiling the List” Competition Mechanism","Real-world Deployment","Pilot-Scale Testing Base","Proof of Concept","Embodied AI Training Ground (Robot Data Collection Center)","Task-Oriented Grasping"]},{"id":"pilot-scale-testing-base","category":"industry","sec":9,"tier":3,"sources":[{"title":"Pilot plant - Wikipedia","url":"https://en.wikipedia.org/wiki/Pilot_plant"}],"as_of":"","related_ids":["embodied-ai-training-ground","proof-of-concept","pilot-small-batch-delivery","engineering-productionization","application-scenario-list-scenario-opening","future-industries"],"name":"Pilot-Scale Testing Base","alt":"中试基地（中试验证平台）","abbr":"","aliases":["Pilot Verification Platform"],"one_liner":"A public platform that validates lab results under near-mass-production conditions before full-scale manufacturing.","explanation":"“Pilot-scale testing” (中试, short for “intermediate-scale trial”) originally refers to the step in chemical and manufacturing industries that scales results up from a small lab trial toward full production. A pilot-scale testing base is a platform that provides exactly these conditions, usually co-built by a government body together with companies and research institutions, offering space, equipment, test standards, and data. In embodied AI, these bases are used to test a robot's reliability and safety under near-real conditions, and to collect training data and run evaluations, helping startups cross the gap between a working prototype and a shippable product. Several regions in China have already established pilot-scale testing or verification platforms focused on embodied AI.","example":"As reported, Alibaba's DAMO Academy has partnered with a National AI Application Pilot-Scale Testing Base (for embodied AI), and Light Wheel AI has also launched simulation and evaluation infrastructure aimed at pilot-scale testing bases.","related":["Embodied AI Training Ground (Robot Data Collection Center)","Proof of Concept","Pilot / Small-Batch Delivery","Engineering / Productionization","Application Scenario List / Scenario Opening","Future Industries"]},{"id":"miit-humanoid-robot-and-embodied-ai-standardization-technica","category":"industry","sec":9,"tier":3,"sources":[{"title":"工业和信息化部人形机器人与具身智能标准化技术委员会成立（工信部，2025-12-27）","url":"https://www.miit.gov.cn/xwfb/bldhd/art/2025/art_25c6077c819d4a77aa7a900a82bbde45.html"}],"as_of":"2025-12","related_ids":["humanoid-robot-intelligence-level-grading","guiding-opinions-on-the-innovative-development-of-humanoid-r","iso-25785-1","embodied-safety","functional-safety"],"name":"MIIT Humanoid Robot and Embodied AI Standardization Technical Committee","alt":"人形机器人与具身智能标准化技术委员会","abbr":"","aliases":[],"one_liner":"A MIIT-established standards committee responsible for humanoid-robot and embodied-AI industry standards.","explanation":"This is an industry standardization technical committee established by the Ministry of Industry and Information Technology on December 26, 2025, with its secretariat housed at the China Institute of Electronics and MIIT chief engineer Xie Shaofeng serving as chair. It's responsible for drafting and revising industry standards for humanoid robots and embodied AI across basic and common technologies, key technologies, components and modules, complete machines and systems, applications, and safety. Early in an industry's life, interfaces, test methods, safety requirements, and intelligence-level grading tend to vary by company; a unified industry standard gives manufacturers, customers, and regulators a shared frame of reference, and can also affect market access and procurement. Newcomers who later see an industry standard project or draft for comment in this space can generally expect this committee to be the body behind it.","example":"Industry standards for things like complete-humanoid-robot safety or joint-module test methods would fall under this committee's review.","related":["Humanoid Robot Intelligence Level Grading","Guiding Opinions on the Innovative Development of Humanoid Robots","ISO 25785-1 (Safety Requirements for Industrial Mobile Robots with Actively Controlled Stability — Part 1: Robots)","Embodied Safety","Functional Safety"]},{"id":"embodied-ai-robot-application-technician","category":"industry","sec":9,"tier":2,"sources":[{"title":"北京人形：具身智能机器人应用技术员进入国家新职业序列","url":"https://www.x-humanoid.com/news-view-330.html"},{"title":"钛媒体：机器人还没学会做家务，卖数据的已经先赚到了钱","url":"https://www.tmtpost.com/8062934.html"}],"as_of":"2026-09","related_ids":["data-collector","embodied-ai-training-ground","teleoperation","data-collection-sop","embodied-ai-data-service-provider","beijing-humanoid-robot-innovation-center"],"name":"Embodied AI Robot Application Technician","alt":"具身智能机器人应用技术员（新职业）","abbr":"","aliases":["Embodied AI Robot Data Collector","Embodied AI Robot Trainer"],"one_liner":"A national occupation category China created in September 2026, covering data collection, model tuning, and field deployment for robots.","explanation":"On September 9, 2026, China's Ministry of Human Resources and Social Security and other departments published the eighth batch of new occupations, including “Embodied AI Robot Application Technician” (occupation code 4-04-05-16), proposed by the China Institute of Electronics with technical backing from the Beijing Humanoid Robot Innovation Center. The official definition broadly covers people who use teleoperation equipment, spatial digitization tools, and simulation-testing platforms to collect and process multimodal robot interaction data, fine-tune and adapt algorithm models, deploy and debug hardware and software on-site, and monitor and optimize operations. It splits into two job types, embodied AI robot data collector and trainer, listed as this term's aliases. Being added to the national occupation classification means official standards, training programs, and skills assessments can now be built around it, and it gives frontline roles like data collector a formal job title.","example":"A data collector wearing a VR headset to teleoperate a robot folding clothes at a data-collection facility falls under the “embodied AI robot data collector” job type.","related":["Data Collector (Teleoperator)","Embodied AI Training Ground (Robot Data Collection Center)","Teleoperation","Data Collection SOP","Embodied AI Data Service Provider","Beijing Humanoid Robot Innovation Center"]}]}