08Simulation & Evaluation
Practicing and testing on a computer: how simulators work, the sim-to-real gap, and the benchmarks used to evaluate. · 210 terms
- 8.1Simulation basics14
- 8.2Physics engines: rigid bodies, contact19
- 8.3Soft bodies, fluids, differentiable sim10
- 8.4Rendering and sensor simulation12
- 8.5Major simulators and frameworks26
- 8.6Scenes, assets, and indoor platforms17
- 8.7From simulation to reality19
- 8.8Evaluation methods and metrics21
- 8.9Tabletop manipulation benchmarks28
- 8.10Household, navigation, and QA benchmarks13
- 8.11RL and locomotion-control benchmarks14
- 8.12Real-world and world-model evaluation17
8.1Simulation basics
The first step to training robots virtually: what a simulator is, how it steps through interaction, and how to make it fast and accurate.
- Simulator仿真器Software that models in a computer how robots, objects, and sensors move and interact.
- Physics Engine物理引擎The software core that computes forces, collisions, and motion step by step according to physical laws.
- Environment (Env; reset/step interface)环境(Env)与 reset / step 接口The standard reinforcement-learning interface: reset starts an episode, step executes one action.
- Rollout推演Running a policy in an environment from start to finish to produce one complete interaction trajectory.
- An episode can end two ways: termination means the task itself is over; truncation means it was cut off by a time limit.
- The amount of simulated time a physics engine advances per step, which trades off accuracy, stability, and speed.
- The number of physics-engine steps that run for every single action a policy outputs.
- Substeps子步Splitting one simulation step into several smaller integration steps, trading extra computation for more stable, accurate physics.
- Real-Time Factor实时因子The ratio of simulated time to real elapsed time, measuring whether a simulation runs faster or slower than reality.
- Running many independent copies of an environment at once to collect interaction data in batches.
- Running thousands of simulated environments at once on a single GPU, shrinking robot training from days to minutes.
- How many environment steps or frames a simulator produces per second, which sets how fast data can be generated and training can run.
- Headless Mode无头模式Running a simulator with no graphical window open, commonly used for batch training and data collection on a server.
- Simulation Fidelity仿真保真度How closely a simulation matches the real world, in both the accuracy of its physics and the realism of its appearance.
8.2Physics engines: rigid bodies, contact
Opening up the physics engine to see how each step computes rigid-body motion, collisions, and contact forces.
- Physics simulation that assumes objects never deform, computing only their translation, rotation, and collisions with each other.
- Simulating a robot as rigid bodies connected by joints, computing how it moves and what forces it feels.
- Generalized (Reduced) Coordinates vs. Cartesian (Maximal) Coordinates关节坐标仿真 vs 笛卡尔坐标仿真(约化坐标 / 最大坐标)Two ways a physics engine can represent a multi-link robot: by joint angles alone, or as separate rigid bodies tied together by constraints.
- Collision Detection碰撞检测(物理引擎)The step where a physics engine works out which objects are touching, where, and how deeply.
- The simplified shape a physics engine uses for collision checks, usually different from what's rendered on screen.
- Cutting a concave mesh into several approximately convex pieces so it can serve as simulation collision geometry.
- A function over space giving each point's distance to the nearest object surface, signed to distinguish inside from outside.
- Collision detection done in two steps: first a coarse bounding-box pass to find likely pairs, then exact contact computation.
- Collision Filtering碰撞过滤(碰撞组)Declaring in advance which object pairs should never count as colliding, saving computation and avoiding contact that shouldn't exist.
- Continuous Collision Detection连续碰撞检测(隧穿问题)Finding collisions along an object's full motion path within a step, preventing fast-moving objects from passing straight through obstacles.
- A simulation or animation glitch where two objects that should block each other pass through one another instead.
- Contact Model接触模型The physics engine's rule for how much force and friction two touching objects generate.
- A contact model that allows a small amount of interpenetration and generates contact force like a spring-damper.
- Constraint Solver约束求解器The physics-engine module that computes contact and joint constraint forces, keeping objects from interpenetrating and joints from coming apart.
- A math problem asking for two sets of non-negative variables where one being positive forces the other to be zero — what contact-force solving reduces to.
- Projected Gauss-Seidel投影高斯-赛德尔求解器A classic iterative method that solves contact and joint forces one constraint at a time, clamping each to a valid range.
- How many rounds a physics engine spends per step correcting contact and joint constraints; more rounds means more accuracy but less speed.
- The numerical method a physics engine uses each step to advance velocity and position forward in time given the current forces.
- A physics simulation diverging numerically — the robot twitches violently, objects fly apart, and the state turns into NaN or huge numbers.
8.3Soft bodies, fluids, differentiable sim
Beyond rigid bodies to deformable ones — cloth, rope, liquid — and simulators that can compute gradients.
- Physics simulation of objects that change shape under force, like cloth, rope, dough, or liquid.
- Cloth Simulation布料仿真Computer simulation of how cloth stretches, bends, wrinkles, and collides — essential for research on manipulating clothing.
- Fluid Simulation流体仿真Numerically simulating how fluids like water or smoke flow and experience force inside a computer.
- A numerical method that cuts an object into many small elements, solves each approximately, and assembles them into the whole.
- Position-Based Dynamics基于位置的动力学A simulation method that satisfies constraints by directly correcting object positions, fast and stable, popular for cloth and soft bodies.
- A continuum simulation method where particles carry the material's state while forces are computed on a background grid.
- Smoothed Particle Hydrodynamics光滑粒子流体动力学A meshless simulation method that breaks a fluid into particles and computes motion by weighting over nearby particles.
- A simulation method that treats a material as a large number of independent particles and computes each one's forces and motion.
- A contact-simulation algorithm that uses barrier potentials to guarantee objects never interpenetrate, well suited to large soft-body deformation.
- A simulator whose output can be differentiated, so gradient descent can optimize actions or physical parameters directly.
8.4Rendering and sensor simulation
The other half beyond physics: rendering scenes into camera images, and simulating readings from sensors like touch.
- The process of computing a camera image from a 3D scene; it produces all of a simulator's camera-based sensor data.
- A rendering method that projects 3D triangles onto the screen and shades them pixel by pixel; fast, and standard for real time.
- Ray Tracing光线追踪A rendering method that traces rays backward from the camera through reflections and refractions; highly realistic but slow.
- Path Tracing路径追踪A rendering method that simulates light bouncing many times through a scene using randomly sampled rays, producing physically realistic images.
- Physically Based Rendering基于物理的渲染A rendering approach that models light and materials according to real optics, so objects look believable under any lighting.
- Photorealistic Rendering照片级真实感渲染Computing an image by simulating real optical behavior, so a simulated picture looks like a photo from a real camera.
- Rendering camera images for hundreds or thousands of parallel environments in one pass, feeding vision policies at high speed.
- Sensor Simulation传感器仿真Generating simulated readings from cameras, lidar, IMUs, force sensors, and other sensors inside a simulator.
- Simulating what a tactile sensor would read inside a physics simulator, so tactile-equipped policies can be trained in sim.
- TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile SensorsTACTOMeta's open-source visuotactile simulator, using PyBullet to render tactile images for sensors like DIGIT.
- An example-based simulation model for GelSight visuotactile sensors, calibrated from a small amount of real data.
- NVIDIA's open-source GPU library for simulating vision-based tactile sensors and training policies that use them.
8.5Major simulators and frameworks
Physics and rendering in practice: the MuJoCo and Isaac families, plus other commonly used simulators.
- An open-source physics engine known for accurate contact simulation, maintained by Google DeepMind.
- MuJoCo XLAMJXA JAX implementation of MuJoCo for batched, parallel simulation on GPU/TPU that also supports differentiation.
- A differentiable rigid-body physics engine from Google written in JAX, bundled with a reinforcement-learning training library.
- A collection of GPU-accelerated robot learning environments led by DeepMind, built for fast single-GPU training and zero-shot transfer to real robots.
- A GPU rewrite of MuJoCo built by DeepMind and NVIDIA on the Warp framework, aimed at massive batched parallelism.
- NVIDIA's open-source real-time physics engine, and the physics core behind Isaac Sim and Isaac Lab.
- NVIDIA's GPU reinforcement-learning simulator launched in 2021; now discontinued.
- NVIDIA Isaac SimIsaac SimNVIDIA's Omniverse-based, open-source robot simulator built for high physical and visual fidelity.
- NVIDIA Isaac LabIsaac LabNVIDIA's open-source, GPU-parallel robot learning framework for training policies at scale in simulation.
- Orbit / IsaacGymEnvs / OmniIsaacGymEnvs (predecessors of Isaac Lab)Orbit / IsaacGymEnvs(Isaac Lab 前身)Three now-discontinued NVIDIA robot-learning frameworks that were later unified into Isaac Lab.
- Lightwheel's open-source Isaac Lab teleoperation data-collection framework, plugging the SO-101 arm into the LeRobot pipeline.
- Newton Physics EngineNewton 物理引擎An open-source GPU physics engine for robotics jointly launched by NVIDIA, DeepMind, and Disney.
- A lightweight GPU robot-learning framework combining an Isaac Lab-style interface with MuJoCo Warp physics.
- A multi-physics GPU simulation platform open-sourced in late 2024, built for speed and unified rigid-body, soft-body, and fluid simulation.
- The Python interface to the Bullet physics engine, and one of the most widely used simulators in early robot reinforcement learning.
- A multi-body physics engine for robotics and reinforcement learning, known for fast and accurate contact solving.
- A UC San Diego robot simulation platform specialized for interacting with articulated objects like drawers and cabinet doors.
- An open-source robot manipulation simulation framework and benchmark from a UC San Diego team, built for GPU parallelism.
- A modular, MuJoCo-based robot-learning simulation framework and benchmark, maintained under the ARISE initiative.
- robomimicRoboMimicA learning-from-demonstration framework from Stanford and UT Austin researchers, with standard datasets and offline learning algorithms.
- The most widely used open-source robot simulator in the ROS ecosystem, maintained by Open Robotics.
- An open-source desktop robot simulator maintained by Switzerland's Cyberbotics, beginner-friendly and popular for teaching.
- A general-purpose robot simulator from Coppelia Robotics in Switzerland, the successor to V-REP.
- Game Engine游戏引擎Software framework for building video games, combining rendering, physics, and animation, and often repurposed for simulation.
- An open-source, Unreal Engine-based self-driving simulator commonly used for closed-loop testing of driving policies.
- A robot-learning platform unifying many simulators behind one interface, with a bundled synthetic dataset and benchmark.
8.6Scenes, assets, and indoor platforms
With a simulator ready, fill it with content: object assets, ready-made indoor scenes, and automated scene generation at scale.
- The robot, object, and scene models used in simulation, which need to look right and also carry physical properties.
- An object made of parts connected by joints that can rotate or slide relative to each other.
- SimReady AssetsSimReady 资产An NVIDIA 3D-asset specification: models ship with physics, materials, and semantic labels already attached, ready to drop into simulation.
- Heightfield Terrain高度场地形Representing uneven terrain with a 2D grid that stores the ground height at each point — the standard approach for legged-robot training.
- The Allen Institute for AI's interactive indoor 3D simulator, widely used for navigation and household tasks.
- VirtualHome: Simulating Household Activities via ProgramsVirtualHome 家庭活动仿真A Unity-based simulation platform where virtual humans perform household chores driven by step-by-step 'programs.'
- Meta's open-source embodied-AI simulation platform, used mainly for indoor navigation, object rearrangement, and human-robot collaboration tasks.
- Stanford's interactive indoor simulation platform, used for navigation and mobile-manipulation research in household settings.
- Stanford's Omniverse-based household simulator, the underlying engine behind the BEHAVIOR-1K benchmark.
- Shanghai AI Lab's city-scale embodied-AI simulation platform built on Isaac Sim, formerly named GRUtopia.
- Automatically producing large numbers of scenes, objects, or terrain from rules plus randomness, instead of building each one by hand.
- Allen Institute for AI's framework for procedurally generating large numbers of interactive 3D houses for embodied-AI training.
- InfinigenInfinigen 程序化世界生成Princeton's open-source procedural 3D world generator, where every asset is generated from scratch by randomized mathematical rules.
- Using foundation models to automatically generate simulation tasks, scenes, and training supervision, producing robot training data at scale.
- HolodeckHolodeck 语言生成三维环境Given a one-sentence description, automatically generates an interactive 3D indoor scene using GPT-4 and Objaverse assets.
- The Allen Institute for AI's open ecosystem of large-scale indoor simulation scenes and robot evaluation tools.
- Genie Sim智元 Genie SimAgiBot's open-source embodied-simulation platform, combining scene generation, synthetic data, and an evaluation benchmark.
8.7From simulation to reality
Moving what you trained in sim onto a real robot: understand the sim-to-real gap, then domain randomization and pulling reality into sim.
- The mismatch between simulation and reality that makes a policy trained in sim perform worse on a real robot.
- Sim-to-Real Transfer仿真到现实迁移Training a robot policy in simulation, then deploying it to work on a real robot.
- Real-to-Sim现实到仿真Recreating a real scene, its objects, and the robot inside a simulator, so the simulation matches reality as closely as possible.
- Real-to-Sim-to-Real真-仿-真闭环Recreating a real scene in simulation, training or generating data there, and then deploying the resulting policy back to the real robot.
- Randomizing appearance and physics in simulation during training so the real world looks like just another variation.
- Dynamics Randomization动力学随机化Randomizing physical parameters like mass and friction during training so a policy can handle a real robot.
- Visual Randomization视觉随机化Randomly varying a simulation's textures, colors, lighting, and camera during training so a vision policy isn't picky about how things look.
- NVIDIA Omniverse ReplicatorOmniverse Replicator 合成数据生成NVIDIA Omniverse's synthetic-data framework that randomizes simulated scenes and automatically outputs labeled training data.
- A domain-randomization method, proposed by OpenAI, that automatically widens its randomization range as the policy gets better.
- Simulating how a real motor's torque output actually responds to a command, to narrow the sim-to-real gap.
- Factory / IndustRealFactory / IndustReal 接触丰富装配仿真NVIDIA's contact-rich assembly simulation and transfer pipeline: simulate tasks like nut-threading fast, then transfer the policy to a real robot.
- Sim-to-Sim Transfer仿真到仿真迁移Running a policy trained in one simulator inside a different simulator, to expose problems before they reach the real robot.
- Connecting a real controller or real control code to a simulated plant for testing — a standard validation step before deploying to a real robot.
- Digital Twin数字孪生A virtual copy of a real object, robot, or factory that's kept in sync using real data.
- Digital Cousin数字表亲A simulated scene that only needs to be geometrically and semantically similar to reality, not an exact one-to-one replica.
- Rendering simulated scenes with 3D Gaussian Splatting reconstructions of real places, paired with a physics engine for motion.
- A robot simulator that reconstructs real scenes with 3D Gaussian Splatting and couples them to a physics engine.
- An open-source real-to-sim-to-real simulation framework combining 3D Gaussian Splatting rendering with MuJoCo physics.
- A high-throughput simulator from Tsinghua and others that uses batched 3D Gaussian Splatting to render photorealistic images quickly.
8.8Evaluation methods and metrics
Training done, now the exam: how success rate is calculated, how to test in sim and reality, and what makes results comparable.
- Success Rate成功率The fraction of attempts at the same task that a policy completes successfully — the most common robotics metric.
- Running a policy on an actual robot repeatedly and recording success rate and other outcomes.
- Running a policy through many tasks in a simulator and tallying success rate, instead of or alongside real-robot testing.
- Letting a policy actually control the robot, act, observe, and act again, scored by whether the task finishes.
- Comparing a model's predicted actions against recorded demonstration actions on offline data, without letting the policy actually control anything.
- Off-Policy Evaluation离线策略评估Estimating how much return a new policy would actually get once deployed, using only data collected by other policies.
- Benchmark基准测试A fixed set of tasks, data, and scoring rules that let different methods be compared under the same conditions.
- The set of rules specifying under what conditions a policy is tested, how many trials, and what counts as success.
- Baseline基线方法An existing or simple method used as a point of comparison to show how much a new method improves.
- The best publicly reported result on a given task or benchmark, commonly abbreviated SOTA.
- Ablation Study消融实验Removing or swapping one component of a method to see how performance changes, testing whether it matters.
- Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds)评测统计显著性(置信区间 / 序贯检验 / 多随机种子)Using statistics to tell whether a gap in success rate between two policies is real or just noise.
- Deliberately changing lighting, positions, objects, or instructions to test how much skill a policy retains outside its training conditions.
- Progress Score进度分数A metric that scores how many steps of a task a robot completed, rather than just recording binary success or failure.
- How often a human has to take over or correct a policy while it runs autonomously; lower means more independent operation.
- How long, on average, a robot can keep working continuously before a human has to step in.
- Sim-to-Real Correlation仿真-真机相关性A measure of whether simulated evaluation scores actually reflect real-robot performance, usually reported as a Pearson correlation coefficient r.
- Mean Maximum Rank Violation平均最大排名违背A metric for whether a simulated evaluation's ranking of policies matches the real-robot ranking; lower is better.
- Elo RatingElo 评分A scoring method that estimates each player's or policy's relative strength from a series of head-to-head wins and losses.
- An evaluation method where a judge, without knowing which model is which, decides only which of two policies performed better.
- When leading models' scores on a benchmark all cluster near the maximum, so it can no longer distinguish good methods from bad ones.
8.9Tabletop manipulation benchmarks
Applying that method to real test sets, starting with the most common tabletop arm-manipulation benchmarks.
- LIBERO BenchmarkLIBEROA simulated benchmark of 130 tabletop manipulation tasks, one of the most commonly reported in VLA papers.
- A benchmark that adds 7 kinds of perturbations to LIBERO specifically to test the robustness of VLA models.
- Adds object, position, instruction, and environment perturbations to LIBERO to test whether a VLA model is just memorizing.
- A simulated evaluation suite that replicates common real-robot manipulation setups for cheap policy scoring.
- A SimplerEnv evaluation setup that makes the simulated image look as close as possible to real-robot camera footage.
- One of SimplerEnv's two evaluation modes: testing a policy across many visually randomized scene variants and averaging the results.
- CALVIN BenchmarkCALVINA tabletop manipulation benchmark testing whether a robot can follow 5 language instructions in a row.
- Average Length (CALVIN)平均完成长度In the CALVIN benchmark, how many consecutive instructions a policy completes on average, out of 5.
- A simulation benchmark of 100 robot-arm manipulation tasks from Imperial College London, built on CoppeliaSim.
- An RLBench-based simulation benchmark that tests a language-conditioned manipulation policy's generalization across four difficulty levels.
- A benchmark applying 14 kinds of systematic environmental perturbations to RLBench tasks to test manipulation policies' generalization.
- A 2D task where a round pusher must push a T-shaped block to a target pose, commonly used to test imitation-learning policies.
- ALOHA Sim (Transfer Cube / Insertion)ALOHA 仿真任务The two bimanual simulated tasks — cube transfer and peg insertion — built in MuJoCo for the ACT paper.
- A bimanual-manipulation simulation data generator and evaluation benchmark from the University of Hong Kong, Shanghai AI Lab, and others.
- A simulation benchmark of 50 robot-arm manipulation tasks, used to test multi-task and meta-reinforcement learning.
- A MuJoCo kitchen scene where a Franka arm completes sub-tasks in sequence, such as opening a microwave and moving a kettle.
- AdroitAdroit 灵巧手任务A benchmark for controlling a 24-DoF simulated five-fingered hand in MuJoCo to open doors, hammer nails, and more, across four tasks.
- A reinforcement-learning benchmark from a Peking University team, using two Shadow dexterous hands in Isaac Gym.
- A deformable-object manipulation benchmark built on NVIDIA FleX, covering cloth, rope, and liquid tasks.
- A tabletop manipulation benchmark driven by interleaved text-and-image prompts, with four levels of generalization testing.
- A manipulation simulation and evaluation platform from Shanghai AI Lab, built on Isaac Sim, that uses a large model to auto-generate tasks.
- A benchmark in Isaac Sim testing whether a robot can manipulate objects to a specified continuous state based on language.
- A Fudan University language-conditioned manipulation benchmark focused on common sense, implicit intent, and long-horizon multi-step reasoning.
- A large simulation benchmark testing planning, reflection, and memory in long-horizon robot manipulation.
- A tabletop-manipulation benchmark specifically testing robot memory, with tasks requiring recall of information that's occluded or gone.
- Peking University's open-source VLA evaluation framework, grading capability boundaries across task, language, and vision axes.
- NVIDIA Isaac Lab-ArenaIsaac Lab-ArenaAn open-source NVIDIA extension to Isaac Lab for composing simulation benchmarks and evaluating robot policies in parallel.
- RoboFinals (Lightwheel industrial-grade simulation evaluation platform)光轮 RoboFinals 工业级仿真评测平台An industrial-grade simulation evaluation platform from Lightwheel, built specifically to test VLA and other robot foundation models.
8.10Household, navigation, and QA benchmarks
Expanding from the tabletop to a whole house: long-horizon chores, language navigation, embodied QA, and LLMs as the brain.
- A large-scale kitchen-household simulation framework and benchmark from UT Austin; its newest version is called RoboCasa365.
- BEHAVIOR-1K (BEHAVIOR Challenge)BEHAVIOR-1KStanford's simulated benchmark of 1,000 everyday household tasks, built on the OmniGibson simulator.
- Open-Vocabulary Mobile Manipulation开放词汇移动操作基准A benchmark that has a robot find any named object in an unfamiliar house and place it on a specified piece of furniture.
- A benchmark that has an agent complete multi-step household chores in a simulated home by following natural-language instructions.
- A text-game version of the ALFRED household tasks, letting an agent learn in text first and then act on the actual visuals.
- Room-to-RoomR2R / VLN-CE 视觉语言导航基准The most widely used vision-language navigation benchmark: follow a human-written route description to reach a destination in an unfamiliar house.
- Navigation Error / Oracle Success Rate / Trajectory Length导航误差 / Oracle 成功率 / 轨迹长度The three standard Vision-and-Language Navigation metrics: how close the agent stops to the goal, whether it ever passed it, and total distance traveled.
- Success weighted by Path Length路径长度加权成功率A navigation metric that credits both reaching the goal and taking a path close to the shortest possible route.
- normalized Dynamic Time Warping归一化动态时间规整A 0-to-1 navigation metric for how closely an agent's path matches the shape and order of a reference path.
- OpenEQA (Open-Vocabulary Embodied Question Answering Benchmark)OpenEQA 开放词汇具身问答基准A Meta benchmark testing whether an agent can answer natural-language questions using what it has observed of a real environment.
- ERQAERQA 具身推理问答基准A 400-question multiple-choice benchmark from Google DeepMind, pairing images and text to test embodied reasoning.
- A benchmark with 4 environments testing how well multimodal large models perform as the “brain” of an embodied agent.
- Embodied ArenaEmbodied Arena 具身评测竞技场An evaluation platform that plugs in benchmarks for embodied question answering, navigation, and task planning under one live leaderboard.
8.11RL and locomotion-control benchmarks
A different kind of task: classic reinforcement-learning testbeds, plus environments for training legged and humanoid robots to walk.
- Arcade Learning Environment (ALE) / Atari 100kAtari 游戏基准(街机学习环境 / Atari 100k)The standard platform for testing reinforcement learning on Atari 2600 games; Atari 100k is its low-sample variant.
- Minecraft EnvironmentsMinecraft 环境(MineDojo / MineRL)Agent training and evaluation platforms built on the game Minecraft, used for open-world research.
- Gym/Gymnasium MuJoCo TasksGym MuJoCo 连续控制任务A classic set of MuJoCo-based continuous-control environments in Gymnasium, a standard test bed for reinforcement-learning papers.
- DeepMind Control SuiteDeepMind 控制套件DeepMind's standard set of MuJoCo-based continuous-control reinforcement-learning tasks.
- D4RLD4RL 离线强化学习基准The most widely used standard datasets and benchmark for offline reinforcement learning.
- OGBench: Benchmarking Offline Goal-Conditioned RLOGBench 离线目标条件强化学习基准A benchmark built specifically to test offline goal-conditioned reinforcement learning algorithms, with 8 environment types and 85 datasets.
- ETH Zürich's open-source reinforcement-learning training environment for legged robots, built on Isaac Gym.
- unitree_rl_gym (Unitree RL Gym)unitree_rl_gymUnitree's official open-source reinforcement-learning example framework, covering everything from simulated training to real-robot deployment.
- robot_lab (RL extension library based on Isaac Lab)robot_lab(Isaac Lab 强化学习扩展库)An open-source library extending Isaac Lab with ready-made reinforcement-learning training environments for many legged and humanoid robots.
- RobotEra's open-source reinforcement-learning framework for humanoid walking, built around zero-shot sim-to-real transfer.
- NVIDIA's open-source framework for training digital humans and humanoid robots to move using GPU physics simulation.
- A simulated humanoid benchmark from Berkeley and others, covering 27 whole-body locomotion and manipulation tasks.
- A MuJoCo-based imitation-learning benchmark for whole-body locomotion, with humanoid, quadruped, and human musculoskeletal models.
- Mean Per-Joint Position Error平均关节位置误差The average distance between predicted and ground-truth joint positions, measuring how accurate a pose estimate or motion tracking is.
8.12Real-world and world-model evaluation
Beyond fixed sim test sets: large-scale real-robot evaluation, using world models to evaluate policies, and evaluating world models themselves.
- A reproducible real-world furniture-assembly benchmark from KAIST and others, testing long-horizon, precise robot manipulation.
- Berkeley's round-the-clock, unattended real-robot evaluation system, which judges success automatically and resets the scene itself.
- A generalist-robot-policy evaluation platform where multiple labs run blind head-to-head real-robot comparisons that get aggregated into a ranking.
- A real-robot online evaluation platform from Dexmal and Hugging Face; its first benchmark is called Table30.
- A unified sim-and-real benchmark for general-purpose manipulation policies, with 42 simulated tasks and 18 real-robot tasks.
- RobotArena Infinity (RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation)RobotArena ∞A benchmark that auto-converts real robot videos into simulated scenes, then ranks VLA models by scoring and human voting.
- Neural Simulator神经模拟器A simulator learned from data with a neural network, predicting what happens next instead of relying on hand-written physics equations.
- Using an interactive video world model in place of a real robot, running the policy in closed loop to estimate its performance.
- Using a video world model instead of a real robot to run rollouts and score and rank robot policies.
- Fréchet Inception Distance弗雷歇初始距离A metric for how close a batch of generated images is to real images in overall distribution; lower is better.
- Fréchet Video Distance弗雷歇视频距离A metric for how close a batch of generated videos is to real videos overall; lower is better.
- Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch SimilarityPSNR / SSIM / LPIPS 图像相似度指标Three widely used metrics for how similar a generated image is to a reference image, each closer to human perception than the last.
- An open-source benchmark that scores video-generation quality separately across 16 dimensions instead of giving one overall number.
- WorldScore: A Unified Evaluation Benchmark for World GenerationWorldScore 世界生成评测Stanford's unified world-generation benchmark that lets 3D and 4D scene-generation models and video models be compared on the same scale.
- A benchmark of real filmed footage that tests whether video-generation models actually understand physics, not just look realistic.
- EWMBenchEWMBench 具身世界模型评测A benchmark from AgiBot and others that specifically evaluates whether robot-manipulation video generation models get the details right.
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsWorldArena 具身世界模型基准A Tsinghua-led embodied world-model benchmark scoring both how good the generated video looks and how useful it actually is.