Embodied AI Glossary中文

Hallucination

幻觉Common

A model confidently producing output that does not match the facts or the scene in front of it.

Hallucination refers to a large model generating content that sounds plausible but does not match the facts or the actual input — first widely discussed in large language models, such as inventing a paper that does not exist; vision models can likewise “see” objects that are not actually in the image. In embodied settings, hallucination turns from “saying the wrong thing” into “doing the wrong thing”: a planner might send a robot to fetch something that is not actually in the room, or plow ahead trying to execute an instruction that cannot be completed. The 2025 HEAL study (EMNLP 2025 Findings) specifically tests large-language-model-driven embodied agents, evaluating 12 models across two simulated environments, and finds that models generally fail to notice when an instruction does not match the scene; on a test set built specifically to probe this, hallucination rates reached up to 40 times those seen with ordinary prompts. Common countermeasures include checking plans against perception results, having the model ask a person when uncertain, and uncertainty estimation.

ExampleThere is no apple in the kitchen, but the user says, “put the apple in the fridge,” and a large-language-model-based planner still generates steps such as “walk to the apple, pick up the apple.”

Related
Large Language Model · Vision-Language Model · Language Grounding · Uncertainty Estimation · KnowNo · Embodied Safety
Sources
A Survey on Hallucination in Large Language Models (arXiv 2311.05232)
HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models (arXiv 2506.15065)
Hallucination (artificial intelligence) - Wikipedia
As of
2025-10

See it in the full glossary →