Embodied AI Glossary中文

Environment (Env; reset/step interface)

环境(Env)与 reset / step 接口EnvEssential

The standard reinforcement-learning interface: reset starts an episode, step executes one action.

In reinforcement learning, the environment is the world an agent interacts with. Originating with OpenAI Gym and now maintained by the Farama Foundation as Gymnasium, it's standardized into a programming interface: reset() starts a new episode and returns the initial observation plus an info dictionary; step(action) executes one action and returns the new observation, the reward, terminated (the episode ended in a meaningful sense, such as success or falling over), truncated (the episode was cut off externally, such as by a timeout), and info. The environment also declares its action_space and observation_space, fixing the format of actions and observations. Frameworks such as Isaac Lab follow a similar convention, which is what lets the same training algorithm be pointed at different environments. The older Gym API had a single done flag; it was split into two because, on a timeout truncation, the value estimate still needs to bootstrap from the next state, and conflating the two cases leads to incorrect updates.

ExampleA typical loop: obs, info = env.reset(), then repeatedly obs, reward, terminated, truncated, info = env.step(policy(obs)), resetting whenever either termination flag comes back true.

Also called
Gym interface, Gymnasium API, Env
Related
Agent–Environment Interaction · Episode · Observation · Action Space · Termination vs. Truncation · Gymnasium
Sources
Gymnasium Documentation: Env API

See it in the full glossary →