Embodied AI Glossary中文

RGB Camera

RGB相机Essential

An ordinary camera that outputs color images, recording red, green, and blue brightness at every pixel.

An RGB camera is simply the most common kind of color camera: its image sensor is covered by a color filter array, most commonly the Bayer filter patented by Kodak engineer Bryce Bayer in 1976, in which half the pixels sense green and a quarter each sense red and blue; a demosaicing algorithm then interpolates a full R, G, B value for every pixel. RGB cameras are cheap, high-resolution, and information-rich, making them the main input for most imitation-learning policies and VLA models. Their limitation is that a single RGB (monocular) camera only captures a 2D projection of the 3D world and has no direct sense of how far away things are; getting depth requires a depth camera, a stereo camera, or a monocular depth estimation model. Other factors to weigh when mounting one include field of view, frame rate, and whether it uses a global shutter (the whole frame exposed at once) or a rolling shutter (exposed row by row, which can skew fast-moving subjects).

ExampleThe original ALOHA dual-arm platform is fitted with 4 ordinary webcams: two mounted on the wrists, one facing forward, and one overhead looking down. ACT, a policy that uses only these color images plus joint angles — no depth — learned to open a semi-transparent condiment cup and insert a battery from about 10 minutes of demonstrations, reaching 80–90% success.

Also called
Color Camera, Monocular Camera
Related
Depth Camera · Stereo Camera · Wrist Camera · Monocular Depth Estimation · Camera Intrinsics · Global Shutter / Rolling Shutter
Sources
Wikipedia: Bayer filter
ALOHA / ACT project page (Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware)
arXiv 2304.13705: Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

See it in the full glossary →