Dynamic / Native Resolution
动态分辨率(原生分辨率输入)AdvancedLetting a vision encoder cut an image at its original size, so bigger images produce more visual tokens instead of being rescaled.
Early Vision Transformers and CLIP-style encoders required rescaling or cropping every image to a fixed size, such as 224×224, which loses detail or distorts aspect ratio. Dynamic resolution instead keeps the image's original size and aspect ratio and cuts it directly into fixed-size patches, so a bigger image simply produces more tokens. Google's 2023 NaViT used 'sequence packing' to fit images of different sizes into the same training batch; Alibaba's 2024 Qwen2-VL introduced Naive Dynamic Resolution, using 2D rotary position encoding inside the vision encoder to record each patch's row and column, letting it handle any resolution. The benefit is that small images save compute while large images keep their detail, at the cost of a token count that grows with resolution. VLA backbones using this usually cap the token count within a min/max range to control inference latency.
ExampleQwen2-VL maps roughly every 28×28 pixels to one visual token, with a default range of 4 to 16,384 tokens per image, adjustable via min_pixels and max_pixels to balance speed against memory.
- Also called
- Native Resolution, Naive Dynamic Resolution
- Related
- Vision Encoder · Vision Transformer · Visual Token · Qwen-VL · Rotary Position Embedding · Inference Latency
- Sources
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution (arXiv:2307.06304)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv:2409.12191)
Qwen/Qwen2-VL-7B-Instruct model card