Socratic Models
Socratic Models(苏格拉底模型)SMsAdvancedA framework that chains several off-the-shelf large models together zero-shot using natural language as the common interface for multimodal tasks.
Socratic Models was released by Andy Zeng, Pete Florence, and colleagues at Google in April 2022. Different foundation models have different strengths: vision-language models (VLMs) understand images, large language models (LLMs) understand commonsense and reasoning, and audio models understand sound. Rather than fine-tuning any of them, Socratic Models uses natural language as a shared interface: one model's output is written into another model's prompt, letting them exchange information as if in conversation and compose new capabilities zero-shot. The paper demonstrates this on first-person video question answering, multimodal assistant dialogue, and robot perception and planning: a vision model first turns tabletop objects into text descriptions, then an LLM breaks the instruction down into a sequence of pick-and-place actions, which are handed to a pretrained language-conditioned policy for execution. Along with SayCan and Code as Policies from the same period, it represents an early approach to using large language models as robot planners.
ExampleA user says 'put all the fruit in the bowl'; a vision model first lists an apple, a banana, and a bowl on the table, and a language model writes this out as 'pick up the apple and put it in the bowl; pick up the banana and put it in the bowl,' which a lower-level pick-and-place policy then executes step by step.
- Also called
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- Related
- Large Language Model · Vision-Language Model · Zero-shot · SayCan · Code as Policies · Inner Monologue
- Sources
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language (arXiv 2204.00598)
Socratic Models project page - As of
- 2022-05