The Wider Lens logoThe Wider Lens
← All topics

Beginner

What is Multimodal AI?

Models that see, hear, and read — not just text.

Multimodal AI handles multiple types of input and output: text, images, audio, and video. GPT-4o, Gemini, and Claude can all 'see' images and discuss them; newer systems generate video and speech too.

Under the hood, different encoders convert each modality into a shared representation the model can reason over — an image becomes a sequence of visual tokens alongside text tokens.

Multimodality unlocks use cases text-only models can't touch: analyzing charts, describing photos, transcribing meetings, narrating video, and controlling robots that perceive the world.

Key points

  • Handles text, images, audio, video
  • Shared representation across modalities
  • Enables vision, speech, and video use cases
  • GPT-4o, Gemini, Claude are multimodal