Definition·
Capabilities

Multimodal

The ability of an AI model to understand and generate multiple modalities (text, image, audio, video) inside a single system.

Detailed explanation

A multimodal model accepts, for example, an image + text and produces a written answer, or text and produces a video. This unification unlocks richer use cases: screenshot analysis, vision-language for robotics, full voice assistants.

Examples

GPT-4o that sees and hears
Gemini multimodal
Claude 3.5 reading illustrated PDFs
Sora for video

Frequently asked questions

Are all LLMs multimodal?

No: some stay text-only. Multimodal models are gaining ground but cost more.

Related terms

Last updated: 7/15/2026

Talent AI

Turn theory into practice

Post a mission or join the community of top AI, Data and Machine Learning experts.