Multimodal AI
Models and Architectures
AI that works with more than one kind of input or output - text, images, audio, and video together.
Early AI models handled one medium each: text models wrote, vision models looked. Multimodal models combine them - you can show a photo and ask questions about it, talk to it out loud, or have it generate an image from a text description.Modern flagship assistants like GPT-4o and Gemini are multimodal. For everyday use it means one tool can read a document, describe a chart, and listen to a voice note.