Course 7, lesson 67 of 100, Ages 10+
AI that sees, hears and talks
One model, many senses
Like I’m 5
Some AIs can look at a picture, listen to your voice and talk back, all in one. They're called multimodal, which means many senses.
The big idea
Multimodal models take in more than text: photos, screenshots, audio and sometimes video. They turn each into tokens or vectors so the same model can reason about all of them together.
That lets you photograph a maths problem and ask for a hint, describe a plant from a photo, or have a spoken conversation in real time. They can still misread images, so check important details.
Examples
- Homework help: Snap a diagram and ask what it shows.
- Accessibility: An app describes surroundings to someone who is blind.
- Voice chat: Talk naturally and hear spoken replies.
How it works
- You give the AI a picture, sound or text.
- Each input is turned into numbers the model understands.
- The model combines them and answers.
Check your understanding
- What does multimodal mean?
- Options: Working with several kinds of input, like images and sound; Having many batteries; Only reading text.
Answer: Working with several kinds of input, like images and sound. Multi means many, and modes are kinds of input. - A multimodal app reads a photo of a sign wrongly. What should you do?
- Options: Double-check important details yourself; Trust it completely; Delete your phone.
Answer: Double-check important details yourself. Image understanding can slip, so verify what matters.
Remember
Multimodal AI combines text, images and sound in one model.
Talk about it
What would you ask an AI that could see through your camera?
Go deeper
Vision encoders such as vision transformers turn image patches into embeddings aligned with text. Models like CLIP showed how to align images and text with contrastive learning.