Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

K-Culture GlossaryWords you meet while using AI

Multimodal

The ability to understand and generate not just text, but images, sound, and video together.

In plain words

Modality refers to a form that information takes — text, images, sound, video. Multimodal means a single model handles several of these forms at once.

Early language models could only read and write text. They were like pen pals with no eyes or ears. Multimodal models can look at a photo and describe what's in it, listen to spoken words and reply in a voice, and make sense of a document that mixes tables, charts, and text all at once.

Today's leading models (the latest GPT, Claude, Gemini) are multimodal by default. That's what makes it possible to snap a photo of your fridge and get a recipe suggestion, and it's why "voice conversation" and "screen understanding" demos are now a standard part of any new model launch.

How it shows up in the news

"The company emphasized multimodal capabilities spanning images, voice, and video" — meaning that abilities beyond text, especially the capacity to "see and hear," have become a key selling point.

Try it yourself

  1. Take a photo of a handwritten note, a receipt, or a whiteboard.
  2. Attach the photo to ChatGPT or Claude and type "Turn what's written here into a table."
  3. Turning something that isn't text (a photo) into text — that's multimodal at work.

See also

Stories using this term

No story has used this term yet. New ones attach here automatically.

Browse every entry