Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

K-Culture GlossaryTechnical words in the news

VLM (Vision Language Model)

VLM

A language model with the ability to "see." Give it a photo, and it reads and explains what's in it.

In plain words

A VLM (Vision Language Model) is a language model with the ability to "see." Feed it a photo, illustration, or document, and it reads the content, describes it in words, and answers questions about it.

If a text-only language model (LLM) is a pen-pal who only knows writing, a VLM is a conversation partner with eyes. It can grasp a whole report full of charts, read an error message in a screenshot, or spot the ingredients in a photo of your fridge.

Think of VLM as the branch of the bigger "multimodal" umbrella that focuses specifically on "vision + language." Lately, derivatives built for robotics and self-driving also show up often — these are VLA models, which understand what a camera sees and translate that into action. Whenever you spot VLM in an article, just read it as "an AI brain that understands images."

How it shows up in the news

"New VLM tops chart and document understanding benchmark" — how accurately a model reads "hard-to-parse" visuals like tables, graphs, and handwriting is the main event in VLM competition.

Try it yourself

  1. Capture a screen showing a chart or graph from a news article or report.
  2. Attach it to a chatbot and ask: "Explain the gap between the #1 and #3 items in this chart, and tell me in one line what happens if this trend continues."
  3. Reading text, bars, and axes together and then reasoning about them — that's "visual understanding" beyond simple image description, and it's what VLMs are judged on. Try it with handwritten notes or complex tables too.

See also

Stories using this term

No story has used this term yet. New ones attach here automatically.

Browse every entry