Early build · scores computed by the deterministic engine from real, sourced developments (30 days) · weights v1.0-prior
LatestLearn › Basics
Vikshy Learn · Basics

Multimodal

An AI that can handle more than just text, like images, audio, or video.

Most early AI chatbots only worked with words. A multimodal model can take in and often produce several kinds of content, such as pictures, sound, and video, not only text. This lets you show it a photo and ask what is in it, or have it read a chart, describe a scene, or listen to a voice note.

For example, You snap a photo of the inside of your fridge and ask a multimodal AI what you could cook for dinner.

← All 42 concepts in the glossary