Most early AI chatbots only worked with words. A multimodal model can take in and often produce several kinds of content, such as pictures, sound, and video, not only text. This lets you show it a photo and ask what is in it, or have it read a chart, describe a scene, or listen to a voice note.
For example, You snap a photo of the inside of your fridge and ask a multimodal AI what you could cook for dinner.