Development · Term

Multimodal

A multimodal model is one that accepts or produces more than one kind of content — combining text with images, audio, or video in a single request rather than handling each in a separate system.

Practically, this means a single model can read a screenshot and answer questions about it, transcribe speech and summarise it, or take a design and return the markup — work that previously required stitching several specialised services together.

The failure modes differ by mode. Models read fine detail in images unreliably — small text, precise counts, exact spatial relationships — so anything consequential extracted from an image should be verified against the source.

Related terms

Have something worth building right?