Skip to content

Multimodal

Models that understand and generate across multiple data types, such as text, images, audio, and video.

For example, you can show a multimodal model a screenshot and ask it to write the code for that UI, or hand it a chart and ask for an analysis. GPT and Gemini are widely used multimodal models.

Related resources