Multimodal
Models that understand and generate across multiple data types, such as text, images, audio, and video.
For example, you can show a multimodal model a screenshot and ask it to write the code for that UI, or hand it a chart and ask for an analysis. GPT and Gemini are widely used multimodal models.