본문으로 건너뛰기

Multimodal

Models that understand and generate across multiple data types, such as text, images, audio, and video.

관련 리소스