Datasets
High-quality open-source AI training datasets
신규Hugging Face Datasets Collection
Explore the vast collection of open-source datasets on Hugging Face, covering high-quality training data across NLP, CV, audio, and more.
신규Kaggle Datasets Platform
Kaggle offers tens of thousands of public datasets spanning machine learning competitions, real-world business, and research, all available for online analysis and download.
신규Papers with Code Datasets
A collection of benchmark datasets used in academic papers, linked to their corresponding SOTA models and code implementations—ideal for research and reproduction.
신규Google Dataset Search
A dataset search engine from Google that lets you find public datasets across sources and quickly locate the data you need for research and training.
신규Common Crawl Web Corpus
An open, petabyte-scale web crawl corpus that serves as a key source of pre-training data for many large language models, available for free.
신규LAION Open Multimodal Datasets
LAION provides large-scale, open-source image-text pair datasets widely used to train multimodal models such as CLIP and Stable Diffusion.
신규UCI Machine Learning Repository
A classic machine learning dataset repository with hundreds of structured datasets, ideal for hands-on practice and algorithm teaching.
신규The Stack Source Code Dataset
A large, permissively licensed source-code dataset from BigCode, widely used to pre-train and evaluate code generation models.
신규ImageNet
A landmark large-scale image dataset with millions of labeled images across thousands of categories, foundational to modern computer vision.