Datasets
High-quality open-source AI training datasets
NUEVOHugging Face Datasets Collection
Explore the vast collection of open-source datasets on Hugging Face, covering high-quality training data across NLP, CV, audio, and more.
NUEVOKaggle Datasets Platform
Kaggle offers tens of thousands of public datasets spanning machine learning competitions, real-world business, and research, all available for online analysis and download.
NUEVOPapers with Code Datasets
A collection of benchmark datasets used in academic papers, linked to their corresponding SOTA models and code implementations—ideal for research and reproduction.
NUEVOGoogle Dataset Search
A dataset search engine from Google that lets you find public datasets across sources and quickly locate the data you need for research and training.
NUEVOCommon Crawl Web Corpus
An open, petabyte-scale web crawl corpus that serves as a key source of pre-training data for many large language models, available for free.
NUEVOLAION Open Multimodal Datasets
LAION provides large-scale, open-source image-text pair datasets widely used to train multimodal models such as CLIP and Stable Diffusion.
NUEVOUCI Machine Learning Repository
A classic machine learning dataset repository with hundreds of structured datasets, ideal for hands-on practice and algorithm teaching.
NUEVOThe Stack Source Code Dataset
A large, permissively licensed source-code dataset from BigCode, widely used to pre-train and evaluate code generation models.
NUEVOImageNet
A landmark large-scale image dataset with millions of labeled images across thousands of categories, foundational to modern computer vision.