DatasetAdvanced
The Stack Source Code Dataset
The Stack is a large, permissively licensed source code dataset from the BigCode project, covering hundreds of programming languages. It is widely used to pre-train and evaluate code generation models, and includes opt-out mechanisms and deduplication tooling to support responsible data practices.
Overview
"The Stack Source Code Dataset" is a "Dataset" resource curated by AI Resource Hub, filed under the Datasets category and suited to Advanced-level learners. It is provided by BigCode, was last updated on 2026-06-22, and holds an editorial score of 4.5/5 from our team. Click "Visit Resource" on the right to open the original page.
Tags
The StackCodeDatasetBigCode
Key Features
- ▹Large corpus of source code
- ▹Many programming languages
- ▹Used to train code models
Pros
- +Foundational for code LLMs
- +Broad language coverage
- +Foundational training data for code models
Cons
- −License filtering is important
- −License filtering is essential before use
- −Huge volume needs serious storage and compute
FAQ
Visit Resource →
Details
- Pricing
- Free / open dataset
- Author
- BigCode
- Editorial score
- ★ 4.5 / 5
- Last updated
- Jun 22, 2026