Skip to content
DatasetAdvanced

The Stack Source Code Dataset

The Stack is a large, permissively licensed source code dataset from the BigCode project, covering hundreds of programming languages. It is widely used to pre-train and evaluate code generation models, and includes opt-out mechanisms and deduplication tooling to support responsible data practices.

Overview

"The Stack Source Code Dataset" is a "Dataset" resource curated by AI Resource Hub, filed under the Datasets category and suited to Advanced-level learners. It is provided by BigCode, was last updated on 2026-06-22, and holds an editorial score of 4.5/5 from our team. Click "Visit Resource" on the right to open the original page.

Tags

The StackCodeDatasetBigCode

Key Features

  • Large corpus of source code
  • Many programming languages
  • Used to train code models

Pros

  • +Foundational for code LLMs
  • +Broad language coverage
  • +Foundational training data for code models

Cons

  • License filtering is important
  • License filtering is essential before use
  • Huge volume needs serious storage and compute

FAQ