Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
NVIDIA's ClimbLab is a 1.2-trillion-token corpus for language model pre-training. It was created by OptimalScale using a semantic clustering method called CLIMB to reorganize and filter data from the Nemotron-CC and SmolLM-Corpus sources into 20 distinct clusters. The dataset was last updated on Hugging Face in May 2025.
License is unknown; users should verify terms of use before downloading.