Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Betterdataset 2M is a mixed pretraining corpus assembled by GODELEV from four open-source datasets. It contains approximately 2 million rows and an estimated 1.59 billion tokens, intended for training language models. The dataset was last updated on Hugging Face on June 1, 2026.
License is unknown, which may impose usage restrictions. The dataset is split into 20 parquet files.