Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A curated collection of text data in English, French, German, Spanish, and Italian, culled from sources including web data, video subtitles, academic papers, digital books, newspapers, and magazines. The dataset was used to pretrain the Lucie-7B foundation LLM and also contains samples of diverse programming languages. It was authored by OpenLLM-France and last updated on the Hugging Face platform in May 2025.
License is unknown; terms of use must be verified before application.