Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
mOSCAR is a multilingual web-crawled text corpus developed by the oscar-corpus organization. The dataset includes additional filtering steps to remove toxic content, a complete Spanish split, and face detection in images to blur them. The dataset page was last updated on 2024-11-23.
License information is unknown and should be verified before use.