Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Aggregating between 100,000 and 1,000,000 non-anonymized news articles and summaries sourced from CNN and Daily Mail. Curated by ccdv and last updated in 2022, it provides paired text for training and evaluating abstractive summarization models.
The 'highlights' column uses <s> and </s> tags as delimiters for individual summary points which may require specific parsing or tokenization strategies before model training.