Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Over 915 million unique tweets, including retweets, were collected from the Twitter Stream, with dedicated gathering starting March 11th yielding over 4 million tweets daily. The dataset includes a cleaned version without retweets, covers multiple languages with higher prevalence of English, Spanish, and French, and provides pre-processed n-gram frequencies for NLP tasks. It offers longitudinal coverage from at least January 27th, with specific additions like Russian language tweets collected between January 1st and May 8th.
License is listed as Open Access (green). The dataset is released in multiple versions; users should note version-specific differences in size and content.