Sign in to view source links and access this dataset
Description
A corpus of legislative proceedings from California, Texas, New York, and one other large U.S. state between 2015 and 2018. The dataset was created by Khosmood, Dekhtyar, Ellwein, and White and presented at Digital Humanities in Washington, DC in August 2024. It is distributed under the CC BY-NC-SA 4.0 license.
Use Cases
Analyzing political discourse and rhetoric based on the text of legislative proceedings.
Training NLP models for topic modeling or sentiment analysis on government text data.
Conducting comparative studies of state-level legislative processes across different U.S. states.
Tracking the evolution of policy discussions over the 2015-2018 time period.
Strengths
Covers proceedings from four of the largest U.S. states, providing a significant geographic scope.
Spans a multi-year period from 2015 to 2018, allowing for temporal analysis.
Has a clear attribution and citation from a 2024 academic presentation.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment for large-scale ML tasks.
The description metadata is limited; actual data quality and completeness require manual inspection.
Provenance
Source
State legislatures of California, Texas, New York, and one other unspecified state.
Collection Method
Likely compiled and processed from official legislative records.
Time Range
2015-2018
Freshness
Last updated 2024-08-07 21:50:10
Geography
United States (specifically California, Texas, New York, and one other large state)
License is CC BY-NC-SA 4.0, which restricts commercial use and requires share-alike distribution of derivatives.