A 2021 commentary by Aziz Khan from Stanford University presents data showing an increase in non-inclusive terms with racial connotations used in life-science literature. The associated source code and data support the analysis and call for action to make science inclusive for all. The dataset is published under an Open Access license.
Use Cases
- Analyze trends in non-inclusive language based on the described corpus of life-science literature.
- Train or benchmark NLP models for detecting biased terminology based on the identified terms.
- Support meta-research on diversity and inclusion in scientific publishing based on the textual analysis.
Strengths
- Data is associated with a peer-reviewed commentary published in eLife.
- Source code is provided, which may allow for reproducibility of the analysis.
- The dataset has a clear Open Access license.
Limitations
- The specific data format, size, and column structure are unknown.
- Row count is unknown, which may limit suitability assessment.
- The last update date is unknown; freshness unverified.
Provenance
- Source
- Aziz Khan, Stanford University
- Collection Method
- Likely extracted and analyzed from life-science literature corpus.
- Time Range
- Analysis pertains to 2021, but the underlying literature corpus time range is unspecified.