Softcite Dataset: 4,971 Scientific Articles with Software Mention Annotations
by James Howison / The University of Texas at Austin
Available on 1 platform
Sign in to view source links and access this dataset
Description
Softcite Dataset Version 2.0 is a gold-standard corpus of 4,971 English-language scientific articles, half in Life Sciences and half in Economics, containing around 46 million tokens. The dataset, created by James Howison of The University of Texas at Austin, includes annotations for software mentions, versions, publishers, URLs, and programming languages. This version 2.0, released in 2023, adds new annotations to the original 2020 corpus through a multi-stage annotation and reconciliation process.
Use Cases
Train named entity recognition models to identify software names based on annotated mentions in full-text articles.
Develop relation extraction models to link software mentions with their versions or publishers based on the annotated relationships.
Benchmark text mining tools for software citation analysis using the provided holdout evaluation set.
Study the prevalence and context of software usage across Life Sciences and Economics disciplines.
Strengths
Corpus contains 4,971 full-text scientific articles, providing substantial textual data.
Annotations are gold-standard, resulting from a multi-stage process with annotator reconciliation.
Includes a dedicated holdout set of 20% of the full texts for evaluation purposes.
Version 2.0 adds over 1,000 new software name annotations compared to version 1.0.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for individual annotations is unknown, which may limit suitability assessment.
Data may reflect a bias towards Open Access articles in Life Sciences and Economics.
Provenance
Source
The University of Texas at Austin
Collection Method
Manual and automatic screening of scientific articles followed by a double annotation process with reconciliation.
Freshness
Version 2.0 was released in 2023, building upon the 2020 version 1.0 corpus.
Data is available under a CC-BY license. The primary format is XML (TEI), with JSON conversions provided; the JSON format uses character offsets which may be less readable.