Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
4,971 full-text scientific articles in English, half in Life Sciences and half in Economics, containing annotations for software mentions. The dataset is a gold-standard corpus created through multi-stage annotation and reconciliation by a team, resulting in over 5,000 annotated software names. It was developed by James Howison at The University of Texas at Austin, with versions released in 2020 and 2023.
Data is available under a CC-BY license. The primary format is XML, with JSON conversions provided; the JSON format uses offsets for annotation spans, which may be less readable.