Research Code Quality and Execution Analysis for 2000+ Datasets, 2010-2020
by A. Trisovic / Harvard University Press
Available on 1 platform
Sign in to view source links and access this dataset
Description
More than 2000 replication datasets with over 9000 unique R files from the Harvard Dataverse repository, published between 2010 and 2020. The dataset accompanies a study by Ana Trisovic, Matthew K. Lau, Thomas Pasquier, and Mercè Crosas that analyzes code quality and execution success. The study found 74% of R files failed to execute without error initially, with failures reduced to 56% after automatic code cleaning.
Use Cases
Analyzing common coding errors in research code based on the execution failure rates mentioned in the description
Studying the impact of journal policy strictness on code re-execution rates as referenced in the abstract
Developing automated code cleaning tools based on the error patterns identified in the large-scale study
Formulating recommendations for code dissemination aimed at researchers, journals, and repositories as described in the article
Strengths
Analysis of over 2000 replication datasets, providing a large-scale view of research code practices
Examines more than 9000 unique R files, offering substantial sample size for code quality assessment
Covers a 10-year time range (2010-2020), allowing for longitudinal analysis of trends
Limitations
Column-level documentation is absent; field semantics must be inferred after download
Row count is unknown, which may limit suitability assessment
Last update date is unknown; freshness unverified
Provenance
Source
Harvard Dataverse repository
Collection Method
Retrieved and analyzed replication datasets published with academic papers
Time Range
2010 to 2020
License is listed as Open Access (green); specific terms should be verified upon download.