Foltz, Taroni, and Greene's data accompanies their manuscript on combining microarray and RNA-seq data for machine learning. The dataset contains all data necessary to recreate the main and supplementary figures from the published study. It was used to evaluate normalization methods like quantile and Training Distribution Matching for supervised and unsupervised model training across platforms.
Use Cases
- Evaluating normalization methods for combining microarray and RNA-seq data based on the study's methodology.
- Training supervised machine learning models on cross-platform normalized gene expression data.
- Training unsupervised machine learning models on cross-platform normalized gene expression data.
- Performing pathway analysis using Pathway-Level Information Extractor (PLIER) on normalized data.
- Recreating figures from the manuscript to validate cross-platform normalization techniques.
Strengths
- Data is directly linked to a peer-reviewed manuscript published in a scientific journal.
- The dataset is described as containing all data necessary to recreate the study's figures, suggesting completeness for its intended purpose.
- The description provides a clear abstract outlining the research goals and findings.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and file size are unknown, which may limit suitability assessment.
- Last update date is unknown; freshness unverified.
Provenance
- Source
- Steven M. Foltz, associated with the manuscript 'Cross-platform normalization enables machine learning model training on microarray and RNA-seq data simultaneously'.
- Collection Method
- Likely generated from RNA-seq titration experiments as part of the referenced research study.