Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Experimental results from a study comparing the performance of Big Data clusters for variant calling data retrieval. The data includes query execution times for different storage formats (VCF and Apache Parquet), input sizes (104 and 1144 individuals), cluster sizes (2 to 150 executor nodes), and HDFS replication factors (3, 5, 7, 9). The dataset was produced by Katerina Boufea at Wageningen University & Research and accompanies the paper 'Managing Variant Calling Datasets The Big Data Way'.
License is listed as Open Access (green), but specific terms should be verified. The dataset likely requires familiarity with genomics data formats (VCF/Parquet) and Big Data frameworks (Hadoop).