Innuendo: Whole and Core Genome MLST Schemas for 2,337 Escherichia coli Genomes
by Mirko Rossi / University of Helsinki
Available on 1 platform
Sign in to view source links and access this dataset
Description
Mirko Rossi from the University of Helsinki provides a reference dataset of 2,337 Escherichia coli genome assemblies and associated metadata. The dataset includes 2,218 public genomes from EnteroBase and 119 Shiga toxin-producing E. coli genomes from the INNUENDO project, downloaded in April 2017. It features curated whole-genome and core-genome multi-locus sequence typing (MLST) schemas containing 7,601 and 2,360 loci, respectively.
Use Cases
Benchmarking whole-genome MLST (wgMLST) allele calling pipelines based on the provided 7,601-locus schema.
Comparative genomic analysis of E. coli strains based on metadata including source, host, country, year, serotype, and pathotype.
Developing core-genome MLST (cgMLST) analysis workflows using the defined schema of 2,360 loci present in at least 99% of genomes.
Studying the population structure of Shiga toxin-producing E. coli (STEC) using the included 119 INNUca-assembled genomes.
Strengths
Includes 2,337 E. coli genome assemblies, providing a substantial reference set.
Offers two curated MLST schemas: a wgMLST schema with 7,601 loci and a cgMLST schema with 2,360 loci.
Metadata is detailed, containing source, host, country, year, serotype, pathotype, and sequence type classifications.
Limitations
Data snapshot is from April 2017; genomic data may not reflect current pathogen diversity.
Column-level documentation for the allele profile files is absent; field semantics must be inferred after download.
The dataset's geographic and temporal representativeness is constrained by the source data available in EnteroBase in 2017.
Provenance
Source
EnteroBase and the INNUENDO Sequence Dataset (PRJEB27020).
Collection Method
Genome assemblies were downloaded from EnteroBase, selected based on ribosomal ST classification, and supplemented with INNUca-assembled genomes. Schemas were curated using the chewBBACA software suite.
Time Range
Isolation years are included in metadata, but the primary dataset snapshot is from April 2017.
Geography
Country of isolation is included in the metadata for each strain.
Genome assemblies from EnteroBase are not included in the package and must be downloaded separately using the provided barcodes.