INNUENDO: Whole and Core Genome MLST Schemas for 331 Yersinia Enterocolitica Strains
by Mirko Rossi / University of Helsinki
Available on 1 platform
Sign in to view source links and access this dataset
Description
331 assembled Yersinia enterocolitica genomes, including 79 novel strains, form the basis for this genomic dataset. The collection provides both a whole-genome MLST schema with 6,344 loci and a core-genome MLST schema with 2,406 loci, curated using the chewBBACA suite. Metadata for each strain includes country and year of isolation, source, host, serotype, biotype, pathotype, and sequence type.
Use Cases
Conducting population structure analysis based on whole-genome or core-genome multi-locus sequence typing (MLST) allele profiles.
Tracing the geographic and temporal spread of strains using metadata fields like country and year of isolation.
Investigating associations between genomic profiles and phenotypic characteristics like serotype, biotype, and pathotype.
Validating and applying the provided chewBBACA-compatible MLST schemas to classify new Yersinia enterocolitica genomes.
Strengths
Includes 331 assembled bacterial genomes, with 79 being novel additions not previously public.
Provides two curated MLST schemas: a wgMLST schema with 6,344 loci and a cgMLST schema with 2,406 loci present in ≥99% of genomes.
Metadata is detailed, containing fields like isolation source, host taxon, and multiple typing classifications for each strain.
Limitations
The data snapshot is from August 2018, which may limit its relevance for tracking recent strains or outbreaks.
Row count for the primary genomic data is unknown, though the description specifies 331 genomes.
Column-level documentation for the allele profile files is absent; field semantics must be inferred from the description or software documentation.
Provenance
Source
European Nucleotide Archive (ENA), NCBI Sequence Read Archive (SRA), and the INNUENDO Sequence Dataset (PRJEB27020).
Collection Method
Genomes were retrieved using getSeqENA, assembled with INNUca v3.1, and MLST schemas were created and validated using the chewBBACA suite.
Time Range
Data retrieved as of August 2018; specific years of isolation for individual strains are included in the metadata.
Geography
Metadata includes country of isolation for each strain, suggesting global coverage.
The schemas and allele profiles are formatted for use with the chewBBACA software suite, which is a required tool for full utilization. Proper citation of the chewBBACA paper is requested when using the schemas.