INNUENDO: Whole Genome MLST Schema and Dataset for Campylobacter Jejuni
by Mirko Rossi / University of Helsinki
Available on 1 platform
Sign in to view source links and access this dataset
Description
6,526 assembled genomes of the foodborne pathogen Campylobacter jejuni form the basis of a curated whole-genome multilocus sequence typing (wgMLST) schema. The dataset includes metadata for each strain, such as country and year of isolation, source, host taxa, and classical MLST sequence type and clonal complex classifications. A final schema of 2,795 loci was created using the chewBBACA pipeline from an initial pangenome of 5,447 loci, with quality control steps applied to remove loci with single alleles, high length variability, or problematic annotations.
Use Cases
High-resolution source attribution and outbreak investigation based on wgMLST profiles.
Studying the population structure and evolution of Campylobacter jejuni using metadata like country, year, and host source.
Validating and benchmarking new wgMLST schema creation and allele calling pipelines against a curated reference.
Comparative genomic analysis of core and accessory genome content across thousands of bacterial isolates.
Strengths
Contains a large, curated collection of 6,526 assembled and quality-controlled Campylobacter jejuni genomes.
Provides a validated, publicly available wgMLST schema of 2,795 loci, created through a documented pipeline (INNUca, Roary, chewBBACA).
Includes detailed per-strain metadata enabling epidemiological and ecological analyses.
Limitations
The primary data retrieval date was April 2017, so the genomic collection is not recent.
Specific row counts for metadata files and exact column names are not provided in the available descriptions.
The dataset's cross-platform presence is limited to a single source type (paperswithcode), suggesting potential for broader dissemination.
Provenance
Source
European Nucleotide Archive (ENA), NCBI Sequence Read Archive (SRA), and the INNUENDO Sequence Dataset (PRJEB27020).
Collection Method
Raw reads were retrieved from public archives, assembled using the INNUca v3.1 pipeline, and annotated to create a pangenome and subsequent wgMLST schema.
Geography
Metadata includes country of isolation, suggesting global coverage, but specific countries are not listed.
License is listed as 'Open Access (green)'; users should verify specific terms for redistribution and commercial use.