Funding period: 2020-2025
Leads: Burton Blais and Catherine Carrillo
Total GRDI funding: $110,000
To help keep Canada's food supply safe, the Canadian Food Inspection Agency (CFIA) uses advanced laboratory methods to find and characterize harmful bacteria in food, specifically targeting Salmonella, Listeria monocytogenes, and Shiga toxin-producing Escherichia coli (STEC). The effectiveness of these tests depends on their ability to identify a wide variety of bacterial strains, including those that are rare, unique, or otherwise difficult to grow or detect. To solve this challenge, this project builds a standardized collection of bacterial "benchmark panels" (standard groups of microbes used to check test accuracy).
The project team gathered a diverse repository of 663 physical bacterial samples, including 161 Salmonella strains, 192 STEC strains, 216 Listeria monocytogenes strains, and 94 "exclusivity" strains. These exclusivity strains consist of non-harmful bacteria used to ensure tests do not give false alarms. By sharing these physical strains alongside their complete genome profiles and standard handling instructions, the project significantly improves test development efficiency. This approach makes test results much easier to compare across different methods, helping scientists identify the absolute best tools for food safety testing. To stay effective, these panels will be regularly updated with new, unusual strains identified by CFIA laboratories.
Digital Testing Capabilities
Beyond physical samples, the project has developed pathogen-specific genome databases designed to reflect the global diversity of the target bacteria. These allow scientists to run rapid, computer-based (in silico) validation tests to see if a new method can successfully detect strains from around the world. This digital collection currently includes 6,730 Salmonella genomes, 9,608 STEC genomes, and 2,716 Listeria genomes.
Operational Impact
The physical Salmonella panel has already been shipped to CFIA testing labs across Canada. Both the physical and genomic benchmark panels were recently deployed to prove the accuracy of a new DNA-based rapid test (qPCR) used to confirm the presence of Salmonella in foods, successfully verifying its reliability before it is used for routine regulatory testing.
Research tool/process
- Smart Genome Selection and Maintenance Workflow: The workflow for selection of diverse genomes from large public datasets involves: (1) recovery of thousands of genomes from public repositories, (2) removing low-quality datasets, (3) removing clonal or closely related datasets. These large-scale datasets support the implementation of rigorous in silico validation protocols ensuring improved performance of gene-based methods for pathogen detection.
- Regular database updates: To ensure these tools remain accurate over time, this selection process is conducted routinely on a bi-annual basis (twice a year). During these updates, novel genomes are added and existing genomes in the core benchmark dataset are evaluated and replaced with higher-quality data as more complete, fully mapped bacterial genomes become available. This ongoing maintenance ensures that DNA-based pathogen detection methods perform reliably against changing global strains.
- Benchmark Strain Panel Development and handling: To ensure that bacterial test results are highly accurate and reproducible across different laboratories, the project follows strict quality control procedures for growing, storing, and tracking the physical bacterial strains. Whole genome sequencing is performed on the strains at the exact same time that the long-term storage (glycerol stocks) is prepared, guaranteeing that the genome perfectly matches the physical sample. Additionally, laboratory teams strictly monitor and record the passage number—how many times the bacteria have been grown or transferred—to document usage and prevent the strains from changing over time. The project also provides standardized instructions for labs to continuously log detailed information about each strain, allowing the network of laboratories to steadily build a richer, more comprehensive profile of the bacteria over time.
Dataset/database
- NCBI Pathogen Genomes Core Reference Datasets: To support large-scale computer testing, the project selected streamlined, highly diverse datasets from massive global public repositories. These standardized groups are built into CFIA digital tools—such as the PrimerValidator on the FoodPort platform—to check the reliability of species-specific test targets and ensure that results remain completely comparable across different laboratories.
- STEC NCBI Database: A set of 10,790 diverse genomes selected from approximately 100,000 published genomes in the NCBI pathogens database.
- Listeria monocytogenes NCBI Database: A set of 3,225 diverse genomes selected from approximately 60,000 published genomes in the NCBI pathogens database.
- Salmonella NCBI Database: A set of 6,730 diverse genomes selected from 500,000 published records in the NCBI pathogens database to ensure broad representation of global strains.
- Complete genomes for Benchmark Strain Panels: This dataset contains the complete, fully mapped genomes for the project's repository of 663 physical bacterial samples. This collection includes 161 Salmonella strains, 192 STEC (E. coli) strains, 216 Listeria monocytogenes strains, and 94 "exclusivity" strains (non-harmful bacteria used to check for false positives). To achieve the highest possible accuracy, these genomes were built using a combination of short-read (Illumina) sequencing for fine detail and long-read (Nanopore) sequencing to seamlessly piece together the full genetic structure.
Contact us
For additional information, please contact:
Genomics R&D Initiative
Email: info@grdi-irdg.collaboration.gc.ca