Detection of Mixed Infection from Bacterial Whole Genome Sequence Data Allows Assessment of Its Role in <i>Clostridium difficile</i> Transmission
收藏资源简介:
Bacterial whole genome sequencing offers the prospect of rapid and high precision investigation of infectious disease outbreaks. Close genetic relationships between microorganisms isolated from different infected cases suggest transmission is a strong possibility, whereas transmission between cases with genetically distinct bacterial isolates can be excluded. However, undetected mixed infections—infection with ≥2 unrelated strains of the same species where only one is sequenced—potentially impairs exclusion of transmission with certainty, and may therefore limit the utility of this technique. We investigated the problem by developing a computationally efficient method for detecting mixed infection without the need for resource-intensive independent sequencing of multiple bacterial colonies. Given the relatively low density of single nucleotide polymorphisms within bacterial sequence data, direct reconstruction of mixed infection haplotypes from current short-read sequence data is not consistently possible. We therefore use a two-step maximum likelihood-based approach, assuming each sample contains up to two infecting strains. We jointly estimate the proportion of the infection arising from the dominant and minor strains, and the sequence divergence between these strains. In cases where mixed infection is confirmed, the dominant and minor haplotypes are then matched to a database of previously sequenced local isolates. We demonstrate the performance of our algorithm with in silico and in vitro mixed infection experiments, and apply it to transmission of an important healthcare-associated pathogen, Clostridium difficile. Using hospital ward movement data in a previously described stochastic transmission model, 15 pairs of cases enriched for likely transmission events associated with mixed infection were selected. Our method identified four previously undetected mixed infections, and a previously undetected transmission event, but no direct transmission between the pairs of cases under investigation. These results demonstrate that mixed infections can be detected without additional sequencing effort, and this will be important in assessing the extent of cryptic transmission in our hospitals.
细菌全基因组测序(bacterial whole genome sequencing)为传染病暴发的快速高精度调查提供了可能。从不同感染病例中分离得到的微生物若呈现紧密的遗传亲缘关系,则提示存在传播的高度可能性;反之,若感染病例间的细菌分离株遗传背景存在显著差异,则可排除二者间的传播关联。然而,未被检出的混合感染——即同一物种的至少2种非关联菌株同时感染,但仅对其中1株完成测序——可能会削弱传播排除结论的确定性,进而限制该技术的应用价值。本研究针对该问题,开发了一种计算高效的混合感染检测方法,无需对多个细菌菌落进行资源消耗高昂的独立测序。鉴于细菌序列数据中单核苷酸多态性(single nucleotide polymorphisms)的密度相对较低,基于当前的短读长测序数据(short-read sequence data)直接重建混合感染的单倍型(haplotypes)往往难以获得稳定可靠的结果。因此,本研究采用基于最大似然估计(maximum likelihood)的两步分析流程,假设每份样本最多携带2种致病菌株。我们将联合估算优势菌株与次要菌株在感染中的占比,以及二者之间的序列差异程度。在确认存在混合感染的样本中,我们会将鉴定得到的优势与次要单倍型与已完成测序的本地菌株分离株数据库进行比对匹配。本研究通过计算机模拟(in silico)与体外(in vitro)混合感染实验验证了算法的性能,并将其应用于一种重要的医院相关性病原菌——艰难梭菌(Clostridium difficile)的传播分析。我们基于此前已发表的随机传播模型,结合医院病房移动数据,筛选出15对疑似存在混合感染相关传播事件的病例。本方法成功检出4例此前未被发现的混合感染,以及1起此前未被识别的传播事件,但未在所研究的病例中发现直接传播关联。上述结果表明,无需额外开展测序工作即可检出混合感染,这对于评估医院内隐秘传播(cryptic transmission)的规模具有重要意义。



