Data from: Next-generation polyploid phylogenetics: rapid resolution of hybrid polyploid complexes using PacBio single-molecule sequencing
收藏资源简介:
Difficulties in generating nuclear data for polyploids have impeded phylogenetic study of these groups. We describe a high-throughput protocol and an associated bioinformatics pipeline (PURC: “Pipeline for Untangling Reticulate Complexes”) that is able to generate these data quickly and conveniently, and demonstrate its efficacy on accessions from the fern family Cystopteridaceae. We conclude with a demonstration of the downstream utility of these data by inferring a multilabeled species tree for a subset of our accessions. We amplified four ~1kb-long nuclear loci and sequenced them in a parallel-tagged amplicon sequencing approach using the PacBio platform. PURC infers the final sequences from the raw reads via an iterative approach that corrects PCR and sequencing errors and removes PCR-mediated recombinant sequences (chimeras). We generated data for all gene copies (homeologs, paralogs, and segregating alleles) present in each of three sets of 50 mostly-polyploid accessions, for four loci, in three PacBio runs (one run per set). From the raw sequencing reads PURC was able to accurately infer the underlying sequences. This approach makes it easy and economical to study the phylogenetics of polyploids, and in conjunction with recent analytical advances, facilitates investigation of broad patterns of polyploid evolution.
多倍体物种的核数据生成难题长期掣肘该类群的系统发育研究。我们开发了一套高通量实验流程与配套的生物信息学分析流程PURC(Pipeline for Untangling Reticulate Complexes),可快速便捷地生成所需核数据,并以冷蕨科(Cystopteridaceae)的种质材料为对象验证了该流程的有效性。 最后,我们通过对部分种质材料构建多标记物种树,展示了该数据集的下游应用价值。 我们针对4个长度约1kb的核基因座进行扩增,并采用PacBio测序平台开展平行标签扩增子测序。PURC通过迭代算法从原始测序读段中推断最终序列,该算法可校正PCR扩增与测序误差,并去除PCR介导的重组序列(嵌合体)。 我们针对3组各50份以多倍体为主的种质材料、4个基因座,通过3次PacBio测序(每组对应1次测序),获取了每份材料中所有基因拷贝的数据,包括部分同源基因(homeologs)、旁系同源基因(paralogs)与分离等位基因(segregating alleles)。 PURC可从原始测序读段中准确推断出目标序列的真实本底序列。 该方法使得多倍体系统发育研究变得简便且经济,结合当前的分析方法进展,可助力解析多倍体演化的宏观格局。



