Data from: Information dropout patterns in RAD phylogenomics and a comparison with multilocus sanger data in a species-rich moth genus
收藏资源简介:
A rapid shift from traditional Sanger sequencing-based molecular methods to the phylogenomic approach with large numbers of loci is underway. Among phylogenomic methods, RAD (Restriction site Associated DNA) sequencing approaches have gained much attention as they enable rapid generation of up to thousands of loci randomly scattered across the genome and are suitable for non-model species. RAD data sets however suffer from large amounts of missing data and rapid locus dropout along with decreasing relatedness among taxa. The relationship between locus dropout and the amount of phylogenetic information retained in the data has remained largely un-investigated. Similarly, phylogenetic hypotheses based on RAD have rarely been compared with phylogenetic hypotheses based on multilocus Sanger sequencing, even less so using exactly the same species and specimens. We compared the Sanger-based phylogenetic hypothesis (8 loci; 6,172 bp) of 32 species of the diverse moth genus Eupithecia (Lepidoptera, Geometridae) to that based on double-digest RAD sequencing (3,256 loci; 726,658 bp). We observed that topologies were largely congruent, with some notable exceptions that we discuss. The locus dropout effect was strong. We demonstrate that number of loci is not a precise measure of phylogenetic information since the number of single-nucleotide polymorphisms (SNPs) may remain low at very shallow phylogenetic levels despite large numbers of loci. As we hypothesize, the number of SNPs and parsimony informative SNPs (PIS) is low at shallow phylogenetic levels, peaks at intermediate levels and, thereafter, declines again at the deepest levels as a result of decay of available loci. Similarly, we demonstrate with empirical data that the locus dropout affects the type of loci retained, the loci found in many species tending to show lower interspecific distances than those shared among fewer species. We also examine the effects of the numbers of loci, SNPs and PIS on nodal bootstrap support, but could not demonstrate with our data our expectation of a positive correlation between them. We conclude that RAD methods provide a powerful tool for phylogenomics at an intermediate phylogenetic level as indicated by its broad congruence with an eight-gene Sanger data set in a genus of moths. When assessing the quality of the data for phylogenetic inference, the focus should be on the distribution and number of SNPs and PIS rather than on loci.
当前,基于传统桑格(Sanger)测序的分子方法正快速向拥有大量基因座的系统发育基因组学方法转变。其中,限制性酶切位点相关DNA(Restriction site Associated DNA, RAD)测序方法已受到广泛关注,因其可快速生成多达数千个随机散布于基因组的基因座,且适用于非模式物种(non-model species)。然而,RAD数据集存在大量缺失数据以及随类群间亲缘关系降低而快速出现的基因座丢失现象。基因座丢失与数据中保留的系统发育信息之间的关联,在很大程度上仍未得到研究。类似地,基于RAD的系统发育假说鲜少与多位点桑格测序得到的系统发育假说进行比较,使用完全相同的物种和标本进行的对比更是少之又少。我们对32种多样化的尺蛾属*Eupithecia*(鳞翅目,尺蛾科)的桑格测序系统发育假说(8个基因座;6172 bp)与双酶切RAD测序(3256个基因座;726658 bp)得到的假说进行了比较。我们观察到二者的拓扑结构整体一致,但也存在一些值得讨论的显著差异。基因座丢失效应显著。我们的研究表明,基因座数量并非系统发育信息的精准衡量指标,因为即使存在大量基因座,在极浅的系统发育层级上,单核苷酸多态性(single-nucleotide polymorphisms, SNPs)的数量可能仍然较低。正如我们所提出的假说:在浅系统发育层级上,单核苷酸多态性(SNPs)和简约信息位点(parsimony informative SNPs, PIS)的数量较低,在中间层级达到峰值,随后由于可用基因座的衰减,在最深层级再次下降。同样,我们通过实证数据证明,基因座丢失会影响保留的基因座类型:在多数物种中存在的基因座,其种间距离往往低于仅在少数物种中共享的基因座。我们还考察了基因座数量、SNPs以及PIS对节点自展支持率的影响,但基于我们的数据,未能证实二者间存在正相关的预期。我们得出结论:正如该尺蛾属中8个基因桑格数据集与RAD方法的广泛一致性所表明的那样,RAD方法为中间系统发育层级的系统发育基因组学研究提供了强有力的工具。在评估用于系统发育推断的数据质量时,研究重点应放在SNPs和PIS的分布与数量上,而非基因座本身。



