遇见数据集

Data from: Telling the whole story in a 10,000-genome world

收藏
DataONE2012-02-10 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

BACKGROUND: Genome sequencing has revolutionized our view of the relationships among genomes, particularly in revealing the confounding effects of lateral genetic transfer (LGT). Phylogenomic techniques have been used to construct purported trees of microbial life. Although such trees are easily interpreted and allow the use of a subset of genomes as "proxies" for the full set, LGT and other phenomena impact the positioning of different groups in genome trees, confounding and potentially invalidating attempts to construct a phylogeny-based taxonomy of microorganisms. Network and graph approaches can reveal complex sets of relationships, but applying these techniques to large data sets is a significant challenge. Notwithstanding the question of what exactly it might represent, generating and interpreting a Tree or Network of All Genomes will only be feasible if current algorithms can be improved upon. RESULTS: Complex relationships among even the most-similar genomes demonstrate that proxy-based approaches to simplifying large sets of genomes are not alone sufficient to solve the analysis problem. A phylogenomic analysis of 1173 sequenced bacterial and archaeal genomes generated phylogenetic trees for 159,905 distinct homologous gene sets. The relationships inferred from this set can be heavily dependent on the inclusion of other taxa: for example, phyla such as Spirochaetes, Proteobacteria and Firmicutes are recovered as cohesive groups or split depending on the presence of other specific lineages. Furthermore, named groups such as Acidithiobacillus, Coprothermobacter and Brachyspira show a multitude of affiliations that are more consistent with their ecology than with small subunit ribosomal DNA-based taxonomy. Network and graph representations can illustrate the multitude of conflicting affinities, but all methods impose constraints on the input data and create challenges of construction and interpretation. CONCLUSIONS: These complex relationships highlight the need for an inclusive approach to genomic data, and current methods with minor alterations will likely scale to allow the analysis of data sets with 10,000 or more genomes. The main challenges lie in the visualization and interpretation of genomic relationships, and the redefinition of microbial taxonomy when subsets of genomic data are so evidently in conflict with one another, and with the "canonical" molecular taxonomy.

研究背景:基因组测序彻底革新了我们对基因组间亲缘关系的认知,尤其揭示了侧向基因转移(lateral genetic transfer, LGT)所带来的混淆效应。系统发育基因组学技术已被用于构建所谓的微生物生命之树。尽管这类树形结构易于解读,且可选取部分基因组作为全基因组集的“替代集”,但侧向基因转移及其他现象会干扰基因组树中不同类群的定位,破坏基于系统发育构建微生物分类学的尝试的有效性,甚至使其完全失效。网络与图论方法能够揭示复杂的关联关系集,但将这些技术应用于大型数据集仍是一项重大挑战。暂且不论其具体代表意义,只有改进现有算法,构建并解读“所有基因组之树或网络”才具备可行性。 研究结果:即便是相似度极高的基因组之间也存在复杂的关联,这表明仅依靠基于替代集的方法简化大型基因组数据集,并不足以解决分析难题。研究人员对1173个已测序的细菌和古菌基因组开展系统发育基因组学分析,为159905个不同的同源基因集构建了系统发育树。从该数据集推断出的亲缘关系在很大程度上取决于其他类群的纳入情况:例如,螺旋体门(Spirochaetes)、变形菌门(Proteobacteria)和厚壁菌门(Firmicutes)等类群,会因其他特定谱系的存在与否,被识别为凝聚类群或是发生拆分。此外,诸如嗜酸硫杆菌属(Acidithiobacillus)、嗜热粪杆菌属(Coprothermobacter)和短螺旋体属(Brachyspira)等已命名类群,展现出多种亲缘关联,这些关联与其生态习性的契合度远高于基于小亚基核糖体DNA的分类学结果。网络与图论表征能够直观呈现大量存在冲突的亲缘关系,但所有方法都会对输入数据施加约束,同时带来构建与解读层面的挑战。 研究结论:这些复杂的关联凸显了对基因组数据采取包容性分析方法的必要性。当前的方法经过小幅调整后,或可扩展至处理包含10000个乃至更多基因组的数据集。当前的主要挑战在于基因组亲缘关系的可视化与解读,以及当基因组数据的不同子集彼此间、与“经典”分子分类学明显相悖时,如何重新定义微生物分类学。

创建时间:
2012-02-10
二维码
社区交流群
二维码
科研交流群
商业服务