遇见数据集

Data from: Accounting for genotype uncertainty in the estimation of allele frequencies in autopolyploids

收藏
DataONE2015-11-20 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Despite the increasing opportunity to collect large-scale data sets for population genomic analyses, the use of high-throughput sequencing to study populations of polyploids has seen little application. This is due in large part to problems associated with determining allele copy number in the genotypes of polyploid individuals (allelic dosage uncertainty–ADU), which complicates the calculation of important quantities such as allele frequencies. Here, we describe a statistical model to estimate biallelic SNP frequencies in a population of autopolyploids using high-throughput sequencing data in the form of read counts. We bridge the gap from data collection (using restriction enzyme based techniques [e.g. GBS, RADseq]) to allele frequency estimation in a unified inferential framework using a hierarchical Bayesian model to sum over genotype uncertainty. Simulated data sets were generated under various conditions for tetraploid, hexaploid and octoploid populations to evaluate the model's performance and to help guide the collection of empirical data. We also provide an implementation of our model in the R package polyfreqs and demonstrate its use with two example analyses that investigate (i) levels of expected and observed heterozygosity and (ii) model adequacy. Our simulations show that the number of individuals sampled from a population has a greater impact on estimation error than sequencing coverage. The example analyses also show that our model and software can be used to make inferences beyond the estimation of allele frequencies for autopolyploids by providing assessments of model adequacy and estimates of heterozygosity.

尽管当前群体基因组分析的大规模数据集采集机遇日益增多,但利用高通量测序(high-throughput sequencing)技术研究多倍体群体的应用却寥寥无几。这在很大程度上源于多倍体个体基因型中等位基因拷贝数测定的难题(等位基因剂量不确定性(allelic dosage uncertainty–ADU)),该问题会使等位基因频率等关键群体遗传学参数的计算变得复杂。本文提出一种统计模型,可利用以读段计数(read counts)形式呈现的高通量测序数据,估算同源多倍体(autopolyploids)群体中的双等位基因单核苷酸多态性(single nucleotide polymorphism, SNP)频率。本研究借助分层贝叶斯模型(hierarchical Bayesian model)构建统一的推断框架,整合基因型不确定性的求和过程,打通了从数据采集(采用基于限制性内切酶的技术,例如GBS、RADseq)到等位基因频率估算的完整链路。本研究针对四倍体、六倍体和八倍体群体,在多种条件下生成模拟数据集,以评估模型性能并为实验数据的采集提供指导。我们还在R语言扩展包polyfreqs中实现了该模型,并通过两个实例分析展示其用法:一是探究预期杂合度与观测杂合度水平,二是评估模型适配性。模拟结果表明,相较于测序覆盖深度,群体采样个体数量对估算误差的影响更为显著。实例分析同样证实,本研究提出的模型与软件不仅可用于同源多倍体的等位基因频率估算,还可通过模型适配性评估与杂合度估算,实现更多维度的群体遗传学推断。

创建时间:
2015-11-20
二维码
社区交流群
二维码
科研交流群
商业服务