遇见数据集

The Power of Gene-Based Rare Variant Methods to Detect Disease-Associated Variation and Test Hypotheses About Complex Disease

收藏
Figshare2016-01-15 更新2026-04-29 收录
官方服务:

资源简介:

Genome and exome sequencing in large cohorts enables characterization of the role of rare variation in complex diseases. Success in this endeavor, however, requires investigators to test a diverse array of genetic hypotheses which differ in the number, frequency and effect sizes of underlying causal variants. In this study, we evaluated the power of gene-based association methods to interrogate such hypotheses, and examined the implications for study design. We developed a flexible simulation approach, using 1000 Genomes data, to (a) generate sequence variation at human genes in up to 10K case-control samples, and (b) quantify the statistical power of a panel of widely used gene-based association tests under a variety of allelic architectures, locus effect sizes, and significance thresholds. For loci explaining ~1% of phenotypic variance underlying a common dichotomous trait, we find that all methods have low absolute power to achieve exome-wide significance (~5-20% power at α=2.5×10-6) in 3K individuals; even in 10K samples, power is modest (~60%). The combined application of multiple methods increases sensitivity, but does so at the expense of a higher false positive rate. MiST, SKAT-O, and KBAC have the highest individual mean power across simulated datasets, but we observe wide architecture-dependent variability in the individual loci detected by each test, suggesting that inferences about disease architecture from analysis of sequencing studies can differ depending on which methods are used. Our results imply that tens of thousands of individuals, extensive functional annotation, or highly targeted hypothesis testing will be required to confidently detect or exclude rare variant signals at complex disease loci.

大型队列的全基因组与全外显子组测序(Genome and exome sequencing)可用于解析罕见变异在复杂疾病中的作用。然而,要顺利推进这项研究,研究人员需检验一系列多样化的遗传学假设——这些假设在潜在致病变异的数量、发生频率与效应量上各不相同。本研究评估了基于基因的关联分析方法(gene-based association methods)检验这类遗传学假设的统计功效,并探讨了其对研究设计的指导意义。本研究依托千人基因组计划(1000 Genomes)数据开发了一套灵活的模拟方案,可完成两项任务:(a) 在至多10000例病例对照样本中生成人类基因的序列变异;(b) 量化多种等位基因结构、位点效应量与显著性阈值下,一组常用基于基因的关联检验的统计功效。针对可解释常见二分性状表型变异约1%的位点,本研究发现:在3000例样本中,所有方法达到全外显子组显著性水平(α=2.5×10^-6时统计功效约为5%~20%)的绝对统计功效均较低;即便在10000例样本中,功效也仅处于中等水平(约60%)。联合应用多种检验方法可提升检测灵敏度,但会以假阳性率升高为代价。MiST、SKAT-O与KBAC在模拟数据集上的个体平均统计功效最高,但各检验方法所检出的单个位点存在显著的结构依赖性差异,这表明基于测序研究分析得出的疾病遗传架构推论,会因所用方法的不同而存在差异。本研究结果表明,若要在复杂疾病位点上可靠检出或排除罕见变异信号,需依托数万例样本、大规模功能注释,或是开展高度针对性的假设检验。

创建时间:
2016-01-15
二维码
社区交流群
二维码
科研交流群
商业服务