遇见数据集

Data from: Genome reannotation of the lizard Anolis carolinensis based on 14 adult and embryonic deep transcriptomes

收藏
DataCite Commons2025-06-01 更新2025-06-15 收录
官方服务:

资源简介:

Background: The green anole lizard, Anolis carolinensis, is a key species for both laboratory and field-based studies of evolutionary genetics, development, neurobiology, physiology, behavior, and ecology. As the first non-avian reptilian genome sequenced, A. carolinesis is also a prime reptilian model for comparison with other vertebrate genomes. The public databases of Ensembl and NCBI have provided a first generation gene annotation of the anole genome that relies primarily on sequence conservation with related species. A second generation annotation based on tissue-specific transcriptomes would provide a valuable resource for molecular studies. Results: Here we provide an annotation of the A. carolinensis genome based on de novo assembly of deep transcriptomes of 14 adult and embryonic tissues. This revised annotation describes 59,373 transcripts, compared to 16,533 and 18,939 currently for Ensembl and NCBI, and 22,962 predicted protein-coding genes. A key improvement in this revised annotation is coverage of untranslated region (UTR) sequences, with 79% and 59% of transcripts containing 5' and 3' UTRs, respectively. Gaps in genome sequence from the current A. carolinensis build (Anocar2.0) are highlighted by our identification of 16,542 unmapped transcripts, representing 6,695 orthologues, with less than 70% genomic coverage. Conclusions: Incorporation of tissue-specific transcriptome sequence into the A. carolinensis genome annotation has markedly improved its utility for comparative and functional studies. Increased UTR coverage allows for more accurate predicted protein sequence and regulatory analysis. This revised annotation also provides an atlas of gene expression specific to adult and embryonic tissues.

背景:绿安乐蜥(Anolis carolinensis)是进化遗传学、发育学、神经生物学、生理学、行为学与生态学领域开展实验室与野外研究的关键物种。作为首个被测序的非鸟类爬行类基因组,该物种亦是用于与其他脊椎动物基因组进行比较研究的核心爬行类模型。Ensembl与NCBI两大公共数据库已基于与近缘物种的序列保守性,完成了安乐蜥基因组的第一代基因注释。而基于组织特异性转录组构建的第二代注释,将为分子生物学研究提供极具价值的研究资源。结果:本研究基于14份成体与胚胎组织的深度转录组从头组装结果,完成了Anolis carolinensis基因组的注释。本次修订后的注释共涵盖59373条转录本,而当前Ensembl与NCBI的注释分别仅包含16533条与18939条转录本,同时预测得到22962个蛋白质编码基因。本次修订注释的一项核心改进在于非翻译区(untranslated region, UTR)序列的覆盖度提升:分别有79%与59%的转录本包含5'端与3'端非翻译区。通过本研究的分析发现,现有Anolis carolinensis基因组组装版本(Anocar2.0)存在基因组序列缺口:共鉴定得到16542条未定位转录本,对应6695个直系同源基因,其基因组覆盖度不足70%。结论:将组织特异性转录组序列整合至Anolis carolinensis基因组注释中,显著提升了该注释在比较与功能研究中的应用价值。非翻译区覆盖度的提升可支持更精准的蛋白质序列预测与调控分析。本次修订后的注释同时提供了成体与胚胎组织特异性的基因表达图谱。

提供机构:
Dryad
创建时间:
2013-03-19
搜集汇总
数据集介绍
Data from: Genome reannotation of the lizard Anolis carolinensis based on 14 adult and embryonic deep transcriptomes 数据集图片
背景与挑战
背景概述
该数据集提供了基于14个成年和胚胎组织深度转录组数据的绿安乐蜥(Anolis carolinensis)基因组重新注释结果,显著提高了基因注释的完整性和准确性。重新注释后包含59,373个转录本,比现有公共数据库的注释更全面,特别改善了非翻译区(UTR)的覆盖,并识别了大量未映射的转录本。这些数据为比较和功能基因组学研究提供了重要资源,包括组织特异性基因表达图谱。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务