遇见数据集

Treehouse compendium of polyA selected RNA-Seq gene expression data from 33 PDX

收藏
NIAID Data Ecosystem2026-05-02 收录
官方服务:

资源简介:

We uniformly analyze sequence data to generate a resource for comparative gene expression studies. Specifically, we obtained access to primary RNA sequence data from repositories and clinical partners, consistently processed the data, harmonized metadata, and released the expression values and metadata without access restrictions The data contains 33 consistently processed gene expression datasets from 3 studies. TPM values are log transformed (specifically, log2(TPM+1)) and are associated with Hugo identifiers. Expected counts are not transformed and are associated with Ensemble gene identifiers. Transcript enrichment was performed via polyA selection. Gene expression in each sample is uniformly quantified using the dockerized TOIL RNA-Seq pipeline versions from 3.2 to 3.4.1 (Vivian et al., 2017); all of these versions produce bitwise identical RSEM gene expression outputs. The pipeline uses RSEM Version 1.2.25 (Li and Dewey, 2011) for quantification after aligning reads with STAR v 2.4.2a (Dobin et al., 2013) using indices generated from the human reference genome GRCh38 and the human gene models GENCODE 23 as described at https://github.com/UCSC-Treehouse/pipelines. Quality is assessed with the MEND pipeline https://github.com/UCSC-Treehouse/mend_qc (Beale et al., 2021).

创建时间:
2025-07-10
二维码
社区交流群
二维码
科研交流群
商业服务