Breast tumour microenvironment structures are associated with genomic features and clinical outcome
收藏资源简介:
Data and code are provided in one directory. This annotation document is divided into notes for those who wish to reuse data, and those who wish to rerun analysis code. Data for reuse: The data comprise three types: Full stack tiff images that contain multiplexed imaging mass cytometry (IMC) images. Image masks that identify image regions associated with cells, epithelium and vessels Processed data that contain measurements taken using the associated images. All full stacks and masks are tiff images. Each IMC acquisition (image) is associated with six images in total: the full stack image itself and five image masks (whole cell, nucleus, cytoplasm, tumour and vessel). The naming convention for these images is MB####_###_ImageType.tiff, where: MB#### is the METABRIC identifier. This can be used to link the data to other METABRIC data in the public domain. ### is the ImageNumber. This links the image to columns in processed data files. It is a sequential integer between one and three digits long. Note that this number, assigned based on file order, is not the same across studies so cannot be used to link images from other METABRIC data sets e.g. Ali et al Nat Cancer 2020. Each image corresponds to a core from a tissue microarray (TMA) slide. ImageType. A descriptive label that identifies the type of image. Notes: The order of image layers in full stack images corresponds to the markerStackOrder.csv file, which identifies each image layer with its corresponding isotope and epitope. Masks are grayscale images where each discrete region is identified by a set of contiguous pixels associated with a single integer value. These tend to be sequential from the top to the bottom of the image (this is why a mask appears as a gradation of gray and white when opened in an image viewer). The processed single cell data ‘ObjectNumber’ column corresponds to whole cell masks, where the integer values of each cell maps to ‘ObjectNumber’, allowing for marker values and other features to be mapped to images. Two processed data files: SingleCells.csv where each row represents a cell, and columns are data associated with each cell. Each observation is uniquely identified by the combination of ImageNumber and ObjectNumber. These data have already been spillover corrected. CellNeighbours.csv where each row represents a cell-cell interaction. The data are in graph format, with columns labelled ‘from’ and ‘to’ meaning from an index cell to a neighbouring cell (despite this convention, the data are undirected); the integers within these columns map to ObjectNumber in SingleCells.csv. Note: The convention for `is_` variables in processed data files is that 0 is FALSE and 1 TRUE. Two column annotation files: Two corresponding annotation files that contain details on the content of each column in processed tabular files are also provided, they are SingleCellsAnnotation.xlsx and CellNeighboursAnnotation.xlsx Other files are annotation and processed data files required by the code in the Code directory; they can be ignored unless you plan to rerun analyses. Code and reproducibility Analysis code and corresponding processed data are also provided in the directory. The code was run within a conda environment, details of which are provided in the file CondaEnv.yml. Processed metadata from the METABRIC study are among the files provided. It is, however, recommended that additional analyses that rely on METABRIC metadata, use data downloaded from their original publications or a public repository as these data are subject to updates, and the user may wish to process them differently. Code is separated by figures. The code must be run in the order figures appear in the paper, as later code relies on derived files created earlier. The code must also be run within the directory as relative paths rely on its structure.
本数据集附带完整的代码与数据目录。本注释文档分为两部分,分别供希望复用本数据集的研究者,以及希望重新运行分析代码的用户参考。 可供复用的数据包含三类: 1. 全栈TIFF图像:包含多重成像质谱流式(imaging mass cytometry, IMC)图像。 2. 图像掩码:用于标识与细胞、上皮组织及血管相关的图像区域。 3. 预处理数据:包含基于关联图像获取的定量测量结果。 所有全栈图像与掩码均为TIFF格式图像。 每一次IMC采集(即单张图像)总计对应6张图像:全栈图像本身,以及5张图像掩码(全细胞、细胞核、细胞质、肿瘤组织与血管)。这些图像的命名遵循统一规范:`MB####_###_ImageType.tiff`,其中: - `MB####` 为METABRIC标识符,可用于将本数据集与公共领域中其他METABRIC数据集进行关联。 - `###` 为图像编号(ImageNumber),用于将图像与预处理数据文件中的列进行关联。该编号为1至3位的连续整数,其分配基于文件的排序顺序。需注意:不同研究中的该编号并不统一,因此无法用于关联其他METABRIC数据集(如Ali等人于2020年发表于《Nature Cancer》的数据集)中的图像。 每张图像对应组织微阵列(tissue microarray, TMA)切片中的一个芯样。 `ImageType`:用于标识图像类型的描述性标签。 注意事项: 1. 全栈图像中的图像层顺序与`markerStackOrder.csv`文件一致,该文件为每个图像层标注了对应的同位素与表位。 2. 图像掩码为灰度图像,每个离散区域通过一组与单个整数值关联的连续像素进行标识。这些整数值通常从图像顶部至底部依次递增,这也是为何在图像查看器中打开掩码时,会呈现灰度过渡至白色的视觉效果。 3. 预处理单细胞数据中的`ObjectNumber`(对象编号)列与全细胞掩码相对应:每个细胞的整数值映射至`ObjectNumber`,可用于将标记物表达值与其他特征关联至对应图像。 本数据集包含两份预处理数据文件: 1. `SingleCells.csv`:每行代表一个细胞,列则为与该细胞相关的各类数据。每条观测数据通过`ImageNumber`与`ObjectNumber`的组合进行唯一标识。该数据已完成溢出校正(spillover correction)。 2. `CellNeighbours.csv`:每行代表一组细胞-细胞相互作用。该数据以图格式存储,列名为`from`与`to`,分别表示从索引细胞指向相邻细胞(尽管采用此命名约定,但数据实际为无向图);两列中的整数值均与`SingleCells.csv`中的`ObjectNumber`相对应。 注意:预处理数据文件中以`is_`开头的变量遵循如下约定:0代表假(FALSE),1代表真(TRUE)。 本数据集同时提供两份列注释文件,用于说明预处理表格文件中各列的详细信息,分别为`SingleCellsAnnotation.xlsx`与`CellNeighboursAnnotation.xlsx`。 其余文件为`Code`目录中代码所需的注释与预处理数据文件,除非你计划重新运行分析,否则可忽略这些文件。 代码与可复现性说明 本目录中同时附带分析代码与对应的预处理数据。本代码基于Conda环境运行,环境详情已在`CondaEnv.yml`文件中给出。 本数据集附带了METABRIC研究的预处理元数据,但我们推荐:若后续分析依赖METABRIC元数据,请使用从原始出版物或公共仓库下载的最新数据,因为此类元数据可能会更新,且用户可能希望采用不同的方式进行处理。 代码按论文中的图表进行了分类,必须按照论文中图表出现的顺序运行代码,因为后续代码依赖于前期代码生成的衍生文件。同时,代码必须在本目录中运行,因为其相对路径依赖于当前目录的结构。



