Sentinel-2云掩模数据集(Sentinel-2 Cloud Mask Catalogue)
收藏资源简介:
该数据集包括513个1022-x1022像素子镜头的云掩模,分辨率为20m,从2018年Level-1C Sentinel-2档案中随机采样。该数据集的设计基于对云掩模的一些观察:(i)整个产品的性能高度相关,因此子尺度比全场景提供更多的每像素值,(ii)当前的云掩模数据集通常关注特定区域,或手动选择使用的产品,这给数据集中引入了一种不代表真实世界数据的偏差,(iii)云掩模性能似乎与表面类型和云结构高度相关,所以测试应包括分析与这些变量相关的故障模式。 使用IRIS工具包对数据进行半自动注释,该工具包允许用户动态训练随机森林(使用LightGBM实现),通过迭代改进其预测来加速注释,但保留了注释器在需要时进行最终手动更改的能力。这种混合方法使我们能够处理比手动更多的掩码,我们认为这对于创建足够大的数据集以近似整个Sentinel-2档案的统计数据至关重要。
This dataset comprises 513 cloud masks of 1022 × 1022 pixels with a spatial resolution of 20 m, randomly sampled from the 2018 Level-1C Sentinel-2 archive. The design of this dataset is based on three key observations regarding cloud masks: (i) The performance of full scenes is highly correlated, so sub-scenes provide more per-pixel samples than full scenes; (ii) Existing cloud mask datasets often focus on specific regions or use manually selected products, which introduces biases that do not represent real-world data into the dataset; (iii) Cloud mask performance appears to be highly correlated with surface type and cloud structure, so testing should include analysis of failure modes associated with these variables. Data was semi-automatically annotated using the IRIS toolkit, which enables users to dynamically train a random forest (implemented via LightGBM) to accelerate annotation by iteratively refining its predictions, while retaining the ability for annotators to make final manual adjustments when needed. This hybrid approach allowed us to process far more masks than manual annotation alone, which we consider critical for creating a sufficiently large dataset to approximate the statistical characteristics of the entire Sentinel-2 archive.




