SIDOD
收藏资源简介:
我们介绍了由NVIDIA深度学习数据合成器生成的新图像数据集,该数据集旨在用于对象检测,姿势估计和跟踪应用程序。该数据集包含从三个逼真的虚拟环境的18个摄像机视点生成的144k个立体图像对,这些虚拟环境具有多达10个对象 (从YCB数据集的21个对象模型中随机选择) 和飞行干扰物。对象和相机姿势,场景照明以及对象和干扰物的数量被随机分配。每个提供的视图都包括RGB、深度、分割和表面法线图像。我们描述了我们的领域随机化方法,并提供了对产生数据集的决策的洞察力。
We introduce a novel image dataset generated by the NVIDIA Deep Learning Data Synthesizer, which is tailored for object detection, pose estimation and tracking applications. This dataset comprises 144k stereo image pairs generated from 18 camera viewpoints across three photorealistic virtual environments. Each of these environments contains up to 10 objects randomly selected from the 21 object models in the YCB dataset, alongside flying distractors. Object and camera poses, scene lighting, as well as the counts of objects and distractors, are all randomly assigned. Each provided viewpoint includes RGB, depth, segmentation, and surface normal images. We describe our domain randomization methodology and provide insights into the decision-making behind the dataset's generation.




