SV2V-RSim
收藏资源简介:
SV2V-RSim是一个全面的车对车(V2V)协同感知基准数据集,旨在推动V2V协同感知研究的发展。该数据集包含四个地图场景,涵盖四种不同的天气条件和从日出到夜晚的六个时间段,总计提供了203K LiDAR帧、402K RGB帧和788K标注的3D边界框,覆盖17个物体类别。数据集以其复杂的交通环境和高度动态的交通模式为特点,提供了更接近真实的渲染质量和精确的几何资产。SV2V-RSim支持多种下游任务,包括协同物体检测、深度估计和语义分割等。数据格式包括LiDAR点云、RGB图像、深度和语义标注信息,以及详细的JSON文件记录传感器位姿和实例信息。数据集还提供了数据处理脚本,用于深度和语义标注的解码,并支持OpenCOOD框架的数据加载。
SV2V-RSim is a comprehensive Vehicle-to-Vehicle (V2V) collaborative perception benchmark dataset designed to advance research in V2V collaborative perception. The dataset includes four map scenarios, covering four different weather conditions and six time periods from sunrise to night, providing a total of 203K LiDAR frames, 402K RGB frames, and 788K annotated 3D bounding boxes covering 17 object categories. The dataset is characterized by complex traffic environments and highly dynamic traffic patterns, offering near-realistic rendering quality and precise geometric assets. SV2V-RSim supports multiple downstream tasks, including collaborative object detection, depth estimation, and semantic segmentation. Data formats include LiDAR point clouds, RGB images, depth and semantic annotations, as well as detailed JSON files recording sensor poses and instance information. The dataset also provides data processing scripts for decoding depth and semantic annotations and supports data loading with the OpenCOOD framework.
SV2V-RSim 数据集概述
基本信息
- 数据集名称: SV2V-RSim
- 用途: 面向自选择车车协同感知的综合基准数据集,支持协同目标检测、深度估计、语义分割等多类下游任务
- 发布状态: 已发布(2026年5月)
数据规模与特性
| 特性 | 数值 |
|---|---|
| 地图数量 | 4 张 |
| 天气条件 | 4 种不同天气 |
| 时间段 | 6 个(从日出到夜晚) |
| LiDAR 帧数 | 203K |
| RGB 帧数 | 402K |
| 标注 3D 边界框数 | 788K |
| 目标类别数 | 17 类 |
数据格式
目录结构
train/{scene_name}/{agent_id}/ ├── {timestamp}_{camera_name}DepthStencil.png # 深度模板渲染图像 ├── {timestamp}{camera_name}ObjectIdentifier.png # 目标标识渲染图像 ├── {timestamp}{camera_name}_RGB.jpeg # RGB图像 ├── {timestamp}.json # 传感器位姿及实例标注 └── {timestamp}.pcd # 激光雷达点云
JSON 标注文件内容
- 传感器位姿(世界坐标系下的朝向和位置)
- 当前智能体类型
- 标注信息:每个标注目标包含:
- 3D包围盒中心、尺寸、朝向
- 目标类别和资产名称
- 模板ID
- LiDAR可见性(基于3D框内点数判断)
- 每个相机的可见性信息(可见性标志、可见像素数、总像素数、可见2D框IoU)
数据处理工具
深度解码
- 脚本:
scripts/depth_extract.py - 输入:
DepthStencil.png图像 - 输出:
*_depth_m.npy(浮点型深度图,单位米)*_depth_cm.png(16位PNG深度图,单位厘米)*_valid_mask.png(有效深度掩码)depth_manifest.csv(统计信息)
语义解码
- 脚本:
semantic_extract.py - 输入:帧前缀、相机名称
- 输出:
*_semantic_id.png(语义ID图,像素值0-17)*_semantic_color.png(彩色语义分割图)*_overlay.png(RGB与语义分割叠加图)*_RGB.jpeg(复制的RGB图像)*_annotation_info.json(类别映射、模板/类型/类别ID、像素统计)
OpenCOOD 数据加载器
- 提供文件:
intermediate_fusion_dataset_lv2v.py,支持在OpenCOOD框架下使用SV2V-RSim数据集
基准测试结果
检测基准基于 OpenCOOD 进行评估,不同模型的性能如下:
| 模型 | AP_M@IoU 0.3 | AP_M@IoU 0.5 | AP_N@IoU 0.3 | AP_N@IoU 0.5 | AP_P@IoU 0.3 | AP_P@IoU 0.5 | AP_S@IoU 0.3 | AP_S@IoU 0.5 | 带宽(MB) |
|---|---|---|---|---|---|---|---|---|---|
| No Fusion | 32.6 | 28.7 | 19.2 | 12.7 | 3.3 | 0.6 | 5.7 | 3.2 | 0 |
| Late Fusion | 48.1 | 42.6 | 25.2 | 13.5 | 4.2 | 1.1 | 10.4 | 7.0 | 12.91 |
| Early Fusion | 57.9 | 56.0 | 30.9 | 23.7 | 4.9 | 1.5 | 13.9 | 8.1 | 25.54 |
| F-Cooper | 56.7 | 53.8 | 45.9 | 38.7 | 15.4 | 4.7 | 32.1 | 21.2 | 27.85 |
| CoBEVT | 57.9 | 55.6 | 52.4 | 44.3 | 23.3 | 12.7 | 43.2 | 30.2 | 28.85 |
| Where2comm | 62.9 | 60.7 | 37.3 | 25.6 | 13.7 | 3.7 | 25.3 | 12.0 | 26.72 |
| SVA (ours) | 65.2 | 62.6 | 49.5 | 38.2 | 21.8 | 11.3 | 41.3 | 27.6 | 25.53 |
待完成事项
- [x] 数据集发布
- [x] 数据处理脚本
- [ ] 评估代码
- [ ] 所有基于SV2V-RSim数据集的基准测试任务




