UNS-Exterior Spatial Sound Events 2023 (UNS-ESSE2023)
收藏资源简介:
This dataset was generated within the UNS2: Localising Audio Events in Crowds use case of the H2020 MARVEL project. The purpose of the dataset is to offer audio samples collected outdoors in an urban area that could be used for development of sound event localisation and detection (SELD) models for acoustic monitoring in urban environments. The dataset was generated within a staged recording process. Audio samples were collected by utilising Infineon Audiohub Nano 8-channel microphone array board with a sampling rate of 48 kHz. The audio was synthetised by mixing target sound events from the FSD50K dataset “gunshot” and “gunfire”, “boom”, and “shatter” with samples of the class “chatter”, that was used as a background noise, extracted also from the FSD50k dataset. The scenario included different SNR values for the events in the mixtures and the measurement of the Sound Pressure Level (SPL) before and after the recording. Sound events were reproduced using eight JBL VP7212MDP10 speakers, which were positioned circularly around the microphone array board in equidistant positions. Data collection was performed for two different distances: 5 m and 10 m. One of the speakers was used to reproduce the target events (overlaid on the background noise), where the selected speaker varied during the data collection, while the others were used to reproduce background noise only. The recording setup was placed outdoors, involving additional ambient noise of the urban city area, which is a case closer to the real-world scenario comparing to the datasets recorded in the laboratory conditions. Together with the audio files, spatiotemporal annotation of the sound events is provided. These annotations include temporal onset and offset of the target events, azimuth, source distance, SNR level and SPL values for background ambience noise. This work was funded by the European Union’s Horizon 2020 research and innovation program MARVEL under grant agreement No 957337. This publication reflects the authors’ views only. The European Commission is not responsible for any use that may be made of the information it contains.
本数据集生成于H2020 MARVEL项目的UNS2:人群音频事件定位用例。本数据集旨在提供城市户外采集的音频样本,用于开发面向城市环境声学监测的声音事件定位与检测(Sound Event Localisation and Detection,SELD)模型。本数据集采用分阶段录制流程生成。音频样本采集依托英飞凌(Infineon)Audiohub Nano 8通道麦克风阵列板完成,采样率为48 kHz。音频通过混合来自FSD50K数据集的目标声音事件(包括"gunshot"(枪声)、"gunfire"(连续枪声/炮声)、"boom"(爆炸声)与"shatter"(碎裂声)),以及来自同一FSD50K数据集的"chatter"(嘈杂交谈声)类别样本生成,后者被用作背景噪声。该录制场景包含混合信号中不同的信噪比(Signal-to-Noise Ratio,SNR)参数,并在录制前后开展声压级(Sound Pressure Level,SPL)的测量工作。声音事件通过8台JBL VP7212MDP10扬声器播放,这些扬声器以等距方式环形布置在麦克风阵列板周围。数据采集针对两个不同的播放距离开展:5米与10米。其中一台扬声器用于播放叠加于背景噪声之上的目标声音事件,数据采集过程中切换使用不同的扬声器,其余扬声器仅用于播放背景噪声。本次录制部署于户外场景,采集了城市区域的额外环境噪声,相较于实验室环境下录制的数据集,更贴近真实应用场景。本数据集随音频文件一同提供声音事件的时空标注信息。此类标注包含目标事件的时间起始与结束时刻、方位角、声源距离、信噪比水平以及背景环境噪声的声压级数值。本工作由欧盟地平线2020研究与创新计划MARVEL项目(资助协议编号957337)资助。本出版物仅代表作者观点,欧盟委员会不对基于其内容的任何使用行为承担责任。



