KOM-Euph
收藏资源简介:
KOM-Euph是一个 keyword-oriented的多模态数据集,由华中农业大学信息学院创建。该数据集基于文本数据集Euph构建,增加了图像和语音模态,包含文本-图像-语音三元组,共86K条数据,覆盖Drug、Weapon和Sexuality三个领域。数据集通过自监督学习框架构建,使用掩码的目标关键词句子及其相关的图像和语音进行训练和验证。该数据集旨在解决多模态数据在隐语识别任务中的应用问题,为研究提供了丰富的多模态资源。
KOM-Euph is a keyword-oriented multimodal dataset developed by the School of Information, Huazhong Agricultural University. It is built upon the original text-only dataset Euph by supplementing with image and speech modalities. The dataset consists of 86K text-image-speech triplets, covering three domains: Drug, Weapon, and Sexuality. Constructed via a self-supervised learning framework, it uses masked target keyword sentences along with their associated images and speech for training and validation. This dataset aims to address the application challenges of multimodal data in implicit language recognition tasks, providing abundant multimodal resources for related research.

- 1Keyword-Oriented Multimodal Modeling for Euphemism Identification华中农业大学信息学院 · 2025年



