MMCR
收藏资源简介:
MMCR数据集是由西北工业大学和阿里巴巴集团联合创建的多模态多轮对话数据集。该数据集包含MMCR-310k和MMCR-Bench两部分,其中MMCR-310k是一个包含310000个对话的数据集,对话覆盖1-4张图片,分为4轮或8轮;MMCR-Bench则是一个诊断性基准,包含8个领域的对话和40个子主题。该数据集通过模拟真实世界的用户聊天机器人交互,强调每轮对话的上下文关联和逻辑推进,旨在提升视觉语言模型的多轮对话上下文推理能力。
The MMCR dataset is a multimodal multi-turn dialogue dataset jointly created by Northwestern Polytechnical University and Alibaba Group. It consists of two parts: MMCR-310k and MMCR-Bench. MMCR-310k is a dataset containing 310,000 dialogues, where each dialogue covers 1 to 4 images and is structured as either 4-turn or 8-turn conversations. MMCR-Bench, on the other hand, is a diagnostic benchmark that includes dialogues across 8 domains and 40 sub-themes. This dataset simulates real-world human-chatbot interactions, emphasizes the contextual relevance and logical progression of each dialogue turn, and aims to enhance the contextual reasoning ability of visual-language models in multi-turn dialogue scenarios.




