LMSYS-Chat-1M
收藏资源简介:
LMSYS-Chat-1M是由加州大学伯克利分校的研究团队创建的大型语言模型对话数据集,包含一百万条真实世界的对话记录。该数据集通过LMSYS项目收集,涵盖了25个最先进的语言模型,并从210,000个独特的IP地址中收集。数据集内容丰富,包括对话的收集过程、基本统计数据和主题分布,强调了其多样性、原创性和规模。该数据集的应用领域广泛,包括开发内容审核模型、构建安全基准、训练指令遵循模型以及创建挑战性基准问题,旨在理解和推进大型语言模型的能力。
LMSYS-Chat-1M is a large language model dialogue dataset created by a research team at the University of California, Berkeley, containing one million real-world conversational records. This dataset is collected via the LMSYS project, encompassing 25 state-of-the-art language models and sourced from 210,000 unique IP addresses. The dataset features comprehensive content, including the dialogue collection process, basic statistical data and topic distribution, highlighting its diversity, originality and scale. It has a wide range of application scenarios, including developing content moderation models, building safety benchmarks, training instruction-following models and creating challenging benchmark questions, aiming to understand and advance the capabilities of large language models.

- 1LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset加州大学伯克利分校 · 2024年



