AnswerCarefully
收藏资源简介:
AnswerCarefully数据集由日本国立情报学研究所开发,旨在提升日语LLM输出的安全性。数据集包含1800对问题及其参考答案,涉及广泛的风险类别,旨在通过人工收集和参考答案的提供,促进LLM的安全微调。数据集不仅有助于提高LLM回答问题的安全性,也方便了自动评估LLM输出安全性的准确性。该数据集为日语LLM的安全评估提供了基准,有助于解决日语LLM在文化背景下的安全性问题。
The AnswerCarefully dataset was developed by the National Institute of Informatics (NII), Japan, to enhance the safety of Japanese Large Language Models (LLMs). It contains 1,800 pairs of questions and their reference answers, covering a wide range of risk categories. The dataset aims to facilitate the safe fine-tuning of LLMs through manual data collection and the provision of reference answers. It not only helps improve the safety of LLMs' responses to questions but also enables accurate automatic evaluation of the safety of LLM outputs. Additionally, this dataset provides a benchmark for safety assessment of Japanese LLMs, assisting in addressing the safety issues faced by Japanese LLMs within their specific cultural context.
AnswerCarefully Dataset 概述
数据集基本信息
- 名称: AnswerCarefully Dataset (AC)
- 最新版本: Version 2.2 (ACv2.2) (发布于2025/5/29)
- 托管平台: Hugging Face
- 语言: 日语(含英语元标签)
- 数据规模:
- ACv1: 946对问答
- ACv2: 1,800对问答
- 用途: 提升日语及其他语言LLM输出的安全性与适当性
数据集特点
- 数据内容:
- 手动创建的日语问答对,涵盖日本社会/文化敏感话题
- 参考Do-Not-Answer数据集的安全分类体系,但样本为原创
- 包含安全参考回答(既安全又尽可能有帮助)
- 分类体系:
- 5个风险领域
- 12种危害类型
- 56个子类别(ACv2调整后)
- 版本更新:
- ACv2.2新增多语言文化适应元数据:
- 问题英文翻译
- 文化特异性标签(0-2级)
- 翻译注释
- 分类标签英文翻译
- ACv2.2新增多语言文化适应元数据:
数据分布
- ACv2数据划分:
- 测试集: 336样本(每个子类别6样本)
- 开发集: 1,464样本
使用条款
- 使用限制: 禁止重新分发
- 注意事项: 包含冒犯性/不安全内容,仅限LLM安全改进用途
- 引用格式:
Hisami Suzuki et al. "AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output." 2025. https://arxiv.org/abs/2506.02372
开发信息
- 开发机构:
- ACv1: Riken AIP(Citadel AI协助)
- ACv2: 国立情报学研究所(LLMC)
- 联系方式: ac_dataset@nii.ac.jp




