遇见数据集

universalner/uner_llm_inst_chinese

收藏
Hugging Face2023-12-20 更新2024-03-04 收录
官方服务:

资源简介:

--- license: cc-by-sa-4.0 language: - zh task_categories: - token-classification dataset_info: - config_name: zh_pud splits: - name: test num_examples: 999 - config_name: zh_gsd splits: - name: test num_examples: 499 - name: dev num_examples: 499 - name: train num_examples: 3996 - config_name: zh_gsdsimp splits: - name: test num_examples: 499 - name: dev num_examples: 499 - name: train num_examples: 3996 --- # Dataset Card for Universal NER v1 in the Aya format - Chinese subset This dataset is a format conversion for the Chinese data in the original Universal NER v1 into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions. The dataset contains different subsets and their dev/test/train splits, depending on language. For more details, please refer to: ## Dataset Details For the original Universal NER dataset v1 and more details, please check https://huggingface.co/datasets/universalner/universal_ner. For details on the conversion to the Aya instructions format, please see the complete version: https://huggingface.co/datasets/universalner/uner_llm_instructions ## Citation If you utilize this dataset version, feel free to cite/footnote the complete version at https://huggingface.co/datasets/universalner/uner_llm_instructions, but please also cite the *original dataset publication*. **BibTeX:** ``` @preprint{mayhew2023universal, title={{Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark}}, author={Stephen Mayhew and Terra Blevins and Shuheng Liu and Marek Šuppa and Hila Gonen and Joseph Marvin Imperial and Börje F. Karlsson and Peiqin Lin and Nikola Ljubešić and LJ Miranda and Barbara Plank and Arij Riabi and Yuval Pinter}, year={2023}, eprint={2311.09122}, archivePrefix={arXiv}, primaryClass={cs.CL} } ```

--- 许可证:CC-BY-SA-4.0 语言: - 中文 任务类别: - Token分类(token-classification) 数据集信息: - 配置名称:zh_pud 划分集: - 名称:测试集 样本数量:999 - 配置名称:zh_gsd 划分集: - 名称:测试集 样本数量:499 - 名称:验证集 样本数量:499 - 名称:训练集 样本数量:3996 - 配置名称:zh_gsdsimp 划分集: - 名称:测试集 样本数量:499 - 名称:验证集 样本数量:499 - 名称:训练集 样本数量:3996 --- # 适配Aya格式的通用命名实体识别v1数据集卡片——中文子集 本数据集系将原始通用命名实体识别v1(Universal NER v1)中的中文语料转换为Aya指令格式的产物,沿用与原始数据集一致的CC-BY-SA-4.0协议及使用许可条款发布。 本数据集包含不同子集及其验证、测试、训练划分集,具体划分规则依语言而定。如需了解更多细节,请参阅: ## 数据集详情 如需获取原始通用命名实体识别v1(Universal NER v1)数据集的相关细节,请访问:https://huggingface.co/datasets/universalner/universal_ner。 如需了解转换为Aya指令格式的具体实现细节,请查看完整版本数据集:https://huggingface.co/datasets/universalner/uner_llm_instructions。 ## 引用说明 若您使用本数据集版本,可引用上述Aya格式转换后的完整数据集(链接:https://huggingface.co/datasets/universalner/uner_llm_instructions),同时请务必引用**原始数据集的发表文献**。 **BibTeX格式引用:** @preprint{mayhew2023universal, title={{Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark}}, author={Stephen Mayhew and Terra Blevins and Shuheng Liu and Marek Šuppa and Hila Gonen and Joseph Marvin Imperial and Börje F. Karlsson and Peiqin Lin and Nikola Ljubešić and LJ Miranda and Barbara Plank and Arij Riabi and Yuval Pinter}, year={2023}, eprint={2311.09122}, archivePrefix={arXiv}, primaryClass={cs.CL} }

提供机构:
universalner
原始信息汇总

数据集概述

数据集信息

配置名称:zh_pud

  • 拆分
    • 测试集:999个样本

配置名称:zh_gsd

  • 拆分
    • 测试集:499个样本
    • 开发集:499个样本
    • 训练集:3996个样本

配置名称:zh_gsdsimp

  • 拆分
    • 测试集:499个样本
    • 开发集:499个样本
    • 训练集:3996个样本

许可证

  • CC-BY-SA 4.0

语言

  • 中文

任务类别

  • 词性标注
搜集汇总
数据集介绍
universalner/uner_llm_inst_chinese 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务