遇见数据集

tasksource/multilingual-zero-shot-label-nli

收藏
Hugging Face2023-06-23 更新2024-03-04 收录
官方服务:

资源简介:

该数据集主要用于改进零样本分类HF管道中的标签理解。数据集包含三个特征:labels、premise和hypothesis,分别表示标签、前提和假设。数据集分为训练集、测试集和验证集,分别包含878967、9400和9400个样本。数据集的总大小为188946124字节,下载大小为104413879字节。数据集的任务类别包括零样本分类和文本分类,任务ID为自然语言推理。数据集是多语言的,并且可以用于改进零样本分类HF管道中的标签理解。

This dataset is primarily used to improve label understanding in zero-shot classification Hugging Face (HF) pipelines. It comprises three features: labels, premise, and hypothesis, which respectively represent label, premise, and hypothesis. The dataset is split into training, test, and validation sets, with 878,967, 9,400, and 9,400 samples respectively. The total size of the dataset is 188,946,124 bytes, and the download size is 104,413,879 bytes. Its task categories include zero-shot classification and text classification, with the task ID being natural language inference. This is a multilingual dataset that can be applied to improve label understanding in zero-shot classification HF pipelines.

提供机构:
tasksource
原始信息汇总

数据集概述

数据集信息

  • 特征:

    • labels: 类别标签,包含以下类别:
      • 0: entailment
      • 1: neutral
      • 2: contradiction
    • premise: 字符串类型
    • hypothesis: 字符串类型
    • task: 字符串类型
  • 数据分割:

    • train: 包含878967个样本,大小为185352754.0字节
    • test: 包含9400个样本,大小为1775890.0字节
    • validation: 包含9400个样本,大小为1817480.0字节
  • 数据集大小:

    • 下载大小: 104413879字节
    • 数据集大小: 188946124.0字节
  • 许可: other

  • 任务类别:

    • zero-shot-classification
    • text-classification
  • 任务ID:

    • natural-language-inference
  • 多语言性:

    • 多语言
二维码
社区交流群
二维码
科研交流群
商业服务