neuralspace/autotrain-data-citizen_nlu_bn
收藏官方服务:
资源简介:
该数据集是为项目citizen_nlu_bn自动处理的AutoTrain数据集,语言为孟加拉语(bn),主要用于文本分类任务。数据集包含文本和目标标签,目标标签共有55个类别,涵盖了从联系真实人物到报告车辆事故等多种意图。数据集分为训练集和验证集,分别包含27146和6800个样本。
This is an AutoTrain dataset automatically processed for the project citizen_nlu_bn, in Bengali (bn). It is primarily used for text classification tasks. The dataset consists of text and target labels, with a total of 55 distinct categories covering various intents including contacting real people and reporting vehicle accidents. The dataset is split into training and validation sets, which contain 27146 and 6800 samples respectively.
提供机构:
neuralspace原始信息汇总
数据集概述
数据集名称
AutoTrain Dataset for project: citizen_nlu_bn
语言
数据集的语言为bn。
数据集结构
数据实例
数据实例包含以下字段:
- text: 文本内容,字符串类型。
- target: 分类标签,共有55个类别。
数据集字段
- text: 文本字段,字符串类型。
- target: 分类标签字段,包含55个类别,如ContactRealPerson, Eligibility For BloodDonationWithComorbidities等。
数据集分割
数据集分为训练集和验证集:
- train: 包含27146个样本。
- valid: 包含6800个样本。



