ComplexDataLab/chai-veracity-double-summarized
收藏资源简介:
该数据集是一个包含45个示例的训练数据集,主要用于处理声明或主张相关的文本数据。每个示例包含多个字段:id(唯一标识符)、_batch(批处理信息)、claim(声明内容)、cluster_id(聚类标识符,用于分组相关声明)、post_count(帖子数量,可能表示与声明相关的讨论次数)、original_ids(原始ID列表,关联到原始数据源)、original_texts(原始文本列表,提供声明的来源或上下文)、date(日期,记录数据生成或收集时间)。数据集大小为115165字节,下载大小为58902字节,适用于自然语言处理任务,如文本分类、聚类分析或信息检索,但具体应用场景未在README中明确说明。
This dataset is a training set containing 45 examples, primarily designed for handling text data related to claims or assertions. Each example includes multiple fields: id (unique identifier), _batch (batch processing information), claim (the content of the claim), cluster_id (cluster identifier for grouping related claims), post_count (number of posts, possibly indicating discussion frequency related to the claim), original_ids (list of original IDs linked to the data source), original_texts (list of original texts providing source or context for the claim), date (date of data generation or collection). The dataset size is 115165 bytes with a download size of 58902 bytes, suitable for natural language processing tasks such as text classification, clustering analysis, or information retrieval, though specific application scenarios are not explicitly mentioned in the README.



