snap-stanford/rl_response_only_gpt5_judge_train_disable_thinking_reddit_global_step_225_post_dist
收藏资源简介:
该数据集包含对话生成相关的结构化数据,主要特征包括:数据集索引(dataset_index)、响应键(response_key)、提示信息(prompt,包含内容、名称和角色)、生成文本(generation)以及额外信息(extra_info,包含评论指标、层次名称、索引、是否顶层、媒体来源、名称、角色、帖子ID、帖子指标、原始提示和分割等)。此外,数据集还包含多种指标(metrics),如内容类型(content_type)、情感(emotion)、敌意(hostility)、幽默(humor)、情感(sentiment)和立场(stance)等,每个指标都有对应的标签和原始文本信息。数据集分为训练集(train),包含3366个样本,总大小为86684154字节。
This dataset contains structured data related to dialogue generation, with main features including: dataset index (dataset_index), response key (response_key), prompt information (prompt, including content, name, and role), generated text (generation), and additional information (extra_info, including comment metrics, hierarchy name, index, is_top_level, media source, name, persona, post ID, post metrics, raw prompt, and split). Additionally, the dataset includes various metrics such as content type (content_type), emotion, hostility, humor, sentiment, and stance, each with corresponding labels and raw text information. The dataset is divided into a training set (train) containing 3366 samples, with a total size of 86684154 bytes.



