遇见数据集

saumyamalik/filtered_wc_sample_500k_decontam_v2

收藏
Hugging Face2025-09-16 更新2025-10-25 收录
官方服务:

资源简介:

这是一个包含会话文本和相关信息的训练数据集,用于模型训练。数据集中的每个会话都包含了文本内容、创建时间、国家信息、IP地址哈希值、请求头部信息等。同时,数据集为每个会话提供了毒性标签,可用于毒性检测和过滤。数据集分为训练集,可供模型训练使用。

This is a training dataset containing conversation texts and related information for model training. Each conversation in the dataset includes text content, creation time, country information, IP address hash, request header information, etc. Additionally, the dataset provides toxicity labels for each conversation, which can be used for toxicity detection and filtering. The dataset is split into a training set for model training purposes.

提供机构:
saumyamalik
二维码
社区交流群
二维码
科研交流群
商业服务