ai4privacy/pii-masking-micro-100k
收藏资源简介:
PII掩码微样本:多语言样本是PII-Masking-3M系列的一个微型分层样本,源自更大的数据集pii-masking-openpii-1.5m。该样本按(source_dataset, language)比例抽样,确保每种语言区域和标签都有代表性,且亚太地区的行优先出现。数据集包含99,990个示例,其中训练集90,000个,验证集9,990个,涵盖30种语言、37个区域和19种个人可识别信息(PII)标签,总标注数为725,798个。数据格式为JSONL,许可证为CC-BY-4.0。该数据集专门用于隐私保护任务,如令牌分类和文本生成,包含合成PII数据,无真实个人数据,适用于多语言NLP应用。
PII Masking Micro: Multilingual Sample is a micro-sized stratified sample of the PII-Masking-3M family, derived from the larger dataset pii-masking-openpii-1.5m. It is sampled proportionally by (source_dataset, language) to ensure representation for every locale and label, with Asia Pacific rows appearing first. The dataset contains 99,990 examples, with 90,000 for training and 9,990 for validation, covering 30 languages, 37 regions, and 19 personally identifiable information (PII) labels, totaling 725,798 annotations. The data is in JSONL format and licensed under CC-BY-4.0. It is designed for privacy-preserving tasks such as token classification and text generation, containing synthetic PII data only, with no real personal data, and is suitable for multilingual NLP applications.




