ptb-text-only/ptb_text_only
收藏资源简介:
这是Penn Treebank项目的第二个版本,包含1989年《华尔街日报》的百万字材料。在这个版本中,稀有词已被替换为<unk>标记,数字被替换为<N>标记。数据集中的文本为美式英语。
This is the second version of the Penn Treebank project, which contains one million words of material from the 1989 issues of The Wall Street Journal. In this version, rare words have been replaced with the <unk> tag, and numbers have been replaced with the <N> tag. The text in this dataset is in American English.
数据集卡片:Penn Treebank
数据集描述
数据集摘要
Penn Treebank项目:包含1989年《华尔街日报》材料的百万字语料库。数据集中的罕见词已被替换为<unk>标记,数字已被替换为<N>标记。
支持的任务和排行榜
- 语言建模
- 掩码语言建模
语言
数据集中的文本为美式英语。
数据集结构
数据实例
[需要更多信息]
数据字段
- sentence: 字符串类型
数据分割
- train: 包含42068个实例,5143706字节
- test: 包含3761个实例,453710字节
- validation: 包含3370个实例,403156字节
数据集创建
策划理由
[需要更多信息]
源数据
初始数据收集和规范化
[需要更多信息]
源语言生产者
[需要更多信息]
注释
注释过程
[需要更多信息]
注释者
[需要更多信息]
个人和敏感信息
[需要更多信息]
使用数据的注意事项
数据集的社会影响
[需要更多信息]
偏见的讨论
[需要更多信息]
其他已知限制
[需要更多信息]
附加信息
数据集策展人
[需要更多信息]
许可信息
数据集仅供研究使用。请检查数据集许可以获取更多信息。
引用信息
@article{marcus-etal-1993-building, title = "Building a Large Annotated Corpus of {E}nglish: The {P}enn {T}reebank", author = "Marcus, Mitchell P. and Santorini, Beatrice and Marcinkiewicz, Mary Ann", journal = "Computational Linguistics", volume = "19", number = "2", year = "1993", url = "https://www.aclweb.org/anthology/J93-2004", pages = "313--330", }
贡献
感谢@harshalmittal4添加此数据集。




