NNE
收藏资源简介:
NNE是一个大规模的嵌套命名实体数据集,由悉尼大学和CSIRO Data61创建,专注于英语新闻稿件中的细粒度命名实体识别。该数据集包含279,795个命名实体提及,涵盖114种实体类型,支持多达6层的嵌套结构。数据集的创建过程涉及定制的注释工具和详细的注释指南,确保了注释的一致性和准确性。NNE数据集的应用领域包括自然语言处理中的下游任务,如共指消解、问答、摘要等,旨在解决现有NER工具在处理嵌套实体结构时的局限性。
NNE is a large-scale nested named entity dataset developed by the University of Sydney and CSIRO Data61, focusing on fine-grained named entity recognition in English news articles. This dataset contains 279,795 named entity mentions, covers 114 entity types, and supports up to 6 layers of nested structures. The dataset construction process involved customized annotation tools and detailed annotation guidelines, ensuring the consistency and accuracy of annotations. Application scenarios of the NNE dataset include downstream natural language processing tasks such as coreference resolution, question answering, text summarization and others, aiming to address the limitations of existing NER tools when handling nested entity structures.




