遇见数据集

allenai/cochrane_dense_oracle

收藏
Hugging Face2022-11-18 更新2024-03-04 收录
官方服务:

资源简介:

该数据集是Cochrane数据集的副本,但训练、验证和测试集的输入源文档被替换为密集检索器。查询是每个示例的`target`字段,语料库是训练、验证和测试集中所有文档的联合,文档是`title`和`abstract`的拼接。检索器使用`facebook/contriever-msmarco`,并通过PyTerrier进行检索,采用默认设置。检索策略为`oracle`,即检索的文档数量`k`设置为每个示例的原始输入文档数量。在训练集和验证集上的检索结果显示,Recall@100和Rprec等指标表现良好。

This dataset is a replica of the Cochrane dataset, where the source input documents for the training, validation, and test splits have been replaced with dense retriever outputs. The query for each example is its `target` field, the corpus is the union of all documents across the training, validation, and test splits, and each document is the concatenation of its `title` and `abstract`. The retriever utilized is `facebook/contriever-msmarco`, and retrieval was performed using PyTerrier with default settings. The retrieval strategy is `oracle`, meaning the number of retrieved documents `k` is set to match the count of original input documents per example. Retrieval results on the training and validation splits demonstrate strong performance on metrics including Recall@100 and Rprec.

提供机构:
allenai
原始信息汇总

数据集概述

基本信息

  • 标注创建者: 专家生成
  • 语言创建者: 专家生成
  • 语言: 英语
  • 许可证: Apache 2.0
  • 多语言性: 单语种
  • 大小类别: 10K<n<100K
  • 源数据集:
    • 扩展自其他-MS^2
    • 扩展自其他-Cochrane
  • 任务类别:
    • 摘要生成
    • 文本到文本生成
  • Papers with Code ID: multi-document-summarization
  • 美观名称: MSLR Shared Task

数据集描述

  • 查询: 每个示例的target字段
  • 语料库: train, validation和test分割中所有文档的联合。一个文档是title和abstract的串联。
  • 检索器: 使用facebook/contriever-msmarco通过PyTerrier,默认设置
  • top-k策略: "oracle",即检索的文档数量k设置为每个示例的原始输入文档数量

检索结果

  • 训练集:
    • Recall@100: 0.7790
    • Rprec: 0.4487
    • Precision@k: 0.4487
    • Recall@k: 0.4487
  • 验证集:
    • Recall@100: 0.7856
    • Rprec: 0.4424
    • Precision@k: 0.4424
    • Recall@k: 0.4424
  • 测试集: 无检索结果,测试集是盲测,没有查询。
二维码
社区交流群
二维码
科研交流群
商业服务