相关数据集
ahmadSiddiqi/fr-retrieval-syntec-s2p
该数据集包含两个主要部分:corpus和queries。corpus部分包含90个例子,总字节数为114096;queries部分包含100个例子,总字节数为8475。数据集的总下载大小为61749字节,总大小为122571字节。每个数据项包含id和text两个字段,分别表示数据的唯一标识和文本内容。
Hugging Face2024-01-26 更新80
FAIRsharing record for: FAST (Faceted Application of Subject Terminology) Topic Facet
This FAIRsharing record describes: FAST (Faceted Application of Subject Terminology) is an enumerative, faceted subject heading schema derived from the Library of Congress Subject Headings (LCSH). The
DataCite Commons2024-01-29 更新60
ssktora/trec_ct_2021-train-bm25-pyserini-20-dev
该数据集包含查询信息及其对应的正例和负例文本段落。每个文本段落都有文档ID、标题和文本内容。数据集被分割为训练集,并提供了相应的配置信息。
Hugging Face2025-04-08 更新60
FrancophonIA/LongEval-Retrieval
LongEval 2024信息检索数据集,包含基于Qwant搜索引擎的用户查询和对应的相关网页文档,用于2024年LongEval信息检索实验室的训练和测试。数据集包含法语文档和英文翻译。
Hugging Face2025-03-30 更新70
phamtungthuy/vanbanlienquan
--- dataset_info: features: - name: question dtype: string - name: field dtype: string - name: relevant dtype: string splits: - name: train num_bytes: 28941457 num_exam
Hugging Face2024-01-05 更新60



