Defects4J
收藏资源简介:
该数据集是一个文本检索数据集,包含从Defects4J来源的单语言文档和查询信息。它由三种配置组成:default配置包含查询和文档的ID信息,正反例文档ID列表,以及类型和元数据;corpus配置包含文档的ID、来源、语言、标题、文本和元数据;query配置包含查询的ID、来源、语言、标题、文本和元数据。数据集分为测试集、语料库和查询三部分,分别包含467个、934个示例。
This dataset is a text retrieval dataset containing monolingual documents and query information sourced from Defects4J. It consists of three configurations: the default configuration includes ID information for queries and documents, lists of positive and negative document IDs, as well as type and metadata; the corpus configuration includes document ID, source, language, title, text and metadata; the query configuration includes query ID, source, language, title, text and metadata. The dataset is divided into three parts: test set, corpus, and queries, which contain 467 and 934 examples respectively.




