BORDIRLINES
收藏资源简介:
BORDIRLINES数据集由宾夕法尼亚大学的研究团队创建,专注于评估跨语言检索增强生成(RAG)系统的鲁棒性。该数据集包含251个涉及地理政治争议的查询,涵盖49种语言,共计720个查询。数据来源于维基百科文章,通过多种信息检索系统进行查询-文章相关性评分,选取相关段落。数据集的创建旨在研究多语言环境下RAG系统的性能,特别是在提供不同语言和来源的上下文时,模型的响应变化。该数据集的应用领域主要是解决跨语言信息检索和生成中的偏见和不一致性问题。
The BORDIRLINES dataset was developed by a research team at the University of Pennsylvania, and it is dedicated to evaluating the robustness of cross-lingual retrieval-augmented generation (RAG) systems. This dataset contains 251 queries involving geopolitical controversies, spanning 49 languages, with a total of 720 queries overall. The data is sourced from Wikipedia articles, where relevant paragraphs are selected after performing query-article relevance scoring via multiple information retrieval systems. The dataset was created to study the performance of RAG systems in multilingual environments, particularly the variations in model responses when provided with contexts in different languages and from various sources. The primary application of this dataset is to address biases and inconsistencies in cross-lingual information retrieval and generation tasks.
BordIRlines 数据集
概述
BordIRlines 是一个用于评估跨语言检索增强生成(Cross-lingual Retrieval-Augmented Generation)的数据集。
下载
数据集可以从 Hugging Face Hub 下载,链接为:https://huggingface.co/datasets/borderlines/bordirlines。
更多信息
有关数据集的更多详细信息和使用说明,请参阅 Hugging Face Hub 上的 README 文件。

- 1BordIRlines: A Dataset for Evaluating Cross-lingual Retrieval-Augmented Generation宾夕法尼亚大学 · 2024年



