HalOmi
收藏资源简介:
HalOmi是一个手动标注的多语言机器翻译幻觉和遗漏检测基准数据集,由FAIR, Meta创建。该数据集涵盖18个翻译方向,包括不同资源水平和脚本,旨在解决机器翻译中的幻觉和遗漏问题。数据集通过严格的标注指南进行手动标注,包含细粒度的句子和词级标注。HalOmi的发布为可靠和可访问的研究提供了基础,以检测和分析翻译病理,并理解其原因。
HalOmi is a manually annotated multilingual machine translation hallucination and omission detection benchmark dataset created by FAIR, Meta. It covers 18 translation directions spanning various resource levels and writing scripts, aiming to address the issues of hallucinations and omissions in machine translation. The dataset is manually annotated in accordance with strict annotation guidelines and includes fine-grained sentence-level and word-level annotations. The release of HalOmi provides a reliable and accessible research foundation for detecting and analyzing translation pathologies, as well as understanding their underlying causes.




