MahaSum
收藏资源简介:
MahaSum数据集是由L3Cube实验室和浦那计算机技术学院共同创建的一个大规模马拉地语新闻摘要数据集。该数据集包含25,374条新闻文章,每条文章都附有高质量的人工摘要。数据集通过从多个在线新闻源抓取文章并手动验证摘要创建而成。MahaSum数据集旨在支持马拉地语的自然语言处理研究,特别是抽象文本摘要任务。该数据集的应用领域包括马拉地语的新闻摘要生成和语言模型训练,旨在解决马拉地语等低资源语言在自然语言处理中的资源匮乏问题。
The MahaSum Dataset is a large-scale Marathi news summarization dataset jointly created by L3Cube Lab and the Pune Institute of Computer Technology. It contains 25,374 news articles, each paired with high-quality human-written summaries. The dataset was constructed by scraping articles from multiple online news sources and manually validating the summaries. The MahaSum Dataset aims to support natural language processing (NLP) research focused on Marathi, particularly the abstractive text summarization task. Its application fields include Marathi news summarization generation and language model training, and it is designed to address the resource scarcity issue of low-resource languages such as Marathi in natural language processing.

- 1L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in MarathiL3Cube实验室,浦那,印度 · 2024年



