SuryaKrishna02/aya-telugu-news-articles
收藏资源简介:
`aya-telugu-news-articles`是一个开源数据集,包含通过爬取泰卢固语新闻网站生成的指令风格记录。该数据集由Cohere For AI的Aya Open Science Initiative创建,包含超过467k条记录,主要用于训练大型语言模型、生成合成数据和数据增强。数据集支持两种任务:根据文章标题生成文章内容和根据文章内容生成标题。数据集为泰卢固语,且为单语种,但可能包含少量英语内容。数据集遵循Apache 2.0许可证,可用于学术或商业用途。
`aya-telugu-news-articles` is an open-source dataset consisting of instruction-style records generated by scraping Telugu news websites. It was created by the Aya Open Science Initiative under Cohere For AI, and contains over 467,000 records. The dataset is primarily used for training large language models (LLMs), generating synthetic data and data augmentation. It supports two tasks: generating full article content from a given article title, and generating an article title from the provided article content. The dataset is primarily in Telugu, though it may contain a small amount of English content. It is licensed under the Apache 2.0 license, and can be used for both academic and commercial purposes.
数据集概述
数据集名称
aya-telugu-news-articles
数据集描述
该数据集是通过网络爬虫从泰卢固语新闻文章网站生成的开放源代码指令样式记录集合。由Cohere For AI的Aya Open Science Initiative创建。
数据集用途
该数据集可用于以下任务:
- 训练大型语言模型(LLMs)
- 合成数据生成
- 数据增强
数据集语言
泰卢固语
数据集版本
1.0
数据集大小
超过467,000条记录
数据集任务
- 给定文章的标题/头条,生成带有该标题/头条的文章。
- 给定文章,生成文章的标题/头条。
数据集字段
inputs:语言模型的提示或输入。targets:语言模型的完成或输出。template_id:在inputs和targets中使用的模板ID。template_lang:在inputs和targets中使用的语言的ISO代码,其中tel指泰卢固语。
数据集模板
用于从爬取的数据创建指令样式提示和完成的模板类别如下:
- 给定文章的标题/头条,生成带有该标题/头条的文章。
- 给定文章,生成文章的标题/头条。
数据集来源
通过从泰卢固语地区的著名新闻文章网站Suryaa Website进行网络爬虫,并进行预处理,如去除不需要的字符,从爬取的数据中去除过长或过短的文章,最后将爬取的数据转换为指令样式提示和完成。
数据集限制
- 数据集内容可能反映网站的偏见、事实错误、政治倾向和敏感问题。
- 尽管尽力保持数据集为单语种,但可能存在一些记录包含泰卢固语和英语混合的情况。
数据集许可证
Apache 2.0
数据集贡献者




