saillab/alpaca_albanian_taco
收藏资源简介:
--- language: - sq pretty_name: Albanian alpaca-52k size_categories: - 100K<n<1M --- This repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: ``` { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } ``` Please refer to the paper for more details: [OpenReview](https://openreview.net/forum?id=02MLWBj8HP) If you have used our dataset, please cite it as follows: **Citation** ``` @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}, url={https://openreview.net/forum?id=02MLWBj8HP} } ``` The original dataset [(Alpaca-52K)](https://github.com/tatsu-lab/stanford_alpaca?tab=readme-ov-file#data-release) was translated using Google Translate. **Copyright and Intended Use** This dataset has been released under CC BY-NC, intended for academic and research purposes only. Please review the licenses and terms and conditions of Alpaca-52K, Dolly-15K, and Google Cloud Translation before using this dataset for any purpose other than research.
--- language: - sq pretty_name: 阿尔巴尼亚语Alpaca-52K size_categories: - 100K<n<1M --- 本仓库包含TaCo论文所使用的数据集。 该数据集遵循TaCo论文中规定的格式,具体如下: { "instruction": "阿尔巴尼亚语指令", "input": "阿尔巴尼亚语输入", "output": "英文指令:英文原指令, 英文回复:英文原回复, 阿尔巴尼亚语回复:阿尔巴尼亚语原回复" } 如需了解更多细节,请参阅该论文:[OpenReview](https://openreview.net/forum?id=02MLWBj8HP) 若您使用了本数据集,请按以下方式引用: **引用** @inproceedings{upadhayay2024taco, title={TaCo: 提升大语言模型(Large Language Model)面向低资源语言的跨语言迁移能力:基于翻译辅助思维链流程}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={第5届面向有限/低资源场景的实用机器学习研讨会,ICLR}, year={2024}, url={https://openreview.net/forum?id=02MLWBj8HP} } 本数据集的原始版本[(Alpaca-52K)](https://github.com/tatsu-lab/stanford_alpaca?tab=readme-ov-file#data-release)通过谷歌翻译完成译制。 **版权与使用意图** 本数据集采用CC BY-NC协议发布,仅用于学术与研究用途。在将本数据集用于研究以外的任何用途前,请务必查阅Alpaca-52K、Dolly-15K以及谷歌云翻译的许可协议与条款细则。
数据集概述
数据集特征
- instruction:数据类型为字符串。
- input:数据类型为字符串。
- output:数据类型为字符串。
- id:数据类型为字符串。
- text:数据类型为字符串。
数据集划分
- 训练集:包含49601个样本,总大小为187725809.16905907字节。
- 测试集:包含12401个样本,总大小为46934290.83094094字节。
数据集大小
- 下载大小:119171373字节。
- 数据集总大小:234660100.0字节。
数据文件配置
- 默认配置:
- 训练集路径:
data/train-* - 测试集路径:
data/test-*
- 训练集路径:



