遇见数据集

nassimjp/tsanga-sharegpt-pashto-english

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: other license_name: community-license license_link: https://creativecommons.org/licenses/by-nc-sa/4.0/ task_categories: - translation - text-generation language: - ps - en tags: - pashto - translation - sharegpt - conversation - chat - sft - instruction-tuning - community-license - open-source - low-resource-language - afghanistan - pashto-ai - bilingual - multi-turn - llm - generative-ai - chatbot - nlp pretty_name: Tsanga ShareGPT Pashto-English Dataset size_categories: - 1k<n<10k --- # 💬 Tsanga ShareGPT Pashto-English Dataset ## د تسنګا شیر-ګپټ پښتو-انګلیسي ډیټاسیټ [![HuggingFace](https://img.shields.io/badge/🤗-HuggingFace-yellow)](https://huggingface.co/datasets/nassimjp/tsanga-sharegpt-pashto-english) [![License](https://img.shields.io/badge/License-Community_License-blue)](LICENSE) [![Language](https://img.shields.io/badge/Language-Pashto-پښتو-green)](https://en.wikipedia.org/wiki/Pashto) [![Language](https://img.shields.io/badge/Language-English-red)](https://en.wikipedia.org/wiki/English_language) [![Size](https://img.shields.io/badge/Size-8.6k-orange)](https://huggingface.co/datasets/nassimjp/tsanga-sharegpt-pashto-english) > **🌟 د پښتو-انګلیسي ژباړې او خبرو اترو لپاره د ټولنې لخوا جوړ شوی ډیټاسیټ** > *A community-built dataset for Pashto-English translation and conversation* دا ډیټاسیټ د **۸،۶۸۲** لوړ کیفیت لرونکو پښتو-انګلیسي جملو جوړو څخه جوړ شوی دی، چې د **ShareGPT** فارمټ کې د خبرو اترو (Conversations) په بڼه چمتو شوی دی. This dataset contains **8,682** high-quality Pashto-English sentence pairs formatted in the **ShareGPT** conversation format. ## 📊 Dataset Statistics | Attribute | Value | |-----------|-------| | **Total Conversations** | 8,682 | | **Language Pair** | Pashto (ps) ↔ English (en) | | **Format** | ShareGPT (conversations) | | **License** | Community License (CC BY-NC-SA 4.0) | | **Average Conversation Length** | 2 messages | | **File Size** | ~2.5 MB | ## 🗂️ Data Structure Each record follows the ShareGPT conversational format: ```json { "conversations": [ { "from": "human", "value": "Translate this to Pashto: Love is the song of the soul" }, { "from": "gpt", "value": "مينه د روح سندره ده" } ] } ``` ### Fields Description | Field | Type | Description | |-------|------|-------------| | `conversations` | List | Array of conversation messages | | `from` | String | Role: "human" or "gpt" | | `value` | String | Message content (Pashto or English) | ## 🔥 Key Features ### 1. 🔄 **Bidirectional Translation** Each conversation is randomly oriented: - Either: Pashto → English - Or: English → Pashto This ensures the model learns translation in both directions. ### 2. 🎭 **Natural Conversation Flow** The ShareGPT format mimics real human-AI interactions, making it ideal for: - Chatbot training - Conversational AI - Virtual assistants ### 3. 🧹 **Clean & Deduplicated** - All duplicate entries removed using MD5 hashing - Empty and short lines filtered out - Whitespace normalized - UTF-8 encoding preserved ### 4. 🌍 **Culturally Rich Content** The dataset includes diverse topics: - History (Pashtun heroes, Ahmad Shah Baba) - Politics (Afghanistan, Pakistan, regional issues) - Literature (Poetry, Ghani Khan, Rahman Baba) - Philosophy (Rumi, life lessons) - Humor (Eid sheep jokes, cultural satire) - Technology (AI, LLMs, datasets) ## 📈 Dataset Breakdown by Category | Category | Estimated Count | Percentage | |----------|-----------------|------------| | 📜 History & Politics | 2,000+ | 23% | | 📖 Literature & Poetry | 1,800+ | 21% | | 💭 Philosophy & Wisdom | 1,500+ | 17% | | 😂 Humor & Satire | 1,200+ | 14% | | 🤖 Technology & AI | 1,000+ | 12% | | 🗣️ General Conversation | 1,182+ | 13% | | **Total** | **8,682** | **100%** | ## 💻 Usage Examples ### Loading with Hugging Face Datasets ```python from datasets import load_dataset # Load the dataset dataset = load_dataset("nassimjp/tsanga-sharegpt-pashto-english", split="train") # Access a conversation conversation = dataset[0] for message in conversation["conversations"]: print(f"{message['from']}: {message['value']}") # Dataset size print(f"Total conversations: {len(dataset):,}") ``` ### Fine-Tuning with LLaMA-Factory ```bash llamafactory-cli train \ --model_name_or_path Qwen/Qwen2-7B \ --dataset tsanga-sharegpt-pashto-english \ --template sharegpt \ --output_dir ./output \ --per_device_train_batch_size 4 \ --learning_rate 2e-5 \ --num_train_epochs 3 ``` ### Using for Pashto-English Translation ```python from transformers import pipeline # After fine-tuning a model on this dataset translator = pipeline("text-generation", model="your-finetuned-model") # Translate from Pashto to English result = translator("لاندې پښتو جمله انګلیسي ته وژباړه: کندهار د پښتنو هویت دی") print(result[0]['generated_text']) ``` ## 📝 Sample Conversations ### Translation Example (Pashto → English) ```json { "conversations": [ { "from": "human", "value": "How do you say this in English: کندهار د پښتنو هویت دی" }, { "from": "gpt", "value": "Kandahar is the identity of the Pashtuns" } ] } ``` ### Translation Example (English → Pashto) ```json { "conversations": [ { "from": "human", "value": "Translate this to Pashto: The rain of exile is bitter" }, { "from": "gpt", "value": "د جلاوطنۍ باران تریخ دی" } ] } ``` ### Philosophical Conversation ```json { "conversations": [ { "from": "human", "value": "Translate this to Pashto: Silence is not emptiness; silence is the place where the tired soul finds its voice again" }, { "from": "gpt", "value": "خاموشي تشه نه ده… خاموشي هغه ځای دی چې ستړی روح بېرته خپل غږ پیدا کوي" } ] } ``` ## 🧹 Curation Process ### Step 1: Collection - Gathered from 102 source text files - Total raw lines: 9,017 ### Step 2: Cleaning - Removed empty lines - Filtered out short lines (<10 characters) - Normalized whitespace ### Step 3: Deduplication - MD5 hash-based duplicate detection - 332 duplicates removed - Final unique lines: 8,685 ### Step 4: Super Shuffle - Lines randomly shuffled to break context dependency - Forces model to focus on content, not sequence bias ### Step 5: Format Conversion - Converted to ShareGPT conversation format - Random direction selection (Pashto↔English) - JSON formatting with UTF-8 encoding ### Step 6: Quality Validation - 3 lines rejected due to format issues - Final quality: 99.97% success rate ## 📊 Processing Statistics ``` ┌─────────────────────────────────────────────────────────────┐ │ Processing Summary │ ├─────────────────────────────────────────────────────────────┤ │ 📁 Raw files processed: 102 │ │ 📄 Total lines read: 9,017 │ │ 🗑️ Empty lines removed: 0 │ │ ✂️ Short lines removed: 0 │ │ 🔄 Duplicates removed: 332 │ │ ✨ Unique lines extracted: 8,685 │ │ ⚠️ Format rejects: 3 │ │ ✅ Final dataset size: 8,682 │ │ 💾 Final file size: ~2.5 MB │ └─────────────────────────────────────────────────────────────┘ ``` ## 🎯 Use Cases ### 1. **Pashto-English Machine Translation** Train models to translate between Pashto and English in both directions. ### 2. **Pashto Conversational AI** Build chatbots and virtual assistants that can understand and respond in Pashto. ### 3. **Low-Resource Language Research** Study techniques for improving NLP performance on low-resource languages like Pashto. ### 4. **Cross-Lingual Transfer Learning** Use as a bridge for transferring knowledge between English and Pashto. ## 🏷️ License This dataset is released under the **Community License** (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International). ### You are free to: - ✅ **Share** — copy and redistribute the material in any medium or format - ✅ **Adapt** — remix, transform, and build upon the material ### Under the following terms: - 🔄 **Attribution** — You must give appropriate credit to the dataset creator (nassimjp) - 📚 **NonCommercial** — You may not use the material for commercial purposes - 🤝 **ShareAlike** — If you remix, transform, or build upon the material, you must distribute your contributions under the same license ## 📚 Citation ```bibtex @misc{tsanga-sharegpt-pashto-english, author = {Nasibullah Nassim (nassimjp)}, title = {Tsanga ShareGPT Pashto-English Dataset}, year = {2025}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/datasets/nassimjp/tsanga-sharegpt-pashto-english}} } ``` ## 🤝 Contributing If you find issues or have suggestions for improving this dataset: 1. Open an issue on the [Hugging Face repository](https://huggingface.co/datasets/nassimjp/tsanga-sharegpt-pashto-english/discussions) 2. Submit a pull request with improvements 3. Contact the dataset creator directly ## 🙏 Acknowledgments - **DeepSeek** for processing assistance - **Palwasha CLI** for data management - **Pashto AI Community** for support and inspiration - All contributors who helped make this dataset possible --- ## 🔗 Quick Links [![HuggingFace](https://img.shields.io/badge/🤗-View_on_HuggingFace-ffd21e)](https://huggingface.co/datasets/nassimjp/tsanga-sharegpt-pashto-english) [![GitHub](https://img.shields.io/badge/GitHub-Repository-black)](https://github.com/spinzar/tsanga-sharegpt-pashto-english) --- **🌟 که دا ډیټاسیټ ستاسو لپاره ګټور و، نو ستوری ورکول مه هېروئ!** *If you find this dataset useful, don't forget to give it a star!* **🇦🇫 د پښتو-انګلیسي AI راتلونکی یوځای جوړوو!** *Building the future of Pashto-English AI together!*

This dataset contains **8,682** high-quality Pashto-English sentence pairs formatted in the **ShareGPT** conversation format.

提供机构:
nassimjp
搜集汇总
数据集介绍
nassimjp/tsanga-sharegpt-pashto-english 数据集图片
构建方式
该数据集立足于低资源语言普什图语与英语之间机器翻译资源匮乏的背景,通过汇集102个源文本文件共9017行原始语句,经历系统化清理与精炼流程构建而成。处理过程涵盖剔除空白行与字符数不足十条的短句、归一化空白字符,继而运用MD5哈希算法识别并移除332条重复条目,最终获得8685条独特语句。在此基础上进行全局随机混洗以削弱序列上下文依赖,随后将平行语料转换为ShareGPT对话格式,并随机指定翻译方向以实现双向对译。经质量校验后三条格式异常样本被拒,最终形成8682条高质量对话记录。
特点
数据集以ShareGPT对话结构封装8682组普什图语与英语句对,每段对话平均包含两轮消息,随机取向的翻译方向使模型得以同时习得双向转换能力。内容涵盖历史政治、文学诗歌、哲学智慧、幽默讽刺、技术人工智能及日常会话等多元主题,占比分别约为23%、21%、17%、14%、12%与13%,兼具文化丰富性与话题均衡性。所有条目均经过MD5去重、空白归一化与UTF-8编码保持,格式校验成功率高达99.97%,整体数据纯净且结构统一,专为低资源语言场景下的对话式翻译与生成任务而设计。
使用方法
该数据集可便捷加载于Hugging Face Datasets库,研究者通过load_dataset函数指定split为train即可获取完整对话集合,并逐条访问conversations字段中的角色与内容。在微调阶段,可借助LLaMA-Factory等框架以sharegpt模板对Qwen2等基座模型进行指令微调,设定批量大小、学习率与训练轮数等超参数。训练完成后,模型可用于普什图语与英语之间的双向机器翻译、对话式AI系统构建、低资源语言迁移学习研究等任务,使用时需遵循CC BY-NC-SA 4.0许可协议,注明出处且不得用于商业目的。
背景与挑战
背景概述
普什图语作为阿富汗及巴基斯坦地区的重要语言,其自然语言处理研究长期受限于双语平行语料的稀缺与质量参差。既有资源多聚焦于正式文本,缺乏贴近日常对话的语用样本,难以支撑对话式人工智能的研发需求。2025年,研究者Nasibullah Nassim(nassimjp)依托普什图人工智能社区,发布了Tsanga ShareGPT普什图语—英语数据集,以ShareGPT多轮对话结构收录8,682组高质量句对,覆盖双向翻译任务。该数据集填补了低资源语言在指令微调与对话生成领域的数据空白,为普什图语机器翻译、跨语言迁移学习及社区化语言技术发展提供了关键基础设施。
当前挑战
该数据集所应对的领域问题在于普什图语机器翻译与对话系统长期受制于平行语料稀缺、方言多样性及英语—普什图语之间显著的形态句法差异,导致模型在低资源场景下泛化能力不足。构建过程中,研究者面临原始文本质量参差、重复条目干扰及短句噪声等数据治理难题,须经由MD5哈希去重、长度过滤与空白归一化等多阶段清洗流程,方能确保语料的纯净度与一致性。此外,如何平衡双向翻译的方向分布、保持文化主题的多样性,并在社区许可框架下实现开源共享,均构成构建与推广过程中的实质性挑战。
常用场景
经典使用场景
在低资源语言机器翻译与跨语言对话系统研究中,该数据集凭借其双向平行语料与ShareGPT对话格式,成为训练普什图语-英语翻译模型及对话代理的经典资源。研究者常将其用于指令微调,使大语言模型习得双语转换能力,并通过多轮交互模板评估模型在真实对话场景下的翻译一致性与流畅度。
解决学术问题
该数据集有效缓解了普什图语作为低资源语言在自然语言处理研究中长期面临的数据稀缺与标注质量参差问题。其经过去重、清洗与格式规范化的高质量平行句对,为跨语言迁移学习、低资源翻译模型鲁棒性评估以及双语语义对齐等学术议题提供了可靠基准,推动了普什图语计算语言学研究的可复现性与公平比较。
衍生相关工作
围绕该数据集,研究者已衍生出多项经典工作,包括基于LLaMA-Factory的指令微调实验、普什图语对话生成模型的基准测试,以及融合文化主题的跨语言情感分析任务。这些工作进一步拓展了低资源语言在生成式人工智能中的边界,并催生了针对普什图语方言变体的数据增强与模型适配研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务