遇见数据集

oumayma03/darija_bible

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

--- license: other license_name: all-rights-reserved-morocco-bible-society-2012 license_link: https://www.bible.com/fr/versions/558 language: - ary - ar pretty_name: Darija Bible (MSTD) size_categories: - 1K<n<10K task_categories: - text-generation - translation - feature-extraction tags: - bible - darija - moroccan-arabic - religion - low-resource - research-only configs: - config_name: default data_files: - split: train path: data/bible_MSTD_ALL.csv - config_name: by_book data_files: - split: MAT path: books/MAT.csv - split: MRK path: books/MRK.csv - split: LUK path: books/LUK.csv - split: JHN path: books/JHN.csv - split: ACT path: books/ACT.csv - split: ROM path: books/ROM.csv - split: b1CO path: books/1CO.csv - split: b2CO path: books/2CO.csv - split: GAL path: books/GAL.csv - split: EPH path: books/EPH.csv - split: PHP path: books/PHP.csv - split: COL path: books/COL.csv - split: b1TH path: books/1TH.csv - split: b2TH path: books/2TH.csv - split: b1TI path: books/1TI.csv - split: b2TI path: books/2TI.csv - split: TIT path: books/TIT.csv - split: PHM path: books/PHM.csv - split: HEB path: books/HEB.csv - split: JAS path: books/JAS.csv - split: b1PE path: books/1PE.csv - split: b2PE path: books/2PE.csv - split: b1JN path: books/1JN.csv - split: b2JN path: books/2JN.csv - split: b3JN path: books/3JN.csv - split: JUD path: books/JUD.csv - split: REV path: books/REV.csv --- # Darija Bible — Moroccan Standard Translation (MSTD) The full **New Testament** translated into **Moroccan Darija** (الترجمة المغربية القياسية, *MSTD*). Moroccan Darija is a **low-resource** spoken Arabic variety. This dataset is published here as a research artifact for language modeling, fine-tuning, machine translation between Darija and other languages, and evaluation of multilingual / Arabic-dialect models. --- ## ⚠️ Copyright notice — please read before using **This dataset reproduces a copyrighted translation.** It is published here in good faith as a research/personal-use artifact, but the underlying text is **not** under any open license. - **Original work:** *الترجمة المغربية القياسية* (Moroccan Standard Translation, *MSTD*) — New Testament. - **Copyright holder:** © 2012 **Morocco Bible Society** / دار الكتاب المقدس بالمغرب. **All rights reserved.** - **Source page:** [YouVersion / bible.com, version 558](https://www.bible.com/fr/versions/558). - **How the data was obtained:** scraped from the public HTML pages of bible.com. - **License of this redistribution:** **none granted.** Posting here does not transfer or imply any rights. You should treat the text as fully copyrighted and consult the original publisher before any non-research use. ### Acceptable use (in the maintainer's understanding, *not* legal advice) - Personal study, academic research, and quantitative/NLP experiments on the text. - Reporting aggregate statistics, training tokenizers, evaluating language models, etc. ### Probably **not** acceptable without explicit permission from the Morocco Bible Society - Re-publishing the text in another product, app, book, website, or downstream dataset. - Commercial use of any kind. - Distributing model checkpoints whose primary purpose is to reproduce this translation verbatim. ### Takedown / rights-holder request If you are the Morocco Bible Society, YouVersion, or any other party with rights to this translation and you would like this dataset **made private or removed**, please open a **Discussion** on this dataset page or contact the dataset owner directly. The dataset will be taken down promptly upon request. --- ## Contents - **27 books** (the entire New Testament — the MSTD translation does not cover the Old Testament). - **7,958 verses**. - One row per verse. ### Files | Path | Description | |------|-------------| | `data/bible_MSTD_ALL.csv` | Combined corpus, all 27 books (default config) | | `books/<USFM>.csv` | One file per book (used by the `by_book` config) | ### Schema | Column | Type | Example | |--------|------|---------| | `book` | string (USFM code) | `MAT`, `JHN`, `1CO` | | `chapter` | string | `"1"`, `"28"` | | `verse` | string | `"1"`, `"23"` | | `text` | string | verse text in Moroccan Darija (Arabic script) | ## Quick start ```python from datasets import load_dataset # Whole NT as a single train split ds = load_dataset("oumayma03/darija_bible") print(ds["train"][0]) # {'book': 'MAT', 'chapter': '1', 'verse': '1', 'text': 'هَادُو هُمَ ...'} # One split per book (book codes starting with a digit are prefixed 'b', e.g. b1CO) ds_by_book = load_dataset("oumayma03/darija_bible", "by_book") print(list(ds_by_book.keys())) # ['MAT', 'MRK', ..., 'b1CO', ...] print(ds_by_book["MAT"][0]) ``` ## Books included (USFM codes) `MAT`, `MRK`, `LUK`, `JHN`, `ACT`, `ROM`, `1CO`, `2CO`, `GAL`, `EPH`, `PHP`, `COL`, `1TH`, `2TH`, `1TI`, `2TI`, `TIT`, `PHM`, `HEB`, `JAS`, `1PE`, `2PE`, `1JN`, `2JN`, `3JN`, `JUD`, `REV`. ## Preprocessing - Cross-reference markers (e.g. `#لوقا 1:27`) embedded in the source HTML have been **stripped** so each row contains only the verse text. - Verses split across paragraph blocks (poetry, etc.) are reassembled into a single row. - Whitespace is normalized to single spaces. ## How to credit The translation itself should always be credited to the **Morocco Bible Society (2012)**: ```bibtex @misc{mstd_darija_nt_2012, title = {الترجمة المغربية القياسية — العهد الجديد}, shorttitle= {MSTD — Moroccan Standard Translation, New Testament}, author = {Morocco Bible Society}, year = {2012}, publisher = {دار الكتاب المقدس بالمغرب}, url = {https://www.bible.com/fr/versions/558}, note = {Copyright © 2012 Morocco Bible Society. All rights reserved.} } ```

The full New Testament translated into Moroccan Darija (Moroccan Standard Translation, MSTD). Moroccan Darija is a low-resource spoken Arabic variety. This dataset is published here as a research artifact for language modeling, fine-tuning, machine translation between Darija and other languages, and evaluation of multilingual / Arabic-dialect models. The dataset includes 27 books (the entire New Testament) with 7,958 verses, one row per verse. Files are provided in CSV format, including a combined corpus and separate files per book. Preprocessing involves stripping cross-reference markers, reassembling verses split across paragraph blocks, and normalizing whitespace.

提供机构:
oumayma03
搜集汇总
数据集介绍
oumayma03/darija_bible 数据集图片
构建方式
该数据集以摩洛哥阿拉伯语(Darija)低资源方言为对象,其构建依托于《摩洛哥标准译本》(MSTD)新约全书。数据源自bible.com公开HTML页面,经网络爬取获得原始文本。预处理环节移除嵌入式交叉引用标记(如#لوقا 1:27),将因诗歌等段落划分而断裂的经文重新拼接为单行,并归一化空白字符至单一空格。最终形成每行对应一节经文的CSV文件,涵盖27卷书、7958节经文,同时提供合并语料与按书卷拆分的两种文件结构。
特点
此数据集聚焦于摩洛哥达里贾语这一资源稀缺的口语阿拉伯语变体,收录完整新约二十七卷,共七千九百五十八节经文。数据以节为单位组织,字段包括书卷USFM代码、章、节及阿拉伯字母书写的达里贾语文本。数据集分为默认合并配置与按书卷拆分配置,后者以书卷代码作为划分标签,对以数字开头的代码前缀‘b’予以区分。文本已经过清洗,去除了源HTML中的交叉引用标记与多余空白,并重组了跨段落经文。版权归属摩洛哥圣经公会,仅限研究用途。
使用方法
研究者和开发者可通过Hugging Face的datasets库加载该数据集。调用load_dataset("oumayma03/darija_bible")可获得包含全部新约的单一训练分割;指定配置名"by_book"则返回以每卷书为独立分割的字典,访问相应书卷代码即可获取对应数据。典型应用包括达里贾语语言建模、微调、与其他语言间的机器翻译,以及多语言或阿拉伯方言模型的评估。使用时应始终将译文归功于摩洛哥圣经公会(2012),并遵守其版权限制,仅用于个人研习、学术研究及文本定量与自然语言处理实验。
背景与挑战
背景概述
在阿拉伯语方言自然语言处理领域,摩洛哥达里贾语(Darija)作为一支资源匮乏的口语变体,长期缺乏公开可用的高质量文本语料,制约了语言建模与机器翻译研究的发展。2012年,摩洛哥圣经公会(Morocco Bible Society)出版了《摩洛哥标准译本》(MSTD)新约全书,为达里贾语提供了罕见的长篇连贯文本。此后,研究者将这一版权文本从bible.com公开页面采集整理,构建了darija_bible数据集,涵盖27卷书、7,958节经文。该数据集已成为达里贾语低资源语言研究的重要基准,推动了多语言模型对阿拉伯语方言的覆盖与评估。
当前挑战
达里贾语本身缺乏标准化正字法与大规模标注资源,口语变体在书写传统上的薄弱使得模型难以捕捉稳定的语言规律,而该数据集仅覆盖新约单一文体,领域偏向宗教文本,限制了模型在通用场景中的泛化能力。构建过程中,源文本受版权严格保护,再分发权利未获授权,研究者须在合法使用与学术共享之间谨慎权衡;原始HTML中嵌入的交叉引用标记、跨段落分割的诗歌体经文以及不规范空白字符,均需经过去噪、重组与归一化处理,以确保每行经文语义完整且格式统一。
常用场景
经典使用场景
在阿拉伯语方言资源稀缺的背景下,darija_bible数据集凭借其完整的《新约》摩洛哥达里贾语译文,成为低资源语言建模与机器翻译研究的经典语料。研究者常将其用于达里贾语与其他语言间的神经机器翻译系统训练,亦可用于微调多语言预训练模型,以评估模型对阿拉伯方言的理解与生成能力。该数据集还频繁现身于语言模型持续预训练任务中,帮助模型捕获达里贾语特有的词汇与句法模式。
实际应用
在实际应用层面,darija_bible数据集可服务于多语种智能助手与聊天机器人的开发,助力系统理解并生成摩洛哥达里贾语,进而提升北非地区用户的交互体验。该数据集亦可用于构建面向达里贾语的拼写检查、文本补全及情感分析工具,并为语言教育软件提供真实语料,支持达里贾语学习者的阅读与翻译练习。此外,在跨语言信息检索与内容推荐系统中,该数据集有助于优化面向达里贾语用户的检索质量。
衍生相关工作
基于darija_bible数据集,研究者已衍生出一系列经典工作,包括摩洛哥达里贾语与标准阿拉伯语、法语之间的神经机器翻译模型,以及针对阿拉伯方言的预训练语言模型微调研究。该数据集还激发了低资源语言分词器的构建与评估,并促进了达里贾语词性标注、命名实体识别等下游任务的探索。部分工作进一步利用该数据集开展多方言语言模型的对比分析,为阿拉伯语方言自然语言处理奠定了重要基础。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务